Project 04 · US Wind Turbine Database and EIA-923 generation
Wind Repowering Uplift
From the first full year after repowering, a repowered plant generates 48.4% more than it did the year before the work, relative to 498 never-repowered plants of the same vintage, with a 95% interval of 36.3% to 61.4%. The design could have detected 14.2%. The obvious regression says 30.4%, and this page shows why that is not the answer.
The question
Repowering is sold as a generation uplift. Measured against never-repowered plants of the same vintage, what is it, and does it differ between projects that grew the rotor, a blade purchase, and projects that only replaced the drivetrain?
The data and its scale
75,727 turbines from the USGS API with a SHA-256 manifest, 8,480 of them flagged as retrofitted; the April 2018 archived release of 57,636 turbines for pre-repowering rotor diameters and capacity; and thirteen EIA-923 archives, about 20 MB each, giving 14,203 wind rows rolled up to 14,077 plant-years across 1,486 plants.
What the source did
- B5On 1,529 repowered turbines across 17 plants the build year has been silently rewritten to the repowering year. Thirteen of the 17 plants are in the treatment group; had they set the age window it would have stretched to 2024 and admitted brand-new plants as controls for twenty-year-old ones. The window is set from the treated plants whose build year was not rewritten: 2001 to 2012.
- ·The turbine id does not survive a repowering. Only 10.5% of retrofitted turbines match the 2018 release by id, and 0% for four cohort years. The project id bridges 92.2% and is the key.
- ·The generation publisher answers a request for a missing file with its landing page at HTTP 200, and answers HEAD with 503 where GET returns 200. A download loop that trusts the status code fills a folder with HTML.
- B16A dot as the missing-value marker on 19 generation rows. The archive also carries -9999 sentinels on 5,137 rotor and 3,041 capacity rows, and the 2025 generation release is provisional, at 624 filers against 1,348.
Seventeen assertion rules with expected-against-found counts, all matching.
Method
An experiment nobody ran, set up in experimentation terms. 90 plants had every turbine retrofitted, 1,208 had none, and 29 had both, which are set aside. Controls are never-repowered plants built 2001 to 2012, 567 of them eligible, each reporting positive generation in every year 2013 to 2024. After dropping the 2015 cohort, one plant, and the 2024 cohort, which has no post-treatment year while 2025 is provisional, 80 treated plants in seven cohorts face 498 controls. The outcome is annual net generation in logs.
Pre-period balance was measured before any outcome: indexed generation for the two groups sits within about two points in every year 2013 to 2019. Treated plants are bigger, 135 MW against 81 MW, and ran a higher capacity factor beforehand, 0.347 against 0.326 capacity-weighted, so the estimate is stated as an upper bound. Treated capacity rose 10.1% between 2018 and today against 1.4% for controls, so about nine points of any uplift is added nameplate rather than performance.
The estimate is a cohort-by-year comparison against never-repowered controls only, the Callaway and Sant’Anna estimator written out by hand so every number can be pointed at, with standard errors from 999 redraws of plants with replacement. Two library implementations reproduce the event-time series and the cohort effects with a largest gap of 0.000000.
Finding
Repowered plants against never-repowered plants of the same vintage
Data behind this chart
| Year | Estimate | 95% interval | Cohorts | Plants |
|---|---|---|---|---|
| −4 | +3.9 | not printed | 7 | 80 |
| −3 | +2.3 | not printed | 7 | 80 |
| −2 | −4.5 | not printed | 7 | 80 |
| −1 | 0 (base) | not printed | 7 | 80 |
| 0 | −6.7 | −14.2 to +1.5 | 7 | 80 |
| +1 | +33.9 | +20.9 to +48.1 | 7 | 80 |
| +2 | +46.6 | +34.4 to +59.8 | 6 | 73 |
| +3 | +54.7 | +40.2 to +70.6 | 5 | 67 |
| +4 | +59.4 | +43.5 to +77.0 | 4 | 61 |
| +5 | +61.6 | +41.6 to +84.4 | 3 | 38 |
| Headline, +1 onward | +48.4% | +36.3% to +61.4% | 7 | 80 |
| Naive regression | +30.4% | +23.7% to +37.5% |
- The groups moved together before the work.Three placebo years read +3.9%, +2.3% and minus 4.5% against the base year. A joint Wald test gives 3.94 on 3 degrees of freedom, p = 0.268.
- The repowering year dips.Minus 6.7%, because the old turbines come down before the new ones go up. Counting that year in, the uplift averages +33.5%.
- The uplift builds over three years.+33.9% at year 1, +46.6% at year 2, +54.7% at year 3, +59.4% at year 4 and +61.6% at year 5, the last on three cohorts and 38 plants. From the first full year on: +48.4%, 0.395 log points, standard error 0.043.
- The naive regression says 30.4%.A two-way fixed-effects regression averages 49 two-by-two comparisons. 96.3% of its weight sits on clean cohort-versus-never-treated ones, 2.3% on earlier-versus-later cohorts, and 1.5% on comparisons that use already-repowered plants as controls, which average minus 3.3%.
The naive two-way fixed-effects regression returns +30.4%, 0.266 log points with a standard error of 0.027, and nothing in its output warns you. Decomposed by hand, the 49 weights sum to 1.000000 and reproduce the coefficient to six decimals. It is nearly right on this panel by luck of the sample, and the decomposition is how you find that out.
The minimum detectable effect at 80% power and a two-sided 5% test is 14.2%, computed from control-plant variance before looking at any treated outcome, and 12.9% from the bootstrap after. Had the true effect been 10%, this design would have missed it more often than not.
The rotor: plants whose rotor grew, 66 of them with a median growth of 18%, gained +44.3%, interval +34.0% to +55.6%; plants whose rotor did not grow, 10 of them, gained +34.5%, interval +27.1% to +42.2%. The gap is inside the uncertainty, so the uplift is not mainly a blade story, or not one this data can see. Four treated plants with no 2018 record showed +198.6% and look like rebuilds; they meet the definition set before any outcome was seen, so they stay in the headline, and without them the figure is +43.0%.
Six sensitivities all land inside the headline interval except by choice of scale: mixed plants counted as treated, +52.8%; levels instead of logs, +108 GWh per plant-year, 31.1% of the pre-repowering mean; controls limited to the 15 states holding a treated plant, +56.8%; an unbalanced panel, +49.9%; a three-year base, +49.6%; and dropping the four rebuilds, +43.0%.
What this cannot tell you
- Anything about blade material. No public source records it.
- A causal effect free of selection. Operators repower the sites they expect to gain most from, the groups were balanced on trend and not on level, and the estimate is an upper bound.
- Efficiency as opposed to output. A capacity-factor outcome needs per-plant-year capacity that neither source provides.
- Local wind. Year effects absorb national weather only.
- Revenue, and the 2024 cohort, which waits on the final 2025 release.
Artifacts
- An 87-cell notebook, run top to bottom on 2026-09-09
- A findings memo written in experimentation vocabulary
- A decision log, D-01 to D-08, each decision with what would reverse it
- A data contract and a source-to-target map
- A four-table star schema exported with a DAX measure layer and a state row-level-security role specified; the Power BI dashboard itself is not yet built
Reproducibility
The first CI run, on 2026-09-07, found that this notebook could not run top to bottom on a clean checkout: an exploratory cell opened two downloads the publisher now answers with a landing page. A one-cell guard fixed it and no number changed. The same run found that the memo’s 42 uncertainty figures, every bootstrap interval, the standard error, the pre-trend test and the after-the-fact minimum detectable effect, were not what the committed notebook printed, while every point estimate held. On 2026-09-09 I ran the notebook by hand on a byte-identical re-pull and re-measured all 42. The figures on this page are from that run.
| Figures in the manifest | 220 |
|---|---|
| Checked on every run | 215 |
| Declared design constants | 5 |
| Notebook run time on the committed baseline | 424 s |
Sources
- US Wind Turbine Database, USGS, through its API, with a SHA-256 manifest.
- The April 2018 archived release of the same database, for pre-repowering rotor diameters and capacity.
- EIA-923 annual generation files, 2013 to 2025, thirteen archives at their committed hashes; the 2025 release is provisional.
- The findings memo, decision log, data contract and source-to-target map in 04-wind-repowering/memo.