Project 06 · IESO hourly Ontario demand, 2002 to 2026
Ontario Demand Forecast with an Honest Baseline
A day ahead, a transparent linear model with no weather input misses hourly demand by 523 MW, 3.0% of demand, against 1,221 MW for the same-hour-last-week forecast, removing 57.2% of its error. A week ahead it removes 15.2%, and at the last hour of the week the two cannot be told apart. Its 80% interval covers 77.5% of hours, both it and the baseline fail in summer, and it cannot pick a top-five peak day.
The question
Issued at midnight for the next 168 hours, how far can a transparent model with no weather in it beat the same-hour-last-week forecast, a day and a week ahead? When it says 80% sure, is it right 80% of the time? And for a large consumer deciding whether to curtail tomorrow, is the day-ahead error small enough to tell a top-five peak day from the day that just misses?
The data and its scale
25 yearly files, 213,456 hours from 2002 to 2026, and 24 zonal files, 204,695 hours, with a per-file SHA-256 manifest carrying the publisher’s own timestamps. Each year exists on the server as a plain and a versioned copy, and all 25 pairs hash identical. Every date carries exactly 24 rows, so the series runs on a fixed standard-time clock and a 168-hour lag always lands on the same wall-clock hour.
What the source did
- B5One hour is missing from a closed year with no marker: 2025-05-01 hour 1, found by a generated spine. Hour 24 of April 30 reads 13,795 MW and hour 2 of May 1 reads 12,352; the gap is filled with the mean of its neighbours, 13,074 MW, and flagged. It is the only gap in 213,456 hours.
- B17The zonal report writes one column with a quoted thousands separator once it reaches four digits, on 129 of 204,695 rows, which a sniffed CSV dialect splits into two fields. The quote character is set explicitly and the separator stripped before the cast.
- B1172,152 zonal rows where the ten zones do not sum to the published total, all but four within 5 MW. Separately, 35 hours across two days in 2016 have every zone at zero, and 43 hours have market demand below Ontario demand.
- B8An outlier check written for data errors finds the twelve hours of the 14 August 2003 blackout, 2,270 MW at hour 17 against 22,380 a week before. It is an event, not an error; it stays, and it sits well before any training window.
Eighteen assertion rules, eight of them blocking, all at their expected values.
Method
Every day from 2024-12-31 to 2026-08-24 is an origin, 602 of them, chosen so that every lead from 1 to 168 is scored the same number of times inside a window ending on a complete month. Three forecasts per origin: the seasonal naive, the same hour and weekday one week earlier; a last-day naive, the origin day repeated; and one ordinary least squares regression per lead on sixteen inputs known at the origin, namely weekly lags, the origin day’s level, the target day’s type and three annual Fourier pairs.
The model refits every 28 origins on the trailing 1,825 origins whose targets were all observed, 5,880 fits, and no fit ever sees an outcome from its own forecast window. The 372 origins before the test window are scored and never reported; their errors build the prediction intervals. There is no weather in it, stated as the floor a weather-fed forecast must beat.
Finding
Mean absolute error by lead, 602 daily origins
Data behind this chart
| Lead | Same hour last week | Last day | Model | Skill |
|---|---|---|---|---|
| Day ahead, leads 1 to 24 | 1,221 | 792 | 523 | 57.2% |
| Two days, 25 to 48 | 1,221 | 1,128 | 780 | 36.1% |
| Days 3 to 6, 49 to 144 | 1,219 | 1,266 | 954 | 21.8% |
| Week ahead, 145 to 168 | 1,214 | 1,214 | 1,029 | 15.2% |
| All leads, 1 to 168 | 1,219 | 1,171 | 878 | 28.0% |
- A day ahead, the gap is wide.523 MW against 1,221 for the seasonal naive and 792 for the last-day naive: skill 57.2%, with a Diebold-Mariano statistic of 12.8 after a Bartlett overlap correction.
- A week ahead, it nearly closes.1,029 against 1,214, skill 15.2%, DM 3.1, p = 0.0017. At the single lead 168 the statistic is 1.4, p = 0.17: after a week the model is a slightly better seasonal naive with a calendar attached.
- The intervals are honest, except in summer.The 80% day-ahead band is 1,563 MW wide and covers 77.5% of hours, against the baseline’s 79.0% at 3,824 MW. Coverage falls to 70.3% in June to August, and 58.6% in July 2025.
- It cannot pick the five peak days.Day-ahead peak error is 646 MW. The median gap between the fifth and sixth highest daily peaks of a base period is 128 MW.
By lead: two days ahead, 780 MW, skill 36.1%; days three to six, 954, skill 21.8%; across all leads, 878 against 1,219, skill 28.0%. The model’s error is nearly flat by day type, 500 to 561 MW, where the baseline’s nearly doubles on holidays, 2,251 against 1,223 MW on weekdays.
Intervals come from each lead’s trailing 365 out-of-sample residuals. Across five nominal levels the model covers 48.5, 77.5, 88.4, 93.8 and 97.3 against 50, 80, 90, 95 and 98: one to three points overconfident at every level, with calibration bought by width, the model’s band being 41% of the baseline’s. Both fail in summer, 70.3% coverage in June to August against 80.5% elsewhere. A 56-day residual window buys three summer points, widens the summer band by 200 MW and loses four points elsewhere, and is reported as not a fix.
The peak and the decision. Day-ahead peak error is 646 MW with a bias of minus 144, and the peak hour is called within an hour 80.2% of the time, against 1,293 MW and 70.8% for the baseline. On the hottest day of the window, 2026-07-14, the actual peak was 25,646 MW and the model called 24,035, with its worst hours 2,368 and 2,389 MW short and 7 of 24 hours inside its 80% band: it knows the shape and not the level.
Ontario’s largest consumers pay according to their draw in the five highest-demand hours of the May-to-April base period. The median gap between the fifth and sixth highest daily peaks over 24 complete periods is 128 MW, smaller than the model’s peak error in 23 of them. In 2025-26 the gap was 148 MW and 10 days had a peak within 646 MW of the fifth, so a consumer relying on this forecast would have curtailed on all 10. The cost is stated in days, because a dollar figure needs a load and a tariff this data does not hold.
What this cannot tell you
- What a weather-fed forecast would do. Every gap here is a weather gap.
- What the operator’s own forecast did over the same hours.
- Anything at intraday origins, since there is one origin a day.
- The dollar cost of an error, which is stated in days curtailed instead.
- The five days themselves. The fifth-to-sixth gap is tens of megawatts in most periods, and no forecast with an error in the hundreds resolves it.
- Whether a different refit interval or training window would score better. Neither sensitivity was run.
Artifacts
- A 40-cell notebook and a findings memo in forecasting vocabulary: baseline, skill, rolling origin, calibration, coverage, sharpness
- A decision log, D-01 to D-14, a data contract and a source-to-target map
- A committed holiday crosswalk of 254 rows
- A seven-table star exported and reconciled, 303,408 forecast rows, with a DAX measure layer and a zone-analyst row-level-security role specified; the Power BI dashboard itself is not yet built
Reproducibility
On the morning of 2026-09-07 the CI came back on 281 of 282 figures, including the full score table, every coverage and reliability figure and the hottest day hour by hour. By the afternoon the IESO had rewritten its two 2026 files with one more day, and the re-run exercised the drift regime for the first time: six whole-series figures moved inside their bands, the series from 213,456 to 213,480 hours and the 2026 mean from 17,086 to 17,080 MW, and the 248 scored figures that read the same file held exactly. The analysis window is fixed at 2026-08-31.
| Figures in the manifest after the 9 Sep corrections | 284 |
|---|---|
| Reproduced exactly, 2026-09-07 afternoon | 275 of 282 |
| Moved inside their band | 6 |
| Declared | 1 |
| Notebook run time on the committed baseline | 84 s |
Sources
- IESO public reports, hourly Ontario demand 2002 to 2026 and zonal demand 2003 to 2026, with a per-file SHA-256 and the publisher’s own created-at stamps in the manifest.
- The committed holiday crosswalk, computed from Easter arithmetic and nth-weekday rules.
- The findings memo, decision log, data contract and source-to-target map in 06-ontario-demand-forecast/memo.

