1. Home
  2. Work
  3. 06

Project 06 · IESO hourly Ontario demand, 2002 to 2026

Ontario Demand Forecast with an Honest Baseline

A day ahead, a transparent linear model with no weather input misses hourly demand by 523 MW, 3.0% of demand, against 1,221 MW for the same-hour-last-week forecast, removing 57.2% of its error. A week ahead it removes 15.2%, and at the last hour of the week the two cannot be told apart. Its 80% interval covers 77.5% of hours, both it and the baseline fail in summer, and it cannot pick a top-five peak day.

Data
25 yearly hourly demand files and 24 zonal files from the IESO public report server
Scale
213,456 hours; 602 test origins; 5,880 model fits
Reproducibility
275 of 282 figures reproduced exactly on the 2026-09-07 fresh pull, 6 moved inside their drift band after the publisher added a day to its 2026 files, 1 declared.
01

The question

Issued at midnight for the next 168 hours, how far can a transparent model with no weather in it beat the same-hour-last-week forecast, a day and a week ahead? When it says 80% sure, is it right 80% of the time? And for a large consumer deciding whether to curtail tomorrow, is the day-ahead error small enough to tell a top-five peak day from the day that just misses?

02

The data and its scale

25 yearly files, 213,456 hours from 2002 to 2026, and 24 zonal files, 204,695 hours, with a per-file SHA-256 manifest carrying the publisher’s own timestamps. Each year exists on the server as a plain and a versioned copy, and all 25 pairs hash identical. Every date carries exactly 24 rows, so the series runs on a fixed standard-time clock and a 168-hour lag always lands on the same wall-clock hour.

03

What the source did

  1. B5One hour is missing from a closed year with no marker: 2025-05-01 hour 1, found by a generated spine. Hour 24 of April 30 reads 13,795 MW and hour 2 of May 1 reads 12,352; the gap is filled with the mean of its neighbours, 13,074 MW, and flagged. It is the only gap in 213,456 hours.
  2. B17The zonal report writes one column with a quoted thousands separator once it reaches four digits, on 129 of 204,695 rows, which a sniffed CSV dialect splits into two fields. The quote character is set explicitly and the separator stripped before the cast.
  3. B1172,152 zonal rows where the ten zones do not sum to the published total, all but four within 5 MW. Separately, 35 hours across two days in 2016 have every zone at zero, and 43 hours have market demand below Ontario demand.
  4. B8An outlier check written for data errors finds the twelve hours of the 14 August 2003 blackout, 2,270 MW at hour 17 against 22,380 a week before. It is an event, not an error; it stays, and it sits well before any training window.

Eighteen assertion rules, eight of them blocking, all at their expected values.

04

Method

Every day from 2024-12-31 to 2026-08-24 is an origin, 602 of them, chosen so that every lead from 1 to 168 is scored the same number of times inside a window ending on a complete month. Three forecasts per origin: the seasonal naive, the same hour and weekday one week earlier; a last-day naive, the origin day repeated; and one ordinary least squares regression per lead on sixteen inputs known at the origin, namely weekly lags, the origin day’s level, the target day’s type and three annual Fourier pairs.

The model refits every 28 origins on the trailing 1,825 origins whose targets were all observed, 5,880 fits, and no fit ever sees an outcome from its own forecast window. The 372 origins before the test window are scored and never reported; their errors build the prediction intervals. There is no weather in it, stated as the floor a weather-fed forecast must beat.

05

Finding

Mean absolute error by lead, 602 daily origins

Mean absolute error by lead, 602 originsPoints joined by lines, mean absolute error in megawatts for four lead buckets plus all leads. Same hour last week: 1,221, 1,221, 1,219, 1,214; all leads 1,219. Last day repeated: 792, 1,128, 1,266, 1,214; all leads 1,171. Model with no weather: 523, 780, 954, 1,029; all leads 878. Skill against the seasonal naive: 57.2%, 36.1%, 21.8%, 15.2%, and 28.0% over all leads. A day ahead the Diebold-Mariano statistic is 12.8; a week ahead 3.1, p = 0.0017.02004006008001,0001,2001,400Mean absolute error, MWDay aheadleads 1 to 24skill 57.2%Two days25 to 48skill 36.1%Days 3 to 649 to 144skill 21.8%Week ahead145 to 168skill 15.2%All leads1 to 168skill 28.0%Day ahead, same hour last week: 1,221 MWTwo days, same hour last week: 1,221 MWDays 3 to 6, same hour last week: 1,219 MWWeek ahead, same hour last week: 1,214 MWAll leads, same hour last week: 1,219 MWDay ahead, last day repeated: 792 MWTwo days, last day repeated: 1,128 MWDays 3 to 6, last day repeated: 1,266 MWWeek ahead, last day repeated: 1,214 MWAll leads, last day repeated: 1,171 MWDay ahead, model, no weather: 523 MWTwo days, model, no weather: 780 MWDays 3 to 6, model, no weather: 954 MWWeek ahead, model, no weather: 1,029 MWAll leads, model, no weather: 878 MW5231,2211,029DM 3.1, p = 0.0017same hour last weeklast day repeatedmodel, no weather
Mean absolute error by lead, 602 originsPoints joined by lines, mean absolute error in megawatts for four lead buckets plus all leads. Same hour last week: 1,221, 1,221, 1,219, 1,214; all leads 1,219. Last day repeated: 792, 1,128, 1,266, 1,214; all leads 1,171. Model with no weather: 523, 780, 954, 1,029; all leads 878. Skill against the seasonal naive: 57.2%, 36.1%, 21.8%, 15.2%, and 28.0% over all leads. A day ahead the Diebold-Mariano statistic is 12.8; a week ahead 3.1, p = 0.0017.04008001,2001,400Mean absolute error, MWDay 157.2%Day 236.1%3 to 621.8%Day 715.2%All28.0%skillDay ahead, same hour last week: 1,221 MWTwo days, same hour last week: 1,221 MWDays 3 to 6, same hour last week: 1,219 MWWeek ahead, same hour last week: 1,214 MWAll leads, same hour last week: 1,219 MWDay ahead, last day repeated: 792 MWTwo days, last day repeated: 1,128 MWDays 3 to 6, last day repeated: 1,266 MWWeek ahead, last day repeated: 1,214 MWAll leads, last day repeated: 1,171 MWDay ahead, model, no weather: 523 MWTwo days, model, no weather: 780 MWDays 3 to 6, model, no weather: 954 MWWeek ahead, model, no weather: 1,029 MWAll leads, model, no weather: 878 MW5231,2211,029p = 0.0017same hour last weeklast day repeatedmodel, no weather
Solid: the model. Dashed: same hour last week. Dotted: the origin day repeated. Skill is the share of the seasonal naive’s error the model removes. The hour-by-hour curve is in the notebook’s chart below.06 · IESO hourly demand
Data behind this chart
Mean absolute error by lead, MW, 06
LeadSame hour last weekLast dayModelSkill
Day ahead, leads 1 to 241,22179252357.2%
Two days, 25 to 481,2211,12878036.1%
Days 3 to 6, 49 to 1441,2191,26695421.8%
Week ahead, 145 to 1681,2141,2141,02915.2%
All leads, 1 to 1681,2191,17187828.0%
  1. A day ahead, the gap is wide.523 MW against 1,221 for the seasonal naive and 792 for the last-day naive: skill 57.2%, with a Diebold-Mariano statistic of 12.8 after a Bartlett overlap correction.
  2. A week ahead, it nearly closes.1,029 against 1,214, skill 15.2%, DM 3.1, p = 0.0017. At the single lead 168 the statistic is 1.4, p = 0.17: after a week the model is a slightly better seasonal naive with a calendar attached.
  3. The intervals are honest, except in summer.The 80% day-ahead band is 1,563 MW wide and covers 77.5% of hours, against the baseline’s 79.0% at 3,824 MW. Coverage falls to 70.3% in June to August, and 58.6% in July 2025.
  4. It cannot pick the five peak days.Day-ahead peak error is 646 MW. The median gap between the fifth and sixth highest daily peaks of a base period is 128 MW.

By lead: two days ahead, 780 MW, skill 36.1%; days three to six, 954, skill 21.8%; across all leads, 878 against 1,219, skill 28.0%. The model’s error is nearly flat by day type, 500 to 561 MW, where the baseline’s nearly doubles on holidays, 2,251 against 1,223 MW on weekdays.

Intervals come from each lead’s trailing 365 out-of-sample residuals. Across five nominal levels the model covers 48.5, 77.5, 88.4, 93.8 and 97.3 against 50, 80, 90, 95 and 98: one to three points overconfident at every level, with calibration bought by width, the model’s band being 41% of the baseline’s. Both fail in summer, 70.3% coverage in June to August against 80.5% elsewhere. A 56-day residual window buys three summer points, widens the summer band by 200 MW and loses four points elsewhere, and is reported as not a fix.

The peak and the decision. Day-ahead peak error is 646 MW with a bias of minus 144, and the peak hour is called within an hour 80.2% of the time, against 1,293 MW and 70.8% for the baseline. On the hottest day of the window, 2026-07-14, the actual peak was 25,646 MW and the model called 24,035, with its worst hours 2,368 and 2,389 MW short and 7 of 24 hours inside its 80% band: it knows the shape and not the level.

Ontario’s largest consumers pay according to their draw in the five highest-demand hours of the May-to-April base period. The median gap between the fifth and sixth highest daily peaks over 24 complete periods is 128 MW, smaller than the model’s peak error in 23 of them. In 2025-26 the gap was 148 MW and 10 days had a peak within 646 MW of the fifth, so a consumer relying on this forecast would have curtailed on all 10. The cost is stated in days, because a dollar figure needs a load and a tariff this data does not hold.

06

What this cannot tell you

  • What a weather-fed forecast would do. Every gap here is a weather gap.
  • What the operator’s own forecast did over the same hours.
  • Anything at intraday origins, since there is one origin a day.
  • The dollar cost of an error, which is stated in days curtailed instead.
  • The five days themselves. The fifth-to-sixth gap is tens of megawatts in most periods, and no forecast with an error in the hundreds resolves it.
  • Whether a different refit interval or training window would score better. Neither sensitivity was run.
07

Artifacts

  • A 40-cell notebook and a findings memo in forecasting vocabulary: baseline, skill, rolling origin, calibration, coverage, sharpness
  • A decision log, D-01 to D-14, a data contract and a source-to-target map
  • A committed holiday crosswalk of 254 rows
  • A seven-table star exported and reconciled, 303,408 forecast rows, with a DAX measure layer and a zone-analyst row-level-security role specified; the Power BI dashboard itself is not yet built
Mean absolute error by lead hour, 1 to 168, for the last-day naive, the seasonal naive and the linear model over 602 origins. All three cycle daily; the model starts lowest at lead 1 and stays below both baselines, the gap narrowing toward lead 168.
Error hour by hour across the week. error_by_horizon.png · original in the repository
Reliability diagram: empirical coverage against nominal coverage at 50, 80, 90, 95 and 98%, for the seasonal naive and the linear model a day and a week ahead. All four lines sit just under the diagonal of perfect calibration; the model’s day-ahead 80% band is 1,563 MW wide against 3,824 MW for the baseline.
Interval calibration by nominal level. interval_calibration.png · original in the repository
08

Reproducibility

On the morning of 2026-09-07 the CI came back on 281 of 282 figures, including the full score table, every coverage and reliability figure and the hottest day hour by hour. By the afternoon the IESO had rewritten its two 2026 files with one more day, and the re-run exercised the drift regime for the first time: six whole-series figures moved inside their bands, the series from 213,456 to 213,480 hours and the 2026 mean from 17,086 to 17,080 MW, and the 248 scored figures that read the same file held exactly. The analysis window is fixed at 2026-08-31.

CI status for this project
Figures in the manifest after the 9 Sep corrections284
Reproduced exactly, 2026-09-07 afternoon275 of 282
Moved inside their band6
Declared1
Notebook run time on the committed baseline84 s
09

Sources

  1. IESO public reports, hourly Ontario demand 2002 to 2026 and zonal demand 2003 to 2026, with a per-file SHA-256 and the publisher’s own created-at stamps in the manifest.
  2. The committed holiday crosswalk, computed from Easter arithmetic and nth-weekday rules.
  3. The findings memo, decision log, data contract and source-to-target map in 06-ontario-demand-forecast/memo.