Project ci · five shipped projects, re-run on fresh pulls · GitHub Actions
Reproducibility CI
Every figure five of the projects quote, 919 in all, is re-run on a fresh pull of its source every Monday and on every push. Same bytes must give the same number to the precision the memo prints it, or the run fails. The first run found a notebook that could not run on a clean checkout, 42 uncertainty figures the committed notebook no longer printed, and eight sentences the tables did not support as written.
Status as recorded, one mark per figure
- reproduced exactly
- drift inside band
- declared
- failed: 0
Data behind this chart
| Project | Listed | Reproduced | Drift in band | Declared |
|---|---|---|---|---|
| 01 Seller risk | 56 | 56 | 0 | 0 |
| 02 Funnel | 22 | 22 | 0 | 0 |
| 03 Defence | 337 | 334 | 0 | 3 |
| 04 Wind | 220 | 215 | 0 | 5 |
| 06 Demand | 284 | 277 | 6 | 1 |
| Total | 919 | 904 | 6 | 9 |
The question
Several hundred figures were each printed by a notebook cell on the day a memo was written, and nothing in the repo said whether they would come back the same on a new machine, or after the publisher refreshed the file. Do the numbers the READMEs and memos quote come back on a clean checkout and a fresh pull, and when a source changes, how far may each kind of figure move before the memo needs a dated update?
What it is
One YAML manifest per project lists every quoted figure with the sentence as printed, the documents that quote it, its kind, the query that reads it back from the tables the project’s own notebook builds, or a regular expression over the notebook’s printed output where the figure never reaches a table, the committed value, and the sources it depends on.
A runner fetches the two Kaggle datasets when they are absent, executes each notebook headless, cell by cell, in its own folder so it re-pulls and rebuilds everything, fingerprints each source, re-measures every figure read-only against the rebuilt DuckDB file, and classifies under two regimes. Identical bytes must give an identical figure at the memo’s precision, or the run fails. Changed bytes are publisher drift, read under a per-source contract of bands by kind. Figures that cannot be checked are declared with a reason and still measured, never dropped.
A GitHub Actions workflow has a plan job that turns the manifest folder into a matrix, one reproduce job per project, and a publish job that writes dated reports, a summary and a badge to an orphan ci-reports branch, so main stays hand-committed.
What the first run found
- 1The wind-repowering notebook could not run top to bottom on a clean checkout. An exploratory cell opened two downloads that the publisher now answers with a landing page; a one-cell guard fixed it and no number changed.
- 2The same memo’s 42 uncertainty statements were not what the committed notebook printed, while every point estimate held. Two clean runs agreed with each other on all 42; the placebo p-value read 0.167 in the memo and 0.268 on a clean run.
- 3Eight memo sentences the tables did not support as written, among them a count read from a three-row display where the table holds 35 such hours, a two that is a three, and two interval bounds off by a point.
- 4Five demand-forecast figures rounded twice from a one-decimal table. They stay as printed, because the manifest checks the table value behind each.
Each was recorded in its manifest entry and in the findings memo. The first three were closed on 2026-09-09: the repowering notebook run by hand and its 42 figures re-measured, the eight sentences corrected, and the manifests updated to 919 figures with 9 declared.
Finding
Fresh-pull run of 2026-09-07, figure by figure
Data behind this chart
| Project | Listed | Reproduced | Drift in band | Declared |
|---|---|---|---|---|
| 01 Seller risk | 56 | 56 | 0 | 0 |
| 02 Funnel | 22 | 22 | 0 | 0 |
| 03 Defence | 337 | 334 | 0 | 3 |
| 04 Wind | 220 | 173 | 0 | 47 |
| 06 Demand | 282 | 275 | 6 | 1 |
| Total | 917 | 860 | 6 | 51 |
- 917 figures, five projects.56, 22, 337, 220 and 282 figures in the five manifests, every one tied to the query that reads it back.
- Nothing failed.All five sources came back byte-identical to the memos’ pulls, so every figure was held to an exact match: 866 of 917 reproduced, 51 declared, 9 of them design constants and notebook properties and 42 bootstrap-derived.
- Drift, seen once on real bytes.That afternoon the IESO rewrote its two 2026 files with one more day. Six whole-series figures moved inside their bands; the 248 scored figures that read the same file held exactly.
- 919 after the corrections.With the 42 re-measured and the sentences fixed, the manifests list 919 figures and declare 9.
The workflow’s first run on GitHub, on 9 Sep 2026, completed green in 14 minutes and created the reports branch. With the two Kaggle secrets stored, a dispatched run reproduced every project and the badge read 860 of 917, 6 drifted, green. The push that afternoon, with the 42 figures re-measured and the manifests at 919, took the badge to 904 of 919, 6 drifted. Notebook run times on the committed baseline: 7.6 s, 15.5 s, 7.8 s, 424 s and 84 s.
What this cannot tell you
- Whether a figure is right. It tests agreement between memo and notebook, and that the agreement survives a re-pull.
- Whether a change in a figure matters. The bands are one person’s judgement, written down.
- How most sources behave when they drift, since the drift regime is untested until a source actually changes.
- Whether a stopped notebook is a real failure: a transient network timeout that survives two attempts reads as the notebook stopping.
- Anything about project 05, which is excluded.
Artifacts
- The runner,
reproduce.py: one file, the standard library plus DuckDB, PyYAML, nbformat and nbclient - Five manifests with committed fingerprint files
- The drift contract as
policy.yml, with its reasons in a data contract memo covering nine sources - A decisions memo of twenty-one decisions, each with what would reverse it, and a findings memo
- The dated reports, the workflow, and the root README badge and paragraph
Sources
- The five project sources as fingerprinted: 14 Kaggle files by size, 1 CanadaBuys hash, 1 USGS API hash, 1 archive hash, 13 EIA-923 hashes and 49 IESO hashes.
- The first fresh-pull run, 2026-09-07, and the first GitHub run, 2026-09-09, with dated reports on the ci-reports branch.
- The findings, decisions and data contract memos in ci/memo.