Forecast accuracy

A public scorecard for the numbers CEAtlas runs on — the resource record measured against meters and independent data, the operational forecast scorecards imported from CompoundVision and CEGridSight with their provenance, and a plain label on every card saying whether it is measured here, measured elsewhere, imported, or not yet measured at all. We publish this because a forecast you can't audit is a forecast you can't sell on.

Last updated: loading… Window: — Backtest version: —
Loading accuracy metrics…

Where the numbers come from

Every figure on a measured card is recomputed weekly from the per-record rows of the validation artefacts we publish in full — per-plant and per-site JSON, not headlines — by etl/run_accuracy_backtest.py. An imported card carries a sibling product's own published backtest — CompoundVision's forecast scorecard and power-curve calibration, CEGridSight's public price track record — with the source URL, the source's own stamp, its window and the time we fetched it printed on the card; those figures are never recomputed, re-rounded or re-worded here. Nothing on this page is typed in: a number that could not be recomputed from an artefact's records is flagged on its card, a check whose artefact is missing is shown as not measured rather than filled from memory, and an import that failed reads as not measured with the reason rather than as an old number posing as current. The raw artefacts: eia923_validation.json (US wind vs EIA-923 metered generation), gb_wind_validation.json (GB wind vs Elexon B1610 settlement meters), merra2_crosscheck.json (wind vs the independent MERRA-2 reanalysis), cfe8760_validation.json (solar vs PVGIS hourly output), and the scorecard itself, accuracy.json. Method, caveats, and what each result changed in the product: methodology — validation and the measured-anchored wind basis.

How we measure

Three forecast classes, three different yardsticks — and an honest label on each saying whether it is measured here, measured on another page, imported from a sibling product's own backtest, or not yet measured.

Resource record — model vs meter (measured here)

"Forecast" on this page means the model's number for a site-period, produced without sight of the meter; "realised" means an independent measurement. What is measured today:

Imported operational scorecards — CompoundVision and CEGridSight

Two sibling products keep their own backtests, so this page imports them rather than rebuilding them. CompoundVision archives every wind and solar fleet forecast it issues and scores it against metered fleet output on a rolling window, per region and on the operating fleet: skill against same-horizon persistence (above zero beats the naive forecast), NMAE as a fraction of installed capacity, and the share of outturns that fell inside its p10–p90 band. CEGridSight freezes each GB day-ahead price forecast vintage before settlement and grades it against the settled Elexon MID outturn — a harder benchmark than the day-ahead auction, because within-day trades price in late wind revisions and plant trips a day-ahead forecast cannot know: skill against yesterday's price and against a seasonal-naive baseline, error by forecast horizon from D+0 outward, tail recall on spikes and negative prices, and a streak of days beating both baselines. The rule for every imported card: we carry the source's numbers with their provenance — the URL, the source's own stamp, its window and the time we fetched it — and never recompute, re-round or re-word them; the server re-fetches each source hourly, a snapshot is committed with the weekly record, and one headline per day accrues here so a longer view than a source's own window is built from what was actually served, never back-filled. CEGridSight's live generation accuracy (per-fuel error over today's elapsed hours) is a same-day view rather than a rolling record, so it is not carried; generation-forecast accuracy on this page comes from CompoundVision only.

Day-ahead grid prices

Measured on the CONUS price scorecard — a separate, daily record in $/MWh (MAE, sMAPE, correlation, peak and spike hit rates per market, day-ahead lead against the market's published day-ahead price). It is carried on this page as a pointer only. GB day-ahead prices in £/MWh are imported from CEGridSight's own public track record — each forecast vintage frozen before settlement and graded against the settled Elexon MID outturn, per settlement period, with skill against yesterday's price and a seasonal-naive baseline and the error by lead day — and carried here with provenance, not recomputed.

Wind & solar generation forecasts (MW)

Imported from CompoundVision's rolling backtest of its own issued forecasts against metered fleet output, per region, on the operating fleet. The yardsticks: skill vs persistence (one minus the ratio of the forecast's error to the same-horizon naive forecast's; above zero beats it), NMAE (mean absolute error normalised by installed capacity, cross-asset comparable) and band coverage (the share of outturns that fell inside the p10–p90 band, which should match the band's width). CRPS is not yet carried: the source's coverage figure comes from its weekly CRPS artefact, but the score itself is not published. The map's own 7-day outlook (Open-Meteo GFS, computed in the browser at view time) is still not measured — no snapshot of it is persisted, so there is nothing archived to score against EIA-930 or Elexon B1610 output.

Site Suitability Scores

Suitability scores are normative — they reflect a stated preference function over land use and resource quality. We don't backtest them against operational outcomes (no obvious counterfactual). Instead we publish input-data provenance and version the score formula. See /methodology.

Method & honesty pledge

The scorecard is recomputed weekly from the committed validation artefacts, and every refresh is published, including the ones that look worse — that's the point. The imported scorecards are re-fetched hourly by the server and a snapshot is committed with the weekly record, so the page carries the newest import while the committed file stays the offline record. If a metric regresses week-on-week we annotate why in the notes. If we discover a bug in a metric we mark the affected runs and re-publish the corrected series. When the weekly job stops publishing for long enough that the record goes stale, the page says so in its notes rather than quietly serving an old record as current.

We will not:

Raw data & reproducibility

The per-record JSONs behind every card are public now — linked above and on each card — and the scorecard file itself is at /data/accuracy.json, readable at /api/accuracy (served with the newest imported snapshot overlaid, and the day-by-day import history beside it). Enterprise subscribers can request the underlying realised-vs-predicted series (Parquet) for the resource record. A methodology paper covering the backtest harness will follow once the operational series exist. Email hello@compoundingenergy.com if you'd like an early read.