Forecast accuracy
A public scorecard for the numbers CEAtlas runs on — the resource record measured against meters and independent data, the operational forecast scorecards imported from CompoundVision and CEGridSight with their provenance, and a plain label on every card saying whether it is measured here, measured elsewhere, imported, or not yet measured at all. We publish this because a forecast you can't audit is a forecast you can't sell on.
Where the numbers come from
Every figure on a measured card is recomputed weekly from the
per-record rows of the validation artefacts we publish in full —
per-plant and per-site JSON, not headlines — by
etl/run_accuracy_backtest.py. An imported card
carries a sibling product's own published backtest — CompoundVision's
forecast scorecard and power-curve calibration, CEGridSight's public
price track record — with the source URL, the source's own stamp, its
window and the time we fetched it printed on the card; those figures
are never recomputed, re-rounded or re-worded here. Nothing on this
page is typed in: a number that could not be recomputed from an
artefact's records is flagged on its card, a check whose artefact is
missing is shown as not measured rather than filled from memory, and an
import that failed reads as not measured with the reason rather than as
an old number posing as current.
The raw artefacts:
eia923_validation.json (US wind
vs EIA-923 metered generation),
gb_wind_validation.json (GB
wind vs Elexon B1610 settlement meters),
merra2_crosscheck.json (wind
vs the independent MERRA-2 reanalysis),
cfe8760_validation.json (solar
vs PVGIS hourly output), and the scorecard itself,
accuracy.json.
Method, caveats, and what each result changed in the product:
methodology — validation
and the
measured-anchored wind basis.
How we measure
Three forecast classes, three different yardsticks — and an honest label on each saying whether it is measured here, measured on another page, imported from a sibling product's own backtest, or not yet measured.
Resource record — model vs meter (measured here)
"Forecast" on this page means the model's number for a site-period, produced without sight of the meter; "realised" means an independent measurement. What is measured today:
- Monthly-shape Pearson r — per plant or site, on metered months only (for EIA-923, the M-frame: monthly respondents, whose months are meter readings rather than annual allocations). Tests whether the model tracks the seasonal shape.
- Metered ÷ modeled ratio, by build vintage — the level check. Metered output includes curtailment and real availability, so the ratio is an upper bound on resource-model error, and it is reported per vintage because fleet age is the largest known driver.
- Worst-year agreement — whether the model's worst wind year is the meter's worst year (for EIA-923, the fleet year), or — against MERRA-2 — falls within that reanalysis's three worst years. The drought question a financier actually asks.
- Hourly r and DC-basis bias for solar — against PVGIS hourly output for fixed-tilt and single-axis-tracker arrays; the AC-convention offset is the inverter-loading design choice and is reported separately from the physics error.
- Power-curve calibration, out of sample — imported from CompoundVision: per-farm RMSE of forecast power against Elexon B1610 (GB) and ENTSO-E A73 (EU) meters, IEC reference curve before vs site-calibrated curve after, re-fit and scored on a day-block holdout. It sits in this group because it tests the curve in isolation (air density pinned, one wind member, no statistical post-processing) — it is not an operational forecast backtest.
Imported operational scorecards — CompoundVision and CEGridSight
Two sibling products keep their own backtests, so this page imports them rather than rebuilding them. CompoundVision archives every wind and solar fleet forecast it issues and scores it against metered fleet output on a rolling window, per region and on the operating fleet: skill against same-horizon persistence (above zero beats the naive forecast), NMAE as a fraction of installed capacity, and the share of outturns that fell inside its p10–p90 band. CEGridSight freezes each GB day-ahead price forecast vintage before settlement and grades it against the settled Elexon MID outturn — a harder benchmark than the day-ahead auction, because within-day trades price in late wind revisions and plant trips a day-ahead forecast cannot know: skill against yesterday's price and against a seasonal-naive baseline, error by forecast horizon from D+0 outward, tail recall on spikes and negative prices, and a streak of days beating both baselines. The rule for every imported card: we carry the source's numbers with their provenance — the URL, the source's own stamp, its window and the time we fetched it — and never recompute, re-round or re-word them; the server re-fetches each source hourly, a snapshot is committed with the weekly record, and one headline per day accrues here so a longer view than a source's own window is built from what was actually served, never back-filled. CEGridSight's live generation accuracy (per-fuel error over today's elapsed hours) is a same-day view rather than a rolling record, so it is not carried; generation-forecast accuracy on this page comes from CompoundVision only.
Day-ahead grid prices
Measured on the CONUS price scorecard — a separate, daily record in $/MWh (MAE, sMAPE, correlation, peak and spike hit rates per market, day-ahead lead against the market's published day-ahead price). It is carried on this page as a pointer only. GB day-ahead prices in £/MWh are imported from CEGridSight's own public track record — each forecast vintage frozen before settlement and graded against the settled Elexon MID outturn, per settlement period, with skill against yesterday's price and a seasonal-naive baseline and the error by lead day — and carried here with provenance, not recomputed.
Wind & solar generation forecasts (MW)
Imported from CompoundVision's rolling backtest of its own issued forecasts against metered fleet output, per region, on the operating fleet. The yardsticks: skill vs persistence (one minus the ratio of the forecast's error to the same-horizon naive forecast's; above zero beats it), NMAE (mean absolute error normalised by installed capacity, cross-asset comparable) and band coverage (the share of outturns that fell inside the p10–p90 band, which should match the band's width). CRPS is not yet carried: the source's coverage figure comes from its weekly CRPS artefact, but the score itself is not published. The map's own 7-day outlook (Open-Meteo GFS, computed in the browser at view time) is still not measured — no snapshot of it is persisted, so there is nothing archived to score against EIA-930 or Elexon B1610 output.
Site Suitability Scores
Suitability scores are normative — they reflect a stated preference function over land use and resource quality. We don't backtest them against operational outcomes (no obvious counterfactual). Instead we publish input-data provenance and version the score formula. See /methodology.
Method & honesty pledge
The scorecard is recomputed weekly from the committed validation artefacts, and every refresh is published, including the ones that look worse — that's the point. The imported scorecards are re-fetched hourly by the server and a snapshot is committed with the weekly record, so the page carries the newest import while the committed file stays the offline record. If a metric regresses week-on-week we annotate why in the notes. If we discover a bug in a metric we mark the affected runs and re-publish the corrected series. When the weekly job stops publishing for long enough that the record goes stale, the page says so in its notes rather than quietly serving an old record as current.
We will not:
- Cherry-pick the window — each check's window is fixed and stated on its card (the metered years, the reanalysis span, the PVGIS year); an operational series carries the rolling window its own card states, read from that series’ record, never chosen here.
- Aggregate selectively — every reported vintage / site type / market is shown, and a check that cannot be computed is shown as not measured.
- Hide a deteriorating series — if accuracy drops, you see it here.
- Recompute, re-round or re-word an imported number — an imported card carries the source's own figure, stamp and window, or it reads not measured with the reason.
Raw data & reproducibility
The per-record JSONs behind every card are public now — linked above
and on each card — and the scorecard file itself is at
/data/accuracy.json, readable at
/api/accuracy (served with the newest imported snapshot
overlaid, and the day-by-day import history beside it). Enterprise subscribers can request the
underlying realised-vs-predicted series (Parquet) for the resource
record. A methodology paper covering the backtest harness will follow
once the operational series exist. Email
hello@compoundingenergy.com
if you'd like an early read.