Methodology & limitations
How CEAtlas computes its proprietary scores — the formulas, the inputs, and where they break. We publish this because a number you can't audit isn't a number worth quoting.
Connection Difficulty Score (CDS)
A single 0–100 number answering one question every developer asks: "is this place hard or easy to plug in?" Higher = harder. The score is sampled at the click point you select.
Formula
+ 0.25 · distance_difficulty
+ 0.25 · constraint_exposure
+ 0.10 · curtailment_activity
Each sub-score is on the 0–100 scale and oriented so that higher = more
difficult. The weights default to the values above; the JavaScript API
accepts a weights override for callers with a different
project profile (e.g. a BESS investor may weight curtailment higher).
Sub-scores
1. Voltage headroom difficulty (40%).
Computed as 100 − headroom_score, where
headroom_score is the existing CEAtlas voltage-headroom
heuristic (see §Voltage Headroom).
Inputs: the highest-voltage substation within the search radius, the
reference asset MW (default 100 MW), and the required voltage class
from CIGRE bands.
2. Substation distance difficulty (25%). Distance to the closest viable-voltage substation, mapped through a hand-tuned curve:
| Distance | Difficulty |
|---|---|
| ≤ 0.5 km | 0 |
| 0.5 – 2 km | 5 – 20 |
| 2 – 5 km | 20 – 40 |
| 5 – 10 km | 40 – 60 |
| 10 – 20 km | 60 – 85 |
| 20 – 30 km | 85 – 100 |
| > 30 km | 100 |
3. Constraint exposure (25%). We sample the nearest GB ETYS boundary at the click point and take its 30-day mean utilisation. Mapping:
| 30-day utilisation | Difficulty |
|---|---|
| < 20% | 10 |
| 20 – 40% | 25 |
| 40 – 60% | 45 |
| 60 – 80% | 70 |
| 80 – 95% | 85 |
| ≥ 95% | 100 |
Outside GB (or in tiles where no ETYS boundary is rendered), we use a
neutral fallback of 30 and note this in the
notes array returned with the score. NESO publishes
day-ahead flow data for ~10 of the 34 ETYS boundaries so blind spots
exist even inside GB.
4. Curtailment activity (10%). Count of BMU constraint actions (Elexon BOAL acceptances) within the search radius over the last reporting period. Mapping:
| BMU actions within radius | Difficulty |
|---|---|
| 0 | 0 |
| 1 – 3 | 15 |
| 4 – 10 | 35 |
| 11 – 25 | 55 |
| 26 – 50 | 75 |
| > 50 | 90 |
Inputs & data sources
- Substations: PyPSA-Eur (EU) + OpenStreetMap (global, incl. 75k US via chunked Overpass)
- GB constraint boundaries: NESO ETYS published flows & limits, 30-day rolling window
- BMU constraint actions: Elexon BMRS BOAL endpoint — point-in-time snapshot (30-day window ending at the layer build date), refreshed manually
- Voltage class thresholds: CIGRE / IEC standard bands
Known limitations
- NESO TEC queue position by GSP — that data is not yet machine-readable; a planned input once we have it
- Power-flow feasibility at the specific connection point — that needs CEForesight (Compounding Energy's 8760 power-flow and dispatch simulator); N-1 contingency and reliability screening is CESentinel
- Bilateral connection costs the TSO will quote you — those depend on works at the substation we can't see from the public record
- Distribution-network operator constraints below 132 kV — we cover transmission, not DNO licensed areas
- Forward-looking ETYS upgrade dates (mostly because publication lag means our "future" is already public)
CDS is best used for relative comparison across sites, not as a go/no-go threshold. A real connection feasibility study still requires a TSO conversation and a power-flow run.
Versioning
Current version: CDS v1.0 (May 2026 launch).
Weight changes and new sub-scores will be released as versioned
updates with a brief changelog. If you cite a CDS value in a
deck or filing, cite the version too.
24/7 CFE Study (real 8760-hour dispatch)
The study pulls a full year of actual hourly weather for the site from CEAtlas's private ERA5 hourly weather store (store-first; the Open-Meteo archive, ERA5 reanalysis family, is the transparent fallback and every result records which one served it), converts it to hourly wind and solar capacity factors, and runs a chronological hour-by-hour dispatch across all 8,760 hours with a stateful battery. That sequencing — not a 12×24 month-hour average — is what makes the 24/7 CFE %, the grid-import residual, and the storage sizing real: it sees multi-day weather lulls that an averaged model cannot.
Resource models.
Wind and solar CFs come from cf_model.py v2.0.0 — the one canonical,
CompoundVision-parity physics engine shared by the map, the API, the studies and the
forecasts (see Capacity factor & the weather store
for the full derivation). Both fuels return net expected-farm CFs:
wind is an IEC 61400-12 smootherstep curve (300 W/m² reference turbine,
rated 10.36 m/s at 100 m) run through Bastankhah wake, 0.97 availability,
thermal, storm-hysteresis and electrical losses; solar is a full
SPA → Erbs → Perez → King chain with an
ILR 1.30 AC clip, for a single-axis tracker (default) or
fixed tilt (solar_array). The battery round-trips at
90% (RTE), applied as a per-leg square-root. These are
single-reference-plant CFs, not a specific turbine or inverter — a project's own kit
shifts the absolute number.
Weather years. Any year from 1940 to the last
complete calendar year is selectable; pre-1950 is ECMWF's preliminary
ERA5 back-extension and is labelled as such. The multi-year mode fixes the
generation mix once and re-runs the dispatch across up to 86 real
weather years (1940–2025) to report a P50 with an
inter-annual band — so a single calm or sunny year can't flatter the
result. Percentiles follow the exceedance convention
(as on the maps): P90 is the year-in-10 downside — the
CFE beaten in nine years of ten, the number an offtaker or lender should
underwrite — and P10 is the one-in-ten good year.
Resource droughts (worst-window analysis)
Every multi-year study also reports resource droughts: the longest contiguous windows in which combined wind + solar output stayed below 10% of the load, with storage deliberately ignored — a drought is a property of the weather and the chosen mix, not of battery sizing. The study returns each year's longest window, the P50 and P90 of those annual maxima ("plan for an N-hour drought nine years in ten" — the sentence a long-duration-storage conversation starts from), and the record's five worst windows with their actual dates. Events of the size that drives seasonal-storage economics occur a handful of times per generation; they are only measurable because the record spans 86 years.
Portfolio diversification (multi-site)
The portfolio mode takes 2–4 sites with per-site nameplates, sums their hourly generation in every weather year, and dispatches the combined profile against one load. As the baseline, each site alone carries the same total nameplate (the portfolio's aggregate wind:solar split) — so the comparison is like-for-like on capacity, and the difference is geography alone. The headline, diversification value, is the portfolio's P90 CFE minus the best single site's: what spatial spread across weather regimes is worth in the bad years, measured on real weather rather than asserted from correlation matrices. Sites in anti-correlated regimes can be worth tens of percentage points at P90; sites in the same regime are honestly reported as worth little.
Validation (published, not asserted). Solar is validated against
PVGIS seriescalc across six globally-stratified sites (2020): hourly
correlation 0.902 (tracker) / 0.952 (fixed), DC-basis
physics bias +0.006 / +0.011 CF, GHI input agreement
0.969. On the served AC convention the CFs sit +0.058 CF
above PVGIS's 1 kWp DC reference — that offset is the ILR 1.30
design (AC nameplate = DC / 1.30), reported separately, not an error.
Wind is validated against independent measured data — the
MERRA-2 cross-validation and the EIA-923 metered-generation study below, plus
the GB settlement-meter check; the earlier PVGIS wind comparison was a
same-family consistency check and is no longer the basis of the claim.
MERRA-2 cross-validation (wind, independent reanalysis). To bound
reanalysis-family risk we cross-checked against NASA's MERRA-2 (GEOS — an
assimilation system independent of ECMWF's) at 36 globally-stratified windy on-land
sites, chain-symmetric: both sides as raw 100 m reanalysis speeds
(ours native from the ERA5 store; MERRA-2's 50 m winds lifted with a per-site
empirical shear exponent from its own 10/50 m ratio). The claims our studies
actually rest on hold up: inter-annual variability agrees — median
annual-speed correlation 0.89 over 1981–2024 (CF-domain
0.88 over 2001–2024) — and the drought years agree:
at 28 of 36 sites our worst wind year of the 44 falls inside MERRA-2's worst three
(~7% would by chance). Where the families part is absolute magnitude:
MERRA-2 runs +28% windier than ERA5 on this sample, a documented
divergence over land where published tall-tower comparisons generally favour ERA5.
We publish that number rather than hide it: it is the honest scale of
between-model uncertainty on absolute levels, it is exactly what
per-site met-data calibration exists to collapse, and
measured-generation validation (below) closes the loop. Full per-site
results: merra2_crosscheck.json.
Sampling frame: windy on-land regions — these statistics are for that sample, not
map-wide.
EIA-923 measured-generation validation (wind, metered reality).
The decisive test: 583 US onshore wind plants (87 GW, 34 states, 2019–2024)
with monthly net generation from EIA-923, each re-modeled twice — once at the EIA
plant point with the standard 100 m class-curve chain, and once
turbine-true: actual USGS-USWTDB turbine positions (45,437 turbines),
capacity-weighted hub heights, and each plant's real specific power setting the
power curve. One honesty detail most comparisons skip: EIA's monthly values from
annual-frequency respondents are allocations, not meter readings — so
every monthly-shape statistic below uses only true monthly-metered plant-years.
What holds up: the shape and the rankings — median metered-month
correlation 0.88 (570 plants; p90 0.97), per-plant
annual correlation 0.61, and both metered reality and the model agree
2023 was the fleet's worst wind year (the fleet-aggregate annual
series correlates at 0.84, though with only six years that worst-year
agreement is the sturdier claim). What runs hot: absolute levels —
and the gap is mostly the old fleet. Metered generation lands at a median
82% of the turbine-true model overall, but split by build year:
pre-2012 plants 74%, 2012–2015 86%, and
2016–2018 plants 92% — for plants resembling new
builds (what a siting screen actually predicts), the model runs ~8% high on the
turbine-true chain, and the gap varies strongly by region — ERCOT
0.75 vs SPP/CAISO 0.91–0.94 — a bundle of curtailment
(forced and economic), regional model bias, and fleet losses that the data does
not let us separate by name. The rest of the old-fleet gap is degradation
and availability we deliberately model as a fixed modern-farm loss. Two upstream
conclusions already actioned: this validation exposed the auto class layer
assigning low-wind machines too widely (its class thresholds are now
fleet-calibrated against the machines actually bought), and MERRA-2 is adjudicated
— metered reality sits below our chain while MERRA-2 sits +28% above it, so the
ERA5 family is the closer one, and still warm. Treat absolute wind CFs as
indicative, use per-site met-data calibration to
collapse site-level bias, and for the vintage/curtailment gap itself switch on
the correction this validation now feeds: the
measured-anchored wind basis
(wind_basis=measured_us, next section). Full per-plant table:
eia923_validation.json. Sampling frame:
US onshore ≥50 MW plants fully in service before 2019, gross-of-storage.
The measured-anchored wind basis
The validation above is now a product control. The single-year and multi-year
24/7 studies (and the API) accept wind_basis=measured_us: an opt-in
basis that scales the modeled wind CF by the empirical cohort
median from the EIA-923 validation, fitted on the IEC-2 chain the
studies actually run. The headline is what it does not change:
for 2016–2018-built plants, metered generation lands at
×0.99 of this chain (95% CI 0.96–1.02) — the
default study chain is measured-true for modern builds, within confidence, at
national scale. The toggle's value is honesty at the edges: older-asset
analysis (×0.88 for 2012–2015 builds, ×0.71
pre-2012 — via the API's basis_vintage), and the few regions that
stay statistically distinct after a CI gate (SPP ×1.05 — metered
above model; ERCOT ×0.58 for the pre-2012 fleet;
non-plains US ×0.95). Four honesty rules govern it.
Cohort-anchored, never extrapolated: every factor is the median of an
observed vintage-by-region cell with a bootstrap 95% CI; a regional cell ships
only when its CI excludes the cohort median, and 12 plants with ratios outside
(0.3, 1.3) — capacity-attribution and repowering artifacts — are excluded, by a
published, re-runnable recipe (etl/build_wind_anchor.py).
All-in, not decomposed: a cell bundles forced and economic curtailment,
regional model bias, and real fleet losses — provably inseparable by name (LBNL
reports SPP curtailed more than ERCOT in 2022, 9.2% vs 4.7%, yet SPP's
modern plants deliver more of the modeled energy). Annual, not hourly:
the factor anchors annual energy and is applied uniformly across hours, while
the losses it bundles concentrate in high-output hours — 24/7 CFE results carry
roughly ±1–2 pp of extra shape uncertainty. Default-off and labeled:
the resource basis remains the default, every measured-anchored result carries
the factor, cell, sample size and CI in its inputs, and the maps are never
silently re-baked. Known frame limits, stated plainly: the fleet is 2023
survivors (factors may flatter plants that retired); partial repowers keep
their original vintage; and "new build" means 2016–2018 machines metered at
ages one to eight, applied unchanged to newer builds — potentially optimistic
for lifetime P50, potentially pessimistic for genuinely better machines. The
full table with CIs is at
wind_measured_anchor.json;
non-US use falls back to labeled US-fitted cohorts — except GB, which now
has its own cells: basis_vintage=gb_offshore (×0.78,
CI 0.69–0.82) and gb_onshore (×0.58, CI 0.56–0.62;
Scottish-weighted, as-delivered including constraint curtailment — a revenue
model bears that risk, and the curtailment decomposition is published
alongside).
The GB check (settlement meters, second continent). The same
comparison run against Elexon B1610 settlement-metered generation — 62 GB wind
BMU-groups (16 offshore farms among them), 13 GW, Feb 2019–Dec 2024
— and reported the way it deserves: as a negative result with structure.
What replicates is temporal shape at a fixed site (median monthly correlation
0.89; 0.94 per-site offshore, 0.86
capacity-weighted). What fails is everything a per-site correction would need:
cross-site correlation is ~zero (offshore +0.07,
onshore +0.25) — the model does not rank these GB sites against
each other — and per-site ratios disperse widely (offshore IQR 0.69–0.83;
onshore 0.52–0.70 with outliers to 1.8). Fleet-level, metered output is
×0.73–0.78 of the IEC-2 chain offshore and ×0.58–0.66 for
the onshore sample — which is entirely Scottish, transmission-metered,
and sits behind the B6 constraint boundary, so it reads as a
constrained-fleet result, not a GB-onshore truth. On matched vintages the
US-GB gap roughly halves (GB 2016–18 builds ×0.79 vs US
×0.99); ratios use per-month REPORTING capacity (Elexon
registered), which runs up to 16% above DUKES nameplate at some sites. And unlike every
other anchor, this one IS partially decomposed: reconstructing constraint
curtailment from Elexon bid-offer acceptances (FPN minus accepted level, the
standard method) shows the Scottish onshore fleet lost
15.5% of its would-be output to the grid saying no —
pre-curtailment its ratio rises to ×0.79 capacity-weighted,
converging on the GB 2016–18 cohort — while offshore curtailment is only
~4%, so the offshore gap (×0.76 pre-curtailment) is genuinely
dense-array wakes, availability and residual model bias. Commissioning
ramps, B1610 reporting gaps and the original sweep's clipped
31 Decembers are all corrected in these figures. Consequences,
stated plainly: the US-fitted measured basis must not be borrowed for
GB, and no GB per-site correction is justified by this data — treat GB
absolute wind CFs as indicative, fleet ratios as context
(gb_wind_validation.json, dispersion
stats included), and per-site truth as the job of
met-data calibration. What the data does
justify has shipped: a GB-fitted fleet anchor built on exactly this
constraint-volume reconstruction — basis_vintage=gb_offshore
(×0.78) and gb_onshore (×0.58),
the GB cells of the measured-anchored basis above,
fleet-scale by construction and deliberately not per-site corrections.
Not modelled: curtailment from grid limits (the Connection Study
covers that separately), panel/turbine degradation over time, and any specific
commercial turbine or inverter. Three physics terms are also deliberately omitted, each
matching CompoundVision's own ensemble fallback — icing, spectral losses (~1–3%), and
snow soiling (a flat 2% soiling is used instead). Wake, availability, thermal, storm
and electrical losses are now modelled (they were not in the retired
resource-curve model). Every result stamps its weather source — the CE ERA5 store or
the Open-Meteo archive fallback — and the cf_model v2.0.0 engine version.
Connection Study (grid screening)
A bidirectional grid-connection screen at a point: for a generation, load or storage asset of a given MW, it returns the required voltage class, the nearest adequate line and its N-1-secure usable headroom, any published DNO headroom, a GB curtailment-risk read, and an indicative upgrade tier + cost.
Multi-year curtailment band (weather-driven)
Supply the generator's wind and/or solar nameplate and the screen adds a physical curtailment distribution: the hourly generation profile is dispatched against the connection's own usable (or published) headroom as a firm export limit, across 20 real weather years (default) from the 86-year archive. The result is an exceedance band — the median year's spilled share and the 1-in-20 bad year (95th percentile; for curtailment the high tail is the downside; with fewer than 20 years the response labels it a near-worst sampled year instead) — plus the median year's monthly spill pattern and a peak-window CF: the fleet's mean deliverable output share (capped at the export limit) during the Nov–Feb 16:00–18:59 UTC winter-evening stress window, a v1 capacity-value proxy for the GB/EU system only (sites outside a GB/EU box return no proxy rather than a meaningless one; it is a documented convention, not a full ELCC study). The band is deliberately un-mitigated — battery and flexible-load mitigation live in the 24/7 generation study. This marries the network screen with real weather: the same connection can spill 3% in a median year and 9% in a windy one, and the band is what a revenue model should carry, not a single draw.
Required voltage (heuristic by MW):
≤20 MW → 33 kV, ≤60 → 66,
≤150 → 132, ≤500 → 275,
>500 → 400 kV.
Usable headroom. A connection must survive the loss of one
circuit, so firm headroom is well below a line's nameplate thermal rating. We
apply an N-1 screening derate of 0.55 to the
thermal rating and show both numbers. Where measured DNO/DSO
published headroom exists at a nearby substation it overrides
the derate estimate (with its RAG rating, licence and constraint note).
Line thermal ratings are measured where available, else estimated
(line_ratings.py).
Indicative upgrade cost tiers (screening bands, not quotes):
Tier 1 bay/transformer at an existing substation £50k–150k/MW;
Tier 2 substation or local circuit reinforcement £150k–500k/MW;
Tier 3 new circuit / deeper reinforcement £0.5M–1.5M/MW
(+ indicative line-build km).
Not modelled: the NESO TEC connection queue, power-flow / load-flow, detailed DNO works, bilateral TSO costs, or consents. This is a first-pass screen to rank and kill sites, not a connection offer.
Deliverability Pre-screen (area → ranked sites)
Given a drawn area (bbox or freeform polygon) and a load, the pre-screen lays
a ~5 km lattice of candidate cells over it, runs the same
real 8760-hour 24/7 dispatch (above) at each cell against the load shape, and
ranks every cell by achievable 24/7 CFE %. The full ranked list is
exportable as CSV; the map highlights cells that clear your target %.
Bounds. The live solve is capped (order of
2,000–4,000 cells and a ~25 s wall budget) so
a warm-cache screen stays interactive; a larger or colder area returns a
partial result flagged as such rather than timing out silently.
Each cell's weather year is fixed and stamped on the result and the CSV.
Interpretation: the pre-screen answers "where in this area could a 24/7 load be met best by on-site renewables?" — it is a resource + dispatch screen, not a grid-deliverability guarantee. Hand a promising cell into the Connection Study for the grid read.
Site Suitability Score
The Site Suitability layer renders 0–100 scores per technology (wind, solar, gas, nuclear, datacenter) on a hex grid. Scores combine resource quality (wind speed / GHI / population proximity), land-use exclusions, grid proximity, and policy zones. The full sub-component breakdown is shown when you click a hex.
Suitability scoring is currently CEAtlas v1.4 — see the in-app Layer info → Site Suitability for sub-component weights and band cutoffs.
Voltage Headroom
Heuristic mapping from required asset MW to required voltage class (CIGRE bands), then a check whether any substation within the search radius has the required class. Returns a VIABLE / UPGRADE / INFEASIBLE verdict with a 0–100 confidence score.
Voltage classes used:
MV 11–33 kV (< 10 MW),
HV 66–132 kV (10–100 MW),
EHV 220–275 kV (100–500 MW),
SEHV 345–400 kV (500–1000 MW),
UEHV 500 kV+ (> 1000 MW).
Local Electricity Price sample
The Local Electricity Price panel samples the LMP zones polygon at your click point and returns 30-day mean / peak / min / volatility plus an annualised cost estimate at the asset's MW and a user- selectable utilisation factor (default 90%). LMP zones are GB bidding zones (NESO) for the GB region and ISO LMP zones for the US.
Capacity factor & the weather store
The wind and solar capacity-factor (CF) explorers are built on CEAtlas's own hourly weather archive — ERA5 reanalysis, 86 weather years (1940–present), stored globally at the native 0.25° grid. Every value comes from real hourly weather, not a single vendor climatology, so you can see not just the long-term average but the spread, the downside, the trend, and the shape of a typical year.
How CF is derived
Every CEAtlas CF surface — the map COGs, the /api/v1/cf endpoint, the
24/7 studies, and the CompoundVision-linked forecasts — runs through one
canonical module, cf_model.py v2.0.0, a line-faithful physics
port of CompoundVision's production engine. Both fuels return net
expected-farm CFs (energy at the meter, after farm and system losses), not a
gross resource curve. Inputs are the store primitives (100 m wind, GHI,
hour-ending UTC) plus two climatology sidecars — a 1991–2020 monthly air-density field
(0.644–1.532 kg/m³) and a month×hour 2 m temperature field.
cut-in 3 / cut-out 25 m/s; Annex-G air-density rescale)
× Bastankhah wake (5-row 7D×5D farm) × 0.97 availability
× thermal derate × storm taper+hysteresis × electrical/blockage loss
Solar CF = SPA → Erbs → Perez 2002 POA → King IAM → King–Sandia cell temp
× (−0.35%/°C) × 2% soiling × ILR 1.30 AC chain, hard AC clip
Wind models a modern-onshore reference farm: a 300 W/m² reference turbine (rated 10.36 m/s at the ERA5-native 100 m hub, cut-in 3, cut-out 25 m/s), an IEC 61400-12 smootherstep curve applied in density-equivalent wind, then a Bastankhah–Gaussian wake for a 5-row 7D×5D farm, 0.97 base availability, temperature derates (×0.85 at ≥38 °C, ×0.60 at ≤−25 °C), a storm taper with shutdown hysteresis (taper 23–25 m/s, re-cut-in below 22), and load-dependent electrical (1.5–2.8%) and blockage (0.5–3.5%) losses.
Solar offers two selectable arrays — a single-axis N–S tracker
(backtracking, GCR 0.35) and a fixed array at optimal tilt
clip(|lat|·0.87 + 3.1, 8°, 45°) facing the equator. The chain runs the
NREL SPA sun position at the −30 min accumulation midpoint (hour-ending labels
verified empirically, corr 0.9978 vs 0.9513), Erbs GHI decomposition, Perez 2002
plane-of-array transposition, King IAM (0.955/0.832 diffuse factors), a King–Sandia
cell temperature (climatology temperature + store wind cooling) at −0.35%/°C, flat 2%
soiling, and an ILR 1.30 inverter chain (0.985 inverter × 1.5% cabling × 0.99
transformer) with a hard AC clip and a −2° night gate. The result is a net AC CF
against AC nameplate.
Three loss terms are deliberately omitted, each matching CompoundVision's own
ensemble fallback: icing (needs humidity/precipitation), spectral (~1–3%, conservative
for c-Si), and snow soiling (a flat 2% is used instead). The old
GHI / 1000 × 0.91 proxy has been retired. These are single-reference-plant
CFs — a specific turbine, hub height, tracking choice, or loss stack shifts the
absolute number, while the relative geography and inter-annual behaviour stay robust.
What the dropdown shows
- Statistics across 86 years: Mean, Median (P50), P90 — the bankable year-in-10 downside a lender underwrites on — Variability (inter-annual σ), and Trend (CF change per decade).
- By year — the annual-mean CF for any single year 1940–2025, with a ▶ animation that sweeps the whole record.
- By season / by month — DJF/MAM/JJA/SON and Jan–Dec climatologies (seasonal pooled by hours, leap-aware).
- Source — the CE weather store (ERA5, 86 yr), the ECMWF IFS HRES ~9 km dataset (2017–present, its own independently computed CF set at 0.1°), or the external DTU Global Wind Atlas / Global Solar Atlas long-term mean as a cross-check.
- Turbine class — IEC Class I / II / III reference
machines, plus Auto: per-cell selection of the
fleet-equivalent class from the record-mean 100 m wind
speed, with class boundaries calibrated to the USWTDB specific-power
distribution of the EIA-923-validated cohort rather than the nominal
IEC 61400-1 Vave values. The nominal bins put most of the
operating US fleet's capacity on the low-wind Class III curve while the
machines actually bought average a Class II specific power — installed
rotor loading tracks machine vintage, not site wind, so the
boundaries are an aggregate fit to the validated fleet, not a per-site
suitability rule — fitted on the US fleet's 5–8 m/s band and applied
globally, so outside that band the class choice is extrapolated. Auto
therefore tracks the fleet-equivalent machine
(modern low-SP builds trend toward the Class III curve — pick
iec3 to model those explicitly). Shown as its own map and used to
resolve
wind_class=autoAPI queries.
Per-site read-out, 12×24, and downloads
Click anywhere for a full read-out (mean / P50 / P90 / variability / trend) plus:
- a 12×24 month-hour heatmap — mean CF for each calendar month × hour of day, the "8760 at a glance" that reveals the diurnal and seasonal signature (solar's midday-summer peak, wind's night/winter lift);
- CSV downloads of the site's full record — monthly (1940–2025) or hourly (paged into API-sized windows and merged), plus the 12×24 grid.
250 m downscaling (optional map overlay)
Map layers can overlay a 250 m ratio texture on the 0.25° field,
w(x) = fine-scale value / ERA5-cell mean. Wind uses the Global Wind Atlas
CF-domain ratio on the map; the API's hourly path instead applies the GWA
speed-domain ratio through the full power chain (the same mechanism as the
EIA-923 ws_mult calibration). Solar uses the Global Solar Atlas PVOUT
ratio. Two further correction layers are baked into every ERA5 surface and
applied identically in the hourly API: a monthly mesoscale ratio
(ECMWF IFS 9 km ÷ ERA5, smoothly month-interpolated with the monthly means
preserved exactly) that restores lake/coastal/convergence structure ERA5's 0.25°
grid cannot resolve, and a 12×24 solar diurnal-shape table
(IFS ÷ ERA5 month-hour climatology, normalised to mean 1 per month so it
reshapes the day without moving monthly energy).
- Temporal statistics are 86-year; sub-cell spatial detail is a modern-era pattern. The 250 m / ~1 km texture comes from the Global Wind Atlas / Global Solar Atlas climatologies (recent-decade reference periods) applied as a static spatial layer on the 86-year 0.25° signal. "86 years at 250 m" unqualified would overclaim — the honest statement is 86-year temporal statistics with modern-era 250 m spatial structure.
- The mesoscale and diurnal-shape corrections are calibrated on 2017–2025 (the ECMWF IFS overlap) and applied as climatological patterns across the whole record — standard downscaling practice that keeps the record homogeneous, but the calibration era is the modern one.
- Pre-1950 ERA5 (1940–1949) is the Copernicus "preliminary" back-extension. Fewer observations existed to assimilate, so 1940–1949 carries somewhat wider uncertainty than 1950-present. It is genuine ERA5 and included in every statistic; treat single-year values from that era with proportionate care.
IFS-informed cell anchoring (textured wind maps & print posters,
since Aug 2026). The anomaly-ratio construction anchors
between-cell magnitude entirely to the 0.25° store means; over extreme relief
(Himalaya, Pamir, Andes) that lattice was visible beneath the 250 m
texture. Displayed wind-mean surfaces therefore blend the cell-scale anchor
with the ECMWF IFS HRES ~9 km means: per cell, IFS
supplies the local between-cell pattern while the smoothed store field sets
the regional level. The blend weight is the per-cell complexity of the
250 m texture — plains cells are unchanged, ~13% of land cells blend,
and the global land mean shifts by +0.16% (relative). Raw store means remain
served unmodified for untextured and temporal views; the
/api/v1/cf time-series paths are unaffected.
Known limitations
- ERA5 is 0.25° (~28 km) — its hourly dynamics are cell-scale even when a 250 m overlay refines the magnitude, so mountain valleys, ridgelines, and near-shore gradients share the cell's temporal shape. For micro-siting, treat CF as a regional estimate.
- Hub height is fixed at 100 m and the reference turbine/farm is generic; solar is one of two reference arrays (tracker or fixed) — a real project's kit differs.
- Solar physics is validated to ~1 pp against PVGIS. Wind
shape and rankings are validated against metered US generation
(583 plants: metered-month r 0.88, worst-year 2023 agreement) and
cross-validated against MERRA-2 (annual r 0.89, drought agreement 28/36) —
but absolute wind levels run hot, mostly for the old fleet:
metered CF is 92% of the turbine-true model for 2016–2018-built plants but 74%
for pre-2012 ones (degradation/availability), with curtailment visible on top
(ERCOT 0.75 vs SPP/CAISO 0.91–0.94). Treat absolute wind CFs as indicative of a
modern new build before curtailment, collapse the site-level uncertainty with
your own met-data calibration, or switch the studies
and API to the shipped measured-anchored basis
(
wind_basis=measured_us, withbasis_vintagecohort cells) when older-fleet or regional realism matters. - The store covers land plus all sea within ~300 km of a coast (mask v2, no Antarctica); open ocean beyond that margin is not sampled.
Percentile context & the annual-summary API
/api/v1/cfannual (free; rate-limited without a token) returns each weather year's annual and
monthly mean CFs for a point across the archive — the distribution
behind the climatology maps. CEAtlas uses it to place live
CompoundVision production forecasts in context: the forecast-window mean CF
is compared with the same calendar month across the 86-year record and
labelled with an exceedance percentile — "P12 vs the 86-yr July
record" means a July this strong occurs in 12% of years. It is an
indicative anomaly read (a days-long forecast window against a monthly
climatology), designed for owner reports, not settlement. The forecast
panel also now surfaces CompoundVision's own P10–P90
uncertainty band alongside the ensemble mean.
Calibrate with your own met data
A met mast or an operating asset knows what a reanalysis chain cannot.
Upload hourly measurements (timestamp,value CSV, UTC) — wind
speeds at a stated mast height, or production MW with a nameplate — and
CEAtlas aligns them with the modeled series over the overlap (minimum
1,000 aligned hours, at least 60% of in-archive rows matching, up to 15
archive years per upload) and returns the evidence: hourly correlation,
mean bias, and a bounded calibration factor. Mast
speeds yield a speed-domain multiplier against the full
modeled 100 m site chain (cell weather × mesoscale × 250 m
terrain, bilinear-blended over the surrounding cells exactly as the
studies compute it; measurements shear-lifted with a power-law
α=0.14 from a stated 10–300 m mast height, with
extrapolation-heavy lifts flagged), so it reshapes hourly CFs through
the power curve; production data yields a CF-domain scale. As a
diagnostic, the hourly correlation is also scanned at lags of ±3
hours — a peak away from zero is the signature of a timezone or
hour-labelling offset in the upload (the mean-ratio factor itself is
shift-robust), and the response says so. Factors are clamped to [0.5, 1.5] — a truthful site outside that
band usually signals a units or capacity mismatch, and the response says
so. Nothing is stored server-side: the factor rides your
subsequent study runs explicitly and appears in their inputs, so every
calibrated result is labelled. This is a bias correction, not a
resource assessment — low hourly correlation is reported and the mean
ratio remains valid, but hourly-shape conclusions deserve care.
Weather-year ensembles for expansion planning
Capacity-expansion models are notoriously biased by their single weather
draw: a plan optimised against one calm winter under-builds firm capacity,
one windy year over-builds wind. CEAtlas selects a small, defensible
ensemble of real weather years from the 86-year record
for a planning region: representative years anchored at the
P25/P50/P75 of a combined resource index (equal-weight z-scores of annual
wind + solar CF, averaged over sample sites across the region), carrying
the bulk of the probability mass, plus extreme years — the
record-worst annual wind, annual solar, and winter (Jan/Feb/Dec) wind
years — at fixed stress weights. The selection is deterministic and
quantile-anchored (no clustering hyper-parameters), and exports directly
in the CENovaSage bundle's weather_weights schema, so
expansion runs execute as weather ensembles and their build-outs return
to the map as Run Results layers. Selection tooling:
etl/select_weather_years.py, driven by the
/api/v1/cfannual distribution endpoint.
Fleet Census
/census is a demographic portrait of the power fleet — an age pyramid by fuel, additions and retirements through time, and the developer-reported build pipeline. It is deliberately per-region, best-registry rather than one blended global source: each region uses the most authoritative registry available and inherits that registry's own frame. Plant popups on the map draw on the same per-plant lifecycle records, with metered-generation sparklines where a public meter exists (EIA-923 in the US, Elexon B1610 in GB, ENTSO-E A73 in the EU).
| Region | Registry | Fleet | Metered generation |
|---|---|---|---|
| US | EIA-860 (2023 filing year) | 1,280.5 GW operating · 203.4 GW retired since 2002 · 208 GW pipeline | EIA-923 per-plant annual net generation |
| GB | DUKES 5.11 (May 2026) + REPD Q1 2026 | 113.9 GW operating · 38.1 GW closures · 162.6 GW pipeline | 94 sites — Elexon B1610 settlement meters |
| Europe | GEM Global Integrated Power Tracker, March 2026 (CC BY 4.0), via the CEAtlas plant corpus | 865.0 GW operating · 67.8 GW retired · 47.5 GW pipeline · 29,507 units, 33 countries (census built 2026-09-11) | 348 plants — ENTSO-E A73 metered generation 2019–2025, 23 control areas |
| World (rest of) | GEM Global Integrated Power Tracker, August 2026 | 7,940.4 GW operating · 612.5 GW retired · 5,119.5 GW pipeline · 30,093 plants ≥ 20 MW (census built 2026-09-03) | — |
- Each region is frozen at its registry's date: the US at the 2023 EIA-860 filing year; Europe at the GEM March 2026 release. Europe was rebased off JRC-PPDB-OPEN 2019 on 29 Aug 2026 — a frame change as well as a refresh, from JRC's 645 GW large-unit (ENTSO-E-visible) fleet to GEM's utility-scale tracker with the wind and solar build-out included (97% of capacity dated, against JRC's 70%). Two gaps are disclosed on the tab: plants the corpus GEM merge failed to key (most of Poland's lignite fleet) are missing, and GEM's oil/gas class is normalised upstream to gas, so European oil peakers sit inside the gas family. Sub-threshold capacity — rooftop solar above all — is out of frame: this is the utility-scale fleet, not every European generator.
- Pipelines are developer-reported (EIA expected COD, REPD consented / under-construction, GEM pre-construction starts) — consented does not mean built, and expect slips.
- Additions are survivor-biased where a registry lists only what still operates (DUKES; EIA pre-2002). The GEM world ledger carries real retirement statuses, with dated cohorts covering 97% of operating capacity.
- "World" means rest-of-world: the US and GB have their own tabs, and European plants appear there only where the Europe index lacks them.
- Europe's retirement ledger is GEM's dated European closures and is partial: it reaches the census through the World extract, which had already deduplicated closures that the old EU index also listed.
World data: Global Energy Monitor, Global Integrated Power Tracker (August 2026 release), used under CC BY 4.0. Full source licences and attributions: /licenses. The aggregates and per-plant lifecycle records behind the page are published as JSON — census: US · GB · EU · World; lifecycle: US · GB · EU · World.
CONUS price forecast accuracy
The CONUS daily cycle publishes a day-ahead price forecast for every US market region each morning (the CENovaSage Run overlay shows the nodal product). Its accuracy record is public and lives in two places drawn from the same numbers: /scorecard, the persisted record re-scored every morning against what actually happened, and the free Price Forecast Accuracy (CONUS) map layer (Markets & Economics) — one marker per scored hub / zone.
Definitions
- Skill vs yesterday's price — the headline: 1 − MAEmodel / MAEpersistence, where persistence is the market's own price at the same hour of the previous market day — the cheapest honest forecast a trader already has — computed on the same hours for both sides. 0 = no better than yesterday's price; −1 = twice its error; +0.5 = half. A negative headline is the record working, not a display error — it is shown as measured, with the number of regions, days and the date span beside it.
- Lead — day-ahead is the run made the day before the scored day (the forecast a subscriber could have bid with — the page's headline view); same-day is the run made on the scored day (a nowcast; the daily calibration gate evaluates it).
- Scored against — the market's day-ahead price
(
da) or its real-time price (rt: hourly, or hourly means of 5-minute intervals). Only day-ahead rows are ever gated. - Before gate — every day-ahead row carries the run's UTC solve stamp and a flag against that market's day-ahead gate closure, measured per row and never asserted: PJM 10:30 EPT, ISO-NE 10:00 EPT, MISO 10:30 EST, SPP 09:30 CPT, ERCOT 10:00 CPT, CAISO 10:00 PPT. NYISO closes at 05:00 EPT, before the 09:30 UTC cycle starts, so its day-ahead rows are never bid-time. Days published after the gate are still scored, and flagged.
- Hours / day — how many hours of the 24-hour local market day were scored. The published hours run through 07:00 UTC of D+2 so every local day completes; days scored before that extension show the truncated evening (19 of 24 hours in the Eastern zones, 18 Central, 16 Pacific) — reported, never silently dropped. Since 26 September 2026 each run also looks a further day ahead (a 96-hour model horizon) so the end of the day-ahead day is not priced at the edge of the model; those look-ahead hours are neither published nor scored.
- Artefact bus-hours — bus-hours inside a hub / zone aggregate priced at a solver construct: the flat $250 export-relief price for 6 or more hours of a day, a price below −$400 (a dual, not a price — every ISO's bid floor is higher), a price above the ISO cap, or a must-run unit backed below its floor where the engine flags it. They are counted on the face and never removed from the score; an ex-artefact MAE is shown beside it as a diagnostic only, because a construct we created is ours to fix, not to exclude.
- Era / system — windows never pool engine eras (the July
workstation era vs the daily runner from 2026-08-25), and every row carries a
systemlabel (era | topology revision | inputs revision) so model generations are never averaged together. - Hub / zone level, on purpose — markers and rows are the market's published pricing points, not individual buses: the review of 2026-09-10 measured that the model's intra-zone price structure carries no skill (correlation ≈ 0 with the observed intra-zone deviations), so per-bus markers would draw a precision the numbers do not have.
What it is scored against
- ISO-NE — 8 load zones; DA and RT hourly LMP (public histRpts files).
- NYISO — zones A–K; DA zonal LBMP and hourly-integrated RT LBMP (public MIS archive).
- PJM — 20 transmission zones + the PJM-RTO aggregate; DA and RT hourly LMP (Data Miner 2 public feed).
- MISO — 8 trading hubs; DA ex-post and RT LMP (public market reports).
- SPP — SPPNORTH_HUB / SPPSOUTH_HUB; DA LMP (Marketplace portal daily file, else the CompoundVision price-store export) and RTBM 5-minute LMP averaged to hours.
- WECC — CAISO trading hubs TH_NP15 / TH_ZP26 / TH_SP15 only (the rest of WECC has no organised market); DA from OASIS PRC_LMP, RT from hourly means of OASIS PRC_INTVL_LMP 5-minute prices.
- ERCOT — DAM settlement point prices at HB_NORTH / HB_SOUTH / HB_WEST / HB_HOUSTON via gridstatus.io (keyed) or the CompoundVision price-store export; no real-time source without a key.
- SERC and FRCC — no organised market, so price columns are null. SERC border sub-areas are compared with the neighbouring PJM DOM / PJM AEP / MISO MS.HUB day-ahead prices as a reference only; FRCC is validated on generation fuel shares against EIA-930 net generation by fuel.
Freshness: the record is re-scored every morning by the CONUS rolling cycle (09:30 UTC); eastern real-time prices publish late and enter the record on a later re-score. The raw feeds are public ISO data used for scoring only and are not redistributed — see /licenses.
Plant deduplication
Wind and solar plants are merged across sources (EIA-860M, MaStR, REPD, PowerPlantMatching + ENTSO-E, WRI Global Power Plant Database, GEM Global Integrated Power Tracker) using a source-priority + spatial-bucket + name-similarity pipeline. Priority order:
- EIA-860M (US, monthly authoritative)
- MaStR (Germany, Marktstammdatenregister)
- UK REPD (GB, authoritative)
- National registers: DK-ENS (Denmark), Vindbrukskollen (Sweden), NVE (Norway, wind + hydro), SEAI (Ireland), RIVM (Netherlands), ANEEL SIGA (Brazil), Geoscience Australia + AEMO (Australia)
- PowerPlantMatching + ENTSO-E (EU)
- PowerPlantMatching (global, less authoritative)
- ODRÉ registre national (France — ranked below PPM/GEM because its coordinates are commune centroids; adds completeness + metered energy)
- WRI Global Power Plant Database (v1.3, fallback)
- GEM Global Integrated Power Tracker (ranked last per site, but the primary registry for regions the others don't cover)
At every tier a GEM-enriched record (source tag +gem)
outranks its unenriched sibling, so GEM attributes survive even where
another registry wins the dedup.
Within a 1 km × 1 km spatial bucket (and 3×3 neighbour scan), a candidate is treated as a duplicate of an already-claimed record if name similarity ≥ 0.85 (difflib SequenceMatcher) or token overlap ≥ 0.75, and capacity ratio is between 0.87× and 1.15×.
Data refresh cadence
| Data source | Cadence |
|---|---|
| BMU constraint actions (Elexon BOAL) | Point-in-time snapshot (30-day window ending at the layer build date), refreshed manually |
| GB ETYS day-ahead flows (NESO) | Daily |
| LMP zone prices | Daily (15-min settlement data, 05:00 UTC refresh) |
| CONUS price scorecard (/scorecard + Price Forecast Accuracy map layer) | Daily — re-scored the next morning by the CONUS rolling cycle (conus_rolling.yml, 09:30 UTC) |
| EIA-860M plants (US) | Point-in-time snapshot, refreshed manually (EIA publishes monthly) |
| REPD plants (UK) | Quarterly per DESNZ release, refreshed manually (current extract: Q1 2026) |
| Plants tile (powerplantmatching / ENTSO-E, national registers, GEM merge, dedup) | Rebuilt nightly by the data refresh (refresh.yml, 05:00 UTC: refresh_data.sh → tippecanoe → Fly volume sync); the registry snapshots it reads refresh on their own cadence — MaStR monthly (mastr_refresh.yml, 2nd of the month), the others manually / per release |
| GEM Global Integrated Power Tracker | Per GEM release, refreshed manually (current: August 2026) |
| Transmission lines (ENTSO-E + OSM) | Quarterly |
| GridFinder ROW grid | Static research dataset (2020; no newer release) |
| OGF US Planned Transmission | Annual |
| Fleet census snapshots (/census + /data) | Per source release (EIA-860 annual · DUKES annual · REPD quarterly · GEM GIPT per release, Europe and World both) |