The intervention: MSE-trained emulation. What moved: how often extremes occur.
tailspec · tail-risk evaluation of six AI weather emulators
Weather-model benchmarks look at ordinary days. A new benchmark evaluating tail behaviour. Preprint and GitHub repository to follow. evaluates the The tails of the distribution: the storm peak, the record heat, the once-a-decade cold..
The six emulators (GenCast, AIFS, GraphCast, Pangu-Weather, FourCastNet 3 and NeuralGCM) run faster than the physics-driven models and match them on the usual scores. Every one falls short on the days the wind is extreme, and the rarer the event, the larger the shortfall.
↓
a single storm
Storm Eunice, 18 February 2022
ERA5
Britain's hardest hit locations, sampled at 00 & 12 UTC.what is plotted
ERA5's 10 m wind, averaged over the 5 British land cells the storm hit hardest.
Each model's 24-hour-ahead forecast for the same cells and times is overlaid, from a
single daily run rather than an ensemble mean.
Storm Eunice crossed southern England on 18 February 2022, closing rail lines and cutting power to about 1.4 million homes. The line plotted to the rightabove is what the wind actually did in the locations the storm hit hardest, as the ERA5 reanalysis reconstructs it.
Each AI model is started fresh from the morning atmospheric condition and asked what the next day holds. The coloured lines join those 24-hour-ahead forecasts across the fortnight, and all six track the reanalysis through the whole run of storms.
At Eunice's own peak, every model underpredicts the actual conditions, by an average of 10.9% and NeuralGCM alone falls 34.4% short. One storm cannot prove a pattern, but consistent one-sided behaviour will show up if we consider every wind extreme in the record.
all wind extremes
How the Shortfall Grows
Broaden the view from one storm to every 24-hour forecast made between 2020 and 2023, worldwide, filed by how extreme the observed wind turned out to be at that place and hour. Even the mildest group on this axis is the windiest 10 per cent of the record, so nothing here is an ordinary day.
Mean 24-hour 10 m wind-speed forecast error compared to ERA5. (Global, 2020 to 2023)what is plotted
Each forecast/observation pair is filed by the percentile the observed wind reached in ERA5's
own time series at that gridcell, then reduced within the group. About 391,000
samples fall in each of the first four, 195,000 in P98-99 and
130,000 in each of the top two.
No model reaches the observation in any extreme group, and the gap widens as the events get rarer. The five on a shared axis fall 0.21 to 0.85 m s⁻¹ short, worse in the rarer bands than the milder ones, while NeuralGCM falls short by an order of magnitude more. What all six share is the trend: a systematic underprediction that grows with rarity.
the mechanism
Why the Extremes Are Not Extreme Enough
the technical version
A point loss is minimised by a functional of the conditional distribution, not a draw from it: squared error by the conditional mean E[y | x], absolute error by the conditional median, either way collapsing p(y | x) to one number per cell. By the law of total variance, a field of those numbers varies less than a field of real draws would, by a gap that scales with the conditional spread, which is what compresses conditional extremes. A proper score for the whole distribution lifts that constraint on individual samples, but not on anything read off their ensemble mean. At a decision threshold the loss is asymmetric anyway (2:1 in our wind-power backtest), so the cost-minimising commitment is a quantile of the predictive distribution, not its centre.
Most of these AI models are trained on a point loss. GraphCast and AIFS minimise squared error, whose optimum is the conditional mean, and Pangu-Weather minimises absolute error, whose optimum is the conditional median. The particular point loss matters less than the fact that it is a point loss at all: either way the target is a summary of every atmosphere still consistent with this morning's observations rather than one of those atmospheres, and a summary falls furthest from a real draw where those atmospheres disagree most, at the top of the distribution.
Training on a probabilistic score eases the constraint but does not remove it. FourCastNet 3 is trained on a CRPS objective and GenCast is a diffusion model that draws samples rather than averaging over them, and both still fall short of the observation in every group on the panel above. Part of this is likely where the number gets read, since headline products are usually taken off the ensemble mean and so put the averaging back in. And real decisions rarely impose symmetric penalties, so what minimises cost is a quantile of the forecast, not the middle of it.
Global RMSE weights every sample by how often it occurs, not by how much a miss there costs. The rarest per cent of the record barely moves it, however large the shortfall, because the ordinary hours that make up the rest of the record dominate the average by sheer count. tailspec scores the tail on its own terms.
the whole roster
Every Model on One Tail Axis
the technical version
First step, the forecasts. Thresholds are each cell's ERA5 95th percentile, matched to the time of day. The count ratio is a model's exceedances divided by ERA5's exceedances of the same thresholds, with a moving-block bootstrap that draws the same initialisations for both, so weather common to both cancels. The ERA5 curve is scored out of sample: a value from year Y passes through the fit that left year Y out. Scored this way, ERA5's own curve ends above the diagonal, so a model is compared with that curve rather than with the theory. Cells whose 95th-percentile wind is below 2 m/s are excluded, 11610 of 1038240 globally.
Second and third steps, the climate models. Their draws pool every stored hour, so the threshold and the fitted tail are taken over all hours too (daily means for HadGEM3). Each climate-model curve on the second step is built from 293 draws, spread evenly over the record and balanced across the hours it stores. Neither step has a count ratio or an ERA5 curve.
Every emulator is scored on the same 293 initialisations and the same 9 lead times, and two questions are asked of its tail. How often does the forecast exceed a cell's own 95th-percentile wind, counted relative to ERA5 on the same dates? And when it does, how far does it go? The second is a quantile-quantile plot. Every value above the threshold is passed through that cell's fitted ERA5 tail, which puts a Greenland fjord and the North Atlantic on one scale. The gold curve is ERA5 itself. For the paired question, whether a forecast of an extreme verifies, see tail calibration below.
extremes produced, relative to ERA5 on the same dates
Past the end of the axis, values pass the reach of ERA5's fitted tail for ERA5 0.42 %, AIFS 0.19 %, GraphCast 0.30 %, Pangu-Weather 0.21 %, NeuralGCM (derived wind) 0.64 %, GenCast 0.33 %, FourCastNet 3 0.40 %.
what is plotted
Every value above a cell's own 95th-percentile wind is passed through that cell's fitted ERA5 tail,
which puts calm and windy places on one scale. The pooled result is plotted against its return
period. On this axis a perfectly reproduced tail lies along the gold ERA5 curve. GenCast was sampled at 12 noise levels rather than the published default of 20. NeuralGCM does not output a 10 m wind. Its wind is derived from the 1000 hPa level, which is
why its count sits far below the rest at every lead.
Forecasts on the Same Dates
How often the extremes occur divides the roster by training objective. At 10 days, the three models trained on a point loss (AIFS, GraphCast and Pangu-Weather) produce clearly fewer extremes than ERA5 on the same dates, and their deficit is larger than at 12 hours. GenCast, a diffusion model, and FourCastNet 3, trained on CRPS, produce slightly more.
That surplus is consistent with the shortfall measured earlier on this page, which scores each forecast at the places and hours where the observed wind was extreme, whereas the count here ignores where and when. GenCast and FourCastNet 3 appear to produce enough extremes overall without always placing them at the right place and time.
Among the values that do exceed the threshold, the roster has no single shape. At 10 days, globally, and at the far end of the axis, AIFS, Pangu-Weather and GraphCast fall short of ERA5's curve, while the bands of GenCast and FourCastNet 3 overlap ERA5's.
Region by region the samples are smaller and the bands wider. The same separation (the point-loss models clearly short of ERA5, the other two not clearly so) holds in Central Mediterranean, Central & southern US and Pacific Northwest, and breaks down in Arabian Peninsula and NW Europe.
Climate Models Without Lead Time
The same axis, now with no date matching, so products that never claimed to forecast a particular day can join: ICON (EERIE), IFS-FESOM2 (EERIE) and CorrDiff (downscaler). The lead selector is gone because these products have no lead time.
Here the axis is ERA5's fitted tail for 11 years spread across 1980-2013, and every product's samples are drawn from 1980-2014, so the models and the reference cover the same era. There is still no ERA5 curve on this step: ERA5 would be scored against a fit to its own years. The curves can be compared with each other and with the diagonal. A model can match the climatology and still be useless on any given day, so nothing on this step or the next is a skill claim.
CorrDiff (downscaler) is different in kind: it sharpens CanESM5 climate-model output rather than simulating the atmosphere, and it was run on 1980, 1990, 2000 and 2010, years held out of its training. Its curve is placed on the same 1980-2013 ERA5 axis as the others, whose 11 years include its four. What it passed against ERA5 from its own years: On four held-out years, CorrDiff driven by CanESM5 inputs matches ERA5 10 m wind from the same years in the geography of its per-cell P95 (high-pass corr >= 0.40), in the censored fraction against ERA5's tail fit (within 1 pp) and in its P99 level (within 10 %). Only just: the censored fraction misses by more than 1 pp in 2 of 5 regions (Pacific Northwest, Arabian Peninsula) and the P99 level by more than 10 % in 2 of 5 (NW Europe, central Mediterranean), where a second ERA5 sample misses in none; its P99 runs high in all six domains, and its P95 geography correlates at 0.84 against 0.99 for two ERA5 samples.
One Model at Two Resolutions
The same axis, now for one climate model run at two resolutions: HadGEM3 at ~60 km and at ~25 km. HadGEM3 stores its 10 m wind over the ocean only, so both curves pool the same 693,221 ocean cells, and each is built from 293 daily means drawn from 1975-2014.
Globally, at the far end of the axis, the two runs' 90 % bootstrap ranges overlap (they are computed but not drawn on this step), so on this measure the finer resolution makes no detectable difference to the shape of the ocean wind's tail. Region by region the ranges separate only in NW Europe, where the ~25 km curve is higher but its range at the far end is the widest on this step, and they overlap in Central Mediterranean, Central & southern US and Pacific Northwest.
A single pair shows what resolution does in HadGEM3 only. As in the step before, there is no ERA5 curve. The Arabian Peninsula is left out of this step because its fitted tails are fragile enough that a gap between two curves there would reflect the fit more than the resolution.
what tailspec measures
Three Ways a Forecast Fails in the Tails
Tail Heaviness
Does the model's own distribution reach the most extreme events, or does it thin out too fast?
the technical version
Generalised Pareto shape parameter ξ, fitted per grid cell by peaks-over-threshold. Too small a ξ means tails that are too thin.
Calibration
When the model says a 1-in-100 event, is it actually 1-in-100?
the technical version
Tail calibration after Allen et al. (2024): an occurrence ratio (exceedances observed against exceedances forecast) and a severity ratio (how far past the threshold, once exceeded).
Value
What a decision made using this forecast is worth, against the same decision made without it.
the technical version
Relative economic value across cost-loss ratios, plus backtested trading desks with real settlement rules (see In practice).
tail shape
Where the Tails Are Too Thin
Red = tails are too thin (the model makes the most extreme winds less likely).
Blue = the model's tail is too heavy.
loading gridded data…
Tail-shape parameter ξ of a Generalised Pareto fit to the peaks over each cell's own
0.95 threshold quantile, one fit per grid cell. Coarsened 4x from the native
721x1440 grid.
the technical version
Peaks-over-threshold fit, per grid cell, to the exceedances above that cell's local P95 (95th-percentile) threshold. The more positive ξ is, the heavier the tail, so a model whose ξ falls short of ERA5's is under-weighting its own worst cases.
Every one of the six forecast models has a thinner wind tail than ERA5, averaged across the globe and at the median grid cell alike. (The CorrDiff downscaler on the map is a different kind of product, run on other years, and is not ranked here.) FourCastNet 3's ξ comes closest, short of ERA5's by 0.047 on the area-weighted mean and by 0.019 at the median, and NeuralGCM is furthest out, short by 0.089 on the mean, roughly 1.9× FourCastNet 3's gap. GenCast, the one diffusion-trained forecast model among the six, is further from ERA5's tail shape than Pangu-Weather and FourCastNet 3, so whatever advantage GenCast has elsewhere in this benchmark does not set the tail shape.
tail calibration
How Often a 1-in-100 Event Happens
Tail heaviness asks whether a model's own distribution reaches its most extreme values. Calibration asks whether a model's stated chance of exceeding a threshold is actually right, and Allen et al. (2024) split the answer into an occurrence ratio (is the exceedance probability itself correct) and a severity ratio (given exceedance, is the size of the overshoot correct). The diagnostic needs full ensemble output, so the comparison here is GenCast and AIFS at full member resolution against the IFS ENS physics ensemble, all three restricted to the same 28 initialisation dates.
the technical version
Reliability curve at the region's 99th-percentile threshold: the empirical CDF of the excess PIT values among exceedances, over u ∈ [0, 1]. A tail-calibrated forecast's curve lies on the diagonal. The occurrence ratio is exceedances observed divided by exceedance probability summed over all forecasts. Above 1 means the forecast under-states how often the threshold is crossed, below 1 means it over-states it.
Occurrence ratio at this threshold (1.00 = exactly the observed exceedance rate): GenCast 1.05, AIFS 0.91, IFS ENS 0.62.
Tail calibration (Allen et al. 2024) at the 99th-percentile threshold, NW Europe, matched on
the same 28 initialisation dates.what is plotted
For every exceedance of the region's 99th-percentile threshold, the excess PIT value (where in its
own predicted distribution the observed excess falls) is pooled into an empirical CDF over u. A
calibrated forecast's curve lies on the diagonal. GenCast and AIFS run here at full member resolution
rather than from the single-trajectory archive the rest of the page uses. They are scored against the
IFS ENS physics ensemble, restricted to the same 28 dates so all three see identical thresholds.
Calibration holds up reasonably at a day ahead and erodes with lead time, in the same direction as every other result on this page. The clearest break is wind in the Pacific Northwest at the longest lead, where both AI ensembles under-state how often the threshold is crossed, and by more than the physics ensemble does.
record-breaking events
Training Objective and Record-Breaking Events
GenCast and AIFS share a closely related architecture, both built on the graph-based encode-process-decode design of GraphCast. What separates them is what they were trained to do. AIFS minimises mean squared error, and the forecast that minimises squared error is the conditional mean. GenCast is a conditional diffusion model, trained by denoising rather than by error minimisation and scored with CRPS, which rewards a single sharp atmosphere that could plausibly have happened over a blurred average of many that could not. A record-breaking event tests this difference directly, because the model has to either commit to an extreme beyond the standing record or fall back towards an average value.
the technical version
Paired bootstrap over 28 initialisations, resampled 10,000 times per model for the intervals shown below. The GenCast-minus-AIFS difference and its p-value come from a permutation test on the same paired initialisations.
Records of heat and cold separate the two clearly, with GenCast ahead on Skill at calling a record before it happens, relative to simply quoting the long-run rate at which records of that kind fall. Zero is that baseline. Below it, a forecast does worse than the naive rate.. Wind records are hard for both models: GenCast still leads, but neither beats the long-run rate at which wind records fall. Mean sea-level pressure, the smoothest of the three fields, shows no advantage either way.
record heat & cold
GenCast ahead by +0.085 [+0.058, +0.122]p < 0.0001
record wind
GenCast ahead by +0.666 [+0.378, +1.036]p < 0.0001
record pressure
no separation (p = 0.32)
Brier skill at calling a record before it happens, over 28 initialisations. Zero is
the skill of quoting the long-run rate at which records fall. Whiskers are bootstrap intervals
(10,000 resamples). The row beneath gives the paired GenCast-minus-AIFS
difference and its permutation p-value.
GenCast and AIFS differ in objective and share a lineage, so the gap between them points to the objective rather than the architecture. How large that gap is depends on the field: widest on wind, clear on heat and cold, and absent on pressure, the field with the least blur to remove.