tailspec · metrics

Tail Metrics for AI Weather Emulators

The measurement layer behind the three findings, computed on forecast-versus-reference pairs. The same metrics apply whichever intervention is under test.

Forecasts are verified date-matched against ERA5 reanalysis, with ECMWF’s operational ensemble as the physics baseline and free-running km-scale climate simulations as a climatological reference. Regridding is conservative, and coarse fields are never interpolated to fine. Extreme-value fits use L-moments by default and maximum likelihood where the uncertainty matters, and every fit is checked with probability plots.

01 · Tail heaviness

Does the model's distribution reach as far as the observed one?

fitted GPD shape ξ per grid cell, 10 m windshare of cellsERA5GenCast
Distribution of the fitted tail shape over every grid cell. The emulator puts its mass at more negative ξ than ERA5 does, a thinner tail almost everywhere.
GEV and GPD shape parameter (ξ)
Extreme-value fits per grid cell, from block maxima and peaks-over-threshold. Too small a ξ is a tail that thins out too fast.
Return level curves
The 1-in-N event each distribution implies, compared as N grows.
Q-Q tail plots
Forecast quantiles against reference quantiles in the top of the distribution.
Percentile bias maps
Where the P99 and P99.9 fields are biased, and by how much.

02 · Tail calibration

When the model says a 1-in-100 event, is it 1-in-100?

threshold quantileoccurrence ratio, 24 hGenCastAIFSIFS ENS
Exceedances observed against exceedances forecast at each threshold, north-west Europe, 24 h lead. A ratio of one is calibrated.
Occurrence ratio
Exceedances observed compared with exceedances forecast, per threshold (Allen et al. 2024).
Severity ratio
Given an exceedance, how far past the threshold the outcome went, compared with the forecast.
Excess distribution and PIT diagnostics
Calibration of the conditional distribution past the threshold, not only of its frequency.

03 · Record-breaking events

Conditioned on the forecast rather than the outcome, so the forecaster's dilemma does not apply.

GenCastAIFSBrier skill for record-breaking 10 m wind, with 95% interval
Skill at forecasting a record-breaking wind day, against climatology, with the bootstrap interval. Both are negative: worse than climatology on records.
Record-conditioned error
Forecast error on the events the model itself flags as record-breaking (Zhang et al. 2026).
Record detection skill
Precision, recall and Brier skill for record-breaking occurrence.

04 · Probabilistic skill in the tail

Proper scores, weighted towards the part of the distribution that matters.

lead time, hspread / skill ratio, 10 m wind, NW EuropeGenCastAIFS
Ensemble spread against ensemble error by lead time. One is a spread that matches the error, and below one is over-confident.
Threshold-weighted CRPS
CRPS with its weight on the tail, the standard remedy for outcome-conditioned verification (Gneiting & Ranjan 2011).
Brier skill at extreme percentiles
Exceedance forecasts scored against climatology at P95 and beyond.
Relative economic value
Forecast value across the full range of cost-loss ratios.
Spread-skill and rank histograms, conditioned on extremes
Whether the ensemble's dispersion matches its error where the outcomes are large.

05 · Spatial structure

Does the model place extremes with the observed spatial coherence, at the resolution it claims?

separation, kmtail dependence χ, 10 m wind, NW EuropeERA5AIFS
How often two points are extreme together, by distance. The reference and the model decay at similar rates here, and the ordering is not significant under the date bootstrap.
Tail dependence χ(u, h) and the F-madogram
How the probability of joint extremeness decays with distance, compared with the reference.
Kinetic-energy and temperature spectra
Spectral power by wavenumber, and the effective resolution it implies.

06 · Physical consistency

Are mass, moisture and energy conserved over a rollout?

lead time, hper-member mass-budget scatter ratioGenCast / AIFS
Ratio of GenCast's to AIFS's member-to-member scatter in the global mass budget, by lead time. A physics model closes that budget identically. One is parity.
Budget drift and residuals
Global mass, moisture and energy accounting across lead time.
Per-member budget scatter
How widely ensemble members disagree on budgets a physics model closes identically.

07 · Decision value

What a decision made with the forecast is worth, scored against what happened.

cost-loss ratiorelative economic value, 10 m windGenCastAIFS
Forecast value across every cost-loss ratio, against climatology. The curves cross at 0.18 [0.08, 0.20].
Backtested delivery desks
A day-ahead wind-power book and a four-city temperature quote game, run with real settlement rules and paired bootstrap inference (the In practice page).
Pricing translations
Excess-of-loss layer mispricing and HDD/CDD contract value implied by each model's tail.

Two Choices the Findings Explain

Rarity is measured by return period and by record-breaking events. A percentile band conditioned on the observed outcome discredits skilful forecasts (Lerch et al. 2017), and the location finding shows that failure mode in a published benchmark. Decision value is scored in money on a backtested book, because a symmetric error metric cannot express an asymmetric cost.

return period

Return Levels Beyond the Record

A return level is the value a variable exceeds once every T years on average, and it is the axis a reinsurance book is written on. Estimating one is a peaks-over-threshold problem: keep everything above a high threshold, fit a generalised Pareto distribution to the excesses, and read the curve off the fitted tail.

The threshold is the local 90th percentile of seasonally standardised anomalies, pooled over a 2°×2° box around each location so the fit has enough exceedances to be stable. Pooling buys precision on the shape and scale parameters but adds no years to the record.

ERA5 here runs 2020 to 2023, four years at six-hourly resolution. The largest value any grid point holds is therefore a 1-in-4-year event, and everything drawn past that is the fitted tail extrapolating beyond its own data. The forecast archives are sparser again: 292 initialisations five days apart, so one grid point’s record reaches about 73 days. Markers are the largest values each record actually holds, the solid line is the span that record supports, and every dashed segment is extrapolation.

at 1-in-10 years ERA5 full record 14.1 m/s · same reanalysis on the archive's dates 13.0 m/s (-8% from sampling alone)

six emulators 10.2 to 14.0 m/s (straddling the control, no ordering resolved)

Peaks-over-threshold GPD fits on seasonally standardised anomalies, pooled over a 2°×2° box, at a 24-hour lead. Markers are the largest values each record actually holds, each curve runs solid only as far as its own record reaches (4 years for ERA5, a fraction of one for the forecast archives), and every dashed segment is the fitted tail extrapolating past its own data. Shaded intervals are 95% block bootstraps over 500 resamples of whole forecast cases, so neighbouring grid cells are not counted as independent. One grid point’s record is counted in dates, never in ensemble members: sixteen IFS ENS members at one initialisation sharpen the fit without adding years of weather. IFS ENS also runs on its own 220 dates, 2020 to 2022, so it is the one curve here not date-matched to the others. ERA5 scored against itself on the archive’s 292 dates falls short of its own full record by more than the emulators do.

The emulator curves overlap one another’s intervals and their ordering changes from one location to the next, so the panel ranks none of them. Its one consistent result is the control: ERA5 restricted to the archive’s 292 dates is 1 to 16 per cent lower than ERA5’s own full record at the 1-in-10-year level, at every location. No model property can explain a gap between the reanalysis and itself, and it is larger than most of the model-to-model differences drawn beside it.

get in touch

The code behind this site is in a private repository for now. If any of it is useful to you, get in touch.