tailspec · metrics
Tail Metrics for AI Weather Emulators
The measurement layer behind the three findings, computed on forecast-versus-reference pairs. The same metrics apply whichever intervention is under test.
Forecasts are verified date-matched against ERA5 reanalysis, with ECMWF’s operational ensemble as the physics baseline and free-running km-scale climate simulations as a climatological reference. Regridding is conservative, and coarse fields are never interpolated to fine. Extreme-value fits use L-moments by default and maximum likelihood where the uncertainty matters, and every fit is checked with probability plots.
01 · Tail heaviness
Does the model's distribution reach as far as the observed one?
- GEV and GPD shape parameter (ξ)
- Extreme-value fits per grid cell, from block maxima and peaks-over-threshold. Too small a ξ is a tail that thins out too fast.
- Return level curves
- The 1-in-N event each distribution implies, compared as N grows.
- Q-Q tail plots
- Forecast quantiles against reference quantiles in the top of the distribution.
- Percentile bias maps
- Where the P99 and P99.9 fields are biased, and by how much.
02 · Tail calibration
When the model says a 1-in-100 event, is it 1-in-100?
- Occurrence ratio
- Exceedances observed compared with exceedances forecast, per threshold (Allen et al. 2024).
- Severity ratio
- Given an exceedance, how far past the threshold the outcome went, compared with the forecast.
- Excess distribution and PIT diagnostics
- Calibration of the conditional distribution past the threshold, not only of its frequency.
03 · Record-breaking events
Conditioned on the forecast rather than the outcome, so the forecaster's dilemma does not apply.
- Record-conditioned error
- Forecast error on the events the model itself flags as record-breaking (Zhang et al. 2026).
- Record detection skill
- Precision, recall and Brier skill for record-breaking occurrence.
04 · Probabilistic skill in the tail
Proper scores, weighted towards the part of the distribution that matters.
- Threshold-weighted CRPS
- CRPS with its weight on the tail, the standard remedy for outcome-conditioned verification (Gneiting & Ranjan 2011).
- Brier skill at extreme percentiles
- Exceedance forecasts scored against climatology at P95 and beyond.
- Relative economic value
- Forecast value across the full range of cost-loss ratios.
- Spread-skill and rank histograms, conditioned on extremes
- Whether the ensemble's dispersion matches its error where the outcomes are large.
05 · Spatial structure
Does the model place extremes with the observed spatial coherence, at the resolution it claims?
- Tail dependence χ(u, h) and the F-madogram
- How the probability of joint extremeness decays with distance, compared with the reference.
- Kinetic-energy and temperature spectra
- Spectral power by wavenumber, and the effective resolution it implies.
06 · Physical consistency
Are mass, moisture and energy conserved over a rollout?
- Budget drift and residuals
- Global mass, moisture and energy accounting across lead time.
- Per-member budget scatter
- How widely ensemble members disagree on budgets a physics model closes identically.
07 · Decision value
What a decision made with the forecast is worth, scored against what happened.
- Backtested delivery desks
- A day-ahead wind-power book and a four-city temperature quote game, run with real settlement rules and paired bootstrap inference (the In practice page).
- Pricing translations
- Excess-of-loss layer mispricing and HDD/CDD contract value implied by each model's tail.
Two Choices the Findings Explain
Rarity is measured by return period and by record-breaking events. A percentile band conditioned on the observed outcome discredits skilful forecasts (Lerch et al. 2017), and the location finding shows that failure mode in a published benchmark. Decision value is scored in money on a backtested book, because a symmetric error metric cannot express an asymmetric cost.
return period
Return Levels Beyond the Record
A return level is the value a variable exceeds once every T years on average, and it is the axis a reinsurance book is written on. Estimating one is a peaks-over-threshold problem: keep everything above a high threshold, fit a generalised Pareto distribution to the excesses, and read the curve off the fitted tail.
The threshold is the local 90th percentile of seasonally standardised anomalies, pooled over a 2°×2° box around each location so the fit has enough exceedances to be stable. Pooling buys precision on the shape and scale parameters but adds no years to the record.
ERA5 here runs 2020 to 2023, four years at six-hourly resolution. The largest value any grid point holds is therefore a 1-in-4-year event, and everything drawn past that is the fitted tail extrapolating beyond its own data. The forecast archives are sparser again: 292 initialisations five days apart, so one grid point’s record reaches about 73 days. Markers are the largest values each record actually holds, the solid line is the span that record supports, and every dashed segment is extrapolation.
at 1-in-10 years ERA5 full record 14.1 m/s · same reanalysis on the archive's dates 13.0 m/s (-8% from sampling alone)
six emulators 10.2 to 14.0 m/s (straddling the control, no ordering resolved)
The emulator curves overlap one another’s intervals and their ordering changes from one location to the next, so the panel ranks none of them. Its one consistent result is the control: ERA5 restricted to the archive’s 292 dates is 1 to 16 per cent lower than ERA5’s own full record at the 1-in-10-year level, at every location. No model property can explain a gap between the reanalysis and itself, and it is larger than most of the model-to-model differences drawn beside it.
get in touch
The code behind this site is in a private repository for now. If any of it is useful to you, get in touch.