tailspec · tail-risk evaluation of AI weather models

Three ways to move a tail

A forecast distribution has a location, a scale and a shape. Every intervention that claims to fix a tail moves one of them. Almost none move the shape, and nobody is required to say which one moved.

Three case files, one per moment: a bias correction, a guided sampler, and the AI weather emulators themselves. Each names the intervention, what it moved, and the number.

the case files

One Intervention per Moment

case file 01 · location

The Bias Correction

A benchmark of eleven forecast systems reports AI models leading at gale-force wind. Its published aggregates carry a mean-bias correction estimated on the verification stations themselves. Remove the flag and the headline reverses.

swing in >P95 wind skill under the correction
±14 pp
correlation of raw mean bias with the skill change
r = −0.988

case file 02 · scale

The Guided Sampler

Classifier guidance steers a generative weather emulator toward rare events and enriches them by a factor of 4 to 90. The enrichment is variance collapse rather than a longer tail, and the method’s estimator fails its own reliability gate.

pooled GPD shape, unguided to guided
ξ −0.265 → −2.005
Read the case file →

case file 03 · shape

The Emulators

Six AI weather emulators match the physics-driven models on the usual scores. Every one of them has a wind tail thinner than ERA5’s, and the shortfall grows with rarity. Priced on a real delivery book, none separates from the operational ensemble.

emulators with thinner-than-observed wind tails
6 of 6
separation from IFS ENS on the backtested book
p = 0.129–0.750
Read the case file →

A fourth case file is reserved: building an event set honestly, scored in EP-curve space.

get in touch

The code behind this site is in a private repository for now. If any of it is useful to you, get in touch.