tailspec · tail-risk evaluation of AI weather models

Three ways to fix a tail

A forecast distribution has a location, a scale and a shape. A bias correction moves the location and a guided sampler moves the scale. The shape sets how rare the worst events are, and all six AI emulators tested draw it too thin.

Read the emulator evaluation: AIFS, GenCast, GraphCast, Pangu-Weather, FourCastNet 3 and NeuralGCM

1-in-101-in-1001-in-1000thresholdfurther into the tail →observed (ERA5)forecast−0.44 at every level

×0.17

forecast mass beyond the observed 1-in-100 level, schematic

−0.60

move of the 1-in-100 level, observed scale units

NeuralGCM has the thinnest wind tail of the six emulators, shape −0.196 where ERA5 fits −0.107. Read more →

Schematic generalised Pareto tails: chance of exceeding (log) against value above the threshold, in observed scale units. Observed shape: the area-weighted ERA5 fit.

Shape, Scale, Location

in practice

Using the Results

metrics

The Metric Catalogue

get in touch

The code behind this site is in a private repository for now. If any of it is useful to you, get in touch.