In practice · the desk

All three

The measurement: one frozen decision rule on a real book. What decides it: the lower tail against the commitment quantile.

tailspec · the trading desk

A One-GPU Emulator at Cost Parity with ECMWF's Ensemble

A desk promises next-day wind power for a book of German and British farms and is paid out on the weather that happened. Promising power that never arrives is penalised twice as hard as withholding power that does. Six emulators and ECMWF's operational ensemble ran that book over 66 delivery days across two seasons, and neither pre-registered emulator separates from the ensemble on any of the four books. Even where one strays furthest, a desk with no forecast is 6 times further from the ensemble than that emulator is. Four of the six land in that band, and the other two are treated separately below.

the strongest case against parity within the pre-registered pair: AIFS against the IFS ensemble, German book, winter 21/22

-0.017[-0.032, -0.003]

p = 0.129

Costs are in capacity-factor units, not currency, and negative means the emulator was cheaper. This is the largest edge either pre-registered emulator had over the operational ensemble on any book, and even here the block permutation test cannot separate the two (p = 0.129). No other comparison in that pair comes as close. On a difference this small over 66 days the bootstrap interval clears zero and the permutation test does not, so the claim above rests on the test alone. What remains open is which days the money is lost on.

the rules

Forecast, Decide, Settle

Forecast, Decide, SettleFORECASTday aheadGenCast · AIFS · IFS ENSFourCastNet 3 · NeuralGCMGraphCast · Pangu-Weather+ climatology, persistenceDECIDEDesk A: how much powerto promiseDesk B: what temperatureto quoteSETTLEagainst what happenedasymmetric penalties
the technical version

Desk B's bias correction is expanding and past-only: each quote day's correction uses only forecast errors from verifications strictly before that day's issue time, so no quote sees its own future. Confidence intervals on both desks come from a circular block bootstrap over days (block length 5, 1,000 resamples), and p-values from a block sign-flip permutation test (10,000 resamples). Every decision rule, penalty and test was pre-registered before any season's forecasts were scored.

Desk A promises next-day wind power for a capacity-weighted book of German and British wind farms. It commits at the ensemble's newsvendor-optimal quantile under a 2:1 penalty for over-promising against under-promising, and delivery settles against the wind power the ERA5 reanalysis implies. Each ensemble has eight members, with the IFS ensemble subsampled to match. GraphCast and Pangu-Weather are deterministic architectures with no native ensemble, so the same quantile rule lands on their single forecast. Desk A reports what that property of the two models costs them.

Desk B quotes fair-value heating-degree-day (HDD) and cumulative-average-temperature (CAT) contracts for a four-city book each day, priced from realised-to-date ERA5 plus the bias-corrected ensemble forecast. The desk then checks whether either model can be picked off against the IFS ENS market it quotes into.

Desk A · day-ahead wind power

Every System on the Wind Book

Every day each system commits the amount of power its own ensemble says it can promise without being caught short more than a third of the time, and the book is paid out on the weather that arrived. The rungs below differ only in the forecast that went in.

the technical version

Commitment is the 2:1 newsvendor quantile of each system's 8-member ensemble, settled against ERA5-derived power at a 100 m hub height. IFS ENS is subsampled to the same member count so the ensembles are matched. Intervals are a circular block bootstrap over days (block 5, 1,000 resamples). The climatology desk is leave-one-year-out over a ±15-day calendar window, built from ERA5 2020-2023, and serves only to set the scale of the axis.

33 delivery days. Best forecast system to worst: 0.0817 CF, or €8.3m a day on a book this size. Operational ensemble down to climatology: 0.0664 CF, €6.8m a day.

Euros are the same capacity-factor numbers times 42.5 GW of installed capacity, 24 hours and a flat €100/MWh, fixed because these months are the gas crisis.

Four of the six emulators land on top of the incumbent. Neither pre-registered model is measurably cheaper or dearer than the IFS ensemble on either book in either season, and GraphCast and Pangu-Weather fall in the same band. The spread across those four is a small fraction of the distance down to climatology, so on this book a desk's profit and loss cannot tell a learned model from a physics ensemble.

The matched rung cuts the IFS ensemble to the same 8 members the emulators have. More members can only sharpen its commitment quantile, so the comparison is rerun against the archive's full 50, and nothing in that group separates from that rung either.

Two of those four, GraphCast and Pangu-Weather, have no ensemble to commit from. The newsvendor quantile that shades a commitment across eight members down for the 2:1 asymmetry has nothing to select from on one member, so for them it lands on the point forecast. The extension that added them expected the price of that missing distribution to be visible, yet on the only comparison this page tests neither deterministic desk is measurably dearer than the operational ensemble. A point forecast cannot widen on the days when the ensemble would have, and 33 delivery days per book are too few for that to surface.

Which pre-registered emulator comes out ahead depends on the country, and does so consistently. AIFS ran the cheaper German book in both seasons and GenCast the cheaper British one in both, though neither ordering is large enough to matter against the interval around it. Across all four systems at parity only the German book keeps the same cheapest system, and the British one changes hands between the seasons. A league table of forecast skill will not tell you which model to put on a given book.

Of the ten systems on this book, the four that reach the incumbent are level with it on the average day. What remains is whether the average day is where the cost falls.

Desk A · where the cost falls

Cost by Realised Wind

A season mean hides which days paid for it. Sorted by how much wind actually turned up, the same delivery days show the cost of the book climbing steadily for most systems, from the calm end of the season to the windy end.

  • GenCast +0.0320 [-0.0073, 0.0708]
  • AIFS +0.0240 [0.0013, 0.0483]
  • FourCastNet 3 +0.1744 [0.0238, 0.2783]
  • NeuralGCM +0.1804 [0.0713, 0.3070]
  • GraphCast +0.0608 [0.0213, 0.1108]
  • Pangu-Weather +0.0409 [0.0131, 0.0821]
  • IFS ENS (8) +0.0485 [0.0225, 0.0745]
  • IFS ENS (50) +0.0365 [0.0208, 0.0538]

Extra cost between the calmest and the windiest day of the season, with a 95% block-bootstrap interval over 33 delivery days. An interval clear of zero means the gradient is not an artefact of which days landed in the sample.

Across the 32 combinations of forecast system and book, 29 gradients point upwards and 20 are separated from flat by the bootstrap. The rise appears in every system, the operational ensemble included, and in both countries, yet no system separates from the others cleanly enough to be ranked by it.

the technical version

A least-squares slope of daily cost on the ranked realised capacity-weighted 100 m book wind of the delivery day, resampled with the same circular block bootstrap as everything else on this page. One slope per book, fitted over that book's 33 delivery days rather than a windiest-decile split, because a decile of 33 days is three days against a block length of five. Severity is realised wind and not realised power: the turbine curve cuts out at 25 m s⁻¹, so a power-ranked axis would file the most violent days under calm.

That gradient is measured in money. A windy day puts more power on the book, so an absolute error has more room to grow whatever the forecast does. Divide each day's cost by the power actually at stake and the gradient inverts: 32 of 32 relative slopes point down, 25 of them clear of zero, so relative to the size of the job these systems are more accurate on windy days. The cost gathers at the windy end because there is more to get wrong there, which matters to a desk but does not show skill failing in the tail.

The pre-registered backtest deliberately excludes every event fortnight and every Monday initialisation, so no named storm of these two seasons is in the sample. The gradient above is therefore a gradient within ordinary weather, and the cost concentrates towards the windy end even with the genuinely extreme days taken out of the book.

Neither parity on the average day nor uneven risk across days says whether these models can represent a genuine extreme. That needs the days this backtest threw out and a measurement of the tail itself, first on a named storm and then on the shape of the whole distribution.

decision value · record events

The Ranking Depends on the Price of a Miss

A probability forecast becomes a decision only once someone names a cost: act when the forecast probability exceeds the ratio of the cost of acting to the loss of doing nothing. That ratio is different for every user, and nothing obliges the same model to win at every value of it. The explorer below prices both ensembles’ record-event forecasts across the whole spectrum.

the technical version

Relative economic value (Richardson 2000) on the same record-event forecasts and the same 28 paired initialisations as the record-breaking analysis: the saving a user with a given cost-loss ratio makes over always acting on climatology, as a fraction of what a perfect forecast would save, maximised over the probability-threshold grid and clipped at zero. Bands are 95 % paired-bootstrap intervals. The crossover and its interval come from the sign of the bootstrapped GenCast-minus-AIFS difference.

Event
record heat & cold · GenCast vs AIFS · paired bootstrap over 28 initialisations
Your stakes

Paying 1 to guard 33: the bootstrap cannot order the two here (GenCast 74%, AIFS 73% of a perfect forecast's saving)

The ordering changes with the cost ratio. On records of heat and cold, AIFS is worth more for miss-averse users and GenCast takes over above a ratio of 0.025 (interval 0.018 to 0.078). On record wind the flip is at 0.18 (0.082 to 0.20), just before the value of either ensemble runs out: past a ratio of about 0.19, neither beats climatology on wind records, the same wind difficulty the record-breaking analysis reports. On record pressure the interval spans most of the axis, so the ordering there is not resolved at all.

Which model is better at record extremes depends on the price, meaning what a miss costs relative to a false alarm. The desk above fixes one such price and pays it out in money, and this section sweeps every price.

Desk B · temperature quote game

Quoting into the Ensemble's Market

If the wind book is right that these emulators are level with the operational ensemble, a second desk quoting into a market priced off that ensemble should find nothing to arbitrage. If either model had a genuine edge over the IFS ensemble, it would show up here as profit.

Each day GenCast and AIFS post their own fair-value quote on the season's contract for a four-city book (HDD in autumn and winter, CAT in summer) into a market quoted by the IFS ENS. A desk trades only when it disagrees with the market by more than the spread, so the flat line at zero is the market itself, quoting at consensus with no edge either way.

After a realistic 2-index-point bid-ask spread, only in Autumn 21 do both desks close ahead. GenCast gives back the whole of its Winter 21/22 edge and finishes behind the market. AIFS holds through Winter 21/22 but slips behind in Summer 22, the one season it quotes alone.

the technical version

The book is equal-weight 4 cities, priced on the (T00Z+T12Z)/2 index against 1991-2019 station normals, base temperature 18°C. The bias correction is expanding past-only day-1 mean bias vs settlement series, so no quote sees its own future. Spread settings of 0, 2 and 5 index points are all computed in the repo, and this panel fixes the realistic 2-point setting throughout. The totals shown are raw cumulative pick-off P&L, not a confidence interval.

Season
Autumn 21 · n = 33 quote days · 2-index-point spread · zero is the IFS-ENS-priced market

GenCast: +20.1 idx ptsAIFS: +9.3 idx pts

No tradable edge survives the spread, as parity predicts. Across two desks and two instruments the verdict is the same, and on ordinary weather neither desk separates these emulators from the operational ensemble.

event chapters

Two Events Outside the Backtest

Storm Eunice and the February 2021 Texas freeze both fall outside the pre-registered 66-day Desk A/B backtest window, so neither settles into a P&L figure. Each card below is a diagnostic reduction.

Storm Eunice

2022-02-11 to 2022-02-24

narrative · no P&L by construction

outside the 66 paired days the frozen Desk A/B backtest runs

Both models' fan-outs saw the same 13 delivery days of the storm fortnight. What the desk needs from them is the cut-out (turbines shutting down above 25 m s⁻¹), and here the two forecasts miss in opposite directions: GenCast put 0.27% of its day-1 forecast samples in cut-out, over-calling the 0.24% of settled farm-samples that actually reached it, while AIFS put only 0.17% in cut-out, under-calling the same reference. The two errors are not equally costly under the newsvendor asymmetry: GenCast's over-call commits to less power than the fleet delivers, the cheap 1x error, while AIFS's under-call risks committing to power the fleet cannot deliver when cut-out actually hits, a shortfall that costs 2x.

5 windiest GB-land ERA5 cells by peak 10 m wind over 2022-02-11..2022-02-24, cells fixed across all models and times; day-1 (24/36 h) leads, 00/12 UTC. ERA5 is drawn against the AIFS day-1 forecast. GenCast's day-1 forecast covers the same window and is not drawn here to keep this compact card to one forecast line (the soft-cap section of the shape finding shows both).

February 2021 Texas freeze

2021-02-08 to 2021-02-21

narrative · no P&L by construction

outside the 66 paired days the frozen Desk A/B backtest runs

For a temperature book the exposure is the tail of heating demand. AIFS is the only model run for this event, since GenCast has no coverage of the Texas freeze. Its day-1 forecast biased HDD by +0.12 against a box-mean observed HDD of 15.3 per day.

mean observed HDD, window (box mean)

15.3deg-days/day

descriptive · no CI

AIFS day-1 HDD bias vs observed (n = 13 days)

+0.12deg-days

descriptive · no CI

where this would break in the real world
  • Settlement throughout is derived from the ERA5 reanalysis. A real desk settles against metered generation and exchange fixes, each with a reporting lag and revision history that ERA5 does not have.
  • Desk B's city contracts settle on a single station thermometer. This book verifies against that station's ERA5 grid cell instead, and the two disagree by a measurable margin (quantified in the repo's basis figures).
  • Desk A's hub-height winds are not measured. They are extrapolated from 10 m ERA5 wind by a fixed power law (α = 0.14), and every downstream cut-out and power-curve number inherits that assumption.
  • Two seasons are enough to reject "no difference" where a result is significant. Pooled, they still give too few independent blocks to size a Sharpe ratio or any other risk-adjusted return. Sixty-six delivery days is also thin for the models that run only one member, whose whole disadvantage would show up on the rare day an ensemble would have widened.
  • Everything on this page beyond the GenCast-versus-AIFS comparison is exploratory. Adding the IFS ensemble to the wind book completes a design that was registered and never wired in. The four extra emulators, the zero-skill floor, the euro figures and the cost-versus-wind gradient were not pre-registered at all, and are recorded as post-hoc in the amendments that introduced them. The confirmatory test is unchanged and was not re-run.
  • "Level with the operational ensemble" is a bounded statement about a named set of systems. It proves no equivalence and says nothing about AI emulators in general. It is asserted of the pre-registered pair, GraphCast and Pangu-Weather are then observed to fall in the same band, and FourCastNet 3 and NeuralGCM, which do not, are reported separately. The intervals behind the claim cross zero, which on its own is no evidence of sameness. The claim rests on those intervals being a small fraction of the distance to a desk with no forecast.
  • The two groups above were separated by a stated rule applied to all four exploratory systems at once. The rule was written after the numbers were seen, so the split describes this backtest only.
  • The climatology baseline is a three-year window climatology from ERA5 2020-2023, coarser than a thirty-year normal, and serves only to put a scale on the cost axis.
  • The euro figures are the same capacity-factor costs multiplied by installed capacity and a fixed reference price, so they reorder nothing. Realised day-ahead prices would inflate the winter book severalfold, because those months are the gas crisis, and would measure the price series more than the forecasts.

get in touch

The code behind this site is in a private repository for now. If any of it is useful to you, get in touch.