tailspec · in practice

We put two AI forecast systems in charge of a trading desk. Which one ran the cheaper book depended on the country.

Same information a real desk would have had, decided a day ahead and settled against what actually happened. Two books: committing next-day wind power in Germany and Great Britain, and quoting temperature contracts against the IFS-ENS-priced market, ECMWF's operational physics ensemble.

cost difference per delivery day, GenCast vs AIFS, German book pooled over autumn 21 and winter 21/22

0.007[0.002, 0.012]

p = 0.0179

Costs are in capacity-factor (CF) units, not currency, and positive means GenCast cost more: on the German book, AIFS ran the cheaper desk. In Great Britain the same test flips, and GenCast ran the cheaper book there instead (-0.012 CF, p = 0.011).

the rules

Forecast, decide, settle

Forecast, decide, settleFORECAST8-member ensemble,day aheadGenCast · AIFSDECIDEDesk A: how much powerto promiseDesk B: what temperatureto quoteSETTLEagainst what happenedasymmetric penalties

Desk A promises next-day wind power for a capacity-weighted book of German and British wind farms, committing at the ensemble's newsvendor-optimal quantile under a 2:1 penalty for over-promising against under-promising. Delivery then settles against the wind power the ERA5 reanalysis implies actually happened.

Desk B quotes fair-value heating-degree-day (HDD) and cumulative-average-temperature (CAT) contracts for a four-city book each day, priced from realised-to-date ERA5 plus the bias-corrected ensemble forecast, then checks whether either model can be picked off against the IFS ENS market it is quoting into.

the technical version

Desk B's bias correction is expanding and past-only: each quote day's correction uses only forecast errors from verifications strictly before that day's issue time, so no quote sees its own future. Confidence intervals on both desks come from a circular block bootstrap over days (block length 5, 1,000 resamples), and p-values from a block sign-flip permutation test (10,000 resamples). Every decision rule, penalty and test was pre-registered before any season's forecasts were scored.

Desk A · day-ahead wind power

Desk A: selling wind power a day ahead

Split by season rather than pooled, the pattern holds. AIFS ran the cheaper German book in both Autumn 21 and Winter 21/22, though only the first season clears significance on its own (p = 0.033 against p = 0.287). GenCast ran the cheaper British book in both seasons too, and here it is the second season carrying the significance (p = 0.046 against p = 0.183).

Country
Season
Germany · Autumn 21 · n = 33 paired days

Δ mean daily cost (GenCast − AIFS): 0.007 CF [0.002, 0.013] permutation p = 0.033

AIFS ran the cheaper German book and GenCast the cheaper British one in both seasons: within each country the direction never flips, only the p-value does.

the technical version

Commitment is penalised 2:1 for over-promising against under-promising, and delivery settles against ERA5-derived wind power at a 100 m hub height. The turbine curve cuts out at 25 m s⁻¹ with hysteresis: it does not re-cut in until wind falls back to 20 m s⁻¹. Confidence intervals and p-values come from the same circular block bootstrap and block sign-flip permutation test as the pooled result above.

Desk B · temperature quote game

Desk B: quoting temperature risk against the market

The market here is the IFS ENS, ECMWF's operational physics ensemble and the incumbent forecast the desk is quoting into. Each day GenCast and AIFS post their own fair-value quote on the season's contract for the four-city book: HDD in autumn and winter, CAT in summer. A desk only trades when its quote disagrees with the market's by more than the spread, so the flat line at zero is the market itself: quoting exactly at consensus, no edge either way.

After a realistic 2-index-point bid-ask spread, only in Autumn 21 do both desks close ahead. GenCast gives back the whole of its Winter 21/22 edge and finishes behind the market. AIFS holds through Winter 21/22 but slips behind in Summer 22, the one season it quotes alone.

Season
Autumn 21 · n = 33 quote days · 2-index-point spread · zero is the IFS-ENS-priced market

GenCast: +20.1 idx ptsAIFS: +9.3 idx pts

Autumn 21 is the only season that puts both desks ahead: no tradable edge survives the spread, and at best these AI ensembles are evenly matched against the market they quote into.

the technical version

The book is equal-weight 4 cities, priced on the (T00Z+T12Z)/2 index against 1991-2019 station normals, base temperature 18°C. The bias correction is expanding past-only day-1 mean bias vs settlement series, so no quote sees its own future. Spread settings of 0, 2 and 5 index points are all computed in the repo, and this panel fixes the realistic 2-point setting throughout. The totals shown are raw cumulative pick-off P&L, not a confidence interval.

event chapters · narrative reductions

Two events, no desk P&L

Storm Eunice and the February 2021 Texas freeze both fall outside the pre-registered 66-day Desk A/B backtest window, so neither carries a settlement figure. Each card below is a diagnostic reduction only.

Storm Eunice

2022-02-11 to 2022-02-24

narrative · no P&L by construction

outside the 66 paired days the frozen Desk A/B backtest runs

Both models' fan-outs saw the storm fortnight. The desk question is whether they called the cut-out (turbines shutting down above 25 m s⁻¹) correctly, and here the forecasts over-called it: AIFS put 0.33% of its day-1 forecast samples in cut-out (n = 13 delivery days) and GenCast 0.29% (n = 8 delivery days), against 0.24% of settled farm-samples that actually cleared it. Over-calling cut-out means committing to less power than the fleet delivers, the cheap error under the newsvendor asymmetry, where an under-commitment that the fleet clears costs 1x and a shortfall against an over-commitment costs 2x.

GB fleet, farm-cell mean, hub-height 100 m, alpha=0.14; day-1 (24/36 h) leads. ERA5 settled against the AIFS day-1 forecast. GenCast's fan-out reaches under two-thirds of this window and is not drawn here (the soft-cap section on the method page shows both).

February 2021 Texas freeze

2021-02-08 to 2021-02-21

narrative · no P&L by construction

outside the 66 paired days the frozen Desk A/B backtest runs

The money question for a temperature book is the tail of heating demand. AIFS is the only model run for this event, since GenCast has no coverage of the Texas freeze. Its day-1 forecast biased HDD by +0.12 against a box-mean observed HDD of 15.3 per day, one day out on a 15-degree-day event.

mean observed HDD, window (box mean)

15.3deg-days/day

descriptive · no CI

AIFS day-1 HDD bias vs observed (n = 13 days)

+0.12deg-days

descriptive · no CI

where this would break in the real world
  • Settlement throughout is reanalysis-derived (ERA5), not metered. A real desk settles against metered generation and exchange fixes, both of which carry their own reporting lag and revision history that ERA5 does not.
  • Desk B's city contracts settle on a single station thermometer. This book verifies against that station's ERA5 grid cell instead, and the two disagree by a measurable margin, quantified in the repo's basis figures.
  • Desk A's hub-height winds are not measured. They are extrapolated from 10 m ERA5 wind by a fixed power law (α = 0.14), and every downstream cut-out and power-curve number inherits that assumption.
  • Two seasons of two-model data are enough to reject "no difference" where a result clears significance, but not enough to size a Sharpe ratio: a pooled two-season backtest has too few independent blocks for a stable risk-adjusted return estimate.

Every one of these is quantified in the private repo.

get in touch

The code behind this site currently lives in a private repository. If any of it bears on what you're working on, I'm happy to talk.