gridcast

Can a day-ahead forecast of French electricity demand publish uncertainty bands that stay honest through an energy crisis? On average yes, month to month only partly.

Backtest 2022-04-01 to 2026-08-31 · 38,732 hourly day-ahead forecasts, each issued at 12:00 Paris time the day before · source code

80.0%coverage of the published 80% interval (target 80%)
95.0%coverage of the published 95% interval (target 95%)
1.95%gridcast MAPE vs 1.64% for RTE's own J-1 forecast (both level-corrected)
50%of 30-day windows within ±5 pp of 80%, vs 12% with a frozen calibration

Calibration through the energy crisis

Conformal prediction guarantees coverage on average, for exchangeable data. Demand forecast errors are not exchangeable: they jump when heating starts in November, shrink in spring, and in autumn 2022 the whole level of French demand dropped. A calibration frozen on April 2021 to March 2022 still covers 82.2% of hours overall, yet its 30-day coverage swings between roughly 50% and 100% and is within ±5 pp of target only 12% of the time. Recalibrating on a rolling 90-day window with conformalized quantile regression (CQR) does most of the repair: 42% of windows on target. Adaptive conformal inference (ACI) at the pre-set step adds a little at 80% (50%) and nothing at 95%, where it widens the bands (5.53 to 5.98 GW) and worsens the interval score (7.14 to 7.33 GW). No method fully absorbs the crisis winter: over 2022-09-01 to 2023-03-31 the published 80% interval covered 77.3%. Click legend entries to compare the other methods; shaded: energy crisis.

The 95% version. Coverage cannot exceed 100%, so here the ±5 pp band is [90%, 100%] and only flags under-coverage: the on-target shares at 80% and 95% are not comparable.

All interval methods, 2022-04-01 to 2026-08-31

80% intervals

Methodcoverage [95% CI]width GWWinkler GW 30-day on target
Quantile LightGBM, uncalibrated55.6%[53.9%, 57.3%]2.015.490%
Split conformal, static82.2%[79.9%, 84.3%]3.395.1712%
Split conformal + ACI80.6%[78.5%, 82.4%]3.194.8938%
Split conformal, rolling 90 d79.7%[77.4%, 81.8%]3.235.0229%
CQR, rolling 90 d80.4%[78.6%, 82.1%]3.354.8542%
CQR rolling + ACI80.0%[78.4%, 81.6%]3.444.8250%

95% intervals

Methodcoverage [95% CI]width GWWinkler GW 30-day ≥ 90%
Quantile LightGBM, uncalibrated81.9%[80.5%, 83.3%]3.798.7811%
Split conformal, static95.8%[94.8%, 96.6%]6.247.9753%
Split conformal + ACI95.3%[94.2%, 96.2%]5.757.5970%
Split conformal, rolling 90 d93.9%[92.6%, 95.0%]5.497.7367%
CQR, rolling 90 d94.7%[93.7%, 95.6%]5.537.1484%
CQR rolling + ACI95.0%[94.1%, 95.8%]5.987.3384%

Coverage CIs from a week-block bootstrap (1000 resamples). Winkler (interval) score: width plus 2/α times the miss distance; lower is better. Highlighted: coverage more than 2 pp from nominal. ACI clips its level when it would ask for an infinite interval (Gibbs & Candès's guarantee assumes it does not); the published 95% interval was capped at the widest calibration score on 26 of the scored days.

By period

PeriodMAPE gridcastMAPE RTE (corrected) 80% coverage, frozenfrozen + ACICQR + ACI
pre-crisis1.83%1.57%86.7%90.0%83.9%
crisis2.09%1.68%76.9%73.4%77.3%
post-crisis1.94%1.64%82.6%80.6%80.0%

Pre-crisis is only five months (April to August 2022), all in the summer, when every method over-covers. Frozen + ACI starts from its burn-in state and, at γ = 0.02 per day, narrows too slowly to catch up, hence its higher coverage there. Too short a window to judge.

Point accuracy: RTE wins, once compared fairly

Scored naively, both forecasts as produced, gridcast beats RTE's published J-1 forecast (2.08% vs 2.62% MAPE against the consolidated demand series). That comparison is confounded by a level gap that public data cannot attribute to definitions or to forecast error. From 2023 on, RTE's forecast runs 2.2% below the consolidated values on average (within 1% in 2015-2021). The gap depends strongly on the hour: largest at 00:00, 13:00, 15:00 (-4.3% to -3.9%), smallest at 06:00 (+0.5%). On the only real-time data in the backtest (1,488 hours, July and August 2026) its bias is +0.3% and its MAPE 1.49%, but two summer months settle nothing. Both forecasts are therefore corrected the same way, by their own mean error over the trailing 28 days at the same local hour, using only errors at least two days old. On that footing RTE is better overall, 1.64% vs 1.95%, and in each period, though not significantly in the five pre-crisis months (difference +0.26 pp, 95% CI [-0.04 pp, +0.58 pp]), where the level correction itself costs gridcast +0.12 pp. gridcast is ahead in 17 of 53 months, all between June and October, when demand is flat and temperature-driven. RTE likely benefits from richer inputs (many more weather stations, cloud cover, embedded solar, operational knowledge) and may issue later than 12:00; gridcast sees one public temperature forecast for ten cities. With observed instead of forecast weather the same model reaches 1.85%. The contribution here is not a better point forecast but a transparent, calibrated uncertainty layer anyone can audit.

A week of the December 2022 cold snap

How it works

Limitations