gridcast
Can a day-ahead forecast of French electricity demand publish uncertainty bands that stay honest through an energy crisis? On average yes, month to month only partly.
Backtest 2022-04-01 to 2026-08-31 · 38,732 hourly day-ahead forecasts, each issued at 12:00 Paris time the day before · source code
Calibration through the energy crisis
Conformal prediction guarantees coverage on average, for exchangeable data. Demand forecast errors are not exchangeable: they jump when heating starts in November, shrink in spring, and in autumn 2022 the whole level of French demand dropped. A calibration frozen on April 2021 to March 2022 still covers 82.2% of hours overall, yet its 30-day coverage swings between roughly 50% and 100% and is within ±5 pp of target only 12% of the time. Recalibrating on a rolling 90-day window with conformalized quantile regression (CQR) does most of the repair: 42% of windows on target. Adaptive conformal inference (ACI) at the pre-set step adds a little at 80% (50%) and nothing at 95%, where it widens the bands (5.53 to 5.98 GW) and worsens the interval score (7.14 to 7.33 GW). No method fully absorbs the crisis winter: over 2022-09-01 to 2023-03-31 the published 80% interval covered 77.3%. Click legend entries to compare the other methods; shaded: energy crisis.
The 95% version. Coverage cannot exceed 100%, so here the ±5 pp band is [90%, 100%] and only flags under-coverage: the on-target shares at 80% and 95% are not comparable.
All interval methods, 2022-04-01 to 2026-08-31
80% intervals
| Method | coverage [95% CI] | width GW | Winkler GW | 30-day on target |
|---|---|---|---|---|
| Quantile LightGBM, uncalibrated | 55.6%[53.9%, 57.3%] | 2.01 | 5.49 | 0% |
| Split conformal, static | 82.2%[79.9%, 84.3%] | 3.39 | 5.17 | 12% |
| Split conformal + ACI | 80.6%[78.5%, 82.4%] | 3.19 | 4.89 | 38% |
| Split conformal, rolling 90 d | 79.7%[77.4%, 81.8%] | 3.23 | 5.02 | 29% |
| CQR, rolling 90 d | 80.4%[78.6%, 82.1%] | 3.35 | 4.85 | 42% |
| CQR rolling + ACI | 80.0%[78.4%, 81.6%] | 3.44 | 4.82 | 50% |
95% intervals
| Method | coverage [95% CI] | width GW | Winkler GW | 30-day ≥ 90% |
|---|---|---|---|---|
| Quantile LightGBM, uncalibrated | 81.9%[80.5%, 83.3%] | 3.79 | 8.78 | 11% |
| Split conformal, static | 95.8%[94.8%, 96.6%] | 6.24 | 7.97 | 53% |
| Split conformal + ACI | 95.3%[94.2%, 96.2%] | 5.75 | 7.59 | 70% |
| Split conformal, rolling 90 d | 93.9%[92.6%, 95.0%] | 5.49 | 7.73 | 67% |
| CQR, rolling 90 d | 94.7%[93.7%, 95.6%] | 5.53 | 7.14 | 84% |
| CQR rolling + ACI | 95.0%[94.1%, 95.8%] | 5.98 | 7.33 | 84% |
Coverage CIs from a week-block bootstrap (1000 resamples). Winkler (interval) score: width plus 2/α times the miss distance; lower is better. Highlighted: coverage more than 2 pp from nominal. ACI clips its level when it would ask for an infinite interval (Gibbs & Candès's guarantee assumes it does not); the published 95% interval was capped at the widest calibration score on 26 of the scored days.
By period
| Period | MAPE gridcast | MAPE RTE (corrected) | 80% coverage, frozen | frozen + ACI | CQR + ACI |
|---|---|---|---|---|---|
| pre-crisis | 1.83% | 1.57% | 86.7% | 90.0% | 83.9% |
| crisis | 2.09% | 1.68% | 76.9% | 73.4% | 77.3% |
| post-crisis | 1.94% | 1.64% | 82.6% | 80.6% | 80.0% |
Pre-crisis is only five months (April to August 2022), all in the summer, when every method over-covers. Frozen + ACI starts from its burn-in state and, at γ = 0.02 per day, narrows too slowly to catch up, hence its higher coverage there. Too short a window to judge.
Point accuracy: RTE wins, once compared fairly
Scored naively, both forecasts as produced, gridcast beats RTE's published J-1 forecast (2.08% vs 2.62% MAPE against the consolidated demand series). That comparison is confounded by a level gap that public data cannot attribute to definitions or to forecast error. From 2023 on, RTE's forecast runs 2.2% below the consolidated values on average (within 1% in 2015-2021). The gap depends strongly on the hour: largest at 00:00, 13:00, 15:00 (-4.3% to -3.9%), smallest at 06:00 (+0.5%). On the only real-time data in the backtest (1,488 hours, July and August 2026) its bias is +0.3% and its MAPE 1.49%, but two summer months settle nothing. Both forecasts are therefore corrected the same way, by their own mean error over the trailing 28 days at the same local hour, using only errors at least two days old. On that footing RTE is better overall, 1.64% vs 1.95%, and in each period, though not significantly in the five pre-crisis months (difference +0.26 pp, 95% CI [-0.04 pp, +0.58 pp]), where the level correction itself costs gridcast +0.12 pp. gridcast is ahead in 17 of 53 months, all between June and October, when demand is flat and temperature-driven. RTE likely benefits from richer inputs (many more weather stations, cloud cover, embedded solar, operational knowledge) and may issue later than 12:00; gridcast sees one public temperature forecast for ten cities. With observed instead of forecast weather the same model reaches 1.85%. The contribution here is not a better point forecast but a transparent, calibrated uncertainty layer anyone can audit.
A week of the December 2022 cold snap
How it works
- Information set. The forecast for day D is issued at 12:00 Paris on D-1. It may use demand up to 11:00 (one hour publication lag), calendar features and temperature forecasts whose model run had finished by then (24 h-ahead values for the first hours of D, 48 h-ahead after that). A test poisons every later value and checks the features, the backtest refit and the live forecast do not change.
- Model. LightGBM on calendar, similar-day demand lags, the latest week-over-week level drift and a population-weighted national temperature (ten largest metro areas) with exponential smoothing for thermal inertia. Refit monthly on all past days. Temperature forecasts are bias-corrected per hour on a trailing 60-day window, and the demand forecast by its own mean error over the trailing 28 days.
- Intervals. Split conformal on absolute residuals, conformalized quantile regression (CQR) on LightGBM quantiles, and adaptive conformal inference (Gibbs & Candès, 2021) on top. Because the forecast for D+1 is issued before D ends, every method learns from its errors with a two-day delay.
- Tomorrow's forecast.
make liveruns the same protocol on today's data: it retrains, forecasts tomorrow with both intervals, appends to a local forecast log and scores past entries once their outcome is published. Days missing from the log are hindcast with the backtest protocol so calibration never runs dry; they are kept out of the track record.
Limitations
- The backtest information set is better than the live one. Demand lags and all feedback use RTE's consolidated series, published months later; at issue time only the real-time vintage existed. Observed temperature before the issue time is ERA5 reanalysis, which arrives about 5.5 days late. Masking those days as the live job must moves MAPE from 1.946% to 1.949% and the published 80% coverage from 80.03% to 80.03%; the vintage effect cannot be measured with public data.
- Only temperature has archived day-ahead forecasts before 2024 in the open archive; cloud cover and wind, which RTE uses, are not in the model.
- Training uses reanalysis (ERA5) temperature, prediction uses forecasts; the conformal layer absorbs the mismatch, the point forecast pays for it. The uncalibrated quantiles ignore weather-forecast error by construction: fed observed weather, they cover 57.6% at nominal 80% instead of 55.6%, so the mismatch explains little of their under-coverage.
- 516 hours (21.5 days) of forecast temperature (30 Dec 2023 00:00 to 20 Jan 2024 11:00 UTC) come from JMA instead of GFS, the only archived model covering them.
- Features were designed knowing that French demand dropped in 2022 (the week-ratio feature exists to absorb such level shifts); hyperparameters were set a priori, not tuned.
- RTE's J-1 forecast may be published later than 12:00 on D-1, so the comparison slightly favours RTE.
- Coverage is marginal over hours, not conditional on weather regime or hour; see the repository for coverage by hour.