Spare parts · intermittent demand · stochastic inventory
The most accurate forecast is not the one that stocks the shelf best
For slow, lumpy spare-parts demand, judge forecasts by the inventory they produce: forecast the distribution, simulate the policy out of sample, then spend the stock budget where it buys the most service.
1. Demand that is mostly zeros
Most months, most parts sell nothing: 55% of part-months in the training window have no demand, and 71% in the hold-out. 79% of the series are intermittent (long gaps, stable sizes) and 19% are lumpy (long gaps, variable sizes); only 1.3% are smooth. This is the regime where the usual forecasting toolkit, built for continuous demand, is least reliable — and where the stock decision depends on the upper tail of demand, not its average.
Syntetos–Boylan–Croston classification on the 1,046 cleaned series (training window only). Dotted lines: ADI = 1.32, CV² = 0.49.
2. Point accuracy and distributional accuracy disagree
By MASE, SES is the most accurate of the 8 methods (0.674), narrowly ahead of TSB (0.678; paired difference 0.004, 95% CI 0.000 to 0.007) and SBA (0.683; paired difference 0.009, 95% CI 0.000 to 0.017); Croston is the least accurate of the smoothing methods (0.741). Yet a forecast of zero every month scores 0.482 (95% CI 0.457 to 0.509), better than every method, and stocks nothing. MASE is minimised by the median forecast, and the median of intermittent demand is usually zero. By pinball loss on lead-time demand — the metric that scores the quantiles a stock policy actually uses — LightGBM wins. Across the 7 methods that reach 95% fill, the Spearman rank correlation with the stock they need is -0.46 for MASE (p = 0.29), +0.50 for RMSSE (p = 0.25) and +0.89 for pinball loss (p = 0.007); with 7 methods, only the pinball loss correlation is significant at 5%.
| Method | MASE | RMSSE | Bias | Pinball | Cover q95 | Stock @ 95% fill (mo) |
|---|---|---|---|---|---|---|
| Naive | 0.717 [0.684, 0.753] | 0.719 | +0.024 | 1.315 [1.248, 1.384] | 60.1% | not reached |
| Moving average | 0.684 [0.659, 0.711] | 0.549 | +0.083 | 0.487 [0.456, 0.520] | 95.3% | 10.1 [9.3, 11.2] |
| SES | 0.674 [0.648, 0.705] | 0.554 | +0.057 | 0.556 [0.523, 0.591] | 92.6% | 12.8 [11.7, 14.0] |
| Croston | 0.741 [0.716, 0.766] | 0.563 | +0.215 | 0.455 [0.430, 0.484] | 98.7% | 9.6 [8.8, 10.6] |
| SBA | 0.683 [0.659, 0.708] | 0.543 | +0.057 | 0.443 [0.416, 0.476] | 97.9% | 9.6 [8.8, 10.6] |
| TSB | 0.678 [0.653, 0.705] | 0.548 | +0.073 | 0.506 [0.475, 0.537] | 95.1% | 11.1 [10.2, 12.2] |
| Bootstrap | 0.956 [0.937, 0.978] | 0.706 | +0.650 | 0.625 [0.606, 0.646] | 99.5% | 10.8 [9.8, 12.1] |
| LightGBM | 0.719 [0.694, 0.744] | 0.545 | +0.154 | 0.417 [0.392, 0.445] | 97.8% | 9.2 [8.4, 10.0] |
| Zero forecast baseline | 0.482 [0.457, 0.509] | 0.585 | -0.616 | 1.661 | 39.2% | stocks nothing 0% fill |
Means over series with 95% series-bootstrap intervals. MASE/RMSSE/bias: one-step forecasts, 12 rolling origins, scaled by in-sample naive errors. Pinball: lead-time demand, averaged over quantiles 0.8/0.9/0.95/0.99, scaled by mean demand. Cover: share of lead-time demand at or below the predicted 95% quantile. LightGBM's MASE, RMSSE and bias come from a one-step Tweedie mean model on the same features and windows; its pinball, cover and stock from one quantile model per service level, trained on lead-time demand. The zero forecast is a baseline, not a stocking method. Teal row: least stock at 95% fill; amber row: best MASE. On narrow screens RMSSE, bias and cover are hidden.
For count data a calibrated quantile q covers at least its nominal level, P(Y ≤ q) ≥ a, while P(Y < q) ≤ a. By that test, over the upper tail that sets the stock (levels 0.8 to 0.99), TSB has the best coverage (mean gap below 0.01 pp), yet it needs 21% more stock than LightGBM, which over-covers slightly (97.8% at the 95% quantile). The test is lenient: TSB's PIT histogram is U-shaped (end bins 1.25 and 1.22), so its full distribution is too narrow even where its upper quantiles cover. SES under-covers at every level from 0.75 (92.6% at the 95% quantile), which is why ordering up to its 95% quantile delivers only 92.3% fill; its PIT histogram is U-shaped (end bins 1.30 and 1.55). Calibration is necessary, not sufficient: at equal achieved fill what separates methods is how well they rank parts (sharpness and resolution). Pinball loss scores both.
Coverage of lead-time-demand quantiles minus the nominal level (lead time 2 months). Shaded: under-coverage, which a calibrated count forecast never shows. Naive is off scale and left out.
Non-randomised PIT histograms for count forecasts (Czado, Gneiting & Held, 2009), lead time 2 months, ten bins. A calibrated forecast gives 1.00 everywhere; high values at both ends mean intervals that are too narrow, a hump in the middle means too wide. LightGBM predicts quantiles only, not a full distribution, so its calibration is read from the coverage chart above.
3. Stocking the shelf: the trade-off curve
Every method traces a curve: more stock buys more service, with diminishing returns. At 95% fill and a 2-month lead time, LightGBM needs 9.2 months of average stock and SES needs 12.8: 39% more (95% CI 29% to 51%). Croston, the worst smoothing method by MASE, needs only 5% more than LightGBM. At this lead time SBA and Croston are statistically tied with LightGBM (paired 95% CI of the gap SBA -0.5% to 9.5%; Croston -0.4% to 9.9%). LightGBM has the lowest stock at every lead time; its lead over every other method is significant only at 3 months (the closest, SBA, needs 7% more, 95% CI 2% to 13%). By class, LightGBM's edge comes from lumpy parts (7.9 vs 9.4 months for SBA); on intermittent parts LightGBM, SBA and Croston need 9.5, 9.5 and 9.6 months. These class figures are point estimates; their bootstrap intervals overlap. Reading each curve at 95% picks each method's service level with hindsight (an efficient-frontier view). As deployed, at the nominal 95% quantile, the ordering holds: SES holds 10.6 months for 92.3% fill, LightGBM 8.9 months for 94.6%. A quantile target is not a fill-rate promise: LightGBM avoids a stock-out in 98.3% of part-months yet serves 94.6% of units, consistent with stock-outs falling on the large orders that carry most units.
Each point is one target quantile (0.50 to 0.99) used as the order-up-to level, re-forecast every month and simulated on the 12 hold-out months with backorders. x: average on-hand stock in months of hold-out demand; y: share of demand served from stock. Dotted markers: stock needed for 95% fill, by linear interpolation along the curve. Switch the lead time above the chart.
How much does the gap depend on tuning?
SES's gap depends on its smoothing constant α. The tuning criterion, pooled in-sample MSE, is flat near its minimum: α = 0.2, 0.25, 0.3 are within 1% of the best, yet SES needs 11.4, 12.8, 14.9 months for 95% fill, a gap to LightGBM of 24%, 39%, 61%. The headline 39% is the gap at the MSE-tuned α = 0.25. The direction does not depend on it: from α = 0.1, where SES needs the least stock (10.0 months, MASE 0.702), to α = 0.3, where its MASE is best (0.674), the forecast gets more accurate by MASE and needs more stock (14.9 months). From α = 0.4 it never reaches 95% fill. SBA is insensitive: its tied constants (α = 0.35, 0.4, 0.5) need 9.6 to 9.8 months, with MASE from 0.674 to 0.688. With the lead-time variance of its own model, ETS(A,N,N), instead of the independent-periods shortcut, SES needs 15.8 months, more than with the shortcut: the shortcut does not explain its stock penalty.
Each point is one smoothing constant α of the tuning grid, run through the same 12 rolling origins and simulation at a 2-month lead time; diamonds mark the MSE-tuned values. Constants whose policy never reaches 95% fill are left out. Dashed line: LightGBM.
4. Spending a stock budget across parts
A uniform rule gives every part the same service target. Marginal analysis (Sherbrooke) instead buys, unit by unit, the stock that removes the most expected backorders per dollar. With dispersed costs independent of demand (σ = 1, ρ = 0), it reaches 95% system fill with 12% less investment: the median over 5 cost draws, which range from 4% to 15%. The 95% interval over parts and cost draws, -2% to 20%, includes zero. At 90% fill the saving is 19% (11% to 26%). The gain needs cost dispersion. With equal unit costs the two rules coincide by construction: marginal analysis then adds units in order of P(X > s), which is the uniform quantile rule. The measured saving, 0%, is a sanity check. The gain also depends on which parts are expensive. When slow movers are the expensive parts (ρ = −0.5) and σ ≤ 1, the saving ranges from -7% to 6% and can be negative: marginal analysis is optimal for the expected-backorder model built on the forecast distribution, not for the fill rate measured out of sample. With very dispersed costs (σ = 1.5) the gain survives even that pattern (median 22%). Overall it is positive in 36 of 46 cost scenarios.
Investment = Σ unit cost × stock level. Static stock levels set once at the start of the hold-out from SBA lead-time-demand distributions, simulated over 12 months. The curves are one synthetic cost draw (seed 2, the median of 5; log-normal, median $50, σ = 1, independent of demand), because the data has no costs.
| Cost dispersion | ρ = -0.5 | ρ = 0 | ρ = +0.5 |
|---|---|---|---|
| σ = 0.5 | -2% -7% to -1% | 0% -1% to 5% | 5% 3% to 8% |
| σ = 1 | 5% -3% to 6% | 12% 4% to 15% | 18% 9% to 22% |
| σ = 1.5 | 22% 12% to 24% | 32% 18% to 38% | 39% 27% to 44% |
Saving at 95% fill by cost dispersion σ and cost–demand correlation ρ (SBA distributions): median over 5 seeds, range on the second line. With equal costs (σ = 0) the two rules coincide by construction; the measured saving is 0%.
| Distribution | 90% fill | 95% fill |
|---|---|---|
| Moving average | 21% 16% to 25% | 14% 8% to 19% |
| SES | 23% 14% to 26% | 18% 9% to 19% |
| Croston | 21% 16% to 23% | 13% 6% to 18% |
| SBA | 19% 15% to 22% | 12% 4% to 15% |
| TSB | 23% 13% to 26% | 13% 5% to 18% |
| Bootstrap | 16% 12% to 21% | 11% 4% to 16% |
Saving by distributional method at a 2-month lead time, σ = 1 and ρ = 0: median and range over 5 cost draws. LightGBM is absent because marginal analysis needs a full distribution to compute expected backorders, and it predicts quantiles only.
5. What this does not show
- One data set, one 12-month hold-out (2001–2002). Intervals come from resampling parts, not time; they do not cover a different demand regime.
- SES's stock gap holds at its MSE-tuned smoothing constant; constants the tuning criterion cannot tell apart give 24%, 39% and 61%. The series bootstrap does not include this tuning uncertainty.
- The literature cleaning rule keeps parts with demand in the last 15 months, which overlap the hold-out: obsolescent parts, where TSB is designed to shine, are excluded.
- Stock at equal fill is read from each trade-off curve by linear interpolation between service levels 1 pp apart, with each method's level chosen on the hold-out; the reference method is selected on the same data.
- Unit costs are synthetic; the allocation result is a sensitivity study, not a cost estimate. Allocation stock levels are static over the hold-out, while the forecast comparison re-forecasts every month.
- The allocation needs a full distribution to compute expected backorders, which LightGBM's quantiles do not give, so it uses SBA, the least-stock distributional method on the same hold-out.
- Lead times are deterministic, backorders are assumed (no lost sales) and every method starts the hold-out with on-hand stock equal to its first order-up-to level.
- Parametric methods assume independent periods (variance = horizon × one-step MSE), which understates SES's lead-time variance under its own model; section 3 reports SES with the ETS(A,N,N) variance.