Delay contagion
How a late flight infects the rest of the day through aircraft rotations, and where a limited budget of turnaround buffer absorbs the contagion best.
The decision payoff (Southwest, 92 held-out days): arrival delay avoided, measured against the padded schedule, for the buffer each policy actually schedules. Bands: 95% bootstrap CI over whole weeks (14 weeks). Details in section 4.
140% of delay comes from the previous flight
The BTS codes every arrival delay of 15 minutes or more into five causes. One of them, late-arriving aircraft, is not a new problem: it is an earlier delay riding the aircraft into its next leg. Over Aug 2025 to Jul 2026 it is 40% of all cause-coded delay minutes, and it builds through the day: 1.2 minutes per departing flight at 6:00, 11.7 by 20:00. Across carriers the share runs from 22% (SkyWest) to 52% (Frontier); cause codes are self-reported, so that spread needs care.
Cause-coded delay minutes per departing flight, by scheduled local departure hour, all reporting carriers.
2The physics of a turn
A turn absorbs an inbound delay as long as the aircraft still has its minimum turn time on the ground. The model is a hinge,
outbound_delay = a + β · max(0, inbound_delay − (scheduled_turn − τ))
where τ is the effective minimum turn time (it includes crew, cleaning and boarding) and β the pass-through rate beyond it. Fitted per carrier on Aug 2025 to Apr 2026 by profiling τ over a 1-minute grid, with 95% CIs from a bootstrap that resamples whole days. τ ranges from 39 to 63 minutes and β from 0.98 to 1.05: past the threshold, delay passes through about one-for-one. The held-out bins show a rounded knee rather than a sharp corner. τ intervals narrower than about a minute reflect the grid, not certainty to the second.
Dots: held-out months (May 2026 to Jul 2026), mean outbound delay in 5-minute bins of ground time left. Lines: hinge fitted on training months. Click legend entries to compare carriers.
Out of sample, the hinge predicts outbound delay with a mean absolute error of 11.67 minutes. Baselines fitted on the same training turns: a constant (22.50), a linear model in inbound delay (14.67), the same with the scheduled turn as a second regressor (14.65), and the hinge with τ pinned at zero (16.80). Slack only helps once it enters through the threshold. A gradient-boosted model on inbound delay, scheduled turn, carrier and hour does better, at 11.23 (better on 12 of 13 carriers with held-out data); the hinge keeps 87% of its improvement over the slack-aware linear model with two interpretable parameters, and it keeps the decision model linear. 0.1% of held-out turns fall outside the delay window used for fitting and scoring (inbound −90 to 600, outbound −60 to 600 minutes).
Is β causal? Several carriers and hubs have β slightly above 1 with intervals that exclude 1. A turn cannot amplify delay, so this points to common shocks: weather or ATC at the station delays the inbound and the outbound's own departure together. The decision model reads β as the effect of slack; section 4 checks how much that matters.
| Carrier | Train turns | τ (min) | β | MAE const. | MAE linear | MAE lin. + slack | MAE boosted | MAE hinge |
|---|---|---|---|---|---|---|---|---|
| Southwest | 801,017 | 48 [47, 48] | 0.98 [0.98, 0.99] | 21.4 | 12.4 | 12.0 | 9.7 | 10.2 |
| Delta | 537,098 | 60 [60, 61] | 1.04 [1.02, 1.05] | 20.1 | 14.3 | 14.4 | 11.0 | 11.5 |
| SkyWest | 464,259 | 41 [41, 42] | 1.02 [1.01, 1.02] | 21.6 | 13.8 | 14.2 | 10.0 | 10.1 |
| American | 463,066 | 57 [57, 57] | 1.01 [1.00, 1.01] | 26.1 | 18.0 | 17.6 | 12.9 | 13.4 |
| United | 393,093 | 63 [62, 63] | 1.02 [1.00, 1.03] | 22.3 | 17.0 | 17.1 | 13.7 | 13.9 |
| Republic | 195,332 | 41 [41, 42] | 1.01 [1.00, 1.02] | 22.3 | 13.9 | 14.1 | 10.8 | 11.0 |
| Envoy | 169,345 | 39 [39, 40] | 1.05 [1.03, 1.07] | 19.3 | 12.8 | 13.0 | 9.0 | 9.3 |
| Alaska | 146,960 | 55 [54, 57] | 0.99 [0.97, 1.00] | 15.8 | 12.1 | 12.3 | 9.3 | 11.1 |
| PSA | 130,844 | 41 [41, 41] | 1.01 [1.01, 1.02] | 29.3 | 17.6 | 17.8 | 13.3 | 13.8 |
| Frontier | 110,747 | 62 [61, 63] | 1.03 [1.01, 1.03] | 29.0 | 16.8 | 16.9 | 13.5 | 14.0 |
| JetBlue | 106,069 | 61 [61, 62] | 1.02 [1.01, 1.03] | 30.2 | 18.4 | 18.5 | 14.9 | 15.3 |
| Spirit | 75,932 | 63 [62, 63] | 1.01 [1.01, 1.03] | 27.5 | 17.7 | 17.2 | 13.9 | 13.8 |
| Allegiant | 65,318 | 62 [61, 63] | 1.02 [1.02, 1.03] | 28.1 | 12.7 | 12.6 | 11.2 | 11.3 |
3Measuring the contagion
To count how much later delay one late departure creates, take clean starts: legs whose aircraft was ready (first leg of the chain, or inbound on time) but which still left late. Sum the arrival delay of every later leg of that aircraft, and subtract the same sum for on-time clean starts of the same carrier, on the same day, in the same part of the day, with the same number of legs left to fly (capped at 4+). Matching on the day removes carrier-wide, day-level conditions, not local weather or a ground stop at one airport; matching on legs left removes exposure.
Extra downstream arrival delay against the primary delay, with 95% day-bootstrap intervals; dashed: one-for-one.
Delays that start in the morning travel furthest, because more legs follow them: the multiplier is 0.99 at 5:00, peaks at 1.13 at 8:00 and falls to 0.33 for 21:00 departures.
Multiplier by local departure hour of the delayed leg, 95% day-bootstrap band.
How robust is the 0.98?
"Clean start" is an approximation of primary: 19% of these events still carry a late-aircraft cause code (28% of their minutes), for instance after a tail swap. The headline survives every check below. Adding the origin airport to the match key (local weather, ground stops) leaves 159,615 events and gives 0.95; exact legs left gives 0.98; dropping late-aircraft-coded legs gives 0.99. Origin matching also explains the odd first bin: slips of 1 to 14 minutes look more contagious (1.22) because they cluster at congested airports; matched within the airport they drop to 1.02. Legs left is counted on the realised chain, after cancellations and swaps the delay may itself cause, so it is partly a post-treatment quantity.
| Matching design | Events (15+ min) | Multiplier | Slips of 1-14 min |
|---|---|---|---|
| Headline: carrier x day x part of day x legs left (4+) | 396,690 | 0.98 [0.96, 1.01] | 1.22 |
| + origin airport in the match | 159,615 | 0.95 [0.93, 0.97] | 1.02 |
| Exact legs left (no 4+ cap) | 381,386 | 0.98 [0.96, 1.01] | 1.17 |
| Without legs that carry a late-aircraft code | 320,338 | 0.99 [0.97, 1.02] | 1.19 |
By carrier, the multiplier runs from 0.68 (United) to 1.91 (Allegiant); Southwest, whose aircraft fly the most legs per aircraft-day (5.4) on a median scheduled turn of 45 minutes (shortest median: 43 min for Envoy and PSA), sits at 1.24.
Super-spreader airports
Raw airport multipliers mostly reflect who flies there: the top of the raw ranking, MDW (Chicago) at 1.42, would score 1.23 on its carrier mix alone. The ranking below compares each delayed start with its own carrier's network multiplier. On that measure the top three are MKE (Milwaukee), ORD (Chicago) and DCA (Washington), at +0.25 to +0.27 downstream minutes per primary minute above their carrier mix; their 95% intervals overlap, so no single airport is singled out.
Airports with at least 1,500 delayed clean starts. Colour: multiplier above (teal) or below (amber) the carrier-mix expectation; size: primary minutes generated there. Hover for intervals.
4Where to put the buffer
Padding turns costs aircraft time, so the question is where a limited budget does the most good. For Southwest (2,930 turns a day in training), a decision is an extra number of scheduled minutes for every turn at a station in a given departure hour (1,565 station-hour cells). The delay recursion along every aircraft chain uses the fitted τ and β, with primary delays and en-route changes taken from the data, so that zero buffer replays each observed day exactly. A sample-average-approximation LP over all 273 training days chooses the buffers that minimise expected arrival delay under a daily budget (1,811,979 constraints, 1,813,543 variables, solved with HiGHS). Policies are then scored on 92 held-out days (the chart at the top).
What "delay avoided" means. Arrival delay is measured against the padded published schedule, the DOT on-time definition airlines are judged on. A buffer never makes an aircraft earlier on the clock: it moves that departure and every later one of the aircraft's day back in the timetable, and absorbs the late inbound inside that slack. Because the shift carries down the chain, the same buffer minutes count as avoided on each later leg that would have been late, which is how one buffer minute can avoid more than one minute of delay. Padding buys predictability, not speed.
Two heuristics show how much of that needs an optimiser. Padding the turns with the most observed propagated delay does worse than uniform padding (0.62 per buffer minute; paired difference to uniform -109 [-129, -89] min/day): those turns are late because their inbound is late, not because their slack is short. A marginal-value greedy that ranks cells by the delay a full buffer would remove there, one cell at a time, gets much closer: 1.32 per buffer minute, so the LP's joint placement adds a further 15% per buffer minute. Rounding the LP's buffers to 5-minute schedule steps costs little (1.51 per buffer minute at 611 min/day). Fractional uniform padding is 12 seconds per turn, which no planner would publish; its implementable version, 5 minutes on randomly chosen station-hours, gets 0.82.
| Policy | Buffer used (min/day) | Delay avoided (min/day) | Per buffer min | On time (A15) | Clock lateness (min/day) |
|---|---|---|---|---|---|
| LP-optimised | 718 | 1,093 [1,015, 1,169] | 1.52 | 69.6% | +751 |
| LP rounded to 5-min steps | 611 | 926 [863, 987] | 1.51 | 69.5% | +672 |
| Marginal-value greedy | 743 | 982 [899, 1,068] | 1.32 | 69.5% | +1,007 |
| Greedy: most-delayed turns | 758 | 473 [407, 542] | 0.62 | 69.4% | +489 |
| Uniform padding (fractional) | 643 | 582 [520, 644] | 0.91 | 69.8% | +750 |
| Uniform, 5 min on random cells | 659 | 538 [482, 594] | 0.82 | 69.4% | +838 |
Budget 600 min/day, 92 held-out days. On time (A15): share of arrivals less than 15 minutes late against the padded schedule (69.2% with no buffer). Clock lateness: change in positive lateness against the original timetable (baseline 71,581 min/day).
On the clock, every policy makes the day later, by a similar amount (+751 min/day for the LP, +750 for uniform padding): the LP does not create time, it places it where it stops the most knock-on delay. On the A15 count it is not the best policy either (69.6% on time against 69.8% for fractional uniform padding): it minimises delay minutes, so it spends buffer on the long cascades. Most of fractional uniform padding's A15 edge is a threshold artefact: a fraction of a minute on every turn moves arrivals of exactly 15 minutes under the line, and the 5-minute version scores 69.4%. An airline that cares about the A15 count should optimise that objective instead.
Does the evaluation flatter the LP?
The held-out replay uses the same hinge the LP optimised against: out of sample in days, but in-model. Re-scoring the same buffers with the replay's τ shifted by ±5 and ±10 minutes, or with β = 0.8, moves the absolute minutes a lot (742 to 1,268 min/day for the LP) but keeps its per-minute advantage over uniform padding in the range 1.5× to 1.9×. The ranking is robust; the absolute numbers depend on the model.
| Replay τ (min) | Replay β | LP avoided | Uniform avoided | LP / uniform per buffer min |
|---|---|---|---|---|
| 48 | 0.98 | 1,093 | 582 | 1.68× |
| 38 | 0.98 | 865 | 414 | 1.87× |
| 43 | 0.98 | 981 | 494 | 1.78× |
| 53 | 0.98 | 1,190 | 678 | 1.57× |
| 58 | 0.98 | 1,268 | 771 | 1.47× |
| 48 | 0.80 | 742 | 399 | 1.67× |
How many scenario days?
An LP fitted on a handful of days learns their noise. Re-solving on random draws of 10 to 273 training days, the in-sample edge over uniform padding shrinks from 157% to 74% while the held-out edge rises from 42% to 88% (85% at 90 days, 86% at 210): beyond about 90 days it levels off. The two curves are not a convergence test: the held-out months are summer and busier than the average training day.
Budget 600 min/day. Faint dots: individual draws (3 per size; the full set is one draw); lines: mean over draws.
Where the buffer goes
The LP concentrates the budget on 185 of 1,565 cells, listed by buffer minutes per day. It does not follow the super-spreader ranking: that ranking measures how far a delay that starts at an airport travels, while the LP looks for turns whose scheduled slack is short relative to τ and to the inbound delay that typically reaches them, with many legs behind them. The last column shows what each cell would be worth padded on its own.
| Station | Departure hour | Turns/day | Buffer per turn (min) | Buffer min/day | Propagated delay per turn (min) | Avoided per buffer min, alone |
|---|---|---|---|---|---|---|
| SFO | 10:00 | 2.0 | 19.0 | 38.7 | 23.1 | 1.28 |
| SFO | 11:00 | 1.9 | 14.2 | 27.4 | 20.4 | 1.07 |
| SFO | 13:00 | 1.3 | 20.0 | 26.9 | 27.5 | 1.26 |
| SAN | 13:00 | 6.2 | 4.1 | 25.6 | 18.5 | 0.71 |
| SFO | 12:00 | 2.0 | 11.2 | 22.2 | 19.2 | 0.92 |
| SAN | 14:00 | 6.0 | 3.0 | 18.1 | 16.3 | 0.67 |
| SAN | 12:00 | 5.7 | 3.1 | 17.6 | 16.3 | 0.69 |
| SFO | 15:00 | 1.7 | 10.2 | 17.6 | 25.2 | 0.92 |
| SEA | 11:00 | 1.5 | 11.0 | 16.0 | 13.5 | 0.98 |
| SFO | 16:00 | 1.4 | 11.0 | 15.1 | 25.6 | 0.93 |
| AUS | 14:00 | 7.0 | 2.0 | 14.3 | 13.9 | 0.64 |
| PDX | 10:00 | 2.0 | 6.0 | 12.2 | 9.6 | 0.71 |
5Method notes and limitations
- Clock. BTS times are local. Every leg is converted to UTC with its airport's IANA zone (first occurrence in the repeated fall-back hour); the arrival date is resolved against the published block time, which handles overnight, date-line and DST legs (tested on hand-made cases).
- Time split. Parameters, hubs, LP buffers and heuristic rankings use Aug 2025 to Apr 2026 only; every performance number is on May 2026 to Jul 2026.
- Contagion is observational. The matched comparison removes carrier-wide day-level conditions and exposure, not aircraft-level confounders; the robustness table shows how far the estimate moves under stricter matching.
- Buffer model. A buffer minute is modelled as slack only: it does not change primary delays, crew legality, gate conflicts or the commercial value of a departure time. The LP is a ranking device for where slack pays, not a full schedule re-optimisation; a budget-neutral re-allocation of existing slack would be needed to gain clock time.
- Uncertainty. CIs on held-out policy results resample whole weeks; hinge and contagion CIs resample days. All are conditional on the model.
- Coverage. Chains break at cancellations, diversions, tail swaps and ground times over 5 hours. Hawaiian reports under Alaska from January 2026 (no held-out turns; 13 carriers are scored) and Spirit stops reporting in May 2026 (207 held-out turns).