Airbus × IBM × AWS hackathon 2026 · aircraft corrosion risk

When the test fleet is not your training fleet

A corrosion-risk model built for a Kaggle-scored hackathon, and the validation work that decided its single most important parameter. Reported result: 6th on the Brier leaderboard, 2nd overall (model + business pitch), climbing from 22nd on the public leaderboard to 6th on the private one. Both boards are small, so that move is consistent with the approach rather than proof of it (section 5). Team project: the event-time code is in the team repository; this is a post-event rebuild and re-analysis of the pipeline I owned as ML lead.

0.86adversarial AUC: test aircraft are easy to tell apart
0.140 → 0.165 → 0.204raw-model Brier: CV overall → CV on test-like aircraft → test-like aircraft left out of training
α = 0.72Brier-optimal shrinkage on test-like aircraft (CV says 1.00)
±0.03195% margin of error of a 143-row public score

1. The problem, stated precisely

Each of 758 training aircraft has a monthly environment history (weather, aerosols, gases, parking time) and one corrosion observation date T. The leaderboard scores only two months per aircraft: T (label 1) and T − 24 months (label 0). That gives 1270 labelled rows (616 positives, 654 negatives) and turns the task into a balanced, within-aircraft question: is this aircraft more exposed now than two years ago? Anything constant per aircraft cancels out. The model is a LightGBM regressor on 0/1 labels (squared error is the Brier score) over 66 strictly causal exposure features.

2. The test fleet is a different population

A classifier trained to tell training rows from test rows, with folds split by aircraft, reaches AUC 0.86. The clearest symptom: 63% of test aircraft have data from the first year of the record (2014), against 0.3% of training aircraft. Their histories are left-censored, so "age" means something else for them.

3. A validation split that looks like the test set

Grouped cross-validation answers "how good is the model on aircraft like the training ones?". To answer the question that matters, the training aircraft were ranked by how test-like the adversarial model finds them; the model is fitted on the 384 least test-like and scored on the 315 most test-like (558 rows). Uncertainty comes from a bootstrap over aircraft, not rows.

The gap between the two validations has two parts. Grouped CV scores 0.140 overall but already 0.165 on the test-like aircraft (against 0.120 on the others): these aircraft are harder in themselves. Leaving them out of training, as the test set does, raises it to 0.204. Part of that last step is the smaller training set (384 aircraft instead of about 559 per CV fold).

The two validations disagree on the one decision that matters. In-distribution, shrinking predictions toward 0.5 only hurts (best α = 1.00). On test-like aircraft the raw model is over-confident and the Brier-optimal shrinkage is α = 0.72. The submission used α = 0.7; on the holdout that improves the Brier by 0.0077 (paired 95% CI 0.0019 to 0.0138). That gain is in-sample for α. With α cross-fitted on random halves of the holdout aircraft (fitted α from 0.61 to 0.84), the gain is 0.0071 (95% CI 0.0018 to 0.0127).

Diamonds are the scores reported on the Kaggle leaderboard; they cannot be recomputed without the hidden labels. The holdout estimate at α = 0.7 is 0.197 (95% CI 0.182–0.211); in-distribution CV would have promised 0.140.

4. What else was tried, judged on the holdout

All candidates share features and splits; differences are paired (same aircraft), negative is better.

ModelBrier at α = 0.795% CIΔ vs submittedAUC
Constant 0.50.2500[0.2500, 0.2500]+0.0532 [+0.0390, +0.0676]0.500
LightGBM squared error (submitted)0.1968[0.1824, 0.2110]0.766
LightGBM log-loss + Platt scaling0.1922[0.1790, 0.2063]-0.0046 [-0.0096, +0.0005]0.774
LightGBM squared error, 500 trees depth 70.2018[0.1867, 0.2171]+0.0050 [+0.0014, +0.0089]0.750
Logistic regression0.2089[0.1953, 0.2240]+0.0121 [+0.0011, +0.0235]0.737
LightGBM on the two clock features0.2045[0.1887, 0.2204]+0.0077 [-0.0067, +0.0226]0.740

Two honest readings. First, a LightGBM that sees only the two clock features (months observed, age) captures most of the full model's gain over a constant forecast on test-like aircraft (point estimate 86%; its paired difference to the full model, +0.0077 [-0.0067, +0.0226], is not significant); the environmental features add little that transfers. Second, the choice of objective (squared error versus log-loss with Platt scaling) is within noise on the holdout.

LeverBrier at α = 0.795% CIΔ vs submittedAUC
Full feature set (submitted)0.1968[0.1824, 0.2110]0.766
Drop the 5 most shift-revealing features0.1984[0.1843, 0.2130]+0.0016 [-0.0014, +0.0046]0.760
Drop the 10 most shift-revealing features0.1932[0.1793, 0.2078]-0.0036 [-0.0077, +0.0005]0.773
Drop the 15 most shift-revealing features0.1925[0.1788, 0.2070]-0.0044 [-0.0091, +0.0004]0.774
10-model ensemble (5 seeds x 2 objectives)0.1929[0.1796, 0.2069]-0.0039 [-0.0067, -0.0010]0.772

Of the 4 levers against the shift, only one has a paired interval that excludes zero, and by a small margin (-0.0039): 10-model ensemble (5 seeds x 2 objectives). With 4 levers tried, treat it as suggestive; it was not part of the submission.

5. The leaderboards were too small to rank on

With 143 rows, a single public Brier score carries a 95% margin of error of ±0.031, treating rows as independent. Two independent scores would need to differ by more than 0.043. Submissions scored on the same rows are correlated, so a paired comparison is tighter: for two similar models (submitted versus log-loss + Platt) the threshold is about 0.011. Public rows probably include T / T − 24 pairs from the same aircraft, which these row-level formulas ignore. Either way, small public gaps carried little information, which is why the submission was chosen on the holdout rather than on public feedback.

The private board is probably no better. Its row count is not known here, but if only T and T − 24 rows are scored, the 142 test aircraft give at most 284 scored rows, leaving at most about 141 for the private board: the same margin of error. The move from 22nd to 6th is consistent with choosing α on a shift-aware holdout, but it does not prove the approach better than the teams ranked nearby.

6. Leak audit

No artefact betrays the corrosion month: missingness at T differs from other months by at most 0.06 percentage points; duplicate rows are about as frequent at T (7.6%) as elsewhere (6.0%); T is the last observed month for only 1.8% of aircraft (median gap 23 months).

Limitations