Portrait of Pierre Chambet

Pierre Chambet

Decisions for operations that run on uncertain data.

I forecast honestly, simulate what can go wrong and optimise what you control, then ship it as software people use. Airline fleets, freight, power grids. Founder of SmartFleetOptim, previously on the Air France Network team, engineer from Télécom SudParis.

The hardest week of the energy crisis, forecast a day ahead

French electricity demand, 10–16 December 2022

51%of hours inside the 80% interval that week
86%inside the 95% interval
80.0% and 95.0%over 38,732 hours, 2022 to 2026
2.9%mean absolute error that week (RTE: 1.9%)

Actual demand against the intervals gridcast issued at noon the day before, scored out of sample. Hours outside the 80% interval. The bands hold on average, not in every week: measuring that gap, and how fast recalibration closes it, is the point of the project. Full report

Selected work

  1. Where should an airline add turnaround buffer to stop delays spreading?

    On 7 million US flights, 40% of cause-coded delay minutes are reactionary: the aircraft came in late. A turn absorbs delay up to a threshold, then passes it on almost minute for minute. Placed by a scenario linear program, a buffer minute avoids 1.7 times as much delay as uniform padding on held-out days; padding the most-delayed turns does worse than uniform.

    dbt and DuckDB over BTS data, time zones done right, a hinge propagation model, matched contagion estimates, an LP solved with HiGHS.

    CodeReport

    Delay avoided against buffer minutes: the LP-optimised line rises fastest, uniform padding and the greedy rule below
  2. Can a demand forecast publish uncertainty bands that stay honest through a crisis?

    On average, yes: over 38,732 out-of-sample hours the 80% and 95% intervals cover 80.0% and 95.0% of French demand. Month to month only partly: a frozen interval is on target 12% of the time, rolling recalibration 50%. Compared fairly, RTE's own point forecast is more accurate, and the report says so.

    LightGBM quantiles, conformalized quantile regression with adaptive conformal inference, a strict day-ahead information set, a scheduled GitHub Actions job.

    CodeReport

    Trailing 30-day interval coverage from 2022 to 2026 for three calibration methods, with the energy crisis shaded
  3. Who should get the marketing e-mail?

    On a 64,000-customer randomised test, the honest answer is everyone: blanket sending earns $155 per 1,000 customers, and the best uplift procedure chosen out of sample falls short of it in 48 of 50 splits. Scored on its own data, the same contest would have promised a gain. At 14 million users the picture changes, and the analysis shows why.

    Experiment readout with CUPED and Holm, S-, T-, X- and doubly robust learners, off-policy value with confidence intervals, Criteo at scale in DuckDB.

    CodeReport

    Incremental profit against share of customers e-mailed for uplift, response and random policies, and Criteo uplift curves
  4. How much spare-parts stock buys a 95% fill rate when demand is mostly zeros?

    On 1,046 car-part series, the method with the best forecast error (SES) needs 39% more stock than a LightGBM quantile model to serve 95% of demand, out of sample (95% CI 29–51%). Pinball loss ranks methods the way the shelf does; MASE does not.

    Croston, SBA, TSB, bootstrap and global LightGBM; order-up-to simulation; Sherbrooke-style budget allocation.

    CodeReport

    Fill rate against average stock for each forecasting method: LightGBM reaches 95% with 9.2 months of stock, SES needs 12.8
  5. Will a corrosion-risk model hold on aircraft it has never seen?

    Adversarial validation showed the test fleet was a different population (AUC 0.86). Validating on test-like aircraft set the one parameter that mattered, a shrinkage of α = 0.72 that grouped cross-validation said was useless. 22nd on the public leaderboard, 6th on the private one, 2nd overall at the Airbus × IBM × AWS hackathon.

    Machine learning lead. LightGBM, grouped and adversarial validation, calibration, paired bootstrap.

    CodeReport

    Brier score against shrinkage: in-distribution validation prefers no shrinkage, test-like aircraft prefer alpha 0.72
  6. How many reserve aircraft should an airline hold, of which type, at which base?

    SmartFleetOptim stress-tests spare-fleet plans across disruption, demand and delivery-delay scenarios, and compares them on tail risk rather than on a single forecast. I ran technical sessions with the operations and data teams of Finnair, SWISS and Air Canada on their own public fleet data.

    Founder and lead engineer. Python engine, FastAPI, React, deployed on a cloud VM. Profiling cut an engine run from 13.6 s to 1.3 s.

    smartfleetoptim.comCode is private

More projects

  • FreightSightLanded-cost allocation for importers: a two-pass Decimal engine, Postgres row-level security, an Odoo connector.
  • Three-ERP consolidationOdoo, an accounting FEC file and a warehouse API reconciled in DuckDB; 22 planted discrepancies, all found, none extra, queried through an MCP server.
  • Fiduciary agentSwiss QR-bills and camt.053 statements: code builds every entry, an agent classifies only what no rule covers, a person approves.
  • HeliosWalk-forward battery arbitrage on French day-ahead prices with model predictive control and Wasserstein-robust optimisation.
  • Differentiable particle filterOptimal-transport resampling in PyTorch, checked against exact Kalman gradients.
  • Facebook100 social structureHomophily, link prediction and communities on all 100 campus graphs.
  • Gait at five walking speeds2,544 knee cycles from 52 adults: re-identification, variance components and grouped cross-validation.
  • Functional data clusteringClustering curves and covariates together, in R; what silhouette tuning costs in ARI.
  • Hidden Markov models from scratchDo hidden states earn their parameters? Rain spells and unseen-speaker speech, judged on held-out data.
  • Mushroom toxicityWhat a multiple correspondence analysis keeps and loses, with honest shuffled cross-validation.
  • From bag-of-words to BERTEvery rung of the NLP ladder on identical splits, with label budgets.
  • MLP against CNN on MNISTPaired tests, calibration, selective prediction and shift robustness.

Experience

  1. SmartFleetOptim, founder and lead engineer

    Airline spare-fleet strategy under uncertainty, from discovery calls with airlines to a production deployment. Supported by IMT Starter.

  2. Airbus × IBM × AWS hackathon, machine learning lead

    Corrosion-risk model. Adversarial validation exposed a train/test shift; calibrated probabilities moved the team from 22nd on the public leaderboard to 6th on the private one, 2nd overall.

  3. NAIST, Japan, visiting research intern

    Self-directed review of optimal transport for deep learning; workshop and ICLR paper review for a computational systems biology lab.

  4. Air France Network, data analyst, program regulation

    Built the Python decision-support tool and Power BI dashboard that Network management used to allocate spare aircraft across long- and medium-haul operations.

  5. Télécom SudParis, Diplôme d’Ingénieur

    Data analysis and pattern classification. MSc-equivalent engineering degree.

Open source

Small fixes, merged upstream, in libraries I use.

Teaching

Deep Learning from Scratch rebuilds neural networks from a single NumPy neuron to convolutional networks, every gradient derived by hand, published as an English series on LinkedIn with notebooks and PDF guides.

My optimal transport notes from NAIST go from Monge and Kantorovich to Wasserstein generative models.