Pchambet / uplift-targeting

Whom should a retailer e-mail? On this experiment, everyone.

A retailer e-mailed 64,000 customers at random: a men's campaign, a women's campaign, or nothing. A campaign team wants to know whom to e-mail, and with which e-mail, so that the campaign earns money it would not have earned anyway. This page answers with out-of-sample, uncertainty-aware estimates, shows why a model contest scored on its own data would have promised more, and checks the method at scale on 13,979,592 advertising users.

At the default economics, no targeting rule beats e-mailing everyone

Incremental profit per 1,000 customers against the share e-mailed: doubly robust estimates; dotted lines are the models' 95% CIs, the grey band random targeting's. The uplift policy picks the e-mail per customer, which is why it ends below the blanket point. The uplift ranking is chosen on one half of the customers and scored on the other. Markers: the three procedures a team could ship, each choosing its own budget on held-out data. Move the sliders to re-price every curve.

Decision memo

  1. Send the men's e-mail. It lifts conversion by +0.68 pp, 95% CI [0.50, 0.86], and spend by +$0.77 per customer, about twice the women's e-mail (+$0.42).
  2. Send it to everyone while an e-mail costs less than about 31 cents all-in (the margin times its spend lift); above that, e-mail no one. Between 0 and 60 cents per e-mail, at no cost level does a targeted procedure beat the better of the two with a 95% CI that excludes zero. At $0.15, blanket sending earns $155 per 1,000 customers, 95% CI [$41, $269]; the best uplift procedure earns $114, a paired difference of -$41, CI [-$121, +$38]. Over 50 random splits it averages -$41 and falls short of blanket in 48 of 50.
  3. Do not trust an in-sample model contest. Choosing among 31 rankings and their budgets on the evaluation data would have recommended personalising which e-mail each customer gets, reporting +$27 per 1,000 over blanket sending. Chosen on one half and scored on the other, the same idea is -$27, 95% CI [-$104, +$49].
  4. Do not assume an uplift model beats a response model: test it on the experiment. On Hillstrom visits the DR-learner does not detectably beat the response model for the women's e-mail and loses to it for the men's e-mail. On 13,979,592 Criteo users the response model's top 20% holds 87% of all incremental conversions, against 73% for the DR-learner.

ADid the experiment work, and what did it do on average?

The randomisation holds: arm sizes match the 1/3 design (chi-square p = 0.90) and every covariate is balanced (largest absolute standardised difference 0.014, against a usual alarm level of 0.1). Both e-mails work on average.

+7.7 ppvisit rate, men's e-mail
+0.68 ppconversion, men's e-mail (+119% relative)
+$0.77spend per customer, men's e-mail
+$0.42spend per customer, women's e-mail

Average effects, three estimators

Variance reduction barely helps here. CUPED on past-year spend narrows the spend interval by 0.02%; the fully interacted regression adjustment of Lin (2013) narrows the visit interval by 1.5%. A year of purchase history explains almost nothing of a two-week window, so the plain difference in means is the honest headline.

EstimatorEffect95% CISE ratio
Mens e-mail · visit
Difference in means7.66 pp[7.00, 8.32]1.0000
CUPED (history)7.64 pp[6.98, 8.30]0.9978
Regression adjustment (Lin)7.61 pp[6.96, 8.26]0.9853
Mens e-mail · conversion
Difference in means0.68 pp[0.50, 0.86]1.0000
CUPED (history)0.68 pp[0.50, 0.86]0.9996
Regression adjustment (Lin)0.68 pp[0.50, 0.86]0.9990
Mens e-mail · spend
Difference in means$0.770[$0.485, $1.055]1.0000
CUPED (history)$0.767[$0.483, $1.052]0.9998
Regression adjustment (Lin)$0.767[$0.482, $1.051]0.9988
Womens e-mail · visit
Difference in means4.52 pp[3.89, 5.16]1.0000
CUPED (history)4.51 pp[3.88, 5.14]0.9980
Regression adjustment (Lin)4.54 pp[3.92, 5.17]0.9857
Womens e-mail · conversion
Difference in means0.31 pp[0.15, 0.47]1.0000
CUPED (history)0.31 pp[0.15, 0.47]0.9997
Regression adjustment (Lin)0.31 pp[0.15, 0.47]0.9992
Womens e-mail · spend
Difference in means$0.424[$0.169, $0.680]1.0000
CUPED (history)$0.423[$0.167, $0.678]0.9998
Regression adjustment (Lin)$0.426[$0.170, $0.681]0.9994

Pre-declared segments

Five segment dimensions were declared before any model was fitted and tested once per e-mail and outcome: 30 heterogeneity tests, Holm-corrected. 2 of 30 survive, all on visits: the men's e-mail lifts the visit rate by +13.4 pp among buyers of both lines, against +6.9 pp to +7.1 pp for the other product-line groups; the women's e-mail lifts the visit rate by +1.1 pp among menswear-only buyers, against +7.1 pp to +7.4 pp for the other product-line groups. Purchases and spend show no detectable heterogeneity at this sample size.

Difference in rate (e-mail minus control) within each segment, 95% CI.

BCan a model find the persuadable customers?

Five meta-learners (S, T, X, cross-fitted DR, transformed outcome), each with a LightGBM and a regularised linear base model, trained with five-fold cross-fitting so that every customer is scored by models that never saw them, and benchmarked against the response model a marketing team would build.

Uplift curves from out-of-fold LightGBM scores: height at x is the incremental outcome per 1,000 customers if the top x% by the model were e-mailed. Dashed: random targeting.

Where the effect is decoupled from the baseline, as with the women's e-mail, the uplift learners find it: the DR-learner's Qini coefficient on visits is 2.7 [1.7, 3.6] (×1,000, bootstrap 95% CI), clearly above random. The plain response model finds much of it too (2.0 [1.1, 2.9]): paired on the same bootstrap draws, the DR-learner does not detectably beat it (+0.7 [-0.2, +1.7]), although 9 of 10 uplift learners score above it on point estimate. Where the effect grows with the baseline, as with the men's e-mail, the DR-learner is no better than random (0.1 [-0.8, 1.1]), loses to the response model (-1.6 [-2.7, -0.6]), and only 1 of 10 uplift learners score above the response model. On spend, 0 of 22 rankings have a Qini interval that excludes zero, and all 11 point estimates for the men's e-mail are negative.

Qini coefficients of every model (×1,000, bootstrap 95% CI)
E-mailOutcomeLearnerBase modelQini95% CI
Mens e-mailvisitS-learnerregularised linear1.740.88 to 2.78
Mens e-mailvisitResponse modelLightGBM1.670.80 to 2.64
Mens e-mailvisitS-learnerLightGBM0.66-0.20 to 1.61
Mens e-mailvisitTransformed outcomeregularised linear0.63-0.29 to 1.64
Mens e-mailvisitT-learnerregularised linear0.54-0.34 to 1.63
Mens e-mailvisitT-learnerLightGBM0.46-0.47 to 1.39
Mens e-mailvisitX-learnerLightGBM0.42-0.54 to 1.44
Mens e-mailvisitX-learnerregularised linear0.35-0.58 to 1.33
Mens e-mailvisitDR-learnerregularised linear0.31-0.63 to 1.28
Mens e-mailvisitTransformed outcomeLightGBM0.22-0.69 to 1.25
Mens e-mailvisitDR-learnerLightGBM0.06-0.77 to 1.13
Womens e-mailvisitT-learnerregularised linear3.722.82 to 4.64
Womens e-mailvisitS-learnerLightGBM3.622.64 to 4.57
Womens e-mailvisitTransformed outcomeregularised linear3.562.72 to 4.42
Womens e-mailvisitDR-learnerregularised linear3.542.71 to 4.37
Womens e-mailvisitX-learnerregularised linear3.542.70 to 4.37
Womens e-mailvisitX-learnerLightGBM2.921.94 to 3.80
Womens e-mailvisitT-learnerLightGBM2.891.94 to 3.81
Womens e-mailvisitDR-learnerLightGBM2.731.72 to 3.61
Womens e-mailvisitTransformed outcomeLightGBM2.611.64 to 3.56
Womens e-mailvisitResponse modelLightGBM1.981.11 to 2.93
Womens e-mailvisitS-learnerregularised linear0.950.08 to 1.85
Mens e-mailconversionS-learnerregularised linear0.19-0.14 to 0.48
Mens e-mailconversionT-learnerLightGBM0.07-0.22 to 0.34
Mens e-mailconversionResponse modelLightGBM0.04-0.24 to 0.32
Mens e-mailconversionX-learnerLightGBM0.04-0.26 to 0.33
Mens e-mailconversionTransformed outcomeLightGBM0.04-0.26 to 0.32
Mens e-mailconversionDR-learnerLightGBM0.02-0.28 to 0.30
Mens e-mailconversionX-learnerregularised linear-0.07-0.36 to 0.16
Mens e-mailconversionTransformed outcomeregularised linear-0.09-0.37 to 0.14
Mens e-mailconversionDR-learnerregularised linear-0.11-0.40 to 0.13
Mens e-mailconversionS-learnerLightGBM-0.16-0.45 to 0.12
Mens e-mailconversionT-learnerregularised linear-0.18-0.49 to 0.07
Womens e-mailconversionS-learnerLightGBM0.400.13 to 0.68
Womens e-mailconversionT-learnerregularised linear0.330.09 to 0.60
Womens e-mailconversionX-learnerregularised linear0.300.07 to 0.58
Womens e-mailconversionDR-learnerregularised linear0.300.06 to 0.57
Womens e-mailconversionTransformed outcomeregularised linear0.300.05 to 0.57
Womens e-mailconversionResponse modelLightGBM0.270.02 to 0.54
Womens e-mailconversionT-learnerLightGBM0.26-0.00 to 0.51
Womens e-mailconversionX-learnerLightGBM0.21-0.06 to 0.47
Womens e-mailconversionTransformed outcomeLightGBM0.16-0.13 to 0.41
Womens e-mailconversionDR-learnerLightGBM0.15-0.13 to 0.41
Womens e-mailconversionS-learnerregularised linear0.05-0.20 to 0.29
Mens e-mailspendTransformed outcomeregularised linear-2.72-49.72 to 44.90
Mens e-mailspendDR-learnerregularised linear-8.09-55.30 to 40.00
Mens e-mailspendX-learnerregularised linear-8.25-55.58 to 40.10
Mens e-mailspendT-learnerregularised linear-8.31-55.63 to 40.04
Mens e-mailspendResponse modelLightGBM-10.45-47.83 to 27.60
Mens e-mailspendDR-learnerLightGBM-11.38-55.82 to 30.37
Mens e-mailspendT-learnerLightGBM-14.56-55.26 to 24.79
Mens e-mailspendX-learnerLightGBM-15.01-55.39 to 21.13
Mens e-mailspendTransformed outcomeLightGBM-15.81-62.22 to 24.94
Mens e-mailspendS-learnerLightGBM-23.20-66.64 to 15.08
Mens e-mailspendS-learnerregularised linear-42.39-82.01 to 1.70
Womens e-mailspendTransformed outcomeregularised linear30.64-10.68 to 67.32
Womens e-mailspendDR-learnerregularised linear29.67-9.70 to 66.09
Womens e-mailspendX-learnerregularised linear29.53-9.71 to 66.12
Womens e-mailspendT-learnerregularised linear29.51-9.74 to 66.11
Womens e-mailspendS-learnerLightGBM24.32-14.12 to 68.76
Womens e-mailspendDR-learnerLightGBM4.99-32.92 to 49.04
Womens e-mailspendResponse modelLightGBM0.82-36.54 to 42.78
Womens e-mailspendTransformed outcomeLightGBM-5.11-44.83 to 39.64
Womens e-mailspendX-learnerLightGBM-13.13-56.09 to 29.09
Womens e-mailspendT-learnerLightGBM-16.31-61.53 to 28.82
Womens e-mailspendS-learnerregularised linear-17.05-55.09 to 18.15

Where the response model wastes contacts

For the women's e-mail, the response model ranks customers by how likely they are to visit, so its top deciles are full of people who would have come anyway. In its top 20%, 71% of the e-mailed visitors would have visited without the e-mail, against 58% in the top 20% of the uplift ranking (T-learner, LightGBM), which buys 8.4 pp of extra visit rate there against 6.8 pp. Reading the control rate as "would have come anyway" assumes the e-mail deters no one.

Women's e-mail; uplift ranking from the T-learner (LightGBM). Grey: visit rate of control customers in the decile. Colour: incremental visit rate caused by the e-mail, 95% CI. Design-based: the models only decide the ranking.

CWhich policy should ship?

The decision is scored on spend, because profit depends on spend. Every policy value is a doubly robust off-policy estimate on customers that neither trained the models nor chose the policy, with a paired standard error for every comparison.

$155per 1,000 customers, e-mail everyone [$41, $269]
$114uplift procedure, e-mails 85% [$9, $219]
$142response procedure, e-mails 93% [$31, $254]
$182what an in-sample model contest would report

30 candidate rankings compete: five learners and two base models, trained on visits, conversions or spend. A model trained on dense visits might rank customers better for revenue than one trained on sparse revenue, so the choice is left to data, but to data the final score never sees. The candidates' merit barely replicates from one random half of the customers to the other: rank correlation 0.24 on spend, against 0.66 on visits. With 578 buyers, the best-looking model is mostly the luckiest one.

Each dot is a candidate ranking, scored on two random halves (average incremental spend per 1,000 customers over budgets). Dotted: equal value on both halves. Half B happens to show a larger spend effect, so the dots sit above the line; what matters is whether their order agrees, and it barely does.

The table values come from one random split of the customers, and their intervals are conditional on the rankings and budgets that split selected. Re-running the whole procedure on 50 random splits, the uplift procedure averages -$41 per 1,000 against blanket sending (from -$89 to +$9), falls short of it in 48 of 50 splits, e-mails a median 80% of customers (52% to 100%), and settles on 16 different rankings along the way.

DWhat changes when the experiment is two hundred times larger?

The Criteo Uplift v2.1 benchmark logs 13,979,592 users from randomised advertising incrementality tests, 85% treated. The same learners were fitted on a hash-based sample of 2.3M users from one half and evaluated on all 6,994,996 users of the other half. Treatment lifts visits by +1.03 pp and conversions by +0.115 pp on average.

Incremental outcomes per 1,000 users when the top x% by each model is treated, pointwise 95% bands. Dashed: random targeting.

At this size every ranking beats random by a wide margin, and the effect is concentrated: the top 10% of users by the response model hold 51% of all incremental visits. Their untreated visit rate is 32% and treatment adds +5.3 pp: here the effect tracks the baseline rate, so ranking by "who will act" is also ranking by "who will be moved". The response model is the best uplift ranker (Qini ×1,000 on conversions: 0.43, against 0.28 for the T-learner and 0.28 for the DR-learner). The meta-learners pay for estimating a small difference between two large, noisy predictions; with 15% controls and a control conversion rate of 0.19%, the DR-learner's final stage returns the same score for 63% of users. The 0.7% of users it scores below zero convert at 5.6% without treatment, against 0.19% overall, and gain +2.0 pp when treated: high-baseline users whose noisy pseudo-outcomes pushed their scores below zero. They hold 13% of all incremental conversions, which is why its curve jumps in the last 1%.

Method and limits