Whom should a retailer e-mail? On this experiment, everyone.
A retailer e-mailed 64,000 customers at random: a men's campaign, a women's campaign, or nothing. A campaign team wants to know whom to e-mail, and with which e-mail, so that the campaign earns money it would not have earned anyway. This page answers with out-of-sample, uncertainty-aware estimates, shows why a model contest scored on its own data would have promised more, and checks the method at scale on 13,979,592 advertising users.
At the default economics, no targeting rule beats e-mailing everyone
Decision memo
- Send the men's e-mail. It lifts conversion by +0.68 pp, 95% CI [0.50, 0.86], and spend by +$0.77 per customer, about twice the women's e-mail (+$0.42).
- Send it to everyone while an e-mail costs less than about 31 cents all-in (the margin times its spend lift); above that, e-mail no one. Between 0 and 60 cents per e-mail, at no cost level does a targeted procedure beat the better of the two with a 95% CI that excludes zero. At $0.15, blanket sending earns $155 per 1,000 customers, 95% CI [$41, $269]; the best uplift procedure earns $114, a paired difference of -$41, CI [-$121, +$38]. Over 50 random splits it averages -$41 and falls short of blanket in 48 of 50.
- Do not trust an in-sample model contest. Choosing among 31 rankings and their budgets on the evaluation data would have recommended personalising which e-mail each customer gets, reporting +$27 per 1,000 over blanket sending. Chosen on one half and scored on the other, the same idea is -$27, 95% CI [-$104, +$49].
- Do not assume an uplift model beats a response model: test it on the experiment. On Hillstrom visits the DR-learner does not detectably beat the response model for the women's e-mail and loses to it for the men's e-mail. On 13,979,592 Criteo users the response model's top 20% holds 87% of all incremental conversions, against 73% for the DR-learner.
ADid the experiment work, and what did it do on average?
The randomisation holds: arm sizes match the 1/3 design (chi-square p = 0.90) and every covariate is balanced (largest absolute standardised difference 0.014, against a usual alarm level of 0.1). Both e-mails work on average.
Average effects, three estimators
Variance reduction barely helps here. CUPED on past-year spend narrows the spend interval by 0.02%; the fully interacted regression adjustment of Lin (2013) narrows the visit interval by 1.5%. A year of purchase history explains almost nothing of a two-week window, so the plain difference in means is the honest headline.
| Estimator | Effect | 95% CI | SE ratio |
|---|---|---|---|
| Mens e-mail · visit | |||
| Difference in means | 7.66 pp | [7.00, 8.32] | 1.0000 |
| CUPED (history) | 7.64 pp | [6.98, 8.30] | 0.9978 |
| Regression adjustment (Lin) | 7.61 pp | [6.96, 8.26] | 0.9853 |
| Mens e-mail · conversion | |||
| Difference in means | 0.68 pp | [0.50, 0.86] | 1.0000 |
| CUPED (history) | 0.68 pp | [0.50, 0.86] | 0.9996 |
| Regression adjustment (Lin) | 0.68 pp | [0.50, 0.86] | 0.9990 |
| Mens e-mail · spend | |||
| Difference in means | $0.770 | [$0.485, $1.055] | 1.0000 |
| CUPED (history) | $0.767 | [$0.483, $1.052] | 0.9998 |
| Regression adjustment (Lin) | $0.767 | [$0.482, $1.051] | 0.9988 |
| Womens e-mail · visit | |||
| Difference in means | 4.52 pp | [3.89, 5.16] | 1.0000 |
| CUPED (history) | 4.51 pp | [3.88, 5.14] | 0.9980 |
| Regression adjustment (Lin) | 4.54 pp | [3.92, 5.17] | 0.9857 |
| Womens e-mail · conversion | |||
| Difference in means | 0.31 pp | [0.15, 0.47] | 1.0000 |
| CUPED (history) | 0.31 pp | [0.15, 0.47] | 0.9997 |
| Regression adjustment (Lin) | 0.31 pp | [0.15, 0.47] | 0.9992 |
| Womens e-mail · spend | |||
| Difference in means | $0.424 | [$0.169, $0.680] | 1.0000 |
| CUPED (history) | $0.423 | [$0.167, $0.678] | 0.9998 |
| Regression adjustment (Lin) | $0.426 | [$0.170, $0.681] | 0.9994 |
Pre-declared segments
Five segment dimensions were declared before any model was fitted and tested once per e-mail and outcome: 30 heterogeneity tests, Holm-corrected. 2 of 30 survive, all on visits: the men's e-mail lifts the visit rate by +13.4 pp among buyers of both lines, against +6.9 pp to +7.1 pp for the other product-line groups; the women's e-mail lifts the visit rate by +1.1 pp among menswear-only buyers, against +7.1 pp to +7.4 pp for the other product-line groups. Purchases and spend show no detectable heterogeneity at this sample size.
BCan a model find the persuadable customers?
Five meta-learners (S, T, X, cross-fitted DR, transformed outcome), each with a LightGBM and a regularised linear base model, trained with five-fold cross-fitting so that every customer is scored by models that never saw them, and benchmarked against the response model a marketing team would build.
Where the effect is decoupled from the baseline, as with the women's e-mail, the uplift learners find it: the DR-learner's Qini coefficient on visits is 2.7 [1.7, 3.6] (×1,000, bootstrap 95% CI), clearly above random. The plain response model finds much of it too (2.0 [1.1, 2.9]): paired on the same bootstrap draws, the DR-learner does not detectably beat it (+0.7 [-0.2, +1.7]), although 9 of 10 uplift learners score above it on point estimate. Where the effect grows with the baseline, as with the men's e-mail, the DR-learner is no better than random (0.1 [-0.8, 1.1]), loses to the response model (-1.6 [-2.7, -0.6]), and only 1 of 10 uplift learners score above the response model. On spend, 0 of 22 rankings have a Qini interval that excludes zero, and all 11 point estimates for the men's e-mail are negative.
Qini coefficients of every model (×1,000, bootstrap 95% CI)
| Outcome | Learner | Base model | Qini | 95% CI | |
|---|---|---|---|---|---|
| Mens e-mail | visit | S-learner | regularised linear | 1.74 | 0.88 to 2.78 |
| Mens e-mail | visit | Response model | LightGBM | 1.67 | 0.80 to 2.64 |
| Mens e-mail | visit | S-learner | LightGBM | 0.66 | -0.20 to 1.61 |
| Mens e-mail | visit | Transformed outcome | regularised linear | 0.63 | -0.29 to 1.64 |
| Mens e-mail | visit | T-learner | regularised linear | 0.54 | -0.34 to 1.63 |
| Mens e-mail | visit | T-learner | LightGBM | 0.46 | -0.47 to 1.39 |
| Mens e-mail | visit | X-learner | LightGBM | 0.42 | -0.54 to 1.44 |
| Mens e-mail | visit | X-learner | regularised linear | 0.35 | -0.58 to 1.33 |
| Mens e-mail | visit | DR-learner | regularised linear | 0.31 | -0.63 to 1.28 |
| Mens e-mail | visit | Transformed outcome | LightGBM | 0.22 | -0.69 to 1.25 |
| Mens e-mail | visit | DR-learner | LightGBM | 0.06 | -0.77 to 1.13 |
| Womens e-mail | visit | T-learner | regularised linear | 3.72 | 2.82 to 4.64 |
| Womens e-mail | visit | S-learner | LightGBM | 3.62 | 2.64 to 4.57 |
| Womens e-mail | visit | Transformed outcome | regularised linear | 3.56 | 2.72 to 4.42 |
| Womens e-mail | visit | DR-learner | regularised linear | 3.54 | 2.71 to 4.37 |
| Womens e-mail | visit | X-learner | regularised linear | 3.54 | 2.70 to 4.37 |
| Womens e-mail | visit | X-learner | LightGBM | 2.92 | 1.94 to 3.80 |
| Womens e-mail | visit | T-learner | LightGBM | 2.89 | 1.94 to 3.81 |
| Womens e-mail | visit | DR-learner | LightGBM | 2.73 | 1.72 to 3.61 |
| Womens e-mail | visit | Transformed outcome | LightGBM | 2.61 | 1.64 to 3.56 |
| Womens e-mail | visit | Response model | LightGBM | 1.98 | 1.11 to 2.93 |
| Womens e-mail | visit | S-learner | regularised linear | 0.95 | 0.08 to 1.85 |
| Mens e-mail | conversion | S-learner | regularised linear | 0.19 | -0.14 to 0.48 |
| Mens e-mail | conversion | T-learner | LightGBM | 0.07 | -0.22 to 0.34 |
| Mens e-mail | conversion | Response model | LightGBM | 0.04 | -0.24 to 0.32 |
| Mens e-mail | conversion | X-learner | LightGBM | 0.04 | -0.26 to 0.33 |
| Mens e-mail | conversion | Transformed outcome | LightGBM | 0.04 | -0.26 to 0.32 |
| Mens e-mail | conversion | DR-learner | LightGBM | 0.02 | -0.28 to 0.30 |
| Mens e-mail | conversion | X-learner | regularised linear | -0.07 | -0.36 to 0.16 |
| Mens e-mail | conversion | Transformed outcome | regularised linear | -0.09 | -0.37 to 0.14 |
| Mens e-mail | conversion | DR-learner | regularised linear | -0.11 | -0.40 to 0.13 |
| Mens e-mail | conversion | S-learner | LightGBM | -0.16 | -0.45 to 0.12 |
| Mens e-mail | conversion | T-learner | regularised linear | -0.18 | -0.49 to 0.07 |
| Womens e-mail | conversion | S-learner | LightGBM | 0.40 | 0.13 to 0.68 |
| Womens e-mail | conversion | T-learner | regularised linear | 0.33 | 0.09 to 0.60 |
| Womens e-mail | conversion | X-learner | regularised linear | 0.30 | 0.07 to 0.58 |
| Womens e-mail | conversion | DR-learner | regularised linear | 0.30 | 0.06 to 0.57 |
| Womens e-mail | conversion | Transformed outcome | regularised linear | 0.30 | 0.05 to 0.57 |
| Womens e-mail | conversion | Response model | LightGBM | 0.27 | 0.02 to 0.54 |
| Womens e-mail | conversion | T-learner | LightGBM | 0.26 | -0.00 to 0.51 |
| Womens e-mail | conversion | X-learner | LightGBM | 0.21 | -0.06 to 0.47 |
| Womens e-mail | conversion | Transformed outcome | LightGBM | 0.16 | -0.13 to 0.41 |
| Womens e-mail | conversion | DR-learner | LightGBM | 0.15 | -0.13 to 0.41 |
| Womens e-mail | conversion | S-learner | regularised linear | 0.05 | -0.20 to 0.29 |
| Mens e-mail | spend | Transformed outcome | regularised linear | -2.72 | -49.72 to 44.90 |
| Mens e-mail | spend | DR-learner | regularised linear | -8.09 | -55.30 to 40.00 |
| Mens e-mail | spend | X-learner | regularised linear | -8.25 | -55.58 to 40.10 |
| Mens e-mail | spend | T-learner | regularised linear | -8.31 | -55.63 to 40.04 |
| Mens e-mail | spend | Response model | LightGBM | -10.45 | -47.83 to 27.60 |
| Mens e-mail | spend | DR-learner | LightGBM | -11.38 | -55.82 to 30.37 |
| Mens e-mail | spend | T-learner | LightGBM | -14.56 | -55.26 to 24.79 |
| Mens e-mail | spend | X-learner | LightGBM | -15.01 | -55.39 to 21.13 |
| Mens e-mail | spend | Transformed outcome | LightGBM | -15.81 | -62.22 to 24.94 |
| Mens e-mail | spend | S-learner | LightGBM | -23.20 | -66.64 to 15.08 |
| Mens e-mail | spend | S-learner | regularised linear | -42.39 | -82.01 to 1.70 |
| Womens e-mail | spend | Transformed outcome | regularised linear | 30.64 | -10.68 to 67.32 |
| Womens e-mail | spend | DR-learner | regularised linear | 29.67 | -9.70 to 66.09 |
| Womens e-mail | spend | X-learner | regularised linear | 29.53 | -9.71 to 66.12 |
| Womens e-mail | spend | T-learner | regularised linear | 29.51 | -9.74 to 66.11 |
| Womens e-mail | spend | S-learner | LightGBM | 24.32 | -14.12 to 68.76 |
| Womens e-mail | spend | DR-learner | LightGBM | 4.99 | -32.92 to 49.04 |
| Womens e-mail | spend | Response model | LightGBM | 0.82 | -36.54 to 42.78 |
| Womens e-mail | spend | Transformed outcome | LightGBM | -5.11 | -44.83 to 39.64 |
| Womens e-mail | spend | X-learner | LightGBM | -13.13 | -56.09 to 29.09 |
| Womens e-mail | spend | T-learner | LightGBM | -16.31 | -61.53 to 28.82 |
| Womens e-mail | spend | S-learner | regularised linear | -17.05 | -55.09 to 18.15 |
Where the response model wastes contacts
For the women's e-mail, the response model ranks customers by how likely they are to visit, so its top deciles are full of people who would have come anyway. In its top 20%, 71% of the e-mailed visitors would have visited without the e-mail, against 58% in the top 20% of the uplift ranking (T-learner, LightGBM), which buys 8.4 pp of extra visit rate there against 6.8 pp. Reading the control rate as "would have come anyway" assumes the e-mail deters no one.
CWhich policy should ship?
The decision is scored on spend, because profit depends on spend. Every policy value is a doubly robust off-policy estimate on customers that neither trained the models nor chose the policy, with a paired standard error for every comparison.
30 candidate rankings compete: five learners and two base models, trained on visits, conversions or spend. A model trained on dense visits might rank customers better for revenue than one trained on sparse revenue, so the choice is left to data, but to data the final score never sees. The candidates' merit barely replicates from one random half of the customers to the other: rank correlation 0.24 on spend, against 0.66 on visits. With 578 buyers, the best-looking model is mostly the luckiest one.
The table values come from one random split of the customers, and their intervals are conditional on the rankings and budgets that split selected. Re-running the whole procedure on 50 random splits, the uplift procedure averages -$41 per 1,000 against blanket sending (from -$89 to +$9), falls short of it in 48 of 50 splits, e-mails a median 80% of customers (52% to 100%), and settles on 16 different rankings along the way.
DWhat changes when the experiment is two hundred times larger?
The Criteo Uplift v2.1 benchmark logs 13,979,592 users from randomised advertising incrementality tests, 85% treated. The same learners were fitted on a hash-based sample of 2.3M users from one half and evaluated on all 6,994,996 users of the other half. Treatment lifts visits by +1.03 pp and conversions by +0.115 pp on average.
At this size every ranking beats random by a wide margin, and the effect is concentrated: the top 10% of users by the response model hold 51% of all incremental visits. Their untreated visit rate is 32% and treatment adds +5.3 pp: here the effect tracks the baseline rate, so ranking by "who will act" is also ranking by "who will be moved". The response model is the best uplift ranker (Qini ×1,000 on conversions: 0.43, against 0.28 for the T-learner and 0.28 for the DR-learner). The meta-learners pay for estimating a small difference between two large, noisy predictions; with 15% controls and a control conversion rate of 0.19%, the DR-learner's final stage returns the same score for 63% of users. The 0.7% of users it scores below zero convert at 5.6% without treatment, against 0.19% overall, and gain +2.0 pp when treated: high-baseline users whose noisy pseudo-outcomes pushed their scores below zero. They hold 13% of all incremental conversions, which is why its curve jumps in the last 1%.
Method and limits
- Out-of-sample by construction. Five outer folds; every score comes from models fitted without that customer. The DR-learner cross-fits its own nuisance models inside the training folds. A test scrambles held-out outcomes and checks that held-out scores do not move.
- Known propensities. Assignment probabilities are fixed by design (1/3 per arm), so IPW and doubly robust policy values are unbiased without a propensity model. Tests check both on synthetic data, the DR one with a deliberately wrong outcome model.
- Selection is part of the model. Ranking and budget are chosen on one random half and applied to the other; the two halves are pooled.
- Uncertainty. Qini intervals are a Poisson bootstrap over customers with scores held fixed: they cover evaluation noise, not retraining noise. Policy intervals are normal approximations on per-customer doubly robust scores, conditional on the rankings and budgets one split selected; the 50-split repeat shows the selection noise they leave out.
- Assumed economics. The data has no margins or costs. The defaults (40% gross margin, $0.15 per e-mail including fatigue and unsubscribe risk) are assumptions; the sliders show the conclusions under others.
- Limits. One two-week campaign from 2008 with 578 buyers; spend is heavy-tailed. Criteo features are anonymised, its labels are visits and conversions (not revenue), and its Qini intervals use 20 random groups of the held-out half instead of a bootstrap. A multiplicative model (response × relative lift) was not tried.