Experiment readout · Criteo Uplift v2.1 · 13,979,592 users

Should everyone keep seeing this campaign?

A decision memo for a growth team: is the ad campaign working, how sure are we, and if we stop showing it to everyone, who should still get it?

The short version

2026-09-29T07:55:42.225863 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 0% 20% 40% 60% 80% 100% Share of users shown the campaign (ranked by predicted uplift) 0% 20% 40% 60% 80% 100% Share of all incremental conversions top 10% of users: 82% of the extra conversions top 30% of users: 89% of the extra conversions Target by model (S-learner) Target at random
The one chart that matters. Users ranked by predicted uplift (S-learner, trained on other users). The curve shows how much of the campaign's total extra conversions you keep if you target only the top X%. The dashed diagonal is random targeting. The line is the mean over five held-out folds; the band spans the folds.

1 · Experiment readout

The data are 13,979,592 users from Criteo incrementality tests. Treated users were eligible to see the campaign's ads, and control users were held out. Two outcomes: whether the user visited the advertiser's site, and whether they converted. Twelve anonymized features were recorded before treatment. Everything below is computed on all rows, with DuckDB aggregates over a parquet file (peak memory under 2 GB).

1.1 Was the split what the design said it was?

The test was designed as an 85/15 split. The observed treated share is 0.8500001 (11,882,655 treated, 2,096,937 control). A chi-square sample-ratio-mismatch test against 0.85 gives χ² = 1.8e-06, p = 0.999. There is no sample ratio mismatch. (Testing against a naive 50/50 expectation would "fail" with a p-value that underflows to zero, which is why the test has to be run against the design ratio.)

1.2 Covariate balance: the check that fails

Randomization should make the two groups look alike before treatment. With groups this large, chance alone produces standardized mean differences (SMD) with a standard deviation of about 0.00075. The observed SMDs reach 0.049. All 12 of 12 features differ significantly after Holm correction. A gradient-boosted model can predict treatment from the features on held-out rows with AUC 0.5075. Twenty label permutations give a maximum AUC of 0.5003.

2026-09-29T07:55:42.298434 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ f0 f1 f2 f3 f4 f5 f6 f7 f8 f9 f10 f11 −0.04 −0.02 0.00 0.02 Standardized mean difference Raw treatment vs control After propensity weighting
Standardized mean difference by feature. The grey band is the ±1.96 SD range that pure chance would produce at this sample size (it is almost invisibly thin). Orange: raw groups. Blue: after reweighting by the cross-fitted propensity score, which cuts the worst imbalance from 0.049 to 0.014, still well above chance.

This fits the dataset's own documentation. The release was assembled from several incrementality tests and, for privacy, was "sub-sampled non-uniformly so that the original incrementality level cannot be deduced." Either could tie assignment to the features. The practical consequence is that a raw treated-minus-control difference is not a clean causal estimate here. The analysis therefore treats assignment as random conditional on the features, with an estimated propensity e(x). Its middle 90% spans 0.832–0.877, so overlap is excellent. The headline estimator is augmented inverse-propensity weighting (AIPW, doubly robust), with nuisance models cross-fitted on two halves of the data.

1.3 Average treatment effect

0.100pp
conversion effect (AIPW)
CI 0.092pp to 0.107pp
48.5%
relative conversion lift
CI 43.6% to 53.6%
0.75pp
visit effect (AIPW)
CI 0.72pp to 0.78pp
18.6%
relative visit lift
CI 17.8% to 19.3%
EstimateVisit95% CIConversion95% CI
Control rate (raw)3.820%0.194%
Treated rate (raw)4.854%0.309%
Difference in means, analytic1.034pp1.006pp – 1.063pp0.1152pp0.1085pp – 0.1219pp
Difference in means, bootstrap (10,000 resamples)1.005pp – 1.063pp0.1083pp – 0.1217pp
Relative lift, naive (delta method)27.1%26.2% – 28.0%59.4%54.4% – 64.7%
AIPW (doubly robust)0.751pp0.725pp – 0.777pp0.0996pp0.0923pp – 0.1068pp
Relative lift, AIPW18.6%17.8% – 19.3%48.5%43.6% – 53.6%
Effect on users actually exposed (AIPW ÷ exposure rate)20.84pp20.11pp – 21.57pp2.76pp2.56pp – 2.96pp

The analytic and bootstrap intervals agree to the displayed precision. The bootstrap resamples rows with replacement; for a binary outcome that is exactly a binomial draw per arm, so it needs no pass over the data. Only 3.6% of treated users were actually shown an ad, and control users never were. Dividing the intent-to-treat effect by that rate (a Wald / instrumental-variable ratio) estimates the effect on the users who saw it. That assumes being eligible but unexposed has no effect of its own.

1.4 Tighter intervals from pre-treatment data

The twelve features predict the outcomes well, so adjusting for them removes noise. CUPED uses one pooled linear coefficient. Lin's estimator lets the slope differ by arm. CUPAC uses a cross-fitted gradient-boosted prediction as the single covariate.

2026-09-29T07:55:42.270678 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 0.6 0.7 0.8 0.9 1.0 Effect of treatment (percentage points) difference in means CUPED, 12 linear features Lin interacted regression CUPAC, cross-fitted GBM score AIPW, cross-fitted GBM Visit rate 0.09 0.10 0.11 0.12 Effect of treatment (percentage points) Conversion rate
Every estimator, point and 95% CI. Grey is the raw difference in means; blue are regression adjustments; orange is AIPW. The shift from grey to the rest is the imbalance correction from §1.2. The narrowing is the variance reduction.
EstimatorVisit effectVariance vs rawConversion effectVariance vs raw
difference in means1.034pp100%0.1152pp100%
CUPED, 12 linear features0.698pp75%0.0916pp89%
Lin interacted regression0.773pp74%0.1002pp89%
CUPAC, cross-fitted GBM score0.638pp70%0.0939pp89%
AIPW, cross-fitted GBM0.751pp85%0.0996pp116%

CUPAC cuts the variance of the visit estimate to 70% of the raw estimator's. That is the same precision as running the test on 1.44× as many users. For the rare conversion outcome the gain is smaller (89%). AIPW is less precise than the pure regression adjustments on conversion (116% of raw variance), because inverse-propensity terms add noise when outcomes are rare. It is still the headline. With assignment tied to the features, it is the only estimator here that corrects through both a propensity model and an outcome model, so it stays consistent if either one is right. The adjusted point estimates disagree with each other by more than their own intervals (visit: 0.64pp to 0.77pp). That is model dependence left over after the imbalance, so the quoted CIs understate the real uncertainty.

1.5 Who responds more? Segment effects

Each feature was cut at its deciles. Several features have one dominant value, so only 9 of 12 split into two or more segments, giving 80 segment-by-outcome effects. Each segment's effect is the mean AIPW score of its users. Correcting for 80 comparisons, 76 remain significant under Holm (family-wise) and 76 under Benjamini–Hochberg (false discovery rate). 0 of them are negative: no segment was measurably harmed. A heterogeneity test per feature (Cochran's Q across its segments, Holm-corrected over 18 tests) rejects "same effect in every segment" for 18 of 18.

2026-09-29T07:55:42.349385 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 1 2 3 4 5 feature value: low to high 0.0 0.2 0.4 0.6 0.8 Conversion effect (pp) by f8 bin 1 2 3 feature value: low to high by f3 bin 1 2 3 feature value: low to high by f9 bin
Conversion effect by segment for the three features with the most heterogeneity. Segments are ordered from the feature's lowest values to its highest. Dashed: the overall AIPW effect. Grey marks are not significant after Holm correction. The pattern shows why targeting can pay: the effect ranges over an order of magnitude between segments.

2 · Power and design

2.1 How big does the next test need to be?

The control conversion rate is 0.194%. At rates that low, small relative lifts need very large samples. With 14.0M users at 85/15, the smallest detectable conversion lift (80% power, two-sided α = 0.05) is 0.0092pp, or 4.8% relative. The 85/15 allocation costs precision. Its variance is 1.96× that of a 50/50 split of the same size, so a test that reserves few users for control needs about twice as many users in total.

2026-09-29T07:55:42.380546 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 2% 3% 5% 7.5% 10% 15% 20% 30% 50% Smallest relative lift in conversion rate you want to detect 100k 1M 10M 100M Users needed (80% power) this dataset: 14.0M users 50/50 split 85/15 split (as run)
Users needed to detect a given relative lift in conversion at the observed control rate (80% power, α = 0.05, unpooled two-proportion z-test).
Relative lift to detectUsers, 50/50Users, 85/15
2.0%40,833,51879,511,898
3.0%18,237,89235,391,364
5.0%6,630,19412,778,863
7.5%2,982,6125,700,589
10.0%1,697,8893,218,445
15.0%772,5431,440,965
20.0%444,637816,473
30.0%206,575368,147
50.0%80,813136,325

An effect as large as the one measured here (49% relative) needs only 144,581 users at 85/15. With CUPAC's variance ratio that falls to about 128,877. The next useful test is a targeted one (§3). The effects it has to detect are smaller, so the table above is the one to plan with.

2.2 Why you cannot stop the test the first time it looks significant

This section simulates 100,000 experiments per setting in which the treatment does nothing, and checks the p-value after each batch of data. Stopping at the first p < 0.05 inflates the false-positive rate well past the promised 5%:

2026-09-29T07:55:42.407696 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 1 2 3 5 10 20 50 100 Number of times the dashboard is checked (A/A test, no real effect) 0% 5% 10% 15% 20% 25% 30% 35% False positive rate promised 5% Stop at first p < 0.05
False positive rate when stopping at the first "significant" look. With 20 looks it is 25.0%; with 100 looks, 37.7%. Monte Carlo standard errors are below 0.15 percentage points.

There are two standard fixes. The table below compares them on 20 equally spaced looks, for a design with 80% power at the planned sample size. Alpha spending (Lan–DeMets with an O'Brien–Fleming-type spending function) fixes the looks in advance and uses very strict thresholds early. mSPRT (a mixture sequential probability ratio test, as used in always-valid p-values) allows checking at any time, with no plan.

MethodFalse positive ratePowerAvg. sample used (if effect is real)
Fixed horizon (look once)4.9%79.9%100%
Naive peeking (z > 1.96 at any look)24.7%88.2%45%
O'Brien-Fleming alpha spending4.9%77.3%74%
mSPRT (always-valid)1.1%49.0%82%

Alpha spending keeps the error rate at 5%, gives up little power, and stops early on average when the effect is real. mSPRT is conservative at the planned horizon (its guarantee holds for unlimited looks). Its cost is sample size: if it may run to 2.0× the planned sample, power reaches 87%, while the average sample used is only 1.11× the plan. Use alpha spending when the review schedule is fixed. Use mSPRT when stakeholders will look at a live dashboard.

Simulation details

The z-statistic path is a Gaussian random walk (exact for known-variance difference in means; an excellent approximation at thousands of conversions per look). Alpha-spending boundaries were calibrated on an independent set of 400,000 null paths. The mSPRT uses a normal mixing prior with variance τ² = 0.392 per look, set to the planned effect size, and rejects when the mixture likelihood ratio exceeds 1/α.

3 · Who should see the campaign? Uplift modeling

An uplift model predicts, for each user, how much the campaign changes their chance of converting. Four standard learners were compared, all using LightGBM as the base model:

Honest evaluation. Users were split into 5 disjoint folds. For each fold, every learner was trained on 1,500,000 users drawn from the other four and scored on all ~2.8M users of the held-out fold. A slice's uplift is measured with the held-out users' AIPW scores, so it is unbiased even with the imbalance from §1.2. Curves, areas and policy values below never touch training rows, and the spread is across the five folds.

2026-09-29T07:55:42.436895 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 0% 20% 40% 60% 80% 100% Share of users treated, ranked by predicted uplift 0.0 0.2 0.4 0.6 0.8 1.0 1.2 Incremental conversions per 1,000 users T-learner X-learner S-learner Transformed outcome Random
Uplift curves, conversion. Treating the top X% of users produces this many extra conversions per 1,000 users in the population. A good model rises steeply and then flattens; random targeting is the straight dashed line. Bands span the five held-out folds.
LearnerConversion Qini coef.Conversion AUUC ×10⁴Visit Qini coef.Visit AUUC ×10³
T-learner 0.513 ± 0.060 7.54 ± 0.97 0.631 ± 0.036 6.12 ± 0.23
X-learner 0.534 ± 0.071 7.64 ± 0.98 0.693 ± 0.012 6.36 ± 0.23
S-learner (best) 0.772 ± 0.053 8.82 ± 1.09 0.837 ± 0.022 6.90 ± 0.24
Transformed outcome 0.329 ± 0.107 6.61 ± 0.92 0.571 ± 0.025 5.89 ± 0.16
Random 0.008 ± 0.061 5.03 ± 0.80 0.009 ± 0.030 3.79 ± 0.24

AUUC is the area under the uplift curve (per-user scale). The Qini coefficient is the model's area above the random-targeting diagonal, divided by the area under that diagonal: 0 means no better than random, 1 means twice the area. Values are mean ± standard deviation across the five folds. On conversion, the S-learner has AUUC 8.82×10⁻⁴ against 5.03×10⁻⁴ for a random ranking (1.75×). It is the simplest learner and it wins by a margin larger than the fold-to-fold spread. That fits a known pattern: when the effect is small next to the baseline, fitting the two arms separately (T-, X-learner) or fitting a high-variance transformed target mostly adds noise. "Best" is picked on the same held-out folds, so its lead over the runner-up is slightly optimistic. The gap between any real learner and random is not.

2026-09-29T07:55:42.469257 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 0% 20% 40% 60% 80% 100% Share of users treated, ranked by predicted uplift 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Incremental visits per 100 users T-learner X-learner S-learner Transformed outcome Random
Uplift curves, visits: same setup, with the visit outcome.

Are the predictions calibrated?

2026-09-29T07:55:42.198793 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 1 2 3 4 5 6 7 8 9 10 Predicted-uplift decile (1 = lowest) 0.0 0.2 0.4 0.6 0.8 Conversion effect (pp) Predicted Measured on held-out users
Predicted vs measured conversion uplift by predicted decile (S-learner, averaged over the five held-out folds; whiskers are 95% CIs). The top decile is where the model is right: predicted 0.750pp, measured 0.818pp. Deciles 2 to 9 are small and roughly ordered. The bottom decile is miscalibrated: the model predicts a negative effect (-0.028pp), but the measured effect is positive, 0.070pp (CI 0.049pp to 0.091pp), larger than deciles 2 to 8. These are not users the campaign harms. They are users on whom the model is unreliable.

What is it worth? Policy value under a cost assumption

The dataset contains no prices, so costs are stated as a ratio. Let r be the cost of targeting one user divided by the value of one conversion. Targeting the top φ of users is worth V·[uplift(φ) − r·φ] per user in the population, where uplift(φ) is the held-out curve above. The best φ depends only on r:

2026-09-29T07:55:42.495477 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 0 5 10 15 20 25 30 Cost per targeted user ÷ value of one conversion (× 10,000) 0% 20% 40% 60% 80% 100% Best share of users to target
Best share of users to target as costs rise (S-learner). Line: median across folds; band: range across folds. The average extra conversion per targeted user is 1.0 per 1,000, so blanket targeting breaks even at r = 10.0 on this axis.
Scenario: cost per targeted userBest share to targetNet value, best policyNet value, target everyone
0.02× avg. incremental value per user100% (100%–100%)+9.76+9.76
0.1×100% (15%–100%)+9.00+8.96
0.5×16% (14%–19%)+7.96+4.98
1.0× (blanket breaks even)10% (7%–13%)+7.26+0.00
1.5× (blanket loses money)7% (7%–11%)+6.83-4.98

Net value is in conversions' worth per 10,000 users in the population (multiply by the value of one conversion to get money), averaged over the five held-out folds. The best-share ranges are across folds.

4 · Recommendation and limitations

For the growth team, in plain terms: the campaign causes real extra conversions, but about four in five of them come from roughly one user in ten, and a model can find those users in advance. For everyone else the campaign does very little. Unless showing the ad is nearly free, stop showing it to everyone and show it to the top-ranked 10–20%. That keeps most of the gain for a fraction of the cost. Do not read the model's bottom-decile "negative" predictions as users to protect from the ad. §3 shows those predictions are wrong. The next step is to test that directly: a new randomized test comparing "target the model's top slice" against "target everyone", sized with §2 and monitored with alpha spending.

Limitations.

Reproduce

Code, tests and this page are generated from one repository: github.com/tachyurgy/uplift-readout. make data downloads and checksums the file and converts it to parquet. make analysis runs the nuisance models, readout, power simulations and uplift models (about 30 minutes on a laptop, under 2 GB of memory). make site renders this page from the result files; no number on it is typed by hand.