Broad Nurture, Propensity, or Uplift?
One trial-to-paid nurture sequence, four targeting policies, and a randomized holdout. Compared on incremental conversions rather than model accuracy, which is the only comparison that changes the answer.
- The decision
- Which trial users should receive the nurture sequence, and which should receive nothing?
- The method
- Randomized 50/50 holdout, T-learner uplift model, policies evaluated from the randomization on a held-out validation cohort.
STEP 01
The business decision
A subscription business sends a five-email nurture sequence to every trial signup. Roughly 120,000 trials a quarter, all of them contacted, nobody excluded.
Two questions arrive together. Is the sequence worth running at all, and should everyone be getting it? The second one is where the money is, because the contact budget is shared with three other programs and every send spends a little of the customer’s patience.
STEP 02
The tempting metric
The team has a conversion model with good discrimination. It ranks trial users by their probability of converting to paid, and it ranks them well.
The obvious move is to point the sequence at the people most likely to convert. Concentrate the budget where the conversions are. It sounds like targeting and it is easy to defend in a meeting.
STEP 03
Why that may be wrong
A conversion model answers “who is likely to subscribe.” The campaign needs an answer to “whose decision would change if we sent this.”
Those are different questions with different answers, and the gap between them is not subtle. The people most likely to convert are frequently the people already converting. Sending them a sequence adds cost and changes nothing.
There is a worse case. Some people react to another five emails by unsubscribing, and a few of those would have subscribed. A model built to find likely converters will rank exactly those people highly, because they were likely converters.
STEP 04
Data available and missing
Available. Product engagement in the first three days of trial, acquisition channel, the plan tier the user looked at, and whether they had subscribed before. All of it observed before the sequence starts, which is what makes it usable for targeting.
Missing. Why the person signed up, what else they are evaluating, and anything about intent that never shows up in product telemetry. That absence is the reason a model can only sort people coarsely, and the reason the estimate below is a group average rather than a person-level prediction.
Critically also available: a randomized holdout. Without one, none of what follows can be estimated at all.
STEP 05
The target quantity
The estimand is the conditional average treatment effect: the change in probability of converting to paid, caused by receiving the sequence, for people with a given set of observed characteristics.
Not the average effect across everyone, which is the thing a single holdout gives you. The whole exercise is about how that effect varies, because variation is what makes targeting possible. If the effect were constant, the only sensible policies would be everyone or nobody.
STEP 06
The strongest feasible design
Randomize the sequence 50/50 across every trial signup for one quarter. Half receive it, half receive nothing. Split the resulting data into a 60% training set and a 40% validation cohort of 48,219 trials that no model ever sees.
Fit two conversion models on the training set, one on the treated arm and one on the control arm. The difference between their predictions is the estimated uplift. On four discrete features that is a saturated model, which means there is nothing to tune and nothing to overfit.
Then compare four policies on the validation cohort:
- Broad nurture. Contact everyone, which is what the program does today.
- Propensity targeting. Contact the top 40% by predicted conversion probability.
- Uplift targeting. Contact the top 40% by estimated uplift. Same budget as the policy above.
- Drop the negatives. Contact everyone the model does not expect to lose, with no budget cap.
STEP 07
Assumptions and failure modes
STEP 08
Worked example and diagnostics
Does the ranking hold up?
Before comparing policies, the ranking has to earn belief. Sort the validation cohort by estimated uplift, split into ten equal groups, and measure the actual effect in each one from the randomization.
- Predicted
- Observed
Show the data
| Period | Predicted | Observed |
|---|---|---|
| D1 | 18.7pp | 19.2pp |
| D2 | 16.6pp | 16.3pp |
| D3 | 12.4pp | 10.7pp |
| D4 | 4.0pp | 2.1pp |
| D5 | 2.2pp | 2.2pp |
| D6 | 0.8pp | -0.2pp |
| D7 | 0.6pp | 0.1pp |
| D8 | 0.5pp | 2.0pp |
| D9 | -0.1pp | 0.6pp |
| D10 | -4.0pp | -1.9pp |
The top decile predicts +18.7pp and measures +19.2pp. The bottom predicts -4.0pp and measures -1.9pp. That is a usable ranking.
It is not a clean one. Deciles eight and nine sit above their predictions, and the bottom decile is less negative than the model expected. Each decile holds about 4,822 people split across two arms, so a point or two of wobble is sampling noise rather than a finding. What matters is that the top of the ranking is genuinely different from the bottom, which it is.
How far down the list is it worth going?
- Ranked by uplift
- Ranked by propensity
Show the data
| Period | Ranked by uplift | Ranked by propensity |
|---|---|---|
| 0% | 0.0 | 0.0 |
| 569.8 | -58.4 | |
| 10% | 928.0 | 7.0 |
| 1,388.9 | 33.2 | |
| 20% | 1,713.7 | 102.2 |
| 2,047.9 | 200.6 | |
| 30% | 2,246.2 | 353.1 |
| 2,306.7 | 632.8 | |
| 40% | 2,346.0 | 995.7 |
| 2,423.9 | 1,399.3 | |
| 50% | 2,453.5 | 1,734.7 |
| 2,466.3 | 2,199.3 | |
| 60% | 2,435.3 | 2,312.4 |
| 2,435.0 | 2,339.1 | |
| 70% | 2,460.9 | 2,377.4 |
| 2,556.0 | 2,408.4 | |
| 80% | 2,554.4 | 2,447.7 |
| 2,604.5 | 2,490.5 | |
| 90% | 2,594.9 | 2,502.7 |
| 2,586.8 | 2,464.5 | |
| 100% | 2,487.3 | 2,487.3 |
The teal curve climbs steeply and then flattens. Past roughly the halfway point it stops climbing at all, because everyone left is either unreachable or actively worth skipping.
The violet curve is the finding. Ranked by conversion probability, contacting more people buys almost nothing until the list is deep enough to start including persuadable users by accident. A model with genuine predictive accuracy produces a targeting policy that is close to worthless.
STEP 09
The estimate, with uncertainty
Conversion lift among the targeted 40%, ranked by estimated uplift
Measured from the randomization on a validation cohort the model never saw. The comparison group is the control arm within the same targeted set, so this is the effect of the policy rather than of the sequence.
- Incremental conversions
- 2,346
- Contacts
- 19,287
- Cost per incremental
- $0.66
All four policies, on the same cohort:
| Policy | Contacts | Incremental conversions | Per 1,000 contacts | Net value |
|---|---|---|---|---|
| Broad nurture | 48,219 | 2,487 | 51.6 | $207,559 |
| Propensity, top 40% | 19,287 | 996 | 51.6 | $83,091 |
| Uplift, top 40% | 19,287 | 2,346 | 121.6 | $197,867 |
| Drop the negatives | 41,087 | 2,599 | 63.3 | $217,656 |
Three things in that table are worth stopping on.
Uplift targeting buys 94% of the conversions using 40% of the contacts. Same sequence, same creative, same budget as the propensity policy. The only difference is the sort order.
Propensity targeting and broad nurture produce the same lift per contact. Both land at 51.6 incremental conversions per thousand. A quarter of modeling work, and the resulting policy performs exactly as well as picking people at random. The exact tie is a coincidence of this simulation, but the direction is not: a propensity ranking mixes people with no upside and people with negative upside in roughly the proportions the population already has.
Contacting fewer people produced more conversions. Dropping everyone with a negative estimated effect removes 7,132 contacts and adds 112 conversions. Those sends were not merely wasted. They were losing subscriptions.
STEP 10
The recommendation
Deploy the drop-the-negatives policy immediately. Move to uplift targeting once the contact budget is genuinely contested.
The reasoning:
- Dropping negatives is nearly free to implement, needs no budget decision, and is the only policy here that beats broad nurture on both conversions and cost. It is strictly better than what the program does today.
- Uplift targeting at 40% produces slightly fewer conversions than broad nurture (2,346 against 2,487) while freeing 28,932 contacts. That trade is only worth making when another program has a better use for them, which is a question about the whole contact budget rather than about this campaign.
- The propensity policy should be retired regardless of what replaces it. It costs modeling effort and delivers the population average.
The honest summary for the program review: the sequence works, a third of the sends do nothing, and about 4,340 of them are actively costing us subscriptions.
STEP 11
What this does not establish
What this does not establish
- That the ranking will hold next quarter. Persuadability moves with price, competition, and the product, and it moves faster than baseline conversion does.
- That the negative-effect group should never be contacted at all. It establishes that this sequence hurts them. A different message might not.
- That uplift targeting is right at any particular budget. The 40% here is arbitrary; the curve in step eight is the input to that decision, not this estimate.
- Anything about long-run value. The outcome is trial-to-paid conversion. A subscriber acquired by nurture may retain differently from one who arrived on their own.
- That the model would work on a different program. It was fitted on this sequence, this population, and these four features.
STEP 12
What to test next
- Keep a permanent holdout. A rolling 5% who never receive the sequence turns this from a project into a number available every quarter, and is the only way to notice the ranking going stale.
- Measure retention, not just conversion. Follow both arms for six months. A nurtured subscriber who churns at month three was not incremental value.
- Test a different message on the negative group. The finding is that this sequence loses them. Whether a shorter one would keep them is a separate experiment and a cheap one.
- Take the contact budget question seriously. These policies were compared at a fixed budget for this campaign alone. The real decision allocates contacts across every program competing for the same customer.
STEP 13
Technical appendix
The estimator, and how the policies were scored Code
Uplift is a T-learner: conversion rate estimated separately for the treated and control arms within each of the 36 strata formed by the four discrete features, with the difference taken as the estimated effect. On discrete features that is a saturated model, so there is no smoothing, no regularisation, and nothing to tune.
Policies are scored from the randomization, never from the simulation’s
latent truth. For a policy targeting set T, the estimate is the
difference in conversion rate between the treated and control members of
T, multiplied by the size of T. Because assignment
was random, those two groups are comparable, and this is a calculation a real
program could run on its own data.
Intervals are the usual two-proportion normal approximation. They understate uncertainty slightly for the targeted policies, because the targeting rule was itself estimated from data; a fully honest interval would account for that, most simply by cross-fitting.
Regenerated from seed 20260824 by a committed script, so
the dataset and every figure above move together or not at all. The source is
private; happy to walk through it. Economics assume
$85.00 first-year
contribution margin and $0.08
per contact.
How close did the estimates get to the truth? Technical
Because this is simulated, the true average effect is known. The data were generated with an average treatment effect of +5.1pp across the cohort. Broad nurture, which treats everyone and is therefore estimating exactly that quantity, measured +5.2pp with an interval of +4.4pp to +5.9pp.
The interval covers the truth and the point estimate is close, which is the expected outcome from a randomized comparison this size rather than evidence that anything clever happened. The harder question, whether the uplift ranking is real, is answered by the decile diagnostic rather than by any single number.
Context
Best suited for
Lifecycle programs with a repeated, controllable contact, enough volume to hold out a meaningful share, and a contact budget that is genuinely scarce.
Data requirements
- A randomized holdout. Uplift cannot be estimated from observational history
- Features observed before the treatment decision, not after it
- Enough volume that a tenth of the cohort still supports a rate estimate
- An outcome that lands inside a window the business will wait for
What changes by context
- B2C subscription
- The default case. Watch that the outcome is retained value rather than the first payment, since nurture can pull forward a conversion that would have happened anyway.
- B2B SaaS
- The unit is an account, which cuts effective sample size by an order of magnitude and usually makes deciles impossible. Rep capacity, not send capacity, is the budget being allocated.
- Product-led
- The strongest features are in-product, and the best intervention is often a product change rather than a message. The same uplift logic applies, but the treatment is harder to randomize cleanly and easier to leak between users.
- Lifecycle-led
- Contact budget is shared across programs, so a per-campaign policy is a local optimum. The real allocation ranks the uplift estimate of every program against one budget.
When it breaks down
- Cohorts below roughly 20,000, where a holdout large enough to fit on leaves too little to validate against.
- Very low baseline conversion, where the effect is smaller than the noise in every decile.
- One-shot campaigns, since the ranking cannot be validated before it has to be used.
- Programs with no holdout and no appetite for one. There is no version of this analysis that works on observational data alone.