Holdout testing: the honest way to measure marketing lift

Every marketing tool grades its own homework, and most grade generously. The standard method — “a customer we messaged bought something within 30 days, so we’ll take credit” — has a flaw you can state in one sentence: some of those customers would have bought anyway. The fix has been known to science for a century. Marketers call it a holdout.

How a holdout works

  1. Before sending anything, set aside a random slice of eligible customers — say 10%. They get no reminders, no cards, nothing.
  2. Market normally to the other 90%.
  3. After a fair window, compare: what fraction of each group ordered?

If reminded customers reorder at 34% and held-out customers at 21%, your marketing caused 13 points of reordering. The 21% is the inconvenient baseline — customers who came back on their own and whom simple attribution would happily claim.

Why randomness is the whole trick

The holdout must be a random slice, decided before sending — never “customers who happened not to open” (they differ systematically from openers) and never chosen after the fact. Random assignment makes the two groups statistically identical twins; whatever gap appears afterward can only have come from the treatment. Deterministic assignment (hashing a customer ID) keeps each customer stably in their group across cycles.

Reading the numbers without fooling yourself

  • Small samples lie enthusiastically. With 20 customers per group, a 34%-vs-21% gap is coin-flip noise. Treat results as directional until each cohort passes roughly a hundred; honest tools say so on the dashboard instead of hiding the caveat in a footnote.
  • Windows should match the behavior. Judge reorder reminders over a purchase cycle (whatever your product’s rhythm is), not a quarter picked for the board deck.
  • Lift concentrates. A 13-point average usually decomposes into big lift on some segments and none on others — which is exactly the information that makes next month’s spend allocation smarter.

The two-speed setup that works in practice

Day to day, per-send attribution with tight rules is the right speedometer: credit an order only when it matches the same customer and the same product, within a bounded window — and count a customer who reordered on their own as an honest zero. Underneath it, a standing holdout runs as the audit, telling you what fraction of that attributed revenue is real lift. When a tool shows you both numbers and the second is smaller, that’s not weakness — that’s the only version of marketing math you can take to a spreadsheet with a straight face.

Common questions

Does holding customers out cost me sales?

A little, briefly — that is the price of knowing. A 10% holdout forgoes the lift on 10% of customers for one period, and in exchange tells you whether the other 90% of spend is doing anything. Most stores find the answer worth far more than the handful of deferred reminders.

How big does the holdout need to be?

Enough customers on each side that the gap could not plausibly be luck — as a rule of thumb, treat comparisons as directional until each group has at least a hundred customers who reached the decision point. Small stores get there by letting the test run longer, not by skipping it.

What is the difference between attributed revenue and incremental revenue?

Attributed revenue counts orders that followed a message within a window — it includes people who would have bought anyway. Incremental revenue is the gap between treated and held-out groups — only what the marketing caused. Attribution is a good day-to-day speedometer; incrementality is the audit.

Keep reading