Holdout test
A holdout test is an experiment in which part of the audience is deliberately excluded from advertising. The test group sees the ads, the holdout group does not, and the gap in purchasing behavior between the two shows how many conversions the advertising actually caused and how many would have happened anyway. It is the cleanest form of incrementality measurement at the audience level.
In short
| What it is | A randomly selected share of the audience kept away from the ads, as a control |
| Where it is used | Platform conversion lift studies, CRM campaigns, retargeting, email |
| What it shows | The share of conversions that would have happened without the campaign |
| Validity condition | Random assignment. Any other split invalidates the result |
How the test works
The procedure in five steps
- Definition. Before launch, set the audience, the campaign, the metric to track (orders or revenue per user) and the test length: one purchase cycle plus a week for conversions to settle.
- Random split. The audience is divided at random, not by behavior and not by whom the algorithm chose to reach. The holdout usually takes 5 to 20% of the audience: large audiences need less, smaller ones need more so that enough conversions occur inside it.
- Running the test. The test group receives the ads; the holdout is excluded from the campaign and left alone for the whole period. Other campaigns, prices and the website stay unchanged during the test.
- Measurement. All conversions in both groups are counted from the same source, the order system or CRM, not the conversions credited by the platform. Conversion rate or revenue per user is compared, because the groups differ in size.
- Calculation. Lift = (test conversion rate minus holdout conversion rate) / holdout conversion rate. Incremental conversions = test conversions minus (test users × holdout conversion rate). Campaign cost divided by incremental conversions gives incremental CPA.
Example calculation
Worked example: e-commerce retargeting of visitors from the last 30 days, four weeks of testing plus a week of settling, 10% holdout.
| Test group | 180,000 users, 7,560 orders, conversion rate 4.2% |
| Holdout | 20,000 users, 720 orders, conversion rate 3.6% |
| Lift = (4.2 minus 3.6) / 3.6 | 16.7% |
| Incremental orders = 7,560 minus (180,000 × 3.6%) | 7,560 minus 6,480 = 1,080 |
| Platform credited 3,000 orders | 1,080 caused (36%) |
| Cost €54,000 | credited CPA €18, incremental CPA €50 |
At an average margin of €45 per order the campaign looked profitable in the report, but according to the test every order it actually caused costs more than it earns. The answer is not to switch the campaign off, but to narrow the audience or lower the budget and repeat the test after the change.
When not to use a holdout
A holdout needs control over who sees the ads and enough conversions in both groups. It is therefore unsuitable in three situations. With small audiences, where the holdout would produce only a few dozen conversions: the gap between the groups then disappears into random variation and the test proves nothing. With channels where an individual user cannot be excluded, such as search, offline media or parts of programmatic: there a geo lift test is used, which splits regions instead of users. And where other campaigns with the same message reach the holdout group: the control group has to be excluded from all of them; otherwise the measured effect shrinks and the campaign comes out worse than it really is.
How a holdout works
The audience is split at random, typically 90/10 or 80/20. The larger part sees the ads, the smaller does not. After a settled period the conversion rates of the two groups are compared. A worked example: if 5 percent of the test group converts against 4 percent of the holdout, the advertising caused only a fifth of the credited conversions: the rest would have come on their own. That ratio is exactly what decides whether a campaign deserves scaling or a budget cut.
From our own practice
The holdout is the standard tool wherever a campaign targets your own audience: retargeting and CRM. That is where the gap between claimed and caused tends to be widest, because the campaign reaches people already on their way to buying. Combined with geo lift tests, which we run through DiagnostIQ for whole channels, the holdout completes the picture at audience level: the geo test answers whether the channel works, the holdout whether the specific audience does.
Common mistakes
- Comparing reached versus unreached without randomisation. People the algorithm chose to reach differ from those it skipped. Without random assignment the comparison is worthless.
- A small holdout on a small audience. When only a handful of people convert in the control group, the result is noise.
- Running it too briefly. Ad effects decay over time. Evaluate after the settling period, not after a week.
Related terms
See also incrementality, geo lift test, remarketing and attribution model.
Frequently asked questions
How large should the holdout be?
Usually 10 to 20 percent of the audience. A smaller holdout preserves campaign reach but extends the time to a conclusive result.
Am I losing revenue by hiding ads from part of the audience?
Temporarily some, that is the price of the information. When the test shows low incrementality, the saved budget repays it many times over.
Where do I set a holdout up technically?
Platforms offer conversion lift studies. In CRM campaigns the holdout is made by splitting the list before sending.
How we can help
We design and evaluate incrementality tests as part of our Performance marketing agency service.