Geo lift test
A geo lift test is an experiment in which advertising is switched on or off only in selected regions, and the revenue trend is compared against control regions where nothing changed. The gap between the groups is the campaign's incremental contribution. It is the most practical way to measure incrementality, because it needs no cookies, no consent and no changes to measurement, just revenue split by region.
In short
| What it is | Switching ads on or off in selected regions against controls |
| What it needs | Revenue that can be split by region, and campaigns that can be geo-targeted |
| What it does not need | Cookies, consent, pixels, site changes |
| Typical length | Four to eight weeks, plus ramp-up and cool-down periods |
How the test works
Phases and their length
The test follows a fixed schedule. Skip a phase and you get a number that cannot be defended.
- Pretest period. At least 8 weeks of weekly revenue by region, preferably 12 or more. It serves to select the control regions and calibrate the model; campaigns run unchanged during it.
- Minimum detectable effect. Before launch, the variance of the historical data is used to calculate the smallest lift the setup can separate from noise. If the expected effect is smaller, the budget change grows or the test gets longer.
- Ramp-up. The first one to two weeks after the change are not evaluated: the campaign is learning and purchases arrive with a delay.
- Evaluated period. A minimum of 4 weeks, usually 4 to 8. A shorter test cannot separate the effect from ordinary weekly fluctuations.
- Cool-down. One to two weeks after the change ends, to capture delayed orders and the regions' return to normal.
The control and the lift calculation
The control is not a single region but a weighted combination of control regions whose history tracks the trajectory of the test group: a synthetic control. The weights are found in the pretest period; during the test the same weights are applied to the current revenue of the control regions, and the result is an estimate of how the test regions would have developed without the change. Lift = (revenue of the test regions minus the estimate) / the estimate. Incremental revenue is that difference in currency terms, and incremental ROAS is the difference divided by the spend in the test regions over the evaluated period.
Worked example
Worked example: an online store launches a new channel in three regions only. In the pretest period these regions averaged €50,000 in weekly revenue, and the synthetic control matches them within 2%. The campaign runs for 7 weeks at €4,000 per week; week one is the ramp-up, weeks 2 to 7 are evaluated.
| Estimate without the campaign (synthetic control) | €52,000 per week, the control regions rose 4% seasonally |
| Actual revenue of the test regions | €58,500 per week |
| Lift | (58,500 minus 52,000) / 52,000 = 12.5% |
| Incremental revenue over 6 weeks | 6,500 × 6 = €39,000 |
| Spend over the 6 evaluated weeks | 4,000 × 6 = €24,000 |
| Incremental ROAS | 39,000 / 24,000 = 1.6 |
Without a control, a comparison with the pretest period would show 17% growth (58,500 against 50,000) and credit the campaign with the season as well. The real effect is 12.5%, and the incremental ROAS of 1.6 is the number to compare with the ROAS the platform credited itself with for the same regions and weeks. Whether the channel pays off is decided by margin: an incremental ROAS of 1.6 means €1.60 of revenue per euro spent, not profit.
How the test is built
Regions are split into test and control groups so that they historically match in revenue development: pairs of regions with similar seasonality and size. In the test regions the campaign is switched off, on, or scaled substantially; in the controls everything runs unchanged. Afterwards the trend of both groups is compared against their shared history. Control regions are usually combined into a synthetic control that mirrors the test regions' history; the lift is the deviation from that forecast.
From our own practice
The geo approach is the foundation of our platform DiagnostIQ, which we use to measure the incremental impact of channels. This is how we measured, for one e-commerce client, that TikTok brings 15 percent additional revenue that would not exist without it, even though attribution undervalued the channel. A small-market rule we design around: fewer regions mean longer tests and larger budget changes, so the effect can be separated from noise.
Common mistakes
- Picking regions at random. Without historical similarity, the gap between groups cannot be attributed to the campaign.
- Testing during a seasonal peak or a promotion. External factors drown out the lift.
- Ending the test early. The first weeks contain the decaying effect of previous advertising. Only the settled state gets evaluated.
- Measuring pixel conversions only. The point of the test is independence from measurement. Revenue is evaluated, not credited conversions.
Related terms
See also incrementality, holdout test, marketing mix modeling and MER.
Frequently asked questions
Does it work on a smaller market?
Yes, just with higher demands: fewer regions mean a longer test and a bigger budget change for the effect to be measurable.
Do I have to switch ads off?
Not necessarily. A scale-up test works the same way: budget doubles in the test regions and the extra revenue is measured.
How often should we test?
Large channels once or twice a year, and always after a major change of budget or strategy.
How we can help
We design and evaluate geo lift tests through DiagnostIQ as part of our Performance marketing agency service.