Pivix Logo
Back to glossary

Incrementality Testing

Incrementality testing measures the additional conversions a marketing campaign actually caused, beyond what would have happened without it.

Key takeaways

  • Lift is the gap between treated and holdout groups, not reported conversions.
  • Randomised assignment is what licenses the causal claim; convenience splits do not.
  • Holdout size trades statistical precision against revenue knowingly left on the table.
  • Baseline rate, sample size and expected effect together determine the run length.
  • A lift figure applies only to the audience, budget and period tested.

In depth

An incrementality test randomises a population into a treated group that can see the campaign and a holdout that cannot, then compares the conversion rate of the two. The difference is incremental lift; dividing incremental conversions by treated conversions gives the share of results the campaign actually caused. Randomisation is what makes the comparison valid: because assignment is independent of intent, both groups carry the same underlying demand, so any gap left over is attributable to exposure rather than to who happened to be in each group.

Detectable lift depends on sample size, baseline conversion rate and the size of the true effect. Small effects on rare conversions need very large populations or very long runs, which is why teams often measure lift on an earlier, more frequent event. Holdout size is the central trade-off: a bigger holdout tightens the confidence interval but forgoes revenue from people you deliberately did not reach. Geo-based splits avoid user-level tracking entirely but sacrifice precision, because regions differ in ways randomisation across individuals would have balanced.

A practical test starts with a written hypothesis, a pre-registered success metric and a run length calculated before launch, not chosen once results look good. Choosing the outcome metric matters as much as the design. Counting clicks or raw form fills rewards campaigns that harvest existing demand; counting high-scoring leads from a qualification scorecard measures whether the campaign brought in buyers who fit. Running the same test structure across channels, one at a time, builds a comparable library of lift estimates over several quarters.

A lift estimate is valid for the audience, budget level and period tested, and does not transfer automatically. A channel that is incremental at a modest spend can saturate at triple the budget, so the result carries an implicit ceiling. Tests also struggle with long consideration cycles, because the holdout must stay suppressed for longer than a typical deal takes, and with brand effects that surface after the measurement ends. Contamination through organic search, referrals or a shared household breaks the separation the design assumes.

Example in practice

Suppose a growth team suspects branded search is taking credit for organic demand. They run a geo-holdout test, pausing those ads in 20% of regions, and qualified Pivix quiz completions there might drop only 4%, implying low incrementality. They would then reallocate that spend to a prospecting channel that could show a 28% lift in net-new high-scoring leads.

How to measure it

The headline number is incremental lift: conversions in the treated group minus the rate-adjusted conversions in the holdout, expressed as a percentage of the treated total. Pair it with a confidence interval, because a lift of twelve percent that ranges from minus five to thirty is not a decision you can act on. Report the interval alongside the point estimate every time.

Then convert lift into a cost figure. Divide campaign spend by incremental conversions to get incremental cost per acquisition, and compare that against the blended figure the platform reports; the gap between them is the size of the over-crediting problem. For qualification funnels, run the same arithmetic on high-scoring leads only, since incremental volume of poor-fit leads is not worth paying for.

Common mistakes

Practitioners routinely stop a test the first day the treated group is ahead. Conversion counts wander early on, so an early peek almost always shows a difference that has not stabilised. Calculate the required sample before launch from your baseline rate and the smallest lift worth acting on, then leave the test alone until it reaches that number. If the answer is needed sooner, shrink the question, not the sample.

The other failure is a leaky holdout. Suppressing one channel while retargeting, email and organic keep reaching the same people measures almost nothing, and the small lift that results gets read as proof the channel is worthless. Before launching, list every route by which the holdout could still be exposed and suppress or document each one. If a route cannot be closed, say so when reporting, because it biases the estimate downward.

Frequently asked questions

What is a holdout group?

A holdout group is a randomized portion of your target audience deliberately not exposed to a campaign, used as a control. Comparing outcomes between the exposed group and the holdout reveals the true lift the campaign produced.

Why do B2B incrementality tests often fail?

Long sales cycles and low weekly conversion volumes make it hard to reach statistical significance, so underpowered tests produce noisy, misleading results. Running the test long enough and protecting the control from contamination are essential for a valid read.

How big should the holdout group be?

Large enough that the smallest lift worth acting on would be statistically detectable, which follows from your baseline conversion rate and the traffic available. A holdout of five to ten percent is a common rule of thumb for high-volume campaigns, but low-volume B2B programmes often need a much larger share, or a longer run, to reach the same certainty.

What is the difference between incrementality testing and attribution?

Attribution divides credit among touchpoints that were recorded; incrementality asks whether the conversion would have happened at all without the campaign. Attribution can never separate harvested demand from created demand, because it only sees people who converted. An incrementality test sees the counterfactual through the holdout, which is why the two methods often disagree on branded search.

How long should an incrementality test run?

Long enough to cover at least one full conversion cycle plus the time needed to accumulate the calculated sample. For a funnel where leads close in three weeks, a two-week test cuts off conversions still in flight and understates lift. Decide the end date before launch, and resist extending it only because the result is not the one you expected.

What is a geo holdout test?

A design that withholds the campaign from a randomly selected set of regions rather than individual users, then compares outcomes between the two sets of regions. It works without user-level identifiers, which suits privacy-restricted channels and offline media. The cost is precision: regions vary in seasonality, competition and mix, so more regions and matched pairs are needed to get a stable read.

Can a campaign show negative incrementality?

Yes, and it is more common than teams expect. Retargeting and branded search often reach people who were already going to convert, so paid spend cannibalises organic conversions and the measured lift lands near zero or slightly below. A negative or zero result is a reason to reduce spend on that placement, not evidence that the test was broken.

How often should incrementality tests be repeated?

Whenever a material input changes: budget level, audience definition, creative strategy or seasonality. A lift estimate ages because saturation and competition move, so a result more than a couple of quarters old should be treated as a hypothesis rather than a fact. Many teams run one channel test per quarter on rotation to keep estimates current without permanent holdouts.

Related terms

Turn glossary theory into qualified leads

Build a scorecard quiz funnel that qualifies and captures leads in minutes — no code required.

Start for free
  • No credit card
  • Free plan
  • Launch in minutes