Writing/Essay/Incrementality
Essay · Incrementality

What incrementality actually measures (and what it doesn't)

A plain-language tour of holdouts, geo tests and the traps that make a "lift" disappear under scrutiny – and why a lift is only ever as good as its design.

AuthorKunal Mirchandani
PublishedJuly 2026
Read time6 min
SeriesMeasuring honestly

Every analytics team has watched it happen. A platform or an agency presents a "lift study": the campaign drove +40% incremental conversions. The chart is beautiful, the methodology page has a diagram with two little groups of people on it, and the number goes straight into the deck.

Three months later, someone reconciles the quarter and the +40% is nowhere to be found. Revenue moved by nothing like it. The lift didn't survive contact with the business – it evaporated on the way from the study to the P&L.

Incrementality is the most abused word in marketing measurement. Used properly, it's the only honest answer to the only question that matters. Used loosely, it's attribution with a lab coat on. The difference is worth understanding properly.

The definition

What the word actually means

Incrementality answers a causal question: what happened because of the spend, that would not have happened anyway?

That's it. Not what happened after the spend. Not what happened near the spend. The difference between the world with the campaign and the world without it.

The only way to see that difference is to create both worlds. Take a group of people, or a set of regions, and split them: one side sees the campaign, one side doesn't, and everything else stays as equal as you can make it. The gap between the two groups is the lift. That's the whole concept – treated versus control, the same logic as a drug trial.

Anything that doesn't involve a real comparison group is not measuring incrementality. It's measuring correlation and hoping.

The toolkit

The toolkit, in plain language

Holdouts are the cleanest version. Randomly withhold the campaign from a small slice of the audience – one in twenty, say – and compare outcomes against the exposed majority. When the randomisation is real, this is as close to truth as marketing measurement gets.

Geo tests exist because you can't withhold a TV ad from one person in Sydney. So you withhold it from Wollongong. Pick matched regions, switch the channel off in some, keep it in others, and compare. The method is simple; the matching is where the work is – regions differ in more ways than you expect, and a badly matched control region quietly poisons the answer.

Platform lift studies – the conversion-lift tools inside the big ad platforms – use the same idea with the platform running the split. Useful, fast, and worth reading with one caveat stapled to the front page: the referee is also a player. The platform grading its own homework has every incentive for the answer to be "keep spending." Not worthless – just never the final word on its own.

And for the portfolio view, marketing mix modelling fills the gaps between experiments – using variation over time rather than a controlled split. Experiments calibrate it; it interpolates between them. (That's a whole essay of its own; this one is about the experiments.)

Failure modes

The traps that make a lift disappear

This is where most "incremental" numbers actually die. The usual suspects:

Five traps
  • Leaky randomisation. The control group sees the campaign anyway – shared devices, logged-out users drifting into logged-in targeting, a sales team that calls everyone regardless of cell assignment. Every leak dilutes the contrast between treated and control, and dilution reads as a smaller, fuzzier lift – or none.
  • Selection drift. You randomised the audience; the platform then optimised delivery within it, showing ads to the people most likely to convert. The treatment is no longer random – it's concentrated on the already-convinced. The "lift" is now measuring who the algorithm found, not what the ads caused.
  • The window is too short. Measure during the promotional spike and the lift looks heroic; annualise it and you've built a budget on a fortnight. Some effects lag – brand search rises weeks later – and some are pulled forward, borrowing next month's sales.
  • No power to detect anything. If the true lift is 3% and your test can only detect 15%, the result will be "no significant effect" – which gets read as proof the channel does nothing. It proves no such thing. Power is decided before the test runs, not explained away after.
  • Spillover. Treated customers tell control customers. Commuters from holdout regions work in treated ones. Word of mouth does not respect your experimental design, and the contamination usually runs in the direction of understating the effect.

None of these are exotic edge cases. They're the default failure modes – which is why a lift number should always arrive with its design attached, the way a forecast should arrive with its assumption.

The honest limits

What incrementality doesn't measure

The second half of the title, because the honest limits matter as much as the method:

  • It doesn't tell you why. A lift is a what, not a why. It won't tell you which creative, which message, which moment did the work. That's a different investigation.
  • It doesn't see the long term. A four-week holdout cannot measure a four-year brand effect. Run only short-window tests and you will systematically undervalue anything that compounds slowly – which is to say, most of brand building.
  • One test is one point. A lift applies to a channel, at a spend level, with a creative, in a period. It does not map the whole response curve. Double the spend and the lift per dollar almost certainly falls – how fast it falls is a separate question, and answering it takes several points, not one.
  • It doesn't stay true. Audiences saturate, creatives wear out, competitors move. Last year's lift is evidence, not entitlement. The number has a shelf life, and the retest is part of the price of knowing.
The standard

Specific is the point

None of this makes incrementality impractical – it makes it specific.

The working rule

A well-run holdout doesn't tell you what works everywhere, always. It tells you what worked here, this quarter, at this spend. And that is exactly the shape a budget decision comes in.

The discipline is the same one I apply everywhere else: publish the design before the result, state the assumption, and score the number against what actually happened.

A lift with a leaky control group is not a measurement – it's a story with a diagram attached. And stories, as a rule, don't reconcile with revenue.

Next playbook

Building an experimentation practice from zero →

Have a lift number that won't reconcile?

An incrementality programme to stand up, a lift study to stress-test, or a role to fill. Happy to talk.