What paid social was actually worth
A geo-holdout incrementality test found that two-thirds of the conversions paid social claimed weren't incremental at all. We moved $6.2M to where the money actually worked, then reconciled the whole thing with finance.
Anonymised version of an engagement delivered in a prior in-house role.
Budgets were being set on a number nobody could defend
The insurer was spending tens of millions a year across paid media, and paid social reported the strongest return of any channel by a wide margin. On the strength of those platform-reported numbers, it kept getting more budget.
The problem was that finance never believed them. The platform's own dashboard was both the salesperson and the scorekeeper, and the return it claimed didn't reconcile with what showed up in actual policy sales. Two parts of the same business were running on two different truths, and the bigger the spend got, the more that gap mattered.
Attributed isn't the same as caused
Last-click and platform attribution answer "who touched this conversion last," not "would this conversion have happened anyway." Retargeting is especially flattering here: it serves ads to people already heading for checkout, then takes the credit when they arrive.
So the question wasn't "how many conversions did paid social get credited with." It was how many of those conversions wouldn't have happened without it: the incremental ones. That's the only number worth setting a budget on.
Measure the gap, don't model around it
The cleanest way to know what advertising caused is to switch it off somewhere comparable and watch what happens. So the core of this was a matched-market geo holdout: turn paid social off in a set of regions, keep it running in a matched set, and read the difference.
Platform conversion-lift tools grade the platform's own homework, and user-level holdouts were decaying as third-party cookies disappeared. A geo design is platform-agnostic and privacy-durable: it doesn't depend on tracking individuals, so it survives the thing breaking everything else.
Markets were paired on baseline conversion rate, size and seasonality, then split. To avoid trusting a single method, I triangulated the bottom-up experiment against a top-down marketing mix model. When an experiment and an MMM independently land in the same place, you can believe the answer.
Design the holdout
Pair and split markets; size the test against a minimum detectable effect.
Run & measure
Hold paid social out for 10 weeks; measure the conversion gap vs control.
Triangulate
Cross-check the experiment against an independent mix model.
Reallocate
Reset budgets on incremental return; reconcile with finance.
The lift was real, just much smaller than claimed
Treated regions did outperform held-out ones, so paid social was genuinely doing something. But the gap was far narrower than the platform's numbers implied: only about a third of attributed conversions were incremental, and the effect sat comfortably above the test's detection threshold rather than anywhere near the reported figure.
Reading incrementality channel by channel changed the picture entirely. Paid social was the most over-credited; non-brand search and video were carrying more real weight than their attributed numbers suggested. The reallocation followed directly from that.
For the first time, marketing's number and the CFO's number were the same number. That, more than the ROAS lift, is what made it stick.
A one-off test became the way budgets get set
The immediate result was a +27% lift in blended incremental ROAS: same total spend, just pointed at what actually moved sales. But the lasting change was structural: incremental return, not platform-reported return, became the metric the media budget is built on.
That framework now governs roughly $42M in annual media spend, run as an always-on cadence of holdouts rather than a single test. And the two-sources-of-truth argument with finance simply went away, because both sides were finally looking at the same defensible number.
What this method can't tell you
Any measurement worth trusting comes with its boundaries stated up front. Here are this one's.
- Geo tests trade precision for validity. The confidence interval is wider than a platform dashboard pretends its number is. We reported a range alongside the point estimate, not false certainty.
- Incrementality is a point-in-time read. It saturates and decays as spend and creative change, which is exactly why this runs as an always-on cadence rather than a one-off verdict.
- The design resolves to channel level, not creative or audience level. Finer questions need their own tests; this result shouldn't be stretched to answer them.
- Matched markets are never identical. The synthetic control reduces residual bias but doesn't erase it, so the number is a strong estimate, not a measurement to three decimal places.
What I actually did
I owned the measurement end to end: designed the holdout and powered it against a minimum detectable effect, built the difference-in-differences and synthetic-control analysis in Python and SQL on BigQuery, and ran the MMM triangulation. I then took the findings and the reallocation recommendation to the CMO and CFO, and codified the whole thing as a standing capability. A small analytics team supported delivery.
Building an experimentation practice from zero →
Have a measurement question worth answering properly?
Attribution you don't trust, an incrementality program to stand up, or a role to fill. Happy to talk.
