Writing/Playbook/Experimentation
Playbook · Experimentation

Building an experimentation practice from zero

Prioritisation, power, and the cultural work of getting an organisation to ship the test before the opinion – no Centre of Excellence required.

AuthorKunal Mirchandani
PublishedJuly 2026
Read time8 min
SeriesMeasuring honestly

Every analytics team I've worked with had the same artefact somewhere in the shared drive: a document called "test backlog", last edited fourteen months ago. Everyone agrees experimentation matters. Almost nobody runs experiments.

The gap isn't knowledge – every marketer has read the same articles about A/B testing – and it isn't tooling, which has never been cheaper. The gap is organisational: nobody's first test ever survives contact with the calendar.

This is the playbook I use to close that gap. It assumes no programme, no platform, no dedicated headcount – just you, a budget, and one quarter of patience.

The first test

Start with a decision, not a hypothesis

The first test's job isn't to produce a number. It's to change a decision – and to be seen to change it. So work backwards from one. Somewhere in the next quarter, a budget conversation is already scheduled: a channel up for renewal, a campaign about to scale, a market about to get more money. That conversation is your test. Design the experiment that would change its outcome.

Two filters make the choice concrete. Big enough to matter: the answer has to move a line item someone will actually notice – a channel's quarterly budget, not a button colour. Small enough to run: it has to fit inside your authority. You don't need sign-off to hold back five per cent of an audience for four weeks; you very much need sign-off to restructure the media plan. Geo holdouts on a mid-size paid channel are the classic first test for a reason: real money, real causality, no engineering tickets.

A concrete first test I like: branded search. Almost every account has it, almost nobody has ever held it back. Switch it off in a handful of matched regions for four weeks and watch what happens to total conversions – not platform-reported conversions, total. Whatever share doesn't disappear was being harvested, not created. That single number has reset more budgets than any dashboard I've built.

What you avoid is the tempting "easy win" – the landing-page A/B test with a friendly tool and no stakes. It teaches the organisation that testing is a CRO hobby. First impressions stick: the first test sets the price of admission for every test after it.

The working rule

A test that can't change a decision is a hobby. Score every proposal by the decision it unblocks, not by how easy it is to run.

The queue

Build the backlog in public

Prioritisation is where most practices quietly die – not for lack of ideas, but for lack of a queue anyone trusts. So keep a single list, visible to the business, where every proposed test carries four fields: the decision it informs, the effect size worth detecting, the cost of running it, and the decision date it must land before.

Then score on decision value, not convenience. The question isn't "what can we test?" – that produces a queue of trivia. It's "what would we do differently if we knew?" Rank by the size of the decision attached and the date it's needed, and let the awkward, important tests beat the easy, empty ones. A proposal that can't name its decision doesn't get a slot; it gets a conversation. That one rule keeps the list short and the quality high.

The visibility is not decoration. When the leadership team can see the queue, two things happen: pet ideas have to compete on the same terms as everything else, and "why isn't my thing being tested?" stops being a corridor conversation. The backlog is where HiPPOs go to be prioritised.

The discipline

Design before results, every time

This is the part that separates a practice from a pile of anecdotes. Every test gets a one-page design, written before launch and shared with the people who will argue about the result afterwards:

The one-page design
  • The decision it informs, and the date it's needed by.
  • The minimum detectable effect – the smallest lift that would change the decision. Not the lift you hope for; the lift that matters.
  • The power calculation – sample size and duration derived from that effect, computed before launch, not explained away after. If you can't detect the effect that would change the decision, don't run the test. Spend the budget on one that can.
  • The design itself – treated and control groups, the holdout mechanism, duration in whole weeks (paydays and weekends corrupt partial ones), and the metric of record.
  • The pre-agreed actions – what you'll do with each possible outcome.

That last item feels bureaucratic and pays off the most. "If the lift is under two per cent, the channel loses a third of its budget" – agreed in advance – is a contract. Without it, every result gets negotiated after the fact, and the test becomes another opinion with a chart attached.

Write the one-pager in plain language. Its audience is the budget owner, not the data team – if they can't repeat the design back to you, it isn't finished. And publish the design before the result. It's the same rule as everywhere else in measurement: the assumption arrives before the number, or the number means nothing.

The politics

The cultural work

Here's the part the testing-platform case studies skip: the first test will kill something someone loves. A channel the performance team defends, a campaign the brand team is proud of, an agency's favourite line item. If the result is going to be survivable, the politics have to be handled before the data arrives.

Socialise the design, not the result. Walk stakeholders through the one-pager while there's still nothing to defend – people are remarkably reasonable about a method before it threatens them. Book the decision meeting before launch, so the result has somewhere to go. One more trick: give the result a number before it has one. "This test will tell us whether $400k of spend is creating demand or collecting it" – said at kickoff – makes the readout an event people attend.

And make the negative result safe. The first time a favourite comes back smaller than its dashboard, the room decides whether experimentation is welcome here. That moment is what the whole setup is for. Celebrate the killed channel, loudly – the money it frees up is the budget for every test that follows.

Field objections

The three objections you'll get

"We don't have enough traffic." Sometimes true for user-level A/B tests; almost never true for geo holdouts, which run on regions, not individuals. If a channel spends enough to matter, it spends enough to test. The objection is usually a political one in a statistical costume.

"The platform already tests for us." Platform lift studies are useful, fast – and graded by the referee. Use them as one input, never the final word; an independent holdout is what keeps them honest.

"We tried testing once and nothing was significant." That's not a failed test; that's an underpowered one. "No significant effect" from a test sized to detect 15% tells you nothing about a 3% effect. Power is decided before the test runs, not after – an underpowered test is a coin flip dressed as evidence.

The cadence

Keep score, and keep it running

A practice is a rhythm, not an event. Four things sustain it:

  • The log. Every test, one row: hypothesis, design, effect detected, decision taken, money moved. This is the scorecard – it turns "we run tests" into "testing moved $1.4M this year", which is the only sentence that defends the practice in a budget review.
  • The rhythm. One meaningful test per quarter beats five abandoned pilots. Quarterly planning sets the decision dates; the backlog fills them.
  • The retest. Lift has a shelf life – audiences saturate, creatives wear out. Last year's number is evidence, not entitlement. Put retest dates in the log.
  • The graduation. Once holdouts are routine, the next layer is always-on measurement: platform experiments where they're honest, mix modelling to fill the gaps, experiments calibrating the model. The practice you've built is exactly the discipline that layer needs.

None of this needs a Centre of Excellence. It needs one decision, one clean design, one result that moved money – and the willingness to do it again next quarter. The practice is the product: every test you ship makes the next one easier to ship. Start where the money is, publish the design, and ship the test before the opinion.

Next essay

The two-sources-of-truth problem, and how to end it →

Ready to ship the test before the opinion?

An experimentation practice to stand up, a measurement programme to repair, or a role to fill. Happy to talk.