The only method that watches what happens without you

Incrementality testing card showing 2.4 percent, the randomised benchmark one observational method reported as 1,306 percent
Contents 8 sections

Incrementality testing withholds your advertising from a comparable group and watches what they do anyway. It is the only method in common use that measures what your advertising caused, rather than counting the outcomes it happened to be nearby for.

Everything else in this cluster describes ways of dividing credit. This describes a way of establishing whether there was credit to divide.

This is for you if you have ever been asked what would have happened without the adverts and had no good answer.

What incrementality testing actually does

It creates a comparison that does not otherwise exist.

Your normal reporting sees only people who were shown adverts. Some of them converted, and you have no observation of what those same people would have done in a world without the advertising, because that world was never created.

An incrementality test creates it deliberately. A comparable group, usually defined by geography or by a randomised split, is not shown the campaigns. Everything else stays the same.

Then you compare. The difference between the two groups is the lift your advertising produced, and it is an observation rather than an assumption.

That is the entire idea, and its simplicity is why it remains the standard against which other methods get judged.

How far the other methods drift

There is one comparison worth knowing in detail, because it puts a number on something usually discussed in the abstract.

Gordon, Zettelmeyer, Bhargava and Chapsky compared observational attribution against randomised experiments across 15 Facebook experiments, 500 million user-experiment observations and 1.6 billion impressions.

In half the studies, the observational estimate was off by a factor of three.

One case is worth stating on its own. In a checkout study, the observational method produced a 1,306% lift where the randomised benchmark was 2.4%. Not a margin of error. An answer of entirely the wrong magnitude.

The data is from 2018 and the mechanism has not changed, so treat magnitudes as indicative and the direction as solid. Every attribution model in common use is an observational method, which means this finding applies to all of them.

Why incrementality testing is not standard practice

Two costs, and both are real rather than excuses.

The first is direct. You stop advertising to part of your market on purpose, so you forgo whatever that group would have bought. That is a genuine loss and somebody has to approve it.

The second is scale. The difference between the two groups has to be large enough to read above normal week-to-week variation. Small accounts do not produce enough conversions for that difference to emerge from the noise, whatever the budget.

So incrementality testing is available to advertisers who have both money and volume, which is a smaller group than the number of people recommending it.

Being honest about that is more useful than pretending it is universally accessible, and it points at what smaller advertisers should do instead.

The underpowered test problem

When these go wrong, they usually go wrong in the same way, and the failure is invisible to most readers.

A test runs. It finds no significant difference. Somebody reports that the advertising is not incremental.

Frequently the test was never capable of finding a difference. If your holdout group is expected to produce a handful of conversions either way, then no result it returns can distinguish a real effect from nothing.

The check is arithmetic and it takes minutes. Work out the base conversion rate, apply it to the size of the withheld group, and ask how many conversions you would expect under no effect. If that number is small, the test cannot answer the question and should not be run in that form.

This applies to tests you commission and to studies people cite at you. A null result from a small sample is the most misread thing in this field.

What smaller advertisers can do instead

I have never run a geographic holdout test. On the accounts I work with, the volume is not there and the cost of withholding would not buy a readable answer.

What I do instead is geographic optimisation, which answers a narrower question honestly. Test locations, watch what each produces, cut the ones that do not work, and concentrate spend where demand and fulfilment actually line up.

That is not incrementality testing and I would not present it as such. It tells you which places perform, not whether the advertising caused the performance.

One worked example of what it costs to learn: on a US flag and banner ecommerce account, an ad set targeting a neighbouring state ran at 0.25 return on ad spend, three purchases at roughly $341 each, about $1,027 spent before it was dropped. That ad set had the highest reach in the account at 17,086 and the most link clicks at 542. Every traffic metric said it was winning. Revenue said it was the only loser.

Roughly a thousand dollars to learn that geography was the constraint. A holdout test answering a bigger question would have cost considerably more.

Reading a test somebody else ran

Three questions, in order, before you accept any incrementality result.

How big was the withheld group, and how many conversions would it be expected to produce under no effect? That establishes whether the test could have found anything.

Who ran it, and what do they sell? Platforms measuring their own effectiveness and vendors selling measurement both have a position. Neither is disqualifying and both belong in your reading.

And what exactly was withheld? A test that pauses one campaign type while others keep running measures something much narrower than the headline usually claims.

Those three questions dispose of most of what circulates as incrementality evidence.

Why incrementality testing keeps not happening

There is a pattern worth naming, because it explains why this remains rare despite everybody agreeing it is correct.

The result is frequently unwelcome. A test that finds your advertising is less incremental than your attribution claimed is a test that makes somebody’s last two years of reporting look generous. Nobody commissions that eagerly.

The cost is also visible while the benefit is not. Withholding advertising has an obvious price that appears in this quarter’s revenue. Better decisions arrive later and are hard to attribute to the test that enabled them.

And the honest version needs pre-registration. Deciding afterwards what counts as success is how a null result becomes a story about something else, and agreeing the criteria in advance removes the escape route people quietly rely on.

None of that is an argument against running one. It is the reason to write down what you will conclude from each possible outcome before the test starts, and to get that agreed by whoever will read the result.

What to change this week

Three steps.

Stop treating any attribution figure as a measure of what your advertising caused, and start describing it as what it is, a division of credit.

Then, if somebody proposes an incrementality test, do the power arithmetic before agreeing the budget. Base rate times holdout size tells you whether the test can produce an answer.

Then, if you are too small for a proper test, run geographic optimisation instead and be clear with yourself about the narrower question it answers.

View through conversions covers an experiment that failed exactly the power check described here, attribution models covers the methods this one is measured against, what conversion numbers actually measure is the counting underneath, and why cost per lead is rising shows what happens when market costs move underneath your measurement. What I do across paid, SEO and the website treats measurement as one problem. To work out what is actually testable on your account, book a teardown.

Frequently asked questions

What is incrementality testing?

It withholds your advertising from a comparable group of people or places, then compares what they do against the group who saw it. The difference is the effect your advertising actually caused, rather than the outcomes it happened to be present for.

How is incrementality testing different from attribution?

Attribution divides credit among the adverts somebody saw before converting. It never asks whether they would have converted anyway. Incrementality testing answers that question directly by observing a group who were not advertised to.

How much do the two methods disagree?

Substantially. A published comparison found observational estimates off by a factor of three in half the studies tested. In one checkout study the observational method reported a 1,306% lift where the randomised benchmark was 2.4%.

Why does not everybody run incrementality testing?

Because it costs money and volume. You deliberately stop advertising to part of your market, so you give up the sales that group would have produced, and you need enough scale for the difference to be readable above normal noise.

What is the minimum size for a useful test?

Larger than most advertisers expect, and it depends on your conversion rate and effect size rather than on a fixed budget number. The arithmetic to check is straightforward, and skipping it is how underpowered tests get run and misread.

What is an underpowered test?

One too small to detect the effect it is looking for. It usually returns no significant difference, which then gets reported as proof the advertising does nothing. That conclusion is unsupported, and the test never had the ability to find anything.

Can I run incrementality testing on a small account?

Usually not properly, and knowing that is worth more than running one badly. On smaller accounts the honest alternative is geographic optimisation: test locations, cut the ones that do not work, and accept you are answering a narrower question.

Do the platforms offer incrementality tools?

Some do, and they are measuring their own effectiveness, which is worth remembering. The independent vendors selling these tests have their own interest in the answer being that attribution overstates. Neither is disqualifying and both belong in your reading of the result.

Want a second pair of eyes on your account?

Book a free teardown

Book a free teardown

Send me your worst-performing campaign.

I will tell you what is wrong on the first call. No obligation, no hard sell, and I will say honestly if I am not the right fit.

Book a free 20-minute account teardown

Replies within one business day, Lahore time.