What Is Incrementality Testing, and Why Your ROAS Cannot Answer It
Here is a conversation I have at least once a month. A founder shows me a brand search campaign reporting 12x ROAS. They want to know whether to double it. My question back: how many of those buyers typed your company name into Google because they already meant to buy?
Nobody in the room knows. Neither does the ad platform. It only knows a click happened before a purchase. That is correlation, and every attribution report is built on it.
Incrementality testing answers a different question: what would have happened if the ad had not run? You deliberately withhold ads from a comparable group, then measure the difference in outcomes. The gap is the incremental effect. Everything else is sales you would have made anyway.
This is not an academic nitpick. When eBay switched off brand keyword ads in a large field experiment, the researchers found that brand-keyword ads had no measurable short-term benefit, and that returns from paid search overall were "a fraction of conventional non-experimental estimates" (Blake, Nosko and Tadelis, Econometrica 2015). A second study using 15 large Facebook experiments concluded that commonly used observational methods often fail to accurately measure the true effect of advertising, even after controlling for extensive demographic and behavioral data (Gordon et al., Marketing Science 2019).
If you have read my breakdown of why attribution models disagree with each other, this is the missing piece. Attribution models split credit. They never tell you whether the credit was earned.
Three Ways to Run Incrementality Testing in Marketing
Incrementality testing in marketing comes in three practical flavors. They differ in who controls the randomization, what data you need, and how much you have to trust the vendor running the test.
| Method | How the control group is built | Best for | Main weakness |
|---|---|---|---|
| Platform lift study (Meta, Google) | Platform randomly withholds ads from a share of eligible users | Single-platform questions, fast setup | Platform grades its own homework; measures only its own ads |
| Geo experiment | You pause or change spend in selected regions and compare to matched regions | Cross-channel questions, offline sales, TV, retail | Needs enough regions and stable regional sales |
| Holdout (on/off) test | You switch a channel off entirely for a defined period or audience | Brand search, retargeting, small channels | Time-based noise; seasonality can swamp the signal |
Platform lift studies
Meta's Conversion Lift splits your eligible audience into a test group that can see ads and a control group that cannot, then compares conversion rates. Meta treats 90 percent or higher confidence as a statistically reliable lift result. Meta incrementality testing is the easiest entry point for most advertisers because the randomization happens at the user level inside the platform, which is statistically cleaner than anything you can build yourself.
Google incrementality testing works similarly. Google Ads Conversion Lift offers two study types: one based on users and one based on geography, and the geo version supports offline data. For the user-based version, Google's documentation states a minimum campaign budget of USD 5,000 and at least 1,000 observed conversions, and warns that studies with long conversion lag that run under 14 days showed up to a 17 percent drop in absolute lift. That minimum is new: Google announced in November 2025 that experiments which once cost upwards of USD 100,000 can now be done for USD 5,000.
For incrementality testing on Google Ads specifically, the geo-based setup covers Search, Shopping, Performance Max, Demand Gen, Display, Video and App, but campaigns must target a single country, and Google does not recommend proceeding when the feasibility check comes back "Low".
My caveat with both: the platform selling you the ads is also the one measuring them. The randomization is sound. But the conversion definition, the attribution window, and which conversions count all come from that platform's tracking. If your pixel double-fires, the lift study inherits it.
Geo experiments
A geo experiment treats regions as the unit of randomization. You increase, decrease, or pause spend in test regions and compare sales against control regions that historically move in step with them. Because you measure outcomes in your own sales data rather than in a pixel, geo tests work for channels where the platform cannot see the conversion: TV, podcasts, retail sales, phone orders.
Meta's open-source GeoLift package uses synthetic control methods to build the comparison group and includes power calculators for market selection before you launch. It is free and runs in R.
Commercial platforms such as Haus and Measured sell managed geo and holdout testing, with the design, market selection and readout handled for you. Whether that is worth paying for depends on how many tests you plan to run and whether anyone in-house can own the model.
Simple holdouts
The bluntest method: turn it off and watch. Pause brand search for three weeks. Exclude a random 20 percent of your customer list from retargeting. It costs nothing to set up, which is why it is often the first test I recommend for brand search and retargeting, the two places where attribution most reliably over-credits ads.
The weakness is time: if a competitor launches a promotion the same weeks you pause brand search, you cannot separate the two effects. Audience-split holdouts (random list exclusions) are far more trustworthy than before/after comparisons.
Can Your Budget Actually Power a Test?
This is where most tests fail before they start. When I scope a test for a client, the first question is not "which tool?" It is "what is the smallest lift you need to detect, and do you have the volume to see it?"
Statistical power depends on three things: your baseline conversion rate, the size of the effect you want to detect, and the number of people or regions in each group. The effect size dominates.
Hypothetical: the same audience, two different questions
Assume a retargeting audience converts at 2 percent without ads. Using the standard two-proportion sample size formula (95 percent confidence, 80 percent power, the same math behind Evan Miller's sample size calculator):
| Lift you want to detect | Conversion rate with ads | Users needed per group | Approx. conversions in control group |
|---|---|---|---|
| 30% relative lift | 2.6% | about 9,800 | about 196 |
| 10% relative lift | 2.2% | about 80,700 | about 1,614 |
Cutting the detectable effect from 30 percent to 10 percent requires roughly eight times the sample. On a channel you already run at scale, the lift worth arguing about is usually the small one, which means the test you actually need is the expensive one.
Practical rules I use when scoping:
- Set the minimum detectable effect from the business decision, not from hope. If you would only cut a channel whose lift is below 15 percent, design the test to detect 15 percent. Nothing smaller matters.
- Run at least two full purchase cycles. Google's own guidance on conversion lag above is a good floor: under 14 days, you undercount delayed conversions.
- Test the whole channel before testing tactics. Detecting whether Meta works at all needs less power than detecting whether Advantage+ beats manual campaigns.
If the power calculation says you need five months of holdout on a channel that spends EUR 3,000 per month, the honest answer is that you cannot test that channel in isolation. Group it with others, or test a bigger spend change.
This is also the stage where broken tracking kills tests quietly. If your conversion counts differ between GA4, the ad platform and your backend (a problem I cover in why GA4 and Google Ads conversions do not match), your baseline is wrong and so is every power estimate built on it. Before I design any lift test, I run a measurement audit to confirm the conversion data is trustworthy. Otherwise you are measuring the lift of your tracking bugs.
How Results Should Recalibrate Attribution and MMM
A lift result sitting in a slide deck changes nothing. The point of incrementality testing is to correct the systems you use for daily and quarterly decisions.
Recalibrating attribution with an incrementality factor
The simplest method: divide incremental conversions from the test by the conversions the platform attributed over the same period. That ratio becomes a multiplier for reported performance.
Hypothetical example: Meta reports 1,000 attributed purchases during a four-week test. The lift study finds 400 incremental purchases. Your incrementality factor is 0.4. A reported 4x ROAS on that channel becomes an incremental ROAS of 1.6x. If your break-even ROAS is 2.5x, you were scaling a channel that loses money on the margin.
Apply that factor in reporting, and consider it in bidding targets. Combined with profit on ad spend instead of revenue ROAS, it gets you close to the number your CFO actually cares about. Factors drift, so I re-test major channels every six to twelve months or after big changes in creative strategy or spend.
This matters most if you have invested in a multi-touch attribution model. MTA distributes credit more intelligently across touchpoints, but it still distributes credit for sales that may have happened anyway. Incrementality factors are the correction layer on top.
Calibrating marketing mix models
MMM and experiments are designed to work together. Google's Meridian lets you set custom ROI priors from past experiments, and Meta's Robyn supports calibration against lift studies and geo tests such as Conversion Lift and GeoLift.
I covered the modeling side in marketing mix modeling for mid-size budgets. The short version: an uncalibrated MMM will happily assign credit to whichever channel's spend happened to rise with seasonal demand. One well-run geo test on your largest channel constrains the model more than another year of data.
The sequence I recommend
- Fix tracking first. Conversion counts must reconcile with your backend within a tolerance you can explain.
- Run a cheap holdout on the most suspicious channel. Usually brand search or retargeting.
- Run a platform lift study on your biggest paid social channel, if you meet the volume thresholds.
- Move to geo testing once you need cross-channel answers or you sell offline.
- Feed results into attribution factors and MMM priors. Then re-test on a schedule.
FAQ
What is incrementality testing in simple terms?
Incrementality testing is a controlled experiment that measures how many sales your ads caused, as opposed to how many sales they were credited with. You withhold ads from a comparable control group and compare its results with the group that saw ads. The difference is the incremental effect of the advertising.
How is incrementality testing different from attribution?
Attribution assigns credit for conversions to the touchpoints that preceded them, regardless of whether those touchpoints changed the outcome. Incrementality testing uses a control group to estimate what would have happened without the ads. Attribution is useful for daily optimization, while incrementality tells you whether the credit is deserved.
What budget do I need for a Google Ads Conversion Lift study?
Google's documentation for user-based Conversion Lift lists a minimum campaign budget of 5,000 US dollars and at least 1,000 observed conversions. Geo-based studies use a feasibility check instead of a fixed floor, and Google advises against running studies rated Low.
Should I use a platform lift study or a geo experiment?
Use a platform lift study when the question concerns a single platform and you meet its volume requirements, because user-level randomization is statistically efficient. Use a geo experiment when you need to measure several channels together, include offline sales, or avoid relying on the platform's own conversion tracking.
How long should an incrementality test run?
Long enough to cover at least two full purchase cycles and to reach the sample size your power calculation requires. Google reports that studies with long conversion lag that run under 14 days can understate lift. Avoid launching during major seasonal peaks unless those peaks are what you want to measure.
Not sure your conversion data is clean enough to run a lift test, or what to do with the result once you have it? Book a marketing measurement audit and I will tell you which channel to test first, whether your volume can support it, and how to feed the answer back into your reporting.