Jun 23, 2026 · 5 min read · GameMantra Team
Incrementality testing: the truth layer attribution lacks
When attribution dashboards answer the wrong question, incrementality testing measures real causal lift. Here is when geo-holdouts beat last-click reports
Your attribution dashboard can look completely right and still be answering the wrong question. It tells you which channel was closest to a conversion. It does not tell you whether that conversion would have happened anyway. Those are different facts, and in 2026 the gap between them is where a lot of marketing budget quietly disappears.
This is the core problem with attribution as a decision tool. It tends to over-credit whatever sat near the moment of conversion — brand search, retargeting, the last ad a user happened to see. A player who was already going to install your game clicks one more ad on the way in, and that ad gets the credit. Cut the ad, and the installs barely move. Attribution said the channel worked. Incrementality says it mostly took credit for demand you already had.
What incrementality actually measures
Incrementality testing measures causal lift against a credible counterfactual. Instead of asking "which touchpoint was last," it asks "what happened to the people we reached versus comparable people we didn't?" The difference between those two outcomes is the lift your spending actually produced.
The cleanest version of this idea is one you already know from product experiments: a control group. You withhold a channel or a campaign from one set of users and run it for another, then compare. If the exposed group converts more, the gap is your incremental effect. If the two groups look the same, the channel was painting over demand that existed regardless.
This is the same logic that makes a control group the backbone of trustworthy measurement anywhere — you can only know what your intervention did if you have a comparable group that didn't get it. Attribution has no counterfactual. It watches conversions happen and assigns blame after the fact. Incrementality builds the comparison in on purpose.
Why geo-holdouts became the practical tool
For most studios, the realistic way to run an incrementality test on user acquisition is a geo-holdout. You pick comparable regions, run your campaign in some and deliberately withhold it in others, and compare results across the two groups. Geography gives you a clean way to split the audience that survives the privacy changes that broke user-level tracking.
This matters because the old measurement spine is gone. With signal loss compounding across platform opt-outs and aggregated attribution, a large share of conversions never tie back to a user-level source. Geo-based experiments are one of the few methods left that produce a genuinely credible read on whether media is working, because they don't depend on tracking individuals — they compare populations.
The discipline is worth stating plainly: a geo-holdout establishes whether a channel deserves investment at all. That is a yes/no question about causation, and it is the question attribution can't answer.
Where geo-holdouts go wrong
Incrementality is not magic, and the failure modes are specific enough to plan around.
The biggest one is budget overflow. When you remove your test regions from a campaign's targeting, the budget that was going there doesn't vanish — it overflows into the control regions, which then spend more than they should. Now your control group is contaminated by extra spending, and the comparison is broken before it starts. This is one of the most common reasons geo-holdouts produce results you can't trust. Designing the test so freed budget is genuinely held back, not redistributed, is half the work.
The second failure is treating one test as a permanent answer. Markets move, creative fatigues, competition shifts. An incrementality read is true for the window you measured it in, not forever. The studios that get value from this run holdouts as a recurring discipline, not a one-off audit they cite for the next year.
The third is confusing two different tests. A holdout answers "does this channel cause lift?" A scale-up test — running a channel harder in some regions than others — answers a separate question: "is there room to spend more here profitably?" The defensible approach treats them as sequential. First the holdout tells you the channel deserves money. Only then does a scale-up test tell you how much. Running the scale-up question first, before you've confirmed the channel even works, gives you a confident answer to the wrong thing.
How to fit this into a real measurement stack
Incrementality doesn't replace your attribution dashboard. It corrects it. The practical pattern is to keep attribution for day-to-day pacing — it is fast, granular, and good enough for "is this campaign delivering installs at roughly the cost I expected." Then run periodic incrementality tests as the truth layer that tells you which channels are actually causing growth versus which are billing you for demand you already owned.
When the two disagree — and they will — incrementality wins, because it has a counterfactual and attribution does not. The channels that look great in last-click and flat in a holdout are exactly the ones to scrutinise. They are usually the ones sitting closest to conversion, harvesting intent rather than creating it.
This same instinct — measure against a real comparison group rather than trusting a number that looks plausible — runs through everything on the monetisation side too. A reported uplift means nothing without a group that didn't receive the change to compare it against. If you want to see how we apply that principle to AI-served offers rather than ad channels, you can read how it works: the players who get an AI offer are always measured against a comparable group who don't, so the lift you see is the lift that's real.
The takeaway: stop asking your attribution dashboard a question it can't answer. It is a good map of where conversions land. It is a bad judge of what caused them. For that, you need a comparison group — and a geo-holdout, run carefully and repeatedly, is the most credible one most studios can build.
Share this post
See what this looks like for your game.
SDK for Unity and Unreal. A 20-minute call to walk you through it.