Aug 5, 2026 · 4 min read · GameMantra Team
Control groups without giving up the revenue they cost
Holding players back from a change is the only way to know it worked. The objection is that it costs money, and that objection is mostly answerable.
A control group — a set of players deliberately excluded from a change so you can compare — is the only reliable way to know whether the change did anything. Without one, you are comparing this week to last week and hoping nothing else moved.
The standard objection is that the control group is revenue you chose not to earn. That objection is real, frequently overstated, and mostly manageable.
What the control actually costs
The intuitive framing is that if a change increases revenue and you exclude ten percent of players from it, you lose ten percent of the increase.
That is the correct arithmetic and it is a much smaller number than it sounds, because it is ten percent of the improvement, not of revenue. If a change improves things modestly, ten percent of a modest improvement is a small amount of money.
Set against that is the cost of not knowing. A change that appeared to work and did not gets kept, built on, and extended. The cost of running a worse game for a year because a change was misattributed is generally much larger than the foregone portion of one improvement — and it compounds, because the next decision is made on top of the wrong conclusion.
The comparison that makes this concrete is between the cost of a control group and the cost of one wrong strategic conclusion. Most teams that have had a wrong conclusion can name what it cost, and it is rarely small.
Where a control costs more than it is worth
There are cases where the objection holds and the honest answer is not to run one.
If the change is a fix for something clearly broken, a control group means deliberately leaving some players with the broken thing. That is a bad trade in almost every case — you do not need a controlled measurement to know that fixing a crash helped.
If the change is required — a compliance obligation, a platform requirement — there is no choice to measure. Excluding a group is not an option regardless of what it would teach you.
And if the population is small enough that the comparison will be inconclusive anyway, the control costs revenue and produces nothing. Being honest about this in advance is better than running an underpowered test and reading the noise.
Outside these, the argument for skipping a control usually reduces to not wanting to wait, which is a schedule preference rather than an economic one.
Making the control cheap
Several adjustments reduce the cost without weakening the measurement much.
Make it small. A control group does not need to be half your players. It needs to be large enough to detect the size of difference you care about, which for most changes is a modest fraction. Teams frequently default to a fifty-fifty split out of habit and pay far more than necessary.
Make it temporary. Once you have a clear answer, roll the change out to everyone. The cost is bounded by the measurement window rather than running forever. Controls that quietly persist for months are a common source of unnecessary cost.
Rotate who is in it. If different players are held back from different changes, no individual player experiences a consistently worse game, and the cost is spread rather than concentrated on an unlucky group.
The one thing not to do is choose the control group by anything correlated with behaviour. A control drawn from a particular market, platform, or acquisition source is not comparable to the treatment group, and the comparison measures the difference between those populations rather than the effect of the change.
See how we structure comparisons in a live game →
The permanent version and why it exists
Separately from per-change testing, some games keep a small group permanently excluded from a category of change — typically automated or personalised decisions — to measure the whole system rather than any individual change.
This answers a different question. Per-change tests tell you whether each change helped. A permanent comparison tells you whether the accumulated effect of the whole approach is positive, which is not the same thing and is not reconstructible from the individual results.
That group should be small, chosen consistently so the comparison stays valid over time, and left alone. Its value comes entirely from being untouched, and every temptation to include them in something "just this once" destroys the measurement it exists to provide.
The reason this matters more than it seems is that a series of individually-positive changes can add up to something worse — each one helped in isolation, and together they produced a game that asks too much or feels too commercial. Only a long-running comparison catches that, and by the time it would show up in aggregate metrics, the cause is untraceable.
Share this post
See what this looks like for your game.
SDK for Unity and Unreal. A 20-minute call to walk you through it.