Data & Attribution

Geo Holdouts: The Most Practical Causal Test

9 min read
Geo Holdouts: The Most Practical Causal Test

Attribution arguments end when someone runs an experiment. Geo testing is the version of experimentation that most teams can actually execute: split the country into regions, change spend in some, hold others constant, and compare outcomes. No cookies, no consent dependency, no platform cooperation required.

Choosing the regions

The single most important step is matching. Pick control regions whose historical weekly sales track the test regions closely over the previous year — correlation of the trend matters far more than matching population or demographics. Tools exist for synthetic control construction, but a careful manual match on trend beats a careless automated one.

DesignBest forMain risk
Turn off in test regionsProving a channel works at allLost revenue during the test
Increase in test regionsFinding saturation pointsNeeds a larger budget
Staggered rolloutLaunches you were doing anywayConfounding with time trends

Duration and carryover

Run for at least four weeks, and longer if your purchase cycle is long. Advertising keeps working after it stops, so a two-week off-test will understate the effect — the first week largely measures residual demand from prior exposure. Include a post-period to observe recovery, which is itself informative.

The most common geo test error is stopping early because the first week looked flat. The first week is supposed to look flat.

Be honest about power

Geo tests have small effective sample sizes — you have a few dozen regions, not a million users. That means they can detect whether a channel contributes meaningfully, but not whether efficiency improved by six percent. Design the test around a question worth a large answer.

Write the analysis plan first: which regions, which metric, which comparison window, and what result would change the budget. Deciding afterwards invites motivated reasoning.

Turn results into a factor

Compare the measured incremental effect with what your attribution reported for the same channel and period. The ratio is a calibration factor you can apply to ongoing reporting until the next test. That is how a one-off experiment keeps paying for itself.

Related articles

The Quarterly Data Quality Audit
Data & Attribution

The Quarterly Data Quality Audit

Most bad decisions made from data are not made from the wrong analysis. They are made from the right analysis on broken inputs.

8 min read