The Exclusion That Silently Failed
A holdout is a promise that a group of people did not see the campaign. Almost nobody verifies the promise.
The usual sequence: markets get assigned, the exclusion list goes into the ad platform, the test runs for six weeks, and the analysis comes back with a lift of about two percent and an interval that comfortably spans zero. Everyone concludes the channel does not work, or that the test was underpowered, and moves on.
Sometimes that is right. Often what actually happened is that the exclusion did not hold. A campaign got duplicated during a creative refresh and the copy did not inherit the geo exclusions. An automated bidding strategy expanded targeting to hit its volume goal. Someone launched a new campaign mid-flight and applied last quarter’s location list. In every one of those cases the held-out group received advertising, the measured gap shrank toward zero, and the test produced a confident wrong answer rather than an obvious error.
That is the dangerous part. A leaking holdout does not fail loudly. It returns a small, plausible, well-behaved estimate, biased toward the conclusion that the channel does nothing.
The fix is unglamorous. Pull delivery by market, weekly, for the whole test window, and confirm that held-out markets received zero impressions. Not “low.” Zero. It takes about ten minutes a week. On a test that costs six weeks and a few hundred thousand dollars of withheld spend, those are the highest-return ten minutes in the project.
Two habits that make it stick:
- Check in week one, not at the end. A leak caught in week one costs you a restart. A leak caught in analysis costs you the entire test.
- Log the check. Someone will question the result six months later, because incrementality estimates are always unwelcome to somebody. The delivery audit is what makes the finding defensible.
The broader point generalizes past geo tests. The gap between the experiment you designed and the experiment that ran is where most measurement credibility is lost, and it is almost never visible in the output of the analysis. The numbers will look fine. They will just be describing a different experiment than the one you think you ran.