A holdout group is a randomly selected portion of the audience that is deliberately excluded from a change, a campaign, or an entire program of changes, so that the cumulative effect on that group can be compared against everyone else. Where a standard experiment compares two versions of one thing for a few weeks, a holdout measures the combined impact of many decisions over months, answering a question that individual tests cannot: did all of this work actually move the business?
The need arises from a well-known accounting problem in optimization and marketing programs. Individual experiments report individual lifts, and those lifts are summed into an annual impact figure that is presented to leadership. But the sum is almost always larger than reality, because winning effects overlap, decay, interact, and are frequently overstated by early stopping or by the winner's curse. A long-running holdout gives a single, direct measurement of the whole program's contribution, which either validates the reported figures or reveals a gap that needs explaining.
Holdouts take several forms. A post-launch holdback keeps a small percentage of users on the control after a winning variant is rolled out, so that the effect can be tracked for months and checked for decay. A program-level holdout excludes a percentage of users from every optimization change for a quarter or longer. A marketing holdout suppresses campaign contact for a random subset to measure incremental rather than attributed response, which regularly shows that a substantial share of credited conversions would have happened anyway. Each answers a different question, and each requires the excluded group to be selected randomly and kept stable over the measurement period.
The design considerations are demanding. The group must be large enough to detect the expected cumulative effect, which is usually easier than for a single test because the effect being measured is larger. Assignment must persist across sessions, devices where possible, and channels, otherwise contamination erodes the comparison. The exclusion must be genuinely enforced, which is harder than it sounds when a dozen teams are shipping changes through different mechanisms. And the organization must accept the deliberate cost of withholding improvements from real customers for an extended period, which is a governance decision rather than an analytical one.
That cost is the main objection, and it deserves a straight answer. Holding out five percent of traffic from a program that genuinely improves conversion does forgo some revenue. The counter-argument is that without a holdout, nobody knows whether the program improves conversion at all, and the amount being spent on it is typically far larger than the forgone revenue from the excluded group. Organizations that have run their first program-level holdout frequently find the measured effect materially smaller than the summed test results implied, which is uncomfortable but is precisely the information needed to allocate budget sensibly.
The internal conversation about a holdout usually goes better when the cost is quantified rather than left as an abstraction. If the programme's estimated effect is a given percentage improvement, then withholding five percent of traffic forgoes five percent of that improvement, which is typically a small figure next to the cost of the programme itself and trivially small next to the risk of continuing to invest in something whose contribution is unknown. It also helps to frame the group as an ongoing measurement instrument rather than a deprivation, since the same mechanism supports post-launch decay monitoring and provides the baseline for any subsequent claim about programme value. Where withholding improvements raises fairness concerns, as it can in services with statutory or contractual obligations, the appropriate response is a smaller group, a shorter period, or a geographic design rather than abandoning measurement entirely.
Establishing a holdout is usually one of the more consequential recommendations in a growth management engagement, because it changes the basis on which the whole program is judged from a stack of individual reports to a single measured outcome. Implementation normally requires the audience segmentation and identity resolution capability maintained by the data analytics function, since a holdout that leaks across channels or devices measures nothing reliable at all.