Four Designs for Testing a Change You Cannot Split by User, and When Each One Misleads You

Four Designs for Testing a Change You Cannot Split by User, and When Each One Misleads You

Some of the changes with the largest revenue consequences cannot be split by user. A free-shipping threshold, a price list, a courier mix, a radio flight, a change to who pays for returns: they apply to everyone or to nobody. Teams usually either skip measurement or run a before-and-after and call the difference a lift.

Both are expensive. In one published geo-experiment case study, the standard ratio estimator put incremental return on ad spend in a 95% interval of [-1.26, 5.69]; a trimmed estimator on the same data returned [0.25, 1.74] (Chen and Au, 2019). One of those contains zero and tells you nothing.

This guide is for growth, CRO and analytics leads who must put a number on a change they cannot randomize per user. By the end you will be able to pick between four designs, know what each needs before it works, and know which will mislead you where demand concentrates in one city.

The four designs in one sentence each

Randomized paired geo experiment (geo holdout). Split geographic units into treatment and control after matching them on pre-period response, then compare the arms. Vaver and Koehler (2011) rank geos by their pre-test metric, partition that list into groups of size M, and randomly pick one geo per group for treatment.

Switchback (time-split). Apply the change business-wide, switch it on and off in blocks, and treat the blocks as randomized units. Bojinov, Simchi-Levi and Zhao (2023) derive optimal designs under assumptions about how long an effect lingers after a switch, plus a data-driven way to estimate that carryover length.

Synthetic control. Build a weighted combination of untreated markets that tracks the treated market's pre-period behaviour, then read the gap after launch. This is Abadie, Diamond and Hainmueller (2010), and the engine inside Meta's GeoLift.

Interrupted time series. Model what the series would have done and read the difference, with no control group. The Bayesian structural time-series version is Brodersen et al. (2015). It is the weakest of the four, and the one most often used by accident: "compare to last month" is this design with the modelling left out.

Compare them on eight criteria

CriterionGeo holdoutSwitchbackSynthetic controlInterrupted time series
Units neededMany comparable geos (source example: 210 US DMAs)One business, many time blocks1-5 treated markets plus a donor pool (GeoLift: 40 cities)One series
Pre-period dataA pseudo pre-period matching test length (14 days in the example)Enough history to estimate carryoverGeoLift's example uses 90 days of daily dataLong, stable history
Survives a national shockYes, it hits both armsNo, a shock inside a block becomes the effectMostly, if donors share itNo
Tolerates one dominant marketPoorlyYes, geography is irrelevantYes, its main use caseYes
Spillover riskReal: cross-border shopping, national mediaNone geographic, high temporalReal, same as geo holdoutNot applicable
Realistic minimum detectable effectMid single digits with enough geosSmall, if the effect is fastGeoLift's default power grid starts at 5%Usually double digits
Time to a usable readWeeks plus response lagDays to weeksWeeksWeeks, never conclusive
What it cannot tell youPer-user behaviour; it tracks no individualsSlow-building effects such as repeat rateWhether the fit is right rather than closeWhether anything else changed too

When each one wins

  • Geo holdout when the change can be targeted geographically, you have a few dozen comparable units, and you need a defensible number for a budget argument. It is the only one of the four with real randomization, so the only one that protects you from biases you have not thought of.
  • Switchback when the change is global by nature (a pricing rule, a search algorithm, a courier allocation) and the effect is fast. It fails once the effect takes longer to appear than your block length.
  • Synthetic control when you only get to change one or two markets, usually because a launch or a regulation forces the split. It is observational: you are asserting the donor pool would have tracked the treated market, not demonstrating it.
  • Interrupted time series only when no comparison unit exists anywhere, reported as an estimate with an explicit list of what else moved.

If the change rolls out to different markets at different times, you are in staggered difference-in-differences territory, where the naive two-way fixed-effects regression is biased. Use the estimators in Callaway and Sant'Anna (2021).

Why regional market structure breaks the default geo split

Geo designs assume demand spread across many units of broadly similar size. In Turkey it is not. The Ministry of Trade's report on 2025 e-commerce puts total volume at 4.57 trillion TL across 5.94 billion transactions, and for the November 2025 campaign period on marketplaces reports Istanbul at 31% of volume by buyer location and 67% by seller location (T.C. Ticaret Bakanligi). Note the limits: one campaign month, marketplaces only. It is the right order of magnitude, not a constant.

That number decides your design. With 81 provinces but roughly a third of demand in one, a random province split produces a result mostly determined by which arm Istanbul landed in. Three responses:

  • Exclude the dominant unit from randomization and analyse it separately. Say so; the result then generalizes to the rest of the country, not the whole book.
  • Split it into sub-units (districts, courier zones) so it contributes several units. Only works if platform and analytics can both target and report at that grain.
  • Roll provinces up into the statistical regions your reporting already uses: Turkey's IBBS Duzey 2, or NUTS-2 for EU markets. Fewer balanced units beat many unequal ones.

The same applies wherever demand concentrates in one metro. A design borrowed from US DMA examples will not transfer.

A worked example: moving the free-shipping threshold

  1. Write the decision first. "If contribution margin per session does not fall more than 2%, we roll it out." A design that cannot detect 2% is the wrong design, and you now know that before building it.
  2. Count your units. After excluding Istanbul and merging provinces below a volume floor, suppose 24 usable regions remain. Too few for a paired geo test on a 2% effect. That is the answer, not a setback.
  3. Switch design. A threshold is a pricing rule, so switchback fits: alternate it weekly site-wide over eight blocks, estimating carryover from how long basket size takes to settle after each switch.
  4. Name the guardrail. Order rate and average basket value will move in opposite directions. Pick the primary metric before launch, not after.
  5. Check the schedule for confounds. If a campaign week lands inside a block, that block is contaminated. Drop it rather than modelling around it.
  6. Verify the plumbing first. Confirm with a one-day pilot that platform and analytics both target and report at the chosen grain, and extend the measurement window past the change by the response lag.

[INTERNAL DATA NEEDED: Switas's own distribution of order volume by province for a representative e-commerce client, so "24 usable regions" can be replaced with a real count and volume floor.]

Where this breaks down

None of these designs measures a per-user effect. They measure what happened to a market or a time block. If the question is "which segment responded", a geo design cannot answer it, and no one should be reading a segment split out of one.

Switchback inference depends on the carryover assumption, and the source paper is explicit that misspecifying carryover order degrades the result. If you cannot estimate it from history, you are guessing.

Synthetic control has no randomization, so its validity rests entirely on pre-period fit, and good fit is necessary rather than sufficient. GeoLift's documentation notes its scaled L2 imbalance metric is not comparable across different KPIs or numbers of periods, so a reassuring number in one study means nothing in another. The same docs state GeoLift is "not a Meta product. For research purposes only."

All geo designs assume accurate geo-targeting and limited cross-boundary movement. Cross-border shopping, VPN use and national broadcast violate that, and there is no clean correction: bound the leakage and say so. This guide does not cover user-level A/B testing or interference between concurrent tests.

FAQ

How many geo units do I need for a geo holdout test?

There is no published minimum, and the primary sources deliberately avoid giving one. Vaver and Koehler recommend simulating the expected confidence-interval width for your actual candidate geo set and test fraction. In practice the answer is usually "more than you have", which is why switchback and synthetic control exist.

Can I just compare this month to last month?

That is interrupted time series with the modelling omitted: it attributes every seasonal, competitive and macro movement in the window to your change. If it is all you have, use a structural time-series model, report an interval rather than a point, and list what else moved.

Does stratifying geos before randomizing actually help?

Yes, measurably. Vaver and Koehler report that grouping geos by size before assignment reduced the confidence interval for return on ad spend by 10% or more against unconstrained randomization. It is close to free, so do it.

Why would a trimmed estimator beat the obvious one?

Geo-level data is heavy-tailed when you only have a few dozen units, so one badly matched pair can dominate the estimate. In the Trimmed Match simulations with 50 pairs and log-normal geos, root mean squared error was 1.96 for the trimmed estimator against 20.09 for the standard ratio estimator, and power 60% against 16%.

How long should a switchback block be?

Long enough that the effect has settled before the block ends, so it is set by carryover length rather than convenience. Estimate carryover from how long the metric takes to stabilize after a historical switch, then make blocks longer than that. Weekly blocks are a common retail start because they also absorb day-of-week effects.

How do I handle a city that is a third of my demand?

Decide before you randomize: exclude it and scope the conclusion to the rest of the country, split it into district or courier-zone units if your stack supports that grain, or use synthetic control with that city as the treated market. All three are defensible. Randomizing 81 provinces as if they were equal is not.

Run this on your next untestable change

Work your next untestable change through the eight criteria above and see which design survives. If the answer is "none, with the units we have", that is worth knowing before you spend a quarter on it. Our CRO and experimentation team designs these tests for e-commerce, travel and health clients in Turkey and the region, including the cases where user-level randomization is off the table.

[INTERNAL DATA NEEDED: one anonymized Switas geo or switchback engagement, with the design chosen and the decision it closed.]

Sources


Inci Dindar
Written by

Inci Dindar

With a background in Software Development, she led design, UX, and product teams across fintech, media, e-commerce, and travel. For the past three years she has consulted full-time on UX and design — including as an external UX consultant for a global management consultancy helping brands and startups with audits, conversion optimization, and design processes.


Related Articles

Switas As Seen On

Magnify: Scaling Influencer Marketing with Engin Yurtdakul

Check Out Our Microsoft Clarity Case Study

We highlighted Microsoft Clarity as a product built with practical, real-world use cases in mind by real product people who understand the challenges companies like Switas face. Features such as rage clicks and JavaScript error tracking proved invaluable in identifying user frustrations and technical issues, enabling targeted improvements that directly impacted user experience and conversion rates.