If you check an A/B test every morning and stop it the first time the dashboard turns green, your real false positive rate is not 5%. In a simulation testing significance after every observation, Evan Miller measured it at 26.1%, more than five times the level most teams believe they are running at (How Not To Run an A/B Test). One in four "wins" that will not replicate is not a statistics problem. It is a roadmap problem.
This guide is for growth and CRO teams running tests on e-commerce or travel traffic too seasonal to sit still for four weeks. By the end you will be able to plan a look schedule before launch, read the threshold for each look off a table, size the test for the small extra sample that planning costs, and recognise where the method stops being valid.
Why peeking breaks a fixed-horizon test
A fixed-horizon test spends its entire 5% error budget once, at one pre-declared sample size. Every extra look is another chance to cross the line. The problem is not looking at the data; it is looking with a threshold priced for a single look. Miller's table gives the reported significance level needed to hold a true 5% rate if you insist on peeking a set number of times:
| Number of looks | Reported significance needed for a true 5% |
|---|---|
| 1 | 2.9% |
| 2 | 2.2% |
| 3 | 1.8% |
| 5 | 1.4% |
| 10 | 1.0% |
Put the other way round: peek ten times and what you read as 1% significance is really 5%. Crude, but honest. The methods below do the same job more efficiently.
Three ways to stop early, and what each charges you
| Approach | Fix before launch | Cost versus a fixed test | Use it when |
|---|---|---|---|
| Fixed horizon | Sample size, one analysis date | None, but no early stop allowed | Traffic is stable, window short |
| Group sequential (alpha spending) | Maximum sample, number and timing of looks | Max sample rises ~3% (O'Brien-Fleming) to ~20% (Pocock), 4 looks | Sample size is estimable |
| Always-valid (mSPRT) | Nothing, in principle | Largest power loss of the three at the same sample | Sample size is not estimable |
The third row matters because it is what most commercial platforms ship. In Georgi Georgiev's simulation at a fixed maximum sample of 7,045 per group, a fixed-horizon test reached 83.3% power, group-sequential AGILE hit its 80% target, SPRT about 73%, and always-valid inference 62.4%. To reach 80% power, always-valid inference needed roughly 85% more sample than the group-sequential design (Analytics-Toolkit, 2022). Spotify's engineers concluded the same independently: group sequential tests are "systematically better or comparable to always valid approaches" when sample size can be estimated (Schultzberg and Ankargren, 2023).
This is worth correcting, because the vendor framing runs the other way. Optimizely's documentation tells you to "skip doing pre-experiment work, such as calculating the sample size before launching the experiment" (Optimizely support). The error control is real and the underlying research is sound (Johari et al., Always Valid Inference); what the framing omits is the power bill. Statsig, implementing the same family of methods, says it plainly: "If you need an accurate measurement of the effect size, wait for the full power that your pre-experiment power calculation estimates" (Statsig docs). Read "no sample size needed" as a claim about validity, never efficiency.
The method, step by step
Worked example: an apparel retailer testing checkout. Baseline session conversion 2.4%, target effect +10% relative (2.64%), two-sided alpha 5%, power 80%, roughly 5,000 sessions a day.
1. Size it as if you will never peek
Two-proportion sizing gives 66,943 sessions per arm, 133,886 total. This is not optional overhead just because you plan to stop early; it is the anchor the schedule is built on.
2. Choose how many looks, and when
Three to five looks captures almost all the available benefit; beyond that each look adds maximum-sample cost for little return. Align looks to whole weeks, because day-of-week composition is a large source of drift in retail traffic. At 5,000 sessions a day, four weekly looks lands near 140,000 sessions, covering the planned maximum.
3. Pick a spending function and write down the thresholds
An alpha spending function decides how much of the 5% budget each look may consume, as a function of how much data has accumulated (Penn State STAT 509). The two classical choices, at four equally spaced looks, two-sided alpha 0.05:
| Look | Share of planned sample | O'Brien-Fleming: stop if p < | Pocock: stop if p < |
|---|---|---|---|
| 1 (end of week 1) | 25% | 0.0001 | 0.0179 |
| 2 (end of week 2) | 50% | 0.0055 | 0.0179 |
| 3 (end of week 3) | 75% | 0.0216 | 0.0184 |
| 4 (end of week 4) | 100% | 0.0411 | 0.0188 |
| Max sample vs fixed test | +3% | +20% |
These were computed with the Lan-DeMets spending functions and match the published characterisation: O'Brien-Fleming uses "very stringent thresholds early on, but the final analysis threshold is near the typical p < 0.05", while Pocock holds a near-constant threshold around p < 0.016 (Ciolino, Kaizer and Bonner, 2023).
O'Brien-Fleming is cheap insurance: it almost never fires at look 1, costs about 3% more maximum sample, and its final threshold is barely tighter than 0.05. Pocock stops early far more often but charges 20% more traffic. Default to O'Brien-Fleming for revenue-linked metrics.
4. Write the rule into the ticket before launch
One sentence, visible to everyone who can stop the test: "We look at the end of weeks 1 to 4, declare a winner only if the two-sided p-value for [primary metric] is below that look's threshold, and otherwise run to week 4." A rule that lives in someone's head is not a rule.
5. Set the futility boundary too
Most teams implement only efficacy boundaries, so they stop winners early and let losers run full length, which is backwards for opportunity cost. A non-binding futility boundary — stop when a meaningful win has become implausible — is where most of the calendar saving comes from. With four O'Brien-Fleming looks, expected sample under a true effect falls to roughly 83% of the fixed-horizon requirement.
6. Do not report the early point estimate as the effect size
A test that stops early stopped because the observed lift was unusually large, so the estimate is biased upward by construction. Use it for the ship decision, not the revenue forecast. Statsig: "Early decisions often produce underpowered lift estimates with high uncertainty."
Pre-launch checklist
- One primary metric, declared in writing before launch
- Maximum sample and calendar end date fixed
- Looks fixed in number and timing, aligned to whole weeks
- Spending function chosen, thresholds written into the ticket
- Futility rule defined, not just the efficacy rule
- SRM checked at every look, before the p-value is read
- Guardrail metrics monitored separately, outside the stopping rule
- Everyone who can stop the test knows unplanned looks are barred
Calibrating this for Türkiye and the region
Türkiye recorded 5.94 billion e-commerce transactions in 2025 on a volume of TL 4.57 trillion, per the Ministry of Trade's E-Commerce Outlook Report presented in May 2026 (Daily Sabah, 12 May 2026). It is a large market made mostly of mid-sized stores, so four-week windows are routine — long enough to collide with a campaign.
The late-November discount period known locally as Efsane Cuma, and the 11.11 event before it, change who is on the site, not merely how many. Discount-led visitors convert on different logic, so a schedule crossing a campaign boundary is not measuring one population. Our rule: finish before the campaign or start after it. If a test must run through one, treat the campaign week as its own stratum and take no stopping decision inside it.
A second adjustment applies across the EU and to KVKK-aligned consent banners in Türkiye: if analytics are consent-gated, part of your traffic is unobserved. Size and schedule on observed sessions, or every look arrives later than the calendar says.
[INTERNAL DATA NEEDED: median daily sessions and median realised test duration across Switas e-commerce clients, to replace the illustrative 5,000 sessions/day figure with an actual distribution.]
Where this breaks down
- Composition shift. Sequential methods assume the incoming stream is exchangeable over time. A campaign or paid-media change breaks that, and stopping early locks in whichever week was unusual.
- Novelty and primacy. A week-1 reading measures reaction, not settled behaviour. If the change is visible to returning customers, put no efficacy boundary on look 1 at all.
- Heavy-tailed metrics. Revenue per session is dominated by rare large orders, and normal approximations are unreliable at small interim samples. Stop on conversion rate and read revenue alongside, or winsorise and say so.
- Multiple metrics. These boundaries control error for one metric. Five metrics with five stopping rules and no further correction puts you back where you started.
- Broken instrumentation. A sequential test will happily cross a boundary on corrupted data. It is no substitute for an SRM check.
- Not covered here: Bayesian stopping rules, multi-armed bandits, and variance reduction such as CUPED. Each interacts with early stopping in ways needing separate treatment.
FAQ
Mark this section with FAQPage schema.
Can I just use a stricter p-value instead of a spending function?
Yes, if you fix the number of looks in advance and there are only a few. Miller's table gives the threshold. It is more conservative than alpha spending, so you pay more sample than necessary, but it needs no new tooling.
Does sequential testing mean I can skip the sample size calculation?
No. Validity is preserved without a sample size estimate; power is not. Both the Georgiev and Spotify simulations show the same pattern: when you can estimate sample size, a group sequential design dominates.
What if I am watching for a regression rather than a win?
That is guardrail monitoring, not hypothesis testing. Watch safety and error metrics continuously under a separate, looser rule and leave the primary stopping rule untouched. Killing a variant that is harming users needs no ceremony.
Is Bayesian testing immune to peeking?
No. Stopping the moment a posterior probability crosses a threshold changes the decision rule's operating characteristics unless the rule was designed for optional stopping. The prior changes the shape of the problem, not its existence.
O'Brien-Fleming or Pocock?
O'Brien-Fleming in almost every commercial case. Choose Pocock only when an early answer has concrete value, you can absorb roughly 20% more maximum sample, and someone is ready to ship on a week-1 result.
Can I add a look I did not plan?
With an alpha spending function, yes, provided the decision to look is not driven by the data you are about to see. Adding a look because the dashboard "looked close" is data-dependent timing and voids the guarantee.
Work with us on this
If you run experiments on a Turkish or regional e-commerce estate and the stopping rule today is "check it every morning", we will rebuild it with you: sizing, look schedule, boundaries, futility rule and SRM gate, written into your experiment template so the next test inherits it. Talk to Switas about your experimentation programme.
Sources
- Evan Miller, How Not To Run an A/B Test
- Johari, Pekelis, Walsh and Koomen, Always Valid Inference: Continuous Monitoring of A/B Tests; published in Operations Research 70(3), 2022
- Georgi Georgiev, Comparison of the statistical power of sequential tests, Analytics-Toolkit, November 2022 (updated March 2023)
- Schultzberg and Ankargren, Choosing a Sequential Testing Framework, Spotify Engineering, March 2023
- Optimizely, Statistical analysis methods overview
- Statsig, Frequentist Sequential Testing
- Ciolino, Kaizer and Bonner, Guidance on interim analysis methods in clinical trials, J Clin Transl Sci 7(1):e124, 2023
- Penn State STAT 509, Alpha Spending Function approach
- rpact, Defining Group Sequential Boundaries
- Türkiye's e-commerce volume jumps more than 52% in 2025, Daily Sabah, 12 May 2026, reporting the Ministry of Trade E-Commerce Outlook Report







