Stop an A/B Test Early in Four Planned Looks Instead of Peeking Every Morning

Stop an A/B Test Early in Four Planned Looks Instead of Peeking Every Morning

If you check an A/B test every morning and stop it the first time the dashboard turns green, your real false positive rate is not 5%. In a simulation testing significance after every observation, Evan Miller measured it at 26.1%, more than five times the level most teams believe they are running at (How Not To Run an A/B Test). One in four "wins" that will not replicate is not a statistics problem. It is a roadmap problem.

This guide is for growth and CRO teams running tests on e-commerce or travel traffic too seasonal to sit still for four weeks. By the end you will be able to plan a look schedule before launch, read the threshold for each look off a table, size the test for the small extra sample that planning costs, and recognise where the method stops being valid.

Why peeking breaks a fixed-horizon test

A fixed-horizon test spends its entire 5% error budget once, at one pre-declared sample size. Every extra look is another chance to cross the line. The problem is not looking at the data; it is looking with a threshold priced for a single look. Miller's table gives the reported significance level needed to hold a true 5% rate if you insist on peeking a set number of times:

Number of looksReported significance needed for a true 5%
12.9%
22.2%
31.8%
51.4%
101.0%

Put the other way round: peek ten times and what you read as 1% significance is really 5%. Crude, but honest. The methods below do the same job more efficiently.

Three ways to stop early, and what each charges you

ApproachFix before launchCost versus a fixed testUse it when
Fixed horizonSample size, one analysis dateNone, but no early stop allowedTraffic is stable, window short
Group sequential (alpha spending)Maximum sample, number and timing of looksMax sample rises ~3% (O'Brien-Fleming) to ~20% (Pocock), 4 looksSample size is estimable
Always-valid (mSPRT)Nothing, in principleLargest power loss of the three at the same sampleSample size is not estimable

The third row matters because it is what most commercial platforms ship. In Georgi Georgiev's simulation at a fixed maximum sample of 7,045 per group, a fixed-horizon test reached 83.3% power, group-sequential AGILE hit its 80% target, SPRT about 73%, and always-valid inference 62.4%. To reach 80% power, always-valid inference needed roughly 85% more sample than the group-sequential design (Analytics-Toolkit, 2022). Spotify's engineers concluded the same independently: group sequential tests are "systematically better or comparable to always valid approaches" when sample size can be estimated (Schultzberg and Ankargren, 2023).

This is worth correcting, because the vendor framing runs the other way. Optimizely's documentation tells you to "skip doing pre-experiment work, such as calculating the sample size before launching the experiment" (Optimizely support). The error control is real and the underlying research is sound (Johari et al., Always Valid Inference); what the framing omits is the power bill. Statsig, implementing the same family of methods, says it plainly: "If you need an accurate measurement of the effect size, wait for the full power that your pre-experiment power calculation estimates" (Statsig docs). Read "no sample size needed" as a claim about validity, never efficiency.

The method, step by step

Worked example: an apparel retailer testing checkout. Baseline session conversion 2.4%, target effect +10% relative (2.64%), two-sided alpha 5%, power 80%, roughly 5,000 sessions a day.

1. Size it as if you will never peek

Two-proportion sizing gives 66,943 sessions per arm, 133,886 total. This is not optional overhead just because you plan to stop early; it is the anchor the schedule is built on.

2. Choose how many looks, and when

Three to five looks captures almost all the available benefit; beyond that each look adds maximum-sample cost for little return. Align looks to whole weeks, because day-of-week composition is a large source of drift in retail traffic. At 5,000 sessions a day, four weekly looks lands near 140,000 sessions, covering the planned maximum.

3. Pick a spending function and write down the thresholds

An alpha spending function decides how much of the 5% budget each look may consume, as a function of how much data has accumulated (Penn State STAT 509). The two classical choices, at four equally spaced looks, two-sided alpha 0.05:

LookShare of planned sampleO'Brien-Fleming: stop if p <Pocock: stop if p <
1 (end of week 1)25%0.00010.0179
2 (end of week 2)50%0.00550.0179
3 (end of week 3)75%0.02160.0184
4 (end of week 4)100%0.04110.0188
Max sample vs fixed test +3%+20%

These were computed with the Lan-DeMets spending functions and match the published characterisation: O'Brien-Fleming uses "very stringent thresholds early on, but the final analysis threshold is near the typical p < 0.05", while Pocock holds a near-constant threshold around p < 0.016 (Ciolino, Kaizer and Bonner, 2023).

O'Brien-Fleming is cheap insurance: it almost never fires at look 1, costs about 3% more maximum sample, and its final threshold is barely tighter than 0.05. Pocock stops early far more often but charges 20% more traffic. Default to O'Brien-Fleming for revenue-linked metrics.

4. Write the rule into the ticket before launch

One sentence, visible to everyone who can stop the test: "We look at the end of weeks 1 to 4, declare a winner only if the two-sided p-value for [primary metric] is below that look's threshold, and otherwise run to week 4." A rule that lives in someone's head is not a rule.

5. Set the futility boundary too

Most teams implement only efficacy boundaries, so they stop winners early and let losers run full length, which is backwards for opportunity cost. A non-binding futility boundary — stop when a meaningful win has become implausible — is where most of the calendar saving comes from. With four O'Brien-Fleming looks, expected sample under a true effect falls to roughly 83% of the fixed-horizon requirement.

6. Do not report the early point estimate as the effect size

A test that stops early stopped because the observed lift was unusually large, so the estimate is biased upward by construction. Use it for the ship decision, not the revenue forecast. Statsig: "Early decisions often produce underpowered lift estimates with high uncertainty."

Pre-launch checklist

  • One primary metric, declared in writing before launch
  • Maximum sample and calendar end date fixed
  • Looks fixed in number and timing, aligned to whole weeks
  • Spending function chosen, thresholds written into the ticket
  • Futility rule defined, not just the efficacy rule
  • SRM checked at every look, before the p-value is read
  • Guardrail metrics monitored separately, outside the stopping rule
  • Everyone who can stop the test knows unplanned looks are barred

Calibrating this for Türkiye and the region

Türkiye recorded 5.94 billion e-commerce transactions in 2025 on a volume of TL 4.57 trillion, per the Ministry of Trade's E-Commerce Outlook Report presented in May 2026 (Daily Sabah, 12 May 2026). It is a large market made mostly of mid-sized stores, so four-week windows are routine — long enough to collide with a campaign.

The late-November discount period known locally as Efsane Cuma, and the 11.11 event before it, change who is on the site, not merely how many. Discount-led visitors convert on different logic, so a schedule crossing a campaign boundary is not measuring one population. Our rule: finish before the campaign or start after it. If a test must run through one, treat the campaign week as its own stratum and take no stopping decision inside it.

A second adjustment applies across the EU and to KVKK-aligned consent banners in Türkiye: if analytics are consent-gated, part of your traffic is unobserved. Size and schedule on observed sessions, or every look arrives later than the calendar says.

[INTERNAL DATA NEEDED: median daily sessions and median realised test duration across Switas e-commerce clients, to replace the illustrative 5,000 sessions/day figure with an actual distribution.]

Where this breaks down

  • Composition shift. Sequential methods assume the incoming stream is exchangeable over time. A campaign or paid-media change breaks that, and stopping early locks in whichever week was unusual.
  • Novelty and primacy. A week-1 reading measures reaction, not settled behaviour. If the change is visible to returning customers, put no efficacy boundary on look 1 at all.
  • Heavy-tailed metrics. Revenue per session is dominated by rare large orders, and normal approximations are unreliable at small interim samples. Stop on conversion rate and read revenue alongside, or winsorise and say so.
  • Multiple metrics. These boundaries control error for one metric. Five metrics with five stopping rules and no further correction puts you back where you started.
  • Broken instrumentation. A sequential test will happily cross a boundary on corrupted data. It is no substitute for an SRM check.
  • Not covered here: Bayesian stopping rules, multi-armed bandits, and variance reduction such as CUPED. Each interacts with early stopping in ways needing separate treatment.

FAQ

Mark this section with FAQPage schema.

Can I just use a stricter p-value instead of a spending function?
Yes, if you fix the number of looks in advance and there are only a few. Miller's table gives the threshold. It is more conservative than alpha spending, so you pay more sample than necessary, but it needs no new tooling.

Does sequential testing mean I can skip the sample size calculation?
No. Validity is preserved without a sample size estimate; power is not. Both the Georgiev and Spotify simulations show the same pattern: when you can estimate sample size, a group sequential design dominates.

What if I am watching for a regression rather than a win?
That is guardrail monitoring, not hypothesis testing. Watch safety and error metrics continuously under a separate, looser rule and leave the primary stopping rule untouched. Killing a variant that is harming users needs no ceremony.

Is Bayesian testing immune to peeking?
No. Stopping the moment a posterior probability crosses a threshold changes the decision rule's operating characteristics unless the rule was designed for optional stopping. The prior changes the shape of the problem, not its existence.

O'Brien-Fleming or Pocock?
O'Brien-Fleming in almost every commercial case. Choose Pocock only when an early answer has concrete value, you can absorb roughly 20% more maximum sample, and someone is ready to ship on a week-1 result.

Can I add a look I did not plan?
With an alpha spending function, yes, provided the decision to look is not driven by the data you are about to see. Adding a look because the dashboard "looked close" is data-dependent timing and voids the guarantee.

Work with us on this

If you run experiments on a Turkish or regional e-commerce estate and the stopping rule today is "check it every morning", we will rebuild it with you: sizing, look schedule, boundaries, futility rule and SRM gate, written into your experiment template so the next test inherits it. Talk to Switas about your experimentation programme.

Sources


Inci Dindar
Written by

Inci Dindar

With a background in Software Development, she led design, UX, and product teams across fintech, media, e-commerce, and travel. For the past three years she has consulted full-time on UX and design — including as an external UX consultant for a global management consultancy helping brands and startups with audits, conversion optimization, and design processes.


Related Articles

Switas As Seen On

Magnify: Scaling Influencer Marketing with Engin Yurtdakul

Check Out Our Microsoft Clarity Case Study

We highlighted Microsoft Clarity as a product built with practical, real-world use cases in mind by real product people who understand the challenges companies like Switas face. Features such as rage clicks and JavaScript error tracking proved invaluable in identifying user frustrations and technical issues, enabling targeted improvements that directly impacted user experience and conversion rates.