You have two weeks to test a checkout flow, no research panel contract, and a stakeholder who has heard that five users is enough. Two of those three constraints are real. The five-user number is not. In Laura Faulkner's 2003 study, random sets of five participants drawn from a pool of 60 found anywhere between 55% and 99% of the known usability problems, and you cannot tell from inside the study which end you landed on.
This is a recruiting playbook for product, UX and CRO teams running moderated usability tests without a research operations function. By the end you will have a participant count you can defend in a planning meeting, a screener that filters out professional testers, over-recruitment math based on published no-show rates, and a consent setup that survives a KVKK or GDPR review.
Step 1: Set the participant count from the decision, not the folklore
The "five users" rule comes from Nielsen and Landauer's 1993 model, popularised in Jakob Nielsen's March 2000 article. The formula is N(1 − L)^n, where L is the share of problems a single user surfaces. Nielsen used L = 31%, which produces the familiar 85% figure at five users.
That 31% is an assumption, not a measurement of your product, so the arithmetic only speaks to problems roughly one user in three hits. A problem that 10% of users hit needs far more sessions before it appears at all. Faulkner's replication with a real pool of 60 users found that ten participants raised the worst-case discovery rate to 80%, and twenty raised it to 95%. Spool and Schroeder's 2001 study, which used open-ended tasks across multiple sites rather than one scripted flow, kept surfacing serious problems after dozens of users.
Nielsen's own article says as much in the caveats that rarely get quoted: three to four users per group for two distinct audiences, 15 for card sorting, 20 for quantitative studies. Set the number from the decision you are about to make.
| What the test has to do | Participants per segment | Why |
|---|---|---|
| Find obvious blockers in one flow, one audience | 5 | Catches high-frequency problems; the floor is around 55% |
| Decide whether a redesign ships | 8–10 | At ten users the worst-case discovery rate in Faulkner's data rises to 80% |
| Build evidence for a contested, expensive change | 15–20 | At twenty the floor reaches 95%, so "we did not see it" means something |
| Two or more distinct audiences (new vs returning, buyer vs approver) | 3–5 each, not 5 total | Different mental models produce different problem sets |
| Produce task success or time-on-task numbers | 20+ | A quantitative study wearing a qualitative costume |
Worked example: a checkout test for a mid-size Turkish retailer, new and returning buyers. Two segments, eight completed sessions each, quota'd 60/40 mobile to desktop, gives 16 sessions to recruit against.
Step 2: Write a screener that filters behaviour, not opinions
In User Interviews' State of User Research 2025, based on 485 researchers surveyed between 25 July and 9 August 2025, time to recruit was still a concern for 54% of respondents, and finding qualified participants remained the top struggle. Screeners that filter on stated intent rather than observed behaviour are the main reason. Six rules we apply to every one:
- Ask about the last 60 days, not intentions. "Which of these did you buy online in the last 60 days?" beats "How likely are you to buy X?"
- Never signal the right answer. Put the qualifying option among plausible decoys, randomised.
- Include one open-ended question, such as "Describe the last thing that went wrong when you ordered something online." Generic, fluent, oddly on-topic answers are the cheapest fraud signal available.
- Ask for one verifiable detail about the last purchase in the category: retailer, delivery method, rough price.
- Cap the panel share. If everyone comes from one panel, you are testing that panel's habitués. Blend in a non-panel channel: your customer list, a partner newsletter, in-store intercepts.
- Quota by device. Mobile checkout failures and desktop checkout failures are different failures.
Reject on two or more of these: a generic or machine-sounding open answer; screening answers that contradict each other; a profile that qualifies for an implausibly wide range of studies; contact details that do not match the claimed profile.
Step 3: Over-recruit against published no-show rates
MeasuringU's review of no-show data puts the average across roughly 14,210 User Interviews studies at 8.6%, ranging from 0% to 34%, with B2B remote sessions at 10.4%. MeasuringU's own 24 studies covering 958 participants came in at 5.0%, with professional recruiting, extra screening, reminders and incentives averaging around $150. Incentive size and no-show rate correlate negatively at around r = −.50.
Plan on 10%, or 15% for B2B and specialist roles, low or non-cash incentives, customer lists with no confirmation call, or in-person sessions in winter. Invitations to send = completed sessions needed ÷ (1 − expected no-show rate). For the 16-session example at 15%, that is 16 ÷ 0.85 ≈ 19 confirmed bookings. Book the extra three as real slots at the end of each day rather than holding them as backfill.
[INTERNAL DATA NEEDED: Switas no-show rate by recruitment channel and incentive level across the last 12 months of moderated studies, to replace these international averages with a regional benchmark.]
Step 4: Split the consent paperwork in two before you record anything
This is where teams in Turkey now get caught. The Turkish Data Protection Authority's principle decision 2026/347, dated 18 February 2026 and published in the Official Gazette on 24 March 2026 (No. 33203), holds that the privacy notice (aydınlatma metni) and the explicit consent text (açık rıza metni) must be separate, with their own headings and their own declarations. Merging them into one signable block is unlawful, as is asking the participant to "approve" the privacy notice at all: a notice is acknowledged, not consented to. Verbatim copies of another organisation's template are also called out. The obligation sits under Article 12 and is sanctionable under Article 18.
Outside Turkey the same rule exists in softer form. GDPR Article 7(2) requires a consent request to be "clearly distinguishable from the other matters, in an intelligible and easily accessible form, using clear and plain language." What that means for a usability session:
- Two documents, or two clearly headed blocks: what you collect and why (acknowledged), and consent to record (given).
- Separate recording consent from consent to reuse clips internally and to show clips to a client. Three asks, three tick boxes, none pre-ticked.
- Pay the incentive whether or not the participant consents to recording. Consent that gates payment is hard to call freely given, and a notes-only session is still usable.
- Write the retention period down and honour it. "Recordings deleted 90 days after the study report" is a sentence you can enforce.
- Keep special-category data out of the screener. Checkout research needs nobody's health information.
The two-week schedule
- Days 1–2. Screener plus both consent documents. Get legal sign-off on the consent split now, not on day nine.
- Day 3. Launch the screener across at least two channels.
- Days 4–6. Screen, schedule, send invites with the privacy notice attached.
- Day 7. Pilot with a colleague, then with one real participant. Fix the task wording.
- Days 8–12. Run sessions, maximum four a day, with reminders 24 hours and two hours before each.
- Day 13. Debrief while the sessions are fresh.
- Day 14. Findings ranked by frequency across participants and by revenue exposure.
Where this breaks down
- Low-incidence B2B roles. Finance approvers at 200 to 1,000-employee companies will not appear in two weeks. Budget three to four and raise the incentive.
- Regulated and public-sector work. Ethics approval, procurement and special-category data handling add weeks no screener trick removes.
- Geographic skew inside Turkey. Panel depth thins out fast outside İstanbul, Ankara and İzmir, so panel-only recruiting over-represents metro users. That matters when the delivery experience differs by region.
- Your customer list is biased by construction. Everyone on it already completed the flow you are testing. The people who failed are not there.
- Open exploration is a different animal. The counts above assume a scripted flow; free browsing behaves like Spool and Schroeder's study.
- This produces no prevalence estimates. "Four of eight struggled" is not "50% of customers struggle." For that number, run a quantitative study or an experiment.
FAQ
Is five participants ever enough?
Yes, when you are looking for high-frequency blockers in one scripted flow with one audience, and you plan to iterate and test again. It is not enough when a single test has to settle a contested decision, because the worst case in Faulkner's data was 55% of problems found.
How many do I need with two distinct audiences?
Three to five per audience, not five in total. Nielsen's own guidance says the same. Distinct groups bring distinct mental models, so their problem sets overlap less than you expect.
Should I recruit from my own customer list or from a panel?
Both, in the same study. Your list has people with real accounts and real context. A panel has people who have not learned your workarounds. One source alone builds a predictable bias into every finding.
What incentive should I offer?
Enough to keep the no-show rate near the published average, since incentive size and no-shows correlate negatively at around r = −.50. Benchmark against local hourly professional rates for the segment rather than copying a US figure.
Do I need explicit consent to record a usability session in Turkey?
Yes, and since principle decision 2026/347 the consent text must be separate from the privacy notice, with its own heading and declaration. Asking someone to approve a combined block is the practice the Authority ruled against.
Do unmoderated tests need the same recruiting rigour?
More, not less. No moderator is there to notice that a participant does not match the screener, so screening quality is your only defence against unusable sessions.
Run this on your own checkout
Pick one flow, size the study against the decision it has to support, and recruit with the math above. If you would rather have the screening, sessions and synthesis handled for you, that is what our user research service does.
Sources
- Nielsen, "Why You Only Need to Test with 5 Users", NN/g, 18 March 2000
- Faulkner, "Beyond the five-user assumption", Behavior Research Methods, Instruments, & Computers 35(3), 2003, 379–383 (PubMed abstract)
- Spool and Schroeder, "Testing web sites: five users is nowhere near enough", CHI '01
- MeasuringU, "A Brief History of the Magic Number 5"
- MeasuringU, "What Is a Typical No-Show Rate for Moderated Studies?"
- User Interviews, "State of User Research 2025"
- KVKK announcement on principle decision 2026/347, 18 February 2026
- Moroğlu Arseven on decision 2026/347, Official Gazette 33203, 24 March 2026
- GDPR Article 7, Conditions for consent







