Most teams pick a usability testing method by budget or habit, then discover the sessions cannot answer the question that prompted them. The method is not a matter of taste; it is set by the decision you are trying to close. This guide gives you a decision-first way to choose between moderated, unmoderated and asynchronous testing: the sample size each supports, the cost and data-quality reality of each, and the data-transfer rules that constrain the choice in Turkey and the EU. It is for product, UX and CRO teams who run a handful of studies a year and cannot afford to waste one. A calibration number first: in Laura Faulkner's 2003 study, random samples of five users found 85.5% of known problems on average, but the worst five-user draw found 55%.
The three methods, defined precisely enough to compare
Moderated. Researcher and participant are in the session together, in person or over a call. The researcher sets tasks, watches and probes. Output: observed behaviour plus explanation.
Unmoderated, recorded. The participant works alone through a scripted task set while a platform captures screen, clicks and usually audio. A researcher analyses the recordings later. Output: behaviour without explanation.
Asynchronous self-report. The participant works alone and writes down what went wrong — a diary study, a feedback form, an annotated task. No recording is analysed. Output: the participant's own account.
Collapsing these three into "moderated vs unmoderated" is how one widely repeated figure gets misapplied. The finding that users report roughly half the problems a trained researcher observes comes from self-report studies, where participants documented problems themselves. That is a fact about asynchronous self-report, not about recorded unmoderated sessions analysed by a researcher. When someone says unmoderated finds half the problems, ask which of the two they mean.
Start from the decision, not the method
Write down the sentence you want to be able to say when the study is over, then find the row.
| The decision you need to close | Evidence that closes it | Method |
|---|---|---|
| Why do people abandon at this step? | Behaviour plus the reasoning behind it | Moderated |
| Is this concept understood at all? | Comprehension probed in the user's own words | Moderated |
| Does this work with a screen reader? | Observed use by people who use one daily | Moderated |
| Can people complete task X unaided? | A completion rate with a confidence interval | Unmoderated, recorded |
| Which layout produces fewer errors? | A comparative metric at a sample size that can separate them | Unmoderated, recorded |
| What does a week of real use look like? | Repeated entries across days | Asynchronous diary |
| Which change makes us more money? | A live experiment on real traffic | None of the three — run an A/B test |
The last row matters more than it looks. Usability testing tells you whether an interface can be used and where it fails, not the revenue effect of a change. A five-person study asked to settle a business-metric question is the wrong instrument, and its answer will not survive contact with the funnel.
The four-question screen
If the table does not contain your question, run these in order. The first "yes" decides.
- Do you need to ask "why" while it is happening? Moderated. You cannot retrofit a probe onto a recording.
- Does the participant need credentials, a sandbox, assistive technology or a device you supply? Moderated. Unmoderated platforms handle a public prototype well and a gated environment badly.
- Do you need a number with a confidence interval attached? Unmoderated. Eight moderated sessions produce a completion rate with an interval so wide it cannot separate 50% from 90%.
- Does the behaviour span more than one sitting? Asynchronous. Onboarding across a week, or a booking abandoned and resumed, will not fit a 60-minute session.
Sample size: what each method can support
The "five users is enough" rule comes from Jakob Nielsen's 2000 article and holds only on average, for problem discovery, in a single flow. Faulkner ran 60 participants on a timesheet application, catalogued 45 problems, then resampled 100 random groups at each size. The spread is what everyone forgets.
| Users | Mean % of problems found | Worst sample in 100 draws |
|---|---|---|
| 5 | 85.5% | 55% |
| 10 | 94.7% | 82% |
| 20 | 98.4% | 95% |
| 30 | 99.0% | 97% |
| 50 | 100% | 98% |
Read it as a risk statement, not a target. Five sessions usually surface most of what is wrong, and roughly one run in twenty leaves you believing a broken flow is fine — acceptable for a prototype, bad for a checkout you are about to ship. Our working split: 5 to 8 moderated sessions for discovery on a single flow, 10 to 12 when it branches by segment or device, 30 or more unmoderated when the deliverable is a rate rather than a problem list.
Cost and calendar: where the money actually goes
Moderated cost is mostly the researcher's calendar — scheduling, no-shows, the session, then analysis. Unmoderated cost is mostly over-recruiting, because you pay for participants you discard, and that discard rate is larger than most budgets assume.
In a 2023 PLOS ONE comparison of five research platforms (N = 2,729), the share of respondents meeting a combined quality bar was 75.4% on Prolific, 75.1% on CloudResearch, 62.7% on SONA, 59.1% on Qualtrics and 54.2% on MTurk, with cost per high-quality respondent ranging from $1.90 to $8.17. A 2026 study in Quality & Quantity is harsher: of 4,453 consenting participants, 15.97% were retained after screening, with attention checks disqualifying 27.13% while explicit bot-detection questions caught only 3.48%.
So size unmoderated studies at the number you need to analyse divided by an expected retention rate, and put at least one open-ended question and one attention check in every script, because cheap bot screens do not work. A moderated session screens itself: someone who joins a call and talks to you for 45 minutes is, whatever else is true, a person.
[INTERNAL DATA NEEDED: Switas median days from recruitment brief to first completed session, and our observed screen-out rate on Turkish-language unmoderated studies, to replace the generic platform figures above.]
The Turkey and EU layer that can override your choice
An unmoderated platform records the participant's screen, often their voice and face, and stores it on vendor infrastructure usually outside Turkey and frequently outside the EEA. That is a transfer of personal data abroad, and since 1 June 2024 Turkish law routes those through a fixed set of mechanisms.
Law 7499 (Official Gazette 32487, 12 March 2024) rewrote Article 9 of Law 6698. Transfers abroad now need an adequacy decision or a listed safeguard: binding corporate rules approved in advance, a standard contract published by the Authority and notified to it, or a bespoke undertaking it approves. The exceptional grounds, explicit consent included, are drafted for incidental, non-recurring transfers — which a standing research programme is not. As the Authority's own page records, no country has yet been declared adequate, so in practice a foreign testing platform sits on a standard contract somebody has to sign and notify.
The rule we apply: if the study captures a participant's own account, real customer data, or special-category information — routine in health and public-sector work — moderated sessions on infrastructure you control are the shorter compliance path even when unmoderated is cheaper. For a public prototype and a synthetic task, the platform route is fine. Readers outside Turkey will recognise the structure from GDPR Chapter V, with one difference that bites: no adequacy decision is in force, so there is no EU-US-style framework to lean on.
Before you book anything
- The decision sentence is written down, and you can name the finding that would change it.
- The method came from the table or the screen above, not from what you ran last time.
- Sample size is justified against the deliverable: a problem list or a rate.
- Unmoderated only: target n divided by expected retention; script has an open-ended question and an attention check.
- Moderated only: tasks written so the participant needs no explanation; probes prepared but not leading.
- Recording, storage location, retention period and transfer mechanism settled before recruitment.
- Someone owns the analysis calendar; unanalysed recordings are how unmoderated studies produce nothing.
Where this breaks down
Faulkner's table comes from one timesheet application with 45 catalogued problems. Discovery curves depend on how likely an average participant is to hit an average problem, and that probability drops in a product with many branching flows. Treat the table as a floor for a simple flow.
This guide does not cover recruitment mechanics, moderation technique, turning findings into hypotheses, or accessibility conformance testing, which is an audit against criteria rather than a usability study.
AI-moderated interviews now sit between moderated and unmoderated, marketed as combining both. We have not seen published method-comparison evidence at the standard of the studies cited here. Until it exists, treat an AI-moderated session as unmoderated testing with a better script, and do not let it carry a decision that depends on skilled probing.
FAQ
Is five participants really enough? On average five find about 85% of problems in a single flow, but the worst of Faulkner's 100 five-person samples found 55%. Reasonable for a prototype, thin for anything you are about to ship.
Can I run a moderated study remotely? Yes, and for most product work you should. Moderated is about a researcher being present, not about a room. You lose environmental context and gain reach.
How many participants do I need for a completion rate? More than a moderated study can afford. Decide the difference you need to detect first — separating 70% from 85% takes a very different sample than separating 30% from 80% — then size for it and expect to discard a quarter of recruits.
Does unmoderated testing find fewer problems than moderated? Recorded unmoderated sessions analysed by a researcher find a comparable set of behavioural problems but no explanations. The "half the problems" figure applies to self-report formats where nobody watches the session.
Do we need explicit consent to record a session? You need a lawful basis and a clear notice whichever method you choose, and in Turkey the transfer-abroad question is separate from the recording question.
How long should an unmoderated task set be? Typically 15 to 20 minutes. Longer scripts increase dropout and reward exactly the participants you want to screen out.
Work with us
If a decision is waiting on evidence and you are not sure which of these three will settle it, we can scope the study with you — method, sample size, recruitment and the data-transfer paperwork. See Switas user research services.
Sources
- Faulkner, L. (2003). Beyond the five-user assumption. Behavior Research Methods, 35(3), 379–383.
- Nielsen, J. (19 March 2000). Why You Only Need to Test with 5 Users. Nielsen Norman Group.
- MeasuringU. Can Users Self-Report Usability Problems? — summarising Castillo (1998), Bruun et al. (2007), Andreasen et al. (2009).
- Douglas, Ewell & Brauer (14 March 2023). Data quality in online human-subjects research. PLOS ONE.
- Battling bots and bad data (11 April 2026). Quality & Quantity, 60(4).
- Kişisel Verileri Koruma Kurumu, Yurt Dışına Aktarım — Law 6698 Art. 9 as amended by Law 7499 (OG 32487, 12 March 2024; in force 1 June 2024).







