Choose Moderated, Unmoderated or Asynchronous Usability Testing by the Decision You Need to Close

Choose Moderated, Unmoderated or Asynchronous Usability Testing by the Decision You Need to Close

Most teams pick a usability testing method by budget or habit, then discover the sessions cannot answer the question that prompted them. The method is not a matter of taste; it is set by the decision you are trying to close. This guide gives you a decision-first way to choose between moderated, unmoderated and asynchronous testing: the sample size each supports, the cost and data-quality reality of each, and the data-transfer rules that constrain the choice in Turkey and the EU. It is for product, UX and CRO teams who run a handful of studies a year and cannot afford to waste one. A calibration number first: in Laura Faulkner's 2003 study, random samples of five users found 85.5% of known problems on average, but the worst five-user draw found 55%.

The three methods, defined precisely enough to compare

Moderated. Researcher and participant are in the session together, in person or over a call. The researcher sets tasks, watches and probes. Output: observed behaviour plus explanation.

Unmoderated, recorded. The participant works alone through a scripted task set while a platform captures screen, clicks and usually audio. A researcher analyses the recordings later. Output: behaviour without explanation.

Asynchronous self-report. The participant works alone and writes down what went wrong — a diary study, a feedback form, an annotated task. No recording is analysed. Output: the participant's own account.

Collapsing these three into "moderated vs unmoderated" is how one widely repeated figure gets misapplied. The finding that users report roughly half the problems a trained researcher observes comes from self-report studies, where participants documented problems themselves. That is a fact about asynchronous self-report, not about recorded unmoderated sessions analysed by a researcher. When someone says unmoderated finds half the problems, ask which of the two they mean.

Start from the decision, not the method

Write down the sentence you want to be able to say when the study is over, then find the row.

The decision you need to closeEvidence that closes itMethod
Why do people abandon at this step?Behaviour plus the reasoning behind itModerated
Is this concept understood at all?Comprehension probed in the user's own wordsModerated
Does this work with a screen reader?Observed use by people who use one dailyModerated
Can people complete task X unaided?A completion rate with a confidence intervalUnmoderated, recorded
Which layout produces fewer errors?A comparative metric at a sample size that can separate themUnmoderated, recorded
What does a week of real use look like?Repeated entries across daysAsynchronous diary
Which change makes us more money?A live experiment on real trafficNone of the three — run an A/B test

The last row matters more than it looks. Usability testing tells you whether an interface can be used and where it fails, not the revenue effect of a change. A five-person study asked to settle a business-metric question is the wrong instrument, and its answer will not survive contact with the funnel.

The four-question screen

If the table does not contain your question, run these in order. The first "yes" decides.

  1. Do you need to ask "why" while it is happening? Moderated. You cannot retrofit a probe onto a recording.
  2. Does the participant need credentials, a sandbox, assistive technology or a device you supply? Moderated. Unmoderated platforms handle a public prototype well and a gated environment badly.
  3. Do you need a number with a confidence interval attached? Unmoderated. Eight moderated sessions produce a completion rate with an interval so wide it cannot separate 50% from 90%.
  4. Does the behaviour span more than one sitting? Asynchronous. Onboarding across a week, or a booking abandoned and resumed, will not fit a 60-minute session.

Sample size: what each method can support

The "five users is enough" rule comes from Jakob Nielsen's 2000 article and holds only on average, for problem discovery, in a single flow. Faulkner ran 60 participants on a timesheet application, catalogued 45 problems, then resampled 100 random groups at each size. The spread is what everyone forgets.

UsersMean % of problems foundWorst sample in 100 draws
585.5%55%
1094.7%82%
2098.4%95%
3099.0%97%
50100%98%

Read it as a risk statement, not a target. Five sessions usually surface most of what is wrong, and roughly one run in twenty leaves you believing a broken flow is fine — acceptable for a prototype, bad for a checkout you are about to ship. Our working split: 5 to 8 moderated sessions for discovery on a single flow, 10 to 12 when it branches by segment or device, 30 or more unmoderated when the deliverable is a rate rather than a problem list.

Cost and calendar: where the money actually goes

Moderated cost is mostly the researcher's calendar — scheduling, no-shows, the session, then analysis. Unmoderated cost is mostly over-recruiting, because you pay for participants you discard, and that discard rate is larger than most budgets assume.

In a 2023 PLOS ONE comparison of five research platforms (N = 2,729), the share of respondents meeting a combined quality bar was 75.4% on Prolific, 75.1% on CloudResearch, 62.7% on SONA, 59.1% on Qualtrics and 54.2% on MTurk, with cost per high-quality respondent ranging from $1.90 to $8.17. A 2026 study in Quality & Quantity is harsher: of 4,453 consenting participants, 15.97% were retained after screening, with attention checks disqualifying 27.13% while explicit bot-detection questions caught only 3.48%.

So size unmoderated studies at the number you need to analyse divided by an expected retention rate, and put at least one open-ended question and one attention check in every script, because cheap bot screens do not work. A moderated session screens itself: someone who joins a call and talks to you for 45 minutes is, whatever else is true, a person.

[INTERNAL DATA NEEDED: Switas median days from recruitment brief to first completed session, and our observed screen-out rate on Turkish-language unmoderated studies, to replace the generic platform figures above.]

The Turkey and EU layer that can override your choice

An unmoderated platform records the participant's screen, often their voice and face, and stores it on vendor infrastructure usually outside Turkey and frequently outside the EEA. That is a transfer of personal data abroad, and since 1 June 2024 Turkish law routes those through a fixed set of mechanisms.

Law 7499 (Official Gazette 32487, 12 March 2024) rewrote Article 9 of Law 6698. Transfers abroad now need an adequacy decision or a listed safeguard: binding corporate rules approved in advance, a standard contract published by the Authority and notified to it, or a bespoke undertaking it approves. The exceptional grounds, explicit consent included, are drafted for incidental, non-recurring transfers — which a standing research programme is not. As the Authority's own page records, no country has yet been declared adequate, so in practice a foreign testing platform sits on a standard contract somebody has to sign and notify.

The rule we apply: if the study captures a participant's own account, real customer data, or special-category information — routine in health and public-sector work — moderated sessions on infrastructure you control are the shorter compliance path even when unmoderated is cheaper. For a public prototype and a synthetic task, the platform route is fine. Readers outside Turkey will recognise the structure from GDPR Chapter V, with one difference that bites: no adequacy decision is in force, so there is no EU-US-style framework to lean on.

Before you book anything

  • The decision sentence is written down, and you can name the finding that would change it.
  • The method came from the table or the screen above, not from what you ran last time.
  • Sample size is justified against the deliverable: a problem list or a rate.
  • Unmoderated only: target n divided by expected retention; script has an open-ended question and an attention check.
  • Moderated only: tasks written so the participant needs no explanation; probes prepared but not leading.
  • Recording, storage location, retention period and transfer mechanism settled before recruitment.
  • Someone owns the analysis calendar; unanalysed recordings are how unmoderated studies produce nothing.

Where this breaks down

Faulkner's table comes from one timesheet application with 45 catalogued problems. Discovery curves depend on how likely an average participant is to hit an average problem, and that probability drops in a product with many branching flows. Treat the table as a floor for a simple flow.

This guide does not cover recruitment mechanics, moderation technique, turning findings into hypotheses, or accessibility conformance testing, which is an audit against criteria rather than a usability study.

AI-moderated interviews now sit between moderated and unmoderated, marketed as combining both. We have not seen published method-comparison evidence at the standard of the studies cited here. Until it exists, treat an AI-moderated session as unmoderated testing with a better script, and do not let it carry a decision that depends on skilled probing.

FAQ

Is five participants really enough? On average five find about 85% of problems in a single flow, but the worst of Faulkner's 100 five-person samples found 55%. Reasonable for a prototype, thin for anything you are about to ship.

Can I run a moderated study remotely? Yes, and for most product work you should. Moderated is about a researcher being present, not about a room. You lose environmental context and gain reach.

How many participants do I need for a completion rate? More than a moderated study can afford. Decide the difference you need to detect first — separating 70% from 85% takes a very different sample than separating 30% from 80% — then size for it and expect to discard a quarter of recruits.

Does unmoderated testing find fewer problems than moderated? Recorded unmoderated sessions analysed by a researcher find a comparable set of behavioural problems but no explanations. The "half the problems" figure applies to self-report formats where nobody watches the session.

Do we need explicit consent to record a session? You need a lawful basis and a clear notice whichever method you choose, and in Turkey the transfer-abroad question is separate from the recording question.

How long should an unmoderated task set be? Typically 15 to 20 minutes. Longer scripts increase dropout and reward exactly the participants you want to screen out.

Work with us

If a decision is waiting on evidence and you are not sure which of these three will settle it, we can scope the study with you — method, sample size, recruitment and the data-transfer paperwork. See Switas user research services.

Sources


Inci Dindar
Written by

Inci Dindar

With a background in Software Development, she led design, UX, and product teams across fintech, media, e-commerce, and travel. For the past three years she has consulted full-time on UX and design — including as an external UX consultant for a global management consultancy helping brands and startups with audits, conversion optimization, and design processes.


Related Articles

Switas As Seen On

Magnify: Scaling Influencer Marketing with Engin Yurtdakul

Check Out Our Microsoft Clarity Case Study

We highlighted Microsoft Clarity as a product built with practical, real-world use cases in mind by real product people who understand the challenges companies like Switas face. Features such as rage clicks and JavaScript error tracking proved invaluable in identifying user frustrations and technical issues, enabling targeted improvements that directly impacted user experience and conversion rates.