Most usability reports are honest about what participants did and silent about what the task script made them do. That is where studies quietly fail. Across 1,189 tasks from 115 tests with 3,472 users, Jeff Sauro found an average task completion rate of 78%, a benchmark that only means something if the tasks could have produced failure. A task that names the button, lists the steps, or asks "how would you" instead of "do this" returns 100% completion and teaches you nothing.
This guide is for researchers, product managers and CRO leads who already run sessions. By the end you will be able to write a task, set its success criterion in advance, leak-check your own wording, pilot it, and read the resulting completion rate without over-claiming.
What a task has to do before it is worth running
Four properties. A task missing one should not reach a participant.
- Failable. There is a plausible route to failure. If you cannot picture a participant failing, you are measuring whether they understood your sentence.
- Realistic. ISO 9241-11:2018 defines usability against specified goals in a specified context of use, so a task inventing a goal nobody holds measures nothing.
- Actionable. You ask for the action, not a description of it. Nielsen Norman Group's guidance is to ask users to do the thing, not to say how they would. The reason is quantitative: across 838 participants and 34 tasks, self-reported completion averaged 93% against 33% verified.
- Label-free. No word in the task appears as a label, button or menu item in the interface.
Write the scenario in five moves
- Name the decision the task informs. Write it above the task: "Do shoppers find the installment option before the payment step or after it?" If no decision changes based on the result, cut the task.
- Write the goal in the user's vocabulary. Source it from support tickets, site-search logs and review text, not the product spec. That is also your defence against label leakage: real users rarely use your menu names.
- Add the minimum context that makes the goal real. Two sentences at most: who they are, what constrains them. Constraints create the decisions you came to watch.
- Strip every interface string. Run the leak check below against the exact wording you will read aloud.
- Define the success criterion and the stop rule. Score binary, 1 or 0, then average across participants. Add a time or attempt cap so one task cannot eat the session.
A travel example. Weak version: "Go to My Bookings, find your Dusseldorf flight, and use the Change Flight button to move it to the next day." It names the control. Everyone passes.
Usable version: "You have a flight to Dusseldorf booked for Friday. A meeting moved and you now need to travel Saturday. Sort it out." Success criterion: the participant reaches a confirmation state for a Saturday departure, or states the change fee correctly and abandons deliberately. Stop at six minutes.
Set the success criterion before the session, not after
Criteria written after you watch the recordings drift toward whatever happened. Fix them in advance, with a target you are willing to miss.
| Task type | Success criterion | Target |
|---|---|---|
| High stakes or regulated | End state reached unaided | 95% or higher |
| Core commerce path | End state reached unaided | 90% or higher |
| Discovery and navigation | Correct destination reached or named | 80% or higher |
| Exploratory | No binary criterion; code the behaviour | Not applicable |
Set targets by the cost of failure, not by habit. Sauro's guidance is that high-stakes tasks sit near 100%, while walk-up consumer tasks are sometimes targeted around 70%. The 78% average is a sanity check, not a target.
Run the leak check on your own wording
Ten checks. Read the task aloud and mark each one. Any hit is a rewrite, not an appendix note.
| # | Leak | Test | Fix |
|---|---|---|---|
| 1 | Interface string | A task word appears as a label, button or menu item | Use the user's word |
| 2 | Step sequence | The task describes an order of operations | Delete the order, keep the outcome |
| 3 | Hypothetical verb | It says "how would you" or "where would you expect" | Rewrite as an imperative |
| 4 | Named path | It points at one route when several exist | State the goal only |
| 5 | Count leak | It reveals how many items or steps exist | Remove the number |
| 6 | Self-report dependency | Scoring relies on the participant claiming success | Verify the end state on screen |
| 7 | Compound task | It carries more than one success criterion | Split it in two |
| 8 | Unfailable task | You cannot describe a realistic failure | Discard the task |
| 9 | Context debt | It needs data you never supplied, such as an order number | Put the data in the scenario |
| 10 | Moderator variance | Two moderators would read it differently | Fix the wording, not the moderator |
Decide task order, and know what order costs you
For a single journey, keep the order fixed: the journey is the thing under test. For independent tasks, rotate the order and log the position each task held. Early tasks are harder, because the participant is still learning the interface, and late tasks suffer fatigue. With eight participants you cannot separate an order effect from task difficulty, so rotation is damage limitation rather than a correction, and the report should say so.
Pilot with two participants and rewrite
Run two pilots with people from the study's own recruit pool, not colleagues who already know your vocabulary. Rewrite anything a pilot asked a clarifying question about: a clarifying question is a leak you missed. Exclude pilots from the analysis and budget them in, rather than borrowing from the sample.
Fit the scenario to the market you are testing in
Turkey's Ministry of Trade put 2025 e-commerce at about 4.57 trillion lira across roughly 5.94 billion transactions, up 52.2% year on year. Dividing one by the other, which is our arithmetic and not a reported figure, implies an average transaction well under 1,000 lira. The typical session in this market is a small, frequent, low-consideration purchase, so a script built around a considered high-value purchase measures a journey most of your traffic never takes. Write the repeat-purchase task too.
Payment choice is also a live decision here, with installments and cash on delivery inside the flow, so a checkout task that hands the participant a saved card skips the step where abandonment happens. [INTERNAL DATA NEEDED: payment-method mix from the client's own checkout analytics, to pick the payment scenario to script.]
One translation trap. If you draft in English and translate, the translator tends to reach for the same word the interface uses, because common interface verbs have fewer everyday synonyms in Turkish. Back-translate, then rerun the leak check on the Turkish wording. For participants using assistive technology, "look at the top right" is a leak and a barrier at once; our WCAG audit tool covers the interface side.
Read the completion rate you actually got
Seven successes out of eight participants is not 88%. Using the adjusted-Wald method published by Lewis and Sauro for small-sample completion rates, the 95% interval around 7 of 8 runs from roughly 51% to 100%. That is the honest width of an eight-person study, which is why a qualitative-sample completion rate belongs in a findings section, not a KPI dashboard.
Two rules follow. Do not declare a target met unless the lower bound clears it. And pair every rate with a difficulty score: the Single Ease Question averages about 5.5 on its 7-point scale across more than 200 tasks, so a task everyone completes at an SEQ of 3.5 is still a finding.
Where this breaks down
- It cannot tell you whether anyone wants the thing. Task success measures reachability, not desire.
- Tight tasks kill exploratory research. If the question is "what do people do here", do not write a task.
- Small-sample rates are directional. Eight to twelve participants give problems to fix, not quarterly metrics.
- Agent-mediated journeys have no human task. If an agent completes the purchase, measurement moves to feed readability.
- Task success is not conformance. A flow every sighted participant completes can still fail accessibility rules.
Frequently asked questions
How many tasks should one session contain?
Budget by time, not count. In a 60-minute session, five to seven tasks with stop rules fit alongside intake and debrief. If the stop rules total over 40 minutes, you have too many tasks.
Can I tell a participant the name of the page they need?
Only if finding that page is outside what you are testing. If navigation is in scope, naming the page destroys the measurement. If you are testing a form on that page, naming it is a legitimate shortcut, recorded as one.
What do I do when a participant asks for help mid-task?
Reflect it back once: "what would you do if I were not here?" If they are still stuck, give the minimum assist, score the task as a failure, and note the assist. Assisted successes scored as successes are how studies produce 95% rates that do not survive launch.
Is self-reported success ever usable?
As a secondary signal only. Self-reported rates averaged 93% against 33% verified across four benchmark studies, so they overstate badly. A self-reported rate below 80% still signals a hard task.
Should unmoderated and moderated tasks be written differently?
Yes. Unmoderated tasks need context front-loaded, because nobody is there to answer a question, and verification built into the end state, such as a code on the confirmation screen. Keep the goal identical if you plan to compare the two.
Do I need to localise task wording for every market?
Localise it wherever you localise the interface. A task translated without a leak check in the target language is a different instrument from the one you validated, so cross-market rates stop being comparable.
Review your script before you run the study
If you have a task script going into field this month, send it to us. We will run it against this leak check and the criterion table and return the rewrite, as part of our user research work.
Sources
- Jeff Sauro, What Is A Good Task-Completion Rate?, MeasuringU, 21 March 2011.
- Jeff Sauro, 10 Benchmarks for User Experience Metrics, MeasuringU, 16 October 2012, SEQ figure updated May 2022.
- Jeff Sauro, How Reliable Are Self-Reported Task Completion Rates?, MeasuringU, 8 December 2015.
- Marieke McCloskey, Turn User Goals into Task Scenarios for Usability Testing, Nielsen Norman Group, 12 January 2014.
- James R. Lewis and Jeff Sauro, When 100% Really Isn't 100%: Improving the Accuracy of Small-Sample Estimates of Completion Rates, Journal of Usability Studies, May 2006, with the completion-rate confidence interval calculator.
- ISO 9241-11:2018, Ergonomics of human-system interaction, Part 11: Usability, definitions and concepts.
- T.C. Ticaret Bakanligi, Turkiye'de E-Ticaretin Gorunumu Raporu 2025, 12 May 2026.







