Task success rate is the proportion of participants who complete a defined task correctly. It is the most fundamental usability metric, because it measures whether an interface actually does what people need it to do, and it is the measure against which everything else, including satisfaction, speed, and aesthetic preference, is secondary. A product people enjoy but cannot use is not a successful product.
Measuring it requires deciding in advance what counts as success, and this definition determines the value of everything that follows. Binary scoring records completion or failure, which is appropriate for tasks with an unambiguous end state such as completing a purchase. Levels of success accommodate partial completion, recording whether the person finished unaided, finished with difficulty, finished with assistance, or failed. What matters is that the criterion is fixed before the sessions, since deciding afterwards whether an outcome counted invites exactly the interpretive drift the metric is meant to prevent.
The metric is unusually persuasive with stakeholders because it is concrete. Reporting that four of ten participants could not find the returns policy, or that half failed to complete registration on a phone, communicates in a way that qualitative description does not. It also survives translation into commercial terms readily, since a failure rate on a revenue-critical task can be connected directly to lost transactions, and that connection is what usually secures resources for remediation.
Interpretation requires care about sample size. Usability studies typically involve five to a dozen participants, which is appropriate for identifying problems but produces very imprecise rates: two failures out of eight is compatible with a wide range of true values. Reporting a percentage from such a sample as though it were a population estimate overstates what the study can support, and the honest presentation gives raw counts alongside any percentage. Where a precise rate genuinely matters, unmoderated testing at larger scale or instrumented analytics on the live product are the appropriate instruments.
The metric is also only as good as the tasks chosen. Tasks that reflect what people actually try to do, phrased in their own vocabulary and situated in a realistic scenario, produce meaningful results. Tasks written in the interface's own terminology measure word matching. Tasks that are artificially constrained, or that no real user would attempt, produce numbers that look rigorous and mean nothing. Selecting the right tasks generally requires prior knowledge of real user goals, drawn from analytics, support data, or earlier qualitative research.
Comparing success rates against a benchmark requires care about what the benchmark represents. Published figures for average task completion across studies are drawn from wildly different products, tasks, and participant populations, and comparing a specific rate against them supports very little. More useful comparisons are internal: the same task measured before and after a change, the same task measured across segments or devices, or the same task compared against a competitor tested under identical conditions. A competitor comparison is particularly persuasive with stakeholders, since a demonstrable gap on a task customers perform routinely is harder to dismiss than an abstract usability concern.
Success rate is at its most useful when paired with the reasons behind the failures, since the rate identifies severity and the observation identifies cause. That pairing is standard practice within a user research programme, and the resulting prioritized list of high-failure tasks is usually the most defensible input to a remediation plan produced by a UX audit, because it ranks problems by measured impact on real goals rather than by expert opinion about their seriousness.