The System Usability Scale is a ten-item questionnaire that produces a single score between 0 and 100 representing perceived usability. Created by John Brooke in 1986, it has become the most widely used standardized usability measure, largely because it is short, technology-agnostic, free to use, and produces results that can be compared across products and over time.
Its structure is deliberate. The ten statements alternate between positive and negative phrasing, which discourages participants from selecting the same response for every item without reading. Each is rated on a five-point agreement scale, and the scoring procedure converts responses into a 0 to 100 figure. That figure is not a percentage, which is the most common misinterpretation: a score of 68 does not mean sixty-eight percent of anything. It is a point on a scale whose meaning comes from comparison with accumulated benchmarks.
Those benchmarks are what give the instrument practical value. Research analyzing large numbers of SUS studies established an average of approximately 68, with scores above that indicating better than typical perceived usability and scores below indicating worse. Subsequent work by Bangor, Kortum, and Miller added an adjective rating scale that maps score ranges to descriptions such as good, okay, and poor, which makes results communicable to stakeholders who have no reference for the raw number. Percentile rankings derived from the same datasets allow a score to be positioned against the wider population of tested systems.
The instrument's strengths are reliability and comparability. It performs consistently even with small samples, typically producing usable results with a dozen or more respondents, and because it is standardized it supports tracking across releases and comparison against a competitor or a previous version. It is also quick enough to append to the end of a usability session or an onboarding flow without imposing meaningfully on the participant.
Its limitations are equally important. SUS measures perception, not performance: a product can score well while people fail tasks, particularly if the interface is pleasant and the failures are attributed by users to their own error. It produces a single number with no diagnostic content, identifying that something is wrong without indicating what, which means it must be paired with qualitative methods to be actionable. It is also sensitive to when and how it is administered, so comparisons are only valid when conditions are held constant, and it is unsuitable for evaluating a specific screen or feature rather than an overall experience.
Alternatives worth knowing about serve slightly different purposes. Shorter instruments with fewer items reduce respondent burden where the questionnaire is appended to a live product experience rather than a research session. Instruments designed specifically for websites or for particular product categories can be more sensitive within their domain. Single-item measures of perceived ease, asked immediately after a specific task, provide task-level diagnosis that a whole-experience score cannot. The practical guidance is to choose one instrument, apply it consistently, and resist the temptation to switch when a score is disappointing, since changing the measure resets the trend and removes the comparability that gave the exercise its value.
Used appropriately, it functions as a tracking metric rather than a diagnostic one, which is how it fits into ongoing work. Within a user research programme, SUS provides a consistent longitudinal measure of whether the experience is improving, while the explanation of any movement comes from session observation and interviews. In product design engagements involving a redesign, a baseline score taken before the work and repeated afterwards is one of the few pieces of evidence that survives the usual disagreement about whether the new version is genuinely better.