Tree testing evaluates whether people can find things in a proposed information structure, by presenting the hierarchy as a text-only set of nested labels and asking participants to locate specific items. Because there is no visual design, no search, and no supporting content, the method isolates the structure itself, which is why it is sometimes described as reverse card sorting.
Its value lies in testing what card sorting cannot. Card sorting shows how people group items when they see them all together; tree testing shows whether someone with a specific goal can navigate to the right place. These are different tasks, and a structure that emerges from a card sort frequently underperforms in a tree test, because grouping and finding rely on different reasoning. Running the two in sequence, generating a structure and then validating it, is standard practice for exactly this reason.
The method produces unusually actionable metrics. Success rate records whether participants reached the correct location. Directness records whether they got there without backtracking, which distinguishes a structure people navigate confidently from one they eventually solve by trial and error. First-click accuracy records whether the initial top-level choice was correct, which is the most diagnostic single number, because a wrong first click usually means the top-level labels do not communicate what lies beneath them. Time to completion adds context but matters less than the others.
The failure patterns it reveals are specific enough to fix. When participants consistently choose the wrong branch for a given task, the labels are ambiguous or the item is in the wrong place. When they oscillate between two branches, the categories overlap and the boundary is unclear to anyone outside the organization. When success is high but directness is low, the structure is navigable but the labels do not inspire confidence. When one category absorbs traffic for unrelated tasks, it is probably a vague catch-all such as "resources" or "solutions" that people select when nothing else looks right.
Tree testing is inexpensive and fast, which makes it suitable for iteration rather than as a single validation exercise. A structure can be tested, revised, and retested within days using unmoderated online tools with modest participant numbers, and the improvement between rounds is usually visible in the first-click data. It is also useful as a diagnostic on an existing site, run against the current structure to establish a baseline before redesign, so that the new structure can be shown to be better rather than merely different.
The method also works well as a comparison tool between a current structure and a proposed one, and framing it that way changes how findings are received. A redesign presented as an improvement on aesthetic or organizational grounds invites argument; the same redesign presented with a measured increase in first-click accuracy and task success against the existing structure is considerably harder to dispute. Running the baseline test on the current site before any design work begins is therefore worth the small additional cost, and it has the secondary benefit of identifying which parts of the existing structure already work, so that the redesign does not discard them along with the parts that do not.
Its limitation is that it deliberately excludes everything except structure. Real sites have search, promotional links, cross-references, and visual cues that help people find things despite an imperfect hierarchy, so a poor tree test result does not always predict poor real-world findability. It is nonetheless the cleanest available signal about the structure itself, which is why it is a standard step in the information architecture stage of a product design engagement and a common component of the diagnostic work in a UX audit when navigation is suspected as a cause of poor performance.