Choose Card Sorting, Tree Testing or First-Click Testing Before You Redraw Your Menu

Choose Card Sorting, Tree Testing or First-Click Testing Before You Redraw Your Menu

Optimal Workshop's own dataset, drawn from millions of tree-test tasks, records 70% task success after a correct first click and 24% after an incorrect one. The first label a person picks carries most of the outcome. What that number does not tell you is which study to run. Teams card sort when they already have a menu and only doubt the wording, tree test a structure whose vocabulary was never checked, then argue about results that were never going to settle the question. This guide is for UX leads, product managers and CRO teams with a navigation problem and a fixed research budget. By the end you will be able to match a navigation decision to one of three methods, size the study, set a pass threshold before you collect data, and state what the result cannot prove.

The three methods answer three different questions

Card sorting is generative. You hand people your content items and let them group and name the groups. It tells you how your audience carves up the domain and what words they use. It cannot tell you whether a particular menu works, because there is no menu in the study.

Tree testing is evaluative. You strip a proposed hierarchy to text-only labels, give people a task and watch where they go. It tells you whether a structure is findable and which branch swallows people. As NN/g notes, the text-only setting deliberately removes placement, colour and imagery, so it isolates label and structure quality.

First-click testing is narrower. One screen, one task, one recorded click. It tells you whether the entry point reads correctly at a glance and nothing about what happens two levels down.

Running the wrong one is the common failure. "Nobody can find the returns policy" is not a card-sorting question. "We do not know whether 'Solutions' means anything to buyers" is not a tree-testing question — a tree test would only measure your guess.

Criteria table: pick by the decision, not the tool you have a licence for

CriterionCard sortingTree testingFirst-click testing
Question it answersHow would people group and name this content?Can people find things in this structure?Does the entry point read correctly at a glance?
TypeGenerativeEvaluative, quantitativeEvaluative, quantitative
Needs an existing structureNoYes, at least a draftYes, a screen or wireframe
Typical participants15 minimum, 30 for high stakes50+ per tree for a quantitative read50+ per screen variant
Primary metricsAgreement and co-occurrence, label frequencySuccess rate, directness, first click, timeFirst-click accuracy, click heat distribution
Output you can act onCandidate categories and user vocabularyA ranked list of broken branches and labelsA verdict on one label or layout choice
Where it misleadsIndividuals group differently; needs judgement to resolveNo visual scent, so it flatters bad layoutsIgnores recovery; a wrong first click is not always a lost task

The decision rules

  • No structure, or one nobody believes in. Card sort first — open if you want vocabulary, closed if the top level is fixed and only placement is in question.
  • Two candidate structures and a meeting that will not end. Tree test both. It is the only one of the three that settles a structural argument with a number.
  • Traffic arrives but the next click is wrong. First-click test that screen. Do not rebuild the architecture to fix one label.
  • Analytics already shows where people drop. Skip discovery. Tree test the leaking branch, then first-click test the replacement label.
  • New market or language. Card sort in that language with that market's participants. A translated menu is an untested menu.

Sample sizes and thresholds you can defend in a review

For card sorting, Tullis and Wood tested 168 users and analysed random subsets. Correlation with the full-sample result was 0.75 at 5 users, 0.90 at 15, 0.93 at 20, 0.95 at 30 and 0.98 at 60. Nielsen reads that as: 15 is where stopping is comfortable, 30 is for well-funded projects. The five-user rule from usability testing does not transfer — a generative method needs enough people for the pattern to stabilise.

For tree testing, NN/g recommends 50 or more participants per tree for a meaningful comparison between trees, and sets pass thresholds from Albert and Tullis's bands across 98 studies: under 40% poor, 41–60% fair, 61–80% good, 80–90% very good, above 90% excellent. Agree the threshold before you see the data. Mission-critical tasks — checkout, refund, appointment booking — should clear 90%; a promotional browse task can live at 70%.

Read directness alongside success, never instead of it. High success with low directness means people arrive by wandering: the label is wrong even though the structure survives. That pattern is the most useful signal a tree test produces.

Worked example

An online pharmacy with a "Health" top-level item and a support article buried under it. The numbers are illustrative, not client data.

  1. Tree test, 60 participants, 8 tasks. "Find out whether your prescription is covered." Success 58%, directness 22% — fair on the bands, and the directness gap points at the label.
  2. Read the first-click table. 41% to "Health", 33% to "Support", 18% to "Account": three plausible homes means the category boundary is not in users' heads.
  3. Closed card sort, 20 participants, using the article titles and the existing top level. If coverage items scatter across three groups, the top level is wrong, not the wording.
  4. Rewrite and retest. New tree, same 8 tasks, 60 fresh participants. Target success above 80% and directness above 60%.
  5. First-click test the new label on the real page before shipping. The tree test had no visual scent; the live page does.

[INTERNAL DATA NEEDED: replace this illustrative sequence with a de-identified Switas tree-test before/after, including success and directness figures and the number of tasks.]

Labels in Turkish, and why the regional case is not cosmetic

Two things make label testing harder in Turkish and several neighbouring languages, and both matter to anyone selling into the region rather than only from it.

First, Turkish is agglutinative: meaning stacks onto a stem as suffixes, so a category that is two short words in English is often one long word in Turkish. Menus that fit in English overflow, and teams fix it by truncating — which is precisely the variable a tree test measures. Test labels at the length they will actually render.

Second, Turkish has four I characters: dotted İ (U+0130) and i (U+0069), dotless I (U+0049) and ı (U+0131). Under Turkish rules, lowercase i maps to İ, not I, so case-insensitive matching written for English fails — a search for "iptal" will not match a stored "IPTAL". If site search, filter chips or breadcrumb casing pass through a locale-naive uppercase call, users get an empty result set for a term that exists. Rule that out before you blame the label.

There is a compliance edge too. WCAG 2.2 Success Criterion 2.4.6, Headings and Labels, is Level AA and requires that headings and labels describe topic or purpose. A label that tested at 40% success is not only a conversion problem; it is evidence against a conformance claim.

What these numbers do not prove

The most quoted figure here is 87% task success after a correct first click against 46% after an incorrect one, from Bailey and Wolfson's work in 2006–2009. Treat it carefully: Optimal Workshop's own, much larger dataset gives 70% and 24% for the same comparison. Same direction, materially different magnitudes — and both are correlations measured inside the instrument, not proof that a click test predicts live-site behaviour.

On that point, Sauro, Schiavone, Du and Lewis ran 130 participants in 2023 comparing first clicks on static images with the live sites: behaviour was similar, around 6% average absolute difference across hotspot regions. They also stated they did not measure downstream task success, and that responsive layouts and hover menus produced real discrepancies. The honest claim is that click tests reproduce where people click, not that they predict whether people finish.

Tree testing has a parallel caveat. Kuric, Demcak and Krajcovic compared three tree-testing variants against high-fidelity prototypes of the same architecture across 180 participants and 1,800 tasks, and the variants did not agree with each other; the tree-visible variant tracked prototype behaviour most closely, and backtracking metrics differed significantly. If your tool renders the tree differently from the one a benchmark came from, the benchmark does not transfer.

Where this breaks down

  • Search-dominant sites. If most sessions start in the search box, you are optimising a path few people use. Check the share of sessions with a search event first.
  • Personalised navigation. Tree testing assumes one fixed hierarchy. If the menu differs by segment, you are testing a fiction.
  • Very small populations. B2B and clinical audiences often cannot supply 50 qualified participants. Report counts rather than percentages and say so.
  • None of the three tests comprehension. They evaluate finding, not understanding once found. That needs usability testing.
  • None of it measures revenue. A navigation fix is a hypothesis for an experiment, not a proven lift. Ship it behind a test.

Frequently asked questions

Can I run a card sort and a tree test on the same participants?

Not in the same session. Sorting teaches people your model, so they then find things a fresh visitor would not. Recruit separate groups, or leave several weeks between.

How many participants do I need for a card sort?

Fifteen gets you a 0.90 correlation with a full-sample result in the Tullis and Wood data, and 30 gets you 0.95. Below 15 the pattern is unstable. Going past 30 buys very little for the cost.

Is 50 participants really necessary for a tree test?

For a quantitative comparison between two trees, yes — that is NN/g's recommendation. For a single tree where you only want to find the obviously broken branches, 20 to 30 will surface the worst offenders, but report counts rather than percentages.

What counts as a passing success rate?

Decide before you run. Use the Albert and Tullis bands as a default: above 80% is very good, 61–80% is good, below 40% is poor. Raise the bar to 90% for tasks tied to money, medication or a legal right.

Success is high but directness is low. What do I fix?

The label, not the structure. People are reaching the right place after exploring, which means the category exists in their model but your wording does not signal it. Rewrite the label and retest the same tasks.

Does an AI tool remove the need for participants?

Not for these methods. They measure what a specific population believes and does; a model's guess about that population is a hypothesis, not data. Use models to draft tasks and cluster open-ended labels faster.

Next step

If you have a navigation argument analytics cannot settle, the cheapest resolution is a tree test on both candidate structures with thresholds agreed in advance. Write the tasks, set the bars and run it — or have us run it and hand you the ranked list of broken branches.

Sources


Inci Dindar
Written by

Inci Dindar

With a background in Software Development, she led design, UX, and product teams across fintech, media, e-commerce, and travel. For the past three years she has consulted full-time on UX and design — including as an external UX consultant for a global management consultancy helping brands and startups with audits, conversion optimization, and design processes.


Related Articles

Switas As Seen On

Magnify: Scaling Influencer Marketing with Engin Yurtdakul

Check Out Our Microsoft Clarity Case Study

We highlighted Microsoft Clarity as a product built with practical, real-world use cases in mind by real product people who understand the challenges companies like Switas face. Features such as rage clicks and JavaScript error tracking proved invaluable in identifying user frustrations and technical issues, enabling targeted improvements that directly impacted user experience and conversion rates.