Skip to main content
DesignKey Studio

Design

User Testing Methods for Better Product Design

Moderated vs. unmoderated testing, tree testing, card sorting, and running a 5-user usability test - a practical guide to user testing methods you can run from.

User Testing Methods for Better Product Design - Design article hero

Every team claims to care about user experience. Fewer teams actually watch real users interact with their product and change decisions based on what they see. User testing is the mechanism that closes that gap - but only if you know which method fits your question, how to run the session well, and how to translate findings into design decisions rather than notes in a document.

This guide covers the core methods used in professional UX practice: moderated and unmoderated usability testing, tree testing, and card sorting. For each method you will get a clear description of what it reveals, when to use it, and what a minimal version looks like so you can run it from this page.

The TL;DR

  • Moderated testing surfaces the "why" behind confusion; unmoderated testing scales the "what."
  • Five users catch approximately 80 percent of usability problems - more is not always better.
  • Tree testing validates navigation structure before any visual design is done.
  • Card sorting surfaces users' mental models for how information should be organized.
  • Findings only matter if they map to design changes - the analysis step is as important as the session.

Why the Method Choice Matters

User testing is not a single technique. It is a category of research methods, each suited to different questions at different stages of design. Choosing the wrong method wastes time and produces misleading data.

The foundational question is: are you trying to understand behavior or measure performance?

Qualitative methods (moderated testing, think-aloud sessions) help you understand why users struggle. They surface mental models, emotional reactions, and the specific moments where the interface breaks down. You get rich data from a small sample.

Quantitative methods (unmoderated testing with completion rate metrics, A/B testing, first-click analysis) help you measure how often a problem occurs across a larger sample. They tell you the scale of an issue, not its cause.

Most design cycles need both. The sequence is usually: qualitative first to identify and understand problems, quantitative later to measure the impact of fixes.

Moderated Usability Testing

What It Is

A moderated test is a live session where a facilitator observes a participant attempt to complete tasks in your product or prototype. The facilitator can ask follow-up questions, probe areas of confusion, and redirect the session in real time.

What It Reveals

Moderated testing excels at surfacing the reasoning behind behavior. When a user clicks the wrong element, you can ask "What were you looking for there?" and get a direct answer. You learn not just where the interface fails, but why it fails and what the user expected instead.

This method is most valuable when:

  • You are testing a prototype before development starts.
  • You are investigating a specific area of the product that is performing poorly in analytics.
  • The tasks involve complex workflows where the failure mode is not obvious from click data alone.

How to Run a Minimal Version

A minimal moderated test requires five participants, one facilitator, and 45-60 minutes per session. For prototype testing, use Figma's presenter mode or any clickthrough prototype tool.

Before the session:

  • Write 3-5 task scenarios in user language ("You want to add a new team member to the project"). Avoid instructions that describe the UI steps ("Click the Settings icon...").
  • Identify what "success" looks like for each task.
  • Brief the participant: tell them you are testing the product, not testing them, and ask them to think aloud as they work.

During the session:

  • Stay quiet. The facilitator's job is to observe and probe, not to guide.
  • When the participant pauses or hesitates, ask: "What are you thinking right now?"
  • Note the exact moment the hesitation occurs, not just the outcome.

After the session:

  • Rate each task: completed successfully, completed with difficulty, or failed.
  • Note the specific points of hesitation or confusion with timestamps.
  • Do not conflate what participants say with what they do - both are data, but observed behavior takes priority.

Five Users Is Usually Enough

The research behind the five-user rule is well established. In a study published by Jakob Nielsen and Tom Landauer (Nielsen Norman Group), five users reliably expose roughly 80 percent of usability problems in a design. Running 10 or 20 users on the same test adds marginal new findings at significant additional cost. The better use of resources is to run five users, fix what you find, then run another five on the updated design.

This is only true for qualitative moderated testing. Quantitative studies measuring completion rates need larger samples to produce statistically reliable results.

Unmoderated Usability Testing

What It Is

Unmoderated tests are conducted remotely without a live facilitator. Participants are given task scenarios and record themselves (or are recorded via screen capture) as they attempt to complete the tasks. Tools like Maze, UserTesting.com, and Lookback support this format.

What It Reveals

Unmoderated testing trades depth for scale. You can recruit 20-50 participants in 24-48 hours and measure completion rates, time on task, and click paths. You lose the ability to ask follow-up questions, which means you often know that a problem exists but need a follow-up moderated session to understand why.

This method is most valuable when:

  • You need to measure the prevalence of a known issue across a broader sample.
  • You want to compare two design directions quantitatively.
  • You have already run moderated testing and want to validate that fixes worked.

What to Watch For

Unmoderated tests are susceptible to quality problems at scale. Participants who rush through tasks, provide low-effort think-aloud commentary, or are not representative of your actual user population can skew results. Screen participants carefully and review session recordings before including them in analysis.

Tree Testing

What It Is

Tree testing evaluates the navigation structure of a product without the influence of visual design. Participants are shown a text-only version of the site hierarchy and asked to find specific items ("Where would you go to update your payment method?"). There is no search, no visual hierarchy, no color to guide them - only the labels and structure.

What It Reveals

Tree testing isolates whether your navigation architecture makes sense to users, independent of whether the UI looks good or the labels are styled attractively. It surfaces two specific problems: items placed in the wrong category, and labels that users interpret differently from the team.

This method is most valuable when:

  • You are redesigning the information architecture of an existing product or site.
  • You are building a new product with more than 20-30 distinct sections or features.
  • Analytics shows users navigating to unexpected sections before finding what they need.

How to Read the Results

Tree testing tools (Optimal Workshop's Treejack is the category leader) produce two key metrics per task: success rate and directness. A high success rate with low directness means users found the right place but explored other branches first - a warning sign that the first guess was wrong. A low success rate points to a labeling or categorization problem.

Run tree tests with at least 20 participants to get statistically useful completion rates. Unlike moderated testing, tree testing is low-effort per participant, so recruiting 30-50 is practical.

Card Sorting

What It Is

Card sorting asks participants to organize a set of topics, features, or items into groups that make sense to them, then optionally name those groups. An open card sort lets participants create their own categories. A closed card sort gives participants predefined categories and asks them to place items.

What It Reveals

Open card sorting reveals users' mental models - how they naturally cluster information. Closed card sorting tests whether a predefined structure (your proposed navigation) matches those mental models.

Card sorting is most valuable during information architecture design, before tree testing. The sequence is: card sort to understand how users think, then tree test to validate that the structure you designed matches that thinking.

This method is also useful when:

  • Feature sets have grown organically and the current navigation reflects build order rather than user logic.
  • Different user segments (e.g., admins vs. end users) may categorize features differently.
  • You are considering merging or splitting product areas and want to know if users will find them intuitively.

Running a Minimal Card Sort

For a basic open card sort: write each item or feature on a separate card (digital or physical). Give participants the full deck and ask them to group cards any way that makes sense, then name each group. Conduct with 15-20 participants for open sorts; fewer for closed sorts.

Analyze results by looking for consistent groupings across participants. Items that participants consistently group together belong in the same navigation category. Items that spread randomly across groups have a labeling or conceptual clarity problem.

Turning Test Findings Into Design Changes

The most common failure mode in user testing is not poor methodology - it is good methodology followed by inaction. Findings accumulate in a report that no one acts on. To avoid this, structure findings as decision-forcing documents rather than observation logs.

For each finding, document:

  1. The specific task and the point of failure (with timestamp or screenshot).
  2. The frequency: how many of five participants hit this issue?
  3. The severity: does it prevent task completion, or just slow it down?
  4. The design change required: a specific, actionable recommendation.

Findings that affect 3 or more out of 5 participants and prevent task completion are critical - fix before shipping. Findings that affect 1-2 participants and cause delay are important but can be scheduled. Findings that affect 1 participant and cause no task failure are observations to monitor.

This triage prevents the common mistake of treating all findings as equal urgency, which leads to large redesigns based on a single participant's comment.

Integrating Testing Into the Design Cycle

User testing is most effective when it is scheduled, not reactive. A product team that tests only when something breaks will always be behind. A team that schedules a moderated test at the end of every design phase, and a tree test or card sort whenever the information architecture changes, stays ahead of usability problems before they reach production.

Our UX/UI design services include usability testing as a standard deliverable, not an optional add-on. If your product is already live and showing signs of usability problems - high drop-off rates, low task completion, support tickets about navigation - a UX redesign may be the right starting point. The first step in that process is always a round of moderated testing to understand what users are experiencing before recommending what to change.

ux-designproduct-designsoftware-developmentux-redesign
Written byDaniel Killyevo8 min read

Share this article

Your next project?

Whether it's an internal tool for your company or a highly available Software-as-a-Service - we help you to get your ideas off the ground!