Find out where your users really fail
Let's clarify in a free initial consultation which testing method fits your question — 30 minutes, a concrete recommendation, no sales pitch.
Every team believes it understands its product. Then, for the first time, you put five real people in front of it, give them a simple task — and watch three of them fail at a spot nobody internally ever questioned. That moment is uncomfortable, healthy, and the reason usability testing may be the single most effective method in the entire UX toolbox: it replaces opinion with observation.
This guide shows you which testing methods exist and when to use which, how a proper test runs through six phases, how to analyse results without fooling yourself — and what AI can genuinely contribute in 2026. As a UX design agency we run tests like these regularly; the examples come from real projects.
In a usability test you observe real users solving realistic tasks with your product. You don't ask what they think they would do — you watch what they actually do. The difference is fundamental: people are poor at predicting or explaining their own behaviour. They are, however, excellent at failing at your navigation right in front of you.
Three distinctions to set expectations:
The most common excuse for not testing: "We don't have the budget for a big study." Good news: you don't need one. Research by Jakob Nielsen and Tom Landauer has shown consistently since the 90s that a single qualitative test round with five users uncovers roughly 85% of a product's usability problems. Each additional user disproportionately re-discovers what the previous ones already found.
The fine print of this rule matters: it applies per test round and per user group. Five users, fix the problems, test again with five new users — that's the cycle that genuinely improves products. Three rounds of five beat one round of fifteen by a wide margin, because you repair between rounds and the next round validates the fixes. And if your product serves fundamentally different audiences (say, clerks and executives in a B2B tool), the five applies per group.
For quantitative claims — "task success rate rose from 62% to 81%" — you need at least 20, better 30+ participants. Which is why the most important fork comes first: do you want to understand or to measure?
Usability testing isn't a single method but a family. Two axes span the field: is a moderator present (moderated vs. unmoderated)? And do you want to understand reasons (qualitative) or collect metrics (quantitative)?
The most important methods in detail:
Moderated remote test. You meet the test user on a video call, they share their screen, you give tasks and ask follow-up questions. The gold standard for most questions: you get thoughts spoken out loud ("thinking aloud"), you can probe deeper, and you need neither a lab nor travel costs. 45–60 minutes per session; five sessions fit into a day.
Unmoderated remote test. The user works through the tasks alone while a tool records screen and voice. Scales better and costs less, but nobody can ask follow-ups — so the tasks must be watertight. Ideal for fast iterations and simple flows.
Lab or on-site test. Necessary when context matters: hardware, point-of-sale systems, medical devices, accessibility setups with screen readers. Slower and more expensive — but irreplaceable when the product is physical.
Guerrilla test. Café, cafeteria, trade-show booth: five passers-by, ten minutes, a prototype on a tablet. Methodologically the messiest variant (not your real target group), but unbeatable for early, rough concept questions on a near-zero budget.
First-click and tree test. Quantitative specialists for navigation and information architecture: where do users click first for task X? Do they find the category in the menu tree? Hundreds of participants, clear percentages — perfect before a navigation relaunch.
5-second test. Users see a page for five seconds and then describe what it was about. Checks whether your value proposition is understood instantly — the most common weakness of home pages.
SUS questionnaire. The System Usability Scale: ten standardised questions, one score between 0 and 100. No substitute for observation, but useful as a long-term benchmark ("have we improved since the relaunch?").
If you're unsure, this decision tree gets you to the right method in three questions:
A test that should produce reliable results follows a clear process. Running the sessions is actually the smallest part — quality is decided in the preparation.
"Let's see how the site lands" is not a test goal. A good goal is a testable assumption: "We believe users abandon the checkout because shipping costs appear too late." Before the test, collect your three to five most important hypotheses — from analytics anomalies, support tickets and the findings of a previous UX audit. If you test without hypotheses you'll still find something — just rarely the most important thing.
The art of task design is naming the goal without revealing the path. Three rules:
Also write a short moderation script: greeting, the reassurance "We're testing the site, not you — you can't do anything wrong", the request to think aloud, and a reminder that the moderator staying silent is normal.
Five wrong users produce worse insights than three right ones. "Wrong" means people who don't represent your target group — first and foremost your own colleagues, who know the product inside out. A short screener (three to five questions) filters for role, prior experience and usage context. You can recruit via the testing tools' panels (fast but generic), your own customer list (authentic but with relationship bias) or social ads with an incentive. Typical compensation is €30–75 per hour, considerably more for hard-to-reach B2B profiles. Always book a sixth participant as backup — one almost always cancels.
The golden rule of moderation: hold out. When the user gets stuck, the urge to help is strong — and that's exactly when you destroy the most valuable moment of the test. Instead: "What would you do now if I weren't here?" More field rules:
After five sessions you have 30–60 observations. Cluster them into problems ("4 of 5 users missed the continue button") and rate each problem on two axes: How severe is it? (cosmetic → hindering → blocking) and How many did it affect? The result is a prioritised list — the same impact/effort logic that makes a UX audit valuable. A blocking problem that hit 4 of 5 users is an emergency. Three cosmetic one-offs are a backlog item.
The report afterwards doesn't need 40 slides: per problem, a 20-second video clip, one sentence of finding, one sentence of recommendation, one severity level. A clip of a real customer despairing at your own checkout convinces any managing director faster than any statistic.
The most frequently skipped step — and the one that makes the difference. A fix is a hypothesis, not a result: whether the new solution works, you only know once five new users have seen it. Teams that establish the test → fix → retest cycle improve their product measurably faster than teams commissioning one big study a year.
A B2B provider was puzzled by many started but few submitted enquiry forms. Analytics showed where the drop-offs happened (the "project budget" field), but not why.
Five moderated remote tests later, the answer was unambiguous: four of five users hesitated at exactly this mandatory field — not because they had no budget, but because they feared committing to a number before they even knew what the project would cost. Two said verbatim they would "rather call first".
The fix: budget as an optional dropdown with ranges plus the microcopy "non-binding estimate". The retest showed no more hesitation, and the form's completion rate rose by a good third. Total effort: two person-days. That's the kind of insight no dashboard in the world delivers.
Hardly any field is currently as flooded with AI promises as testing. An honest assessment:
What genuinely works today: AI-assisted analysis is a real lever. Automatic transcription of all sessions, clustering of similar observations, sentiment markers where users hesitate or curse, automatic highlight clips — this halves the time spent in phase 5 and makes eight instead of five sessions manageable. A language model is also a decent sparring partner for task design and screener drafts.
What to treat with caution: "synthetic users" — LLM agents that simulate personas and operate your interface. They sound strikingly real but have a structural problem: they are trained on plausible behaviour, not on the real failures of your specific audience. They don't trip over jargon your actual customers don't know, they have no impatience, no bad glasses, and no phone screen in glaring sunlight. As a quick pre-check to weed out gross logic errors before the real test: useful. As a replacement for real users: no. Anyone selling you "usability testing without users" is selling you the removal of the insight.
You can start this week with this:
The honest answer: less than most people think — if you start small. A guerrilla test costs an afternoon. An unmoderated remote test with five participants lands in the low four figures including tool and incentives. A professionally moderated test cycle (concept, five sessions, analysis, report) ranges from €4,000 to €12,000 depending on recruiting effort; complex lab studies sit above that.
Compare that to the most expensive alternative: building on assumption. A single overlooked blocking problem in the checkout costs more per month — at any relevant traffic level — than the test that would have found it. Our rule of thumb from project practice: always test before expensive decisions — before the relaunch, before the big feature, before the campaign that sends traffic to an untested funnel.
For qualitative tests: five per round and target group (finds ~85% of problems). For quantitative metrics like success rates: at least 20–30. Three small rounds beat one big one.
The most important ones: moderated and unmoderated remote tests, lab/on-site tests, guerrilla tests, first-click tests, tree tests, 5-second tests, eye tracking and the SUS questionnaire. The choice depends on whether you want to understand reasons or measure values.
Yes — remote tests are the standard today. Moderated via video call with screen sharing, unmoderated via testing platforms with recording. Only for hardware, physical context or special accessibility setups does the lab have the edge.
Five people from the target group are asked to find and order a product in the shop while thinking aloud. The moderator observes where they hesitate or fail and afterwards clusters the patterns by severity — as in the practice example above.
AI significantly accelerates analysis, transcription and test design. Synthetic AI users can flag gross logic errors up front, but they don't replace real people — they simulate plausible behaviour, not the real failures of your target group.
The audit is an expert analysis without users (fast, affordable, finds violations of proven principles); the test observes real users (slower, more expensive, finds the problems no expert predicts). The ideal is the combination: first the audit, then the test.
UAT verifies before go-live that the software meets the functional requirements ("does it work as specified?"). Usability testing verifies that real users can cope with it ("is it understandable and efficient?"). A piece of software can pass every UAT and still be unusable.
Usability testing replaces the most expensive habit of all — building on assumption — with the observation of real users. Five people per round are enough to find roughly 85% of the problems; what matters is clean tasks without hints, moderation that holds out instead of helping, and analysis by severity and frequency. The biggest mistake isn't testing wrongly — it's not testing at all, or only once. Test, fix, retest: that cycle makes products measurably better. And where AI helps (analysis, preparation), use it — but don't let anyone optimise real users away.
Let's clarify in a free initial consultation which testing method fits your question — 30 minutes, a concrete recommendation, no sales pitch.