In short
For finding usability problems in one interface with one type of user, around five participants surfaces most of the issues, and further sessions mostly repeat what you already saw. That rule does not apply to distinct user groups, which each need their own small round, or to quantitative questions such as conversion rates, which need far larger numbers.
Where the number comes from
The five-user guidance comes from research showing that usability problems are not evenly distributed. A small number of issues affect most people, and a long tail affects few.
So the first participant reveals a large share of the problems. The second reveals fewer new ones, having repeated much of what the first found. By the fifth, new discoveries have slowed considerably.
The practical conclusion was not that five is a magic number. It was that testing five people twice is more valuable than testing ten people once, because you fix things between rounds and test the fixes.
When it holds
The rule applies well to a specific situation: finding usability problems, in one interface, with one reasonably homogeneous group of users, on a defined set of tasks.
That covers a lot of real work. Can people complete signup? Do they understand the pricing page? Can they find the thing they came for?
In those cases, five sessions will find the serious problems, and a sixth will mostly confirm them.
When it does not
Distinct user groups. If your product serves administrators and end users, or beginners and experts, those are different tests. Five of each, not five in total — their problems genuinely differ.
Quantitative questions. "Which version converts better" is statistics, not observation. Five people tell you nothing about a conversion rate; you need enough traffic for the difference to be distinguishable from noise, which is usually orders of magnitude more.
Rare but severe problems. If a failure affects one person in twenty and matters enormously — a payment error, an accessibility barrier — five sessions will probably miss it.
Preference questions. "Which design do people prefer" needs many more people, and is a weak question anyway. What people say they prefer and what they can use are frequently different, which is why observing beats asking.
Accessibility. Testing with disabled users is not covered by a general sample. It needs deliberate recruitment of people who use assistive technology, and it finds problems no general session will.
Rounds beat sample size
The most useful reframing: do not think about how many people, think about how many rounds.
Five people, fix what you found, five more, fix again. Three rounds of five is dramatically more valuable than one round of fifteen, because the later rounds test improved designs rather than re-confirming known problems.
It is also easier to arrange. Finding five people is a week of light effort; finding fifteen is a project.
Recruiting well matters more than counting
The wrong five people are worse than useless, because they produce confident conclusions about the wrong audience.
Recruit people resembling your actual users in the ways that affect the task — domain familiarity, technical confidence, whether they have used the product before.
Do not test with colleagues. They know too much, and they are trying to be helpful, which is exactly what you do not want.
Include people who have never seen it. First-time comprehension is where most problems live, and you can only test it once per person.
Running them so the data is worth having
Give tasks, not instructions. "Book an appointment for next Tuesday" rather than "click the booking button."
Do not help. Watching someone struggle is uncomfortable and is the entire point. The moment you intervene, that data is gone.
Ask what they expected, when something surprises them. The gap between expectation and behaviour is where the design failure sits.
Watch what they do, discount what they say. People rationalise, and are polite about work they know you did.
Test one thing properly rather than everything superficially.
The honest summary
For finding usability problems with one group: about five, then fix, then five more.
For comparing conversion between designs: far more, and it is an analytics question rather than a research one.
For distinct audiences: five each.
For accessibility: recruit specifically.
And if you have tested nobody, testing one person tomorrow is worth more than planning a rigorous study you will not run. The most common failure here is not an inadequate sample size — it is testing with nobody at all.
If you want a round run properly on something before it ships, book a call.
Common questions
Is five users really enough for usability testing?
For finding usability problems in one interface with one reasonably similar group of users, yes — a small number of issues affect most people, so the first few participants surface most of them and later sessions largely repeat what you saw. The insight was that five people twice beats ten people once.
When do you need more than five test participants?
When you have distinct user groups, which each need their own round; when the question is quantitative, such as which version converts better; when a problem is rare but severe; when you are asking about preference; and for accessibility, which requires deliberately recruiting people who use assistive technology.
How many users do I need to compare two designs?
Far more than five, because that is a statistical question rather than an observational one. Comparing conversion rates requires enough traffic for the difference to be distinguishable from noise, which is usually orders of magnitude more than usability testing needs.
Should I test with colleagues?
No. They know too much about the product and are trying to be helpful, which is exactly what invalidates the session. Recruit people who resemble your actual users in the ways that affect the task, and include people who have never seen the interface before.
What is the biggest mistake in usability testing?
Helping. Watching someone struggle is uncomfortable and is the entire purpose — the moment you intervene, the data is gone. The second biggest is trusting what people say over what they do, since people rationalise and are polite about work they know you did.
