Research · Usability Testing · Insights

User Research

We research your users and test your product with them - uncovering needs and friction points alike. Clear insights, not just opinions, to guide every design decision.

There is a reliable pattern in product work. A team debates a decision for weeks. Opinions harden. Seniority starts to matter more than argument. Then someone puts the product in front of five real users, and within an hour the answer is obvious and nobody is arguing any more.

The debate was never about the product. It was about the absence of evidence.

User research is how you stop having that debate.

What is user research?

User research is the systematic study of how people behave, what they need, and where your product fails them. It replaces assumption with observation, and it produces evidence that decisions can be based on rather than defended with.

The distinction that matters most: research studies behaviour, not opinion. What people say they would do and what they actually do diverge consistently and predictably. Research designed around stated preference produces confident conclusions that are wrong — and wrong in a way that feels validated, which is worse than having no data at all.

Two broad categories, answering different questions:

  • Generative research — conducted before or early in design, to understand needs, context, and behaviour. Answers *what should we build?*
  • Evaluative research — conducted on something that exists, to find where it fails. Answers *does this work?*

Most teams do too little of the first and too much of the second, and then wonder why they're optimising something nobody wanted.

Why five users is usually enough

The most persistent objection to user testing is sample size. It's also the one most often misapplied.

For qualitative usability testing, five to eight participants per user segment typically surfaces the large majority of significant problems. This is a well-established finding, and the reason is structural: serious usability problems affect most users, so they appear quickly. The tenth participant rarely reveals a problem the fifth didn't.

For quantitative research — measuring how often something occurs, comparing designs statistically, sizing a market — five is nowhere near enough. Those questions need hundreds or thousands of responses.

Confusing the two is the most common research mistake in either direction. Teams either dismiss qualitative testing as unscientific, or draw statistical conclusions from eight interviews. Matching the method to the question is most of the skill.

Research methods and what each is for

In-depth interviews. Understanding context, motivation, and how a task fits into someone's wider work or life. Best for generative research, when you need to understand a problem space rather than evaluate a solution.

Contextual inquiry. Observing people doing the actual task in their actual environment. Consistently produces findings interviews miss entirely — workarounds, interruptions, and constraints people don't think to mention because they've stopped noticing them.

Usability testing. Participants attempt real tasks while you observe where they hesitate, fail, or take unexpected routes. The single most efficient method for finding problems in an existing product or prototype.

Moderated vs. unmoderated. Moderated sessions allow follow-up questions and catch the unexpected. Unmoderated tests are cheaper and faster, and suit well-defined tasks where you already know what you're measuring.

Surveys. For establishing prevalence — how many people experience something you've already identified qualitatively. Poor for discovery, because you can only ask about things you already thought of.

Card sorting and tree testing. Specifically for information architecture. Card sorting reveals how users group and label content; tree testing verifies whether a proposed structure lets people find things.

Diary studies. Tracking behaviour over days or weeks. Necessary when the experience unfolds over time — onboarding, habit formation, recurring tasks — and invisible in a single session.

Analytics and behavioural data. Not a substitute for research, but the cheapest available input. Shows where problems are, at scale, and tells you where to point qualitative methods.

Session recordings and heatmaps. Behavioural observation without recruitment. Excellent for spotting patterns worth investigating; limited for understanding why they occur.

The part that determines everything: recruitment

Research quality is capped by participant quality. Eight sessions with the wrong people produce confident conclusions pointing in the wrong direction — and there's no way to detect the error from the findings themselves.

Screening matters more than volume. A tight screener with a handful of qualified participants beats a loose one with twenty. Recruit for behaviour ("people who booked travel online in the last month") rather than demographics.

Existing users and prospective users answer different questions. Current users can tell you where your product fails; they can't tell you why people don't sign up. Churned users are frequently the most informative group and the least often contacted.

Beware convenience samples. Testing with colleagues, friends, or your most engaged power users produces flattering, unrepresentative results. Engaged users have already adapted to your product's problems.

Segment separately. If you have distinct user types, each needs its own participants. Five sessions spread across three segments is not a study of three segments.

What good testing looks like

Tasks, not tours. Ask participants to accomplish something real. Don't walk them through the interface — that tests your explanation, not the product.

Silence. The hardest discipline in moderation. When someone hesitates, the instinct is to help. Resisting it is where the finding is.

Ask what, not whether. "What are you trying to do?" produces information. "Was that confusing?" produces politeness. Participants want to be helpful, and leading questions are how research gets contaminated.

Watch behaviour over commentary. What someone does carries more weight than what they say about what they did. Where the two conflict, trust the behaviour.

Include the failure paths. Errors, empty states, and edge cases are where products lose people, and they're routinely excluded from test scripts because they're less pleasant to script.

Record and review. Live observation misses things. Recordings let findings be verified rather than remembered — and let stakeholders watch, which changes minds far faster than a report does.

From findings to decisions

Research that doesn't change anything is expensive entertainment. The synthesis step is where value is created or lost.

Group by pattern, not by session. A finding observed once is an anecdote. Observed in four of six sessions, it's a pattern worth acting on.

Separate observation from interpretation. "Three participants scrolled past the pricing section" is an observation. "Users don't care about pricing" is an interpretation, and possibly wrong. Keeping them distinct lets others check your reasoning.

Prioritise by severity and frequency. How badly does it affect the user, and how many users hit it? Together these rank the findings; either alone misleads.

Recommend, don't just report. A list of problems transfers work rather than reducing it. Every significant finding should carry a proposed direction.

Let stakeholders watch. Two minutes of a real user failing at a task moves an organisation further than twenty pages of findings. Where it's possible, it's the highest-leverage thing research can do.

Store it so it's findable. Most organisations research the same questions repeatedly because nothing was retained. A searchable repository turns individual studies into accumulating knowledge.

Common mistakes

Asking users what they want. People are experts in their problems and unreliable designers of solutions. "Would you use a feature that…" reliably produces enthusiastic agreement that never becomes behaviour.

Leading questions. "Did you find that easy?" gets a yes. The phrasing of a question determines its answer more than most teams realise.

Testing too late. Research after the build is complete produces findings nobody can afford to act on. Testing a prototype costs a fraction and arrives while change is still possible.

Researching to confirm. Studies commissioned to justify a decision already made are theatre. If the conclusion is fixed, skip the research and keep the budget.

Over-researching. Past the point of saturation, additional sessions stop producing new information. Recognising that moment and moving to synthesis is a skill; missing it burns budget.

Treating findings as mandates. Users report symptoms accurately and prescribe badly. What they struggled with is data. What they suggest building usually isn't.

One-off studies. Isolated research answers one question. Research as a habit compounds — and the second study always costs less than the first, because the infrastructure exists.

How we approach research

We design the study around the decision. The first question is always what you'll do differently depending on the result. If no answer changes anything, the study isn't worth running — and saying so is part of the job.

We study behaviour, not opinion. Sessions are built around real tasks in realistic conditions. Where we ask questions, they're about what happened, not what someone imagines they'd do.

We're rigorous about recruitment. Screening is where studies are made or ruined, and it's the step most often rushed. Getting the wrong participants invalidates everything downstream, invisibly.

We separate observation from interpretation. Findings are reported so you can see the evidence and evaluate the reasoning, rather than accepting conclusions on authority.

We bring your team into the room. Watching a user struggle changes minds in a way that reading about it does not. Where stakeholders can observe, they should.

We're honest about what the method can support. Eight sessions don't produce statistics. Where you need numbers, we'll say so and design for it rather than dressing qualitative findings in quantitative language.

The approach is grounded in certification from the Nielsen Norman Group — the institution that shaped modern UX practice — in usability testing, journey mapping, information architecture, UX leadership, and analytics, and thirteen years across agencies, startups, and corporate teams in insurtech, fintech, health, education, enterprise, and transportation.

Every engagement is run personally. No account layer between you and the person doing the work.

Frequently asked questions

How many users do we need to test with?

For qualitative usability testing, five to eight per distinct segment surfaces most significant problems. For quantitative questions — measuring frequency, comparing designs statistically — you need hundreds. The number depends entirely on the type of question, and conflating the two is the most common error in both directions.

What's the difference between user research and usability testing?

Usability testing is one research method, focused on whether people can use something. User research is the broader field, including interviews, contextual inquiry, surveys, diary studies, and more. Testing evaluates a solution; research more broadly can also explore a problem.

Can we test before we've built anything?

Yes, and it's usually the best time. Prototypes, wireframes, and even paper sketches can be tested. The earlier the test, the cheaper the change — and the more likely findings are to actually be acted on.

Do you handle recruitment?

Yes. Recruitment and screening are where studies succeed or fail, so it isn't a step worth delegating casually. Where you have access to your own users, that's often faster and better — we'll design the screener either way.

How long does a study take?

A focused usability test — planning, recruitment, sessions, synthesis — typically runs two to three weeks. Generative research with multiple segments runs longer, usually four to six. Recruitment is the most common bottleneck, particularly for specialist audiences.

What if the research contradicts what we believe?

That's the most valuable outcome available, and the cheapest place to find it out. Research that only ever confirms existing assumptions either wasn't needed or wasn't designed properly.

Can we run research ourselves?

Yes, and teams with the capability often should — it's faster and builds internal knowledge. What external practitioners add is study design, moderation discipline, and independence. If you want to build the capability rather than buy the study, that can be the engagement.

Isn't analytics enough?

Analytics tell you what happened and how often, at scale. They can't tell you why, and they can't tell you about the people who never arrived. The combination is far stronger than either alone: analytics locate the problem, qualitative research explains it.

Where to start

The diagnostic question: what's the most consequential assumption your product currently rests on, and how do you know it's true?

If the answer involves the word "obviously," that's where to start.