In this guide
- Start with a decision your team can make
- Choose the method around the question
- Recruit for the work, not just the job title
- Write realistic tasks and prepare the environment
- Ask for evidence you can use
- Scope markets and accessibility deliberately
- Compare agency scope and costs
- Turn findings into changes and retesting
Your product team has redesigned an approval workflow. Stakeholders like the cleaner screens, but nobody has watched a customer submit a request, correct a mistake and check whether the request reached the right person. The unresolved question is practical: can the intended user complete the work without your team explaining the interface?
SaaS usability testing services help teams observe that interaction and decide what to change. A well-scoped engagement brings together representative participants, realistic tasks, a prepared test environment and analysis that connects observed problems to product decisions.
This guide is for SaaS founders and product leaders choosing a research partner in the US, UK, UAE or Dubai. It explains what to buy, what your team needs to provide and how to distinguish a useful study from a collection of opinions about screens.
Start with a decision your team can make
Write down the release or investment decision first. You might need to choose between two navigation structures, decide whether a new account setup is ready for development, or identify why administrators abandon a reporting workflow. Each question implies different tasks, participants and product access.
Keep the first study bounded. Testing every feature across every role often leaves too little time to investigate the difficult moments. Choose an important journey and include its exceptions: missing information, an unavailable action, a correction or a handoff. Agree which team can act on the findings and when that decision must happen.
An expert review and a usability study answer related but different questions. A SaaS product UX audit can identify likely issues from an experienced review. Testing observes what selected users actually do. Neither automatically establishes how common a problem is across your whole customer base. If your decision depends on population-level comparisons, ask for a separate measurement plan.
For SaaS teams, it helps to distinguish the buyer from the person doing the work. A procurement lead may approve the subscription while an operations specialist uses it every day. Testing only with the buyer can miss the workflow the product must support.
Choose the method around the question
Ask the agency to explain why its proposed format suits the uncertainty. Moderated sessions let a researcher observe the task and ask follow-up questions. Unmoderated sessions can work for clearly defined activities that participants can attempt independently. A prototype study helps assess a proposed flow; a live-product study can expose behaviour that a simplified prototype cannot represent.
| Decision | Possible approach | What to clarify |
|---|---|---|
| Understand confusion in a complex workflow | Moderated task sessions | How the moderator avoids coaching and records assistance |
| Check a focused, self-contained journey | Unmoderated task study | Whether instructions and recordings provide enough context |
| Evaluate a proposed design before implementation | Interactive prototype testing | Which states are simulated and which conclusions remain tentative |
| Compare performance over time | Defined benchmark study | Recruitment consistency, measures and the basis for sample size |
Do not treat an A/B experiment as interchangeable with task-based research. An experiment can compare a defined outcome under its own conditions; usability sessions can help explain how a journey becomes confusing. If both are proposed, request separate questions, methods and deliverables.
Agree the role of each observer. Product colleagues can learn from watching, but interruptions and explanations change the session. Ask for a moderator protocol, an observer channel and a way to capture questions for later. The purpose is to understand the interface from the participant's perspective.
Recruit for the work, not just the job title
A recruitment screener should establish whether a participant performs the relevant activity. “Finance manager” is less useful than knowing whether they review expense submissions, set approval rules or export monthly reports. Capture product familiarity, task frequency, decision authority and relevant tools without collecting information the study does not need.
Distinguish new users, experienced customers and people switching from another product. They bring different expectations. If you mix them, require the analysis to preserve those differences rather than summarising everyone as a single audience. The same applies to administrators, contributors and read-only users.
Ask how the supplier will find participants and verify fit. Existing customers may be easiest to reach, but a group nominated by account managers can overrepresent enthusiastic power users. A general panel may struggle to supply niche enterprise roles. Discuss those limits before committing to dates.
There is no useful universal participant count for every SaaS study. Ask the agency to justify its proposed sample against the number of distinct user groups, the complexity of the tasks and the kind of conclusion you need. A small qualitative round can reveal specific problems; it should not be presented as a precise adoption or success-rate estimate for the market.
Make recruitment responsibilities explicit: who sends invitations, confirms eligibility, handles incentives, reschedules no-shows and approves replacements. Ask what happens if a required role cannot be recruited. Reducing scope deliberately is preferable to quietly substituting unsuitable participants.
Write realistic tasks and prepare the environment
Give participants a situation and a goal, then observe their route. “A colleague needs access to view this month's report without editing it” is more revealing than “Click Settings, open Members and choose Viewer.” The latter tells the participant how to use the interface.
Nielsen Norman Group's task-scenario guidance recommends realistic activities that encourage action without giving away the interface steps. Ask the agency to share its discussion guide before recruitment closes so your product team can check the scenarios for accuracy.
Define what completion looks like and distinguish independent completion from success after assistance. Record wrong turns, hesitation, errors and recovery as well as the end state. A participant who reaches the right screen but assigns the wrong permission has not necessarily completed the intended task.
Prepare realistic sample content, account roles and relevant error states. An empty dashboard cannot reveal whether customers understand a crowded one. At the same time, a test should not require participants to expose private customer records or make irreversible changes to a real workspace. Use synthetic records and clearly bounded test accounts where appropriate.
Pilot the session before the main study. Confirm that links open, logins work, the prototype supports the intended route and recording captures the necessary detail. Decide how to handle a technical failure separately from a usability problem. Otherwise the report may attribute a broken test setup to the product design.
Ask for evidence you can use
A deliverable should connect each important finding to the task, participant context and observed behaviour. Ask for a concise issue register with supporting notes or appropriately shared clips, the consequence for the user, a confidence statement and a proposed next action.
Separate observation from interpretation. “The participant returned to the overview three times while searching for an export” is an observation. “The navigation label may not match their mental model” is an interpretation. “Rename the section” is a design option that still needs evaluation. Keeping those layers clear makes the report easier to challenge constructively.
Agree a prioritisation framework before the final presentation. Consider whether an issue blocks the task, causes a harmful mistake, affects an important role or has a workable recovery path. Frequency within a small study is one input, not a population estimate. Include serious one-off observations when their consequences warrant investigation.
Require the report to describe its limits: missing user groups, prototype shortcuts, moderator assistance, incomplete sessions and recruiting constraints. A short limitations section makes decisions more reliable. It should not be hidden behind a single confidence score.
AI-assisted transcription or summarisation may be part of the supplier's workflow. Ask which material is processed, where approved recordings are stored and who checks generated summaries against the original evidence. Product recommendations should remain traceable to what actually happened during the sessions.
Scope markets and accessibility deliberately
US, UK and UAE participants do not automatically need separate studies for every feature. Identify which differences could change the task: vocabulary, date formats, organisational roles, device use, language or working practices. Then recruit and design scenarios around those differences.
For a Dubai launch with Arabic in scope, specify whether research will run in Arabic, English or both. Include an appropriately fluent moderator and review the actual right-to-left experience, including mixed-direction content such as email addresses or product codes. Translating an English findings deck does not replace observing Arabic-language use.
Plan access needs during recruitment. Ask whether the agency can support participants' assistive technology, preferred communication methods and remote-session requirements. A prototype may not expose the semantics or keyboard behaviour of the eventual implementation, so be explicit about what can be evaluated at that stage.
Usability research with disabled participants and a technical accessibility audit can inform each other, but they are different scopes. Do not accept a small research study as a certificate that the product meets every accessibility requirement. Include implementation review where the launch decision depends on it.
Compare agency scope and costs
Send suppliers the same priority journey, user roles, product maturity, target markets and decision deadline. Ask for a workstream estimate rather than comparing session prices alone. A low headline fee may leave recruitment, analysis or a second round outside the proposal.
- Planning: research questions, stakeholder alignment, screener and discussion guide.
- Recruitment: sourcing, eligibility checks, incentives, scheduling and replacements.
- Preparation: prototype or test-account setup, sample data and a pilot.
- Fieldwork: moderation, observation arrangements, recording and language support.
- Analysis: evidence review, issue register, prioritisation and stakeholder workshop.
- Follow-through: design collaboration, retesting and delivery of agreed research assets.
Ask which assumptions would change the quote. Niche enterprise roles, several languages, a difficult prototype and multiple stakeholder review rounds can all change the work involved. Request named dependencies and milestone dates rather than a delivery promise that assumes recruitment will always go smoothly.
For agency selection, request a redacted example showing how a finding led to a decision. Discuss who will conduct the research and who will analyse it. A polished presentation is less useful than clear reasoning, careful moderation and a report your team can turn into work.
Turn findings into changes and retesting
End the study with a decision workshop involving product, design and engineering. For each material issue, record an owner, a proposed response and whether more evidence is needed. Some findings need a content change; others expose a permission model or workflow dependency that visual redesign alone cannot solve.
Connect findings to the backlog with enough context to survive handover. Preserve the task, observed difficulty, intended outcome and acceptance condition. Avoid turning a nuanced finding into a ticket that merely says “improve navigation.”
Reserve time to retest the changed journey. Use the same underlying goal where comparison is useful and account for differences in participants or product state. Check whether the fix solves the original problem and whether it creates a new one elsewhere.
Makreate's UX design services include research, prototypes and usability testing. Bring your priority journey, target roles and release decision to the first conversation. If the team is already rebuilding a trial experience, use the SaaS onboarding UX guide alongside this brief to connect research questions to the work being designed.
Frequently asked questions
What should SaaS usability testing services include?
A defined engagement should cover research questions, participant criteria, tasks, test preparation, sessions, analysis and a decision workshop. Confirm recruitment, incentives, language support and retesting explicitly in the proposal.
Can we test a prototype before development?
Yes, if the prototype supports the tasks and states the study needs. Document simulated behaviour and missing interactions, and avoid treating prototype findings as proof that the implemented product will behave the same way.
How much should usability testing cost?
Compare the work required for recruitment, user groups, languages, preparation, sessions, analysis and retesting. Ask for itemised scope and assumptions; a session-only price does not describe the full research engagement.
Planning a usability study?
Bring your product journey, target users and release decision to a scoping conversation with Makreate.
