Stop Guessing Whether the Design Works
Every design review contains the same unresolved argument: is this grip comfortable, is this button obvious, will people understand the icon. Nobody in the room is a user, and the loudest opinion usually wins. Usability testing replaces that argument with eight hours of observation and a ranked list of defects.
For a physical product it is cheap. Two days of sessions with eight to ten participants, run on printed or foam models, typically costs $3,000 to $9,000 including recruitment and incentives, and routinely catches problems that would cost a tooling revision at $8,000 to $30,000 later.
Write Tasks, Not Questions
The single biggest mistake is asking people what they think. Opinions are polite, unreliable, and shaped by whoever is holding the model. Behavior is not.
Convert every design question into a task with an observable outcome. Instead of "is the lid easy to open," the task is "open the container and pour 200 ml into this cup." Instead of "is the interface clear," the task is "set the device to run for 15 minutes on the low setting." Then you watch and count.
Good task sets share a few properties: they start cold with no explanation, they cover first-time use as well as routine use, they include at least one recovery scenario such as clearing a jam or replacing a battery, and they include the unglamorous operations people actually do, such as cleaning and storage. Keep it to six to nine tasks so a session fits in 45 to 60 minutes.
Who to Recruit and How Many
Five participants per distinct user group surfaces the large majority of severe problems. Two distinct groups, say first-time buyers and professional users, means ten sessions, not five. If you need numbers you can quote in a pitch deck, such as an average task time or a comparative score, plan for 15 to 20 per group instead.
Recruit on behavior, not demographics. "Owns a competing product and uses it weekly" is a useful screener; "aged 25 to 45" is not. Screen out designers, engineers, and anyone who has seen the product before. Expect $75 to $200 per participant in incentives for consumers and $200 to $600 for professionals such as nurses, electricians, or chefs. Where the product's dimensions themselves are in question, pair the sessions with the body-measurement work described in anthropometry in product design, and make sure the recruit spans the small and large ends of the range rather than clustering in the middle, a point reinforced in inclusive design.
What to Test With
Match the prototype to the question. Grip, weight, and reach questions need a solid, correctly weighted model, not a hollow print, and the approach is covered in ergonomic test models. Interface and sequence questions need something that responds, even if a person behind a curtain is triggering the lights. Deciding which build you need is the subject of looks-like vs works-like prototypes.
Build two or three copies. Participants break things, and a broken model ends a test day.
How to Run the Session
Two people run it: a moderator who talks and a notetaker who does not. Record hands and the product, not faces, which also simplifies consent.
- Hand over the product with no instructions. The first 30 seconds of unguided handling is the most valuable data in the session. Note which way up they hold it and what they touch first.
- Ask for think-aloud. "Tell me what you are looking for" produces the reasoning behind the error.
- Do not rescue. When a participant struggles, wait. Counting to ten in silence is uncomfortable and necessary. Log an assist only when they have genuinely stopped.
- Never explain the design. The moment you say "that button is for the timer," the session is over as a source of data.
- Ask about the last time, not the future. "When did you last need to do this?" gives facts; "would you buy this?" gives flattery.
- Debrief at the end only. Save all explanation and preference questions for after the tasks.
What You Actually Measure
Record a small, consistent set so results are comparable across participants and across design rounds:
Task success scored as completed unaided, completed with an assist, or failed. This is the headline number. Time on task, useful mainly as a comparison between design versions. Errors, logged with what the person did and what they expected instead. Assists and abandonments. Physical measures where relevant, such as the force needed to open a closure, measured with a gauge rather than estimated. A short subjective scale at the end, which is fine as a trend indicator and worthless as a single number.
A useful pass bar for a consumer product is 90 percent unaided success on core tasks and no single error appearing in more than two participants. For medical and other safety-critical products the bar and the method are formalized, as described in usability engineering for medical devices.
From Finding to Design Change
Within a day of the last session, list every observed problem with the number of participants affected and a severity rating: critical if it caused failure or a safety risk, major if it caused an assist or significant delay, minor if it caused hesitation only. Sort by frequency times severity and draw a line. Everything above it changes before design freeze; everything below it goes on a backlog with a date.
Write each fix as a design change with an owner, not as an observation. "Users could not find the release" is a note. "Move the release to the front face, increase to 0.6 in (15 mm), add a molded arrow" is a change. Then retest the changed items with five fresh participants, because fixes create new problems roughly a third of the time. Session logistics and consent details are covered further in user testing with a prototype, and the underlying design principles in ergonomics in product design.
Run a Test Before You Cut Steel
Projects House plans and runs usability sessions as part of product development: task design, recruiting criteria, test models, moderation, and a prioritized change list your engineering team can act on. Send your prototype status and target user through our contact form.