Definition
An evaluative method in which representative users perform realistic tasks with a service or interface while observers record measures of effectiveness (task success), efficiency (time and effort) and satisfaction to identify usability problems and inform design improvements.

Principle

Principle
Direct observation of representative users performing tasks reveals functional and interaction problems that specifications and expert review may miss; task performance metrics direct prioritization of fixes by frequency and severity of failure modes.

Demonstration

Demonstration
Illustrative scenario: In a moderated test, five public‑library patrons are asked to locate and download an article from a catalogue. Observers record task completion, steps taken, errors and time to completion; recurring navigation failures on a specific form field lead designers to simplify that field and re‑test.

Misapplication

Misapplication
Using unrepresentative participants (staff or convenience samples) or single‑session anecdotal notes to claim broad usability characteristics commits a semantic error: the method’s validity depends on participant representativeness, realistic tasks and systematic metrics, not on isolated impressions.

Consequence

Consequence
Properly conducted usability testing produces actionable, prioritized insights that reduce user errors and support interface changes; it requires investment in recruitment, scenario design and analysis and can expose trade‑offs between discoverability, efficiency and feature richness.

Reversal

Reversal
For large‑scale behavioral patterns, passive analytics or A/B testing may be more appropriate to measure frequency and magnitude of issues across populations; usability testing is most powerful for diagnosing why specific tasks fail rather than measuring aggregate incidence.

Boundary

Boundary
Clearly within: task‑based moderated or unmoderated sessions with representative users and recorded task metrics. Boundary case: expert heuristic evaluation diagnosing potential problems without users. Clearly outside: purely quantitative analytics that lack task‑level qualitative observation.

Semantic Tension

Semantic Tension
Depth of Insight ↔ Scale of Evidence — moderated usability testing provides deep diagnostic insight from few participants, while large‑scale analytics provide statistical coverage; both are complementary and must be balanced by project goals.

Synthesis

Synthesis
Usability testing converts observed user behavior into prioritized design actions: it trades breadth for diagnostic depth, making it the method of choice when the objective is to understand and fix concrete interaction failures rather than to estimate population rates.