Author: AILesson9 min setupTested with:ChatGPTReviewed: 2026-08-28
Quick answer
Turn session observations into evidence-counted usability issues without blaming participants or overstating prevalence. Provide: Test notes and outcomes, Study and task definitions, Issue and decision rules. Expected result: A deduplicated issue register with task denominators, evidence, impact, hypotheses, severity rationale, and follow-up tests.
1
Add your context
Your text stays in this browser. AILesson Prompts does not send it to a model or server.
2
Your prompt
Unfilled fields remain visible as placeholders, so you can still copy and edit the prompt
Synthesize usability issues from the supplied test record.
Session notes:
[notes]
Study and task definitions:
[plan]
Issue and decision rules:
[decision]
Audit eligible sessions, task attempts, prototype versions, order, missing notes, moderator prompts, and completion definitions before coding. Separate direct observation, participant statement, calculated count, interface-cause hypothesis, and design recommendation. Group observations only when the user goal, obstacle, and consequence are materially the same; preserve counterexamples and variant differences. For each issue show affected task, n/eligible attempts, unaided versus prompted outcome, evidence references, behavior, consequence, suspected mechanism labeled as hypothesis, recurrence, impact, recoverability, severity rationale, confidence, and privacy-safe evidence. Do not diagnose participants, call behavior user error, infer population prevalence, average incompatible tasks, or count repeated events by one person as independent participants. Exclude prototype-only failures from product issues while retaining them in a study-quality log. Produce prioritized issues, non-issues and conflicting evidence, study limitations, design questions, and a verification plan.
Private by defaultPrompt assembly happens locally in your browser. Avoid placing confidential information into any AI service unless your organization allows it.
From input to outcome
A worked example
See how concrete context turns this recipe into a usable result
Actual input
Test notes and outcomes
Mobile membership-pause study, P01–P08. Task 1 pause two months; Task 2 edit resume date. P01 T1 opened Cancel first, returned, found Pause unaided, completed but said charge stopped today; T2 completed. P02 T1 completed unaided and explained all terms; T2 completed. P03 T1 found Pause after generic location prompt, then correct; T2 completed. P04 T1 found and completed but missed automatic resume; T2 tried tapping displayed date, then found Edit unaided. P05 T1 selected Cancel and would have confirmed; moderator stopped before final fictional action; T2 not attempted. P06 T1 correct; T2 prototype Edit link did not respond in version B, session stopped. P07 T1 correct but said two months meant sixty days; confirmation displays calendar dates; T2 completed. P08 T1 correct; T2 needed generic prompt. P01, P03, P05 saw Cancel above Pause; P02, P04, P06, P07, P08 saw Pause above Cancel. Quotes are paraphrased. No sensitive data.
Study and task definitions
Eight purposeful participants. T1 eligible P01–P08; success requires pause route plus correct start, no-charge period, current access, and auto-resume explanation. Partial if path complete but one term wrong. Critical error is unintended cancellation or believing charging stops immediately. T2 eligible only after T1 and working Edit control; P05 did not reach T2, P06 prototype failure makes T2 ineligible. Generic prompt is not unaided. Version B failed once only. Study cannot estimate population prevalence, retention, production performance, or accessibility.
Issue and decision rules
Product team decides what to revise before another prototype round. Severity considers task blockage, financial or account consequence, recurrence across participants, and recovery. Critical only for plausible irreversible or financial harm; high for major goal failure with weak recovery; medium for delay or misunderstanding; low cosmetic. Report n/eligible participants, not event counts. Minimum two participants for a recurrent theme, but single critical error remains visible. No participant labels beyond IDs. Do not prescribe final UI; provide testable design questions.
Example output
Usability issue synthesis
Pause and cancel are not reliably distinguished. Evidence: P01 entered Cancel then recovered; P05 selected Cancel and would have confirmed; P03 needed a generic location prompt. This affects 3/8 eligible T1 attempts, with 2/8 taking the wrong route and 1/8 prompted. P05 represents a critical-error scenario because account cancellation is materially different and recovery before final confirmation was not evidenced. Hypothesis: relative placement and labels may not make the temporary-versus-permanent distinction salient; the version pattern is suggestive because all wrong-route observations occurred with Cancel above Pause, but n=3 for that ordering and order was not isolated experimentally. Severity high, confidence medium. Design question: can hierarchy and consequence preview distinguish temporary pause before route selection?
Billing start is misunderstood after successful navigation. P01, 1/8, completed the pause path but believed charges stopped immediately. Single critical financial misunderstanding; recurrence unknown. Severity high due to consequence, confidence low-medium. Test whether the next-renewal start and current access can be paraphrased before confirmation.
Resume terms are incompletely understood. P04 missed automatic resume and P07 interpreted two months as sixty days: 2/8 partial T1 outcomes with different manifestations. Keep them under one “resume timing” family but separate subcodes. Severity medium, confidence medium. Test explicit calendar start/resume summary and comprehension.
Resume-date editing has discoverability friction. Among six eligible working-control T2 attempts, P08 needed a prompt and P04 first tapped the date before finding Edit: 2/6 friction, 1/6 prompted. Severity medium, confidence medium. The P06 failure is excluded from product findings and logged as a version-B prototype defect; P05 is not in the denominator.
Counterevidence: P02 completed and explained all terms; P06 and P08 completed T1 correctly; several participants recovered without help. Limits: purposeful sample, small version cells, no production behavior or population rates. Next round should repair version B, standardize ordering or deliberately test it, instrument route choice, require term paraphrase, and retest cancellation recovery without executing a real action.
Why this works
1
Task denominators and evidence references prevent memorable moments from becoming unsupported prevalence claims
2
Separating behavior from mechanism hypotheses keeps findings useful without pretending the interface cause is proven
Check the result
Does each issue cite eligible attempts and distinguish unaided, prompted, and prototype-failure outcomes?
Were observations merged only when goal, obstacle, and consequence match?
Are severity, mechanism, prevalence, recommendation, and confidence bounded by the study evidence?
Use it with confidence
Frequently asked questions
Practical answers about when to use this recipe, what to provide, and where human review still matters
What should I prepare before using “Synthesize usability issues from test notes”?
For “Synthesize usability issues from test notes,” prepare Test notes and outcomes, Study and task definitions, and Issue and decision rules. Replace placeholders only with information you can verify. If a detail is unknown, preserve that uncertainty explicitly instead of asking the model to infer it.
When is the “Synthesize usability issues from test notes” result not ready to use?
The result is not ready if it does not yet deliver the stated outcome—A deduplicated issue register with task denominators, evidence, impact, hypotheses, severity rationale, and follow-up tests—from the supplied evidence, or if it relies on unresolved assumptions, missing approvals, or invented details. Use the checks as release gates: revise the source inputs or assign a named, authorized reviewer instead of polishing an unsupported output.
Which AI tools have recorded tests for “Synthesize usability issues from test notes”?
The published test record for “Synthesize usability issues from test notes” lists ChatGPT as of 2026-08-28. This confirms recorded runs, not guaranteed compatibility or identical results in later product versions. For another tool or version, keep every constraint visible and repeat the result checks before use.