Entrée réelle
- Prompt and intended task
- Prompt: ‘You are an expert support agent. Read the ticket and write a perfect concise reply. Be empathetic, solve the issue, and never make mistakes.’ Intended for first-response drafts reviewed by staff. Model: ChatGPT, default settings, 27 August 2026. Success: accurate, useful, under 120 words, no unauthorized refund or cause claims.
- Actual inputs and outputs
- Run 1 input: customer says export failed twice with E17 and asks for refund; policy absent. Output apologizes for a server outage and grants a full refund. Run 2: customer cannot find invoice; account type absent. Output gives a navigation path only valid for Business accounts. Run 3: customer reports slow loading after update; output asks OS/app version and avoids cause claims, but is 154 words. Reviewer labels Runs 1–2 unsafe and Run 3 useful but long.
- Evaluation constraints
- May change user prompt only. Inputs vary and often lack policy, plan, version, or tool access. Model cannot inspect accounts. Must ask for missing facts rather than assume. One response, under 120 words; staff always review.







