Diagnose why a prompt performs poorly

Автор: AILesson7 мин на настройкуПроверено на:ChatGPTПроверено: 2026-08-28

Быстрый ответ

Compare intent, inputs, instructions, and real outputs to identify evidenced failure modes instead of guessing. Укажите: Prompt and intended task, Actual inputs and outputs, Evaluation constraints. Ожидаемый результат: A ranked diagnosis, minimal fixes, and rerun tests linked to observed failures.

1

Добавьте контекст

ваш текст остаётся в этом браузере. AILesson Prompts не отправляет его ни в модель, ни на сервер.

2

Ваш промпт

Незаполненные поля остаются видимыми как заполнители, поэтому вы всё равно можете скопировать и отредактировать промпт

Diagnose why the supplied prompt performs poorly using the actual runs as evidence.

Exact prompt, surrounding instructions, intended use, model/settings, and success criteria:
[prompt]

Representative inputs and outputs with dates:
[runs]

Allowed changes, limits, data/tools, required behavior, and known constraints:
[constraints]

Trace each observed failure to the smallest plausible prompt, input, context, evaluation, tool, or model cause. Check task ambiguity, missing evidence, conflicting instructions, overloaded steps, misplaced priority, undefined terminology, poor delimiters, underspecified output, absent unknown handling, examples that teach the wrong pattern, and success criteria that cannot be observed. Distinguish confirmed cause from hypothesis; do not blame randomness or the model without comparison evidence. Avoid rewriting until diagnosis is complete. Return: intended-versus-observed table; ranked failure modes with quoted evidence, confidence, and alternative explanation; prompt clauses that help or hurt; input and evaluation defects; minimal edit for each high-confidence issue; revised prompt; and rerun matrix that isolates one change at a time.
Попробовать в Playground
Конфиденциально по умолчаниюПромпт собирается локально в вашем браузере. Не вводите конфиденциальную информацию в сервисы ИИ, если ваша организация этого не разрешает.

От исходных данных к результату

Разобранный пример

Посмотрите, как конкретный контекст превращает этот рецепт в полезный результат

Реальный ввод

Prompt and intended task
Prompt: ‘You are an expert support agent. Read the ticket and write a perfect concise reply. Be empathetic, solve the issue, and never make mistakes.’ Intended for first-response drafts reviewed by staff. Model: ChatGPT, default settings, 27 August 2026. Success: accurate, useful, under 120 words, no unauthorized refund or cause claims.
Actual inputs and outputs
Run 1 input: customer says export failed twice with E17 and asks for refund; policy absent. Output apologizes for a server outage and grants a full refund. Run 2: customer cannot find invoice; account type absent. Output gives a navigation path only valid for Business accounts. Run 3: customer reports slow loading after update; output asks OS/app version and avoids cause claims, but is 154 words. Reviewer labels Runs 1–2 unsafe and Run 3 useful but long.
Evaluation constraints
May change user prompt only. Inputs vary and often lack policy, plan, version, or tool access. Model cannot inspect accounts. Must ask for missing facts rather than assume. One response, under 120 words; staff always review.

Пример вывода

Intended versus observed: Runs 1 and 2 violate the no-authorization/no-assumption criteria; Run 3 meets factual safety but violates length.

Ranked diagnosis: (1) High confidence—“solve the issue” conflicts with missing policy/account data and pressures the model to complete a resolution; evidence is two different unsupported completions. Minimal fix: require supported next steps and explicit unknown handling. (2) High—“expert” and “never make mistakes” do not define permissions or checks; they add no observable control. Remove. (3) High—“concise” conflicts with no numerical limit inside the prompt; Run 3 shows 154 words. Put ≤120 words in the output contract. (4) Medium—the absence of structured input fields makes missing plan/policy easy to overlook; add an evidence block or missing-information instruction. Alternative model variability is possible but not demonstrated.

Revised prompt: “Draft one first-response support message of no more than 120 words using only the ticket and supplied policy/account facts. Acknowledge the observed issue without naming an unconfirmed cause. Do not approve refunds, credits, replacements, or account changes unless an explicit supplied rule authorizes them. Do not claim to inspect tools or account state. If a needed plan, version, policy, or result is missing, ask at most two specific questions and give only safe steps supported by the input. Preserve exact error codes. Output the message only.”

Reruns: first add unknown/authority rules against Runs 1–2; then add length contract against Run 3; finally test an input with explicit refund authorization to ensure the safeguard does not block permitted action.

Почему это работает

  1. 1

    Actual run evidence separates prompt defects from missing input and subjective dissatisfaction.

  2. 2

    One-change reruns make improvements attributable instead of bundling several guesses.

Проверьте результат

  • Does every high-confidence diagnosis cite a repeated or directly traceable failure?

  • Are prompt, input, evaluator, tool, and model causes kept distinct?

  • Can the rerun matrix tell which edit caused an improvement or regression?

Используйте уверенно

Часто задаваемые вопросы

Практические ответы о том, когда использовать этот рецепт, что нужно предоставить и где по-прежнему важна проверка человеком

What should I prepare before using “Diagnose why a prompt performs poorly”?

For “Diagnose why a prompt performs poorly,” prepare Prompt and intended task, Actual inputs and outputs, and Evaluation constraints. Replace placeholders only with information you can verify. If a detail is unknown, preserve that uncertainty explicitly instead of asking the model to infer it.

When is the “Diagnose why a prompt performs poorly” result not ready to use?

The result is not ready if it does not yet deliver the stated outcome—A ranked diagnosis, minimal fixes, and rerun tests linked to observed failures—from the supplied evidence, or if it relies on unresolved assumptions, missing approvals, or invented details. Use the checks as release gates: revise the source inputs or assign a named, authorized reviewer instead of polishing an unsupported output.

Which AI tools have recorded tests for “Diagnose why a prompt performs poorly”?

The published test record for “Diagnose why a prompt performs poorly” lists ChatGPT as of 2026-08-28. This confirms recorded runs, not guaranteed compatibility or identical results in later product versions. For another tool or version, keep every constraint visible and repeat the result checks before use.

Продолжайте работу

AILesson · Рекомендуемые курсы

ваш следующий шаг: примените ИИ на практике

Перейдите от понимания ИИ к выполнению задач. Практикуйтесь в составлении запросов, проверке и улучшении результатов с помощью интерактивных уроков для работы и повседневной жизни.

Просмотреть все курсы