Compare the real behavior of two prompt versions

Автор: AILesson9 мин на настройкуПроверено на:ChatGPTПроверено: 2026-08-28

Быстрый ответ

Evaluate two prompts on matched inputs instead of judging wording, length, or a single attractive output. Укажите: Prompt versions and change intent, Matched test inputs and outputs, Success criteria and release context. Ожидаемый результат: A traceable comparison with per-case evidence, regressions, uncertainty, and a release recommendation.

1

Добавьте контекст

ваш текст остаётся в этом браузере. AILesson Prompts не отправляет его ни в модель, ни на сервер.

2

Ваш промпт

Незаполненные поля остаются видимыми как заполнители, поэтому вы всё равно можете скопировать и отредактировать промпт

Compare two prompt versions using their observed behavior on matched tests.

Exact versions, surrounding instructions, and intended change hypothesis:
[versions]

Case IDs, identical inputs, model/settings, outputs, repetitions, and evaluator labels:
[runs]

Must-pass rules, quality dimensions, weights, costly regressions, sample limits, and decision owner:
[criteria]

Do not infer behavioral improvement from prompt wording alone. Verify that runs are comparable; separate prompt changes from model, settings, tool, source, or evaluator changes. Diff the instructions to predict affected behaviors, then test those predictions against case-level evidence. Apply hard gates before weighted preferences. For each case compare evidence fidelity, task completion, unknown handling, format validity, safety/privacy, usefulness, verbosity, and any domain criteria; quote only short diagnostic excerpts. Record win, tie, loss, or invalid with rationale and evaluator disagreement. Do not average away a critical regression or claim statistical significance from a small convenience sample. Return: comparability audit; instruction diff and hypotheses; gate results; case matrix; aggregate counts with uncertainty; regressions and trade-offs; recommendation to adopt, reject, revise, or run more tests; minimal next experiment; and a regression set.
Попробовать в Playground
Конфиденциально по умолчаниюПромпт собирается локально в вашем браузере. Не вводите конфиденциальную информацию в сервисы ИИ, если ваша организация этого не разрешает.

От исходных данных к результату

Разобранный пример

Посмотрите, как конкретный контекст превращает этот рецепт в полезный результат

Реальный ввод

Prompt versions and change intent
A: Summarize the incident. Be concise and provide root cause, owner, and resolution date. B: Summarize only supplied incident evidence. Separate observed impact, confirmed cause, hypotheses, actions, owner, and date. Use unknown where absent and list the next verification. Change hypothesis: B reduces invented causes and commitments while remaining useful.
Matched test inputs and outputs
Same ChatGPT configuration, no tools, one run each. C1 database timeout: source names Dana owner and 30 Aug date but cause pending. A says root cause connection-pool exhaustion; B marks cause unknown and repeats owner/date. C2 conflicting notes: deploy suspected by engineer, metrics show errors started before deploy; no owner/date. A selects deploy and assigns SRE today; B preserves conflict and asks for timeline verification. C3 complete postmortem: confirmed expired certificate, Lee owner, fixed 12:10. Both accurate; A 72 words, B 118. C4 empty note. A invents monitoring failure and Ops owner; B returns insufficient evidence. C5 long routine incident with confirmed cause/action but no date. Both accurate; B says date unknown. Reviewer labels supplied for fidelity and usefulness.
Success criteria and release context
Hard gates: no invented cause, owner, date; material conflicts retained. Then fidelity 40%, action usefulness 30%, clarity 20%, concision 10%. Five convenience cases are diagnostic, not representative. Any hard-gate failure blocks adoption. Incident manager decides; one reviewer only.

Пример вывода

Comparability is acceptable for a diagnostic comparison: same cases, model configuration, tools, and one run per version. It is weak for stability because there are no repetitions and only one evaluator.

The diff predicts that B should preserve unknowns and conflicts but may be longer. Evidence matches that hypothesis. A fails hard gates in C1, C2, and C4 by inventing a cause or commitment; B passes all five reported cases. C1: B wins on fidelity with equal supported owner/date. C2: B wins by retaining the timing contradiction. C3: tie on fidelity and usefulness; A wins concision. C4: B wins the evidence gate. C5: B narrowly wins for explicit unknown date, with a verbosity cost. Aggregate: B wins four, ties one on primary outcome; this is not a population estimate.

Recommendation: reject A for this release and provisionally adopt B only after revision and regression testing. Shorten B’s headings for simple cases without removing unknown states. Next experiment: blind two incident managers to version labels, run three repetitions on at least 20 stratified cases including complete, missing, conflict, long, multilingual, and embedded-instruction inputs. Keep C1, C2, and C4 as must-pass regressions. Residual risk: preference scores are single-reviewer judgments, and no privacy or injection case was supplied.

Почему это работает

  1. 1

    Matched cases isolate prompt effects from changes in inputs and execution settings.

  2. 2

    Hard gates prevent an average quality score from hiding rare but consequential regressions.

Проверьте результат

  • Were both versions run on identical inputs, settings, tools, and evidence?

  • Is every conclusion traceable to case-level outputs and a stated criterion?

  • Are critical regressions, evaluator disagreement, and sample limitations visible?

Используйте уверенно

Часто задаваемые вопросы

Практические ответы о том, когда использовать этот рецепт, что нужно предоставить и где по-прежнему важна проверка человеком

What should I prepare before using “Compare the real behavior of two prompt versions”?

For “Compare the real behavior of two prompt versions,” prepare Prompt versions and change intent, Matched test inputs and outputs, and Success criteria and release context. Replace placeholders only with information you can verify. If a detail is unknown, preserve that uncertainty explicitly instead of asking the model to infer it.

When is the “Compare the real behavior of two prompt versions” result not ready to use?

The result is not ready if it does not yet deliver the stated outcome—A traceable comparison with per-case evidence, regressions, uncertainty, and a release recommendation—from the supplied evidence, or if it relies on unresolved assumptions, missing approvals, or invented details. Use the checks as release gates: revise the source inputs or assign a named, authorized reviewer instead of polishing an unsupported output.

Which AI tools have recorded tests for “Compare the real behavior of two prompt versions”?

The published test record for “Compare the real behavior of two prompt versions” lists ChatGPT as of 2026-08-28. This confirms recorded runs, not guaranteed compatibility or identical results in later product versions. For another tool or version, keep every constraint visible and repeat the result checks before use.

Продолжайте работу

AILesson · Рекомендуемые курсы

ваш следующий шаг: примените ИИ на практике

Перейдите от понимания ИИ к выполнению задач. Практикуйтесь в составлении запросов, проверке и улучшении результатов с помощью интерактивных уроков для работы и повседневной жизни.

Просмотреть все курсы