Compare the real behavior of two prompt versions

Author: AILesson9 min setupTested with:ChatGPTReviewed: 2026-08-28

Quick answer

Evaluate two prompts on matched inputs instead of judging wording, length, or a single attractive output. Provide: Prompt versions and change intent, Matched test inputs and outputs, Success criteria and release context. Expected result: A traceable comparison with per-case evidence, regressions, uncertainty, and a release recommendation.

1

Add your context

Your text stays in this browser. AILesson Prompts does not send it to a model or server.

2

Your prompt

Unfilled fields remain visible as placeholders, so you can still copy and edit the prompt

Compare two prompt versions using their observed behavior on matched tests.

Exact versions, surrounding instructions, and intended change hypothesis:
[versions]

Case IDs, identical inputs, model/settings, outputs, repetitions, and evaluator labels:
[runs]

Must-pass rules, quality dimensions, weights, costly regressions, sample limits, and decision owner:
[criteria]

Do not infer behavioral improvement from prompt wording alone. Verify that runs are comparable; separate prompt changes from model, settings, tool, source, or evaluator changes. Diff the instructions to predict affected behaviors, then test those predictions against case-level evidence. Apply hard gates before weighted preferences. For each case compare evidence fidelity, task completion, unknown handling, format validity, safety/privacy, usefulness, verbosity, and any domain criteria; quote only short diagnostic excerpts. Record win, tie, loss, or invalid with rationale and evaluator disagreement. Do not average away a critical regression or claim statistical significance from a small convenience sample. Return: comparability audit; instruction diff and hypotheses; gate results; case matrix; aggregate counts with uncertainty; regressions and trade-offs; recommendation to adopt, reject, revise, or run more tests; minimal next experiment; and a regression set.
Private by defaultPrompt assembly happens locally in your browser. Avoid placing confidential information into any AI service unless your organization allows it.

From input to outcome

A worked example

See how concrete context turns this recipe into a usable result

Actual input

Prompt versions and change intent
A: Summarize the incident. Be concise and provide root cause, owner, and resolution date. B: Summarize only supplied incident evidence. Separate observed impact, confirmed cause, hypotheses, actions, owner, and date. Use unknown where absent and list the next verification. Change hypothesis: B reduces invented causes and commitments while remaining useful.
Matched test inputs and outputs
Same ChatGPT configuration, no tools, one run each. C1 database timeout: source names Dana owner and 30 Aug date but cause pending. A says root cause connection-pool exhaustion; B marks cause unknown and repeats owner/date. C2 conflicting notes: deploy suspected by engineer, metrics show errors started before deploy; no owner/date. A selects deploy and assigns SRE today; B preserves conflict and asks for timeline verification. C3 complete postmortem: confirmed expired certificate, Lee owner, fixed 12:10. Both accurate; A 72 words, B 118. C4 empty note. A invents monitoring failure and Ops owner; B returns insufficient evidence. C5 long routine incident with confirmed cause/action but no date. Both accurate; B says date unknown. Reviewer labels supplied for fidelity and usefulness.
Success criteria and release context
Hard gates: no invented cause, owner, date; material conflicts retained. Then fidelity 40%, action usefulness 30%, clarity 20%, concision 10%. Five convenience cases are diagnostic, not representative. Any hard-gate failure blocks adoption. Incident manager decides; one reviewer only.

Example output

Comparability is acceptable for a diagnostic comparison: same cases, model configuration, tools, and one run per version. It is weak for stability because there are no repetitions and only one evaluator.

The diff predicts that B should preserve unknowns and conflicts but may be longer. Evidence matches that hypothesis. A fails hard gates in C1, C2, and C4 by inventing a cause or commitment; B passes all five reported cases. C1: B wins on fidelity with equal supported owner/date. C2: B wins by retaining the timing contradiction. C3: tie on fidelity and usefulness; A wins concision. C4: B wins the evidence gate. C5: B narrowly wins for explicit unknown date, with a verbosity cost. Aggregate: B wins four, ties one on primary outcome; this is not a population estimate.

Recommendation: reject A for this release and provisionally adopt B only after revision and regression testing. Shorten B’s headings for simple cases without removing unknown states. Next experiment: blind two incident managers to version labels, run three repetitions on at least 20 stratified cases including complete, missing, conflict, long, multilingual, and embedded-instruction inputs. Keep C1, C2, and C4 as must-pass regressions. Residual risk: preference scores are single-reviewer judgments, and no privacy or injection case was supplied.

Why this works

  1. 1

    Matched cases isolate prompt effects from changes in inputs and execution settings.

  2. 2

    Hard gates prevent an average quality score from hiding rare but consequential regressions.

Check the result

  • Were both versions run on identical inputs, settings, tools, and evidence?

  • Is every conclusion traceable to case-level outputs and a stated criterion?

  • Are critical regressions, evaluator disagreement, and sample limitations visible?

Use it with confidence

Frequently asked questions

Practical answers about when to use this recipe, what to provide, and where human review still matters

What should I prepare before using “Compare the real behavior of two prompt versions”?

For “Compare the real behavior of two prompt versions,” prepare Prompt versions and change intent, Matched test inputs and outputs, and Success criteria and release context. Replace placeholders only with information you can verify. If a detail is unknown, preserve that uncertainty explicitly instead of asking the model to infer it.

When is the “Compare the real behavior of two prompt versions” result not ready to use?

The result is not ready if it does not yet deliver the stated outcome—A traceable comparison with per-case evidence, regressions, uncertainty, and a release recommendation—from the supplied evidence, or if it relies on unresolved assumptions, missing approvals, or invented details. Use the checks as release gates: revise the source inputs or assign a named, authorized reviewer instead of polishing an unsupported output.

Which AI tools have recorded tests for “Compare the real behavior of two prompt versions”?

The published test record for “Compare the real behavior of two prompt versions” lists ChatGPT as of 2026-08-28. This confirms recorded runs, not guaranteed compatibility or identical results in later product versions. For another tool or version, keep every constraint visible and repeat the result checks before use.

More ways to explore

Where this recipe fits

Keep the work moving