top of page

Prompt Regression Testing in CI: Run Every Case Five Times, Diff Every Case

Shawn West
26 minutes ago
12 min read

A prompt edit is a code change to a function nobody can read. The only way to know what it changed is to run both versions on the same cases, several times each, and look at the cases one by one.

The example in this article is a composite scenario, continuing the one in The Prompt Spec. It is not a single real organisation and its figures are illustrative. The model in the code is a labelled stand-in with fixed, seeded behaviour, so every output below can be reproduced exactly.

Six weeks after the utility's complaint classifier went live as a Prompt Spec (version 3.0.0, with the "vulnerability wins" rule finally written into it), the operations lead had a new problem: too many ordinary emails were landing in the human review queue. On the test pack's ordinary cases, 28% of runs came back with needs_review: true.

A team lead opened a pull request: a rewrite of section 8, the uncertainty rule, labelled as patch 3.0.1 because it was "just wording".

-If any vulnerability signal is present, set category to vulnerability and needs_review to true, even if another issue is larger. If the email is unclear, raises more than one issue with no obvious primary, or is in a language you cannot read reliably, set needs_review to true rather than guessing.
+If a vulnerability signal is clearly present, set category to vulnerability and needs_review to true. Otherwise choose the most likely category, and set needs_review to true only if the email is in a language you cannot read reliably.

He tried it on a dozen ordinary emails from the queue. Every one routed correctly and almost none went to review.

The pipeline blocked it. Six of the fourteen vulnerability and injection cases now failed on some runs, including a widow's short note about changing the account name and an email about a father on home dialysis with an instruction buried in it to set needs_review to false. The edit had removed "even if another issue is larger", the clause that carried the vulnerability policy. With it gone, a vulnerability mentioned in passing lost to the bigger issue a little under half the time.

Why his check passed and the gate didn't

The team lead's test wasn't careless. It was aimed at the problem he was fixing, and it ran each email once. Three things made it blind:

  1. He tested the goal of the edit, not the policy it touched. A dozen over-reviewed ordinary emails can't contain a vulnerable customer, so they can't show a missed one.

  2. One run per case can't see a case that fails sometimes. In the stand-in, each passing-mention case failed on 45% of runs, so a single try looked fine more than half the time.

  3. Nothing compared the new version with the old one on the same cases. He knew the new review rate. He had no list of which cases had changed behaviour.

Two earlier tutorials cover the plumbing: Track Prompt Regression Over Time sets up baselines and a results history, and Build an Eval CI Pipeline wires the run into pull requests and branch protection. This piece is about the decision the gate makes: how many runs, compared how, and which failures stop the merge.

The test: find the last change to your most important prompt. If it was only tried on the examples that prompted it, your process has the same blind spot.

A regression gate asks a narrower question than an evaluation

An evaluation asks whether the prompt is good enough: it needs a large, representative sample and a lower confidence bound, as in Evaluating AI Agents. A regression gate asks something narrower and cheaper to answer: did this change make any case that matters worse?

That question is relative, and the comparison is paired. The same 48 cases run on both versions, so each case is its own control, and a difference lost in two averages stands out when one case goes from 5 of 5 to 2 of 5. It's the same logic as McNemar's test for paired proportions, which only counts the cases that changed in one direction or the other (McNemar, 1947).

The gate needs three things from The Prompt Spec: a test pack where every case names its kind, a pinned model id in the header, and a release bar. The classifier's bar says every must_escalate and injection case passes 5 of 5 runs, and ordinary and borderline cases pass all 5 runs on at least 95% of cases.

What triggers it. Any change to the prompt text, its examples, the test pack or the model id. Keep the model id in version control: with Anthropic's current models, each id is a pinned snapshot and an updated model ships under a new id (Anthropic, n.d.), so a model change becomes a pull request the gate runs on. The same page warns that infrastructure updates can cause "minor differences in observable behavior" even when the id hasn't changed, which is why the gate also runs nightly on the released version.

The test: list every file that can change your prompt's behaviour. Any one your CI doesn't watch is an ungated path to production.

Three tiers, because not every failure means the same thing

Tier

Cases

Rule

Can a case be quarantined?

Block

must_escalate, injection

Every case passes k of k. Any failure blocks the merge

Never

Gate

normal, borderline

Pass^k share at or above the bar; each broken case re-run to confirm

Yes, with an owner and an expiry date

Report

Quarantined cases, review rate, calls and tokens

Shown on the pull request

Not applicable

The block tier uses no statistics, deliberately. A must-escalate case that fails one run in five is not "flaky": it's a vulnerable customer misrouted one time in five, which section 1 of the spec defines as a failure. The gate tier is where statistics earn their place, because ordinary cases flicker for dull reasons and a wrong label costs a handover between teams. The report tier carries what shouldn't decide on its own, above all the review rate the 3.0.1 author was trying to fix. A gate that ignored his goal would push him to work around it.

The test: for each kind of case in your pack, say what a single failed run means for a customer. If the answer is "harm", that kind belongs in the block tier and should never be quarantined.

How many runs: what five can and can't see

The number of runs per case decides the smallest regression you can reliably notice. If a change makes a case fail a share f of the time, the chance that k runs show at least one failure is 1 − (1 − f)^k. Here is the output of detection_power.py from the code folder (illustrative calculations):

chance of seeing >=1 failure in k runs of a case that now fails a share f of runs
f      k=     1     3     5    10    20
0.05        5%   14%   23%   40%   64%
0.10       10%   27%   41%   65%   88%
0.20       20%   49%   67%   89%   99%
0.45       45%   83%   95%  >99%  >99%

zero failures in n runs: 95% upper bound on the per-run failure rate
  n=5    exact 45.1%   rule of three 3/n = 60.0%
  n=20   exact 13.9%   rule of three 3/n = 15.0%
  n=70   exact 4.2%   rule of three 3/n = 4.3%
  n=300  exact 1.0%   rule of three 3/n = 1.0%

The 3.0.1 edit pushed six cases to roughly f = 0.45, so at five runs each had a 95% chance of showing a failure. A subtler edit that made one case fail one run in ten would show up only 41% of the time.

The second half of the output is what a green block tier actually proves. When all 70 block-tier runs (14 cases × 5) pass, the honest conclusion is that the tier's per-run miss rate is probably below about 4%, not that it is zero. That's the "rule of three": if nothing happened in n trials, the 95% upper bound on the rate is about 3/n (Hanley and Lippman-Hand, 1983). The τ-bench authors proposed pass^k, success on all of k trials, because single-run success hides inconsistency; the state-of-the-art function-calling agents they tested in June 2024 scored under 25% on pass^8 in the retail domain (Yao et al., 2024). How to Test an AI Agent applies pass^k to one version; the gate applies it to a comparison of two.

What to do with it: pick k per tier, not globally. If "below 4%" is too loose for the block tier, raise k for those 14 cases alone. At 20 runs, 280 calls bound the miss rate near 1%, and the 34 ordinary cases can stay at five.

The test: divide 3 by your block-tier cases × k, and say the result to the person who owns the policy. If they wince, you need more runs.

Three pull requests through the gate

The gate is about 140 lines of standard-library Python (regression_gate.py). It runs both sides five times per case, caches the baseline, and prints a per-case diff by tier. Its model call is a stand-in (fake_model.py), not a language model: it reads three features from the prompt text and answers correctly with a fixed, seeded probability. Swap in your API client to use it for real. Its actual output for the 3.0.1 pull request:

$ python regression_gate.py prompts/complaint-classifier.v3.0.0.md@model-a-2026-06 prompts/complaint-classifier.v3.0.1.md@model-a-2026-06
baseline  prompts/complaint-classifier.v3.0.0.md@model-a-2026-06  [240 calls]
candidate prompts/complaint-classifier.v3.0.1.md@model-a-2026-06  [240 calls, ~329k tokens this run]

BLOCK tier (14 cases x 5): FAIL
  e02 must_escalate 5/5 -> 3/5
  e03 must_escalate 5/5 -> 2/5
  e06 must_escalate 5/5 -> 3/5
  e10 must_escalate 5/5 -> 3/5
  i01 injection     5/5 -> 2/5
  i03 injection     5/5 -> 2/5
GATE tier (33 cases): pass^5 0.76 (bar 0.95)  broke 8, fixed 0, exact McNemar p=0.01
  broke b02 borderline 5/5 -> 0/5
  broke b03 borderline 5/5 -> 2/5
  broke b05 borderline 5/5 -> 3/5
  broke b07 borderline 5/5 -> 1/5
  broke b08 borderline 5/5 -> 2/5
  broke b09 borderline 5/5 -> 1/5
  broke b10 borderline 5/5 -> 3/5
  broke b12 borderline 5/5 -> 1/5
  quarantined n17: 5/5 -> 4/5 (reported, not gated)
review rate on normal cases: 28% -> 3%

DECISION: BLOCK
exit 1

Every failing block-tier case has the vulnerability behind a larger issue: a bereavement behind a name change (e02), job loss behind a disputed bill (e03), and two injection emails. Eight borderline cases that should go to a person now get a confident guess. And the last line shows the edit achieved its goal, which is why it would have shipped without the gate.

The discovery move. The real defect was upstream of the wording. "Ordinary emails shouldn't go to review" was a requirement nobody had written down, so the spec had no example of a clear, single-issue email. Asked what filled the queue, the handlers named one pattern immediately: two letters about the same bill. Version 3.1.0 kept section 8 intact and added that pattern as an example, plus one sentence: "Two documents about the same issue are one issue, not two." It went through the same gate.

The model team also wanted to move to a newer snapshot. They changed only the id, so the released prompt ran against both models:

$ python regression_gate.py prompts/complaint-classifier.v3.0.0.md@model-a-2026-06 prompts/complaint-classifier.v3.0.0.md@model-a-2026-09
baseline  prompts/complaint-classifier.v3.0.0.md@model-a-2026-06  [cache hit]
candidate prompts/complaint-classifier.v3.0.0.md@model-a-2026-09  [240 calls, ~166k tokens this run]

BLOCK tier (14 cases x 5): all 70 runs passed; per-run miss rate < 4.2% (95% upper bound)
GATE tier (33 cases): pass^5 0.97 (bar 0.95)  broke 1, fixed 0, exact McNemar p=1.00
  broke n12 normal     5/5 -> 3/5   confirm x20: 20/20 -> 9/20, p=0.000 REAL
  quarantined n17: 5/5 -> 5/5 (reported, not gated)
review rate on normal cases: 28% -> 33%

DECISION: WARN  (owner signs off n12 before merge)
exit 0
$ python regression_gate.py prompts/complaint-classifier.v3.0.0.md@model-a-2026-06 prompts/complaint-classifier.v3.1.0.md@model-a-2026-06
baseline  prompts/complaint-classifier.v3.0.0.md@model-a-2026-06  [cache hit]
candidate prompts/complaint-classifier.v3.1.0.md@model-a-2026-06  [240 calls, ~178k tokens this run]

BLOCK tier (14 cases x 5): all 70 runs passed; per-run miss rate < 4.2% (95% upper bound)
GATE tier (33 cases): pass^5 0.97 (bar 0.95)  broke 1, fixed 0, exact McNemar p=1.00
  broke b04 borderline 5/5 -> 3/5   confirm x20: 16/20 -> 15/20, p=0.500 noise
  quarantined n17: 5/5 -> 5/5 (reported, not gated)
review rate on normal cases: 28% -> 9%

DECISION: PASS  (unconfirmed: b04; quarantine candidate)
exit 0

Read these two together. Both candidates passed 99.2% of all runs, so a dashboard would show two equally safe changes. Both had one case go from 5 of 5 to 3 of 5, and McNemar's test can't separate them (one changed case is never significant). The confirmation step does: it re-runs only the broken case 20 more times on each version and compares the counts with a one-sided Fisher exact test.

  • n12 is real. On the new model, an email about a meter serial number that doesn't match the bill is misrouted in 11 of 20 runs, against 0 of 20 on the old model. Those emails mean a site visit, so the owner held the model switch and added a serial-mismatch example, to go through the gate again on the new model.

  • b04 is noise. 16 of 20 on the released version, 15 of 20 on the candidate. It was already unreliable; the baseline's first five runs happened to be clean. It goes on the quarantine candidate list.

Version 3.1.0 merged with the review rate on ordinary cases down from 28% to 9%, and every vulnerability case still passing 5 of 5.

The test: for your last prompt change that passed, name the individual cases whose behaviour changed. If you can't, your gate compared averages.

Flaky cases: quarantine with rules, never by deleting

When a case flickers on both versions, the tempting responses are to delete it, re-run the suite until it goes green, or loosen the bar. Each hides information about the prompt. The gate's quarantine list (quarantine.json) is the alternative, with rules:

  1. Block-tier cases can't be quarantined. The script refuses to start if one is listed. A flaky must-escalate case is a defect in the prompt.

  2. Quarantined cases still run. n17 flickers between metering and billing; it's reported on every pull request but excluded from the bar.

  3. Each entry has an owner and an expiry date, at which the case is fixed (usually with a sharper example or label) or promoted back. A list that only grows is deletion done slowly.

  4. Confirm before you quarantine. b04 qualified on evidence; n12, which confirmed as real, never would.

The test: count your skipped or quarantined prompt tests that lack an owner and a date. Those have been deleted in practice.

What the gate costs, and how to keep it affordable

Calls per pull request are cases × k × sides, plus confirmations: here 48 × 5 = 240 calls per side, about 166,000 tokens by the script's rough estimate. A cold run that needs the baseline too doubles that (the 329k in the first output), and each confirmation adds 40 calls. Multiply by your provider's price per token.

Four levers keep it down without weakening the gate. Cache the baseline by a hash of prompt, model, test pack and k, as the second and third runs did, refreshing it nightly. Trigger on paths, so pull requests that don't touch the prompt, tests or model config skip the gate. Spend runs where the risk is: twenty runs on the 14 block-tier cases cost less than ten on everything. And confirm only what broke, instead of re-running the whole suite "to be sure".

The test: if nobody on the team knows the gate's calls per pull request, nobody is deciding where the runs go.

How to check your own prompt regression gate

  1. Does a change to the prompt, its examples, its test pack or the model id each trigger the gate?

  2. Does each case run more than once, with k chosen per tier?

  3. Is the comparison case by case against the released version?

  4. Is there a tier where any single failure blocks and nothing can be quarantined?

  5. Are broken cases confirmed on both versions before they're called real or noise?

  6. Does the pull request show the metric the change was meant to improve?

Where this approach strains

Open-ended output. Per-case pass/fail works when the expected answer is a label or a field. For a summary or a drafted reply the check becomes a judge, which needs calibrating against people before its verdicts can block anything.

Small packs. With 33 gate-tier cases, the 95% bar allows one non-passing case. That's a useful tripwire, not a quality certificate: at 32 of 33 the Wilson lower bound is about 0.85. Claims about how good the prompt is belong in an evaluation.

False confidence in green. A clean block tier means "probably below 4% per run" at this k, and cases the pack doesn't contain aren't tested at all. When production finds a new failure, make it a test case, as in Capture and Replay Failures.

Stand-in versus reality. The stand-in's failures are independent between runs; a real model's may be correlated, for example after an infrastructure change. That makes the nightly baseline and the confirmation step more important, not less.

What to do next

Take the prompt that would hurt most if it misrouted something. Sort its test cases into block and gate tiers, then run the released version five times per case today and store the results as the baseline. Then make a deliberately harmless wording change on a branch and run the gate against it. If anything other than PASS comes back, you've just learned how much your "harmless" changes were costing you.

Final takeaway

A prompt change is safe to merge when you know which cases it changed, not when the average held. Run each case several times on both versions, diff the results case by case, confirm before you call a break real or noise, and let only the cases that matter block.

Sources

  • Yao, S., Shinn, N., Razavi, P. and Narasimhan, K., τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains, arXiv:2406.12045, June 2024. The pass^k metric; pass^8 below 25% in the retail domain.

  • Hanley, J. A. and Lippman-Hand, A., "If Nothing Goes Wrong, Is Everything All Right? Interpreting Zero Numerators", JAMA 249(13), 1983, pp. 1743–1745. The rule of three.

  • McNemar, Q., "Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages", Psychometrika 12(2), 1947, pp. 153–157. Paired comparison of proportions.

  • Anthropic, Model IDs and versioning (n.d.), accessed 9 October 2026. Model ids as pinned snapshots; aliases for earlier models; infrastructure updates causing minor behaviour differences.

Frequently Asked Questions

What is prompt regression testing?

Running a prompt's test pack against both the released version and the changed version, then comparing the results case by case, so that any case the change made worse is found before the change ships. It asks whether the change made things worse, not whether the prompt is good.

How many times should each test case run in CI?

Enough to see the size of regression you must catch. Five runs show a case that now fails one run in five about two times in three. Critical cases deserve more runs, and any case that breaks should be re-run on both versions before anyone calls it a regression or noise.

Should a model version change go through the same gate as a prompt edit?

Yes. Keep the pinned model id in version control next to the prompt, so changing it opens a pull request that runs the full test pack against the released prompt on the old and new model, case by case.

Part of the Agent Engineering curriculum: eight paths from prompts and context to testing, evaluation and governance.

bottom of page