Home / Blog / Evaluating and observing agents / Regression Testing for LLM Applications: The CI…

Evaluating and observing agents

Regression Testing for LLM Applications: The CI Gate That Holds

Regression testing for LLM applications needs a gate that survives noise: a paired rule on flipped cases, three verdicts and a triage list. Copy the spec.

By Enrique Gutiérrez · Published · 23 min read

Regression testing for LLM applications means rerunning a fixed bank of graded cases on every change to a prompt, a tool definition or the model, then letting an enforced rule decide the merge. The rule that holds compares baseline and candidate on the same cases and judges only the cases that flipped.

This post gives you that rule with its arithmetic, a third verdict for unclear results, a spec to paste into a design document, and a procedure for the morning the gate goes red. The book sets the principles and I quote them. The numbers are mine, illustrative starting values, and each is marked where it appears.

What is regression testing for LLM applications, in the book’s words?

Regression testing for LLM applications is one habit with one rule, and Chapter 16 of the book states it in a sentence: “every change to a prompt, a tool definition, or the model reruns the suite, and the numbers decide.” The suite is an eval set, a bank of tasks with grading logic attached.

A suite turns into a gate at a precise moment. The chapter puts it this way: “A suite becomes a gate the day a threshold is enforced instead of glanced at: below this pass rate on the regression suite, the change does not merge; any failure on the safety cases, the pipeline stops, with no override wired into the normal path.”

The book’s free glossary defines the regression gate as “An eval set wired into release machinery as an automatic barrier: the change ships only if the bank of graded tasks still passes.” A CI regression gate for LLM agents differs from an ordinary test job in three questions an ordinary test never raises: how many runs, against what baseline, with what tolerance.

Which suite does the gate read? Capability evals are the hard tasks the agent mostly fails. Regression evals are the tasks it reliably passes, and “their job is to stay near one hundred percent and scream when anything slips.”

What does the book fix, and what does this post add?

The book fixes six things about a gate and leaves every number open. This post fills the open slots with one defensible rule, and the table keeps the two apart.

Element of the gate Source What it says
An enforced threshold Book, Chapter 16 “a threshold is enforced instead of glanced at”
Zero tolerance on safety cases Book, Chapter 16 “any failure on the safety cases, the pipeline stops”
Repeated runs Book, Chapter 16 “give each task a few trials”
Small shifts are noise Book, Chapter 16 “a two-point shift on a fifty-task suite is well inside the noise”
Smoke subset and full set Book, Chapter 16 “a smoke subset per commit and the full battery nightly and before release”
Fail closed Book, Chapter 16 “fail closed: uncertainty blocks rather than waves through”
Baseline against candidate on the same cases This post The paired design
Margin, significance level, flip cap This post Illustrative values your team must set
Pass, block, inconclusive This post The three-verdict rule
Majority of three, all runs for safety This post Aggregation of repeated runs
Flaky list and triage order This post Procedure for flaky cases and a red gate

Why does a fixed threshold on a raw pass rate cry wolf?

A fixed threshold on a raw pass rate cries wolf because the pass rate moves when nothing in your system changed. In regression testing for LLM applications, three sources of noise move it: the sample of cases, the variation between runs of the same input, and the grader when the grader is a model.

Take a gate that fails the build under 90 percent, on a suite that truly sits at 92. On 50 cases, a single run of that unchanged suite lands under 90 percent about one time in five (binomial, 44 or fewer passes of 50: 0.21). Such a gate fires on changes that touched nothing.

How much does the sample of cases move the number?

The sample of cases moves the number by an amount that shrinks only with the square root of the set size. The book’s own line is that one run per task “makes the gate itself flaky,” and that “a two-point shift on a fifty-task suite is well inside the noise.”

The post on how many eval examples you need derives the intervals and the sample sizes. The gate needs one consequence only: on the few hundred cases a CI job can afford, gaps of two to four points between two separate runs are ordinary.

Does temperature 0 remove run-to-run noise?

Temperature 0 reduces run-to-run noise and does not remove it on hosted models. The book says so directly: “even with the randomness dialed to zero a hosted model does not promise identical replies.” Two papers measured it, and each states limits that matter for how far you carry the result.

Atil and twelve co-authors (2025, arXiv:2408.04667v5, a preprint) ran five models on eight multiple-choice tasks, ten runs each, with temperature 0, top-p 1 and a fixed seed. Their abstract reports “accuracy variations up to 15% across naturally occurring runs”, and that “none of the LLMs consistently delivers repeatable accuracy across all tasks”. The paper states its limits: the answer extraction “has many hard-coded parts, which reduces the generality of the system”, and the models sat behind APIs, so “we can only speculate about the reason”.

Ouyang, Zhang, Harman and Wang (2024, arXiv:2308.02828v2) asked one provider’s hosted chat models for code on 829 problems from three benchmarks, five requests per problem. At temperature 0 they found “43.64% (CodeContests), 27.40% (APPS), and 18.29% (HumanEval) of problems with no equal test output among the five code candidates.” The authors place their threats to external validity in the datasets, the model versions and the prompt design of their study.

Neither paper gives a general flake rate. What transfers is the direction: non deterministic LLM tests are the normal case, and a gate that assumes identical reruns contradicts the measurements. The Atil paper names the remedy in passing: “An alternative might be regression testing that tolerates variability.”

On Hacker News in April 2024, the user mg sent one prompt twice with temperature 0 and a fixed seed and wrote: “And I got two different replies.”

What does a model-based judge add?

A model-based judge adds a second layer of noise, because the same output can receive different verdicts on different calls. A gate that uses a judge is measuring the system and the grader at once.

The cleanest way to separate them came from a maintainer of an open-source eval framework, answering a September 2026 issue from the user Abelo9996. Repeating whole runs will not isolate the judge. In the maintainer’s words, repeats “re-run the solver as well as the judge, so disagreement across epochs mixes model variance with judge variance”. The fix is to store one output and judge that fixed output several times.

Calibrating the judge against human labels is a separate job, covered in the post on whether an LLM judge is reliable. Prefer a deterministic check wherever one can decide the case.

What do teams that run gates say about the threshold?

Teams that publish their practice describe the gate and leave the number out. I read two first-hand accounts for this post, both cited by the book, and neither states a numeric threshold or a run count.

Anthropic’s engineering guide (January 2026) says: “Because model outputs vary between runs, we run multiple trials to produce more consistent results.” It adds that regression evals “should have a nearly 100% pass rate.” Hamel Husain’s 2024 essay says: “Your pass rate is a product decision, depending on the failures you are willing to tolerate.”

In an October 2025 issue on another open-source eval framework’s tracker, the user denis-snyk wrote: “I want to assert that each test passes 8/10 of those repeats.” Nobody there asked where 8 of 10 comes from. I name the two trackers linked in this post, promptfoo and Inspect, only as dated examples of a category.

What is the paired design?

The paired design runs the baseline and the candidate on the same fixed cases and compares them case by case. For regression testing for LLM applications it is this post’s addition: the book frames the gate as a floor on a pass rate and does not describe a stored baseline.

Every case then lands in one of four cells: passed under both, failed under both, flipped to fail, or flipped to pass. The two versions agree in the first two cells, so all the information sits in the flips.

Evan Miller’s 2024 paper on error bars for evals (arXiv:2411.00640v1) makes the statistical case. Question scores tend to be positively correlated across versions, so in his words “paired differences represent a ‘free’ reduction in estimator variance when comparing two models.” He recommends the paired analysis “wherever practicable.” His is a methods paper built on large-sample approximations.

Pairing has one precondition: both arms see the same cases, grader and tool fixtures, and only the change under test differs.

The gate rule: how do you decide on the flipped cases?

You decide with an exact binomial test on the flipped cases, plus two guards, and you get one of three verdicts. If the change did nothing, each flipped case was equally likely to flip either way, like a fair coin.

Call b the number of cases that flipped to fail and c the number that flipped to pass. The one-sided p-value is the chance of seeing b or more failures among b + c coin tosses. By hand: add the binomial coefficients from b up to b + c, then divide by 2 raised to the power b + c.

This is McNemar’s test in its exact form. The NIST reference page describes it for “two paired variables where each variable has exactly two possible outcomes” and uses the binomial statistic when b + c is 20 or fewer. NIST prints a two-sided test. The one-sided reading, and its use at any count of flips, are my choices for a gate that asks only “did it get worse?”

Here is the rule, with illustrative values your team must replace with its own:

Verdict Condition Action
Block Any safety case failed, or the one-sided p is at or below 0.05 No merge; run the triage list
Pass Otherwise, when the net drop b − c is within the margin and b is at most the flip cap Merge
Inconclusive Everything else Escalate once, then fail closed

The margin is the drop you would ship without a second look: 1 point here, which is 2 cases on 200. The flip cap bounds how many cases may break even when as many others get fixed; I set it at 10 on 200 cases here. The 0.05 is a convention. Choose all three before the first red build, because a threshold chosen while the gate is red will be whatever lets the change through.

What does the worked example show?

The worked example shows a 4-point drop that the rule blocks, and the same 4-point drop that it cannot call. Take 200 cases. The baseline passes 184 and the candidate passes 176.

First split: b = 10 and c = 2. The cells are 174 passed under both, 14 failed under both, 10 flipped to fail and 2 flipped to pass; 174 + 10 = 184 and 174 + 2 = 176.

Twelve cases flipped. The ways to get 10, 11 or 12 failures among 12 tosses are 66 + 12 + 1 = 79, and 2 raised to the power 12 is 4,096. So p = 79 / 4,096 = 0.019, and the verdict is block.

Second split: b = 20 and c = 12, the same net drop of 8 cases. Now 32 cases flipped, and the same sum gives p = 0.108, above 0.05, so no block. The net drop of 8 cases exceeds the 2-case margin, so no pass either. The verdict is inconclusive.

A gate that reads only the pass rate cannot tell these two changes apart.

For quick checks, this table gives the smallest b that blocks at one-sided 0.05 for a given c, computed from the same sum.

Flipped to pass (c) Smallest b that blocks One-sided p at that b
0 5 0.031
1 7 0.035
2 9 0.033
3 10 0.046
4 12 0.038
5 13 0.048

The table does not depend on the size of the set.

What happens on an inconclusive verdict?

An inconclusive verdict triggers one escalation that was written down before the run, and then the gate fails closed. The escalation has two forms. If the job ran the smoke subset, run the full set. If it ran the full set, rerun every flipped case with more runs on both arms.

Rerunning only the cases that broke would give the candidate extra chances and tilt the result toward green. Rerun the fixed ones as well, under the same aggregation rule, recount b and c, and apply the rule a second time.

If the second verdict is still inconclusive, the change does not merge on the automatic path. A named reviewer reads the broken cases and records a decision with a reason. That applies the book’s rule to my third verdict: “fail closed: uncertainty blocks rather than waves through”.

What does the unpaired reading of the same numbers say?

The unpaired reading says the 4-point gap “could easily be noise”, which is why the paired count is worth the bookkeeping. The calculator below opens on the worked example, 184 of 200 against 176 of 200.

With JavaScript on, the Eval sample-size calculator runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Eval sample-size calculator on its own page to share a result by link.

This tab treats the two runs as independent samples, as if each version had been run on its own 200 cases. It displays −4.0 points, a two-sided p of 0.18, and about 882 tasks per version to detect a 4.0-point difference reliably. It does not compute the test on flipped cases; that step is the hand arithmetic above.

The unpaired test charges both runs for the variation between easy and hard cases, although the same cases were used twice. On the same cases, the 188 that agree cancel out, and only the 12 flips are left to explain.

Which gate design decisions are yours to make?

Nine decisions define a gate, and each has a default that depends on which suite it governs. The options and conditions are this post’s, except where a row quotes the book. Filter by suite to see the rows that apply.

Decision Options When Suite
What is compared A floor on the pass rate, or baseline against candidate on the flips A floor only while the suite is tiny and at 100%; the paired rule once a baseline is stored smoke, full
Baseline source A cached run of the main branch, or a fresh run in the same job Cached while the cache key is unchanged; fresh when any part of the key moved smoke, full
Runs per case One; three with a majority; every run must pass One for deterministic checks on a stable step; three for agent runs; every run for safety cases smoke, full, safety
Decision rule A fixed point threshold, or the exact test with margin and flip cap A fixed threshold only after same-version runs show the noise sits well under it smoke, full
Safety rule Zero tolerance, counted apart from the rate Always; the book: “no override wired into the normal path” safety
Scope per trigger Smoke subset per change; full set nightly and before release The book’s accommodation for cost smoke, full
Infra failures A separate status, retried, with a cap Always; too many infra failures fail the job and leave the change unjudged smoke, full, safety
Inconclusive One escalation, then a named reviewer Always defined in advance smoke, full
Second metrics Budgets for cost, latency and step count Always on the full set; the book: “track cost, latency, and step count on the same fixed bank” full

The infra row comes from the book’s instruction to “keep infrastructure failures (timeouts, crashed sandboxes, rate limits) out of the quality number entirely, tracked under a separate status”. Counting a timeout as a failed case teaches the team that red means rerun. The last row is the book’s as well, and the post on AI agent evaluation metrics covers which of those metrics to report first.

How many runs per case, and how do you aggregate them?

Three runs per case with a majority verdict is a reasonable start for agent runs, and the aggregation rule must be fixed before the run. The book asks for “a few trials” at the gate and, in general, “a handful at minimum, a couple of dozen when the stakes justify it; the exact count is an economics question, and any figure is illustrative”. The count of three and the majority rule are mine.

The flip test needs one pass-or-fail verdict per case, so repeated runs must collapse to one. An ordinary case passes when most of its runs pass. A safety case passes only when every run passes.

A case that truly passes 90 percent of the time passes a majority of three with probability 0.972 and all three with probability 0.729. The strict rule fails a decent case more than a quarter of the time, which is acceptable only where strictness is the point. The post on pass@k and passk owns that distinction, and the pass@k calculator computes it.

More runs have a ceiling. In Miller’s stylized example, two runs per question cut the variance of the score by a third, four runs by half, and no number of runs by more than two thirds. He also warns that pooling all runs as if they were separate cases “will be inconsistent”. The case stays the unit: collapse each case first, then count flips across cases.

How do you handle flaky cases?

You handle flaky cases by measuring them with a same-version run and giving each one an owner. This procedure is this post’s own. The book’s one line on the subject concerns a bloated suite of low-signal checks: “Prune it the way you would prune flaky tests, and for the same reason.”

Run the baseline against itself a few times before trusting any threshold. Nothing changed between those runs, so every flip is noise. Set the flip cap above the largest one-direction count you see.

Any case that flips between identical baseline runs goes on a flaky list with an owner and a date, in a separate, reviewed change. A listed case still runs and appears in the report, and it is left out of b and c until repaired. Repair means fixing an ambiguous criterion, fixing the grader, or deleting the case in a reviewed change.

Cap the list at an illustrative 5 percent of the suite, so a growing list forces repair. Never put a safety case on it: a safety case that flips is a finding about the agent.

Why do safety cases get zero tolerance?

Safety cases get zero tolerance because the book says so and because statistics cannot excuse a never-event. The sentence again: “any failure on the safety cases, the pipeline stops, with no override wired into the normal path.”

A safety case asserts something the agent must never do, such as issuing a refund above its limit. The flip test asks whether a change made things worse on balance, and nobody trades one leaked record against two fixed formatting cases.

So the safety rule runs first. One failed run on one safety case blocks the change, whatever the flip counts say. The post on agent trajectory evaluation shows how to write these checks against the path the agent took.

What does the gate cost per change?

The gate costs cases times runs per case times model calls per run, for each arm you execute. The book states the principle: “Multi-trial evaluation multiplies your compute bill by the trial count.” The inputs below are illustrative, and I leave prices out because they date within months.

Take 200 cases, 3 runs each and an agent that averages 12 model calls per run. One arm is 200 × 3 × 12 = 7,200 model calls, and a fresh baseline beside the candidate doubles it to 14,400. A 40-case smoke subset at 3 runs is 40 × 3 × 12 = 1,440 calls, a fifth of one full arm.

Three levers cut the bill. First, the book’s split: “the standard accommodation, familiar from slow test suites, is a smoke subset per commit and the full battery nightly and before release.” Second, trigger the job only when a prompt, a tool definition, the model or the harness code changed. Third, cache the baseline run of the main branch, which halves every job.

A cached baseline is valid only while its key is unchanged: model version, prompt and tool-definition revisions, judge version and eval-set revision. If any part moved, the cache describes a different system.

Never cache the candidate’s runs or repeated runs of one case: they are the measurement. Safety-case runs are always live.

What does the gate spec look like on paper?

The gate spec is one page of plain text that states the trigger, the sets, the runs, the baseline, the rule and what is forbidden. Every number in it is an illustrative starting value. Replace the numbers before the first red build.

REGRESSION GATE SPEC  (all numbers are starting values; set your own)

TRIGGER
  Any change to: prompts, tool definitions, model version, harness code.

SETS
  smoke   : 40 cases, on every triggering change
  full    : 200 cases, nightly and before release
  safety  : all safety cases, in both smoke and full

RUNS AND AGGREGATION
  ordinary case : 3 runs, passes if at least 2 pass
  safety case   : 3 runs, passes only if all 3 pass
  infra failure : separate status, retry up to 2 times, never counted
                  as a fail; more than 5% infra failures fails the JOB

BASELINE
  Cached run of the main branch, same cases, same aggregation.
  Cache key: model version, prompt revision, tool-definition revision,
             judge version, eval-set revision.
  Key changed outside this change -> rerun the baseline in this job.

FLAKY LIST
  A case that flips between identical baseline runs. Added only in a
  separate, reviewed change. Capped at 5% of the suite.

COUNT (on cases not on the flaky list)
  b = cases that pass on baseline and fail on candidate
  c = cases that fail on baseline and pass on candidate
  p = P(X >= b), X ~ Binomial(b + c, 1/2)      (exact, one-sided)

VERDICT (first matching line wins)
  1. any safety case failed                      -> BLOCK
  2. p <= 0.05                                   -> BLOCK
  3. b - c <= margin  AND  b <= flip cap         -> PASS
  4. otherwise                                   -> INCONCLUSIVE
  margin  : 2 cases on the full set (1 point), 1 case on the smoke set
  flip cap: 10 on the full set, 4 on the smoke set

ON INCONCLUSIVE (once per change)
  smoke -> run the full set and apply the rule again
  full  -> rerun EVERY flipped case (b and c) on both arms with 5 runs,
           majority verdict, recount, apply the rule again
  still inconclusive -> no automatic merge; a named reviewer reads the
           broken cases and records the decision and the reason

SECOND METRICS (full set)
  cost, latency, step count: each with its own budget

OUTPUT OF EVERY RUN
  verdict, b, c, p, the list of flipped cases, infra count,
  cache key, flaky-list size

FORBIDDEN
  rerunning until green; raising a threshold in the change under test;
  deleting or editing a case in the change under test; adding a case
  to the flaky list in the change under test; caching candidate runs;
  putting a safety case on the flaky list

What do you do when the gate fires?

When the gate fires you work through a fixed list that sorts the cause into one of five classes before anyone touches the change. This triage order is this post’s own. It rests on one sentence of the book: “you can only attribute a regression if you can name the change, so prompts and tool definitions are versioned in the repository alongside the code”.

  1. Infra or quality? Timeouts, rate limits and crashed sandboxes are retried under their own status and do not count.
  2. Safety case? If a safety case failed, the block stands. Reproduce it to understand it; no rerun lifts it.
  3. List the flipped cases, both directions. Cases already on the flaky list are not counted.
  4. Changed dependency? Compare cache keys: model version behind an alias, tool backend, fixture data, eval-set revision. If something moved outside the change, rerun the baseline and apply the rule again.
  5. Judge moved? Re-judge the stored baseline outputs with the current judge. Different verdicts on identical outputs mean the grader changed; fix it first.
  6. Stale case? Read the transcripts of the broken cases. If the expected answer is out of date or unclear, repair the case in a separate, reviewed change.
  7. Flaky case? Rerun the baseline against itself on every flipped case, in both directions. A case that flips between identical baseline runs goes on the flaky list in a separate, reviewed change, never a safety case, and the verdict is recomputed once without those cases.
  8. Real regression? Rerun the remaining broken cases with more runs on both arms. If the candidate’s pass rate on a case stays lower, the change caused it. This step explains the block and does not lift it.
  9. Record the cause class: real regression, flaky case, judge, dependency or stale case.

Rerunning until green is absent from the list: a case that truly passes half the time goes green at least once in three tries 87.5 percent of the time, so retrying turns a gate on reliability into a gate on luck.

How do three example changes come out?

Three invented changes run through the spec and the triage list get one verdict each, with no step contradicting another. All use the 200-case full set and the illustrative values above.

Change Flipped to fail (b) Flipped to pass (c) One-sided p Verdict
Prompt wording edit 9 3 299 / 4,096 = 0.073 Inconclusive
Model-version upgrade 14 14 0.575 Inconclusive
Refactor of a tool wrapper 1, a safety case 0 0.5 Block

The prompt edit. No safety case failed, and p = 0.073 is above 0.05, so lines 1 and 2 do not fire. The net drop is 6 cases, 3 points, above the 2-case margin. The verdict is inconclusive: all 12 flipped cases rerun on both arms with 5 runs, and the rule is applied once more.

The model upgrade. The pass rate did not move, and a gate on the raw rate would wave this through. The flip cap catches it: 14 cases that used to work now fail, above the cap of 10. The verdict is inconclusive, and the 28 flipped cases rerun. The model version is the declared change here, so triage step 4 asks for no rebaseline; after a merge the cache key changes and the baseline is rebuilt. The refactor needs no walk-through: line 1 fires before any counting.

Where does the offline gate stop and online checking start?

The offline gate stops at the edge of the dataset: it answers whether this change broke something you had already thought to test. On offline vs online evals for LLM apps, the book gives the offline side the question of whether a change helped. Offline, it says, “is the only place that question has a clean answer, because the dataset holds still”.

Eval-driven development as plumbing.
Figure 16.4 Eval-driven development as plumbing. Every change runs against two suites: the capability evals it is meant to climb, and the regression evals it must not break; tasks that stabilize are promoted from the first into the second. The change reaches a gate (in accent) where a threshold is enforced, not glanced at—the numbers, not moods, decide whether it merges or is blocked. Reuse this diagram

Production asks what is happening to real users right now, and the glossary marks the boundary: “The gate tests the failures you already imagined; the rest of the rollout ladder exists because reality imagines better”. Sampling live runs, alert rules and drift checks belong to the post on monitoring AI agents in production. Between the gate and full traffic sit two more rungs, where teams shadow deploy an agent and then canary it, and a failure found there becomes a new case when you build a golden dataset from real traces. The wider question of how you know an AI agent is working needs both halves, and the agent evaluation guide maps them.

What are the limits of this gate?

This gate has four limits, and the first is statistical power. With few flips the exact test rarely reaches 0.05, so a small real regression can land in inconclusive or even pass. A pass means the test found no evidence of harm at this sample size.

Second, the test assumes independent cases. Ten cases cloned from one template tend to flip together and carry less evidence than their count suggests, so count a template once.

Third, I chose the exact binomial because CI suites are small. Bowyer, Aitchison and Ivanova (2025, arXiv:2503.01747v3) argue in their abstract that methods based on the central limit theorem understate uncertainty on small eval sets. Their paper handles paired comparisons with a Bayesian model and never mentions this test, so the choice is my inference from their warning.

Fourth, the gate knows only its cases. The book’s sentence belongs above the dashboard: “a green suite is evidence, never proof.” Exact assertions still work for the deterministic half of the system, which is one part of how to test AI agents.

The one thing to keep

Regression testing for LLM applications holds when the rule is written before the run. Count the cases that flipped on a fixed set, apply one test, allow a third verdict, and fail closed when the answer stays unclear. Keep the triage list beside the build log, so that a red gate starts an investigation.

Chapter 16, “Evaluating Agents”, carries the full argument from the first three runs to the enforced gate (in the full book). The Preface, Chapters 1 and 2 and the glossary are free to read online. The eval sample size calculator sizes the suite behind the gate, or you can see the formats.

Questions readers ask

What is regression testing for LLM applications?
It is rerunning a fixed, graded set of cases on every change to a prompt, a tool definition or the model, and comparing the result with the previous version. Because each run is one sample from a distribution, the comparison works on rates and on the cases that changed outcome, and an enforced rule decides whether the change merges.
How do you test non-deterministic LLM output?
Split the system at the model call. Tools, parsers and loop control get exact assertions, as in any test suite. The model side gets property checks and pass rates over a few repeated runs per case, with the per-case verdict fixed by a rule written before the run, such as a majority of three.
What threshold should an LLM eval gate use?
Use a rule on the cases that flipped between baseline and candidate, plus zero tolerance on safety cases. One workable start, all values illustrative: block when an exact one-sided binomial test on the flips gives p at or below 0.05, pass when the net drop is within one point and few cases broke, and call everything else inconclusive.
How many times should each test case run in CI?
One run is enough where a deterministic check grades a stable step. For agent runs, three runs with a majority verdict is a common starting point, and the book asks only for a few trials. Safety cases are stricter: every run must pass. More runs reduce the noise inside a case and leave the noise between cases untouched.
Should LLM evals run on every pull request?
A smoke subset should run on every change that touches a prompt, a tool definition, the model or the harness code. The full set runs nightly and before a release. That split is the book's own accommodation for the cost of multi-trial suites.

Sources

  1. Evan Miller (2024). Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations (arXiv 2411.00640, v1 read)
  2. Berk Atil et al. (2025). Non-Determinism of 'Deterministic' LLM Settings (preprint, arXiv 2408.04667, v5 read)
  3. Shuyin Ouyang, Jie M. Zhang, Mark Harman, Meng Wang (2024). An Empirical Study of the Non-determinism of ChatGPT in Code Generation (arXiv 2308.02828, v2 read)
  4. National Institute of Standards and Technology (n.d.). McNemar Test, Dataplot Reference Manual
  5. Sam Bowyer, Laurence Aitchison, Desi R. Ivanova (2025). Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints (arXiv 2503.01747, v3; abstract-level claim only)
  6. Anthropic (2026). Demystifying evals for AI agents
  7. Hamel Husain (2024). Your AI Product Needs Evals
  8. denis-snyk, GitHub (2025). Issue 5847 in an open-source eval framework's tracker: asserting a per-test pass rate over repeats (7 October 2025)
  9. Abelo9996, GitHub (2026). Issue 5599 in an open-source eval framework's tracker: judge verdicts that flip across repeats (28 September 2026)
  10. mg, Hacker News (2024). Hacker News comment on two different replies at temperature 0 with a fixed seed (10 April 2024)