Home / Blog / Evaluating and observing agents / How to Test AI Agents When Every Run Is Different

Evaluating and observing agents

How to Test AI Agents When Every Run Is Different

How to test AI agents: exact tests for tools, parsers and stop rules, pass rates on a fixed world for the model side, and a flaky-case procedure. Copy it.

By Enrique Gutiérrez · Published · 20 min read

The short answer to how to test AI agents is to split the suite at the model call. The code you wrote (tools, parsers, stop rules, the loop) keeps ordinary tests with exact assertions. The model’s output gets checks on properties and end state, run several times on a fixed world and reported as a pass rate.

This guide is the map for an engineer who knows unit tests and has watched asserts on agent output go red with no change in the diff. It shows where each check belongs, what a case looks like, how a production failure becomes one, and what to do with a case that flakes. The statistics behind gates and sample sizes live in the specialised posts, and I link to them instead of deriving them again.

Why do asserts on agent output flake?

Asserts on agent output flake because they demand one value from a component that produces a distribution. Chapter 16 of the book puts the mismatch in one sentence: “The practice of running fixed tests over deterministic code assumes the code is the behavior.” With an agent, the behavior sits in a model that samples.

An essay by Saurav Bhattacharya (2026) follows the decay of one such test: “It passes on Tuesday. It fails on Wednesday because the model reworded one sentence.” A trim follows, then a regular expression, and at the end the test is deleted. On Hacker News in April 2026 the user BeetleB described a suite that asks a model which tool it would call and asserts on the reply: “The tests are flaky.”

Ordinary software has a name for this. John Micco’s 2016 post on the Google Testing Blog defines it: “We define a ‘flaky’ test result as a test that exhibits both a passing and a failing result with the same code.” In that company’s corpus of conventional tests, the post reports that about 1.5% of all test runs were flaky. Those figures describe deterministic code in 2016 and say nothing about agents.

For an agent, the definition changes meaning. A check on live model output that passes and fails with the same code may be reporting the truth: the agent succeeds some of the time. The work is to tell that case apart from a badly placed assertion.

Which parts of an agent still get ordinary tests?

Every part you wrote still gets ordinary tests, and that is the first half of how to test AI agents: the tools, the parsers that read model output, the stop rules, the retry logic and the code that assembles the next prompt. Chapter 15 of the book, in the section “Reproducing Nondeterministic Failures,” calls the line between the two halves the “deterministic–nondeterministic seam.”

Of the side you wrote, the chapter says: “That side is ordinary software, and it deserves ordinary software’s tests, exact assertions and all.” Two instructions follow: “Unit-test each tool with normal and hostile inputs, no model anywhere. Feed your parsers the malformed, truncated, half-JSON output a model will eventually produce.”

One question places any check on the correct side. I call it the placement question, and it is this post’s own wording: would two correct runs produce the same bytes at this point? If yes, assert exactly and run the test once. If no, you are on the model’s side of the seam.

Part of the agent Same bytes on two correct runs? Kind of test Model in the test
A tool’s function Yes Exact assertions on normal, hostile and error inputs None
Parser of model output Yes Exact assertions on valid, truncated and malformed replies None
Stop rules and budgets Yes Exact assertions on when the loop ends and with what status Scripted stub
Retry and error handling Yes Exact assertions on what is retried and what is surfaced Scripted stub
Context assembly Yes Exact or snapshot assertion on the rendered prompt None
Which tool the model picks No Property check, several runs, a pass rule Live
The final answer’s wording No Property check or a calibrated judge, several runs Live
The state the run leaves behind No End-state check, several runs, a pass rule Live
A never-event (a forbidden action) No Check on the recorded tool calls, every run must pass Live

A tool with an unreliable network dependency is still on the exact side. Its flakiness is the ordinary kind, and the fix is ordinary too; the post on idempotent tools and safe retries covers the contract such a tool should keep.

How do you test the loop without a live model?

You test the loop with a scripted model: a stub that returns a fixed sequence of replies while the test asserts on what your harness did. The book describes the stub as returning “a fixed sequence of responses, including the error cases and edge cases that are awkward to coax from a live model on demand.”

The assertion targets are in the same passage: the harness “stops when it should, retries what it should, surfaces what it must.” A script that never emits a final answer proves the step budget fires. A script that returns a malformed tool call proves the parser’s error reaches the model as a readable message. Each of the stop conditions for agent loops is control logic, so each one gets a test like this, which the chapter calls “fast, free, and fully deterministic.”

A recording is a scripted model with real content. In replay, the chapter says, your own code runs again while every external call is served from the recording.

Record and replay across the deterministic–nondeterministic seam (the dashed line): your code on the left, the model and the world on the right.
Figure 15.4 Record and replay across the deterministic–nondeterministic seam (the dashed line): your code on the left, the model and the world on the right. In record mode the agent runs live, and an interception layer journals every crossing of the seam into the recording. In replay mode your own code runs again—live and debuggable—but every external call is served from that recording; the live model and tools (drawn as ghosts) are never called, so a failure can be re-run as many times as the diagnosis needs, at no cost and with no side effects. The two journals, in accent, are what make the snowflake a specimen. Reuse this diagram

The book’s advice is to keep a library of golden traces, “recorded runs whose behavior you consider correct,” and replay them against every new version of your code. The post on how to debug an AI agent covers the recording itself. The chapter credits golden traces with catching the prompt regression, an edit that helps one scenario and breaks others. My caveat for testing: the recorded replies answered the old prompt. A replay can show that your code now sends a different prompt, because a sound replay engine, in words the chapter quotes, “fails loudly rather than silently falling through to a live system.” It cannot say how the model will answer the new one; that takes a live run.

What should you mock, and what must stay live?

Mock the model to test your harness, and keep a thin live layer to test the model against your contract. A stub proves that your code handles a given reply. Whether the live model still produces that reply is a separate question, and no stub can answer it.

An Ask HN post from January 2026 shows the cost of forgetting this. The user tom1337 wrote: “At the moment, all LLM calls are mocked in unit and end-to-end tests. Recently, we updated a model version and it started returning responses that no longer matched our expected schema”, and then: “Because all tests used mocked responses, this only surfaced in production.”

The check that incident needed is small. Send each prompt to the real model, verify that the reply parses against the schema your parser expects, repeat a few times, and run the set whenever a prompt, a tool definition or the model version changes. This is a property check: it reads the shape of the reply and ignores the wording. The post on structured output for tool calls explains why the shape can still break.

The same split applies to tools. Stub a tool’s backend to keep the world fixed, and keep the tool’s contract (its name, parameters and result format) identical to production. A test where the agent sees a simplified tool definition measures a different agent.

What do you assert on the model’s side of the seam?

On the model’s side you assert properties and end state, and you stop asserting text; this is the half of how to test AI agents that unit-test habits do not cover. The book’s instruction is four words long: “assert properties, not strings.” A property is something true of every acceptable output whatever its wording.

Three kinds of check do the work, listed from cheapest to dearest.

Properties of the output. The reply parses. It names the customer ID that was in the input and no other. It stays under a length limit. Code decides each of these.

End state. Chapter 16 says to “grade the state, not the prose,” and gives the reason: “The transcript can claim success while the world is unchanged, and only one of them is your product.” If the task was to schedule a meeting, query the calendar fixture after the run. If it was to fix a bug, run the tests.

Never-events. The chapter asks you to “include negative cases: tasks asserting what the agent should not do.” Its example is an agent asked to deactivate an account that deletes it and reports the account “no longer active.” An output check gives that run full marks. A check on the recorded tool calls catches it.

Resist asserting the full path. The book’s default is to “grade the outcome, and stay out of the agent’s route,” and it reserves path checks for efficiency and safety. The post on agent trajectory evaluation shows how to write those two.

Some outputs have no property that code can decide, such as whether a summary is faithful. Those go to a model-based judge, with the caution in the book’s line “the best judge is the one you did not need.” The post on whether an LLM judge is reliable covers calibrating one.

Why run each case several times, and how do you report it?

You run each case several times because one run is one sample, and you report the count with its denominator: 17 of 20, never “passes.” Chapter 16 asks for “a handful at minimum, a couple of dozen when the stakes justify it,” and adds that “any figure is illustrative.”

Anthropic’s engineering guide (2026) states the same practice: “Because model outputs vary between runs, we run multiple trials to produce more consistent results.” Hamel Husain’s 2024 essay removes the expectation that comes with unit tests: “unlike traditional unit tests, you don’t necessarily need a 100% pass rate. Your pass rate is a product decision, depending on the failures you are willing to tolerate.”

A worked example with illustrative numbers shows why a single run misleads. Suppose a suite has 20 cases and the agent truly passes each one 90% of the time. Run once per case and require all green: the build is green with probability 0.9 raised to the power 20, which is 0.12. About seven builds in eight go red with no change in the code, assuming the cases fail independently.

Repeats help and do not rescue an all-green rule. With three runs per case and two passes required, one such case passes 97.2% of the time, and all 20 pass together 56.7% of the time. That is why a suite on the model’s side is read as a rate against a baseline, which is the subject of the post on regression testing for LLM applications.

What does 17 of 20 mean for a pass rule?

A case that passes 17 of 20 runs behaves very differently under different pass rules, and the calculator below opens on that case. A single run is green 85% of the time.

With JavaScript on, the pass@k and passk calculator runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the pass@k and passk calculator on its own page to share a result by link.

The tool shows 99.9% for at least one pass in three runs and 59.6% for three passes in three. The second number is the count of ways to draw three passing runs from the 17, divided by the ways to draw any three from the 20: 680 / 1,140. So a strict rule that needs every one of three runs goes green about 60% of the time on this case.

A two-of-three rule is more forgiving. At 85% per run it passes 93.9% of the time (my arithmetic from the same rate; the tool does not display it). The rule you choose decides what “flaky” will mean in your CI, so write it into the case before the first run. The post on pass@k and passk owns the two statistics.

Twenty runs give a rough rate with a wide interval around it. How wide, and how many cases a suite needs, is worked out in how many eval examples you need.

What do you hold fixed: the model or the world?

Hold the world fixed, and treat the model as the one thing you measure. Sampling settings and seeds on a hosted model narrow the variation and give no guarantee. The book says “even with the randomness dialed to zero a hosted model does not promise identical replies,” and the regression post cites two papers that measured it.

The world is yours to freeze. The book’s list of what a run silently consumes is “clock reads, random seeds, generated identifiers,” with a warning: “Time is the one everyone forgets.” A harness that stamps today’s date into the prompt runs a different test every day.

The items below extend that list; the additions are mine.

  • Clock. Inject one fixed time per case.
  • Identifiers and randomness you own. Seed them, so that two runs create the same record IDs.
  • Data. Load the same fixture (database rows, files, inbox) before every run.
  • Tool backends. Serve search results and third-party responses from fixtures, including the errors.
  • Isolation. Start every run from a clean environment.
  • Versions. Pin and record the model identifier, the prompt revision and the tool-definition revision with each result.

Isolation has outside support. Anthropic’s guide says: “Each trial should be ‘isolated’ by starting from a clean environment. Unnecessary shared state between runs (leftover files, cached data, resource exhaustion) can cause correlated failures due to infrastructure flakiness rather than agent performance.”

With the world fixed, any variation left in the results belongs to the model or the grader.

How to test AI agents in layers: what runs, cheapest first?

An agent test suite has seven layers in this post’s arrangement, and the first three make no model calls at all. Most of the practical answer to how to test AI agents is deciding which layer a check belongs to. The ordering follows an old rule. Ham Vocke’s 2018 article on the test pyramid puts it this way: “Write lots of small and fast unit tests. Write some more coarse-grained tests and very few high-level tests that test your application from end to end.”

Layer What runs Model calls What it asserts When it runs
1. Unit tests A tool or a parser alone 0 Exact values and errors Every commit
2. Loop tests Harness with a scripted stub 0 Exact: stops, retries, surfaced errors Every commit
3. Replay Harness against recorded runs 0 Structure: same tools, same key facts, same conclusion Every commit
4. Live single step One real model call per run 1 per run Properties: parses, picks an allowed tool On a change to a prompt, a tool definition or the model
5. Live end to end, graded by code The whole agent on a fixed world Many per run, times the runs End state and never-events, as a rate Smoke subset per change, full set nightly
6. Live end to end, graded by a judge The same, plus judge calls The most A rubric verdict from a calibrated judge Nightly and before a release
7. Production sampling Real traffic None extra for the agent Online signals and sampled grading Continuously

Cadence follows cost. Husain’s essay describes three levels of his own and says: “I often run Level 1 evals on every code change, Level 2 on a set cadence and Level 3 only after significant product changes.” The book’s version is “a smoke subset per commit and the full battery nightly and before release.”

The cost gap is large. With illustrative inputs of 40 smoke cases, 3 runs each and 12 model calls per run, layer 5 costs 40 × 3 × 12 = 1,440 model calls per change. Layers 1 to 3 cost none. So push every check down to the lowest layer that can decide it.

Layer 7 belongs to the post on monitoring AI agents in production. Which numbers to report from layers 5 and 6 is the subject of AI agent evaluation metrics.

How do you turn a production failure into a test case?

You turn a production failure into two tests: an exact test for the code that changed, and a sampled case that reproduces the input on a frozen world. Chapter 16 names the cycle: “traces reveal failures, failures become tasks, tasks guard against regression.”

A user on Hacker News described life without it in May 2026: “teams fix a failure in an agent, change the prompt or model a week later, and the same failure quietly comes back. Nobody catches it until a user does.”

The steps, in order:

  1. Keep the trace. It holds the input, the tool results and the versions. The post on LLM tracing with OpenTelemetry lists what a trace must record for this to work.
  2. Find the first wrong step and decide which side of the seam it sits on.
  3. Write the exact test if the fault was in a tool, a parser or the harness.
  4. Write the sampled case: the input, the world fixture built from the trace, a check on the end state, the never-events and a reference solution. The book requires the last item: “every task must be solvable, and the way to know is to include a reference solution that proves it.”
  5. Run the case against the old version. It should fail at a stated rate. A case that the broken version passes has not captured the failure. This step is my addition.
  6. Fix, rerun, and file the case in the regression set once the fixed version passes it reliably.

Strip personal data from the fixture before it enters the repository. A trace is user data.

What does one case look like on paper?

One case fits in a short block of plain text that states the input, the frozen world, the checks, the runs and the pass rule. The run counts are the illustrative starting values from the regression post: three runs, two passes for an ordinary case, and three of three for a never-event.

AGENT TEST CASE  (numbers are starting values; set your own)

ID            : <short-name>
SOURCE        : <trace id or ticket>   | written from imagination: yes/no
SIDE OF SEAM  : exact (same bytes on two correct runs)  |  sampled

INPUT
  <the user request or task, verbatim>

WORLD (frozen before every run)
  clock        : <fixed timestamp>
  ids / random : <seed>
  data fixture : <name and revision>
  tool backends: <stubbed from fixture | live sandbox>
  environment  : clean per run

VERSIONS RECORDED WITH EACH RESULT
  model identifier, prompt revision, tool-definition revision

CHECKS
  properties   : <e.g. reply parses; names only IDs present in the input>
  end state    : <e.g. calendar fixture holds exactly one event at 14:00>
  never-events : <e.g. no call to delete_*; no message to outside addresses>
  grader       : code  |  calibrated judge (only if code cannot decide)

REFERENCE SOLUTION
  <a known output or action sequence that passes every check>

RUNS AND PASS RULE (written before the first run)
  exact case   : 1 run, must pass
  sampled case : 3 runs, passes if at least 2 pass
  never-events : checked on every run; one violation fails the case
  infra failure: separate status, retried, never counted as a fail

BASELINE RATE
  <passes> of <runs> on <date>, version <...>

OWNER         : <name>

What do you do with a flaky case?

You find the cause of a flaky case before you change it, because a case that flips can have seven causes and only the last is the agent. The procedure below is this post’s own. It is compatible with the flaky list in the regression post, which defines a flaky case as one that flips between identical baseline runs.

  • Infrastructure? A timeout, a rate limit or a crashed sandbox is not a failed case. The book says to keep these “out of the quality number entirely, tracked under a separate status.”
  • Wrong side of the seam? Ask the placement question of each assertion. An exact match on model text gets rewritten as a property or an end-state check.
  • Exact-side test that flakes? Then it is an ordinary flaky test. Look for the clock, test order, shared state or a live network call, and fix that.
  • World not fixed? Compare two runs’ fixtures, timestamps and identifiers. Any difference there is yours to remove.
  • Measure the rate. Run the case 20 times on the fixed world and write down passes of 20. Until this step, “flaky” is an impression.
  • Grader noise? Store one output and grade it several times. Different verdicts on the same output mean the grader needs repair first.
  • Unfair failures? Read the failing transcripts. The book’s test: “If a failing case leaves you confused about what was wanted, the case is the bug.” Sharpen the criterion or fix the reference solution.
  • None of the above? The agent fails this case at the measured rate. Record the rate, then either fix the agent or accept the rate with a pass rule that states it.
  • Never-event case? It is never excused as flaky. One violation is a finding about the agent.

Two responses are missing from the list on purpose. Retrying until green turns a reliability check into a check on luck: a case that passes half the time goes green at least once in three tries 87.5% of the time. Loosening a check until it always passes removes the case while leaving its name in the report.

A case that fails for a real reason and that nobody plans to fix can still earn its place, as a capability case with a recorded rate. The book warns about the alternative, a suite of low-signal checks that “trains the team to ignore red,” with the instruction: “Prune it the way you would prune flaky tests, and for the same reason.”

How do three example cases come out?

Three invented cases, none used above, pass through the placement question, the template and the checklist with one answer each.

A due-date tool that fails on the last day of the month. Two correct runs return the same bytes, so this is an exact case with one run and no model. The checklist stops at the third item: the test read the system clock. Inject a fixed clock, add the month-end input as a second exact test, and no rate is involved.

“Agent stops after eight steps,” asserted against a live model. The test flakes because the model sometimes finishes in five. The stop rule is control logic, so the placement question sends it to layer 2: a scripted stub that never returns a final answer, and an exact assertion that the harness halts at step eight with a budget-exhausted status. One run.

“Move my 2 p.m. meeting to Thursday,” which passes 17 of 20. The wording varies, so the case is sampled. The world is a calendar fixture with a fixed clock. The end-state check reads the fixture: one event on Thursday, none in the old slot. The never-event is a message to an outside attendee. Checklist items one to four, six and seven come back clean, and item five gives the rate: 17 of 20. The three failing runs left the old event in place without sending any message. So item eight applies: the agent fails this case about 15% of the time. Under the two-of-three rule the case would pass 93.9% of the time, which hides a failure a user would meet roughly one time in seven. The rate goes into the case as its baseline, and the failing transcripts go to whoever owns the prompt.

Where does this approach stop working?

This way of answering how to test AI agents has four limits, and the first is the stub. A scripted model and a recording both describe how the model behaved once. They test your code and go stale when the model changes, which is why layer 4 exists.

Second, properties are necessary and not sufficient. A reply can parse, cite the right ID, stay under the length limit and still be wrong. Property checks catch broken outputs cheaply; they do not certify good ones.

Third, a fixed world is narrower than production. Freezing the clock and the data removes noise and also removes the odd inputs that real users supply. The offline suite, in the book’s words, holds “the failures you have already imagined.”

Fourth, repeated runs cost money in proportion to their number. The chapter is direct about it: “Multi-trial evaluation multiplies your compute bill by the trial count.” Below some scale, a small suite of believed cases read carefully beats a large one run blindly.

One sentence from Chapter 16 belongs above every dashboard: “a green suite is evidence, never proof.”

The one thing to keep

Knowing how to test AI agents comes down to placing each check. Ask whether two correct runs would produce the same bytes. Where they would, write the test you have always written and run it once. Where they would not, freeze the world, check what the run left behind, run it several times and report passes over runs, under a rule you wrote first.

The seam and the scripted model are in Chapter 15, “Observability and Debugging”; the eval set, the grading ladder and the regression gate are in Chapter 16, “Evaluating Agents” (both in the full book). The Preface, Chapters 1 and 2 and the glossary are free to read online. The guide to evaluating and observing agents collects the rest of this cluster, and the pass@k calculator is there when you have your own counts, or you can see the formats.

Questions readers ask

How do you test an AI agent?
Split the system at the model call. Tools, parsers and loop control get ordinary unit tests with exact assertions and no model. The model side gets cases that run on a fixed world several times each, are graded on properties and on the end state the agent left behind, and are reported as a pass rate under a rule written before the run.
Can you unit test an AI agent?
You can unit test every part you wrote: each tool, each parser of model output, the stop conditions, the retry logic and the code that assembles the next prompt. Drive the loop with a scripted stub in place of the model. What a unit test cannot do is assert the exact text a live model returns.
Should you mock the LLM in agent tests?
Mock it to test your own harness, and never mock it to test the model's behavior. A stub proves that the loop handles a given reply. It says nothing about whether the live model still produces that reply, so keep a thin layer of live checks that runs whenever the prompt, a tool definition or the model version changes.
Does temperature zero make agent tests deterministic?
No. Chapter 16 of the book says that “even with the randomness dialed to zero a hosted model does not promise identical replies.” Lower randomness narrows the variation. The dependable move is to freeze everything around the model (clock, identifiers, data, tool results) and measure what is left as a rate.
What should you do with a flaky agent test?
Find its cause before changing it. Check for an infrastructure failure, an exact assertion placed on model output, an unfixed clock or shared state, a noisy grader and an ambiguous criterion, in that order. If none applies, the agent fails that case some of the time, and the rate is a finding to record and act on.

Sources

  1. Anthropic (2026). Demystifying evals for AI agents
  2. Hamel Husain (2024). Your AI Product Needs Evals
  3. Saurav Bhattacharya (2026). Stop Asserting Equality: How to Test Agents When Every Run Is Different
  4. John Micco (2016). Flaky Tests at Google and How We Mitigate Them (Google Testing Blog, 27 May 2016)
  5. Ham Vocke (2018). The Practical Test Pyramid (martinfowler.com, 26 February 2018)
  6. tom1337, Hacker News (2026). Ask HN: How do you integration-test AI / LLMs? (6 January 2026)
  7. thiht and replies, Hacker News (2024). Ask HN: What's the consensus on “unit” testing LLM prompts? (20 July 2024)
  8. BeetleB, Hacker News (2026). Hacker News comment on flaky tool-choice tests (10 April 2026)
  9. 1taimoorkhan0, Hacker News (2026). How do you catch AI agent regressions after prompt or model changes? (28 May 2026)