To debug an AI agent, stop re-running it and work from a recording. Freeze the failing run, replay it through your own code, measure how often the failure recurs across N live runs, read forward to the first wrong step, decide who wrote that step, change one thing, and keep the run as a test.
This post is about how to debug an AI agent when you hold exactly one bad run and a ticket. It assumes the run was recorded; what a recording has to contain is the subject of the companion post on agent observability, and I don’t repeat its tables here. The material comes from Chapter 15, “Observability and Debugging,” in the sections “Reproducing Nondeterministic Failures” and “A Taxonomy of Common Agent Bugs.” The ordering into six steps and the arithmetic are mine.
Why can’t you debug an agent the way you debug a service?
A service bug reproduces when you recreate its inputs, and an agent’s inputs include things you never set: the model’s sampling, live tool responses and the clock. Run the same request again and those three change, so the failure may simply not appear.
The complaint is easy to find in the reader’s own words. One engineer on Hacker News described a run that “fails on step 17, then 41, then step 9,” and added: “now you can’t reproduce it because it’s probabilistic” (wayy, November 2025). Another asked how to debug workflows where “the final output is incorrect even though nothing technically fails,” with “no runtime errors - just wrong results” (terryjiang2020, February 2026).
The first instinct is to switch the randomness off. One essay on replay calls that instinct, setting temperature to zero and calling the system deterministic, “a myth that has cost teams countless debugging hours” (Tian Pan, April 2026). The book’s version, in “Reproducing Nondeterministic Failures,” is the sentence this whole post leans on: “You cannot reproduce a failure by recreating its inputs, because the inputs were never fully yours to set.”
So the answer to how to debug an AI agent replaces “recreate” with “record,” and it replaces “it reproduces” with a number.
How to debug an AI agent: what are the six steps?
The six steps are freeze, reproduce, bisect, classify, fix and keep, and each one ends with something written down. The table gives the output of each step, because a step with no written output is the one that gets skipped under pressure.
| Step | What you do | What you write down |
|---|---|---|
| 1. Freeze | Pull the recording of the failing run and confirm it is complete | Run id; what is missing, if anything |
| 2. Reproduce | Replay the recording through your code; then run the input live N times | Same wrong output, yes or no; failures out of N |
| 3. Bisect | Read forward to the first wrong step | Step number and kind |
| 4. Classify | Decide who wrote the first wrong bytes | One of four classes, and the test that decided it |
| 5. Fix | Change one thing; re-ask the step, then re-run the task | Failures out of N, before and after |
| 6. Keep | Turn the run into tests | The exact test and the sampled case |
- Freeze: the failing run’s recording is saved somewhere it will not expire, and nobody has “just tried it again” as the investigation.
- Reproduce, part one: the recording replayed through the current code produces the same wrong output, or the step where it diverges is noted.
- Reproduce, part two: the same input has run live N times against frozen tool data, and the result is written as failures out of N.
- Bisect: the first wrong step has a number, and every step before it has been checked against something other than the model’s own words.
- Classify: the fault has one of four names (harness, tool, context, model) and a sentence saying which test decided it.
- Fix: one change was made, and the first wrong step was re-asked N times before and after it.
- Fix: the whole task ran N times after the change, with N chosen from the failure rate measured in step two.
- Keep: an exact test covers the code that changed, with no model in the loop.
- Keep: the failing input, its frozen tool data and a written check are stored as a sampled case with a threshold.
One run carries the rest of the post, as a worked example with illustrative numbers. An internal agent answers warehouse questions. A planner asks, “Can we ship 40 units of KB-204 from Lisbon this week?” and the agent replies, “Yes, Lisbon has 52 units.”
Lisbon had 52 units on the shelf and 40 of them were already allocated to another order. Twelve could ship.
Step 1: What do you freeze before touching anything?
Freeze the recording of the failing run: every model request and response, every tool call and result, the clock and generated identifiers, and the versions of code, prompt and model. A fresh run is a different run, and it can overwrite the only evidence you have.
The book describes the result of recording as a change in what you hold: “Record once, and the snowflake becomes a specimen.” A specimen can be examined as many times as the diagnosis takes. The check here is short: can you read the full text the model was sent at each step, and the full result each tool returned? If the answer is no, write down what is missing. That gap is your first finding, and the observability post’s replay checklist tells you what to add.
In the warehouse run, the recording has nine steps: three tool calls with their results, a drafting call, and the model outputs between them. Nothing is missing.
Step 2: What does “reproduce” mean when every run differs?
Reproducing an agent failure means two separate results: the recording replayed through your code gives the same wrong output, and the same input run live N times fails at a rate you can state. The first result is a yes or a no. The second is a fraction with its denominator.
Replay serves every model response and tool result from the recording while your own code runs again. It tests the parts you wrote. One property makes it trustworthy: when the code asks for something outside the recording, the engine “fails loudly rather than silently falling through to a live system” (Tian Pan, in the essay cited above). In the warehouse run, replay produces the same reply, “Yes, Lisbon has 52 units.” That says the recording is complete and the code is doing today what it did then.
Live repetition answers a different question: how often does this input go wrong? The model is live, and the tools answer from a frozen snapshot of the data as it stood at the time of the failure, so that the stock level cannot move between runs. The warehouse input, run 20 times, gives the wrong answer in 6. The reproduction is “6 of 20, or 30%.”
Published advice differs here. One vendor’s field guide answers the question of a bug that “only happens 30% of the time” with “Set temperature=0 for the replay” (Respan, 2026). That re-ask is quick and worth doing, and the guide itself grants that it may not reproduce the failure exactly. My addition is that it yields one sample of a setting production may not use. The rate is what tells you how many users are affected and how many clean runs a fix will need. A lab’s guide to agent evals states the general principle: “Because model outputs vary between runs, we run multiple trials to produce more consistent results” (Anthropic, January 2026).
How many runs is enough?
Enough that luck stops being a likely explanation, and the count depends on the failure rate. This is the post’s own arithmetic, from standard probability. If a failure appears in a fraction f of runs and nothing has changed, the chance of n clean runs in a row is (1 − f) raised to n.
| Failure rate before the fix | Clean runs needed for that chance to fall below 5% | Chance of that many clean runs by luck |
|---|---|---|
| 50% | 5 | 3.1% |
| 30% | 9 | 4.0% |
| 10% | 29 | 4.7% |
| 5% | 59 | 4.9% |
| 1% | 299 | 4.95% |
Two cautions keep the table honest. The rate before the fix is an estimate too: 6 failures in 20 runs is compatible with a true rate anywhere from 14.5% to 51.9% (a 95% Wilson interval). At the low end of that range, the count needed is 20 clean runs, and 9 is no longer enough. And clean runs never prove zero. Twenty clean runs out of twenty put the failure rate below 13.9% with 95% confidence, by the exact one-sided bound, and no lower.
The arithmetic is one multiplication repeated. An unfixed agent with a 30% failure rate passes a single run 70% of the time, so nine clean runs in a row have a chance of 0.7 to the ninth power, which is 4.0%. The pass@k calculator gives the same figure as passk when the per-attempt success rate is set to 70% and k to 9. The formula assumes the runs are independent, and live runs of one input share whatever makes that input hard.
Step 3: How do you bisect to the first wrong step?
Read the recording forward from the first step and stop at the earliest one whose output is false, unsupported or mishandled. That step is the bug. The book states the rule in “A Taxonomy of Common Agent Bugs”: “read the trace forward to the first step where something is wrong.” It gives the reason too: “because that first step is the bug, and everything after it is consequence.”
Checking a step means comparing it with something outside the model’s own account. A tool result is compared with the system of record. A model output is compared with the prompt it was given: every claim in it should trace to the user’s text or to an earlier result. A harness action is compared with what the code is specified to do.
Two shortcuts help. If you have a passing run of the same input, and after step two you have fourteen, compare the two recordings and stop at the first difference that matters. One guide to production debugging calls this finding “the earliest divergence point between successful and failing runs” (Latitude, 2026). If the run is long, halve it: pick the middle step and ask whether everything shown and said so far is true. A 200-step run takes at most 8 such checks. Halving assumes that an error persists once it is in the context, which holds often and not always, so confirm the step you land on by reading its neighbors.
In the warehouse run, the symptom is at step 9, the reply. Step 4, the stock lookup, returned {"on_hand": 52, "allocated": 40}, and both numbers match the warehouse database. Step 7 is a model output with a working note: “Lisbon has 52 units, enough for 40.” Step 7 is the first wrong step, two steps before anyone saw anything.
Step 4: Which of four faults is it?
Classify the fault by who wrote the first wrong bytes: the harness, a tool, the context assembly or the model. The four names are this post’s sorting, built on the book’s observation that most agent bugs “live in the harness, the interfaces, and the context rather than inside the model.”
Start from the kind of step you landed on. A harness action and a tool result each have one answer. A model output has three candidates, checked in a fixed order, and the model comes last.
| First wrong step is | Question | If yes, the class is | The test that decides |
|---|---|---|---|
| A harness action (parsing, executing, retrying, stopping) | Did your code mishandle a correct model output or tool result? | Harness | Replay reproduces it every time |
| A tool result | Is the result false against the system of record? | Tool | Compare the result with the source |
| A model output | Was something the model needed missing, stale, cut or contradicted in the prompt for this step? | Context | Read the rendered prompt |
| A model output | Was the prompt complete, while a tool’s name, description, schema or result left the meaning open? | Tool | Re-ask with only that wording changed |
| A model output | Was the input complete, true and unambiguous? | Model | Re-ask N times; the failures persist |
The class names the writer. It does not always name where the fix goes. A context fault is repaired in the code that assembles the prompt. A model fault is repaired around the model, with a check in the harness, a tighter instruction or a different model, because the model’s weights aren’t yours to edit.
The warehouse run lands on a model output. The prompt for step 7 contains the step 4 result in full, so the context is clean. The result is true, and yet nothing in it or in the tool’s description says which number can be promised to a customer. That is the second model-output row: a tool fault. The deciding test is a re-ask of step 7 from its recorded prompt, 20 times: 7 wrong. Then 20 more with one change, the result rewritten to carry "available_to_ship": 12: 0 wrong.
Re-ask from the recorded prompt and never from the user’s question. The Respan guide gives the reason: “The user’s query went through your context construction code, which is part of what produced the bug. Skip that layer when reproducing.”
How do four other failures sort?
Each of these four gets exactly one class from the table, which is the test of whether the table is usable.
A parser drops a valid tool call because the model added a sentence after it. The first wrong step is a harness action, and replay reproduces it every time: harness.
A user asks a follow-up and the agent answers as if the first turn never happened. The first wrong step is a model output, and its prompt lacks the first turn: context. The book says of this family, “these are systems bugs, not model bugs, and a trace catches them in the act.”
An agent calls the same failing tool 14 times. The first wrong step is the second identical call. Its prompt is complete, and the error it was shown doesn’t say whether retrying can help: tool. The missing repeat counter in the loop is a second finding for the report, and the class stays the same.
An agent declares the task done with half of it finished, on a prompt that listed every item plainly, in 3 of 20 re-asks. Nothing upstream is wrong: model. The fix is a completion check in the harness.
When the first wrong step is a call to the wrong tool, the diagnosis has its own post: why agents call the wrong tool. For matching a symptom to one of the book’s six named shapes, the agent bug bestiary asks the questions in order.
Step 5: How do you fix it and know the fix worked?
Change one thing, re-ask the first wrong step N times, then run the whole task N times, and report both counts. The book’s instruction for the tight loop is to “change one variable at a time,” and the two counts answer different questions: whether the step is repaired, and whether the task is.
The warehouse fix moves the subtraction into code. The stock tool now returns available_to_ship, and its description says that this is the number to promise against. The step-level result is already in hand: 7 of 20 wrong before, 0 of 20 after. The full task, run live 20 times against the same snapshot, fails 0 times.
Twenty was chosen on purpose. It is the count the same formula gives for the low end of the measured rate, 14.5%. The honest wording of the result is “0 of 20, down from 6 of 20; the failure rate is now below 13.9% at 95% confidence.”
A changed prompt or tool result also changes what replay can tell you. A recorded model response is an answer to the recorded prompt, so replay with frozen responses can’t test a new tool result. It still tests code: if the fix had been in a parser, replay would be the right check, and it would be exact.
Step 6: How does the run become a regression test?
The run becomes two tests, one on each side of what the book calls the deterministic–nondeterministic seam. The book says it in five words: “recorded runs are regression tests.” Which kind of test depends on which side of the seam the fix sits.
The first test is ordinary. It asserts, with no model anywhere, that the stock tool returns available_to_ship equal to on-hand minus allocated. It runs in milliseconds on every commit.
The second is a sampled case. It stores the planner’s question, the frozen snapshot and a written check: the reply must not confirm a shipment larger than available_to_ship. It runs N times and compares the pass rate with a threshold. One essay on testing agents puts the instruction this way: “run the case N times and assert on the pass rate, not on a single result” (Saurav Bhattacharya, 2026).
Choose the threshold knowing what it costs. A gate that demands 20 passes of 20 will fail about one build in three when the true failure rate is 2% (the pass chance is 66.8%). Either accept that and re-run, or raise N and allow a stated number of failures.
This check can be written in code because the answer contains a number. When the property is fuzzier, such as whether a summary is faithful to its source, the check needs a judge, and a planned post on LLM-as-a-judge versus human evaluation covers how far to trust one. The wider method, from single cases to a suite, belongs to a planned post on how to test AI agents. A case promoted this way joins the eval set.
What does the finished bug report look like?
It is one page with a count on every line where a count is possible. The template below is this post’s own; the filled values for the warehouse run follow it.
AGENT BUG REPORT
Run id / link to recording:
Symptom, in the reporter's words:
Expected, and the source that says so:
1. FREEZE
Recording complete (model calls, tool results, clock and ids, versions): yes / no
Missing:
2. REPRODUCE
Replay through current code gives the same wrong output: yes / no / diverged at step __
Live runs, same input, frozen tool data: __ failed of __
Check used to call a run failed:
3. BISECT
First wrong step: #__ kind: harness action / tool result / model output
Earlier steps were checked against:
4. CLASSIFY
Class: harness / tool / context / model
Test that decided it:
Other findings (not the class):
5. FIX
The one change:
First wrong step re-asked: __ wrong of __ before, __ wrong of __ after
Full task after the change: __ failed of __
Stated bound on the remaining failure rate:
6. KEEP
Exact test (no model):
Sampled case (input, frozen data, check, N, threshold):
Not explained:
| Line | Warehouse run (illustrative) |
|---|---|
| Replay | Same wrong output |
| Live runs | 6 failed of 20 |
| First wrong step | #7, model output |
| Class | Tool: the result did not say which quantity can be promised |
| Deciding test | Re-ask with only the result changed: 7 of 20 wrong, then 0 of 20 |
| Full task after the change | 0 failed of 20; rate below 13.9% at 95% confidence |
| Exact test | Stock tool returns on-hand minus allocated |
| Sampled case | Reply never confirms more than available_to_ship; 20 runs |
| Not explained | Why the model read 52 as available in 7 of 20 re-asks and not in the other 13 |
The last line stays on the page for a reason. The replay essay is plain about the limit: “It doesn’t explain why the model said what it said.”
Where does this procedure break?
This account of how to debug an AI agent breaks when the recording is incomplete, when tools can’t be frozen, when the failure is rare, and when there is no check to count with.
The replay essay names the first limit: “It can’t replay what it didn’t record.” A run that failed before recording existed can only be approached by live repetition, which may never show the failure. The fix for that is the next recording.
Live repetition needs frozen tool data. An agent that writes to real systems, sends messages or reads sources that change by the minute needs stubs or a sandboxed copy first, and building those can take longer than the diagnosis.
Rare failures are expensive. At a 1% rate the table asks for 299 clean runs, and each run of a long task costs model calls and minutes. For those, fix the step, re-ask the step many times, and let production monitoring carry the rest of the evidence.
The four classes are a sorting aid. Two faults can coincide, as in the retry loop above, and the class of the first one does not excuse the second. The Respan guide reports that five bug shapes cover “roughly 90%” of what its team sees; that is one team’s count from its own practice, and I’d expect a new kind of agent to produce runs that fit no row.
Recordings hold users’ data in full. Keeping a failing run forever as a test means keeping its content, so scrub or synthesize the inputs before a case enters a suite that many people can read.
The one thing to keep
A failure you can state as “6 of 20” is one you can fix and then show fixed, and that is most of how to debug an AI agent. The book ends its list of bug shapes on the reassurance behind this procedure: “Debugging an agent is, most days, debugging the system you built around the model.” Three of the four classes are code and text you own, and the fourth is repaired from outside.
The record-and-replay material and the six bug shapes are in Chapter 15, Observability and Debugging (in the full book); measuring pass rates properly is Chapter 16, Evaluating Agents (in the full book). The guide to evaluating and observing agents collects the related posts and tools, and you can see the formats.
Questions readers ask
- Can you make an AI agent deterministic by setting temperature to zero?
- No. A lower temperature narrows the variation and does not remove it, and hosted inference is not fully deterministic even at zero: one essay on replay traces the remaining variation to request batching and floating-point arithmetic. In a multi-step run a small difference at one step changes every later step. Record the run and replay it instead of trying to recreate it.
- How many times should you re-run a failing agent task?
- Enough to state a rate with a denominator. Twenty runs is a workable start: it separates a failure that happens half the time from one that happens one time in twenty. To accept a fix, the number of clean runs depends on the rate before it: 9 for a 30% failure, 29 for 10%, 59 for 5%.
- What is the first wrong step in an agent run?
- It is the earliest step whose output is false, unsupported by what came before it, or mishandled by your code. It is rarely the final answer. Read the recording forward from the start, or compare it with a passing run and stop at the first difference that matters.
- Is the bug in the model or in my code?
- Check your code first. If the first wrong step is something the harness did, or a tool returned something false or ambiguous, or the prompt for that step was missing what the model needed, the fault is in a part you own. Call it a model fault only when the input was complete, true and unambiguous.
- How do you turn an agent bug into a regression test?
- Keep two tests. The first is an ordinary exact test on the code you changed, such as a tool's output or a parser, with no model in the loop. The second is the failing input with frozen tool data and a written check, run N times and compared with a pass-rate threshold.
Sources
- Tian Pan (2026). Deterministic Replay: How to Debug AI Agents That Never Run the Same Way Twice
- Respan (2026). AI Agent Debugging
- Latitude (2026). The complete guide to debugging AI agents in production
- Saurav Bhattacharya (2026). Stop Asserting Equality: How to Test Agents When Every Run Is Different
- Anthropic (2026). Demystifying evals for AI agents
- terryjiang2020 and replies, Hacker News (2026). Ask HN: How do you debug multi-step AI workflows when the output is wrong?
- wayy, Hacker News (2025). Hacker News comment on debugging agents that fail at a different step each run