What is an agent bug bestiary?
An agent bug bestiary is a field guide to the failures agents actually produce, sorted so you can start from what you see and end at a fix. Each card gives four things: the symptom, the cause, the clue that shows in a trace, and the fix. The core of this one is the taxonomy in Chapter 15 of the book, which calls itself “the bestiary, six specimens long.” Around those six, the tool collects the other failures the book names chapter by chapter, from the loop that forgets its own results to the approver who stops reading.
The premise is that agent failures repeat. “What looked, in the first week, like an inexhaustible variety of weirdness resolves into a short list of shapes wearing different clothes,” the chapter says. One field guide puts a number on the compression, five shapes covering “roughly 90% of what you will see.” The book offers that percentage as one team’s experience; the shortness of the list is the part that generalizes. The AI agent failure modes post walks through the same taxonomy at essay length. This page is the working tool you keep open beside a trace.
What are the six specimens?
The six specimens are the stuck loop, hallucinated tool arguments, lost and poisoned context, the swallowed error with its cousin silent truncation, the wrong stop condition, and silent degradation across deploys. Each has a signature you can find with one query or one reading.
| Specimen | What you see | Where the trace shows it |
|---|---|---|
| The stuck loop | Forty tool calls where healthy runs take nine | The same tool called many times in a row |
| Hallucinated tool arguments | A customer ID, date or table from nowhere | An argument with no source in the user’s text or prior results |
| Lost and poisoned context | A forgotten turn, or an early wrong fact treated as truth | The failing step’s rendered prompt |
| The swallowed error, and silent truncation | A confident wrong answer with a clean status | A clean tool call followed by a clean reply; result byte counts |
| The wrong stop condition | A run that never stops, or one that stops too soon | Step counts pinned at the cap, or early exits with the goal unfinished |
| Silent degradation across deploys | Worked in May, worse in June, no code changed | Requested vs responding model, prompt version, prompt hash |
The swallowed error is the one the chapter most wants you to fear, “because it is the one that ships wrong answers with a straight face.” Its signature is the absence of a signature: “There is no error, just a clean tool call followed by a clean assistant reply,” and the user gets an answer in the same tone as a right one. That is why the wizard asks about result sizes and cut-off outputs even when every status looks green.
How does the wizard rank suspects?
The wizard asks the questions Chapter 15 teaches you to ask of a trace, in the order you would open it: the session summary first (how many steps, how it ended), then the tool spans, the arguments, the rendered prompt, the result sizes, and last the versions recorded on every span. Each card lists the answers that point at it. A “yes” counts for a card, a “no” counts against it, and the suspects are ranked by the total, with Chapter 15’s specimens first when two cards tie.
The ranking is the tool’s arithmetic, not the book’s. Its job is to send you to the right card quickly. The book’s method then takes over, and it never varies: “read the trace forward to the first step where something is wrong—not the last step, where the wrongness became visible—because that first step is the bug, and everything after it is consequence.” With a good trace and a bad one for the same input, the reading becomes a diff.
Most questions need attributes your traces may not carry yet. The rendered prompt matters most. “The rendered prompt is what the model actually saw. The template plus variables is what you intended. Bugs live in the difference.” If you cannot answer a question, leave it as “not checked” and treat that as the first finding: the trace is missing something it should record.
Why do most fixes point at the harness?
Most fixes point at the harness because that is where most agent bugs live. Look at where Chapter 15’s fixes land: tool descriptions, error surfaces, context assembly, stopping logic, version pins. “Chapter after chapter, the treatment lived in the parts you own.” The model appears in the bestiary “mostly as an amplifier, the machinery that turns your ambiguous description into an invented argument and your silent cut into a confident lie.”
The cards from other chapters follow the same pattern. The loop that requests the same call forever usually has a dropped append in the harness: the result never made it back into the history, so the model, which remembers nothing between passes, asks again. The duplicate record after a retry is a tool that was never made idempotent. The approver who waves through the one dangerous action is a gate placed on too many routine ones. “Debugging an agent is, most days, debugging the system you built around the model, and that is the good news, because that system is the part you can fix before lunch.”
The layer tag on each card (engine, harness, context, tools, multi-agent, operations, security, evaluation) is the tool’s way of grouping them. Filter by layer when you already suspect where the problem is, and by chapter when you want the full treatment of one area.
How does a fixed bug become an eval case?
A fixed bug becomes an eval case when you keep the failing trace instead of closing the ticket. Chapter 15 calls the failing trace “two assets wearing one file format.” It is the reproduction, the recorded run you replay against the fix and keep replaying so this failure can never ship quietly again. And it is the seed of an eval case: “the input, the observed wrong behavior, and your judgment of what right looks like.”
The notes panel above writes that stub for you. Pick the card that matches, describe the run in three fields, and copy the result into your eval set or the ticket. The chapter’s arithmetic for the habit is blunt: skip it “and you will fix failures one at a time forever, an arrangement your bug supply can sustain indefinitely.” Adopt it, read a few dozen real runs, sort what went wrong into named categories, and turn each distinct failure into a task. “Do that for a quarter and you own something no afternoon of imagination could have written: a measuring instrument built from your agent’s actual failures rather than your feared ones.”
What does the bestiary not do?
The bestiary does not read your traces. It cannot see a missing turn or a truncated result; it can only tell you where to look and what the book says to do about what you find there. The ranking counts your answers, so a wrong answer gives a wrong suspect, and many real incidents are two bugs at once: a hallucinated argument that a swallowed error then hid, for example. The cards are the starting set the chapter hands you, not a closed list. When a failure fits none of them, name it yourself.
Nothing here replaces instrumentation. If you are still deciding what to record, Chapter 15 is the place to start, and the compounding error calculator shows why small per-step failure rates matter so much over a long run. For the security cards, the lethal trifecta audit checks whether a design is exposed in the first place. For the evaluation cards, the eval sample size calculator and the judge agreement calculator cover the measurement side. The full treatment of traces, monitoring and reproduction is in Chapter 15, Observability and Debugging (in the full book).
Questions readers ask
- Why do most fixes point at the harness and not the model?
- Because most agent bugs live in the parts you built around the model: tool descriptions, error surfaces, context assembly, stopping logic and version pins. Chapter 15 describes the model in this bestiary as mostly an amplifier, the machinery that turns an ambiguous description into an invented argument and a silent cut into a confident lie. That is good news, since the system around the model is the part you can fix.
- What is the one reading method the book uses for every bug?
- Read the trace forward to the first step where something is wrong, not the last step where the damage became visible. That first step is the bug and everything after it is consequence. With one good trace and one bad trace for the same input, the reading becomes a diff: find the first step where the two disagree.
- How does a fixed bug become an eval case?
- The failing trace is two assets in one file. It is the reproduction you replay against the fix and keep in your replay library, and it is the seed of an eval case: the input, the observed wrong behavior, and your judgment of what right looks like. Chapter 15 calls the habit of promoting every resolved failure into a test the flywheel.
- What do I need in my traces for the wizard to work?
- The attributes Chapter 15 asks every span to carry: the rendered prompt rather than the template, every tool call's arguments and result, the result's size, tokens and latency, the requested and the responding model, the prompt version and a hash of the rendered prompt, plus a session summary with steps taken and final status. Without them most questions can only be answered “not checked”.
- Is this list complete?
- No. Chapter 15 reports one field guide's estimate that five shapes cover roughly 90% of what you will see, and treats the number as one team's experience rather than a constant. The shortness of the list is the general lesson. When a failure fits no card, name it, and add it to the categories you sort your own traces into.
Sources
- Respan (2026). AI Agent Debugging
- Gabriel Anhaia (2026). Tool-Result Truncation: The Silent Bug That Makes Agents Lie
- Braintrust (2026). Best AI agent debugging tools