Home / Blog / Security, reliability and cost / AI Agent Failure Modes: A Field Taxonomy

Security, reliability and cost

AI Agent Failure Modes: A Field Taxonomy

AI agent failure modes sorted for diagnosis: symptom, cause, trace signature and fix for each, cross-checked against MAST. Keep it open beside your traces.

By Enrique Gutiérrez · Published · 15 min read

AI agent failure modes are the recurring ways an agent run goes wrong, and there are fewer of them than a week of incidents suggests. Six shapes cover most of what a production trace will show you: the stuck loop, hallucinated tool arguments, lost and poisoned context, the swallowed error, the wrong stop condition, and silent degradation across deploys.

This post is a field reference. Each failure gets a symptom, a likely cause, what to look for in the trace, the fix, and the chapter of AI Agents, Engineered where the fix is worked out. The bestiary comes from Chapter 15, “Observability and Debugging” (in the full book), and the responses from Chapter 18, “Reliability, State, and the Harness” (in the full book). Where published research taxonomies add something, I cross-reference them, chiefly the MAST study of more than 1,600 annotated multi-agent traces.

What are the AI agent failure modes worth naming?

The AI agent failure modes worth naming are the ones that repeat, leave a distinct signature in a trace, and point to a distinct fix. Chapter 15 calls its list “the bestiary, six specimens long,” and the names below are the book’s own. A name is useful only if it tells you where to look next.

Chapter 15 is candid about where the list comes from. One field guide it draws on claims that five shapes cover “roughly 90% of what you will see,” and the book treats that number as one team’s report rather than a constant: “The percentage is one team’s experience; the shortness of the list is everyone’s.” I agree with the second half more than the first. The exact share varies by system; the fact that you keep meeting the same six animals does not.

Naming failures has a second payoff: the name routes the fix. The chapter says so directly: “The taxonomy is a diagnostic index to the book.” Almost every fix points back to tool descriptions, error surfaces, context assembly, stopping logic or version pins, the parts of the system you own.

The field taxonomy at a glance

The table below is the whole reference in one place: read the symptom, confirm it against the trace signature, then apply the fix. The first six rows are Chapter 15’s bestiary. The last three are failures Chapter 18 treats as reliability problems, which show up in the same traces and are worth classifying with the same discipline.

Failure mode Symptom Likely cause What to look for in the trace Fix Chapter
The stuck loop A run with forty tool calls where healthy runs take nine A tool returned an error or empty result, and nothing told the model retrying was hopeless Group tool spans by name per session; flag long runs of consecutive identical calls; read the first repeated span A structured retryable field on errors; harness counts identical calls and intervenes after the third; step and cost budgets underneath 15, 18; design in 3, 5
Hallucinated tool arguments An argument “from nowhere”: an ID the user never gave, a future date, a table that does not exist An ambiguous description or a permissive free-string field; the model fills the gap plausibly Diff each argument value against the user’s text and prior span results; a value with no source was invented Tighter types and enums; say when not to use the tool; “If you do not know X, do not guess. Ask the user”; trigger and near-miss tests 15; design in 5
Lost and poisoned context The agent forgets an earlier turn, or treats an early wrong fact as ground truth Truncation, a code path that forwards only the last message, over-eager compaction, or a stale record entering early Open the failing step’s rendered prompt and read what the model actually saw: the turn is missing, or the poison is visible Context assembly fixes: compaction that keeps what matters, trim-and-say-so returns 15; theory in 7
The swallowed error, and its cousin silent truncation A confident wrong answer; the right fact was in the tool result A layer between tool and model ate an error or cut the result, and nothing marked the cut A clean tool call followed by a clean reply; result size far below that tool’s norm, or an empty result where data should exist Surface errors as structured observations; truncate-and-say-so; record result byte count on every span; paginate at source 15, 18
The wrong stop condition Either never stops, or stops early with a half-finished answer presented as done (false completion) Stopping logic in the harness that is wrong or untested Step counts pinned at the cap, or early exits with the goal’s checklist visibly unfinished Explicit stop conditions; test the loop with a model stub that says “done” too early or never 15; design in 3
Silent degradation across deploys Quality sags; no code changed A model alias rolled forward, a prompt edit, or the data behind a prompt variable moved Diff a good trace and a bad one with the same input shape; find the first step where they disagree; compare responding model, prompt version, prompt hash Pin models explicitly; version prompts in source control; record both on every span; alert on week-over-week trends 15
Retry storm One outage becomes many; costs and latency spike Immediate, unbounded or nested retries across layers Bursts of retries per failure across client, tool, step and run layers; same dependency failing repeatedly Backoff with jitter; retry only transient errors; a global retry budget; a circuit breaker 18
Duplicated side effect after resume The email goes out twice; the card is charged twice A timeout on a write (the ambiguous failure), or a resume from a checkpoint taken before the action was recorded A write call with no matching receipt, then the same logical write again after a crash or retry Idempotency keys; record intent, execute, record a receipt; sagas with compensating actions 18
Silent partial failure A half-completed transaction reported as success A failed action-taking step absorbed “gracefully” instead of halting A failed write step followed by a success message, with no escalation span Fall back on reads, fail loud on writes; escalate with goal, attempts, error and a trace link 18

How do you read a trace to classify a failure?

You read the trace forward from the start and stop at the first step where something is wrong, then classify that step. Chapter 15 puts the rule plainly: “read the trace forward to the first step where something is wrong—not the last step, where the wrongness became visible—because that first step is the bug, and everything after it is consequence.”

One agent run drawn as a span waterfall.
Figure 15.2 One agent run drawn as a span waterfall. Each bar is a span: where it starts is when it happened, how long it runs is how long it took, and the indentation of its row records who called whom—the tool spans nest under the model call that asked for them, and everything nests under the root run. Every span carries the same record, the attributes bracketed over the first model call: the rendered prompt, the tokens, the cost, the latency. The bug is not the final answer but the first wrong span—here the long database call that erred, in accent—and a waterfall makes it the thing your eye lands on. The timings and ordering are illustrative. Reuse this diagram

The rule matters because AI agent failure modes chain. Picture a hypothetical run: a swallowed error at step four can become a stuck loop at step twelve and a false completion at step thirty. If you classify by the last visible symptom, you will fix the loop (lower the budget) and leave the cause in place.

The research literature has the same intuition: Zhu and colleagues describe agents as vulnerable to “cascading failures, where a single root-cause error propagates through subsequent decisions,” and build their debugging method around isolating that root cause. The arithmetic of why a single early error does so much damage is the subject of why compounding errors make small agent mistakes snowball, so I won’t repeat it here.

Two pieces of trace hygiene make the reading possible. Record the rendered prompt of each model call, so you can see what the model actually saw. And attach a session summary span with steps, cost and final status, so that “show me yesterday’s worst runs” is one query. If you have neither yet, the step-by-step version of this reading is in how to debug an AI agent from its traces.

Worked example: is it a loop or a swallowed error?

The run below is illustrative, built for this post rather than taken from a real incident, but every line in it is a shape the table above describes. A support agent is asked whether order 4471 can still be refunded. Try classifying it before you read my reading.

step  span                         result (abridged)
1     llm: plan                    "look up order, then check refund policy"
2     tool: get_order(id=4471)     {"items": [...], "status": "deliv     <- 4,096 bytes; median for this tool: 11,200
3     llm: reason                  "order delivered 12 days ago; within 30-day window"
4     tool: get_policy(sku=...)    {"error": "timeout"}
5     tool: get_policy(sku=...)    {"error": "timeout"}
...   (same call, 14 more times)
20    llm: answer                  "Yes, the order is eligible for a full refund."
summary: 20 steps (p50 for this task: 6), status=success

The symptom you notice first is the stuck loop: sixteen identical get_policy calls. The fix for that row is real and worth doing (mark the timeout as retryable with a limit, intervene after the third identical call). But the loop is not the first wrong step.

Step two is. The order record arrived at 4,096 bytes when that tool normally returns about 11 KB, and the JSON stops mid-field: silent truncation.

The part that was cut held the line saying the order was a final-sale item. Step three reasoned over the half it received, exactly as Chapter 15 warns: “This is the class I most want you to fear, because it is the one that ships wrong answers with a straight face.” And step twenty is a false completion: the policy check never succeeded, yet the run reports success.

So this one run contains three AI agent failure modes. Step two is the first wrong step and the cause of the wrong answer; the timeout loop is a second, independent fault, and the false completion is the stop logic failing to notice it. The classification for the incident ticket is “swallowed error / silent truncation, step 2.” The loop and the false completion get their own fixes, but the eval case you write from this trace should reproduce step two.

What does each failure look like up close?

Each failure has one detail that distinguishes it from its neighbors, and that detail is what you should check before you commit to a classification. The notes below add that detail to the table and stay close to how Chapter 15 describes each specimen.

Why does an agent keep looping on the same call?

An agent keeps looping because a tool returned an error or an empty result and nothing in the result told the model that trying again was hopeless. The cause is “almost always visible in the first repeated span.” The book’s fix has three layers: make retryability a structured field, add a harness backstop that notices repeated identical calls, and keep budgets underneath as the last line. The long version, including how to tell a productive retry from a stall, is in what to do when an AI agent keeps looping.

Where do hallucinated tool arguments come from?

Hallucinated tool arguments come from gaps in the tool interface that the model fills with something plausible. The test is mechanical: every argument value should trace back to the user’s text or a prior span’s result. Enumerated types cannot be hallucinated outside their enum, which is why the book calls these fixes “all interface work” rather than model work. This is where AI agent hallucination does the most practical damage.

How do lost context and poisoned context differ?

Lost context means a needed turn is missing from what the model saw; poisoned context means a wrong fact is present and treated as true. The book groups them because one diagnosis covers both: read the rendered prompt of the failing step. Both are related to context rot, and the chapter’s verdict is the useful one: “these are systems bugs, not model bugs, and a trace catches them in the act.”

What makes a stop condition wrong?

A stop condition is wrong when the harness lets the agent run past the point where it should halt, or halt before the goal is met. The trace separates the two at a glance: step counts pinned at the cap, or early exits with work left. The book’s point is that stopping logic is ordinary harness code, so you can test it with a scripted stub “at zero model cost, in milliseconds, forever.”

How do you catch silent degradation across deploys?

You catch it by diffing one good trace against one bad trace with the same input shape and checking, in order, the responding model, the prompt version and the prompt hash. One of the three almost always differs. The fix is the pinning discipline, “because ‘it worked yesterday’ is only a solvable mystery if yesterday was written down.” Catching it before users do is the job of a staged rollout; see shadow-deploying an agent before a canary release.

Which response does each failure want?

Each failure wants the response that matches its recovery path, and Chapter 18 sorts errors by asking what the failure wants done about it. Transient failures want a pause and a bounded retry. Model-recoverable failures (wrong tool, malformed argument, unparseable output) want to go back to the model as a structured observation, not a blind retry. Permanent failures want to fail fast; policy failures want a loud halt; ambiguous failures, like a timeout on a write, want idempotency before anyone retries.

One tool failure, two fates.
Figure 18.1 One tool failure, two fates. Swallowed into an empty result, it feeds the model a hole it fills with plausibility—a fluent, confident summary of a database never reached. Surfaced instead as a structured observation (in accent) carrying a code, a retryable flag, and a hint, the same failure becomes something the model can read and recover from. The model can only recover from a failure it can see. Reuse this diagram

The common thread is visibility. Chapter 18 boils the section down to one sentence: “the model can only recover from a failure it can see.” That is the swallowed-error row of the table seen from the builder’s side. A structured error with a stable code, a retryable flag and a hint turns a laundered failure back into something the model, or a human, can act on.

The chapter also draws a line I find myself quoting in design reviews: “Fall back on words; fail loud on deeds.” A stale cache with a warning is a fine degraded answer to a question. There is no acceptable degraded version of charging a card. That rule is what separates a graceful fallback from the silent partial failure in the table’s last row. The full playbook of retries, budgets, checkpoints and fallbacks is in how to make AI agents more reliable.

How does this compare with research taxonomies like MAST?

Research taxonomies agree with the bestiary on the big shapes and add finer cuts for multi-agent systems. The best known is MAST, the Multi-Agent System Failure Taxonomy of Cemri and colleagues (2025), built from 150 traces with expert annotators (inter-annotator agreement κ = 0.88) and applied to a dataset of more than 1,600 traces across seven frameworks.

MAST has 14 modes in three categories: system design issues, inter-agent misalignment, and task verification. Several map directly onto the bestiary.

Its most frequent mode, step repetition (FM-1.3, 15.7% of observed failures in the MAST dataset), is the stuck loop. Loss of conversation history (FM-1.4, 2.8%) is lost context. Unawareness of termination conditions (FM-1.5, 12.4%) and premature termination (FM-3.1, 6.2%) are the two halves of the wrong stop condition. No or incomplete verification (8.2%) and incorrect verification (9.1%) are how false completions get through.

What MAST adds is the coordination layer: conversation reset, failure to ask for clarification, task derailment, information withholding, ignoring another agent’s input, and reasoning-action mismatch (13.2%, the second most common mode). If you run several agents, those deserve their own rows. The study reports failure rates of 41% to 86.7% on the seven open-source systems it examined, a snapshot of those systems at that time rather than a property of multi-agent designs in general.

Two other taxonomies are worth knowing by name. The Microsoft AI Red Team’s whitepaper on agentic failure modes (2025) takes the security view and calls cross-domain prompt injection “potentially the most significant failure mode for agentic AI systems”; that is poisoned context with an adversary behind it. And an empirical study of 385 faults drawn from 40 agent repositories (Shah et al., 2026) found root causes such as data schema mismatches, dependency drift and state management complexity, which reads like a list of harness problems.

Is the failure in the model or in the harness?

Most of the time the failure is in the harness, the interfaces or the context, and the model is the amplifier. Chapter 15 ends its bestiary on exactly this: “Debugging an agent is, most days, debugging the system you built around the model, and that is the good news, because that system is the part you can fix before lunch.”

The research points the same way. The MAST authors conclude that “MAS failure is not merely a function of challenges in the underlying model; a well-designed MAS can result in performance gain when using the same underlying model,” and report that clarifying agent roles alone raised one system’s task success by 9.4%. A 2026 taxonomy by Raj and colleagues makes the question part of the label: each of its 41 modes carries “a fault side indicating where the repair belongs,” model, harness, or environment.

I’d adopt that habit even without the paper. When you file a failure, write down which side owns the fix. If the answer is “the model” more than occasionally, check again; in my reading of the bestiary, only hallucinated arguments sit close to the model, and even there the fix is a tighter interface.

Where does this taxonomy stop?

This taxonomy stops at failures you can see in a trace, and some important ones you cannot. A run that is wrong in a way no step reveals (a plausible answer to the wrong question, a subtle factual error with no source to diff against) needs evaluation, not observability. Chapter 15 borrows a line for this boundary: monitoring without evaluation “is like watching a plane’s instruments without knowing where it should be flying.”

The list is also a starter set, not a closed one. Your system will produce a failure with no name yet, and the right response is to name it, add a row, and promote the trace to your eval set. The book’s version of the habit is to “read a few dozen real runs, sort what went wrong into named categories (this section just handed you the starter set), and turn each distinct failure into a task in your eval set.” And if you want to see how quickly unchecked steps erode reliability before any of these failures appear, the compounding error calculator puts numbers on it; the agent loop explainer shows where in each turn these failures enter.

The one thing to keep

Classify by the first wrong step, and write down which side owns the fix. Six names and three reliability rows will cover most of your incidents, and the table is meant to sit open beside your trace viewer while you work. Every failure you classify and promote to an eval case is one more signal you can verify the agent against.

The bestiary in full, with the tracing, monitoring and replay machinery behind it, is in Chapter 15, “Observability and Debugging” (in the full book); the responses are in Chapter 18 (in the full book). Chapters 1 and 2 are free to read online, and you can see the formats.

Questions readers ask

What are the most common AI agent failure modes?
Six shapes cover most production incidents: the stuck loop, hallucinated tool arguments, lost or poisoned context, the swallowed error (and silent truncation of tool results), the wrong stop condition including false completion, and silent degradation after a model, prompt or data change.
Why do AI agents fail more often than single LLM calls?
An agent chains many dependent model decisions, and each one reads the results of the earlier ones. A small error early in the run becomes context for every later step, so failures compound instead of staying isolated, and most of them come from the harness and interfaces around the model.
How do I tell which failure mode caused an incident?
Open the trace and read forward from the start to the first step whose input or output is wrong. Match that step against the trace signatures: repeated identical tool spans, an argument with no source, a missing turn in the rendered prompt, a truncated or empty result, an early or capped exit, or a version change between a good and a bad run.
Is hallucination the main reason AI agents fail?
It is the mechanism behind several failure modes, but rarely the root cause. A model fills gaps plausibly, so an ambiguous tool description, a swallowed error or a truncated result turns into a confident wrong answer. Fixing the gap in the interface usually fixes the hallucination.
Do multi-agent systems have different failure modes?
They inherit the single-agent modes and add coordination failures. The MAST taxonomy (Cemri et al., 2025) lists inter-agent misalignment modes such as conversation reset, information withholding and ignoring another agent's input, alongside verification failures.

Sources

  1. Mert Cemri, Melissa Z. Pan, Shuyi Yang, et al. (2025). Why Do Multi-Agent LLM Systems Fail?
  2. Kunlun Zhu, Zijia Liu, Bingxuan Li, et al. (2025). Where LLM Agents Fail and How They can Learn From Failures
  3. Harsh Raj, Vipul Gupta, Anas Mahmoud, et al. (2026). Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
  4. Mehil B. Shah, Mohammad Mehdi Morovati, Mohammad Masudur Rahman, Foutse Khomh (2026). Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes
  5. Pete Bryan, Giorgio Severi, et al. (Microsoft AI Red Team) (2025). Taxonomy of Failure Mode in Agentic AI Systems (whitepaper)