Home / Blog / Agent fundamentals / Compounding Errors in AI Agents: Why Small Mist…

Agent fundamentals

Compounding Errors in AI Agents: Why Small Mistakes Snowball

Compounding errors in AI agents are worse than pⁿ: self-conditioning, snowballing and correlated errors. See the evidence and the redesign that fixes it.

By Enrique Gutiérrez · Published · 12 min read

Compounding errors in AI agents are the reason a step that works 95 percent of the time produces a twenty-step run that fails almost two times in three. The textbook model multiplies per-step success rates, pⁿ, and that model is the optimistic case. Real agents fail worse, because each mistake becomes context that makes the next mistake more likely.

The arithmetic itself is the easy part, and the compounding error calculator will do it for any numbers you give it. This post is about what the arithmetic leaves out: why measured agents fall off the pⁿ curve faster than independence predicts, what the long-horizon studies actually found, and how to redesign a realistic task so that the exponent works for you instead of against you.

What are compounding errors in AI agents?

Compounding errors in AI agents are the multiplication of small per-step failure rates into large whole-run failure rates. An agent runs a loop of dependent steps, each model call reasoning over the results of the calls before it, and the run succeeds only if every step does. Per-step reliability is therefore a ceiling on the run, never a forecast of it.

Chapter 2 of the book, free to read, puts the baseline plainly: “Suppose each step succeeds with probability p, independently, and a task needs n steps; the whole run then succeeds with probability roughly pⁿ.” At 95 percent over twenty steps that is about 36 percent. At 92 percent over thirty steps it is about 8 percent. “The per-step number would have earned a bonus; the end-to-end number is a coin you would never bet on. Exponentials do not negotiate.”

Compounding error, drawn.
Figure 2.6 Compounding error, drawn. Whole-run success is per-step reliability raised to the number of chained steps, so it decays exponentially: 95% per step is about 36% across twenty steps, and 92% per step is about 8% across thirty (the accented points). An excellent per-step number becomes a coin you would never bet on; exponentials do not negotiate. Reuse this diagram

If you have not yet pinned down what counts as a “step,” the post on what an agent loop is and how each turn works draws the cycle. For this post, a step is one model decision whose output the rest of the run depends on: a tool choice, an argument, an edit, a conclusion.

Why do real agents fail worse than pⁿ?

Real agents fail worse than pⁿ because the formula assumes each step fails independently, and an agent’s steps share a transcript. A wrong result does not vanish after it happens; it sits in the context window where every later step reads it. The book’s phrase for this is exact: “the clean multiplication is the optimistic version.”

Three separate mechanisms break the independence assumption, and they are worth keeping apart because each one calls for a different fix:

Mechanism What happens Evidence What it breaks
Self-conditioning Errors already in the history raise the error rate of later steps Sinha et al. (2025), controlled error injection The constant p in pⁿ
Hallucination snowballing The model defends an early false claim with more false claims Zhang et al. (2023) The model’s ability to catch its own mistake in-session
Correlated errors Retries and second models fail in the same way, on the same inputs Kim et al. (2025), 350+ models The independence of retries and checkers

The first two make each run degrade from the inside. The third is quieter and, in my experience, more often overlooked: it makes the obvious remedies (try again, ask another model) weaker than they look on paper.

What is self-conditioning, and how was it measured?

Self-conditioning is the tendency of a language model to make more mistakes when its context already contains its own earlier mistakes. It was isolated by Sinha, Arun, Goel, Staab and Geiping in a 2025 study of long-horizon execution, which handed models the knowledge and the plan for a long, simple task so that any failure had to come from carrying it out.

The finding that matters for builders came from a controlled experiment. The researchers rewrote the model’s history to contain errors at chosen rates, then measured accuracy at a later turn. With a clean history, accuracy at turn 100 was already lower than at turn 1, which they attribute to long-context degradation. But “as we increase the rate of injected errors into the context, accuracy at turn 100 consistently degrades further.” In their words, “models become more likely to make mistakes when the context contains their errors from prior turns.”

Two further results are worth carrying around. Scaling the model largely fixed the long-context part but not the self-conditioning part: “Self-conditioning does not reduce by just scaling the model size.” And models that reason before answering were much less affected; the authors report that thinking “mitigates self-conditioning.” The book compresses the consequence into one sentence: “A confused agent stays confused, and left alone, it gets more so.” That is also the mechanism behind context rot when the rot is the agent’s own output.

Why can’t an agent catch its own snowballing hallucinations?

An agent often can’t catch its own snowballing hallucination because, inside the same session, it is conditioned on the false claim and tends to elaborate on it rather than retract it. Zhang and colleagues named the pattern in 2023: “an LM over-commits to early mistakes, leading to more mistakes that it otherwise would not make.”

The striking part of that study is the second half of the finding. When the false supporting claims were shown to the models separately, the two chat models tested identified 67 and 87 percent of their own mistakes (2023 models; the exact rates will have moved). The knowledge to reject the error was there. What was missing was a context in which the error was not already a premise.

That one result carries most of the design advice in this post. A check run in the same context is a weak check. A check run in a fresh context, or by something that is not a model at all, gets a far better chance at the same error. The post on why AI agents hallucinate in the first place covers the engine side; the point here is what happens after the first wrong token lands.

What do long-horizon studies show about agent reliability?

Long-horizon research finds that agents lose reliability with task length faster than their single-step skill would suggest, and that the gap is about execution rather than knowledge. METR, a research group that measures autonomous AI capabilities, put it this way in 2025: “AI agents often seem to struggle with stringing together longer sequences of actions more than they lack skills or knowledge needed to solve single steps.”

METR’s metric is the time horizon: the length of task, measured in how long it takes a skilled human, that a model completes with a given success rate. In their March 2025 data, models succeeded almost always on tasks taking humans under four minutes and less than 10 percent of the time on tasks over about four hours. Those specific numbers have since moved, and METR itself marks them as dated. What has not moved is the gap between reliability levels: the paper reports that “models’ 80% time horizons are 4-6x shorter” than their 50 percent horizons.

Here is a rough comparison of my own, and I hold it loosely. If failures came at a constant per-step rate, the 80 percent horizon would be about a third of the 50 percent horizon, because ln 0.8 divided by ln 0.5 is about 0.32. A gap of four to six times is wider than that, and it is not what self-conditioning alone would produce; a rising error rate would pull the two horizons closer together. The likelier reading is that METR’s tasks vary in difficulty as well as length, so a single per-step rate can’t describe them. What the gap does say is that “usually works” and “reliably works” are very different lengths of task.

The same paper also notes that agents struggle most in “environments without clear feedback loops,” which is the verification argument seen from the other side.

How many unverified steps can your agent afford?

Your agent can afford roughly ln(1/target) ÷ ε unverified steps, where ε is the per-step error rate. That falls straight out of pⁿ: set pⁿ equal to your target and solve for n. Sinha et al. derive the same shape for the 50 percent point, a horizon of about 0.69 ÷ ε steps.

The rule is worth computing once by hand, because it changes how you think about model upgrades. With independent steps (the optimistic case):

Per-step success Error rate ε Steps to a coin flip (50%) Unverified steps at a 90% target
90% 10% about 7 1
95% 5% about 14 2
99% 1% about 69 10
99.9% 0.1% about 693 105

Two things jump out. First, near the top, small gains in per-step accuracy buy enormous horizon gains, which is the point of Sinha et al.’s title: what looks like diminishing returns on a single-step benchmark is exponential growth in task length.

Second, anywhere in the 90-to-99 percent range, the affordable unverified run is short: one to ten steps, not dozens. That number is the spacing for your checks. A check must fire at least that often, and a stop condition must end the run when the checks keep failing.

When is pⁿ too pessimistic?

pⁿ is too pessimistic when a step’s error is either harmless, caught, or repaired. The formula counts every failure as fatal and invisible. Real runs contain redundant searches whose misses cost nothing, steps the agent notices are wrong and redoes, and steps guarded by a test that rejects bad output and triggers a retry.

The honest picture has errors pushing in both directions:

pⁿ is too optimistic when pⁿ is too pessimistic when
Earlier errors sit in the context (self-conditioning) A step’s mistake doesn’t affect the outcome
The model defends a wrong claim (snowballing) An external check catches the failure and a retry fixes it
Retries fail the same way (correlated errors) The agent observes a failed tool result and recovers
Tasks lack a feedback signal The run restarts from a clean checkpoint, resetting the drift

Notice that everything in the right-hand column is something you build. The left-hand column is what you get by default. That asymmetry is the whole argument for treating reliability as a design property rather than a model property. The field taxonomy of AI agent failure modes names the specific ways a run goes wrong; this table is about which side of the curve each one pushes you.

How do you redesign a task so errors stop compounding?

You redesign a task in three moves, the three levers Chapter 2 names: “Shrink n,” “Raise the effective p,” and “cut the price of dying.” The worked example below applies them in that order to a task a backend engineer might plausibly hand an agent. The numbers are illustrative, chosen to be realistic rather than measured.

The task. Add a nullable region field to an orders API: migration, model, serializer, eight call sites, four test files, docs and a pull request. Done naively, the agent makes about 24 model-driven steps: read the ticket, four separate codebase searches, the migration, model and serializer, eight call-site edits, four test edits, run and interpret the tests, fix, docs, and the PR description. Assume 95 percent per step.

Baseline. 0.95²⁴ ≈ 29 percent. Add a small illustrative drift for self-conditioning (each step 0.2 points worse than the one before) and it falls to about 16 percent. Roughly one run in six lands, and you can’t tell which from the agent’s own report.

Move 1: shrink n. The four searches become one call to a deterministic code-search tool that returns every hit. The eight call-site edits become two steps: the model writes a codemod (a small script that rewrites code mechanically), and ordinary code applies it, “which does not roll dice.” Tests come from one template. The task drops to 10 model-driven steps: 0.95¹⁰ ≈ 60 percent, and about 54 percent with the same drift. Fewer steps also means less transcript for errors to poison.

Move 2: raise the effective p. Put an oracle after every step that has one: apply the migration to a scratch database, type-check, run the unit tests, compile after the codemod. Suppose the checks catch 90 percent of failed steps and a failed step gets one retry. If the retry is an independent draw at 95 percent, each step behaves like a 99.3 percent step and the run reaches about 93 percent. If the retry tends to repeat the same mistake (say it succeeds only 60 percent of the time), the effective step is 97.7 percent and the run reaches about 79 percent.

That gap between 93 and 79 percent is the correlated-errors problem made concrete. The way to close it follows from the snowballing result: retry in a fresh context with the failed attempt summarized as a fact (“the previous migration failed with this error”), not replayed as a transcript to be defended.

Move 3: cut the price of dying. Split the run into three phases (schema, callers, tests and docs) and take a checkpoint after each green gate. Checkpoints don’t change the odds. They change what a failure costs: a broken test file reruns the third phase, not the migration. Each phase also starts with a fresh context, which resets the self-conditioning drift.

Design Model-driven steps Effective per-step success Run success (illustrative)
Naive 24 95% ~29% (~16% with drift)
Shrink n 10 95% ~60% (~54% with drift)
+ verified retry, correlated 10 ~97.7% ~79%
+ verified retry, fresh context 10 ~99.3% ~93%
+ checkpoints per phase 10 same same odds; a failure costs one phase

What are the limits of this model?

The limits are that the inputs are estimates and the checker has its own error rate. Per-step reliability is a guess until you run each step many times, and whole-run reliability is a guess until you run the whole task many times; the pass@k and passk statistics describe that second measurement.

The checks themselves can be wrong, and a model used as a checker inherits the correlated-errors problem: Kim, Garg, Peng and Garg found in 2025 that “larger and more accurate models have highly correlated errors, even with distinct architectures and providers.” A second model reviewing the first is not an independent witness. Where you can, check with something that is not a model: tests, compilers, schemas, counts.

Finally, some tasks have no cheap oracle at all. For those, the full book’s Chapter 13, on writing the outer loop, gives the decision rule: if “done” is not machine-verifiable, wire a human in or build a judge (and calibrate it), and accept the weaker guarantee. Its one-line test applies to every design in this post: “A loop is exactly as trustworthy as its oracle.”

The one thing to keep

Compounding errors in AI agents are not a property of a bad model. They are a property of a long, unverified chain, and the chain is yours to design. Count the model-driven steps, put an outside check after each one that can have one, retry in a fresh context, and checkpoint between phases. As Chapter 2 puts it, “Per-step accuracy is a ceiling, not a forecast.”

The full argument, with the failure catalog behind it, is in Chapter 2, “Limitations and Failure Modes”, free to read. Chapter 13 on the outer loop is in the full book; see the formats.

Questions readers ask

Why do AI agents fail on long tasks when each step looks reliable?
Because whole-run success is roughly the product of per-step success, and in practice per-step reliability falls as the run gets longer. A 95%-reliable step gives a 20-step run about a 36% chance under independence, and less once earlier errors start conditioning later steps.
Is pⁿ ever too pessimistic for an agent?
Yes. It assumes every error is fatal and uncaught. Steps whose mistakes are harmless, steps the agent notices and repairs, and steps guarded by a real check with a retry all push actual reliability above pⁿ.
How many unverified steps can an agent afford?
Roughly the natural log of one over your target success rate, divided by the per-step error rate. For a 90% target that is about 0.105 divided by the error rate: two steps at 95% per step, ten at 99%. A check has to fire at least that often.
Does asking the model to double-check its work stop errors compounding?
Not reliably. In the same session the model is conditioned on its own mistake and tends to defend it. A check helps when it comes from outside that context: a test, a schema, a compiler, or a fresh session that sees only the spec and the output.
Do newer models fix compounding errors?
They raise per-step accuracy, which lengthens the horizon a lot, but the structure stays. Sinha et al. (2025) found that scaling model size did not reduce self-conditioning, though models that reason before acting were much less affected.

Sources

  1. Akshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab, Jonas Geiping (2025). The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
  2. Muru Zhang, Ofir Press, William Merrill, Alisa Liu, Noah A. Smith (2023). How Language Model Hallucinations Can Snowball
  3. Thomas Kwa, Ben West, Joel Becker, et al. (METR) (2025). Measuring AI Ability to Complete Long Software Tasks
  4. METR (2025). Measuring AI Ability to Complete Long Tasks (blog post)
  5. Elliot Myunghoon Kim, Avi Garg, Kenny Peng, Nikhil Garg (2025). Correlated Errors in Large Language Models