What does a compounding error calculator show?
A compounding error calculator shows how per-step reliability turns into whole-run reliability. Compounding errors in AI agents are the way small per-step failure rates multiply into large whole-run failure rates. An agent is a chain of dependent steps, and the chain succeeds only if every step does. At 95 percent reliability per step, a twenty-step task succeeds about 36 percent of the time, because 0.95 multiplied by itself twenty times is roughly 0.36.
Chapter 2 of the book, which is free to read, puts it in one line: “Suppose each step succeeds with probability p, independently, and a task needs n steps; the whole run then succeeds with probability roughly pⁿ.” Its second example is harsher. At 92 percent per step over thirty steps, the run succeeds about 8 percent of the time. “The per-step number would have earned a bonus; the end-to-end number is a coin you would never bet on. Exponentials do not negotiate.”
Why real agents fail even faster than this formula predicts is the subject of the article Compounding errors in AI agents. The calculator puts your own numbers into that formula and answers the question that follows from it: how reliable does each step need to be for the run to be reliable?
Why does 95 percent per step turn into 36 percent per run?
Ninety-five percent per step turns into 36 percent per run because probabilities of independent successes multiply. Each step that must succeed shaves five percent off whatever chance remained. After one step, 95 percent of runs are still alive; after two, 90 percent; after ten, 60 percent; after twenty, 36 percent.
Reliability engineers have known this for a long time as Lusser’s law: a system of components in series is only as reliable as the product of their reliabilities. The book’s contribution is to apply it to an agent’s loop, where the “components” are model calls and each one reasons over the output of the last. The consequence is a curve that looks gentle for the first few steps and then falls hard:
| Per-step success | 5 steps | 10 steps | 20 steps | 30 steps |
|---|---|---|---|---|
| 99% | 95% | 90% | 82% | 74% |
| 95% | 77% | 60% | 36% | 21% |
| 92% | 66% | 43% | 19% | 8% |
| 90% | 59% | 35% | 12% | 4% |
The calculator draws the same curve for your inputs, marks a coin flip at 50 percent, and draws your target as a line so you can see where the two cross.
How reliable does each step need to be?
Each step needs to be reliable enough that its reliability raised to the number of steps still meets your target. To succeed 90 percent of the time over twenty steps, each step must succeed about 99.47 percent of the time. The tool solves that equation for you as the twentieth root of the target, and solves the inverse too: at your current per-step reliability, the longest chain that still meets the target.
The second number is often the more useful one. At 95 percent per step and a 90 percent target, the longest acceptable chain is two steps. That is a design constraint, not a statistic: it tells you how much of the task the model can be trusted to chain on its own before something else has to step in.
This is the arithmetic behind the book’s first bearing, the compass of Chapter 1, which puts verifiability first. A long chain with no checks is a bet on an exponent. A chain whose steps are checked as they go is a different system with different odds.
Is the real failure rate worse than pⁿ?
The real failure rate is usually worse than pⁿ, because the formula assumes independence and agents do not fail independently. The book is explicit: “the clean multiplication is the optimistic version.” An agent’s errors land in its own transcript, where they become context for every later step. “A run that has started to go wrong is, from that point forward, reasoning over a partially wrong desk, and its per-step error rate drifts upward as the run lengthens.”
Researchers measuring long tasks report exactly that signature. Sinha and colleagues, in a 2025 study of long-horizon execution, found per-step accuracy declining as a model’s own earlier errors accumulated in its context, an effect they call self-conditioning. The same mechanism produces what the book calls hallucination snowballing: one wrong fact, once written down, becomes the premise of the next step. The hallucination entry covers the factual version.
The levers panel lets you add an illustrative drift: each step loses a fixed amount of reliability relative to the one before. It is a toy model, not a measurement, and it is there to make one point visible. Even a small drift bends the curve down faster than the clean exponential does.
Which levers improve whole-run reliability?
Three levers improve whole-run reliability, and the book names them together: “Shrink n,” “Raise the effective p,” and “cut the price of dying.” The levers panel shows the effect of the first two on your numbers, and the third explains what the panel cannot show.
- Shrink n. Fuse steps, precompute what can be precomputed, and “push bulk mechanical work into ordinary code, which does not roll dice.” Halving a twenty-step chain at 95 percent per step lifts the run from 36 to 60 percent without touching the model.
- Raise the effective p. “Verify at each step and recover on failure, so that an error costs a retry instead of the run.” If a real check catches a failed step and the step is retried once, a 95 percent step behaves like a 99.75 percent step, and the twenty-step run succeeds about 95 percent of the time. The catch is the word real: the check has to come from outside the model’s own judgment.
- Cut the price of dying. Checkpoint the run so that “when a run does fail it resumes from the last good state instead of from zero.” A checkpoint does not change the odds; it changes what a failure costs.
The book’s summary of the whole section is worth keeping on a card: “Per-step accuracy is a ceiling, not a forecast.”
How should you use this before building an agent?
Use it to price autonomy before you build it, while changing the design is still cheap. Write down the steps your agent will need, estimate a per-step success rate from a handful of trial runs, and see what the chain gives you. If the answer is a coin flip or worse, the design needs fewer steps, checks between them, or a human at the points where a failure would cost the most.
The estimate also tells you what to measure. A per-step success rate is a guess until you run the steps many times, and a whole-run success rate is a guess until you run the whole task many times. The pass@k and passk statistics describe the second measurement, and the eval sample-size calculator tells you how many runs you need before either number means something.
Questions readers ask
- If every step is 95% reliable, why does a 20-step run fail almost two times in three?
- Because the run succeeds only if every step does. Multiplying 0.95 by itself twenty times gives about 0.36, so the whole run succeeds roughly 36 percent of the time.
- Is the pⁿ estimate pessimistic?
- No, it is the optimistic case. It assumes each step fails independently, but an agent's errors land in its own transcript and condition later steps, so real per-step reliability tends to drift down as a run gets longer.
- Which lever helps most: fewer steps, better steps, or checkpoints?
- Fewer steps and verified retries change the odds; checkpoints change the cost of losing. Cutting a chain in half or checking each step can move a coin flip to a near-certainty, while a checkpoint lets a failed run resume instead of restarting.
- Does a model checking its own work count as a verified retry?
- Not reliably. The retry lever assumes something outside the model can tell a failed step from a good one, such as a test, a schema check or a comparison with ground truth.
Sources
- Sinha et al. (2025). The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- Wikipedia. Lusser's law
- Anthropic (2024). Building effective agents