Learn / Explainers / Why errors compound: 95% per step, 36% per run

Why errors compound: 95% per step, 36% per run

Animated explainer · 3:07 · from Chapter 2

Space plays and pauses · ← → step · F fullscreen

A three-minute animated explainer: a hundred runs fall away down a chain of 95% steps, why pⁿ is the optimistic case, and the three levers that fight back.

Transcript

Every step of the animation, in the words it uses on screen. Select a line to jump to it.

From Chapter 2

The Engine: How Language Models Work

This explainer condenses a few pages of Chapter 2. The chapter itself traces the ideas through a worked example and is free to read online.

Questions

Does a better model fix compounding error?
It raises p, but it does not remove the exponent. At 95% per step a twenty-step run finishes about 36% of the time; even 99% per step leaves about 82% across twenty steps and about 74% across thirty. A better model buys you longer chains before the curve bites, not immunity. Try your own numbers in the compounding-error calculator.
Why is independence the optimistic assumption?
Because pⁿ assumes each step fails on its own. In an agent, a mistake is written into the transcript and becomes context for every later step, so the per-step error rate drifts upward as a run lengthens. Researchers measuring long tasks report reliability decaying faster than independence predicts. The evidence, and why retries help less than the arithmetic suggests, is in why agent errors compound.
Which lever does the book reach for first?
Shrinking n is the cheapest: every step you fuse, precompute, or move into ordinary code stops rolling dice at all (at 95% per step, eight steps finish about 66% of the time against 36% for twenty). Raising the effective p by verifying each step and recovering on failure comes next, and checkpointing caps what a failure costs. The calculator shows how much each lever buys for your chain.
Is pⁿ ever too pessimistic?
Yes. It treats every error as fatal and uncaught. Steps whose mistakes are harmless, errors the agent notices and repairs, and steps guarded by a real check with a retry all push actual success above pⁿ. That is the second lever at work: checking and recovering raise the effective p above the raw per-step number.