What is a retry storm simulator?
A retry storm simulator shows what a group of clients does to a failing dependency under different retry policies, so you can see why the standard defenses exist before you need them. This one runs three small models from Chapter 18 of the book: a crowd retrying, layers of retries multiplying, and a circuit breaker in front of a dependency that goes down.
The chapter introduces the subject with a warning about its size. “The retry looks too small to deserve engineering. It is the most common way agents turn one failure into many.” An agent makes many tool calls per run, often through several layers of software that each retry on their own, so the arithmetic on this page applies to it more than to most programs.
How does a retry storm happen?
A retry storm happens when the clients that an outage affected all retry together and their retries keep the service down. The chapter gives the mechanism step by step: a service is briefly overloaded and refuses requests, every failed client retries immediately, and the service “now facing its normal load plus the entire wave of retries, stays down”. Its summary: “The clients trying to recover from the outage have become the outage”.
The first row of the chart is that case. With the default settings, a hundred clients fail at the same moment and each has five tries. Retrying at once, all four hundred retries arrive in the first tenth of a second, on a dependency that will be down for ten seconds. Every client uses up its tries and gives up, having sent five hundred calls for nothing.
The model is kinder than a real storm in one respect and harsher in another. The outage here has a fixed length, so the retries do not keep the service down longer, which is the heart of the chapter’s mechanism. And calls take no time, so immediate retries all land in the same instant.
Why do backoff and jitter both matter?
Backoff and jitter both matter because they fix different things: backoff gives the dependency time, and jitter stops the clients from arriving together. The chapter defines them in plain terms. “Backoff means waiting longer between successive attempts—exponential backoff doubles the wait each time, giving the struggling dependency room to breathe.”
The second row shows what backoff alone does. The waves move apart, to one, three, seven and fifteen seconds, and in the default run the last one lands after the dependency is back, so every client gets through. But each wave is still a hundred calls in the same tenth of a second. A service that has just come back up gets its whole crowd at once.
That is the job of jitter. “Jitter adds a small random offset to each wait, so that a thousand clients that failed in the same instant do not sleep the same second and reconverge as a synchronized spike.” In the third row each client waits a little less or a little more than the schedule, by a random amount, and the largest wave falls from a hundred to about a quarter of that.
| Policy, default run | Peak retries in a tenth of a second | Calls in all | Clients that got through |
|---|---|---|---|
| Retry at once | 400 | 500 | 0 of 100 |
| Back off | 100 | 500 | 100 of 100 |
| Back off with jitter | about 25 | 500 | 100 of 100 |
The policy that the fields start with is the chapter’s own illustration: “start around a second, double per attempt, cap the wait at about a minute, and give up after a handful of tries.” The chapter calls these numbers “calibrated to nothing but common practice”, and says they matter less than two rules: “retry only the transient class, since retrying a bad credential or a model’s malformed output burns money on a certainty; and bound the attempts, always”.
How much jitter to add is the tool’s choice. The chapter says only “a small random offset”; the simulator spreads each wait by 50% by default, half earlier and half later, so the schedule is no longer on average, and lets you change it. The “new random draw” button reshuffles the offsets, and the peak moves between about 20 and 33.
How do nested retries multiply?
Nested retries multiply because each layer that retries on its own repeats everything the layers beneath it did. The chapter lists a typical stack: “The HTTP client inside your tool retries three times; your tool wrapper retries the tool three times; the loop retries the step; a workflow engine above it retries the run.”
Then the arithmetic. “Four modest policies multiply into dozens of calls per failure—a self-inflicted storm arriving from inside the house.” The chapter gives a count for the first two layers only. Taking three tries at each of the four, which is the tool’s reading, the second panel shows 3 × 3 × 3 × 3 = 81 calls for one underlying failure, from a single client. No layer did anything unreasonable.
The fix is to stop treating retries as each layer’s private business: “make the retry budget global: cap the total time or attempts a run may spend retrying, across all layers, so that persistence in the small cannot add up to a siege in the large.” Type a budget into the panel and the worst case drops to that number, whatever the layers would do on their own. The agent cost-per-task estimator shows what the same overhead does to a run’s token bill.
What does a circuit breaker do?
A circuit breaker stops calls to a dependency that keeps failing, so callers fail at once and the dependency gets no extra load while it is down. The chapter describes it as a small state machine with three states. Closed is normal: “calls flow through and failures are counted.” Past a threshold it opens, and then “every call now fails instantly, without touching the dead dependency—no thirty-second timeout paid per attempt, no extra load piled on a service trying to get up.” The third state is the way back: “After a cooldown the breaker goes half-open and lets one trial call through; success closes the circuit, failure re-opens it.”
The third panel runs a steady stream of calls through that machine and beside it the same stream with no breaker. In the default run the dependency is down for sixty of the hundred and twenty seconds, and each call to a dead dependency costs a thirty-second timeout.
| Default run | No breaker | With the breaker |
|---|---|---|
| Calls that reached the dependency | 60 | 32 |
| Calls refused at once | 0 | 28 |
| Caller time in timeouts, summed over calls | 900 s | 120 s |
The strip above the table shows the states. The breaker opens on the third failure, tries once after thirty seconds, fails, opens again, and closes on the next trial once the dependency is back. Two calls are refused after recovery while the last cooldown runs out, which is the breaker’s price, and the result says so.
These savings are the best case. The breaker in the model learns each outcome at once, while a real call to a dead dependency takes its full timeout to fail, so with a call every two seconds a real breaker would let more through before it opened.
For an agent the chapter’s use is specific: “a breaker around each flaky external tool keeps one dead integration from stalling every run that touches it”. While the circuit is open, the run can take a fallback path.
The panel’s last switch demonstrates the tuning mistake the chapter warns about. “a breaker should trip on failures, and a merely disappointing result is an answer.” A lookup that finds nothing has worked. Tick the box that counts those results as failures and the default run changes: the trial call after the dependency is back happens to return “no such record”, the breaker re-opens, and fifteen calls are refused on a working dependency where two were before. Set both outage times to zero and it still opens, on a dependency that never went down. Which calls return a disappointing result is a random draw, so the count varies. The chapter’s line for the mistake: “a breaker that counts it as damage will amputate a healthy dependency.”
Which failures should be retried at all?
Only transient failures should be retried mechanically; the other four classes in the chapter each want something else. The classification comes from an earlier section of the same chapter, which sorts failures by what they want done about them, and the reference card at the foot of the tool lists all five.
| Class | The chapter’s examples | What it wants |
|---|---|---|
| Transient | Rate limits, overload responses, network timeouts on reads | A pause and another attempt |
| Model-recoverable | A wrong tool chosen, a malformed argument, output that will not parse | To go back to the model as an error it can read |
| Permanent | An authentication error, a record that does not exist, a request the service will never accept | To fail fast |
| Policy | A guardrail trip, an action the run should never have attempted | A loud halt |
| Ambiguous | A timeout on a write | Idempotency before any retry |
A missing record appears twice on this page, and both readings hold: it is permanent for the request that asked for it, and it is still an answer from a working dependency, so it is no reason for a breaker to open.
The second row is where agents differ from ordinary clients. A failure the model caused wants to return to the model, and such failures “pointedly do not want a blind mechanical retry”, since the same prompt tends to produce the same mistake. The last row is the dangerous one: “a timeout on a write is the ambiguous case—the work may have happened”. Idempotent tools and safe retries covers the key that makes such a retry safe, and Run the loop lets you make these calls yourself on a scripted run.
The chapter’s instruction is to sort before reaching for any mechanism: “Classify first, then route each class to its one primitive (retry, replan, fail fast, halt), and resist the beginner’s instinct to stack every mechanism on every step.”
What does the simulator leave out?
The simulator leaves out almost everything that makes a real outage messy, and it says so beside the results. Its model has every client fail in the same instant, calls that take no time, a dependency that is completely down for a fixed window and then completely up, and a breaker that counts failures in a row and learns each outcome immediately. Real services degrade partially and recover slowly, real clients fail at scattered times, and production breakers usually watch a failure rate over a window. The seconds are illustrative throughout.
It also leaves out cost, which the chapter does not. “Every retry adds latency and token spend; every breaker and budget is code you now test and maintain”. The rule for how much of this to build is to match it to what a failure can damage: “the read-only lookup path can fail cheap and often, and gold-plating it buys you nothing; the path that charges cards deserves the full apparatus.”
Use the three panels to see why each mechanism exists, then decide which paths in your own agent deserve them. The agent bug bestiary has the retry storm and the ambiguous write timeout as cards, with the trace clue for each. The guide to agent security and operations places reliability beside guardrails and cost, and the whole section, with timeouts, state and recovery, is in Chapter 18, Reliability, State, and the Harness (in the full book).
Questions readers ask
- Why does jitter matter if I already back off?
- Because backoff changes when clients retry and not whether they retry together. Clients that failed in the same instant and use the same schedule wake at the same moments, so each wave is still every client at once. Jitter adds a small random offset to each wait, which spreads a wave across time. In the simulator's default run it cuts the largest wave to about a quarter.
- Why should a 'no such record' result never trip the breaker?
- Because it is an answer, not a failure. The dependency did its job and reported that the record is absent. Chapter 18 says a breaker should trip on failures, and a merely disappointing result is an answer. A breaker that counts such results as damage will cut off a dependency that is working, which the chapter calls amputating a healthy one.
- What is the ambiguous failure, and why can't I just retry it?
- It is a timeout on a write: the request went out and nothing came back, so you do not know whether the work happened. Retrying may do it twice, which for a payment or a message is a real second charge or a second email. The chapter's answer is idempotency: a key that lets the receiving service recognize the repeat and return the first result.
- What is a retry storm?
- A retry storm is an outage kept alive by the retries of the clients it affected. A service is briefly overloaded and refuses requests; every client retries at once; the service now faces its normal load plus the whole wave and stays down. The clients trying to recover from the outage have become the outage.
- Are the numbers a forecast for my system?
- No. The model is deliberately simple: every client fails at the same instant, the dependency is completely down and then completely up, and the breaker counts consecutive failures. Real outages are partial and real breakers usually watch a failure rate. The simulator shows the shape of each effect so the reason for each mechanism is visible.
Sources
- Martin Fowler (2014). CircuitBreaker
- Marc Brooker (Amazon Builders' Library) (2019). Timeouts, retries, and backoff with jitter
- Marc Brooker (AWS Architecture Blog) (2015). Exponential Backoff And Jitter
- Alejandro Forero Cuervo (2016). Handling Overload (Site Reliability Engineering, chapter 21)