The short answer to how to make AI agents more reliable is a build order of six mechanisms: classify failures and show them to the model, bound timeouts and retries, cap the run, record receipts for writes, checkpoint long runs, and choose the fallback before the outage. Each one depends on steps before it. None of them touches the prompt.
I wrote this for the backend engineer who inherited an agent that works in the demo. It gives the sequence, what each step buys, what it costs and the test that proves it. The taxonomy of what goes wrong is a separate reference, the field guide to AI agent failure modes; here the subject is what to build, and in which order.
Why is making an agent reliable a harness problem?
It is a harness problem because the failures you can remove with engineering happen in the code around the model: a tool call that times out, a process that dies, a write that runs twice. The model’s own mistakes are a different problem with different tools, and this post leaves them to evaluation.
The published record supports the split. A survey of teams running agents in production (Pan and colleagues, Measuring Agents in Production, last revised 2026) reports that “Reliability (consistent correct behavior over time) remains the top development challenge, which practitioners currently address through systems-level design.” A 2026 single-author review of coding agents finds that “many apparent model failures originate elsewhere in the system” (Jarmak, 2026).
How large is the gap being closed? The τ-bench paper (Yao and colleagues, 2024) tested tool-using agents on simulated customer-service tasks and reported that the strongest agents it tested succeeded on under 50% of tasks, and that a retail task was solved in all eight of eight trials less than 25% of the time. Those are 2024 figures for one benchmark. They mix model errors with everything else, which is the reason to remove the part that is plain engineering first.
Chapter 18 of the book (in the full book) frames the subject as the unhappy path. “For an agent that runs dozens of steps against real systems, tool failure is the ordinary operating condition,” it says.
How to make AI agents more reliable: the sequence in one table
Build the rows in order and skip the ones your agent does not need: rows 4 and 5 are conditional, the other four are not. The order is mine; the mechanisms are Chapter 18’s, in close to the order the chapter presents them. The last column says which runs need the row, and the buttons narrow the table to one kind.
| Step | What you add | What it buys | What it costs | Applies to |
|---|---|---|---|---|
| 1. Classify and surface | Every failure gets one of five classes; tool errors return to the model as a structured result | The model stops reasoning over holes; every later step knows what it is handling | Hours of wrapper code; one extra pass per fed-back error | every run |
| 2. Timeouts and bounded retries | A deadline on every call; retries for the transient class only, owned by one layer, drawn from one run-wide allowance | Transient failures stop killing runs; no call hangs; no retry storm from your own layers | A few percent more calls; latency on the bad path; a policy to maintain | every run |
| 3. Run budgets | Ceilings on passes, tokens, wall clock and money, with a status that says “stopped” | A bound on the worst run; a loop that ends | Some good runs cut off; a cap to re-derive when the model or tools change | every run |
| 4. Receipts for writes | An idempotency key per logical action; intent recorded before, receipt after | A retry or a resume cannot repeat a side effect | A durable store; a wrapper for each write tool; a rule for unknown outcomes | runs that write |
| 5. Checkpoints | State saved at step boundaries under a stable run id; a tested restore | A crash, a deploy or an approval pause costs a resume, not a rerun | Storage, serialization, version pinning, a restore drill | runs that outlive a process |
| 6. Fallbacks and escalation | A chosen substitute for each critical read; a loud halt for failed writes; an escalation package | The failure that stays ends in a labeled answer or a short human decision | A second path to keep alive; staleness to label; someone to receive escalations | every run |
“Every run” rows are the base. The other two buttons show what a property adds on top. A circuit breaker belongs to step 2 and only earns its place when many runs share one dependency; it has its own section below.
The order follows dependencies. A retry policy needs the class from step 1, or it retries things that cannot succeed.
A run budget has to count the retries and fed-back errors that steps 1 and 2 create. A checkpoint needs a step boundary and a status to save, which step 3 defines, and it is unsafe for writes until step 4 exists. A fallback is what runs after step 2 has given up.
Cost rises in the same direction. Steps 1 to 3 are code in the tool wrapper and the loop. Steps 4 and 5 need durable storage and a drill. Step 6 needs a second path and a person.
Step 1: which failures go back to the model, and which don’t?
Failures the model caused go back to the model as a structured result, and transient ones are retried by code. Permanent ones fail fast, policy ones halt the run, and an ambiguous one waits for step 4. Chapter 18 sorts them by recovery path and gives the rule: “Classify first, then route each class to its one primitive (retry, replan, fail fast, halt), and resist the beginner’s instinct to stack every mechanism on every step.”
| Class | Examples from the chapter | The one response |
|---|---|---|
| Transient | Rate limits, overload responses, network timeouts on reads | Pause and retry, bounded (step 2) |
| Model-recoverable | A wrong tool chosen, a malformed argument, output that will not parse | Return to the model as a structured error |
| Permanent | An authentication error, a record that does not exist | Fail fast; tell the model the call cannot succeed |
| Policy | A guardrail trip, an action the run should never have attempted | Halt loudly |
| Ambiguous | A timeout on a write | No resend without an idempotency key (step 4) |
The second row is where service habits mislead. The service reflex on a failed call is to send it again, and for a failure the model produced that is the wrong primitive. The chapter’s reason: “replaying the identical prompt re-rolls the same weighted dice, and a model that just emitted invalid output has decent odds of emitting it again.” Show the model the specific mismatch instead.
A structured error, in the chapter’s shape, carries a stable code, a retryable flag and a hint. This one is illustrative:
{
"error": {
"code": "invalid_argument",
"message": "start_date must be YYYY-MM-DD; got '10/07'",
"retryable": false,
"hint": "Reformat the date and call again."
}
}
The figure shows the alternative. When a wrapper catches the exception and returns an empty result, the model concludes there was nothing to find, and the chapter’s verdict is “The failure has been laundered into an answer.” The sentence to keep from this step is the chapter’s: “the model can only recover from a failure it can see.”
Outside sources agree on the mechanism and on its limit. One vendor’s account of building a research agent says that “letting the agent know when a tool is failing and letting it adapt works surprisingly well” (Anthropic, 2025). The 12-factor agents guide warns that “if you do this TOO much, your agent will start to spin out and might repeat the same error over and over again,” and suggests a counter “to limit to ~3 attempts of a single tool” (Horthy, HumanLayer).
So a fed-back error is bounded twice: it costs one pass of the run budget, and a count of consecutive errors on one tool ends the run when it passes a small number. The policy class is the subject of the post on AI agent guardrails, where the default on a trip is to halt.
Buys: the swallowed error disappears, and every later step has a class to act on. Costs: wrapper code for each tool and one pass per fed-back error. Proof: make a tool raise, and assert that the next model input contains the code and the hint.
Step 2: what should the retry policy be?
Retry the transient class only, give every call a deadline, let one layer own the retries, and draw them all from one allowance per run. Chapter 18 warns against skipping the design: “The retry looks too small to deserve engineering. It is the most common way agents turn one failure into many.”
Start with what a bounded retry buys. This is a worked example with illustrative numbers. Take a run of 12 tool calls where each call has an independent 5% chance of a transient failure, and a wrapper that crashes the run on any failure. The share of runs that get through is 0.95¹² = 54.0%.
Now allow three attempts per call. A call fails all three with probability 0.05³ = 0.0125%, and the run survives with probability (1 − 0.000125)¹² = 99.85%. The expected number of calls rises from 12 to 12 × (1 + 0.05 + 0.0025) = 12.63, which is 5.25% more. The why of the first number is the subject of the post on compounding errors in AI agents.
What do retries do in an outage?
In an outage, retries buy nothing: the arithmetic above assumes independent failures, and during an outage every attempt fails for as long as it lasts. The simulator below opens on that case, and its numbers are the tool’s own model, not measurements.
Forty runs fail on the same tool at the same moment, and the tool stays down for 20 seconds. The retry policy is the chapter’s illustration: a first wait of about a second, doubling, five tries.
The waits of 1, 2, 4 and 8 seconds put the last retry at 15 seconds, still inside the outage. All 40 runs give up after 200 calls between them, whether they retry at once, back off, or back off with jitter. Backoff and jitter did cut the largest burst, from 160 retries in a tenth of a second to 40 and then to 13. They could not outlast the outage.
With JavaScript on, the Circuit breaker and retry simulator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Circuit breaker and retry simulator on its own page to share a result by link.
Why should one layer own the retries?
One layer should own the retries because layers that each retry multiply. In the simulator’s second panel, an HTTP client that tries three times, inside a tool wrapper that tries three times, inside a loop that tries the step twice, makes 3 × 3 × 2 = 18 calls for one failure. Marc Brooker gives the service-side version in the Amazon Builders’ Library: with a five-deep stack and three retries per layer, “the load on the database will increase 243x.” For low-cost operations, he writes, “our best practice is to retry at a single point in the stack.”
The chapter’s fix is to “make the retry budget global: cap the total time or attempts a run may spend retrying, across all layers.” In the simulator a budget of 6 turns the 18 into 6. Google’s SRE book describes the same two-level idea for services: “a per-request retry budget of up to three attempts” plus a per-client budget under which a request “will only be retried as long as this ratio is below 10%” (Forero Cuervo, 2016).
Timeouts complete the step. A call with no deadline can hold a run forever, so every call gets one, and crossing it is a failure that goes through the classes like any other. A timeout on a read is transient. A timeout on a write is the ambiguous class, and nothing in step 2 makes resending it safe.
Buys: blips stop ending runs, no call hangs, and your own layers cannot multiply a failure. Costs: a few percent more calls on average, more latency on the bad path, and a policy that has to be kept in one place. Proof: inject two failures and then a success, and assert three calls; inject a permanent failure, and assert exactly one.
When does a circuit breaker earn its place?
A circuit breaker earns its place when many runs share one dependency, because its job is to stop the whole fleet paying a timeout each on something that is known to be down. Martin Fowler’s description of the pattern, which he credits to Michael Nygard’s Release It!, is that once failures reach a threshold “all further calls to the circuit breaker return with an error, without the protected call being made at all” (Fowler, 2014).
The simulator’s third panel runs the same 20-second outage against a stream of one call every 3 seconds for 2 minutes, 40 calls in all, with a 30-second timeout. Seven calls fall inside the outage.
With no breaker, all seven wait out the timeout: 210 seconds of caller time. With a breaker that opens after three failures in a row, three calls time out and four are refused at once: 90 seconds. One trial call after the cooldown finds the tool back and closes the circuit.
Two cautions. The chapter’s: “a breaker should trip on failures, and a merely disappointing result is an answer.” And Brooker’s: circuit breakers “introduce modal behavior into systems that can be difficult to test.” A single agent with one run at a time gets most of the benefit from the run-wide retry allowance alone.
Step 3: what does a run budget add once retries are bounded?
A run budget bounds the whole run, where a retry policy bounds one call: it puts a ceiling on passes, tokens, wall-clock time and money, and it ends the loop with a status that says the run was stopped. Chapter 3 defines it as “a hard ceiling enforced by the harness, indifferent to the model’s opinion.”
Retries and fed-back errors are why it comes third. Step 1 lets the model try again after a bad argument, and step 2 lets code try again after a blip. Both are bounded per call, and neither bounds a model that keeps finding new calls to make. The budget does, and it should count what the first two steps spend: a fed-back error is a pass, and a retry is wall-clock time.
Three details make the cap useful to the caller. It needs its own status, so that a stopped run is never reported in the voice of a finished one. The remaining allowance should be checked before each call and not after the overrun. And the number should come from your own verified runs, which is the method worked through in the post on stop conditions for agent loops.
Buys: a known worst case per run, and an end to the loop that never converges. Costs: some good runs are cut off at whatever rate you chose, and the caps go stale when the model, the tools or the prompt change. Proof: stub a model that asks for the same tool forever, and assert that the run ends with the stopped status at the cap.
Step 4: why do receipts come before checkpoints?
Receipts come first because a checkpoint without them makes duplicates more likely: a resumed run repeats whatever was in flight when the process died. The chapter states the limit exactly: “A checkpoint records position. It does not, and cannot, record whether the outside world did the work.”
The window is the gap between performing a side effect and recording it. The email service accepts the send, the process dies before the state is saved, and the resumed run sends again. The same gap opens in step 2 when a write times out. One discipline closes both: an idempotency key derived from the action’s logical identity, the intent recorded before the call, and a receipt recorded after.
The measured effect is large in the one study I know of. In a 2026 sandbox experiment of 25,930 agent episodes, offering an idempotency key on every write “lowers the duplicate rate from 28% to 4%,” and “agents reported success in 90% of the episodes in which they had duplicated an effect” (Li, 2026). It is a single-author preprint in a controlled sandbox, so read the percentages as indicative. How to derive the key and what to do when a provider accepts none is covered in idempotent tools and safe retries.
Buys: a retry or a resume returns the first result and does nothing twice. Costs: a durable store that outlives the run, a wrapper for each tool call that writes, and a decision about what to return when the outcome is unknown. Proof: kill the process after the downstream call succeeds and before the receipt is written, resume, and count the side effects. The count must be one.
A read-only agent skips this step.
Step 5: when is a checkpoint worth building?
A checkpoint is worth building when a run can outlive the process that started it: a long task, a deploy in the middle of work, or a pause while a person approves a step. The chapter defines it as “a snapshot of the working state, saved at a deliberate boundary, keyed by a stable run identifier, so that a fresh process can reload it and continue.”
What it buys is the price of dying. Take the 12-step run again, and a crash that is equally likely after any step; the unit here is steps, and the numbers are illustrative.
With no checkpoint the restart redoes 6.5 steps on average and 12 at worst. With a checkpoint every third step it redoes 1.0 on average and 2 at worst, for four saves per run. With one after every step it redoes none, for twelve saves.
The chapter’s guidance on cadence matches that curve: “checkpoint every few units of work: every step is wasteful, only-at-the-end is catastrophic, and for a task that finishes in one sitting the honest cadence is none at all.” What to save follows its split between state that is cheap to recompute and state that is not: “Persist the durable part aggressively and the ephemeral part lazily or never.”
There are two ways to implement it. One serializes the working state at each boundary. The other journals the result of every model and tool call and, on restart, replays the code against the journal, which is the idea behind durable execution: “The code replays; the world does not.” Temporal, Restate and Inngest are examples of engines in that category. When one is worth adopting is the subject of the companion post on durable execution for AI agents.
Either way, pin what the saved state depended on. If the prompt, the model or a tool’s schema changed between the crash and the restore, the resumed run is a different program reading old state. An approval is durable state too, and the chapter asks that it be recorded with who granted it, when, and a fingerprint of what they saw.
Buys: an interruption costs a resume. Costs: storage, serialization of everything in the working state, version pinning and a drill. Proof: the chapter’s habit, “crash on purpose.” Kill the process right after a model call, restore, and watch what the resumed run does, because “a checkpoint you have never restored from is a guess.”
Step 6: what happens when the failure stays?
When the retries are spent and the dependency is still down, a read falls back to a worse source that says so, a write halts loudly, and a person receives enough to decide in seconds. The chapter compresses the rule into one line: “Fall back on words; fail loud on deeds.”
For reads, the design work is choosing the substitute before the outage: a second model, plain text search behind the polished one, a replica or this morning’s cache “marked as stale” so the label travels with the answer. In the simulator’s run, this is what the 40 runs should have had after their fifth try.
For writes there is no worse version worth having. A charge, a send, a delete or a merge that cannot complete should stop the run and say so. A model left to improvise around a failed write may report work that never happened.
The bottom of every ladder is a person, and the chapter lists what the handoff should carry: the goal in one line, what was attempted, what failed (in the structured shape from step 1), what was already tried, a recommended next step and a link to the full trace. Every one of those is a field the earlier steps already produce.
Buys: an outage becomes a labeled answer or a short human decision. Costs: a second path that must be kept working, and people on the receiving end. Proof: force the primary tool to fail permanently, and assert a stale-labeled answer on a read path and a halted status with a complete package on a write path.
How far down the sequence do three agents go?
How far down depends on two questions the chapter ends its state section with: “what here is irreversible, and how long does the run live?” For step 4, read irreversible as any write a retry or a resume could repeat: a duplicate note can be deleted, but it has already been sent. Three illustrative agents show the range.
| Agent | Writes it could repeat? | Outlives a process? | Steps it needs |
|---|---|---|---|
| Answers questions over documents in about thirty seconds, read-only | No | No | 1, 2, 3, 6 |
| Support agent that issues refunds; a person approves each one | Yes | Yes, while it waits for the approval | 1 to 6 |
| Nightly migration agent, forty parallel runs filing tickets through one API | Yes | Yes | 1 to 6, plus a breaker on the ticket API |
The second row is easy to misjudge by clock time. The run takes two minutes of work, but an approval can sit for a day, and a run parked at a gate outlives its process as surely as a long one. Remove the approval and the same agent needs steps 1, 2, 3, 4 and 6.
The first row matches the chapter’s own sizing: “the thirty-second read-only summary needs none of it,” it says of state engineering, and “Reliability engineering, like insurance, is priced per what it protects.”
How do you prove each step works?
Prove each step with a test that has no model in it, in the same order you built them. A stub that returns scripted model outputs is enough for all of them.
- Step 1. A tool that raises produces a tool result with a code, a retryable flag and a hint, and the next model input contains it.
- Step 1. A tool that fails the same way on every call ends the run at the consecutive-error cap, with a status that is not “done.”
- Step 2. Two injected transient failures and a success produce exactly three calls; a permanent failure produces exactly one.
- Step 2. A tool that hangs is ended by its timeout, and a timed-out write is not resent without a key.
- Step 2. With every layer failing, the total calls for one step equal the run-wide retry allowance, not the product of the layers.
- Step 2, only if many runs share a dependency. The breaker opens on failures, stays closed on “no such record,” and closes after one successful trial.
- Step 3. A stub model that requests the same tool forever is stopped at the cap, and the caller sees the stopped status.
- Step 4, skip if the agent never writes. A crash between the downstream success and the receipt, followed by a resume, leaves exactly one side effect.
- Step 5, skip if no run outlives its process. A process killed after a model call resumes from the last checkpoint and redoes no more steps than the cadence allows.
- Step 5. A restore against a changed prompt, model or tool schema is detected and refused or flagged.
- Step 6. A permanently failing read returns the fallback’s answer with a staleness label; a permanently failing write halts with the escalation package filled in.
Each test also names a field to record: the error class, the attempt number, which budget fired, the checkpoint id, the fallback used. Those fields are what makes the next incident readable. They tell you what the machinery did and not whether the run’s answer was right, which is the argument of the post on agent observability.
What does this sequence not fix?
The sequence does not fix a run that completes cleanly and is wrong. The chapter closes on the point: “None of it makes the model smarter.” A model that picks the wrong customer, misreads a policy or declares success early passes through all six steps untouched, and catching it takes a check on the result.
The costs are real and they add up. The chapter’s own accounting is that an agent that “retries, checkpoints, double-checks, and reviews itself is an agent that spends several times what its naive twin spends.” Build the rows you need and no others.
Three more limits. The retry arithmetic above holds for independent blips and fails for correlated outages, as the simulator shows. The breaker and retry numbers come from a deliberately simple model with instant calls and a clean outage. And I found no neutral measurement of how much each step raises a production agent’s success rate, so this post’s answer to how to make AI agents more reliable is arithmetic and tests, with no promised percentage.
The two questions to carry out of this
The practical answer to how to make AI agents more reliable is to build steps 1 to 3 and step 6 for every agent, then ask the chapter’s two questions about repeatable writes and lifetime. A run that writes adds receipts. A run that outlives its process adds checkpoints. Then crash it on purpose, because the tests are the only evidence that any of it works.
Chapter 18, Reliability, State, and the Harness (in the full book) develops each mechanism, including sagas, state machines and the debate over how much harness to build; Chapter 3, The Agent Loop (in the full book) covers budgets and stop conditions. Try the outage yourself in the retry storm simulator, find this post’s neighbors in the AI agent security and operations guide, or see the formats.
Questions readers ask
- How reliable are AI agents?
- Less than a single success suggests, on the published evidence. A 2024 benchmark of tool-using agents in simulated customer service reported that the strongest agents it tested solved under half the tasks, and that a retail task was solved in all eight of eight trials less than a quarter of the time. A survey of production teams, last revised in 2026, names reliability as the top development challenge.
- Should I retry a failed model call?
- Retry it when the failure is transient: a rate limit, an overload response, a network timeout. Do not replay it when the model itself produced the failure, such as a malformed argument or unparseable output. Return that error to the model as a structured tool result so it can correct the call, and cap how many times in a row it may fail.
- What is the difference between a checkpoint and durable execution?
- A checkpoint is a saved snapshot of a run's working state that a fresh process can reload. Durable execution is a way of running the code so that every model and tool result is journaled and a restart replays the code against the journal. Both record the run's position, and neither records whether the outside world did the work.
- Do I need a durable execution engine to make an agent reliable?
- Not for the first three steps, which are code in the tool wrapper and the loop. An engine starts to pay when runs outlive a process: long tasks, pauses for human approval, deploys in the middle of work. Whichever way the state is saved, tools that write still need idempotency keys and receipts.
- Can I make an agent more reliable without changing the prompt?
- Yes. Every step in this sequence lives in the harness, the code around the model: error results, retry policy, budgets, receipts, checkpoints and fallbacks. They do not make the model's decisions better. They stop ordinary failures of tools, networks and processes from becoming wrong answers, duplicate actions or lost work.
Sources
- Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik Narasimhan (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Melissa Z. Pan et al. (2025 (v4 2026)). Measuring Agents in Production
- Stephanie Jarmak (2026). Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model
- Anthropic (2025). How we built our multi-agent research system
- Dex Horthy, HumanLayer (undated; read October 2026). 12-Factor Agents, factor 9: Compact Errors into Context Window
- Marc Brooker, Amazon Builders' Library (2019). Timeouts, retries, and backoff with jitter
- Alejandro Forero Cuervo (2016). Handling Overload (Site Reliability Engineering, chapter 21)
- Martin Fowler (2014). CircuitBreaker
- Jiapeng Li (2026). Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents