Idempotent tools for AI agents are tools you can call twice with the same intent and get the effect of calling once. That property is what makes an agent’s retries safe. Agents retry on timeouts, on malformed output and on resume after a crash, and a tool that isn’t idempotent turns each of those ordinary events into a second charge or a second email.
The evidence that this matters is no longer anecdotal. In a 2026 sandbox study of 25,930 agent episodes, offering an idempotency key on every write cut the duplicate rate from 28% to 4%, and agents reported success in 90% of the episodes where they had in fact duplicated an effect (Li, 2026). The agent won’t tell you it charged the customer twice, so your tool contract has to make that impossible.
This post covers why agents retry, what a correct key looks like, which retry policy fits which tool, and a worked crash-and-resume example you can check against your own code.
Why do AI agents retry tool calls at all?
Agents retry because almost every failure in a loop looks the same from the inside: a call went out and no clean answer came back. The harness retries transient errors, the model reissues calls it is unsure about, and a resumed run re-executes steps whose completion was never recorded.
The book’s chapter on tools puts it in one sentence: “Agents retry constantly (on timeouts, on ambiguous output, on a connection dropped mid-operation), and they retry without the human glance that would notice a duplicate.” A person at a terminal who sees a hung request will check the inbox before pressing send again. An agent won’t, unless you built that check into the tool.
There are at least four distinct sources of repetition, and it helps to name them because each needs a slightly different defense:
- Transport retries. An HTTP client, a tool wrapper or a workflow engine resends after a timeout or a 5xx.
- Model reissues. The model, given a vague error or a long context, calls the same tool again with slightly different arguments.
- Resume after a crash. The process dies mid-run, and the resumed run repeats any step whose result was not saved.
- Parallel duplicates. Two branches or subagents decide, independently, to do the same thing.
If you are new to how the loop produces these calls, the agent loop explainer walks the cycle step by step.
What goes wrong when a retried tool is not idempotent?
A retried non-idempotent tool repeats its side effect: the card is charged twice, the welcome email goes out twice, the ticket is filed twice. The danger concentrates in one case, the ambiguous write, where a timeout leaves you unable to tell whether the server did the work before going quiet.
Stripe’s 2017 engineering post on idempotency lists the three places a call can break: the connection fails before the request lands, the call fails midway, or the call succeeds and the connection breaks before the client hears back (Leach, 2017). Only the first is safe to retry blindly. The third is the trap, because the work is done and the client believes it isn’t.
Chapter 18 of the book calls this the ambiguous failure and gives it the example that makes it memorable: “Charge a card, hear nothing back, and ‘just try again’ becomes a question of whether your customer gets billed twice.” The same chapter traces a second route into the trap. A crash between performing a side effect and recording it leaves the resumed run convinced the step never ran.
What makes a tool idempotent?
An operation is idempotent when doing it twice has the same effect as doing it once. Reads are idempotent by nature. Some writes are too: setting a field to a value, upserting by a natural key, deleting by id. Writes that create, charge or send are not, and need an idempotency key to become safe.
So the first design move is to prefer naturally idempotent shapes wherever the business meaning allows it. set_status(ticket_id, "closed") can run ten times and the ticket is closed once. create_branch("feature/issue-123") can return the existing branch, flagged as already existing, on the second call, because the name is unique. The chapter on tools states the rule for creates directly: “a retried create should return the existing record, flagged as pre-existing, instead of minting a second one.”
For operations that must happen exactly once and have no natural key, such as sending a message or moving money, you need the service to recognize a repeat. Most idempotent tools for AI agents are built from these two ingredients: a naturally idempotent shape where one exists, and a key where it doesn’t. That is what the idempotency key is for.
How do idempotent tools for AI agents use idempotency keys?
An idempotency key is a unique token attached to a mutating call. The service stores the first result under that key and returns the stored result for any repeat instead of doing the work again. The key must be the same on every retry of the same logical action, or the mechanism does nothing.
Payment APIs are the reference implementation of the category. Stripe’s API reference, as one example, says it saves “the resulting status code and body of the first request made for any given idempotency key, regardless of whether it succeeds or fails,” compares later parameters to the original and errors if they differ, and may remove keys once they are “at least 24 hours old” (Stripe API reference). Amazon’s Builders’ Library describes the same idea as a caller-provided client request identifier (Featonby).
The part agents get wrong is derivation. Chapter 18 gives the rule: derive the key “from the action’s logical identity—run, step, tool, arguments—so that every retry of the same logical action carries the same key.” A fresh random value per attempt looks diligent and defeats the whole point.
Random keys, such as the UUIDs Stripe suggests, are fine only if the key is generated once per logical action and saved with the intent, not regenerated on each attempt. In Li’s study, every remaining duplicate under a key-offering contract involved an agent that sent the first attempt without a key or changed the key when it retried.
Amazon’s write-up adds a caution about the opposite shortcut. Hashing the request parameters to detect duplicates “doesn’t work in all cases,” because a caller may really want two identical resources. The key should express intent, which is why the run and the step belong in it.
Why does tool shape decide whether a key can work?
Tool shape matters because a key derived from arguments is only stable if the arguments are stable. A generic send_email(to, subject, body) tool lets the model regenerate the body on every attempt, and a reworded body produces a different hash or trips the service’s parameter check.
This is where idempotency meets the book’s broader advice on tool design. Chapter 5 argues that “The right unit for a tool is a task; the endpoints are plumbing, and plumbing belongs inside.” A task-shaped send_invoice(invoice_id) tool fixes the identity of the action in code. The tool renders the email from the invoice, derives the key from the run, the step and invoice_id, and the model never touches the parts that must stay byte-identical.
I find this the single most useful reframing for engineers coming from API work. When you design idempotent tools for AI agents, you aren’t only making the endpoint safe to repeat. You’re narrowing what the model can vary, so that “the same action” has one representation. The companion post on how to design tools for LLM agents covers the rest of that craft, and the post on why agents call the wrong tool covers what happens when the shapes overlap.
Worked example: an invoice email, a crash, and a resume
Here is an illustrative run, simplified enough to trace by hand. An agent closes out a customer order in three steps: create the invoice, send the invoice email, then mark the order billed. The process is killed by a routine deploy after the email provider accepted the send but before the agent saved its checkpoint.
Try to predict what the resumed run does before reading on. Its checkpoint says step 1 finished and step 2 did not. That is all it knows.
The naive path. The resumed run reloads the checkpoint, sees step 2 as pending and calls send_email again. The model, working from a reloaded transcript, writes a slightly different body.
The customer receives two invoice emails a few minutes apart, and the run reports success. Nothing in the logs looks wrong, because each send succeeded.
The receipt path. The tool follows the three verbs Chapter 18 asks you to memorize: “record intent before the call, execute through the idempotent wrapper, record a receipt after—and on any retry, check for the receipt before doing anything.”
| Moment | What the tool does | State it leaves |
|---|---|---|
| Before the send | Writes intent row: key run-81f2:step-2:invoice-4471, status pending |
Intent recorded |
| The send | Calls the provider with the same key | Provider stores result under key |
| Crash | Process dies before receipt or checkpoint | Intent pending, no receipt |
| Resume, step 2 | Finds pending intent, retries with the same key |
Provider returns stored result, no second email |
| After | Writes receipt with the provider’s message id | Receipt is the system of record |
The book’s sentence for why this works is worth keeping: “A checkpoint records position. It does not, and cannot, record whether the outside world did the work.” And the corollary: “The receipts are the system of record; the transcript is the agent’s memory of writing them.”
What if the provider doesn’t accept keys?
Then the tool has to fall back on a read-back, and that is only safe under conditions you should check. Before resending, it searches the provider’s sent log for a message tagged with invoice-4471. If the log is immediately consistent and the original request cannot still be in flight, a hit means “already sent” and a miss means “safe to send.”
If either condition fails, the honest move is to stop. Li (2026) proves that no verification-only policy is exactly-once under late commits without a bound on in-flight time, and measured frontier models duplicating in 56% of episodes where the request was still in flight. The paper adds an uncomfortable corollary: a read-back path that cannot see lagging effects was worse than none, because agents trusted it instead of escalating. So the tool should return outcome: unknown and let the run escalate, which the book’s rule for deeds already covers: “Fall back on words; fail loud on deeds.”
Which retry policy fits which tool?
Retry policy should follow the tool’s class, not a global default. Reads retry freely within a budget; naturally idempotent writes retry the same way; creates, charges and sends retry only with the same key; and writes with unknown outcomes and no key support should not auto-retry at all.
| Tool class | Examples | Automatic retry? | What makes it safe |
|---|---|---|---|
| Read | search, get, list | Yes, transient errors only, bounded | Nothing extra needed |
| Natural idempotent write | set field, upsert by id, delete by id | Yes, bounded | Stable target identifier |
| Keyed write | send, charge, create | Yes, with the same key only | Idempotency key plus receipt |
| Unkeyed write, ambiguous outcome | provider with no key support | No | Consistent read-back, or escalate as unknown |
| Multi-step irreversible | reserve, charge, issue | No whole-workflow retry | Per-step keys plus compensating actions (a saga) |
Writing these classes into your tool descriptions helps the model too. Protocols for tool servers have started to carry the property as metadata. The Model Context Protocol’s idempotentHint, as one example, means “calling the tool repeatedly with the same arguments will have no additional effect,” but the spec is explicit that annotations are hints and should not be trusted from untrusted servers (MCP schema). Treat the hint as documentation of a guarantee your code enforces, never as the guarantee.
Why surface errors to the model instead of retrying blindly?
Because some failures want a retry and others want a decision. A malformed argument or a wrong tool choice needs the model to see the specific mistake and fix it; replaying the identical call re-rolls the same dice. As Chapter 18 puts it, “the model can only recover from a failure it can see.”
The chapter sorts failures by recovery path, and the sorting is what keeps retries in their lane:
| Failure class | Example | Right response |
|---|---|---|
| Transient | rate limit, timeout on a read | Back off with jitter, retry, bounded |
| Model-recoverable | malformed argument, unparseable output | Return a structured error to the model |
| Permanent | bad credential, record does not exist | Fail fast |
| Policy | guardrail trip | Halt loudly |
| Ambiguous | timeout on a write | Key-protected retry, or escalate |
A structured error carries a stable code, a retryable flag and a hint, such as “connection refused; retryable; the database may be unreachable from outside the VPN.” That shape also answers the question many teams arrive with, why does my AI agent keep looping? Often it’s because a vague error invites the model to retry with small variations, forever. The post on why agent errors compound shows how one unrecovered failure of this kind spreads through the rest of a run.
How do retry budgets and circuit breakers stop retry storms?
Retry budgets cap total retry effort across every layer, and circuit breakers stop calling a dependency that keeps failing. Both exist because retries nest: three retries in the HTTP client, three in the tool wrapper and three in the loop multiply into dozens of calls per failure.
The arithmetic is worth doing once. Three layers that each make up to three attempts can produce 3 × 3 × 3 = 27 calls for one failing operation (an illustrative count, but a real pattern). If that operation is a model call, each of those attempts also costs tokens; the agent cost per task estimator will show you what a retry multiplier does to a per-task bill. Chapter 18’s fix is to make the budget global, capping the time or attempts a whole run may spend retrying.
For a dependency that has stopped answering, the circuit breaker, popularized by Michael Nygard’s Release It! and described by Martin Fowler, trips open after a failure threshold, fails calls instantly during a cooldown, then lets one trial call through.
One tuning note from the book matters for agents in particular: count failures, not disappointing answers. A lookup returning “no such record” is a normal result, and a breaker that counts it as damage will cut off a healthy dependency.
Does durable execution make idempotent tools unnecessary?
No. Durable execution journals each step’s result and replays the code after a crash, so completed steps return their recorded value instead of running again. But a step that crashed before its result was journaled is retried, and if that step had a side effect, the retry repeats it unless the tool is idempotent.
Engines in this category (Temporal, Restate and Inngest are examples) describe the promise as crash-proof execution (Wheeler, 2025). Their own documentation is candid about the boundary. Temporal’s docs recommend that activities be idempotent and explain that an activity which “fails to report to the server at all” will be retried (Temporal docs). Chapter 18 compresses the idea into one line: “The code replays; the world does not.”
So the engine and idempotent tools for AI agents are complements. The engine guarantees your checkpoint and your run’s position survive the process; the tool guarantees the outside world sees each action once. The follow-up post on durable execution for AI agents covers when the engine is worth buying.
Where does this advice stop applying?
It stops paying for itself on short, read-only runs. A thirty-second research agent that only searches and summarizes needs bounded retries and nothing else. Keys, receipts and journals earn their cost on tools that create, charge, send or delete, and on runs long enough to meet a crash.
Three limits are worth stating plainly. Keys have a retention window, and Stripe documents that keys may be removed once they are “at least 24 hours old”; a run parked at an approval gate for three days can outlive its own protection, so your receipt store must keep records at least as long as your longest resume. Keys protect a single step from repeating itself and say nothing about a sequence that must be undone; that is the saga’s job. And Li’s results come from a controlled sandbox in a single-author preprint, so read the exact percentages as indicative rather than as rates for your system.
Retrofitting idempotency onto command-line programs and internal services built for humans has its own wrinkles, covered in agent tools for software built for humans.
How do you audit your tool table in an afternoon?
List every tool your agent can call, classify each one by the table above, and check four properties for every write: a stable key derived from logical intent, a receipt written after success, a defined behavior on unknown outcomes, and a place in the global retry budget.
This is the cheapest way to find out how many of your tools are already idempotent tools for AI agents and how many only look that way. In practice, the audit is a spreadsheet with one row per tool and these columns:
- Class. Read, natural idempotent write, keyed write, unkeyed write, or multi-step.
- Key source. Which fields form the key, and is it the same on retry? Grep for random key generation inside retry loops.
- Receipt. Where the record of completion lives, and whether a resumed run checks it first.
- Unknown outcome. What the tool returns when it can’t tell whether the write landed.
- Retry owner. Which single layer retries this tool, under which budget.
Then do what Chapter 18 calls the habit that separates state engineering from state theater: crash on purpose. Kill the process after the provider accepts a send and before the receipt is written, resume, and watch what the run does. A harness whose recovery you have never exercised is a guess; the post on what an agent harness is covers the rest of what that layer owns.
The one thing to carry away: an agent’s retries are inevitable, so make every tool call that changes the world safe to repeat before the agent ever repeats it. The full treatment, including the saga pattern and state machines, is in Chapter 18, Reliability, State, and the Harness (in the full book), with the tool-design side in Chapter 5, Tools and the Action Space (in the full book). Part I is free to read, and the book is available in several formats.
Questions readers ask
- What is an idempotent tool for an AI agent?
- A tool whose repeated calls with the same logical intent leave the world in the same state as a single call. Reads are idempotent by nature. Writes become idempotent through a natural key (set, upsert, delete by id) or an idempotency key the service uses to return the stored result instead of acting twice.
- Should an agent generate a new idempotency key on every retry?
- No. A fresh key per attempt defeats the mechanism, because the service sees each retry as a new request. Derive the key from the action's logical identity, run id, step, tool and arguments, as Chapter 18 puts it, so every retry of the same action presents the same key.
- Is checking whether the email was sent before resending enough?
- Only when the read path is immediately consistent and the original request cannot still be in flight. If the first attempt might commit late, the check finds nothing and the retry duplicates. In that case use a key, or report the outcome as unknown and escalate instead of resending.
- Does durable execution remove the need for idempotent tools?
- No. Durable execution stops completed steps from re-running on replay, but a step that crashed before reporting its result is retried. If that step sent an email or charged a card, the retry repeats it unless the tool is idempotent.
- Why is my AI agent looping on the same failing tool call?
- Usually because the error it receives is vague, so it retries with small variations, or because several retry layers are stacked. Return a structured error that says what failed and whether it is retryable, and cap retries with one global budget across all layers.
Sources
- Jiapeng Li (2026). Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents
- Stripe (2026). Idempotent requests (API reference)
- Brandur Leach (Stripe) (2017). Designing robust and predictable APIs with idempotency
- Malcolm Featonby (Amazon Builders' Library). Making retries safe with idempotent APIs
- Model Context Protocol (2025). Model Context Protocol schema reference: ToolAnnotations
- Temporal (2026). Activity Definition: Idempotency
- Tom Wheeler (Temporal) (2025). The definitive guide to Durable Execution
- Martin Fowler (2014). CircuitBreaker