The stop conditions for agent loops come in three kinds, which Chapter 3 of the book calls exits: the run is done, a budget fires, or an error halts it. The model decides only the first, and only when no program can check its word. Everything else belongs to the code around it.
That much fits on a slide. This guide to stop conditions for agent loops is for the backend engineer about to ship a loop, who needs the parts a slide leaves out: what triggers each exit, what the caller receives, how to pick the number for a cap, and how a caller tells “done” from “gave up” from “was stopped.” The source is the section “Stop Conditions, Budgets, and Frameworks” in Chapter 3 (in the full book). The loop itself is covered in what an agent loop is.
What are the three stop conditions for agent loops?
They are a finished run, a fired budget and a halting error, in that order in the chapter. The first is the happy path, the second is “the one I called mandatory,” and “Errors are the third family of exits.” The book’s glossary entry for stop condition counts the same way: “The rules that end a run. Three families”.
The chapter’s figure draws the same ground a little differently, and the difference is worth seeing before you build on it.
That drawing splits the first exit into two doors ranked by trust, a verified check and the model’s own “done,” and leaves error halts to the prose. So the prose counts three families and the figure draws three ways out, and they overlap on two. The post on what an agent loop is counts as the figure does (verified, judged, capped) and treats errors as a further family; the ground is the same. Put together, they give four distinguishable ways a run can end. That is the number of statuses the rest of this post uses: three exits, four statuses.
| Exit (Chapter 3) | What triggers it | Who decides | Status the caller receives | What the caller can assume |
|---|---|---|---|---|
| 1. Done, checked | The model offers a final answer and a program’s check passes | The check, in the harness | done_verified |
The checked condition holds |
| 1. Done, claimed | The model offers a final answer and no check exists | The model | done_claimed |
The model believes it is finished |
| 2. Budget | A ceiling on passes, tokens, wall clock or money is reached, or (this post’s addition) a limit on repeated actions | The harness | stopped_budget + the limit |
Nothing about the work; the loop ended |
| 3. Error | A tool or guard reports a failure the model should not get another try at | The harness | halted_error + the error class |
The run went somewhere it should not have |
The status names are mine. Any four names will do, provided a caller can branch on them without reading prose.
Who decides that the run is over, the model or the harness?
The model decides one thing: when to stop asking for tools and offer an answer. The harness decides the rest, and where a check exists, the model’s offer does no more than trigger the check.
Chapter 3 explains why the offer cannot be the only door. The model’s judgment “fails in both directions”: it can announce the fix “while two tests still fail, in the same assured tone as a true report,” and it can go on “finding one more thing to polish, then another, while the bill runs.” The first failure has the same root as the one described in why AI agents hallucinate: the model produces the plausible sentence, and “the task is complete” is always plausible.
The chapter’s instruction is short: “Where such a check exists, wire it into the loop and let that check pronounce the run finished.” Its reason is one line: “The model’s “done” is testimony; a green test is evidence.”
What the chapter leaves to the builder is the case where the model says done and the check fails. My reading is that this is not an exit. The check’s output goes into the message history as one more result, the loop goes again, and the budgets bound how long that can last. The same holds for an empty or malformed final answer, which the usual rule (no tool calls means finished) would otherwise accept as done.
A reply with no tool calls is the usual signal. One vendor’s SDK documentation, as an example, states its rule for a final output as text of the expected type where “there are no tool calls.” That is a fact about one model call. The status of the run is a separate value, and your code is what sets it.
What does a loop with all three exits look like?
It is the chapter’s ten-line loop with four return statements where the original has two. The sketch below is this post’s pseudocode, in the book’s style; the chapter prints the shorter version.
run(goal, policy):
record = open_record(status = "running") # written before the first call
history = [standing_instructions, goal]
repeat:
limit = first_limit_reached(record, policy) # passes, tokens, wall clock, money, repeats
if limit:
return finish("stopped_budget", reason = limit)
response = model(history) # under a per-call timeout
append response to history
if response is a final answer:
if response.text is empty or malformed:
append "final answer missing or malformed" to history
continue # not an exit
if policy.check is none:
return finish("done_claimed") # testimony, labeled as such
verdict = policy.check() # a program; the model has no vote
if verdict.passed:
return finish("done_verified", evidence = verdict.output)
append verdict.output to history # not an exit: go again
continue
result = execute(response.tool_call) # under a per-call timeout
append result to history
if result is a halt-class error:
return finish("halted_error", reason = result.error_class)
# every other error is now in the history, where the model can read it
Three details carry weight. The budget test runs before every model call, so an exhausted run never pays for one more. The check is called by the harness and is never a tool the model may choose to skip. And finish is one function, so every exit writes the same report.
The book’s own skeleton ends with return "stopped: step budget exhausted", which is the right size for an index card. In a service, that sentence should travel in a status field.
What should each exit report to the caller?
Every exit should return the same envelope: a status, a reason, the counts, a pointer to the transcript and the partial work. The status tells the caller which branch to take. The other fields are for the person who has to decide what happens next.
The failure this prevents has a well-known shape. In a 2023 issue on one framework’s tracker, a user reports getting “Agent stopped due to iteration limit or time limit” (in their words, “as the error”) while needing “some particular output from the model.” A stop reason delivered where the answer belongs looks like an answer to any caller that does not parse it.
Some frameworks now raise a named error when the cap fires. One vendor’s SDK documentation, for example, says it raises an exception when the turn limit is exceeded, and one graph framework’s documentation has an error page for the case where a graph “reached the maximum number of steps before hitting a stop condition.” An exception separates stopped from done. It carries the partial work only if you attach it.
| Field | done_verified |
done_claimed |
stopped_budget |
halted_error |
|---|---|---|---|---|
| Reason | Name of the check | “no check” | Which limit, its value and the amount used | Error class and the call that raised it |
| Answer | The model’s answer | The model’s answer, marked unverified | None, or a wrap-up marked as testimony | None |
| Evidence | The check’s output | None | The last check output, if a check ran | The error as the tool returned it |
| Partial work | Not applicable | Not applicable | What changed and what did not | What changed before the halt |
| Counts and transcript | Always | Always | Always | Always |
| Caller’s next move | Use the result | Have a person or a later check read it | Read the transcript, then retry with a larger budget or fix the cause | Stop and page the owner; do not retry blindly |
Chapter 3 is explicit about the third column: “Fail loudly. Save the full transcript, surface whatever partial work exists, and say plainly that the run was stopped rather than finished.” It also says what silence costs: “a budget that fires and silently returns nothing converts a recoverable situation into a mystery.”
Should the model get a last word when the budget fires?
It may, as long as the status does not change. A 2024 issue shows a user asking for exactly this: an option the user’s own code comment describes as a “final pass to generate an output if max iterations is reached.”
A wrap-up call with tools switched off, paid from a small allowance reserved for it, gives a reader a summary of where things stand. It is still the model describing unfinished work. Return it in the envelope under a label that says so, next to stopped_budget.
How do you pick the number for a budget?
Pick it the way timeouts are picked: decide what share of good runs you accept cutting off, then read the matching percentile from runs that a check verified. A cap on passes is a timeout measured in passes.
The chapter defines the budget: “A budget is a hard ceiling enforced by the harness, indifferent to the model’s opinion: on passes through the loop, on tokens spent, on wall-clock time, on money.” On the value it offers “ten or twenty passes for a focused task, a figure I offer as illustration rather than doctrine,” and says the right one is “an empirical matter, read from real traces once you have them.” It stops there, and Chapter 15 takes up the traces. The method below is mine.
It borrows from service engineering. Writing about remote calls in the Amazon Builders’ Library, Marc Brooker describes the practice: “we choose an acceptable rate of false timeouts (such as 0.1%). Then, we look at the corresponding latency percentile on the downstream service (p99.9 in this example).” He adds a pitfall that applies here too: where the high percentile sits close to the median, “adding some padding” avoids a small shift causing many timeouts.
Agent runs give the method something to grip.
In the SWE-agent paper (Yang et al., 2024), a section headed “Agents succeed quickly and fail slowly” reports, for runs that did not exhaust their budget, that resolved instances finished at a median of 12 steps and a median cost of $1.21, against a mean of 21 steps and $2.52 for unresolved ones. The budget was $4 per instance.
It also reports that 93.0% of resolved instances were submitted before the budget ran out, against 69.0% of all instances. Those figures describe one agent, one model and one benchmark, at 2024 prices.
The authors drew the conclusion a budget-setter needs: “we suspect that increasing the maximum budget or token limit are unlikely to substantially increase performance.” Good runs cluster early. A cap placed just past the cluster costs few of them.
The procedure has five steps.
- Record runs of the real task under a generous provisional cap, with a check deciding which succeeded.
- Confirm the provisional cap is not hiding anything: few or no successes should have needed all of it.
- Take the verified successes only, and compute percentiles for passes, tokens and seconds.
- Choose a false-stop rate, read the matching percentile, multiply by a padding factor and round up.
- Replay the recorded runs against the new caps and count who gets cut.
Step 3 needs the first exit. With no check, “successful” means the model said so, and the distribution you measure includes the runs that declared victory early.
A worked example: caps from 400 recorded runs
This is a worked example with illustrative numbers. The 400 runs are synthetic, generated by a seeded script for one focused task with a check wired in, and recorded under a provisional cap of 60 passes. Nothing below measures a real system.
Of the 400 runs, 310 passed the check (77.5%). Another 53 got stuck repeating one identical call, 26 wandered through different calls without ever passing, and 11 ended on a halt-class error. No successful run needed all 60 passes, so the provisional cap was not truncating the successes.
| Verified successes (310 runs) | p50 | p90 | p95 | p99 | Max |
|---|---|---|---|---|---|
| Passes | 8 | 13 | 14 | 21 | 33 |
| Tokens, summed over passes | 49,082 | 114,645 | 169,062 | 376,035 | 738,978 |
| Wall clock, seconds | 69 | 123 | 150 | 211 | 254 |
Accept a false-stop rate of about 1%, read the 99th percentile, pad by 1.2 and round up: passes to the next whole pass, tokens to the next 10,000, seconds to the next 30. That gives 26 passes (21 × 1.2 = 25.2), 460,000 tokens (376,035 × 1.2 = 451,242) and 270 seconds (211 × 1.2 = 253.2).
Look at the token row before moving on. The median successful run used 49,082 tokens and the 99th percentile used 376,035, almost eight times as many, while passes rose from 8 to 21. Every pass resends the whole history, so tokens grow faster than passes. A token cap derived from the pass cap by multiplication would be wrong.
What happens when the recorded runs are replayed?
Replayed under the three caps, 307 of the 310 successes still finish, and 3 are cut off: 2 by the pass cap and 1 by the token cap. That is 0.97% of good runs, close to the rate chosen. The longest success took 33 passes, and this policy stops it at 26.
Total spend falls from 198.4 million tokens under the provisional cap to 53.6 million, a drop of 73.0%. The reason is in the baseline: 89.2% of the provisional spend went to the 90 runs that never passed. The percentage depends on the arbitrary 60, so read it as a direction.
Adding a repeat rule (stop when the same tool, arguments and result appear three times in a row) changes the stuck runs only. All 53 now stop on the repeat rule, at a median of pass 12 where the caps alone stopped them at a median of pass 26. No success trips it, because the longest identical streak in any successful run was two. Total spend falls to 36.7 million tokens, 81.5% below the baseline.
| Final status under the derived policy | Runs | Share of 400 |
|---|---|---|
done_verified |
307 | 76.75% |
stopped_budget: repeats |
53 | 13.25% |
stopped_budget: passes |
25 | 6.25% |
stopped_budget: tokens |
3 | 0.75% |
stopped_budget: wall clock |
1 | 0.25% |
halted_error |
11 | 2.75% |
Two results support the chapter’s remark that “Mature agents run several at once, because the failure each catches is slightly different.” All three token stops fired while the run was still under its pass cap, which is the chapter’s case of “a single pathological tool result” in miniature. And the wandering runs passed the repeat rule untouched, since each call differed from the last. The pass cap caught 23 of the 26.
What did the budgets not fix?
They fixed the bill and left the failures in place: 82 of 400 runs (20.5%) still end as stopped_budget, 79 of them failures and 3 the good runs cut off. The chapter’s sentence for this is “A cap that trips rarely is insurance working as intended. A cap that trips often is a diagnosis.”
One run in five is often. The chapter lists what the diagnosis can be: “the task is bigger than you budgeted, a tool is failing in a way the model cannot route around, or the model is wandering.” Here the reason field already sorts the cases: 53 repeats point at a tool or at the history, and 25 pass-cap stops point at the task or the model. Working out which, for a loop that is already stuck, is the subject of the post on what to do when an AI agent keeps looping.
The example has limits. A 99th percentile from 310 runs rests on about three of them, so with fewer runs use a lower percentile and more padding. Recompute after any change to the model, the tools or the prompt, since each one moves the distribution.
Which errors should stop the run, and which go back to the model?
Halt on errors that are evidence about the run, and feed back the ones the model can act on. Chapter 3 states the test as a question “you can apply mechanically: is this an error the model can act on, or evidence that the run itself is off the rails? Feed back the first kind; halt on the second.”
Feeding back is the default. A failing tool should return its failure as text, the chapter says, “because models recover from well-described errors with surprising competence.” Its examples are “the missing file, the malformed argument, the empty search.” None of these is an exit. Each costs one pass, which the budget counts.
The halting list is short: “A permission denied on a resource the agent should never have approached, a guardrail trip, a spending limit nearly reached.” For these, the chapter says to stop “rather than hand the model another chance to be creative.” A model shown “permission denied” may look for another route to the same resource.
| The harness sees | Class | What it does |
|---|---|---|
| Missing file, bad argument, empty search result | The model can act on it | Append as a result; continue |
| Timeout on a read | Transient | Retry a fixed number of times per call, counted against a run-wide retry allowance; then append the failure |
| Permission denied outside the task’s scope | Evidence about the run | halted_error |
| Guardrail trip | Evidence about the run | halted_error |
| Account or tenant spending limit near | Evidence about the run | halted_error |
| Timeout on a write that carries no idempotency key | Ambiguous | halted_error; a person checks whether it happened |
| Run-wide retry allowance used up | Evidence about the environment | halted_error |
The first, third, fourth and fifth rows are the chapter’s. The other three are my additions from the failure classes of Chapter 18, which the chapter points to for “the whole discipline of failing well.” The ambiguous write is the subject of idempotency keys, and the retry storm simulator shows why the retry allowance has to cover the whole run.
One overlap needs a rule. Money appears in the chapter both as a budget and, as “a spending limit nearly reached,” as a halting error. I read the first as the run’s own ceiling, which this run’s counter enforces, and the second as a limit outside the run that a gateway or tool reports. The first returns stopped_budget and the second halted_error, because the second says something is wrong beyond this run.
What goes on a stop-policy card?
One page per agent and task type holds all the stop conditions for agent loops of that kind: the check, the limits and where each number came from, the halting errors, the report fields and the alerts. The card below carries the worked example’s numbers. Replace every one of them.
STOP POLICY: <agent> / <task type> measured on <date>, from <n> verified runs
(the numbers are this post's illustrative example; replace every one)
EXIT 1 DONE
check: <the program that pronounces the run finished, or "none">
check passes -> done_verified carry the check's output
no check exists -> done_claimed carry the answer, marked unverified
check fails -> not an exit append the check's output, continue
empty answer -> not an exit append the complaint, continue
EXIT 2 BUDGET -> stopped_budget, reason = the limit that fired
passes: 26 p99 of verified successes (21) x 1.2
tokens: 460,000 p99 (376,035) x 1.2, rounded up
wall clock: 270 s p99 (211 s) x 1.2, rounded up
money: <token cap at the current price, or a fixed amount>
repeats: 3 identical (tool, arguments, result) in a row
exempt: <tools that are meant to poll>
per-call timeout: <seconds> on the model call and on every tool
checked: before every model call, in the order listed
scope: the whole run, including sub-agents and retries
wrap-up call: <yes/no>; tools off; the status stays stopped_budget
EXIT 3 ERROR -> halted_error, reason = the error class
halt on: permission denied outside the task's scope
guardrail trip
account or tenant spending limit near
timeout on a write with no idempotency key
run-wide retry allowance used up
feed back: everything else, as a result the model can read
EVERY REPORT CARRIES
status, reason, passes, tokens, seconds, transcript id,
answer (if any), check output (if any), partial work
ALERTS
stopped_budget above <x>% of runs in <window>
any halted_error for permission or guardrail
a record still "running" past the wall-clock cap: treat as halted_error, reason lost
The last alert covers a case that none of the three exits does. If the process dies mid-pass, no exit fires and the caller receives nothing. A record opened before the first call, plus a caller that treats an overdue “running” as a halt, turns that silence into a status.
The per-call timeout is on the card for a similar reason. A wall-clock budget tested between passes cannot interrupt a tool that hangs. The timeout turns the hang into an error, and the error table decides what follows.
How do you audit a loop you already have?
Run each exit on purpose and read what the caller gets. The list below is the card turned into tests; most items need a stub model and no real one.
- A stub model that requests the same tool forever ends in
stopped_budget, with the limit named. - A stub tool that returns one very large result trips the token limit while the run is under its pass limit.
- A stub tool that hangs is cut by a per-call timeout, and the run still ends inside its wall-clock limit.
- The caller can tell all four statuses apart by a field, without parsing the answer text.
- A final answer with a failing check does not end the run, and the check’s output appears in the history.
- An empty or malformed final answer does not return a done status.
- A run that hits a limit returns the transcript id and the partial work, including what was changed.
- A permission error outside the task’s scope halts the run, and the model gets no further call.
- Each limit has a recorded source: which runs, which percentile, what padding, what date.
- The limits cover sub-agents and retries, so nested work cannot spend outside them.
- Killing the process mid-run leaves a record the caller can find, and an overdue record raises an alert.
- Someone is alerted when the share of stopped runs passes a threshold, and reads the transcripts.
To practice the judgment calls first, the run-the-loop simulator puts you in the harness’s seat for a scripted run. For the money row, the agent cost-per-task estimator converts a token cap into a price at your current rates.
Where does this approach fall short?
It falls short wherever the first exit is weak, wherever the task has no stable distribution, and against a failure it was never meant to catch. Each deserves a sentence of honesty.
No check, no clean distribution. Without a program to verify success, the percentile method measures runs that claimed success. You can still set caps from them. They will be caps on how long the agent takes to say it is done.
Open-ended tasks. A research task with no natural length has a wide, flat distribution, and any percentile cuts real work. A budget the caller chooses per request, with the same four statuses, serves better than one number for the task type.
The repeat rule is narrow. It catches the identical call with the identical result. A model that alternates between two calls walks past it, and a tool meant to poll trips it unless exempted. The March 2026 issue on progress-aware termination states the gap in general form: step limits “only cap total steps. They do not detect stuck states.” A fuller progress test belongs to the outer loop.
A cap is blunt. A 2026 feature request on another framework’s tracker calls a step counter “blind” and says it “kills valid 15-step workflows at step 11.” That is a cap set by guess. A measured cap still cuts some good runs, 3 of 310 in the example, and the report is what makes that recoverable.
Stopping late is not harmless. By the time a cap fires, the run has acted on its own mistakes for some passes, which is the mechanism behind compounding errors in AI agents. A budget limits how far that goes. It does not undo a write.
An operator’s stop is outside the three. Chapter 3 does not count a person or an orchestrator cancelling the run. Give it its own reason under halted_error, and build the switch outside the agent’s process.
The exit to build first
Build the budget first, because the chapter says so (“Set the cap before you write anything else”) and because it is the only exit that guarantees an end. Then build the report, since a cap with no report hides what happened. Add the check next, and the caps can be measured from verified runs. Sort the errors last.
A 2024 engineering essay describes the common practice as including “stopping conditions (such as a maximum number of iterations) to maintain control.” A 2025 study of multi-agent traces (Cemri et al.) suggests how often stopping goes wrong in both directions: among the failures it annotated, 12.4% were agents unaware of termination conditions and 6.2% were premature terminations. Stop conditions for agent loops are a small amount of code, and the report they return is what lets anyone verify the run.
Chapter 3, The Agent Loop (in the full book) has the full argument, including the framework question that follows it. The agent fundamentals guide places this post beside its neighbors and the agent loop explainer, or you can see the formats.
Questions readers ask
- What are the stop conditions for an agent loop?
- Chapter 3 of the book names three exits. The run is done, pronounced by a check a program runs where one exists and by the model's own final answer where none does. A budget on passes, tokens, wall-clock time or money fires, enforced by the harness. Or an error that the model cannot act on halts the run. The budget is the only one that is mandatory.
- How many steps should an agent loop be allowed to take?
- Enough for nearly every run that would have succeeded, and no more. Record runs under a generous provisional cap, keep the ones a check verified, pick the share of good runs you accept cutting off, and read the matching percentile of their pass counts. Add padding. The book offers ten or twenty passes for a focused task as an illustration, not a rule.
- What should an agent return when it hits its step limit?
- A result with a status that says it was stopped, the name of the limit that fired, the counts, a pointer to the full transcript, and whatever partial work exists. The book's wording is to fail loudly and say plainly that the run was stopped rather than finished. A sentence in the answer field is not enough, because a caller cannot branch on it.
- How do I stop an agent from repeating the same tool call?
- Count identical calls in the harness. Fingerprint each pass as the tool, its arguments and its result, and stop the run when the same fingerprint appears several times in a row; three is this post's illustrative threshold. Then read the transcript. The cause Chapter 3 names is a tool result that never reached the history, so the model, blind to the outcome, asks again.
- Should the model get one last call to summarize when a budget fires?
- It can, with tools switched off and inside a small reserved allowance, and it is useful for a person reading the report. The status must stay stopped. The summary is the model's account of unfinished work, so it is testimony and should be labeled that way.
Sources
- Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents
- John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, Ofir Press (2024). SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (arXiv:2405.15793, version 3)
- Marc Brooker (Amazon Builders' Library) (undated; read October 2026). Timeouts, retries, and backoff with jitter
- Mert Cemri et al. (2025). Why Do Multi-Agent LLM Systems Fail? (arXiv:2503.13657, version 3)
- LangChain issue tracker (2023). Agent stopped due to iteration limit or time limit (issue 8493)
- LangChain issue tracker (2026). Progress-aware termination: detect no-progress loops in agent tool execution (issue 36139)
- CrewAI issue tracker (2026). Native deterministic guardrail to prevent infinite agent delegation and tool loops (issue 6414)
- OpenAI (undated; read October 2026). Running agents (Agents SDK documentation)
- LangChain (undated; read October 2026). GRAPH_RECURSION_LIMIT (LangGraph documentation)