When an AI agent keeps looping, the transcript shows one of five kinds: the same call repeated, a short cycle of calls, a retry on an error the model cannot fix, drift that never reaches “done,” or agents handing a task back and forth. Each kind has its own cause and its own exit. A budget caps the cost of all five.
This is a troubleshooting checklist for the backend engineer whose agent is looping now, or was last night. It covers what to do first, how to tell the kinds apart from a few lines of transcript, the fix for each, and what each guard would have saved. The design of the exits themselves is a separate post, on stop conditions for agent loops, and this one uses its statuses without re-deriving them. The book material is Chapter 3, “Stop Conditions, Budgets, and Frameworks,” and Chapter 18 (both in the full book), “Surface Errors Back to the Model” and “Retries, Timeouts, and Budgets.”
What do you do in the first ten minutes when an AI agent keeps looping?
Stop the run from outside, keep the evidence, and find out what it changed. That order matters because the two instinctive moves, restarting the task and editing the prompt, destroy the transcript or repeat the damage.
- Cancel from outside the agent’s process. A model that is looping will not decide to stop, and a stop message appended to its history is one more input it may ignore. Kill the worker, or flip the switch your orchestrator provides.
- Save the full transcript before anything restarts. The message history and the tool results are the only record of which kind of loop this was. Chapter 3’s instruction for a budget that fires applies to a manual stop too: “Save the full transcript, surface whatever partial work exists, and say plainly that the run was stopped rather than finished”.
- List the writes. A loop over a read tool costs tokens. A loop over a write tool sent the email, posted the message or created the record once per pass. Count the duplicates before you look at anything else, because they may need undoing while the cause can wait.
- Do not resume on the same history. The history contains the loop, and a model shown twelve identical calls has a strong example of what comes next.
- Set a low provisional cap on passes before the task runs again, even a guessed one. Chapter 3 gives the reason in one line: “Without a cap, that bug is a bill with no ceiling; with one, it is a log entry.”
The rest of this post is the diagnosis of why the AI agent keeps looping, and it starts from the saved transcript.
Which kind of loop is in your transcript?
Read the last ten or so passes and compare three things between them: the call (tool and arguments), the result, and whether the result was an error. The combination that repeats names the kind. The five-way split is this post’s own arrangement; the causes trace to the two chapters and to the reports quoted below.
| Kind | What repeats in the transcript | Usual cause | Exit | Transcript shows |
|---|---|---|---|---|
| Same call repeated | Identical tool, arguments and result, several passes in a row | The result never reached the history, or reached it in a form the model cannot use | Repair the append or the result; a repeat rule stops the run | identical calls |
| Short cycle (ping-pong) | A block of two to four passes, then the same block again | Two requirements that cannot both hold, or two actions that undo each other | Remove the conflict; give the model a stated way to report it | alternating calls |
| Retry on an error it cannot fix | The same tool returning the same error; arguments may change a little | A permanent or policy failure fed back as if the model could repair it | Classify errors; permanent ones fail fast, policy ones halt the run | errors |
| “Almost done” drift | Nothing identical; new text every pass; the task’s check never improves, or there is no check | No definition of done that a program can test | Wire a check; without one, only a budget ends it | new text each pass |
| Sub-agent re-delegation | The same brief handed from one agent to another and back | No agent owns the task; a worker cannot do it and cannot say so | Count hops; share one budget; give workers a “cannot do” result | hand-offs, alternating calls |
If two rows seem to fit, take the more specific one. An identical call that returns the same error three times is a retry on an error, and the error is the lead to follow; a cycle made only of hand-offs is re-delegation.
These kinds are common enough to have names in the research literature. Cemri et al. (2025) annotated failures in the traces of multi-agent systems and labeled 15.7% of them step repetition, defined as “Unnecessary reiteration of previously completed steps in a process,” and 12.4% as not recognizing task completion. Those are shares of the failures the authors annotated in multi-agent systems, and they say nothing about how often a single agent loops.
What causes each kind of loop, and what is the exit?
When an AI agent keeps looping, each kind points to a different layer: the harness for the first, the task definition for the second and fourth, the error handling for the third, and the orchestration for the fifth. Only the fourth is mainly about the model’s judgment.
Why does the agent repeat the same call when the call succeeded?
Because, from where the model sits, the call has not been answered. One open-source agent’s tracker has the plain form of the report: the agent is “repeatedly executing the same bash command even though it receives successful results each time.”
Check three things in this order. First, whether the result is in the history at all. Chapter 3 calls this bug the dropped append, and its symptom is the same call requested forever. Second, whether the result arrived whole: a result truncated to nothing, or replaced by a placeholder when the context was trimmed, leaves the question open. Third, whether the result answers the question. A lookup that returns an empty list, with no instruction on what an empty list means, invites another try.
The exit is in the harness or the tool, and the prompt is not involved. Fix the append, return a result that states what it found (“0 orders for customer 88” beats []), and add the repeat rule described below so the next bug of this shape ends in three passes.
What is a ping-pong loop?
A ping-pong loop is a short cycle: the agent does A, then B, then A again, with each step sensible alone. A familiar form is an edit and a test that undo each other. Setting a value one way fails one test, setting it the other way fails another, and the agent alternates.
The cause is a pair of requirements that cannot both hold, in the instructions or in the checks. A model has no native way to say so and keeps trying to satisfy the most recent complaint. The exit is to remove the conflict once you have found it, and to give the agent a stated outcome for the general case: a final answer of the form “these two requirements conflict, here is the evidence,” which the caller treats as a result.
The repeat rule misses this kind entirely, because no two consecutive passes match. It needs its own rule, which compares blocks of passes.
Why does the agent retry an error it cannot fix?
Because the harness handed it the error as if it were fixable. Chapter 18 sorts failures by recovery path, and two of its classes must never go back to the model for another try. Permanent failures “want to fail fast; retrying a bad credential wastes money on a certainty.” Policy failures, such as a guardrail trip, want a loud halt.
A report from September 2026 shows what happens when a policy block is fed back instead. The agent, unable to edit a file, “enters a persistent loop trying alternative workarounds (spawning multiple subagents, attempting to write python scripts to /tmp, trying sed -i, etc.).” That is the loop Chapter 3 warns of when it says the safe response to such errors is to halt “rather than hand the model another chance to be creative.”
The transcript signature has a variant worth knowing. The arguments often change slightly from pass to pass while the error stays the same, so a rule that compares whole calls misses it. Compare the tool and the error code. Chapter 18’s dividing sentence is the fix in general form: “the model can only recover from a failure it can see.” Give every tool error a stable code and a retryable flag, feed back the ones the model can act on, and end the run on the others. The build order for that is in how to make AI agents more reliable.
One neighbor of this loop lives below the model. If the HTTP client, the tool wrapper and the loop each retry, Chapter 18’s arithmetic applies: “Four modest policies multiply into dozens of calls per failure”. The transcript shows one slow pass; the tool’s own logs show the storm. The retry storm simulator lets you count it.
Why does the agent never say it is done?
Because nothing it can test tells it so. Chapter 3 describes the behavior: the model can “fail to declare victory at all, finding one more thing to polish, then another, while the bill runs.” The transcript shows new text on every pass, so no comparison of calls finds anything.
Two versions exist and they have different exits. Where the task has a check a program can run (tests, a schema, a count), the drift shows up as a score that stopped improving: the failing-test count goes three, two, two, three, two. Wire the check into the loop and stop when its best value has not improved for several passes. Where no check exists, the drift is invisible to any rule. A commenter on Hacker News described the surface as “a jitter loop (rephrasing the same query slightly each time)” and “stalling with apologetic non-progress text.” The exit there is a budget, plus work on the task: Chapter 3’s rule is “The model’s ‘done’ is testimony; a green test is evidence,” and a task with no evidence available will keep producing this loop.
What is a re-delegation loop between agents?
It is the ping-pong loop with agents in place of tools. A feature request on one framework’s tracker names it “The Delegation Ping-Pong”: “Agent A delegates a task to Agent B, but Agent B gets confused and delegates it back to Agent A.”
The causes are structural. No agent owns the task, so each can pass it on. A worker that cannot do the job has no result type for “I cannot,” so it hands off. And each agent often runs under its own step limit, which means a system of three agents with a cap of 20 each has no cap of 20 on the task.
The exit has three parts. Every hand-off carries a hop count, and the count has a ceiling. The run’s budgets are shared, so a sub-agent spends from the parent’s remaining allowance. And a worker may return “cannot do, because,” which the parent must treat as a result and may not re-delegate unchanged.
How can a harness detect that an AI agent keeps looping?
With four rules that compare recent passes and need no model call. Record a fingerprint for each pass: the agent, the tool, the arguments in a canonical form, a hash of the result and the error code if there was one. Then, after every pass, test these in order.
- Error retry. The last three results are the same error code from the same tool.
- Repeat. The last three fingerprints are identical, and the tool is not one that is meant to poll.
- Cycle. The last block of two, three or four passes equals the block before it, and the block holds at least two different fingerprints. If every pass in the block is a hand-off, call it re-delegation.
- No progress. A progress number has not improved for six passes, not counting passes by a tool that is meant to poll. Use the check’s score where a check exists; otherwise count the distinct results seen so far in the run.
The thresholds are illustrative, and three for the repeat rule matches the stop-conditions post. A rule that fires ends the run with the status that post calls stopped_budget, with the rule as the reason (no_progress: cycle, say), the fingerprints that matched and the transcript id.
As a worked example, here is what a short script implementing the four rules returns for ten synthetic transcripts. The transcripts are made up for the purpose, one per kind plus five that test the edges.
| Transcript | Passes | Rule fires at pass | Kind reported |
|---|---|---|---|
| A. Same order lookup, same result, four times after two other calls | 6 | 5 | repeat |
| B. Customer lookup and order lookup alternating after one other call | 7 | 5 | cycle |
| C. Write refused with “permission denied,” path varied each time | 5 | 4 | error retry |
| D. Ten distinct edits; failing tests go 3, 2, 2, 3, 2, 2, 3, 2, 2, 2 | 10 | 8 | no progress |
| E. Two agents handing one brief back and forth | 5 | 4 | re-delegation |
| F. Healthy run that ends with a passing test | 5 | never | none |
| G. A job-status tool polled four times while the job runs | 7 | never | none (tool exempt) |
| H. Search rephrased eight times, same empty result | 9 | 8 | no progress |
| J. Edit, test, other edit, test, repeated as a block of four | 9 | 9 | cycle |
| I. Fourteen passes of new polishing edits, no check | 14 | never | none |
Three rows deserve a second look. In C the arguments differ on every pass, which is why rule 1 ignores them. In H the queries differ and the results are errors to nobody, so only the progress count sees it. And I is the miss: with new text every pass and no check, no rule fires, and a cap on passes is the only thing that ends the run.
That last row is why detection does not replace budgets. The March 2026 request for progress-aware termination on one framework’s tracker puts the complementary point: limits on steps “only cap total steps. They do not detect stuck states.” You need both. The rules stop a stuck run early, and the cap stops the run the rules cannot see.
What would a budget have saved on the run that looped?
A cap on passes bounds the bill, and the bound is looser than it looks, because the cost of a pass grows as the history grows. Each pass re-reads everything before it, so input tokens rise with the square of the pass count. A looping run is long, and it pays for its own length on every pass.
A worked example, with illustrative numbers. An agent has a fixed prefix of 3,000 tokens (instructions and tool definitions) and each pass adds 1,000 tokens to the history: 300 written by the model and 700 of tool result. A healthy run takes 12 passes. A broken run behaves the same for 8 passes and then repeats one call for as long as it is allowed. Input tokens for n passes are 3,000 × n + 1,000 × n(n − 1)/2.
| How the run ends | Passes paid for | Input tokens | Against the healthy run |
|---|---|---|---|
| Healthy run, for reference | 12 | 102,000 | 1.00× |
| Repeat rule fires (passes 9, 10 and 11 identical) | 11 | 88,000 | 0.86× |
| Cap of 25 passes | 25 | 375,000 | 3.68× |
| Cap of 60 passes | 60 | 1,950,000 | 19.12× |
Sixty passes is 5 times the healthy run’s 12 and costs 19.12 times its input tokens. Against the run the repeat rule stopped, the cap of 25 costs 4.26 times as much and the cap of 60 costs 22.16 times as much. At 60 passes, 90.8% of the input is the model re-reading earlier passes, most of them identical. The estimator below opens on the 60-pass run; move the steps to 25, 12 and 11 to reproduce the other rows.
With JavaScript on, the Agent cost-per-task estimator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Agent cost-per-task estimator on its own page to share a result by link.
The same arithmetic answers a reader whose AI agent keeps looping only in the sense of running long, and who arrived asking why AI agents are too expensive or why an agent is so slow. Wall-clock time scales with passes, and tokens scale faster than passes. A run does not need to be stuck to be costly; the healthy 12 passes above already spend 64.7% of their input on re-reading. That case has its own levers, which the planned post on how to reduce AI agent costs ranks, and it should not be confused with a loop. The test is the status: a run that finished late is slow, and a run that was stopped was stuck.
Passes are one budget of four. Chapter 3 lists “passes through the loop, on tokens spent, on wall-clock time, on money” and explains why several run at once: “a loop can sit under its step cap while a single pathological tool result blows the token budget.” A report from July 2026 shows a second way a pass cap alone falls short. After one user message, the agent “ran 45 model turns automatically” and those turns produced “1498 total tool calls.” A cap on model turns counts 45 there, and a cap on tool calls is the one that would have fired early. Choosing the values is the subject of the stop-conditions post, which sets them from the pass counts of runs a check verified.
The checklist: diagnose the loop, then cap the next one
Work the list top to bottom against the run that looped; it applies whenever an AI agent keeps looping, whichever kind it turns out to be. The first part is about that run, the second fixes its cause, and the third is what the harness should have had.
- The run was cancelled from outside the agent’s process, and nothing restarted it on the same history.
- The full transcript is saved, with every tool result as the model received it.
- Every write the run made is listed, and duplicates are undone or queued for a person.
- The last ten passes are classified as one of the five kinds, by comparing calls, results and error codes.
- Same call repeated: the result is confirmed present, whole and readable in the history that the next model call received.
- Short cycle: the two requirements or actions that conflict are named, and the agent has a stated way to report a conflict as its answer.
- Error retry: the error has a class; permanent errors end the attempt, policy errors halt the run, and only errors the model can act on are fed back.
- Drift: the task has a check a program runs, or the task is marked as having none and relies on budgets.
- Re-delegation: hand-offs carry a hop count with a ceiling, and a worker can return “cannot do.”
- A cap on passes, tokens, wall-clock time and tool calls applies to the whole run, sub-agents and retries included.
- The four no-progress rules run after every pass, with polling tools exempt from the repeat and no-progress rules, and bounded by their own timeout.
- Every model call and tool call has its own timeout.
- A stopped run returns a status a caller can branch on, the limit or rule that fired, the counts and the transcript id.
- The looping transcript is kept as a test: replayed against the harness, it ends at the pass the rule predicts.
- Someone is alerted when the share of stopped runs rises, and reads the transcripts.
The fourteenth item is the one that turns an incident into a guard. A stub model that replays the recorded calls costs nothing to run, and the run-the-loop simulator is a way to practice the harness’s decisions on a scripted run first. The general procedure for a failing run, of which this is one branch, is in the post on how to debug an AI agent.
Can someone make an agent loop on purpose?
Yes, wherever the agent reads text an outsider can write. A tool result is input to the next model call, and a web page, a ticket or a document can contain instructions to fetch one more page, try again or hand the task to another agent. The mechanism is the one behind indirect prompt injection, aimed at your bill instead of your data.
This is an argument for where the guards live, and it is this post’s own point, drawn from no measurement. A rule in the prompt can be argued with by anything else in the context. A counter in the harness cannot be reached by text. Budgets and the four rules hold whether the loop came from a bug, a confused model or a hostile page, and they need no knowledge of which.
Where does this checklist fall short?
It falls short on tasks with no check, on loops that are not exact, and on everything a stopped run leaves behind.
Drift without a check is only capped. Row I above ends at the budget, having spent all of it. If that is your common case, the work is on the task: find something a program can test, even a weak proxy.
Fingerprints are exact, and loops are often approximate. A search rephrased each pass with slightly different results defeats rules 1 to 3, and rule 4 catches it only if the results repeat. Comparing results by similarity is possible and adds a threshold to tune and a new way to stop a good run.
The rules can stop a healthy run. A tool that legitimately polls must be exempt. A long task that revisits a file trips nothing here, but a tighter threshold would catch it. The feature request that named the delegation ping-pong also complains that a blind step counter “kills valid 15-step workflows at step 11.” Every rule and cap should be replayed against recorded good runs before it ships.
Stopping is late for writes. By the third identical pass, a write tool has acted three times. The repeat rule limits the count. Only an idempotency key on the write makes the repeats harmless.
A stop is not a fix. Chapter 3 says “A cap that trips often is a diagnosis.” A dashboard that shows stopped runs falling after you raised the cap has measured the cap.
The one guard to add tonight
Add the cap on passes and tool calls tonight, with a status that says the run was stopped and a saved trace. That bounds every kind of loop in this post, including the one no rule can see. Add the four rules next, since they turn a bill of 375,000 input tokens into one of 88,000 in the example above. Then fix the cause the transcript named.
A 2024 engineering essay describes the practice in one clause: “it’s also common to include stopping conditions (such as a maximum number of iterations) to maintain control.” When an AI agent keeps looping, control means the code around the model ends the run, and the report it returns is what lets you verify why.
Chapter 3, The Agent Loop (in the full book) builds the exits and Chapter 18, Reliability, State, and the Harness (in the full book) builds the error classes, retries and timeouts under them. The guide to agent security, reliability and cost places this post beside its neighbors and the agent loop explainer, or you can see the formats.
Questions readers ask
- Why does my AI agent keep looping?
- Because nothing outside the model ends the run, and something inside the run keeps the model from finishing. The transcript says which: the same call repeated (the result is missing or unusable), a short cycle (two requirements that conflict), a repeated error (a failure the model cannot fix), drift (no checkable definition of done), or hand-offs between agents (no owner).
- How do I stop an AI agent from looping right now?
- Cancel the run from outside the agent's process, save the full transcript before anything restarts it, and list the writes the run made so duplicates can be undone. Then set a low provisional cap on passes before running it again. Do not restart the same task on the same history, because the history contains the loop.
- Does a max-steps limit fix a looping agent?
- It ends the run, which is the part that protects the bill, and it repairs nothing. The same task will loop again to the same limit. A step limit also fires late: it counts passes and cannot tell a stuck run from a long one. Pair it with rules that detect no progress, and read the transcript of every stopped run.
- Will telling the model not to repeat itself stop the loop?
- Not reliably, and it is the wrong layer. In four of the five kinds the model repeats because of something the harness or a tool did: a result that never reached the history, an error with no usable description, a check that cannot pass. An instruction does not change those. A counter in the harness works whatever the model does.
- Why is my AI agent slow and expensive when it is not looping?
- Because each pass re-reads the whole history, so input tokens grow with the square of the number of passes. In this post's illustrative run, a healthy 12 passes already spend 64.7% of their input tokens re-reading earlier passes. A long run that makes progress is a cost problem with different levers: fewer passes, smaller tool results, a cached prefix.
Sources
- Mert Cemri et al. (2025). Why Do Multi-Agent LLM Systems Fail? (arXiv:2503.13657, version 3)
- Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents
- LangChain issue tracker (2026). Progress-aware termination: detect no-progress loops in agent tool execution (issue 36139)
- CrewAI issue tracker (2026). Native deterministic guardrail to prevent infinite agent delegation and tool loops (issue 6414)
- mistral-vibe issue tracker (2025). Agent gets stuck in infinite loop repeating the same tool call (issue 83)
- OpenClaw issue tracker (2026). Agent tool call infinite loop + model.completed usage always 0 (issue 108056)
- jules-ai-agent issue tracker (2026). Agent enters infinite retry loop when editing tools are missing or blocked by policy (issue 8)
- aura-guard (Hacker News) (2026). Comment on 'Ask HN: Best practices for AI agent safety and privacy'