An agent loop is the cycle at the center of every AI agent: the model reads the run’s history, chooses one next action, your code executes it and appends the result, and the cycle repeats until a stop condition fires. So, what is an agent loop in practice? About ten lines of ordinary code, and the reason an agent can work through a task it was never scripted for.
I want to make that definition precise enough to use. By the end of this post you should be able to read any agent’s run, whatever framework produced it, as four beats around a growing history, and when a run misbehaves, to name which beat failed. The material comes from Chapter 3 of AI Agents, Engineered, “The Agent Loop” (in the full book), and pairs with the three-minute agent loop explainer.
What is an agent loop, exactly?
An agent loop is a bounded repeat in which a language model supplies one decision per pass and plain code supplies everything else: the tool execution, the memory and the stop button. The people who build agents for a living describe them with a shrug. Anthropic’s Erik Schluntz and Barry Zhang wrote that agents “are typically just LLMs using tools based on environmental feedback in a loop” (Schluntz and Zhang, 2024).
Thorsten Ball, building a working coding agent in a widely read walkthrough, put it more bluntly: “It’s an LLM, a loop, and enough tokens” (Ball, 2025). Both are right, and both leave out the part that decides whether the loop is safe to run, which is how it stops. Here is the whole machine in pseudocode, the same shape the book writes on its index card:
history = [standing_instructions, user_goal]
repeat up to MAX_STEPS times:
response = model(history) # reason over everything so far
if response is a final answer:
return response.text # exit: the model judges it done
result = execute(response.tool_call) # act: your code runs the tool
append response.tool_call to history # record what the model did
append result to history # record what the world said back
return "stopped: step budget exhausted" # exit: the safety net
Notice what is missing. There is no planning module, no branch that reads “if the tests fail, then…”, no strategy anywhere in the code. The book compresses the point into one line: “you wrote the loop; the model writes the path.”
What are the four beats of one pass?
Every pass through an agent loop has four beats: observe (assemble the instructions, goal, history and latest result), reason (one model call that decides the next move), act (the decision arrives as a structured tool call or a final answer) and result (your code runs the tool and appends whatever came back, failures included).
Hold onto the division of labor. The model owns exactly one beat out of four. The other three are yours, which means three of the four places a run can go wrong are in code you wrote and can test.
| Beat | Who does it | What happens | Typical failure |
|---|---|---|---|
| Observe | Your code | Lay the instructions, goal and full history in front of the model | Stale or bloated context; the key fact buried or missing |
| Reason | The model | Weigh where things stand; pick a tool and arguments, or declare done | Wrong tool, invented argument, premature “done” |
| Act | Your code | Validate and execute the requested tool call | Unsafe side effect, unvalidated arguments, a retry that doubles an action |
| Result | Your code | Capture the output or the error and append it to the history | Silent failure, truncated output, the append that never happens |
People underrate the fourth beat. As the book observes about its own trace, the error messages were “the most valuable things the agent read all afternoon”: a loop’s diet is feedback, especially bad news. A tool that swallows its error and returns an empty string starves the loop of the one thing it runs on.
Why is the history the agent’s only memory?
The history is the agent’s only memory because the model is stateless: each call sees exactly what you send and remembers nothing afterward. Continuity lives entirely in the message history, the growing list of goal, model actions and tool results that your code lays back in front of the model on every single pass.
The book’s phrasing is worth quoting whole: “whatever your code appends is the agent’s memory, and whatever it fails to append never happened, in the strictest sense available.” That makes the two append lines in the pseudocode load-bearing. Drop the tool result and the model, blind to the outcome, requests the same call again. Drop the model’s own action and it loses track of what it has already tried.
If you ever debug an agent that calls the same search five times in a row, start there. The book names the dropped append as the failure a step cap guards against, “the same call requested forever,” and it is a bug in your code, not a mystery inside the model. (The glossary entry for message history has the short version.)
What does a short agent loop run look like?
A short run makes the four beats concrete: each pass adds one action and one result to the history, and the model re-reads all of it before choosing the next move. The trace below is my own illustration, not a recorded run; the job, file names and error text are invented to show the shape.
The goal: “Why did last night’s invoice export job fail?” The standing instructions list three read-only tools: search_logs, read_file and run_query. The history starts with two items.
| Pass | Items on the desk | Model decides | Your code appends |
|---|---|---|---|
| 1 | 2 | search_logs(job="invoice-export", since="last night") |
A stack trace: column tax_region does not exist, raised from exports/invoices.sql |
| 2 | 4 | read_file("exports/invoices.sql") |
The query text, which selects tax_region |
| 3 | 6 | run_query("describe invoices") |
The schema: the column is now region_tax_code |
| 4 | 8 | read_file("migrations/latest.sql") |
A migration, applied yesterday, that renamed the column |
| 5 | 10 | Final answer | The diagnosis and a one-line fix to the query |
Look at what drove the run. Nothing in the loop knew to check the schema after reading the query; that step came out of the model on pass 3, conditioned on the six items before it. The path emerged one decision at a time, which is exactly why the same loop can, on another afternoon, wander.
Now count the reading. The history ends at ten items, but across five passes the model read 2 + 4 + 6 + 8 + 10 = 30 item-loads, because every pass re-reads everything. The stored history grows linearly with the number of steps while the total reading grows roughly with its square. That is the arithmetic behind the cost of long runs, and behind the book’s warning that a long-running loop “is always drifting toward the overflowing desk.”
A five-pass diagnosis is cheap. A fifty-pass run is a different object, and not only because of cost: each unverified step is another chance to go wrong, and those chances multiply. The compounding error calculator lets you put numbers on that for your own step count.
Where does ReAct fit in the agent loop?
ReAct (short for Reason + Act) is one way to write the reasoning beat: the model produces a brief written thought, then one action, then reads the observation before thinking again. It was introduced by Shunyu Yao and colleagues at Princeton and Google in a paper first posted in 2022 and published at ICLR 2023 (Yao et al., 2022).
The paper’s account of why interleaving works explains both halves at once: “reasoning traces help the model induce, track, and update action plans as well as handle exceptions, while actions allow it to interface with external sources, such as knowledge bases or environments, to gather additional information.” On two interactive benchmarks, ALFWorld and WebShop, prompted ReAct beat imitation and reinforcement learning methods “by an absolute success rate of 34% and 10% respectively, while being prompted with only one or two in-context examples.” Those figures belong to the models and benchmarks of that paper and should be read as evidence for the idea, not as a current performance number.
Read the result honestly, too. On question answering and fact verification the same paper found the best approach overall was a combination of ReAct and plain chain-of-thought reasoning, and the book points out that interleaving has permanent costs: every action is a full model call, and a model reasoning one step ahead can wander on long tasks. The durable idea is grounding, reasoning kept tied to fresh evidence at every step. In the book’s words, “Grounding, it turns out, is a schedule.”
The original recipe parsed literal Thought:, Action: and Observation: lines out of the model’s text. Current models are typically trained to reason before acting and to emit actions as structured tool calls, so you will rarely type that format yourself. As Chapter 3 puts it: “The format was a costume; the interleaving is the idea.” More on the trade-offs in the glossary entry on ReAct.
How does an agent loop stop?
An agent loop stops through one of three exits, and the book ranks them by how far each can be trusted: a verified check that a program can run (evidence), the model’s own judged “done” (testimony), and a capped budget on steps, tokens, time or money (the safety net that guarantees the loop ends).
The model’s judgment cannot be the only door out, because it fails in both directions. An agent can announce a fix complete while two tests still fail, in the same assured tone as a true report, or it can keep polishing forever while the bill runs. The book’s rule: “a loop whose only exit is the model’s judgment is a loop with no guaranteed exit at all.”
| Exit | Who decides | What it proves | Use it when |
|---|---|---|---|
| Verified | A program: tests pass, file parses, response validates | That the work meets a checkable condition | Any time such a check exists; let it pronounce the run finished |
| Judged | The model returns a final answer | That the model believes it is done | The goal has no programmatic check; have a person or a second check read the result |
| Capped | The harness enforces a ceiling on passes, tokens, time or money | Nothing about the work, only that the loop ended | Always, set before anything else is written |
Go back to the invoice trace and ask which exit it used. It was judged: no program can confirm that a diagnosis is correct, which is why a human should read it. Change the goal to “fix the export” and a verified exit appears, because the harness can run the export against staging and read the exit code. The book’s line for this is the one I would put on a sticky note: “The model’s “done” is testimony; a green test is evidence.”
The cap is not optional. “Set the cap before you write anything else,” the chapter says, offering “ten or twenty passes for a focused task” as an illustration rather than doctrine.
When a cap fires, fail loudly: save the transcript, surface partial work and say the run was stopped, not finished. Errors form a further family the book handles separately, with one mechanical test: “is this an error the model can act on, or evidence that the run itself is off the rails? Feed back the first kind; halt on the second.” The companion post on stop conditions for agent loops works through budgets and error halts in detail.
What is the harness, and which part is yours?
The harness is everything you build around the model to turn it into a working agent: the loop, the tool implementations and what they may touch, the care of the history, the stop rules, and later the logging, recovery and guardrails. The book sums it up: “An agent, in one line, is a model plus a harness.”
Any answer to what is an agent loop is incomplete without this word, because it changes where you look when something breaks. The model is the judgment you rent; the harness is the part you own and can test. Since the loop itself is about ten lines, the difference between a demo and a dependable agent lives almost entirely in the harness, which is the argument of the companion post on what an agent harness is and what you actually build.
Lilian Weng’s widely cited overview describes the model as functioning “as the agent’s brain, complemented by several key components” such as planning, memory and tool use (Weng, 2023). I find the loop view more useful for debugging, because it tells you the order in which those components touch each run. Frameworks fit the same picture: agent frameworks (LangGraph and smolagents are two examples of the category) wrap these four beats in persistence, retries and tracing, and none of that changes what the model does on a single pass.
Which beat failed? A diagnostic map
Most agent misbehavior traces back to one of the four beats, and naming the beat narrows the fix. The map below is my own synthesis of the chapter’s argument, written as symptoms you would see in a trace.
| Symptom in the trace | Beat | First thing to check |
|---|---|---|
| Same tool call repeated with the same arguments | Result | Is the result being appended? Does the tool fail silently? |
| Model ignores an instruction from the start of the run | Observe | Is the history so long that the instruction is buried? |
| Plausible but wrong tool or argument | Reason | Is the tool’s description clear about when to use it? |
| An action happened twice after a retry | Act | Is the tool safe to repeat? See idempotent tools and safe retries |
| Confident “done” with the work unfinished | Exit | Is there a verified check, or only the model’s word? |
| Run hits the cap every time | Exit | Read the saved transcript: task too big, tool failing, or wandering? |
When is an agent loop the wrong tool?
An agent loop is the wrong tool when your code, not the model, should decide the next step: when the path is known in advance, a fixed workflow with a model inside is cheaper, faster and easier to test. The loop earns its keep only when the number and order of steps genuinely depend on what the run discovers.
It also has limits even where it fits. Every action costs a full model call that re-reads the whole history, so long runs pay in latency and tokens. One-step-ahead reasoning can wander on long tasks, which is why plan-first designs exist and why the book spends its next chapter on planning. And the loop has no way to know it is right except the checks you give it; where none exists, its “done” stays testimony.
Here is the takeaway I would keep from all of this. If a colleague asks you what is an agent loop, the short answer is four beats, a growing history and three exits. The loop is simple, and the simplicity is the point: once you can see those parts, every agent run you read becomes legible, and every failure has an address.
The full treatment, including the minimal agent built from four parts and the framework question, is in Chapter 3 of the book (in the full book); Chapter 1 and Chapter 2, which set up the vocabulary, are free to read online. If you want the book itself, the guide to choosing between the Kindle and paperback editions compares the formats.
Questions readers ask
- What is an agent loop in simple terms?
- An agent loop is a short program that calls a language model, runs whatever tool the model asks for, appends the result to a running history, and calls the model again. It repeats until a check confirms the goal is met, the model declares it is done, or a budget on steps, tokens, time or money runs out.
- Is the agent loop the same thing as ReAct?
- Not quite. The agent loop is the control structure: call, act, append, repeat. ReAct (Reason + Act, Yao et al., 2022) is a way of writing the reasoning step inside that loop, alternating a brief thought with each action. Current models usually reason internally and emit structured tool calls, so the interleaving survives while the old Thought/Action text format has mostly disappeared.
- How many steps should an agent loop be allowed to take?
- There is no universal number. Set a hard cap before anything else, start modest for a focused task (the book offers ten or twenty passes as an illustration, not a rule), then tune it from real traces. A cap that fires rarely is insurance; a cap that fires often is a diagnosis worth reading.
- Why does my agent keep calling the same tool over and over?
- The most common cause is a result that never reaches the history: the tool ran, but its output was not appended, so the model, seeing no outcome, asks again. Check the append path first, then check whether the tool fails silently or returns an error the model cannot act on.
- Do I need a framework to build an agent loop?
- No. The core loop fits in about ten lines of ordinary code. Frameworks add plumbing around it, such as persistence, retries, tracing and multi-agent orchestration, which is worth adopting when you hit a concrete need you would otherwise build yourself.
Sources
- Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao (2022). ReAct: Synergizing Reasoning and Acting in Language Models
- Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents
- Thorsten Ball (2025). How to Build an Agent
- Lilian Weng (2023). LLM Powered Autonomous Agents