Home / Tools / Run the loop

Free tool · runs in your browser · from Chapter 3

Run the loop

An agent loop simulator: play the harness for a scripted run, choose append, retry, stop or escalate on each pass, and see which failures you cause.

The tool

Your inputs stay in this tab. Share a result by copying the page address: the state lives in the URL. The rules, the quotes and the failures described are the book's (Chapters 3, 12, 15 and 18). The scripted scenario, the four-button menu, two failure names (never done, a duplicate side effect), the points, the bands and the pass costs are the tool's own.

What is an agent loop simulator?

An agent loop simulator is an exercise in which you play the code around a language model for one scripted run: on each pass the model asks for something, a result comes back, and you decide whether to append it, retry the call, stop the run or hand it to a person. The model’s part is fixed in advance. Every decision that remains is yours, which is the point.

Chapter 3 of the book describes the loop as four beats (observe, reason, act, observe the result) and then gives the division of labor in a sentence meant to be carried out of the chapter: “the model supplies the judgment—which tool, with what arguments, and whether the job is done—and your code supplies everything else: the hands, the memory, and the stop button.” The game above gives you the hands, the memory and the stop button for seven passes, and nothing else. If you want the idea at essay length first, what is an agent loop covers it; this page is the practice ground.

The premise is the book’s own. Chapter 15 points out that stopping logic is ordinary software and can be tested without a model at all: “script a model stub that returns”done” too early, or never, and assert your loop does the right thing”. The simulator is that stub, with you standing where the assertions would be.

What does the harness do on each pass?

On each pass the harness runs what the model asked for, captures whatever came back, appends it to the history, and decides whether the loop goes round again. Chapter 3 defines the harness as “everything you build around the model to turn it into a working agent”, and it lists the loop, the tools, the care of the message history and the stop rules as its parts.

The simulator reduces that work to four moves. Three of them map onto distinctions the book draws; the menu itself does not appear in the book.

Move What it means here Where the book draws the line
Append and continue Put what came back into the history and call the model again Chapter 3: the appends are “load-bearing”
Retry the call Run the same call again without telling the model Chapter 18: “retry only the transient class”
Stop the run End the loop, saving the transcript Chapter 3: budgets, checks and loud halts
Escalate to a person Pause and hand over a packaged decision Chapters 12 and 18: escalation triggers and the handoff
The loop has three ways out, and they are not equally trustworthy.
Figure 3.5 The loop has three ways out, and they are not equally trustworthy. A verified check—a suite that goes green, a file that parses—is evidence, and where one exists it should pronounce the run finished (in accent). The model’s own “done” is only testimony, believed with caution. And the budget cap is the safety net: it settles nothing about the work, but it guarantees the loop stops, so it must always be there. Reuse this diagram

The scenario is a support agent refunding a duplicate charge and confirming to the customer. It is the tool’s own composite, built so that one run contains a read, a correctable error, a transient timeout, a repeated empty result, a write that goes quiet, and two claims of being finished.

Why does a dropped result make the agent repeat itself?

A dropped result makes the agent repeat itself because the model remembers nothing between passes: the history your code maintains is its only memory, so a result that never reaches the history is a result the model never saw. Chapter 3 puts it as strongly as it can: “whatever your code appends is the agent’s memory, and whatever it fails to append never happened, in the strictest sense available.”

The consequence follows mechanically. “Drop the tool result and the model, blind to the outcome, requests the same call again”. On the first pass of the game, choosing Retry on a read that succeeded does exactly this: the result is thrown away, and the model asks again. In a real system the cause is usually a code path that forwards only the last message, or a truncation that cuts the result out. The agent bug bestiary files it as the dropped append.

The same rule covers bad news. An error is a result, and the model can only correct what it can read. On the second pass the tool rejects a malformed order id and says so. The right move is to append the error, because Chapter 3’s dividing question has a clear answer here: “is this an error the model can act on, or evidence that the run itself is off the rails? Feed back the first kind; halt on the second.”

When should a harness retry, and when is retrying dangerous?

A harness should retry only failures that are transient, and it should treat a timeout on a write as a different thing altogether. Chapter 18 sorts failures by what they want done about them and states the first rule plainly: “retry only the transient class, since retrying a bad credential or a model’s malformed output burns money on a certainty”.

The game sets up the contrast with two timeouts. On pass three a read goes quiet, and Retry is the best move, because reading twice changes nothing in the world. On pass five a refund goes quiet, and the same button causes the worst outcome in the run. The chapter draws the line in one sentence: “a timeout on a read is transient, retry it; a timeout on a write is the ambiguous case—the work may have happened”.

That ambiguity is what produces duplicate side effects. “Charge a card, hear nothing back, and”just try again” becomes a question of whether your customer gets billed twice.” The book’s fix is idempotency: a key derived from the action’s logical identity, so a repeat returns the stored result instead of doing the work again. The scenario deliberately gives you a harness without keys, which leaves the loud options. Chapter 18’s rule for deeds that fail is short: “Fall back on words; fail loud on deeds.” The post on idempotent tools and safe retries builds the keyed version.

How do you stop a loop that will not stop, or one that stops too soon?

You stop a loop with exits your code owns: a hard cap on passes, a counter for repeated calls, and a completion check that a program can run. Chapter 3 gives the reason the cap cannot be optional: “a loop whose only exit is the model’s judgment is a loop with no guaranteed exit at all.”

Pass four shows the loop that will not stop. The model has asked for the same search three times and received an empty list each time. Chapter 15 describes the cause (“the model, seeing no reason to believe retrying is hopeless, retries”) and the backstop: “count repeated calls in your loop, and after the third identical attempt inject a message that says so and asks the model to change course or report the obstacle.” Appending the empty result once more lets the pattern run until the budget fires. That firing is the budget doing its job. Chapter 3 wrote the line about the dropped append, and it holds for any loop that repeats: “Without a cap, that bug is a bill with no ceiling; with one, it is a log entry.”

Pass six shows the opposite failure. The model announces that the refund is issued and the customer told, while the history holds no sent message. Chapter 15 names this false completion: “the agent that stops too soon, returning a half-finished answer delivered as if it were complete”. The defense is a stop condition the model does not control. “Where such a check exists, wire it into the loop and let that check pronounce the run finished.” In the game the best move is to append the failed check and continue, so the model reads what is missing and finishes the job.

When is escalating to a person the right move?

Escalating is the right move when the run is wedged or when a deed has become uncertain, and it is the wrong move when the model or a retry could resolve the problem for free. Chapter 12 gives a short list of triggers, beginning with “Repeated failure on the same step”, and Chapter 18 places a person at the bottom of every fallback ladder.

Two passes in the game reward it. After the third empty search, a person answers the policy question in a line. After the refund times out, a person checks the payment system’s own record and finds the refund went through once. Chapter 18 describes what good escalation buys: “escalation stops being the agent’s failure state and becomes its most valuable move, the one that converts a wedged run into a fifteen-second human decision.”

The game also charges for overuse. Paging someone about a malformed argument or a network blip scores lower than letting the loop handle it, because human attention is the scarcest resource in the system. The consequence tier classifier is the tool for deciding ahead of time which actions should wait for a person.

What does the game simplify?

The game simplifies three things, and each is marked in the interface. First, the append is not a choice in the book. Both of Chapter 3’s listings append on every pass; the dropped append is a bug in code, which the game turns into a button so the cost is visible. A real harness also appends a timeout or an empty result before it halts or escalates. The game makes you pick one move, so on those passes “Append and continue” is wrong because of the continuing, not the appending.

Second, one Retry button stands for two different actors. A harness retries mechanically, and a model re-requests a call after reading an error. Chapter 18 keeps them apart, noting that failures the model can fix “pointedly do not want a blind mechanical retry”. Third, on pass four the book’s first remedy is an injected message telling the model to change course, which the game folds into Escalate.

The scoring is entirely the tool’s. The book assigns no points, and its own numbers for budgets are offered as illustration. A score here tells you which rules you applied on seven scripted passes, and nothing about a harness you have built. For that you need traces of real runs, which is where the failure modes of AI agents and the bestiary pick up. The agent fundamentals guide places the loop next to the engine and the tools, what is an agent harness covers the rest of the code around the model, and the full argument is in Chapter 3, The Agent Loop (in the full book).

Questions readers ask

What is an agent loop, in one sentence?
An agent loop is ordinary code that calls a language model, runs the tool the model asks for, appends the result to a growing history, and repeats until something stops it. Chapter 3 of the book splits the work this way: the model supplies the judgment, and your code supplies the hands, the memory and the stop button.
What does the harness decide that the model does not?
Everything outside the judgment: what gets appended to the history, which failures are retried, when the run ends, and when a person is asked. The model proposes a tool call or says it is finished. Your code decides what to do with that, and the simulator puts you in the position of that code.
Why is retrying a failed tool call sometimes the wrong move?
Because only one class of failure wants a retry. A timeout on a read is transient and safe to repeat. A malformed argument fails the same way every time, so it should go back to the model as a result. A timeout on a write is ambiguous, since the work may have happened, and repeating it can refund, charge or send twice.
How do I stop an agent from saying it is done when it is not?
Do not let the model's statement end the run. Where a program can check the goal, such as a passing test, a receipt or a sent message in the history, wire that check into the loop and let it pronounce the run finished. The book's line is that the model's statement is testimony and a green test is evidence.
Is the four-button menu how real harnesses work?
No. In a real harness these are code paths, not choices made on each pass, and the append is unconditional. The menu is a teaching device that lets you feel each decision. The page marks the points where the game simplifies the book.

Sources

  1. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao (2022). ReAct: Synergizing Reasoning and Acting in Language Models
  2. Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents
  3. Malcolm Featonby (Amazon Builders' Library). Making retries safe with idempotent APIs
  4. Birgitta Böckeler (2026). Harness engineering for coding agent users