Modeling an AI agent as a state machine means writing the loop around the model as named states and legal transitions in ordinary code: find work, assign it, check the result, record what happened, decide what is next. The model chooses the path inside one state. Your code owns every transition, including the one into done.
That last transition is where most of this post is spent. After it you can write your agent’s outer loop as a state table with four distinct exits, state “done” as a predicate that a program other than the worker evaluates, and estimate how fast a lenient check spoils the work the loop builds on. Six pages on this topic that I read in October 2026 all model the task or the conversation as states, and none shows what a too-kind completion check costs.
What does it mean to treat an AI agent as a state machine?
It means separating two loops and making the outer one explicit. The inner loop is one run of the agent, the model choosing tool calls until a stop condition fires, and what an agent loop is covers it; the run-the-loop tool lets you step through one. The outer loop is the code that decides what the task is, launches that run, checks what came back and chooses the next move. A backend engineer has usually written the second kind already: order lifecycles, payment states, job queues.
AI Agents, Engineered uses the term for an agent’s loop shape in Chapter 18 (in the full book), as “a fixed set of named states” with defined transitions. Its account of the gain is the reason to bother: “What the state machine buys is legibility: illegal transitions cannot happen, loops-forever becomes a state you can see and bound, debugging gains a map, and the current state is a compact, honest thing to checkpoint.” The same paragraph says where production designs land: “a deterministic skeleton of states with model judgment doing the deciding inside each one”.
Chapter 13 (in the full book), the chapter on the outer loop, never calls that loop a state machine and prints no pseudocode for it. The framing in this post is mine, built on that chapter’s verbs.
What are the outer loop’s five verbs?
The book gives the outer loop five verbs: find, assign, check, record, decide. Chapter 13 says that under any outer loop, from a one-line shell loop to a phase framework, “the same five verbs are underneath: find work, assign it to an agent, check the result, record what happened, decide what is next”.
A trigger starts the cycle: a schedule, an event, or a message from another system. Find turns the sources into candidate work, and the chapter is blunt about it: “Nothing here requires intelligence.” Assign gives one task a fresh context window and “a brief that would survive a stranger”; the orchestrator worker pattern post covers that brief, so I add nothing about it here.
Record is the memory: “A loop that records only successes has a diary; a loop that records failures, dead ends, and the reason a task was abandoned has a memory the next pass can actually use.” Decide reads the record: “Done means halt; blocked means escalate to a human; otherwise pick the next item and go again.” Around the cycle the chapter adds brakes: budgets, stuck detection, human gates and a kill switch.
The outer loop as a state-transition table
The table below has eight working states, four terminal ones and 25 legal transitions; the restart row near the end stands for seven of them. It is this post’s construction: the state names, the counters and the rule order are mine, and the activity in each state comes from Chapter 13’s verbs and brakes. The numeric thresholds are illustrative placeholders that you set.
| State | What happens there | Event or check, in ordinary code | Next state | Written before the next state is entered |
|---|---|---|---|---|
IDLE |
Nothing runs; the loop waits | A schedule, an event or a message arrives, and the kill switch is off | FINDING |
Run id, trigger, start time, this run’s budgets |
FINDING |
Code surveys the sources and the record, and lists the workable items (open, and not parked for a person) | The list has at least one item; one is picked by a fixed rule | ASSIGNING |
The list, with a stable id per item, and the pick |
FINDING |
Same | The list is empty | DECIDING |
The empty list |
ASSIGNING |
Code builds the brief (item, spec, last recorded fault) for a fresh window and an isolated workspace | Brief complete | RUNNING |
The brief verbatim, item marked in flight, attempt number |
ASSIGNING |
Same | A required field of the brief needs a person’s decision, such as an item with no acceptance check | RECORDING |
The missing decision, as the reason to park the item |
RUNNING |
Code launches the worker as the first act, once the record names RUNNING. The model chooses the path; the outer loop only meters |
The worker process returns a result | CHECKING |
Where the artifact lives, the cost of the pass, receipts for any side effect, and the worker’s message stored as a note |
RUNNING |
Same | A per-pass cap fires (steps, tokens, wall clock), the launch fails, or a restart finds this state with no result | RECORDING |
Verdict “no result” and its cause |
CHECKING |
The oracle runs on the artifact: mechanical checks first, then a read-only verifier. The worker’s “done” is not an input | Verdict fail, the check does not return within its time limit, or a second restart finds it unfinished | RECORDING |
Verdict fail, the check that failed or the cause, its output whole |
CHECKING |
Same | Verdict pass, and publishing this item is ungated | RECORDING |
Verdict pass and its evidence: which checks ran, what each returned |
CHECKING |
Same | Verdict pass on the mechanical checks, and a person must approve before publishing: the item is gated (a protected branch, production, an external message) or part of its goal has no machine check | AWAITING_HUMAN |
Verdict and evidence, a fingerprint of exactly what awaits approval, wait start time |
AWAITING_HUMAN |
The item is parked at the gate. No process has to stay alive; one that sleeps here writes a clean-stop line first, and one that stays alive is counted on every host restart | An approval arrives for that fingerprint | RECORDING |
Who approved, when, and the fingerprint |
AWAITING_HUMAN |
Same | A denial arrives, or the wait’s time limit passes | RECORDING |
The denial with its reason, or the timeout; the item is marked for parking |
RECORDING |
Code appends the outcome: closed, failed with the fault, or parked with the reason. A closed item is published under an idempotency key. Counters move: passes used (failed ones included), attempts on this item, passes since the last close. An item at its attempt cap is parked | The store acknowledges the append | DECIDING |
The append: item, attempt, verdict, evidence pointer, receipt, time, counters |
DECIDING |
Code reads the record, runs the goal oracle and applies the next five rules in this order | The goal predicate is true | DONE |
Reason “goal met” and the oracle’s output |
DECIDING |
Same | No workable item is left | ESCALATED |
Reason “needs a person” and the parked items |
DECIDING |
Same | Passes since the last closed item have reached the stuck threshold (illustrative: 5) | STUCK |
Reason “no progress” and the failing checks of those passes |
DECIDING |
Same | A loop budget is spent: passes, tokens, money or wall clock | OUT_OF_BUDGET |
Which cap fired, and each clause of the goal at that moment |
DECIDING |
Same | None of the four rules above fired | FINDING |
Reason “go again” |
FINDING, ASSIGNING, RUNNING, CHECKING, AWAITING_HUMAN, RECORDING, DECIDING |
After a crash, the restarted process finds this state named in the record and adds one to its restart count before anything else | It is the third restart in this state (illustrative limit: 2) | STUCK |
Reason “restarts”, with the state’s name |
| Any state | An operator sets the kill switch | The process sees the switch and stops | No transition; the record keeps naming the same state | A clean-stop line, “stopped by operator”, so the next start is not counted as a crash |
DONE, ESCALATED, STUCK, OUT_OF_BUDGET |
Terminal for this run. The report is written and nothing moves | None; the next trigger opens a new run with a new id | None | The final report, with the reason code |
A plain-text version to paste into a design document:
OUTER LOOP: states and legal transitions
(illustrative thresholds: 3 attempts per item, stuck after 5 passes with nothing closed,
2 restarts allowed per state)
IDLE trigger arrives, kill switch off -> FINDING write: run id, trigger, budgets
FINDING a workable item exists; pick one -> ASSIGNING write: candidate list, the pick
FINDING no workable item -> DECIDING write: empty list
ASSIGNING brief complete; launch comes next -> RUNNING write: brief, in flight, attempt no.
ASSIGNING brief needs a person's decision -> RECORDING write: reason to park
RUNNING worker returns a result -> CHECKING write: artifact, cost, receipts, note
RUNNING cap fires, launch fails, or no result -> RECORDING write: verdict "no result", cause
CHECKING fail, timeout, or unfinished twice -> RECORDING write: verdict, failing check, output
CHECKING verdict pass, item ungated -> RECORDING write: verdict, evidence
CHECKING pass, and a person must approve -> AWAITING_HUMAN write: evidence, fingerprint, start
AWAITING_HUMAN approval for that fingerprint -> RECORDING write: who, when, fingerprint
AWAITING_HUMAN denial, or time limit passes -> RECORDING write: denial or timeout, park item
RECORDING store acknowledges the append -> DECIDING write: outcome, counters, receipt
DECIDING 1. goal predicate true -> DONE write: reason, oracle output
DECIDING 2. no workable item left -> ESCALATED write: reason, parked items
DECIDING 3. passes since last close >= threshold -> STUCK write: reason, failing checks
DECIDING 4. a loop budget is spent -> OUT_OF_BUDGET write: which cap, goal clauses
DECIDING 5. none of the above -> FINDING write: "go again"
FINDING, ASSIGNING, RUNNING, CHECKING, AWAITING_HUMAN, RECORDING, DECIDING
third restart in the same state -> STUCK write: reason "restarts", the state
Terminal for the run: DONE, ESCALATED, STUCK, OUT_OF_BUDGET. No exits.
Not in the table on purpose: RUNNING -> DONE. The worker's "done" is a note.
The worker is launched inside RUNNING, after the record names RUNNING. A restart never relaunches.
The kill switch stops the process in any state and writes a clean-stop line;
the record says where to resume, and only the first start after a clean stop goes uncounted.
How do you read the table?
Read it for the edges that are missing, because the value of drawing an AI agent as a state machine lies in what the table forbids. No row leads from RUNNING to DONE, so the worker’s final message cannot end the run; it is stored as a note and nothing reads it as a verdict. A terminal state is reached in two ways only: a rule in DECIDING, which reads the record and the oracle, or the restart limit.
Two further properties are deliberate. Every worked pass decrements the budget in RECORDING, whatever its verdict; a bug report from September 2026 describes the opposite, a step limit that counted only successes: “So a failed run is never counted”. And a check that errors or times out is a fail.
The inner loop is one opaque state. The agent harness is the code around the model inside RUNNING; this table is the layer above it.
The kill switch is the one stop outside the transitions: it halts the process in any state, and the record says where to resume. The table assumes one outer-loop process per run. Parallel workers each get their own row in the record and their own pass through these states, which makes the table the first thing I would write before deciding how to build a multi-agent system.
What does one run look like, traced through the table?
A run is a path through the table that ends in a terminal state. Here is the path for the case this post cares about most, a worker that reports success while a test still fails. The item number, the check name and the clock time are invented.
state event or check written
IDLE 06:00 schedule fires; kill switch off run 41 opened, budgets set
FINDING one workable item: #88 list [#88], pick #88
ASSIGNING brief complete brief, #88 in flight, attempt 1
RUNNING launched; returns "Done. All tests pass." artifact, cost; message kept as a note
CHECKING suite: 1 failed (test_export_empty) verdict FAIL, the test output whole
RECORDING append acknowledged #88 failed; passes 1; since last close 1
DECIDING goal false; 1 workable; 1 < 5; budget left "go again"
FINDING one workable item: #88 list [#88], pick #88
ASSIGNING brief complete, with the failing test brief, #88 in flight, attempt 2
RUNNING launched; worker returns a result artifact, cost
CHECKING suite: 0 failed; lint clean; verifier pass verdict PASS, evidence
RECORDING append acknowledged #88 closed, published, receipt
DECIDING goal predicate true DONE, reason "goal met"
The false report cost one pass, because the report had no edge to follow.
I traced three other runs the same way. They were a clean success, a crash in the middle of an item, which launches nothing on restart, and a loop that closes nothing for five passes and ends in STUCK.
How does the loop know it is done? The goal function
The loop knows it is done when a program, run after every pass, evaluates a stated goal as true. The worker’s word is the thing to distrust. A commenter wrote in December 2024 of two open-source agent frameworks that they “often times will say a task is completed, but it actually has not done the action”.
The book’s glossary names the pair of goal and check the goal function: “a persistent goal held outside the model (the spec, the checklist, on disk, because the agent forgets between runs) paired with a machine-checkable stopping condition evaluated after each pass”.
Its evaluator is the oracle, which the glossary defines as “a source of truth outside the thing being judged, which pronounces an answer right or wrong”. Chapter 13 makes the consequence a rule, “A loop is exactly as trustworthy as its oracle.” It then separates the checker from the maker, so that “the author of the work holds no vote on whether the work is finished”.
The predicate below is this post’s pseudocode, in the book’s style; the book prints none for the outer loop.
# The goal lives on disk, outside the model, and holds still for the run.
GOAL: every item in the plan is closed
and the test suite passes on the main branch
and the linter reports nothing
# The oracle: a program the worker did not write and cannot edit.
goal_is_met(record):
return no item in record.plan is open or parked
and run(test_suite).failed is 0 # measured now, by code
and run(linter).warnings is 0
# Decide reads the record and the oracle. The worker's "done" is not an input.
decide(record, budget):
if goal_is_met(record): return DONE
if no workable item is left in record: return ESCALATED
if record.passes_since_last_close >= STUCK_AFTER: return STUCK # illustrative: 5
if budget.passes_left is 0 or budget.tokens_left <= 0
or budget.money_left <= 0 or budget.time_left <= 0:
return OUT_OF_BUDGET
return GO_AGAIN
A goal that is partly unmeasurable shows up as a clause you cannot write, which is the book’s point in one line: “A goal you can loop toward is a goal whose satisfaction a machine can check.”
Which oracle is strong enough?
An oracle is strong to the degree that the worker cannot influence it and that it measures the world directly. The ladder below is my summary of Chapters 10 and 13, weakest first; the rungs are the book’s and the arrangement is mine.
| Oracle | What it is | Who can fool it | The book’s words |
|---|---|---|---|
| The worker’s own “done” | The model that did the work says it has finished | The worker, without trying | “a poor oracle for that work” |
| A critic model with a rubric | A separate call, ideally a different model, judging written criteria | Anything outside the rubric | “an estimate where the verifier’s is a fact” |
| A checklist | Items ticked off | Whoever ticks, unless each item is itself checked | “an oracle only if its items are themselves checkable” |
| A program that measures the world | Tests, a build, a schema validator, a query against the system of record | Gaps in what it measures | “A test suite is a strong oracle” |
Reach matters as much as the rung. In METR’s June 2025 report on one frontier model, reward hacking appeared in 30.4% of runs on one task suite and 0.7% on another, “more than 43× more common”. The authors suggest a cause with a “perhaps”: on the first suite “the model was able to see the entire scoring function”. That is one model on one lab’s suites in mid-2025.
A 2025 preprint, ImpossibleBench, names the behavior a table can prevent: an agent with access to unit tests “may delete failing tests rather than fix the underlying bug”. In state terms, CHECKING must use assets that RUNNING cannot write.
What if “done” cannot be machine-checked?
Then the loop needs a person or a calibrated judge in the CHECKING state, and you accept a weaker guarantee. Chapter 13 gives the rule as a question to ask of the task, whether “done” is machine-verifiable. If it is, a loop can drive it. If it is not, “wire a human into the loop, or build a judge” and “accept the weaker guarantee”.
The chapter ends that rule with a prohibition: “what you must never do is hand an unverifiable goal to a bare while”. In the table, a person in the checking seat is the AWAITING_HUMAN path: the mechanical checks that do exist still run first, and the person’s approval is the rest of the verdict.
A status check deserves the same suspicion as a model. In a bug report from August 2026, a payment handler treated two unsettled statuses as sent and marked the task complete: “The record says the payment completed; nothing revisits it.” No model was involved. The oracle was lenient by two enum values.
Why does a lenient check compound in the wrong direction?
A lenient check compounds in the wrong direction because the loop keeps what the check accepts and then builds on it. Chapter 10 states the mechanism for the evaluator-optimizer, one of the four agentic workflow patterns it catalogs: “the loop optimizes exactly what the evaluator measures, and nothing else”.
The same chapter draws the conclusion: “A weak critic is worse than none: the loop will optimize something, and you would rather it optimize nothing than optimize the wrong thing while billing you for progress.” The book makes this argument in prose and attaches no arithmetic to it. The example that follows is mine, and every number in it is illustrative: the inputs are invented and round, and nothing was measured.
A nightly loop works a backlog of 40 items. Each pass produces a correct change 70% of the time. The check wrongly rejects good work 5% of the time and passes bad work at a false-pass rate f. A failed item is retried, and after three failed attempts it is parked for a person; I leave the stuck rule and the budget out so the arithmetic stays visible.
Compare a strict check, f = 0.02, with a lenient one, f = 0.50, such as a thin test file or a reviewer that shares the worker’s context.
| Per attempt | Arithmetic | Strict check | Lenient check |
|---|---|---|---|
| Accepted and right | 0.70 × 0.95 | 0.665 | 0.665 |
| Accepted and wrong | 0.30 × f | 0.006 | 0.150 |
| Accepted in total | sum of the two rows above | 0.671 | 0.815 |
| Rejected, r | 1 − accepted | 0.329 | 0.185 |
| Share of accepted work that is wrong | wrong ÷ accepted | 0.006 ÷ 0.671 = 0.9% | 0.150 ÷ 0.815 = 18.4% |
What does the dashboard show?
The dashboard shows the lenient loop winning. Over up to three attempts, an item is closed unless all three are rejected, which happens with probability r³.
| Over up to three attempts | Arithmetic | Strict check | Lenient check |
|---|---|---|---|
| Attempts per item | 1 + r + r² | 1.44 | 1.22 |
| Items closed | 1 − r³ | 96.4% | 99.4% |
| Items parked for a person | r³ | 3.6% | 0.6% |
| Closed and wrong | closed × wrong share | 0.9% | 18.3% |
| Closed and wrong, of 40 items | 40 × the row above | 0.3 | 7.3 |
The lenient loop closes more, escalates less and spends fewer passes: about 49 for the backlog against about 57. Those are the three numbers a dashboard shows, and all three favor the loop that ships seven wrong items a night.
Retries do not rescue it. Each retry is another draw against the same check, so the wrong share of accepted work stays at 18.4% however many attempts you allow.
What happens when the loop builds on its own output?
The wrong share multiplies across accepted items. Each accepted item joins the state the next pass starts from: the repository, the plan, a shared store. If items were independent, the chance that everything accepted so far is right after n items would be (1 − wrong share) to the power n.
That is the multiplication from Chapter 2 (free to read), applied one level up, to accepted items across passes. Why agent errors compound derives it for steps inside one run, so I only apply it here, with the right shares rounded to 0.991 and 0.816. Unrounded, the strict column gives 84% at 20 items and 77 in the last row.
| Accepted items the loop has built on | Strict check: 0.991n | Lenient check: 0.816n |
|---|---|---|
| 5 | 96% | 36% |
| 10 | 91% | 13% |
| 20 | 83% | 1.7% |
| 40 | 70% | 0.03% |
| Most items with the odds still above 50% | 76 | 3 |
The last row is ln 0.5 ÷ ln p, rounded down: 76.7 for the strict check and 3.4 for the lenient one. With the lenient check, by the fourth accepted item the foundation more likely holds a wrong piece than not.
The calculator below opens on the lenient column at 20 accepted items. Read its “step” as “accepted item” and its “per-step success” as the share of accepted items that are right. It shows 1.7%, says each item would have to be right 96.59% of the time to keep even odds over 20 items, and gives 3 as the longest chain that meets a 50% target.
With JavaScript on, the Compounding error calculator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Compounding error calculator on its own page to share a result by link.
What does the example leave out?
It leaves out three things, and two of them make reality worse. Independence is the optimistic assumption: a wrong accepted item raises the error rate of later passes that read it. A worker that optimizes against the check pushes f up over time, and Chapter 13 describes an unwatched loop accumulating stubs unless the gate is strong enough “that a stub fails the gate”. In the other direction, some wrong items are never built on.
A store that several loops share widens the reach, and Engin Diri’s 2026 essay Stop Prompting. Design the Loop. puts the thesis in one sentence: “Without an oracle, the loop compounds confidently wrong work, faster than you can read it.”
How should the loop stop when it is not done?
It should stop through one of three other terminal states, each with its own reason code, so that a finished run and an exhausted run can never be confused. A 2024 engineering essay states the baseline: “it’s also common to include stopping conditions (such as a maximum number of iterations) to maintain control”. A cap alone is one exit; the table has four.
| Terminal state | Fires when | The report holds | Who acts next |
|---|---|---|---|
DONE |
The goal predicate is true on the record | The oracle’s output, the closed items, the evidence | A person reads the report |
ESCALATED |
No workable item is left and the goal is false | Each parked item, its attempts, its last verdict or the denial | A person decides the parked items |
STUCK |
Several passes in a row closed nothing (illustrative: 5), or the process restarted a third time in one state | The failing checks of those passes, or the state’s name | A person inspects the check, the spec or the environment, and stops any worker left alive |
OUT_OF_BUDGET |
A cap on passes, tokens, money or wall clock fired | Which cap, and each clause of the goal at that moment | A person decides whether to fund another run |
Why are stuck and out of budget separate states?
They are separate because they call for different responses. The budget is the floor. Chapter 13 lists “a hard cap on iterations, a token-and-money ceiling, a wall-clock limit”, set per pass and again per loop, and gives the reason: “This is why a hard iteration cap is the floor beneath every goal function: it catches the case where “done” never comes true.”
Stuck detection fires earlier and says more. The chapter’s version is “the same test failing across several consecutive passes means the loop has found a wall, and the correct response is to stop and notify rather than spend the rest of the budget confirming the wall exists”. My counter, passes since the last closed item, catches that case and also a loop that fails a different way each time.
The rule has to be code. A March 2026 issue describes a no-progress rule that lived only in prompt text, with “no programmatic enforcement”. Parked status and attempt counts live in the record and carry into the next run, so an always-failing item does not earn three fresh attempts every night. I would also hold the next trigger after a STUCK run until someone clears the reason, since the chapter warns that “a loop that fails cheaply every night still fails every night”.
Where does the human gate sit?
The human gate sits between a passing check and a consequential action, as the AWAITING_HUMAN state. Chapter 13 places approval gates at “merges to protected branches, anything touching production, any external communication”, and calls them “checkpoints where the loop stops and a person decides”.
At a gate, the work passed and a person authorizes its effect; the approval is bound to a fingerprint of what was shown. An escalation hands over work the loop could not finish. The wait is a saved state, so the process can exit and the approval message can wake it.
How does the loop resume after a crash?
It resumes by reading the newest line of the record, which names the state it was in, and applying a fixed rule for that state. Chapter 13 gives the reason the record exists: “State lives on disk, because the model forgets and the files do not.”
| State named in the newest record | What the restarted process does | Why that is safe |
|---|---|---|
FINDING, DECIDING |
Run the state again | It reads the sources, the record and the oracle, and changes nothing else |
ASSIGNING |
Rebuild the brief from the record | The launch happens only after the record names RUNNING, so no worker exists yet |
RUNNING |
Stop any worker left from this attempt, discard the workspace, look up receipts for any side effect, record “no result” and count the attempt. Never relaunch | The pass is ambiguous, and an isolated workspace makes discarding it free |
CHECKING |
Run the check again on the stored artifact; on a second restart in this state, record a fail with the cause “check did not complete” | The check is read-only; a suite that writes to a shared system needs the same receipts as a worker |
AWAITING_HUMAN |
Keep waiting; the time limit runs from the stored start | The wait is data. A process that slept here left a clean-stop line, so its next start is not counted |
RECORDING |
Repeat the append and the publish under the same keys | Keyed by run, item and attempt, a repeat changes nothing, provided the receiving system honors the key; where it does not, look up the receipt before publishing. A run that ends STUCK in this state may already have published, so look up the receipt before the item is reopened in a later run |
A terminal state, or no run open (IDLE) |
Nothing; wait for the next trigger | A terminal state never changes |
What stops a crash loop, or a second worker?
Two rules do, and both are mine. The worker is launched in one place only, on the normal entry to RUNNING, so no attempt can have two workers. And unclean restarts are counted per state in the record: a third one in the same state ends the run as STUCK with the state’s name.
# First act of every process start. Thresholds are illustrative.
resume(record):
state = record.newest_state
if state is terminal: return state # nothing to redo
if the newest line is a clean-stop line: # kill switch, or asleep at a gate
write a start line # a clean stop excuses one start
else:
record.restarts_in[state] += 1 # written before any work
if record.restarts_in[state] > RESTART_LIMIT: return STUCK # any state; limit: 2
if state is RUNNING:
stop any worker left from this attempt # never launch here
write verdict "no result"; go to RECORDING
if state is CHECKING and record.restarts_in[state] is 2:
write verdict fail, "check did not complete"; go to RECORDING
run state again # AWAITING_HUMAN: keep waiting
RUNNING and CHECKING usually leave sooner, because their own rows move the run on to RECORDING at the first or second restart. The limit is for a restart that dies before it can write that move. A STUCK report that names RUNNING means a worker may still be alive, so stop it before the next run. The count and the stop it triggers go into the record as one write, a state’s count returns to zero when the run leaves it, and a clean-stop line excuses only the next start.
A failed launch is a “no result” pass, and a source that is down is treated as a crash under the same limit. The limit has a floor. A store that is down cannot record anything, and neither can a start that dies before its first write; that alarm belongs to whatever supervises the process.
The record is a checkpoint, and Chapter 18 states its limit: “A checkpoint records position. It does not, and cannot, record whether the outside world did the work.” For a worker that sent an email before dying, idempotency and receipts are the fix, and idempotent tools and safe retries works through that crash.
Kill the process in the middle of a pass and watch the restart before you trust any of this, because, as Chapter 18 puts it, “a checkpoint you have never restored from is a guess”.
Review an existing loop: a checklist
Run these eleven statements against the script or service that drives your agent today, whether or not anyone has drawn that AI agent as a state machine.
- The run cannot end on the model’s final message: a check in code stands between the worker’s result and done.
- The check runs where the worker cannot reach it; tests, scorers and rubrics are read-only to the worker.
- Every status the check counts as success is a settled status, with pending and processing excluded.
- A check that errors or times out is recorded as a fail.
- A failed or crashed pass decrements the pass budget, and attempt counts survive into the next run.
- Done, needs a person, no progress and out of budget are four exits with four reason codes.
- The no-progress rule is a counter in code, and the prompt does not carry it.
- The record is written before the next state is entered, and it includes failures and the reason for abandoned work.
- The worker is launched after the record names the attempt, and a restart never launches a second one.
- Someone has killed the process mid-pass and watched the restart, and a process that keeps dying in the same state stops the run.
- A named person reads every terminal report, and a stuck or exhausted run alerts them.
When is a workflow engine worth it?
A workflow engine is worth it when a single pass is long, has side effects in the middle, or must wait days for a person. Modeling an AI agent as a state machine does not require one. A table and a state file are enough when each pass is short, starts from a fresh window and writes its record before the next one begins.
That second kind of loop already resumes at pass granularity. Chapter 18’s test for adding machinery is two questions: “what here is irreversible, and how long does the run live?” Hand-written machines do run in production; one practitioner reported in October 2025 an agent state machine written without libraries, “Running successfully in prod with very few issues.”
Two categories of tooling exist, and I name two examples of each as they stood in October 2026. Durable-execution runtimes, often sold as workflow engines, implement what the glossary calls durable execution. Temporal and Restate are two, and both resume a run by replaying a recorded history.
Agent graph frameworks are the second category, with LangGraph and Apache Burr as examples. Either category gives you durable waits, timers, retries and an inspectable history, at the price of integration work and rules about how your code is written.
No engine supplies the transition into done. A weak check inside a durable runtime produces false completion that survives restarts.
Where does the state-machine framing stop helping?
Treating an AI agent as a state machine stops helping at the check, at open-ended work and at the edge of my evidence. Four limits are worth stating plainly.
The machine makes control flow legible, and the check decides reliability. A 2025 taxonomy of multi-agent failures, Cemri et al., attributes shares of its annotated failures to “premature termination (FM-3.1, 6.20%), no or incomplete verification (FM-3.2, 8.20%), or incorrect verification (FM-3.3, 9.10%)”. Those are shares of annotated failures in the systems it studied, and the paper warns that “the presence of a verifier is not a silver bullet”.
A deterministic skeleton repeats a bad action faithfully. If the action inside a state is wrong, the table runs it again tomorrow with full audit logs.
Fixed states cost adaptability. The table fits a countable backlog and a checkable goal. For exploratory work, Chapter 18 names the price as “the open-ended adaptability that made an agent attractive in the first place”.
Even a strong oracle has a ceiling. Chapter 13: “A green gate is a fact about the checks you thought to write, and about nothing you did not think to write.” The strict loop in my example still ends 40 items at 70%, so the morning’s reading remains a cost.
A table like this also sits beside the other shapes on an AI agent design patterns cheat sheet; it wraps whichever pattern runs inside RUNNING.
The line to find this week
Open the code that drives your agent and find the line where a run ends. If the model’s final message can reach that line, your loop has one illegal edge, and the first fix is to put a check the worker cannot write between the two. Then give the loop its other three exits, and write the record before each move, the worker’s launch included.
The five verbs, the goal function and the brakes are in Chapter 13, Writing the Outer Loop (in the full book), and the evaluator-optimizer loop is in Chapter 10, Workflows and Composition Patterns (in the full book). The glossary entries linked above are free to read. The free guide to agent patterns collects the related posts and tools, and you can see the formats.
Questions readers ask
- Is an AI agent a state machine?
- The model's own loop is not one: its path is chosen at run time and written nowhere in your code. The loop you write around it can be, and for unattended work it should be: named states, legal transitions, a current state saved to a durable record, and separate exits for done, escalated, stuck and out of budget.
- How does an agent loop know when it is done?
- It should not decide for itself. A goal held outside the model, plus a check that a program evaluates after every pass, decides. The book calls that pair the goal function and the checking program the oracle. The worker's own report that it has finished is stored as a note and is never an input to the decision.
- How do I stop an agent loop from running forever?
- Give the loop more than one exit, each written as a transition in code. Use the goal check, a hard cap on passes and spend that also counts failed passes, a stuck rule that fires after several passes with nothing closed, and an escalation path for work that needs a person. A stop that exists only as a sentence in the prompt is a request.
- Do I need a workflow engine or a graph framework to do this?
- Not for a scheduled loop whose passes are short and whose record is written between passes: a state table and a state file already resume at pass granularity. Consider an engine when a single pass is long, has side effects in the middle, or must wait days at a human gate. No engine decides what done means; that check stays yours.
- What should the outer loop persist between steps?
- For every transition: the next state, the item, the attempt number, the verdict with its evidence, the reason code for the decision, a receipt for any side effect, and a count of unclean restarts in the current state. Record failures and abandoned work as well as successes, and write the record before the next state is entered, and before a worker is launched, so a restarted process can read where it is.
Sources
- Engin Diri (Pulumi) (2026). Stop Prompting. Design the Loop.
- Erik S. and Barry Zhang (Anthropic) (2024). Building Effective Agents
- Sydney Von Arx, Lawrence Chan, Beth Barnes (METR) (2025). Recent Frontier Models Are Reward Hacking (one lab's evaluation suites, June 2025)
- Zhong, Raghunathan, Carlini (2025). ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases (preprint)
- Cemri et al. (2025). Why Do Multi-Agent LLM Systems Fail? (preprint, version 3)
- Temporal documentation (2026). Workflow Execution (documentation; one example of a durable-execution runtime, read October 2026)
- Restate documentation (2026). Durable Execution (documentation; a second example of a durable-execution runtime, read October 2026)
- LangChain documentation (2026). LangGraph overview (documentation; one example of an agent graph framework, read October 2026)
- Apache Burr (2026). Apache Burr README (a second example of an agent graph framework, read October 2026)
- Hacker News (2024). Hacker News comment on agents that report completion without acting (December 2024)
- Hacker News (2025). Ask HN: Do you roll your own agent or use a framework? (comment, October 2025)
- markmhendrickson/ateles issue tracker (2026). Issue #575: a payment handler that treats unsettled statuses as sent (August 2026)
- zi-yue-1129/DATAGEN issue tracker (2026). Issue #59: a step limit that never counts a failed run (September 2026)
- Jamie-BitFlight/claude_skills issue tracker (2026). Issue #1093: a no-progress rule that exists only as prompt text (March 2026)