Home / Blog / Agent fundamentals / AI Agent Architecture Guide: Four Parts, One Lo…

Agent fundamentals

AI Agent Architecture Guide: Four Parts, One Loop, Many Harnesses

This AI agent architecture guide reduces any design to four parts and a loop, orders fifteen decisions by stage, and gives you a doc skeleton to copy.

By Enrique Gutiérrez · Published · 29 min read

This AI agent architecture guide starts from four parts and one loop. An agent is a model client, a tool table, a message history and a loop that calls the model, runs the tool it asks for and appends the result. Everything else on an architecture diagram is harness, added one decision at a time.

The rest of the page is for the person who has to write the design document. It reduces any proposed diagram to those four parts, lists fifteen AI agent architecture decisions in the order a team is forced to answer them, and ends with a document skeleton to copy. The parts and the trade-offs come from AI Agents, Engineered. The ordering and the stage column are this post’s synthesis, and I say so again where they appear.

What are the four parts of an AI agent architecture?

The four parts are a model client, a set of tools, a message history and a loop. Chapter 3 of AI Agents, Engineered (in the full book) makes the claim without hedging: “An agent has exactly four parts, and you already understand every one of them.”

A minimal agent has exactly four parts, and all four live in the harness—the frame here.
Figure 3.4 A minimal agent has exactly four parts, and all four live in the harness—the frame here. The model client speaks to the provider, the tools are the actions on offer, the history is the run’s whole memory, and the loop (in accent) ties them together. The model itself sits outside the frame: the judgment you rent, wired in through the client. The harness is everything else—the part you own and build. Reuse this diagram

The same chapter gives the name for everything else: “An agent, in one line, is a model plus a harness.” The harness is the code you build around the model, and the post on what an agent harness is lists its components and what breaks without each. Here I need one property of it. Every harness part answers a question somebody had to decide.

Other counts exist. One vendor’s guide to building agents (PDF created April 2025) names three core components: model, tools and instructions. The loop arrives later in that guide as the “run”. I prefer four, because the history and the loop are the two parts your team writes, owns and debugs.

The book writes the loop in pseudocode, and so does this post: there is no code here in any language or SDK. A reader who wants to type the four parts can follow build an AI agent from scratch.

How do you reduce someone else’s diagram to four parts?

Take each box and ask which of the four parts it is. A box that is none of them is harness, and the useful follow-up is which decision put it there.

Box on the slide What it is The decision that would justify it
“Reasoning engine”, “LLM core” The model, reached through the model client None; it is the part you rent
“Planner”, “planning module” A prompt, and sometimes a loop shape 6
“Short-term memory” The message history 7
“Long-term memory”, “knowledge base” Tools that save and search, plus a store 7
“Tool layer”, “action layer”, “integrations” The tool table 4
“Orchestrator”, “supervisor”, “router” The loop, or several loops 6, 8
“Specialist agents” More loops, each with its own history 8
“Guardrails”, “safety layer” Harness: checks in code, outside the model 5, 9
“Observability”, “evaluation” Harness 11

Chapter 3 says of its minimal agent that there is “no planning module to install, no reasoning engine to configure, no memory system beyond the list you append to”. Long-term memory, in the same chapter, “arrives with no new machinery: one tool to save a note, another to search past notes.” So a twelve-box slide can hold as few as four parts; the rest are claims about harness.

The explainer pages I fetched while searching for an AI agent architecture guide on 6 October 2026 mostly present layers named perception, reasoning and action, or agent types named reactive and deliberative. Those labels give a builder nothing to choose between, so I leave them out. The book’s version of the same point applies to any framework: “When a framework’s agent misbehaves, there are only four places to look”.

AI agent architecture guide: which fifteen decisions, and in what order?

This AI agent architecture guide puts fifteen decisions in one table, starting with who chooses the next step and ending with what you build or adopt. The table gives each one its options, its trade-off, what to write down, the chapter, and the latest stage at which it can be made.

The provenance matters. Each trade-off is the book’s, quoted or closely paraphrased, with the chapter beside it. The order of the rows and the “decide by” column are this post’s synthesis; the book makes each decision in its own chapter and assigns no stages. Each “what to write in the doc” cell is also mine.

The stages are defined by who is watching:

  • Prototype: before the first run on a real task, with the builder reading every transcript.
  • Pilot: before anyone other than the builder relies on the output.
  • Production: before runs happen that nobody watches as they go, or before more people use it than the builder can sit beside.

When does a row come due earlier than its stage?

A row comes due early when either of two conditions fails; each stage is the latest safe moment only while both hold. First, every irreversible or externally visible tool is stubbed, confined or pointed at a test account. Second, every run finishes while a caller waits for it, the caller being the person or program that asked. A scheduled run has no caller, so it fails the second condition from its first night.

When either condition fails, rows 9 and 12 are due before that tool goes live or that run starts. A run that no caller waits for also needs the serving half of row 14 at that moment. Those conditions are Chapter 18’s two questions turned into a rule: “what here is irreversible, and how long does the run live?”

With JavaScript on, buttons above the table narrow it to the rows due at one stage; the earlier stages’ rows are assumed done. Chapter 1 is free to read online; the other chapters cited are in the full book.

# Decision Options Trade-off, in the book’s terms What to write in the doc Chapter Decide by
1 Who chooses the next step: your code or the model? Plain code · one model call · a workflow · an agent · an agent that proposes to a person “Each rung up buys adaptability and pays for it in money, latency, and predictability.” The test: “can you draw the flowchart before the request arrives?” The rung, and the evidence that the rung below fails Ch. 1, Ch. 14 prototype
2 What ends a run, and what says this run worked? The model’s own “done” · a check a program can run · a named reviewer · a budget · an error halt “The model’s “done” is testimony; a green test is evidence.” The success check, or the reviewer where no program can check; what each exit returns Ch. 3 prototype
3 What are the ceilings? Caps on passes, tokens, wall-clock time and money; one or several. For an agent that writes, a cap on records or actions changed per run (this post’s addition) “Set the cap before you write anything else.” And: “Without a cap, that bug is a bill with no ceiling; with one, it is a log entry.” The numbers, labeled as starting values; what is returned when one fires Ch. 3 prototype
4 Which tools, at what grain? One tool per endpoint · task-shaped tools · one general code-execution tool “The right unit for a tool is a task”. A general tool “offers no semantic guardrails, demands more skill from the model, and is harder to permission and audit” The tool table: name, read or write, reversible or irreversible, and what makes it reversible. Each irreversible tool is stubbed, confined by row 5 or pointed at a test account until rows 9 and 12 are decided Ch. 5 prototype
5 Where does model-written code run, queries included, and what can untrusted text reach? Nothing to isolate · the host process · a container · stronger isolation; network open or allow-listed “never run model-written code with access to anything you would not hand a stranger”. The cost: “Code execution is infrastructure” The boundary, the network rule and what is mounted or reachable; or one line saying no code runs and which inputs are untrusted. This row owns where a model-written query executes; what its credential may do is row 9’s Ch. 5, Ch. 17 prototype
6 What shape is the loop? A free loop · a plan first, then execution with replanning · named states with model judgment inside each A free loop pays a model call per action and can wander. On planning: “Global coherence is a thing you buy.” On states: “What the state machine buys is legibility”, paid for in adaptability The shape, the named states if any, the rule for replanning Ch. 3, Ch. 4, Ch. 18 pilot
7 What enters the window on each pass? Carry everything · write out, select, compress, isolate; knowledge carried, retrieved or trained “The trade is always recall against focus” The assembly order, a soft ceiling (illustrative), what is trimmed first, where knowledge comes from Ch. 7, Ch. 8 pilot
8 One agent or several? One continuous context · an orchestrator with isolated workers “Start with one agent.” Split only if the answer is yes to “does this work decompose into branches that can run without seeing each other?” If split, “writes stay single-threaded while intelligence fans out” The answer to that one question; the brief format if split Ch. 11 pilot
9 What may it do without a person, and where is that enforced? A prompt instruction · a gate in code by consequence tier: none, a log, a review queue or a signature · scoped credentials “An instruction in the prompt is a request.” The alternative is to “classify everything the agent can do into consequence tiers—read-only, reversible, externally visible, irreversible—and attach the oversight to the tier” Each tool’s tier and gate; who clears the queue when nobody is present at run time; each credential’s scope and lifetime; how many of private data, untrusted content and an outbound channel are present Ch. 12, Ch. 17 pilot
10 What happens when a call fails? Crash · swallow the error · surface it to the model · retry · halt “the model can only recover from a failure it can see”. The bill: “Reliability engineering, like insurance, is priced per what it protects.” The error shape; per tool, whether a failure is retried, surfaced or halts; a retry budget for the whole run. Until row 12 gives a write a key, a write that times out halts Ch. 18 pilot
11 How will you know what it did, and whether a change made it worse? A returned transcript, which the prototype stage already assumes · a durable trace · a fixed case bank with a merge gate · a staged rollout “every change to a prompt, a tool definition, or the model reruns the suite, and the numbers decide”. The gate’s limit: “the gate tests the failures you already imagined” Trace fields; where the cases come from; the gate; rollout rungs, which may read “deferred until production” Ch. 16, Ch. 20 pilot
12 Where does the run live, and what does a resume do? The memory of one process · a periodic checkpoint · a journal and replay; keys and receipts on writes; compensations for sequences A checkpoint records position: “It does not, and cannot, record whether the outside world did the work.” And: “agents that persist everything pay for it, and agents that persist nothing pay more” Durable and ephemeral state; the cadence, which may be none; which writes carry a key; the crash test you ran Ch. 18, Ch. 20 production
13 What does it do when the primary path is gone? Crash · improvise · a fallback ladder that ends in a person, at once or through a ticket read later “Fall back on words; fail loud on deeds.” The ladder for each critical dependency; what the escalation hands the person Ch. 18 production
14 Which two of accuracy, cost and latency, and how is it served? Any two; a synchronous request · a queued job that returns a task id · a scheduled run “accuracy, cost, and latency form a triangle, and you usually get to pick two” The two, per feature; cost per task and the term that makes it grow; the serving shape; for a scheduled run, who reads the result and when Ch. 19, Ch. 20 production
15 What do you build, what do you adopt, and how much harness? Your own loop · a framework · a durable-execution runtime; a thin harness · a fuller one “Adopt for a named need, never for the feeling that serious systems use frameworks.” Custody: “never store what you own inside what you rent” For each adopted part, the row whose need it meets; what stays in your repository; the audit on each model change Ch. 3, Ch. 18, Ch. 27 prototype, pilot, production

The nearest published neighbor I know is 12-Factor Agents (repository created March 2025). It is twelve principles with no build order and no “what to write down” column, and several of its factors map onto rows 7, 9 and 12.

What do you decide before the prototype runs?

Five things: the rung, the stop rule, the ceilings, the tool table, and where anything dangerous runs. Each takes minutes for a small system, and each is expensive to discover from an incident.

Row 1 comes first because it can shrink the list. A workflow whose steps you can draw in advance has no loop to shape (row 6) and no agent count to settle (row 8); it still needs a success check, ceilings, a context policy, authority and evidence. Schluntz and Zhang’s 2024 essay Building effective agents says the same from a vendor’s side: “we recommend finding the simplest solution possible, and only increasing complexity when needed.” The should this be an agent? tool turns Chapters 1 and 14 into ten questions, and the post on the agents versus workflows decision rule covers the ladder.

Rows 2 and 3 are the stop condition and the budget. A check a program can run matters for the same reason covered in why AI agents hallucinate: the model’s report of success is generated the way everything else it says is generated. Where no program can check, as with a research summary, write the reviewer’s name in that cell.

For the ceiling, Chapter 3 mentions ten or twenty passes for a focused task and labels the figure as illustration. The cap on records changed per run is my addition to the chapter’s four, for agents that write. What an agent loop is ranks the exits, and run the loop lets you practice them.

Which tools exist, and where does risky code run?

Rows 4 and 5 fix the action space: the tool table, and the boundary around anything the model writes. The tool table doubles as the inventory for row 9, which is why I ask for the read-or-write and reversible columns this early. How to design tools for LLM agents covers the contract, and the tool contract linter checks one definition at a time.

Row 5 is a prototype decision because a builder’s laptop holds credentials. An agent that runs no code answers it in one line. A model-written database query is code for this purpose: row 5 says where it executes and what that process can reach, and row 9 says what its credential may do.

What do you decide before anyone else relies on it?

Six things: the loop’s shape, the context policy, one agent or several, authority, the failure path, and evidence. All six can be deferred while one person reads every transcript, and none can once a second person trusts the output.

Shape (row 6). The default free loop is the ReAct agent pattern: reason, act, observe, repeat. Named states suit a run that must pause for an approval, since the current state is small and honest to save. The catalog of orchestration shapes belongs on an AI agent design patterns cheat sheet; this table gives them one row.

Context (row 7). Write down what enters the window and what leaves it first. The context window budget planner prices one pass. A dated first-hand report shows why this row needs a reopening signal. The author of one agent product’s 2025 account writes, “we’ve rebuilt our agent framework four times, each time after discovering a better way to shape context.”

One or several (row 8). Two vendor architecture guides, read as examples of the category, agree with the book’s default. The Google Cloud guide (last reviewed 28 May 2026) says “we recommend that you start with a single agent”, and the OpenAI guide cited earlier says to “maximize a single agent’s capabilities first”. One vendor’s 2025 write-up of its own research system prices the alternative: “multi-agent systems use about 15× more tokens than chats”. Single agent vs multi-agent has the full verdict.

Who signs, what fails, and how will you know?

Rows 9, 10 and 11 answer those three questions, and each has a home in code that the prompt cannot replace.

Authority (row 9). The approval gate sits in code, in Chapter 12’s words “after the model has chosen, before anything runs”. The chapter attaches one of four gates to each tier: none, a log, a review queue or a signature. An agent that runs while nobody is present still gets an answer: its externally visible and irreversible writes land in a queue, as a proposed change that a person applies or approves later.

The consequence tier classifier grades each tool. The lethal trifecta audit counts the three legs the doc cell asks about.

Failure and evidence (rows 10 and 11). Row 10 carries a clause that keeps the table consistent with itself: a write without a key halts on a timeout, because nobody knows whether it happened. Row 11 needs a trace and a regression gate before the pilot, since the first complaint from a user will be about a run nobody watched.

What do you decide before runs go unwatched?

Three things: where the run lives, what happens when a dependency is gone, and how the feature is priced and served. They can wait because a short, watched, read-only run loses little when it dies. A system that a person always watches, such as an assistant that only drafts, reaches this stage by headcount.

Row 12 is the one most often answered by accident. A checkpoint tells a resumed run where it was. Whether the email went out or the card was charged is a separate fact, which is what idempotency keys and receipts record; idempotent tools and safe retries covers the mechanics. Chapter 18 supplies the test to write into the doc: “a checkpoint you have never restored from is a guess”.

Adopted machinery leaves this row open. In September 2026 a user on the issue tracker of a widely used agent framework asked: “When a tool call times out or the worker dies mid-tool, what is the intended recovery contract?” That question is row 12, asked after the framework was chosen. The 2025 write-up cited above says of its own system that “we built systems that can resume from where the agent was when the errors occurred.”

“None” is a legitimate answer here. Chapter 18 says of checkpoint cadence that “for a task that finishes in one sitting the honest cadence is none at all”.

What happens when a dependency fails, and what does a run cost?

Row 13 separates advice from action. A stale answer with a warning label is an acceptable fallback for a read; a refund has no acceptable weaker version, so it stops and goes to a person.

Row 14 sets the economics, and the agent cost-per-task estimator shows the term that grows with step count. Chapter 20 adds the serving rule: once a run outlasts the connection, the caller gets a task id, and the run’s identity lives in a database row that no process owns. A scheduled run has no caller, so its entry names who reads the result and when.

Should you build the loop or adopt a framework?

Decide it last at each stage, after the rows above have named what the system needs, and adopt when the need is one you would otherwise have to build. Row 15 is the only row marked for all three stages, because each stage names new needs. The post on what an agent harness is has the table of what a framework supplies and what each part costs; this section adds the evidence and the entry to write.

Readers usually ask this first. An Ask HN post from October 2025 puts it plainly: “I am curious if people here roll their own agents from scratch or use frameworks.” Replies in that thread run both ways. One reads “We used frameworks in the past […] Now we roll our own, generally a much better and cleaner experience.” Another comes from someone content with a framework who had rolled their own before. Of the hand-written version: “it worked well for very simple tasks”, and “I would be concerned about complexity if there were more steps, tool calls, or the need to compose multiple agents”.

Has anyone measured a framework against a hand-written loop?

I looked for a public measurement that compares a framework with a hand-written loop on the same task and the same model, and as of October 2026 found none.

The nearest thing is a demo paper by Milev and Kanagala at ACM CAIS, 29 May 2026, of which I read the abstract only. It holds one model and one tool server fixed across six frameworks and SDKs and scores three scenarios. The abstract reports “empirical evidence that scenario-specific orchestration adds no measurable benefit over generic agentic loops driven by prompts.” It compares frameworks with one another, includes no hand-written baseline, and is one team’s result on one model.

Two older papers measure neighboring things. Kapoor et al. (2024) argue in their abstract that “SOTA agents are needlessly complex and costly, and the community has reached mistaken conclusions about the sources of accuracy gains”, which concerns agent designs on benchmarks.

Cemri et al. (2025) annotate 1600+ traces across 7 multi-agent frameworks into 14 failure modes; the study does not isolate the framework as a cause.

So the rule has to work without that evidence, and the book’s does. Chapter 3 states the trade as symmetrical: “minimal code gives you total visibility and no magic, at the price of writing your own plumbing; a framework supplies the plumbing, at the price of opacity, lock-in, and abstraction layers to debug through.”

What does a fair reading of both sides say?

Both sides agree on the list of plumbing and disagree only about who writes it. A framework vendor’s own essay (Harrison Chase, April 2025) says agentic systems “all benefit from the same set of helpful features, which can be provided by a framework, or built from scratch.” Chapter 3 gives the list: persistence, retries, tracing, memory management and orchestration. All five change how dependable a run is, and none changes what the model does on a single pass.

Adopting early is reasonable when the team already knows the framework, can point to the four parts inside it, and can name the row it serves. Building is reasonable while one loop is being learned or prototyped. Either choice is safe on one condition, which Schluntz and Zhang’s essay also sets: the team understands the code underneath.

Durable-execution runtimes show what adoption does and does not settle. Two examples of the category are Temporal, whose 2025 explainer defines the term, and Restate. Restate’s documentation (read October 2026) says the runtime “replays the journal, skipping completed steps and resuming from exactly where it left off.” Either meets the need in row 12 for long runs. Which writes carry a key is still your entry to make.

Lock-in inverted.
Figure 27.4 Lock-in inverted. The durable core you own—prompts, tool schemas, the eval set—sits at the center (in accent), and the three parts you rent (the model, the framework, the runtime) plug in around it through identical, thin standard sockets. Because the sockets are standard, each rented module has a swappable twin waiting: the ecosystem can churn and a swap becomes an edit behind an interface rather than a rewrite. Reuse this diagram

How much harness is enough?

Enough to match how long the system lives, how many people trust it and how unattended it runs. That is Chapter 18’s scaling law: “harness investment should scale with how long the codebase must live, how many people must trust the agent’s output, and how unattended the agent runs.”

The chapter endorses one audit for every team, which goes in row 15’s cell: “on every model change, reread your standing machinery and delete what the new model has made redundant”. The harness grid builder lays your controls out so the empty cells show, and the harness post linked above covers the debate behind the law.

Worked example: a support agent that issues refunds

Here is one system walked through all fifteen rows. The system is hypothetical: an agent for an online store that reads a ticket, looks up the order and may issue a refund. Every number in it is illustrative.

# What the doc says Decided at
1 A workflow handles refunds whose rule is a lookup. The agent takes tickets where the next step depends on what the order history shows. Evidence: a sample of past tickets the flowchart could not route prototype
2 Done means the payment system shows one refund against the order and the reply quotes its id; the harness reads that back. On a reply-only ticket the check is the support agent who reads the draft and sends it. A budget or error exit returns the transcript, marked stopped prototype
3 12 passes, 90 seconds of working time (the wait for a signature is not counted), a token cap and a money cap per ticket, and one refund per ticket. When one fires, the ticket goes to a person with the transcript prototype
4 get_order_context (read), draft_reply (write, reversible: a draft can be discarded; a support agent reads it and sends it), refund_payment (write, irreversible, pointed at the payment provider’s test mode until rows 9 and 12 are decided) prototype
5 No code runs. Ticket text and attachments are untrusted and reach the model as data prototype
6 Named states: gathering, proposing, awaiting approval, acting, done. The model decides inside gathering and proposing. No replanning: a rejected proposal returns to gathering once, then goes to a person pilot
7 In order: the refund policy (standing, carried), the compiled order context (retrieved per ticket), the thread. Soft ceiling at 40% of the window; long threads are summarized, oldest turns first pilot
8 One agent. The answer to the one question is no: the lookup, the decision and the refund share state pilot
9 Gates: get_order_context none, draft_reply a log, refund_payment a signature every time, from the support lead on shift. The credential can refund and nothing else, and is issued per ticket. All three trifecta legs are present; order lookups are scoped in code to the customer who opened the ticket, which narrows the private-data leg pilot, before the first live refund
10 Errors return a code, a message, a retryable flag and a hint. Read timeouts retry twice; a refund timeout halts and escalates; six retries per ticket in total pilot
11 Trace per step: state, tool, arguments, result, tokens, time. Cases: 50 past tickets whose outcome a person decided. Gate: no change to a prompt or tool merges if the pass rate on those cases falls. Rollout rungs: deferred until production pilot
12 Durable: the approval (who, when, what they saw) and the refund receipt. Ephemeral: the gathered context. Checkpoint at each state change. Key on refund_payment: ticket, step, tool, arguments. Crash test: killed after the provider accepted the refund; the resumed run found the receipt and did not refund again pilot, pulled forward from production
13 Payment provider down: no refund is promised and the ticket goes to a person. Order system down: the drafted reply says a person will follow up. The handover carries the goal, what was tried, the failure and the trace production
14 Accuracy and cost. Cost per ticket grows with passes, since the order context is re-read on each one. Served as a queued job: the help desk sends its fixed acknowledgment at once, and the decision follows the signature; the ticket id is the task id production; the serving half at pilot
15 Prototype: our own loop. Pilot: an approval can wait for hours, so the run must survive a restart (row 12). A table of our own or a durable-execution runtime would both meet it; prompts, tool contracts, cases and traces stay in our repository either way. On each model change, the standing refund instructions are reread each stage

Three entries moved. The pilot issues real refunds and waits hours for a signature, so both conditions fail: row 9 was settled before the first live refund, and row 12 and the serving half of row 14 came forward from production. During the prototype the refund tool pointed at test mode, so a refund that timed out cost nothing.

I walked four other systems through the table. A read-only research assistant answers rows 5, 8 and 12 in one line each and puts a named reviewer in row 2. A coding agent that runs commands spends most of its effort on rows 2, 5 and 9: a test suite as the check, a container with the network denied, and a gate on anything that leaves the sandbox.

The last two came from this post’s reviewer and changed the table. A nightly data-cleanup agent with write access has no caller and nobody present, which is why row 3 caps records changed, row 9 names the queue and row 14 lists a scheduled run. An assistant that only drafts replies is watched forever, which is why production has a second trigger.

What goes in an architecture document for an agent?

One section per decision, each stating what you will do, the alternative you rejected, the consequence you accept and the signal that would prove you wrong. The skeleton below is original to this post; As of October 2026 I found no published template specific to agents.

Its form borrows from two well-known sources on the genre. Michael Nygard’s 2011 decision records state each decision in full sentences that begin “We will”. Malte Ubl’s 2020 account of design docs calls the alternatives section “one of the most important ones”.

ARCHITECTURE: <agent name>     Status: proposed | accepted     Date:

The four parts
  Model client:    <the one place every model call passes through>
  Tools:           <count; the table is in section 4>
  Message history: <what is appended; where it is stored>
  Loop:            <where it lives; shape is in section 6>

Each section below uses the same four lines.
  We will:   <the decision, in a full sentence>
  Rejected:  <the alternative, and why>
  Accepted:  <the consequence we take on>
  Wrong if:  <the signal that reopens this decision>
A deferred section reads: "Deferred until <signal>. Until then: <default>."
The stage headings hold under two conditions: every irreversible or externally visible tool is
stubbed, confined or on a test account, and every run finishes while a caller
waits for it. If either fails, sections 9 and 12 are due now, and so is the
serving line of section 14 for a run that no caller waits for.

BEFORE THE PROTOTYPE RUNS
 1. Who chooses the next step
    We will:   <code / one call / workflow / agent / agent proposing to a person>
    Rejected:  <the rung below, with the evidence that it fails>
    Accepted:  <the cost and unpredictability this rung adds>
    Wrong if:  <e.g. every run takes the same path>
 2. Stop rule and success check
    We will:   <the check a program runs, or the named reviewer>
    Rejected:  Accepted:  Wrong if:
 3. Ceilings (starting values)
    We will:   passes __  tokens __  wall-clock __  money __
               records or actions changed per run __   (if it writes)
               When one fires we return: transcript, partial work, "stopped".
    Rejected:  Accepted:  Wrong if:  <e.g. a cap fires on most runs>
 4. Tool table
    We will:   name | read or write | reversible? and what makes it so
               Irreversible tools stay stubbed, confined or on a test account
               until sections 9 and 12 are decided.
    Rejected:  Accepted:  Wrong if:
 5. Isolation
    We will:   <boundary; network rule; what is mounted or reachable>
               or <no code runs>. Model-written queries execute: <where>
               Untrusted inputs: <list>
    Rejected:  Accepted:  Wrong if:

BEFORE ANYONE ELSE RELIES ON IT
 6. Loop shape
    We will:   <free loop / plan then execute / named states: ...>
    Rejected:  Accepted:  Wrong if:
 7. Context policy
    We will:   <what enters each pass, in order; soft ceiling; trimmed first>
               Knowledge: <carried / retrieved / trained>
    Rejected:  Accepted:  Wrong if:
 8. One agent or several
    We will:   <one / orchestrator and workers, with the brief format>
    Rejected:  Accepted:  Wrong if:
 9. Authority
    We will:   per tool: tier | gate (none / log / review queue / signature)
               Who clears the queue when nobody is present: <...>
               Credentials: <scope and lifetime>
               Private data, untrusted content, outbound channel: __ of 3
    Rejected:  Accepted:  Wrong if:
10. Failure path
    We will:   Error shape: code | message | retryable | hint
               Retried: <list>   Surfaced: <list>   Halts the run: <list>
               Retry budget per run: __
    Rejected:  Accepted:  Wrong if:
11. Evidence
    We will:   Trace fields: <...>   Cases from: <...>   Gate: <...>
               Rollout rungs: <...>
    Rejected:  Accepted:  Wrong if:

BEFORE RUNS GO UNWATCHED
12. State and recovery
    We will:   Durable: <...>   Ephemeral: <...>   Cadence: <... or none>
               Writes that carry a key: <list>
               Crash test on <date>: killed after ___; resumed run did ___.
    Rejected:  Accepted:  Wrong if:
13. Fallback
    We will:   <per critical dependency: the ladder, ending in a person,
               at once or through a ticket read later>
               The escalation hands over: <goal, attempts, failure, next step, trace>
    Rejected:  Accepted:  Wrong if:
14. Cost and serving
    We will:   <the two of accuracy, cost, latency, per feature>
               Cost per task: __   Grows with: __
               Served: <sync / queued with a task id / scheduled>
               If scheduled: who reads the result, and when
    Rejected:  Accepted:  Wrong if:

AT EVERY STAGE, LAST
15. Built, adopted, and how much harness
    We will:   adopted part | the section whose need it meets | what stays ours
               On each model change we reread and delete: <...>
    Rejected:  Accepted:  Wrong if:

Once the sections are filled, the agent verifiability scorecard is a useful second read. It asks twelve questions about the signals the system exposes.

What can you write as “deferred until X”?

Any row whose stage has not arrived, provided the entry names the signal that reopens it and the default in force until then. A deferred entry with no signal is an undecided one.

Three examples, all in the form the skeleton gives:

  • Row 8: “Deferred until one task’s reading overflows a single window on more than a few runs. Until then: one agent.”
  • Row 12: “Deferred until a run has no caller waiting or a write goes live. Until then: process memory, and a failed run is rerun.”
  • Row 15: “Deferred until a section names a need we would spend more than a sprint building. Until then: our own loop.”

The rows that resist deferral are 1 through 5. A prototype without them still has a rung, a stop rule, a ceiling, a tool set and a place where code runs. It has them by default, and nobody chose the defaults.

Where does this order not hold?

It fails wherever the two conditions fail, and it has four further limits that the table cannot show.

Irreversibility and run length reorder it. A team whose first tool charges cards meets rows 9 and 12 on day one. The stage column describes the common case.

Many systems leave rows empty. Chapter 18’s own example is that “the thirty-second read-only summary needs none of it”, speaking of the recovery apparatus. Filling every section at full depth for a watched prototype is the over-building the scaling law warns about.

The stages and the skeleton are untested as a method. In this AI agent architecture guide they are a synthesis of the book’s chapters and two documentation formats. Nobody has measured whether teams that fill this document ship better agents.

A thin seam protects code and leaves behavior exposed. Chapter 27 recommends one place through which all model calls pass. A commenter on a June 2025 thread, arguing against frameworks, wrote that “the ability to swap out APIs just isn’t the bottleneck.. like ever.” Both points stand: the seam absorbs an interface change, and row 11’s gate is what catches a change in behavior.

The harness depreciates. Some of what you write in rows 6 and 7 compensates for the current model’s weaknesses, and the audit in row 15 exists to remove it.

The review to hold this week

Take the proposal on your desk and sort its boxes into the four parts and the harness. Then open the skeleton, fill rows 1 through 5, and write “deferred until” on every other row with a signal beside it. A review that argues about those entries is arguing about the architecture, and that argument is what an AI agent architecture guide is for. Appendix A (in the full book) closes on the same thought about the minimal agent: “what separates your afternoon from their operation is harness, applied one named need at a time.”

The four parts and the stop rules are in Chapter 3, The Agent Loop (in the full book), the action space is in Chapter 5, Tools and the Action Space (in the full book), and failure, state and harness sizing are in Chapter 18, Reliability, State, and the Harness (in the full book). Chapter 1, What Is an Agent? is free to read online and covers row 1. The agent fundamentals guide collects the related posts and tools, and you can see the formats.

Questions readers ask

What are the components of an AI agent architecture?
Four: a model client that sends messages and tool definitions to the model, a set of tools, a message history that holds everything the run has seen and done, and a loop that ties them together. Budgets, recovery, traces, gates and sandboxes are harness, added around those four parts for a named consequence.
Should I use a framework or write my own agent loop?
Write down the need first. A framework supplies plumbing such as persistence, retries and tracing, at the price of opacity and lock-in; your own loop gives full visibility at the price of writing that plumbing. Adopt when you can name the need the framework meets and can still find the four parts inside it. No public head-to-head measurement was found for this post.
Where should an AI agent's state live?
For a short, read-only run, in the memory of the process, and a failed run is simply rerun. Once a run has no caller waiting for it or performs an irreversible action, its state belongs in storage that no single process owns, keyed by a run id, with approvals and receipts for completed writes saved first.
How does an agent recover from a crash in the middle of a run?
It reloads a checkpoint or replays a journal of recorded steps. Neither can know whether a write that was in flight at the moment of the crash reached the outside world, so each consequential write carries an idempotency key and leaves a receipt that the resumed run checks before acting. Test it by killing the process on purpose.
When is a multi-agent architecture worth it?
When the work divides into branches that can run without seeing each other, with no shared mutable state and no ordering between them, and the task is valuable enough to pay the multiplied token bill. If the steps share state or must agree on decisions, one agent with one continuous context is the better design.

Sources

  1. Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents
  2. Dex Horthy (HumanLayer) (2025). 12-Factor Agents (repository created March 2025)
  3. Anthropic (2025). How we built our multi-agent research system (one vendor's engineering write-up, cited for its own figures)
  4. Yichao "Peak" Ji (2025). Context Engineering for AI Agents: Lessons from Building Manus
  5. OpenAI (2025). A practical guide to building agents (PDF created April 2025; one example of a vendor architecture guide)
  6. Google Cloud Architecture Center (2026). Choose a design pattern for your agentic AI system (last reviewed 28 May 2026; a second example of a vendor architecture guide)
  7. Harrison Chase (LangChain) (2025). How to think about agent frameworks (a framework vendor's essay, 20 April 2025)
  8. Roberto Milev and Uday Kanagala (2026). Arena: Benchmarking AI Agent Frameworks Under Fixed-Model Conditions (demo, ACM CAIS 2026, 29 May 2026; abstract read)
  9. Sayash Kapoor et al. (2024). AI Agents That Matter (preprint, arXiv:2407.01502; abstract read)
  10. Mert Cemri et al. (2025). Why Do Multi-Agent LLM Systems Fail? (preprint, arXiv:2503.13657; abstract read)
  11. Michael Nygard (2011). Documenting Architecture Decisions
  12. Malte Ubl (2020). Design Docs at Google
  13. Tom Wheeler (Temporal) (2025). The definitive guide to Durable Execution (one example of a durable-execution runtime's own description)
  14. Restate. Durable Execution, documentation (a second example of the category; read October 2026)
  15. langchain-ai/langgraph issue tracker (2026). Issue #9006: the intended recovery contract when a worker dies mid-tool (September 2026)
  16. Hacker News (2025). Ask HN: Do you roll your own agent or use a framework? (October 2025, with two replies quoted)
  17. Hacker News (2025). Hacker News comment on swapping model APIs (June 2025)