This AI agent design patterns cheat sheet is one table of 22 patterns, from a single model call to an unattended loop. Every row has the same six fields: its shape, when to use it, its cost in calls, its characteristic failure, the check that catches that failure, and the simpler pattern to try first.
It is a reference page for a design review, so the prose is short. The rows and their names come from Chapters 10 to 13 of the book. Each pattern has a longer post or chapter section, linked from the index below the table.
How do you read this AI agent design patterns cheat sheet?
Read a row left to right as a review question: what is the shape, what will it cost, how will it break, what will tell us, and why was the simpler row not enough. The last column is the family, and the buttons above the table filter on it.
Shape is one line on who decides the next step: your code, the model, or a person. Calls counts model calls for one request, relative to one call. The letters are the same in every row: k stations, n votes, N rounds (the first attempt is round one), s steps in one agent run, W workers, P passes of a loop.
Those counts are arithmetic, not measurements. They assume each call costs about what a single call costs, which undercounts any loop: an agent re-reads its history on every step, so tokens grow faster than calls. Oversight rows cost no model calls; they are priced in human decisions and waiting time.
Failure is the one that belongs to the pattern, the fault it adds that a single call did not have. Check is what a reviewer can ask to see. Try first points at a cheaper row, which makes the table a ladder.
The cheat sheet: 22 patterns, six fields each
The table is my arrangement of the patterns that Chapters 10 to 13 name. The patterns and their definitions are the book’s, and where the book takes them from a source the source is credited below. A few row labels are mine where the book uses another name: Chapter 11 calls the specialists the reasoning consultant and the search specialist and the third archetype the reviewer, and Chapter 13 calls the fresh-context loop by its practitioners’ nickname and the shared store a store of signals. The call counts and the “try first” column are this post’s own.
| Pattern | Shape | Use when | Calls vs one call | Characteristic failure | The check that catches it | Try first | Family |
|---|---|---|---|---|---|---|---|
| One augmented call | One model call with retrieval, tools and memory attached; no arrow | The job is to classify, extract, summarize or answer over given documents | 1 | Output that is well formed and wrong; nothing raises | A labeled set rerun on every change; a gate in code on the output | A better prompt and the right inputs; nothing sits below this row | workflow |
| Prompt chain | Calls in a fixed order, a gate in code at each seam | One call must do several jobs in sequence and each handoff can be checked | k calls, about k× latency | An early error elaborated by every later station; stages invented so that there is a chain | A gate per seam; rejections counted per seam; a seam with no checkable contract is merged back | One call with a gate | workflow |
| Routing | A classifier picks one of several branches you wrote | Inputs come in kinds that want different prompts, tools or models | 2 in sequence; 1 if the router is a rule in code | A confident answer from the wrong branch | The router measured alone as a classifier on labeled inputs; an “other” label with a fallback | One call; a rule in code where the categories are crisp | workflow |
| Cascade | A cheap model answers first; a dearer one is called only if the answer fails a quality gate | Routine inputs are common, and the path is batch or background work | 1 when the cheap answer stands, 2 in sequence when it escalates; on average 1 + the escalated share | A weak answer that passes the quality gate; doubled latency on every escalation | Escalation rate; error rate of the answers that did not escalate, on a labeled sample | One model for everything | workflow |
| Parallelize by sectioning | Independent subtasks run at once; code joins the results | One step holds several separate concerns | k at once, about 1× latency; add 1 if a model merges | Pieces that each settled a shared decision their own way | Merge fixtures with inconsistent pieces; a rule for a missing branch | One call; a chain if the parts share a decision | workflow |
| Parallelize by voting | The same task run several times; code counts the answers | One verdict matters too much for one draw, and answers compare by equality | n at once, about 1× latency | Unanimous and wrong; “keep the best” quietly needs a scorer | Unanimous-and-wrong rate on a labeled sample; a threshold set from the two error costs | One call with a gate | workflow |
| Evaluator-optimizer (generator-critic) | Generate, check, feed the failure back, revise; the checker is not the generator | Something can say what is wrong: tests, a schema, a calibrated critic | Up to N with a check in code; up to 2N with a model critic | Rubber stamp, oscillation, over-edit, or a climb toward the critic’s blind spot | A round cap in code; a no-progress stop; keep the best candidate; known-bad candidates must fail | A gate with one retry | workflow |
| Iterated self-review (Rule of Five) | The same model reviews its own work several times, at different altitudes | No verifier, no second model and no rubric exist yet | 1 + one call per pass; Chapter 10 reports two to three passes on small tasks, four to five on large | “Converged” means the critic ran out of complaints; findings are testimony | A real verifier after the passes; a cap on passes | The evaluator-optimizer row, wherever a check in code exists | workflow |
| Single agent loop | One model decides each next step in one continuous context until an exit fires | You cannot write the subtask list before the input arrives | s calls for s steps; tokens grow faster than calls | “Done” declared early, or never; errors that compound across steps | Exits in code: a verification signal, a budget, a step cap; a trace of every step | A workflow row | single agent |
| Orchestrator-worker | A lead model writes the subtask list at run time, briefs workers with their own contexts, then synthesizes | Writing the list takes judgment, and the branches can run without seeing each other | 1 + W × s + 1 for each wave | Duplicated work and gaps from thin briefs; pieces built on conflicting assumptions | A brief a stranger could execute; every brief and return recorded; one writer per artifact | A single agent; sectioning if the list is fixed | multi-agent |
| Read-only specialist on call | A stronger reasoner or a search subagent, exposed to the main agent as a tool | A hard sub-problem, or a search whose output is not needed afterward | Add the specialist’s own steps for each consultation | Consulted on everything; its summary trusted as fact | Consultations counted per run; the return treated as tool output, untrusted until checked | The main agent with the right context loaded | multi-agent |
| Fresh-window reviewer | A read-only agent that did not write the work judges it against the standard | A mistake hides in the author’s own context, and no check in code covers it | Add 1 call for each review, or its own steps if it drives the running system | Author and reviewer circle; “works” accepted with no evidence | Read-only access; a cap on rounds; proof required: observed behavior, not the maker’s claim | A check in code: tests, lint, types | multi-agent |
| Human-orchestrated parallel agents | A person splits the work, briefs several agents and judges results; a competition runs one task several times and keeps the best | The pieces are small and independent and each result is cheap to verify | Agents × s; a competition of k attempts costs k × s | The operator’s attention saturates; drift goes unnoticed; a command lands in the wrong workspace | A work-in-progress limit; a status board; an isolated checkout for each agent | One agent you steer | multi-agent |
| Swarm with a merge queue | Many workers write in parallel from one baseline; merges are serialized | The work is inherently parallel and each piece is specified in advance | W × s, plus a merge pass for each worker | The merge wall: later work rests on a baseline that earlier merges changed | A merge queue; triage of serial against parallel work before the fan-out | Orchestrator-worker with a single writer | multi-agent |
| Approval gate | The run halts before an action until a person approves, rejects or edits | The action is externally visible or cannot be undone | 0 model calls; one human decision and one wait for each gated action | Gating so much that approval becomes reflex; a gate the model can talk itself past | The gate keyed on the action, in code, between proposing and executing; consequence tiers written down | Make the action reversible and logged, so it needs no signature | oversight |
| Ask-a-human tool | The agent can choose to consult a person, like any other tool | A question of preference or intent that only a person can settle | 0 model calls; one question and one wait for each ask | Asks at every step, or pushes through and misreads the intent | Questions counted per run; each one read: was it a fact the agent could have looked up? | Put the answer in the brief | oversight |
| Escalation triggers | Conditions in code around the loop hand control to a person mid-run | Repeated failure, an input outside policy, signs of injection, a budget nearly spent | 0 | Self-reported confidence as the only trigger; a raw transcript handed over | The trigger list reviewed; the handoff packaged as a decision with what it touches and whether it can be undone | Budgets and stop conditions on the run | oversight |
| Plan-level approval | A person approves the plan and scope once, keeps the right to interrupt, then reviews intent and evidence | Tasks above the trivial with mostly reversible steps | 0 model calls; one human approval of the plan and one review for each task | A good plan and a drifting execution; a skimmed diff that certifies presence | A declared scope, so anything outside it is a flag; evidence of outcome a stranger could check | Per-action gates on the top tiers only | oversight |
| Autonomy dial | A per-task setting with four positions: signature on each consequential action, standing permissions, plan-level approval, monitored autonomy | Every task; the position moves on evidence | 0; a setting | One position for the whole system; positions kept after the model changes | A record for each task class with the evidence behind its position; re-earned on each model change | The tightest position | oversight |
| Fresh-context loop | The same prompt run pass after pass with a wiped window; state lives on disk; a phase loop adds named phases | A well-specified job whose “done” a machine can verify | P × s; each pass re-reads the spec from the start | Rebuilds what already exists; stubs that pass; runs until an outside limit stops it | Back-pressure: compiler, types and tests gate every commit; a cap on passes | One supervised run of the same task | outer loop |
| Outer loop (five verbs) | A trigger wakes a system that finds work, assigns it, checks it, records it and decides what is next | Recurring work with a goal held outside the model and a check a machine can run | For each wake: s for every assigned task, plus the checker | “Done” decided by the maker; merged work nobody read | An oracle outside the maker; budgets per pass and per loop; stuck-detection; a tested kill switch | The same five verbs run from a backlog by a person | outer loop |
| Shared signal store | Several loops read and write one on-disk store of typed, deduplicated records | Two or more loops keep meeting the same facts | A few reads and writes for each pass | One wrong record steers every loop that reads it; duplicates pile up | A dedupe rule (raise a count, never add a twin); a provenance link on every record | One loop with its own state file | outer loop |
Where do these patterns and names come from?
The workflow rows and the orchestrator were named in one 2024 essay, and the rest were named by practitioners whom the book’s chapters cite. The essay’s definitions are short enough to quote: “Prompt chaining decomposes a task into a sequence of steps, where each LLM call processes the output of the previous one”, and in the evaluator-optimizer workflow “one LLM call generates a response while another provides evaluation and feedback in a loop” (Schluntz and Zhang, 2024).
Chapter 10 adds the cascade and the Rule of Five to that essay’s catalog, and gives the review loop its four failure shapes. Chapter 11, in the section “Named Subagent Archetypes”, gives the reason for keeping names at all: “having names for them does the same work design-pattern vocabulary has always done: it compresses a design conversation into a sentence.”
Two things on other lists are left out on purpose. ReAct, reflection, planning and tool use describe what one agent does inside its loop; here they sit inside the single agent row. “Handoff” does not appear as a named pattern in Chapters 10 to 13, so it has no row.
How do you choose a row?
Start at the simplest row that fits the task, and move down only when a measured failure points at the next row. Four questions place a task in a family, and the “try first” column does the rest.
- Can you write the subtask list before the input arrives? If yes, stay in the workflow rows. Chapter 11, “The Orchestrator–Worker Pattern”: “if you can write the subtask list before seeing the input, use the workflow.” If no, start at the single agent loop.
- Can the branches run without seeing each other, and without writing to the same artifact? If yes, the multi-agent rows are allowed. If no, keep one continuous context. The chapter’s structural rule is “keep writes single-threaded”.
- Which actions cannot be taken back? Those get an oversight row, whatever the wiring. Chapter 12, “Approval Gates and Escalation”: “the gate must live in your code, never in the agent’s judgment”.
- Can a machine verify “done”? Only then is an outer-loop row on the table. Chapter 13, “The Goal Function”: “A loop is exactly as trustworthy as its oracle.”
The source essay puts the starting point in one sentence: “we recommend finding the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all” (Schluntz and Zhang, 2024). Chapter 10 says the same of the first row: “Exhaust the box before you reach for the wiring.”
That first row is the unit every other row arranges. In Chapter 10’s words, “what an arrangement adds is structure, never new intelligence”. The longer treatments of questions 1 and 2 are the agents versus workflows decision rule and the signs for when not to use AI agents.
What does a stack of rows cost? A worked example
A design is a stack of rows, and its cost is the sum of their call counts. The example below is illustrative: the task is invented and the 94% per-call figure is a round number chosen for the arithmetic, not a measurement of any model.
A proposal for handling billing disputes reads “an orchestrator, three specialist workers, a critic”. In the table’s terms that is orchestrator-worker plus a fresh-window reviewer. With five steps per worker it costs 1 + 3 × 5 + 1 = 17 calls, and the reviewer makes 18.
The simpler row that fits is a router with a four-station chain behind it, because the dispute types and the stations can be listed in advance. That is 1 + 4 = 5 calls. The proposal costs 18 ÷ 5 = 3.6 times as many calls.
Now suppose every call must be right for the answer to be right, and each is right 94% of the time, independently. The 18-call design succeeds 0.94¹⁸ ≈ 33% of the time and the 5-call design 0.94⁵ ≈ 73%. One retry behind a real check lifts each call to 1 − 0.06² = 99.64%, and the two designs to about 94% and 98%.
With JavaScript on, the Compounding error calculator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Compounding error calculator on its own page to share a result by link.
Both assumptions are simplifications. Independence is optimistic, because an early error sits in the context and makes later steps worse. Requiring every call to be right is pessimistic for an agent that notices a bad result and recovers. The comparison still shows the shape of the choice: the longer stack needs a check at every seam to reach what the shorter one reaches with fewer.
Call counts are also not the token bill. One team that built a multi-agent research system reported of its own workload that “agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats” (Anthropic, 2025). That is one team’s figure from 2025, quoted for direction. Chapter 11 draws the lasting point from it: “Parallelism buys wall-clock speed and breadth of coverage; it never buys efficiency.”
Where is each pattern explained?
Each family has a chapter in the book and one or more posts on this site, and this cheat sheet is the index to them. The guide to agent patterns collects the same posts in reading order.
| Family | Book chapter and sections | Longer posts |
|---|---|---|
| Workflow | Chapter 10: “The Augmented LLM as the Base Unit”, “Prompt Chaining, Routing, and Parallelization”, “The Evaluator–Optimizer Loop” | Agentic workflow patterns, with a gate at each seam; a post on the evaluator optimizer pattern is planned |
| Single agent | Chapter 3 for the loop; Chapter 11’s verdict starts from it | What an agent loop is; stop conditions for agent loops; the AI agent architecture guide |
| Multi-agent | Chapter 11: “The Orchestrator–Worker Pattern”, “Named Subagent Archetypes”, “Human-Orchestrated Parallel Agents”, “The Merge Wall and the Single-versus-Multi Verdict” | The orchestrator worker pattern and its brief; single agent vs multi agent; multi-agent system overhead |
| Oversight | Chapter 12: “Approval Gates and Escalation”, “Reviewing Intent and Outcomes, Not Lines”, “Interruptibility, Steering, and the Autonomy Dial” | Approval gates by consequence; the autonomy slider; review theater |
| Outer loop | Chapter 13: “A Short History of the Loop”, “The Goal Function”, “Loop Engineering”, “When Loops Compound”, “Keeping the Loop Safe” | An AI agent as a state machine |
Glossary entries for the terms a review will use: augmented LLM, evaluator-optimizer, orchestrator-worker, approval gate, autonomy dial, outer loop and goal function.
A review card to paste into a design document
The card below has the table’s fields as blanks, one block for each row in the proposed stack. A proposal that cannot fill the “check” line for a row is not ready to build that row.
DESIGN REVIEW CARD: one block per pattern in the stack
Pattern (row name from the cheat sheet):
Who decides the next step (code / model / person):
Why this row: the property of the task that calls for it:
Simpler row tried first, and the measured failure that ruled it out:
Calls per request (fill in k, n, N, s, W or P):
Characteristic failure we expect:
The check that catches it, and where it runs:
What happens when the check fails (retry / fall back / stop / ask a person):
Actions that cannot be undone, and the gate on each:
Who reads the output, and how long that takes per day:
Whole stack
Total calls per request (sum of the blocks):
Calls for the simplest stack that fits:
Ratio of the two:
Two of the lines come from readers’ own complaints. One forum commenter asked: “Is there any proof that these multi agent orchestrators with fancy names actually do anything other than consuming more tokens?” (Hacker News, January 2026). The “simpler row tried first” line is where a proposal gives that proof.
Another described a chain of planning, implementation and review agents: “where did the bug originate? Which agent’s context was polluted? Good luck tracing that” (Hacker News, February 2026). The “check, and where it runs” line is the answer a design owes to that question.
Where does the cheat sheet fall short?
The table compresses, and four things are lost in the compression.
The costs are counted. No row’s call count was measured, and none includes tokens, tool time or the person’s hours. A real bill comes from traces of your own runs.
The failure column lists the characteristic failure, not all of them. For multi-agent systems alone, a study of annotated traces “across 7 popular MAS frameworks” identified “14 unique modes, clustered into 3 categories: (i) system design issues, (ii) inter-agent misalignment, and (iii) task verification” (Cemri et al., 2025). The table has room for one or two.
The multi-agent rows are contested. An essay from a team that builds a coding agent argued that “in 2025, running multiple agents in collaboration only results in fragile systems” (Yan, 2025); the research-system report above describes a multi-agent design its builders kept. Chapter 11 reads the disagreement as a property of the workload: read-heavy work with independent branches suits a fan-out, and work that mutates one shared artifact does not.
The rows are not exclusive. A router’s branch can be a chain, a worker is a single agent loop, and an outer loop launches any of them. The table names the parts; the stack is yours to draw.
The takeaway
Take the table to the next design review and ask the row questions in order: which rows are in the stack, what each costs in calls, what its characteristic failure is, and which check catches it. Then ask the one that the AI agent design patterns cheat sheet exists for: which simpler row was tried, and what measured failure ruled it out. A proposal with good answers has earned its extra calls.
The patterns are worked through in Chapters 10 to 13, starting with Chapter 10, “Workflows and Composition Patterns”, in the full book. The free guide to agent patterns gathers the related posts, and you can see the formats.
Questions readers ask
- What are the main AI agent design patterns?
- Chapters 10 to 13 of the book give five families. Fixed workflows: one augmented call, prompt chain, routing, cascade, parallelization by sectioning and by voting, the evaluator-optimizer loop and iterated self-review. Then the single agent loop. Multi-agent: orchestrator-worker, read-only specialists, a fresh-window reviewer, human-orchestrated parallel agents and a swarm with a merge queue. Oversight: approval gates, an ask-a-human tool, escalation triggers, plan-level approval and the autonomy dial. Last, the outer loop.
- Which agent design pattern should I start with?
- Start with one model call that has the retrieval, tools and memory it needs, measured on a labeled set of real inputs. Move to a chain, a router or a fan-out only when the failures you measured match that shape, and to a single agent only when you cannot write the subtask list before the input arrives. Multi-agent rows come after a single agent, for work whose branches run without seeing each other.
- How much more does an agent pattern cost than a single call?
- Counted in model calls: a chain of k stations costs k, a router and its branch 2, a vote of n draws n, a review loop capped at N rounds up to N calls with a check in code or 2N with a model critic, an agent run of s steps s calls, and an orchestrator with W workers 1 + W × s + 1 for each wave. Tokens grow faster than calls in any loop, because each call re-reads the history.
- Is ReAct or reflection an agent design pattern?
- They are patterns of a different kind. ReAct, reflection, planning and tool use describe what one agent does inside its loop. The rows in this cheat sheet describe how calls and agents are wired together, who decides the next step, and where a person or a check sits. A single agent loop that reasons and then acts is one row here, and a review loop is the evaluator-optimizer row.
- Do oversight patterns such as approval gates belong in a list of agent design patterns?
- Yes, because they are design decisions with a cost and a failure like any other. An approval gate costs no model calls, it costs one human decision and a wait for each gated action, and it fails by being applied so widely that approvals become reflex. It stacks on any wiring: a chain, a single agent or an unattended loop can all carry a gate on their irreversible actions.
Sources
- Erik Schluntz and Barry Zhang (Anthropic) (2024). Building Effective Agents (page prints "Published Dec 19, 2024"; read 7 October 2026)
- Anthropic engineering (2025). How we built our multi-agent research system (page prints "Published Jun 13, 2025"; one team's report on its own workload)
- Mert Cemri et al. (2025). Why Do Multi-Agent LLM Systems Fail? (arXiv:2503.13657, submitted 17 March 2025)
- Walden Yan (Cognition) (2025). Don't Build Multi-Agents
- Hacker News commenter agluszak (2026). Forum comment asking for proof that orchestrators do more than consume tokens (11 January 2026)
- Hacker News commenter vincentvandeth (2026). Forum comment on tracing a bug across handoffs (23 February 2026)