Agentic workflow patterns are fixed arrangements of model calls whose control flow lives in your code. The four shapes are a pipeline, a dispatcher, a scatter-gather and a review loop, with a model at each station. Because every station is probabilistic, each seam between stations needs a check you would never put between two functions.
So yes, it is a pipeline. If you already build staged jobs, dispatchers, fan-out and bounded retries, you know the arrows. This post covers what changes when the box is a model: the check between stations, the cost in calls, the failures, and which shape to try first when a single call is right about four times in five.
Whether you need a workflow at all is a separate decision, and the agents vs workflows decision rule owns it: draw the flowchart before the request arrives, and stop at the lowest rung that works.
What are agentic workflow patterns?
Agentic workflow patterns are the small set of ways to wire several model calls together when your code, not the model, decides which call runs next. A 2024 engineering essay named the set, and Chapter 10 of AI Agents, Engineered (in the full book) catalogs the same shapes.
The unit they arrange is one model call with three things attached: retrieval (documents fetched for this request), tools (functions the model can ask your code to run) and memory (what is carried over from earlier turns). The essay calls it the augmented LLM: “The basic building block of agentic systems is an LLM enhanced with augmentations such as retrieval, tools, and memory” (Erik S. and Barry Zhang, 2024). I’ll call one such call a station, and the handoff between two stations a seam.
Some boxes are model calls and every arrow is still yours. The chapter’s opening says why that matters: “The arrows are ordinary programming: testable, replayable, immune to persuasion.”
Is an agentic workflow just a pipeline?
In structure, yes: each of the four shapes has a counterpart in long-standing integration patterns. What is new is the box. A model station can return well-formed output that is wrong, without raising, and a retry is a fresh draw, not a replay.
The mapping is this post’s. The middle column quotes the Enterprise Integration Patterns catalog, from its pages on Pipes and Filters, the Content-Based Router and Scatter-Gather. The last column is mine.
| Shape | You already call it | The integration pattern says | What a model in the box changes |
|---|---|---|---|
| Chain | A staged pipeline with validation between stages | Pipes and Filters: “divide a larger processing task into a sequence of smaller, independent processing steps (Filters) that are connected by channels (Pipes)” | A stage can return well-formed, wrong output; a retry is a new draw |
| Route | A dispatcher: classify, then switch |
Content-Based Router: “route each message to the correct recipient based on message content” | The routing function is a classifier with an error rate, not a field lookup |
| Parallelize | Scatter-gather, or fan-out and fan-in | Scatter-Gather: “broadcasts a message to multiple recipients and re-aggregates the responses back into a single message” | Branches settle unstated shared decisions differently; copies of one model are not independent witnesses |
| Evaluate | A bounded retry with the failure passed back | No single catalog entry that I found | The checker may itself be a model with an error rate |
Three of the four are acyclic. The review loop is the one cycle, which is why it is the one that needs a cap.
Which names does this post use?
This post uses four verbs: chain, route, parallelize, evaluate. The essay lists five workflows. The book treats four as fixed workflows and moves the fifth to its multi-agent chapter, because there a model writes the subtask list.
| The 2024 essay | Chapter 10 of the book | This post |
|---|---|---|
| “Building block: The augmented LLM” | The augmented LLM, with three “ports” | A station |
| Prompt chaining, with “programmatic checks” between steps | Prompt chaining, with the gate as a defined term | Chain |
| Routing | Routing, plus a variant it calls the cascade | Route |
| Parallelization: sectioning and voting | The same names; its recap calls the shape “the fan-out” | Parallelize: section or vote |
| Evaluator-optimizer | Evaluator–optimizer, also “generator–critic” and “propose-and-check” | Evaluate, or the review loop |
| Orchestrator-workers | Moved to Chapter 11, as the seam into multi-agent systems | Not covered here |
The list of agentic workflow patterns is a convention, not a law. The same lab’s March 2026 post says “three patterns cover the vast majority of use cases: sequential, parallel, and evaluator-optimizer”, and has no routing section. Lists of reflection, tool use, planning and multi-agent collaboration are a different thing under a similar name: they describe what an agent does, not how calls are wired.
The four agentic workflow patterns in one table
Each of the four agentic workflow patterns answers one way a single call gets overloaded, introduces one characteristic failure, and has a price you can count before you build it. The table is this post’s synthesis, the fixed-workflow rows of a longer AI agent design patterns cheat sheet: the cost column is counted for this post, not measured.
| Pattern | Backend equivalent | Fits when | The failure it introduces | Calls and latency vs one call | How to test the seam |
|---|---|---|---|---|---|
| Chain, k stations | Staged pipeline with validation | One call must do several jobs in a fixed order, and each handoff can be checked | An early error elaborated by every later station | k calls, about k× latency. Three stations: 3 calls, about 3× | A contract test per station on recorded inputs; a gate per seam; rejections counted per seam |
| Route | Dispatcher | Inputs come in kinds that need different prompts, tools or models | A confident answer from the wrong branch | 2 calls in sequence (classifier, then branch), up to 2× latency; 1 call if the classifier is code | A labeled set, a confusion matrix, an “other” label with a fallback, a gate on the label |
| Parallelize: section, k branches | Scatter-gather over a list | One step has several independent parts, each asking a different question | Pieces that disagree on a decision nobody stated | k calls at once, about 1× latency; add 1 call and 1× if a model merges them | Merge fixtures with inconsistent pieces; a rule for a missing branch |
| Parallelize: vote, n draws | Redundancy with a quorum | One verdict matters too much for one draw, every draw asks the same question, and answers compare by equality | Unanimous and wrong | n calls at once, about 1× latency. Three votes: 3 calls | Agreement rate; unanimous-and-wrong rate on a labeled sample; threshold set from the two error costs |
| Evaluate, cap of N rounds | Bounded retry with feedback | A check in code or a calibrated critic can say what to fix, and one retry was not enough | A loop that approves everything, alternates or never exits | Code evaluator: up to N calls, about N×. Model critic: up to 2N calls, about 2N×. Cap of 3: 3 or 6 calls | Known-bad candidates must fail; a round cap; a no-progress stop; a held-out check |
How are the multiples counted?
They are counted by hand, on one assumption: every call takes about as long and costs about as much as the single call it replaces. For a chain that is a ceiling, since narrower stations read less. A fan-out’s wall-clock time is about its slowest branch, but each branch re-reads whatever you send it, so the bill follows the branch count.
One unit serves the whole post. A round is one generate plus one check, and the first attempt is round one. A cap of N rounds is therefore N − 1 retries, and a gate with one retry is a cap of 2.
Composition adds up. Take a router (1 call) whose branch is a three-station chain (3), with a three-way vote on one station (2 more) and a two-round loop with a model critic on the last (generate, evaluate, generate, evaluate: 3 more). That is 1 + 3 + 2 + 3 = up to 9 calls for one request, and 7 if the first candidate passes.
What is a gate, and what can it check?
A gate is the check that stands between two stations: code that inspects one station’s output before the next station reads it. Chapter 10 defines it in its section “Prompt Chaining, Routing, and Parallelization”: “A gate is a check written in ordinary code and stationed at a seam”.
A gate can be one of three things, cheapest first:
- A schema. The output parses, required fields are present, and a label is one of the allowed values (structured output makes this cheap).
- A rule. A length limit, a required disclaimer, a total that equals the sum of its lines.
- A check against the world. The order id exists, the cited policy exists, the query runs, the tests pass.
Where no code can check the property (tone, faithfulness to a source), the seat goes to a second model with a rubric. That is no longer a gate in the chapter’s sense: its verdict is an estimate with its own error rate, and it needs the calibration described under the review loop. This post calls it a critic.
A gate sees form more easily than substance. The chapter: “a schema check certifies form and never substance, and a subtly wrong reply sails through a length check.” It can also be wrong the other way and reject good output, so keep known-good outputs it must pass beside the known-bad ones it must reject.
So each station also needs an eval set: recorded inputs with expected outputs, or for free text expected properties, rerun on every prompt change.
What happens when a gate fails?
The bad output stops at the seam, and one policy decides what comes next. The book lists the options: “A gate that fails can retry the step, fall back to something safe, or stop and summon a human; what matters is that the bad intermediate never reaches the next box.”
This post uses one policy everywhere: on fail, retry once with the failure message attached; if that fails too, fall back or stop and ask a person. A retry with nothing attached is, in the chapter’s words, “a fresh draw from the same lottery”. One retry is still a gate. Allow a second and you have a loop, the evaluate pattern, which needs the loop rules below.
I would add a third outcome to pass and fail: could not check, for a check that did not finish, such as a lookup service that was down. Send it to the fallback or to a person. If the output is what ran out of time, as with a generated query that exceeds its time limit, score it a fail.
Two cautions carry over from any backend retry. A station that called a tool with a side effect repeats it when retried, so retry only stations whose tools are read-only or idempotent. A check that executes the output (a query, a test run) executes model-written code, so run it read-only, in a sandbox, with a timeout.
Chain: what do unchecked seams cost?
An unchecked chain multiplies its stations’ success rates, so it can be less reliable than the single call it replaced. The essay’s definition: “Prompt chaining decomposes a task into a sequence of steps, where each LLM call processes the output of the previous one.”
Take a call that is right 80% of the time, and suppose splitting it into three narrower stations lifts each to 90% (illustrative). With nothing between them the chain succeeds 0.9 × 0.9 × 0.9 = 73% of the time, below where it started and at about three times the latency. The chapter’s sentence for this: “A chain without gates is just a slower way to be wrong; the gates are what you are actually building.”
The calculator below opens on four stations at 95% each with a 90% target. It shows 81% whole-run success (0.954 = 0.8145). Its verdict says each step would need to succeed 97.40% of the time, and that at 95% per step the longest chain that meets the target is 2 steps.
With JavaScript on, the Compounding error calculator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Compounding error calculator on its own page to share a result by link.
Open its “Try the levers” panel for the row that matters here: with one verified retry per step, the per-step rate becomes 99.75% and the whole run 99.0%. Read that as a ceiling. The tool assumes every step has the same rate and fails independently, and that the check is perfect: it catches every failure, and the retry is an independent draw.
Why long agent runs often do worse than the multiplication, and when a checked chain does better, is the subject of why agent errors compound. A chain is the wrong shape when you are inventing stages, and the book names the sign: “the tell is gates you cannot write, because a seam with no checkable contract is a seam the task did not actually have.”
Route: how do you test a router?
Test a router as what it is, a classifier: on a labeled set of real inputs, on its own, before you tune any branch. The essay defines routing this way: “Routing classifies an input and directs it to a specialized followup task.” It adds a condition that is easy to skip, that routing works “where classification can be handled accurately”.
A wrong route sends the input to a specialist prompt for the wrong specialty, and Chapter 10 draws the consequence: “so the router’s error rate is a floor under everything downstream, and it deserves its own measurement, its own eval set, and its own monitoring as traffic drifts.”
An illustration: if the router is right 95% of the time and the right branch succeeds 90% of the time, the pair succeeds 0.95 × 0.90 = 85.5% of the time, plus whatever a wrong branch happens to get right.
I looked for a measured misroute rate for a production router, with a stated label set and a denominator, and found none. Your own labeled set is the only number available.
What goes in the router’s test plan?
Four things go in, and none needs a model to judge it:
- A labeled set of real inputs, with inputs that belong to two categories included on purpose. The eval sample-size calculator helps decide how many.
- A confusion matrix, the table of true label against predicted label. One accuracy number hides which pairs get mixed up.
- An “other” label and a fallback, for inputs that fit no category or that the router cannot place. Send them to a general prompt or a person, and track the share.
- A gate on the label. Constrain the output to the label set and reject anything else. Then watch the mix of labels over time, because traffic drifts while prompts stand still.
An input about two things is a different case from “other”. If such inputs are common, run both branches as a fan-out instead of choosing one. When inputs differ in difficulty and not in kind, Chapter 10 covers a variant of routing, the cascade, in the same section; this post does not price it.
Parallelize: how do you combine the answers?
Combine parallel answers in code, by a rule you chose before the fan-out ran: assembly for sectioning, counting for voting. The essay’s two definitions: “Sectioning: Breaking a task into independent subtasks run in parallel. Voting: Running the same task multiple times to get diverse outputs.”
The line between them is the question each call answers. If every call answers the same question, however it is worded, you have a vote and you count answers. If each call answers a different question (a different section, a different policy clause, a different concern), you have sectioning, even when the merge rule is “flag if any call flags”.
Sectioning’s merge is assembly: concatenate in a fixed order, or fill the fields of one record. Its failure is hidden coupling. The book: “If the parallel subtasks secretly share a decision—a style, a data format, an assumption about the API—each call will settle the shared question its own way”.
So state every shared decision in the same words in every branch, or chain the coupled work.
What does voting buy, and what limits it?
Voting buys confidence in one verdict, at n times the bill, and only when the answers can be counted. Chapter 10 sets the boundary: “counting votes needs answers comparable by equality”. A label, a boolean, a number or a query’s result set can be counted; three paragraphs cannot.
If each draw is right 80% of the time and draws were independent, a majority of three is right 3 × 0.8² × 0.2 + 0.8³ = 89.6% of the time. At 40% per draw the same formula gives 35.2%: voting amplifies what the single call already is.
One measurement shows both the gain and the gap. On the 1,319 problems of a grade-school math benchmark, one chat model from early 2023 scored 76.7% with a single response and 82.5% by majority vote over three of its own responses (Huang and colleagues, 2023, Table 7). If each of the three responses was right 76.7% of the time and the draws were independent, the majority would be right at least 3 × 0.767² × 0.233 + 0.767³ = 86.2% of the time. The paper measured 82.5%. My reading, not the paper’s, is that the draws were not independent: the same problems were hard each time. That is one model on one benchmark.
Different models are not independent either. A 2025 study of over 350 models found that on one leaderboard, with 71 models answering about 14,000 multiple-choice questions, “pairs of models agree on average about 60% of the time when both models are incorrect (choosing between incorrect answers uniformly at random would lead to an agreement rate of 1/3)” (Kim, Garg, Peng and Garg, 2025). Agreement is weaker evidence than it looks, and the book prices it: “five votes cost five times the tokens, so voting belongs on the few verdicts whose stakes pay for it”.
Evaluate: why won’t my review loop stop?
A review loop fails to stop when its only exit is a verdict it cannot reach, and it approves everything when the reviewer is the author’s twin. The fourth shape, the evaluator optimizer pattern, is defined in the essay in one sentence: “one LLM call generates a response while another provides evaluation and feedback in a loop.”
The book’s rule for the evaluator–optimizer loop is separation: “The generator proposes; something that is not the generator disposes”. It also names four failure shapes. The rubber stamp passes every candidate on round one, the oscillation has each round undo the last, the over-edit polishes a good answer into a worse one, and the wrong hill climbs toward the evaluator’s blind spot.
What do the measurements say about self-review?
Three measured results bear on it. All three come from benchmark questions; none is a study of a production review loop.
Accuracy fell after self-review. A 2023 paper had models review and revise their own answers with no outside signal, and reports that “after self-correction, the accuracies of all models drop across all benchmarks” (Huang and colleagues, 2023). One frontier model of that year, on 200 sampled math problems, went from 95.5% to 91.5% after one round and 89.0% after two, at 1, 3 and 5 calls (the paper counts a round as one review call plus one revision call after the first answer). When the correct label decided when to stop, the same model rose to 97.5%. The gain came from the outside signal.
The feedback said everything was fine. The authors of a 2023 preprint on a refinement method report that on their math-reasoning task, one chat model’s feedback for 94% of instances was “everything looks good” (Madaan and colleagues, 2023). It reports gains on its other tasks.
A second model is not an independent judge. The 2025 correlated-errors study scored models on leaderboard questions with another model’s own answers standing in for the key. It found that such a judge “systematically inflates the accuracy of models that are less accurate than itself”, because it marks a wrong answer correct when it would have given the same one.
Which rules make a loop end?
Four rules do it, and they are the book’s, in my order:
- Seat an external check where one exists. “Wherever a verifier exists, seat it: the test suite, the compiler, the schema validator, the query that executes or fails to.” Pass its failure back whole.
- Use a critic where none exists. A separate call with a separate prompt and an explicit rubric, ideally a different model, which lowers the overlap and does not remove it. Calibrate it first; whether LLM-as-a-judge is reliable is its own subject, and the glossary entry has the short version.
- Cap the rounds in code. The chapter gives “three to five rounds” as “the common figure, offered as illustration”. The cap is a stop condition that no model’s opinion can move.
- Stop on no progress. End early when the candidate repeats an earlier version or the verdict stops improving, and “keep the best candidate seen, by the evaluator’s own scoring, rather than the last one produced”.
On cost, the chapter counts in rounds: “every round is at least one generate and one evaluate, so an N-round loop costs roughly N times the calls and N times the latency of a single shot.” That matches the count of model calls when code does the evaluating: N generate calls plus the check’s run time. With a model critic each round is two model calls, so the same N rounds is 2N.
Which pattern should you try first?
Try the shape that matches how your single call is failing, and before any of them, put a gate on the single call. The book’s choosing rule, from the section “Prompt Chaining, Routing, and Parallelization”: “A single call gets overloaded in three distinguishable ways”. It names them in six words: “Sequence, category, stakes. Chain, route, parallelize”.
The sequence below is this post’s arrangement of that rule, with steps added before it and the loop’s entry test after it. Stop where it says stop.
- Measure the one call on a labeled set of real inputs. Good enough: stop. Use one call.
- Check what the call was given. If its failures lacked the document, record or passage they needed, fix retrieval or the input before changing any wiring.
- Gate the one call. If code can tell a bad output from a good one, add the gate and the retry policy: retry once with the failure message attached, then fall back or stop and ask a person. Measure again. Good enough: stop.
- Check that the subtask list is fixed. You can write it in advance, or plain code can compute it from the input (one call per section, per row, per file). If only a model could write the whole list, stop: this is not a fixed workflow. If only one step needs a model to decide, keep the workflow and isolate that step: a router when it picks among branches you wrote, a small agent when it does not (the decision rule linked at the top owns that call).
- Read the failures that remain and sort them into the four bins below.
Which bin does a failure belong in?
Each failure goes in the bin whose description it matches, and two tie-breaks settle the ones that match two. The bins, each with the shape it points to:
- Fixing one kind of input breaks another: route.
- The call skips or botches one of several jobs it must do in order: chain.
- Some of several independent parts get shortchanged: parallelize by sectioning. One verdict flips between runs and a miss is expensive: parallelize by voting.
- The output breaks a rule that a check in code or a calibrated critic can name, and the one retry does not fix it: evaluate.
Build for the fullest bin first, then measure again. Route wins a tie with any other bin, because it decides which prompt the other shapes sit behind. Between chain and evaluate, pick evaluate when a check in code can name the fault, and chain when the second job needs its own prompt. The shapes nest, so a router’s branch can be a chain and one station of a chain can be a vote.
What does the gate-first step buy?
On paper it buys most of the gap, for a fraction of a call. A call that is right 80% of the time, behind a gate that catches every failure, with one retry that succeeds at the same rate, is right 1 − 0.2² = 96% of the time at an expected 1.2 calls. Both assumptions are generous, so treat 96% as a ceiling. It is still the cheapest first move.
One task, six builds
The same task can be built as one call and in each of the shapes. The example is constructed for this post: write weekly release notes from 40 merged change descriptions spread over 5 components, of which 10 are internal and should not appear. The counts are arithmetic, not logs.
| Build | What runs | The check at the seam | Model calls |
|---|---|---|---|
| One call | All 40 descriptions in, notes out | Every change id in the notes exists in the input | 1 |
| Chain | Extract user-visible changes, group them by component, write the notes | Every id exists; every extracted item sits in exactly one group; required sections present | 3, about 3× latency |
| Route | Classify each change as feature, fix, breaking or internal; each kind has its own writing prompt; internal changes are dropped | The label is one of the four | 40 classify + 30 write = 70 |
| Parallelize: section | One writing call per component, at once; code concatenates in a fixed order | Every component returned; same heading format in each | 5, about 1× latency |
| Parallelize: vote | Three differently worded calls ask the same question, “does this change break existing callers?”; flag on any yes | Each answer is yes or no | 3 per change voted on |
| Evaluate | Draft, check, revise with the failures, capped at 3 rounds, keep the best | Code: ids exist, links resolve, sections present | Up to 3; up to 6 with a model critic for tone |
Per-item builds are the expensive ones: 70 calls for the router, and the vote reaches 90 if it runs on all 30 user-visible changes. And the first row already has a gate, which is why the chooser gates the one call before it considers any shape.
The step contract card
The step contract card is one station’s interface written down: what goes in, what must come out, what checks it and what happens when the check fails. It is this post’s template, in no product’s syntax. Fill one per station, router and merge step included, and delete the lines that do not apply.
STATION <name> Kind: <model call | code>
INPUT <what it receives, and from which station or retrieval step>
OUTPUT <what it must return> Schema: <fields, types, allowed values>
IF ROUTER Label set, including "other": <labels>
IF MERGE Rule: <concatenate | fill fields | count votes, with threshold>
Missing branch: <wait | proceed without | fail the request>
GATE Checks, in code: <schema | rule | check against the world>
Outcomes: PASS / FAIL / COULD NOT CHECK (never fold the third into PASS)
If the check executes the output: <read-only | sandbox | time limit>
CRITIC Only if no code check exists: <rubric, and how it was calibrated>
ON FAIL Rounds, at most: <N, first attempt included; 2 means one retry>
Retry with the failure message attached.
Safe to retry: <yes | no: which tool writes>
If N is above 2, stop early if: <the output repeats | the verdict stops improving>
If N is above 2, keep: <the best candidate seen, by the checker's score | the last>
Then: <fall back to what | stop and send to whom>
ON COULD NOT CHECK <fall back to what | send to whom>
EXAMPLES Labeled set: <where the recorded inputs and expected outputs live>
Known-bad outputs the gate must reject: <where they live>
Known-good outputs the gate must pass: <where they live>
Rerun on: <every prompt, model or schema change>
BUDGET Calls per request: <max> Seconds per request: <max>
LOG Gate rejections and could-not-check counts, per seam
If neither the GATE line nor the CRITIC line can be filled after an honest attempt, the task may not have a seam there, and the two stations belong in one call.
What is going wrong? A symptom table
A bad run can be read backward, from what you see to the seam that lacks a check. The last column is the pattern the row is about. Rows that describe a single call also carry “one call”, so that button shows everything that applies before you have built a workflow. The mapping is my analysis; with scripts off the whole table is still here.
| What you see | What is wrong | The fix | Pattern |
|---|---|---|---|
| One call is right most of the time, and its failures break a rule you could state in code | There is no gate on the call’s output | Add a gate; on fail, retry once with the failure message attached, then fall back or stop | one call |
| The output is wrong, and the station was never given the document, record or passage it needed | Retrieval or the input is at fault, not the wiring | Log what each station received; fix what it is given before adding a station | one call |
| The workflow is slower than the single call and no more accurate on your labeled set | Stages were invented; no seam had a contract | Merge back to one call | one call |
| One call skips or botches one of several jobs it must do in order | The jobs crowd each other in one prompt | Split into stations, with a gate at each handoff | one call, chain |
| The final output is wrong, and the first station’s output was already wrong | Nothing checks the first seam, so later stations elaborated the error | Add a gate after station one; count rejections per seam | chain |
| A station’s output is well-formed and the next station fails on it, or a step was silently skipped | The gate checks shape, not the property the next station needs | Gate on the postcondition: the language changed, every category is present, every id exists | chain |
| A check could not run and the item went through anyway | “Could not check” was folded into “pass” | Three outcomes; send the third to a fallback or a person | one call, chain |
| You cannot say what the gate between two stations should check, in code or as a rubric | The task has no seam there | Merge the two stations | chain |
| Improving the prompt for one kind of input makes another kind worse | One prompt is serving several kinds | Split by kind behind a classifier | one call, route |
| A branch answers well, about the wrong topic | A misroute | Measure the router alone: labeled set and confusion matrix | route |
| A request about two things gets an answer about one | A blended input met a single-label router | Put blended cases in the labeled set; if they are common, run both branches (a fan-out) instead of choosing one | route |
| Quality fell and no prompt changed | The mix of traffic drifted, or the model behind the call changed | Monitor the label distribution; rerun the eval set; relabel a fresh sample | one call, route |
| One call covers several independent parts and shortchanges some | Too much scope for one prompt | One call per part, merged in code | one call, parallelize |
| One call’s verdict flips between runs, and a miss is expensive | One draw is carrying too much | Vote on that verdict; set the threshold from the two error costs | one call, parallelize |
| Parallel pieces contradict each other on format, terms or assumptions | The branches shared a decision nobody stated | State it in the same words in every branch, or chain them | parallelize |
| Latency is fine and the bill grew with the number of branches | Every branch re-reads the whole input | Send each branch only its slice | parallelize |
| All votes agree and the answer is wrong | The draws are not independent | Put a check that is not a model behind the vote; measure the unanimous-and-wrong rate on a labeled sample | parallelize |
| You are keeping “the best” of several candidates | That is scoring, not voting | Vote only on outputs comparable by equality; scoring needs an evaluator | parallelize |
| One call breaks a rule a check or a critic can name, and one retry with the failure attached does not fix it | One retry cannot use up what the failure says | A loop: round cap, no-progress stop, keep the best | one call, evaluate |
| Every candidate passes on round one | The evaluator is too lenient, or is the generator rereading itself | Feed it known-bad candidates; it must fail them. Track the round-one pass rate | evaluate |
| Round three undoes round two | Oscillation | Detect a candidate that repeats an earlier one; stop and escalate | evaluate |
| The loop always runs to its cap | The exit is unreachable, or the feedback names no located fault | Make the exit reachable; feed back the failing case; add a no-progress stop | evaluate |
| The evaluator’s score rises and a held-out check sees no change | The loop is climbing the evaluator’s blind spot | Keep a check the loop never sees; calibrate a model critic against human labels | evaluate |
Do you need a framework for agentic workflow patterns?
No: the four shapes are ordinary control flow, and Chapter 10 says so in its section “The Augmented LLM as the Base Unit”: “Three calls and two if statements are a chain; a dictionary lookup is a router. Write them by hand at least once.” The essay agrees: “many patterns can be implemented in a few lines of code”.
The other side, from a June 2025 forum comment: “the first benefit you get from a good framework is the easy ability to try out different (and cross-vendor) LLMs” (Hacker News, 2025). A framework vendor’s own post, from an interested party, makes the same complaint as the essay about one kind of framework, the kind that hands you a ready-made agent: “Agent abstractions can make it easy to get started, but they can often obfuscate and make it hard to make sure the LLM has the appropriate context at each step” (Chase, 2025).
A framework can add plumbing such as durable execution, tracing and model swapping, and it can hide the prompt and the response that a gate and an eval set need. The book’s rule covers both: “adopt a framework for a named piece of plumbing you would otherwise build yourself, never for the feeling that serious systems use one”.
What is counted here and not measured?
Every cost in this post is counted, not measured, and the evidence is narrower than it first looks.
The costs are arithmetic. I found no latency measurement of the four shapes and no clean comparison of one call against a chain on the same task at the same budget.
The measurements are benchmark results from 2023 to 2025, with models of those years. They show direction and mechanism, not rates for your loop.
A fixed workflow is the wrong tool when the subtask list hides in the input. The chapter’s boundary: “If you cannot write the subtask list before seeing the input, no shape in this chapter will write it for you.” That case belongs to the orchestrator worker pattern and to the single agent vs multi agent decision. Long or cyclic runs want explicit states and transitions; modeling an AI agent as a state machine is the same move for the case where the model picks the transition.
The takeaway
If you have one call that is right four times in five and a request to “make it agentic”, don’t start by adding stations. Label real inputs, check what the call was given, put a gate on it, and read the failures that remain. They tell you which of the four agentic workflow patterns to build, and a seam where you can name neither a check in code nor a calibrated critic is one you shouldn’t build. Chapter 10 closes its composition example on the same division of labor: “The arrows are why the system debugs like a program. The shaded boxes are why it needs the gates.”
The full catalog, with the cascade and the loop’s convergence rules, is in Chapter 10, “Workflows and Composition Patterns”, in the full book. The free guide to agent patterns collects the related posts, and you can see the formats.
Questions readers ask
- Is an agentic workflow just a pipeline?
- In structure, yes. A prompt chain is a staged pipeline, a router is a dispatcher, parallelization is a scatter-gather and the evaluator loop is a bounded retry with the failure passed back. What changes is the box: a model station can return well-formed output that is wrong without raising, so each handoff needs a check in code, and each decision a model makes needs a measured error rate.
- How do I check an LLM's output and retry if it is wrong?
- Put a gate after the call: code that checks the output against a schema, a rule or the outside world (does the id exist, does the total add up). Give the gate three outcomes, pass, fail and could-not-check. On fail, retry once with the failure message attached; if that fails too, fall back or stop and ask a person. Retry only stations whose tools are read-only or idempotent.
- Why does my review loop never stop, or approve everything?
- It never stops when its exit depends on a verdict it cannot reach, such as zero findings, and nothing else ends it. It approves everything when the evaluator is too lenient or is the same model rereading its own work. Use a check in code, or a calibrated critic that is not the generator, cap the rounds in code, stop when the candidate repeats or the verdict stops improving, and keep the best candidate seen.
- How do I know my router is sending requests to the right branch?
- Measure the router alone, as a classifier. Label a set of real inputs, including ones that belong to two categories, run the router over them and read the confusion matrix to see which pairs it mixes up. Add an "other" label with a general fallback, gate on the label being one of your branches, run both branches when inputs about two things are common, and watch the mix of labels over time.
- Do I need a framework to build agentic workflow patterns?
- Not for the four shapes, which are ordinary control flow: calls in sequence, a switch, concurrent calls with a merge, and a bounded loop. The essay that named them recommends starting with direct model API calls. A framework can still earn its place for a named piece of plumbing such as surviving restarts, tracing or swapping models, provided you can still read the prompt and the response at every station.
Sources
- Erik S. and Barry Zhang (Anthropic) (2024). Building Effective Agents (page dated Dec 19, 2024; the page prints the first author as "Erik S." and carries a note that its tooling has changed since; read 6 October 2026)
- Claude blog (Anthropic) (2026). Common workflow patterns for AI agents—and when to use them (no byline; dated March 5, 2026)
- Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet (arXiv:2310.01798v2; ICLR 2024)
- Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, et al. (2023). Self-Refine: Iterative Refinement with Self-Feedback (arXiv:2303.17651; preprint as fetched)
- Elliot Kim, Avi Garg, Kenny Peng, Nikhil Garg (2025). Correlated Errors in Large Language Models (arXiv:2506.07962v1; accepted to ICML 2025)
- enterpriseintegrationpatterns.com. Enterprise Integration Patterns: Pipes and Filters (pattern page, read 6 October 2026)
- enterpriseintegrationpatterns.com. Enterprise Integration Patterns: Content-Based Router (pattern page, read 6 October 2026)
- enterpriseintegrationpatterns.com. Enterprise Integration Patterns: Scatter-Gather (pattern page, read 6 October 2026)
- Harrison Chase (LangChain) (2025). How to think about agent frameworks
- Hacker News commenter smoyer (2025). Forum comment on what a framework gives you first (June 2025)