The difference in agents vs workflows comes down to one question: who decides the next step? In a workflow, your code does, along paths you drew before any request arrived. In an agent, the model does, at runtime, in response to what it just observed. Cost, latency, risk and the way you test the system all follow from that one answer.
That question is worth asking before you write any code, because the label on a system tells you very little. Gartner predicted in June 2025 that over 40% of agentic AI projects will be canceled by the end of 2027, citing “escalating costs, unclear business value or inadequate risk controls,” and estimated that only about 130 of the thousands of vendors selling agentic AI are real (Gartner, 2025). In the same release, Gartner analyst Anushree Verma put the problem plainly: “Many use cases positioned as agentic today don’t require agentic implementations.” This post gives you the rule I use to tell which ones do, then walks a real-shaped task through it.
Agents vs workflows: what is the one question?
The one question is whether the model or your code owns the control flow, meaning the decision about which step runs next. Everything else people list when they compare agents vs workflows (autonomy, flexibility, reasoning, tool use) is a consequence of that ownership, not a separate property you can mix and match.
The cleanest published version of this line comes from an engineering essay that much of the field now cites. Workflows, it says, are “systems where LLMs and tools are orchestrated through predefined code paths,” while agents are systems where models “dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks” (Schluntz and Zhang, 2024). A shorter definition has since become common jargon: “An LLM agent runs tools in a loop to achieve a goal” (Willison, 2025).
That loop is the giveaway in agents vs workflows. Somebody has to decide whether to go around again, and in an agent that somebody is the model.
What makes this distinction useful, rather than academic, is that both systems can be built on exactly the same model. Swapping the model changes neither the category nor the bill’s shape. Moving the control flow does. For the fuller definition, see what an AI agent actually is, and the glossary entries for agent and workflow.
Can you draw the flowchart before the request arrives?
If you can draw the complete control-flow diagram before the first request arrives, you are building a workflow; if the diagram can only be drawn afterward, from the trace of what the model chose, you are building an agent. That test turns an abstract definition into something you can check at a whiteboard in five minutes.
The book states it as a litmus test in Chapter 1: “If you can confidently draw the control-flow diagram before the request arrives, you are looking at a workflow. If the diagram can only be drawn in hindsight, once the model has reacted to what it found, you are looking at an agent.” The test is not unique to the book; some vendor guides now use a similar flowchart check (Wallace, 2026). What most of them skip is what to do with the answer, which is the rest of this post.
Try it on a system you know: take a pen and draw boxes for every model call and arrows for what happens between them. If every box gets a name, you hold a workflow. If a region of the page stays blank because its steps depend on what the model finds, mark that region. It is the only part of the system that might need an agent. To run the same test as a checklist, the Should this be an agent? decision tool walks you through the book’s signs one question at a time.
A convention from the book’s workflow chapter makes the drawing honest: “the arrows belong to your code, and the shaded boxes belong to the model.” In a workflow, the arrows are ordinary programming, testable and replayable. In an agent, the model draws the arrows as it goes.
Agents vs workflows vs chatbots: how do they compare?
A chatbot, a workflow and an agent differ along one axis, who decides the next step, and that single difference sets their cost shape, their latency, the kind of damage a mistake does, and how expensive it is to know whether the system worked. The table below is my compression of the four-line “quote” the book prices in Chapter 14 (in the full book).
| Chatbot or single call | Workflow | Agent | |
|---|---|---|---|
| Who decides the next step | The human, one message at a time | Your code, on paths drawn in advance | The model, at runtime |
| When the flowchart can be drawn | No flowchart: one call, then wait | Before the request arrives | Only afterward, from the trace |
| Running cost | Flat: one call per request | Adds up: a known number of calls | Multiplies: unknown step count, growing context each turn |
| Latency | One round trip | A sum over a known count, quotable before launch | Not quotable; varies run to run |
| What a mistake looks like | Words: one wrong answer | Contained: an error stuck in its step, caught at the seam | Deeds: an action, then more actions built on it |
| Cost of knowing it worked | Sample and grade against an eval set | Test each step against its contract, plus end-to-end checks | A full evaluation instrument, traces and human review |
| Fits | Lookup, classification, transformation | Known steps, audit requirements, tight budgets | Open-ended goals with a verifier |
Two rows deserve a second look. The “mistake” row is why an agent’s risk is different in kind, not just in degree: the bounded failure of a chatbot is exactly what an agent gives up. The “knowing” row is the one teams leave off their estimates.
On the lower rungs, proof of correctness comes almost bundled with the artifact; at the top, the proof is a second construction you build alongside the agent, and it is often the larger of the two. The compounding error calculator shows why: chain enough unverified steps and even a high per-step success rate decays fast.
The chatbot column has its own post, AI agent vs chatbot, which uses the same question to separate the two.
What is the escalation ladder, and why start at the bottom?
The escalation ladder is a four-rung ordering of designs (plain code, a single model call, a workflow, an agent) in which each rung buys adaptability and pays for it in money, latency and predictability. The discipline is to start at the bottom and climb only when a rung provably fails on real inputs.
At the bottom sits a regular expression, a SQL query or an if statement. Free, instant, deterministic, auditable. The second rung is one well-crafted model call, perhaps with retrieval and a few examples in the prompt. The book is blunt about how often teams stop too high: “a surprising fraction of ‘we need an agent’ dissolves into ‘we needed a better prompt.’” The same essay quoted above agrees that “optimizing single LLM calls with retrieval and in-context examples is usually enough” for many applications.
The third rung is a workflow, several calls on paths your code drew. The common shapes are chaining, routing, parallelization and an evaluator loop with a gate, which the companion post on agentic workflow patterns walks through. The top rung is the agent, which needs a harness around the loop: tools, context management, stop rules, approval gates, tracing and evaluations.
Most agents vs workflows comparisons skip the two rungs below both, and that omission is expensive. The rungs are destinations, not waypoints. For a problem whose steps you can draw, a lower rung is the permanent answer, and staying there is the engineering choice, not a timid one.
Worked example: does a refund inbox need an agent?
A refund inbox is a good test case because it sounds agentic and is mostly not: run it through the rule and nearly every step lands on the bottom three rungs, with one narrow branch that might earn an agent. Here is the walk, step by step, with an illustrative policy.
Suppose the request on your desk reads “build an agent that handles refund requests from the support inbox.” The policy, for illustration: unused items returned within 30 days of delivery are refunded automatically below a set amount; anything above it goes to a person. Before choosing an architecture, write down what has to happen to one email, then ask the question of each step.
Which steps are plain code?
The eligibility checks are plain code: whether the order is within the window, whether the amount is under the threshold, whether the item was already refunded. Each is date arithmetic or a database lookup. No model should touch them, because a model can only make them slower, costlier and occasionally wrong.
Which steps need one model call?
Turning a free-text email into structured fields (order number, reason, claimed condition of the item) needs a model, because the input is unstructured language. It needs one call with an output schema, not a loop. Drafting the reply is a second single call. Neither step decides what happens next; each fills in content and hands it back to code.
Can you draw the whole flowchart?
Yes, for the main path. Extract the fields (model), look up the order (code), check the policy (code), draft the reply (model), check the draft against a short rubric (a gate), then either send it or queue it for a person. Every arrow is drawn before the first email arrives, and every arrow belongs to your code. That is a workflow, and for the large majority of emails it is the whole system.
Is any step genuinely agentic?
One branch might be. Picture the email that says the carrier marked the parcel delivered, the customer never received it, and a previous support conversation already promised a refund. Nobody can list in advance which records to open: carrier tracking, earlier tickets, a second order, a payment dispute. That blank corner of the diagram is the only candidate for an agent, and the book’s three conditions decide whether it qualifies: a success criterion a machine can check, a feedback signal at each step, and a sensible place for a person to look.
Here the feedback signal exists (each lookup returns real data) and a place for a person is easy to add. A fully machine-checkable success criterion does not exist, and money moves, so I would keep the agent in a propose-only posture: it investigates and recommends, and a person approves the refund.
Version one should not build even that. Send the exceptions to a human queue, count them, and build the agentic branch only when the volume justifies it and a measured comparison shows the agent does better. The agent cost-per-task estimator helps you price that branch before you commit.
What comes out is the arrangement the book calls the dominant production shape: “a workflow shell with one or two genuinely agentic steps inside it: predictable structure wherever you can predict, model-owned autonomy only at the junctures you cannot enumerate.” The request said “agent.” The rule said workflow, with one supervised agentic branch, deferred until evidence asks for it.
What is agent-washing, and what does it cost in each direction?
Agent-washing is labeling a system an agent because the word sells, regardless of who owns the control flow underneath. Gartner defines it as “the rebranding of existing products, such as AI assistants, robotic process automation (RPA) and chatbots, without substantial agentic capabilities” (Gartner, 2025). The book adds a point the industry discussion usually misses: the mislabel bills you in both directions.
| Direction | What it looks like | What you pay |
|---|---|---|
| Workflow sold as an agent | A fixed pipeline with “agent” on the slide | Oversight ceremony, governance review and harness work sized for autonomy the system lacks; evaluation against the wrong bar |
| Agent sold as a pipeline or “automation” | A model quietly owns the control flow somewhere inside | Budgets, stop conditions, security review, an eval set and human checkpoints never get bought, because the label said they were unnecessary |
That second row is the dangerous one. An undercharged agent ships without the disciplines autonomy demands, and its failures are deeds rather than words, bounded only by whatever blast radius you happened to enforce. The overcharged workflow merely wastes money, and it can make a rigid pipeline look miraculously reliable for the wrong reasons.
The cure is the same question, asked of the system in front of you regardless of its branding. Chapter 14 ends its treatment of the trap this way: “The label decides which chapter of this book you believe applies. The control flow decides which one actually does.”
In practice, I ask for two traces from two different inputs and compare them. If the sequence of steps can differ and the model chose the difference, there is an agent in there; if not, there is a workflow, whatever the brochure says. The glossary entry on agent-washing collects the book’s other references to the term.
Where does the rule get blurry?
The rule gets blurry wherever the model makes some decisions but not all of them: routers, orchestrators and loops that code bounds. In those cases the honest answer is a position on a dial rather than a box, and the question becomes how much of the control flow the model owns.
A router is the easy case. One model call picks among branches you already wrote, so every branch is on the diagram before the request arrives, and the book places it one small step up the autonomy dial from a fixed pipeline. See the glossary entry for routing. An evaluator loop that regenerates a draft until a gate passes, capped at a fixed number of rounds, is also drawable in advance; Chapter 10 catalogs it among the workflow patterns.
The orchestrator is the genuine edge. The essay that popularized the workflow-agent split files orchestrator-workers under workflows, yet says its subtasks “aren’t pre-defined, but determined by the orchestrator based on the specific input” (Schluntz and Zhang, 2024). By the flowchart test, you can draw the frame of that system in advance but not its boxes, so the decomposition step is agentic.
I don’t think this breaks the rule; it shows why some in the field prefer treating systems as “agent-like to different degrees” rather than forcing a binary (Ng, 2024). Ask the question per step, not per system.
The guides also disagree in one useful place. A widely read vendor guide lists “heavy reliance on unstructured data” as one of three signs that a task warrants an agent (OpenAI, 2025). Under the control-flow rule, unstructured input justifies a model call, which is rung two; it says nothing about who should own the next step. The refund example shows the difference: the email is unstructured, and the system is still a workflow.
How do you defend a workflow-first design in review?
You defend a workflow-first design by showing the flowchart, pricing each rung, and naming the specific class of inputs that would justify climbing higher. That turns an agents vs workflows argument from a matter of taste into a claim someone can test.
In practice, I bring four things to the review. The first is the drawn diagram, with any blank regions marked. The second is the four-line quote per rung: what it costs to build, to run, to be wrong, and to know whether it worked. The third is the climbing criterion, stated in advance: “if more than an agreed share of exceptions defeat the workflow and a measured agent does better on them, we add the branch.”
The fourth is the descent plan, because the ladder runs both ways. Chapter 14 describes using an agent as an instrument of discovery: run it while the flowchart is unknown, read its traces, and when run after run visits the same steps in the same order, freeze that stretch into workflow steps.
The strongest objection is “it might need to adapt someday.” The book treats that as a hypothesis, cheap to test honestly: build the simpler version, measure it on real cases, and climb only when you can point at inputs it provably fails on. Engineers who also own the backend will find more of this framing in AI agents for backend developers, and the case for staying low has its own checklist in when not to use AI agents.
There is a limit to all of this. The rule tells you which architecture a task wants; it does not tell you whether the task is worth doing, and it cannot rescue a request that is not yet a task. “Automate our customer support” has no flowchart because, as Chapter 14 puts it, “It is a department wearing the grammar of a task.” Decompose it first, then ask the question of each piece. The agent loop explainer shows what the top rung looks like once you get there.
The takeaway
Near the end of its first chapter, the book offers a sentence I’d pin above any backlog: “An agent, then, is a cost you pay for adaptability you can name. If you cannot name the adaptability, keep your money.” For any agents vs workflows decision, ask who decides the next step, try to draw the flowchart, and stop at the lowest rung that works. If you want a second opinion before the design review, run the task through the decision tool.
Chapter 1, “What Is an Agent?” is free to read online and contains the full taxonomy, the autonomy dial and the ladder. Chapter 10 (workflow patterns) and Chapter 14 (the full cost of each rung and the agent-washing trap) are in the full book; see the formats.
Questions readers ask
- What is the difference between an AI agent and a workflow?
- In a workflow, model calls and tool calls run along paths your code defines in advance; the model fills in content but code decides what happens next. In an agent, the model decides the next step at runtime, based on what it just observed, and keeps looping until it judges the goal met or a stop rule ends the run.
- What is an agentic workflow?
- The phrase usually means a workflow that contains model calls, sometimes with one or two steps where the model makes a decision, such as a router choosing between fixed branches. By the control-flow test, it is still a workflow as long as your code owns the arrows between steps.
- When should you use an AI agent instead of a workflow?
- When the steps depend on what the system discovers along the way, so the flowchart cannot be drawn in advance, and when three conditions hold: a success criterion a machine can check, a feedback signal at each step, and a sensible place for a human to review. Otherwise a workflow or something simpler is cheaper and easier to verify.
- What is agent-washing?
- Agent-washing is labeling a system an agent for its sales value rather than its architecture. Gartner used the term in June 2025 for vendors rebranding assistants, robotic process automation and chatbots without substantial agentic capabilities. The cure is to ask whether the model or the code decides the next step.
- Is a router or an orchestrator an agent?
- A router makes one model decision among branches you already wrote, so the full diagram exists in advance and the system is a workflow. An orchestrator that decides at runtime how to split a task sits further along the autonomy dial: you can draw its frame in advance but not its boxes.
Sources
- Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents
- Gartner (2025). Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027
- Andrew Ng (2024). Welcoming Diverse Approaches Keeps Machine Learning Strong (The Batch, issue 253)
- OpenAI (2025). A practical guide to building agents
- Simon Willison (2025). I think “agent” may finally have a widely enough agreed upon definition to be useful jargon now
- Jim Allen Wallace (Redis) (2026). AI Agents vs Workflows: When to Use Each