Home / Tools / Should this be an agent?

Free tool · runs in your browser · from Chapter 14

Should this be an agent?

When to use AI agents, and when plain code, one model call or a workflow does the job better. Answer ten questions and see where your task sits on the ladder.

The tool

Your inputs stay in this tab. Share a result by copying the page address: the state lives in the URL. Every question is a sign or condition from Chapters 1 and 14; the order in which the tool asks them is the tool's own.

When should you use an AI agent?

You should use an AI agent when the steps of a task depend on what is discovered along the way, a machine can check whether the result is right, each step gets a feedback signal, and a person has a sensible place to look. When you can draw the flowchart before the request arrives, plain code, a single model call or a workflow does the job more cheaply and more reliably.

That is the decision this tool walks you through. The questions come from two chapters of the book: Chapter 1, which introduces the escalation ladder and is free to read, and Chapter 14, which returns to the decision with every cost priced. The tool asks only the questions that still matter for your answers, places the task on a rung, and explains why.

What is the escalation ladder?

The escalation ladder is the book’s ordering of designs from simplest to most autonomous: plain code, a single model call, a workflow, and an agent. “Each rung up buys adaptability and pays for it in money, latency, and predictability,” and the discipline that goes with it is “to stay on the lowest rung that solves your problem.”

The escalation ladder.
Figure 1.4 The escalation ladder. Each rung up—from plain code to a single model call to a workflow to an agent—buys adaptability and pays for it in cost, latency, and unpredictability. The discipline is to stand on the lowest rung that solves the problem (in accent), and to climb only when a real case proves the rung below cannot reach. Reuse this diagram

The rungs differ in who decides the next step. In plain code and in a workflow, your code decides. In an agent, the model decides. That single difference is why the book’s litmus test works: “can you draw the flowchart before the request arrives?” If yes, the decision about the next step can live in code, where it is cheap, fast and auditable.

Rung Who decides the next step Typical task What it costs when it is wrong
Plain code Your code, fully specified Flag transfers above a threshold The same bug every time, fixed once
Single model call Your code; the model makes one judgment Classify, extract, summarize One wrong answer, nothing built on it
Workflow Your code, on paths drawn in advance Route an email, then draft a reply Contained to the step that failed
Agent The model, step by step Make a failing test pass in an unknown codebase Errors compound across steps

Which signs point down the ladder?

Five signs point down the ladder, and Chapter 14 says “any one of them comes close to deciding alone.” The tool asks about each one:

  1. The steps can be enumerated. “If you can enumerate the steps, the flowchart exists, and drawing it in code is cheaper than renting a model’s judgment to rediscover it on every request.”
  2. Correctness must be guaranteed or audited. “A code-chosen path is a document you can hand to a regulator, while a model-chosen path is a probability you must defend.”
  3. A person is waiting, or the volume is high and the margin thin. Then “every multi-call shape is priced out before quality even enters the conversation.”
  4. The work is lookup, classification or transformation. These “are single-call shapes however sophisticated the judgment inside the call.”
  5. The stakes are high and verification is absent. Then “the sound design withholds autonomy altogether”: the co-pilot posture, in which “the agent proposes, a person approves.”

The first question the tool asks comes before all of these. Chapter 14 opens with four requests, and the fourth, a request to automate customer support, “should have refused to be sorted at all. It is a department wearing the grammar of a task.” When the answer is “a whole area of work,” the verdict is to decompose it before choosing any design.

What has to be true before you build an agent?

Before you build an agent, the goal must be open-ended, judgment must be required “at branches you cannot enumerate,” and three conditions must hold: “A success criterion a machine can evaluate,” “A feedback signal at each step,” and “A sensible place for a human to look.” The book’s conclusion is direct: “If the three conditions hold, and the task’s value clears the four-line quote, buy the agent with a clear conscience.”

The four-line quote is Chapter 14’s way of pricing a rung: “what it costs to build, what it costs to run, what it costs when it is wrong, and what it costs to know whether it worked.” The fourth line is the one teams forget, and it is the line that grows fastest as you climb. An agent’s correctness is a rate that has to be measured for as long as it runs; the eval sample-size calculator shows how many runs that measurement takes.

When some conditions hold and others do not, the tool keeps a person in the loop and names what is missing. Coding is the domain where all three conditions come free, the book observes, which is why coding agents arrived first. In most other domains someone has to build the criterion, the signal and the review point before autonomy is safe.

Which design ships most often?

The design that ships most often is neither a pure workflow nor a pure agent. Chapter 1 and Chapter 14 both describe it: “a workflow shell with one or two agentic steps at the junctures you cannot predict.” The tool returns this verdict when most of the flowchart exists and only one or two branches need judgment, or when a sign pointing down caps a task that otherwise looks agentic.

The ladder is climbed in both directions. Running an agent over a problem whose flowchart is unknown can be a way of discovering it: its traces show the paths it takes, and “when the paths stabilize, when run after run visits the same steps in the same order, the diagram now exists.” At that point the book’s advice is to “freeze the stable stretch into workflow steps, keep model judgment only at the branches that stayed wild, and pocket the difference on every line of the quote.”

Which traps distort the decision?

Two traps distort the decision, and they push in opposite directions. The first is the argument for building more than you need: “‘it might need to adapt someday.’ That is a hypothesis, and cheap to test honestly: build the simpler version first, measure it against real cases, and climb the ladder only when you can point at a class of inputs the simple system provably fails on.”

The second is agent-washing: “because the word sells, scripted pipelines get relabeled as agents.” Chapter 14 adds that the mislabel bills you in both directions. A workflow sold as an agent carries oversight ceremony it does not need. An agent sold as an “automation” is worse, because “the disciplines autonomy demands (budgets and stop conditions, the security audit… an eval set, a place for a person) were never bought, because the label said they were unnecessary.”

The question that cuts through both traps is the one the tool is built on: “does the model decide the next step, or does code?”

How does the tool reach its verdict?

The tool reaches its verdict by asking the book’s questions in an order that settles the easy cases first. A whole area of work is sent back for decomposition. A fully specifiable task lands on plain code. A one-judgment-per-item task lands on a single call. A task whose flowchart you can draw, or one capped by an audit requirement or a waiting user, lands on a workflow. Only an open-ended task reaches the three conditions, and only one that meets all three gets an unqualified agent.

The book calibrates the decision on a pair of examples, and the tool reproduces both. “Route each incoming email to billing, technical, or sales and draft a first reply” has knowable steps and high volume, “and wants a workflow.” “Make this failing test pass in a codebase you have never seen” has unknowable steps and a built-in verifier, “and is a genuine agent problem.” Other guides make the same point in their own words. Anthropic’s engineering guide advises “finding the simplest solution possible, and only increasing complexity when needed,” and repeats that “you should consider adding complexity only when it demonstrably improves outcomes.”

A verdict from a tool is a starting point, not an architecture review. Use it to frame the conversation with your team, then test the cheapest design that the verdict allows against real inputs before climbing.

Questions readers ask

What is the flowchart test?
It asks whether you can draw the flowchart of the task before the request arrives. If you can, code should hold the control flow and the design belongs at a workflow or below. If the steps depend on what is discovered along the way, the task may need an agent.
The tool put me on a workflow, but my demo was an agent. Is that a downgrade?
No. The book calls stepping down the ladder as respectable as climbing it. An agent's traces can show you the flowchart; once the paths stabilize, freezing them into workflow steps cuts cost, latency and risk on every request.
What is the co-pilot posture?
An agent that proposes while a person approves each consequential action. The book recommends it when the stakes are high and there is no cheap way to verify the agent's work.
Can a task need an agent and still fail this test?
Yes, if one of the three conditions is missing: a machine-checkable success criterion, a feedback signal at each step, or a place for a person to look. The tool then keeps a person in the loop until you supply the missing condition.

Sources

  1. Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents
  2. OpenAI (2025). A practical guide to building agents