The short answer to when not to use AI agents: when you can list the steps before the request arrives, when the work is one judgment per item, when the path must be reproducible for an audit, when latency or cost per task has a hard ceiling, or when a wrong action is expensive and nothing checks it. Each case has a cheaper design.
This post is for the tech lead with a proposal on the desk that says “build an agent for X.” It gives you nine things the proposal has to show, a ten-minute check for each, and a named thing to build instead when a check fails. The alternatives matter as much as the checks. A bare “no” in a design review reads as obstruction; “this is a single call with three rules around it, and here is why” reads as engineering.
Practitioners voice the suspicion behind the question in public. One Hacker News commenter put it this way in March 2026: “Almost every time I have an idea for AI Agent, I end up just making a script/binary that does the same, but so much faster that adding AI to it feels silly” (comment). The essay that organized much of this field says the same thing in a calmer voice: “we recommend finding the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all” (Schluntz and Zhang, 2024).
When not to use AI agents: what are the signs?
The signs are the ones Chapter 14 of the book (in the full book) lists in its section “Signs You Need One, Signs You Don’t”: the steps can be enumerated, correctness must be guaranteed or audited, a person or a thin margin is waiting, the task is lookup, classification or transformation, or the stakes are high with no cheap verification.
The chapter is direct about their weight: “Five signs point down the ladder, and any one of them comes close to deciding alone.” The ladder is the book’s ordering of designs from plain code, to a single model call, to a workflow, to an agent. The rule that separates the top two rungs (who decides the next step, your code or the model) has its own post, the agents vs workflows decision rule, and I won’t re-derive it here.
What this post adds is the form a reviewer needs. I have turned the five signs into six checks you can run against a written proposal (the first sign, steps you can enumerate, becomes Checks 2 and 4), and added three of my own. One comes before the five: whether the request is a task at all. One comes after: whether irreversible actions wait for something. The last is about volume: whether the task runs enough times to repay the cost of checking it.
At the time of writing (October 2026), several of the top search results for this question were registration forms in front of an analyst report. One open page offers a whiteboard test, “If you can write out the complete decision tree on a whiteboard and every input is predictable, you don’t need an agent” (Vora, 2026), and a decision map whose alternatives include robotic process automation, a direct integration, a direct model call and a human process. The map has no row for a fixed workflow of model calls, or for one bounded model step inside it, and those middle rungs are where the work lands.
What does a proposal have to show before it gets an agent?
A proposal has to show nine things, and they are stated below so that a tick is good news for the agent. Checks 1 to 6 set the rung: read them in order and stop at the first one you cannot tick. Checks 7 to 9 apply only to a proposal that passed the first six, and each unticked one adds a requirement or a veto.
- Check 1. The proposal names one task, with a defined input and a defined output.
- Check 2. At least one step needs judgment that cannot be written down as rules.
- Check 3. The work is several steps that depend on each other, not one judgment per item.
- Check 4. The steps cannot be listed before the request arrives, and the proposal names the class of inputs for which they cannot.
- Check 5. Nobody needs the path guaranteed, or reproducible for an audit.
- Check 6. Nobody is waiting on a screen, and there is no hard ceiling on latency or cost per task.
- Check 7. A machine can tell whether a run succeeded, each step returns a signal, and a person has a place to look.
- Check 8. Every expensive or irreversible action waits for a machine check or a person.
- Check 9. The expected number of runs repays the hours it takes to build the checks.
Nine ticks mean the proposal has made its case. Checks 2, 3 and 4 are the ones proposals get wrong in good faith, because a demo hides them. A demo shows the model choosing steps; it does not show whether the same steps would have been chosen every time, in which case code could have chosen them.
The order of the first six is my own, and it follows the order in which the book’s ladder is climbed: a failed check near the top of the list lands on a lower rung, so you stop at the cheapest design that fits. The Should this be an agent? decision tool settles its verdict in the same order, though it asks about the flowchart second, and it covers Checks 1 to 8; Check 9 is not in it.
What do you build instead, check by check?
Each failed check has a named alternative, and the table below pairs them. The “ten-minute check” column is what to do with the proposal in front of you; the last column is what the design lands on, and the buttons filter by it.
| Check | The proposal must show | A ten-minute check | If it can’t, build this | Lands on |
|---|---|---|---|---|
| 1 | One task, defined input and output | Write one example input and the output you would accept. If you need a paragraph to say what “the output” is, it fails | Nothing yet. Split the request into pieces shaped like tasks and run each piece through this table | decompose |
| 2 | A step that needs judgment rules can’t express | For each step, try to write the rule. A threshold, a query, a pattern match and a lookup table all count as rules | A function, a query or a scheduled job. A model may help write it; no model runs in it | plain code |
| 3 | Several dependent steps, not one judgment per item | Ask whether each item is handled alone: classify it, extract from it, rewrite it, score it | One model call per item with a fixed output format, then rules in code that act on the answer | single call |
| 4 | Steps that can’t be listed in advance, and the inputs that make it so | Draw the boxes and arrows for five real inputs. If the five drawings match, the flowchart exists | A fixed workflow: model calls in the boxes, your code on the arrows, a check at each seam. If the drawings match except at one branch, keep the workflow and give that branch one bounded model step | workflow |
| 5 | No guaranteed or audited path required | Ask who would have to explain a single run to an auditor, a regulator or a customer | A workflow, with any step the model must choose confined to one bounded box whose output is checked | workflow |
| 6 | Nobody waiting; no hard ceiling per task | Write down the latency a user tolerates and the cost per task the margin allows. Compare them with a loop of unknown length | A workflow with a known number of calls, or a single call if the ceiling is one round trip | workflow, single call |
| 7 | A machine-checkable result, a signal at each step, a place for a person to look | Name the check that would fail on a wrong result. “Someone will look at it” is a place to look and not a check | The agent proposes and a person approves each action; if mistakes are cheap, one model-chosen step inside a workflow, until the missing condition is supplied | person approves, workflow |
| 8 | Expensive or irreversible actions wait for a check or a person | List every tool that writes, sends, pays or deletes, and what each call waits for | The same agent with an approval gate on those tools | person approves |
| 9 | Runs that repay the cost of checking | Divide the hours to build the checks by the hours saved per run (worked below) | No agent. A person does the task, with a single model call as an assistant | no build |
Two rows need a note. Row 2’s “a model may help write it” is the point a commenter made about a team that had wired an agent to a daily data pull: “they could just have the model write a small script that does the job perfectly every time” (comment, July 2026). A model at build time is ordinary tooling. A model at run time is a cost you pay on every request.
Rows 5 and 6 cap the design even when Check 4 is honestly ticked. A task can have unpredictable steps and still sit behind an audit requirement, and then the unpredictable part gets one box. The book’s wording for the audit case: “a code-chosen path is a document you can hand to a regulator, while a model-chosen path is a probability you must defend.”
What does each alternative look like, and what does it cost to check?
Each alternative is a rung of the ladder, and the lower the rung, the more of the proof of correctness comes with the thing you built. The figure compresses the book’s pricing of the four rungs.
A function or a scheduled job
Plain code is the answer when every branch can be specified. You pay in specification and get certainty back: a unit test written once holds until the requirements change. Chapter 14 has the line I’d quote in the review: “An if statement has never hallucinated.”
A single model call with rules around it
A single call is the answer when the work is one judgment per item. The book: “Tasks shaped like lookup, classification, or transformation are single-call shapes however sophisticated the judgment inside the call.” The model reads the item and returns a label or a set of fields in a fixed format; code decides what the label triggers.
That is a classifier plus rules, and it is still not free to check. From this rung on, in the chapter’s words, “correctness is a rate, and a rate must be measured,” so an eval set “belongs on the invoice from day one.” The eval sample-size calculator sizes it.
A fixed workflow
A workflow is the answer when there are several steps and you can draw them. Its costs add where an agent’s multiply, and an error stays in the step where it happened until a check at the seam catches it. The shapes (chain, route, parallelize, evaluate) and what each seam can check are in the post on agentic workflow patterns.
A workflow with one bounded model step
This is the design for a task that is mostly drawable with one branch nobody can enumerate. One Hacker News commenter described the need exactly: “a semi-deterministic workflow where nodes in the workflow are agentic, i.e., powered by LLMs, but the workflow itself remains deterministic” (comment, August 2026). The model owns the inside of one box; code owns when the box runs, what it may touch and what happens to its output.
An agent that proposes while a person approves
The trade calls this the co-pilot posture. Chapter 14 reserves it for high stakes without cheap verification: “the sound design withholds autonomy altogether,” and “the agent proposes, a person approves.” The person is the check that the domain did not supply. How to decide which actions wait, and what the approver should see, is the subject of approval gates by consequence; the glossary entry for approval gate has the short definition.
Worked examples: four proposals through the checklist
Four proposals follow, each walked through the checks until one fails or all pass. The scenarios are illustrative. Each ends with what the checklist gives and what the decision tool returns for the same answers.
“An agent that pulls our metrics every morning and loads them into a table”
This is the shape of the story quoted above. Check 1 passes: one input, one output. Check 2 fails: fetch, reshape and load are all rules, and none needs judgment. Build a scheduled job. The tool, given “one task,” “yes, every step” and “yes” to writing the logic out completely, returns “Rung 1: plain code.”
“An agent that reads supplier invoices and enters them in the ledger”
Check 2 passes, because reading a scanned invoice in an unknown layout needs judgment. Check 3 fails: each invoice is handled alone, and the judgment is one extraction. Build a single call that returns the fields in a fixed format.
Put rules after it: the line items must sum to the total, the supplier must exist, the purchase order must match. Hold entries above a set amount for a person. The tool returns “Rung 2: a single model call.”
“An agent that investigates production alerts and applies the fix”
Checks 1 to 4 pass. The steps depend on what the logs show, and nobody can list them. Suppose no audit rule applies and minutes of latency are acceptable, so 5 and 6 pass as well.
Check 7 fails on its first clause. An alert that clears does not prove the fix was right; a restart clears an alert and hides its cause. Each step does return a signal (logs, metrics, command output), and an on-call engineer is a place to look, so two of the three conditions hold.
Check 8 fails too, since a change to production is expensive to get wrong and the proposal has it waiting for nothing.
Build the co-pilot: the agent investigates and proposes the change, and the engineer approves it. The tool opens below on these answers and returns “Rung 4: an agent, as a co-pilot,” with one missing condition named: “a success criterion a machine can evaluate.”
With JavaScript on, the Should this be an agent? runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Should this be an agent? on its own page to share a result by link.
“An agent that upgrades a dependency across our services until the tests pass”
Checks 1 to 6 pass for the same reasons as the alert case. Check 7 passes this time: the test suites are a machine-checkable criterion as far as their coverage reaches, each build and test run is a signal, and a pull request is a place to look. Check 8 passes because the agent works on a branch and merging waits for review. Check 9 passes if there are many services and upgrades recur.
Nine ticks. The tool returns “Rung 4: an agent.” This is the yes case, and it is no accident that it is a coding task: the book observes that coding is the domain where all three conditions come with the territory.
When is volume too low to pay for the checking?
Volume is too low when the number of runs you expect is smaller than the hours it takes to build the checks divided by the hours each run saves. This check is my addition to the book’s signs; it follows from a sentence in Chapter 14: “a quote is per task, and nobody buys one task. Multiply every line by the volume you actually expect before you sign.”
The sum has three inputs. F is the hours to build the checking: the eval tasks, the graders, the traces, the review step. m is the hours a person spends doing one task with no agent. c is the hours a person spends checking one agent run. Each run saves m − c hours, so the checking is repaid after F ÷ (m − c) runs.
A worked example with illustrative numbers: F = 60 hours, m = 3 hours, c = 1 hour. Each run saves 3 − 1 = 2 hours, and 60 ÷ 2 = 30 runs to break even.
A quarterly report runs 4 times a year, so 30 runs take 7.5 years. A task that runs 20 times a week reaches 30 runs in 1.5 weeks. The architecture is the same in both cases, and the answers are opposite.
Two cautions keep the sum honest. If checking a run takes as long as doing the task (c = m), the saving per run is zero and no volume repays anything. And F is not paid once: the eval set is rerun on every change to the prompt, the tools or the model, so the sum flatters the agent.
The check does not say that low volume rules out agents. The book’s own example of a sound purchase is “an agent an engineer launches a dozen times a day.” There, the engineer reads every result, so F is close to zero: nothing separate was built to do the checking. Check 9 bites on unattended work, where the checks are a construction of their own.
What are the signs you do need an agent?
You need an agent when the goal is open-ended, judgment is required at branches you cannot enumerate, and three conditions hold: “A success criterion a machine can evaluate,” “A feedback signal at each step,” and “A sensible place for a human to look.” Those are Chapter 14’s words, and they are Checks 4 and 7 read as a yes.
The outside sources agree on the first half. Agents suit “open-ended problems where it’s difficult or impossible to predict the required number of steps, and where you can’t hardcode a fixed path,” in the essay quoted earlier, which adds the cost in the next breath: “The autonomous nature of agents means higher costs, and the potential for compounding errors” (Schluntz and Zhang, 2024). A widely read vendor guide ends its own list of criteria with the same fallback: “Otherwise, a deterministic solution may suffice” (OpenAI, 2025).
The book’s verdict on a proposal that passes is unambiguous: “If the three conditions hold, and the task’s value clears the four-line quote, buy the agent with a clear conscience.” The four lines are what it costs to build, to run, to be wrong and to know whether it worked.
Passing the checklist starts the work. An agent that earned its place still needs a harness: budgets, exits and a done check the worker cannot vote on. The post on the AI agent as a state machine writes that outer loop as a table, and the agent loop explainer shows the inner one.
Which objections will you hear, and what answers them?
Three objections come back in design reviews, and each has an answer you can test instead of argue.
“It might need to adapt someday.” Chapter 1 calls this a hypothesis and prescribes the test: “build the simpler version first, measure it against real cases, and climb the ladder only when you can point at a class of inputs the simple system provably fails on and confirm that an agent actually does better.” Check 4 asks the proposal to name that class of inputs now.
“The demo worked.” A demo is one run. Chapter 14: “A green run is an anecdote about one path through a distribution.” Ask for the five drawings in row 4 of the table, made from real inputs.
“The models keep getting better.” A better model lowers the error rate inside each box. It does not change who must explain a run to an auditor (Check 5), what latency a waiting user tolerates (Check 6), or whether a wrong payment can be recalled (Check 8). Those checks describe your organization, and a model release does not move them.
Where does this checklist fall short?
The checklist judges the shape of a design; it does not tell you whether the task is worth doing, and it depends on honest answers to Checks 2 to 4. A proposer and a reviewer can disagree in good faith about whether steps can be listed. Settle it with inputs: collect real ones, draw their paths, and count how many drawings differ.
A “no” is also not permanent, and a “yes” is not either. The book treats an agent as a legitimate instrument of discovery when the flowchart is unknown: run it, read its traces, and when the paths repeat, “freeze the stable stretch into workflow steps, keep model judgment only at the branches that stayed wild.” A prototype agent whose stated purpose is to find the flowchart passes this review, as long as it runs where its blast radius is small and the proposal says when it will be frozen.
The break-even sum is crude. It counts hours and ignores what a wrong result costs, which for some tasks is the larger number. Treat it as a floor: a proposal that fails it has failed, and one that passes it has more to show.
The one thing to keep
Deciding when not to use AI agents comes down to asking a proposal what it can show, in order, and answering each gap with a design. A script, a single call with rules, a fixed workflow, one bounded model step, a person approving: each is a destination, and the review that sends a proposal to one of them has done its job. Chapter 14 gives the reason in two sentences: “The sticker price of an agent is a prompt. The cost of ownership is a harness, a ledger, and a standing claim on your attention.”
Chapter 1, “What Is an Agent?” is free to read online and introduces the ladder and the flowchart test. Chapter 14, with the full pricing of each rung and the signs in the book’s own order, is in the full book. The free guide to agent patterns collects the related posts, and you can see the formats.
Questions readers ask
- When should you not use an AI agent?
- Do not use an AI agent when you can list the steps before the request arrives, when the work is one judgment per item such as classifying or extracting, when the path must be guaranteed or audited, when latency or cost per task has a hard ceiling, or when a wrong action is expensive and nothing checks it. Each of those cases has a cheaper design.
- What should I build instead of an AI agent?
- Build the lowest rung that does the job. Fully specified logic is plain code. One judgment per item is a single model call with rules around it. Known steps are a fixed workflow. Known steps with one unpredictable branch are a workflow with one bounded model step. Open-ended work with no machine check is an agent that proposes while a person approves.
- Does calling a language model make a system an agent?
- No. A system is an agent when the model decides the next step at runtime. A single call that classifies a ticket or extracts fields from an invoice uses a model and decides nothing about control flow: your code receives the answer and chooses what happens next.
- Can a low-volume task justify an agent?
- It can when a person already checks each run and no separate checking has to be built, as with an engineer supervising a coding agent. For unattended work, do the sum: divide the hours needed to build the eval set, traces and review step by the hours saved per run. With illustrative numbers, a task that runs four times a year can need years to break even.
- When is an AI agent the right choice?
- When the goal is open-ended, so the steps depend on what is discovered along the way, and three conditions hold: a success criterion a machine can evaluate, a feedback signal at each step, and a sensible place for a person to look. The task's value must also cover the cost of building, running and checking the agent.
Sources
- Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents
- OpenAI (2025). A practical guide to building agents
- Ashit Vora (RaftLabs) (2026). When to use AI agents (and when not to)
- Hacker News commenter (2026). Hacker News comment on a daily metrics agent
- Hacker News commenter (2026). Hacker News comment on scripts replacing agent ideas
- Hacker News commenter (2026). Hacker News comment on deterministic workflows with model-powered nodes