The difference in AI agent vs chatbot is who decides the next step. In a chatbot, the person does, one message at a time. In an agent, the model does, in a loop, until a stop rule ends the run. The same model can sit inside both, so the label on the product tells you very little.
You are probably here because you have to put one word on a slide, a landing page or a roadmap: chatbot, assistant, copilot, agent. I’d treat that word as a claim. It commits you to what the system touches, who reviews it and what a mistake looks like.
Chapter 1 of AI Agents, Engineered, which is free to read online, draws the line. The four-row table, the “chatbot with tools” column, the promise table, the one-line descriptions and the placement test below are this post’s own work. The book prints none of them, and “chatbot” there is a working term with no glossary entry.
AI agent vs chatbot: where is the line?
The line sits at one question, which is who chooses what happens next after each step. Chapter 1 answers it for three systems: the person chooses in a chatbot, your code chooses in a workflow, and the model chooses in an agent.
The chapter sets up three examples: an assistant that answers and waits, a ticket pipeline with four fixed steps, and a system told to fix failing tests. Then it says what separates them. “All three may be built on exactly the same underlying model.” And, a few lines on:
“In the first system, you do, one message at a time. In the second, your code does; the model only fills in content along a route that was fixed before the first ticket ever arrived. In the third, the model decides: which file to read, whether to edit or rerun, when to stop.”
Three words in this post need a plain gloss. A tool is, in the book’s glossary, “A function the model can ask your program to run”, such as a search or an order lookup. A loop means your program calls the model again after each result, so the model can react to what it just saw. A workflow is several model calls and tool calls on a route your code fixed in advance.
The chapter describes the chatbot in three short sentences: “A chatbot is turn-by-turn conversation. The human drives every step; the model generates text and waits. There is no loop and nothing is executed.”
What do the four columns look like side by side?
Four columns cover most products that ship: a chatbot, a chatbot with tools, a workflow and an agent. Each is defined by who decides the next step, and the other three rows follow from that answer.
| Chatbot | Chatbot with tools | Workflow (automation) | Agent | |
|---|---|---|---|---|
| Who decides the next step | The person, one message at a time | The person. Inside one turn the model may ask for one call; then it answers and waits | Your code, on a route fixed before the request arrived | The model, at runtime, until it judges the goal met or a stop rule ends the run |
| What it touches | Nothing outside the conversation. It may read documents your code fetched for it | What its tools reach, one call per turn: usually a read, sometimes one change the person confirmed | The same systems, in the same order, on every run | Anything its tools permit, in an order nobody wrote down |
| How it fails | A wrong answer, which your company has still said | A wrong answer resting on a wrong or stale lookup, or one wrong change | An error at a known step, in the same place each time, visible where two steps meet | A wrong action, then more actions built on it |
| What to verify | Sampled answers against the source they should rest on; the list of things it may promise on your behalf | The chatbot cell, plus: the answer matches what the tool returned, any change waits for a confirmation the person saw, and any change that spends your money has its cap in code | Each step against its contract; the whole route against a test set | The trace of each run; limits enforced in code (budget, permissions, approval gate); an outcome check that does not come from the model |
A trace is the step-by-step record of one run, and a contract is the written statement of what a step takes in and must hand on.
The second column is my reading of a gap the chapter leaves open. Its parenthesis reads: “(A chatbot, in this picture, is the bare model call with the augmentations mostly unused.)” The word “mostly” leaves room for a chat assistant that looks one thing up, answers and waits.
Most AI agent vs chatbot arguments on forums are about what belongs between the two. One Hacker News commenter wrote in June 2026: “Without agent features, you have just a chatbot” (locknitpicker). Others put the weight on the path the system takes. The table takes that second view and gives the first its own column.
I use “workflow” and “automation” for the same column. Classic automation with no model in it lands there too, which settles most AI agent vs automation arguments: if code fixed the steps, the column is the third one.
Do published definitions draw the same line?
Most published definitions draw the line on control, and they disagree about where a chatbot falls. Four of them, read side by side, show it.
| Source and date | Its words | Where a chatbot lands |
|---|---|---|
| Russell and Norvig, 1995, as reproduced by Franklin and Graesser (1996) | “An agent is anything that can be viewed as perceiving its environment through sensors and acting upon that environment through effectors.” | Inside: it perceives a message and acts by replying |
| Schluntz and Zhang, December 19, 2024 | “Customer support combines familiar chatbot interfaces with enhanced capabilities through tool integration.” | An interface that an agent may sit behind |
| OpenAI guide, undated | “Applications that integrate LLMs but don’t use them to control workflow execution—think simple chatbots, single-turn LLMs, or sentiment classifiers—are not agents.” | Outside, when it is “simple” |
| NIST notice, Federal Register, January 8, 2026 | “AI agent systems are capable of planning and taking autonomous actions that impact real-world systems or environments.” | Out of scope unless “orchestrated to act autonomously” |
Two disagreements matter to a product owner. The textbook sense is so wide that a chatbot qualifies as an agent. And the first engineering source treats the chat window as a surface, so the same window can sit in front of any column.
The last row adds a second property. That notice is a request for information that sets its own scope, with no force as a standard. It asks about systems “capable of taking actions that affect external state, i.e., persistent changes outside of the AI agent system itself.” That property is my “what it touches” row.
What is agentic AI, and how does it differ from generative AI?
Agentic AI is a label for systems in which a model chooses some of the steps; generative AI names what the model produces. In agentic AI vs generative AI, an agent has a generative model inside it, and a plain chatbot is generative without being agentic.
The label is loose. The NIST notice says, “Other terms used to refer to AI agent systems include AI agents and agentic AI.” So when someone asks what is agentic AI, I answer with the column and skip the adjective.
For AI agent vs LLM, the answer is shorter. An LLM (a large language model) turns text into more text, and the agent is the arrangement around it: tools, a loop, a goal and limits.
What changes once the model decides the next step?
Three things change together: the bill, the kind of mistake, and the amount of machinery you must build around the model. Chapter 1 prices all three in its last section.
Cost. The chapter says of money that “an agent makes many model calls per task instead of one, and each call re-sends the growing conversation, so cost scales twice over—once in the number of calls, again in the size of each.” It adds an illustration and labels it as one: “A ten-step loop can cost an order of magnitude more than the single call it replaced (the multiplier is illustrative; the direction is what matters).”
Risk. A chatbot’s failure is bounded: “the worst a chatbot can produce is a wrong answer, because acting on the world is out of its reach.” An agent gives that bound up. “Its mistake is a wrong action, followed by ten more actions built on top of it”, the chapter says. The damage a wrong action can do before something stops it has a name, blast radius.
Oversight. The fourth column’s build list is in the chapter too: “a chatbot needs a prompt, while an agent additionally needs tool plumbing, context management, error recovery, stop conditions, approval gates, tracing, and evaluations.” That surrounding machinery is the harness. An approval gate is a pause where a person approves a proposed action before it runs.
Human approval keeps a system in the agent column. The chapter places much of production at the point where an agent “loops freely over its tools but pauses for human approval before anything consequential”. Inside each pass the model reasons and then acts, a rhythm usually called the ReAct agent pattern.
Which column is your product in?
Six yes-or-no questions place a product: the first five lead to exactly one column, and the sixth sizes the checking. An AI agent vs chatbot comparison earns its keep when you can run it on your own system. Answer from the code or with the engineer who wrote it, because all four columns look the same through a chat window.
- Does the model’s output ever decide what your program runs next? That means the model chooses whether to call a function (a tool call), or which of several branches to take. A step your code runs on every request does not count, even when the model wrote its input: a search your code always runs, before or after a model call, belongs to a fixed route. No: go to 2. Yes: go to 3.
- Does the model’s text go only to the person who asked, for them to read and act on? Keeping a transcript does not change the answer. Yes: chatbot. No, your code feeds the text into another step or writes it into another system: workflow.
- After a result comes back, can the model choose another action in the same run, with no person choosing in between? A person clicking “approve” on a step the model proposed still counts as the model choosing. No: go to 4. Yes: go to 5.
- Does the reply go to a person in a conversation, who then sends the next message? Yes: chatbot with tools. No, the result goes on to a step your code had already written: workflow.
- Could two different requests produce two different sequences of actions that nobody wrote down in advance? Yes: agent. No, the model only repeats a step your code scripted: workflow.
- Can any action change something outside the system: send, spend, book, delete? A yes means you owe the whole of your column’s “what to verify” cell; the column stays where questions 1 to 5 put it.
Question 3 is where products move without anyone deciding it. A chat assistant starts with one lookup per message. Then someone lets the model call again after reading the result. The product has now left the second column, and question 5 will usually put it in the fourth. The slide still says “assistant.”
Placing a product in a column is a coarser cut than testing what an AI agent is check by check from a trace.
Which promise needs which architecture?
Every promise on a slide has a least architecture that makes it true, and each architecture brings a build list with it. Pick your column in the filter to see the promises for which it is the smallest sufficient build.
| The promise on the slide | What has to be true inside | What it obliges you to build | Least architecture that makes it true |
|---|---|---|---|
| “Answers questions about your help docs or your policy” | A model call that reads documents your code fetched | A list of what it may say on your behalf; sampled answers checked against the source | chatbot |
| “Tells you where your order is” | A chat turn with one read-only lookup the model asks for (if your code always fetches the order first, a chatbot keeps this promise) | The lookup result shown in the answer; a fallback when the lookup fails | chatbot with tools |
| “Books it, files it or cancels it when you say so” | A chat turn with one change, made after a confirmation the person saw | A confirmation step; a change that is safe to retry; a log or an undo; a cap in code when the change spends your money | chatbot with tools |
| “Processes every request the same way, start to finish” | A fixed sequence of calls | A contract for each step; a test set for the whole route | workflow |
| “Works while you sleep” | A schedule or a trigger in front of a fixed sequence | Monitoring and alerts for runs nobody watched | workflow |
| “Resolves the issue end to end”, when the steps depend on what it finds | A model choosing among tools in a loop | Budget, stop rule, permissions, an approval gate on consequential actions, traces, an outcome check | agent |
| “Learns from every conversation” | A separate claim about memory or retraining, which any of the four may or may not have | Evidence that behavior changes, and a review of how | none of the four |
| “Replaces your support team, your lawyer or your analyst” | A performance claim, which no architecture makes true by itself | Evidence from testing against the people it is said to replace | none of the four |
A column further right can usually keep a promise listed further left, and you pay for the difference. “Works while you sleep” is the row people misread most: running unattended describes a schedule, and a nightly script is still a workflow.
The last two rows are claims about results, the family that the regulator actions below concern, and they sit outside all four columns.
How does the test run on real products?
Run on four real-shaped products, the test puts each in exactly one column, and the promise table gives the same answer. Two are quick; two are written out.
A support widget that answers from help-center articles stops at question 2: your code fetches the articles, the model writes, the person reads. It is a chatbot. A nightly invoice pipeline with fixed steps and one model call per step also stops at question 2, on the other answer. Code feeds each output into the next step, so it is a workflow.
The same widget with an order lookup
Give that widget one read-only call that looks up an order. Question 1 is now yes: the model asks for the lookup. Question 3 is no, provided the code allows one call per message. Question 4 is yes, so the widget is a chatbot with tools.
The column’s rows hold. It touches the order system, read-only, once per turn. It fails by answering from a wrong or stale lookup. What to verify is that the answer matches what the lookup returned, and the promise table agrees: “Tells you where your order is” needs exactly this much.
Ask your engineer one thing before you rely on that placement: can the model call the lookup again after reading the result? If it can, question 3 flips to yes and the test continues to question 5.
An assistant that takes a ticket and fixes it
Now take an assistant that receives a ticket, reads logs, changes a setting and checks the result. Questions 1 and 3 are both yes, since it reads a log and then chooses what to open or change next. Question 5 is yes, because two tickets will produce two different sequences. It is an agent.
Question 6 is also yes: changing a setting is a persistent change outside the system. So the whole fourth-column cell is owed. That means a trace of each run, a budget and permissions held in code, an approval gate before the change, and a check that the problem is gone. “Resolves the issue end to end” is the matching promise.
That last check has to come from outside the model: the error rate dropped, the test passes, the customer confirmed.
When is a chatbot enough?
A chatbot is enough when the person can act on the answer themselves, a wrong answer is cheap to catch, and nothing has to change in another system. Chapter 1 is plain about it: “Chatbots are genuinely useful and often the right tool”.
The same goes one column over. If the task needs one lookup, or one change the person confirms, the second column keeps the promise with a short build list.
Whether a task deserves the third or fourth column is a separate decision with its own post, the agents vs workflows decision rule. If your placement came out “agent”, the Should this be an agent? tool tests whether it should stay there.
Can a chatbot that only answers still cost its owner?
Yes: a tribunal has ordered a company to pay for what its website chatbot told a customer. The chapter says a chatbot’s worst output is a wrong answer; that a wrong answer can cost money is this post’s addition. What follows is one case, with no rate behind it.
In Moffatt v. Air Canada (2024 BCCRT 149, issued February 14, 2024), British Columbia’s Civil Resolution Tribunal heard a small claim about bereavement fares. The decision records: “The chatbot suggested Mr. Moffatt could apply for bereavement fares retroactively.” The system only answered. The airline “did not provide any information about the nature of its chatbot”, so the technology behind it is unknown.
The tribunal’s reasoning is the part to keep. “It makes no difference whether the information comes from a static page or a chatbot.” And: “I find Air Canada did not take reasonable care to ensure its chatbot was accurate.”
The order reads, “I order Air Canada to pay Mr. Moffatt a total of $812.02”. That is $650.88 in damages, $36.14 in pre-judgment interest and $125 in tribunal fees. The decision prints those sums with a plain dollar sign, and the forum is Canadian. It is a small-claims ruling in one province, on its own facts (Civil Resolution Tribunal, 2024).
Two more answer-only cases, each from its original report:
- A support email bot, April 2025. A Hacker News poster reported on April 14, 2025 that a developer-tools company’s support had described a login bug as policy, calling it “a hallucinated excuse from a support bot” (scaredpelican). A reply signed “(Cursor cofounder)” on April 16 said, “Any AI responses used for email support are now clearly labeled as such”, and “We’ve made sure this user is completely refunded”. The bot only answered. The poster reported cancellations; no count is documented.
- A parcel firm’s support chatbot, January 2024. BBC News reported on January 19, 2024: “DPD has disabled part of its online support chatbot after it swore at a customer.” No sum is reported (BBC News, 2024).
For contrast, one case where the system acted. On July 21, 2025, The Register reported a user’s account of a coding service that “deleted a database despite his instructions not to change any code without permission”. The same report says a rollback later worked (Sharwood, 2025). The report puts no figure on the loss.
Read together, these four say something narrow. Where the system only answered, the documented costs were a small order, a refund, a labeling change and a feature switched off. Where it acted, a user reported a deleted database, which a rollback later restored. Both kinds reached the company.
What can you honestly call your product?
Call it by the verbs it performs and name the point where it stops for a person. Each column of the table has earned some words and has no claim on others. A line that stays inside its column is one a buyer’s engineer will accept.
| Column | Words it has earned | Words it has no claim on |
|---|---|---|
| Chatbot | answers, explains, drafts, summarizes | looks up, books, changes, resolves |
| Chatbot with tools | answers; one or two named actions, each “when you ask” | resolves, end to end, by itself, autonomous |
| Workflow | runs, processes, routes; a count of fixed steps | decides, chooses, adapts |
| Agent | chooses its own steps, works toward a goal, with its limits stated | any of those without the limits; “replaces” without test evidence |
These four lines are templates. Fill the brackets and delete the columns you are not in.
Chatbot: [Product] answers questions about [scope] from [source].
It replies in text and leaves your account as it is.
Chatbot with tools: [Product] answers questions and can [look up / book / file]
[one thing] when you ask. Each action is a single step,
and any change waits for your confirmation and stays
under [limit].
Workflow: [Product] runs [task] through [n] fixed steps: [verbs].
A person reviews [which outputs].
Agent: [Product] works toward [goal] by choosing its own steps
among [tools]. It stops for approval before [consequential
actions] and is capped at [budget].
For a deck about where the product is going, write both columns and the evidence between them: today the second column; next the fourth, once a named check passes on a stated number of real cases.
The opposite habit has a name, agent-washing. Chapter 1 describes it this way: “because the word sells, scripted pipelines get relabeled as agents.” A Gartner release of June 25, 2025 uses “agent washing” for the rebranding of existing products, “such as AI assistants, robotic process automation (RPA) and chatbots, without substantial agentic capabilities”. Its headline is a forecast by that firm, which I leave with it.
What have regulators acted on?
In the United States, two agencies have acted on claims about what an AI product does, and neither action below concerns the word “agent.” I read only U.S. agency pages, and nothing here is legal advice.
On September 25, 2024, the Federal Trade Commission announced “Operation AI Comply.” One of its cases involved DoNotPay. In the release’s account, the company “claimed to offer an AI service that was “the world’s first robot lawyer,” but the product failed to live up to its lofty claims that the service could substitute for the expertise of a human lawyer.” The release continues: “The complaint alleges that the company did not conduct testing to determine whether its AI chatbot’s output was equal to the level of a human lawyer”.
By the same release, the company “has agreed to a proposed Commission order settling the charges”, and “The settlement would require it to pay $193,000”, a sum in U.S. dollars (FTC, 2024). That is the release’s wording on its date.
On March 18, 2024, the Securities and Exchange Commission announced settled charges against two investment advisers over statements about their use of AI. Its Chair is quoted: “Such AI washing hurts investors.” The release says, “The firms agreed to settle the SEC’s charges and pay $400,000 in total civil penalties”. They did so “Without admitting or denying the SEC’s findings” (SEC, 2024).
That release is about investment advisers and says nothing about a startup’s pitch. The two share the shape of the claim at issue: this product does X, or replaces Y. A line that names the verbs and the limit is the practical answer to that shape.
Where does this comparison stop working?
The four columns are a coarse map, and they blur in three places. Chapter 1 warns about the first one itself: “It is tempting to treat chatbot, workflow, and agent as three boxes. The truth is a dial”.
Mixed systems. The chapter also says, “Almost nothing that ships is a pure workflow or a pure agent.” A chat window in front of a fixed pipeline of several model calls lands in the workflow column at question 2, and its users still meet a chatbot. It owes two cells: the workflow column’s per-step checks and the chatbot column’s list of what it may promise.
What the system touches. Question 6 only says whether anything changes outside. Grading each action by what a wrong one costs is a finer job, and the consequence tier classifier does it for a list of actions. The full set an agent’s tools permit is its action space.
The evidence. The incidents are four cases I could read at their source, and the regulator actions are two from one country. Neither set supports a rate. The definitions carry dates because the vocabulary is still moving.
An AI agent vs chatbot placement also says nothing about whether the product is good. Deciding what the system may do without a person, and what evidence would let it do more, is the core of working with AI agents as a product manager.
The takeaway
For AI agent vs chatbot, ask who decides the next step, then ask what the system touches, and write the product line from those two answers. Chapter 1 gives the question to carry into every review of the result. It asks, “what signal tells you it worked?”
Chapter 1, “What Is an Agent?” is free to read online and holds the three working terms, the dial and the costs quoted above. The agent fundamentals guide collects the related posts and tools, and the What is an agent? explainer shows the loop in motion. The rest of the book builds the harness the fourth column needs, and you can see the formats.
Questions readers ask
- What is the difference between an AI agent and a chatbot?
- Who decides the next step. In a chatbot the person does, one message at a time, and the model writes a reply and waits. In an AI agent the model does: it picks an action, reads the result and picks again, in a loop, until it judges the goal met or a stop rule ends the run.
- Is a chatbot with tools an AI agent?
- A chat assistant that makes one call, answers and waits for your next message is a chatbot with tools, because you still choose every next step. It becomes an agent when the model reads the result of one action and chooses another by itself, in a sequence nobody wrote down in advance.
- AI agent vs LLM: what is the difference?
- An LLM, or large language model, is the component that turns text into more text. An AI agent is an arrangement around that component: tools, a loop, a goal and limits. One LLM can sit inside a chatbot, a workflow or an agent, so changing the model leaves the column unchanged.
- Agentic AI vs generative AI: are they the same thing?
- They name different properties. Generative says what the model produces: text, images, code. Agentic says how much of the path the model chooses. An agent has a generative model inside it, and a chatbot built on the same model is generative without being agentic.
- Can a chatbot that only answers still cost a company money?
- Yes. In Moffatt v. Air Canada (2024 BCCRT 149, issued February 14, 2024) a small-claims tribunal in British Columbia ordered the airline to pay $812.02 after its website chatbot gave a customer wrong information about bereavement fares. It is one case, decided on its own facts.
Sources
- Civil Resolution Tribunal (British Columbia) (2024). Moffatt v. Air Canada, 2024 BCCRT 149
- U.S. Federal Trade Commission (2024). FTC Announces Crackdown on Deceptive AI Claims and Schemes (press release)
- U.S. Securities and Exchange Commission (2024). SEC Charges Two Investment Advisers with Making False and Misleading Statements About Their Use of Artificial Intelligence (press release 2024-36)
- Gartner (2025). Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (press release)
- National Institute of Standards and Technology, Center for AI Standards and Innovation (2026). Request for Information Regarding Security Considerations for Artificial Intelligence Agents (Federal Register 91(5), docket NIST-2025-0035)
- Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents
- OpenAI (undated). A practical guide to building agents (PDF)
- Stan Franklin and Art Graesser (1996). Is it an Agent, or just a Program?: A Taxonomy for Autonomous Agents
- locknitpicker (Hacker News) (2026). Hacker News comment 48425736
- scaredpelican (Hacker News) (2025). Cursor IDE support hallucinates lockout policy, causes user cancellations (story 43683012, with replies 43700931 and 43701403)
- BBC News (2024). DPD error caused chatbot to swear at customer
- Simon Sharwood (The Register) (2025). Vibe coding service Replit deleted user’s production database, faked data, told fibs galore