What is an AI agent? It is a program in which a language model, given a goal and a set of tools, chooses its own next step, sees the result, and chooses again until the goal is met or a stop rule ends the run. The model owns the sequence of steps, and your code owns the limits.
That sentence is my own compression of Chapter 1 of AI Agents, Engineered, and I suspect it is only half of what you came for. Most people who ask this question have a specific thing in front of them: a product that was sold as an agent, a project the team has started calling one, or something built over a weekend that needs a name.
So this post is a test. It has five checks, you run them in order against one concrete system, and the place where the system stops is your verdict. If you cannot see inside the system, skip to the three questions further down.
What is an AI agent, in one sentence?
An AI agent is a language model placed in a loop with tools and a goal, where the model decides what happens next. The book’s first chapter says it in one short sentence, “An agent is a language model placed in a loop,” and then spells out the motion: “It looks at the state of the task, picks an action, sees the result, and goes again, until it judges the goal met or a stopping rule ends the run” (Chapter 1, “The Agent Idea”).
Four terms in there need a plain gloss. A language model (an LLM) is a program that takes text in and produces text out. A tool is an ordinary function, such as “search this” or “run this command,” that your program executes when the model asks for it. A loop means the model is called again after each result.
The fourth term is control flow, the order in which a program’s steps run, and it is where an agent differs from every program you have used before. Chapter 1 again: “In an agent, the sequence is decided by the model, at runtime, in response to what it observes.”
If you need a line short enough for a meeting, the field has largely converged on one, which the book adopts: “An LLM agent runs tools in a loop to achieve a goal” (Willison, 2025). My own version, for when someone asks what makes it safe to run: an agent is a model that chooses its own next step, in a loop, within limits set by code.
How do you test whether a system is an AI agent?
You test a system with five checks, run in order. Each one assumes the one before it: a language model is in the path, its output decides something, the result comes back and it decides again, the path is not in your code, and code sets the limits.
Pick one system and get as close to it as you can. The best evidence is a trace, the recorded list of steps from a single run. An architecture diagram or ten minutes with the person who built it will also do. Then work down the list.
Tick a check only when the system passes it, and stop at the first one you cannot tick.
- A language model is in the path. Find the place where a language model is called while the system runs. If every decision comes from rules, thresholds, a schedule, or a learned model of another kind (a spam filter, a recommender), stop here.
- Its output decides something. Find one thing the program does differently because of what the model returned: a branch taken (one of several paths already written in the code), or a function called with arguments the model wrote. If the model’s text is only shown, stored or sent as content, and the same code runs next whatever it says, stop here.
- The result comes back and it decides again. Follow whatever ran because of that decision. Its result must return to the model, which then picks the next action. Find a run where that return trip happens at least twice, with no person and no fixed code step choosing in between. If the result goes to the user or to the next stage of a pipeline instead, stop here.
- The path is not in your code. Run two different inputs and compare the two traces. Look at which actions ran and in what order. The model should have chosen that sequence from several possible actions. Two calls to the same tool with different arguments the model wrote count as different actions. If the only thing it chose was how many times to repeat one step whose content your code fixed, or if you could have written the sequence of named steps before either run, stop here. If two traces match, try a third input that is really different before you conclude anything.
- Code sets the limits. Find the stop rule, the budget on steps or cost, the permissions, and any approval gate (a pause for a person before a consequential action). Confirm that code enforces each one, where the model cannot argue its way past. Limits that exist only as sentences in a prompt do not count. To pass, code must hold at least one cap (on steps, time or cost) that ends the run whatever the model does; write down any of the others that are missing.
What does each stopping point mean?
The last check you ticked is the answer, so here is the mapping in plain text.
| Last check passed | What you have | The honest name for it |
|---|---|---|
| None | Rules, thresholds, a schedule or a learned model of another kind | Automation without a language model |
| 1 | A model that writes content into a slot your code prepared, once or at each stage of a fixed route | A model call, or a fixed pipeline of them (a chatbot or a workflow) |
| 2 | A model that makes one decision, after which code or a person takes over | A workflow with one model decision (a router, which picks among branches your code already wrote), or a chatbot with one tool call |
| 3 | A model that decides, round after round, whether to repeat a step your code scripted | A workflow with a loop |
| 4 | A model that chooses the path, with only its own judgment to stop it | An unbounded agent, which nobody should run |
| 5 | A model that chooses the path inside limits it cannot move | A bounded agent |
Two notes on where this comes from. Check 4 is a narrower field version of the book’s litmus test: “If you can confidently draw the control-flow diagram before the request arrives, you are looking at a workflow. If the diagram can only be drawn in hindsight, once the model has reacted to what it found, you are looking at an agent.” And the idea of ordering systems by how much the model’s output controls is older than this post. A 2024 Hugging Face article ranks “Simple processor,” “Router,” “Tool call” and “Multi-step Agent” as rising levels of agency (Roucher et al., 2024).
Check 5 is my addition. Chapter 1’s definition asks only for a stopping condition: a goal “is what ends the loop—an agent has a stopping condition, which is what separates it from a runaway process.” The check goes one step further and asks that code hold the stop, the budget and the permissions, so that the model cannot talk its way past them. That is why a system that passes check 4 and fails check 5 is still an agent in the table above.
The mechanics of the loop, and the three ways it can exit, have their own post: what an agent loop is. Inside each pass the model reasons and then acts, a rhythm usually called the ReAct agent pattern.
What do the five checks look like on real systems?
On real systems the checks place eleven familiar designs on the mapping above: one stops before check 1, six stop at check 1 or 2, one stops at check 3, and three reach check 4. The table is my own classification, made by running the five checks as written. Read ✔ as pass, ✘ as fail, ~ as “depends on the build” and – as “not reached.” The verdict column gives the last check passed and its name from the mapping.
| System | 1 | 2 | 3 | 4 | 5 | Verdict | Why |
|---|---|---|---|---|---|---|---|
| Thermostat, web crawler, rules-based automation bot | ✘ | – | – | – | – | None: automation without a language model | It senses and acts on rules and thresholds; no language model is called |
| Plain chatbot: question in, answer out | ✔ | ✘ | – | – | – | 1: a model call | The text is shown to you, and you choose every next step |
| Retrieval pipeline: fetch documents, then generate an answer | ✔ | ~ | ✘ | – | – | 1: a fixed pipeline; 2 if the model writes the search query | Code always retrieves, then calls the model; the documents go to the next stage and never return for a second decision |
| Scheduled job with one model call: a nightly summary or digest | ✔ | ✘ | – | – | – | 1: a model call | The summary is stored or sent as content; the same code runs whatever it says |
| Classifier or router: label the request, send it down one of several fixed branches | ✔ | ✔ | ✘ | – | – | 2: a workflow with one model decision | The label picks a branch; the result goes to the next stage |
| Chat assistant that makes one tool call, then answers and waits | ✔ | ✔ | ✘ | – | – | 2: a chatbot with one tool call | The result returns to the model once; a second trip needs your next message |
| Draft-and-check loop capped at a few rounds | ✔ | ✔ | ✔ | ✘ | – | 3: a workflow with a loop | Each new draft returns to the model, which only chooses “again” or “done” for one scripted step |
| Support bot, scripted build: classify, look up policy, draft, queue for review | ✔ | ~ | ✘ | – | – | 1: a fixed pipeline; 2 if the label selects the policy | The same named steps in the same order for every ticket |
| Support bot, tool-loop build: the model chooses among lookup, refund and escalate until resolved | ✔ | ✔ | ✔ | ✔ | ~ | 4 or 5: an agent; ask about check 5 first | The model picks which action comes next after each result, and when to stop |
| Coding agent: search, read, edit, run tests, repeat until they pass | ✔ | ✔ | ✔ | ✔ | ~ | 4 or 5: an agent; ask about check 5 first | Which action runs next depends on what it finds; bounded only if the build caps steps and limits what it can touch |
| Deep-research tool: the model chooses among searching, opening a page and writing up | ✔ | ✔ | ✔ | ✔ | ~ | 4 or 5: an agent; ask about check 5 first | What it searches or opens next depends on what the earlier results said |
The two support bots are the rows to remember. They carry the same product label, sit behind the same chat window, and may call the same model. One is a workflow, “four steps, in that order, every time” in the words of Chapter 1’s own ticket example. The other hands the model the control flow, and a refund tool along with it.
Does running on its own make something an agent? The scheduled job answers that question, which I see asked often. It doesn’t. Unattended describes the trigger, and agent describes who chooses the path.
A retrieval pipeline (often called RAG, retrieval-augmented generation) follows the same logic. It reaches check 3 only when the results return to the model and the model decides whether, what and how often to retrieve again.
For a picture of what passing check 4 feels like, consider a story one developer told on a forum in 2026. A connection the agent normally used was down, so it wrote and debugged a script of its own to reach the same data (Hacker News comment). Nobody had written that path. That is the property, and it is also why check 5 exists.
Is an AI agent just a model with tools?
No: tools are necessary and they are not enough. An agent is the arrangement around a model, meaning the loop, the tools, the goal and the limits, and the same model can sit inside a chatbot, a workflow or an agent without changing. Chapter 1 makes the point about its three example systems: “All three may be built on exactly the same underlying model.”
This clears up the “AI agent vs LLM” confusion, and it is the part most short answers to “what is an AI agent?” leave out.
The LLM is a component that turns text into text. Give it a tool and you have a model that can ask for one thing to happen. The agent begins when the result comes back and the model chooses again, which is check 3.
A chatbot is the contrast case, and the book describes it flatly: “The human drives every step; the model generates text and waits. There is no loop and nothing is executed.” Those two sentences are the whole of the AI agent vs chatbot distinction: a chatbot stops at check 1 or 2 because you choose each next step. A three-minute animation, the chatbot, workflow and agent explainer, shows the same line drawn.
“Agentic AI vs generative AI” is the other pair that gets tangled, because the two words describe different things about the same technology. Generative says what the model does: it produces content. Agentic says how much of the control flow the model owns. An agent built on a language model is generative AI placed in a loop with tools, and that sentence is the short version of what agentic AI means.
Why does everyone’s definition sound different?
Definitions sound different because the word carries two technical senses about thirty years apart, plus a marketing sense, and people rarely say which one they mean. The older sense covers anything that perceives and acts. The current engineering sense requires a language model choosing actions.
The classic AI textbook definition, from 1995, calls an agent “anything that can be viewed as perceiving its environment through sensors and acting upon that environment through effectors” (Russell and Norvig, as quoted by Franklin and Graesser, 1996). The same paper quotes the textbook on what that notion was for: “a tool for analyzing systems, not an absolute characterization that divides the world into agents and non-agents.” The authors meant a lens, and the industry later used the word as a category.
| Definition | Language model required? | Loop required? | Who owns the control flow? | Binary or spectrum? |
|---|---|---|---|---|
| Russell and Norvig, 1995 (via Franklin and Graesser) | No | No | Not addressed | A lens for analysis |
| Franklin and Graesser, 1996 | No | Acts “over time” | The agent’s “own agenda” | A broad class |
| Schluntz and Zhang, 2024 | Yes | Yes | The model; this is the line | A line inside a wider family |
| Roucher et al., 2024 | Yes | Only at the upper levels | The model, to a degree | A spectrum |
| Ng, 2024 | Yes | At the clear end | Not the criterion | A spectrum |
| Willison, 2025 | Yes | Yes | The model, through the loop | Binary enough to be jargon |
| AI Agents, Engineered, Chapter 1 | Yes | Yes | The model | A line drawn on a dial |
Each cell is my reading of the source, and the engineering definitions agree on more than their wording suggests. One draws the line at systems where models “dynamically direct their own processes and tool usage” (Schluntz and Zhang, 2024). Another prefers degrees, saying “agent” is “not a discrete, 0 or 1 definition” (Roucher et al., 2024), a view shared by the suggestion to treat systems as “agent-like to different degrees” (Ng, 2024). All of them put the decision about the next step in the model’s hands.
None of this arguing is new. In 1995 two researchers warned that “‘agent’ might become a ‘noise’ term, subject to both abuse and misuse” (Wooldridge and Jennings, 1995). A year later came a paper titled “Is it an Agent, or just a Program?” Three decades on, one writer crowdsourced 211 definitions, had a model sort them into 13 groups, and only then judged the word usable (Willison, 2025).
How much autonomy does it take to count?
Less than the marketing implies: an agent has to choose its own sequence of steps, and it does not have to work without supervision. A system that loops over its tools and pauses for a person before anything consequential can still pass all five checks: approving a step is not choosing it, so check 3 holds, and an approval gate is one of the limits check 5 looks for.
Chapter 1 draws this as a dial, and puts the supervised agent on it by name: “an agent loops freely over its tools but pauses for human approval before anything consequential—sending the email, spending the money, deleting the data; you will find much of production here.”
So “I still have to guide it” does not demote a coding agent to a chatbot. The thermostat question has a two-part answer. The authors of that 1996 paper tested their own broad definition on extreme cases and conceded that “a thermostat satisfies all the requirements of the definition.” Under the engineering sense it fails check 1, because no model is choosing anything.
Which three questions expose a mislabeled agent?
Three questions, asked of a vendor or a colleague, locate the check a system fails without your reading any code. Each one asks for something that can be shown, which is what separates them from asking whether the product is “autonomous.”
- “Show me two traces from two different inputs. Where do the steps differ, and what chose the difference?” Identical step sequences on clearly different inputs suggest the system stops at check 3 or earlier; ask for a third. A difference chosen by an
ifstatement (a fixed rule in code) means the model decided nothing, which is check 1 at most. A difference the model chose once, at the start, is a router, which is check 2. An agent shows a difference chosen step after step, each time after a result came back. - “After a tool runs, who sees the result and picks the next action?” If the answer is “the next stage of the pipeline” or “the user,” the result never returns to the model as a decision. That is check 3.
- “What stops a run, what caps its cost, what is it allowed to touch, and where is each of those enforced?” “It’s in the system prompt” is a failed check 5. You are looking for a stop condition, a budget, permissions and any approval gate, held in code.
A vendor who cannot produce a trace has answered a different question for you. Whatever the system is, you would have no way to see what it did.
Where does the test give a fuzzy answer?
The test gets fuzzy wherever the model owns some of the decisions, which describes most systems in production. Chapter 1 says so directly: “Almost nothing that ships is a pure workflow or a pure agent.” The verdict is a line drawn on the autonomy dial, and six cases sit close to it.
The chat assistant with tools. Many chat products now search, read and search again inside a single reply. Within that turn the system passes checks 3 and 4, and between turns you are back in charge. The honest description is a small agent inside each turn of a chatbot.
Learned models that are not language models. A spam filter, a recommender or a classic voice assistant built on an intent classifier uses a learned model that is not a language model. Under the engineering sense used here it stops before check 1; under the 1995 sense it is an agent.
Plan first, then execute. A system where the model writes a whole plan and code runs it without reporting back fails check 3, even though nobody could have drawn its path in advance. The book’s flowchart test would call that step agentic; this test calls it one large model decision. Report both.
Checks 3 and 4 usually pass together. Most definitions fold them into the single word “loop.” I keep them apart for one case, the capped draft-and-check loop, where the model decides again and still chooses nothing except whether to repeat one scripted step.
Mixed systems. A workflow can contain one agentic step. Run the checks per step and report the highest verdict along with where it lives. The post on agents vs workflows covers those blurry cases, among them the orchestrator (a model that splits a task into subtasks and hands them out).
Closed products. You cannot run check 4 on a system that shows you no traces. You can still ask the three questions, and weigh the answers as claims.
What does the test leave out?
The test leaves out two things: whether the system should be an agent, and whether its work is right. The first question has its own procedure in the Should this be an agent? decision tool. On the second the test is silent, and the model’s fluency cannot tell you either.
The takeaway
When someone says “agent,” ask who chooses the next step, then ask to see two traces and run the five checks on them. If the model chose the difference step after step, each time after seeing a result, you have an agent. The conversation should then move to the question Chapter 1 puts first on the compass: “what signal tells you it worked?”
So the full answer to “what is an AI agent?” has two halves: a model that chooses its own path, and a check from outside that tells you where the path led. That check matters more for agents than for anything lower on the list. The book illustrates why with a number it labels as an illustration: at 95% success per step, twenty chained steps finish correctly about 0.9520 ≈ 36% of the time.
The test needs no code and no machine-learning background, which is the bar I set for anything about AI agents for beginners. Chapter 1, “What Is an Agent?” is free to read online and contains the three example systems, the dial and the ladder of simpler alternatives; the rest is in the full book, and you can see the formats.
Questions readers ask
- What is an AI agent in simple terms?
- An AI agent is a program in which a language model is given a goal and a set of tools, chooses its own next step, sees the result, and chooses again until the goal is met or a stop rule ends the run. The model owns the sequence of steps; code owns the limits.
- What is the difference between an AI agent and an LLM?
- The LLM is the component that produces text. The agent is the system built around it: the loop, the tools, the goal and the limits. The same model can sit inside a chatbot, a workflow or an agent, so swapping the model does not change which of the three you have.
- Is a chatbot with tools an AI agent?
- Only if the model keeps going on its own: it calls a tool, reads the result and decides the next call, more than once, with nobody choosing in between. A single tool call followed by waiting for your next message is a chatbot with a tool.
- Is a script that calls a model on a schedule an AI agent?
- No. It runs unattended, but its code fixed every step before the run began, and the model only fills in content. Running unattended describes the trigger; being an agent describes who chooses the path.
- Is agentic AI the same as generative AI?
- They describe different things about the same technology. Generative says what the model does, which is produce content. Agentic says how much of the control flow the model owns. An agent built on a language model is generative AI placed in a loop with tools.
- Does an AI agent have to be fully autonomous?
- No. An agent that loops over its tools and pauses for human approval before anything consequential is still an agent, because the model still chooses the steps. The book's first chapter places much of production at that position on the autonomy dial.
Sources
- Simon Willison (2025). I think “agent” may finally have a widely enough agreed upon definition to be useful jargon now
- Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents
- Stan Franklin and Art Graesser (1996). Is it an Agent, or just a Program?: A Taxonomy for Autonomous Agents
- Michael Wooldridge and Nicholas R. Jennings (1995). Intelligent Agents: Theory and Practice
- Aymeric Roucher, merve and Thomas Wolf (Hugging Face) (2024). Introducing smolagents, a simple library to build agents
- Andrew Ng (2024). Welcoming Diverse Approaches Keeps Machine Learning Strong (The Batch)