AI Agents, Engineered

Chapter 1

What Is an Agent?

16 min read · 4 figures

On this page
  1. 1.1 The Agent Idea
  2. 1.2 Agents, Workflows, and Chatbots
  3. 1.3 Do You Even Need an Agent?

This chapter keeps a promise made in the preface: before anything gets built, we pin down what an agent is. The need is practical, and immediate. Vendors label scripted bots “agents”; researchers reserve the word for systems that run unattended for hours; and somewhere in between, you have to make actual decisions with actual budgets. If the term stays fuzzy, you will reach for an autonomous agent where a three-step pipeline would have been cheaper and easier to debug, or you will hand-wire a rigid pipeline for a problem whose steps you genuinely cannot predict. The chapter does three things. It states the agent idea itself, explains why it became practical when it did—deliberately without dates—and lays out the compass the book will steer by. It then draws the working distinction among chatbots, workflows, and agents that every later chapter leans on. And it closes with the question too few people ask out loud: whether you need an agent at all.

The Agent Idea

Strip away the demos and the marketing, and the idea is small enough to hold in one hand. An agent is a language model placed in a loop. It is given a goal and a set of tools—functions it can ask your program to run: search this, read that, execute this command. It looks at the state of the task, picks an action, sees the result, and goes again, until it judges the goal met or a stopping rule ends the run. After years of talking past itself, the practitioner community has, at the time of writing, largely converged on a one-sentence version: an agent runs tools in a loop to achieve a goal.1 Simon Willison, “I think ‘agent’ may finally have a widely enough agreed upon definition to be useful jargon now,” simonwillison.net (September 2025). The definition quoted—“An LLM agent runs tools in a loop to achieve a goal”—is his, defended there phrase by phrase. Every word of that sentence is load-bearing. “Tools” is what connects the model to the world; “loop” is what lets it react to what it finds; “a goal” is what ends the loop—an agent has a stopping condition, which is what separates it from a runaway process. Chapter 3 unpacks the loop in code.

What is genuinely new here is where the control flow lives. In every program you have written until now, the sequence of operations was decided by you, in advance, in code. In an agent, the sequence is decided by the model, at runtime, in response to what it observes. That single change is what makes agents worth a book, and also what makes them expensive, slow, and occasionally alarming. The rest of the book is about both halves of that sentence.

You may reasonably ask why this became practical when it did. Language models were fluent long before anyone trusted one with a loop; fluency was never the missing piece. Three things had to mature. Each will return later in the book, first as machinery and then as a failure mode when it runs short.

The first: models learned to take instructions. A model fresh out of its initial training is a pure continuer of text—ask it a question and it may reply with more questions, because that is a plausible way for text to continue (Chapter 2 shows why). A later stage of training, which the field calls post-training, turns that continuer into something you can give a job to: it follows directions, works toward the stated objective, and stops when done. An agent is a delegation, and delegation requires a worker that accepts the assignment.

The second: text acquired a bridge to action. Function calling—the model emits a machine-readable request, “call this function with these arguments,” which your code executes before feeding the result back—gave the model a disciplined channel to the world. The model proposes; your code disposes. Chapter 2 covers the mechanism in detail. The consequence is the point here: a system that could only ever write about the world became one that can ask, in a form a program can safely obey, for things to happen in it.

The third: the model’s working memory grew. Everything an agent knows in the moment must fit inside the model’s context window (the bounded stretch of text a model can consider in a single call). When windows held a few pages, a loop starved: the instructions, the history, and the latest tool result would not fit together. Windows grew to hold something closer to a working session, and loops became sustainable.

Each of these matured as a slope quietly crossing a threshold, which is convenient for both of us: the explanation of “why now” does not expire. It also tells you what to check when some new capability is announced: ask which of the three ingredients just got better.

Before the taxonomy and the machinery, I want to hand you the compass this book steers by: four bearings that will reappear in nearly every chapter. I state them here without their full defense; the defense is the book.

Verifiability comes first. The preface stated the thesis and it bears repeating: an agent is only as trustworthy as the signal you can use to verify it. The model’s own confidence carries no information about whether it is right, so trust must come from outside—a passing test, a validating schema, a source you can open, a human who approves. The first question to ask of any agent design is the one this book will ask over and over: what signal tells you it worked?

Next comes the scarcity of context. The context window is finite, and the part that surprises people is that the model’s attention degrades before the space runs out: a fact buried in the middle of a long pile is, in effect, half-forgotten. Treat context as a budget to be spent deliberately. Part III of this book is about spending it well.

Compounding error is the third. An agent chains many steps, and reliability multiplies across them. Suppose, purely as an illustration, that each step succeeds 95% of the time; a run of twenty chained steps then succeeds about 0.9520 ≈ 36% of the time. The particular numbers vary; the arithmetic does not, and it is merciless. Long loops survive only with recovery, checkpoints, and verification along the way. This book spends whole chapters on that machinery.

And the last bearing is the simplest thing that works. The engineering advice that has aged best in this young field is blunt: find the simplest solution possible, and add complexity only when it demonstrably pays—which may mean not building an agent at all.2 Erik Schluntz and Barry Zhang, “Building Effective Agents,” Anthropic engineering blog (2024): “we recommend finding the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all.” The essay’s workflow–agent distinction anchors the next section as well. Sometimes the simplest thing is a single prompt. Sometimes it is a regular expression. The last section of this chapter takes this bearing seriously enough to give it its own discussion. Figure 1.1 sets the four bearings on a single rose, with the needle resting on the first of them.

The book’s compass.
Figure 1.1 The book's compass. Four bearings decide the design forks in almost every later chapter, and the needle rests on the first of them—verifiability—because an agent is only as trustworthy as the signal you can use to check it. The other three, the scarcity of context, compounding error, and the simplest thing that works, are read alongside it, not instead of it.

A compass is a modest instrument; it will not design your system for you. But whenever the book reaches a fork—more autonomy or less, another agent or a plain script, richer context or leaner—you will watch the same four bearings decide it, and by the end I hope they will be deciding your forks too.

Agents, Workflows, and Chatbots

Let me put three systems in front of you. Decide which of them is the agent before reading on.

The first is a question-and-answer assistant. You paste in a function and ask what it does, and it explains; you paste a diff and ask for a draft commit message, and it writes one; then it waits for your next message. The second processes every incoming support ticket the same way: a model call classifies the ticket, ordinary code fetches the matching policy document, a second model call drafts a reply, a third checks the draft against a rubric, and the result lands in a human review queue—four steps, in that order, every time. The third is told “the checkout tests are failing—fix them,” and, unsupervised, it searches the codebase, reads three files, edits one, runs the tests, reads the new failure, edits again, and stops when the suite is green.

Most people’s instinct sorts these instantly: the third one feels different in kind, and the instinct is right. The useful work is saying precisely why. All three may be built on exactly the same underlying model. What differs is the question the last section called the heart of the idea: who decides what happens next. In the first system, you do, one message at a time. In the second, your code does; the model only fills in content along a route that was fixed before the first ticket ever arrived. In the third, the model decides: which file to read, whether to edit or rerun, when to stop. That question—who owns the control flow—is the entire distinction, and it gives us the three working terms this book uses everywhere. Figure 1.2 draws all three side by side, differing only in who decides the next step.

Chatbot, workflow, and agent differ in one thing only: who decides the next step.
Figure 1.2 Chatbot, workflow, and agent differ in one thing only: who decides the next step. In a chatbot the human drives every turn; in a workflow the path is fixed in code and the model merely fills in each step; in an agent the model owns the loop, choosing tools and reading results until it reaches the goal. The model-owned loop (in accent) is the whole subject of this book.

A chatbot is turn-by-turn conversation. The human drives every step; the model generates text and waits. There is no loop and nothing is executed. Chatbots are genuinely useful and often the right tool, and their failure surface is small in a way I want on the record: the worst a chatbot can produce is a wrong answer, because acting on the world is out of its reach.

A workflow is several model calls (and tool calls) wired together along paths your code defines in advance. The model fills in the content at each step—classify this, draft that—but your code decides which step comes next. The recurring shapes have names you will meet again in Chapter 10: chaining (one call’s output feeds the next), routing (classify, then branch), and parallelization (fan out several calls, then aggregate). The ticket pipeline above is a workflow, and you could have drawn its flowchart before switching it on.

An agent hands the model the control flow. One widely cited engineering essay draws the line exactly here: workflows are “systems where LLMs and tools are orchestrated through predefined code paths,” while agents are systems where models “dynamically direct their own processes and tool usage.”3 Erik Schluntz and Barry Zhang, “Building Effective Agents,” Anthropic engineering blog (2024); both phrases are quoted from the essay, which also names the augmented LLM building block described in the next paragraph. An earlier survey—Lilian Weng, “LLM Powered Autonomous Agents,” lilianweng.github.io (2023)—describes the same anatomy from the inside: a model “brain” supported by planning, memory, and tool use. The number of steps, and which step comes when, exists nowhere in your code; the model chooses as it goes, reacting to whatever it just observed. That is why the test-fixing system feels different: its transcript reads like the log of a colleague’s afternoon, and only distantly like the execution of a program.

Underneath the last two sits the same atom, worth naming because the patterns you will meet in Chapter 10 are assembled from it. The augmented LLM is a model call enhanced with three things: retrieval (pulling relevant information in at request time), tools (functions it can request), and memory (some way of carrying forward what matters). A workflow is several of these atoms wired together by your code along a fixed path. An agent is one of these atoms placed inside a loop, deciding for itself what to do on each pass. (A chatbot, in this picture, is the bare model call with the augmentations mostly unused.) Same atom, two arrangements—keep that picture and the taxonomies stop feeling mysterious.

Here is the litmus test I will use for the rest of the book. If you can confidently draw the control-flow diagram before the request arrives, you are looking at a workflow. If the diagram can only be drawn in hindsight, once the model has reacted to what it found, you are looking at an agent.

It is tempting to treat chatbot, workflow, and agent as three boxes. The truth is a dial (Figure 1.3). At one end, the model generates text and a human decides everything. Near that same end sits the fixed pipeline (the ticket system from a moment ago), where code owns every step and the model only fills in content. A small step up, a workflow grants the model one decision: a router picks which hardcoded branch handles the request. Further along, an agent loops freely over its tools but pauses for human approval before anything consequential—sending the email, spending the money, deleting the data; you will find much of production here. At the far end, the model owns the path end to end for many steps without per-step confirmation: maximum capability, maximum blast radius—the damage a wrong action could do before anything stops it, a term Chapter 17 will make precise. What moves along this dial is a single quantity, how much of the control flow the model owns, and mature systems choose a position deliberately rather than inheriting one from a demo.

Autonomy is a dial, not a switch.
Figure 1.3 Autonomy is a dial, not a switch. The quantity that varies along it is how much of the control flow the model owns. The labels are stops worth naming; real systems usually sit between them.

One more observation before we leave the taxonomy, because it forecasts where this book ends up. Almost nothing that ships is a pure workflow or a pure agent. The dominant production shape is a workflow shell with one or two genuinely agentic steps inside it: predictable structure wherever you can predict, model-owned autonomy only at the junctures you cannot enumerate. You are setting a dial, step by step, and the next section is about how far to turn it.

Do You Even Need an Agent?

Now that you can tell an agent from a workflow, you face the question this section exists to make respectable: do you need one at all? I put it this early on purpose. The mistake this field warns its newcomers about most often is reaching for an autonomous agent when a single prompt, a fixed pipeline, or fifty lines of ordinary code would have been cheaper, faster, and easier to debug. An agent demo is a seductive thing. You watch a model plan, act, recover, and finish, and the natural conclusion is that this power belongs everywhere. Resist the conclusion long enough to price it.

Compared with a single model call, an agent adds several recurring costs, and they arrive as a package. Money, first: an agent makes many model calls per task instead of one, and each call re-sends the growing conversation, so cost scales twice over—once in the number of calls, again in the size of each. A ten-step loop can cost an order of magnitude more than the single call it replaced (the multiplier is illustrative; the direction is what matters). Latency, second: the calls are sequential, each waiting on the last tool result, so a sub-second answer becomes many seconds or minutes; if a human is waiting on the other end, this alone can disqualify the design. Nondeterminism, third: the same input can take a different path and produce a different answer on each run, which quietly breaks your habits for testing, debugging, and auditing.4 One published analysis of agent reliability reports single-run success rates around 60% falling to roughly 25% when the same task must succeed on eight consecutive runs—figures illustrative of the consistency gap, not permanent facts. Fiddler AI, “AI Agent Failure Rate: Why 70–95% Fail in Production.” And the failure surface changes in kind, which is the cost I most want you to feel: the bounded failure the last section credited to the chatbot is exactly what an agent gives up. Its mistake is a wrong action, followed by ten more actions built on top of it, compounding down the loop exactly as the compass arithmetic predicts.

Two further costs are easy to miss when the demo is glowing. Granting a model tools and autonomy widens your attack surface; the security community has a name for over-granting it: excessive agency, meaning more permissions, functionality, or freedom than the task requires.5 OWASP, “Top 10 for LLM Applications”—the entry on Excessive Agency. The danger also extends past what you grant deliberately: text an agent merely reads can steer what it does next, a family of attacks called prompt injection that Chapter 17 examines. Finally, the operational burden: a chatbot needs a prompt, while an agent additionally needs tool plumbing, context management, error recovery, stop conditions, approval gates, tracing, and evaluations. That surrounding machinery is called the harness, and building it is most of the work; a brilliant model in a flimsy harness fails in production.

So the honest default is a ladder, which is the compass’s fourth bearing made practical, and the discipline is to stay on the lowest rung that solves your problem. The bottom rung is plain code: a regular expression, a SQL query, an if statement—free, instant, deterministic, auditable; if the logic is fully specifiable, no model is required. The next rung is a single well-crafted model call, perhaps with retrieval and a few examples in the prompt; a surprising fraction of “we need an agent” dissolves into “we needed a better prompt.” The rung above that is a workflow: several calls on paths you drew in advance. The top rung is an agent. Each rung up buys adaptability and pays for it in money, latency, and predictability. For a problem whose steps you can draw in advance, a lower rung is the correct answer permanently; staying there is the discipline. Figure 1.4 draws the ladder, with adaptability rising as you climb and cost rising with it.

The escalation ladder.
Figure 1.4 The escalation ladder. Each rung up—from plain code to a single model call to a workflow to an agent—buys adaptability and pays for it in cost, latency, and unpredictability. The discipline is to stand on the lowest rung that solves the problem (in accent), and to climb only when a real case proves the rung below cannot reach.

The working test is the one from the previous section, now used as a decision aid: can you draw the flowchart before the request arrives? If yes, stop at a workflow or below—known steps, auditable correctness requirements, tight latency or cost budgets, and simple lookup-classify-transform tasks all point down the ladder. What points up the ladder is the genuinely open-ended goal: the steps depend on what the model discovers, and judgment is required at branches you cannot enumerate. Even then, insist on the compass’s first bearing in concrete form—a clear success criterion, a feedback signal at each step, and a sensible place for a human to look. A concrete pair to calibrate on: “route each incoming email to billing, technical, or sales and draft a first reply” has knowable steps and high volume, and wants a workflow; “make this failing test pass in a codebase you have never seen” has unknowable steps and a built-in verifier, and is a genuine agent problem.

One trap deserves its own paragraph, because it argues so eloquently for the wrong design: “it might need to adapt someday.” That is a hypothesis, and cheap to test honestly: build the simpler version first, measure it against real cases, and climb the ladder only when you can point at a class of inputs the simple system provably fails on and confirm that an agent actually does better. The reverse trap is agent-washing: because the word sells, scripted pipelines get relabeled as agents. The question that cuts through both is the litmus test you already own, asked about the system in front of you, whatever its label says: does the model decide the next step, or does code? Chapter 14 returns to this decision with the benefit of everything in between; Chapter 16 shows how to do the measuring.

An agent, then, is a cost you pay for adaptability you can name. If you cannot name the adaptability, keep your money. And if you can—if the steps genuinely cannot be drawn in advance and a verifying signal exists—then an agent is the right tool, and the next several chapters are about building one well. First, though, you need to understand the engine we will be looping: Chapter 2 is about what a language model actually is.