AI Agents, Engineered

Preface

Preface

This book exists because of a particular moment, and I suspect you have already had it. You hand a language model a task that would take a person half an hour: find out why the build broke, say, or reconcile a messy spreadsheet against the database. Instead of an answer, you watch it work. It searches, reads, tries something, hits an error, corrects itself, and comes back a few minutes later, finished. The first feeling is delight. The second feeling arrives a beat later and lasts longer: I didn’t see every step. How much of this can I trust?

That second feeling is the subject of this book.

messy work in, finished work out, the steps between behind frosted glass
Figure 1 Messy work goes in one side of a frosted-glass wall and finished work comes out the other; the steps between are only silhouettes. The delight is immediate. The doubt outlasts it, and it is the subject of this book.

I wrote it for a practicing engineer—someone comfortable with software but not necessarily with machine learning—who is evaluating agents, or has just started building with them, and wants to get past both the hype and the dread to the engineering underneath. You will find definitions built from first principles, patterns that recur across every domain where agents do useful work, examples in plain pseudocode, and a persistent interest in what things cost and how they fail. The only prerequisite is ordinary engineering literacy; whenever a term of art appears, I define it on the spot, in plain words.

Two promises shape every chapter, and I want to make them explicit, because together they explain what you will and will not find here.

The first promise is that the book is time-agnostic. This field improves at a pace that embarrasses anyone who writes things down; while I was gathering material for these chapters, notes I had captured only months earlier were already going stale, their products renamed and their benchmark numbers overtaken. So I have refused to anchor anything on the season it was written in. No claim in this book depends on which model is currently best, what a token costs this quarter, or how large this year’s context windows are. Where a number appears, it is marked as an illustration. The aim is a book that reads as well a few years from now as it does today.

The second promise is that it is solution-agnostic. The durable knowledge in this field lives in concepts and patterns—the loop, the tool contract, the context budget, the evaluation harness—and those transfer across every vendor and framework. The syntax of any one product does not transfer, and it ages in months. So this book teaches the pattern, and names concrete tools only in passing, as labeled examples of a category. When you finish it, you should be able to sit down with whatever harness (the software shell an agent runs inside) your team has adopted, or with none at all, and know what to build.

One idea recurs so often in these pages that it amounts to a thesis, and you should meet it now: an agent is only as trustworthy as the signal you can use to verify it. A language model produces fluent, confident text whether it is right or wrong; its tone tells you nothing you can rely on. The trust you place in an agent therefore has to come from somewhere outside the model: a test suite that passes, a schema that validates, a source you can open and check, a human who approves the irreversible step. Where such a signal exists, you can grant autonomy generously—that, as we will see, is most of why coding became the first domain where agents earned their keep, since code arrives with its own verdict: compile it, run the tests. Where no signal exists, autonomy is a leap of faith, and this book will keep saying so. Nearly every design question in the chapters ahead resolves to the same first move: ask what signal tells you it worked.

A word about vocabulary. “Agent” may be the most overloaded term in modern software; it is applied with equal confidence to a scripted bot reading from a phone tree and to a system trusted to work alone for an afternoon. Jargon is only useful when speaker and listener share a definition; otherwise two people can passionately discuss entirely different things.1 Simon Willison, “I think ‘agent’ may finally have a widely enough agreed upon definition to be useful jargon now,” simonwillison.net (September 2025). The observation about shared jargon, and the working definition this book adopts in Chapter 1, are his. So the first chapter does the unglamorous work of pinning the terms down, and the rest of the book holds to them.

As for how to read it: the book is divided into seven parts, ordered as a course. Part I establishes what an agent is and how the engine underneath—the language model—actually works. Parts II and III build one: the loop, the tools, the reusable skills, and the careful management of the model’s working memory. Part IV is about architecture: workflows, multi-agent systems, the dial between oversight and autonomy, the craft of writing the harness’s outer loop, and the discipline of asking whether the simplest thing that works needs an agent at all. Parts V and VI cover what most of real agent engineering turns out to be: evaluation, observability, security, reliability, cost, and deployment. Part VII turns to applications, and to what seems durable on the road ahead. The appendices hold a minimal agent in annotated pseudocode, a glossary, and a map for further reading. Read the parts in order for a grounding, or, once oriented, raid them as a reference; the chapters cross-reference one another and try to stand on their own.

I should be honest about what a book like this can promise. Some details here will be superseded, probably sooner than I would like; I have tried to write only the parts that outlast their examples, and to flag my uncertainty where I have it. The bet this book makes is that principles outlive products. If, a few years from now, the examples feel quaint but the questions—what signal verifies this? what is the simplest thing that works?—are still the questions you ask, the book will have done its job.