Home / Glossary

Glossary

The vocabulary of building agents, defined in plain language. Each term comes from the book’s appendix and links to the chapter where it is explained.

A

action space
The universe of actions available to an agent: the sum of what its tools permit, no more. You design an agent's competence and its risk in the same stroke, by deciding what goes in this set (Chapter 5).
agent
A language model placed in a loop and given tools. It looks at the state of the task, picks an action, sees the result, and goes again until it judges the goal met or a stopping rule ends the run; the field's one-sentence version is that an agent runs tools in a loop to achieve a goal (Chapter 1). The defining property, distinguishing it from a workflow: the model, and no flowchart of yours, decides what happens next.
agent-legible code
Code organized so an agent can navigate and safely change it: small files, explicit contracts, fast tests, conventions written down where a fresh context can find them. This book's coinage, flagged as such where it is introduced (Chapter 21).
agent protocol
See protocol.
agent-washing
Labeling a system an agent because the word sells, whatever the architecture underneath. Both directions of the mislabel cost you: a workflow sold as an agent buys oversight ceremony it does not need; an agent shipped as an “automation” skips the disciplines autonomy demands. The cure is one question: does the model decide the next step, or does code? (Chapters 1 and 14.)
appropriate reliance
Trust calibrated to demonstrated reliability: leaning on the agent exactly as much as its track record supports. The two failure directions predate agents by decades in automation research—over-trust (misuse: accepting what should be checked) and under-trust (disuse: re-doing what the machine got right) (Chapter 24).
approval gate
A point where the run halts until a person approves, rejects, or edits the proposed action. The basic instrument of oversight; place gates by consequence (what a wrong action costs), never by convenience (Chapter 12).
augmented coding
Working with a coding agent while continuing to care about the code itself—its design, its tests, its maintainability—with the agent as a collaborator whose output you read and shape. The opposite end of the spectrum of intent from vibe coding (Chapter 21).
augmented LLM
A single model call with three ports attached: retrieval, tools, and memory. The basic building block from which workflows and agents are assembled; every pattern in this book arranges copies of this unit (Chapters 1 and 10).
autonomy dial
How much an agent may do between moments of your attention, set per task rather than built into the system. Four working positions: every consequential action signed; standing permissions with escalation at the envelope's edge; plan-level approval; monitored autonomy against budgets. The same instrument seen from the product side is the autonomy slider, running from a human in the loop to a human on it (Chapters 12 and 25).
autonomy slider
See autonomy dial.
autoregressive generation
How a model writes: one token at a time, each conditioned on everything before it, prompt and output alike, with no draft held in reserve. The property behind both the power (each step can react to the last) and the bills (see prompt caching) (Chapter 2).

B

blast radius
Everything a step, a run, or an agent could break or leak if it went as wrong as possible. The book's permanent vocabulary for the cost half of every deployment bet; three dials set it—what the agent is permitted, where it runs, and which of its actions wait for a human (Chapter 17).
budget
A hard ceiling enforced by the harness, indifferent to the model's opinion: on passes through the loop, tokens, wall-clock time, or money. Mandatory, because a loop whose only exit is the model's judgment has no guaranteed exit at all; mature agents run several budgets at once (Chapter 3).

C

canary
A rollout rung: route a small slice of real traffic to the new version, watch the cohorts side by side, widen only while the numbers hold. Two disciplines make it real: consistent assignment (a user stays on one version for the session) and automated rollback with thresholds chosen in daylight (Chapter 20).
chain-of-thought
Prompting a model to write intermediate reasoning steps before committing to an answer. The written steps measurably improve performance on multi-step problems; they are also generated text, to be read as a debugging aid rather than a sworn account of the computation (Chapter 4).
checkpoint
A durable record of a run's position, taken so a crashed run can resume rather than restart. A checkpoint records position and never whether the outside world did the work; that gap is why checkpoints need idempotency beside them (Chapter 18).
chunk
The unit of retrieval: a passage of a few hundred words, small enough to be precise, large enough to be understood alone. Chunking policy quietly shapes everything downstream of a retrieval system (Chapter 8).
classifier
A system that assigns each input to one of a fixed set of categories. This book's version is built on a pretrained model configured by written class definitions—the definitions are the model—and developed with the three-set discipline (Chapter 26).
compaction
Compressing a long history—decisions taken, open problems, current plan—into a short summary that replaces it, so a long run can continue on a clean desk. The coarse grade of the compress operation (alongside surgical pruning and tool-result trimming); lossy by design, so what the summary drops is a design decision (Chapters 7 and 9).
compass, the
This book's device: four bearings stated in Chapter 1 and leaned on throughout. The first and most-used is what signal tells you it worked?—verifiability as the design question that precedes all others; the last is the simplest thing that works.
compound engineering
See compound step.
compound step
The small ritual that closes a task: before moving on, ask what the agent should have known at the start, and write it into the standing project-instructions file, one line per lesson. The habit behind what some teams call compound engineering—each run leaves the system permanently better briefed. The model forgets and cannot be trained by you; the harness can learn, and this is how (Chapter 22).
confused deputy
Classical security's name for a trusted party tricked into spending its authority on someone else's behalf. An injected agent is exactly this: every tool call it makes is signed, in effect, with your credentials, whether the instruction behind it came from you or from a page it read (Chapter 17).
constrained decoding
Enforcing an output schema inside the generation machinery: the schema becomes a filter that strikes invalid tokens from the ranked list before each draw, so the model chooses freely among valid continuations and cannot emit an invalid one. The strong grade of structured output; asking nicely in the prompt is the weak one (Chapter 2).
context engineering
The discipline of deciding what fills the model's context window at every step of a run—the right information and tools, in the right format, at the right time, decided again on each pass. Prompt engineering writes a document; context engineering operates a system (Chapter 7).
context rot
The tendency of a model's ability to use any given fact to decay as the window around it fills. The reason a bigger window does not repair a crowded one, and the standing argument for curation (Chapter 7).
context window
The bounded stretch of text a model can consider in a single call—instructions, history, tool results, and the answer being written, all sharing one token budget. This book pictures it as a desk (see desk, the) (Chapter 2).

D

dark launch
See shadow deploy.
desk, the
This book's picture of the context window: the work surface on which everything the model consults during one call must fit at once, swept bare between calls. Introduced in Chapter 2 and used to the last page; like any metaphor it is flagged where it thins.
dumb zone
The degraded state a model slides into as its window fills past a soft limit well below the hard maximum: it stops engaging with your corrections and reflexively agrees (“you're absolutely right”) while the work stops improving. The practitioner's name for context rot felt at the keyboard; the remedy is to keep window utilization low and start fresh rather than argue it back (Chapter 7).
durable execution
Writing agent code as if failure did not exist and letting a runtime make the fiction true: the result of every model call and tool call is recorded in an append-only journal, and on recovery the code re-executes from the top, with journaled steps returning their recorded values instead of running again. The code replays; the world does not (Chapter 18).

E

embedding
A long list of numbers assigning a text its position on a map of meaning, produced by a model trained so that similar meanings land at nearby coordinates. The machinery under semantic search: “dog” and “canine” are neighbors despite sharing no letters (Chapter 8).
eval
A test for an AI system: an input, plus grading logic applied to whatever comes out. The unit from which eval sets are built (Chapter 16).
eval set
A graded bank of test tasks you rerun on every change, so that “did this edit make the agent better or worse?” has an answer you would defend with a number. Grown most honestly from real failures found in traces (Chapters 15 and 16).
evaluator–optimizer
The workflow whose defining arrow points backward: one call generates a candidate, an evaluator (ideally a real verifier—a test suite, a schema check) grades it, and the failures feed back for revision, looping until the check passes or a cap fires. Also heard as generator–critic (Chapter 10).

F

feature flag
Configuration deciding which version serves which slice of traffic, so promotion and rollback are config flips—instant, reversible, no redeploy. The control plane of the rollout ladder (Chapter 20).
fine-tuning
Continuing a model's training on your own examples, so that the weights themselves change. In this book's division of labor: fine-tuning for how the model should respond, retrieval for what it should know, because knowledge belongs where you can update it, inspect it, and cite it (Chapter 8).
function calling
See tool call.

G

generator–critic
See evaluator–optimizer.
goal function
What an outer loop needs that the inner loop does not supply: a persistent goal held outside the model (the spec, the checklist, on disk, because the agent forgets between runs) paired with a machine-checkable stopping condition evaluated after each pass. A loop is exactly as trustworthy as the check that stops it—see oracle (Chapter 13).
grounding
Reasoning kept tied, step by step, to facts checked against the world, rather than allowed to elaborate on its own recollections. What the interleaved schedule of ReAct buys; the measured effect is less fabrication and less error compounding (Chapter 3).
guardrail
A check that lives outside the model's reasoning and runs deterministically, whatever the model has decided: ordinary code (occasionally a second, separate model) stationed at the agent's boundaries, empowered to allow, block, modify, or escalate what passes. The model supplies judgment; the harness supplies limits; only one of the two can be sweet-talked (Chapter 17).

H

hallucination
The field's term for output that is fluent, specific, confident, and wrong—an invented citation, a function that does not exist. The mechanism is ordinary generation doing its job where a fact is missing; the operational danger is that fabrication arrives in exactly the tone of knowledge (Chapter 2).
harness
Everything you build around the model to turn it into a working agent: the loop itself, the tool implementations and what they are permitted to touch, the care of the message history, the stop rules, and, as the system matures, the logging, recovery, and guardrails. An agent, in one line, is a model plus a harness. The field uses the word with elastic scope; this book anchors the definition in Chapter 3 and reconciles the wider senses in Chapter 18.
human in the loop / on the loop
Two ends of the autonomy dial's range, in vocabulary older than language models: in the loop means inside the run, approving actions as they happen; on the loop means above it, reviewing what an autonomous run proposes and produces (Chapters 12 and 25).

I

idempotency
The property that doing an operation twice has the same effect as doing it once, which makes retrying safe by construction. Reads have it naturally; writes are the trap, and the standard fix is the idempotency key—a token derived from the action's logical identity, under which the downstream service stores the first result and returns it on any repeat (Chapter 18).
in-context learning
Teaching a genuinely new behavior through examples placed in the window rather than through training. The reason a handful of well-chosen demonstrations can stand in for a training set—and the mechanism the classifier of Chapter 26 is built on (Chapter 2).
inner loop
See outer loop.

K

kill switch
One obvious, fast, tested way to stop every running loop and agent at once. Test it before the first unattended run, the way you test a fire alarm before the building fills (Chapter 13).

L

lethal trifecta
The three-way combination that makes an agent a data-theft risk: access to private data, exposure to untrusted content, and a channel for external communication. Any two can be lived with; all three in one agent means an injected instruction can read the secrets and mail them out. Named in the security literature and adopted by this book as a checklist (Chapter 17).
LLM (large language model)
Mechanically, a function: text in, and for every possible next token, a probability that it comes next, with the surrounding system drawing one token at a time (see autoregressive generation). Called like any API; stateless between calls (Chapter 2).
LLM-as-a-judge
Using a language model, armed with a written rubric, to score or compare outputs the way a human reviewer would—the standard escape from graders that read at human speed and bill at human rates. A judge is an instrument that must itself be calibrated against human judgment before it is trusted (Chapter 16).

M

memory
An agent's ways of carrying forward what matters, sorted in this book by the office furniture it resembles: episodic memory is the logbook (append-only records of what occurred), semantic memory is the filing cabinet (facts, needing curation because they go stale), and procedural memory is the procedures manual (how to behave—the sharpest object in the room, because an entry is the agent's future behavior). Short-term memory is simply the message history (Chapter 9).
message history
The growing, role-tagged list of everything said and done in a run: instructions, the goal, the model's replies, every tool call, every result. Because the provider is stateless, this list is the agent's entire short-term memory: whatever your code fails to append never happened (Chapters 2 and 3).
model client
The function that sends the message history plus tool definitions to the provider and returns the model's next response. One of the four parts of the minimal agent, and the only one that touches a vendor (Chapter 3; Appendix A).

O

oracle
The testing trade's word, borrowed by this book: a source of truth outside the thing being judged, which pronounces an answer right or wrong. A test suite is a strong oracle; a compiler is a crude but incorruptible one; the model that did the work is a poor oracle for that work. A loop is exactly as trustworthy as its oracle (Chapters 4 and 13).
orchestrator–worker
The multi-agent shape: a lead agent decomposes the task and delegates the pieces to workers, each run in quarantine with a fresh window and a self-contained brief, returning distilled results. Nearly every coordination failure traces to a violation at one of the two doors—the brief in, the return out (Chapters 9 and 11).
ordinal scale
A score made of ordered levels, like grades, rather than measured quantities, like weights. This book's advice for LLM scorers: prefer a small ordinal scale with each level defined like a mini-class over a continuous score the model cannot actually resolve, and recover resolution by averaging discrete judgments (Chapter 26).
outer loop
The machinery wrapped around the agent (the inner loop): it decides what the task should be, launches the agent with a fresh window, evaluates the outcome against the goal function, keeps score somewhere durable, and chooses the next move. You already run one by hand and call it the backlog; writing it down as a system is what lets it run unattended (Chapter 13).

P

pass@k and passk
Two summary statistics for a system that is a distribution rather than a function. pass@k: the probability that at least one of k attempts succeeds—the right number when a cheap verifier picks the winner from a pile. passk: the probability that all k attempts succeed—the right number for an agent that must be correct every time it acts unattended. The same agent can be a triumph on the first and a catastrophe on the second (Chapter 16).
progressive disclosure
Revealing information in layers, each layer pulled in only when the task demonstrates it is needed: a name always visible, a procedure loaded on trigger, reference files read on demand. The design principle that lets a large library of skills cost almost nothing until used (Chapter 6).
prompt caching
A provider optimization exploiting autoregressive generation: the reading-state of a prompt's front depends on that prefix alone, so a stored prefix state can be reused by any later call whose opening tokens match byte for byte, at a steep discount. The design consequence is a rule: stable content first, variable content last (Chapter 19).
prompt engineering
The craft of writing the instruction slices of the window—the wording of the system prompt, the shape of the examples. A real craft, and a subset: context engineering answers for the whole desk (Chapter 7).
prompt injection
The attack family in which instructions are smuggled into text the model reads, exploiting the fact that the model sees one undifferentiated token stream with no structural boundary between trusted instructions and merely-read content. Direct: the attacker is the user, typing. Indirect: the instruction hides in content the agent encounters doing legitimate work—a web page, an email, a code comment—and this is the main event for tool-using agents, whose job is reading things other people wrote. Every known defense is statistical; this book's working rule is to treat any claim of prevention as false and to engineer the blast radius instead (Chapter 17).
protocol
An agreement about how two independently written programs talk—message shapes, sequences, promises—so that strangers' code cooperates without bespoke glue per pairing. Agent practice uses two kinds: tool protocols connect an agent to capabilities and data (the Model Context Protocol is the prominent labeled example), and agent protocols connect it to other agents that reason, plan, and answer in their own time (Agent2Agent is the labeled example). Either way the payoff is arithmetic: adapters stop multiplying as connectors × agents (Chapter 6).

Q

quarantine
Isolating work in its own context window: delegate a self-contained subtask to a subagent with a fresh window and a small toolset, let it burn what it burns, accept back a distilled result. Two doors, strictly kept: what may enter is the brief you wrote, and what may come back out is conclusions—treated as evidence to weigh, never instructions to obey. Also the wall between untrusted content and consequential action (Chapters 7, 9, and 17).

R

RAG (retrieval-augmented generation)
The pattern behind most systems that answer from your documents: at request time, retrieve the most relevant chunks from an index and place them in the window beside the question, so the model answers from evidence on the desk rather than from its frozen training data (Chapter 8).
ReAct
The interleaved schedule at the heart of the agent loop—a short written thought, one action, one observation, then the next thought conditioned on what just came back—named in the research literature as Reason + Act. The rigid textual format of the original has dissolved into model training; the schedule survives (Chapter 3).
regression gate
An eval set wired into release machinery as an automatic barrier: the change ships only if the bank of graded tasks still passes. The gate tests the failures you already imagined; the rest of the rollout ladder exists because reality imagines better (Chapters 16 and 20).
regressor
A system that assigns each input a number rather than a category. For LLM-based scorers, see ordinal scale (Chapter 26).
reliability envelope
The spread between pass@k and passk as k grows—a picture of an agent's temperament. A narrow envelope is a steady agent; a wide one is a talented gambler, capable of the task but not to be trusted with it unattended (Chapter 16).
review theater
Diligence performed at a volume where it can no longer be real: scrolling four hundred generated lines with the care draining out, while the green checks beside the diff certify internal consistency and say nothing about whether this was the right thing to build (Chapter 12).
routing
Classify, then branch: a cheap decision step sends each input down the path built for its kind. One of the three basic workflow shapes, and, usefully, a classifier wearing work clothes (Chapters 10 and 26).

S

sandbox
A hard boundary around where the agent runs, enforced by the operating system or hypervisor rather than by anyone's judgment. The logic is subtraction: a credential never mounted inside the boundary cannot be read by any injection, because the target is absent. Build it from battle-tested primitives; the weakest layer is the one you built yourself (Chapters 17 and 20).
shadow deploy
Running a candidate version against real traffic while showing users only the incumbent's output, so the new behavior is measured at production scale with zero user exposure. Also called a dark launch; the rollout ladder's second rung (Chapter 20).
skill
A bundle of instructions, reference files, and scripts an agent loads on demand—packaged expertise: the judgment about the verbs, where a tool is a verb. Loaded through progressive disclosure, versioned like code, and audited like a dependency, because an instruction file is behavior (Chapter 6).
span
One unit of work in a trace: a single operation with a start time, an end time, and structured attributes. Every model call is a span, every tool call is a span, every subagent is a subtree; the trace is the tree they form (Chapter 15).
standing project-instructions file
The file a mature agent reads at the start of every session: build commands, conventions, the rules of the house—procedural memory in its plainest working clothes, persisting what the agent cannot infer from the code alone. This book also calls it the constitution. Kept ruthlessly small (for each line: would removing it cause mistakes? if no, cut), and improved by the compound step (Chapters 9 and 22).
stop condition
The rules that end a run. Three families: the model's own judgment that the goal is met (testimony, and the least trustworthy exit); the budget, mandatory; and error exits, split by one question—can the model act on this error, or is it evidence the run is off the rails? Feed back the first kind; halt on the second (Chapter 3).
structured output
Model output constrained to a machine-readable shape—typically JSON against a schema—so ordinary code can branch on what the model decided. The bridge that makes tool use, and so agents, possible; comes in two grades (see constrained decoding) (Chapter 2).
subagent
Another stateless engine whose entire briefing is whatever you put in its delegation. The definition matters more than the anthropomorphism: it was not in the room for anything, knows only the brief, and returns only what its author disciplined it to return (see quarantine) (Chapters 7 and 11).
supervision surface
An interface built for judging work in progress rather than conversing about it: the run list, the approval queue, the trace, the evidence, the recovery controls—a board, where chat is a scroll (Chapter 24).
system prompt
The message slot carrying the standing instructions: who the model is supposed to be, what it may and may not do, the shape its answers must take. Written once, resent by the harness on every call, and—a fact worth tattooing somewhere—a request rather than an enforcement mechanism; see guardrail (Chapters 2 and 17).

T

three-set discipline
Developing an LLM classifier without fooling yourself, using three disjoint sets: the demonstrations packaged into the skill, the development set you iterate against, and a held-out test set consulted rarely and retired when overused. The training set could be dispensed with; the discipline that grew up around it could not (Chapter 26).
token
The unit a model actually reads and writes: a chunk of text from a fixed vocabulary—often a word, sometimes a piece of one; on average, think three-quarters of an English word. Tokens are what the context window is measured in and what the meter bills (Chapter 2).
tool
A function the model can ask your program to run, described to it as a name, a description, and an argument schema; the function itself never leaves your process. The model requests; your code executes—the division of authority the whole architecture rests on (Chapters 3 and 5).
tool call
The model's structured request to run a tool: the tool's name plus arguments fitting its schema, emitted as structured output instead of prose. Also called function calling, interchangeably (Chapter 2).
trace
The complete, structured record of one run—every model call, tool call, and result, as a tree of spans carrying timings, token counts, and costs. The same record is a transcript when read as a conversation and a trajectory when treated as one sampled path through a task's possibilities: one object, three trades looking at it (Chapter 15).
trajectory
See trace.
transcript
See trace.
trust calibration
See appropriate reliance.

V

vector database
A store built to answer one question quickly: given a new embedding, which stored ones sit nearest? The index behind semantic retrieval, though plenty of ordinary databases have learned the trick (Chapter 8).
verification gap
The distance between how convincing an output looks and how cheaply it can be checked. A property of the domain, and the reason research agents fail differently from coding agents: for open-ended factual claims there is no compiler for facts, only partial substitutes with limits (Chapter 23).
vibe coding
The mode in which you do not read the code before running it: you judge the behavior, and the code is a means you decline to inspect. Legitimately useful for throwaways and prototypes; the term was never meant to cover all AI-assisted coding—see augmented coding (Chapter 21).

W

workflow
Several model calls and tool calls wired together along paths your code defines in advance: the model fills in content at each step, and your code decides which step comes next. You could draw its flowchart before switching it on—the property that separates it from an agent (Chapters 1 and 10).
write / select / compress / isolate
The four context operations—the whole repertoire: write context down outside the window (files, notes, memory); select what comes in (retrieval, tool loadout); compress what stays (trimming, compaction); isolate what belongs elsewhere (quarantine, subagents). Few enough to memorize, and this book has yet to meet a context remedy that is not some combination of them (Chapter 7).