This AI agents study guide reviews the working vocabulary of agent engineering in one evening. It has three parts: the 87 defined terms in the glossary of AI Agents, Engineered, grouped by the seven parts of the book; a 35-question self-test with a separate answer key; and the one question all 27 chapters keep asking: what signal tells you it worked?
I wrote the book, so read this page knowing that. I built this guide for the evening before a date: an exam on a course that uses the book, an interview, or a meeting where the vocabulary will go past quickly. You have probably read and watched plenty already. What you lack is a way to find out, tonight, which terms you can explain and which you only recognize.
It is a long page because it holds the whole term table. With scripts on, the table can be narrowed to one part and the self-test questions can be ticked, and your browser remembers the ticks. With scripts off, the full table and every question are still on the page.
What does this AI agents study guide cover, and what does it leave out?
The guide covers the vocabulary and a gap-finder: 87 terms mapped to chapters, 35 recall questions, an answer key and an order to work in. It leaves out the arguments behind the terms, which are in the 27 chapters, and it cannot be your first exposure to ideas you have never read about.
The numbers make the limit concrete. The 87 definitions run to about 4,500 words (about 4,600 for the whole appendix, with its introduction and the 11 synonym lines), which is roughly 21 minutes at the web reader’s rate of 220 words per minute. The 27 chapters run to about 176,000 words by my count of the manuscript, which is about 800 minutes, a little over 13 hours, at the same rate. An evening is enough for the first and nowhere near enough for the second.
One more boundary, stated before you invest the time. At the time of writing, the Preface, Part I (Chapters 1 and 2) and the glossary are free to read online; Parts II to VII are in the full book. All 87 definitions are free, so every answer below can be checked without buying anything. The book is in pseudocode throughout, so nothing here or there is code you can run.
What is the one question?
The one question is “what signal tells you it worked?” The Preface of AI Agents, Engineered states the thesis behind it, “an agent is only as trustworthy as the signal you can use to verify it,” and the glossary files the question under the compass, the book’s name for its four recurring design bearings.
The Preface gives the reason in two sentences. A language model “produces fluent, confident text whether it is right or wrong; its tone tells you nothing you can rely on.” So trust has to come from outside the model, and the Preface lists where: “a test suite that passes, a schema that validates, a source you can open and check, a human who approves the irreversible step.”
Chapter 1 names the four bearings as verifiability, the scarcity of context, compounding error and the simplest thing that works. The question belongs to the first. If you remember one sentence from this page in an interview, make it that question, asked about whatever design is on the table. The agent verifiability scorecard applies it to a design step by step.
Where does each part restate it?
Six of the seven parts restate the question in those words, and each gives it the vocabulary of its own subject. Part VI does not repeat it verbatim; I searched Chapters 19 and 20 and found the idea applied, through the regression gate, without the sentence. The wording below is verbatim from the manuscript, with chapter cross-references written as numbers.
| Part | Chapter | The book’s wording |
|---|---|---|
| Preface | Free | Nearly every design question in the chapters ahead resolves to the same first move: ask what signal tells you it worked. |
| I | Ch. 1, free | The first question to ask of any agent design is the one this book will ask over and over: what signal tells you it worked? |
| II | Ch. 3 | The model’s “done” is testimony; a green test is evidence. This is the book’s compass—what signal tells you it worked?—applied to the humblest question an agent faces |
| III | Ch. 7 | The first question this book asks of any agent is Chapter 1’s: what signal tells you it worked? The standing question of Part III sits right behind it: what exactly was on the desk when it failed? |
| IV | Ch. 13 | So the compass question this book keeps asking—what signal tells you it worked?—becomes, here, a literal component. In the next paragraph: A loop is exactly as trustworthy as its oracle. |
| V | Ch. 16 | I have asked one question of every design in this book—what signal tells you it worked?—and evaluation is that question given a budget and a schedule |
| VI | Ch. 20 | No verbatim restatement. The nearest line makes the rollout ladder’s first rung “Chapter 16’s regression gate, the suite whose numbers decide the merge.” |
| VII | Ch. 27 | What stays is the question every chapter turned out to be asking, the one I hope now sits at the front of your mind whenever someone shows you an agent, including your own: what signal tells you it worked? |
Read down the column and you have the answer to the judgment question in each part’s own words: an exit a program can check, the contents of the desk, an oracle, an eval set, a regression gate.
How are the 27 chapters organized?
The 27 chapters sit in seven parts that the Preface describes as “ordered as a course”: what an agent is, how to build one, how to manage its context, which architecture to choose, how to make it reliable, how to ship it, and where it is applied. Appendix B, the glossary, indexes all of it.
Term counts come from a script run over the glossary source, which files each term under the chapter named in its entry’s closing pointer (the first, where the pointer names two). Chapter minutes are my estimates from manuscript word counts at 220 words per minute, so treat them as approximate.
| Part | Chapters | What the part is about | Terms | Chapter reading (approx. min) | Free or locked |
|---|---|---|---|---|---|
| I Foundations | 1–2 | What an agent is and how the language model underneath works | 17 | 57 | Free |
| II Building Agents | 3–6 | The loop, planning and self-correction, tools, skills and protocols | 13 | 125 | Full book |
| III Context Engineering | 7–9 | Managing the context window, retrieval, memory across long runs | 16 | 81 | Full book |
| IV Patterns and Choices | 10–14 | Workflows, multi-agent systems, oversight, the outer loop, when not to build an agent | 9 | 137 | Full book |
| V Making Agents Reliable | 15–18 | Observability, evaluation, security and guardrails, reliability and state | 17 | 135 | Full book |
| VI Shipping and Operating | 19–20 | Cost and latency, deploying and scaling | 4 | 55 | Full book |
| VII Applications and the Road Ahead | 21–27 | Coding agents, research and business agents, agent UX, recipes, classifiers, the frontier | 11 | 210 | Full book |
| Total | 27 | 87 | about 800 |
Four chapters are home to no glossary term under that rule: 11, 14, 25 and 27. Their vocabulary is filed under an earlier chapter that introduces it, so a short row in the table does not mean a thin chapter. Three appendices follow the parts: a minimal agent in annotated pseudocode, the glossary and a reading map.
If you want a plan for reading the chapters themselves, that is a different page: an AI agents learning roadmap in four stages. The post on what the free chapters cover, section by section maps Part I in detail.
Which terms belong to which part?
Each of the 87 defined terms belongs to the part whose chapter treats it properly, and the table below lists all of them with a link to the free definition. Grouping by part is the point: an alphabetical list hides which ideas go together, and a term is easier to recall beside its neighbors.
A commenter on Hacker News described the need well in May 2026, after a meeting that left them feeling lost: they wanted “a little 5 mins flash-card type markdown list” of what each thing is (comment by keyle). This AI agents study guide is that list, with one addition: where each term lives.
The table is sorted by part, and with scripts on you can pick one part to hide the others. The first time through, say what each term means before you open its page, then check yourself against the definition.
| Term | Home chapter | Part |
|---|---|---|
| agent | Ch. 1 | Part I |
| agent-washing | Ch. 1 | Part I |
| augmented LLM | Ch. 1 | Part I |
| The compass | Ch. 1 | Part I |
| workflow | Ch. 1 | Part I |
| autoregressive generation | Ch. 2 | Part I |
| constrained decoding | Ch. 2 | Part I |
| context window | Ch. 2 | Part I |
| hallucination | Ch. 2 | Part I |
| in-context learning | Ch. 2 | Part I |
| LLM (large language model) | Ch. 2 | Part I |
| message history | Ch. 2 | Part I |
| structured output | Ch. 2 | Part I |
| system prompt | Ch. 2 | Part I |
| The desk | Ch. 2 | Part I |
| token | Ch. 2 | Part I |
| tool call | Ch. 2 | Part I |
| budget | Ch. 3 | Part II |
| grounding | Ch. 3 | Part II |
| harness | Ch. 3 | Part II |
| model client | Ch. 3 | Part II |
| ReAct | Ch. 3 | Part II |
| stop condition | Ch. 3 | Part II |
| tool | Ch. 3 | Part II |
| chain-of-thought | Ch. 4 | Part II |
| oracle | Ch. 4 | Part II |
| action space | Ch. 5 | Part II |
| progressive disclosure | Ch. 6 | Part II |
| protocol | Ch. 6 | Part II |
| skill | Ch. 6 | Part II |
| compaction | Ch. 7 | Part III |
| context engineering | Ch. 7 | Part III |
| context rot | Ch. 7 | Part III |
| dumb zone | Ch. 7 | Part III |
| prompt engineering | Ch. 7 | Part III |
| quarantine | Ch. 7 | Part III |
| subagent | Ch. 7 | Part III |
| write / select / compress / isolate | Ch. 7 | Part III |
| chunk | Ch. 8 | Part III |
| embedding | Ch. 8 | Part III |
| fine-tuning | Ch. 8 | Part III |
| RAG (retrieval-augmented generation) | Ch. 8 | Part III |
| vector database | Ch. 8 | Part III |
| memory | Ch. 9 | Part III |
| orchestrator–worker | Ch. 9 | Part III |
| standing project-instructions file | Ch. 9 | Part III |
| evaluator–optimizer | Ch. 10 | Part IV |
| routing | Ch. 10 | Part IV |
| approval gate | Ch. 12 | Part IV |
| autonomy dial | Ch. 12 | Part IV |
| human in the loop / on the loop | Ch. 12 | Part IV |
| review theater | Ch. 12 | Part IV |
| goal function | Ch. 13 | Part IV |
| kill switch | Ch. 13 | Part IV |
| outer loop | Ch. 13 | Part IV |
| eval set | Ch. 15 | Part V |
| span | Ch. 15 | Part V |
| trace | Ch. 15 | Part V |
| eval | Ch. 16 | Part V |
| LLM-as-a-judge | Ch. 16 | Part V |
| pass@k and passk | Ch. 16 | Part V |
| regression gate | Ch. 16 | Part V |
| reliability envelope | Ch. 16 | Part V |
| blast radius | Ch. 17 | Part V |
| confused deputy | Ch. 17 | Part V |
| guardrail | Ch. 17 | Part V |
| lethal trifecta | Ch. 17 | Part V |
| prompt injection | Ch. 17 | Part V |
| sandbox | Ch. 17 | Part V |
| checkpoint | Ch. 18 | Part V |
| durable execution | Ch. 18 | Part V |
| idempotency | Ch. 18 | Part V |
| prompt caching | Ch. 19 | Part VI |
| canary | Ch. 20 | Part VI |
| feature flag | Ch. 20 | Part VI |
| shadow deploy | Ch. 20 | Part VI |
| agent-legible code | Ch. 21 | Part VII |
| augmented coding | Ch. 21 | Part VII |
| vibe coding | Ch. 21 | Part VII |
| compound step | Ch. 22 | Part VII |
| verification gap | Ch. 23 | Part VII |
| appropriate reliance | Ch. 24 | Part VII |
| supervision surface | Ch. 24 | Part VII |
| classifier | Ch. 26 | Part VII |
| ordinal scale | Ch. 26 | Part VII |
| regressor | Ch. 26 | Part VII |
| three-set discipline | Ch. 26 | Part VII |
The glossary also has 11 synonyms with no page of their own. Each points to a term in the table:
- agent protocol: see protocol
- autonomy slider: see autonomy dial
- compound engineering: see compound step
- constitution: see standing project-instructions file
- dark launch: see shadow deploy
- function calling: see tool call
- generator–critic: see evaluator–optimizer
- inner loop: see outer loop
- trajectory: see trace
- transcript: see trace
- trust calibration: see appropriate reliance
One entry shows why the rule reads the closing pointer. In-context learning mentions Chapter 26 in its text and closes with a pointer to Chapter 2, so it sits in Part I. The glossary itself says which words are the field’s and which are the book’s own devices: the compass and the desk are flagged as the book’s.
Can you answer these from memory?
The self-test in this AI agents study guide has 35 questions, five per part, and every answer can be checked against a free glossary entry or a free chapter. Answer aloud or on paper before you look at anything. Tick a question only when you answered it without help; what stays unticked is your revision list.
The answers are in a separate section further down, on purpose. One list of 30 agent interview questions prints each answer directly under its question (Sankrityayan, Analytics Vidhya, 2026), which makes it a reading exercise. The research below says the effort of retrieving is what does the work. That effort is also the difference between an AI agents study guide and an AI agents glossary: the glossary tells you the answer, and the guide makes you produce it.
Part I, Foundations
- I-1. What one question separates a chatbot, a workflow and an agent?
- I-2. What quick test tells you a workflow will do the job?
- I-3. What is a hallucination, and what makes it dangerous in operation?
- I-4. If each step succeeds 95% of the time (an illustrative figure), how often does a run of 20 chained steps succeed?
- I-5. Is an instruction in the system prompt a guarantee?
Part II, Building Agents
- II-1. Why is a hard budget on an agent loop mandatory?
- II-2. Who runs a tool: the model or your code?
- II-3. What are the three families of stop condition, and which is the least trustworthy?
- II-4. What is an oracle, and why is the model that did the work a poor one?
- II-5. How does a skill differ from a tool?
Part III, Context Engineering
- III-1. What is context rot, and why does a bigger window not fix it?
- III-2. Name the four context operations.
- III-3. How does context engineering differ from prompt engineering?
- III-4. What does fine-tuning change, and which job does the book give it as opposed to retrieval?
- III-5. Which three kinds of long-term memory does the book name, and which piece of office furniture stands for each?
Part IV, Patterns and Choices
- IV-1. By what do you decide where an approval gate goes?
- IV-2. What is the difference between a human in the loop and a human on the loop?
- IV-3. What is review theater?
- IV-4. What two things make up a goal function?
- IV-5. What is a kill switch, and when do you test it?
Part V, Making Agents Reliable
- V-1. What is a trace, and what is a span?
- V-2. What is the difference between pass@k and passk, and which one fits an agent that acts unattended?
- V-3. What must happen before a model used as a judge is trusted?
- V-4. What are the three legs of the lethal trifecta?
- V-5. What does idempotent mean, and why does it matter for retries?
Part VI, Shipping and Operating
- VI-1. What is the layout rule that follows from prompt caching?
- VI-2. What is a shadow deploy?
- VI-3. What two disciplines make a canary real?
- VI-4. What does a feature flag turn promotion and rollback into?
- VI-5. What does a regression gate test, and why does a rollout need more rungs after it?
Part VII, Applications and the Road Ahead
- VII-1. What is the checkable criterion for vibe coding?
- VII-2. What is the compound step?
- VII-3. What is the verification gap, and why do research agents fail differently from coding agents?
- VII-4. What do over-trust and under-trust each produce?
- VII-5. Which three sets does the three-set discipline use for a classifier built on a language model?
Answer key
Each answer names its source so you can check the checker. Quoted words are verbatim from the glossary entry named, unless the line says Chapter 1. “Home” is the chapter the entry points to; for Parts II to VII that chapter is in the full book, and the glossary entry itself is free.
Part I. Both home chapters are free.
- I-1. Who owns the control flow: you (a chatbot), your code (a workflow) or the model (an agent). Chapter 1: “That question—who owns the control flow—is the entire distinction.”
- I-2. Chapter 1: “can you draw the flowchart before the request arrives?” If yes, a workflow or something simpler will do. The glossary entry for workflow says the same thing: “You could draw its flowchart before switching it on.”
- I-3. Glossary, hallucination: “The field’s term for output that is fluent, specific, confident, and wrong.” The danger, in the same entry, is that “fabrication arrives in exactly the tone of knowledge.” Home: Chapter 2.
- I-4. About 36%, because 0.9520 ≈ 0.36. Source: Chapter 1, which marks the 95% as an illustration. Try other numbers in the compounding error calculator.
- I-5. No. Glossary, system prompt: it is “a request rather than an enforcement mechanism.” Enforcement is the job of a guardrail, a Part V term.
Part II. Home chapters 3 to 6, in the full book.
- II-1. Glossary, budget: “a loop whose only exit is the model’s judgment has no guaranteed exit at all.” A budget caps passes, tokens, wall-clock time or money. Home: Chapter 3.
- II-2. Your code. Glossary, tool: “The model requests; your code executes.” Home: Chapter 3.
- II-3. The model’s own judgment that the goal is met, the budget, and error exits. The entry for stop condition calls the first “testimony, and the least trustworthy exit.” Home: Chapter 3.
- II-4. Glossary, oracle: “a source of truth outside the thing being judged, which pronounces an answer right or wrong.” The entry adds that “the model that did the work is a poor oracle for that work.” Home: Chapter 4.
- II-5. A tool is a verb. A skill is “packaged expertise: the judgment about the verbs,” a bundle of instructions, reference files and scripts loaded on demand. Glossary, skill. Home: Chapter 6.
Part III. Home chapters 7 to 9, in the full book.
- III-1. Glossary, context rot: “The tendency of a model’s ability to use any given fact to decay as the window around it fills.” The entry continues: “The reason a bigger window does not repair a crowded one, and the standing argument for curation.” Home: Chapter 7.
- III-2. Write, select, compress, isolate. Glossary, write / select / compress / isolate. Home: Chapter 7.
- III-3. Glossary, context engineering: “Prompt engineering writes a document; context engineering operates a system.” Home: Chapter 7.
- III-4. It changes the weights. Glossary, fine-tuning: “fine-tuning for how the model should respond, retrieval for what it should know.” Home: Chapter 8.
- III-5. Episodic memory is the logbook, semantic memory is the filing cabinet, procedural memory is the procedures manual. Glossary, memory. Home: Chapter 9.
Part IV. Home chapters 12 and 13, in the full book.
- IV-1. By consequence. Glossary, approval gate: “place gates by consequence (what a wrong action costs), never by convenience.” Home: Chapter 12.
- IV-2. In the loop means inside the run, approving actions as they happen. On the loop means above it, reviewing what an autonomous run proposes and produces. Glossary, human in the loop / on the loop. Home: Chapter 12.
- IV-3. Glossary, review theater: “Diligence performed at a volume where it can no longer be real.” Home: Chapter 12.
- IV-4. A persistent goal held outside the model, and a machine-checkable stopping condition evaluated after each pass. Glossary, goal function. Home: Chapter 13.
- IV-5. Glossary, kill switch: “One obvious, fast, tested way to stop every running loop and agent at once.” Test it before the first unattended run. Home: Chapter 13.
Part V. Home chapters 15 to 18, in the full book.
- V-1. Glossary, trace: “The complete, structured record of one run.” A span is one unit of work inside a trace: “a single operation with a start time, an end time, and structured attributes.” Glossary, span. Home: Chapter 15.
- V-2. pass@k is the probability that at least one of k attempts succeeds; passk is the probability that all k succeed. The entry calls passk “the right number for an agent that must be correct every time it acts unattended.” Glossary, pass@k and passk. Home: Chapter 16. The pass@k calculator shows how far apart the two can sit.
- V-3. Glossary, LLM-as-a-judge: “A judge is an instrument that must itself be calibrated against human judgment before it is trusted.” Home: Chapter 16.
- V-4. Glossary, lethal trifecta: “access to private data, exposure to untrusted content, and a channel for external communication.” Home: Chapter 17.
- V-5. Glossary, idempotency: “doing an operation twice has the same effect as doing it once, which makes retrying safe by construction.” Writes are the trap, and the standard fix is an idempotency key. Home: Chapter 18.
Part VI. Home chapters 19 and 20, in the full book.
- VI-1. Glossary, prompt caching: “stable content first, variable content last.” Home: Chapter 19.
- VI-2. Running a candidate version against real traffic while users see only the incumbent’s output, so it is measured “with zero user exposure.” Glossary, shadow deploy. Home: Chapter 20.
- VI-3. Glossary, canary: “consistent assignment (a user stays on one version for the session) and automated rollback with thresholds chosen in daylight.” Home: Chapter 20.
- VI-4. Configuration flips. Glossary, feature flag: “instant, reversible, no redeploy.” Home: Chapter 20.
- VI-5. Glossary, regression gate: “The gate tests the failures you already imagined; the rest of the rollout ladder exists because reality imagines better.” The term is filed under Part V (home: Chapter 16) and its entry also points to Chapter 20.
Part VII. Home chapters 21 to 26, in the full book.
- VII-1. Glossary, vibe coding: “The mode in which you do not read the code before running it.” Home: Chapter 21.
- VII-2. The ritual that closes a task. Glossary, compound step: “before moving on, ask what the agent should have known at the start,” and write it into the standing project-instructions file, one line per lesson. Home: Chapter 22.
- VII-3. Glossary, verification gap: “The distance between how convincing an output looks and how cheaply it can be checked.” For open-ended factual claims “there is no compiler for facts.” Home: Chapter 23.
- VII-4. Over-trust produces misuse, “accepting what should be checked.” Under-trust produces disuse, “re-doing what the machine got right.” Glossary, appropriate reliance. Home: Chapter 24.
- VII-5. The demonstrations packaged into the skill, the development set you iterate against, and a held-out test set consulted rarely. Glossary, three-set discipline. Home: Chapter 26.
In what order should you review in one evening?
Work in ten steps: the thesis first, then the seven term groups in book order, then the self-test, then only the entries behind your misses. The whole pass takes roughly 75 to 100 minutes. That figure is approximate, and the table shows how I arrived at it.
Reading minutes use the web reader’s own formula, words divided by 220, applied to my word count of each group’s entries. Self-test minutes are an assumption of about one minute per question, and the last step is an assumption too. I have not timed anyone.
| Step | What you do | Approx. minutes | Basis |
|---|---|---|---|
| 1 | Read the thesis paragraph of the Preface and the compass passage in Chapter 1 | 5–15 | Chapter 1 whole is about 16 minutes by the reader’s label |
| 2 | Part I terms (17): hide the definition, say it, then check | 4 | 796 words ÷ 220 |
| 3 | Part II terms (13) | 3 | 704 words ÷ 220 |
| 4 | Part III terms (16) | 4 | 849 words ÷ 220 |
| 5 | Part IV terms (9) | 2 | 452 words ÷ 220 |
| 6 | Part V terms (17) | 4 | 972 words ÷ 220 |
| 7 | Part VI terms (4) | 1 | 182 words ÷ 220 |
| 8 | Part VII terms (11) | 2–3 | 546 words ÷ 220 |
| 9 | Self-test, 35 questions, from memory | about 35 | Assumed: one minute per question |
| 10 | Reread only the entries behind your misses | 10–15 | Assumed |
| Total | roughly 75–100 | Counted steps sum to about 70–86; pausing to recall each term adds the rest |
If you have less time, do steps 1 and 9 and let the misses choose your reading. If you have a week, repeat step 9 on the misses a day or two later. Do not spend the evening improving the plan; the plan is ten rows.
Why does testing yourself beat rereading?
Answering from memory produces better delayed recall than rereading, in controlled experiments, and rereading feels better than it works. In Roediger and Karpicke’s 2006 study, students who took a recall test on a prose passage remembered 56% of it a week later; students who restudied it remembered 42% (Psychological Science, 2006).
The same paper carries the caveat. Five minutes after studying, the restudy group was ahead, 81% against 75%. Rereading wins when the test is minutes away and loses by two days (54% against 68%). The abstract adds that repeated studying “increased students’ confidence in their ability to remember the material,” which is the trap: the method that feels safer the night before is the weaker one by the morning after next.
Two later results sharpen the advice. In Karpicke and Roediger’s 2008 vocabulary experiment, students’ predictions of their performance “were uncorrelated with actual performance” (Science, 2008). Rowland’s 2014 meta-analysis found “initial recall tests yielding larger testing benefits than recognition tests” (Psychological Bulletin, 2014). So produce the definition yourself; do not just nod at it.
Honest limits apply. These studies used prose passages and vocabulary pairs. Nobody has run them on engineering concepts, and nobody has tested this guide on students.
A 2016 replication found that part of the 2008 gap was spacing: with spacing controlled, “both repeated testing and restudying improved learning” (Soderstrom, Kerr and Bjork, 2016). Spacing is cheap to add: for a test a week away, one study of fact learning put the best gap before a review at about 20 to 40% of that week (Cepeda et al., 2008).
The connection to the book’s thesis is my own framing; the book does not make it. Feeling fluent after a reread is testimony. An answer you retrieved and then checked against the source is evidence. An AI agents study guide earns its name only if it produces the second kind.
What do interviewers say they listen for?
Four people who run interviews for AI engineering roles have written about it. None of the four describes asking for definitions alone: one asks concept questions and wants the trade-offs with them, one stresses evals, and two describe watching whether a candidate checks the model’s output. That is four accounts, dated 2024 to 2026. They do not tell you how common each kind of interview is, what a given company asks, or what an examiner will set.
Jonathan Pedoeem’s 2025 account of an agent system-design interview lists concept questions on fine-tuning, embeddings, vector databases and long-term memory, which are all Part III terms here. It also warns: “Candidates should be ready to discuss practical applications and trade-offs, not just memorized definitions” (PromptLayer blog, 2025). His company sells prompt and evaluation tooling, which is worth knowing when you weigh the emphasis.
The other three write about checking, of the system or of the model’s output. Eugene Yan wrote in 2024 that “having a basic understanding of evals is key for anyone building ML-powered products” (eugeneyan.com). James Brady of Elicit reported in 2026 that “Candidates often accept model output uncritically” (Elicit blog). Sebastian Duerr and Hagay Lupesko of Cerebras wrote in 2026 that what matters is whether a candidate can “use AI thoughtfully, verify its output, and stay accountable for the result” (Cerebras blog).
This AI agents study guide is built for both kinds of interview: the one that asks concept questions and the one that watches whether you check. The term table serves the first. The one question serves the second, because “what signal tells you it worked?” is what “verify its output” means once you have to name the signal. Brady’s observation, incidentally, already has a name in the Part IV group: review theater.
Where are the syllabus and the lecture notes?
The syllabus and the lectures are in the site’s free teaching kit, which carries a 13-week course plan built on the book. Each week of the plan lists its key concepts, and most of those concepts are glossary terms, so the term table above doubles as an index to the course.
Only three of the thirteen weeks have a lecture at the time of writing, and I would rather you knew that before clicking: lecture 1, on what an agent is, lecture 2, on the engine’s failure modes and the loop, and lecture 3, on planning, tools and the action space. Each has a slide deck, instructor notes and exercises. If you are looking for AI agents lecture notes for the later weeks, the syllabus tells you which chapters and terms each week covers. If your course uses the book as its AI agents textbook, treat this guide as the revision sheet that sits beside the syllabus.
For a first pass before any of this, the short chatbot, workflow and agent explainer covers question I-1. If you are choosing how to study in the first place, see the page on the best way to learn AI agents; if you are starting cold, the one on learning AI agents from scratch comes before this guide.
What can’t an evening with this guide do?
An evening with this AI agents study guide cannot teach you a part of the field you have never read about, cannot stand in for building something, and cannot predict a specific exam or interview. It tells you which definitions you hold and which you don’t, and that is all it claims.
Three limits deserve plain statement. First, a definition is the compressed end of an argument. The glossary says so itself: the pointer to a chapter “is part of the definition, because most of these terms earn their keep only inside an argument.” For Parts II to VII that argument is in the full book.
Second, the self-test checks recall of definitions. The final exam in the site’s own syllabus also covers arithmetic and design decisions with justification, and five questions per part cannot cover that. Third, the term-to-part mapping is one rule applied by a script. Twenty of the 87 entries point to more than one chapter (appendices not counted), so some terms could fairly sit in two groups.
Anything that names a product is deliberately absent. Product knowledge ages within months, and the vocabulary here is the part that carries over from one tool to the next.
The one thing to keep
Every term in this guide is a piece of an answer to one question. When a term slips your mind in the room, ask what signal would tell you the design worked, and work back to the term. A budget, an oracle, a trace, an eval set and a regression gate are all answers to it.
Start with what is free: Chapter 1, “What Is an Agent?” for the compass, and Appendix B, the glossary for all 87 definitions. The chapters behind Parts II to VII are in the full book; when your misses tell you which part you need, see the formats.
Questions readers ask
- Is a glossary enough to pass an exam or an interview on AI agents?
- No. A glossary gives you recognition: the term looks familiar. Exams and interviews ask for recall and use. Cover each definition, say it from memory, check it, then answer questions that make you apply the term. The four interviewer accounts cited in this guide, dated 2024 to 2026, describe listening for trade-offs, evals or verification of the model's output, and one says outright that memorized definitions are not enough.
- How long does it take to review AI agents with this study guide?
- Roughly 75 to 100 minutes for the vocabulary and the self-test. Reading time uses the web reader's formula (words divided by 220), which puts the glossary at about 21 minutes; the self-test assumes about one minute per question for 35 questions. Reading all 27 chapters is a different job: about 13 hours at the same rate.
- Do I need the full book to use this guide?
- No. All 87 glossary definitions are free online, and so are the Preface and Part I (Chapters 1 and 2). Every answer in the self-test can be checked against a free glossary entry or a free chapter. The home chapters for Parts II to VII are in the full book, so the guide cannot be your first exposure to those ideas.
- How many terms are in the AI agents glossary?
- The glossary in Appendix B of AI Agents, Engineered has 98 entries. Of those, 87 carry a definition and 11 are synonyms that say 'See' another entry, such as function calling (see tool call) and trajectory (see trace). The site builds one page for each of the 87 defined terms.
- Does the guide or the book include code I can run?
- No. The book is written in pseudocode throughout, by design, and this guide contains no code at all. The Preface promises 'examples in plain pseudocode', and the minimal agent in Appendix A is annotated pseudocode meant to be translated into your own language.
Sources
- Henry L. Roediger III and Jeffrey D. Karpicke, Psychological Science 17(3) (2006). Test-enhanced learning: taking memory tests improves long-term retention
- Jeffrey D. Karpicke and Henry L. Roediger III, Science 319(5865) (2008). The critical importance of retrieval for learning
- Nicholas C. Soderstrom, Tyson K. Kerr and Robert A. Bjork, Psychological Science 27(2) (2016). The Critical Importance of Retrieval—and Spacing—for Learning
- Christopher A. Rowland, Psychological Bulletin 140(6) (2014). The effect of testing versus restudy on retention: a meta-analytic review of the testing effect
- Nicholas J. Cepeda, Edward Vul, Doug Rohrer, John T. Wixted and Harold Pashler, Psychological Science 19(11) (2008). Spacing effects in learning: a temporal ridgeline of optimal retention
- Jonathan Pedoeem (PromptLayer) (2025). The Agentic System Design Interview: How to evaluate AI Engineers
- Eugene Yan (2024). How to Interview and Hire ML/AI Engineers
- James Brady (Elicit) (2026). Engineering interviews in the era of agents
- Sebastian Duerr and Hagay Lupesko (Cerebras) (2026). Hiring Engineers for an AI-Native World
- Vasu Deo Sankrityayan (Analytics Vidhya) (2026). 30 Agentic AI Interview Questions: From Beginner to Advanced
- keyle (2026). Hacker News comment on wanting a flash-card glossary