Engineering AI Agents · Week 1
What an agent is, and the engine underneath
Chatbot, workflow, agent · the compass · tokens, the desk and the dice
Slides CC BY 4.0 · from the teaching kit of AI Agents, Engineered by Enrique Gutiérrez · aiagentsengineered.com/teach/
Welcome to the course. This first week answers two questions: what exactly an agent is, and what kind of machine sits at its center.
Everything today comes from the free part of the book, the preface and Chapters 1 and 2, so nobody needs to buy anything to keep up.
One promise about the whole course: we will name concepts and patterns, not products. Any tool or vendor I mention is an example of a category.
The moment this course is about
“I didn’t see every step. How much of this can I trust?”Preface
Thesis: an agent is only as trustworthy as the signal you can use to verify it.
The preface opens with a moment most of you have had: you hand a model a half-hour task and watch it search, try, fail, correct itself and finish.
The first feeling is delight; the second, which lasts longer, is this question about trust.
The book’s answer is its thesis, and it is also this course’s grading philosophy. A model sounds equally confident when it is right and when it is wrong, so trust has to come from a signal outside the model: a test, a schema, a source you can open, a human approval.
You will be graded on the signals you ship, not on demos.
By the end of today you can
Four objectives
Classify a system as chatbot, workflow or agent, by who owns the control flow
State the four compass bearings and a design fork each one decides
Explain tokens, the context window as a finite desk, and why two identical runs differ
Compute pⁿ : whole-run success from per-step reliability
These are the four things you should be able to do on paper by the end of the session, and the reading quiz checks them.
Notice that each one is a verb you can test: classify, state, explain, compute.
The first half of the lecture covers the first two, from Chapter 1. The second half opens the hood, from the first three sections of Chapter 2.
Part 1 · Chapter 1 (free)
What is an agent?
We start with vocabulary, because “agent” is one of the most overloaded words in software.
It gets applied both to a scripted phone-tree bot and to a system trusted to work alone for an afternoon. If speaker and listener mean different things, the conversation goes nowhere, so we pin the term down first.
The working definition
An agent runs tools in a loop to achieve a goal .Simon Willison’s one-sentence definition, adopted in Chapter 1
Tools connect the model to the world
Loop lets it react to what it finds
Goal ends the loop: a stopping condition
Strip away the demos and the idea is small: a language model placed in a loop, given a goal and a set of tools, which are functions it can ask your program to run.
The book adopts this one-sentence definition because every word carries weight.
The goal matters more than people think: an agent has a stopping condition, and that is what separates it from a runaway process.
Ask the room: what would a loop with tools but no goal look like? A process that never knows when to stop.
What is genuinely new
“In an agent, the sequence is decided by the model, at runtime , in response to what it observes.”
In every program you have written so far, you decided the sequence, in advance, in code.
Chapter 1, “The Agent Idea”
This is the one idea to carry out of the room. In ordinary software, the order of operations is fixed by you before the program runs.
In an agent, the model picks the next step after seeing the result of the last one.
The book puts it plainly: that single change is what makes agents worth studying, and also what makes them expensive, slow and occasionally alarming. The rest of the course is about both halves.
Why it became practical: three ingredients
Fluency was never the missing piece
1 · Post-training
Models take instructions An agent is a delegation; it needs a worker that accepts the job
2 · Function calling
A bridge to action “The model proposes; your code disposes.”
3 · Context window
Working memory grew Instructions, history and the latest result fit together
Chapter 1 · deliberately without dates: ask which ingredient just got better
Language models were fluent long before anyone trusted one with a loop. Three things had to mature.
First, post-training turned a pure text continuer into something that follows instructions and stops when done. Second, function calling gave the model a disciplined, machine-readable way to ask your code to do things. Third, the context window grew large enough to hold a working session.
The book explains this without dates on purpose, so the explanation does not expire. When a new capability is announced, ask which of the three ingredients just got better.
Decide before the next slide
Which one is the agent?
A
Q&A assistant: paste a function, get an explanation; then it waits
B
Ticket pipeline: classify → fetch policy → draft reply → check against rubric → human queue. Same four steps, every time
C
“The checkout tests are failing. Fix them.” It searches, reads, edits, reruns, and stops when green
All three may run on exactly the same model.
These are the three systems from Chapter 1. Give the room thirty seconds and take a show of hands for each.
Almost everyone picks C, and the instinct is right. The useful work is saying precisely why.
Point out the last line: the model can be identical in all three. So whatever distinguishes them, it is not the model.
The whole distinction
Who decides what happens next?
Chatbot : you do, one message at a time
Workflow : your code does, on paths drawn in advance
Agent : the model does, as it goes
Figure from Chapter 1
Here is the answer: the three systems differ in who owns the control flow.
In the assistant, the human drives every turn and nothing is executed; the worst it can produce is a wrong answer. In the ticket pipeline, code decides the route and the model only fills in content. In the test fixer, the model chooses which file to read, whether to edit or rerun, and when to stop.
Workflow shapes have names we will meet later: chaining, routing and parallelization.
The animated explainer at aiagentsengineered.com/learn/explain/what-is-an-agent/ walks through this same contrast in about three and a half minutes; it works well before the classification exercise.
The litmus test for the rest of the course
“If you can confidently draw the control-flow diagram before the request arrives, you are looking at a workflow. If the diagram can only be drawn in hindsight, once the model has reacted to what it found, you are looking at an agent.”Chapter 1, “Agents, Workflows, and Chatbots”
Use it on any system, whatever its label says. Relabeling a pipeline as an “agent” is agent-washing .
This test is the most useful sentence in the chapter, and you will use it all semester.
Try it on the ticket pipeline: you could draw its flowchart before the first ticket arrives, so it is a workflow. The test fixer’s path only exists after the fact, so it is an agent.
It also protects you from agent-washing, where scripted pipelines get called agents because the word sells. The question that cuts through it: does the model decide the next step, or does code?
Same atom, two arrangements
The augmented LLM
model call + retrieval + tools + memory
Workflow : several atoms, wired together by your code along a fixed path
Agent : one atom inside a loop, deciding what to do on each pass
Chatbot : the bare model call, augmentations mostly unused
Underneath workflows and agents sits one building block, which the engineering essay the book cites calls the augmented LLM.
It is a model call enhanced with retrieval, which pulls relevant information in at request time; tools, which are functions it can request; and memory, some way of carrying forward what matters.
Keep this picture and the taxonomies stop feeling mysterious: it is the same atom, arranged by code or placed in a loop.
Not three boxes: a dial
How much of the control flow does the model own?
Figure from Chapter 1
Much of production: an agent that pauses for human approval before anything consequential
The far end buys capability and pays in blast radius
The common shape: “a workflow shell with one or two genuinely agentic steps inside it”
Chatbot, workflow and agent are useful labels, but the truth is a dial, and a single quantity moves along it: how much of the control flow the model owns.
Walk the stops from left to right: text only, fixed pipeline, a router that picks one branch, an agent that pauses for approval before sending the email or spending the money, and finally a model that runs many steps unconfirmed.
Blast radius means the damage a wrong action could do before anything stops it. Mature systems choose a position on the dial deliberately instead of inheriting one from a demo.
The compass the course steers by
Four bearings
Verifiability
Scarcity of context
Compounding error
The simplest thing that works
Figure from Chapter 1
Chapter 1 hands you a compass with four bearings that come back in nearly every later chapter.
Verifiability: trust comes from a signal outside the model. Scarcity of context: the window is finite and attention degrades before it fills. Compounding error: reliability multiplies across chained steps. The simplest thing that works: add complexity only when it demonstrably pays.
The needle rests on the first one deliberately. The book states the bearings now and defends them over the rest of the book; we will do the same over the semester.
A compass decides forks
One design fork per bearing
Bearing The fork it decides The question to ask
Verifiability More autonomy or less What signal tells you it worked?
Scarcity of context Richer context or leaner Does this deserve the desk?
Compounding error A long free run or checkpoints What is pⁿ for this chain?
Simplest thing An agent or a plain script Which rung solves it?
Three forks quoted from Chapter 1: “more autonomy or less, another agent or a plain script, richer context or leaner”; the checkpoint fork follows from its compounding-error bearing
This slide is objective two. The book says that whenever it reaches a fork, you will watch the same four bearings decide it.
Where a strong verifying signal exists, as with code and its test suite, you can grant autonomy generously; that is why coding became the first domain where agents earned their keep.
The compounding row pairs with the arithmetic on the next slide, and the last row with the escalation ladder after it.
On the exam you will be given a fork and asked which bearing decides it and why.
Worked example · compounding error
Reliability multiplies
0.9520 ≈ 36%
p per stepn whole run
99% 10 90%
95% 10 60%
90% 10 35%
Figure from Chapter 2 · numbers illustrative
Here is the arithmetic behind the third bearing. If each step succeeds with probability p, independently, a run of n steps succeeds about p to the n.
The book’s illustration: 95 percent per step sounds excellent, yet twenty chained steps succeed only about 36 percent of the time. “The particular numbers vary; the arithmetic does not, and it is merciless.”
Have students compute the table rows by hand before you reveal them; it takes a minute and it makes the point stick.
The defense of this formula, and its three levers, is next week’s material.
The fourth bearing made practical
Do you even need an agent?
More calls, each re-sending a growing history: money
Sequential calls: latency
Different path each run: nondeterminism
A wrong action , then ten more built on it
Wider attack surface; a whole harness to build
Figure from Chapter 1 · stay on the lowest rung that works
The mistake the field warns newcomers about most often is reaching for an agent when a prompt, a pipeline or fifty lines of code would do.
Compared with one model call, an agent adds a package of costs: money, latency, nondeterminism, and a failure surface that changes in kind, from a wrong answer to a wrong action. It also widens the attack surface and requires a harness, the surrounding machinery of tools, stop conditions, approval gates, tracing and evaluation.
So the ladder: plain code, a single call, a workflow, an agent. Each rung up buys adaptability and pays for it. Staying low is a discipline, not a failure of ambition.
Calibrate on the book’s pair
Workflow
“Route each incoming email to billing, technical, or sales and draft a first reply” Knowable steps, high volume
Agent
“Make this failing test pass in a codebase you have never seen” Unknowable steps, built-in verifier
“An agent, then, is a cost you pay for adaptability you can name. If you cannot name the adaptability, keep your money.”Chapter 1, “Do You Even Need an Agent?”
Chapter 1 gives a calibration pair. The email router has steps you can draw in advance and runs at high volume, so it wants a workflow.
The failing test has steps that depend on what the model discovers, and it comes with a verifier: run the tests. That combination is a genuine agent problem.
Even when an agent is justified, insist on three things: a clear success criterion, a feedback signal at each step, and a sensible place for a human to look.
Beware the argument “it might need to adapt someday.” Build the simpler version first and climb only when you can point at inputs it provably fails on.
In class · pairs · 15 minutes
Should this be an agent?
Each pair writes down three requests from their own work or studies
Run each through the browser tool: aiagentsengineered.com/tools/should-this-be-an-agent/
For each verdict, name who owns the control flow and the verifying signal
Copy the result as Markdown and submit all three
Runs in the browser · no account, no API key
Now you use the litmus test and the ladder on your own problems. Pairs pick three real requests, ideally one they suspect is an agent and one they suspect is not.
The tool asks questions drawn from the book’s signs and conditions, places the task on a rung and explains why.
The submission is the tool’s Markdown export plus one line each: who owns the control flow, and what signal would tell you it worked.
When time is up, ask two pairs whose verdict surprised them to explain why; disagreements with the tool are the most instructive part.
Part 2 · Chapter 2, sections 1–3 (free)
The engine underneath
We are going to put a language model in a loop, hand it tools, and let it make decisions while we are elsewhere. When it behaves strangely, “it’s magic” will not narrow down the bug.
So the second half builds the minimum mental model of the engine: what it is, its working memory, and how it writes. No mathematics beyond arithmetic, and nobody trains anything.
LLMs as next-token predictors
“The capital of France is ___”
In: text so far. Out: a probability for every possible next token
Pick one, append it, run again: autoregressive generation
“the model never drafts its answer in advance”
Figure from Chapter 2
Ask the room to finish the sentence. “Paris” arrives without anyone consulting a mental atlas. That reflex is the best intuition for what the model does; unlike you, it has no other mode.
Mechanically, the model is a function from text to a ranked list of next-token probabilities. The surrounding system picks one, appends it, and runs the model again until a stop condition fires.
The property to keep: there is no plan or outline in reserve. Its whole method is a local guess, repeated.
Keep these two apart
Training vs inference
Training Once, by the model’s makers Sets the weights Ends at a training cutoff
Inference Every use: every chat, every agent step Weights frozen “the model does not learn from your conversations”
Parametric knowledge (in the weights) vs context knowledge (on the desk now): the model is more reliable with the second.
Training is the one-time, expensive process that sets the weights. Inference is everything you do afterward, with the weights frozen.
Two consequences: nothing you type changes a weight, so what feels like memory in a session is just the transcript being re-read on every call. And the model’s knowledge stops at a training cutoff, and it will not reliably tell you so.
Post-training adds manners, not facts, and it does not stop fabrication. In the book’s words, “Plausible and true are correlated,” but they part company exactly where the model lacks the fact. That is why so much engineering moves facts into the context.
Tokens and the context window
The model’s world has a grain
annoying + ly · un + happi + ness
Subword tokenization : common words whole, rare ones in pieces
A fixed lookup table: “the network never sees your characters”
So counting letters or spelling backwards is a memory task
Code, JSON and many non-English languages cost more tokens
Chapter 2 · example splits as in the book; actual splits differ per tokenizer
A token is a chunk of text from a fixed vocabulary; on average, about three-quarters of an English word. Whole-word vocabularies explode and break on typos; single characters make texts enormously long. Subword tokenization is the compromise.
The tokenizer is a frozen lookup table that turns text into integer IDs. Whatever happens inside a chunk was erased before the model saw anything.
That is why models have stumbled at counting the letters in a word: the letters inside a chunk are exactly the information the table threw away. Your homework tests this.
When a token count matters, measure your actual text with the provider’s own tokenizer.
The context window is a desk
“There is no drawer, no shelf, no second desk.”
One call, tokens on the desk
instructions 500
conversation 8,000
2 documents 40,000
tool result 3,000
Over 50,000 tokens read, and paid for, before the first word
Stateless: the whole transcript goes back on the desk every call
Too big: an error, or silent truncation
Mid-pile facts get lost in the middle
Chapter 2 · token counts are the book’s made-up but realistic sketch
The book’s picture for the context window is a desk: everything the model consults during one call must fit on it at once, including the answer it is writing.
The API is stateless, so the software re-lays the whole transcript every turn. With the book’s illustrative numbers, the model reads over fifty thousand tokens before writing a word, and you pay for each one.
When the pile no longer fits, the kind outcome is an error; the treacherous one is silent truncation, where your original instructions quietly fall off. “When a long session starts forgetting, check for overflow before you blame the model.”
And even what fits is used unevenly: material in the middle is recalled worse than the edges. The window is a ceiling, not a target.
How text is generated
“By default, the machinery rolls dice.”
The cat sat on the …
mat 0.42
floor 0.18
sofa 0.13
rug 0.09
Probabilities invented, shape typical (Chapter 2)
Greedy : always the top token; deterministic, flat, loops
Sampling : a weighted lottery, one draw per token
Temperature : low buys consistency, high buys variety; correctness at neither end
Top-p : the pool sizes itself to the model’s confidence
Ask the room to commit: which word does the model write? Most say “mat.” For most systems you will call, that is wrong, because the default is sampling.
Greedy decoding always takes the top token. It sounds responsible but produces flat, repetitive text that can loop. Sampling draws from the list weighted by probability, so “mat” wins 42 times in a hundred and “rug” about nine.
Temperature and top-p only reshape or trim the list before the draw. Use low temperature when a program will parse the output or a test will check it, higher for brainstorming.
Why two identical runs differ
80
distinct answers from 1,000 runs of one prompt at temperature zero ; identical for the first 102 tokens
He et al., Thinking Machines Lab (2025), as reported in Chapter 2 · illustrates the effect, not a constant
Source 1: sampling, any temperature above zero
Source 2: hosted endpoints batch many users; numerics shift with batch size (batch invariance missing)
Near-ties at the top flip; the text forks for good
Habit: test properties of the output, never exact strings
A prediction exercise first: one prompt, a thousand times, temperature zero. How many distinct answers? Your arithmetic says one. The experiment the book reports got eighty.
The fuller explanation is batch invariance, or its absence. A hosted endpoint batches requests from many customers, the batch size moves with traffic, and low-level routines can return minutely different numbers at different batch sizes. When two tokens are nearly tied, that wobble decides the winner.
The deciding input is how many strangers were using the same servers at that instant, which you cannot see or replay. So write tests that assert properties: it parses, it names the right customer, it passes the checker. Make the outcome checkable rather than the output identical.
Recap
Five things to keep
An agent runs tools in a loop to achieve a goal; the model owns the control flow
Draw the flowchart before the request arrives? Then it is a workflow
Compass: verifiability, scarce context, compounding error, simplest thing
The model predicts tokens on a finite desk it re-reads every call
It writes by lottery, so check outcomes, not strings
If students leave with only these five lines, the week did its job.
Ask one student to apply the litmus test to a system named by another student, as a quick check before they leave.
Every one of these lines returns next week, when we chain the engine’s failure modes into the compounding arithmetic and build the loop itself.
Before next week
Homework and reading
Tokenizer experiment
Ask any free chat interface or local model to count a letter in a word and to spell a word backwards
Explain the outcome with subword tokenization, in one page
Reading quiz on Chapters 1 and 2 (§1–3)
This week’s reading, all free: Preface · Ch. 1 · Ch. 2 §1–3
The homework uses any model you can reach for free; no paid API is needed. The point is the explanation, not whether the model passes, since such gaps get patched and newer models may succeed.
If the model you try gets it right, explain what it had to do differently: reconstruct the spelling from memory, or reach for a tool.
Next week’s reading finishes Chapter 2, which is free, and adds Chapter 3 and Appendix A from the full book. In lab 1 you will type in the minimal agent and run it against a scripted model client.
Engineering AI Agents · Week 1
What signal tells you it worked?
Slides from the teaching kit of AI Agents, Engineered by Enrique Gutiérrez, CC BY 4.0 · aiagentsengineered.com/teach/
Reuse, adapt and translate with attribution · creativecommons.org/licenses/by/4.0/
End on the question the whole course keeps asking. Every design choice from here on starts with it.
These slides are CC BY 4.0: instructors may reuse, adapt and translate them, keeping the attribution line on this slide.