How to run the session
This lecture turns the engine of week 1 into an agent. The first half finishes Chapter 2, which is free: prompting in one slide, structured output and function calling, the six failure modes, and the defense of compounding error with its three levers. The second half is Chapter 3 and Appendix A, from the full book: the loop, ReAct, the four parts of a minimal agent, the harness, and the three exits. The two halves meet in one idea, which the end slide puts as a question: the model’s “done” is testimony, so what is your evidence?
The deck has 28 slides, split by two section slides. Press N for presenter notes, F for fullscreen; slide numbers are deep links (deck.html#18). The PDF is the same deck, one slide per page, without notes. All the pseudocode on the slides is the book’s, except slide 24, which is the lab’s scripted client and is labeled as such.
Four slides are built to be answered before you advance. On slide 6, students sketch a parser for three phrasings of the same fact; give them a full minute. On slide 8, ask at which of the five steps anything happens in the world (step 3). On slide 20, ask for a guess at how many lines a working agent needs. On slide 22, ask what the model does on pass 2 when nothing in the program connects “greeting” to greetings.py. Take a show of hands each time.
Work the arithmetic on slides 13 and 14 on the board, not on the slide. Students who derive pⁿ themselves, and then the retry formula, keep it.
Timing (80 minutes)
| Segment | Slides | Minutes |
|---|---|---|
| Opening: the bridge from week 1, objectives | 1–3 | 4 |
| Prompting, structured output, constrained decoding | 4–7 | 9 |
| Function calling: the exchange, the tool definition, unhappy paths | 8–9 | 8 |
| Failure modes: hallucination, the jagged frontier, the six-row table | 10–12 | 9 |
| Compounding error derived, the three levers | 13–14 | 7 |
| In-class exercise: explainer, then the calculator in pairs, debrief | 15 | 15 |
| The loop: four beats, the index card, ReAct | 16–19 | 11 |
| The minimal agent: four parts, the harness, Appendix A’s program and run | 20–22 | 9 |
| Exits, the scripted client, frameworks | 23–25 | 6 |
| Recap, lab, reading | 26–28 | 2 |
For a 90-minute slot, play the three-minute agent loop explainer on slide 17 and give the extra time to slide 18, tracing the chapter’s seven-pass checkout run on the board against the pseudocode. For a 75-minute slot, skip slide 5 (students read prompting on their own) and take slide 25 as a one-line answer.
Common misconceptions and how to address them
“The model runs the tool.” It never does. It writes a structured request and stops; your code runs the function at step 3 of the exchange. Make students say it back: the model proposes; your code disposes. The practical consequence is that every authorization, confirmation and limit lives in code you control, and the security weeks build on exactly this.
“Structured output means the answer is correct.” Constrained decoding guarantees the shape, not the content. If the model does not know the order total, the schema makes it produce a well-typed number. Ask: what does your code check after the parse? Ranges, cross-field rules, whether the customer ID exists.
“Tool results go back to the user.” This is the classic first bug, and Chapter 2 names it. The result goes back to the model at step 4, because only the model knows why it asked. Students meet it in the lab if they wire the loop by hand.
“Asking the model ‘are you sure?’ is a check.” Sycophancy says otherwise: push skeptically and it caves, ask approvingly and it endorses. Run the demonstration from Chapter 2 live if a free model is at hand: a correct answer, then “I don’t think that’s right.” Then point forward to the verifier argument in week 3.
“pⁿ is pessimistic; real agents do better.” In one way it is pessimistic: it treats every error as fatal and uncaught, and checking with recovery pushes success above it (that is the second lever). In another way it is optimistic: it assumes independence, and an agent’s errors land in its own transcript and raise the error rate of later steps. Students should be able to state both.
“The agent has a plan and a memory system.” It has a history list and a loop. The appearance of method comes from one decision per pass, each conditioned on everything appended so far. Slide 22’s run, entry by entry, is the antidote; the memory is the entries.
“The step cap is for bad models; a good model knows when to stop.” The cap guards against harness bugs as much as model behavior: a dropped append makes even a perfect model repeat a call forever. The chapter’s rule is to set the cap before writing anything else.
“Real agents need a framework.” A framework adds plumbing: persistence, retries, tracing, memory management, orchestration. None of it changes what the model does on one pass. Build the bare loop once, which is what the lab is for.
Materials and notes for instructors
The diagrams are the book’s own (Chapters 2 and 3 and Appendix A), also on the diagrams page. All seven figures the syllabus lists for this week are on the slides, plus the Chapter 2 prompt-anatomy figure on slide 5. The in-class exercise uses two pages on this site: the Why errors compound explainer (about three minutes) and the compounding-error calculator. The agent loop explainer animates slide 17’s figure and works as a warm-up for the second half.
Numbers on the slides are labeled. The 36 percent and 8 percent examples are the book’s illustrations. The 19-point result is from one 2023 field experiment, reported in Chapter 2, and is cited with its date. The lever numbers on slide 14 (60 percent with ten steps, 95 percent with one checked retry) are standard algebra on the book’s 95%-by-20 baseline and assume a check that always catches a failed step; the calculator’s retry lever makes the same assumption and says so. No slide depends on a particular model, vendor, SDK or price; constrained decoding is described by mechanism, and provider names for it are mentioned only as varying.
The lab replaces the model with a script, so it costs nothing and runs offline. That is also the testing approach Chapter 15 recommends for the deterministic side of the harness (“test the loop itself against a scripted model”), so the lab is an early look at week 7’s observability material, not a workaround.
Exercises
In class: How reliable must each step be? (pairs, 15 minutes)
Play the Why errors compound explainer for the room. Then pairs open the compounding-error calculator, which runs in the browser with no account and no API key, and answer four questions.
- For a 90% end-to-end target over 20 steps, what per-step reliability p* is needed?
- At 95% per step, what is the longest chain that still meets the 90% target?
- At 95% per step over 20 steps, what does each lever buy: halving the steps, and one retry guarded by a check that catches failures?
- For a multi-step task of the pair’s choosing (from work or study), which lever is cheapest, and why?
Submit: the calculator’s “Copy result as Markdown” export and the answers, at most half a page.
Acceptance criteria:
- p* is given as the twentieth root of 0.9, about 99.47% (99.5% to one decimal is accepted).
- The longest chain at 95% per step is two steps, with one sentence on why that is a design constraint rather than a statistic.
- The lever comparison shows the baseline (about 36%) against both levers (about 60% for ten steps; about 95% with one checked retry) and states the assumption the retry figure depends on.
- The chosen task names its steps, and the lever choice names the check that would make a retry possible, or explains why no such check exists and so the pair shrinks n instead.
Lab 1: the minimal agent, with a scripted model client (individual, about 3 hours)
Type in the program from Appendix A, “The Program in Full,” in a language of your choice. Replace provider.send with a scripted model client: a function with the same signature that ignores the network and returns the next response from a fixed list. The list replays the appendix’s run, “The greeting in this project still says ‘Hello’. Find where it is defined and change it to ‘Welcome’.”: list_files → read_file("greetings.py") → edit_file (Hello to Welcome) → read_file("greetings.py") → a final answer with no tool calls. Run it against a scratch directory holding a greetings.py with GREETING = 'Hello'. No paid API is needed or allowed; running the same code against a local open-weight model is an optional extra.
Requirements:
- The four parts of Chapter 3 are visible as separate pieces of code: the model client (scripted), the tool registry, the message history, and the loop.
- Three tools,
list_files,read_fileandedit_file, with the descriptions and argument schemas from Appendix A;edit_fileenforces the exact-match contract (old_textmust match exactly one place; an emptyold_textcreates a file). executeturns every failure into a result tagged with the call’sid: an unknown tool name, an exception in the tool, anold_textthat matches zero or several places. Nothing a tool does can crash the loop.MAX_STEPSis a named constant checked by the loop, and an exhausted budget returns a distinct, labeledstoppedoutcome with the history, never a silent empty value (Chapter 3, “Stop Conditions, Budgets, and Frameworks”).- A run of the scripted client ends
finishedafter five passes, andgreetings.pythen containsGREETING = 'Welcome'.
Two tests that break the loop on purpose (Chapter 3 §“Stop Conditions”; Chapter 15, the deterministic–nondeterministic seam):
- A model that never says done. A stub client that requests
list_fileson every call. The test asserts that the run returnsstopped, after exactlyMAX_STEPSmodel calls, with the history attached. - A dropped append. A reactive stub that behaves like the model Chapter 3 describes: its script is two responses, a
list_filesrequest and then a final answer with no tool call. On its first call it returns the first scripted response. On every later call it checks whether the history holds a result for its previous request (matched byid): if not, it repeats that request unchanged (same name, arguments andid); once the result for its previous request is present, it returns the next scripted response. A correct harness therefore makes exactly two model calls and reachesfinisheddeterministically. Run it against your correct harness, where the test assertsfinished, and against a deliberately broken copy whose loop omits appending the tool result, where the test assertsstoppedat the cap instead of a hang. Include the broken copy in the submission (a flag or a subclass is fine).
Acceptance criteria:
- Source code, a README with the command that runs everything, and test output showing all tests passing.
- The happy-path test, both break-the-loop tests, and at least one test of
executethat feeds it an unknown tool name and asserts an error result, not an exception. - No network access in any test; tests pass with the machine offline.
- A short paragraph (at most 150 words) answering: on pass 4 of the run, what does the agent’s “verification” actually establish, and what fourth tool would turn it into evidence? (Appendix A, “A Run, End to End.”)
Reading quiz (before week 3)
A short quiz on Chapter 2 sections 4–6, Chapter 3 and Appendix A, covering the four objectives: mark the side-effect step in an unlabeled function-calling diagram; match six symptoms to the six failure modes and give a remedy for each; compute pⁿ and p* for two given pairs; and name the four parts of a minimal agent and its three exits, ranked by trust.
Reading for week 3: Chapter 4, Planning, Reasoning, and Self-Correction and Chapter 5, Tools and the Action Space (both in the full book).