How to run the session
This is the opening lecture, and it carries two jobs: it pins down the vocabulary the rest of the course depends on, and it opens the hood on the language model just far enough that later failures stop looking like magic. Everything it uses is free to read: the preface, Chapter 1 and the first three sections of Chapter 2 (“LLMs as Next-Token Predictors”, “Tokens and the Context Window”, “How Text Is Generated”). Students can start the course without buying anything.
The deck has 28 slides in two halves, split by section slides. Press N for presenter notes, F for fullscreen; slide numbers are deep links (deck.html#15). The PDF is the same deck, one slide per page, without notes, for handouts and for instructors who present from a PDF viewer.
Run it as a conversation, not a recital. Three slides are built to be answered by the room before you advance: slide 8 (which of three systems is the agent?), slide 20 (finish “The capital of France is…”) and slide 24 (which word does the model write?). Slide 25 opens with a prediction question too. Take a show of hands each time; the wrong answers are the teaching moment.
Timing (80 minutes)
| Segment | Slides | Minutes |
|---|---|---|
| Opening: the trust question, the thesis, objectives | 1–3 | 7 |
| What an agent is: definition, control flow, why now | 4–7 | 10 |
| Chatbot, workflow, agent: the three systems, the litmus test, the augmented LLM, the dial | 8–12 | 14 |
| The compass and its forks; compounding error worked by hand | 13–15 | 10 |
| Do you need an agent? The ladder and the calibration pair | 16–17 | 6 |
| In-class exercise: Should this be an agent? in pairs, then debrief | 18 | 15 |
| The engine: next-token prediction, training vs inference, tokens | 19–22 | 9 |
| The desk, the lottery, nondeterminism | 23–25 | 7 |
| Recap, homework, reading | 26–28 | 2 |
For a 90-minute slot, give the exercise 20 minutes and spend the other five on slide 15, having students compute the table rows themselves. For a 75-minute slot, shorten the debrief and merge slides 11 and 12 into one pass.
Common misconceptions and how to address them
“An agent is any app with an LLM in it.” This is the agent-washing reflex. Answer with the litmus test on slide 10: if you can draw the control-flow diagram before the request arrives, it is a workflow, however it is marketed. Ask students for a product they believe is an agent and test it live.
“More autonomy is always better; workflows are a beginner’s version.” Chapter 1 says the opposite: staying on the lowest rung that solves the problem is a discipline, and the dominant production shape is “a workflow shell with one or two genuinely agentic steps inside it.” The email-routing example on slide 17 is a workflow permanently, not temporarily.
“The model remembers our earlier chats” or “it learns from my corrections.” Inference never changes a weight. What feels like memory in a session is the transcript laid back on the desk at every call (slides 21 and 23). This misconception returns in week 5 on memory; fix it now.
“Temperature zero makes the output deterministic.” It removes the lottery, not the variation. Slide 25 carries the reported experiment (80 distinct answers from 1,000 temperature-zero runs) and its explanation through batch invariance. The engineering response is to test properties of the output instead of exact strings.
“If the model counts letters correctly, the tokenizer argument is wrong.” Newer models may pass, because gaps get patched by training or the model reaches for a tool; Chapter 2 says as much. The mechanism the failure exposes is still there. The homework asks students to explain either outcome.
“95% per step is good enough.” Have them compute 0.95²⁰ before you show it. The surprise is the lesson, and next week’s compounding-error material builds on it.
Materials and notes for instructors
The diagrams on the slides are the book’s own (Chapters 1 and 2), also available individually on the diagrams page for your own slides. The syllabus lists two more Chapter 2 figures for this week, the context window as a desk and sampling as a weighted lottery; those are illustrations not published as SVG on the site, so slides 23 and 24 draw the book’s own numbers as simple bars instead. The syllabus also names a short explainer, What is an agent?, which animates slides 8–10 in about three and a half minutes; play it before the classification exercise or assign it beforehand. The agent loop explainer is a good preview for students who want to look ahead to week 2.
All numbers on the slides are the book’s illustrations and are labeled as such. No slide depends on a particular model, vendor or price.
Exercises
In class: Should this be an agent? (pairs, 15 minutes)
Each pair writes down three requests from their own work or studies, ideally one they suspect needs an agent and one they suspect does not. They run each request through the Should this be an agent? browser tool, which places the task on the escalation ladder (plain code, single call, workflow, agent, agent as co-pilot) and explains why. It needs no account and no API key.
Submit: the tool’s “Copy result as Markdown” export for all three requests, plus two lines per request.
Acceptance criteria:
- Three distinct requests, each with the tool’s verdict export attached.
- For each request, one line naming who owns the control flow in the recommended design (the human, the code or the model), consistent with the litmus test from Chapter 1.
- For each request, one line naming the signal that would tell you it worked (a test, a schema, a source, a human approval), or stating that none exists and what that implies for autonomy.
- At least one request where the pair disagrees with the tool or was surprised by it, with one sentence on why.
Homework: the tokenizer experiment (individual, about 1 hour)
Using any model you can reach for free (a free chat interface or a locally runnable open-weight model; no paid API is needed), ask it (a) how many times a given letter appears in three ordinary words of your choice, and (b) to spell two words backwards. Record the exact prompts and answers. Then write one page explaining the outcomes using subword tokenization, as described in Chapter 2, “Tokens and the Context Window”.
Acceptance criteria:
- The prompts and the model’s verbatim answers are included, with the kind of model used named as a category (for example “a free hosted chat model” or “a local open-weight model”), not as a product recommendation.
- The explanation states that the tokenizer is a fixed lookup table and that the model sees token IDs, not characters.
- Every outcome is explained, including correct answers: what the model must do to succeed (reconstruct the spelling from what it absorbed in training, or use a tool).
- One paragraph applies the idea beyond spelling: why code, JSON or another language may cost more tokens than English prose.
- At most one page; no paid API used.
Reading quiz (before week 2)
A short quiz on the preface, Chapter 1 and Chapter 2 sections 1–3, covering the four objectives: classify five systems by who owns the control flow, match each compass bearing to a fork, explain the desk and the lottery, and compute pⁿ for two given pairs of p and n.
Reading for week 2: Chapter 2 sections 4–6 (free), Chapter 3, The Agent Loop and Appendix A, A Minimal Agent, Annotated (both in the full book).