How to run the session
Week 3 has two halves that meet in one idea. Chapter 4 (Planning, Reasoning, and Self-Correction) asks how a goal becomes steps, how a model thinks harder about one step, and whether it can catch its own mistakes; its answer, “a real verifier beats the model grading itself,” points the course out of the model and into the environment. Chapter 5 (Tools and the Action Space) is the craft of that environment: the tools that are the agent’s entire reach, and the surfaces (descriptions, schemas, results, errors) through which the model holds them. Both chapters are in the full book; the slides quote them verbatim where it matters, so students who have not bought the book can still follow the lecture.
The deck has 29 slides in two halves, split by section slides. Press N for presenter notes, F for fullscreen; slide numbers are deep links (deck.html#24 is the exercise). The PDF is the same deck, one slide per page, without notes.
Four slides are built to be answered by the room before you advance: slide 2 (what would you do first with the migration goal?), slide 8 (which row does a given task belong in?), slide 11 (does “review your answer” improve it on average? take a show of hands and make everyone commit) and slide 18 (would wrapping the three calendar endpoints work?). Slide 9 opens with the chapter’s 38 × 27 mental-arithmetic exercise; give it the full fifteen seconds, because the felt strain is the explanation of chain-of-thought.
Have the in-class handout ready (the seven flawed definitions below, printed or posted), and check before class that the linter link on slide 24 opens with the set loaded; the link carries the JSON in the page address, so it works offline once the page is cached.
Timing (80 minutes)
| Segment | Slides | Minutes |
|---|---|---|
| Opening: the migration goal, objectives | 1–3 | 5 |
| Planning: decomposition, interleaved vs plan-then-execute, replanning, choosing | 4–8 | 13 |
| Reasoning: chain-of-thought, test-time compute, self-consistency | 9–10 | 7 |
| Self-correction: the prediction, two columns, find vs fix, the verifier rule and ladder | 11–15 | 12 |
| Tools: the action space, the calendar worked example, too many tools | 16–19 | 9 |
| The interface: anatomy, errors, retrofitting for agents (two slides) | 20–23 | 8 |
| In-class exercise: lint, fix, re-lint, then debrief | 24 | 20 |
| Code execution and computer use | 25–26 | 4 |
| Recap, Lab 2, reading | 27–29 | 2 |
The plan fits the same 80-minute slot as weeks 1 and 2. The exercise is the tightest segment: 20 minutes leaves about 13 for lint-fix-re-lint and 7 for the debrief, so tell teams to fix the three “Fix first” findings and the delete tool before anything else. It sits before code execution and computer use on purpose: it applies slides 18–23 while they are fresh, and slides 25–26 can be compressed if the debrief runs long. If your slot is 85 minutes or more, give the exercise the extra five minutes first; that is the variant the exercise was first timed at. For a 75-minute slot, cut the exercise to 17 minutes and cover slides 25–26 in one pass of two minutes, pointing students to the chapter. For a 90-minute slot, give the exercise 30 minutes and have one team present its whole redesigned set.
Common misconceptions and how to address them
“Plan-then-execute means the agent follows a fixed script.” That is the blind executor Chapter 3 rejected. The chapter’s plan is a working document inside a loop that still observes, and the replan path is what makes the design serious: “A plan with no replanning step is a guess with a schedule.” Slide 6’s figure draws the replan path as its widest arc, from the executor back to the planner; point at it, then let slide 7 carry the quotation.
“More planning is always safer.” Planning is insurance bought in proportion to the cost of the accident, and decomposition has a bottom: every extra step is another call and another chance to lose the thread. If the steps are known, the right answer is not a better planner but a workflow (week 1’s litmus test).
“The reasoning trace explains why the model answered as it did.” A trace is prose from the same machinery as the answer. It is a map of places to dig, never a proof. A fluent chain to a wrong answer reads exactly as confident as one to a right answer.
“Asking the model to double-check its work catches errors.” This is the misconception slide 11 is built to surface, and many students will vote yes. The evidence splits: revision helps with style, tone and missing requirements; intrinsic self-correction of reasoning errors often makes things worse, and early positive results had smuggled in an oracle through the stopping rule. Do not overcorrect into “reflection is useless”: reflection downstream of a real signal (a failing test fed back whole) is the workhorse of practical reliability. The rule is about where the verdict comes from.
“Self-consistency is a verifier.” It is evidence gathered without an oracle, and it improves the attempt, but agreement is not correctness: paths can converge on the same wrong answer, and the technique needs answers comparable for equality. It sits with the reasoning techniques, not on the verifier rung.
“A good tool set mirrors the API.” The calendar example on slide 18 is the antidote; one-for-one wrapping works, which is why it is a trap. Ask “what task is the model trying to do?” and move the deterministic work inside the tool.
“If the model misuses a tool, we need a smarter model.” The chapter’s debugging order is the reverse: suspect the contract before you blame the intelligence. Its two vendor anecdotes (a benchmark gain from description refinements alone; a search tool that appended the current year until its parameter description was fixed) are on the notes for slide 20.
“A clean lint means a good tool.” The linter checks the shape of the contract with heuristics; it cannot tell that two tools should be one, or that a description is untrue. Step 3 of the exercise (find two flaws the linter missed) exists to make this concrete.
Materials and notes for instructors
All seven figures named in the syllabus for this week are on the slides (Chapter 4: plan vs interleaved, self-consistency vote, verifier vs self-judgment; Chapter 5: task-shaped tools, tool interface anatomy, code sandbox, UI reliability ladder) and are available individually on the diagrams page. Numbers on the slides are the book’s and are labeled as dated or illustrative (the Self-Refine average, the 150,000-to-2,000-token example). No slide depends on a particular model, vendor or framework; where the chapter names a vendor’s experience, the slides present it as one reported example.
The structured error shape required in Lab 2 (code, message, retryable, hint) comes from Chapter 18, “Reliability, State, and the Harness”; slide 21 says so. Students meet it fully in week 10.
Instructor key for the exercise. A good redesign of the seven tools: read_logs becomes search_logs (pattern, service enum, ISO 8601 since, limit with a maximum, truncation announced); search and find_records merge into one search_records that returns readable labels with each record_id; create_ticket takes user_id, a severity enum, a due_date with a stated format and an idempotency_key, and returns the ticket’s id and URL; deleteUser becomes deactivate_user with dry_run, confirm, a human approval and a stated undo window; get_customer is split into a read-only get_customer and a separate send_survey (with dry_run and the ticket as idempotency key); run becomes a sandboxed run_python. Linted, that set reports no “Fix first” findings, one “Should fix” and two “Consider” (verified on 2026-10-05 by running the linter’s own lint() function, from the published app.js, on the full key below; the original seven-tool set lints at 3 “Fix first”, 33 “Should fix” and 30 “Consider”). The one “Should fix” is a set-level warning that the tools may complete the lethal trifecta (private data, untrusted content, a way out). The linter matches it by keywords, but the legs are real: customer records, ticket text written by customers, and an email that leaves the building. That finding cannot be fixed by rewording, which makes it a good debrief question and a preview of week 9; the Lethal Trifecta Audit is the follow-up tool. The two “Consider” findings are worth a minute too: run_python is flagged as a general-purpose execution tool (the description already confines it to a sandbox, so the answer is “yes, and it is sandboxed”), and deactivate_user is flagged as having no verb because “deactivate” is not in the linter’s word list, a heuristic miss that shows why a lint is not a review. The two flaws the linter does not flag in the original set are the overlap between search and find_records (two tools for one choice) and the one-for-one, endpoint-shaped design of the set as a whole; accept any other genuine miss that a team can tie to a chapter rule. The full key, all seven redesigned tools (paste it into the linter to reproduce the result):
Key: the seven redesigned tool definitions (JSON)
[
{
"name": "search_logs",
"description": "Searches one service's logs for lines matching a pattern and returns the matching lines with their timestamps, newest first. Use when diagnosing an incident or checking whether an error occurred; prefer search_records for user or customer data. Returns at most limit lines and says when results were truncated, so narrow the pattern or the since window rather than raising the limit. Read-only. Errors: UNKNOWN_SERVICE (not retryable; hint: use one of the listed services), INVALID_SINCE (not retryable; hint: use ISO 8601).",
"input_schema": {
"type": "object",
"properties": {
"pattern": {
"type": "string",
"description": "Text or regular expression to match, for example \"timeout\" or \"HTTP 5\\\\d\\\\d\"."
},
"service": {
"type": "string",
"enum": [
"api",
"billing",
"auth",
"worker"
],
"description": "The service whose logs to search."
},
"since": {
"type": "string",
"format": "date-time",
"description": "Only lines at or after this time, ISO 8601, for example 2026-10-05T09:00:00Z."
},
"limit": {
"type": "integer",
"minimum": 1,
"maximum": 200,
"default": 50,
"description": "Maximum number of lines to return."
}
},
"required": [
"pattern",
"service"
],
"additionalProperties": false
}
},
{
"name": "search_records",
"description": "Searches user and customer records by name, email address or company and returns up to limit matches, each with a readable label (full name plus email address) and its record_id. Use this to find the record_id that other tools need; never guess an id. Do not use it for log lines; use search_logs instead. Read-only. If several records match, show the labels to the user and ask which one is meant. Errors: NO_MATCH (not retryable; hint: try a shorter query or ask the user).",
"input_schema": {
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "A name, email address or company name, for example \"Ana Ruiz\"."
},
"limit": {
"type": "integer",
"minimum": 1,
"maximum": 25,
"default": 10,
"description": "Maximum number of matches to return."
}
},
"required": [
"query"
],
"additionalProperties": false
}
},
{
"name": "create_ticket",
"description": "Creates a support ticket for one user. Use when the user reports a problem that needs follow-up by the support team. Pass an idempotency_key derived from the request so a retried call returns the existing ticket instead of a duplicate. Returns the ticket_id and the ticket's URL. If you do not know the severity, do not guess; ask the user. Errors: USER_NOT_FOUND (not retryable; hint: look the user up with search_records), INVALID_ARGUMENT (not retryable; hint: the message names the field and its valid values).",
"input_schema": {
"type": "object",
"properties": {
"user_id": {
"type": "string",
"description": "The user's record_id, as returned by search_records."
},
"title": {
"type": "string",
"maxLength": 120,
"description": "One-line summary of the problem."
},
"severity": {
"type": "string",
"enum": [
"low",
"med",
"high"
],
"description": "How urgent the problem is, as stated by the user."
},
"due_date": {
"type": "string",
"format": "date",
"description": "Date the ticket should be resolved by, YYYY-MM-DD."
},
"idempotency_key": {
"type": "string",
"description": "A stable key for this logical request; reusing it returns the existing ticket."
}
},
"required": [
"user_id",
"title",
"severity",
"idempotency_key"
],
"additionalProperties": false
}
},
{
"name": "deactivate_user",
"description": "Deactivates a user account so the user can no longer sign in. Use only when an admin explicitly asks to remove a user. Reversible for 30 days (restore from the admin console); permanent after that. Call first with dry_run true: it returns what would change. The real call requires confirm true and an approval from a human. Returns the user_id and the date the deactivation becomes permanent. If you are unsure which user is meant, do not guess; ask. Errors: USER_NOT_FOUND (not retryable), APPROVAL_REQUIRED (not retryable; hint: request approval).",
"input_schema": {
"type": "object",
"properties": {
"user_id": {
"type": "string",
"description": "The user's record_id, as returned by search_records."
},
"dry_run": {
"type": "boolean",
"default": true,
"description": "If true, report the consequences without changing anything."
},
"confirm": {
"type": "boolean",
"default": false,
"description": "Must be true for the real deactivation."
}
},
"required": [
"user_id",
"dry_run"
],
"additionalProperties": false
}
},
{
"name": "get_customer",
"description": "Gets one customer's profile by record_id. Use when you need a customer's plan, contact details or open tickets before acting. Returns the name, email address, plan and open ticket ids. Read-only: it changes nothing and contacts no one; to contact the customer, use send_survey instead. If you do not know the customer_id, do not guess; look it up with search_records or ask the user. Errors: CUSTOMER_NOT_FOUND (not retryable; hint: look the customer up with search_records).",
"input_schema": {
"type": "object",
"properties": {
"customer_id": {
"type": "string",
"description": "The customer's record_id, as returned by search_records."
}
},
"required": [
"customer_id"
],
"additionalProperties": false
}
},
{
"name": "send_survey",
"description": "Sends a customer the satisfaction survey for one closed ticket, by email. Use only after a ticket is closed and the user asks for feedback to be collected. Call first with dry_run true to see the recipient and the message. The ticket_id acts as the idempotency key: a second call for the same ticket returns the earlier send instead of emailing again. A sent survey cannot be undone. Returns the survey_id and the send time. If you do not know the ticket_id, do not guess; ask the user. Errors: TICKET_NOT_CLOSED (not retryable; hint: wait until the ticket is closed).",
"input_schema": {
"type": "object",
"properties": {
"customer_id": {
"type": "string",
"description": "The customer's record_id, as returned by search_records."
},
"ticket_id": {
"type": "string",
"description": "The closed ticket the survey is about; also the idempotency key."
},
"dry_run": {
"type": "boolean",
"default": true,
"description": "If true, return the recipient and message without sending."
}
},
"required": [
"customer_id",
"ticket_id",
"dry_run"
],
"additionalProperties": false
}
},
{
"name": "run_python",
"description": "Runs a short Python script in an isolated sandbox with no network access and no credentials, and returns stdout, stderr and the exit code. Use for calculations and data transformations on data you already have. Do not use it to reach company systems; use search_records or get_customer instead. Output is truncated to 10,000 characters, and the run is stopped after timeout_seconds. Errors: TIMEOUT (retryable once; hint: simplify the script or raise timeout_seconds), SANDBOX_ERROR (retryable; hint: try again).",
"input_schema": {
"type": "object",
"properties": {
"code": {
"type": "string",
"description": "The Python source to run."
},
"timeout_seconds": {
"type": "integer",
"minimum": 1,
"maximum": 60,
"default": 20,
"description": "Wall-clock limit for the run."
}
},
"required": [
"code"
],
"additionalProperties": false
}
}
]
Exercises
In class: lint seven flawed tools, fix them, re-lint (teams of three, 20 minutes)
Each team opens the seven deliberately flawed tool definitions below in the browser tool-contract linter (the link loads the set; teams can also paste the JSON). The linter runs in the browser, needs no account and no API key, and quotes the chapter rule behind every finding.
[
{
"name": "read_logs",
"description": "Reads the service logs.",
"input_schema": {
"type": "object",
"properties": { "service": { "type": "string" } }
}
},
{
"name": "search",
"description": "Searches for things in the system and returns stuff.",
"input_schema": {
"type": "object",
"properties": { "q": { "type": "string", "description": "What to search for." } },
"required": ["q"]
}
},
{
"name": "find_records",
"description": "Finds records. Returns a list of matching records with their UUIDs.",
"input_schema": {
"type": "object",
"properties": { "query": { "type": "string" } },
"required": ["query"]
}
},
{
"name": "create_ticket",
"description": "Creates a support ticket for a user. Returns ok.",
"input_schema": {
"type": "object",
"properties": {
"user": { "type": "string", "description": "The user." },
"severity": { "type": "string", "description": "How bad it is." },
"due": { "type": "string" },
"title": { "type": "string" }
}
}
},
{
"name": "deleteUser",
"description": "Deletes a user account from the database. Use when an admin asks to remove someone.",
"input_schema": {
"type": "object",
"properties": { "id": { "type": "string" } },
"required": ["id"]
}
},
{
"name": "get_customer",
"description": "Gets a customer profile by email and then sends them a satisfaction survey. Returns the profile.",
"input_schema": {
"type": "object",
"properties": { "customer": { "type": "string" } },
"required": ["customer"]
}
},
{
"name": "run",
"description": "Runs a command.",
"input_schema": {
"type": "object",
"properties": { "options": { "type": "object" } }
}
}
]
Teams fix the “Fix first” findings first, then the “Should fix” ones, renaming, merging or splitting tools where the design calls for it. Then they look for flaws the linter did not report, and re-lint.
Submit: the fixed JSON, the linter’s “Copy result as Markdown” export of the final run, and a short note.
Acceptance criteria:
- The final lint reports zero “Fix first” findings.
- Every remaining “Should fix” finding is either fixed or answered in one sentence (for example, a set-level warning the team judges acceptable, with the reason).
- The redesign removes at least one overlap or one-for-one wrapper, with one sentence naming what the model no longer has to do in its context.
- Every mutating tool has either an idempotency mechanism or a dry run, and the destructive one states whether it can be undone.
- The note names two flaws the linter missed and, for each, the Chapter 5 rule it breaks.
Lab 2: a verifier, a task-shaped tool and structured errors (individual, about 4 hours)
Extend your Lab 1 agent (Appendix A’s minimal agent with a scripted model client and the three file tools). No paid API is needed: the model remains a scripted client that returns a fixed sequence of responses. Optionally, also run it against a local open-weight model.
- Add
run_tests, a real verifier: it runs the project’s test command in a subprocess with a timeout and returns pass/fail, the number of failures and the first failure’s message, never the full log. - Add a task-shaped
search_logs(pattern, limit)(or an equivalent search over files) that returns matching lines with a little surrounding context, capped bylimit, and says when results were truncated, in place of reading whole files. - Make every tool error structured:
code(a stable identifier),message(what happened and what a correct call looks like),retryable(boolean) andhint(what to try next). No tool raises out ofexecute. - Give the mutating tool (
edit_file)dry_runsemantics: withdry_run: trueit validates the arguments and returns the diff it would apply, changing nothing. - Write the tool descriptions and schemas for all five tools and lint them with the tool-contract linter.
Acceptance criteria:
- A scripted-client test in which the model’s first call to a tool has a malformed argument (a missing required field or a value outside an enum); the test asserts that the tool returns a structured error with
retryableandhint, that the error is appended to the message history as a tool result, that the next scripted call is corrected, and that the run reaches its final answer. - A test that
run_testsreports failure on a deliberately broken file and success after the scripted edit, and that the loop’s “done” depends onrun_testspassing, not on the model saying so (keep the checker separate from the maker). - A test that
edit_filewithdry_run: trueleaves the file unchanged and returns the diff. - A test that
search_logsrespectslimitand reports truncation. - The Lab 1 guarantees still hold:
MAX_STEPSwith a loudstoppedoutcome, and every failure turned into a result. - The linter’s Markdown export for the five tools, with no “Fix first” findings, is committed alongside the code.
- No network calls and no paid API anywhere in the test suite.
Reading quiz (before week 4)
A short quiz on Chapters 4 and 5 covering the four objectives: choose a planning strategy for four described tasks; explain the find-versus-fix asymmetry and name a verifier for a given task; redesign a three-endpoint tool set into one task-shaped tool; rewrite a bad error message into a structured one.
Reading for week 4: Chapter 6, Skills, Protocols, and Interoperability and Chapter 7, Managing the Context Window (both in the full book).