Engineering AI Agents · Week 3
Planning, tools, and the action space
Plan globally, revise locally · a real verifier beats self-judgment · tools shaped to the task
Slides CC BY 4.0 · from the teaching kit of AI Agents, Engineered by Enrique Gutiérrez · aiagentsengineered.com/teach/
Last week we built the loop: a model, a message history, tools, and a stop condition. This week we ask what the model does inside that loop, and what it can reach.
The first half is Chapter 4: how a goal becomes steps, how a model thinks harder about one step, and whether it can catch its own mistakes.
The second half is Chapter 5: the tools themselves, which are the agent's entire reach into the world. Both chapters are in the full book; the slides quote them where it matters.
A goal nobody has turned into steps
“Migrate this service off the deprecated payments API .”
Who invents the steps, and in what order?
Who notices when a step went wrong?
What can the agent actually touch ?
Chapter 4 opens with this goal; the three questions are today’s three parts
Chapter 4 opens with this goal. Between that sentence and a finished migration lies a sequence of steps that nobody has written down, in an order that matters, through a system that will surprise the agent at least twice.
The three questions on the slide map onto today. Who invents the steps is planning. Who notices a wrong step is verification. What the agent can touch is the action space, its tools.
Ask the room for thirty seconds: what would you do first on Monday morning with this goal? Most people say “read a little, then make a list.” Hold onto that answer.
By the end of today you can
Four objectives
Contrast plan-then-execute with interleaved planning and say when each wins
Explain why “a real verifier beats the model grading itself”
Redesign a one-for-one API wrapping into a task-shaped tool
Write a tool description, schema and error message that pass the tool-contract linter
Four verbs you can test: contrast, explain, redesign, write.
The first two come from Chapter 4 and are checked on the reading quiz. The last two come from Chapter 5, and you will practice both in today’s exercise and in Lab 2.
Objective four has a concrete pass condition: the linter you will use in class reports no “Fix first” findings, and every “Should fix” is either fixed or answered in writing.
Part 1 · Chapter 4 (full book)
Planning, reasoning, self-correction
Chapter 4 does three things in order: planning, which turns a goal into steps; reasoning, which deepens a single step; and self-correction, which asks whether the model can check its own work.
The third part ends with the sentence this whole course leans on, so it gets the most time.
Task decomposition
The list you would write on Monday
inventory call sites → write adapter → move low-traffic endpoints → run integration tests → move the rest → delete old client
Ask for it, shape the task so it falls out, or hand the agent the plan
Lightest planning: the opening lines of a chain-of-thought
A step out of order is destructive: its wreckage lands in the history
Chapter 4, “Task Decomposition and Planning” · steps as in the book
This is the list from the chapter. It is rough, some entries will turn out wrong, but nobody starts a week-long migration without one. That list is a plan, and producing it is task decomposition.
Getting a list out of a model is the easy part: ask for subgoals, ask for an outline first, or simply give it the steps. When you already know the right steps, dictating them costs one message and removes the largest source of variance in the run.
Why care about order? Last week’s compounding arithmetic collects on every link, and a step taken out of order, migrating endpoints before the adapter exists, leaves wreckage in the history that conditions everything after.
Two ways to plan
Interleaved vs plan-then-execute
Interleaved (ReAct): one step at a time, chosen after each observation. Adaptive, myopic
Plan-then-execute : a planner drafts the whole plan, an executor works through it
Planning first attacks the missing step
“Global coherence is a thing you buy.”
Figure from Chapter 4
The loop from last week already plans: one step at a time, chosen fresh on each pass. The chapter calls this interleaved planning. Nothing outlives one observation, so it adapts; nothing forces the whole task into view, so it wanders. The book’s example: inventory the call sites beautifully, then drift into refactoring an adjacent module.
Plan-then-execute spends the first phase producing an explicit plan. The research behind it targeted a specific failure, the missing step: a necessary action silently skipped.
The split also lets an expensive model plan while cheaper models execute, and it gives you an artifact to read before any tool runs. The honest accounting: run naively, it spends more calls than the loop it replaces.
The part that separates finishing from flailing
“A plan with no replanning step is a guess with a schedule.”Chapter 4, “Task Decomposition and Planning”
Replan on surprise Return plan, completed steps and the surprise to the planner: finish, revise, or add steps
Write the plan down A todo list in the history keeps the goal in view, and lets you approve it before execution
A plan made before the first observation is a hypothesis about a world the model has not seen. Step 2 will fail, the changelog will describe the wrong version, a teammate’s commit will land mid-run.
An interleaved agent replans by construction. A plan-driven agent replans only if you build the moment in, and you must, because models “struggle to adjust plans when faced with unexpected errors,” in the words of the survey the chapter quotes. Left alone, a model bends the observation to fit the plan.
Writing the plan down helps two readers. The model: the goal stays in view instead of sinking into the poorly attended middle of the desk. You: six checkbox lines tell you whether the run is on course, and approving five proposed steps before they run is cheap oversight at the best moment.
Choosing a planning strategy · two questions
What does a wrong step cost to undo?
Task Strategy
Short, cheap steps No explicit plan; chain-of-thought is enough
Next move depends on the last result Interleaved loop
Long, dependent, expensive mistakes Plan-then-execute with replanning
Independent subtasks, latency matters A dependency graph, run in parallel
You already know the steps Draw the flowchart: a workflow
Chapter 4 · “plan globally, revise locally” · decompose to the coarsest steps that reliably succeed
The chapter reduces the choice to two questions. First: what does it cost to take a wrong step back? Backtracking a bad search costs one call; backtracking a half-executed migration costs an afternoon. Up-front planning is insurance, bought in proportion to the cost of the accident.
Second: do you already know the steps? If you can draw the flowchart, draw it and let the model work inside each box. That is week 1’s litmus test again.
One warning from the chapter: decomposition has a bottom. Every extra step is another call and another dice roll, so decompose to the coarsest steps that reliably succeed, never the finest you can imagine.
Quick check: ask the room which row a “rename this variable everywhere” task belongs in, and which row the migration belongs in.
Reasoning techniques · try it: 38 × 27, in your head
“Thinking, for a language model, is writing”
Roughly the same computation per token, whatever is at stake
Chain-of-thought : intermediate steps on the desk, before the answer
First reach on anything multi-step; it costs one sentence
Order matters
Text before the answer can steer it; text after it can only rationalize it
Read the trace as
“a map of places to dig, never as a proof of correctness”
Start with the exercise from the chapter: multiply 38 by 27 in your head, fifteen seconds. You feel the strain around the second partial product. With a pencil the strain vanishes; your arithmetic did not change, only where intermediate results were allowed to live.
The model spends roughly the same computation on every token. Demand a one-word answer and the whole problem must be solved in one token’s worth of work. Let it write out its steps first and each step lands on the desk, where it conditions everything after. That is chain-of-thought.
Two cautions. Ask for the work before the verdict, never after. And a trace is prose from the same machinery as the answer: a fluent chain to a wrong answer reads exactly as confident as one to a right answer.
Buying more thinking, and what it costs
Test-time compute and self-consistency
Test-time compute : deliberate longer before answering; a dial to set per task
Self-consistency : sample several paths, take the majority answer
Evidence without an oracle ; needs answers comparable for equality
10 samples = 10× the tokens, whether or not you show them
Figure from Chapter 4
If thinking is writing, thinking harder means writing more before committing. Models can be configured to deliberate privately at length; the field calls this test-time compute. The study the chapter cites found the payoff depends on difficulty: generous on problems hard but within reach, wasteful on easy ones. So treat reasoning effort as a dial per task.
Self-consistency turns sampling noise into an instrument. Wrong paths scatter, each wrong in its own way; right paths converge. An oracle is the testing trade’s word for a source that already knows the right answer, and the vote needs none. It does need answers you can compare: a number, a label, a date.
The bill: ten samples cost ten times the tokens, and private deliberation is billed and occupies the desk even when hidden. And note the hinge: all of these improve the attempt; none tells you whether the attempt succeeded.
Predict before the next slide
You append: “Review your answer carefully. Identify any mistakes, and produce a corrected version.”
Averaged over many tasks, does the corrected version beat the original?
The experiment Chapter 4, “Reflection and Self-Correction”, hangs on
Take a show of hands: yes, no, depends. Make everyone commit.
“Yes” is a respectable answer: rereading and revising is how all good writing gets made. The device has a name, reflection or self-critique.
The honest report needs two columns, and the line between them is the fact to carry out of this chapter.
The answer has two columns
Works · flaws in the rendering
Style, tone, missing requirement Draft–critique–revise: about 20 points better on average across seven tasks (Self-Refine)
Fails · flaws in the reasoning
Intrinsic self-correction Models “struggle to self-correct their responses without external feedback, and at times, their performance even degrades”
Reflexion’s famous gains: the post-mortem came after “task feedback signals”. The verdict walked in from outside
Early “successes” stopped revising at the known answer: a smuggled-in oracle
Chapter 4 · Madaan, Huang, Shinn et al. (2023) · figures dated, mechanism durable
Left column: revision shines where the standards were available all along and the first pass failed to apply them all at once. Tone, a requirement addressed nowhere, a buried point. The chapter summarizes a later paper: self-correction helps “style and quality,” while attempts to correct “logical or reasoning errors often cause correct answers to become incorrect.”
Right column: the landmark study isolated intrinsic self-correction, correcting without any external feedback. On reasoning, performance sometimes degraded. Asked to find mistakes, the model obliges whether or not there are any; that is the sycophancy from week 2.
And the celebrated reflection results? Look at where the signal came from: a test failed, an interpreter threw. The earlier positive experiments used ground-truth labels to decide when to stop, which smuggles in an oracle.
Why a good reasoner is a poor reviewer
Finding is hard; fixing is easy
Find the error
Weak Same weights, same blind spots: “The model grades its own homework and likes what it sees.”
Fix a located error
Strong “step four is wrong” → a well-posed repair, improved on every task tested
The division of labor: let the world find; let the model fix.
Chapter 4 · Tyen et al., “LLMs Cannot Find Reasoning Errors, but Can Correct Them Given the Error Location” (2024)
The cleanest explanation comes from an experiment that measured the two halves of reviewing separately. Finding the error: models struggle, even in unambiguous cases. Fixing it once told where it is: strong, across every task.
Why? Fixing a located error is shaped like everything the model does well: here is a flaw, produce the repair. Finding an error in your own output means rereading with the very weights that wrote it. No new information enters; the second pass is another draw from the same lottery.
A detail worth mentioning: in the same work, a small classifier outperformed a large prompted model at finding mistakes. A humble external check beat a brilliant internal one.
The sign the rest of the book points back to
“A real verifier beats the model grading itself.”Chapter 4, “Reflection and Self-Correction”
Verifier : an external check that exercises the work: tests, compiler, schema validator, calculator, the query runs
Figure from Chapter 4 · evidence vs opinion
Every road in the section arrives here. A verifier exercises the work instead of contemplating it. Its verdicts are measurements: narrow, since a passing check certifies exactly what it examines, but within that scope they cannot be argued with or flattered.
The critical survey the chapter cites says it plainly: self-correction “works well in tasks that can use reliable external feedback.”
And notice why this rule ages well. Models may get better at spotting their own slips, and the percentages will move; the asymmetry of trust will not. Where a measurement exists, an opinion is the wrong thing to bet the run on. This is the week 1 thesis read from the other end.
When there is no compiler
A ladder of checkers
Rung Its verdict is Use it
Real verifier a measurement of fact wherever one exists
Independent critic (LLM-as-a-judge)an estimate of quality separate call, rubric, ideally another model
Self-review opinion on the same desk never voluntarily
Keep the checker separate from the maker. Feed failures back whole: message, stack trace, expected vs actual.
Much of what agents produce has no compiler: a summary’s faithfulness, an argument’s quality. The fallback is an independent model critic with a separate prompt and an explicit rubric. Week 8 treats it properly; here we only place it on the ladder.
It sits below the verifier, because it estimates where the verifier measures, and above self-review, because independence gives it a fresh desk and the rubric anchors the judgment.
What to wire tomorrow: reflection downstream of a signal. When the test fails, feed the whole failure back; the model receives a located fault, which it repairs well. Bound every revision loop with a budget. And the next time you type “are you sure?” at a model, go find a check that cannot be charmed.
Part 2 · Chapter 5 (full book)
Tools and the action space
Chapter 4 ended by pointing outward: the best signals come from the environment, and an agent reaches the environment through its tools.
The second half is the craft of choosing those tools and making them usable by a caller that is capable, tireless, and not quite trustworthy.
Take this literally
“The set of tools you expose is the complete inventory of what your agent can ever do.”Chapter 5, opening
Traditional API
Deterministic code calling deterministic code; the caller cannot improvise
Tool
Deterministic code called by a non-deterministic caller : right tool with wrong arguments, wrong tool, invented tool
The action space is the universe of actions available to the agent. A model that cannot search your codebase will guess at its contents; a model with no way to ask a clarifying question will charge ahead on a wrong assumption. Both failures were settled before the first token, by what was on the tool list.
The chapter’s picture: you are outfitting a workshop for a skilled worker you will never meet. Everything must arrive in writing: the labels, the manuals, the note on a jammed machine.
So design defensively, for a caller that reasons in natural language. The good news: tools comfortable for a model tend to be comfortable for humans too.
Worked example · “find thirty minutes for me and the design team”
Wrap the endpoints, or shape the task?
list_users + list_events + create_event: directory and six calendars onto the desk; the model intersects by reading
schedule_event(participants, duration, window): code intersects; the model sees one line back
The right unit for a tool is a task ; the endpoints are plumbing
Figure from Chapter 5
The chapter’s design exercise. Ask the room first: would wrapping the three endpoints one for one work? It would, and that is what makes it a trap.
Walk the run: the whole company directory onto the desk, then pages of calendar entries per person, then the scheduling done as an act of reading. Every byte occupies scarce context, every timestamp is another draw from the lottery, and one misread entry poisons the booking.
Redesigned, the tool resolves participants, fetches calendars, intersects, books, and returns “booked, Thursday, two o’clock.” The model keeps the part that needs judgment. The same move gives you search_logs instead of read_logs, which is exactly what Lab 2 asks you to build.
The second reason one-for-one wrapping fails
Too many tools confuse the model
Every definition rides on the desk on every pass , and is an option to weigh
Overlap (read vs fetch) presents a choice where none should exist
Add capability without a tool: an entry point , or delegation to a subagent
Unavoidably large? Namespace it: crm_contacts_search, billing_invoices_search
Read transcripts; revisit the set on every model change. Subtraction is a design move
Chapter 5 · “More tools don’t always lead to better outcomes” (vendor guidance quoted there)
A model shown a tool feels a documented pull to use it, relevant or not. Past a certain count, additions subtract; week 4 gives the measurement behind this.
Two moves add capability without adding a tool: give the model a pointer it can follow, so detail loads only when needed, and delegate a messy job to a subagent that returns only the answer.
How do you know your set is right? Expose it, run real tasks, read the transcripts. Hesitations, detours and ignored tools are usually defects in your design. And expect the answer to move: the chapter tells of a todo-list tool that rescued early models and later became a cage. The chapter’s name for the habit: you learn to see like an agent.
The tool interface
“Every one of these channels is a prompt”
Description : for a capable new hire; when to use it, when not
Names : user invites four readings; user_id one
Schema : enums, required, no extra properties
Results : the relevant slice; names over UUIDs; say when truncated
Figure from Chapter 5
The definition is reread on every pass, the results become the context for the next decision, and the errors arrive exactly when the model does not know what to do. All three are prompts, whether or not you wrote them as prompts.
The evidence that this is not housekeeping: one vendor reported a benchmark result after refining tool descriptions, wording only; and a search tool kept appending the current year to queries until the parameter description was clarified. The chapter’s debugging order: when a model misuses a tool, suspect the contract before you blame the intelligence.
On results: return what informs the next action, build in pagination and limits, and when you truncate, say so, because a silent cut reads as a complete answer. On a successful mutation, return the new record’s id and URL.
The error channel is a design surface
The highest-signal message a tool sends
Corrects itself next pass
severity must be one of: low, med, high (got "urgent")
Lab 2’s structured error (Ch. 18)
{
"code": "INVALID_ARGUMENT",
"message": "severity must be one
of: low, med, high (got 'urgent')",
"retryable": false,
"hint": "ask the user which
severity they mean"
}
Chapter 5, “The Tool Interface” · field shape from Chapter 18 (code, message, retryable, hint)
An error arrives at the exact moment the model does not know what to do next, and most tools waste it. A bare status code teaches nothing; a raw stack trace teaches less than its bulk suggests.
The good message says what went wrong, the valid set, and the fix implied, and the model corrects itself on the next pass with no human in sight. Write every error as instructions to a capable colleague who cannot see your face.
Chapter 18 adds the structure we will require in Lab 2: a stable code rather than free prose, whether the failure is retryable, and a hint.
Read the example carefully. retryable says whether the same call might succeed if repeated, as after a timeout; this one never will, so it is false. A corrected call is a new call. And because “urgent” came from the user, the hint does not pick a value on the user’s behalf: it sends the model back to ask.
Last week you made execute turn every failure into a result; this week you make those results worth reading.
Retrofitting for agents · 1 of 2
Make explicit what a human used to supply implicitly
Never block on a prompt; fail fast when there is no terminal
Destructive: confirm by default, --yes to bypass
Structured output on stdout; diagnostics on stderr
Help says when to use a command, and what comes next
Chapter 5 · “Crashes are tolerable; hangs are problematic” (field notes quoted there)
Much of what an agent drives was built for a forgiving human: she types “y” at a prompt, remembers the help text for years, notices the duplicated row. The agent cannot answer a prompt nobody surfaced, rereads help on every call, and retries tirelessly. Every gap human resourcefulness papered over becomes a hang, a burned retry, or a duplicate record.
The silent hang is the most expensive failure: the agent learns nothing and recovers slowly. That is why the first rule is never to wait on a prompt nobody will answer, and why output the next step must parse goes to stdout while diagnostics go to stderr.
Retrofitting for agents · 2 of 2
Assume every call will be retried
Idempotent : a retried create returns the existing record
--dry-run on every mutation
Predictable vocabulary: posts list, posts create
Harden inputs as if they were hostile
Chapter 5, retrofitting rules for agent-facing programs
Idempotency matters because agents retry on timeouts and ambiguous output without the human glance that would notice a duplicate. A dry run lets a constructed call be validated before the one execution that matters.
A predictable vocabulary means the model can guess the next command instead of rereading help, and hostile-input hardening is there because the arguments come from a model that may have read untrusted text.
None of these rules hurt human users. They are the design we always meant to do. Lab 2 asks for dry-run semantics on your mutating tool.
In class · teams of three · 20 minutes
Lint seven flawed tools, fix them, re-lint
Open the set in the browser linter: seven flawed tool definitions (also in your handout)
Fix every “Fix first” finding; then the “Should fix” ones
Find two flaws the linter missed ; name the chapter rule each breaks
Re-lint; copy the result as Markdown and submit with your fixed JSON
read_logs search find_records create_ticket deleteUser get_customer run
aiagentsengineered.com/tools/tool-contract-linter/ · runs in the browser · no account, no API key
Each team gets the same seven definitions, deliberately flawed: a log reader with no limit, two overlapping search tools, a ticket creator with an ambiguous user field and no enum, a delete with no dry run, a getter that secretly sends email, and a generic “run” tool.
The link opens the linter with the set already loaded. The first lint reports dozens of findings; teams work through them in severity order.
Step 3 matters most. The linter checks the contract, not the design: it does not notice that two of the tools should be one, or that the getter should be split in two. When you redesign, you will rename and merge tools, which is the point.
Debrief: ask two teams to show their before and after for the delete tool, and one team to read its two misses.
Code execution as a universal action
“Discretion on the outside, determinism on the inside”
One meta-tool , run_code(source), composes all the rest
10,000 rows stay in the runtime; five lines come back
Sandbox : own filesystem, capped CPU, memory and time; network denied by default; fresh per session
A script that ran proves nothing about what it computed
Figure from Chapter 5
The chapter’s scenario: filter a ten-thousand-row export, join it with an at-risk list, email the owners. Which tool filters, which joins? You can add a tool per operation, or let the model filter by reading. Both are bad. The third option is one tool that writes and runs programs.
The model decides what to compute; code computes it. The chapter reports one vendor’s illustrative example, about 150,000 tokens down to about 2,000, by keeping intermediate results inside the execution environment. The number will vary; the direction will not.
The non-negotiable: never run model-written code with access to anything you would not hand a stranger. The code comes from a probabilistic author who may have been steered by untrusted text. Honest edges: a sandbox is infrastructure, and for a single lookup it is overkill.
When there is no tool end at all
Computer use: descend only as far as forced
API : a deterministic contract
Structured browser : the accessibility tree ; click by meaning
Pixels : universal, slowest, most fragile
The page is input written by strangers: indirect prompt injection
Figure from Chapter 5
The supplier portal with no API, only a login page and a “Download CSV” button: computer use hands the model a screen, a cursor and a keyboard. It is the same loop, specialized: perceive, reason, act, look again. The model requests the click; your harness performs it, and that gap is where safety checks live.
Perception is the central decision. Pixels work on anything that renders but need coordinates and carry no meaning. The accessibility tree names every element, so a layout shift does not misdirect the click, and your harness knows what is being clicked. Strong deployments go hybrid.
The headline risk: hostile instructions hidden in pages the agent reads anyway. Week 9 treats it fully; until then, least privilege, human confirmation before consequential actions, and the working assumption that every byte from the open web is hostile.
Recap
Five things to keep
Plan globally, revise locally ; no replanning step means a guess with a schedule
Thinking is writing: work before the answer, never after
A real verifier beats the model grading itself: let the world find, let the model fix
The action space is everything the agent can do; shape tools to tasks, keep them few
Descriptions, results and errors are prompts; write them for a colleague who cannot see you
If students leave with these five lines, the week did its job.
Quick check before they leave: name one verifier for a task from your own work, and one task where you would have to fall back to an independent critic.
Next week builds on the last line: tools, protocols and skills are three layers, and the context window they all compete for is the scarce resource.
Before next week
Lab 2 and reading
Lab 2 · extend Lab 1
run_tests: a real verifier
search_logs: a task-shaped tool
Every error: code, message, retryable, hint
dry_run on the mutating tool
Scripted-client test: recovery from a malformed argument
This week’s reading: Ch. 4 · Ch. 5 (full book) · no paid API needed for the lab
Lab 2 extends the minimal agent from Lab 1, still with a scripted model client, so nobody needs a paid API.
You add a real verifier, run_tests, and a task-shaped search_logs that returns matches with context instead of whole files. Every tool error becomes structured, and the file-editing tool gets dry-run semantics.
The key test: the scripted model sends a malformed argument, your tool returns a structured error, and the next scripted turn sends a corrected call. The run must still finish. That is today’s chapter in one test.
Reading for next week is Chapters 6 and 7, both in the full book.
Engineering AI Agents · Week 3
Let the world find; let the model fix.
Slides from the teaching kit of AI Agents, Engineered by Enrique Gutiérrez, CC BY 4.0 · aiagentsengineered.com/teach/
Reuse, adapt and translate with attribution · creativecommons.org/licenses/by/4.0/
End on the division of labor from Chapter 4. It is the shape of every reliable agent loop we will build this semester: the environment supplies the verdict, the model supplies the repair.
These slides are CC BY 4.0: instructors may reuse, adapt and translate them, keeping the attribution line on this slide.