Engineering AI Agents · Week 3

Planning, tools,
and the action space

Plan globally, revise locally · a real verifier beats self-judgment · tools shaped to the task

Slides CC BY 4.0 · from the teaching kit of AI Agents, Engineered by Enrique Gutiérrez · aiagentsengineered.com/teach/

A goal nobody has turned into steps

“Migrate this service off the deprecated payments API.”

Chapter 4 opens with this goal; the three questions are today’s three parts

By the end of today you can

Four objectives

  1. Contrast plan-then-execute with interleaved planning and say when each wins
  2. Explain why “a real verifier beats the model grading itself”
  3. Redesign a one-for-one API wrapping into a task-shaped tool
  4. Write a tool description, schema and error message that pass the tool-contract linter

Part 1 · Chapter 4 (full book)

Planning, reasoning, self-correction

Task decomposition

The list you would write on Monday

inventory call sites→write adapter→move low-traffic endpoints→run integration tests→move the rest→delete old client

Chapter 4, “Task Decomposition and Planning” · steps as in the book

Two ways to plan

Interleaved vs plan-then-execute

  • Interleaved (ReAct): one step at a time, chosen after each observation. Adaptive, myopic
  • Plan-then-execute: a planner drafts the whole plan, an executor works through it
  • Planning first attacks the missing step
  • “Global coherence is a thing you buy.”
Two ways to plan. Left: a tight interleaved cycle of decide, act, observe. Right: a planner drafts the plan, an executor works through it, and a wide replan path returns surprises to the planner.
Figure from Chapter 4

The part that separates finishing from flailing

“A plan with no replanning step is a guess with a schedule.”Chapter 4, “Task Decomposition and Planning”

Replan on surprise

Return plan, completed steps and the surprise to the planner: finish, revise, or add steps

Write the plan down

A todo list in the history keeps the goal in view, and lets you approve it before execution

Choosing a planning strategy · two questions

What does a wrong step cost to undo?

TaskStrategy
Short, cheap stepsNo explicit plan; chain-of-thought is enough
Next move depends on the last resultInterleaved loop
Long, dependent, expensive mistakesPlan-then-execute with replanning
Independent subtasks, latency mattersA dependency graph, run in parallel
You already know the stepsDraw the flowchart: a workflow

Chapter 4 · “plan globally, revise locally” · decompose to the coarsest steps that reliably succeed

Reasoning techniques · try it: 38 × 27, in your head

“Thinking, for a language model, is writing”

  • Roughly the same computation per token, whatever is at stake
  • Chain-of-thought: intermediate steps on the desk, before the answer
  • First reach on anything multi-step; it costs one sentence

Order matters

Text before the answer can steer it; text after it can only rationalize it

Read the trace as

“a map of places to dig, never as a proof of correctness”

Buying more thinking, and what it costs

Test-time compute and self-consistency

  • Test-time compute: deliberate longer before answering; a dial to set per task
  • Self-consistency: sample several paths, take the majority answer
  • Evidence without an oracle; needs answers comparable for equality
  • 10 samples = 10× the tokens, whether or not you show them
Self-consistency: one hard question sampled several times; right routes converge on one answer while wrong routes scatter, and a majority vote recovers the answer most paths agree on.
Figure from Chapter 4

Predict before the next slide

You append: “Review your answer carefully. Identify any mistakes, and produce a corrected version.”

Averaged over many tasks, does the corrected version beat the original?

The experiment Chapter 4, “Reflection and Self-Correction”, hangs on

The answer has two columns

Works · flaws in the rendering

Style, tone, missing requirement

Draft–critique–revise: about 20 points better on average across seven tasks (Self-Refine)

Fails · flaws in the reasoning

Intrinsic self-correction

Models “struggle to self-correct their responses without external feedback, and at times, their performance even degrades”

Chapter 4 · Madaan, Huang, Shinn et al. (2023) · figures dated, mechanism durable

Why a good reasoner is a poor reviewer

Finding is hard; fixing is easy

Find the error

Weak

Same weights, same blind spots: “The model grades its own homework and likes what it sees.”

Fix a located error

Strong

“step four is wrong” → a well-posed repair, improved on every task tested

The division of labor: let the world find; let the model fix.

Chapter 4 · Tyen et al., “LLMs Cannot Find Reasoning Errors, but Can Correct Them Given the Error Location” (2024)

The sign the rest of the book points back to

“A real verifier beats the model grading itself.”Chapter 4, “Reflection and Self-Correction”

Verifier: an external check that exercises the work: tests, compiler, schema validator, calculator, the query runs

The generate, check, revise loop wired two ways: to a real verifier whose verdict is evidence, and, dashed, back to the model's own judgment, whose verdict is only opinion.
Figure from Chapter 4 · evidence vs opinion

When there is no compiler

A ladder of checkers

RungIts verdict isUse it
Real verifiera measurement of factwherever one exists
Independent critic (LLM-as-a-judge)an estimate of qualityseparate call, rubric, ideally another model
Self-reviewopinion on the same desknever voluntarily

Keep the checker separate from the maker. Feed failures back whole: message, stack trace, expected vs actual.

Part 2 · Chapter 5 (full book)

Tools and the action space

Take this literally

“The set of tools you expose is the complete inventory of what your agent can ever do.”Chapter 5, opening

Traditional API

Deterministic code calling deterministic code; the caller cannot improvise

Tool

Deterministic code called by a non-deterministic caller: right tool with wrong arguments, wrong tool, invented tool

Worked example · “find thirty minutes for me and the design team”

Wrap the endpoints, or shape the task?

list_users + list_events + create_event: directory and six calendars onto the desk; the model intersects by reading

schedule_event(participants, duration, window): code intersects; the model sees one line back

The right unit for a tool is a task; the endpoints are plumbing

The same scheduling job in two action spaces: one endpoint per tool, where every intermediate byte crosses the model's desk, versus one task-shaped tool that does the deterministic work inside and returns one line.
Figure from Chapter 5

The second reason one-for-one wrapping fails

Too many tools confuse the model

Chapter 5 · “More tools don’t always lead to better outcomes” (vendor guidance quoted there)

The tool interface

“Every one of these channels is a prompt”

  • Description: for a capable new hire; when to use it, when not
  • Names: user invites four readings; user_id one
  • Schema: enums, required, no extra properties
  • Results: the relevant slice; names over UUIDs; say when truncated
The anatomy of a tool across the boundary between your process and the model's context: name, description and schema cross to the model and steer it; the function stays in your process; only results and errors travel back.
Figure from Chapter 5

The error channel is a design surface

The highest-signal message a tool sends

Teaches nothing

Error 422

Corrects itself next pass

severity must be one of: low, med, high (got "urgent")

Lab 2’s structured error (Ch. 18)

{
  "code": "INVALID_ARGUMENT",
  "message": "severity must be one
    of: low, med, high (got 'urgent')",
  "retryable": false,
  "hint": "ask the user which
    severity they mean"
}

Chapter 5, “The Tool Interface” · field shape from Chapter 18 (code, message, retryable, hint)

Retrofitting for agents · 1 of 2

Make explicit what a human used to supply implicitly

Chapter 5 · “Crashes are tolerable; hangs are problematic” (field notes quoted there)

Retrofitting for agents · 2 of 2

Assume every call will be retried

Chapter 5, retrofitting rules for agent-facing programs

In class · teams of three · 20 minutes

Lint seven flawed tools, fix them, re-lint

  1. Open the set in the browser linter: seven flawed tool definitions (also in your handout)
  2. Fix every “Fix first” finding; then the “Should fix” ones
  3. Find two flaws the linter missed; name the chapter rule each breaks
  4. Re-lint; copy the result as Markdown and submit with your fixed JSON
read_logssearchfind_recordscreate_ticketdeleteUserget_customerrun

aiagentsengineered.com/tools/tool-contract-linter/ · runs in the browser · no account, no API key

Code execution as a universal action

“Discretion on the outside, determinism on the inside”

  • One meta-tool, run_code(source), composes all the rest
  • 10,000 rows stay in the runtime; five lines come back
  • Sandbox: own filesystem, capped CPU, memory and time; network denied by default; fresh per session
  • A script that ran proves nothing about what it computed
Code execution as a universal action: the model's judgment stays outside a sealed sandbox; it writes a program that runs against bulk data inside, with the network denied by default; only a small result crosses back.
Figure from Chapter 5

When there is no tool end at all

Computer use: descend only as far as forced

  • API: a deterministic contract
  • Structured browser: the accessibility tree; click by meaning
  • Pixels: universal, slowest, most fragile
  • The page is input written by strangers: indirect prompt injection
Three ways to reach another system drawn as a ladder: API, structured browser through the accessibility tree, and raw pixels; reliability falls and reach widens as you descend.
Figure from Chapter 5

Recap

Five things to keep

Before next week

Lab 2 and reading

Lab 2 · extend Lab 1

  • run_tests: a real verifier
  • search_logs: a task-shaped tool
  • Every error: code, message, retryable, hint
  • dry_run on the mutating tool
  • Scripted-client test: recovery from a malformed argument

Reading for week 4

This week’s reading: Ch. 4 · Ch. 5 (full book) · no paid API needed for the lab

Engineering AI Agents · Week 3

Let the world find; let the model fix.

Slides from the teaching kit of AI Agents, Engineered by Enrique Gutiérrez, CC BY 4.0 · aiagentsengineered.com/teach/

Reuse, adapt and translate with attribution · creativecommons.org/licenses/by/4.0/