Engineering AI Agents · Week 2

The engine’s failure modes,
and the loop

Function calling · six ways the engine fails · compounding error · the loop, its four parts and its exits

Slides CC BY 4.0 · from the teaching kit of AI Agents, Engineered by Enrique Gutiérrez · aiagentsengineered.com/teach/

Where Chapter 2 hands over to Chapter 3

“Part II takes this engine—brilliant, forgetful, overconfident, dice-driven—and puts it in a loop.”

Chapter 2, closing paragraph

By the end of today you can

Four objectives

  1. Describe the function-calling exchange and point to the one step where side effects happen
  2. List the six engine failure modes and name the engineering remedy for each
  3. Derive compounding error, pⁿ, and its three levers
  4. Implement the four-part minimal agent with a scripted model client and a mandatory budget

Part 1 · Chapter 2, sections 4–6 (free)

From text to action

Prompting fundamentals, in one slide

In an agent, no one types the final prompt

The anatomy of a chat prompt: standing instructions in the system message, the current request in the user message, the model's own answers accumulating as assistant messages; the software resends the whole stack on every call.
Figure from Chapter 2

Sketch the parser before the next slide

Three runs, three correct answers

Name: Ada, Total: $42
The customer is Ada Lovelace and she spent forty-two dollars.
Happy to help! Ada's order total is $42.00 (invoice #1041).

Chapter 2, “Structured Output and Function Calling”

Two grades of guarantee that sound alike

Asking vs constrained decoding

Weaker · “JSON mode”

It parses

A field can still be missing, invented, or a string where the number goes

Stronger · constrained decoding

It fits the schema

The schema becomes a filter that strikes invalid tokens before every draw
“A schema, misused, makes fabrication easier to consume, not rarer.”Chapter 2 · a guarantee about shape says nothing about truth

Function calling · where does anything happen?

The model proposes; your code disposes

Function calling, step by step, in two lanes: your code sends the question and the list of tools; the model replies with a structured request to call one; your code runs the real function, step 3, the only place a side effect happens; the result goes back to the model, which answers or asks again.
Figure from Chapter 2 · schematic; wire formats vary by provider

All the model will ever know about a tool

Name, description, argument schema

name:        create_ticket
description: Open a support ticket.
             Use when the user reports
             a problem that needs human
             follow-up.
arguments:   { title:    string,
               severity: one of
                 "low" | "med" | "high",
               body:     string }
  • Misused tool? Suspect the description before the model
  • Expect zero, one or many calls per turn
  • Bad call: return the error as a result, cap the retries
  • “A flawlessly schema-conformant call to delete_account still deletes the account.”

Chapter 2 · pseudocode, as in the book; each vendor’s syntax differs

Failure mode 1 · the master symptom

Hallucination: fluent, confident, wrong

“language models are optimized to be good test-takers, and guessing when uncertain improves test performance.”Kalai et al., “Why Language Models Hallucinate” (2025), quoted in Chapter 2

Failure mode 2 · the jagged frontier

−19 points

accuracy for consultants with model help, on one task chosen to sit just outside the model’s competence

Dell’Acqua et al., HBS Working Paper 24-013 (2023), 758 consultants, preregistered; as reported in Chapter 2

  • Inside the frontier, the same group finished more tasks, faster, at higher rated quality
  • The failing task did not look harder
  • Competence next door proves nothing about competence here
  • Remedy: “test your model on your task”

Respond with engineering, not sterner prompts

Six failure modes, six remedies

Failure modeWhat it looks likeEngineering remedy
HallucinationConfident invented factsGround in sources; verify outside the model
Jagged frontierFails an easy-looking taskTest your model on your task
Phrasing sensitivityReworded prompt, new answerVersion and re-test prompts like code
SycophancyCaves when you push backA real verifier, not “are you sure?”
Exact computationArithmetic and counting wobbleRoute exact work to code or a calculator
Instructions = dataFetched text steers it: prompt injectionDefend in the harness (Ch. 17)

Chapter 2, “Limitations and Failure Modes” · properties of current models; verify where yours stand

The master entry, now defended

Reliability multiplies: pⁿ

0.9520 ≈ 36%

0.9230 ≈ 8%

  • Independence is the optimistic case
  • Errors land in the transcript and condition every later step
Compounding error: whole-run success decays exponentially with the number of chained steps; 95% per step is about 36% across twenty steps, and 92% per step about 8% across thirty.
Figure from Chapter 2 · numbers illustrative

“Per-step accuracy is a ceiling, not a forecast.” · Chapter 2

Three levers

1 · Shrink n

Fewer dice

Fuse steps, precompute, push bulk work into code

0.9510 ≈ 60%

2 · Raise effective p

Verify and recover

An error costs a retry, not the run

(1 − 0.05²)20 ≈ 95%

3 · Cut the price of dying

Checkpoint

A failed run resumes from the last good state, not from zero

same odds, cheaper failure

Levers from Chapter 2 · the worked numbers are illustrative algebra from a 95% × 20-step baseline (36%), assuming a check that always catches a failed step

In class · pairs · 15 minutes

How reliable must each step be?

  1. Watch: Why errors compound (3 min explainer)
  2. Open aiagentsengineered.com/tools/compounding-error-calculator/
  3. Target 90% end to end over 20 steps: find p*, the per-step reliability you need
  4. At 95% per step, what is the longest chain that still meets 90%?
  5. Pick the cheapest lever for a task of your choice, and defend it in two sentences

Runs in the browser · no account, no API key · copy the result as Markdown to submit

Part 2 · Chapter 3 and Appendix A (full book)

The loop

Observe · reason · act · observe the result

Four beats to a pass

One pass through the loop has four beats circling a growing history; the model supplies only the reason beat, and your code supplies the other three: lay the desk, run the tool, record the result.
Figure from Chapter 3 · explainer: the agent loop in three minutes

The index card · Chapter 3, “The Loop”

history = [standing_instructions, user_goal]

repeat up to MAX_STEPS times:
    response = model(history)              # reason over the desk
    if response is a final answer:
        return response.text               # stop: the model judges it done
    result = execute(response.tool_call)   # act: your code runs the tool
    append response.tool_call to history   # record what the model did
    append result to history               # record what the world said back

return "stopped: step budget exhausted"    # stop: the safety net

ReAct: interleaving reasoning and acting

Why one small step at a time?

  • Plan up front, then execute blindly: “all of its judgment at the point of maximum ignorance”
  • Reasoning alone: a sealed room; it elaborates a fabrication
  • Acting alone: a sleepwalker; it repeats gestures
  • Interleave: thought, action, observation. Grounding “is a schedule”
Three schedules for a model with tools: reasoning alone drifts from the world into fabrication; acting alone repeats a gesture that no longer fits; interleaving re-anchors the reasoning to an observation at every step.
Figure from Chapter 3 · Yao et al., ReAct, ICLR 2023

A minimal agent from scratch

Four parts inside a harness

  1. Model client: sends messages and tool definitions; stateless
  2. Tools: name, description, schema, and the function that runs
  3. Message history: the run’s entire memory
  4. The loop: ties the three together
A minimal agent has exactly four parts, all inside the harness frame: the model client, the tools, the history and the loop; the model itself sits outside the frame.
Figure from Chapter 3
“An agent, in one line, is a model plus a harness.”Chapter 3, “A Minimal Agent from Scratch”

Worked example · Appendix A

Under seventy lines, no magic

function execute(call):
  if call.name not in TOOLS:
    return result(call.id,
      error: "no such tool: " + call.name)
  outcome = try TOOLS[call.name].fn(call.arguments)
  if the call failed:
    return result(call.id,
      error: plain_description(outcome))
  return result(call.id, outcome)
A minimal agent in one picture: user input is appended to a growing message history; the model reads it and either answers or asks for a tool; the tool's result is appended back and control returns to the model; that return arrow is the loop.
Figure from Appendix A

Appendix A · goal: change ‘Hello’ to ‘Welcome’

Watch the history grow

PassThe model asks forThe result appended
1list_files()README.md, app.py, greetings.py, tests/
2read_file("greetings.py")GREETING = 'Hello' …
3edit_file(…'Hello' → 'Welcome')ok: 1 replacement made
4read_file("greetings.py")GREETING = 'Welcome' …
5no tool calls: “Done. …”the loop returns finished

“Your program is dumb on purpose; the transcript only looks smart.”

Stop conditions and budgets

Three exits, ranked by trust

  • Verified check: suite green, file parses. “The model’s ‘done’ is testimony; a green test is evidence.”
  • Budget: steps, tokens, time, money; mandatory, set first. Fires? Fail loudly, save the transcript
  • Errors: can the model act on it? Feed it back. Evidence the run is off the rails? Halt
The loop's three exits ranked by trust: a verified check is evidence and should pronounce the run finished; the model's own done is testimony, believed with caution; the budget cap is the safety net that must always be there.
Figure from Chapter 3

How you will test the harness · Chapter 15

Replace the model with a script

SCRIPT = [ call list_files(),
           call read_file("greetings.py"),
           call edit_file(...),
           call read_file("greetings.py"),
           say  "Done. GREETING now 'Welcome'." ]

function scripted_send(request):
    # replaces provider.send
    return next response in SCRIPT
  • The loop, the tools, the stop rules are ordinary code: exact assertions
  • Test 1: a stub that never says done → stopped at MAX_STEPS
  • Test 2: drop an append → the test must catch it
  • Fast, free, deterministic; no API key

Lab pseudocode, not from the book; the testing approach is Chapter 15’s

The day-one question

Shouldn’t I just use a framework?

  • What a framework adds is plumbing: persistence and resume, retries, tracing, memory management, multi-agent orchestration
  • Every item is a harness part; none changes a single pass
  • So it “can make an agent more dependable and cannot make it smarter”
“Build from scratch to understand. Adopt a framework to scale—and only once you can name what it is saving you.”Chapter 3, “Stop Conditions, Budgets, and Frameworks”

Recap

Five things to keep

Before next week

Lab 1 and reading

Lab 1: the minimal agent, no API

  • Type in Appendix A in your language
  • Swap provider.send for a scripted model client replaying the Hello → Welcome run
  • Three tools, errors as results, MAX_STEPS with a loud stopped
  • Two tests that break the loop on purpose

Reading for week 3

This week’s reading: Ch. 2 §4–6 (free) · Ch. 3 · App. A (full book)

Engineering AI Agents · Week 2

The model’s “done” is testimony. What is your evidence?

Slides from the teaching kit of AI Agents, Engineered by Enrique Gutiérrez, CC BY 4.0 · aiagentsengineered.com/teach/

Reuse, adapt and translate with attribution · creativecommons.org/licenses/by/4.0/