Home / Teach

For instructors · CC BY 4.0

Teach AI agents: a free, vendor-neutral course kit

A 13-week AI agents course syllabus for software engineering students, with browser-only exercises, explainers and the book's diagrams, free under CC BY 4.0.

Maintained by Enrique Gutiérrez · Last reviewed

Who is this kit for?

This kit is for instructors who want to teach engineering AI agents to software engineering students without depending on a vendor, a framework or a paid API. It is built on AI Agents, Engineered and follows its thesis: an agent is only as trustworthy as the signal you can use to verify it. Students are graded on the eval sets, traces and gates they ship, not on demos.

Many university courses on agents state that there is no required textbook (the syllabi of Michigan’s EECS 498 and Texas A&M’s ECEN 689 are two examples), and most free agent courses are paired with a particular framework or provider. This kit takes the opposite position. It names concepts and patterns, mentions tools only as labeled examples, and lets every lab run against a scripted model client so that a student with no credits is never at a disadvantage.

What is in the kit?

The kit has four parts, all free and all usable in a browser:

Part What it gives you Where
Syllabus 13 weeks mapped to the book’s 27 chapters, with objectives, readings, in-class exercises and labs Below on this page
Browser tools Calculators that compute what the chapters teach: compounding error, agent cost per task, eval sample sizes Tools
Explainers Short animated explainers of the core ideas, with transcripts Learn
Diagrams The book’s diagrams as SVG with captions, ready for slides Diagrams

Lecture decks with instructor notes are being added week by week and appear in this page’s list as they are published.

How does the course run?

The course runs as one lecture and one lab a week for thirteen weeks. The first two weeks cover what an agent is, how a language model works, and why errors compound across a chain of steps; both weeks can be taught entirely from the free chapters. The middle weeks build the loop, tools, context, memory and multi-agent patterns. The last third is the part most courses skip: observability, evaluation, security, reliability, cost and deployment.

Assessment follows the same emphasis. Labs carry 30 percent, a verification-centered project 40 percent, an exam 20 percent and reading responses 10 percent. In the project, a team builds a small agent and must ship an eval set, a trace review and a regression gate alongside it; a project without evidence that it works does not pass, however good the demo looks.

Three in-class exercises show the style:

  1. Week 2. Students estimate a per-step success rate for a ten-step task, then use the compounding error calculator to see what the whole run gives them, and redesign the task until it clears a target.
  2. Week 8. Teams score their agent on fifty tasks, then use the eval sample-size calculator to decide whether a prompt change actually moved the pass rate.
  3. Week 11. Students price their own agent run with the agent cost-per-task estimator and find the step at which the re-read history overtakes the fixed prompt.

How do I use the book with my students?

Start with the free chapters. The preface, Chapters 1 and 2 and the glossary are open to anyone at the online reader, so students can begin the course before deciding whether to buy the book. For the remaining chapters, the book is sold as PDF and EPUB on Leanpub, on Kindle and in paperback; the formats section on the home page has the links.

Each week below lists the chapters to read and flags the free ones. The readings are short by design: one or two chapters a week, with the lab doing the rest of the teaching.

Lecture decks

  1. Week 1: What an Agent Is, and the Engine Underneath ch00, ch01, ch02
  2. Week 2: The Engine’s Failure Modes and the Loop ch02, ch03, appA
  3. Week 3: Planning, Tools, and the Action Space ch04, ch05

The syllabus

A one-semester course for upper-undergraduate and master’s students in software engineering, built on AI Agents, Engineered (27 chapters + Appendix A, “A Minimal Agent, Annotated”). No machine-learning background is assumed; the book explains the engine from first principles (Chapter 2) and never requires training a model.

License for this syllabus: CC BY 4.0—adapt freely with attribution to the book.

Design principles

  • Verification-centered. The book’s compass bearing—what signal tells you it worked? (Ch. 1)—is the course’s grading philosophy: students are assessed on the eval sets, traces, and gates they ship, not on demo videos.
  • Time- and solution-agnostic, like the book. Lectures name concepts and patterns; vendors and models appear only as labeled examples. No assignment depends on a particular provider, and no assignment requires a paid API: labs run against a scripted model client (a stub that returns a fixed sequence of responses, the “deterministic–nondeterministic seam” of Ch. 15) and, optionally, any locally runnable open-weight model as a category. Students who have credits may use a hosted model; grading never depends on it.
  • Browser-only in class. In-class exercises use the companion browser tools and animated explainers (tools, explainers), which run without a server or an account.
  • Free reading flagged. The preface, Part I (Ch. 1–2) and the glossary (App. B) are free online; other chapters are in the book (PDF/EPUB/paperback). Weeks 1–2 are fully coverable with the free material.

Prerequisites

  • One programming language with comfort in functions, dictionaries/maps, HTTP clients, and JSON (the book’s pseudocode assumes exactly this: “If you can read a function definition and a dictionary, you can read this”, App. A).
  • Introductory probability (independent events, expected value)—enough to follow pⁿ, pass@k and a confidence interval.
  • Basic software-engineering practice: version control, unit tests, CI.
  • No ML, no linear algebra, no training of models.

Course-level outcomes

By the end of the course a student can:

  1. Distinguish chatbot, workflow and agent by who owns the control flow, and place a problem on the escalation ladder with the four-line cost quote (Ch. 1, 14).
  2. Explain, from first principles, why an LLM is a next-token predictor whose failure modes (hallucination, sycophancy, nondeterminism, prompt injection) are properties of the engine, and derive the compounding-error arithmetic (Ch. 2).
  3. Implement the minimal agent (four parts, one loop) with structured tool calls, fed-back errors, and mandatory budgets, against a scripted model client (Ch. 3, App. A).
  4. Design task-shaped tools and skills, and manage the context window with the four operations (Ch. 5–9).
  5. Choose and justify workflow patterns, multi-agent decomposition, consequence-tier gates and an autonomy-dial position for a given task (Ch. 10–14).
  6. Instrument an agent with traces and spans, reproduce a failure by record-and- replay, and build an eval set with a calibrated judge and a regression gate (Ch. 15–16).
  7. Run the lethal-trifecta audit, set the blast-radius dials, and engineer retries, idempotency, checkpoints and a harness grid (Ch. 17–18).
  8. Estimate cost and latency, design a serving shape with backpressure, and plan a rollout ladder (Ch. 19–20).
  9. Apply the book’s transversal recipes and build an LLM classifier with the three-set discipline (Ch. 21–27).

Assessment scheme

Component Weight What is graded
Labs (8, best 7 count) 30% Working code + a one-page “what signal tells you it worked?” note per lab
Verification project (team of 2–3) 40% See the capstone brief: an agent plus an eval set, a trace review, a security audit and a rollout plan
Final exam 20% Concepts, arithmetic (pⁿ, cost staircase, pass@k, kappa), design decisions with justification
Reading & in-class tool exercises 10% Short pre-class quizzes on the reading; in-class exercise submissions (tool “Copy as Markdown” exports)

Grading rule inherited from the book: a project is graded on passᵏ, not pass@k (Ch. 16). A single impressive run earns nothing; the eval set’s numbers, read as a distribution over several runs, do.


Weekly plan

Format per week: Title · chapters · learning objectives (measurable) · key concepts (glossary terms) · figures · in-class exercise (browser tool / explainer) · homework or lab · reading (free chapters flagged ★).

Week 1: What an Agent Is, and the Engine Underneath

  • Chapters: Preface ★, Ch. 1 ★ What Is an Agent?, Ch. 2 ★ The Engine (first three sections).
  • Objectives: (1) Classify five described systems as chatbot / workflow / agent and justify each by who owns the control flow; (2) state the four compass bearings and give one design fork each decides; (3) explain tokens, the context window as a finite desk, and why two identical runs differ; (4) compute pⁿ for given p, n.
  • Key concepts: agent, workflow, augmented LLM, autonomy dial, token, context window / the desk, autoregressive generation, compass, the.
  • Figures: fig-ch01-chatbot-workflow-agent, fig-ch01-compass-four-bearings, fig-ch01-autonomy-dial, fig-ch02-context-window-desk, fig-ch02-weighted-lottery.
  • In-class: watch explainer what-is-an-agent; then, in pairs, run should-this-be-an-agent on three requests of the students’ own and submit the Markdown export.
  • Homework: Reading quiz; “tokenizer experiment”—count letters / spell a word backwards with any available model and explain the outcome using subword tokenization (Ch. 2 “Tokens and the Context Window”). No paid API needed (any free chat interface or local model).
  • Reading: Preface ★, Ch. 1 ★, Ch. 2 ★ §1–3.

Week 2: The Engine’s Failure Modes and the Loop

  • Chapters: Ch. 2 ★ (remaining sections), Ch. 3 The Agent Loop, App. A A Minimal Agent, Annotated.
  • Objectives: (1) Describe the function-calling exchange and locate the one step where side effects happen; (2) list the six engine failure modes and name the engineering remedy for each; (3) derive compounding error and its three levers;
    1. implement the four-part minimal agent with a scripted model client and mandatory budgets.
  • Key concepts: structured output, constrained decoding, tool call / function calling, hallucination, harness, message history, model client, ReAct, grounding, stop condition, budget.
  • Figures: fig-ch02-function-calling-exchange, fig-ch02-compounding-error-curve, fig-ch03-agent-loop-cycle, fig-ch03-react-interleaving, fig-ch03-minimal-agent-four-parts, fig-ch03-three-exits, fig-app-a-minimal-agent-anatomy.
  • In-class: explainers why-errors-compound and the-agent-loop; then compounding-error-calculator: find p* for a 90% end-to-end target over 20 steps and discuss which lever is cheapest.
  • Lab 1 (local code, no API): Type in Appendix A’s program in your language. Replace provider.send with a scripted model client that returns a fixed sequence of responses (list_files → read_file → edit_file → read_file → final answer) for the “Hello → Welcome” run. Requirements: the three tools, execute turning every failure into a result, MAX_STEPS with a loud stopped outcome, and two tests that break the loop on purpose (drop an append; a model stub that never says done) and assert the harness behaves (Ch. 3 §“Stop Conditions”; Ch. 15 §“the seam”). Optional: run the same code against a local open-weight model.
  • Reading: Ch. 2 ★ §4–6, Ch. 3, App. A.

Week 3: Planning, Tools, and the Action Space

  • Chapters: Ch. 4 Planning, Reasoning, and Self-Correction, Ch. 5 Tools and the Action Space.
  • Objectives: (1) Contrast plan-then-execute with interleaved planning and name when each wins; (2) explain why “a real verifier beats the model grading itself”;
    1. redesign a one-for-one API wrapping into a task-shaped tool; (4) write a tool description, schema and error message that pass the tool-contract linter.
  • Key concepts: chain-of-thought, oracle, verifier vs self-judgment, action space, tool, task-shaped tools, code execution as a universal action, sandbox, computer use.
  • Figures: fig-ch04-plan-vs-interleaved, fig-ch04-self-consistency-vote, fig-ch04-verifier-vs-self-judgment, fig-ch05-task-shaped-tools, fig-ch05-tool-interface-anatomy, fig-ch05-code-sandbox, fig-ch05-ui-reliability-ladder.
  • In-class: tool-contract-linter on a provided JSON of seven deliberately flawed tool definitions; teams fix the findings and re-lint.
  • Lab 2: Extend Lab 1 with run_tests (a real verifier) and a search_logs-style task-shaped tool; make every error structured (code, message, retryable, hint); add --dry-run semantics on the mutating tool. Scripted-client tests must show the loop recovering from a malformed argument via the fed-back error.
  • Reading: Ch. 4, Ch. 5.

Week 4: Skills, Protocols, and Context Engineering

  • Chapters: Ch. 6 Skills, Protocols, and Interoperability, Ch. 7 Managing the Context Window.
  • Objectives: (1) Distinguish tool / protocol / skill and diagnose which layer a given gap belongs to; (2) compute the M×N → M+N collapse and state three costs of protocols; (3) name the five context failure modes from symptoms; (4) apply the four operations to a mid-run desk inventory.
  • Key concepts: skill, progressive disclosure, protocol, context engineering, context rot, dumb zone, write / select / compress / isolate.
  • Figures: fig-ch06-skills-tools-protocols, fig-ch06-progressive-disclosure, fig-ch06-nxm-protocol, fig-ch06-skill-patterns, fig-ch07-context-assembly, fig-ch07-attention-lamp, fig-ch07-five-failure-modes, fig-ch07-four-operations.
  • In-class: explainer context-rot-and-the-four-operations; then context-window-budget-planner with Ch. 2’s sketch (500 / 8,000 / 40,000 / 3,000) and a growth of 1,000 per step—report steps to the illustrative 40% ceiling and one operation per component. Quick skill-token-budget run for 10 vs 50 skills.
  • Lab 3: Package one procedure as a skill folder (name + description, body under a few hundred lines, one reference file, one script executed without being read). Add tool-result trimming and a progress file (write) to the Lab 2 agent; show with the scripted client that the history stays bounded over a 30-step run.
  • Reading: Ch. 6, Ch. 7.

Week 5: Retrieval, Memory, and Workflow Patterns

  • Chapters: Ch. 8 Retrieval and Knowledge, Ch. 9 Memory, Ch. 10 Workflows and Composition Patterns.
  • Objectives: (1) Describe the RAG pipeline and two ways beyond naive retrieval;
    1. sort memories into episodic / semantic / procedural and state the quarantine rule’s two doors; (3) choose among chaining, routing, parallelization and the evaluator–optimizer loop for given tasks and state each pattern’s failure mode.
  • Key concepts: RAG, chunk, embedding, vector database, fine-tuning, memory (episodic / semantic / procedural), standing project-instructions file / constitution, compaction, quarantine, routing, evaluator–optimizer.
  • Figures: fig-ch08-rag-pipeline, fig-ch08-three-ways-to-know, fig-ch09-memory-tiers, fig-ch09-quarantine-two-doors, fig-ch10-prompt-chain-gates, fig-ch10-routing-paths, fig-ch10-evaluator-optimizer-loop.
  • In-class: whiteboard a support pipeline as a prompt chain with gates, then as a router; identify the single step that genuinely needs a loop. Glossary scavenger hunt (free App. B) for the week’s twelve terms.
  • Lab 4 (browser or local): Build a tiny keyword + “semantic” (any local embedding library, or TF-IDF as the solution-agnostic stand-in) retriever over a folder of course notes; wire it as a select tool into the agent; write a constitution file and a compaction routine; demonstrate quarantine by delegating a noisy sub-question to a second scripted agent that returns a one-paragraph memo.
  • Reading: Ch. 8, Ch. 9, Ch. 10.

Week 6: Multi-Agent, Oversight, the Outer Loop, and When Not To

  • Chapters: Ch. 11 Multi-Agent Systems, Ch. 12 Oversight and Autonomy, Ch. 13 Writing the Outer Loop, Ch. 14 Choosing Your Approach.
  • Objectives: (1) Decide single vs multi-agent with the decomposition question and write a worker brief with the four required fields; (2) classify an agent’s tools into the four consequence tiers and write the gate policy; (3) specify a goal function (persistent goal + machine-checkable stop) with a hard cap and a kill switch; (4) price a request on the ladder with the four-line quote.
  • Key concepts: orchestrator–worker, subagent, approval gate, review theater, autonomy dial, outer loop / inner loop, goal function, oracle, kill switch, agent-washing.
  • Figures: fig-ch11-orchestrator-worker, fig-ch11-worker-memo-desk, fig-ch11-merge-wall, fig-ch12-consequence-tiers, fig-ch12-autonomy-dial, fig-ch13-outer-loop-verbs, fig-ch13-loop-oracle-gauge, fig-ch14-ladder-revisited.
  • In-class: explainers consequence-tiers and orchestrator-worker-and-the-merge- wall; consequence-tier-classifier on each team’s project tool list; mast-explorer brief-linter on a one-line delegation, then on a corrected one.
  • Lab 5: Add a code-side approval gate between tool selection and invocation for Tier 3–4 tools, with the handoff package (action, why, touches/undoable, before→after, “reject with edits”). Write a minimal outer loop (find / assign / check / record / decide) that runs the inner agent against a checklist on disk with a machine-checkable stop and an iteration cap. Project proposal due: the problem, its rung on the ladder, and the verification signal.
  • Reading: Ch. 11, Ch. 12, Ch. 13, Ch. 14.

Week 7: Observability and Debugging

  • Chapters: Ch. 15 Observability and Debugging.
  • Objectives: (1) Instrument an agent so every model call and tool call is a span carrying the rendered prompt, arguments, results, tokens, latency and the responding model; (2) implement record-and-replay across the seam and replay a recorded failure; (3) diagnose the six bug specimens from traces; (4) turn a fixed failure into an eval case.
  • Key concepts: trace, span, transcript / trajectory, record and replay, the deterministic–nondeterministic seam, drift, golden traces, the trace-to-dataset flywheel.
  • Figures: fig-ch15-flight-recorder, fig-ch15-trace-span-waterfall, fig-ch15-drift-monitoring, fig-ch15-record-replay, fig-ch15-trace-to-dataset-flywheel.
  • In-class: trace-waterfall-viewer on three provided traces (a stuck loop, a silent truncation, a requested≠responding model); students find the first wrong span, not the last. agent-bug-bestiary wizard on a fourth, unlabeled trace.
  • Lab 6: Add a minimal tracer (wrap the client, wrap the dispatcher, parent links, session summary span) and a record/replay layer that fails loudly on anything not in the recording. Record five runs, replay one with a deliberately changed parser, and show the structural diff. Store result byte counts per tool span. Deliver one “golden trace” per project tool.
  • Reading: Ch. 15.

Week 8: Evaluating Agents

  • Chapters: Ch. 16 Evaluating Agents.
  • Objectives: (1) Write testable acceptance criteria and a 20–50 task eval set grown from observed failures, every task passable with a reference; (2) compute and interpret pass@k, passᵏ and a confidence interval for a pass rate; (3) build and calibrate an LLM-as-a-judge (rubric, reason-first, reference-anchored) with Cohen’s kappa against hand labels and test for position and verbosity bias; (4) wire a regression gate with capability and regression suites.
  • Key concepts: eval, eval set, pass@k and passᵏ, reliability envelope, LLM-as-a-judge, regression gate, capability vs regression evals, saturation, offline vs online evaluation.
  • Figures: fig-ch16-reliability-envelope, fig-ch16-three-runs-ritual, fig-ch16-judge-grades-homework, fig-ch16-regression-gate.
  • In-class: explainer pass-at-k-vs-pass-hat-k; pass-at-k-plotter for each team’s measured p; eval-sample-size-calculator to size the project’s suite (and to see why a two-point change on fifty tasks is noise); judge-agreement-calculator on a provided 60-label sample.
  • Lab 7 (project milestone): The project eval set v1: 20–50 tasks with reference solutions and graders on the lowest rung that works (programmatic → binary rubric → judge); run each task ≥ 3 times with the scripted client (or a local model) and report pass rates with intervals; calibrate any judge on ≥ 30 hand labels and report kappa, precision/recall and level bias; add negative cases; keep infrastructure failures out of the quality number.
  • Reading: Ch. 16.

Week 9: Security, Safety, and Guardrails

  • Chapters: Ch. 17 Security, Safety, and Guardrails.
  • Objectives: (1) Explain direct vs indirect prompt injection and why every model-layer defense is statistical; (2) run the lethal-trifecta audit and the two-legs-plus-sharp-tool audit on a design and propose the cheapest amputation;
    1. place guardrails in the four families and justify spending most of the budget on action guards; (4) set the three blast-radius dials and audit a skill or tool server as a supply-chain dependency.
  • Key concepts: prompt injection, confused deputy, lethal trifecta, guardrail (input / output / action / structural), fail closed, blast radius, least privilege, sandbox, approval fatigue, tool poisoning, rug pull.
  • Figures: fig-ch17-confused-deputy, fig-ch17-lethal-trifecta, fig-ch17-guardrail-families, fig-ch17-blast-radius, fig-ch17-tool-poisoning.
  • In-class: explainers the-lethal-trifecta and prompt-injection; lethal-trifecta-audit on each project; tool-contract-linter supply-chain pass on a provided tool server description seeded with a poisoned instruction.
  • Homework (red team): Each team writes three indirect-injection payloads for another team’s agent (hidden in a document the agent reads) and a one-page containment review: which leg to cut, which action guard to add, which dial to turn. Run with the scripted client: the payload’s effect is simulated by the stub honoring the injected instruction, so the exercise tests the harness, exactly as the chapter argues.
  • Reading: Ch. 17.

Week 10: Reliability, State, and the Harness

  • Chapters: Ch. 18 Reliability, State, and the Harness.
  • Objectives: (1) Classify failures (transient, model-recoverable, permanent, policy, ambiguous) and route each to one primitive; (2) implement backoff with jitter, a global retry budget, timeouts and a circuit breaker; (3) make side-effecting tools idempotent with keys and receipts and checkpoint a run so it survives a kill; (4) fill the harness grid for a project and apply the scaling law.
  • Key concepts: idempotency, checkpoint, durable execution, saga / compensation, circuit breaker, fallback ladder, graceful degradation, guides vs sensors, computational vs inferential, harness.
  • Figures: fig-ch18-error-surfacing, fig-ch18-circuit-breaker, fig-ch18-receipts, fig-ch18-harness-grid, fig-ch18-two-rooms.
  • In-class: circuit-breaker-simulator (book defaults; nested 3×3×3×3 = 81); harness-grid-builder for each project with the three scaling dials.
  • Lab 8: “Crash on purpose”: add checkpoints keyed by run id, idempotency keys and receipts on the mutating tool; kill the process after the side effect and before the receipt, restore, and prove the action does not repeat. Add a breaker around one flaky scripted tool and a fallback ladder with a “marked as stale” rung. Tests run with the scripted client at zero cost.
  • Reading: Ch. 18.

Week 11: Cost, Latency, Deployment, and Scaling

  • Chapters: Ch. 19 Cost, Latency, and Performance, Ch. 20 Deploying and Scaling.
  • Objectives: (1) Derive the input-token staircase F·n + g·n(n−1)/2 and identify the five cost levers; (2) separate time-to-first-token from completion time and choose streaming vs parallelism vs caching for a given complaint; (3) design the async job pattern with task states, idempotent submission, polling and backpressure; (4) write a rollout plan with shadow, canary, automated rollback and an immutable artifact version.
  • Key concepts: token economics, prompt caching, the accuracy–cost–latency triangle, prefill vs decode, the thirty-second wall, durable execution at scale, backpressure, shadow deploy, canary, feature flag.
  • Figures: fig-ch19-cost-compounding-steps, fig-ch19-prompt-caching-prefix, fig-ch19-streaming-perceived-wait, fig-ch19-accuracy-cost-latency-triangle, fig-ch20-thirty-second-wall, fig-ch20-sync-vs-async-serving, fig-ch20-backpressure-queue, fig-ch20-rollout-ladder.
  • In-class: explainers why-agent-cost-compounds and the-rollout-ladder; agent-cost-per-task-estimator with each team’s measured F, g, n from their traces; prompt-caching-savings (move the timestamp); queue-backpressure- simulator (runs/min ceiling from tokens per run); rollout-plan-generator.
  • Homework: From the project’s traces, compute cost per run, latency percentiles and step-count distribution; set the step budget from the distribution plus margin; write the one-paragraph triangle statement (“which corner this feature sacrifices, how far it may slide, which meter triggers a rethink”, Ch. 19).
  • Reading: Ch. 19, Ch. 20.

Week 12: Coding Agents and Classifiers

  • Chapters: Ch. 21 Coding Agents, Ch. 22 The Coding Workflow in Practice, Ch. 26 Agents as Classifiers and Scorers.
  • Objectives: (1) Explain why coding led (“the domain grades its own work”) and transfer the ground-truth-loop template to another domain; (2) run the brainstorm → spec → plan → execute loop with TDD/TCR and the compound step; (3) build an LLM classifier from written definitions (essence, inclusions, exclusions, boundary rules), reason-first rubric, abstain flag, and an ordinal scorer; (4) develop it with the three-set discipline and report kappa and the confusion pairs.
  • Key concepts: vibe coding vs augmented coding, agent-legible code, spec-driven development, TCR, compound step, classifier, ordinal scale, three-set discipline, in-context learning.
  • Figures: fig-ch21-ground-truth-loop, fig-ch21-augmented-vibe-spectrum, fig-ch22-brainstorm-spec-plan-execute, fig-ch22-tdd-tcr-guardrail, fig-ch26-examiner-marking-guide, fig-ch26-worksheet-before-verdict, fig-ch26-three-sets-sealed-envelope, fig-ch26-screen-then-detail.
  • In-class: three-set-splitter on a provided 200-ticket labeled CSV (50/50, demos per class, disjointness check); judge-agreement-calculator on the dev set after one guide revision; cascade-cost-estimator at 50 vs 50,000 items/day.
  • Lab (optional, replaces lowest lab grade): Build the ticket-triage classifier as a skill folder (GUIDE, examples, schema, CHANGELOG); grade it on the sealed test set once; log every revision with before/after dev scores. Graders may use the scripted client with a frozen rubric-to-label mapping, or a local model.
  • Reading: Ch. 21, Ch. 22, Ch. 26.

Week 13: Research, Business, UX, Recipes, the Frontier—and Capstone Demos

  • Chapters: Ch. 23 Research and Business Agents, Ch. 24 Agent UX and Human Trust, Ch. 25 Transversal Recipes, Ch. 27 The Frontier and How to Keep Learning.
  • Objectives: (1) Explain the verification gap and the four partial substitutes;
    1. frame ROI against risk with the cost of checking subtracted; (3) place a use case on the autonomy slider and pick its recipe; (4) describe appropriate reliance and a supervision surface; (5) name the durable layers of the ecosystem and the generator-and-gate pattern.
  • Key concepts: verification gap, grounding vs truth, the privileged-employee rule, appropriate reliance, supervision surface, autonomy slider, human in the loop / on the loop, transversal recipe, property-based testing as a gate, portable core.
  • Figures: fig-ch23-no-compiler-for-facts, fig-ch23-partial-substitutes, fig-ch23-roi-vs-risk, fig-ch24-calibration-line, fig-ch25-recipes-slider, fig-ch27-generator-and-gate, fig-ch27-durable-layers.
  • In-class: roi-vs-risk and recipe-picker on each project; agent-verifiability-scorecard as the team’s own pre-demo self-assessment. Capstone demonstrations (see brief): the demo shows the eval dashboard and a trace review, not a single lucky run.
  • Reading: Ch. 23, Ch. 24, Ch. 25, Ch. 27 (App. C for further reading).

Capstone brief: the verification-centered project

Build an agent that you can prove works, and show exactly how far. Teams of 2–3. The agent may be modest; the verification must be real.

Deliverables (all required; weights within the 40% in brackets):

  1. Design note (5%)—the problem, its rung on the ladder with the four-line quote (Ch. 14), the verification signal (Ch. 1), the consequence-tier table and dial position (Ch. 12), and the decomposition verdict (single vs multi-agent, Ch. 11).
  2. Working agent (8%)—four parts, structured tools with structured errors, budgets, code-side gates, idempotent side effects with receipts, checkpoints (Ch. 3, 5, 12, 18). Runs end-to-end against the scripted client; optionally against a local or hosted model.
  3. Eval set (10%)—20–50 hand-checked tasks grown from observed failures, each with a reference solution; graders on the lowest rung that works; negative cases; ≥ 3 runs per task; pass rates with intervals and the reliability envelope; any judge calibrated (≥ 30 labels, kappa, precision/recall, bias checks); a regression gate in CI with capability and regression suites (Ch. 16).
  4. Trace review (7%)—full tracing with the span inventory; a written review of the worst ten runs sorted by cost or step count, each diagnosed by the first wrong span and mapped to the bug bestiary; two golden traces; one recorded failure replayed and fixed, promoted to an eval case (Ch. 15).
  5. Security audit (5%)—lethal-trifecta and two-legs audits, the guardrail families present, the three dials, and a supply-chain review of every tool/skill (Ch. 17). Includes the red-team results from Week 9.
  6. Cost, latency and rollout (5%)—measured cost per run and the staircase coefficients F, g; latency percentiles; step budget set from the distribution; the triangle statement; a rollout plan with shadow/canary/rollback and an immutable artifact version (Ch. 19–20).

Demo day rules: no cherry-picked runs. The demo opens on the eval dashboard, shows a trace, and ends with the verifiability scorecard. Grading follows passᵏ.

Suggested project shapes (each a Ch. 25 recipe, scoped down): a queue triage and router over a synthetic ticket stream; an unstructured-to-structured extractor with field-level citations and a confidence-gated lane; a deep-research analyst over a fixed local corpus with openable citations; a digital coworker running one SOP against mock systems; a coding agent that makes failing tests pass in a toy repository (the book’s canonical agent problem).

Adopt the book

  • The chapter structure maps cleanly onto a 13-week term (above) or onto two quarters (Ch. 1–14 “Building”, Ch. 15–27 “Making it Dependable and Applying it”). A six-week professional short course can run Weeks 1–2, 7–9, and 11.
  • Part I (Ch. 1–2), the preface and the glossary are free online, so the first two weeks and every glossary lookup cost students nothing; the complete book is available as PDF/EPUB and in paperback.
  • The companion browser tools and explainers are designed as in-class exercises: each has a “Copy as Markdown” export students can submit, and each links back to the chapter and glossary entries it rests on.
  • The book is time- and solution-agnostic by design, so the syllabus does not need revising when models or vendors change; swap the examples, keep the concepts.
  • Instructors are welcome to adapt this syllabus under CC BY 4.0; please credit AI Agents, Engineered and link to the book’s site.

Questions readers ask

Do students need a paid model API to take this course?
No. Every lab runs against a scripted model client, a stub that returns a fixed sequence of responses, so grading never depends on a provider. Students who have access to a hosted or local model can use it, but nothing requires it.
Do students need a machine-learning background?
No. The course assumes programming, basic probability and ordinary software-engineering practice. The book explains how language models work from first principles in Chapter 2 and never asks anyone to train a model.
How much of the book is free?
The preface, Part I (Chapters 1 and 2) and the glossary are free to read online, which covers the first two weeks of the syllabus completely. The other chapters are in the full book, sold as PDF and EPUB, on Kindle and in paperback.
Can I adapt the syllabus and the diagrams for my own course?
Yes. The syllabus, the exercises and the book's diagrams on this site are licensed CC BY 4.0: reuse them, adapt them and translate them, with attribution to the book.
Is the course tied to a particular agent framework?
No. Like the book, the course teaches concepts and patterns. Frameworks and vendors appear only as labeled examples of a category, so the material stays valid as tools change.

License and attribution

The syllabus, lecture decks, exercises and the book’s diagrams on this site are licensed CC BY 4.0. Use them in your course, adapt them, translate them. Attribution line: “From the teaching kit of AI Agents, Engineered by Enrique Gutiérrez, CC BY 4.0, aiagentsengineered.com/teach/”.