Home / Blog / Teaching AI agents / How to Teach AI Agents: Sequencing and Assessin…

Teaching AI agents

How to Teach AI Agents: Sequencing and Assessing the Course

An AI agents course syllabus for instructors: why the loop comes first, how to grade verification over demos, and how peer courses compare. Get the kit.

By Enrique Gutiérrez · Published · 13 min read

An AI agents course syllabus is a sequence before it is a topic list. Mine runs the loop before tools, measurement before patterns, security before autonomy, and multi-agent systems last, and it grades students on the verification they ship (an eval set, a trace review, a regression gate) rather than on a demo.

This post is for instructors designing that course. It argues for the order, shows an assessment design that survives students having capable models, compares seven public university courses, and works one week through lecture, lab and grading. It deliberately does not repeat the week-by-week plan: that lives in the free teaching kit for AI Agents, Engineered, with a syllabus, labs, browser exercises and diagrams; the syllabus, exercises and diagrams are CC BY 4.0. Read this for the reasoning, and use the kit for the details.

Why does the order of an AI agents course syllabus matter so much?

The order matters because each topic in an agents course is a tool for checking the next one. Students who learn workflow and multi-agent patterns before they can measure an agent have no way to tell whether a pattern helped, so they learn to judge by the demo, which is the habit the course should break.

The book behind the kit states its thesis in one line: “an agent is only as trustworthy as the signal you can use to verify it.” Chapter 1 turns that into the first bearing of the compass, a question students should hear in week one and every week after: “what signal tells you it worked?” A syllabus is where that question either becomes a habit or stays a slogan.

I opened the public syllabi of the courses instructors most often benchmark against, and in the build-centered ones a pattern repeats: formal evaluation tends to arrive mid-to-late term. Stanford’s CS329Z teaches evaluation in weeks 7 and 8 of an eleven-week schedule; Michigan’s EECS 498 holds its evals lecture in week 7 and its regression-gates lecture in week 13 of 15; a popular sixteen-week online syllabus places evaluation in week 12 and safety in week 14 (codewithowais). Michigan is the strongest of these, because its students run an eval suite on their own agent by week 7.

Still, the default order across the field is “build, then test,” inherited from courses where the build was deterministic.

What are the four sequencing rules?

The four rules are: teach the loop before tools, measurement before patterns, security before autonomy, and multi-agent systems last. Each puts the instrument for judging a capability before the capability itself, so students meet every new piece of complexity with a way to decide whether it earned its place.

Why teach the loop before tools?

The loop comes first because every later topic is a change to one of its four parts. Chapter 3 of the book shows that a working agent is a model client, a set of tools, a message history and a loop with a step cap, short enough to write on an index card. Students who build that loop by hand can locate any later concept in it.

A minimal agent has exactly four parts, and all four live in the harness—the frame here.
Figure 3.4 A minimal agent has exactly four parts, and all four live in the harness—the frame here. The model client speaks to the provider, the tools are the actions on offer, the history is the run’s whole memory, and the loop (in accent) ties them together. The model itself sits outside the frame: the judgment you rent, wired in through the client. The harness is everything else—the part you own and build. Reuse this diagram

Building the loop first also teaches the first verification lesson for free. The chapter’s line is “The model’s ‘done’ is testimony; a green test is evidence,” and its rule for budgets is blunt: “Set the cap before you write anything else.” A student who writes a stop condition and a test that breaks the loop on purpose has done verification before the word appears on a slide. Stanford makes the same choice: students “first build core components (RAG, tool use, agent loops) from scratch,” then see how a framework abstracts them (CS329Z). The companion post on building an AI agent from scratch and the agent loop explainer can serve as pre-class material.

Why measurement before patterns?

Measurement comes before patterns because a pattern is a hypothesis about improvement, and without an eval set the hypothesis is never tested. Routing, evaluator–optimizer loops and orchestrator–worker designs all add cost and failure surface; students should compare each against a baseline on numbers, not on how the transcript reads.

The starting point is cheap. Chapter 16 opens with a ritual: run the same task three times, read all three outputs, and ask of each only “would I accept this?” It scales into a graded set quickly, and the chapter is specific about size: “Twenty to fifty tasks you have hand-checked and believe in beat a few hundred synthetic ones nobody has read.” That fits in a two-week unit.

The cheapest eval you can run this afternoon.
Figure 16.2 The cheapest eval you can run this afternoon. Give the agent the same task three times, read all three outputs, and ask of each only “would I accept this?” Here runs 1 and 2 land on the same shape of answer while run 3 diverges (in accent); the spread you read by hand is the reliability envelope of the previous figure, measured with your own eyes. Repeated trials, a human grader, an acceptance criterion—the whole discipline in miniature. Reuse this diagram

Why security before autonomy?

Security comes before autonomy because the autonomy dial is only safe to turn once students can see what an injected instruction does to their own agent. Teaching approval gates and unattended loops first invites students to grant permissions they cannot yet reason about, which mirrors exactly how production incidents start.

The unit that works best here is an audit of the student’s own design against the lethal trifecta (private data, untrusted content and a way to send data out), followed by a red-team exercise between teams. UW-Madison’s CS 839 already puts indirect prompt injection in week 5 of its paper seminar (CS 839), well before most courses reach the topic.

Why put multi-agent systems last?

Multi-agent systems belong last because they multiply every problem the earlier units taught students to measure: more calls, more cost, more places for a brief to lose context, and failures that cross agent boundaries. A student who can already trace, evaluate and secure one loop can judge whether a second agent pays; one who cannot will simply add agents.

How does the free teaching kit apply these rules?

The teaching kit at /teach/ applies the first rule exactly and the other three through habits rather than calendar position. Its published weeks follow the book’s chapter order, which puts multi-agent and oversight in week 6, observability and evaluation in weeks 7 and 8, and security in week 9.

I want to be plain about that trade-off, because instructors will notice it. Chapter order lets students build more of the system before they formally measure it, and it keeps the readings in sequence.

The kit compensates in two ways: every lab carries a one-page “what signal tells you it worked?” note, and from the first lab students test their loop against a scripted model client. If you prefer the stricter order argued here, the weeks are modular, and adapting the AI agents course syllabus takes an afternoon. Move the observability and evaluation weeks to 4 and 5, security to 6, and push context, memory, patterns and multi-agent later. The agentic AI curriculum post works through that reordering phase by phase.

How should you assess an AI agents course?

Assess it with a verification-centered project: students build a modest agent and must ship an eval set, a trace review and a regression gate alongside it. The agent may be small; the evidence must be real. In the kit this project carries 40 percent of the grade, and a project without evidence that it works does not pass.

Three deliverables form the core of the kit’s brief, which also asks for a design note, a security audit and a cost and rollout plan. Each of the three is hard to produce without running and reading your own system:

Deliverable What students submit Why a model can’t fake it cheaply
Eval set 20–50 hand-checked tasks with reference solutions, graders, negative cases, and pass rates over at least three runs per task The numbers must reproduce when the grader reruns the suite
Trace review The worst ten runs diagnosed by the first wrong span, one failure replayed, fixed and promoted to an eval case It requires the student’s own traces, which nobody else has
Regression gate A CI check that blocks a merge when the regression suite drops below a threshold The grader can push a breaking change and watch the gate fire

The grading rule follows Chapter 16’s distinction between pass@k and passk (pass@k and passk). A single impressive run earns nothing. What counts is the distribution over repeated runs, read through the reliability envelope: an agent that succeeds 90 percent of the time per attempt succeeds on all three of three runs only about 73 percent of the time (0.9³; an illustrative figure). The pass@k calculator makes this concrete in class.

The motivation is not hypothetical. An experience report from the University of Washington Bothell opens with the sentence “Large language models can complete most of the assignments in an introductory artificial intelligence course,” and describes rebuilding assessment around “tasks that resist unattributed automation” (Pisan, 2026).

Two peer courses add a check on understanding. Stanford follows each homework with “a 10-minute HW-based quiz where students explain their design decisions and tradeoffs,” and those quizzes carry 15 percent of the grade (CS329Z Logistics). Michigan’s syllabus puts the principle in two sentences: “Understanding is what gets graded here, not generation. If you can’t explain it, you didn’t build it.”

A defense of the eval set and the trace review is a natural fit for that format. The sibling post on AI agents assignments for students goes deeper into individual assignment designs.

Do students need a paid API?

No. Labs can run against a scripted model client, a stub that returns a fixed sequence of responses, so every student can build and test the harness at zero cost. The book places this stub at what Chapter 15 calls the deterministic–nondeterministic seam: the loop, tools and parsers are ordinary software and deserve exact tests.

The scripted client has one limit worth naming to students: a fixed script never varies, so three runs give three identical results and the pass rate measures the harness and the graders, not a model.

Two remedies work. One is a seeded “noisy” script that chooses among several recorded responses, including malformed and wrong ones, with set probabilities; that is my own lab design, not a standard. The other is a locally served open-weight model (llama.cpp and Ollama are two examples of the category). Michigan takes the second route for nearly all of its coursework and states that “the expected cost of this course is $0” (EECS 498).

How do current university AI agents courses compare?

Current university courses split into three designs: research seminars built on paper reading, guest-lecture series, and build-centered engineering courses. Only the build-centered ones grade students on measuring their own agent, and even there formal evaluation tends to arrive mid-to-late term. The table summarizes what each public syllabus says, as opened in October 2026.

Course Design Prerequisites Where evaluation sits How students are assessed
Berkeley CS294 LLM Agents (Fall 2024) Guest-lecture series with weekly readings ML and deep learning coursework strongly encouraged A listed topic; safety lecture last Variable units; at 3–4 units, project milestones, report and a 20% implementation component
Berkeley CS294 Agentic AI (Fall 2025) Guest-lecture series plus team project ML and deep learning experience strongly encouraged Fifth lecture of the term Agent track: teams first build “green agents” that evaluate tasks, then compete
Stanford CS329Z Build from scratch, then frameworks; quarter project An NLP course Weeks 7–8 of 11 Project 40%, homework 15%, HW quizzes 15%, paper video 10%, peer review 20%; one homework is “Evaluate an Agent”
Michigan EECS 498 AASE Apply, Analyze, Create around one coding agent; local models EECS 281 and EECS 201, or ULCS standing or instructor permission Evals week 7; regression gates week 13 No exams; eval suite of at least eight tasks, each run at least three times; hackathons
UW-Madison CS 839 Paper seminar Not listed in the course README Benchmarks read as papers in week 2 Reports, reviews and project presentations
UIUC CS598 Research-driven seminar on software-engineering agents Research background, NLP/ML coursework A benchmarks module mid-term Homework 20%, paper presentation 20%, participation 10%, project 50%; no exam
NYU Stern Foundations of AI Agents Six-session business course None formal Week 4 deliverable Weekly deliverables; demo day and peer review 40%
AI Agents, Engineered kit Build-centered, vendor-neutral, 13 weeks Programming, basic probability, no ML Every lab note; graded eval set in week 8 Labs 30%, verification project 40%, exam 20%, reading 10%

Two observations stand out. The research seminars are excellent at what they do, which is preparing students to read and write papers, and an engineering course should not imitate them. And Berkeley’s 2025 project, in which students build evaluating agents before competing ones, is the clearest public example of evaluation treated as a first-class build artifact. I would borrow it.

Worked example: one measurement week, lecture to grade

This worked example is the first measurement week in the stricter order, so it falls right after tools. Its single objective is that every team leaves with a twenty-task eval set for its own agent, run three times against a scripted or local client, with pass rates the instructor can reproduce. The numbers below are illustrative.

The lecture (75 minutes)

The lecture has four segments. The first 15 minutes are on compounding error: if each step succeeds 95 percent of the time, a 20-step task finishes cleanly about 36 percent of the time (0.95²⁰), which students verify live in the compounding error calculator. The next 20 minutes cover the three-run ritual and testable acceptance criteria, using Chapter 16’s contrast between “Summarize the meeting well” and an extraction rule two experts could not disagree about.

The third segment, 25 minutes, climbs the grading ladder from programmatic checks to binary rubrics to an LLM-as-a-judge, ending on the chapter’s warning: “A judge you have not calibrated is an opinion you have automated.” The last 15 minutes show the regression gate pipeline as the destination for week two of the unit.

Eval-driven development as plumbing.
Figure 16.4 Eval-driven development as plumbing. Every change runs against two suites: the capability evals it is meant to climb, and the regression evals it must not break; tasks that stabilize are promoted from the first into the second. The change reaches a gate (in accent) where a threshold is enforced, not glanced at—the numbers, not moods, decide whether it merges or is blocked. Reuse this diagram

The lab (two hours)

Each team writes twenty tasks drawn from failures it has already seen in its own agent, each with a reference solution proving it can be passed and at least three negative cases. They run the suite three times through a seeded noisy scripted client, keep infrastructure errors out of the quality number, and report pass rates per task. The run costs nothing because no live model is called.

The lab closes with a sizing exercise. A team reporting 80 percent on twenty tasks has a 95 percent confidence interval of roughly ±18 points (normal approximation, illustrative), and even on fifty tasks it is still about ±11. The eval sample-size calculator shows why Chapter 16 says “a two-point shift on a fifty-task suite is well inside the noise.”

The assessment (10 points)

Criterion Points What earns full marks
Every task passable 3 Each task has a reference solution the grader can run
Grader on the lowest rung that works 2 Programmatic checks wherever an objective answer exists
Negative cases 2 At least three tasks assert what the agent must not do
Reproducible numbers 2 The instructor’s rerun matches the reported pass rates
Signal note 1 One page answering “what signal tells you it worked?”

Where does this design fall short?

The design falls short in three places. It asks more of instructors, because reproducing student eval runs takes infrastructure; it underweights research skills, so a course meant to feed a PhD pipeline should keep a paper-reading strand; and scripted clients can’t teach how a live model actually behaves under pressure.

Each has a mitigation. Rerun a random sample of teams rather than all of them, and add a weekly paper response alongside the labs, in the spirit of the CS 839 seminar. Then schedule at least one lab against a local model so students meet real variance before the project. If you teach a business or policy audience rather than engineers, the NYU Stern format of six sessions and weekly deliverables is a better starting point than any thirteen-week plan.

Where to go from here

The one idea to keep is that the order of an AI agents course syllabus is your assessment philosophy in disguise: whatever students learn to measure first becomes the standard they hold everything else to. Teach the loop, then the instrument, then the risks, then the complexity.

The full week-by-week plan, labs and exercises are in the free teaching kit, and the companion AI agents lecture slides add instructor notes as they are published. Chapters 1 and 2 are free: start with Chapter 1 for the compass, then read Chapter 3 on the loop and Chapter 16 on evaluation in the full book, available in the formats listed on the home page.

Questions readers ask

What should an AI agents course syllabus cover?
It should cover what an agent is and how the underlying model fails, the agent loop and its stop conditions, tools and context, evaluation and observability, security and oversight, workflow and multi-agent patterns, and reliability and deployment. The order matters as much as the list: students should be able to measure an agent before they learn the patterns that make it more complex.
When should evaluation be taught in an AI agents course?
Early. A three-run check and scripted-client tests belong in the same week as the first loop, and a graded eval set belongs right after tools, before workflow and multi-agent patterns. In the build-centered public courses, formal evaluation tends to arrive mid-to-late term, which means students build for weeks with no instrument for telling better from worse.
Do students need a paid API to take an AI agents course?
No. Labs can run against a scripted model client that returns a fixed sequence of responses, which tests the harness, the tools and the graders exactly and at zero cost. Students who want live behavior can use a locally served open-weight model; Michigan's EECS 498 runs nearly all of its coursework on models students serve themselves and states that the expected cost of the course is $0.
How do you grade AI agents assignments when models can do the homework?
Grade evidence the student had to produce by running and reading their own system: an eval set with reference solutions, pass rates over repeated runs, a written review of the worst traces, and a regression gate in CI. Pair it with a short individual quiz in which students explain their own design decisions, as Stanford's CS329Z holds after each homework.
Is there a textbook for a university AI agents course?
Some current university courses, Michigan's EECS 498 among them, state that they have no textbook, and many others assemble papers instead. AI Agents, Engineered was written to fill that gap without tying a course to one vendor; its preface, first two chapters and glossary are free online, and the teaching kit maps its 27 chapters onto a 13-week term.

Sources

  1. University of Michigan EECS (2026). EECS 498-016 Applied Agentic Software Engineering: Syllabus (Fall 2026)
  2. Diyi Yang, Michael Ryan and John Yang (Stanford University) (2026). CS329Z Engineering AI Agents
  3. Stanford University (2026). CS329Z Engineering AI Agents: Logistics
  4. Dawn Song and Xinyun Chen (UC Berkeley) (2024). CS294/194-196 Large Language Model Agents (Fall 2024)
  5. UC Berkeley RDI (2025). CS294/194-196 Agentic AI (Fall 2025)
  6. UC Berkeley RDI (2025). CS 294/194-196 Agentic AI: Introduction slides
  7. Aws Albarghouthi (UW-Madison) (2026). CS 839 AI Agents (Spring 2026)
  8. Lingming Zhang (UIUC) (2026). Software Engineering with LLM Agents (CS598)
  9. Srikanth Jagabathula and Ilan Lobel (NYU Stern) (2026). Foundations of AI Agents: Syllabus (Spring 2026)
  10. Yusuf Pisan (2026). Teaching Intro AI When the Tools Can Do the Homework: A Course Redesign and a Student Bill of Rights
  11. Scaler (2026). Agentic AI Syllabus: Complete Curriculum for Autonomous AI Systems
  12. Owais (codewithowais) (2026). Agentic AI Syllabus