An agentic AI curriculum is a sequence of topics in which each unit arrives after the instrument students need to judge it. Mine runs the loop, tool design, tracing and evaluation, security, context, workflow patterns and autonomy, then multi-agent systems. That sequence reorders the chapters of the book this site belongs to.
I wrote this for an instructor with a course proposal open and a box labeled “rationale” still empty. By the end you can place any topic against a dependency table and show a committee where seven public courses put the same five topics. You can also paste a one-page rationale into the form.
The companion post on the AI agents course syllabus states the four ordering rules and the assessment design. This one supplies the argument underneath them and the chapter-by-chapter reordering that post promised.
Why is the order the hard part of an agentic AI curriculum?
The order is the hard part because the topic list is nearly settled and the sequence is wide open. Most plans I read draw on the same handful of topics: loops, tools, memory, evaluation, security and multi-agent systems. They disagree on what comes first, and almost none of them says why.
In the threads I read, the people asking for an order in public were learners, and I found no instructor asking. The opening post of one thread on r/AI_Agents ends with the question “If you were starting from zero today, what would you learn in what order?” (u/Money-Designer-9724, 5 October 2026). The answers in such threads mostly agree on the easy part: one loop before a framework, one agent before several.
They split on the rest. A reply in a second thread says “Evaluation and security matter, but they are deeper topics you can pick up later” (u/Luvena21, 22 September 2026).
That reply is the position this post argues against. I found no published study that compares two orderings of an agentic AI curriculum. I also found no instructor describing a committee’s objection in public, so the committee here is my inference about your situation. What follows is an argument from prerequisites: what a student must already hold before a topic can be judged at all.
What are the four ordering rules, and what does each depend on?
The four rules are the loop before tools, measurement before patterns, security before autonomy, and multi-agent systems last. Each rule names a dependency: the later topic can only be judged with something the earlier topic builds. Here is the dependency behind each, with the book’s line where one exists.
The loop before tool design
Tool design depends on the loop because a tool description only fails visibly inside a loop a student can read. That reason is mine. The book supplies the neighboring one, about frameworks, in Chapter 3: “Stay minimal while you are learning, prototyping, or running one well-scoped loop; the bare version teaches you what every part is for”.
The rule needs one clarification. A minimal loop already contains tools as one of its parts, so students call a tool in the session where they write the loop. What comes afterward is tool design as a craft: names, descriptions, error messages. The agent loop explainer works as pre-class material.
Measurement before patterns
Patterns depend on measurement because a pattern is a claim that the system improved, and the claim needs an instrument. Chapter 1 states the dependency in its advice on when to add complexity: “build the simpler version first, measure it against real cases, and climb the ladder only when you can point at a class of inputs the simple system provably fails on”. The same paragraph says a later chapter “shows how to do the measuring”, and that chapter is Chapter 16.
One pattern leans on evaluation directly. In the evaluator–optimizer loop of Chapter 10, a model critic sits where a test would sit, and the book says of that critic: “Its verdict is an estimate where the verifier’s is a fact”. Making such estimates trustworthy is the subject of Chapter 16, six chapters later. A student who meets the critic first has no way to ask whether it can be believed.
The instrument is also cheap to start. Chapter 16 puts its smallest version ahead of everything: “Before any infrastructure, any dashboard, any framework decision, there is a ritual you can run this afternoon”. The ritual is three runs of one task, read end to end.
Security before autonomy
Autonomy depends on security because widening what an agent may do alone is only safe once students have seen what injected text does to their own agent. Chapter 13, on unattended loops, states the prerequisite with a forward reference to the security chapter. The lethal trifecta is private data, untrusted content and an outbound channel in one agent. The chapter’s instruction is to “audit the combination, per loop, before the first unattended run”.
Autonomy also depends on evaluation. Chapter 12 sets the default (“Start every new task class tight.”) and then loosens it only on evidence, which includes evaluation scores. Its figure of the autonomy dial carries the same condition in its caption.
The Preface gives the same idea as a warning: “Where no signal exists, autonomy is a leap of faith, and this book will keep saying so.” A course that teaches the dial before the signal and the attack asks students to turn it on faith.
Multi-agent systems last
Multi-agent systems depend on nearly everything else, which is why I place them last among the design topics. The word “last” is mine. The book’s contribution is skepticism: Chapter 11’s verdict begins with the instruction “Start with one agent.”
The chapter also warns against the status the topic carries: “Treating multi-agent as a maturity badge, something a serious system graduates into, misreads what it is. It is an architecture with a habitat.”
The dependency comes from the conditions the chapter attaches to any split. Among them: “observability runs across the whole system from day one; the bill gets measured rather than assumed”. Tracing and measurement are Chapters 15 and 16. The chapter also flags blast radius, a security estimate from Chapter 17, as what “sets the sensible degree of parallelism”.
One commenter makes the same point in a Hacker News thread on multi-agent work in production: “most of the pain at scale isn’t the agents themselves, it’s observability” (kaihwang, 14 September 2026). It is one anonymous opinion, offered as an illustration. The engineering decision itself belongs to the post on single agent versus multi-agent.
Which topic needs which one first?
Twelve topics cover an agentic AI curriculum at this grain, and each needs between zero and four earlier ones. The table below is this post’s own argument, and no source prints it. Where the “why” is the book’s, the line is quoted; where it is mine, the cell says “my reasoning”.
| # | Topic | What it needs first | Why | Book chapter |
|---|---|---|---|---|
| 1 | What an agent is, and whether you need one | Nothing | The book: Chapter 1 draws “the working distinction among chatbots, workflows, and agents that every later chapter leans on” | 1 (free) |
| 2 | How the model works and fails | 1 | My reasoning: nondeterminism and compounding error explain why ordinary testing habits break | 2 (free) |
| 3 | The loop, stop conditions, budgets, a first three-run check | 1, 2 | The book: “With the engine understood, we can build the machine around it” | 3, App. A |
| 4 | Planning and tool design | 3 | My reasoning: a weak tool description shows up only inside a loop students can read | 4, 5 |
| 5 | Traces | 3, 4 | My reasoning: there is nothing to trace until a loop calls tools; the trace also records tokens per step | 15 |
| 6 | Evaluation: an eval set, repeated runs, a calibrated judge | 5 | The book: “the traces this chapter teaches you to read turn out to be the raw material the next chapter’s eval sets are grown from” | 16 |
| 7 | Security: injection, the lethal trifecta, least privilege | 4, 6 | The book: tools turn text “into a potential set of orders”; the model guard is the Chapter 16 judge “wearing its second uniform” | 17 |
| 8 | Skills, context, retrieval, memory | 3, 4, 6 | The book, on context settings: the evaluation habits of Chapter 16 “catch the regressions your intuition misses” | 6–9 |
| 9 | Workflow patterns; when to build no agent | 6 | The book: “climb the ladder only when you can point at a class of inputs the simple system provably fails on” | 10, 14 |
| 10 | Oversight, autonomy, unattended loops | 6, 7 | The book: “Part V builds the state, evaluation, and security engineering that gates and dials rest on” | 12, 13 |
| 11 | Multi-agent systems | 5, 6, 7, 9 | The book: “observability runs across the whole system from day one; the bill gets measured rather than assumed” | 11 |
| 12 | Reliability, cost, deployment, applications | 6, 7 | The book: the harness sensors of Chapter 18 and all of Chapter 26 “all rest on the judge this chapter builds” | 18–27 |
Read a row as a test of your own schedule. Suppose workflow patterns sit in week 5 and the first eval set is due in week 9. Row 9 is violated, and for four weeks students judge patterns by how the transcript reads. Row 6 needs row 5 for a practical reason: the first eval tasks come from failures students found by reading a trace.
One row is looser than the others. Row 8 can move: context and memory need evaluation, and nothing in the table forces them before or after security. Rows 8, 10, 11 and 12 are needed by no other row, so a short course can drop any of them cleanly. Dropping row 9 means dropping row 11 with it.
How does the book’s own chapter order differ, and why?
The book’s chapters do not follow this order. AI Agents, Engineered builds an agent (Parts I to III), then covers architecture (Part IV), then makes the result reliable (Part V). Patterns in Chapters 10 to 14 precede evaluation in Chapter 16. Oversight and the outer loop in Chapters 12 and 13 precede security in Chapter 17, and multi-agent systems are Chapter 11 of 27.
The Preface describes that arrangement in its own words: “the book is divided into seven parts, ordered as a course.” Any agentic AI textbook has to commit to one order, and this one gives reasons for its choice. Part IV decides what gets built before Part V measures it. Chapter 14 closes Part IV by saying that “whatever you build, you will need to see exactly what it did, and you will need to measure whether it worked”.
Chapter 16 then opens on a finished system: “You have built the loop, curated the desk, wired up the tools; the agent runs, and it is often impressive. Is it good?” For a reader working alone, that is a natural moment to ask.
On autonomy the book is explicit about why design precedes machinery. Chapter 12 says “the design of the working relationship belongs here, before the machinery, because the machinery only automates decisions you have to make first”. Inside Chapter 11, the debate about whether to build multi-agent systems is held back on purpose: “I have ordered the chapter so their debate comes last, after you have seen the evidence it turns on”.
Why reorder it for a course?
A graded calendar changes what the order does. A reader can flip forward when Chapter 10 mentions the judge. A cohort cannot, and whatever is graded first becomes the standard students hold.
The Preface allows the move: “Read the parts in order for a grounding, or, once oriented, raid them as a reference; the chapters cross-reference one another and try to stand on their own.” The dependencies in the table are mostly the book’s own forward references, followed in the direction they point.
How do the chapters reorder, phase by phase?
Nine phases map this agentic AI curriculum onto the 27 chapters, and three blocks of chapters move. Chapters 15 and 16 come forward to sit right after tools, Chapter 17 follows them, and Chapter 11 moves behind Chapters 12 to 14.
| Phase | Table rows | Chapters, in teaching order | Change from the book’s order |
|---|---|---|---|
| 1 Foundations | 1, 2 | 1, 2 | None |
| 2 The loop | 3 | 3, App. A | None |
| 3 Planning and tools | 4 | 4, 5 | None |
| 4 The instrument | 5, 6 | 15, 16 | Moved up, ahead of Chapters 6 to 14 |
| 5 Risk | 7 | 17 | Moved up, ahead of Chapters 6 to 14 |
| 6 Context and memory | 8 | 6, 7, 8, 9 | Now after evaluation and security |
| 7 Patterns and autonomy | 9, 10 | 10, 12, 13, 14 | Now after evaluation and security; Chapter 11 lifted out |
| 8 Multi-agent | 11 | 11 | Moved behind Chapters 12 to 14 |
| 9 Operating and applications | 12 | 18, 19, 20, then 21–27 | None |
This agrees with the one-sentence version in the syllabus post: “Move the observability and evaluation weeks to 4 and 5, security to 6, and push context, memory, patterns and multi-agent later.” That post’s free kit follows chapter order, with multi-agent and oversight in week 6, observability and evaluation in weeks 7 and 8, and security in week 9. The syllabus post also puts a three-run check in the same week as the first loop, and I keep that habit in phase 2.
What does the reordering cost?
The reordering costs four chapter openings that point at chapters the class has not read, and each needs a short bridge from you. Chapter 15 opens by recalling how Chapter 14 closed Part IV. Chapter 16 opens with “curated the desk”, and the desk is Part III, which now comes two phases later. Chapter 17 opens by settling a debt from Chapter 7 on context poisoning, and its action guardrails mention an approval gate, the subject of Chapter 12.
The fourth is in phase 7. Chapter 12 begins “The previous chapter ended on a question it deliberately left open”, and the previous chapter is Chapter 11, which the class now reads afterward. Tell students the question in one sentence: how does a person keep judging work that has outgrown their ability to read it?
There is a pedagogical cost as well, and the syllabus post already names it: chapter order “lets students build more of the system before they formally measure it, and it keeps the readings in sequence.” In the reordered course, students measure a small agent in phase 4 and grow it afterward. I accept that trade, because the eval set written for the small agent judges every later addition.
Where do seven public courses put the same five topics?
Seven public course pages, read on 7 October 2026, place the loop, tools, evaluation, security and multi-agent in seven different arrangements, and none follows the order proposed here. Each row records only what the page’s schedule lists, by the week or session number the page uses. “None titled” means no session title names the topic; the course may still teach it.
| Course, as its page titles it (term) | Loop or agent basics | Tools | Evaluation | Security or safety | Multi-agent | Design |
|---|---|---|---|---|---|---|
| University of Michigan, “EECS 498-016 · Applied Agentic Software Engineering” (Fall 2026) | Week 4 (L08, with tools); built in Lab 03, week 5 | Week 4 (L08); the tool interface again in week 11 (L21) | Week 7 (L13, Lab 05), after a measuring lab in week 6; regression gates in week 13 (L24) | Approval layer week 5 (L09); guardrails week 7 (L14); permission policy and prompt injection week 10 (L18, L19) | None titled (the string “multi-agent” does not occur on the syllabus page) | build-centered |
| Stanford, “CS 329Z: Engineering AI Agents” (Fall 2026) | Week 1 (foundations); agent design patterns in week 4 | Week 3 | Weeks 7–8 (slots 13–14 of 19) | Week 8 (slot 15 of 19) | Week 5 (slot 8 of 19) | build-centered |
| UC Berkeley, “Agentic AI”, CS294/194-196 (Fall 2025) | Lecture 2 of 13 | None titled | Lecture 5 of 13 | Lecture 13 of 13 | Lectures 7 and 11 of 13 | guest-lecture series |
| UC Berkeley, “Large Language Model Agents”, CS294/194-196 (Fall 2024) | Lecture 2 of 12 | None titled | Lecture 11 of 12 (“Measuring Agent capabilities…”) | Lecture 12 of 12 | Lecture 3 of 12 (a frameworks lecture; one of its two listed readings is a multi-agent paper) | guest-lecture series |
| UW-Madison, the AI agents class, CS 839 (Spring 2026) | Weeks 0–1 | Week 1 | Week 2 (benchmarks, read as papers) | Week 5 | Week 9 (second-to-last topic week) | seminar |
| UIUC, “Software Engineering with LLM Agents” (sessions dated 01/20 to 05/05; the README prints no year) | Modules I–II (LLM basics 01/29; coding agents from 02/10) | None titled | Module III, from 03/05 (benchmarks) | None titled | None titled | seminar |
| NYU Stern, “Foundations of AI Agents”, “Syllabus (2026)” (second half of Spring 2026) | Day 1 of 6 | Day 1 | Day 4 | Day 3 on the course site (“security considerations”); absent from the syllabus PDF | Day 2 | short course |
The Design labels follow the syllabus post’s three categories, and “short course” is my label for the six-day NYU format. The Michigan, Stanford and NYU Stern pages each say their schedule or contents may still change, and the UIUC README heads its table “Tentative Schedule”. An eighth page, Texas A&M’s ECEN 689 for Fall 2026, lists only an assignment order, so it has no row. Its last two assignments of seven are “Multi-Agent Orchestration” and “Final Multi-Agent Workflow Optimization”.
What do the seven schedules show?
They show the first rule as common practice, and no schedule that follows the other three together. All four schedules that title both agent basics and tools put the basics first or in the same session. The counts below are mine, from the table.
- Evaluation and multi-agent. Five schedules list both, counting the Berkeley 2024 frameworks lecture as multi-agent. Evaluation comes first in two (Berkeley 2025 and UW-Madison) and second in three (Stanford, NYU Stern and Berkeley 2024).
- Security. It is the final lecture in both Berkeley terms and slot 15 of 19 at Stanford. Only UW-Madison places it before its multi-agent week.
- Multi-agent. It is the last topic in none of the seven.
- Autonomy. Michigan’s Analyze phase, weeks 4 to 7, carries the line “Take the human out of the loop”, with an approval layer taught in week 5. Its prompt-injection lecture is in week 10, and Lab 08 in week 11 is “Harden the approval layer you built in week 5”.
Michigan is the closest of the seven to teaching security alongside autonomy. Its prompt-injection lecture still comes after that phase.
Two plans from outside universities, as examples of that category, also differ from the proposed order. A 16-week individual syllabus puts multi-agent in week 9, evaluation in week 12 and safety in week 14. An 18-lesson free vendor course lists “Building Trustworthy AI Agents” sixth, its multi-agent lesson eighth and “Securing AI Agents” eighteenth. Its README tells readers to “start wherever you like”, so that list is a catalog more than a sequence.
One detail fits the skeptical half of the fourth rule. Stanford’s multi-agent lecture lists “Why Do Multi-Agent LLM Systems Fail?” among its additional readings, beside an essay titled “Don’t Sleep on Single-agent Systems”. UW-Madison’s week 9 lists the same essay as supplementary reading.
Why might a course order it differently?
A course might order its agentic AI curriculum differently for at least five reasons, and your committee will raise some of them. None of the seven schedules matches mine, so “peer courses do it this way” is unavailable to you as a defense.
Students need something worth measuring. This is the book’s own order, and the reasoning is in its Chapter 16 opening. A student who has built retrieval, memory and two patterns has a system whose failures are interesting. My answer is the three-run ritual, which needs only a loop, and no data settles it.
Design decisions come before machinery. Chapter 12 makes this case for oversight, as quoted above. It is right that an approval gate is a policy decision first. My narrower claim concerns the unattended end of the dial, where the book itself asks for the audit first.
What does the one published experience report say?
The one report I found favors multi-agent projects and early frameworks, and it supports early evaluation. Mello and Maher (CSEDU 2026) describe an undergraduate course of 59 students and analyze 24 final project reports. Their abstract says “Higher-performing projects consistently integrated explicit evaluation strategies and multi-agent architectures.” Their first implication is that “agentic frameworks should be introduced early and positioned as the primary vehicle for project work.” Both statements go against this post.
The same paper supports the second rule: “Higher-performing projects consistently treated evaluation as a first-class design concern rather than as an afterthought.” Its limits cut both ways. It is one course in one term, graded by its own instructors, and it reports an association without testing an ordering. Its framework is a small one written for the course, described as “deliberately small enough to be understood end-to-end within a semester”, which differs from adopting a production framework in week one.
Which practical constraints push the other way?
Two constraints do: project lead time and course format.
A project needs lead time. If teams choose an architecture in week 3, a multi-agent unit in the final third arrives too late to inform the choice. In my order the project starts as one agent, and a team adds a second only after a measured baseline exists.
The format decides. A guest-lecture series has to fit its speakers’ calendars, and a seminar follows the literature. That is my inference, since no page states its reasons. A closing safety lecture also suits a survey: by then the audience has seen every capability the lecture warns about.
Worked example: a five-day intensive for working engineers
Laid out in ten half-days, the sequence fits a five-day intensive with no topic ahead of something its row needs. I walked two course shapes through the table: this one, and a 14-week semester course for CS seniors with a project. Both pass. I print the intensive, the tighter fit of the two.
| Half-day | Table rows | What participants do |
|---|---|---|
| Day 1, morning | 1, 2 | Classify three of their own tasks as chatbot, workflow or agent; compute how per-step error compounds |
| Day 1, afternoon | 3 | Write the loop by hand with a step cap; run one task three times and read the outputs |
| Day 2, morning | 4 | Rewrite two tool descriptions and watch the loop’s behavior change |
| Day 2, afternoon | 5, 6 | Read traces of the morning’s failures; turn them into a small eval set |
| Day 3, morning | 7 | Audit their agent for the trifecta; attempt a prompt injection through a tool result |
| Day 3, afternoon | 8 | Add retrieval or a notes file; rerun the Day 2 set to see whether it helped |
| Day 4, morning | 9 | Try one workflow pattern against the single-loop baseline, on numbers |
| Day 4, afternoon | 10 | Set a gate and a dial position; run one loop unattended inside the Day 3 scopes |
| Day 5, morning | 11 | Split one task across workers; compare cost and pass rate with one agent |
| Day 5, afternoon | 12 | Retries, cost per task and a rollout plan, as a survey |
Two tight spots showed up in the walk. Day 2 afternoon carries two rows, so the eval set will be small and the judge of row 6 stays uncalibrated. An LLM-as-a-judge therefore cannot be trusted for grading that week, and programmatic checks have to carry it.
Row 11 asks for a measured bill on Day 5, while the cost chapter belongs to row 12 that afternoon. The token counts recorded in Day 2’s traces cover it, which is why row 5 mentions them.
In the semester walk I found one constraint worth passing on. A project proposal due before the patterns phase cannot require an architecture choice. It can require a task, a success signal and a first eval set, and the architecture can be argued later from results. The lethal trifecta audit serves as the row 7 exercise in either format.
A sequence rationale to paste into a course proposal
The template below restates the dependency table row for row and says plainly that it reorders the textbook. Fill the brackets, delete rows your course omits, and keep the paragraph on costs.
SEQUENCE RATIONALE: [COURSE CODE] [COURSE TITLE]
Principle. Each unit comes after the instrument students need to judge it.
Order and prerequisites (unit <- what it needs first: reason)
1. What an agent is; whether one is needed <- nothing: every later unit
uses the chatbot / workflow / agent distinction. [weeks __]
2. How the model works and fails <- 1: nondeterminism and compounding
error explain why ordinary testing habits break. [weeks __]
3. The loop, stop conditions, budgets <- 1, 2: students hand-build a
model call, tools, a message history and a capped loop, then run
one task three times and read the outputs. [weeks __]
4. Planning and tool design <- 3: a weak tool description shows up only
inside a loop students can read. [weeks __]
5. Traces <- 3, 4: nothing to trace until a loop calls tools; traces
also record tokens per step. [weeks __]
6. Evaluation (repeated runs, a hand-checked eval set, a calibrated
judge) <- 5: eval tasks are grown from traced failures. [weeks __]
7. Security (injection, private data + untrusted content + an outbound
channel, least privilege) <- 4, 6: tools create the risk; a guard
model is a judge and needs calibrating. [weeks __]
8. Skills, context, retrieval, memory <- 3, 4, 6: each is a change
whose benefit has to be measured. [weeks __]
9. Workflow patterns; when to build no agent <- 6: a pattern is a claim
of improvement and needs a baseline. [weeks __]
10. Oversight, autonomy, unattended loops <- 6, 7: permissions widen on
evaluation evidence, after a security audit. [weeks __]
11. Multi-agent systems <- 5, 6, 7, 9: they need tracing across agents,
a measured cost and a single-agent baseline. [weeks __]
12. Reliability, cost, deployment, applications <- 6, 7: harness checks
and guardrails reuse the judge. [weeks __]
How this differs from peer courses. In [N] public schedules reviewed on
[DATE] ([LIST]), evaluation is taught in [POSITIONS] and security in
[POSITIONS]. This course moves both earlier on purpose. The order is
argued from prerequisites. I know of no study comparing orderings.
Text and readings. [TEXT] is assigned out of chapter order. This sequence
reorders the textbook's chapters: [MAPPING, e.g. Ch. 1-5, 15-17, 6-9,
10, 12-14, 11, 18-27].
What the order costs. [e.g. students measure a small agent before they
build a large one; four chapter openings refer to material not yet read.]
A self-taught reader needs a different document: an AI agents roadmap with a proof at the end of each stage, since nobody grades a self-learner. The tool policy that keeps a course vendor-neutral is a separate decision again, part of teaching AI agents to computer science students without tying them to one product.
Each unit’s exercises still have to be something a grader can check. Designing AI agents assignments for students is its own problem, because the rows above say when a unit comes and nothing about how to grade it.
What are the limits of this analysis?
The limits are a small sample, a moving target and an untested claim. Seven courses is a convenience sample of English-language pages whose schedules were public and readable on one day. It supports “these seven differ” and nothing about what most courses do.
Schedules change each term, and four of the seven pages mark theirs as tentative or subject to change. Check the live page before you cite a course in a proposal, and record your own date.
A session title is also a coarse measure. A course with no session titled “evaluation” may assess it in every lab, so the table compares calendars and says nothing about what students learn.
Finally, the order itself is unproven. It rests on forward references in one book and on my reading of what each topic requires. A literature-led order may serve researchers in training better, and a six-session course for managers has no room for twelve rows.
Where to go from here
An agentic AI curriculum earns its order one dependency at a time. Before each unit, ask what students would need in hand to tell whether the new thing helped. If the answer is a later week, move one of the two.
The week-by-week plan is in the free teaching kit. As an AI agents textbook for a university course, the book works when assigned in the phase order above.
Start with Chapter 1, which is free online with the Preface, Chapter 2 and the glossary. There the compass sets the question every unit returns to: “what signal tells you it worked?” Chapter 16 on evaluation and Chapter 11 on multi-agent systems are in the full book, and the available formats are listed on the home page.
Questions readers ask
- In what order should the topics of an agentic AI curriculum be taught?
- One defensible order is: what an agent is and how the model fails; the loop; planning and tool design; tracing and evaluation; security; context, retrieval and memory; workflow patterns, oversight and unattended loops; multi-agent systems; then reliability, cost, deployment and applications. Each unit follows the instrument students need to judge it. The order is argued from prerequisites, and no study comparing orderings was found.
- Should an AI agents course start with a framework?
- The case for the hand-built loop first is that a framework hides the four parts every later topic changes. Stanford's CS 329Z sets a first homework to be built 'with no agent frameworks'. One experience report recommends the opposite for a small teaching framework written for its course: Mello and Maher (CSEDU 2026) say agentic frameworks 'should be introduced early'.
- When should evaluation be taught in an AI agents course?
- Before the patterns it has to judge. A three-run check fits the same week as the first loop, and a graded eval set fits right after tool design. In seven public schedules read on 7 October 2026, evaluation sits anywhere from week 2 (as benchmark papers) to the second-to-last lecture of the term.
- Does multi-agent belong in a first course on AI agents?
- Yes, late in the design material and with the single-versus-multi question attached. The verdict in Chapter 11 of AI Agents, Engineered begins with the instruction 'Start with one agent.' Two public courses already pair their multi-agent session with a reading that argues for single agents. One experience report found that its higher-performing student projects used multi-agent architectures, so the placement is a judgment call.
- Is there an agentic AI textbook that follows this teaching order?
- AI Agents, Engineered does so only when its chapters are assigned out of order. Its chapters run build, then architecture, then reliability, so a course that wants evaluation and security early assigns Chapters 15 to 17 ahead of Chapters 6 to 14. The Preface, Chapters 1 and 2 and the glossary are free online; the other chapters are in the full book.
Sources
- University of Michigan EECS (2026). EECS 498-016 Applied Agentic Software Engineering: Syllabus (Fall 2026; read 7 October 2026)
- Stanford University (2026). CS 329Z: Engineering AI Agents, schedule and logistics pages (Fall 2026; read 7 October 2026)
- UC Berkeley RDI (2025). Agentic AI, CS294/194-196 (Fall 2025; read 7 October 2026)
- UC Berkeley RDI (2024). Large Language Model Agents, CS294/194-196 (Fall 2024; read 7 October 2026)
- UW-Madison course repository (2026). The AI agents class, CS 839 (Spring 2026): course README (read 7 October 2026)
- UIUC course repository (2026). Software Engineering with LLM Agents: course README (read 7 October 2026)
- NYU Stern (2026). Foundations of AI Agents: Syllabus (2026), with the course site aiagents.stern.nyu.edu (read 7 October 2026)
- Texas A&M University (2026). ECEN 689, LLMs for Agentic Hardware Design and Security: syllabus (Fall 2026; read 7 October 2026)
- Chad Mello and James Maher (2026). FairLLM: A Pedagogical Framework for Teaching Agentic Large Language Model Systems in an Undergraduate Artificial Intelligence Course (CSEDU 2026, vol. 3, pp. 2239–2246, DOI 10.5220/0014911800004021)
- codewithowais (2026). Agentic AI Syllabus (a 16-week individual syllabus; one example of a non-university plan)
- Microsoft (2026). AI Agents for Beginners: README (an 18-lesson free vendor course; a second example)
- u/Money-Designer-9724 (r/AI_Agents) (2026). Reddit thread: I’m trying to learn AI Agents, but I’m getting really confused. How should I approach it? (5 October 2026)
- u/Luvena21 (r/AI_Agents) (2026). Reddit comment on what to learn first and what can wait (22 September 2026)
- kaihwang (Hacker News) (2026). Hacker News comment on observability in multi-agent work (14 September 2026)