This AI agents learning roadmap has four stages: a minimal loop, an agent with real tools inside a sandbox, an eval set with a calibrated judge, and a coding or research agent you have operated for a week. Each stage names the chapters to read, one build, one proof you can show, and the check that ends it.
I wrote the book those chapters come from, AI Agents, Engineered, so weigh the reading list with that in mind. The order of the stages is my proposal. It rests on the book’s chapter order and on what learners ask in public, and nobody has tested it against another order.
The question behind it was posted on r/AI_Agents on 5 October 2026: “If you were starting from zero today, what would you learn in what order?” The same post gave the reason for asking. Its author did not want to keep “jumping from one technology to another without actually understanding how to build agents properly.”
What does this AI agents learning roadmap promise?
The roadmap promises one thing: at any point you can say which stage you are in and what would have to be true to leave it. A stage ends when its check passes, whatever the hours spent or the topics covered.
That promise comes from what learners ask. For this post I collected 18 questions and answers from three r/AI_Agents threads and six Hacker News items, all read on 6 October 2026.
Seven asked for one order: what first, what next. Six asked what comes after a first agent. Three were about the churn of roadmaps and tutorials, and two about how to know that something is finished.
Eighteen rows from two sites is a convenience sample, and I claim nothing about learners in general. The full table sits in my research notes and is unpublished; the threads I quote are in the sources. It does say where a roadmap for this reader should spend its length: on the order, and on the step from a working loop to an agent that can do damage.
What does it leave out?
This AI agents learning roadmap leaves out calendars and framework choices. I found no measurement of how long any stage takes a learner, so the only time figures on this page are reading minutes, labeled approximate. The roadmap says what each stage lets you show, and it makes no claim about what any employer checks.
Two facts about the book shape every build below. It is written in pseudocode throughout, Appendix A included; the Preface promises “examples in plain pseudocode”, so each build is your own program in your own language. And at the time of writing, only the Preface, Part I (Chapters 1 and 2) and the glossary are free to read online. Every other chapter cited here is in the full book, and the exit checks work without it.
Which stage are you in?
You are in the first stage whose exit list you cannot complete today, whatever you have read or watched. An AI agents learning roadmap is only useful if you can find yourself on it, so take the four questions in order. Stop at the first one you cannot answer with something you can open.
- Stage 1. Lower your agent’s step cap to 2 and give it a task that needs more steps than that. Does the run end with a result labeled “stopped”, with the transcript saved?
- Stage 2. Do you hold a log showing that a script, running where your agent’s code runs, was refused when it tried to read a file outside its workspace?
- Stage 3. Do you hold a pass rate from tasks you checked by hand, and has any judge model you use been scored against your own labels?
- Stage 4. Do you have a log of every run from seven consecutive days of real work, with the worst runs diagnosed?
Each question is one item from that stage’s exit list, chosen because it is quick to try. The full list, further down, is what decides. If a quick question passes and the list has an unticked item, you are still in that stage. If all four lists are complete, you have finished what this page maps.
Why place yourself by a list?
A list places you by what exists on your disk today, and a demo is one run that ended well. Each question above asks for something a single run cannot supply. My guess is that many readers with a demo stop at the first or second, and the lists will tell you better than I can.
Placing yourself this way has one more use. It removes the argument about whether a course or a tutorial “counts”. If the Stage 1 list is ticked and you can explain each item, the first stage is behind you, whatever you read or watched to get there.
What is the smallest thing you can do today?
The smallest thing is one unticked item from your stage’s list, and each stage has one that fits in a single sitting. A developer on Hacker News in July 2025 described the opposite condition: “However, I’m not exactly sure where to start. There’s just too much noise.”
- Stage 1. Set the step cap to 2, run a task that needs more steps, and save the transcript.
- Stage 2. Write a script that tries to read one file outside the workspace, and run it where the agent’s code runs.
- Stage 3. Write three tasks with a reference solution each, one of them a task where the right result is no change.
- Stage 4. Make your harness append one line to a log file for every run it starts.
You do not need the rest of this page for that. Read your own stage’s section, do the one item, and leave the other roadmaps closed until it is done.
What are the four stages of the roadmap?
The four stages are a minimal loop, real tools in a sandbox, an eval set with a calibrated judge, and a week of operating a coding or research agent. Each row below gives the reading, the build, the proof and the exit check. Minutes are approximate; “free” means readable online today, and “full book” means locked.
| Stage | Chapters (approx. reading minutes) | Build | Proof you can show | The check that says the stage is done |
|---|---|---|---|---|
| 1. A minimal loop | Ch. 1 (16, free), Ch. 2 (43, free), Ch. 3 (25, full book), App. A (15, full book). Total 99, of which 59 free | The four parts in one file, translated from the book’s pseudocode: model client, a few tools, message history, a loop with a step cap | The file and three saved transcripts: one finished, one stopped by the cap, one with a tool failure returned to the model | With the cap at 2, on a task that needs more steps, the run returns “stopped”. With a tool unplugged, the agent reports the failure and invents nothing |
| 2. Real tools and a sandbox | Ch. 5 (35, full book), Ch. 17 (38, full book). Total 73 | Tools with consequences, each with a description, a schema and structured errors; execution in a sandbox; scoped credentials; one approval gate | A one-page contract per tool, a denied-actions log from your own escape attempts, and your three settings in three lines | A test script inside the sandbox fails to read outside the workspace, reach an unlisted host or outrun the time limit. The gated action does not run without you. Each refusal is logged |
| 3. An eval set with a calibrated judge | Ch. 15 (38, full book), Ch. 16 (29, full book). Total 67 | Traces on every run; 20 to 50 hand-checked tasks, three runs each; a judge only where code cannot decide, measured against your labels | The eval table, the pass rate with its confidence interval (the range the true rate plausibly sits in), the judge’s agreement table with its rubric, if you use a judge, and one rerun before and after a change | Every task has a reference solution that passes its own check. Any judge you use was scored on a few dozen of your labels, with kappa beside accuracy |
| 4. A coding or research agent operated for a week | Coding: Ch. 21 (32) and Ch. 22 (42), total 74. Research: Ch. 23 (28). Both: Ch. 12 (26) and Ch. 19 (30), total 56. All full book | Seven days of real work of your own, every run logged: a repository with a test suite, or questions answered from sources you can open | A one-page operating report: runs per day, how each ended, cost and steps, the worst runs diagnosed | The harness wrote a log entry for every run, the bad ones included. Each of the worst runs has a named cause. The Stage 3 suite grew from the week’s failures |
How long is the reading, and which chapters are free?
Reading the core chapters of this AI agents learning roadmap takes about 369 minutes when the path ends in a coding agent (99 + 73 + 67 + 74 + 56) and about 323 minutes when it ends in a research agent. Both figures use the web reader’s own formula: a chapter’s word count divided by 220, rounded. They estimate reading only, and they are approximate.
Each build takes longer than its reading. I know of no measurement of how much longer, so I give none. What I can size is the work itself: about sixty lines of pseudocode to translate, 20 to 50 tasks, three runs each, a few dozen hand labels, seven days of logged use.
Every chapter on the path is in the table below, the optional ones included. With scripts on you can narrow it to one stage; with scripts off, every row stays on the page.
| Chapter | What it gives the stage | Approx. minutes | Access | Role | Stage |
|---|---|---|---|---|---|
| Ch. 1, What Is an Agent? | The definition, the control-flow test, the autonomy dial | 16 | Free | Core | Stage 1 |
| Ch. 2, The Engine: How Language Models Work | Function calling and the engine’s failure modes | 43 | Free | Core | Stage 1 |
| Ch. 3, The Agent Loop | The loop, the four parts, stop conditions and budgets | 25 | Full book | Core | Stage 1 |
| App. A, A Minimal Agent, Annotated | The whole program in annotated pseudocode | 15 | Full book | Core | Stage 1 |
| Ch. 5, Tools and the Action Space | The tool set, the tool interface, code execution and the sandbox | 35 | Full book | Core | Stage 2 |
| Ch. 17, Security, Safety, and Guardrails | Prompt injection, the lethal trifecta, least privilege, sandboxing, human approval | 38 | Full book | Core | Stage 2 |
| Ch. 18, Reliability, State, and the Harness | Errors surfaced to the model, retries, timeouts, state and recovery | 34 | Full book | Optional | Stage 2 |
| Ch. 15, Observability and Debugging | Traces and spans, monitoring, reproducing failures, a taxonomy of agent bugs | 38 | Full book | Core | Stage 3 |
| Ch. 16, Evaluating Agents | Building an eval set, LLM-as-a-judge, regression gates | 29 | Full book | Core | Stage 3 |
| Ch. 26, Agents as Classifiers and Scorers | Three disjoint sets and the held-out test set | 25 | Full book | Optional | Stage 3 |
| Ch. 21, Coding Agents | Why coding led, the coding-agent loop, agent-legible code | 32 | Full book | Core, coding path | Stage 4 |
| Ch. 22, The Coding Workflow in Practice | Spec, plan and execute; tests first; git as the safety net; review | 42 | Full book | Core, coding path | Stage 4 |
| Ch. 23, Research and Business Agents | The verification gap and its partial substitutes | 28 | Full book | Core, research path | Stage 4 |
| Ch. 12, Oversight and Autonomy | Approval gates, reviewing intent and outcomes, the autonomy dial | 26 | Full book | Core, both paths | Stage 4 |
| Ch. 19, Cost, Latency, and Performance | Where tokens and money go, and the cost levers | 30 | Full book | Core, both paths | Stage 4 |
| Ch. 13, Writing the Outer Loop | Goal functions for loops that run unattended | 33 | Full book | Optional | Stage 4 |
| Ch. 20, Deploying and Scaling | Sandboxed execution at scale and staged rollout | 27 | Full book | Optional | Stage 4 |
Stage 1: can you build a minimal loop and make it stop?
Stage 1 is done when you have written the smallest agent yourself and can make it stop on purpose. The build is four parts in one file, translated from the book’s pseudocode into your own language: a model client, a few tools, the message history, and the loop with a step cap.
Chapter 3 introduces the build with a promise: “An agent has exactly four parts, and you already understand every one of them.” Everything in that file except the model is the harness, the ordinary code that runs the loop and executes the tools. Read the free chapters first, Chapter 1, “What Is an Agent?” and Chapter 2, “The Engine”, then Chapter 3, The Agent Loop (in the full book) and Appendix A, A Minimal Agent, Annotated (in the full book).
Run your translation in a throwaway folder. The appendix says “The minimal agent is minimal in its safety, too” and means that the edit tool runs with your own permissions.
How do you know Stage 1 is done?
Stage 1 is done when two tests pass: the cap fires and says so, and a broken tool produces a report where an invented result would be easier. Chapter 3 is direct about the first: “Set the cap before you write anything else.” A budget is one kind of stop condition, and the first exit check below tests yours.
Learners supplied the second test. An answer in the same r/AI_Agents thread put it this way: “One useful test of your understanding: unplug a tool. If the agent loses access to a folder, does it report that or invent a result?” Another answer said to break the agent on purpose once it works, and to watch whether it “reports the failure or invents an answer”.
Keep the file and three saved transcripts as the proof. The browser exercise Run the loop puts you in the harness’s seat for one scripted run, which helps before you write your own.
Two other routes reach the same two tests. If you want a second pass at this stage, build an AI agent from scratch in one file with real code, and put that file through the cap test and the unplugged tool. If the pace here is too fast, slow down and learn AI agents from scratch, the loop first, then tools, then evals, which is the order of the first three stages.
- I can point to the four parts in my own file: model client, tools, message history, loop.
- With the step cap lowered to 2, on a task that needs more than two steps, the run ends with a result labeled “stopped” and does not crash.
- With a tool unplugged or made to fail, the failure returns to the model as a result, and the agent reports it and invents nothing, in three runs out of three.
- I have three saved transcripts: one that finished, one stopped by the cap, one with a tool failure.
- For each transcript I can name the exit it took, and say what outside the model’s own message shows the finished run did its job.
Stage 2: can your agent hold real tools inside a sandbox?
Stage 2 is done when your agent holds tools with consequences and you have tried, and failed, to get out of the room it runs in. A sandbox, in Chapter 5’s words, “is the sealed room the interpreter runs inside: its own filesystem, capped processor, memory, and time, and no network except what you deliberately allow.”
Published roadmaps mostly skip this stage: of the four compared further down, one mentions sandboxing. An answer to a first-time builder on r/AI_Agents in October 2026 named what the stage prevents: “the fails you were avoiding all have one shape: the agent held a write it shouldn’t have had.” The same thread began with a builder who had been avoiding long-running agents for that reason.
The reading is Chapter 5, Tools and the Action Space (in the full book) and Chapter 17, Security, Safety, and Guardrails (in the full book). Chapter 5 opens on the stakes: “The set of tools you expose is the complete inventory of what your agent can ever do.”
What do you build in Stage 2?
You build three things onto the Stage 1 agent: tools that change something, a boundary around where they run, and one gate. Give the agent a tool that runs code or commands, one that writes files, and one call to a real service. Write each a description, a schema and structured errors, meaning failures returned as data the model can read and act on. Chapter 5 says where to start: “Begin with the description, because it is the cheapest thing to fix and the most common thing to get wrong.”
Then put execution in the sealed room. Deny the network by default and name the few hosts it may reach, which is the allow-list. Keep credentials out of the room, and follow the chapter’s second setting: “And give each session a fresh sandbox, torn down afterward, so nothing leaks between runs.” Scope every credential to the task, and put one approval gate on the single action you cannot undo.
Chapter 17 calls the agent a deputy, because it acts with authority you lent it. It calls these three settings dials and sums them up: “Least privilege caps what the deputy’s keys open; the sandbox caps where it can stand; approval puts a human in front of the few moves that remain expensive.” Together they bound the blast radius, which is the most damage a wrong action can do.
What does the Stage 2 proof look like?
The Stage 2 proof is paperwork plus a log: a one-page contract per tool, your three settings in three lines, and a record of every escape you tried. Two tools check the paperwork. Paste a tool definition into the tool schema linter and it reports findings on the name, description, schema and side effects.
The lethal trifecta is a name from the security literature, adopted by the book, for three capabilities that are dangerous together: access to private data, exposure to untrusted content, and a channel for external communication. The lethal trifecta audit shows which of the three your design now holds.
Your log of refusals will have the shape below, with illustrative lines written for this post; your fields will differ.
attempt result logged by
read /home/me/.ssh/id_key refused sandbox filesystem
connect example.org:443 refused network allow-list
loop past the time limit killed sandbox timer
send_email, no approval given held approval gate
delete with the read-only key refused the service
A limit belongs here. A sandbox assembled by a learner is an exercise. Chapter 17 advises assembling the boundary from isolation primitives that have survived years of attack, and it quotes a vendor’s postmortem of a correctly restricted sandbox: “The sandbox worked perfectly, and yet the data was exfiltrated.”
- From inside the sandbox, my test script tries to read a file outside the workspace; the attempt is refused and logged.
- The same script tries to reach a host that is off the allow-list; the attempt is refused and logged.
- A run that goes past the time limit is stopped and logged.
- A search of the sandbox’s files and environment variables finds no credential or secret.
- A second session cannot see the first session’s files.
- The gated action, attempted without my approval, does not run, and the attempt is logged.
- My credential, used for something outside the task, is refused.
- I can name the worst thing the agent could do with what it holds.
Stage 3: can you measure it with an eval set and a calibrated judge?
Stage 3 is done when you hold a pass rate from tasks you checked by hand, and any judge you rely on has been measured against your own labels. An eval set is the collection of tasks you grade the agent on, and each task pairs an input with a way to score the result. An LLM-as-a-judge is a model asked to grade what code cannot.
The reading is Chapter 15, Observability and Debugging (in the full book) and Chapter 16, Evaluating Agents (in the full book). Start by recording every run as a trace, the structured record of what the agent did. Chapter 15 says how to read one: “Read the testimony; verify against the record.”
Then write 20 to 50 tasks you have checked by hand, each with a reference solution, programmatic checks first, three runs per task. Include negative cases, which are tasks where the right behavior is to do nothing. The sibling post on the best way to learn AI agents specifies that task set, its template and its interval arithmetic for a first project, so I will not repeat them. The eval sample size calculator gives the confidence interval, the range the true pass rate plausibly sits in, for any count.
How do you calibrate the judge?
You calibrate a judge by labeling a sample yourself and measuring how far the judge agrees with you. Chapter 16 states the method: “Before a judge grades anything unsupervised, you calibrate it the way you would calibrate any instrument: against ground truth, which here means you.” It asks for a few dozen outputs at minimum, some of them bad.
Report Cohen’s kappa beside accuracy. Kappa is an agreement statistic corrected for chance: it asks how much better than lucky guessing the judge did. In the worked example of this site’s judge agreement calculator (illustrative numbers), 35 of 50 hand-labeled outputs are good. A judge that always says “pass” scores accuracy 0.70 and kappa 0.00.
Sample size matters, and the chapter is blunt: “Kappa computed on ten examples is a random number; you need a few dozen labels before it means anything.” Its verdict on skipping the step is the line I would pin above the desk: “A judge you have not calibrated is an opinion you have automated.” The post on whether LLM-as-a-judge is reliable goes further into the biases.
Set your bar before you measure. The chapter’s footnote offers a rough yardstick: human graders on such tasks tend to agree at “a kappa around 0.7 to 0.8”, which it calls “a sensible bar for a judge meant to stand in for a human”. The calculator’s page uses the same neighborhood.
What if you need no judge?
If code can decide every criterion in your set, you need no judge, and Stage 3 ends without one. Chapter 16 brackets its whole method with the rule that “the best judge is the one you did not need” and tells you to use a code check wherever one can decide. Write down, per task, that code decides it, and tick the judge item on that basis.
The item reopens the day you add a criterion code cannot check, such as whether an explanation is correct. From then on the judge is an instrument in your set, and it gets calibrated like one.
- My set has 20 or more hand-checked tasks, and every task has a reference solution that passes its own check.
- Every run is traced, and each task ran at least three times.
- Infrastructure failures, such as a timeout or a crashed harness, are counted apart from quality failures.
- I report the pass rate with its confidence interval.
- Either code decides every criterion in my set and I have written that down per task, or my judge was scored against my own labels on a few dozen outputs that include bad ones.
- If I use a judge, I report Cohen’s kappa beside accuracy, and it meets the bar I wrote down before measuring. If I use none, this item is ticked with the one above.
- I can show one failure that became a new task, and one rerun of the suite before and after a change.
Stage 4: what does a week of operating a coding or research agent teach?
Stage 4 is done when you have pointed the agent at real work of your own for seven days, had the harness log every run, and written down what the worst runs taught you. The week is the length of the exercise by definition. It says nothing about how long the stage takes to reach.
“Operated for a week” is this post’s framing. The book has no chapter by that name; the practice is spread across its chapters on monitoring, cost and oversight. Pick one of two paths, and read Chapter 12, Oversight and Autonomy (in the full book) and Chapter 19, Cost, Latency, and Performance (in the full book) for either.
What changes on the coding path?
On the coding path the agent works in a repository of yours that has a test suite, and the suite is the verifier. Chapter 21, Coding Agents (in the full book) explains why coding went first: “The signal arrives in seconds, costs nothing, and is nobody’s opinion.” Chapter 22, The Coding Workflow in Practice (in the full book) turns that into a working routine.
Two rules hold for the seven days. Nothing merges without a green suite. And Chapter 22’s instruction is to “Treat any diff that touches a test file as an escalation”, so you read every such diff before it merges.
The suite has limits, and Chapter 21 names them: “green is a fact about the tests, and agents left unwatched exploit the gap from both sides”. A passing run tells you about the tests you wrote. Whether the change fits the codebase is still yours to judge.
What changes on the research path?
On the research path the agent answers questions you need answered, from sources you can open, and no program can grade the result. Chapter 23, Research and Business Agents (in the full book) calls this the verification gap: the distance between how convincing an output looks and how cheaply it can be checked.
Your instrument is a sample. The chapter describes it: “The third substitute is the spot-check: you, opening citations, on a sample you choose by consequence.” Count the share of sampled citations that support the sentence they are attached to, and write the number down.
The cost is your reading time, which the chapter puts in one line: “There is no compiler for facts; there are only inspectors, and inspectors bill by the hour.”
What goes in the operating report?
The operating report is one page: runs per day, how each run ended, cost and steps per run, the worst runs by cost or step count, and what you stopped letting the agent do. Diagnose each of the worst runs to its first wrong step. Chapter 19 gives the reason to sort by cost: “A run that fails after forty steps is billed for forty steps.”
Three tools help. The agent cost-per-task estimator estimates what one run costs in tokens, the agent bug bestiary goes from a symptom to ranked suspects, and the consequence tier classifier sorts each action into a tier with the gate it needs. The post on AI agent failure modes is a longer companion for the diagnosis.
End the report with a setting. Say where on the autonomy dial the agent ran this week and what would have to be true to move it. One builder, describing what held up over months, wrote in an October 2026 thread: “Autonomy came last, per area, once I knew its bounds.”
One week gives a baseline and no trend. Chapter 15’s rule for drift is that “A week-over-week regression rule catches a slow slide that no absolute threshold will” and that rule needs a second week.
- My harness writes the log, and the number of entries matches the number of runs I started in the seven days, the bad ones included.
- Each of the worst runs by cost or step count has a named cause: its first wrong step.
- The Stage 3 suite grew by at least one task from this week’s failures.
- My path’s rule held. Coding: nothing merged without a green suite, and I read every diff that touched a test. Research: I counted the share of sampled citations that support their sentence.
- I can say where on the autonomy dial the agent ran, and what would have to be true to move it.
What do the common roadmaps order, and what do they leave out?
The common AI agents roadmap orders topics, and it leaves out the point where a topic is finished. I read four on 6 October 2026. All four are dated examples of one category, the page that puts topics in an order, and the table notes the form each takes. I rank none of them.
| Roadmap, as read on 6 October 2026 | Form | Steps | Where evaluation appears | Does a step end in a test you can pass or fail? |
|---|---|---|---|---|
| roadmap.sh, AI Agents (published 2025, modified March 2026) | Topic map | 101 topic files in its public content folder | Several topic files, on testing, human-in-the-loop evaluation and metrics | No, in the file list and the five topic files I opened |
| NovelVista (Akshad Modi, updated June 2026) | Step list | 7 steps | Nowhere: the strings “eval”, “sandbox” and “monitor” occur 0 times | No. Project ideas by name, with no check |
| Simplilearn (Sayantoni Das, July 2026) | Step list, for training leads | 7 stages | One bullet in Stage 3 and a metrics bullet in Stage 6 | No. “7 real-world projects and 1 capstone”, unspecified |
| Ron Paul, GaaS (May 2026) | Advice essay | 5 sections | In the fourth of five sections, with guardrails, cost and deployment | No. Names two milestones, with no test for either |
What do they do well?
Each does something worth having. The topic map on roadmap.sh is the broadest free index of the subject I know, and its file list includes tool sandboxing and several kinds of testing. I read that list and five of its topic files; where those topics sit in its visual order I did not check.
NovelVista’s step list goes as far as naming project ideas. The Simplilearn piece names a real problem for people who plan training, in its heading “Why Most Companies Have a Course Catalog, Not a Learning Roadmap”. Ron Paul’s essay is the closest in spirit to this one. It puts evaluation in the path and warns: “This is where many learners stop too early, mistaking a working demo for real competence”.
What is missing from all four?
All four order what to learn, and none closes a step with an artifact and a test the reader can pass or fail. For roadmap.sh that statement rests on its file list and the five topic files I opened. Sandboxing appears in one of the four, as a single topic among 101.
Two of them do describe testing as a practice. The roadmap.sh topic file on unit testing says to “Run the tests every time you change the code.” Ron Paul’s essay says to “write test cases that measure how often your agent succeeds”. Evaluation is absent from a third page and two bullets in the fourth.
Two of the four give a time estimate, and neither cites a measurement behind it, which is why this AI agents learning roadmap attaches none. If you want a map of everything, use a topic map. If you want to know what you owe yourself next, use an exit check.
What do you do between stages, and when do you go back?
Between stages, keep the same project and carry every failure forward as a test. Go back a stage whenever a later change breaks an earlier check. One answer in the same r/AI_Agents thread that opened this post gave the first half as plain advice: “Keep one concrete project the whole time”.
Going back is routine. A new tool or a new credential in any stage reopens the Stage 2 escape script. Changing the prompt or a tool description reruns the Stage 3 suite. And a bad run in the Stage 4 week becomes a Stage 3 task, and sometimes it sends you to the sandbox.
So the stages stack, and none is left behind. By the fourth, the first three checks run on every change, and the agent verifiability scorecard is a quick way to see which signals your project still lacks.
When do frameworks, memory and multi-agent systems come in?
They come in when a check fails for a reason they fix. The Preface says “the book is divided into seven parts, ordered as a course”, and this roadmap follows that order for the first three stages and then skips most of two parts: context management, retrieval, memory, workflow patterns and multi-agent systems. I left them off the core path because none is needed to produce the four proofs.
Read them when the need shows up in a transcript. Context and memory become necessary when a run outgrows the context window. Multi-agent designs become necessary when one agent demonstrably cannot do the job, and the post on single-agent versus multi-agent systems has the test. An answer in the opening thread ordered it the same way: learn “state, memory, retrieval, error handling, and verification before moving on to multi-agent systems.”
A framework fits after Stage 1. Chapter 3’s rule, quoted in the FAQ on this page, is to build from scratch to understand and to adopt a framework once you can name what it saves you. One answerer in that thread advised the opposite, a simple framework first, and still put multi-agent work last.
What can you show for each stage?
Each stage ends in an artifact a stranger could read or rerun: saved transcripts with a budget stop, a denied-actions log, a pass rate with a judge’s agreement table, and a one-page operating report. Showing them means saying what each proves and where it stops.
| Stage | What you can show | What it shows | What it does not show |
|---|---|---|---|
| 1 | The file and three transcripts | You wrote the loop and it stops when told | That the agent is good at anything |
| 2 | Tool contracts and a denied-actions log | The boundary refused the escapes you tried | That it refuses the ones you did not try |
| 3 | Pass rate, interval, judge agreement table | A measured result on your tasks, with an instrument you checked | Behavior on tasks outside the set |
| 4 | A one-page operating report | A week of real use, with failures diagnosed | A trend, which needs a second week |
Passing an exit check shows that the artifact exists, which is a smaller claim than competence. I would distrust any roadmap that promised more.
For the vocabulary behind these artifacts, the AI agents study guide maps 87 glossary terms to chapters and has a self-test. Use it beside the stages: the study guide checks what you can explain, and the exit lists check what you have built.
Who should not follow this roadmap?
Four kinds of reader should not follow this AI agents learning roadmap as written. Each has a better route.
- A reader with a framework and a date already fixed. When a job or a course has chosen both, Stage 1’s from-scratch build is a detour you cannot afford yet. Meet the date, then come back and write the loop yourself.
- A reader who needs a working automation this week. Adopt an existing product. The stages teach you to build and check an agent, which is slower than using one.
- A reader whose results only a person can grade. Stage 3 builds its set on code checks first. If every output of yours is prose, a judge carries the whole set from the first task, and calibrating it becomes the stage. The Stage 4 research path describes the instruments for that harder road.
- A reader heading for model training or ML research. That path needs a different roadmap. This one assumes a hosted model and, in the Preface’s words, “ordinary engineering literacy”.
What limits apply to everyone else?
Three limits apply to every reader. The stage order is this post’s proposal, built on one book’s chapters and a small sample of public questions. No study has compared it with another order. And three of the four stages cite locked chapters, so the free material covers most of the reading for Stage 1 and none for the later stages.
It is also one author’s reading list. If you want to weigh other books before choosing, I wrote an honest comparison of the best books on AI agents, my own included.
If you learn better in a taught sequence, the site’s teaching kit has a 13-week syllabus that asks its capstone for much the same evidence. Its first three weeks come with AI agents lecture slides, three open lectures with instructor notes.
The one thing to keep
A stage you can leave is a stage with a check. Chapter 1 puts the book’s question in one sentence: “The first question to ask of any agent design is the one this book will ask over and over: what signal tells you it worked?” This AI agents learning roadmap asks the same question of your learning, four times.
So pick the first unticked item in your stage’s list and make it this week’s work. The last chapter of the book says why the build matters more than more reading: “Reading about tools informs you. Building against them calibrates you.”
As of October 2026, two chapters on this path cost nothing to open: Chapter 1, “What Is an Agent?” and Chapter 2, “The Engine”, 59 of Stage 1’s 99 reading minutes. Chapter 3, Appendix A and the chapters for Stages 2 to 4 are in the full book. When your stage’s list sends you to one of them, see the formats.
Questions readers ask
- In what order should I learn AI agents?
- The loop first, then real tools inside a boundary, then measurement, then a week of real use. Each of the four stages ends in an artifact and a check you can run. Topics such as retrieval, memory and multi-agent systems come in when a check fails for a reason they fix.
- What should I learn after building my first agent?
- Give the agent one tool with consequences and a sandbox to hold them, then prove the sandbox by trying to get out of it. Write a script that tries to read a file outside the workspace, reach a host that is off the allow-list (the list of hosts the sandbox may contact) and run past the time limit. All three attempts should be refused and logged.
- How long does it take to learn AI agents?
- I found no measurement of it, and this roadmap gives no calendar. What can be counted is the reading: about 369 minutes for the core chapters when the path ends in a coding agent, by the web reader's formula of words divided by 220, which is approximate. The builds are sized in tasks, runs and labels, and they take longer than the reading.
- Do I need a framework, machine learning or math to follow the roadmap?
- No. The book's Preface says "The only prerequisite is ordinary engineering literacy", and Stage 3 uses arithmetic plus one agreement statistic. A framework fits after Stage 1. Chapter 3 gives the rule: "Build from scratch to understand. Adopt a framework to scale—and only once you can name what it is saving you."
- Can I follow the roadmap with only the free chapters?
- Partly. As of October 2026 the Preface, Chapters 1 and 2 and the glossary are free online, which covers 59 of the 99 reading minutes of Stage 1. The exit checks and the site's tools need no book. Chapter 3, Appendix A and the chapters for Stages 2 to 4 are in the full book.
Sources
- u/Money-Designer-9724 (2026). I'm trying to learn AI Agents, but I'm getting really confused. How should I approach it? (r/AI_Agents thread)
- u/Moonsteroid (2026). Bit late but i just built my first Agent (r/AI_Agents thread)
- u/Money_Ad4075 (2026). Building a personal AI agent on my Mac what would you prioritize first? (r/AI_Agents thread)
- roadmap.sh contributors (2026). AI Agents roadmap (published 2025, modified 2026)
- erkanerol (2025). Ask HN: What is your strategy for staying up-to-date with AI developments?
- roadmap.sh contributors (2026). developer-roadmap, roadmaps/ai-agents/content (topic file list)
- Akshad Modi (NovelVista) (2026). Agentic AI Roadmap 2026: Step-by-Step Learning Path
- Sayantoni Das (Simplilearn) (2026). How to Create an Agentic AI Learning Roadmap for Your Workforce
- Ron Paul (2026). How to Build a Learning Roadmap for Agentic AI