Home / Blog / Teaching AI agents / The Best Way to Learn AI Agents Is to Build One…

Teaching AI agents

The Best Way to Learn AI Agents Is to Build One You Can Verify

The best way to learn AI agents is to build one small agent a program can check. Get the first project, a 20-task eval plan and the arithmetic to copy.

By Enrique Gutiérrez · Published · 18 min read

The best way to learn AI agents is to build one small agent whose result a program can check, and to write the checks in the same week as the loop. You finish holding a pass rate with an honest margin and a set of failures you can explain, which is more than a demo gives you.

That is an opinion, and I will argue for it and then say what it costs. It starts from a moment I kept finding when I read what learners post. Someone has spent months on tutorials and framework tours, has a demo that ran once, and has nothing that shows it works. A newcomer who could see that outcome coming wrote, on r/AI_Agents in April 2026: “I just don’t want to spend 3 months on the wrong thing.”

What is the best way to learn AI agents, and why this one?

The best way to learn AI agents is to pick a project where success is checkable by a program, build the smallest agent that attempts it, and measure it on tasks you wrote yourself. The measuring is the half that most learning paths postpone, and it is the half that tells you whether you understood the first half.

I should be honest about the evidence for what learners want. For this post I collected 26 questions and answers from public threads on Hacker News and Reddit, all read on one day in October 2026. Seven were about the order to learn things in, six about getting hired, six about being stuck after a lot of material, four about what to build first. Only three were about how to know an agent works, and one of those three advised leaving evaluation for later.

So most readers did not come here asking for evals. My claim is that one practice answers the other questions too.

A check fixes the order: loop and check together. A failing check is a to-do list no course hands you. The need for a check picks the project. And a measured result is something another person can question, where a demo can only be watched.

The book this site belongs to describes the demo’s evidence without mercy, in Chapter 16: “It is one run, on one input, graded by the warm feelings of five people who wanted it to work.” A learner alone at a desk is in the same position with an audience of one.

Why do tutorials and framework tours leave you with a demo?

Tutorials and framework tours leave you with a demo because most of them stop at the point where the agent runs. That is a reasonable choice of scope, and it leaves the question “does it work?” for the reader to discover alone.

Two widely read from-scratch essays are good examples of the build half. Thorsten Ball’s “How to Build an Agent” (2025) and Thomas Ptacek’s “You Should Write An Agent” (2025) each show how little code an agent is. When I searched the text of both pages in October 2026, neither contained the string “eval”. A reader who had just been pointed to Ball’s walkthrough posted the next question on Hacker News in 2025: “how do I evaluate the correctness of output produced by LLM given the inputs.”

Roadmaps and courses tend to place the answer late. One detailed AI agents learning roadmap (Zafar, 2026) puts evaluation at step 10 of 11, after frameworks.

Two free vendor-hosted courses, named here as dated examples of a category, show the same shape as of October 2026. The Hugging Face AI Agents Course ends with an agent scored on a benchmark the course chose, and lists “Agent Observability and Evaluation” as a bonus unit. Microsoft’s AI Agents for Beginners lists 18 lessons, none titled evaluation, with exercises on that vendor’s own stack.

There is also a reason the smooth path feels right. In a randomized comparison in college physics, students in active classrooms scored higher on a test of learning, yet “their perception of learning, while positive, was lower than that of their peers in passive environments” (Deslauriers et al., 2019). A polished video produces the feeling. This is classroom evidence, the nearest I could find, and it was not measured on self-taught adults.

How do you choose a first agent project?

Choose a first agent project with one test: a program must be able to tell whether the run worked, without asking the model. If it can, you can count successes. If the only judge is the agent’s own closing message, you have a demo generator.

Chapter 3 puts the distinction in one line: “The model’s ‘done’ is testimony; a green test is evidence.” An agent can announce a finished job in the same assured tone whether or not the job is finished. A file that contains the new value, or a test that exits clean, does not depend on tone. The chapter ranks the ways a loop can end by how far each can be trusted.

The loop has three ways out, and they are not equally trustworthy.
Figure 3.5 The loop has three ways out, and they are not equally trustworthy. A verified check—a suite that goes green, a file that parses—is evidence, and where one exists it should pronounce the run finished (in accent). The model’s own “done” is only testimony, believed with caution. And the budget cap is the safety net: it settles nothing about the work, but it guarantees the loop stops, so it must always be there. Reuse this diagram

The test rules out more first projects than you might expect. A research summarizer fails it, because judging prose needs a person or a calibrated judge model. A browser agent on someone else’s website fails it in a different way. One learner who spent two months there was told, in a thread on r/AI_Agents in August 2026, to start with “a task where you own both ends, your own repo or your own API.”

Owning both ends means you create the starting state and you can inspect the final state. Most advice on AI agents for beginners sorts concepts into an order. The test sorts projects, and it has an answer for each candidate you bring to it.

What should the first agent be?

The first agent should be a file agent with three tools (list files, read a file, edit a file) that works inside small throwaway project folders you create. Its outcome is files on disk, so a short script can check every run.

This is the book’s own minimal agent. Appendix A, A Minimal Agent, Annotated (in the full book) gives the whole program and counts it at “About sixty lines, blank lines and comments included” before taking it apart. Chapter 3 says of the design: “An agent has exactly four parts, and you already understand every one of them.” The four parts are a model client, tools, a message history and the loop. Everything around the model is the harness, and it is ordinary code.

One thing to know before you plan the evening. The appendix is pseudocode, on purpose: “The program is in pseudocode, and the choice is deliberate.” You translate it into Python, JavaScript or whatever you write, against whatever capable model you can reach.

If you searched for AI agents in Python from scratch, this is how the book serves you: the shape, annotated line by line, with the typing left to you. Doing that translation yourself is what it means to learn AI agents from scratch.

Two cautions. The edit tool runs with your permissions, so keep the agent in scratch folders. And the appendix is frank about what the agent cannot do: with these three tools its own confirmation is re-reading the file it just wrote. The list of what is deliberately missing includes the line this post is built on: “Nothing measures whether runs succeed.”

So the measuring lives outside the agent, in a script of yours. If you want to feel the loop before writing one, Run the loop lets you play the harness for a scripted run in the browser.

What goes in the 20-task set?

The 20-task set holds four kinds of task: single-file changes, find-then-change tasks, multi-file changes, and negative cases where the right behavior is to change nothing. Each task is a folder (the starting state), a goal sentence, a check script and a reference solution.

An eval set is a curated collection of tasks, each an input plus a way to grade the output. Twenty is the low end of the range the book recommends for a start: “Twenty to fifty tasks you have hand-checked and believe in beat a few hundred synthetic ones nobody has read.” Anthropic’s engineering guide to agent evals (2026) gives the same range from practice: “20-50 simple tasks drawn from real failures is a great start.” The split below is illustrative.

Kind Tasks (illustrative) Example goal What the check script asserts
Single-file change 6 “Change the greeting from Hello to Welcome” The constant has the new value; the folder’s small test passes
Find-then-change 6 “The timeout is defined somewhere in this project; set it to 30” The new value is in the right file; no other file changed
Multi-file change 4 “Rename the function fetch to fetch_user everywhere” The old name is absent from every file; the test passes
Negative case 4 “Change the greeting to Welcome” in a folder that has no greeting; “Summarize README.md” No file changed (compare hashes before and after)

The negative cases matter more than their count suggests. A file agent that always edits something will pass every positive task. Chapter 16 asks for them directly: “Whatever the rung, include negative cases: tasks asserting what the agent should not do.”

The reference solution is the edit done by your own hand, saved next to the task. It proves the task can be passed, and the book treats that as non-negotiable: “every task must be solvable, and the way to know is to include a reference solution that proves it.” Writing it is also how you find out that your goal sentence was ambiguous.

Write the first three tonight, one of them a negative case. The other seventeen come from what those three teach you.

What counts as a pass?

A run passes when three things hold: the check script exits clean on the final state of the folder, no file outside the task’s allowed list changed, and the loop ended by finishing, inside its step cap. A task passes when all three of its runs pass.

Notice what the criterion leaves out. The agent’s closing sentence is never consulted. A run that hits the cap counts as stopped, which is a failure with its own label, because the stop condition fired before the goal was met. Chapter 3’s advice on the cap applies from the first line of your translation: “Set the cap before you write anything else.”

This criterion also answers a common worry, that unit tests with fixed input and output pairs cannot grade a model whose wording varies. The check inspects the state a run leaves behind, so two runs that word things differently can both pass.

Requiring all three runs is the strict reading. The book calls it passk, “the probability that all k attempts succeed,” and sets it beside the forgiving pass@k, where one success in k is enough. The glossary entry on pass@k and passk has both definitions. A learner should choose the strict one, since it is the number that describes an agent left to work alone.

The criterion grades the folder, and it stays out of the agent’s route. Whether the agent listed files before reading them is its own business. There are cases where the path itself needs grading, such as forbidden actions, and the post on agent trajectory evaluation and when to grade the path covers them.

Why three runs each, and what do you log?

Three runs each, because an agent’s output varies from run to run and a single run shows you one sample. Each run starts from a fresh copy of the folder, so the first project is 20 × 3 = 60 runs.

Three is the book’s ritual, borrowed from practice: “give the agent the same task three times, read all three outputs end to end.” The arithmetic shows why it changes the picture.

If each run of a task passed independently with probability 0.9 (an illustrative figure), all three would pass with probability 0.9 × 0.9 × 0.9 = 0.729. A task that is “usually fine” fails the strict criterion about one time in four. The pass@k calculator does this for any rate and any k.

The cheapest eval you can run this afternoon.
Figure 16.2 The cheapest eval you can run this afternoon. Give the agent the same task three times, read all three outputs, and ask of each only “would I accept this?” Here runs 1 and 2 land on the same shape of answer while run 3 diverges (in accent); the spread you read by hand is the reliability envelope of the previous figure, measured with your own eyes. Repeated trials, a human grader, an acceptance criterion—the whole discipline in miniature. Reuse this diagram

At least one university course grades this shape. The syllabus of the University of Michigan’s EECS 498, read in October 2026, requires an “Eval suite: at least eight tasks, each run at least three times, reported as pass rates.”

Log one line per run and save the full transcript beside it. The line holds the task, the kind, the trial number, the outcome, the passes through the loop, the tool calls, which check failed, and a one-line cause you fill in by hand after reading. Keep infrastructure failures such as timeouts and rate limits under their own status, outside the pass rate. Chapter 16’s reason is blunt: “an agent marked wrong because the harness fell over is your metric lying to you.”

The eval plan as a template

The plan below is a template in no particular tool’s syntax, filled in with this post’s example. Copy it, replace the values, and implement each part as plain code in the language you translated the agent into.

project: file agent, three tools (list_files, read_file, edit_file)
workspace: throwaway folders only; a fresh copy per run
step_cap: 20                      # illustrative; passes through the loop
trials: 3                         # runs per task

task:                             # one block per task, 20 in all
  id: find-timeout-02
  kind: find-then-change          # single-file | find-then-change | multi-file | negative
  goal: "The timeout is defined somewhere in this project; set it to 30"
  start_state: tasks/find-timeout-02/start/
  allowed_changes: [config/settings.txt]     # empty list for a negative case
  check: tasks/find-timeout-02/check          # a script; exit 0 means pass
  reference_solution: tasks/find-timeout-02/reference/   # your own edit; must pass the check

task_mix:                         # illustrative
  single-file: 6
  find-then-change: 6
  multi-file: 4
  negative: 4                     # the agent should change nothing

run_passes_when:                  # the agent's final message is not consulted
  - check exits 0 on the final state of the folder
  - no file outside allowed_changes differs from start_state
  - the loop returned finished within step_cap (a stopped run fails)

task_passes_when: all 3 runs pass

log_per_run:                      # one line per run, plus the saved transcript
  - task id, kind, trial number
  - outcome: pass | fail | stopped | infra_error   # infra_error is kept out of the pass rate
  - passes through the loop, tool calls
  - tokens in and out, if your provider reports them
  - which check failed
  - path to the transcript
  - cause, one line, written by hand after reading

report:
  - tasks passing all 3 runs, out of 20, with a 95% interval
  - runs passed, out of 60, as a second number (never as its own interval)
  - results by kind
  - every failing task, with its cause
  - the versions of the prompt and tool descriptions that produced this result

The last report line is easy to skip. Chapter 16 notes that you can only attribute a change in results if you can name the change, so the prompt and the tool descriptions belong in version control next to the code.

How do you read the result?

Read the result as a range, by kind, with the failures named. Suppose (every number in this section is illustrative) that 17 of the 20 tasks pass all three runs.

Kind Tasks Passed 3 of 3 Runs passed
Single-file change 6 6 18 of 18
Find-then-change 6 5 17 of 18
Multi-file change 4 3 10 of 12
Negative case 4 3 9 of 12
Total 20 17 54 of 60

The headline is 17 ÷ 20 = 85%. Around it sits a 95% Wilson score interval, a range built so that it would contain the true pass rate in about 95 of 100 repeats of the experiment. For 17 of 20 it runs from 64.0% to 94.8%, a band 31 points wide. An honest sentence is “17 of 20, so somewhere between about two-thirds and nineteen in twenty.”

Both the interval and the comparison further down are standard statistics (they do not come from the book), and the calculator opens on these numbers.

With JavaScript on, the Eval sample-size calculator runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Eval sample-size calculator on its own page to share a result by link.

Can 20 tasks show that a change helped?

Twenty tasks can show that a change removed a large defect, and they cannot confirm a modest gain. Suppose an earlier version of your prompt had passed 14 of 20, and the edit moved it to 17. That is 70% to 85%, a gain of 15 points (still illustrative).

Switch the calculator to its third tab, which is seeded with this comparison. The two-sided p-value is about 0.26, meaning that if the edit had changed nothing, chance alone would produce a gap at least this large about one time in four (by this approximate test; with only 20 tasks the exact figure is higher). Detecting a real 15-point difference four times out of five would take about 121 tasks per version. Chapter 16 says the same in words: “a two-point shift on a fifty-task suite is well inside the noise.”

Two rules keep the report honest. Do not pool the 60 runs into one interval, because three runs of one task are not independent trials; report 54 of 60 as a second number.

And read the three failing tasks until each has a cause. In this example the negative case that failed 0 of 3 is the finding: the agent edits when it should decline. Why small sets are still worth running, and when you need hundreds, is worked through in the post on how many eval examples you need.

What does this cost, and when should you not do it?

This path costs time at the start and model calls throughout, and it is the wrong path for some readers. A tutorial puts a demo on screen the same afternoon; twenty tasks with check scripts and reference solutions are unglamorous work that comes before the agent looks any better.

The calls add up. The appendix’s worked run takes 5 passes through the loop, so 60 runs is about 60 × 5 = 300 model calls, with a ceiling of 60 × 20 = 1,200 at an illustrative step cap of 20. Chapter 16 states the general cost plainly: “Multi-trial evaluation multiplies your compute bill by the trial count.” One mitigation: build and debug the harness and the checks against a stub that returns a fixed sequence of responses, as the labs in this site’s teaching kit do, and spend real calls only on measuring.

It will also feel worse than a course while it works. That is the Deslauriers finding again, and a wall of failing checks is less pleasant than a smooth video.

Who should not start here?

Three kinds of reader should not start here: one who cannot yet program, one who needs a framework for a job next week, and one whose project has no programmatic check.

  • A reader who cannot yet program. Kirschner, Sweller and Clark (2006) argue that “minimally guided instruction is less effective and less efficient” than guided instruction, and that the advantage of guidance recedes only once learners have high prior knowledge. Follow one worked walkthrough end to end first; “just build” with no guide is the weak form of this advice.
  • A reader who needs a framework for a job next week. Learn that framework. The book concedes frameworks their place: “Build from scratch to understand. Adopt a framework to scale—and only once you can name what it is saving you.”
  • A reader whose project has no programmatic check. Prose outputs need an LLM-as-a-judge, and calibrating one is a project in itself.

Some learners advise the opposite order. One wrote, in a September 2026 thread on r/AI_Agents: “Evaluation and security matter, but they are deeper topics you can pick up later.” For security at depth, I agree. For a first check on a first agent, I do not, and the book grants the limit of my side too: “There is a scale below which full statistical rigor costs more than the occasional production failure it would prevent.”

How do the learning paths compare?

The three common paths differ less in what they teach than in what you hold when you finish. The table describes each as fairly as I can; the course row rests on the two courses named above, as they stood in October 2026.

Path What you do What you hold at the end Where it is the right choice
Course-first Follow a structured sequence of lessons and labs Broad vocabulary, completed labs, sometimes a certificate or a score on the course’s own benchmark You want a map of the field, or you learn best with a schedule
Framework-first Learn one library’s abstractions and build with them A working app, and fluency in that library A job or team already uses it, or you need state, retries and tracing now
Build-and-verify Translate a minimal agent, write 20 tasks, run each three times A small agent you understand, a pass rate with its interval, and failures with causes You can program, your task is checkable, and you want to know why it fails

These paths combine. A walkthrough is guidance for the build, and guidance is what the Kirschner paper says novices need. A framework is a sensible second step once you can name what it saves you; Chapter 3 says a framework “can make an agent more dependable and cannot make it smarter.”

What does the evidence support?

The evidence supports a narrow claim: doing and measuring beat watching in taught courses, and at least some instructors grade the measuring. I found no primary source on what employers reward here, so I make no claim about hiring.

In one undergraduate course, 16 of 24 final projects (67%) included explicit evaluation, and the authors report: “Higher-performing projects consistently treated evaluation as a first-class design concern rather than as an afterthought” (Mello and Maher, 2026). That is one course, with no control group and a rubric that rewarded evidence.

In a meta-analysis of 225 studies of undergraduate STEM classes, students under traditional lecturing were 1.5 times more likely to fail than students under active learning (Freeman et al., 2014). Those were instructor-run courses, and nobody has measured a self-taught adult choosing between a video course and a project with checks. I am extrapolating, and you should weigh the argument accordingly.

Instructors face the same choice from the other side of the desk, and teaching AI agents to CS students raises it for a whole class at once.

The one thing to keep

The best way to learn AI agents is to finish with something that can be checked. Twenty tasks, three runs each, a log and a range: that is a small result, and it is yours to defend because you can say what the 17 means and why the 3 failed. When a new failure turns up, it becomes task 21.

Chapter 16 closes on the habit: “The teams that ship reliable agents started measuring early and never stopped; the elaborateness of the machinery matters far less.” The loop is in Chapter 3, “The Agent Loop”, the measuring in Chapter 16, “Evaluating Agents”, and the minimal agent in Appendix A, all in the full book. Chapters 1 and 2 are free to read; see the formats.

Questions readers ask

What is the best way to learn AI agents?
Build one small agent whose result a program can check, and write the checks while you write the loop. Follow one annotated walkthrough for the build, then write about twenty tasks of your own, run each three times, and report how many passed with the interval around that number.
Should I learn an agent framework first?
Usually no, unless a job or course next week requires a specific one. Chapter 3 of AI Agents, Engineered gives the rule: "Build from scratch to understand. Adopt a framework to scale—and only once you can name what it is saving you." A framework supplies plumbing around the same four parts you would write yourself.
What should my first AI agent project be?
One where you own both ends and the outcome is state a program can inspect. A file agent with three tools (list, read, edit) working in throwaway folders qualifies: after each run, a script checks the files. Avoid first projects whose output is prose, or whose environment resists you with logins and captchas.
How do I know my agent works?
The agent's own report is not evidence. Write about twenty hand-checked tasks, each with a starting state, a programmatic check on the final state and a reference solution. Run each task three times and report the tasks that passed every run, with a confidence interval.
Do I need machine learning or math to learn AI agents?
No, for building and checking a first agent on a hosted model. You need to program, to call an API, and to do arithmetic on pass rates. Training models and research are a different path, and the mathematics that roadmaps list belongs to that path.

Sources

  1. Scott Freeman et al. (2014). Active learning increases student performance in science, engineering, and mathematics
  2. Louis Deslauriers, Logan S. McCarty, Kelly Miller, Kristina Callaghan, Greg Kestin (2019). Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom
  3. Paul A. Kirschner, John Sweller, Richard E. Clark (2006). Why minimal guidance during instruction does not work
  4. C. Mello, J. Maher (2026). FairLLM: A Pedagogical Framework for Teaching Agentic Large Language Model Systems in an Undergraduate Artificial Intelligence Course
  5. University of Michigan (2026). EECS 498 Applied Agentic Software Engineering, syllabus
  6. Anthropic (2026). Demystifying evals for AI agents
  7. Thorsten Ball (2025). How to Build an Agent
  8. Thomas Ptacek (2025). You Should Write An Agent
  9. Hugging Face. AI Agents Course, unit 0
  10. Microsoft. AI Agents for Beginners
  11. Aqsa Zafar (2026). How to Learn AI Agents in 2026: A Step-by-Step Roadmap
  12. Wanting to get into AI agent dev but completely overwhelmed (r/AI_Agents thread)
  13. u/honeyplumeRomi (2026). Looking for ideas on where to start with AI agents (r/AI_Agents thread)
  14. u/lberdy (2026). How can I effectively learn and master AI agents? (r/AI_Agents thread)
  15. hhimanshu (2025). Hacker News comment on evaluating LLM output