Home / Blog / Teaching AI agents / Learn AI Agents From Scratch: The Loop, Then To…

Teaching AI agents

Learn AI Agents From Scratch: The Loop, Then Tools, Then Evals

Learn AI agents from scratch in three stages: the loop, then tools, then evals. Get the build, the exit check and the too-early signs for each stage.

By Enrique Gutiérrez · Published · 21 min read

To learn AI agents from scratch, take three subjects in order: the loop, then tools, then evals. Each stage gives you the skill you need to debug the next one. Write the loop and you can read a run; shape a tool and you can repair a run; write evals and you can tell whether the repair held.

For each stage, this page gives the build, the one thing you must be able to explain before moving on, the checks that end the stage, and the mistake that shows you left too early. Then it takes the reverse orders in turn: framework first, multi-agent first, evals last or never.

The order is my proposal. It follows the chapter order of the book this site belongs to, and nobody has tested it against another order.

Why does the order matter when you learn AI agents from scratch?

The order matters because each of the three subjects is the instrument for debugging the next. A tool problem shows up as entries in the message history, so you need the loop to see it. An eval failure has to be traced to a cause, so you need the loop and the tools to read it.

In public, the question arrives as one about products. A newcomer on Hacker News in December 2025 asked: “what is the current go-to standard in industry when it comes to tools/frameworks for building agents ‘from scratch’?” The first reply moved the question down a level: “I’d focus on the underlying service provider APIs first, in the same manner as any other cloud resource.”

Chapter 3 supplies the next step. It ends its from-scratch build with the observation that connects the first two stages: “the intelligence lives in the model and in the words you wrote describing the tools, and none of it lives in the control flow.” Once the loop holds no mystery, the words describing the tools are the next thing to study.

Chapter 5 then hands tools over to measurement: removing a tool is a legitimate design move, it says, and knowing when to remove one takes evidence, which it leaves to Chapter 16.

“The loop first” does not mean a loop with no tools, because such a loop has nothing to do. It means the tools stay trivial while the loop is the subject. Likewise, every stage has a check, and “evals third” means that Stage 3 is where checking becomes a subject of its own.

Which stage are you in?

You are in the first stage whose question you cannot answer about your own agent, with its code and one saved run open in front of you. What you have read or watched does not place you. Take the three questions in order and stop at the first one that defeats you.

  • Loop. What exactly does the model receive on the third pass of that run, item by item, and what are all the ways the run could have ended?
  • Tools. Find one call in your transcripts where the model used a tool wrongly. Which surface misled it: the definition, an earlier result, or an error message?
  • Evals. What does your pass rate count: which tasks, how many runs of each, checked against what?

Each question is the quickest test of its stage, and the stage’s checklist, further down, is what decides. If a question passes and that list has an unticked item, you are still in that stage.

A reader with a framework demo can still pass the first question: log the full request the framework sends on each pass, and answer from the log. The second question needs a wrong call to examine, and if your transcripts hold none, you have not yet run enough tasks to leave Stage 2. For the third, “it works when I try it” is one run.

The site’s AI agents learning roadmap in four stages places readers by the same rule and covers more ground: a sandbox, a calibrated judge and a week of operation. This page takes the first three subjects more slowly and argues for their order.

Stage 1: what do you build to learn the loop?

Build the smallest agent there is: a function that calls a model, three trivial tools, a list that records everything, and a loop with a step cap. The subject of this stage is the list and the loop. Keep the tools dull on purpose.

Chapter 3 gives the inventory: “An agent has exactly four parts, and you already understand every one of them.” The four are a model client, tools, the message history and the loop. Everything except the model is the harness, which Chapter 3 defines as “everything you build around the model to turn it into a working agent.”

One pass through the loop has four beats circling a growing history.
Figure 3.2 One pass through the loop has four beats circling a growing history. The model supplies a single beat—reason, in accent: it weighs the desk and decides the next move. Your code supplies the other three—lay the desk, run the tool, record the result—and the result is appended to the history before the desk is laid again. The model’s own judgment that the work is done is the one exit that leaves the cycle. Reuse this diagram

There are two ways in. Translate the pseudocode of Chapter 3, The Agent Loop (in the full book) into the language you write, or build an AI agent from scratch in one file, a post that prints a complete Python program with a scripted stand-in for the model, so it runs with no account.

To see the cycle before you type it, watch the three-minute agent loop explainer or play the harness’s part in Run the loop, a scripted run in the browser.

Give the agent one task whose result a program can check, such as changing a value in a file. That single check is all the evaluation Stage 1 needs.

What must you be able to explain before leaving Stage 1?

You must be able to say what the model receives on any pass, and why that is all it knows. The model keeps nothing between calls. Your code resends the whole history every time, so the history is the agent’s entire memory of the run.

Chapter 3 puts the consequence in one sentence: “whatever your code appends is the agent’s memory, and whatever it fails to append never happened, in the strictest sense available.” It also names what follows from a missing line: “Drop the tool result and the model, blind to the outcome, requests the same call again.”

The second half of the explanation is the exits. A run can end because the model says it is finished, because a budget fired, or because an error halted it (the kind that shows the run itself is unsafe, such as a refused permission; ordinary tool failures go back to the model). Chapter 3 is blunt about relying on the first alone: “a loop whose only exit is the model’s judgment is a loop with no guaranteed exit at all.” A budget is the simplest stop condition, and the chapter’s instruction is “Set the cap before you write anything else.”

  • I can list, in order, every item my code sends to the model on pass 3 of a saved run, and I checked my list against the logged request.
  • I removed one of the two appends, predicted what the run would do, and the run did it.
  • My loop has a step cap that I set before the first real run, a run that reaches it returns “stopped” with the transcript saved, and for any saved run I can say which exit it took.
  • A tool that fails returns its failure to the model as a result, and my program does not crash.
  • For one finished run, I can point to something outside the model’s final message that shows the task was done.

What shows you left Stage 1 too early?

The sign is that you respond to a misbehaving run by rewording the prompt before you have read the history. A reader who skipped this stage sees an agent “going in circles” and concludes that the model is weak. A reader who finished it opens the history and looks for the result that is missing or useless.

A second sign is that you cannot say why a run ended. If “finished” and “stopped” look the same in your output, every later count will mix the two.

Stage 2: what changes when tools become the subject?

In Stage 2 the loop stays as it is and you rewrite the tools as text the model reads. A tool has three surfaces that reach the model: its definition, what it returns, and how it fails. Chapter 5 says of all three: “Every one of these channels is a prompt, whether or not you wrote it as one.”

The build is small. Add one tool that does something your three trivial ones could not, such as a search across files or a call to a service you control. Then write its three surfaces with care, and rewrite the surfaces of the tools you already had.

The anatomy of a tool across the boundary between your process and the model’s context.
Figure 5.3 The anatomy of a tool across the boundary between your process and the model’s context. Three parts cross to the model’s side and steer it like prompt text—the description above all, in accent, the highest-leverage surface you write. The function itself never leaves; only its results and errors travel back, as still more text the model has to read. Reuse this diagram

The set matters as much as each member. Chapter 5 opens with the stakes: “The set of tools you expose is the complete inventory of what your agent can ever do.” It calls that inventory the action space, and it warns against growing it casually: “treat every addition as a cost that must argue for itself.”

Before you give the agent a tool that can do damage, put a boundary around where it runs. That is the sandbox work in Stage 2 of the roadmap linked above, and I do not repeat it here.

What must you be able to explain before leaving Stage 2?

You must be able to say, for any tool, everything the model can know about it and where it learns each thing. The model sees a name, a description and an argument schema on every pass. It sees results and errors when they happen. It never sees your function.

That explanation gives you a debugging order, which Chapter 5 states as a rule: “when a model misuses a tool, suspect the contract before you blame the intelligence.” Its example is a web-search tool that kept adding the current year to its queries. The chapter’s source, Ken Aizawa’s “Writing effective tools for agents” (Anthropic, 2025), reports that the repair was a better tool description.

The method for finding such defects is Chapter 5’s: “expose a tool, run real tasks, and read the transcripts.” You can read transcripts because of Stage 1. For a second reader, paste a definition into the tool schema linter.

  • For each tool, I can show the name, description and argument schema exactly as they appear in a logged request.
  • I changed only a description or a parameter name, and saved transcripts from before and after show the model’s call change.
  • Every error my tools return says what went wrong and what a correct call looks like.
  • No tool returns more than the next decision needs, and a result that was cut short says so.
  • I can name the task each tool exists for, and no two tools could reasonably answer the same request.

What shows you left Stage 2 too early?

The sign is that you fix tool misuse by adding a sentence to the system prompt, or by adding another tool. Both leave the misleading surface in place. The first also changes a text that every pass reads.

A quieter sign is a tool that returns everything it has, which the model must then read on every later pass.

Stage 3: what is the smallest eval that counts?

The smallest eval that counts is one task, run three times from the same starting state, with all three results read to the end and a program checking the final state. Chapter 16 gives it as a ritual: “give the agent the same task three times, read all three outputs end to end, and ask of each one only ‘would I accept this?’”

The cheapest eval you can run this afternoon.
Figure 16.2 The cheapest eval you can run this afternoon. Give the agent the same task three times, read all three outputs, and ask of each only “would I accept this?” Here runs 1 and 2 land on the same shape of answer while run 3 diverges (in accent); the spread you read by hand is the reliability envelope of the previous figure, measured with your own eyes. Repeated trials, a human grader, an acceptance criterion—the whole discipline in miniature. Reuse this diagram

Then make “would I accept this?” precise, and hand the judgment to code wherever code can decide. Chapter 16’s rule is to “grade the state, not the prose”: open the file, query the record, run the test.

From there, grow the ritual into an eval set, a curated collection of tasks, each with a way to grade the result. For this stage three tasks are enough, provided one of them is a task where the right result is no change. Chapter 16 asks for those “negative cases” by name, and for a reference solution per task that proves the task can be passed.

Three tasks end the learning stage because they make you use every part of the instrument once. They are not a measurement of your agent. The 20-task set, its template and its interval arithmetic are in the post on the best way to learn AI agents. The roadmap’s Stage 3 exit asks for 20 tasks or more, so finishing here leaves you at the start of the roadmap’s Stage 3.

What must you be able to explain before leaving Stage 3?

You must be able to say what your number is a number of. A pass rate has a numerator, a denominator, a count of runs per task and a definition of “pass.” Leave one out and two people will compute different numbers from the same runs.

You must also be able to say why a few runs prove little. Suppose (an illustrative figure) that an agent passes a task on 70% of runs. Three passes in a row then happen with probability 0.7 × 0.7 × 0.7 = 0.343, so about one time in three, an agent that fails three runs in ten shows you a clean sweep.

The pass@k calculator does this for any rate, and the glossary entry on pass@k and passk defines the strict and the lenient reading.

  • I ran one task three times, each from a fresh copy of the starting state, and read all three transcripts to the end.
  • My acceptance criterion is written down, and a program checks it against the final state and ignores the agent’s final message.
  • I have at least three tasks, one of them a task where the right result is no change, and each has a reference solution that passes its own check and was run three times.
  • I report a count with its denominator: tasks that passed every run, out of tasks, with the runs per task stated.
  • After one change to a prompt or a tool description, I reran every task and read each failure until it had a cause.

What shows you left Stage 3 too early?

The sign is the sentence “it feels better since I changed the prompt.” Chapter 16 describes where that leads: you fix refunds, cancellations break, you fix cancellations, refunds wobble. Its explanation is that “a prompt is a single shared artifact read at every step, so any edit is global.”

The other sign is adding a part (a framework, a memory store, a second agent) and being unable to say whether it helped. Chapter 16 has a four-word line for this: “Intuition proposes; measurement decides.”

A worked example: how does one failure look at each stage?

One failed run needs all three skills in turn, which is the argument for the order in miniature. What follows is a worked example with illustrative details; the cap of 10 is illustrative too.

A file agent gets the goal “The timeout is defined somewhere in this project; set it to 30.” The run ends after ten passes. Each pass shows the same request, to read a file called settings, and the same result: Error: 2.

Stage 1, the reading. The run ended by its step cap, so it stopped and did not finish. The result was appended every time, so this is not the dropped append. The model repeated itself because the result in the history gave it nothing to act on.

Stage 2, the repair. The misleading surface is the error. Rewrite it to say what went wrong and what would work: no file has that name, here are the files in the folder, pass a path relative to the project root. Chapter 5’s standard for an error is “what happened, what a correct call looks like, what to try next.”

Stage 3, the check. Run the task three times before the change and three times after, from fresh copies. Say the result is 0 of 3 before and 3 of 3 after, where a pass means the file holds the new value and no other file changed. That is a large effect, and three runs can show one. By the 0.343 arithmetic above, 3 of 3 is still weak evidence about the rate, so the task joins the set and runs again with every later change.

Skip Stage 1 and this failure reads as “the model is confused.” Skip Stage 2 and the fix is a longer system prompt. Skip Stage 3 and the fix is declared a success after one run.

Which stage does this symptom belong to?

Each beginner symptom points back to the stage whose question was skipped. The table lists fifteen, five per stage, with what to check first. With scripts on you can narrow it to one stage; with scripts off, every row stays on the page.

What you see What to check first Stage
The model asks for the same call again, and the last result is absent from the history Both appends: the model’s request and the tool’s result Loop
The run never ends, or ends only when you interrupt it Whether a step cap exists and is enforced by your code Loop
The program crashes when a tool raises an exception Whether failures are returned to the model as results Loop
The run says “finished” and the file is unchanged What, other than the model’s message, ends the run Loop
You cannot say what was sent to the model on a given pass Whether you log the full request, with or without a framework Loop
The model picks the wrong one of two similar tools Whether the two overlap, and whether one can go Tools
The model passes a name where an identifier was expected The parameter’s name, type and description Tools
The model repeats a failing call after reading the error Whether the error says what a correct call looks like Tools
The agent loses track after a tool returns a very large result What the tool returns, and whether it truncates and says so Tools
Your fix for a misused tool was a new sentence in the system prompt The tool’s definition, result and error, in that order Tools
You fixed one case and an older one broke Whether a fixed set of tasks reruns after every change Evals
“It worked when I tried it” How many runs, of how many tasks Evals
A pass rate is quoted with no denominator Tasks, runs per task and the definition of a pass Evals
Your check reads the agent’s final message Whether a program inspects the final state Evals
You added a framework, memory or a second agent and cannot say whether it helped Whether you have the same tasks measured before and after Evals

Two rows look alike and are not. “The run says finished and the file is unchanged” is a Stage 1 problem when it concerns how one run ends, and a Stage 3 problem when your grader believes the same message. The repair in both is a program that looks at the state.

Why not start with a framework?

Start without a framework because a framework assembles the request for you, and the request is what Stage 1 teaches you to read. A widely cited vendor essay, Schluntz and Zhang’s “Building effective agents” (Anthropic, 2024), says frameworks “often create extra layers of abstraction that can obscure the underlying prompts and responses, making them harder to debug.”

The same essay names the failure that follows: “Incorrect assumptions about what’s under the hood are a common source of customer error.” Its advice to developers is to “start by using LLM APIs directly.” A learner asking for resources on Hacker News in April 2026 got the same advice from a commenter in fewer words: “Use the APIs directly and avoid as many abstraction layers as possible.”

Chapter 3 explains why the detour through raw code is short. A framework is plumbing around the same four parts, and “When a framework’s agent misbehaves, there are only four places to look.” Someone who has built the four knows where those places are.

The other side deserves its hearing. A reply in the December 2025 thread quoted at the top of this page advised that newcomer that a popular framework “is a decent place to start if starting from scratch,” and a framework is the right first step when a job or a team already depends on one. Chapter 3’s rule covers both cases: “Adopt for a named need, never for the feeling that serious systems use frameworks.”

Why not start with multi-agent systems?

Start with one agent because every agent in a multi-agent system is the same loop, so a design with five of them asks you to debug five of the thing you have not yet learned to debug once. Thomas Ptacek’s from-scratch essay “You Should Write An Agent” (2025) describes a sub-agent as “just a new context array, another call to the model.”

The costs multiply as well. Anthropic’s account of “How we built our multi-agent research system” (June 2025) reports from its own data that “agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats.” Dividing the two figures, 15 ÷ 4 = 3.75, so in that data a multi-agent system used roughly four times the tokens of a single agent. The division is mine, and the figures are one vendor’s, dated.

Practitioners disagree about when several agents pay. Walden Yan’s “Don’t Build Multi-Agents” (Cognition, 2025) calls the parallel sub-agent architecture “very fragile,” while the Anthropic account describes a product built on one. A learner cannot referee that from a demo. With a Stage 3 set, you can run both designs on your own tasks and count.

Why not leave evals for last, or put them first?

Do not leave evals for last, because everything you add after Stage 2 is a claim that the agent got better, and without a set of tasks the claim is never tested. Anthropic’s “Demystifying evals for AI agents” (2026) says it directly: “Evals get harder to build the longer you wait.”

Hamel Husain’s “Your AI Product Needs Evals” (2024) makes the strong version of the claim from his own experience of such products: “unsuccessful products almost always share a common root cause: a failure to create robust evaluation systems.” That is one practitioner’s observation about products; I extend it to learning by analogy only.

Learners ask, and one thread I read got no answer. An Ask HN post from November 2025 asked what works for agent evals “in practice.” When I read the thread in October 2026 it held six replies, and none described an eval method; the discussion had moved to product features.

Why not first, then? In one sense evals are first: Stage 1 already ends on a programmatic check. But an eval set studied before the loop and the tools produces failures you cannot read. Chapter 16 says the discipline it would keep above all others is to “read the failures,” and reading a failure means reading a history and a tool exchange.

The same Anthropic guide grants the other side, and so do I: teams starting out “can get surprisingly far through a combination of manual testing, dogfooding, and intuition.” Evals as the third subject is a claim about study order, with a small check present from Stage 1.

What do other from-scratch paths order differently?

Three free paths that a search for “learn AI agents from scratch” returns each use a different order, and they differ most in where evaluation sits. The table records what each page said on 6 October 2026. They are examples of a category, and pages change.

Path Order of the first topics Where evaluation sits
“How to learn AI agents, from scratch”, a hosting company’s guide (Connor, 2026) Concepts, the loop, tools, memory, loop design, safety No stage. Stage 5 asks how a result is verified; the word “eval” does not appear on the page
agents-from-scratch, a 12-lesson repository Chat, system prompt, structured output, routing, tools (lesson 5), the loop (lesson 6) Lesson 11 of 12
AI Agents for Beginners, an 18-lesson vendor course Use cases, frameworks (lesson 2), design patterns, tool use (lesson 4); multi-agent is lesson 8 No lesson has evaluation in its title
This page The loop, tools, evals Third of three, with a check in every stage

Each does something this page does not. The first guide covers the shell skills agents depend on and says plainly not to start with a framework. The repository runs on a local model. The vendor course is the broadest, and its README claims no order: “Each lesson covers its own topic so start wherever you like!”

If you are earlier than any of this, start with AI agents for beginners: what to learn first and what to skip.

What do you do this week?

This week, do the first unticked item in your stage’s checklist, and nothing from a later stage. Each list has one item that fits in a sitting.

  • Stage 1. Log the full request your code sends on each pass of one run, and compare pass 3 with what you expected.
  • Stage 2. Rewrite one tool’s error messages so that each says what a correct call looks like, then rerun a task that used to trip on it.
  • Stage 3. Run one task three times from fresh copies and read all three transcripts to the end.

When a stage’s list is complete, explain its one thing aloud, without notes, to someone who programs. Instructors grade that ability, and the post on AI agents assignments for students shows the formats they use.

The stages are a study order for the first pass. After it, the three alternate for as long as the agent lives: a failure is read, a tool is repaired, the set reruns.

What are the limits of this order?

The order is argued, not measured. I found no study that compares learners who took these subjects in different sequences, so the case rests on a dependency argument and on the book’s chapters.

It also assumes you can program, and it stops early. Context management, security, cost and operation come after these three, and the roadmap covers where they fit. Three tasks are the exit of a learning stage; they do not make an agent safe to rely on.

The reader evidence is thin: three Hacker News threads read on one day, chosen because they were about order, so they say nothing about how common the question is. And Chapters 3, 5 and 16 are in the full book, though the stages, the checklists and the symptom table work without them.

The one thing to keep

To learn AI agents from scratch is to earn three explanations in order: what the model receives, what the model can know about a tool, and what your number counts. Each one is a way to check something you would otherwise take on trust, and each makes the next one possible.

The loop is in Chapter 3, “The Agent Loop”, tools in Chapter 5, “Tools and the Action Space”, and measurement in Chapter 16, “Evaluating Agents”, all in the full book. Chapter 1 and Chapter 2 are free to read online, the teaching kit has a syllabus with labs, and you can see the formats.

Questions readers ask

How do I learn AI agents from scratch?
Take three subjects in order. First write the loop yourself, with trivial tools, until you can say what the model receives on every pass and how a run ends. Then make tools the subject: definition, result and error. Then write a small eval: a few tasks, three runs each, checked by a program.
Should I learn the agent loop or tool calling first?
The loop, with tools kept trivial. A first loop needs a tool or two to have anything to do, so the two arrive together in code. The order is about which one you study. Tool problems are diagnosed by reading the message history, and the history is what Stage 1 teaches you to read.
When should I start writing evals for my agent?
A single programmatic check belongs in Stage 1, as the thing that says a run worked. The eval as a subject, with a set of tasks, repeated runs and a count, comes third, once you can read a failing run and repair a tool. Third of three is early: it comes before frameworks, memory and multi-agent designs.
Do I need a framework to learn AI agents?
No. Chapter 3 of AI Agents, Engineered says to "Build from scratch to understand. Adopt a framework to scale—and only once you can name what it is saving you." If a job needs a specific framework now, learn it, and log the requests it sends so that you can still answer the Stage 1 question.
How long does it take to learn AI agents from scratch?
I know of no measurement, so this page gives no schedule. Each stage ends when its checks pass, whatever the hours. One published guide estimates an evening for a first loop and about two months of evenings for a bounded, useful agent; that is its author's estimate, given without data.

Sources

  1. Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents
  2. Ken Aizawa (Anthropic) (2025). Writing effective tools for agents
  3. Anthropic (2026). Demystifying evals for AI agents
  4. Anthropic (2025). How we built our multi-agent research system
  5. Walden Yan (Cognition) (2025). Don't Build Multi-Agents
  6. Thomas Ptacek (2025). You Should Write An Agent
  7. Hamel Husain (2024). Your AI Product Needs Evals
  8. Matt Connor (SSD Nodes) (2026). How to learn AI agents, from scratch
  9. pguso. agents-from-scratch (repository README)
  10. Microsoft. AI Agents for Beginners (repository README)
  11. realberkeaslan (2025). Ask HN: What framework/tool should I use for building agents?
  12. 7e10 (2026). Ask HN: Learning resources for building AI agents?
  13. akira_067 (2025). Ask HN: Agent evaluations, what is everything I should know?