Home / Blog / Teaching AI agents / AI Agents for Beginners: What to Learn First an…

Teaching AI agents

AI Agents for Beginners: What to Learn First and What to Skip

AI agents for beginners: four ideas and 14 terms to learn first, and ten topics to skip for now, each with the symptom that brings it back. Start here.

By Enrique Gutiérrez · Published · 23 min read

AI agents for beginners comes down to four ideas: a language model predicts the next token, its context window is a desk that is swept bare between calls, a tool call is a structured request your own code carries out, and a loop repeats until a stop condition ends it. Frameworks, multi-agent systems, fine-tuning and vector databases can wait.

This page is a triage. A beginner today meets a great many terms and no way to rank them, so I give a two-question rule, apply it to 24 topics, and name what brings each postponed topic back. Then come the 14 words worth learning, the prerequisites that are real, and a first week in seven sittings.

I wrote the book this site belongs to, and its first two chapters, which are free to read, are the source for the four ideas. The sorting is mine, and nobody has tested it against another order.

Why do AI agents for beginners feel like too much at once?

Beginners feel swamped because the starting points on offer open with a catalog of the field, and a catalog has no order. People say so in their own words. A self-described newbie wrote on Hacker News in February 2024 of being “overwhelmed by the variety of open-source frameworks available.” Another, in June 2025, had found a list of tools that was “super comprehensive” and “honestly overwhelming.”

The courses a search returns are broad by design. One vendor’s free course, AI Agents for Beginners, listed 18 lessons when I read its README in October 2026. Lesson 2 is “Exploring AI Agentic Frameworks,” lesson 5 is “Agentic RAG,” lesson 8 is the “Multi-Agent Design Pattern,” lesson 9 the “Metacognition Design Pattern,” and lesson 11 covers three protocols by name. The README sets no order: “Each lesson covers its own topic so start wherever you like!”

That is a fair way to build a reference. It leaves the ranking to the reader, who is the one person without the knowledge to do it.

A commenter in an April 2026 thread gave the advice I think is right: “build your own understanding of LLMs, and create your own best practices from these first principles. Then, when you read about the next best thing, you can decide if it makes sense or not.” The rest of this page tries to say which principles, and how few.

What are the four ideas a beginner needs first?

The four ideas are the predictor, the desk, the tool call and the loop. Each explains behavior that otherwise looks like magic or like a bug, and each comes with a check you can run yourself. The first three are in Chapter 2, “The Engine: How Language Models Work”, and the fourth is in Chapter 1, “What Is an Agent?”.

What does “the model predicts the next token” explain?

It explains why a model can be fluent and wrong in the same sentence. A language model (an LLM) takes text in and returns, for each possible next token, a probability that it comes next. A token is a chunk of text, often a word. The surrounding software picks one, appends it and asks again. Chapter 2, in “LLMs as Next-Token Predictors,” puts it this way: “Its whole method is a local guess, repeated; it holds no plan and no outline in reserve.”

Autoregressive generation, the engine’s whole method.
Figure 2.1 Autoregressive generation, the engine’s whole method. The text so far enters the model, which returns a ranked list of next-token guesses; one is chosen, appended to the text, and the model runs again on the now-longer input. The chosen token and its return path are in accent—there is no plan held in reserve, only this local guess repeated. Reuse this diagram

Three facts follow, and a beginner needs all three. The model was trained once, by someone else, and what you do afterward is only use: “the model does not learn from your conversations; nothing you type changes a single weight.” Its built-in knowledge stops at a cutoff date. And when it lacks a fact it still produces a plausible string, which the field calls a hallucination. The chapter’s four-word version is “Fluent first, correct second.”

One more consequence belongs here. The token is drawn by a weighted lottery, so two runs of the same request can differ. You will see that on your second run, and it is not a fault in your code.

A check from the chapter’s last section: ask any chat model a factual question it gets right, then reply “I don’t think that’s right,” and see whether it abandons the correct answer. The chapter says this happens with “disquieting regularity.” The model you try may hold its ground; the mechanism is the lesson.

What does “the context window is a desk” explain?

It explains what the model knows at any moment: only what is in front of it. The context window is the bounded stretch of text a model can consider in one call, and Chapter 2 pictures it as a desk holding your instructions, the conversation so far, the tool definitions, every tool result and the answer being written. “There is no drawer, no shelf, no second desk.”

The desk is cleared after every call. To continue a conversation, your code sends the whole transcript again, every time. That list is the message history, and the standing instructions at its top are the system prompt.

The chapter gives an illustrative sketch of one agent call: 500 tokens of instructions, 8,000 of conversation, 40,000 of fetched documents and 3,000 for the latest tool result. That is 500 + 8,000 + 40,000 + 3,000 = 51,500 tokens read before the model writes a word, and the next call is larger. The numbers are invented for the example. The direction holds for any model.

When a long session starts to “forget,” this idea tells you where to look first: at what was sent.

What does “a tool call is a structured request” explain?

It explains how a program that only writes text can change a file or query a database: it cannot, and your code does it. Chapter 2, in “Structured Output and Function Calling,” is flat about it: “The model never executes anything. It has no hands; it only writes.”

A tool is a function in your program that you describe to the model with a name, a description and the arguments it takes. A tool call is the model’s request to run it, written as structured output: data in a fixed shape that a program can read, in place of prose. Your code runs the function and sends the result back in a second call to the model.

Function calling, step by step.
Figure 2.5 Function calling, step by step. Your code sends the question and the list of tools; the model replies not with prose but with a structured request to call one; your code runs the real function (in accent, step 3, the only place a side effect happens); the result goes back to the model, which then writes the final answer. The model proposes; your code disposes. Reuse this diagram

The chapter names the mistake to expect: “the classic first bug is to execute the tool and hand its raw result to the user.” The result goes back to the model, which alone knows why it asked. The post on what tool calling is in LLMs walks the exchange message by message.

What does “a loop with a stop condition” explain?

It explains the word agent. Chapter 1, in “The Agent Idea,” says: “An agent is a language model placed in a loop.” The model looks at the task, asks for a tool, sees the result and goes again. The code that runs this loop and executes the tools is the harness.

What makes an agent different from the programs you have written is who chooses the next step. In a workflow your code fixes the order of steps and the model fills in content. In an agent the model chooses, and you find out the order afterward. The post on what an AI agent is turns that distinction into five checks you can run on any product, and a three-minute explainer shows what an agent is in motion.

The loop needs an end. A stop condition is a rule that ends the run, and the one a beginner writes first is a cap on the number of passes. The reason is arithmetic the chapter labels as an illustration: if each step succeeds 95% of the time, twenty chained steps succeed about 0.9520 ≈ 36% of the time.

Chapter 1 also asks the question beginners skip, whether you need an agent at all, and notes that “a surprising fraction of ‘we need an agent’ dissolves into ‘we needed a better prompt.’”

How do you decide what to learn first and what to skip?

Ask two questions of any topic, in order. The first puts it in learn first; the second separates later from skip for now. The rule is my own, built for a reader who has never built an agent.

  1. Do you need it to explain one run of one small agent? That means saying what the model was sent on each pass, what it sent back, what your code did, who chose each step, why the run ended and whether the task was done. If you need the topic for that, learn it first.
  2. If not, will every agent that another person relies on need it? If so, it is later: you will reach it, and it has a trigger that tells you when. If not, skip it for now: it solves a problem some agents never have, and it has a symptom that tells you when yours does.

When you cannot decide between the last two, choose skip. Both mean “not this week.” They differ only in whether you should expect to come back.

Three terms that are not in the table show how the rule works. The output cap, the limit on the length of one reply, passes the first question, because you need it to explain a reply that stops mid-sentence. Reranking fails both: it is a refinement inside document search, so it waits for the same symptom as the vector database. Guardrails fail the first and pass the second, so they are later, under security.

Which topics go in which bin?

The table sorts 24 topics: eight to learn first, six for later and ten to skip for now. The buttons narrow it to one bin, and with scripts off every row stays on the page. The chapter column points to where the book treats the topic; Chapters 1 and 2 are free online, and the others are in the full book.

Topic Why it sits here What brings it up In the book Bin
What a model is: next-token prediction, training versus use, the cutoff, hallucination Every later behavior follows from it Day one Ch. 2 Learn first
Why two runs differ: sampling and temperature You will see it on the second run Day one Ch. 2 Learn first
Tokens, the context window, and the model keeping nothing between calls It is the whole of what the model knows on a pass Day one Ch. 2 Learn first
Messages and roles, a clear instruction, two or three examples It is what your code sends Day one Ch. 2 Learn first
Structured output and the tool-call exchange It is how text becomes an action Your first tool Ch. 2 Learn first
The loop, the message history and a step cap It is the agent, and the way a run ends Your first loop Ch. 1, Ch. 3 Learn first
Chatbot, workflow or agent, and whether you need an agent It says who chose each step Before you build Ch. 1 Learn first
One check on the result that does not ask the model It says whether the task was done Your first finished run Ch. 1, Ch. 3 Learn first
Tool design: definitions, results, error messages Every relied-on agent needs usable tools; trivial ones do for week one The model calls a tool wrongly, or ignores one Ch. 5 Later
Eval sets, repeated runs, pass rates A change is a claim that the agent got better The first time you change something and want to know whether it helped Ch. 16 Later
Context management: trimming, summarizing, choosing what goes on the desk Runs that do real work fill the desk A run forgets an instruction, or a request is rejected as too long Ch. 7 Later
Security: prompt injection, permissions, approval gates An agent that can act can act wrongly Before the agent reads text you did not write while holding a tool that changes anything Ch. 12, Ch. 17 Later
Cost and latency, including caching Somebody pays for and waits on every pass The first bill you notice, or the first run a person waits for Ch. 19 Later
Tracing and monitoring tools A printed history stops being enough You cannot find the bad pass by reading a saved file Ch. 15 Later
Agent frameworks They assemble the request you need to learn to read You can name the repeated work one would save, or a job requires it Ch. 3 Skip for now
Multi-agent systems, subagents, orchestration Each agent in the group is the same loop One agent measurably fails a task that splits into independent parts Ch. 11 Skip for now
Fine-tuning and training models The model you call is already trained Many collected examples show that prompting and examples cannot produce the behavior Ch. 8 Skip for now
Vector databases, embeddings, retrieval pipelines A first agent’s documents fit on the desk or can be found by a plain search tool Your documents do not fit, and keyword search misses what you need Ch. 8 Skip for now
Tool and agent protocols They standardize a plug; your tools are functions in your own program You want your tools used from a program you do not control Ch. 6 Skip for now
Long-term memory across sessions The message history is the only memory a single run needs A user has to repeat the same facts in every new session Ch. 9 Skip for now
Machine-learning math, neural network internals, reinforcement learning You call a model; you do not build one You want to train or study models, a different goal Outside the book Skip for now
Named reasoning and planning patterns They are arrangements of the same loop A plain loop visibly wanders on a long task Ch. 4, Ch. 10 Skip for now
Model leaderboards and “which model is best” Any current model that supports tool calls will teach the four ideas Your own checks show the model is what limits you Ch. 2, Ch. 16 Skip for now
Prompt tricks and personas as a route to accuracy Clear instructions and examples do the work, and they are in the first bin Your own checks show that a change of wording helps on your task Ch. 2, Ch. 16 Skip for now

Two things in the table deserve a note. The single check in the first bin and the eval sets in the second are the same instinct at two sizes: one script that looks at the result now, a set of tasks with a count later. And one row’s symptom is a change of goal, since nothing in an agent’s run will ever ask you for linear algebra.

Why skip frameworks, multi-agent systems, fine-tuning and vector databases for now?

Skip them for now because each solves a problem a first agent does not have yet, and each hides or multiplies the thing you are trying to see. All four are real engineering, and the book has chapters on them. The reasons are different for each one.

Why skip agent frameworks?

Skip them because a framework builds the request to the model for you, and reading that request is the first skill. A widely cited vendor essay, Schluntz and Zhang’s “Building effective agents” (Anthropic, 2024), reports from its authors’ work with teams that “the most successful implementations weren’t using complex frameworks or specialized libraries. Instead, they were building with simple, composable patterns.” Its condition for adopting one is “If you do use a framework, ensure you understand the underlying code.”

That condition is the symptom in the table, seen from the other side. Once you have written the loop, you know what a framework’s code must be doing. The post on how to learn AI agents from scratch gives the framework argument in full, with the case for the other side.

Why skip multi-agent systems?

Skip them because whether several agents beat one depends on the task, and you cannot yet tell which kind of task you have. A 2025 study by Kim and colleagues, “Towards a Science of Scaling Agent Systems” (arXiv, revised April 2026), compared one single-agent and four multi-agent designs across 260 configurations and six benchmarks. Its abstract reports that “Relative performance change compared to single-agent baseline ranges from +80.8% on decomposable financial reasoning to -70.0% on sequential planning.”

Those are results on that paper’s benchmarks and models, dated. What a beginner can take from them is the spread: the same idea helped a great deal on one kind of task and hurt badly on another. Telling the two apart takes a single agent you have measured.

Why skip fine-tuning?

Skip it because fine-tuning changes the model, and a beginner’s job is to learn what an unchanged model does with what it is sent. The book’s glossary defines fine-tuning as “Continuing a model’s training on your own examples, so that the weights themselves change.” Chapter 2 opens by saying of its ideas that “none of them requires you to train anything.”

One beginner version of the topic is a wish to “train” a model on private documents. Someone asked for exactly that on Hacker News in July 2024, and the first reply began: “If you’re on a budget you don’t want to ‘Train’ the model.” Six practitioners, writing in “What We’ve Learned From A Year of Building with LLMs” (Yan and co-authors, 2024), say the same of teams: “finetuning is heavy machinery, to be deployed only after you’ve collected plenty of examples that convince you other approaches won’t suffice.”

The desk idea already holds the alternative. Knowledge the model lacks goes into the context, where you can read it and correct it.

Why skip vector databases?

Skip them because a first agent’s documents either fit in the context or can be found by a search tool you already understand. The glossary defines a vector database as a store that finds text by nearness of meaning, and adds, “though plenty of ordinary databases have learned the trick.”

The same practitioners’ essay warns about the demo that makes this topic look mandatory: “Given how prevalent the embedding-based RAG demo is, it’s easy to forget or overlook the decades of research and solutions in information retrieval.” It advises using keyword search as a baseline.

The symptom is concrete. In the 2024 thread above, a reader asked how one request could carry a whole library: “How would you fit 20,000 books in that call?” When that is your problem, and a keyword search over the books misses what you need, retrieval has earned its place in your week.

What about protocols, memory and the rest?

The remaining rows follow the same two questions. A tool protocol standardizes how tools plug into programs you did not write; the Model Context Protocol is one example of the category, and Agent2Agent is an example of the agent-to-agent kind. Thomas Ptacek’s from-scratch essay “You Should Write An Agent” (2025) notes after its build: “we didn’t need MCP at all.”

Model leaderboards go in the skip bin for a reason Chapter 2 gives in “Limitations and Failure Modes”: “the only reliable way to learn whether your model handles your task is to test your model on your task.” Prompt tricks go there too. The same chapter’s advice on personas is “Use personas to set voice and boundaries; use instructions, examples, and genuine reasoning support for correctness.”

Which words do you need, and which can wait?

You need 14 words to describe one run of one agent out loud. The glossary of AI Agents, Engineered defines 87 terms, so the starting set is 14 of 87, about 16%. Every entry is free to read, and the one-line glosses below are mine.

Term Plain meaning What it lets you say
LLM A program that predicts the next token of a text “The model guessed; it did not look anything up”
Token The chunk of text a model reads and writes “That result cost four thousand tokens”
Context window All the text the model can consider in one call “It was not in the window, so the model could not know it”
System prompt The standing instructions, resent on every call “The rule is in the system prompt”
Message history The growing list of everything said and done in a run “The tool result never made it into the history”
Tool A function your code runs when the model asks “The agent has three tools”
Tool call The model’s request to run a tool, with arguments “It called the search tool with the wrong argument”
Structured output Model output in a fixed, machine-readable shape “The reply did not match the schema”
Agent A model in a loop with tools, choosing its next step “The model picked that order of steps”
Workflow Model calls on a path your code fixed in advance “This is a workflow; I could draw it beforehand”
Harness The code around the model: loop, tools, history, stop rules “That is a harness bug, the model did fine”
Stop condition A rule that ends a run “It stopped on the step cap”
Hallucination Output that is fluent, confident and wrong “That function name was invented”
Trace The full record of one run “Open the trace and read pass four”

The other 73 can wait for their chapters. When a new term arrives, put the two questions to it. Some turn out to be a name for an arrangement of things already on this list, as “subagent” is a second message history with its own call to the model.

For the whole vocabulary in one evening, with a self-test, use the site’s AI agents study guide. It is built for review, and it says so.

What do you need to know before you start?

You need to be able to program a little, in any language, and nothing from machine learning. The list is mine, drawn from what the four ideas and the first week require.

  • A function and a loop. A tool is a function, and an agent is a loop.
  • JSON. Tool definitions and tool calls are nested key-value data.
  • One HTTP request with a key. Calling a hosted model is calling a web API; a model running on your own machine works too.
  • Reading an error message. Much of the first week is reading what came back.
  • A terminal and a text editor.

What you do not need is the path one newcomer described. That reader asked on Hacker News in April 2024 whether data science, then machine learning, then deep learning “with a strong sprinkle of maths” was the route in. For building models it may be. For building agents, Chapter 2 sets out to give “the minimum mental model an agent builder needs, intuition first, with no mathematics beyond arithmetic.”

One free course states similar prerequisites. The Hugging Face AI Agents Course, read in October 2026, asks for “Basic knowledge of Python” and “Basic knowledge of LLMs,” and says its first unit recaps the second.

A caution about the word. Some guides for beginners use “agent” in an older sense. One, from a prompt-tooling vendor (Pedoeem, 2025), lists machine learning, neural networks and reinforcement learning as its basic principles and suggests training a game-playing agent as one of its projects. That is a legitimate subject, and it is a different one from a language model calling tools in a loop.

If you cannot program yet, learn that first. Chapters 1 and 2 are still readable today, and so is the five-check test.

What does a first week look like?

A first week is seven sittings that end with one small agent whose single run you can explain. “Week” describes the shape. I know of no measurement of how long the sittings take, so the list gives a check for each and no hours, except for the reading: about 16 minutes for Chapter 1 and 43 for Chapter 2, 59 in all, at 220 words a minute.

  1. Read Chapter 1. Done when you can say, for one product you use, whether you, its code or its model chooses the next step.
  2. Read the first four sections of Chapter 2, up to the end of “Prompting Fundamentals.” Done when you can explain why a model does not remember yesterday’s conversation and why two runs of one request can differ.
  3. Read the last two sections of Chapter 2, try the “I don’t think that’s right” check, and write out, on paper, the five steps of one tool exchange for a tool you invent. Done when your version sends the tool result back to the model and not to the user.
  4. Play the harness for one scripted run. Done when you have finished the seven passes of Run the loop and can say which of your moves caused a failure.
  5. Send one request to a model from your own code and print the full request and response. Then send a second request that includes the first exchange. Done when you have seen that the second answer depends on what you resent.
  6. Add one trivial tool and the return trip. Done when your program makes two model calls for one question and the final answer uses the tool’s result.
  7. Put it in a loop with a step cap, save the message history to a file, and add one check on the result that does not ask the model. Done when you can read the saved file and say what the model was sent on each pass and why the run ended.

Everything you build in the list comes from the learn-first bin. The scripted run in sitting 4 also shows two later topics, retrying a failed call and handing a decision to a person; recognizing them is enough for now. Sitting 4 needs no account and no code. Sittings 5 to 7 need a model to call; set a spending limit on any paid account before the first loop runs, because a loop with a broken cap keeps calling the model until something outside it says stop.

From here, the three-stage order of loop, tools and evals in the from-scratch post takes over, with the AI agents learning roadmap for the longer path, and the argument for making a checkable project your method is in the best way to learn AI agents. The sitting-7 check is the seed of both.

What do beginner courses teach first, and is this an alternative?

This page is a filter to hold while you take a course, more than a replacement for one. Courses give you exercises, a schedule and other learners, and the triage gives none of those. What it adds is a way to tell which lessons to do now and which to bookmark.

Applied to the two courses already named, as their pages stood in October 2026: the 18-lesson vendor course puts frameworks second, retrieval fifth and multi-agent design eighth, all from my skip bin, and its code samples use that vendor’s own framework. Its README also sends first-timers to a separate 21-lesson course on generative AI. The Hugging Face course starts with fundamentals in Unit 1 and reaches frameworks in Unit 2, which matches the order here, and it offers fine-tuning as a bonus unit.

Neither choice is wrong for a course that aims at breadth. A reader who wants an AI agents book for beginners in place of a course can start with the two free chapters. Instructors sorting the same topics for a whole class will find the sequencing questions in teaching AI agents to CS students.

Where does this triage not apply?

It does not apply when someone else sets your order, and it stops being a guide once you ship. Four limits are worth stating plainly.

  • It is argued, and nobody has measured it. I found no study comparing beginners who took these topics in different orders. The case rests on what one run requires.
  • A course or a job overrides it. If the framework is on the syllabus or in the codebase, learn it now, and log the requests it sends so the four ideas stay visible.
  • Skip means skip for now. The security row is in the later bin for learning. For anything another person relies on, it is a requirement before launch.
  • The reader evidence is thin. I quote a handful of Hacker News posts from 2024 to 2026, chosen because they voice the problem. They say nothing about how common it is.

The second question of the rule also carries a judgment about what every relied-on agent needs. Readers who disagree with where I put a row should move it; the tie-break keeps the cost of that small.

The one thing to keep

Learn what one run requires, and let symptoms in your own runs call in everything else. That is the whole of AI agents for beginners as I would teach it: four ideas, 14 words, seven sittings, and a list of ten things you are allowed to ignore until your own agent gives you a reason.

The habit to take from the first week is the last item on its list, a check that does not ask the model. Chapter 1 puts the question in the form the book repeats everywhere: “what signal tells you it worked?”

Chapter 1, “What Is an Agent?” and Chapter 2, “The Engine: How Language Models Work” are free to read online, and so is the glossary. The teaching kit has a syllabus with labs. The other chapters named in the table are in the full book, and you can see the formats.

Questions readers ask

Do I need machine learning or math to learn AI agents?
No, not to build and understand one. An agent calls a model that someone else trained, so the work is ordinary programming around that call. Chapter 2 of AI Agents, Engineered builds its mental model of the model with no mathematics beyond arithmetic. Linear algebra and training matter if you want to build models, which is a different subject.
Do I need to know how to code to learn AI agents?
To build one, yes: a function, a loop, JSON and an HTTP request are the working minimum, in any language. To understand what an agent is and judge a product, no. Chapter 1 of the book and the five-check test on this site need no code at all.
Should a beginner start with an agent framework?
Not as the first step, unless a course or a job requires one. A framework assembles the request to the model for you, and that request is the thing a beginner most needs to see. Pick one up when you can name the repeated work it would save you, and log the requests it sends.
What should a beginner skip when learning AI agents?
Skip, for now, frameworks, multi-agent systems, fine-tuning, vector databases, tool and agent protocols, long-term memory, machine-learning math, named reasoning patterns, model leaderboards and prompt tricks. None is needed to explain one run of one agent. Each comes back when a specific symptom shows up in your own runs.
How long does it take to learn the basics of AI agents?
I know of no measurement, so this page gives no schedule for the building. The reading can be sized: Chapters 1 and 2 of the book take about 59 minutes at 220 words a minute. The first week on this page is seven sittings, each with a check that says it is done.

Sources

  1. Microsoft (undated; read October 2026). AI Agents for Beginners (repository README)
  2. Hugging Face (undated; read October 2026). Welcome to the AI Agents Course (Unit 0)
  3. Jonathan Pedoeem (PromptLayer) (2025). Building Your First AI Agent: A Beginner's Guide
  4. Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents
  5. Eugene Yan, Bryan Bischof, Charles Frye, Hamel Husain, Jason Liu and Shreya Shankar (2024). What We've Learned From A Year of Building with LLMs
  6. Yubin Kim et al. (2025; v3 April 2026). Towards a Science of Scaling Agent Systems
  7. Thomas Ptacek (2025). You Should Write An Agent
  8. diefunction (2024). Ask HN: Curious about Open-Source AI Agent Frameworks
  9. shivc (2024). Ask HN: How do I get into AI engineering?
  10. crazymoka (2024). Ask HN: Have AI learn my own business knowledge for chat bot
  11. nadis (2025). Ask HN: Resources for building AI agents for software development?
  12. jdw64 (2026). Ask HN: May be a basic question, but how can I use AI well?