To build an AI agent from scratch you need four things in one file: a function that calls a model, a table of tools, a list that records everything said and done, and a loop of about two dozen lines that knows how to stop. This post prints that file. It runs with only Python installed.
Read and run the whole program before you choose a provider or a framework: the loop is small, and what it leaves out is the rest of the engineering. The model here is replaced by a scripted stand-in, and the real one sits behind a single function, call_model, with a written contract. I will call that function the seam.
The file is a Python rendering of the pseudocode in Appendix A, A Minimal Agent, Annotated (in the full book), with my additions labeled where they appear. The loop’s mechanics are in the post on what an agent loop is and the three-minute agent loop explainer.
What will you have when you finish?
You will have one Python file of 282 lines that you have run, read in the order it executes and broken six ways, plus a written list of what you must supply to connect a real model. Nothing is trained or fine-tuned, and nothing is installed.
If you are working out how to learn AI agents from scratch, this is a reasonable first exercise; the wider path is in the free guide to agent fundamentals.
Which model do you need, and will this code still run?
You need no model to start, and the code has nothing in it that a provider can retire: it imports only Python’s standard library, opens no network connection, needs no account or key, and names no model.
When you do choose, the requirement is a model whose interface accepts tool definitions and returns structured tool calls, whether a hosted service billed by usage or an open-weight model (one whose weights you can download) on your own machine. To check a candidate, look in its provider’s documentation for a page on tool calling, which some call function calling, and confirm that the model is listed as supporting it.
I have tested none of them with this file. The loop’s two budgets are there for a model that produces malformed calls or does not stop.
How do you build an AI agent from scratch in one file?
You write four parts, in the order a run uses them: the model client, the tool registry, the message history and the loop. Save the block below as agent.py.
#!/usr/bin/env python3
"""agent.py: a complete minimal agent harness in one file, with a scripted
stand-in where the model goes. Standard library only.
python3 agent.py # demo: scripted stand-in model, temp folder
python3 agent.py DIR "the goal" # your folder, after you fill in call_model
Four parts: a model client (one function), a tool registry, a message
history, and a loop with two budgets. Read it top to bottom.
"""
import json
import sys
import tempfile
from pathlib import Path
MAX_STEPS = 12 # illustrative; set it from real traces
MAX_TOKENS = 20_000 # illustrative; the unit is whatever call_model reports
STANDING_INSTRUCTIONS = (
"You are a careful assistant working inside the user's project folder. "
"Look before you touch: list or read before you edit. Report failures "
"plainly. When the goal is met, reply with a short summary and make no "
"tool calls."
)
WORKSPACE = None # the only folder the tools may touch; main() sets it
# ---- 1. the model client: THE ONE SEAM -------------------------------------
def call_model(history, tool_definitions):
"""Send the whole history and the tool definitions; return the next message.
This is the only function that knows which model you use. To plug in a
real one, replace the body with a request to your provider and translate
both ways, to and from this file's own shapes:
history entries
{"role": "system", "text": str}
{"role": "user", "text": str}
{"role": "model", "text": str, "tool_calls": [call, ...]}
{"role": "tool", "call_id": str, "name": str, "text": str,
"is_error": bool}
a call {"id": str, "name": str, "arguments": dict}
return value {"text": str, "tool_calls": [call, ...], "tokens": int}
"tokens" is the usage your provider reports for this request, input
plus output. The model is stateless: it sees only what is in `history`
on this call.
"""
return scripted_model(history, tool_definitions)
# ---- 2. the tool registry ---------------------------------------------------
def inside_workspace(path):
if WORKSPACE is None:
raise RuntimeError("no workspace set; run this file through main()")
target = (WORKSPACE / path).resolve()
if target != WORKSPACE and WORKSPACE not in target.parents:
raise ValueError(f"path is outside the workspace: {path}")
return target
def list_files(path="."):
entries = sorted(inside_workspace(path).iterdir())
return "\n".join(e.name + ("/" if e.is_dir() else "") for e in entries)
def read_file(path):
return inside_workspace(path).read_text(encoding="utf-8")
def edit_file(path, old_text, new_text):
target = inside_workspace(path)
if old_text == "":
if target.exists():
raise ValueError("file exists; pass the exact text to replace")
target.write_text(new_text, encoding="utf-8")
return f"ok: created {path}"
text = target.read_text(encoding="utf-8")
count = text.count(old_text)
if count != 1:
raise ValueError(f"old_text must match exactly once; matched {count}")
target.write_text(text.replace(old_text, new_text), encoding="utf-8")
return "ok: 1 replacement made"
def text_args(required, optional=()):
names = list(required) + list(optional)
return {"type": "object", "required": list(required),
"properties": {name: {"type": "string"} for name in names}}
TOOLS = {
"list_files": {
"description": "List the files and folders under a relative path. "
"Call it with no path to see the project root.",
"schema": text_args([], ["path"]),
"fn": list_files,
},
"read_file": {
"description": "Return the full text of one file, by relative path. "
"Use list_files first if you are unsure of the path.",
"schema": text_args(["path"]),
"fn": read_file,
},
"edit_file": {
"description": "Replace old_text with new_text in the file at path. "
"old_text must match exactly one place in the file. "
"To create a new file, pass an empty old_text.",
"schema": text_args(["path", "old_text", "new_text"]),
"fn": edit_file,
},
}
# The only part of a tool the model ever sees. The function stays home.
TOOL_DEFINITIONS = [
{"name": name, "description": tool["description"], "schema": tool["schema"]}
for name, tool in TOOLS.items()
]
def execute(call):
"""Run one requested call, which must carry an id and a name. An unknown
tool, wrong arguments or anything a tool raises comes back as a result."""
def result(text, is_error=False):
return {"role": "tool", "call_id": call["id"], "name": call["name"],
"text": text, "is_error": is_error}
if call["name"] not in TOOLS:
return result(f"no such tool: {call['name']}", is_error=True)
try:
return result(TOOLS[call["name"]]["fn"](**call["arguments"]))
except Exception as error: # a failed tool is news for the model
return result(f"{type(error).__name__}: {error}", is_error=True)
# ---- 3 and 4. the history and the loop -------------------------------------
def run_agent(goal):
history = [{"role": "system", "text": STANDING_INSTRUCTIONS},
{"role": "user", "text": goal}]
show(history, 0)
passes = tokens = 0
def outcome(status):
return {"status": status, "passes": passes, "tokens": tokens,
"history": history}
while passes < MAX_STEPS:
if tokens >= MAX_TOKENS:
return outcome("stopped: token budget exhausted")
passes += 1
print(f"-- pass {passes} " + "-" * 50)
seen = len(history)
response = call_model(history, TOOL_DEFINITIONS)
tokens += response["tokens"]
reply = {"role": "model", "text": response["text"],
"tool_calls": response["tool_calls"]}
history.append(reply) # the model's message, calls included
if not response["tool_calls"]:
show(history, seen)
return outcome("finished")
for call in response["tool_calls"]:
history.append(execute(call)) # one result per call, tied by id
show(history, seen)
# Also the label when both budgets ran out on the same pass.
return outcome("stopped: step budget exhausted")
def show(history, start):
"""Print the entries added since `start`: the trace, as it grows."""
for index in range(start, len(history)):
entry = history[index]
text = " ".join((entry["text"] or "").split()) # None prints as ""
if len(text) > 58:
text = text[:55] + "..."
label = "ERROR " if entry.get("is_error") else ""
print(f"history[{index}] {entry['role']:<6} {label}{text}")
for call in entry.get("tool_calls", []):
print(f"{'':18}-> {call['name']}({json.dumps(call['arguments'])})")
# ---- a stand-in model, so the file runs with no account and no network -----
class RequestRejected(Exception):
"""What a model interface answers when the history is malformed."""
def scripted_model(history, tool_definitions):
"""Not a model. A few rules that read the history, which is the point:
like a real model, it knows only what the loop appended. It guesses a
file name, recovers from the error, then fixes the demo's greeting."""
check_pairing(history)
results = [entry for entry in history if entry["role"] == "tool"]
last = results[-1] if results else None
number = 1 + sum(len(e["tool_calls"]) for e in history if e["role"] == "model")
def reply(text, name=None, **arguments):
calls = [{"id": f"call_{number}", "name": name, "arguments": arguments}]
used = sum(len(json.dumps(entry)) for entry in history) // 4
return {"text": text, "tool_calls": calls if name else [],
"tokens": used + len(text) // 4} # a rough size estimate
if last is None:
return reply("I'll open the greeting file.", "read_file",
path="greeting.py")
if last["is_error"]:
return reply("That failed. I'll look at the folder.", "list_files")
if last["name"] == "list_files":
names = [n for n in last["text"].split() if "greet" in n]
return reply(f"{names[0]} looks likely.", "read_file", path=names[0])
if last["name"] == "read_file" and "'Hello'" in last["text"]:
return reply("Found it. Changing the greeting.", "edit_file",
path=path_of(history, last),
old_text="GREETING = 'Hello'",
new_text="GREETING = 'Welcome'")
if last["name"] == "edit_file":
return reply("Let me confirm the change.", "read_file",
path=path_of(history, last))
return reply("Done. GREETING now says 'Welcome'; greet() uses it.")
def path_of(history, result):
"""The path argument of the call that `result` answers."""
for entry in history:
for call in entry.get("tool_calls", []):
if call["id"] == result["call_id"]:
return call["arguments"]["path"]
def check_pairing(history):
"""Each result must answer a call the model made, and each call must be
answered before the model is asked again."""
waiting = set()
for entry in history:
if entry["role"] == "model":
if waiting:
raise RequestRejected(f"calls never answered: {sorted(waiting)}")
waiting = {call["id"] for call in entry["tool_calls"]}
elif entry["role"] == "tool":
if entry["call_id"] not in waiting:
raise RequestRejected(
f"result for {entry['call_id']}, which no model message "
"in the history requested")
waiting.discard(entry["call_id"])
if waiting:
raise RequestRejected(f"calls never answered: {sorted(waiting)}")
def main():
global WORKSPACE
demo = len(sys.argv) == 1
if demo:
WORKSPACE = Path(tempfile.mkdtemp(prefix="agent-demo-")).resolve()
(WORKSPACE / "tests").mkdir()
(WORKSPACE / "README.md").write_text("# demo project\n")
(WORKSPACE / "app.py").write_text(
"from greetings import greet\nprint(greet('Ada'))\n")
(WORKSPACE / "greetings.py").write_text(
"GREETING = 'Hello'\n\ndef greet(name):\n"
" return f'{GREETING}, {name}!'\n")
goal = ("The greeting in this project still says 'Hello'. Find where "
"it is defined and change it to 'Welcome'.")
elif len(sys.argv) == 3 and Path(sys.argv[1]).is_dir():
WORKSPACE, goal = Path(sys.argv[1]).resolve(), sys.argv[2]
else:
sys.exit('usage: python3 agent.py or python3 agent.py DIR "goal"')
try:
run = run_agent(goal)
except RequestRejected as error:
sys.exit(f"REQUEST REJECTED: {error}")
print(f"\n{run['status']} after {run['passes']} passes, "
f"about {run['tokens']} tokens, {len(run['history'])} history entries")
Path("agent-transcript.json").write_text( # in the folder you ran it from
json.dumps(run["history"], indent=2), encoding="utf-8")
print("full history saved to agent-transcript.json")
if demo:
first_line = (WORKSPACE / "greetings.py").read_text().splitlines()[0]
print(f"greetings.py, line 1: {first_line}")
print(f"demo folder: {WORKSPACE}")
if __name__ == "__main__":
main()
The file is 282 lines, blank lines and comments included, against “About sixty lines” for Appendix A’s pseudocode (section “The Program in Full”). run_agent, the function that is the agent, is 27 lines without its blank ones, and the while loop inside it is 17. About 66 lines are the stand-in model and its history check.
The return arrow in that figure is the loop. To make a loop’s decisions by hand, try Run the loop, a scripted seven-pass run in which you play the code around the model and choose whether to append, retry, stop or escalate to a person.
How do you run it, and what should you see?
Run python3 agent.py in an empty folder that holds only the file, and you should see thirteen history entries appear across six passes, ending with the word finished. This is the complete output on my machine (Python 3.12.3 on Linux; Python 3.14.7 gave the same):
history[0] system You are a careful assistant working inside the user's p...
history[1] user The greeting in this project still says 'Hello'. Find w...
-- pass 1 --------------------------------------------------
history[2] model I'll open the greeting file.
-> read_file({"path": "greeting.py"})
history[3] tool ERROR FileNotFoundError: [Errno 2] No such file or directory:...
-- pass 2 --------------------------------------------------
history[4] model That failed. I'll look at the folder.
-> list_files({})
history[5] tool README.md app.py greetings.py tests/
-- pass 3 --------------------------------------------------
history[6] model greetings.py looks likely.
-> read_file({"path": "greetings.py"})
history[7] tool GREETING = 'Hello' def greet(name): return f'{GREETING}...
-- pass 4 --------------------------------------------------
history[8] model Found it. Changing the greeting.
-> edit_file({"path": "greetings.py", "old_text": "GREETING = 'Hello'", "new_text": "GREETING = 'Welcome'"})
history[9] tool ok: 1 replacement made
-- pass 5 --------------------------------------------------
history[10] model Let me confirm the change.
-> read_file({"path": "greetings.py"})
history[11] tool GREETING = 'Welcome' def greet(name): return f'{GREETIN...
-- pass 6 --------------------------------------------------
history[12] model Done. GREETING now says 'Welcome'; greet() uses it.
finished after 6 passes, about 1778 tokens, 13 history entries
full history saved to agent-transcript.json
greetings.py, line 1: GREETING = 'Welcome'
demo folder: /tmp/agent-demo-346le5lo
The model lines are the stand-in’s scripted text, not a model’s. Two lines depend on the machine. The last names the demo folder, created fresh in your system’s temporary folder on each run and left behind.
The token figure is the stand-in’s estimate from character counts: one entry contains that folder’s path, so a longer path gives a larger number, and the estimate leaves out the tool definitions, which a real interface counts on every request.
One file is written outside the demo folder: agent-transcript.json, the full history, in the folder you ran the command from. The trace on screen cuts each entry to 58 characters; the transcript does not. A run that ends in a rejection or a traceback writes none, so an older transcript may still be sitting there; each successful run overwrites it. With your own folder and a real model, the transcript holds the full text of every file a tool read, so treat it like those files.
Each entry in the trace is one append. On pass 1 the stand-in guesses a file name and is wrong, and entry 3 is that failure, returned as an ordinary result and marked ERROR. Pass 4 is the only write, and pass 6 returns text with no tool calls. Appendix A traces the same goal in five passes and eleven entries (section “A Run, End to End”), in a transcript it calls “hypothetical and lightly tidied for the page”; mine opens with a wrong guess, which adds a pass.
The second form, python3 agent.py DIR "the goal", is for the day a real model sits behind the seam; the stand-in knows only the greeting task.
What does the stand-in not show you?
It does not show a model reasoning. scripted_model is six rules I wrote: it picked greetings.py because the name contains “greet”. It reads only the history, which keeps the experiments below honest, but its choices are mine.
A real model chains tools on its own judgment, and no script can show that. A system with no language model in the path is not an AI agent, and until you plug one in, this file has none. The appendix says of its trace: “Your program is dumb on purpose; the transcript only looks smart.” With a stand-in, the transcript is not smart either.
What goes in the history, and what exactly do you send back?
After the model asks for a tool, you append two entries in this order: the model’s own message, request included, and then one result for each call, carrying the id of the call it answers. Then you send the whole list again. Here are entries 2 and 3 of the run above, as the transcript holds them, with the folder’s path shortened:
{"role": "model", "text": "I'll open the greeting file.",
"tool_calls": [{"id": "call_1", "name": "read_file",
"arguments": {"path": "greeting.py"}}]}
{"role": "tool", "call_id": "call_1", "name": "read_file",
"text": "FileNotFoundError: [Errno 2] No such file or directory: '<workspace>/greeting.py'",
"is_error": true}
The first entry carries a tool call: a request, in data, that names a tool and its arguments. The second is its result, and call_id binds the two: “the id is what ties each result to the request it answers” (Appendix A, “Executing a Request”).
This step is where I found the hardest evidence of people stuck. One Stack Overflow asker wrote “However, I am stuck here” after an interface answered that the call ids “did not have response messages” (question 77882437, 25 January 2024). Another sent a user question followed directly by a tool result: “It seems to completely ignore the”tool” prompt” (question 79473550, 27 February 2025).
Two more things belong to your code. It runs the function: the model only returns a request, and a third asker expected otherwise (“Why does it not execute the function, but it just prints the call”, question 77461857, 10 November 2023). The appendix’s answer, in “The Tool Registry”: “The function stays home.” And it keeps the conversation: the model is stateless, meaning it keeps nothing between calls, so run_agent passes the full history on every pass.
What is the one function you rewrite for a real model?
It is call_model, the model client: the only function in the file that would know which provider you use. Its docstring is the contract, and nothing else in the program changes when the provider does.
In goes the whole history, as entries with one of four roles (system, user, model, tool), plus the tool definitions: a name, a description and an argument schema (a machine-readable description of the arguments) for each tool. Out must come one dictionary with three keys: text (the model’s prose, possibly empty), tool_calls (a list, possibly empty, of calls that each have an id, a name and arguments already parsed into a dictionary) and tokens (the usage reported for that request, input plus output, counted in tokens, the fragments of text a model reads and writes).
The design is the appendix’s: “It is the only line in the program that touches a provider. Swap vendors, and this is the function you rewrite” (section “The Model Client”). The written contract and the stand-in are mine. If you are looking for an LLM API for beginners, this contract is the part I would expect to carry from one provider to the next.
One commenter on Thomas Ptacek’s essay about writing an agent called building on a provider’s client library “like saying you should implement a web server, you will learn so much, and then you go and import http” (Hacker News comment, 6 November 2025). The part someone else built is the model; the client can at least be one replaceable function.
Why are the tools confined to one folder?
They are confined because a real model behind the seam will read and edit files on its own judgment, and the only limit on its reach is the one your code enforces. Every path passes through inside_workspace, which resolves it and refuses anything outside the WORKSPACE folder.
The three tools can list a folder, read a text file, create a file and replace one exact passage in a file, all inside that folder. They cannot run a command, delete a file or reach the network. They can empty a file or remove any passage from one, and they can edit a file that something else runs later, so a folder of scripts is not a harmless target.
The check is my addition; the appendix warns, in “Executing a Request”, that its own edit_file “will cheerfully modify anything you can”. I tried five escapes: a path that climbs out, an absolute path, symbolic links to a file and to a folder outside, and creating a file through a dangling link. Each came back as ValueError: path is outside the workspace. A path check is still not a sandbox: the program runs with your permissions.
Why is there no shell tool?
There is no shell tool because one general tool removes every limit the three narrow ones set. Simon Willison gave the reason: “Since the most powerful coding agent tool is”run this command in the shell” a rogue agent can do anything that you could do by running a command yourself” (Willison, 30 September 2025).
For a first agent, my answer is to not offer the tool, and never to pass model output to eval.
How does the loop stop?
It stops through one of three labeled exits: finished, when the model replies with no tool calls; stopped: step budget exhausted, after MAX_STEPS passes; and stopped: token budget exhausted, when the running token count reaches MAX_TOKENS. Each returns the status, the counts and the full history.
A pass is: check the token budget, call the model with the whole history, append its message, return if it asked for nothing, otherwise run each call and append each result. The while condition is the step cap, and it is the one reported if both budgets run out on the same pass.
The first exit is the model’s opinion: “the run is over when the model stops asking for things” (Appendix A, “The Loop”), and the last sentence of STANDING_INSTRUCTIONS (the system prompt) tells the model to stop that way. The other two exits do not ask the model.
Two things here are mine: a cap of 12 where the appendix has 20 (both labeled illustrative), and the token budget, which the appendix lists among its omissions (“The budget counts passes and never money or tokens”). These are not the only stop conditions for agent loops; a check the model cannot vote on and a time limit are two more, and this file has neither.
What did the well-known loops leave out?
The loop code in three of the best-known write-ups prints no cap on passes: Thorsten Ball’s (“It’s an LLM, a loop, and enough tokens”, 15 April 2025), Philip Zeyliger’s (“the core idea is the above 9 lines”, 15 May 2025) and Ptacek’s (“Clearly, this is a toy example”, 6 November 2025). They are why the small loop is common knowledge.
Ball’s loop is a bare for {, Zeyliger’s is while True, and Ptacek’s is a while that continues for as long as the model keeps asking for tools. All three sit inside a chat loop that waits for a person’s next message, so someone is at the keyboard; what none of them bounds is how many passes the model can take between two of that person’s messages.
Chapter 3 puts the cap first, in “Stop Conditions, Budgets, and Frameworks”: “Set the cap before you write anything else.” A stop condition belongs in the first version.
What happens when you break it on purpose?
Each of six one-line changes produces a different, specific failure that shows what the changed line was protecting. Make one change, run the file, then restore the line before the next.
- In
run_agent, changecall_model(history, TOOL_DEFINITIONS)tocall_model(history[:2], TOOL_DEFINITIONS), so the model sees only the instructions and the goal. Does it need the rest? Run it and watch what the loop does. - Change
history.append(execute(call))toexecute(call), so the tool runs and its result is thrown away. Run it and watch what the loop does on pass 2. - Put a
#in front ofhistory.append(reply), so the model’s own message is never recorded. Run it and watch what the loop does on pass 2. - In
execute, replace the last line, thereturn result(...)underexcept, withraise. What does one wrong file name cost now? Run it and watch what the loop does. - Change
MAX_STEPS = 12toMAX_STEPS = 3. Run it and watch what the loop does when the cap arrives first. - Change
MAX_TOKENS = 20_000toMAX_TOKENS = 600. Run it, watch what the loop does, then check what line 1 ofgreetings.pysays.
These are my results, from the file exactly as printed above.
| # | The line to look for | What it shows |
|---|---|---|
| 1 | stopped: step budget exhausted after 12 passes, about 1212 tokens, 26 history entries; the file still says 'Hello' |
Every pass repeats the same wrong read_file call. The model knows only what it is sent, and the cap is what ended the run. |
| 2 | REQUEST REJECTED: calls never answered: ['call_1'] on pass 2, exit status 1 |
A request with no result is a malformed history. |
| 3 | REQUEST REJECTED: result for call_1, which no model message in the history requested on pass 2, exit status 1 |
A result with no request is malformed too. |
| 4 | A Python traceback ending in FileNotFoundError, on pass 1, exit status 1 |
With errors as results, the same wrong guess cost one pass. Without, it ends the run. |
| 5 | stopped: step budget exhausted after 3 passes, about 537 tokens, 8 history entries; the file still says 'Hello' |
The cap fails loudly and says how far the run got. |
| 6 | stopped: token budget exhausted after 4 passes, about 868 tokens, 10 history entries; the file says 'Welcome' |
The budget fired after the edit and before the confirming read. |
The token figures in rows 5 and 6 depend on the length of your demo folder’s path, as the default run’s does; row 1’s does not. Row 6 has a margin: my pass 3 ended at 537, 63 under the limit. With a 151-character path it ended at 601, and the run stopped after 3 passes with the file unchanged. If that happens, set MAX_TOKENS between your pass-3 and pass-4 totals.
What do the six results show?
Row 1 is the same call requested forever, the failure Chapter 3 means by: “Without a cap, that bug is a bill with no ceiling; with one, it is a log entry.” Row 4 reverses the appendix’s policy for execute: “Three cases, one policy: everything becomes a result.”
Rows 2 and 3 produce the two failures from the Stack Overflow questions earlier. The first asker had appended the tool’s output, but under the model’s role and without the call’s id, which an interface cannot tell from no result at all. Here the rejection comes from check_pairing, a check I put in the stand-in.
Row 6 is the one I would not have predicted. The run reported stopped, and the file on disk had already changed: a budget stop is not a rollback. The budget is checked before each model call, which is why 868 passed a limit of 600.
How do you plug in a real model?
You replace the body of call_model with a request to your provider, translating the history and tool definitions into its format on the way in and its response into the three-key dictionary on the way out. That body is your adapter. This post prints one for no provider, and it has not run the seam against a live model: I had none to call, so your first real run is where you find out whether your adapter honors the contract.
What happened to the model names in older tutorials?
The dated snapshots behind some of them have been retired or scheduled for removal by the providers themselves; whether the short names the tutorials print still resolve, I did not test. That is why this file names none.
On 6 October 2026 I read two providers’ deprecation pages. One lists the dated snapshot of the model generation named in Ball’s code and Zeyliger’s prose as retired on February 19, 2026: “Requests to retired models will fail.” The other lists the dated snapshot of the model named in Ptacek’s code for removal on December 11, 2026. None of the three prints a dated identifier.
A reader of one provider’s own course did meet the failure: “you have hardcoded in a model name that is deprecated resulting in a 404” (issue #71, 17 February 2025). None of this faults the writing: a tutorial that calls a real model has to name one, and names expire.
What does call_model have to translate?
It has to translate the few points at which provider interfaces differ. The table compares categories from the documentation of four interfaces (three hosted, one local runtime), read on 6 October 2026. It has no syntax to copy; your provider’s reference has that.
| What differs at the seam | Variants I found | What your call_model must do |
|---|---|---|
| How a tool request is represented | A typed block inside the model’s message; a list of calls attached to the message; a separate typed item in a flat list | Collect every request into tool_calls |
| How a result is paired with its request | By the request’s id in three of the four; one runtime’s example pairs by the tool’s name | Send each call_id back in the place the interface expects. If the interface gives a call no id, make one up in call_model and send the tool’s name back |
| Where a result goes | Inside a user-role message directly after the model’s message; a message with its own tool role; a typed item appended to the input; a typed item sent with a reference to the previous exchange | Turn each tool entry into that form, in order |
| How arguments arrive | Already parsed, or as text you must parse | Return arguments as a dictionary either way |
| Where the system instruction goes (two of the four checked) | A separate field outside the message list; a message with its own role inside the list | Move the system entry to where it belongs |
| Who keeps the conversation | You, in most documented examples; the server, by reference, as the main route in one and as an option in at least one other | Send the whole history, or only what is new |
One hosted interface states the pairing rule outright: “Tool result blocks must immediately follow their corresponding tool use blocks in the message history” (provider documentation, read 6 October 2026).
What does a contract violation look like?
It looks like one of five symptoms, and a sixth looks like a violation and is not one. I produced the last four by wrapping the stand-in; the first two need a real interface.
- The second request is rejected. Your translation lost a
call_idor put a result under the wrong role; pairing mistakes surface on pass 2. - The request is accepted and the model acts as if the result is not there. Check that the model’s own message, calls included, went back too.
- Every tool result is the same
TypeError(argument after ** must be a mapping, not str), and the run ends at the step cap after 12 passes. You leftargumentsas text. KeyError: 'tokens'on pass 1. The returned dictionary is missing a key. Returning 0 fortokensis worse: with the budget set to 5, the token budget never fired.- A blank
modelline above a tool call. Your adapter returnedNoneas the text of a message that carried only calls. The run finished as usual, but the contract says a string. Return an empty one. finished after 1 passes, and the text tells you what to do instead of doing it. This is not a contract violation. The model answered in prose, and the loop cannot tell that from success. Check that your adapter sent the tool definitions; then fix the words (tool descriptions, standing instructions) before the code.
One habit helps with the first two: keep check_pairing(history) as the first line of your new call_model.
What should your first real run look like?
It should be small enough that a mistake costs almost nothing. Set MAX_STEPS to 4 and MAX_TOKENS to a few times the size of one request, and raise both only after you have read agent-transcript.json in full. If your provider offers a spending limit on the account or the key, set it too; that limit sits outside the program.
Point the agent at a scratch folder that holds copies. Whatever a tool reads is sent to the model’s provider, so keep credentials and anything private out of that folder, and do not keep agent.py in it. Start with small text files: one read returns the whole file, whatever its size, and the token budget is checked only before the next call.
A real model will choose its own reads and edits inside that folder, and you will have approved none of them in advance. Give it the greeting task first, and watch for what the stand-in could not show, the moment Chapter 3 names in “A Minimal Agent from Scratch”: “It chained the tools itself.”
What does this file leave out?
It leaves out almost everything that makes an agent dependable, and Appendix A lists the omissions itself in its section “What Is Deliberately Missing”. The quotations in the table are the appendix’s; the chapters are all in the full book.
| What the minimal agent lacks, in the appendix’s words | Chapter that adds it | In this post’s file |
|---|---|---|
| Context management: “the history grows without limit” | 7, with 8 and 9 | Missing |
| “no retry, no timeout, no idempotency when a tool fires twice, no way to resume a run the process crashed out of” | 18 | Missing |
| Observability beyond “a returned transcript” | 15 | Missing; the trace is printed and saved once |
| “no sandbox, no allow-list, no defense when a file the agent reads turns out to contain instructions” | 17 | A path check only, added here |
| “The budget counts passes and never money or tokens” | 19 | A token budget, added here; no money |
| “Nothing measures whether runs succeed” | 16 | Missing |
| “nothing gates the edit behind a human’s approval” | 12 | Missing |
| “nothing runs the agent unattended toward a machine-checkable goal” | 13 | Missing |
| “there is exactly one desk, one loop, one agent” | 11 | Missing |
One gap deserves a sentence outside the table. finished means the model stopped asking, not that the greeting works: the demo’s tests/ folder is empty, and no tool here could run tests if it held any. The appendix says of its own trace that the model’s “confirmation is that the edit landed, and silence on whether anything still works.”
Do you need a framework to build an agent?
No, not to build an AI agent from scratch or to understand one, and possibly yes to operate one. Everything in the table above is harness, the code around the model, and the post on what an agent harness is weighs building it against adopting it.
Ptacek, asked in the discussion of his essay why not use a framework, replied: “Because you won’t learn as much using an agent framework, and, as you can see from the post, you absolutely don’t need one” (Hacker News reply, 7 November 2025). A commenter in the same discussion gave the other side: “sometimes I’m happy plugging my tools and data into someone else’s platform because they are spending orders of magnitudes more time than me doing the janitor work to keep up with whatever’s trendy” (Hacker News comment, 7 November 2025).
Chapter 3 reconciles the two in “Stop Conditions, Budgets, and Frameworks”: “Adopt for a named need, never for the feeling that serious systems use frameworks.”
Limits of this tutorial
This is a page on how to build an AI agent from scratch that never calls a model, so it proves less than one that does.
The seam is untested against a live model. The contract and the category table come from documentation and the stand-in’s behavior; some interface may need something the contract does not carry.
The stand-in is a script. It ignores the tool definitions, so the schemas are unexercised until a real run, and it says nothing about how well any model uses these tools or what a run costs.
The safety is thin. One path check, tested on five inputs, stands between a model and the files your account can write. There is no approval step, no limit on the size of one read, and no defense against instructions hidden in a file.
The reader evidence is a sample. I found the quotations by searching for people who were stuck, so they say nothing about how common the stuck points are. I ran the file on Linux only.
The takeaway
Running the file gives you the harness: one function where the model goes, three tools, one list and a loop with three exits. The one function is what makes it an agent, and that step is still yours to write and test. It is also a concrete answer to what agentic AI is: software in which a model’s output chooses the next step, inside a loop with limits the model cannot overrule.
The four parts and the framework question are in Chapter 3, “The Agent Loop”, and the annotated program is in Appendix A; both are in the full book. Chapter 1 and Chapter 2, which set up the vocabulary, are free to read online, and you can see the formats.
Questions readers ask
- Do I need a framework to build an AI agent?
- No. An agent is a function that calls a model, a set of tools, a message history and a bounded loop, and all four fit in one file with no dependencies beyond the standard library. A framework adds plumbing such as persistence, retries and tracing. Chapter 3 of AI Agents, Engineered gives the rule: adopt one for a named need, and stay minimal while you are learning.
- Is an AI agent really just a while loop?
- The control flow is. The loop calls the model, records what it said, runs the tools it asked for, records the results and repeats until the model stops asking or a budget runs out. The judgment is in the model and in the words that describe the tools. Everything a production agent adds, such as retries, sandboxing and context management, is built around that loop.
- What do I send back to the model after a tool call?
- The whole history again, with two new entries at the end: the model's own message containing the tool request, then a result entry that carries the id of the call it answers. A result without its request, or a request without its result, is a malformed history, and model interfaces either reject it or produce an answer that ignores the result.
- Can I build an AI agent without paying for an API?
- You can run and study the whole loop for free: the file in this post replaces the model with a scripted stand-in and needs no account, key or network. To watch a model make its own choices you need a real one behind the call_model function, which can be a hosted service or an open-weight model run on your own machine.
- How does an agent know when to stop?
- By default it stops when the model replies without asking for a tool, which is the model's opinion that the work is done. Because that opinion can be wrong or never arrive, the loop also needs limits the model cannot overrule. The file in this post has two: a cap on passes and a budget on tokens, each with its own labeled outcome.
Sources
- Thorsten Ball (2025). How to Build an Agent, or: The Emperor Has No Clothes
- Philip Zeyliger (2025). The Unreasonable Effectiveness of an LLM Agent Loop with Tool Use
- Thomas Ptacek (2025). You Should Write An Agent
- Simon Willison (2025). Designing agentic loops
- Stack Overflow (2024). Stack Overflow question 77882437: errors in the last step of appending messages in function calling (asked 25 January 2024)
- Stack Overflow (2025). Stack Overflow question 79473550: results from a tool ignored (asked 27 February 2025)
- Stack Overflow (2023). Stack Overflow question 77461857: function calling does not execute the function (asked 10 November 2023)
- Hacker News commenter teiferer (2025). Forum comment comparing a client import to importing a web server (6 November 2025)
- Thomas Ptacek (Hacker News, as tptacek) (2025). Forum reply on learning without an agent framework (7 November 2025)
- Hacker News commenter z2 (2025). Forum comment on platforms that absorb interface churn (7 November 2025)
- anthropics/courses issue tracker (2025). Issue #71: implement dynamic model name selection in course notebooks (opened 17 February 2025)
- Anthropic documentation (2026). Model deprecations page of a hosted model provider (read 6 October 2026)
- OpenAI documentation (2026). Deprecations page of a second hosted model provider (read 6 October 2026)
- Anthropic documentation (2026). Tool-use documentation: handling tool calls (read 6 October 2026)
- OpenAI documentation (2026). Function calling guide (read 6 October 2026)
- Google documentation (2026). Function calling guide of a third hosted provider (read 6 October 2026)
- Ollama documentation (2026). Tool calling in a local model runtime (read 6 October 2026)
- Anthropic documentation (2026). Messages reference, on where the system instruction goes (read 6 October 2026)
- OpenAI documentation (2026). Text generation guide, on instructions and message roles (read 6 October 2026)