What is an agent harness? It is everything you build around a language model to turn it into a working agent: the loop that calls the model again and again, the tools and what they may touch, the message history, the stop rules, and, as the system matures, the logging, recovery and guardrails. An agent is a model plus a harness. The model supplies judgment; the harness is the part you own.
That definition comes from Chapter 3 of AI Agents, Engineered, and I want to defend it in practical terms, because the word is new enough that every vendor page defines it slightly differently. If you are a backend engineer who has just been handed an agent project, the useful news is that the harness is mostly the kind of software you already know how to write: a loop, a dispatcher, a state store, timeouts, retries, logs. This post lists the components, names the failure each one prevents, and ends with a rule, and a worked example, for deciding when a framework earns its place.
What is an agent harness, exactly?
An agent harness is the deterministic code around a nondeterministic model. The model reads what you send it and returns either a final answer or a request to call a tool; the harness executes that request, records the result, decides whether to go around again, and enforces every limit the model cannot be trusted to enforce on itself.
The book puts the division of labor in one sentence: “the model supplies the judgment—which tool, with what arguments, and whether the job is done—and your code supplies everything else: the hands, the memory, and the stop button.” Every word of that is literal. The model provider’s service is stateless; it remembers nothing between calls. So the conversation, the tool execution and the decision to stop all live in your process, whether you wrote that process or imported it.
The figure above, from Chapter 3, draws one turn of the loop: observe, reason, act, observe the result. Only one of those four beats belongs to the model. If the loop itself is new to you, the companion post on how the agent loop works, beat by beat and the short agent loop explainer cover it from scratch.
Where did the term “agent harness” come from?
The word is borrowed and its scope is still elastic. Software testing has had the test harness for decades, and benchmark suites run models inside evaluation harnesses. Applied to agents, the idea that an agent is the model plus its scaffold appears in a UK government discussion paper from 2023, as the Wikipedia entry on agent harnesses records, and the phrase agent = model + harness spread through practitioner writing in early 2026.
Attribution of harness engineering, the discipline, is genuinely contested, so I will report the sources rather than crown anyone. In a February 2026 post, Mitchell Hashimoto wrote that he had “grown to calling this ‘harness engineering’”: the idea that “anytime you find an agent makes a mistake, you take the time to engineer a solution such that the agent never makes that mistake again” (Hashimoto, 2026). He added that if an established term existed, he would adopt it.
A month later, Vivek Trivedy of LangChain published a widely repeated formulation: “Agent = Model + Harness. If you’re not the model, you’re the harness” (Trivedy, 2026). Some writers credit him with coining the discipline’s name (Osmani, 2026), but the phrase was spreading within days of Hashimoto’s post: OpenAI’s Ryan Lopopolo published “Harness engineering” on February 11 (Lopopolo, 2026), and Trivedy’s own “Improving Deep Agents with Harness Engineering” followed on February 17.
The most careful treatment I know is the essay by Birgitta Böckeler of Thoughtworks, published on martinfowler.com, which opens by noting that the term “has emerged as a shorthand to mean everything in an AI agent except the model itself” and then narrows it (Böckeler, 2026). She separates the harness a tool’s builder ships from the outer harness a user wraps around it: instruction files, checks, hooks. The senses nest. Whichever layer you own is your harness, and the engineering is the same.
What are the components of an agent harness?
An agent harness has four core components and four that arrive as the system matures. The core four are the model client, the tools, the message history and the loop with its stop rules; Appendix A of the book builds exactly these and calls the result “the thinnest harness that works.” The other four are recovery, observability, guardrails and verification.
The table below is the one I wish the vendor pages included: each component, what it does, and the specific thing that goes wrong when it is missing.
| Component | What it does | What breaks without it |
|---|---|---|
| Model client | Sends the history and tool definitions; returns the next response | Nothing runs; also the one place a provider swap touches, so a scattered client makes portability expensive |
| Tools (definition + implementation + permissions) | Tells the model what it can request; executes requests within limits | The model can only talk; with loose permissions, it can do anything your credentials can |
| Message history | Holds every instruction, action and result; the agent’s entire memory of the run | Drop a tool result and the model re-requests the same call; drop its own action and it forgets what it tried |
| Loop with stop conditions | Decides whether to go around again; checks a programmatic success signal | The agent declares victory with tests still failing, or never stops polishing |
| Budgets (steps, tokens, time, money) | Hard ceilings the model cannot argue with | One repeated call becomes an unbounded bill |
| Error handling and recovery | Returns recoverable errors to the model; retries transient failures; resumes from a checkpoint | A blip kills an hour of work, or a resumed run sends the same email twice |
| Observability (traces) | Records every step, call, token count and timing | A failure you cannot reproduce and cannot explain |
| Guardrails and verification | Checks outside the model’s reasoning: allow-lists, approvals, tests, validators | Unsafe actions reach production; wrong output reaches a human who trusts it |
Two rows deserve emphasis because the popular explainers skip them. First come the stop rules: the book’s line is “The model’s ‘done’ is testimony; a green test is evidence,” and a loop whose only exit is the model’s judgment has no guaranteed exit at all.
Second is the recovery row. A checkpoint records where the run was; as Chapter 18 puts it, “A checkpoint records position. It does not, and cannot, record whether the outside world did the work.” That gap is why idempotency sits beside every checkpoint.
Why does the harness matter as much as the model?
An agent harness matters because it decides what the model sees, what its requests are allowed to do, and when its work counts as finished. Hold the model fixed, change only the harness, and measured performance moves; practitioners who have run that experiment report gains large enough to reorder a leaderboard.
LangChain published the cleanest public number I found. In February 2026, Trivedy’s team held one model fixed and changed only the system prompt, tools and middleware of their coding agent on the Terminal Bench 2.0 suite of 89 tasks. The score rose from 52.8% to 63.6% from the prompt and middleware changes alone (a build-then-verify loop, environment context, loop protection, timeout warnings), and to 66.5% once reasoning settings were tuned as well (Trivedy, 2026). Treat it as one vendor’s run on one model at one date.
The diagnosis behind it is more durable than the figure: “The most common failure pattern was that the agent wrote a solution, re-read its own code, confirmed it looks ok, and stopped.” That is a stop-condition failure, and the fix lived entirely in the harness.
The same pattern shows up from the opposite direction. Anthropic’s write-up on long-running agents describes a later session that “would look around, see that progress had been made, and declare the job done” (Anthropic, 2025). Their remedy was again structural: a progress file, an initial setup run, incremental commits. None of it touched the model.
Arithmetic explains why. A loop multiplies per-step reliability: an agent that is right 98% of the time per step finishes a 30-step task cleanly only about 55% of the time (0.98 to the 30th power; an illustration, not a measurement). The compounding error calculator lets you try your own numbers. The model sets the per-step rate; the harness decides how many of those failures get caught, fed back and corrected before they compound.
How much agent harness should you build?
Build as much harness as your conditions demand and no more. The book’s rule is a scaling law: “harness investment should scale with how long the codebase must live, how many people must trust the agent’s output, and how unattended the agent runs.” A watched prototype needs almost none; an overnight production agent needs most of the table above.
Chapter 18 sets two camps against each other honestly. Solo practitioners working in the foreground argue that most scaffolding is charade, because a capable model given a plain request does the job and the human watching catches the drift. Teams running agents unattended argue for deliberate guides and sensors, because nobody is watching. The chapter’s resolution is that both are right about their own room.
One rule it endorses without reservation: “scaffolding is a depreciating asset.” Some of every harness compensates for the current model’s weaknesses, and that part rots silently; on each model change, reread your standing instructions and delete what the new model has made redundant.
For the controls that do stay, Böckeler’s grid is the most useful map I know, and Chapter 18 adopts it. Guides steer before the agent acts; sensors check after and feed the result back into the loop. Each is either computational (a type checker, a test, a schema validator: cheap and deterministic) or inferential (a model acting as reviewer: slower, probabilistic, and in need of calibration). The ranking rule follows: wherever code can decide the question, let code decide it.
Framework or build your own?
Build your own first, and adopt a framework when you hit a concrete need you would otherwise have to build. The book’s version: “Adopt for a named need, never for the feeling that serious systems use frameworks,” and “Build from scratch to understand. Adopt a framework to scale—and only once you can name what it is saving you.”
The minimal version is genuinely small. Thorsten Ball’s walkthrough builds a working code-editing agent in “less than 400 lines of code, most of which is boilerplate,” and summarizes the machine as “an LLM, a loop, and enough tokens” (Ball, 2025). Anthropic’s engineering guidance, written for teams it had watched ship agents, says the same: “start by using LLM APIs directly,” and if you do use a framework, “ensure you understand the underlying code,” because “incorrect assumptions about what’s under the hood are a common source of customer error” (Schluntz and Zhang, 2024).
What a framework adds is plumbing, and Chapter 3’s catalog is short. The table pairs each item with the signal that you need it.
| What a framework supplies | The named need that justifies it | Cost you accept |
|---|---|---|
| State persistence and resumption | Runs outlive a process: human approvals, deploys, hour-long tasks | A runtime to operate; replay rules to respect |
| Retries, timeouts, backoff | Rate limits and transient failures show up in your traces | Defaults you must read, or they retry the wrong things |
| Tracing | You debug runs you did not watch | Another data store; sometimes a vendor dependency |
| Memory and context management | Histories outgrow the model’s working window | Summaries you did not write, deciding what the model forgets |
| Multi-agent orchestration | One loop genuinely cannot hold the task | Coordination failures that a single loop never has |
None of these changes what the model does on a single pass, which is why, in the book’s words, “a framework can make an agent more dependable and cannot make it smarter.” Crash-safe resumption is the item most likely to justify outside machinery, and it has its own category: durable execution engines (Temporal, Restate and Inngest are examples of the category) that journal every step so a recovered run replays instead of repeating. One such vendor’s definition is blunt: “Durable Execution is crash-proof execution” (Wheeler, 2025). The deeper mechanics, journals, replay and receipts, are the subject of the sibling post on durable execution for AI agents.
Worked example: a ticket-triage agent from week one to month six
A support-ticket triage agent shows how the harness grows one named need at a time. It reads incoming tickets, looks up the customer, applies a label from a fixed set, and drafts a reply. The scenario is a hypothetical composite, and the numbers in it are illustrative.
Week one: the thinnest harness that works
You start with four parts and about a page of code. The tools are search_tickets, read_account, apply_label and draft_reply; the last one writes a draft and never sends. The loop looks like this:
history = [standing_instructions, ticket]
repeat up to 15 times: # step budget: mandatory
response = model(history, TOOLS)
append response to history
if response has no tool calls:
result = parse(response) # programmatic stop check:
if result.label in ALLOWED_LABELS: # the label must be valid
return result
append "label must be one of ..." to history
continue
for call in response.tool_calls:
if call.name not in TOOLS: append error; continue
output = try TOOLS[call.name](call.args) # errors become results
append (call.id, output) to history
return escalate("step budget exhausted", history) # fail loudly, keep transcript
Notice where the decisions sit. The stop rule is a validator, so the model’s “done” has to pass a check you wrote. The step cap is 15, a number you will later replace with one read from real traces. A failed tool returns its error as text the model can read, and the budget exit hands the transcript to a human instead of returning nothing.
Month two: the first named need
The agent now runs on the queue overnight, and the logs show two problems: the ticketing API rate-limits bursts, and one run in a few dozen dies on a timeout. That is a named need: retries with backoff on transient errors, timeouts on every tool, and a trace store so you can read the runs nobody watched. You could adopt a framework here; for one loop, a retry wrapper and structured logs are an afternoon’s work, and building them keeps every behavior visible. This is also the moment to estimate what a run costs now that retries can multiply calls; the agent cost-per-task estimator gives the order of magnitude.
Month six: the need that justifies outside machinery
Product asks for the agent to send replies for low-risk categories, with human approval for the rest. Approval means a run pauses for hours, survives deploys, and resumes.
Resumption means the danger the book warns about: the process dies after the send succeeds but before the checkpoint lands, and the resumed run sends again. Now you need durable state, an idempotency key on send_reply, and probably a durable-execution engine or a framework with real persistence. You can name exactly what it saves you, which is the test.
Look back at the sequence. The loop never changed. Every addition answered a failure you could point to in a trace.
How do you audit a framework’s harness before trusting it?
Audit a framework by asking where each harness component lives and what it does by default. A README describes the happy path; your questions should target the unhappy one. If a question has no clear answer in the documentation or the source, treat that as the finding.
- History: Can I see the exact messages sent to the model on every call, unedited?
- Stop rules: Is there a default step cap? Can I plug in a programmatic success check, or only trust the model’s “done”?
- Budgets: Can I cap tokens, wall-clock time and money, and what happens when a cap fires?
- Errors: Are tool failures returned to the model as readable text, retried, or swallowed?
- Recovery: If the process dies mid-run, what resumes, and can a resumed run repeat a side effect?
- Permissions: Where do I restrict what each tool may touch, and is the default open or closed?
- Tracing: Can I replay a failed run step by step without the vendor’s dashboard?
- Exit cost: How many of my tool definitions and prompts would survive leaving?
Where this advice stops applying
This framing has limits worth stating. If your problem has a fixed sequence of steps, you may not need an agent harness at all; a workflow with model calls in predefined places is cheaper, easier to test and easier to explain, and the decision between the two is covered in the AI agent architecture guide. If you are a user of a finished coding agent rather than its builder, your harness is the outer layer of instruction files, checks and hooks, and Böckeler’s essay is the better guide than the loop-building advice here. And a harness regulates what it can sense: as Chapter 18 notes, no sensor catches a misunderstood requirement.
The term itself may drift, and the answer to “what is an agent harness” could, two years from now, mean only the outer layer or only vendor runtimes. The object it points to will not change: the deterministic code that decides what a probabilistic model sees, may do, and must prove before its work counts.
The takeaway
The model is rented judgment, and the harness is where your engineering lives: that is the short answer to what is an agent harness, and the one to carry into a design review. Start with the four parts and a step cap, let each failure in a trace name the next component, and adopt outside machinery only when you can say what it saves you.
You will find the harness definition, the minimal agent and the stop rules in Chapter 3; harness engineering, the guides-and-sensors grid and the “how much harness” debate live in Chapter 18. Both are in the full book, and the formats are compared in Kindle or paperback: which edition of the book suits you. For the wider cluster, start at the agent fundamentals pillar.
Questions readers ask
- What is an agent harness in simple terms?
- An agent harness is the ordinary software around a language model that turns it into an agent: the loop that calls the model repeatedly, the tools it can request, the record of what has happened, the rules that stop the run, and the logging, recovery and guardrails added as the system matures. The model decides; the harness does everything else.
- Is an agent harness the same as an agent framework?
- No. The harness is a role in the system, and a framework is one way to obtain part of it. Every agent has a harness, whether you wrote it in a few hundred lines or imported it. A framework supplies prebuilt harness parts such as persistence, retries and tracing; you still own the tools, the permissions and the stop rules.
- Can you build AI agents without a framework?
- Yes. A minimal agent is a model client, a set of tools, a message history and a loop with a step cap, and practitioners have built working coding agents in a few hundred lines. Building without a framework is the best way to learn what each part does; adopt one when you hit a concrete need you would otherwise build yourself.
- What is harness engineering?
- Harness engineering is the practice of improving an agent by changing the system around the model instead of the model: adding a check, a rule, a tool or a recovery path each time the agent fails in a way that could recur. Mitchell Hashimoto described it in early 2026 as engineering a fix so the agent never makes the same mistake again.
- Does a better model make the harness unnecessary?
- It makes some of it unnecessary. Guidance written to work around an older model's weaknesses depreciates and should be deleted on each model change. The stop rules, permissions, persistence and verification remain, because they answer questions about your system that no model improvement can answer for you.
Sources
- Addy Osmani (2026). Agent harness engineering
- Birgitta Böckeler (2026). Harness engineering for coding agent users
- Mitchell Hashimoto (2026). My AI Adoption Journey
- Vivek Trivedy (2026). The Anatomy of an Agent Harness
- Vivek Trivedy (2026). Improving Deep Agents with Harness Engineering
- Erik Schluntz and Barry Zhang (2024). Building effective agents
- Thorsten Ball (2025). How to Build an Agent
- Anthropic (2025). Effective harnesses for long-running agents
- Ryan Lopopolo (2026). Harness engineering: leveraging Codex in an agent-first world
- Tom Wheeler (2025). The definitive guide to Durable Execution
- Wikipedia contributors (2026). Agent harness