What is agent harness engineering?
Agent harness engineering is the deliberate design of the controls around an AI agent: the things that steer it before it acts and the things that check what it did and feed the result back. The worksheet above takes the controls you already have, sorts them into the two-by-two grid from Chapter 18 of the book, names the gaps, and gives a reading of how much of this machinery your situation calls for.
The chapter starts from a problem of trust. A model is nondeterministic and does not know your context, and reviewing everything it produces by hand does not scale. An engineered harness attacks that from both sides. The chapter quotes the practitioner framework it follows: a harness “increases the probability that the agent gets it right in the first place, and it provides a feedback loop that self-corrects as many issues as possible before they even reach human eyes.”
If the word itself is new, what is an agent harness covers the definition. This page is about putting one together on purpose, which the chapter contrasts with letting it grow “by accretion.”
What is the difference between a guide and a sensor?
A guide acts before the agent does and a sensor acts on what the agent did. The chapter keeps the names because they are chosen well: “Guides are feedforward controls: they steer the agent before it acts.” Its examples are a conventions file, a how-to document, a rule, a skill, and “a code-generation template that makes the right structure the path of least resistance”.
“Sensors are feedback controls: they observe what the agent did and route the signal back into the loop so the agent corrects itself.” The test suite, the linter, the type checker, a dependency rule and a review agent are the chapter’s list.
The second half of the sensor’s definition is the part most often missed, and the worksheet asks about it for every control that acts after the fact. The chapter calls it the defining clause: “a check whose output lands on a dashboard for humans is monitoring; it becomes a sensor when its output lands back in the agent’s context.” A coverage report nobody feeds to the agent is useful to the team and does nothing to correct the agent. The tool puts such checks in a separate box under the grid.
You need both directions. The chapter describes what happens with one: “all sensors and no guides gives an agent that keeps making the same mistakes and keeps getting corrected; all guides and no sensors gives a growing rulebook nobody ever tests against behavior.” Those two failure modes are the first thing the diagnosis looks for.
What are the four cells of the harness grid?
The four cells come from crossing direction with what executes the control: a guide or a sensor, run by deterministic code or by a model. The chapter calls the second axis computational against inferential: “deterministic code, fast, cheap, reliable every time” against “run by a model, slower, costlier, probabilistic.”
| Guide: before the act | Sensor: on what it did | |
|---|---|---|
| Computational (code) | A codemod that rewrites deprecated calls | The type checker |
| Inferential (a model) | The conventions file | A review agent feeding verdicts back |
Those four examples are the ones the chapter places itself. The conventions file surprises people, since it is only a text file. It sits in the inferential row because “it steers by being read, and its effect is only as reliable as the model’s attention to it”.
The inferential sensor is the cell that needs the most care. It is a model judging another model’s work, which makes it the judge of the evaluation chapter in another role, and the same obligations apply: “the inferential sensor is an instrument, it carries documented biases, and its verdicts deserve calibration before they gate anything”. The tool marks that cell when you fill it, and the judge agreement calculator is the place to do the calibrating. The rule for choosing between the rows is the chapter’s too: “wherever code can decide the question, let code decide it, and spend the model where judgment is genuinely required.”
The chapter’s figure places a few more: tests and linters beside the type checker, code templates beside the codemod, skills and rules beside the conventions file. Eight of the thirteen one-click examples are placed by the chapter in this way. The other five (the how-to document, the dependency rule, mutation testing, the drift watcher and the team dashboard) the chapter names by direction or by category only, and the row the tool gives them is its own reading. The two groups are labeled above.
Where should each sensor run?
Each sensor should run at the point in the delivery pipeline that its cost allows: cheap ones on every change, expensive ones after integration, and a third group on a schedule. The chapter puts it in one sentence: “fast, cheap sensors run beside the agent on every change; expensive ones—mutation testing, broad architectural review—run after integration; and a third class runs on a schedule against the whole system, watching for drift”.
The worksheet asks for a stage for every sensor and lays them out in those three groups. It adds two notes of its own. One fires when a model-run sensor or mutation testing sits on every change, where its cost is paid most often. The other fires when nothing runs on a schedule, since the slow decay the chapter lists (dead code, aging dependencies, thinning coverage) builds up between changes.
How a sensor words its output also matters, because that output becomes context for the agent. The chapter’s advice is to write failure messages “as instructions to a capable colleague who cannot see your face.” The grid cannot check that for you, but it is worth doing for every sensor you list.
The habit that keeps a grid current is short: “any issue arriving more than once should leave behind a guide or a sensor that makes its next arrival less likely.” When the agent makes the same mistake twice, the question to ask is which control was missing. That is also the test for filling an empty cell. The tool suggests examples for each empty one and says to add them when an issue repeats, not before.
How much harness do you need?
You need as much harness as your situation’s lifetime, audience and lack of supervision demand, and for a watched prototype that is very little. The chapter gives a scaling law in place of a verdict: “harness investment should scale with how long the codebase must live, how many people must trust the agent’s output, and how unattended the agent runs.”
It reaches that law by taking the opposing view seriously. Practitioners who work alone at high speed argue that most scaffolding is wasted effort, and the chapter grants the strongest part of their case: “scaffolding is a depreciating asset.” Some of any harness exists to compensate for a weakness the current model has, and that part goes stale when the model changes. The two camps, it concludes, “are describing different rooms”. In the solo room, “the human is the sensor suite”. In a team’s room, with agents running while nobody watches, that attention does not stretch.
The two poles are stated plainly. “A solo prototype, watched line by line, wants almost none”. And “A team’s production codebase, worked by agents overnight, wants the full grid of the last section”.
The worksheet turns the three variables into dials with three levels each and adds them up, counting unattendedness twice because the chapter calls it “the steepest of the three multipliers”. The levels, the sum and the cut-offs between its three readings (almost none, a middle course, the full grid) are the tool’s, and so is the comparison it then makes with what you listed. Adding is a simplification of what the chapter calls multipliers, and the reading of almost none is kept for an agent that is watched in the foreground. A long list of controls in a situation that calls for almost none gets the same attention as empty cells in one that calls for the full grid.
One answer triggers a banner of its own. If the model has changed since you last reread your standing instructions, the chapter’s instruction applies: “on every model change, reread your standing machinery and delete what the new model has made redundant.”
What can a harness not do?
A harness cannot tell you that the software does what you meant, and it cannot vouch for its own sensors. The chapter is direct about both limits, and the worksheet inherits them.
The first is that a green light can mean nothing. In the experiment the chapter cites, a file with full statement coverage had no unit tests at all, and mutation testing found planted bugs the suite never caught. “A sensor can glow green while verifying nothing”. The grid counts controls. It has no way to know whether one works, which is why mutation testing is among the examples, as the sensor that checks the sensors.
The second is scope. What a harness regulates well is internal quality: structure, style, boundaries. “No sensor catches a misunderstood requirement, and correctness is not even defined if you never clearly specified what you wanted.” The goal the chapter quotes from the framework is modest for that reason: “A good harness should not necessarily aim to fully eliminate human input, but to direct it to where our input is most important.”
Two limits are the tool’s. Its examples come from software work because the chapter’s do, so an agent in another field needs its own equivalents of a linter and a type checker. And a reading of “almost none” is not permission to drop everything: even in the solo room the chapter keeps a standing instructions file and sizes tasks by blast radius, which the consequence tier classifier helps with. The agent verifiability scorecard asks the broader question of what signal tells you the work is right, the guide to agent security and operations places the harness beside guardrails and reliability, and both sections are in Chapter 18, Reliability, State, and the Harness (in the full book).
Questions readers ask
- What turns a check into a sensor?
- Where its output goes. Chapter 18 defines a sensor as a control that observes what the agent did and routes the signal back into the loop so the agent corrects itself. A check whose result lands on a dashboard for people is monitoring. The same check becomes a sensor when its output lands back in the agent's context.
- Why is a conventions file an inferential control?
- Because a model has to read it for it to do anything. The chapter says the conventions file steers by being read, and its effect is only as reliable as the model's attention to it. A codemod or a template enforces the same kind of rule in code, the same way every time, which makes it computational.
- How much harness should a solo prototype have?
- Almost none. The chapter says a solo prototype, watched line by line, wants almost none, because the person watching is the sensor suite. What still matters in that room is the operator's attention and sizing each task by what it could break. The need grows with the codebase's lifetime, the number of people relying on the output, and how unattended the agent runs.
- What is the difference between a guide and a guardrail?
- A guide steers the agent toward good work before it acts: a conventions file, a template, a skill. A guardrail blocks a harmful action at a boundary. The chapter describes the harness as the same instruments and guardrails pointed at quality instead of safety, and wired into a loop instead of a report.
- Does a full grid mean the agent's work is correct?
- No. The grid counts controls and cannot tell whether one verifies anything. The chapter reports a file with full statement coverage and no unit tests, where planted bugs went unnoticed by a green suite. It also notes that no sensor catches a misunderstood requirement. A harness clears mechanical findings so human attention can go to intent.
Sources
- Birgitta Böckeler (2026). Harness engineering for coding agent users
- Birgitta Böckeler (2026). Maintainability sensors for coding agents
- Peter Steinberger (2025). Just Talk To It: the no-bs Way of Agentic Engineering
- Peter Steinberger (2025). My Current AI Dev Workflow