Home / Learn / Agents at work

Guide

AI Agents in Practice: Coding, Research, Business and Trust

AI agents in practice: why coding led, the spec-plan-execute workflow, the verification gap in research, trust, recipes and classifiers. Start the guide.

By Enrique Gutiérrez · Last reviewed

AI agents in practice are the same loop pointed at a domain, and the domain decides how cheaply the output can be checked. Coding agents led because code has a compiler and tests. Research and business agents face a verification gap, since there is no compiler for facts. Product decisions (scope, error cost, gates, what users see) come first.

That is the subject of Part VII of the book, Chapters 21 to 27 (all in the full book), and this guide walks it chapter by chapter. Each section ends with one move you can make on Monday. The thesis underneath is the book’s: an agent is only as trustworthy as the signal you can use to verify it. Every domain in this part is a different answer to the compass’s first question, set out free in Chapter 1: what signal tells you it worked?

Why did coding agents lead?

Coding agents led because, in Chapter 21’s phrase, “the domain grades its own work.” Compilers and test suites give verdicts that are cheap, fast and incorruptible, so an agent can be wrong twice and right the third time at the cost of a few tokens.

The book calls each of these verdicts an oracle, a source of truth outside the thing being judged. Anthropic’s agent guidance lists the same properties: code solutions are “verifiable through automated tests,” and agents “can iterate on solutions using test results as feedback” (Schluntz and Zhang, “Building Effective Agents,” December 2024). It adds that “human review remains crucial for ensuring solutions align with broader system requirements.”

What you still own depends on where you sit on a spectrum. Vibe coding judges behavior and never reads the code. Augmented coding keeps the hand-coding bar while a machine types; Kent Beck’s definition opens “In vibe coding you don’t care about the code, just the behavior of the system” (Beck, “Augmented Coding: Beyond the Vibes,” 2025). Both are legitimate; the failure is drift, a vibed prototype that quietly becomes load-bearing.

The last lesson runs the adaptation backward. Agent-legible code is, in the chapter’s words, “code that survives being understood a few files at a time”: greppable names, permission checks next to the action, explicit errors, fast and truthful builds. The case study is a coding agent in a legacy codebase.

The transferable template, in one contrast.
Figure 21.5 The transferable template, in one contrast. On the left a loop closes because a ground-truth check (in accent) sits inside it: generate, check, iterate, and ship only what the check passes. On the right the same loop cannot close, because nothing plays the checker’s role, and generation trails off into ship-and-hope. The corrective move is the long arrow beneath—before pointing an agent at a domain, build the signal that lets the loop close. Reuse this diagram

The portable rule sits near the chapter’s close: “Before you point an agent at a domain, ask what the ground truth signal of success is; if the answer is ‘there isn’t one,’ then building one is the project, and everything else is decoration on an unverifiable loop.”

What does the coding workflow look like for a team?

The coding workflow that teams copy runs brainstorm, spec, plan, execute. The agent interviews you one question at a time, the transcript becomes a spec file in the repository, and a plan cuts it into small tested chunks executed one at a time. Chapter 22 sums up the day: “The agent took over the typing. The engineer kept the paperwork.”

The canonical loop and its paperwork.
Figure 22.1 The canonical loop and its paperwork. Four phases hand off left to right, each depositing a checked-in artifact—spec.md, plan.md, todo.md—so the project’s state survives every context window rather than living in a conversation; a standing project-instructions file feeds all four. The accent arrow is the compound step: before moving on, whatever a run taught you is written back into that file, so the next run starts smarter than this one did. Reuse this diagram

Spec-driven development generalizes the interview: write intent before code, implement against it, and when something is wrong, fix the spec and regenerate. It fails when the spec goes stale, so mature teams wire each acceptance criterion to a test. The chapter is blunt that most working teams live in the middle, keeping a real spec for one slice of the system (spec-driven development with AI agents works through the details).

Tests do two jobs with an agent. Red/green TDD proves the code works, and one failing test at a time paces an over-eager generator. Commits become savepoints, which is why the chapter can say “you can afford exactly as much boldness as your undo is boring.” Guard the verifier itself: any diff that touches a test file goes to a human, because agents under pressure sometimes delete the failing test.

Review then demands receipts. Have the agent exercise the change as a user would, record real command output, and never file a pull request you have not read. The last debt is understanding: an explainer packet and a five-question quiz before you ship keep the person in the loop a participant rather than a signature. An agentic coding workflow for teams turns this into a checklist, and each lesson the agent should have known goes into the standing project instructions file, the book’s compound step.

Why do research and business agents fail differently?

Research and business agents fail differently because one produces words and the other produces deeds. A research agent’s typical failure is a wrong sentence dressed exactly like a right one; a business agent’s is an action, such as a refund, a changed record or a sent email, that does not wait to be reviewed. Chapter 23 builds verification for the first and governance for the second.

The verification gap is the distance between how convincing research output looks and how cheaply it can be checked. It is measurable. The DeepTRACE audit of public search and deep-research systems found “large fractions of statements unsupported by their own listed sources,” with citation accuracy ranging from 40 to 80% across systems; deep-research configurations reached high citation thoroughness and still carried unsupported statements (Venkit et al., arXiv:2509.04499, September 2025). More citations, in other words, do not mean more truth.

Four partial substitutes for a compiler of facts.
Figure 23.3 Four partial substitutes for a compiler of facts. A draft passes through openable citations, grounding, spot-checks, and a judge; each catches some errors (the stubs into the trays below). None is a sieve without gaps: a thin residual of wrong-but-fluent claims (in accent) passes through all four and reaches the checked draft. Verification here is engineered and partial, never automatic and complete. Reuse this diagram

The field’s four partial substitutes stack: openable citations at the level of claims, grounding in sources placed before the model, your spot-check of the load-bearing claims, and a calibrated judge (is LLM-as-a-judge reliable? covers the calibration). The chapter’s summary: “There is no compiler for facts; there are only inspectors, and inspectors bill by the hour.”

The business agent, which the book calls the clerk, needs the governance owed a privileged employee: a defined role, scoped access, an audit trail and escalation triggers designed before go-live. Its job description is also the lethal trifecta assembled as a workflow (untrusted tickets, private records, outbound replies), so its safety comes from the scoping around it. Start internal and low-stakes, where, in the chapter’s words, “a bad week costs apologies, not customers.” That sequencing is most of why an agent works in the demo and fails in production.

How do research and clerk agents pay off differently?

Research and clerk agents pay off on the same three numbers, weighted differently: what the work is worth (value per instance times volume), what an error costs when it lands, and what it costs to check the work. The third is the one spreadsheets forget, and the chapter states the consequence plainly: an agent whose output costs as much to verify as to produce by hand returns approximately nothing.

The two hires earn their keep in different places. The research agent pays off on high-value breadth work that a person was always going to review. The clerk pays off on volume against a baseline you record before deploying (cost per ticket, handle time, resolution rate), which makes it the easiest agent to justify to a finance lead. The risk column inverts the intuition: deeds can be capped with permissions, tiers and gates, while a wrong fact costs nothing when written and an unknown amount when believed. In the chapter’s words, “Deeds fail loudly at a size you chose. Words fail silently at a size the future chooses.” Framing AI agent ROI against risk and AI agent pricing models carry the arithmetic further.

How do users come to trust an agent?

Users come to trust an agent well when their reliance tracks its actual reliability, a property called trust calibration. Chapter 24 opens with two colleagues on the same agent: one approves a confident summary unread and ships a mistake, the other re-derives every figure and switches the agent off. Both failures were built in the interface.

Over-trust produces misuse; under-trust produces disuse. The target is appropriate reliance: a reliable agent relied on more, an unreliable one less. Confident tone, polished summaries and a prominent green check raise trust without touching reliability. A confidence percentage is the model’s self-report in a numeric costume.

What moves people along the diagonal is substance made visible. Show the plan before execution, because “A user who can see the plan is supervising a strategy. A user who cannot is supervising a stream.” Attach evidence to every claim at the moment of the claim: a claim with proof costs a glance to check, and one without costs an investigation that busy people skip. The chapter’s rule is four words: “Design for the glance.”

Steering needs its own engineering. Checkpoints, an editable plan and an inbox the agent polls between tool calls let a correction land without the abort button’s cost of losing correct work. Approvals belong in a queue of typed review objects, with interruptions reserved for the top consequence tier, the approval gate logic of the oversight chapters. Together these make a supervision surface, serving the shift from doing the work to directing and verifying it. The chapter adds the counterweight: a surface that only displays is half-designed, because the person signing must also keep understanding the system. AI agents for product managers maps which of these decisions belong to the product owner.

Which recipes cut across industries?

A transversal recipe is a reusable shape of work that recurs across industries, composes into an ongoing capability, and maps onto patterns you already know; swap the industry and only the nouns change. Chapter 25 catalogs eight of them and lays every one along the autonomy slider, from a human in the loop to a human on it (human in the loop, on the loop).

Andrej Karpathy’s talk “Software Is Changing (Again)” (2025) compresses the demo-to-product gap into “Demo is works.any(), product is works.all().” And the far end of the slider has its own rule: “full autonomy is a fact about the sandbox, whatever the model’s competence.” A recipe earns its way loose on evidence, re-earned on every model change. AI agent use cases for startups applies the catalog; the agents vs workflows decision rule explains why several recipes are mostly workflow inside.

The eight recipes on the autonomy slider.
Figure 25.1 The eight recipes on the autonomy slider. Position marks where each recipe’s economics naturally let it run, and the ordering tells the chapter’s story: recipes that commit deeds sit tighter than recipes that hand a human a draft, and the only position of full autonomy is bought with a sandbox, where every outcome is disposable. The positions are starting points; each deployment moves along the slider on evidence, per Chapter 12’s rules for the dial. Reuse this diagram

How do the eight recipes compare?

The eight recipes differ most in where they sit on the slider and in how they typically fail.

Recipe What it does Slider position Typical failure
Autoresearch Runs an experiment loop overnight against a scalar metric Unattended, only because everything it touches is disposable Reward hacking a weak metric
Self-maintaining knowledge base Keeps an interlinked wiki current between you and your sources Assisted Quietly “resolving” a real contradiction
Deep-research analyst Plans, searches, synthesizes and cites Advisory Convincing and wrong; no compiler for facts
Ambient watcher Wakes on events in a stream and notifies, asks or proposes On the loop Injection delivered by the stream itself
Unstructured-to-structured Turns documents into records with field-level citations Acting, gated per field Silent corruption at volume
Queue triage and router Classifies, prioritizes, enriches and routes Routing loose, auto-resolve gated Confident misroutes nobody sees
Digital coworker Runs a documented process end to end against real systems Per action, by consequence tier Deed-level autonomy before the plumbing is proven
Premortem and red team Narrates or hunts failures in your plans Advisory Pessimism by construction

Can an agent be a classifier?

Yes: a classifier built on a language model is one well-configured call that reads an input and returns one label from a closed list, and it needs no training set. Chapter 26 builds a support-ticket triage instrument this way, and the finished classifier is a document you can read, diff and review.

The frame is an examiner and a marking guide. The model arrives educated but knows nothing of your categories, so “the definitions are the model”: each class gets an essence, inclusions, exclusions and boundary rules for known collisions, and the output schema closes the menu and gives it an exit (other, or cannot-classify). Examples work as a textbook rather than a database: a prototype per class, edges near boundaries, and contrastive pairs that differ only in the deciding feature.

Reason first, then the verdict. Make the model answer a few rubric questions drawn from the definitions before it names the class; text written before the label can inform it, text written after can only defend it. The worksheet also localizes failures and surfaces ambiguity through a flag. Scores follow the same craft: a small ordinal scale with each level defined like a class, since a bare 0-to-1 number has no definition and so can never be wrong.

The chapter’s development discipline made spatial: one pool of labeled tickets divided into three disjoint homes.
Figure 26.4 The chapter’s development discipline made spatial: one pool of labeled tickets divided into three disjoint homes. A handful live inside the guide as demonstrations; the development set is the sparring bench you re-run and study freely; and the sealed test set (in accent) is the envelope opened rarely, the only number you report. A ticket that tries to live in two sets at once is a leaked exam answer—struck out here, because disjointness is the whole point. Reuse this diagram

Then measure without fooling yourself, using the three-set discipline: in-guide examples, a development set you iterate against, and a sealed test set that alone reports accuracy. The chapter’s line: “The training set could be dispensed with. The discipline that grew up around it cannot.” At volume, put a cheap screen in front of the detailed pass and send anything uncertain onward. Build an LLM classifier is the tutorial; LLM classifier vs fine-tuned classifier weighs the classical road.

What lasts at the frontier, and should you build or adopt?

What lasts is the map rather than the weather: a few durable layers and one principle, pairing a generative model with a verifier it cannot argue with. Chapter 27 takes the thesis to its sharpest form. A model drafts specifications, invariants, property tests or proof scripts; a deterministic checker renders the verdict; the counterexample goes back into the context.

The limits are the lesson. A spec that passes can be the wrong spec, model-written properties tend toward the obvious, and models sometimes recite the textbook version of a protocol instead of modeling yours. The machine now writes the verification artifact cheaply; deciding what belongs in it stays with you.

The map of products the chapter draws has six layers: foundation models, frameworks, protocols, skills, runtimes and sandboxes, and evaluation and observability, with the human surface and governance running alongside all of them. Build versus adopt has a different default at each layer, summed up as a custody rule: “never store what you own inside what you rent.”

Layer Default What stays yours
Foundation model Rent A thin access layer so the engine can be swapped
Framework Rent or adopt once you can name what it saves you Your loop logic, kept in ordinary code so it stays portable
Runtime and sandbox Adopt The policy for what the sandbox may reach
Evaluation and observability Adopt the plumbing The eval set, rubrics and traces
Skills, prompts, tool contracts Own All of it; this is your system’s judgment
Protocols Implement the standard Nothing bespoke

The section ends on three short sentences: “Rent the engine. Own the judgment. Keep the receipts in your own drawer.” Build vs buy AI agents applies the rule to a real purchase.

Worked example: when has a triage classifier earned autonomy?

A triage classifier has earned autonomy when its sealed test set gives an accuracy interval whose worst case you can live with for the action it drives. Here is the arithmetic for an illustrative support team; the numbers are invented, the method is from Chapters 25 and 26.

The team wants incoming tickets sorted into billing, bug, feature-request and other. The should-this-be-an-agent tool puts the task on the rung “one judgment per item”: a single model call, not an agent. The team writes a two-page guide and labels 200 real tickets. Since no training set is needed, all 200 go to evaluation: 100 for development, 100 sealed.

After a few revisions the development set reads 94 of 100. The team opens the envelope once: 86 of 100. The gap is the overfitting the chapter warns about; the sealed number is the one to report. The eval sample-size calculator gives a 95% Wilson interval of about 77.9% to 91.5% for 86 of 100, so the honest claim is “somewhere between roughly 78% and 92%.” Narrowing that to ±5 points would take a sealed test set of about 185 independent labeled tickets at this accuracy, and about 385 in the worst case near 50%, the arithmetic behind how many eval examples you need.

Now set the slider per action, as the queue-triage recipe prescribes. Routing a ticket is reversible, so it can run on the lower bound of 78% with the ambiguity flag sending hard cases to a person. Auto-closing a ticket is a deed whose bad direction is invisible, so it stays gated until the team measures the false-confident rate (how often the classifier was sure and wrong) against the baseline it recorded before launch. Same instrument, two settings, each justified by a number.

Where does each idea live in the book and on this site?

Each idea in this part has one chapter where the book defends it and, usually, one page here that works it through.

Concept Chapter Go deeper Glossary
Ground truth and why coding led Chapter 21 Coding agent in a legacy codebase oracle
Augmented vs vibe coding Chapter 21 Agentic coding workflow for teams augmented coding
Brainstorm, spec, plan, execute Chapter 22 Spec-driven development compound step
The verification gap Chapter 23 Demo vs production verification gap
ROI against risk Chapter 23 AI agent ROI blast radius
Trust calibration Chapter 24 AI agents for product managers supervision surface
Autonomy slider and recipes Chapter 25 Should this be an agent? autonomy dial
Classifiers and the three sets Chapter 26 Eval sample-size calculator three-set discipline
Durable layers, build vs adopt Chapter 27 Build vs buy AI agents skill

What are the limits of this advice?

The advice in this part rests mostly on practitioner reports rather than controlled studies, and the book says so: the field is too young for better. Treat the workflows as reasoned starting points, and judge the reasoning, not the provenance.

Three limits apply. Percentages such as the 40-to-80% citation accuracy are dated snapshots of the systems audited at the time; the durable finding is that citation thoroughness and truth are decoupled. Recipe maturity varies, and several entries (the untended knowledge base, the always-on coworker) are still directions of travel rather than shipping capabilities. And the ROI frame needs your numbers: a baseline nobody recorded cannot be beaten, only claimed.

What should you carry from AI agents in practice?

Carry one question from this part into every domain: what signal tells you it worked? Coding answers it for free, research has to engineer it, business agents govern their way to it, and a classifier answers it with a sealed envelope. Where the answer is “nothing,” building the signal is the project.

The chapters of Part VII are in the full book. Read Chapter 1 free for the compass this part keeps steering by, and when you want the coding workflow, the verification gap and the classifier build, see the formats.

The chapters behind this guide

  1. Chapter 21: Coding Agents In the full book
  2. Chapter 22: The Coding Workflow in Practice In the full book
  3. Chapter 23: Research and Business Agents In the full book
  4. Chapter 24: Agent UX and Human Trust In the full book
  5. Chapter 25: Transversal Recipes: Big Patterns That Cut Across Industries In the full book
  6. Chapter 26: Agents as Classifiers and Scorers In the full book
  7. Chapter 27: The Frontier and How to Keep Learning In the full book

Tools and explainers for this topic

Tool

Eval sample-size calculator

A free eval sample size calculator for AI agents: confidence intervals for a pass rate and the number of tasks you need, computed in your browser.

Tool

Should this be an agent?

When to use AI agents, and when plain code, one model call or a workflow does the job better. Answer ten questions and see where your task sits on the ladder.

Explainer · 3 min

The agent loop: four beats and three exits

A three-minute animated explainer of the agent loop: the four beats of every pass, the history that is the agent's only memory, and three exits ranked by trust.

Questions readers ask

Should my team use coding agents?
Yes, if the codebase has a test suite the agent can run and someone reviews intent before merging. Coding agents work because tests give them a verdict at every step. Without a suite, the agent’s first job is to write one, and an unreviewed diff only moves the reviewing onto a colleague.
Which domains are ready for AI agents?
Domains whose output can be checked cheaply are ready first. Code leads because compilers and test suites grade it automatically. Research has to engineer its checks, such as openable citations and spot-checks, and business work needs governance around every action. Where no ground-truth signal exists, building one is the project, before any agent is.
What is the verification gap?
The verification gap is the distance between how convincing an agent’s output looks and how cheaply it can be checked. It is widest in research, where there is no compiler for facts. Claim-level citations, grounding, spot-checks of the load-bearing claims and a calibrated judge each narrow it, and none of them closes it.
What is a transversal recipe for AI agents?
A transversal recipe is a reusable shape of work, such as queue triage, document-to-record extraction or a deep-research analyst, that recurs across industries with only the nouns changed. Each recipe sits somewhere on the autonomy slider, from a human in the loop to a human on it, and earns looser settings on evidence.
When has an LLM classifier earned autonomy?
When its sealed test set gives an accuracy interval whose worst case you can live with for the action it drives. Set autonomy per action: a reversible step such as routing can run on the lower bound, while an action whose errors stay invisible, such as auto-closing a ticket, stays gated until the false-confident rate has been measured.