Home / Blog / Agents at work / An Agentic Coding Workflow for Teams: Brainstor…

Agents at work

An Agentic Coding Workflow for Teams: Brainstorm, Spec, Plan, Execute

An agentic coding workflow for teams: the artifact, signer and meaning of done at each stage, plus a size rule and an owner rule. Copy the agreement.

By Enrique Gutiérrez · Published · 29 min read

An agentic coding workflow for teams takes the four-stage loop one engineer runs with a coding agent (brainstorm, spec, plan, execute) and decides, for each stage, the artifact it leaves in the repository, who signs it and what “done” means. It then adds two rules for the change that comes out: a size limit and a named owner.

Half of that sentence is borrowed and half is mine. The four stages and their artifacts come from Chapter 22 of the book this site belongs to, which follows one engineer through one day on mostly new code. The author of the recipe that chapter builds on is blunt about its reach: “These workflows are not easy to use as a team. The bots collide, the merges are horrific, the context complicated” (Reed, 2025).

The team layer on this page is this post’s extension of the book: who signs each stage, the size rule, the change classes and the ownership rule. I label it each time it appears. Nobody has measured it, and no team I know of has run it. Treat it as a hypothesis with a review date.

What is an agentic coding workflow for teams?

An agentic coding workflow for teams is a written agreement about four things: what gets written down before an agent writes code, who approves it, how large a change may be, and which named person answers for it. It is what a group needs, beyond a tool, to use coding agents in a software team.

The loop underneath is the book’s. In the brainstorm the agent interviews the engineer, one question at a time; the spec is compiled from that interview and checked into the repository. A plan cuts the spec into steps and is saved with a checklist. Then comes execution, where, in the book’s words, you “feed the plan’s prompts to a coding agent one at a time, reviewing between chunks” (Chapter 22, “The Canonical Loop”).

The canonical loop and its paperwork.
Figure 22.1 The canonical loop and its paperwork. Four phases hand off left to right, each depositing a checked-in artifact—spec.md, plan.md, todo.md—so the project’s state survives every context window rather than living in a conversation; a standing project-instructions file feeds all four. The accent arrow is the compound step: before moving on, whatever a run taught you is written back into that file, so the next run starts smarter than this one did. Reuse this diagram

The files are the portable part: “the value lives in the files, and the tools are interchangeable labor.” A team can therefore agree on the workflow without agreeing on a tool, and this page teaches none. It assumes what the book calls augmented coding, where merged code is held to the hand-written standard, and not vibe coding, which judges only behavior.

The one-page workflow: who signs what at each stage

One table holds the stages: each has an artifact, a signer and a definition of done. The first four rows are the canonical loop. The fifth comes from the chapter’s section on review, and the sixth from the same section as the loop.

The stages, the artifacts and the two approvals are the book’s: “specify, then plan, then decompose into tasks, then implement, then verify, with a human gate between each phase. Approve the spec before planning begins; approve the plan before tasks are cut” (Chapter 22, “Spec-Driven Development”). Nowhere does the book say who approves or how, so the last two columns are this post’s extension.

Stage Artifact in the repository Who signs, and how “Done” means
Brainstorm The interview transcript, kept or thrown away once the spec exists Nobody; the owner runs it Every question the agent raised is answered in the spec or listed there as open
Spec (feature, core or legacy) A spec file of a page or less One person other than the owner, by approving the pull request that adds the file Each acceptance criterion names the test that will decide it
Plan (feature, core or legacy) A plan file with a checklist; each step is one revertible commit Feature: the owner, by committing it before the first implementation commit. Core or legacy: one person other than the owner, by approving the pull request that adds it Each step names its files, its test and an estimate
Execute One commit per step, test first; the captured output of the change running The owner The suite is green, the owner has run the change, the output is captured
Review and merge A pull request with the eight fields of the agreement below One person other than the owner The reviewer can say it is what was intended and only that
Compound One line in the repository instruction file, in the same pull request That pull request’s owner and reviewer The agent’s mistake cannot recur silently

Approving the spec is the cheapest review in the loop, and the book says why: the gates are “placed where they are cheapest (you are reviewing a page of prose, at reading speed, while everything is still editable).”

The last row is the book’s compound step: “before moving on, ask what the agent should have known at the start, and write it into the file.” That file is the repository instruction file, the standing document that tells any agent how to build, test and behave in this codebase.

The working agreement, on one page

Here is the whole agentic coding workflow for teams as one page you can paste into a team channel. Every rule in it is this post’s proposal; the sections after it give the reasons and the sources. Set [N], [M], the core list and the review date before you post it.

WORKING AGREEMENT: changes written with coding agents
Applies to every change, whoever or whatever typed it.

Team settings
  [N]  most changed lines in one pull request, tests included
  [M]  most changed lines in a one-line fix, tests not counted
  Core list: [file in the repository] naming (a) core paths, where
    a mistake costs most, and (b) legacy paths, where we do not
    trust the tests
  This agreement is reviewed on: [date]

1. OWNER
Every pull request has one named person as owner: the one who asks
for review. Before asking, the owner has run the change, has read
every changed line, and can walk a reviewer through every changed
file without the agent. A pull request opened by an agent has no
owner until a person takes it on those terms.

2. CLASS (take the first that fits)
  throwaway       never merged, and never run against a production
                  system or production data (copies and exports
                  count as production data)
  core or legacy  touches any path on the core list
  one-line fix    at most [M] changed lines, described in one
                  sentence; or a mechanical change of any size;
                  or documentation or tests only, up to [N] lines
  feature         everything else
Not covered: code that is never merged but does run against a
production system or production data.

3. BEFORE CODE
  throwaway       nothing; label it; it never gets a pull request
  one-line fix    nothing; the one sentence is the spec
  feature         a spec file, approved; a plan file, committed by
                  the spec's owner in the first implementation
                  pull request, ahead of any implementation commit
  core or legacy  a spec file, approved; a plan file, approved. On
                  a legacy path the plan's first step is tests
                  that pin current behavior, in a pull request of
                  their own, merged before any behavior changes.
"Approved" means a person other than the owner approved the pull
request that added the file, and it merged before the first
implementation pull request opened. Spec and plan pull requests
hold only those files and sit outside the classes and the limits.
A spec has one owner. Each pull request under it has its own
owner and takes the spec's class. If one turns out to touch a
core-list path the spec did not name, stop: the spec is approved
again as core or legacy. Spec and plan live in the repository
that gets the first pull request; the others link to them.

4. SIZE
One idea per pull request, at most [N] changed lines.
A file regenerated by a command because of a hand-made change in
the same pull request (a lockfile after a dependency bump) stays
with that change, is not counted, and is listed with its command.
A mechanical change (the whole diff is the output of a command: a
rename, a reformat) goes alone in its own pull request, with the
command, and has no line limit.

5. PULL REQUEST DESCRIPTION (eight fields, none left empty)
  Owner: [name]. I ran this, read every changed line, and can
    walk a reviewer through every changed file without the agent.
  Class: [one-line fix | feature | core or legacy]. Core list:
    [touches none | touches: ...]. What it does, in one sentence:
    [...]. Spec: [link | none]. Plan: [link | none].
  Size: [n] changed lines. Regenerated files and their commands:
    [... | none]. Mechanical: [the command | no].
  Ran: [command and output as captured, or a screenshot; never a
    summary. A configuration value: the change seen working
    outside production. No executable code: "nothing to run".]
  Tests: added [...] / edited [..., why] / deleted [..., why] /
    none. Seen to fail first: [output | no new tests]. (A test
    that pins existing behavior: seen to fail when I broke that
    behavior on purpose.)
  Beyond the plan (one-line fix: beyond the one sentence):
    [... | nothing]
  Other uses of what I changed: searched for [names]; found [...]
  Instruction-file line for a mistake the agent made: [the line,
    added in this pull request | none needed]

6. REVIEWER
One person other than the owner. An automated reviewer may run
first; it does not count as the reviewer.
  a. Check the class: changed paths against the core list, line
     count against [M] and [N], and whether a change classed as
     documentation or tests only is that. Wrong class: return it.
  b. Return it unread if a field is empty, the size is over the
     limit for its class (a mechanical change has none), or a
     spec or plan the class needs is missing or was not approved
     in time.
  c. Read the spec and plan, then every test change line by line,
     then the diff:
       one-line fix    line by line
       feature         for intent and scope; any new entry point
                       (a route, a handler, a query) line by line
       core or legacy  line by line, after the owner has walked
                       you through the change
     For a mechanical change or a regenerated file, in any class,
     re-run the command instead and read only what differs.
  d. "Line by line": you judge whether each changed line is
     correct. "For intent and scope": you look at every changed
     block and trace it to a plan step; the tests and the captured
     run carry the question of whether each line is correct.
  e. In any class you may ask the owner for a walk-through. An
     owner who cannot give one gets the pull request back.

How big may an agent-written pull request be?

An agent-written pull request should hold one idea and stay under a line limit the team sets, tests included; section 4 of the agreement has the two exceptions. Over the limit, it is split or returned unread.

The book’s version is about commits: “size tasks so that each lands as its own isolated, revertible commit,” because “the unit of revert is the unit of courage” (Chapter 22, “Git as the Safety Net”). A pull request is the same unit one level up, and an agent makes splitting one cheap.

I can’t give you the number. One large company’s public review guidance, undated and written before agents, says: “100 lines is usually a reasonable size for a CL, and 1000 lines is usually too large, but it’s up to the judgment of your reviewer” (Google Engineering Practices; a CL is that company’s word for one change under review). Start inside that range, time two or three careful reviews on your own code, and set [N] at what one person reads well in one sitting.

Take the arithmetic to a manager who calls the limit strict. The rate below is a placeholder to replace with your own measurement: if a careful reviewer on your codebase covered 300 changed lines an hour, a 2,000-line pull request would be nearly seven hours of one person’s attention, and a 9,000-line one thirty. One engineer asked a forum in October 2025 how to review exactly that: “It span across 9000 LOC, 63 new files.”

Refusal on these grounds has precedent. In November 2025 a compiler project closed a generated pull request of 13,323 added lines across 147 files without reviewing it on its merits. The maintainer who closed it listed several reasons: the generated files credited another person as their author, nobody had discussed the design, the change was hard to review and lightly tested, and reviewers were scarce. Two lines from those comments: “We would appreciate people discussing design before they dump 13K-lines PRs on us,” and the closure was “not a value judgment of the code” (ocaml/ocaml pull request 14369).

Approving a diff nobody could have read is the coding form of human-in-the-loop review theater. Chapter 12 of the book defines review theater as “diligence performed at a volume where it can no longer be real.”

Do you review the decisions or the diff?

You review both, in a fixed order: the spec and plan first, then the tests, then the diff. The order and the two reading modes in section 6 are this post’s extension; the ingredients are the book’s.

One engineer asked in June 2026 whether to keep reading code or only decisions, because “you can’t read 8x the volume by reading harder. the math just doesn’t work.” The order answers by moving judgment to where it is cheapest and keeping the diff small enough to read.

The book describes what a well-formed agent-written pull request carries: “the git discipline supplies small, one-idea commits; the spec supplies the context and the declared scope; the demo document supplies the evidence” (Chapter 22, “Agentic QA and Review”). The section’s argument ends on a line worth pasting into a pull-request template: “Proof travels; assertion doesn’t.”

Reading for intent and scope means every changed block is looked at and traced to a plan step, while the tests and the captured run carry the question of whether each line is correct. Reading line by line means the reviewer answers that question too. The lighter mode is legitimate only when the artifacts are present: an approved spec, a plan, tests that decide the acceptance criteria, and output captured from a real run. When one is missing, the pull request goes back; nobody skims it.

Two things are read line by line in every class. The first is any test change, on the book’s rule: “Treat any diff that touches a test file as an escalation,” meaning “a thing a human looks at, every time, however routine the rest of the change.” The book also advises having tests “written or reviewed in a separate session”; here the reviewer is that second reader. The second is a new entry point, because the book warns that an agent adding a route tends to omit a permission check kept elsewhere, so “code layout is quietly a security control” (Chapter 21, on agent-legible code).

Who owns code that an agent wrote?

One named person owns every pull request: the person who asks for review. The owner has run the change, has read every changed line, and can walk a reviewer through every changed file without the agent. This ownership rule is this post’s extension of the book, built on published norms.

The book quotes a practitioner’s rule for the author’s half: “Don’t file pull requests with code you haven’t reviewed yourself” (Willison, Anti-patterns). The same writer has the plainest version of the principle: “A computer can never be held accountable. That’s your job as the human in the loop” (Willison, 2025).

Three open-source projects have published contribution policies that say the same thing about reading first and about accountability. I read each in October 2026 and found no date on any of them; they are examples of a norm, and I don’t cite them as authorities.

  • One project’s policy: “Contributors must read and review all LLM-generated code or text before they ask other project members to review it. The contributor is always the author and is fully accountable for their contributions” (LLVM AI Tool Use Policy).
  • A second makes the human submitter responsible for “Reviewing all AI-generated code” and “Taking full responsibility for the contribution” (Linux kernel documentation).
  • A third sets the bar at explanation: “If you can’t explain what your changes do and how they interact with the greater system without the aid of AI tools, do not contribute to this project” (Ghostty AI policy). It applies to outside contributors and exempts the project’s own maintainers; the rule on this page applies it to everyone, which is this post’s choice.

Running the change is in none of the three. It follows a practitioner’s rule the book quotes: “Never assume that code generated by an LLM works until that code has been executed.” All three also ask for disclosure of AI assistance, which the agreement omits because it applies to every change.

The third policy answers the reviewer who asked in August 2026: “Am I supposed to go through the code line by line and quiz him on what it does?” The poster was not the author’s manager (“I’m just another dev on the team”), which is why the rule has to be agreed before the review and not invented during it. The agreement’s mechanism is a walk-through the owner gives, and the reviewer’s remedy is to return the change.

Sign-off by seniority sounds stronger and is weaker. A March 2026 news report said one large company, after outages, would have senior engineers sign off AI-assisted changes by junior and mid-level engineers; the company called its review “part of normal business” (Ars Technica, from the Financial Times). One forum reader replied: “how can they possibly properly review reams of code to a sufficient degree they can personally vouch for it?” A senior signature on the diff adds no reading hours and does not make the author able to explain it.

What does the owner check before asking for review?

The owner checks eight things, one for each field of the pull request description in section 5 of the agreement, and each can be verified without anyone else’s help. The list is this post’s.

  1. Owner. I ran the change, I have read every changed line, and I can walk a reviewer through every changed file without the agent.
  2. Class. I compared the changed paths with the core list and stated the class; any spec or plan the class needs is linked, and approved where the class says so.
  3. Size. The change is one idea and under the limit for its class; regenerated files are listed with their commands; a mechanical change is alone in its pull request.
  4. Ran. The command and its output are attached as captured (a configuration value: the change seen working outside production; no executable code: nothing to run).
  5. Tests. Every added, edited or deleted test is listed with a reason, and the output shows each new test failing first.
  6. Beyond the plan. Nothing is in the diff that the plan (or, for a one-line fix, the one sentence) did not ask for, or I have listed it.
  7. Other uses. I searched for other uses of every function, field and configuration key I changed.
  8. Instruction-file line. If the agent made a mistake worth preventing, the line is in this pull request.

Item 5 is the book’s: “a test you never saw fail has proven nothing” (Chapter 22, “Test-Driven Development with Agents”). A test that pins existing behavior passes on its first run, so its proof is a failure when the owner breaks that behavior on purpose.

What do you write down before the code?

Before the code, a feature or a core or legacy change gets two short files in the repository: a spec of a page or less and a plan whose steps are each one revertible commit. Five lines is a legitimate spec. The fields are the book’s artifacts; the approval lines are this post’s.

The spec exists because of a cost the book states exactly. A prompt that describes a feature and asks for it “asks the model to fill every unstated requirement with a guess, and it schedules your discovery of the wrong guesses for after the code exists.”

For the plan, the book quotes the original recipe’s sizing rule: steps “small enough to be implemented safely with strong testing, but big enough to move the project forward,” with “no hanging or orphaned code” left between them. Its own addition is to “let the estimate double as a tripwire: a run that has blown far past the files or minutes you guessed is a run to stop and probe.”

SPEC: [short name]                        (a page or less)

Owner: [one named person]
Class: [feature | core or legacy]
Paths this will touch: [...]   On the core list: [none | which]

What and why: [two or three sentences. What changes for whom,
  and why now.]
Why this way, and what we rejected: [the alternative we declined
  and the reason. No test can keep this honest; the owner does.]
In scope: [...]
Out of scope: [...]
Must not change: [behavior, interfaces or data that stay exactly
  as they are]
Acceptance criteria (each one names the test that will decide it):
  1. [criterion] -> [test name or file]
  2. [criterion] -> [test name or file]
Open questions: [list, or "none". No open question may change the
  scope.]

Approved by: [name], in [link to the pull request that added
  this file]

PLAN: [same short name]

Sizing rule: each step is one commit that can be reverted alone,
small enough to test, large enough to move the work forward. No
step leaves unused code behind.

Steps (tick each as it lands):
  [ ] 1. [what] | files: [...] | test to write first: [...]
         | estimate: [n files, n minutes]
  [ ] 2. [what] | files: [...] | test to write first: [...]
         | estimate: [n files, n minutes]
  (Legacy path: step 1 is tests that pin current behavior, in a
   pull request of their own, merged before any behavior changes.)

Stop and ask a person when: a run goes far past its estimate; the
agent changes or deletes a test; the agent builds something the
plan did not ask for; the agent repeats the same failed attempt.

Pull requests: [one | one per step | steps grouped as: ...], each
at most [N] changed lines, each with its own owner: [names].
Unfinished work stays behind: [flag name | not needed]

Feature: committed by the spec's owner before the first
  implementation commit.
Core or legacy: approved by [name], in [link to the pull request
  that added this plan]

The “why this way” field is there because of a limit the book admits. Tests can keep the what honest, it says, but the why “reduces to no assertion, and goes stale silently. Write it down anyway, and assign someone to care.” Three of the stop conditions follow warning signs that Chapter 21 quotes from a practitioner: loops, unrequested functionality, and tests disabled or deleted.

Isn’t this waterfall again?

It is waterfall whenever the same pile of documents is demanded for a typo and for a subsystem, so the agreement demands different amounts for different changes. Writing intent down first is the core of spec-driven development with AI agents, and the objection to it deserves a straight answer.

The critics are describing something real. A consultant who tried spec-first tools on a small bug reported that one’s requirements document “turned this small bug into 4 ‘user stories’ with a total of 16 acceptance criteria,” and concluded: “To be honest, I’d rather review code than all these markdown files” (Böckeler, 2025).

The book’s defense is narrow, and I think correct: “Waterfall failed because validation arrived months after the decisions it should have corrected.” It applies “only while the loop stays fast and the gates keep firing.” The book also gives the two ends of the scale: “a throwaway script needs none of this; the system your company will still run in three years needs all of it.”

How much ceremony does each kind of change need?

Each change needs the ceremony of its class, and section 2 of the agreement decides the class from facts a reviewer can check: whether the change will be merged, whether it touches a path on the core list, how many lines it changes, and whether it is mechanical or touches only documentation or tests. The classes and the table are this post’s proposal.

The core list does the work. Common core entries are money, authentication, migrations and public interfaces. The legacy half names directories whose tests the team does not trust; on an old codebase, start it with the ones that have already burned you, and take a directory off when its pinning tests have landed. A configuration value is classed like anything else, so a one-line timeout change off the list is a one-line fix.

A throwaway that someone wants to keep is rewritten under the class it then fits. The book is strict here: “What is illegitimate is an artifact living under one value system after being built under the other” (Chapter 21).

Kind of change Before code Tests and evidence How the reviewer reads it Change class
A small fix, a configuration value or a dependency bump with its regenerated lockfile, off the core list, at most [M] lines; or documentation or tests only, up to [N] Nothing; the one sentence is the spec Suite green; the output of what the owner ran Line by line; a regenerated file by re-running its command one-line fix
A mechanical change off the core list (a rename, a reformat), any size, alone in its pull request Nothing; the one sentence and the command Suite green; the command Re-runs the command and reads only what differs one-line fix
New or changed behavior off the core list, in one pull request Spec approved; plan committed by the spec’s owner before the first implementation commit Each new test seen to fail first; suite green; output captured Spec and plan, tests line by line, then the diff for intent and scope; new entry points line by line feature
The same, spread over several pull requests, people or repositories One spec with one owner, approved; one plan; unfinished work behind a flag As in the row above, in each pull request, which has its own owner and stays under [N] As in the row above, for each pull request feature
Any change that touches a core path Spec approved; plan approved Each new test seen to fail first; suite green; output captured After the owner’s walk-through: spec and plan, then tests and diff line by line core or legacy
Any change that touches a legacy path Spec approved; plan approved; the plan’s first step is pinning tests in their own pull request, merged first Pinning tests seen to fail when the behavior is broken on purpose; then as for a core path As for a core path core or legacy
A prototype built to settle a requirement or a design question Nothing; labeled throwaway None; judged by its behavior No pull request, no reviewer throwaway
A one-off script or analysis that reads no production data Nothing The owner runs it and reads the output No pull request, no reviewer throwaway

A mechanical change that touches the core list is core or legacy like any other change there; only the reading differs, as section 6 says.

Does the workflow survive a legacy codebase?

It survives with different proportions: a shorter spec, tests before the change, smaller steps and a line-by-line read. Running a coding agent in a legacy codebase changes how much each stage weighs more than it changes the stages. I found no controlled comparison of agents on legacy versus new code within one team, so what follows is practice, not finding.

The book says outright that “the loop as described is a greenfield loop,” and that on established code “The unit of planning shrinks from the project to the task, and getting the right context in takes over the job the big spec did.”

A tech lead on what he called “an ancient codebase” wrote in March 2026: “I have to go over its generated code line-by-line and verify that earlier decisions I had already rejected aren’t slipping into the code again.” That is the job the “must not change” field and the instruction file are for.

The order of work I’d propose, as practice, matches the legacy row of the table:

  1. Questions before edits. Chapter 21 notes that when an agent only answers questions about a codebase, “Nothing is written, no risk is taken.”
  2. A short spec scoped to the task, with “must not change” filled in.
  3. A plan for that task, with smaller steps than you’d allow on new code.
  4. Pinning tests as the plan’s first step, merged before anything changes.
  5. For a deep refactor, the old implementation kept alive beside the new one until the end, then the closing question the book takes from a practitioner: “did I miss anything from the old implementation?”
  6. A search for other callers, and one line in the instruction file for every trap the agent fell into.

Do coding agents make a team faster?

On single, well-defined tasks, controlled studies have reported speedups, though neither of the two randomized single-task results on speed below (rows three and six) reaches statistical significance; at repository level, the published data show more code and lasting complexity. None of these studies tests the four-stage loop, or any team workflow. They describe the conditions the workflow responds to, and they do not confirm that it works.

The book says the same of its own chapter: “Nearly everything here comes from practitioners’ first-hand reports rather than controlled studies (the field is too young for better).” For each paper below I read the abstract and, in the full text on arXiv, the passages behind every figure quoted; the second row is the authors’ own post.

Study, year Who ran it Sample What was measured Result One limit
Becker and colleagues, 2025 An independent research nonprofit 16 developers, 246 tasks Time per task on their own mature open-source projects, AI allowed or not at random Tasks took 19% longer with AI allowed; beforehand the developers forecast 24% less time, afterwards they estimated 20% less 16 people; the authors “do not claim” they represent most software work
The same group’s update, 2026 The same nonprofit 57 developers (10 from the first study, 47 new); the authors call the data “an unreliable signal” A second experiment begun in August 2025 The authors restate the 2025 result and “believe it is likely that developers are more sped up from AI tools now” “only very weak evidence for the size of this increase”
Paradis and colleagues, 2024 A large technology company, on its own engineers and its own tools (vendor-run) 96 engineers Time on one “complex, enterprise-grade task,” with or without three AI features “about 21%” less time, “although our confidence interval is large”; in the full text the estimate is not statistically significant at the 5% level (p = 0.086) Assistance features, not agents; the authors “cannot assume” the effect applies more broadly
Borg and colleagues, 2025 Authors from a university, a consultancy and a code-analysis company whose commercial metric is one of the quality measures (vendor-affiliated; conflict of interest declared); preregistered 151 participants across two phases, 95% professional developers; 75 in the randomized phase, short of the planned 128 Whether other developers could evolve AI-assisted code as easily, without AI “no significant differences” in completion time or code quality One feature in one web application; the authors call the randomized phase “underpowered for small to moderate effects”
He and colleagues, 2025 University researchers 806 open-source repositories that adopted one agentic code editor, against matched controls Lines added, static-analysis warnings and code complexity before and after An estimated 281.3% more lines added in the first month, with gains gone after two months; warnings up 30% and complexity up 41%, persisting Open-source projects; adoption inferred from a configuration file; not teams with review gates
Shen and Tamkin, 2026 Researchers at an AI lab (vendor-run) 52 participants A quiz after learning a new programming library with or without AI Quiz scores 4.15 points lower on a 27-point quiz with AI, which the authors report as “a reduction in the evaluation score by 17% or two grade points”; no statistically significant time saving Learning something new, not daily work on familiar code

Read as a set, the table cuts both ways. The fourth row is the one pessimists skip: the study detected no difference in how easily other people changed code written with AI assistance, and was too small to rule out a modest one. The second row cuts the other way. The authors of the best-known slowdown result have not withdrawn it; they say it describes early 2025, that developers are probably sped up more now, and that their new data cannot show by how much.

The first row’s lasting finding is the gap between felt and measured speed. Developers who took 19% longer estimated afterwards that AI had cut their time by 20%, so a team that measures its speed by asking people inherits that gap. The fifth row is the clearest published evidence of technical debt from agent-written code: a burst of output, then lasting complexity.

Two vendor-published summaries, which I cite without figures because I read only the summaries, point the same way at team level: more and larger pull requests with longer review (Faros AI, 2025), and instability when change volume rises without strong tests and fast feedback (Google Cloud, 2025).

What should a team measure?

Measure what those studies say moves, plus the one queue this workflow creates, and choose the measures before the trial starts. This list is my synthesis:

  • How long a change waits for review.
  • Changed lines per pull request.
  • The share of changes reverted or followed by a fix within a window you set.
  • Lead time from first commit to production.
  • How long a spec waits for its approval.

Leave out lines written, pull requests opened and tokens spent. The first two rise whether or not anything improved. Token spend belongs in the budget conversation, where it depends on AI agent pricing models and on how many runs a change takes; it says nothing about delivery.

For a trial, apply the owner rule and the size rule to every pull request for a month, and run one upcoming feature through the whole loop. Write down the five numbers for the month before and the month of the trial. Two months prove nothing, but they show whether review wait and change size moved.

Where does this workflow cost more than it saves?

It costs more than it saves on small changes if [M] is set too low, on unsettled requirements, and on any team that approves artifacts without reading them. It also leaves several problems where it found them.

Ceremony that outgrows the change. A change one line over [M] is a feature with an approved spec. If five-line specs grow into pages, people will squeeze work under [M] to avoid them. A long generated spec also invites what the book calls false precision: a document that “feels rigorous while quietly dropping the detail that mattered.”

A second queue. Every feature now waits twice, once for the spec’s approval and once for review. A page of prose is quick to read, but it still waits for a reader, which is why the fifth measure is there.

Unsettled requirements. When nobody knows what is wanted, an approved spec records a guess. Build a throwaway, decide, then start again at the spec.

Outside the agreement. Code that is never merged but runs against a production system or its data is not covered: what it may read or write is a decision about the consequence of each action, and the consequence tier classifier sorts those. A throwaway whose output drives a decision gets no second reader either.

What it does not solve. An approval is only as good as the reading behind it, and nothing here measures the reading. Two changes that each pass alone can still conflict in meaning; one checkout per agent, which the book recommends (“give each its own checkout or worktree”), stops file collisions only. And understanding still decays. The book’s last section opens on an engineer who passed every gate and finds, three weeks later, that “the understanding is still gone. It was never anyone’s job to keep.” The owner’s walk-through narrows that gap on the day of the merge and promises nothing after it.

The takeaway

Agree who signs what before the next large pull request arrives, because a rule agreed in advance is a working agreement and the same rule invented during a review is an accusation. The book’s own summary of the loop is “Briefs down, receipts up.” A team’s part is to put a name beside each one.

Chapter 22, “The Coding Workflow in Practice,” develops the loop, the tests, the version-control habits and the explainer packet in full, and Chapter 21, “Coding Agents,” covers why coding suits agents and what agent-legible code looks like (both in the full book). The agents at work guide places this post beside its neighbors, or you can see the formats.

Questions readers ask

How do you review a 9,000-line AI-generated pull request?
You don't. Return it and ask for it split, one idea per pull request and each under the team's size limit, with the spec and plan linked so the split follows the plan's steps. Refusing has precedent: one compiler project closed a 13,323-line generated pull request in November 2025 without judging the code, asking for design discussion first, among other reasons.
When a team's output rises, do you still review the code or just the decisions?
Both, in that order. Decisions are reviewed first, in a spec approved before any code exists, where a wrong one costs a sentence to fix. The diff is still read, and a size limit keeps it readable. Reading the diff for intent and scope is legitimate only when the artifacts are present: an approved spec, a plan, tests that decide the acceptance criteria, and recorded evidence. If one is missing, the pull request goes back.
Who is accountable for code an AI agent wrote?
The person who asks for review. Three open-source projects have published contribution policies on this, and two of the three say so in nearly the same words: the human contributor is the author and is accountable. A workable team rule adds what the owner must be able to do: run the change, read every changed line, and walk a reviewer through every changed file without the agent.
What do you do when a developer submits AI-generated code they don't understand?
Ask the owner to walk you through the change, and return the pull request if they can't. Explaining is the owner's duty; reverse-engineering the change is not the reviewer's. A team that agrees this rule before the first incident can apply it between peers, without anyone needing a manager's authority.
Does a brainstorm, spec, plan, execute workflow work on a legacy codebase?
Not as written. The book and the recipe's own author both describe the four-stage loop as a greenfield loop. On old code, planning shrinks to the task, tests that pin current behavior are merged first, and the reviewer reads the diff line by line. I found no controlled comparison of agents on legacy versus new code, so this is practice and not a finding.

Sources

  1. Harper Reed (2025). My LLM codegen workflow atm
  2. Simon Willison (n.d., read October 2026). Anti-patterns (Agentic Engineering Patterns guide)
  3. Simon Willison (2025). Your job is to deliver code you have proven to work
  4. Google (n.d., read October 2026). Small CLs (Google Engineering Practices documentation)
  5. LLVM Project (n.d., read October 2026). LLVM AI Tool Use Policy
  6. Linux kernel developers (n.d., read October 2026). AI Coding Assistants (Linux kernel documentation)
  7. Ghostty project (n.d., read October 2026). AI Usage Policy (AI_POLICY.md)
  8. ocaml/ocaml maintainers (2025). Pull request 14369, "DWARF support for macOS and Linux" (closed unmerged; figures and comments read from the GitHub API)
  9. Ars Technica, syndicating the Financial Times (2026). After outages, Amazon to make senior engineers sign off on AI-assisted changes
  10. Birgitta Böckeler (2025). Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl
  11. Joel Becker, Nate Rush, Elizabeth Barnes and David Rein (METR) (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
  12. METR (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (METR's summary post)
  13. METR (2026). We are Changing our Developer Productivity Experiment Design
  14. Elise Paradis and colleagues (Google) (2024). How much does AI impact development speed? An enterprise-based randomized controlled trial
  15. Markus Borg and colleagues (2025, v3 2026). Echoes of AI: Investigating the Downstream Effects of AI Assistants on Software Maintainability
  16. Hao He, Courtney Miller, Shyam Agarwal, Christian Kästner and Bogdan Vasilescu (2025). Speed at the Cost of Quality: How Cursor AI Increases Short-Term Velocity and Long-Term Complexity in Open-Source Projects
  17. Judy Hanwen Shen and Alex Tamkin (Anthropic) (2026). How AI Impacts Skill Formation
  18. Faros AI (2025). The AI Productivity Paradox Report 2025 (the vendor's blog summary)
  19. Google Cloud (2025). Announcing the 2025 DORA Report: State of AI-assisted Software Development