Home / Blog / Agents at work / Spec Driven Development With AI Agents: The Spe…

Agents at work

Spec Driven Development With AI Agents: The Spec as Source of Truth

Spec driven development with AI agents on one backend feature: the interview, the full spec, criteria turned into tests, and drift. Copy the worked spec.

By Enrique Gutiérrez · Published · 31 min read

Spec driven development with AI agents means writing what a change must do, and why, in a short file kept with the code before an agent writes any of it, then implementing against that file. The spec is the source of truth about what was intended and agreed. The running code remains the truth about what the system does.

That is a smaller claim than “the spec is the source of truth.” The split into two truths is my reading of where the book behind this site lands. Chapter 22 draws the strong version, in which “Intent lives in the spec; the code is a build artifact generated from it,” and then labels it: “the fully regeneratable end state is largely promissory as I write.” It places working teams short of that: “keeping a real spec for a defined slice of the system and touching it first when that slice changes.”

The rest of this page is one spec for one ordinary backend feature: the interview behind it, its criteria run as tests against implementations that are wrong on purpose, and the rule that keeps the file true after the code ships.

What does “source of truth” mean here, and what does it not?

Here “source of truth” means the place you look to learn what a change was meant to do, and the file you edit first when that intention changes. It does not mean the spec describes production; only the code that runs does that.

Code cannot tell you which of its behaviors were decided and which are accidents. The book’s figure draws the strong version: the spec generates the code, every repair goes back up through “fix the spec,” and patching the code in place is crossed out. The change rule further down departs from it in two rows, and says so there.

Spec-driven development inverts the usual hierarchy.
Figure 22.2 Spec-driven development inverts the usual hierarchy. Intent lives in the spec; the code is a build artifact generated from it. When something is wrong, the honest repair (in accent) loops back up to fix the spec and regenerate—not patch the output in place, the crossed-out path, which quietly makes the code the source of truth again and leaves the spec to rot. Reuse this diagram

The habit this reverses, in the chapter’s words: “the doc drifts, everyone quietly stops trusting it, and truth retreats into whatever ships” (Chapter 22, “Spec-Driven Development”). So the claim has three limits: the spec is the truth about intent, for the slice it covers, and you can rely on it as a description of the code only while its tests pass.

What is spec driven development with AI agents?

Spec driven development with AI agents is, in the book’s definition, to “write down what you want and why, precisely, in a reviewable artifact, before code exists; let the agent implement against the artifact; and when something is wrong, fix the artifact and regenerate rather than patching the output” (Chapter 22).

The most careful outside definition of the artifact comes from a consultant who reviewed three tools built for the practice. A spec, she writes, is a “structured, behavior-oriented artifact - or a set of related artifacts - written in natural language that expresses software functionality and serves as guidance to AI coding agents” (Böckeler, 2025).

Who approves a spec and how much ceremony each kind of change deserves are team questions. The agentic coding workflow for teams answers them and holds a blank spec-and-plan template. This page writes out the body of one spec in full, with two sections that template does not have (the decisions with their deciders, and the behavior), under the same header.

What would an agent build from a one-line request?

From a one-line request, an agent builds one of several defensible features and does not tell you which. Take a ticket that says: make invitation links expire after 7 days. Today a workspace admin invites a person by email and the link works forever. Every number in this example is illustrative.

What follows is a constructed illustration: nothing in this post reports what a coding agent did when given this request. A plausible first implementation adds an expiry time seven days after creation and passes a test that opens a link on day eight.

That build decided four things nobody asked about: whether invitations already pending die on deploy day, whether a link opened at the exact deadline is valid, whether an expired link looks different from a mistyped one, and what a person holding a dead link can do next. The book calls this “encoding a decision nobody made.” A practitioner gives the mechanism: “An agent does not refuse to guess. It guesses instantly and commits” (Sharda, 2026).

A benchmark paper, accepted at a 2026 machine-learning conference, agrees on a narrower question. It gave agents underspecified issues from an issue-resolution benchmark, with a simulated user available to answer, and reports that models “default to non-interactive behavior without explicit encouragement, and even with it, they struggle to distinguish between underspecified and well-specified inputs” (Vijayvargiya and colleagues, 2025). Two of the models it tested were exceptions on the second point.

Which questions does the interview have to answer?

The interview has to answer five questions about any feature: what happens to what already exists, what each number means exactly, what each new action does to old state, what each refusal answers, and what stays untouched or unbuilt. The list is this post’s arrangement; the interview itself is the book’s.

Chapter 22’s canonical loop opens with a prompt that, in its words, “asks the model to ask you.” The recipe it comes from prints the prompt in full: “Ask me one question at a time so we can develop a thorough, step-by-step spec for this idea. Each question should build on my previous answers, and our end goal is to have a detailed specification I can hand off to a developer. Let’s do this iteratively and dig into every relevant detail. Remember, only one question at a time” (Reed, 2025).

The book’s account of why it works: “the requirements come out of you, one forced question at a time, including the ones you did not know you had.” Here are the five questions asked of the invitation feature. The engineer runs the interview and writes the file; for product managers and the others in the third column, this is where their requirements enter the repository.

The general question Asked of this feature Who answers Where the answer lands
What happens to what already exists? Invitations pending on the day this goes live: expired at once, given a fresh period, or valid forever? Product owner Decisions (D5)
What does each number mean exactly? Seven days from creation or from the last send, and is “sent” the issue or the delivery? Hours or calendar days, on which clock? Valid or expired at the exact deadline? Product owner; tech lead Decisions (D1, D2)
What does each new action do to old state? A dead link needs a way back, so an admin can resend. Does the old link still work afterwards? Can an expired or an accepted invitation be resent? What does a replaced link answer once the invitation is accepted? Product owner Decisions (D3, D6); Behavior
What does each refusal answer, and to whom? Does an expired link look different from one that never existed? Support lead; security reviewer Decisions (D4)
What stays untouched, and what is not being built? An accepted link opened again; who may resend; reminders; per-workspace periods; resend limits; a resend log; cleanup Tech lead; product owner Must not change; Out of scope

When is the interview over?

The interview is over when each of the five questions has an answer written down or, for something that will not be built, has been moved by name to “out of scope.” A question about what already exists cannot be moved; it has to be answered. That stop rule is my proposal; the team post’s is looser (every question answered or listed as open), and I found no outside source that gives one.

Expect to discard most of what the agent asks. One forum commenter who has the agent write its questions into a file reports: “The questions are usually 80% useless but those 20% often do point me to stuff I have not considered.” The percentages are a figure of speech.

The worked spec, complete

This is the whole spec for the feature, compiled from the interview’s answers: a filled example to overwrite, with six sections in an order that is this post’s choice, under the team post’s header. It takes about three minutes to read.

SPEC: invitation expiry              (all numbers illustrative)

Owner: [one named person, who rereads the reasons]
Class: [feature | core or legacy]
Paths this will touch: [...]
Approved by: [name], in [the pull request that added this file]

Purpose
  An invitation link stops working a fixed period after it was
  last sent (D2), and a workspace admin can resend it. Today a
  link that was forwarded or leaked months ago still adds a
  member.

Decisions
  D1  The clock starts at the last send. Sent means the instant
      the server issues the link, whether or not the message is
      delivered.
      Rejected: at creation (a resend on day 6 would buy one day).
      Decided by: product owner.
  D2  The period is 168 hours from the send time, both read from
      one server clock in UTC. At exactly 168 hours the link is
      expired, so the stated period is never exceeded.
      Rejected: "the end of the 7th day" (needs a time zone for
      every workspace). Decided by: product owner (the period),
      tech lead (hours, clock, boundary).
  D3  A resend issues a new link for any invitation not yet
      accepted, expired or not, and the old link stops working
      at that instant.
      Rejected: the same link with a new deadline (a resend is
      how an admin recovers a link sent to the wrong person).
      Decided by: product owner.
  D4  An expired link answers "expired", always with the text
      "This invitation has expired. Ask for a new one." A link
      that was never issued answers "not found".
      Rejected: one answer for both (support must tell them
      apart). Decided by: support lead, security reviewer.
  D5  An invitation still pending at go-live gets 168 hours from
      go-live. Go-live is the instant, recorded once, at which
      this change is switched on in production.
      Rejected: counting from creation (links sent last week
      would die unannounced). Decided by: product owner.
  D6  Once an invitation is accepted, every link ever issued for
      it answers "already used", and it cannot be resent.
      Rejected: a replaced link goes on answering "expired" (its
      holder would ask for an invitation they no longer need).
      Decided by: product owner.

Behavior
  Opening a link answers exactly one of "accepted" (the member
  is added), "expired", "not found", "already used".
  Pending means not yet accepted.
  A resend of a pending invitation returns the new link. A
  resend of an accepted invitation, or of one that does not
  exist, returns no link and changes nothing.
  A link replaced by a resend answers "expired" while the
  invitation is pending, and is remembered for as long as the
  invitation exists.

Must not change
  A link that was already accepted answers "already used",
  however old it is.
  -> pinned by test_ac8_accepted_link_still_answers_already_used

Out of scope
  Reminder emails. A different period per workspace. A limit on
  how often an admin may resend, including a double submit. A
  record of who resent, and when. Deleting expired invitations.
  Who may resend: whoever may invite today; that check is not
  changed here.

Acceptance criteria (each names the test that decides it)
  AC1  Given an invitation sent 167h59m59s ago, when its link is
       opened, then the answer is "accepted" and the member is
       added.
       -> test_ac1_link_works_one_second_before_168_hours
  AC2  Given an invitation sent exactly 168 hours ago, when its
       link is opened, then the answer is "expired" with the
       text of D4 and no member is added.
       -> test_ac2_link_expired_at_exactly_168_hours
  AC3  Given an invitation first sent 10 days before it was
       resent, and resent 167h59m59s ago, when the new link is
       opened, then the answer is "accepted" and the member is
       added.
       -> test_ac3_resent_link_works_one_second_before_168_hours
  AC4  Given an invitation sent 1 hour ago and resent just now,
       when the old link is opened at the instant of the resend,
       then the answer is "expired" and no member is added.
       -> test_ac4_old_link_dead_at_the_instant_of_resend
  AC5  Given an invitation created 30 days before go-live and
       still pending, when its link is opened 167h59m59s after
       go-live, then the answer is "accepted" and the member is
       added.
       -> test_ac5_pending_at_go_live_works_one_second_before_168_hours
  AC6  Given an invitation created 30 days before go-live and
       still pending, when its link is opened exactly 168 hours
       after go-live, then the answer is "expired" and no member
       is added.
       -> test_ac6_pending_at_go_live_expired_at_exactly_168_hours
  AC7  Given a link that was never issued, when it is opened,
       then the answer is "not found" and no member is added.
       -> test_ac7_unknown_link_is_not_found
  AC8  Given an invitation accepted 30 days ago, when its link is
       opened again, then the answer is "already used" and the
       member list is unchanged.
       -> test_ac8_accepted_link_still_answers_already_used
  AC9  Given an accepted invitation, when an admin resends it,
       then no link is returned and the member list is unchanged.
       -> test_ac9_resend_of_accepted_invitation_changes_nothing
  AC10 Given an invitation sent 169 hours ago, when its link is
       opened, then the answer is "expired" and no member is
       added.
       -> test_ac10_link_still_expired_one_hour_after_168_hours
  AC11 Given an invitation resent exactly 168 hours ago, when
       the new link is opened, then the answer is "expired" and
       no member is added.
       -> test_ac11_resent_link_expired_at_exactly_168_hours
  AC12 Given an invitation that was resent and then accepted
       through the new link, when the old link is opened, then
       the answer is "already used" and the member list is
       unchanged.
       -> test_ac12_replaced_link_answers_already_used_after_acceptance
  AC13 Given no invitation with the id asked for, when an admin
       resends it, then no link is returned and the member list
       is unchanged.
       -> test_ac13_resend_of_unknown_invitation_returns_nothing

Below the header, no table, column, file or library is named, because how to build it belongs to the plan. One hole is left on purpose: the spec does not say what happens when the invited address already belongs to a member. Decide it before you copy the file.

At about 120 lines this is longer than the “page or less” the team post asks for. It carries thirteen criteria and six reasons for a feature with a deadline and existing data, and four of those criteria exist only because of the second round of tests described below. Five lines is still a legitimate spec for a smaller change.

What is a spec not?

A spec is not the plan, the standing instructions, the ticket or the product requirements document, and most bad specs are one of those four wearing its name. The table is this post’s synthesis.

Document What it is for How long it lives Who reads it
Spec What one change must do and why, in a form a test can decide The criteria and decisions: as long as the behavior does The approver, the agent in sessions that touch this feature, whoever changes it next
Plan The steps, files and order of the work Until the change ships The agent executing it and its owner
Standing project-instructions file What is true of every change: build and test commands, conventions, house rules As long as the repository The agent, in every session
Ticket Who does a unit of work, and when Until it closes The team and its tracker
Product requirements document Why something is worth building, and for whom Until the product bet is settled Product, design and whoever funds it

The tool review quoted above draws the line between the spec and the instructions file. General context documents, she writes, “are relevant across all AI coding sessions in the codebase, whereas specs only relevant to the tasks that actually create or change that particular functionality” (Böckeler, 2025; the grammar is the source’s). The book describes the same file as “always loaded, carrying what the agent cannot infer from the code.”

How does an acceptance criterion become a test?

An acceptance criterion becomes a test by direct translation: the given is the fixture, the when is one call, the then is the assertions, and the criterion’s identifier goes into the test’s name so that a failure names the sentence that stopped being true.

In the book’s words, the mature practice “wires each criterion to a test, so that the spec’s health shows up in the build.” Here are AC2 and AC4 in illustrative Python, using only the standard library, as they were run.

# Illustrative Python. T0 is a fixed instant (10:37:21 UTC); H is one hour.
def test_ac2_link_expired_at_exactly_168_hours(self):
    _, link = self.inv.send(EMAIL, T0)
    answer = self.inv.open_link(link, T0 + 168 * H)
    self.assertEqual(answer.status, "expired")
    self.assertEqual(answer.message, EXPIRED_TEXT)
    self.assertNotIn(EMAIL, self.inv.members)

def test_ac4_old_link_dead_at_the_instant_of_resend(self):
    inv_id, old = self.inv.send(EMAIL, T0)
    self.inv.resend(inv_id, T0 + H)
    answer = self.inv.open_link(old, T0 + H)
    self.assertEqual(answer.status, "expired")
    self.assertNotIn(EMAIL, self.inv.members)

The translation forces two design constraints on the code: time has to be passed in or replaceable, or the fixture cannot say “168 hours ago,” and go-live has to be an instant a test can set. Both go in the plan.

Which tests pass against a wrong implementation?

More than I guessed. The invitation rule became a small Python module, and tests were run against implementations that are wrong on purpose, in two rounds. Each wrong implementation was built around one named defect. None is what a coding agent produced when given the spec, and nothing here reports how an agent behaves.

Round one. Three tests drafted the way loose criteria suggest (expired on day 8; works on day 6; old link dead after a ten-day-old invitation is resent), then the nine criteria of this spec’s first version, against three wrong implementations. The draft tests passed two of the three. The nine criteria caught all three.

Round two. Seventeen more variants: fifteen wrong by that first version and two that chose differently where it was silent. Eleven of the seventeen passed all nine tests. The spec printed above is the revision, and two variants still pass its thirteen.

Wrong on purpose 3 draft tests First 9 criteria Revised 13 criteria
Every pending invitation is expired 1 fails fails AC1, 3, 5, 8, 9 fails AC1, 3, 5, 8, 9, 12
Valid at exactly 168 h (> for >=) pass fails AC2, 6 fails AC2, 6, 11
A resend never cancels the old link pass fails AC4 fails AC4
Clock runs from creation; a resend does not restart it pass fails AC3 fails AC3
Old link works for 30 s after a resend pass pass fails AC4
Expired links answer “not found” 2 fail fails AC2, 4, 6 fails AC2, 4, 6, 10, 11
Pending at go-live counted from creation pass fails AC5 fails AC5
Every run of the go-live step renews every pending invitation pass pass pass
Expires one second early pass fails AC1, 5 fails AC1, 3, 5
Seven calendar days in UTC, not 168 hours pass pass fails AC1, 3, 5
Send time stored truncated to the hour pass pass fails AC1, 3
“expired” carries the wrong text pass pass fails AC2
Resend of an accepted invitation removes the member pass pass fails AC9
Only the latest replaced link is remembered pass pass pass
Wall clock used; the time passed in is ignored 1 fails fails AC2, 6 fails AC2, 6, 10, 11
Resend keeps the same link with a new deadline 1 fails fails AC4 fails AC4
Resend of an unknown invitation raises an error (first version silent) pass pass fails AC13
A replaced link answers “expired” after acceptance (first version silent) pass pass fails AC12
A resent link never expires pass pass fails AC11
Expired only at the exact instant (== for >=) 1 fails pass fails AC10

Run with Python 3.14 and python3 -m unittest, each test file once per implementation. The first three rows are round one. The correct implementation passes all three files.

What closed the gaps?

Five changes closed nine of the eleven gaps, and the table’s last row is the one to remember. An implementation that expires a link only at the exact instant passes the boundary pair and fails the draft test “expired on day 8.” So AC10 joins AC1 and AC2: the pair does not replace the far test.

  • A pair on every deadline. AC1 and AC2 sit one second apart at 168 hours. The first version had no pair after a resend, and a resent link that never expires passed. AC3 and AC11 are that pair.
  • An example at the stated instant. D3 says “at that instant.” The first AC4 looked one minute later, and a 30-second grace period passed.
  • A starting state that leaves one explanation. AC4 resends an invitation one hour old, so only the resend can explain “expired.” The tests’ reference instant was midnight, where seven calendar days and 168 hours agree; moving it to 10:37:21 caught two variants without changing a criterion.
  • Every stated fact asserted. AC2 now checks the text of D4, and AC9 checks the member list.
  • Silences decided. D6 and one Behavior line say what the first version left open; AC12 and AC13 test them.

What still gets through?

Two wrong implementations pass all thirteen tests, and I am leaving them. One renews every pending invitation each time the go-live step runs; the spec says go-live is “recorded once,” and no criterion runs it twice. The other remembers only the latest replaced link, so after two resends the first link answers “not found”; no criterion resends twice.

Both are prose with no test, the weakness this page warns about. Each would cost one more criterion about an unusual sequence, and a spec nobody finishes reading protects less than one with two known holes. Thirteen passing tests do not mean these are the only two.

What do passing tests not prove?

Passing tests prove that the assertions held, and nothing about what no assertion mentions. Chapter 21 says it of the test suite: “A test suite is a verdict on behavior: these assertions held, or this one failed, with this trace.” It also gives the caveat: “green is a fact about the tests.”

A change can pass and still be wrong in at least three ways.

The behavior nobody wrote down. None of the thirteen tests says what happens when the invited address already belongs to a member. The book’s sentence covers the case: “A patch can pass every test and still be wrong: wrong for the architecture, wrong for the conventions the team keeps, wrong in a corner no assertion reaches” (Chapter 21, “Why Coding Led”).

The sentence that is only prose. D4’s reason, that support must tell an expired link from an unknown one, becomes no assertion.

The test that was changed. A test is an oracle, which the book defines as “a source of truth outside the thing being judged,” and it stays outside only if the implementer cannot rewrite it.

Test-driven development wired to the test-commit-or-revert guardrail.
Figure 22.3 Test-driven development wired to the test-commit-or-revert guardrail. A failing test steers the agent, and running the suite decides. A pass commits the green state and licenses the next test; a fail (the accent path) reverts the working tree to the last green, so a bad attempt costs exactly one discarded try. The bracket marks the tests as guarded—a diff that touches them is a human’s call, never the agent’s escape. Reuse this diagram

Do agents really satisfy tests the wrong way?

On tasks built to tempt them, yes, and the published rates vary widely with the task. A 2025 benchmark preprint mutated the unit tests of existing coding tasks so that they contradicted the written specification, and told the agents to prioritize the specification. Any pass then, as the authors put it, “necessarily implies a specification-violating shortcut.”

One frontier model passed 76% of one such set built from multi-file repository tasks and 2.9% of a set built from algorithmic problems (Zhong, Raghunathan and Carlini, 2025; I read the abstract and the passages behind each figure). Tests and specification conflict there by construction, which ordinary work does not do, and the figures belong to that model and year.

Their recommendation transfers: “either hiding test files entirely or restricting them to read-only access during implementation, when feasible,” though read-only access “does not eliminate other cheating methods such as special-casing or operator overloading.”

How do you guard the tests?

Guard them with structure, in three ways the book gives: see each test fail, keep its author apart from the implementer, and have a person read every change to a test file. A test must fail first because “Red is the moment the verifier itself gets verified.” And: “Treat any diff that touches a test file as an escalation,” glossed as “a thing a human looks at, every time, however routine the rest of the change” (Chapter 22, “Test-Driven Development with Agents”).

A 2004 essay on tests as specifications made the same point before agents existed: “The value of the double check is very much tied into using different methods for each side of the double check” (Fowler, 2004). The weakness reappears one level up, in how to test AI agents that are themselves the product: a check the thing under test can influence has stopped being independent.

How does a spec go stale?

A spec fails in three ways the book names, and only the first is staleness in the strict sense: it is shelved, it was never as precise as it looked, or it is frozen into a handoff. Each ends with an agent building from a document that does not say what anyone wants.

The stale spec. “The classic death is the stale spec: written at kickoff, admired, shelved; and now an agent is confidently building from a document that lies, which is strictly worse than building from no document at all, because the confidence is machine-speed.”

False precision. A long interview is summarized into “a document that feels rigorous while quietly dropping the detail that mattered.”

One practitioner describes the ordinary route to the first: “someone patches the code directly to stop the bleeding, and the spec is now quietly wrong.” Later an agent reads the old spec to orient itself. “The model did not hallucinate. The spec did” (Sharda, 2026).

When does the spec change, and when does the code?

The spec changes first whenever the wanted behavior changes or turns out to have been wrong or unstated; the code changes alone when the spec was right and the code was not. The table is this post’s wording of the chapter’s position.

Chapter 22 states the strong form, “fix the artifact and regenerate rather than patching the output,” and the working form, a spec that teams keep “touching … first when that slice changes.” Rows 3 and 5 permit a patch to the code with no change to the spec’s decisions. That departs from the strong version the figure draws, and it is my reading.

First find out which row you are in: reproduce the report and read what the spec says about it.

# What you found What changes first Then
1 Nothing is wrong: the code does what the spec says, and nobody a decision names wants it changed Nothing Answer the report
2 The people a decision names want different behavior (the period becomes 72 hours) The decision and its criteria in the spec, changed by whoever the decision names; then ask the first interview question again, about what was created under the old rule The tests, now red; then the code
3 The code does not do what the spec says (a link works at 168 hours) A failing test that shows it, if none already fails The code. The decisions do not change; if the failing test is new, add its criterion to the spec, because an example was missing
4 The spec is silent, or says something its deciders never meant (nobody decided what an invitation to an existing member does) Stop. The spec gets the decision, with who made it A criterion and its test; then the code
5 A production fix already went in directly Nothing; it is done The criterion and its test, in the same change if it can be, and otherwise as the next change, with the spec marked out of date until then
6 A test fails and the behavior is as the spec says (the test was wrong, or leaned on a library) The test, read by a person Neither the spec nor the code
7 A refactor Neither the spec nor any assertion If an assertion had to change, the behavior changed: go to row 2. A test that changes only the name it calls is still a refactor, and a person reads that diff

Where does the spec live, and what happens after the change ships?

The spec lives in the repository beside the code, and after shipping its parts age at different speeds. Filed there, the book says, it “exists outside any chat session, survives every context window, and can be handed to a different model, a different tool, or a colleague.” It still has to be put in front of each session that touches the feature; a spec lost when the context was cleared is, for that run, no spec at all.

After the change ships, the acceptance tests join the regression suite and keep running. The purpose paragraph will age in its “today” sentence, and I would let that happen; it must not carry a number or a rule that a decision also states.

A failing acceptance test is the alarm for a criterion. Prose with no test has no alarm, which is the book’s point about reasons: the why “reduces to no assertion, and goes stale silently. Write it down anyway, and assign someone to care.” That someone is the owner in the spec’s header.

Isn’t the code the real source of truth?

For what the system does, yes, and the people who say so are right about that. “There is only one source of truth and that is the source code,” one forum commenter wrote in February 2026. The other camp answers with a working rule: “if something’s off, I don’t go and modify the code, this is when specs and code start to drift” (a September 2026 comment).

They answer different questions: code records behavior, and a spec records intent. The skeptics’ practical point also stands. A commenter who writes “I rarely if ever update the specs if it finds bugs as there’s no need” is in row 3 of the table whenever the spec was already right, and in row 4 whenever it was not.

Review your own spec before planning starts

Run these ten checks on your draft before any plan is written. Each can be answered yes or no by reading the document (item 3 also needs your list of interview answers). The first version of the worked spec failed items 5, 6 and 7 while I believed it passed; the revision passes all ten on my reading, with one note on item 5 and one exception to item 9: the hole left on purpose above, which you must close before the checklist is true of your copy. The list is this post’s.

  1. Purpose. It says what changes, for whom and why, in three sentences or fewer, and repeats no number or rule that a decision states.
  2. Deciders. Every decision names who made it and one rejected alternative with its reason.
  3. Nothing lost. Every answer from the interview appears as a decision, a behavior line, a must-not-change line or an out-of-scope entry.
  4. Numbers. Every number has a unit, the event it is measured from and the clock it is read on.
  5. Boundaries. Every limit or deadline has an acceptance criterion on each side of it, as close as the unit allows, and every number that sets a limit has one example well past it.
  6. Answers. Every answer the feature can give a caller, each success, each refusal and each fixed text, appears in at least one acceptance criterion.
  7. One behavior. Every acceptance criterion has one starting state, stated in full, one action and an outcome someone outside the code can observe.
  8. Tests named. Every acceptance criterion and every must-not-change line names the test that decides it.
  9. Scope. “Out of scope” has at least one entry, and nothing left open in the spec could change what gets built.
  10. Nothing borrowed. No line below the header names a file, table, library or step, and no line would be true of every change in the repository.

The note on item 5: D3’s instant has one criterion, AC4, on it; the near side is any link that was never replaced. On item 9, an open question that cannot change what is built may be listed as open, as the team post’s template allows; one that can must be decided. On item 10, the header’s list of paths classes the change and plans nothing.

Spec smells: what to look for in each section

A spec smell is a visible feature of the document that predicts a wrong build. This table lists nineteen, by the section of the worked spec where each appears. The grouping and wording are this post’s; three rows follow rules a practitioner states as “One behavior per criterion,” “Numbers, not adjectives” and “State the unhappy path explicitly” (Sharda, 2026).

What you see in the spec Why it fails The fix Section
A solution where the reason should be (“add an expiry column”) Nobody can judge whether the build serves the goal Say what changes for whom, and why now Purpose
Paragraphs on how the existing code works Ages first; the agent can read the code Cut to three sentences; the agent reads the code Purpose
Silence on something that already exists (pending rows, old clients) The agent decides and does not say Ask in the interview; record the answer Decisions
A decision with no decider or rejected alternative Later, nobody knows whether it still holds or who may change it Name who decided, what was rejected and why Decisions
An interview answer that never reached the spec The spec “feels rigorous” and is incomplete Check the spec against the list of answers Decisions
An adjective where a number belongs (“after a reasonable time”) The agent picks the number Number, unit, event, clock Decisions
Tables, columns or libraries named Constrains the plan and goes out of date first Move it to the plan Behavior
Only the successful path is described The agent invents the refusals State every answer the feature can give Behavior
“Must not change” is empty Adjacent code gets unrequested improvements List the behavior, interfaces and data that stay as they are Must not change
House rules (“run the linter”, “never edit old migrations”) True of every change; absent whenever there is no spec Move them to the standing instructions file Must not change
A line with no test that pins it Nothing notices when it changes Name the pinning test; write it first Must not change
“Out of scope” is empty Reasonable next steps get built List what could plausibly be wanted and is not Out of scope
An open question parked here that changes what gets built (“period per workspace: to be decided”) The agent answers it Decide it now, or state the fixed behavior for this change Out of scope
Two behaviors in one criterion (“expires and can be resent”) Half of it passes review Split it Acceptance criteria
A limit with an example on one side only, or none well past it An implementation that always refuses passes; so does one that refuses only at the limit One example on each side, as close as the unit allows, and one far past Acceptance criteria
An outcome only the code can see (“marked internally”) No test outside the code can decide it Name the response or the message a caller receives Acceptance criteria
A starting state that allows two explanations The test passes for the wrong reason Choose a state where only this behavior explains the outcome Acceptance criteria
A sentence that states more than its test asserts (“changes nothing”) The unasserted half can be wrong Assert every fact the sentence states Acceptance criteria
No test named Nothing fails when it stops being true Name the test and see it fail Acceptance criteria

Does the same shape hold for other features?

It held for the other backend features I walked through the interview, the sections, the checklist and the change rule; two are shown. This is a desk check: nothing here was built or run.

Feature Decision the interview has to surface Boundary pair A bug after release, and where the rule sends it
Rate limit on an endpoint, per customer Per account or per API key; fixed or sliding window; does a refused request count The request that reaches the limit and the one after it A customer with two keys gets double. If the spec said “per account,” the code is wrong: row 3
New rounding rule on an invoice Which invoices the rule applies to; per line or per total; who in finance decides Half a cent on each side; issued just before and just after the effective instant Totals differ from the ledger by a cent. The agreed rule was not what finance meant: row 4

Three strains showed. A change that touches every read path, such as soft delete, needs more criteria than one sitting can hold, which is a reason to split it. A criterion about time under load names a test whose verdict depends on the machine. And a feature with no exact answer per input, such as sorting incoming support mail into queues, cannot name the right queue for every message, so its criterion names the labeled examples, the measure and the threshold, which is the evidence you would also want when choosing between an LLM classifier and a fine-tuned classifier.

Does it work on existing code, and is it worth the time?

On existing code the “must not change” section carries most of the weight, and whether the practice pays is unmeasured.

Each line of that section needs a test that pins current behavior before anything changes. For a refactor, the book notes, “The old code, still present, serves as a second specification (one that provably ran).” The order of work for a coding agent in a legacy codebase is laid out in the legacy section of the team workflow post.

As for worth: I found no controlled study of spec driven development with AI agents against a comparison condition. The nearest thing I found is a registered report whose protocol has been peer reviewed and which has no results yet: participants will each solve one benchmark programming problem by working through specification, tests and function in sequence, all in the same workflow, so even its results will not compare spec-first work with anything (Rosa and colleagues, 2026).

What exists is testimony about the cost, and it points both ways. One forum commenter reports “2-3 hours writing a ‘spec’ focusing on acceptance criteria” and a tested feature by the end of the day. Another spent “many many hours” before seeing any code and found it “hadn’t built the right thing.”

Where does a spec cost more than it saves?

A spec costs more than it saves when the change is smaller than its interview, when nobody yet knows what is wanted, and when the document is approved without being read.

Small changes. A fix you can state in one sentence has its spec already. The reviewer quoted earlier, arguing for small iterative steps, is “very skeptical that lots of up-front spec design is a good idea, especially when it’s overly verbose” (Böckeler, 2025). The team post covers the waterfall charge and the changes that need no spec.

Unsettled requirements. If the product owner cannot answer the interview, the spec records guesses with a decider’s name attached.

What it does not solve. Outside the criteria you thought of, a green build is convincing and unchecked, which is the verification gap in miniature. And a spec in the repository cannot hold decisions that were made somewhere the repository cannot see.

The takeaway

Before you hand an agent the next ticket, write beside each thing you want the observation that would tell you it is wrong. Where that observation is a named test, the spec is the source of truth for that line, and the test tells you the day the code stops agreeing with it. Where there is none, you have a wish, and the agent will grant some version of it.

Chapter 22, “The Coding Workflow in Practice,” develops the spec, the tests and the version-control habits around them, and Chapter 21, “Coding Agents,” explains why tests make coding the domain where agents work best (both in the full book). The agents at work guide places this post beside its neighbors, or you can see the formats.

Questions readers ask

Is the spec the source of truth, or the code?
Each is the truth about a different thing. The running code is the only truth about what the system does. The spec is the truth about what the change was meant to do, for the slice it covers, and you can rely on it as a description of the code only while its acceptance tests pass. When the two disagree, a person decides which one is wrong; the tests tell you that they disagree.
How detailed should a spec for a coding agent be?
As detailed as its boundaries and no longer than one sitting. Every number gets a unit and the event it is measured from, every limit gets an example on each side and one well past it, and every refusal gets its answer. How to build it stays out: files, tables and libraries belong in the plan. The worked spec on this page is about 120 lines for one feature with a deadline and existing data; five lines is a legitimate spec for a smaller change.
Why did the agent build something else when it had the spec?
Three causes can be checked. The spec was silent on a decision, so the agent made it. The spec said it in prose that no test decides, so nothing failed. Or the spec was not in the context of the session that did the work. A benchmark paper first posted in 2025 also found that the models it tested did not ask clarifying questions unless explicitly prompted to.
How do you keep a spec from going stale?
Tie every acceptance criterion to a named test, so that a criterion that stops being true turns the build red, and change the criterion before the code when the wanted behavior changes. That protects the criteria only. The reasons behind the decisions have no test, so they need a named person who rereads them.
Is a product requirements document the same as a spec?
No. A product requirements document argues that something is worth building and for whom. A spec, in the sense used here, states what one change must do in a form a test can decide. The requirements document is input to the interview that produces the spec, and the people who wrote it are the ones who can answer most of the interview's questions.

Sources

  1. Birgitta Böckeler (2025). Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl
  2. Harper Reed (2025). My LLM codegen workflow atm
  3. Kunal Sharda (2026). I changed how I write acceptance criteria, and my AI agent stopped building the wrong thing
  4. Kunal Sharda (2026). Spec-driven development with AI is real now. The stale spec is the part nobody fixed
  5. Ziqian Zhong, Aditi Raghunathan and Nicholas Carlini (2025). ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases (preprint)
  6. Sanidhya Vijayvargiya, Xuhui Zhou, Akhila Yerukola, Maarten Sap and Graham Neubig (2025, v3 2026). Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
  7. Giovanni Rosa, David Moreno-Lumbreras, Gregorio Robles and Jesús M. González-Barahona (2026). Understanding Specification-Driven Code Generation with LLMs: An Empirical Study Design (Stage 1 registered report)
  8. Martin Fowler (2004). Specification By Example
  9. Forum comment (2026). Hacker News comment 47199116
  10. Forum comment (2026). Hacker News comment 49740011
  11. Forum comment (2026). Hacker News comment 48511237
  12. Forum comment (2026). Hacker News comment 46897138
  13. Forum comment (2025). Hacker News comment 45935925
  14. Forum comment (2025). Hacker News comment 45936358