Home / Blog / Agents at work / A Coding Agent in a Legacy Codebase: Make It Ag…

Agents at work

A Coding Agent in a Legacy Codebase: Make It Agent-Legible First

A coding agent in a legacy codebase fails where a new hire struggles. Audit one service, fix it in a safe order and count before and after. Copy the audit.

By Enrique Gutiérrez · Published · 23 min read

A coding agent in a legacy codebase fails for the reason a new hire struggles there. What a change requires sits in people’s heads, in distant files and behind a slow test suite. Audit one service for those gaps, fix them in a safe order, and count wrong-file edits, failed runs and review minutes before and after.

The worked case in this post is a composite. I assembled a made-up service from the complaints engineers post about agents on old code and from the list in Chapter 21 of the book. No real team stands behind it, so it reports no result, and its before-and-after table is printed empty for your own numbers.

I looked for a study that measured an agent’s success rate, or a team’s review load, on one service before and after a legibility change. I found none. What has been published sits beside the case, each study with its number, its setup and its limits.

You leave with an audit of thirteen checks, six steps in an order I defend, a short conventions skeleton and a measurement sheet. The properties come from the book. The order, the audit and the composite are this post’s own.

Why does a coding agent in a legacy codebase fail where a new hire also struggles?

A coding agent fails on old code because it reads the codebase a few files at a time, with no memory of earlier sessions. Anything a change depends on has to be findable in the text. A new hire has that problem for a month, and the agent has it at the start of every session.

A coding agent is a model working in a loop: it searches and reads files, edits them, runs a check and reads the verdict. Two examples of the category as of 2026 are Claude Code and Aider. I name them as examples and recommend neither.

Chapter 21 of the book turns the usual question around. It says “the codebase is the agent’s working environment, and environments can be hospitable or hostile”.

The chapter then sends one small task into two codebases. In the first, a search for the domain word lands on a plainly named function with its permission check three lines above. In the second, the search returns a re-export file that hides the real module, the logic sits behind a factory configured at runtime, and the permission check lives in another package.

A strong engineer gets through the second one by holding a map in her head. The book’s next lines are the argument of this post: “The agent has no head to hold a map in. It reads six files, picks the plausible-looking wrong layer, and edits it—fluently.”

The coding-agent loop.
Figure 21.2 The coding-agent loop. Each pass understands the task, locates the code, makes an edit, and runs the tests—the ground-truth beat, in accent—then reads the failures and iterates, leaving the circle only when the suite is green. The verifier the domain supplies for free is what closes the cycle; strip out that one beat and the loop is a text generator with commit access. Reuse this diagram

The figure is Chapter 21’s loop. Legibility serves two of its beats: locating the code and running the check. The chapter says of the first that “when the locating fails, everything after it is an agent confidently editing the wrong thing”.

Engineers describe the same failure in their own words. One wrote in February 2026: “We have AI agents contribute in a large legacy codebase, and without proper guidance, the agents quickly get lost or reimplement existing functionality.” Another, on a team of eight, called the tools “very myopic”: “they prefer very local changes (adding new helper functions all over the place, even when such helpers already exist)”.

What does “agent-legible” mean in the book, and what does the book leave out?

Agent-legible code is the book’s name for code an agent can navigate and change safely. The book calls the term its own and credits the practice to others: “Agent-legible code (the coinage is mine; the practice is the field’s) is code that survives being understood a few files at a time.”

The glossary, which is free to read, compresses it to four items: “Code organized so an agent can navigate and safely change it: small files, explicit contracts, fast tests, conventions written down where a fresh context can find them.”

Chapter 21’s section itemizes more, and that fuller list is the one this post works from:

  • Findable names. One mapping from where a thing is declared to where it is imported, with suspicion of barrel files, wildcard re-exports and aliases.
  • Plain constructs. Descriptive names, composition, the query visible in the code.
  • Checks beside the action. A model adding a route copies what is in front of it, “because the pattern it imitates is the pattern it can see”. The chapter concludes that “code layout is quietly a security control”.
  • Explicit errors. Failures visible at the call site, with no catch-all handler around them.
  • Stable diffs. Formatters and one construct per line, so edits land where intended.
  • Tooling. The requirements “compress to three words: fast, truthful, uniform”.
  • Few boundaries. Each hop between processes or languages costs the model something.

Chapter 22 adds the written half: conventions go into a standing file, one line per lesson.

What does the book not give you?

The book gives a list of properties and stops there. It offers no order of steps, no success-rate or review-load number, and no legacy case study. Its section on legible code never states a file-size rule, although the glossary line says “small files”.

Chapter 22 says of itself: “Nearly everything here comes from practitioners’ first-hand reports rather than controlled studies”. Chapter 21 marks its point about boundaries as “observed pattern rather than measurement”.

The book also declines to claim that the two readers always agree. “Agent legibility is mostly software craft rediscovered, under a new reader with no ego and no patience.” The word doing the work is “mostly”, and a later section here covers the exceptions.

What does the published evidence show, and what has nobody measured?

The published evidence shows associations and partial effects: healthier code breaks less under model refactoring, and instruction files change cost more clearly than success. I found no study that measured success rate or review load for a coding agent in a legacy codebase before and after a legibility change. Each result below comes with its limits.

Does healthier code break less often when a model refactors it?

One study says it does, on small files. Borg, Hagatulah, Tornhill and Söderberg (2026, arXiv:2601.02200v1) asked models to refactor 5,000 Python files from competitive programming and counted how often the tests still passed. They report “significantly lower break rates on code in the Healthy CodeHealth range (CH ≥ 9), with corresponding risk reductions of 15-30%”.

The files are single-file contest solutions, far from a multi-file service. The result is an association: nobody cleaned the code and measured again. Two of the authors work at the company that sells the code-health metric.

The authors also warn that “no quick gains can be expected in the hairiest of old legacy parts”. That sentence argues for pinning behavior before any cleanup.

Does a conventions file raise the agent’s success?

Two studies of instruction files point in different directions on cost and agree that no success gain is shown. An instruction file is a document the agent reads at the start of each session. Two examples of the convention as of 2026 are AGENTS.md, an open format, and CLAUDE.md, one vendor’s equivalent.

Gloaguen and colleagues (2026, arXiv:2602.11988v3) tested such files on benchmark tasks, including 138 issues from 12 repositories with developer-written files. Their finding: “providing context files does not generally improve task success rates, while increasing inference cost by over 20% on average”. Developer-written files improved performance “by 2.4% on average (p=21%)”, which is short of statistical significance.

Their stated limit is the reader’s own case: “The current evaluation is focused on Python.” Open-source Python tooling is well represented in what models have already seen. A private service with house rules may gain more from a file, and nobody has measured that.

Lulla and colleagues (2026, arXiv:2601.20404v2) ran one coding agent with and without the file on 124 pull requests from 10 repositories. The file was “associated with a lower median runtime (Δ28.64%) and reduced output token consumption (Δ16.58%), while maintaining a comparable task completion behavior”. The pull requests were capped at 100 changed lines and 5 files, so the result covers small tasks. The study did not score whether the output was correct, which it calls “beyond the scope of this paper”, so it shows no gain in success.

The two disagree on cost and neither shows a success gain. The conclusion they support is narrow: a conventions file is cheap to try and unproven as a way to raise success. The first paper recommends that such files “only include instructions required for coding agents that are not already present in the README”.

What do studies on mature repositories add?

They name the problem and stop short of testing the fix. Becker, Rush, Barnes and Rein (2025, arXiv:2507.09089v2) ran a randomized trial with 16 developers on 246 tasks in mature projects. Their analysis lists “Implicit repository context (Limits AI performance, Raises developer performance)” among the factors behind the slowdown they measured. One developer said the tool “made some weird changes in other parts of the code that cost me time to find and remove”.

That study varied nothing about the repositories. It cannot say whether writing the implicit context down would have changed the result. Its headline figures sit in the study table of the post on an agentic coding workflow for teams.

On AI agent technical debt, He and colleagues (2025, arXiv:2511.04427v3) compared repositories that adopted one agentic editor with matched controls. Their abstract reports a “large, but transient increase” in velocity and “a substantial and persistent increase in static analysis warnings and code complexity”. These are open-source projects, and the study describes adoption, so it says nothing about a legibility fix.

What do companies report about themselves?

Company write-ups describe heavy investment in the check and report no before-and-after comparison. A December 2025 post from Spotify Engineering describes verifiers that return “only the most relevant error messages on failure and return a very short success message otherwise”. It also says some agents were “trying to solve problems that weren’t strictly in their prompt, like refactoring code or disabling flaky tests”.

Those are the company’s own words about its own system, with no comparison condition. A Google report on code migrations labels itself plainly (Nikolov and colleagues, 2025, arXiv:2501.06972v1): “It is not a research study, in the sense that we do not carry out comparisons against other approaches or evaluate research questions/hypotheses.” Both accounts concern large-scale automated code changes.

Is legacy code always the harder case?

Some readers report the opposite, and their reasoning is sound. One wrote in May 2026: “I have exactly the inverse findings on my end. The bigger and more legacy the codebase, the more accurate the patches become. The harness itself seems to be the most important part.”

The explanation given: “If I was greenfield, there would be nothing to query or constrain against.” Existing code and history give an agent something to query and to copy, provided it can find them. That is the commenter’s point about the harness, and it is the reason the audit below starts with the check.

How do you audit one service in an hour?

Pick one service and one module inside it, then answer thirteen questions by doing something and looking at the result. Each line is a check with a way to see the answer. The step in brackets is where a failed check sends you.

The audit is this post’s own, built from the book’s list. Running it needs no agent, although a read-only question session with one is a fair way to do items 8 to 10. The book says of such sessions: “Nothing is written, no risk is taken”.

  • 1. Pinned behavior. In the module the first tasks will touch, break one line on purpose and run the suite. A test fails. (Step 1)
  • 2. Guarded tests. Look at the last ten merged changes that touched a test file. A person read each test diff, under a written rule. (Step 1)
  • 3. Stable verdict. Run the check three times on an unchanged checkout. All three verdicts match. (Step 2)
  • 4. One verdict. Introduce a type error or a lint failure. No command the team treats as “passing” stays green. (Step 2)
  • 5. One way to run it. Ask three engineers how to run the tests for this module. They name the same command. (Step 2)
  • 6. Fast verdict. Change one line and time the check from edit to verdict. The wait is short enough to repeat many times an hour, and the first lines of output name what failed. (Step 3)
  • 7. Written conventions. Ask a reviewer for three rules they enforce in review that the code does not show. Each one is written where an agent’s session will read it. (Step 4)
  • 8. Search lands. Search for one domain word. You reach the code that does the work within a few file opens. (Step 5)
  • 9. Symbols resolve. Follow one symbol from a use to its declaration. No re-export file or alias sits in between. (Step 5)
  • 10. Checks beside the action. For one route or entry point, find its permission or validation check. It is in the same file as the action. (Step 5)
  • 11. Visible errors. Open three calls that can fail. The failure is handled or returned at the call site, with no catch-all around it. (Step 5)
  • 12. Quiet diffs. Add one item to a list or one argument to a call, then run the formatter. The diff shows only that change. (Step 5)
  • 13. Few boundaries. Trace one typical change. It stays inside one process and one language, with no implementation chosen from configuration at runtime. (Step 6)

Write down what you saw for each failed line.

Which step comes first for your service?

Start at the lowest-numbered step that has a failed audit item, and do the steps in number order from there. The rows and what each one changes are the book’s. The order is mine, and I lean on one line from Chapter 21: “What this chapter adds to that arithmetic is an ordering for your effort”. The chapter means verifier before agent, and I apply it to the repository.

Pick your starting condition to see its row. With several conditions, take the earliest row.

Step What you change What the agent gains Why it sits here Audit items Start here when
1. Pin and guard Tests that fix current behavior in the target module; a rule that a person reads every test diff A verdict that fails when behavior changes, and no escape through editing the tests Every later step changes code or tooling and needs a net under it 1, 2 no pinned tests
2. Make the verdict honest One command, one result; flaky tests fixed or quarantined by name; type and lint failures fail the check A red or green it can believe Speeding up a verdict nobody trusts buys nothing 3, 4, 5 verdict not trusted
3. Make the verdict fast A check scoped to the module, with terse output; the full run kept before merge Many attempts per hour inside the loop Speed multiplies whatever the verdict says, so honesty comes first 6 slow check
4. Write the conventions A short file, one line per lesson; any rule broken repeatedly becomes a lint or structural check The rules the code does not show The file’s first lines are the commands from steps 2 and 3, and its effect is unproven, so a baseline has to exist 7 conventions ignored
5. Make the code findable and local Remove re-exports and aliases on the paths the agent took; move checks beside actions; replace catch-alls; add a formatter Search lands on the code that does the work Code changes are the riskiest step and need steps 1 to 3 under them 8, 9, 10, 11, 12 agent gets lost
6. Reduce boundaries Collapse an indirection or a seam, with the old implementation alive beside the new one Fewer hops per change The largest change with the least evidence behind it, so it goes last 13 many boundaries

Two rules keep the table honest. A step with no failed audit item is skipped. Steps 5 and 6 are done only on the paths where counted runs show the agent struggling, which requires the baseline from the measurement section.

The composite case: what does the audit find in an old monolith?

In the composite, a coding agent in a legacy codebase meets an audit that fails most of its lines, and the team starts at step 1. To repeat the label: this service is invented. I built it from the patterns in the reader complaints quoted here and from the book’s list. Its findings are plausible and none of them is data.

The service is the invoicing module of an old monolith. The suite is slow and sometimes fails for no reason. The conventions live in review comments and in senior engineers’ heads.

The team runs the audit on that module. Search for “invoice” lands quickly, symbols resolve, and everything runs in one process, so lines 8, 9 and 13 pass. Line 12 fails: no formatter runs on this module, and one added argument rewraps the whole call.

Lines 1 to 7 fail. Breaking the rounding rule leaves the suite green, and nobody has a rule about test diffs. An unchanged checkout does not always return the same verdict. The application starts with type errors present, and engineers name different test commands.

Lines 10 and 11 fail too. Permission checks are registered in a central list far from the routes, and most calls sit inside catch-all handlers.

Before any change: the baseline

The team first picks a handful of past tasks from the module’s history and replays each from its starting commit. They record the counts described in the measurement section. I print none of them, because an invented count would read as a finding.

Step 1: pin and guard

The team has the agent write tests that fix the module’s current behavior, which is the book’s advice: “No suite? The agent’s first task is to write one”. They adopt the book’s guard as a written rule: “Treat any diff that touches a test file as an escalation”.

What goes wrong: one generated test pins a known rounding bug as correct behavior. The reviewer catches it only because the rule forces a read of every test diff. One reader describes the unguarded version in September 2026: “The agent keeps removing our tests to replace them with tests that are easier to pass”.

The new tests also lengthen a suite that was already slow, and step 3 has not happened yet. Pinning tests come from what the code does today; deriving tests from acceptance criteria is covered in the post on spec driven development with AI agents.

Step 2: make the verdict honest

The team settles on one command that runs the tests, the type check and the linter and reports one result. The book’s reason is that a toolchain that lets code run while type checks fail “can gaslight the agent”. The flaky tests are quarantined by name, each with an owner.

Two costs follow. The honest command is red, because years of type errors now count, and the team has to record the existing ones as a known baseline. Quarantine also removes coverage from the retry logic, which is where the flaky tests lived.

Step 3: make the verdict fast

The team adds a check scoped to the invoicing module that skips the application boot and prints only failures. The book gives the reason: the verifier “is consulted inside the loop”.

What goes wrong: a change passes the scoped check and breaks a caller in another module. The full run before merge catches it. A fast check is a less complete one, so the team keeps both and writes down which one the agent runs.

Step 4: write the conventions

The first attempt is the one I would expect to disappoint. The team asks the agent to generate an overview of the service and commits a long file. I print no replay result for it, here or at any other step.

What the published evidence predicts is no gain at a higher cost: Gloaguen and colleagues report that repository overviews in such files were not helpful. On that evidence the team deletes the overview and keeps only lines the code does not show, using the skeleton below.

Suppose one rule keeps being broken anyway. A reader who builds a structure linter, and so has an interest, describes the pattern: the file “is not deterministic or actually enforced in any way”. The book agrees in Chapter 18, where a conventions file is “an inferential guide” whose effect “is only as reliable as the model’s attention to it”. The team turns that rule into a lint check.

Step 5: make the code findable and local

The rule for this step is to change code only on paths where replayed runs record wrong-file edits. Within that rule, permission checks for the invoicing routes move beside the actions, catch-all handlers come off the calls on those paths, and a formatter settles the diffs.

What it costs: a senior engineer objects that the central permission list existed so no route could be forgotten. The team keeps a test that fails for any route with no check, which preserves the guarantee in a form the loop can run.

Step 6 is skipped, since line 13 passed. The composite ends there with no result. One reader reports the outcome it cannot claim, as an impression with a caveat: “Once the legacy codebase is “LLMified”, the coding agents seem to perform more predictably. YMMV here, as it’s hard to do large refactors without tests for correctness.” (January 2026).

A second service: tests fine, agent lost

A different service takes a different path through the same table. Picture a notification service, also invented, with a fast and trusted suite. Its code is heavily abstracted: barrel files, import aliases, a factory that picks the sender from configuration, and several processes behind one feature.

Lines 1 to 7 pass. Lines 8, 9 and 10 fail, since the first search returns re-export files and the validation sits in middleware in another package. Line 13 fails as well, while lines 11 and 12 pass.

The starting condition is “agent gets lost”, so the team begins at step 5 and reaches step 6 afterward. For step 6 the book’s device applies: keep the old implementation alive, because it “serves as a second specification (one that provably ran)”.

What goes in the conventions file?

The file holds what the agent cannot infer from the code, one line per lesson, and nothing else. The book’s glossary calls it the standing project-instructions file and gives the pruning rule: “for each line: would removing it cause mistakes? if no, cut”.

It grows by the compound step, which Chapter 22 describes as “before moving on, ask what the agent should have known at the start, and write it into the file”. The chapter’s examples set the grain: “The environment flag that cost twenty minutes, the build quirk, the convention the agent kept violating: one line each.”

The skeleton is mine and uses plain text only. The first five headings match steps in the table.

CONVENTIONS FOR: <service or module>      last pruned: <date>

THE CHECK (steps 2 and 3)
- One command for the full verdict: <command>
- One command for this module only: <command>
- Which of the two to run while working: <which>

TESTS (step 1)
- Any change to a test file stops for a person to read.
- Behavior that must not change: <one line each>

CONVENTIONS THE CODE DOES NOT SHOW (step 4, three to five lines)
- <rule> (enforced by: <check name>, or "review only")

WHERE THINGS LIVE (step 5)
- <domain word> -> <directory or file>
- Checks that must sit beside the action: <which>

SEAMS (step 6)
- A change here may cross: <boundaries allowed>
- Old implementation kept at: <path>, until <condition>

TRAPS ALREADY HIT (one line each, with a date)
- <date>: <what the agent should have known at the start>

PRUNING RULE
- For each line: would removing it cause mistakes? If no, cut.

A study of 2,303 such files in 1,925 repositories found that test procedures appear in 75.9% of them and security in 14.8% (Chatlatanagulchai and colleagues, 2025, arXiv:2511.12884v2, abstract). The “checks beside the action” line is there to cover that gap.

A conventions file steers by being read, which makes it a guide. A lint check reports after the fact, which makes it a sensor. The harness grid builder sorts a team’s controls into those cells and shows which rules have no sensor behind them.

How do you know whether it helped?

You know by counting the same things on the same tasks before and after each step, which is the only evidence available for a coding agent in a legacy codebase today. The book names the counts: “the agent’s struggle shows up as extra tool calls, wrong-file edits, and failed runs you can count”. The sheet below adds the two review measures, and its cells are empty on purpose.

Count How to collect it Before After
Runs that end with the check green, out of runs started Replay each task from its starting commit; the remainder are the failed runs
Tool calls per run Read them from the agent’s run log
Wrong-file edits per run Edits outside the files the original human change touched, or reverted within the run
Runs that touched a test file Search each run’s diff for test paths
Review minutes per merged change Reviewers note the time on live changes over a fixed window
Changed lines per pull request Read from version control over the same window

The first four rows come from replay: pick 5 to 10 real past tasks in the module, each with a known starting commit and a human change that solved it. Run each task several times, since one run is one sample.

Two rules keep the columns comparable. Steps 1 to 3 change the check itself, so score every replayed run, before and after, with one yardstick: the tests from the human change plus the full command from step 2. A past commit does not contain your fix, so apply each finished step on top of the starting commit before replaying. Where no check exists yet, the first row has no “before”.

The last two rows cannot be replayed, because a person’s review happens once. Collect them on live changes for a fixed window before the first step and again after the last.

Ten tasks will reveal a large change and cannot confirm a small one, and the eval sample size calculator shows how wide the uncertainty is. If a step moves nothing you can see, say so and stop investing in it.

Tool calls are also the count that survives a change of vendor. What a replayed run costs in money depends on how the tool bills, which is the subject of the post on AI agent pricing models.

Where do agent legibility and human legibility conflict?

They conflict where a person’s memory makes an abstraction cheap and the agent’s lack of one makes it expensive. The book says the two “genuinely conflict” at times and gives the example: “a little duplication can beat a beautiful indirection for a reader with no spatial memory”.

I see three recurring cases, and this list is my own reading:

  • Duplication against a shared helper. People prefer one helper in one place. An agent that cannot find it writes a second one.
  • Central registries against local checks. A central list makes omissions visible to a person auditing it. A local check is the one the agent copies.
  • Short names behind aliases against long qualified names. Aliases shorten lines for people who know the map. They hide the origin from a reader who searches.

The composite’s permission argument is the second case, and neither side was wrong. The book’s practical point is that such an argument can now be settled by counting. Replay the tasks under both layouts and keep the one that the counts and the reviewers can both live with.

For the rest, the book quotes Armin Ronacher, writing in 2025: “When an agent struggles, so does a human.”

What are the limits of this post?

This post rests on a composite, a list from the book and a handful of studies that each measure something adjacent. Four limits follow from that.

The composite proves nothing. It shows how the audit and the table fit together on one invented service. Whether the steps raise success on your service is an open question until you count.

The order is a judgment. I put verification first on the book’s general argument. A team with a trusted suite could reasonably write its conventions before anything else.

The evidence is narrow. The code-health result covers single contest files. The instruction-file studies cover open-source Python and small pull requests. None covers a private service with years of history, and ten replayed tasks will miss modest effects.

A dated snapshot. The studies are from 2025 and 2026, and the agents they tested will be replaced.

Software with no source to edit needs a different move, covered in the post on agent tools for software built for humans.

One module, one hour, one sheet

Before you change a prompt or switch tools, run the thirteen checks on one module and fill the “before” column. A coding agent in a legacy codebase is a reader with no memory, and the audit tells you what that reader cannot find. The sheet tells you whether your fix reached it.

The list behind the audit comes from Chapter 21, Coding Agents, and the conventions habit from Chapter 22, The Coding Workflow in Practice, both in the full book. The glossary entries linked above are free to read. The agents in practice guide collects the neighboring posts, or you can see the formats.

Questions readers ask

Why do coding agents fail on legacy code?
An agent starts every session knowing nothing about the codebase and builds its picture from text search and a handful of file reads. Old code often keeps what a change depends on in people’s heads, in distant files or behind a slow, unreliable test suite. The agent then edits a plausible wrong place and has no trustworthy verdict to correct it.
Does an instruction file make an agent follow our conventions?
It can help, and the published evidence does not show that it reliably does. Gloaguen and colleagues (2026) found that context files did not generally improve task success and raised inference cost by over 20% on average, in Python open-source repositories. Write only what the code does not show, and move any rule that keeps being broken into a check the agent’s loop runs.
Can a coding agent refactor legacy code safely?
Only as safely as the tests that pin the current behavior. Pin the module first, read every change to a test file yourself, and for a deep restructure keep the old implementation alive beside the new one. Borg and colleagues (2026), working on single contest files, found fewer breaks on healthier code and warn that no quick gains can be expected in the hairiest legacy parts.
How do I know whether making the code agent-legible helped?
Pick 5 to 10 real past tasks, replay each from its starting commit several times, and count runs that end with the check green, tool calls, wrong-file edits and runs that touch a test file. Repeat after each step. Track review minutes and changed lines per pull request on live changes over a fixed window.
What is agent-legible code?
It is the book’s name for code an agent can navigate and safely change a few files at a time. The free glossary defines it as code with small files, explicit contracts, fast tests and conventions written down where a fresh context can find them. Chapter 21 gives a longer list and says the practice belongs to the field.

Sources

  1. Markus Borg, Nadim Hagatulah, Adam Tornhill, Emma Söderberg (2026). Code for Machines, Not Just Humans: Quantifying AI-Friendliness with Code Health Metrics (arXiv 2601.02200, v1 read; accepted at FORGE 2026)
  2. Gloaguen, Mündler-Sasahara, Müller, Raychev, Vechev (2026). Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? (arXiv 2602.11988, v3 read)
  3. Lulla, Mohsenimofidi, Galster, Zhang, Baltes, Treude (2026). On the Impact of AGENTS.md Files on the Efficiency of AI Coding Agents (arXiv 2601.20404, v2 read)
  4. Becker, Rush, Barnes, Rein (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv 2507.09089, v2 read)
  5. Chatlatanagulchai et al. (2025). Agent READMEs: An Empirical Study of Context Files for Agentic Coding (arXiv 2511.12884, v2 abstract)
  6. He, Miller, Agarwal, Kästner, Vasilescu (2025). Speed at the Cost of Quality (arXiv 2511.04427, v3 abstract)
  7. Nikolov, Codecasa, Sjovall, Tabachnyk, Chandra, Taneja, Ziftci (2025). How is Google using AI for internal code migrations? (experience report, arXiv 2501.06972, v1 abstract)
  8. Spotify Engineering (2025). Background coding agents, part 3: feedback loops (company engineering blog, self-reported)
  9. hallh (Hacker News) (2026). Hacker News comment on agents getting lost in a large legacy codebase (15 February 2026)
  10. synthc (Hacker News) (2026). Hacker News comment on very local changes and duplicate helpers (15 March 2026)
  11. aleqs (Hacker News) (2026). Hacker News comment on an instruction file that is not enforced (6 August 2026)
  12. Yoric (Hacker News) (2026). Hacker News comment on an agent replacing tests in a legacy rewrite (13 September 2026)
  13. jumploops (Hacker News) (2026). Hacker News comment on agents after a legacy codebase is reworked (27 January 2026)
  14. bob1029 (Hacker News) (2026). Hacker News comment reporting more accurate patches on larger, older codebases (24 May 2026)