Home / Blog / Context engineering and memory / What Is Context Rot? Symptoms and the Four Fixes

Context engineering and memory

What Is Context Rot? Symptoms and the Four Fixes

What is context rot? Why an AI agent degrades as its context window fills, the five failure modes behind it, and four fixes. Learn to diagnose a run.

By Enrique Gutiérrez · Published · 11 min read

Context rot is the decay of a language model’s ability to use any given fact as the context window around that fact fills. The fact is still there, nothing in the window needs to be wrong, and the window may be far from full. The model simply uses what it has less well, which is why a bigger window does not fix a crowded one.

So what is context rot in practice? It is the reason your agent was sharp for the first twenty minutes and then started re-running tests it had already run, ignoring an instruction it followed an hour ago, and agreeing with every correction you made without getting any better. This post gives the symptoms, the five failure modes Chapter 7 of AI Agents, Engineered names, and the four operations that fix them. By the end you should be able to read a degraded run’s transcript and say which disease it has.

What is context rot, exactly?

Context rot is a slope, not a cliff: as input length grows, a model’s accuracy on a task it handles easily at short length drifts downward, unevenly and model by model. Chapter 7 defines it as “the tendency of a model’s ability to use any given fact to decay as the window around it fills.”

The best evidence for what is context rot comes from a 2025 technical report from Chroma by Kelly Hong, Anton Troynikov and Jeff Huber. Its design is what makes it useful: the authors held task complexity constant and varied only the length of the input, across 18 models and 194,480 model calls. Their headline finding: “Across all experiments, model performance consistently degrades with increasing input length.” (The phrase itself predates the report; a Hacker News commenter floated “context rot” in June 2025 to describe output quality falling off as a context fills with “distractions and dead ends.”)

Chapter 7 draws the practical conclusion in one line: “The decay is a slope rather than a cliff, it varies by model, and it begins long before the window is full.” That last clause is the one engineers miss. The advertised window is a fact about storage. It promises nothing about attention.

Why does a full context window hurt an agent?

A full context window hurts because the model’s attention is a fixed budget spread over everything on the desk, and an agent’s loop keeps adding to the pile on every step. Each tool result, each intermediate thought and each dead end is laid back in front of the model on the next call, relevant or not.

The book’s picture for the window is the desk: “the single work surface on which everything the model consults during a call must fit.” Chapter 2 (free to read) introduces it and adds a second, older finding: even what fits is used unevenly. Liu et al. (2023), in “Lost in the Middle”, found that accuracy “is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle.” In one striking case, a model given the answer-bearing document in the middle of its input scored lower than it did with no documents at all.

Chroma’s report adds two results that matter for agents. On a conversational memory benchmark, models did markedly better with a focused prompt of about 300 tokens than with the full prompt of about 113,000 tokens containing the same answer. And, against intuition, every one of the 18 models did better when the filler text was shuffled nonsense than when it was coherent prose. Chroma’s conclusion: “Whether relevant information is present in a model’s context is not all that matters; what matters more is how that information is presented.”

In the book’s plainer words: “Nothing in the pile needs to be false. Scale alone dilutes.” If you want the companion argument for why the window must be curated at all, the post on what context engineering is and why the desk is scarce makes it.

What does the dumb zone look like from outside?

The dumb zone is the practitioners’ name for the far side of a window’s effective limit, where the agent’s output has visibly degraded even though the run continues without errors. Its most recognizable tell is reflexive agreement: the agent answers each correction with “you’re absolutely right” while the work gets no better.

Dex Horthy of HumanLayer popularized the term in a 2025 AI Engineer talk. His written guide describes keeping “utilization in the 40%-60% range (depends on complexity of the problem),” and the talk places the start of the dumb zone around 40 percent of the window for the coding agent he used as an example. Treat that number with care. It is a rule of thumb from particular models and tasks, and its own author frames it as varying with both. The book says the same: “Treat the number as an illustration; treat the practice as sound.”

What transfers is the practice. Plan work so it never needs the whole desk, curate early rather than at the limit, and when an agent has slid into the dumb zone, start fresh instead of arguing with it. In the book’s words, “The emptier the desk, the sharper the model.”

What are the five ways context fails?

Agent context fails in five named ways: poisoning, distraction, confusion, clash, and fighting the weights. Chapter 7 calls telling them apart “most of the debugging,” because each has its own symptom and its own cure. The first four come from Drew Breunig’s essay “How Long Contexts Fail” (2025); the fifth is his later addition.

A bedside chart for context failures.
Figure 7.3 A bedside chart for context failures. Each observable symptom points to one named mode, and naming the mode is most of the debugging. Four of the five are diseases of the desk and clear when the desk is cleared; the fifth—fighting the weights, set apart in accent—is the only one that survives a fresh, minimal context, which makes a clean-desk retry the cheapest diagnostic in the chapter. Reuse this diagram

Where does context rot itself sit? Under distraction. The book is explicit: “Distraction, the study’s context rot, and the practitioners’ dumb zone are one decay under three names.” The other modes are related but different, and confusing them sends you to the wrong fix.

Symptom you see in the transcript Failure mode What is on the desk First fix (operation)
Repeats past actions, re-runs the same test, recombines its own log instead of planning Distraction (context rot, dumb zone) A history so long it outweighs training Compress: compact the history; write the plan to a file and restart
Calls wrong or pointless tools, uses a tool when none is needed Confusion Irrelevant material, often too many tool definitions Select: a smaller tool loadout for this task
Builds on a “fact” it invented earlier; corrections don’t stick Poisoning A false statement repeated on every pass Compress (prune the poisoned turn) or restart clean; don’t just append a correction
Contradicts itself, clings to an early answer against later evidence Clash Two versions of the truth with nothing marking which wins Compress: prune superseded instructions and stale guesses
Stays stubborn on one format or behavior, even in a fresh session Fighting the weights Nothing wrong with the desk; the request fights training Route around: adopt the model’s preferred format, or change models

Two of these deserve a closer look. Clash has the best measurement behind it: Laban et al. (2025) delivered the same task information piecemeal across conversation turns instead of in one prompt and found “an average drop of 39% across six generation tasks.” Their diagnosis is pure clash: “when LLMs take a wrong turn in a conversation, they get lost and do not recover.” Poisoning is the hardest to clear, because, as the book puts it, “Appending a correction later does nothing to remove the poison; it merely places truth and falsehood on the same desk and asks the model to referee.” Poisoning is also how one early error compounds through a run, which the post on why agent errors compound and the compounding error calculator quantify.

How do you tell context rot from a model limit?

You tell them apart by reproducing the failure on a clean desk: a new session with only the one instruction and the one input. If the failure vanishes, the old context caused it and one of the first four modes is at work. If it persists, the context was never the problem.

Chapter 7 calls this “the cheapest diagnostic in this book,” and the reason it works is structural: “Fighting the weights is the only failure of the five that survives a clean desk.” Two cautions keep the test honest. Run it several times, because model output is sampled and one pass proves little. And strip the desk fully: a contradiction baked into your standing instructions or tool definitions rides into every fresh session and still counts as context.

This test also answers the most common complaint I hear from engineers new to agents, that the agent “ignores instructions.” Often the instruction is present but buried mid-pile, contradicted by a later one, or outweighed by a long history. The post on why AI agents ignore instructions works through that case in detail.

What are the four fixes for context rot?

The four fixes are the four operations of context engineering: write context outside the window, select only what the current step needs, compress what stays, and isolate messy subtasks in their own window. Lance Martin’s 2025 essay for LangChain introduced this grouping, and the book adopts it because every remedy it has met reduces to some combination of the four.

The four operations, arranged around the desk.
Figure 7.4 The four operations, arranged around the desk. Write moves material out of the window into durable storage; select brings in only what the current step needs; compress shrinks what stays; isolate gives a messy subtask its own window and accepts back a distilled summary. Every context-management tactic in Part III is one of these four, or a combination. Reuse this diagram

The book maps them onto a real desk. Write is moving reference material into a filing cabinet with a sticky note saying where it is; select is pulling out the one folder today’s work needs. Compress is replacing a stack of interview notes with a one-page summary; isolate is handing a gnarly sub-problem to a colleague at another desk and asking for a memo of conclusions.

The full treatment, with costs, is in the deep dive on write, select, compress, isolate.

Operation What it does Common forms Cures Cost
Write Saves information outside the window so it survives resets Progress file, scratchpad, memory store Lost plans after compaction or restart Extra read/write steps; a store to maintain
Select Brings in only what this step needs Retrieval (RAG), small tool loadout, just-in-time file reads Confusion Latency; a poorly guided agent may wander
Compress Keeps the information, sheds the tokens Compaction, pruning, tool-result trimming Distraction, clash, poisoning Lossy; dropped details surface later
Isolate Splits work across windows Subagent in quarantine returning a short summary Distraction in the main thread Token spend multiplies; coordination is on you

For context compaction in AI agents, the book’s rule is blunt: compression is lossy by design, so “files survive what summaries forget.” Compact with recall in mind first, and pair it with a written plan.

Worked example: diagnosing a run that degraded

Here is a hypothetical but representative run, drawn so you can practice the vocabulary. Try to diagnose each symptom before reading my answer.

An agent is asked to migrate a service from one logging library to another. The standing instructions say: “Do not modify files under vendor/.” It has 31 tools available, including three search tools with overlapping descriptions. For about forty steps it works well. Then, scrolling the transcript, you find:

  1. Step 44. The agent runs the full test suite. It ran the same suite, with the same failure, at steps 38 and 41.

  2. Step 47. It calls a web search tool to find the new library’s configuration format, which a tool result at step 12 already printed in full.

  3. Step 52. It writes “since the old library is also used by the billing module, we must keep both.” Nothing earlier says this; you trace it to a misread grep result at step 19, after which the agent has repeated it four times.

  4. Step 55. It edits a file under vendor/. You correct it. It replies “you’re absolutely right,” reverts, and two steps later edits another file under vendor/.

Step 44 is distraction: the agent is replaying its own log instead of planning. Step 47 is confusion: with 31 tools and overlapping descriptions, it reached for one it did not need. Step 52 is poisoning: a false claim entered at step 19 and now rides along on every pass. Step 55 is the dumb zone’s signature, reflexive agreement, and also a likely clash or buried-instruction problem: the vendor/ rule sits at the top of a long window, far from the end and buried under forty steps of history that now outweigh it.

Now run the clean-desk test on step 55: a fresh session, the standing instructions, and one file to migrate. In this illustration the agent leaves vendor/ alone, so the context was the cause, not the model.

The repair uses all four operations. Write: have the agent record done, remaining and decisions in a progress file.

Compress: compact the history, explicitly pruning the step-19 claim and trimming old test logs to one line each. Select: give the migration a loadout of five or six tools instead of 31, and restate the vendor/ rule near the end of the prompt. Isolate: send the question “where else is the old library used?” to a subagent that returns a file list, not three hundred lines of grep output. Then restart on the clean desk.

When does this advice not apply?

This advice weakens when the task is genuinely short, when the window is mostly retrieval rather than reasoning, or when curation itself starts dropping facts the model needs. Over-curation is a real failure: prune too eagerly and the model lacks the fact it needed.

Chapter 7 frames the trade honestly: “The trade is always recall against focus,” and there is no universal setting, only measurement.

Most tasks need only the two cheap moves, trimming tool results and writing the plan to a file. Every number in this post (40 percent, 113,000 tokens, 39 percent) is a measurement from one study or one practitioner’s setup. New models move the thresholds; the shapes have held so far. To manage agent context windows well, find your own soft limit by measuring your own task.

The one question to keep

Whatever else you take from this answer to what is context rot, keep the question Chapter 7 closes on, borrowed from Breunig: “Is everything in this context earning its keep?” Ask it of every tool definition, every stale payload and every superseded plan, and the four verbs are how you act on the answer.

The full argument, with the failure-mode evidence and the four operations at depth, is in Chapter 7, “Managing the Context Window” (in the full book). The desk itself and the lost-in-the-middle effect are in Chapter 2, free to read. For the rest of this cluster, start at the context engineering pillar, or see the formats.

Questions readers ask

What is context rot in simple terms?
Context rot is the tendency of a language model to use any given fact less well as the amount of text around that fact grows. Nothing has to be wrong in the window. More material dilutes the model's attention, so quality slopes downward well before the window is full.
Does a bigger context window prevent context rot?
No. A bigger window raises how much fits, not how well the model attends to it. Controlled studies find degradation as input length grows even far below the advertised limit, so a larger window mostly gives you more room to rot in.
What is the dumb zone?
The dumb zone is a practitioner's name for the part of a run where the window is full enough that output quality has visibly dropped. A common tell is reflexive agreement with every correction while the work stops improving. The often-quoted 40 percent utilization ceiling is a rule of thumb that varies by model and task, not a constant.
How do I know whether my agent has context rot or the model just can't do the task?
Reproduce the failure in a fresh, minimal context: a new session with only the one instruction and the one input, run several times. If the failure disappears, the old context caused it. If it persists, the problem is the model or the request, and trimming context will not help.
Is compaction enough to fix context rot?
Compaction helps but it is lossy: the summary can silently drop the one constraint that mattered later. Pair it with writing the plan and key decisions to a file the agent re-reads, and trim tool results as soon as they have been used, so there is less to compact.

Sources

  1. Workaccount2 (Hacker News) (2025). Hacker News comment coining "context rot"
  2. Kelly Hong, Anton Troynikov, Jeff Huber (Chroma) (2025). Context Rot: How Increasing Input Tokens Impacts LLM Performance
  3. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang (2023). Lost in the Middle: How Language Models Use Long Contexts
  4. Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer Neville (2025). LLMs Get Lost In Multi-Turn Conversation
  5. Drew Breunig (2025). How Long Contexts Fail
  6. Lance Martin (LangChain) (2025). Context Engineering (write, select, compress, isolate)
  7. Anthropic (2025). Effective context engineering for AI agents
  8. Dex Horthy (HumanLayer) (2025). Advanced Context Engineering for Coding Agents
  9. Dex Horthy, AI Engineer (2025). No Vibes Allowed: Solving Hard Problems in Complex Codebases (talk page)