Home / Tools / Context window budget planner

Free tool · runs in your browser · from Chapter 7

Context window budget planner

A free context window budget planner: see how full your agent's window is, when it crosses the 40% rule of thumb, and which of four fixes to apply.

The tool

Your inputs stay in this tab. Share a result by copying the page address: the state lives in the URL. The nine components, the 40% heuristic, the five failure modes, the four operations and the Chapter 2 sketch come from the book; the utilization arithmetic, the smallest-window and steps-to-ceiling formulas, the 15% flag threshold and the mapping from component or symptom to operation are added by the tool.

What does a context window budget planner do?

A context window budget planner adds up what an agent carries into each model call, divides it by the window size, and tells you how close the run sits to the point where quality starts to sag. Utilization is the fraction of the context window you actually fill. The planner compares it with an editable ceiling, near 40 percent by default, and suggests which operation would shrink the largest pieces.

The idea comes from Chapter 7 of the book, “Managing the Context Window,” which argues that context is scarce in a way raw size cannot repair. The chapter takes inventory of the desk the model works on: “the standing instructions; the user’s actual request; the conversation so far; whatever longer-term memory the system carried in from past sessions; the documents retrieved for this task; the definitions of the tools on offer; the results those tools have already returned; any worked examples you supplied; and the required shape of the output. Nine components, give or take.” The planner has one input for each.

You type your own window size. The tool never assumes one, because window sizes change with every model generation and the book refuses to treat any of them as permanent. What stays true across generations is the shape of the problem, and that shape is what the planner draws.

Why plan a budget below the hard limit?

You plan below the hard limit because a model’s ability to use a fact decays as the window around it fills, and the decay starts long before the window is full. The book puts it in two sentences: “The advertised window size is a fact about storage: how much fits. It promises nothing about attention.”

The evidence is a controlled study. Hong, Troynikov and Huber at Chroma (2025) tested 18 models, held the task fixed, varied only the length of the input, and found performance degrading as input grew, often on tasks the same models handled near-perfectly when the input was short. Practitioners adopted the study’s name for the effect: context rot. The article What is context rot? Symptoms and the four fixes walks through it in detail.

Somewhere below the hard limit sits a soft limit, and its location moves with the model and the task. Harder reasoning saturates sooner than simple lookup, which is why a model that finds a needle in an enormous haystack can still lose the thread of a long agent run. The book is plain about how to find yours: “The only trustworthy way to find the soft limit for your model and your task is to measure it. The spec sheet will not tell you.”

Where does the 40 percent line come from?

The 40 percent line is a practitioners’ rule of thumb, not a measured constant: it circulates among engineers who build coding agents and keep utilization low, and the book repeats it only as an illustration, so the planner treats it as a default you should replace with your own measurement. Chapter 7 reports that a heuristic circulates alongside the name dumb zone: keep utilization low, “with figures in the neighborhood of forty percent circulating as the ceiling.” The book then adds the instruction this tool follows: “Treat the number as an illustration; treat the practice as sound.”

That is why the ceiling in the planner is an input, pre-set to 40 and labeled illustrative. The source of the heuristic, Dex Horthy’s 2025 writeup on context engineering for coding agents, itself describes keeping utilization “in the 40%-60% range,” depending on the complexity of the problem. If your own measurements say your agent stays sharp at 55 percent on lookup tasks and sags at 30 percent on long refactors, set the line where your evidence puts it.

The dumb zone is what lies past the line. Its tell, in the book’s words: “the agent stops engaging with your corrections and starts reflexively agreeing with every one of them—‘you’re absolutely right’—while the work gets no better.” When you see that, the chapter’s advice is to prefer starting fresh over arguing. “The emptier the desk, the sharper the model.”

How do you read the planner’s result?

Read the result as three zones and one projection: under the ceiling means room to work, past the ceiling means expect degraded quality before the hard limit, overflow means the pile no longer fits, and the projection counts how many loop passes remain before the ceiling and the hard limit. The stacked bar shows each component’s share of the desk, a dashed line at the ceiling, and a solid frame at the hard limit. The verdict below it names the zone.

Zone Utilization What it means What to do
Under the ceiling below your ceiling Room to work, though hard reasoning can sag earlier Keep the two cheap moves in place
Past the ceiling from the ceiling up to 100% Expect the dumb zone before the hard limit Curate now: compact, trim, select
Overflow 100% or more An error, silent truncation, or a framework summary Find out which one your stack does

The overflow row deserves the most attention because it fails quietly. Chapter 2 lists what happens when the pile no longer fits: the request is rejected, the input is truncated by dropping the oldest turns, or a framework summarizes the history before sending it. “The outright error is the kind outcome. Silent truncation is the treacherous one.” A run that starts forgetting its instructions may simply have lost them off the front of the transcript.

The projection is plain arithmetic the tool adds. Every pass of an agent loop appends to the history, so the pile grows. Given a growth per step, the planner divides the remaining headroom by it: steps to the ceiling equal the ceiling’s token count minus the current pile, divided by the growth, rounded down. The same formula against the window gives steps to the hard limit.

Which operation fixes which part of the desk?

Every remedy is one of four operations: write context outside the window, select what comes in, compress what stays, and isolate work in a separate window. Chapter 7 borrows the grouping from Lance Martin’s 2025 essay and adds: “I have yet to meet a context remedy that is not some combination of these.” The glossary entry write, select, compress, isolate has the short version.

The four operations, arranged around the desk.
Figure 7.4 The four operations, arranged around the desk. Write moves material out of the window into durable storage; select brings in only what the current step needs; compress shrinks what stays; isolate gives a messy subtask its own window and accepts back a distilled summary. Every context-management tactic in Part III is one of these four, or a combination. Reuse this diagram

The planner badges any component that holds 15 percent or more of the desk, and it always badges tool results. The threshold and the assignment of operations to components are the tool’s; each cure is the chapter’s.

  • Retrieved documents → select or isolate. Keep identifiers on the desk and load content just in time, or send the reading to a subagent in quarantine that reports back a page of conclusions instead of the pile.
  • Conversation so far → write and compress. Keep a progress file the agent re-reads after any reset, and use compaction to distill the history to decisions, open problems and the current plan. Prune dead ends; hide errors once they are resolved.
  • Tool results → compress. Once a result has been acted on, clear the payload and keep a one-line fact plus a reference. The book calls tool payloads “the single biggest source of bloat.”
  • Tool definitions → select. Expose only the tools the task needs. The chapter’s measurement is almost comic: a small model failed a task with 46 tools in view and passed it with the 19 relevant ones, though all 46 fit in its window.
  • Memory carried in and standing instructions → select. Preload a small stable core and fetch the rest on demand.
  • Worked examples → compress. Chapter 2 notes that “two or three well-chosen ones capture most of the available gain.”

The principle under all six is the sentence the book borrows from Drew Breunig: “if you put something in the context the model has to pay attention to it.” Nothing on the desk is free, including what is merely irrelevant.

Worked example: Chapter 2’s pile in two windows

Chapter 2 sketches a pile “with made-up but realistic numbers”: standing instructions, 500 tokens; conversation so far, 8,000; two fetched documents, 40,000; the latest tool result, 3,000. That is 51,500 tokens before the model writes a word, and the planner’s preset loads exactly those four numbers.

With no window entered, the planner answers a different question: what is the smallest window that keeps this pile under a 40 percent ceiling? The answer is 51,500 divided by 0.40, or 128,750 tokens. Now try two hypothetical windows, chosen only to be round. At 100,000 tokens, utilization is 51.5 percent, already past the ceiling. At 250,000 tokens, it is 20.6 percent, comfortably under.

Add growth. Suppose each loop pass adds about one more tool result, 3,000 tokens. In the small window, the pile hits the hard limit in 16 steps. In the large window, it crosses the ceiling, at 100,000 tokens, in 16 steps too. The bigger window bought the run the same sixteen passes of good work; it moved the line, and the line still arrives.

The operations column tells the rest. Documents hold 78 percent of the desk and get select or isolate; the history holds 16 percent and gets write and compress; the tool result gets compress. Moving the two documents behind just-in-time retrieval would drop the pile to around 11,500 tokens, which fits under the ceiling of a window far smaller than either hypothetical. That is the lesson Chapter 7 draws in its own words: good context engineering means finding “the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome.”

What if the agent misbehaves under the ceiling?

If the agent misbehaves under the ceiling, the problem may be what is on the desk rather than how much: a confusing, poisoned or contradictory piece of context can derail a run at any utilization, and a fresh, minimal context tells you whether the context is to blame. The symptom check below the planner uses the book’s bedside chart to name one of five failure modes, each with its own cure.

A bedside chart for context failures.
Figure 7.3 A bedside chart for context failures. Each observable symptom points to one named mode, and naming the mode is most of the debugging. Four of the five are diseases of the desk and clear when the desk is cleared; the fifth—fighting the weights, set apart in accent—is the only one that survives a fresh, minimal context, which makes a clean-desk retry the cheapest diagnostic in the chapter. Reuse this diagram

An agent that repeats its own past actions is distracted. One that calls pointless or wrong tools is confused. One that fixates on a false fact of its own invention is poisoned. One that contradicts itself, or clings to a premature answer, suffers a clash. And one that stays stubborn on a clean desk is fighting its weights. None of these needs a full window. The 46 tools in the confusion example fit comfortably, a clash needs only two pieces that disagree, and one wrong “fact” in the transcript is enough for poisoning.

The cheapest test in the chapter separates them: reproduce the misbehavior in a fresh, minimal context. “Fighting the weights is the only failure of the five that survives a clean desk.” If the failure vanishes, the context was the cause; if it persists, curation will not fix it. Run the test more than once, since a single pass of a probabilistic model proves little. The article on AI agent failure modes covers the wider family of failures this one belongs to.

Where does this planner stop being useful?

The planner stops being useful where token counts stop being the problem: it measures how full the desk is, not whether a document is relevant, whether two instructions contradict each other or whether a summary dropped the constraint that mattered, so a nearly empty desk can still fail. A desk at 15 percent utilization can still fail if the 15 percent is wrong.

Over-curation is the opposite risk, and the chapter names it: prune too eagerly and the model lacks the fact it needed; compress too hard and the summary silently drops what mattered. “The trade is always recall against focus; there is no universal setting.” The advice that keeps this honest is the chapter’s own default: “Most tasks need only the two cheap moves—trim the tool results, write the plan to a file.” Heavier machinery should wait for an observed failure.

Two related tools fill in what this one leaves out. Every token on the desk is billed on every pass, and the agent cost per task estimator prices that re-reading. A long run on a crowded desk also compounds its own mistakes, which the compounding error calculator makes visible. The context engineering guide gathers the rest of the cluster.

The question to keep

One question carries the whole chapter, and the book borrows it from Breunig: “Is everything in this context earning its keep?” The planner gives you numbers to ask it with. The answer still comes from reading what is on the desk.

The full argument, with the evidence for each failure mode and the four operations at depth, is in Chapter 7, “Managing the Context Window” (in the full book). The desk itself, and what happens when the pile overflows, is in Chapter 2, free to read.

Questions readers ask

Why does the tool warn at 40% when my window is far from full?
Because quality sags long before the window is full. Practitioners report a soft limit well below the hard one, and a rule of thumb near 40% utilization circulates as the ceiling. The book offers that number as an illustration, not a constant, so the line is editable: measure your own soft limit and set it.
Which of the five failure modes am I seeing, and how do I tell?
Use the bedside chart. Repeating its own past actions is distraction; calling pointless or wrong tools is confusion; fixating on a false fact it invented is poisoning; contradicting itself is clash; staying stubborn on a clean, minimal desk is fighting the weights. Reproduce the failure in a fresh context to separate the first four from the fifth.
What are the two cheap moves the book says most tasks need?
Trim the tool results and write the plan to a file. Trimming clears a payload once it has been acted on and keeps a one-line fact plus a reference; writing the plan lets the agent reorient after compaction or a fresh start. Heavier machinery should wait for an observed failure.
Should I just use a model with a bigger context window?
A bigger window raises the hard limit, and the ceiling moves with it, but it does not make a crowded desk sharper. Measured performance degrades as input grows even when everything fits, so a bigger window buys headroom, not focus. Curation is still the work.
How do I count tokens for each component?
Most model providers return the input token count with every response, and most expose a tokenizer or a counting endpoint. Log one representative mid-run call, split the assembled input by component, and type the counts in. A rough split is enough for a budget.

Sources

  1. Kelly Hong, Anton Troynikov and Jeff Huber (Chroma) (2025). Context Rot: How Increasing Input Tokens Impacts LLM Performance
  2. Drew Breunig (2025). How Long Contexts Fail
  3. Drew Breunig (2025). How to Fix Your Context
  4. Dex Horthy (HumanLayer) (2025). Getting AI to Work in Complex Codebases (Advanced Context Engineering for Coding Agents)
  5. Anthropic (2025). Effective context engineering for AI agents
  6. Lance Martin (LangChain) (2025). Context Engineering for Agents