Home / Blog / Context engineering and memory / Write Select Compress Isolate Context Engineeri…

Context engineering and memory

Write Select Compress Isolate Context Engineering, Measured

Write select compress isolate context engineering, one operation at a time: how each works, how it fails, what it costs and how to measure it. Start here.

By Enrique Gutiérrez · Published · 16 min read

Write select compress isolate context engineering sorts every context technique into four operations: write state outside the window, select what comes back in, compress what stays, and isolate separable work in its own window. Each operation has its own failure mode and its own cost, so none should ship without a measurement that shows it helped.

I wrote this for the engineer who owns the traces and the evals, and who has just been asked to “add compaction” or “use subagents” because a long run got slow, expensive or confused. The post on context rot and its four fixes diagnoses the run and names the failure modes. This one takes each operation apart: how it works, how it fails, what it costs, and which number tells you it worked. It ends with an order to apply them in and a worked budget.

What are the four operations, and where did the grouping come from?

The four operations are a taxonomy of context engineering moves: write saves information outside the window, select pulls in only what the current step needs, compress keeps information while shedding tokens, and isolate splits work across windows. Lance Martin introduced the grouping in a June 2025 post: “I group approaches into 4 buckets,” he wrote, and named them write, select, compress and isolate. LangChain republished it on its blog the following month.

Chapter 7 of AI Agents, Engineered (in the full book) adopts it, and notes a parallel six-tactic inventory from Drew Breunig’s “How to Fix Your Context”: retrieval, tool loadout, quarantine, pruning, summarization and offloading, “the same moves at a finer grain.” The book’s claim for the taxonomy is strong: “I have yet to meet a context remedy that is not some combination of these.”

The LangChain version adds the two things to build before any of it, and they are this post’s spine: “First, ensure that you have a way to look at your data and track token-usage across your agent.” And: “be sure you have a simple way to test whether context engineering hurts or improve agent performance.” Each section below therefore ends with a metric.

The glossary entry has the one-line definitions. If you need the case for curating the window at all, the post on what context engineering is makes it.

What should you record before you change anything?

Record a baseline: the assembled input of every call split by component, plus success, input tokens, turns and cost for each task on a fixed set of real tasks. Without it, a change that quietly hurts success looks the same as one that saved the run.

Chapter 9, Memory: Working State Across Long Runs (in the full book) puts the risk plainly: “a compression that quietly hurts task success looks identical, from the inside, to one that saved the run.” The component split is what a trace gives you if each call’s input is logged whole, and it usually tells you which operation to try first, because the largest component is where tokens go.

The task set does not need to be large to start. Anthropic’s team, describing its multi-agent research system, reported starting “with a set of about 20 queries representing real usage patterns,” because early changes had large effects. Small sets detect only large differences, though; the eval sample size calculator shows how many tasks a smaller difference needs. Model output is sampled, so run each configuration more than once.

  1. Every model call’s full input is logged, and I can split it into instructions, tool definitions, history, tool results, retrieved documents and memory.
  2. I have a fixed eval set of real tasks, each with a pass or fail check that does not depend on my reading the transcript.
  3. For each task I record success, total input tokens across every thread, number of turns, wall-clock time and cost.
  4. If my provider reports how much input was served from cache, I record that share per task.
  5. For each task I have listed the facts and constraints that must survive a reset or a compaction.
  6. Each configuration runs on the whole set more than once, and I change one operation at a time.
  7. I keep the failing traces, not just the scores, so I can see which operation dropped what.

How does each technique work, fail, cost and get measured?

Each technique below sits under one operation and has a mechanism, a way it fails, a cost and one metric that shows whether it helped. Pick the symptom you see in your traces to narrow the table. The symptom names are this post’s own plain-language labels; the context rot post maps symptoms to the named failure modes. The cost and metric columns are this post’s judgment unless a source is named in the sections that follow.

Technique Operation Mechanism How it fails Cost Metric that shows it helped Symptom
Progress file or scratchpad Write The agent records plan, progress and decisions in a file and re-reads it after any reset Notes go stale, or are never re-read A few read and write steps per run Success after a forced reset; calls spent re-finding what the notes held lost plan, repeats work
Memory store across sessions Write Facts and procedures persist between runs and are loaded later A wrong memory is persisted and returns every session A store to maintain, plus a selection problem to load from it Success on tasks that need a past fact; rate of wrong memories found in audit lost plan, false fact
Tool loadout Select Only the tools this task needs are defined in the call The needed tool is left out; changing the set mid-run breaks the cache Selection logic; a less capable agent if too narrow Tool-choice accuracy; share of tasks whose needed tools were in the loadout wrong tool, bloat
Retrieval (RAG) Select Relevant passages of a corpus are fetched into the window at question time The retriever misses, and nothing announces the miss Index, latency per query Recall of the needed passage in what was retrieved; success missing fact, bloat
Just-in-time references Select The window holds paths, IDs and queries; content is loaded when needed The agent wanders, or never loads what it needed Latency; extra tool calls Tool calls per task; success bloat, missing fact
Tool-result trimming (observation masking) Compress A result already acted on is replaced by a one-line trace and its reference The agent needed the raw result again and cannot re-fetch it Close to none if the reference is kept Input tokens per task; success unchanged bloat, repeats work
Compaction Compress History is summarized and a fresh window starts from the summary The summary drops a constraint; the summarizer reads the rotted context One summarization call each time; recall loss Must-survive facts present after compaction; success; turns per task bloat, repeats work
Pruning Compress Named stale items are deleted, the rest kept verbatim The rule for “stale” deletes something that mattered A rule per category Contradictions per run; success contradiction, false fact
Start fresh from notes Compress and write Abandon the window; reopen seeded with clean notes Notes were incomplete, so the restart loses state Re-reading the notes; any lost work Success after restart versus continuing false fact, contradiction, repeats work
Subagent in quarantine Isolate A worker with a fresh window and small toolset explores and returns a short result The brief is too thin; sibling workers make conflicting choices; a poisoned source poisons the summary Each worker’s own reading; coordination Total tokens per task across threads; main-thread peak input; success bloat, untrusted input
Sandbox or state fields Isolate Large objects live in an execution environment or a state field the model does not read The model needs a detail it cannot see Infrastructure Input tokens per task; success bloat

How does write work, and how does it fail?

Write moves state out of the window into a file or store so it survives compaction, truncation and resets; the window carries a pointer or a short version. Its failure is quiet: notes that go stale, notes the agent never re-reads, and, for stores that persist across sessions, a wrong memory that returns every time.

The book’s case for write is survival: “The transcript is mortal: it will be compacted, truncated, or reset, and every plan that lived only in the transcript dies with it.” Practitioners report the failure in those words. One Hacker News commenter described an agent that read its plan.md at the start of a session and lost it in compaction: “By step 4 it needs to compact” (October 2026). A file that exists but is never reloaded is not write; it is a log.

A team that builds an agent product, Manus, describes a variant that serves attention as well as survival: the agent rewrites a todo.md as it works, “reciting its objectives into the end of the context,” which pushes the plan into what the post calls “the model’s recent attention span.”

Cost. Chapter 7: “The costs are modest: extra read and write steps, and the obligation to manage the store.” The store becomes its own selection problem once it grows past what fits, which is where agent memory architecture begins.

How to know it helped. Force the reset. At a fixed step in each test task, discard the window, restart with only the standing instructions and the notes, and compare success with an uninterrupted run. Count calls that fetch something the notes already held; that number should fall toward zero.

How does select work, and how does it fail?

Select decides which slice of everything available earns a place in this call: passages from a corpus through retrieval, a small tool loadout, or references loaded just in time. Its failure is the silent miss: the needed passage, tool or file never arrives, and the model proceeds without it.

The best measurement concerns tools. Gan and Sun’s RAG-MCP paper (2025) retrieved tool descriptions before each query instead of listing them all, and report that it “more than triples tool selection accuracy (43.13% vs 13.62% baseline)” while cutting prompt tokens “by over 50%.” That is the cure for the confusion mode, where overlapping tool descriptions lead the model to call the wrong one.

The Manus team argues the other side from production. Their rule: “avoid dynamically adding or removing tools mid-iteration,” because tool definitions sit near the front of the input, so a change invalidates the cache for everything after it, and earlier steps that refer to a removed tool confuse the model. My reading of the two is that both hold: choose the loadout once per task or phase, at the start of a fresh window, not on every step.

Cost. Chapter 7: “The cost is latency, plus the risk that a poorly guided agent wanders.” Changing the prefix also costs cache reuse, which the prompt caching entry explains.

How to know it helped. Label, for each test task, which tools and documents it needs. Measure the share of tasks where all of them were selected (selection recall), tool-choice accuracy, and success. If tokens fall and success falls with them, the selector is missing things.

How does compress work, and how does it fail?

Compress keeps the information a run needs while shedding tokens, at three grades: trimming old tool results, compacting the whole history into a summary, and pruning named stale items. Its failure is loss without an error: the dropped detail is the one that matters later, and nothing says it is gone.

Chapter 7 states it exactly: “compression is lossy by design, and lossy in a treacherous way: the detail you dropped is precisely the one whose importance surfaces later, and no error message announces the loss.” Chapter 9 adds the trap in automatic compaction: “near the limit, the summarizing is necessarily done by the same model reading the same rotted context that made compaction necessary.” Its remedy is to summarize deliberately, at milestones, while the context is still healthy.

The cheapest grade may be enough. Lindenbauer and colleagues (2025) compared strategies for a coding agent and found “a simple environment observation masking strategy halves cost relative to the raw agent while matching, and sometimes slightly exceeding, the solve rate of LLM summarization.” They also found summaries lengthened runs, suggesting summaries “act as a reinforcing signal, encouraging the agent to keep going.”

Pruning has its own risk. In Chapter 9’s words, “pruning needs a rule for ‘stale,’ and a wrong rule deletes the thing that mattered.”

Cost. One summarization call per compaction, and recall loss you cannot see without a probe. Anthropic’s guidance on compaction prompts is to “Start by maximizing recall,” then improve precision. The craft of the prompt is the subject of the post on context compaction for AI agents.

How to know it helped. Check each task’s must-survive list against the summary, then compare success, turns and cost per task with and without compression.

Should resolved errors stay in the context?

Practitioners disagree, and the answer depends on whether the error can still teach. Chapter 7 relays the “12-Factor Agents” advice to hide errors “once they are resolved”; the Manus team argues to “leave the wrong turns in the context,” because “Erasing failure removes evidence,” and a visible failure steers the model away from repeating it.

My reconciliation, which is this post’s own: keep a failure while it is unresolved or recurring, and prune it once it is fixed and has not come back. Measure it as the rate of repeated identical failures per run. If pruning errors raises that rate, the errors were still doing work.

How does isolate work, and how does it fail?

Isolate gives separable work its own window: a subagent explores with a fresh context and a small toolset, and only a distilled result returns. Its failures sit at the boundary: a brief too thin for the worker to do the right task, sibling workers making choices that conflict, and a summary that carries a poisoned source’s influence back.

Anthropic’s context engineering essay describes workers that use “tens of thousands of tokens or more” and return “a condensed, distilled summary of its work (often 1,000-2,000 tokens).” Walden Yan of Cognition argues the opposite default: “Share context, and share full agent traces, not just individual messages,” because “Actions carry implicit decisions, and conflicting decisions carry bad results.” Chapter 9 draws the line between them: “Quarantine is for work that separates.”

Chapter 9, Memory: Working State Across Long Runs (in the full book) also gives the second reason to isolate, risk. Reading untrusted pages or documents in a separate worker with no privileges is quarantine; it “reduces exposure; it does not neutralize it.”

A main window and an isolated worker window are separated by a wall with two narrow slots: a brief goes in through one, a distilled memo comes out through the other, and the worker's exploratory mess stays behind the wall.
Figure 9.5 The quarantine rule made architectural: a sealed workroom whose exploratory mess cannot cross the wall, joined to the calm main office by exactly two deliberately narrow slots—a brief going in, a distilled memo coming out. Reuse this diagram

Cost. Chapter 7 calls isolation’s costs “the steepest of the four.” Anthropic reported that “agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats.” What the worker is told is the subject of the post on the orchestrator worker pattern.

How to know it helped. Sum tokens per task across every thread, record the main thread’s peak input, and compare success. Count rework at merge, where one worker’s output had to be redone to fit another’s.

How do you apply write select compress isolate context engineering in order?

Apply them cheapest and safest first, and stop when the metric stops moving: trim tool results, write the plan to a file, narrow selection, compact or prune deliberately, start fresh from notes when the context is poisoned, and isolate only work that separates. Untrusted content is the exception: quarantine it from the start.

The order follows Chapter 9’s treatment protocol, which runs trimming, offloading, just-in-time loading, compaction, pruning and a fresh start, “cheapest and safest first,” then isolation. Chapter 7 sets the floor: “Most tasks need only the two cheap moves—trim the tool results, write the plan to a file—and no paging memory system should exist before an observed failure asks for it.”

  1. Baseline. Run the checklist above and note the largest component.
  2. Trim tool results (compress). Keep a one-line trace and the reference.
  3. Write the plan and progress to a file the agent re-reads (write).
  4. Narrow selection: a loadout chosen at task start, references instead of payloads (select).
  5. Compact or prune deliberately, at milestones, with a must-survive list (compress).
  6. Start fresh from notes when a false fact or contradiction persists after pruning.
  7. Isolate a subtask that separates, and any untrusted reading from step 1.

Stop when success on the task set stops improving or begins to fall. Chapter 7’s warning is that over-curation is real: “The trade is always recall against focus; there is no universal setting, and the only honest arbiter is measurement.”

Two cases not in this post show how the order applies. A support agent with dozens of overlapping tools picks the wrong one: the symptom is “wrong tool,” the table points to a loadout, and the metric is tool-choice accuracy at step 4. A research agent that reads arbitrary web pages has an “untrusted input” symptom: quarantine applies from the start, whatever its token count.

Worked example: what do trimming, a loadout and a subagent change?

In an illustrative 40-call run, trimming old tool results, narrowing the loadout and adding a progress file cut the tokens the run reads to 48.2% of the unmanaged baseline, and moving one separable 15-call exploration into a subagent cut the total further, to 69.5% of the trimmed run. Every number here is illustrative, computed from the stated assumptions.

The assumptions: standing instructions of 2,000 tokens; 36 tool definitions of 250 tokens each (9,000); each step adds 600 tokens of the model’s own text and a 1,400-token tool result. Unmanaged, the input on call k is 11,000 + 2,000 × (k − 1), so the 40th call reads 89,000 tokens, and the run reads 2,000,000 in all. On that last call, tool results are 54,600 tokens (61.3%), the model’s own history 23,400 (26.3%), definitions 9,000 (10.1%) and instructions 2,000 (2.2%).

Configuration Last call input Whole run reads Share of baseline run
A. Unmanaged 89,000 2,000,000 100%
B. Trim: last 5 results kept, older ones as 30-token traces 42,420 1,184,850 59.2%
C. B plus a 12-tool loadout and a 500-token progress file 36,920 964,850 48.2%
D. C with calls 21–35 sent to a subagent 28,970 (main) 670,550 (all threads) 33.5%

In D, the subagent starts with a 1,500-token brief and four tools, trims the same way, and returns a 1,500-token memo; it reads 185,850 tokens and the main thread 484,700. Isolation lowered the total here because one sequential exploration left the main desk entirely. It raises the total when many workers fan out in parallel, each paying for its own brief and history, which is the setting of Anthropic’s multi-agent research system: its 15× figure is relative to a chat interaction, against about 4× for a single agent. Count tokens across threads to see which regime you are in.

Note what the table cannot show. C adds 500 tokens per call for the progress file and buys survival, not savings; its metric is the reset test. None of the configurations is better until success on the task set says so. To lay out your own components against a window, the context window budget planner does the arithmetic.

Where does this advice stop?

The order and the metrics in this guide to write select compress isolate context engineering fit long, tool-heavy runs. A short task with a small window needs none of the machinery, and Chapter 7’s “two cheap moves” are often the whole answer. The numbers from outside sources come from particular agents, benchmarks and dates: the masking result from coding tasks in 2025, the token multipliers from one vendor’s research system. Treat them as directions, and measure on your own tasks.

The table’s costs and metrics are this post’s judgment, and two disagreements remain open: whether errors should stay visible, and whether multi-agent designs pay for their coordination. When an agent “forgets” a rule that is still in the window, the cause may be ordering rather than volume; the post on why AI agents ignore instructions separates those causes.

The question to keep

Every operation in write select compress isolate context engineering throws information away or defers fetching it, so every one needs a number that shows it did not throw away the wrong thing. So the question to ask of each change is not whether it saved tokens but whether success held while it did. The metric is how you find out.

The four operations at depth are in Chapter 7, “Managing the Context Window” (in the full book), and the treatment protocol for a long run and the quarantine rule are in Chapter 9, Memory: Working State Across Long Runs (in the full book). The context engineering guide maps the rest of this cluster, or you can see the formats.

Questions readers ask

What are write, select, compress and isolate in context engineering?
They are the four operations every context-management technique reduces to. Write saves state outside the window, as a progress file or memory store. Select brings in only what the current step needs, through retrieval, a small tool loadout or references loaded on demand. Compress sheds tokens from what stays, through trimming, compaction or pruning. Isolate gives separable work its own window, usually a subagent that returns a short summary.
Which context engineering operation should I try first?
Trim tool results first, then write the plan and progress to a file the agent re-reads. They are the cheapest and safest moves, and tool results are usually the largest component of a long run's input. Narrow selection next, then compact or prune, and isolate last, except for untrusted content, which belongs in quarantine from the start.
Is LLM summarization better than dropping old tool outputs?
Not necessarily. A 2025 study of a coding agent found that replacing old tool outputs with a placeholder halved cost relative to the unmanaged agent and matched the solve rate of LLM summarization, and that summaries made runs longer. Measure both on your own tasks before adding a summarizer.
Do subagents save tokens?
Sometimes. A subagent keeps its exploration off the main desk, which lowers the main thread's input, but each worker pays for its own brief, tools and history. One vendor reported multi-agent research runs using about 15 times the tokens of a chat interaction, against about 4 times for a single agent. Count tokens per task across all threads and compare success, not just the main thread's size.
How do I know whether compaction lost something important?
Write down, for each test task, the facts and constraints that must survive, then check each one against the summary or ask the resumed agent for it. Compare success, turns and cost per task with and without the compaction on the same task set. A drop in success with no error message is the usual sign of a lost constraint.

Sources

  1. Lance Martin (2025). Context Engineering for Agents
  2. LangChain (2025). Context Engineering (republication of Lance Martin's post)
  3. Drew Breunig (2025). How to Fix Your Context
  4. Anthropic (2025). Effective context engineering for AI agents
  5. Anthropic (2025). How we built our multi-agent research system
  6. Tobias Lindenbauer, Igor Slinko, Ludwig Felder, Egor Bogomolov, Yaroslav Zharov (2025). The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
  7. Tiantian Gan, Qiyao Sun (2025). RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation
  8. Yichao 'Peak' Ji (Manus) (2025). Context Engineering for AI Agents: Lessons from Building Manus
  9. Walden Yan (Cognition) (2025). Don't Build Multi-Agents