Home / Blog / Context engineering and memory / A Context Engineering Checklist for Long Agent Runs

Context engineering and memory

A Context Engineering Checklist for Long Agent Runs

A context engineering checklist of 40 pass-or-fail lines for long agent runs, with a budget that adds up and a policy block. Audit one recorded run.

By Enrique Gutiérrez · Published · 25 min read

A context engineering checklist is a set of true-or-false statements about what enters an agent’s context window on every call: which sources, in what order, under what token budget, and what happens when the budget runs out. This one has 40 lines in ten groups, each tagged with its operation and the chapter that explains it.

I wrote this context engineering checklist for a tech lead whose team already runs an agent and manages its context by folklore. One engineer restarts at a size they have learned to distrust, another lets the harness summarize, a third never compacts. By the end you should have three things to post in the team channel: the ticked list, a budget that adds up, and a filled-in policy block.

How do you use a context engineering checklist on one recorded run?

Take one recorded run of your own system, one hour, and tick only what the evidence shows. A line passes when the place it names (the trace, the prompt assembly code, the memory store or the policy block) shows the statement to be true. “We intend to” and “I can’t tell” both count as no. When a run delegates, tick the list for the parent window, then repeat the first group and the tool-results group for one worker, with that worker’s own ceiling.

Pick a long run: one that reached a late step, used tools and, if your system compacts, was compacted at least once. A line about an event the run never contained gets “not applicable” with the reason written beside it, and “we keep no durable memory, by decision” is a passing reason. Lines that fail become tickets, each with a chapter to read.

Every line carries one of the four operations of context engineering, which are write, select, compress and isolate, or the tag “none”. The tags, the grouping and the pass conditions are my arrangement; the practices come from Chapters 2, 5, 7, 8, 9 and 19 of AI Agents, Engineered unless a line says otherwise.

Group Lines Operations Where you look Chapter
Own the window 4 none trace, policy 7, 9
Order 2 none assembly code, two consecutive calls 2, 19
Standing instructions 4 select, compress, write the always-loaded file, its history 7, 9
Tool definitions 3 select tool configuration, trace 7
Retrieval 6 select, compress, none retrieval code, index job 8
Memory 7 write, select, compress, none memory store, write path 7, 9
Tool results 3 compress, write tool code, a late call 5, 7, 9
History and compaction 4 compress, write compaction events in the trace 9
Isolation 4 isolate delegation calls in the trace 9
Verification 3 none test suite, incident notes 7, 9

Eleven of the 40 lines in this context engineering checklist carry no operation. Nine cover ownership, order, the budget and verification, one is the grounding instruction for retrieval and one asks where the memory store came from. A list built only from the four verbs would miss all eleven.

Can you see and count what the model read?

You can if any call’s assembled input prints from a stored trace with a token count per component. Every later group depends on this one, because a line you cannot check against a recorded window is an opinion. Chapter 7 calls the assembled input “a first-class artifact of your system, worth inspecting, versioning, and testing like code”.

The same chapter says where the ceiling comes from: “The only trustworthy way to find the soft limit for your model and your task is to measure it. The spec sheet will not tell you.” The inventory of a single call is a separate job, and the per-call manifest in context engineering vs prompt engineering is the form for it.

  • Any call’s exact assembled input can be printed from a stored trace. Where to look: print a late step of the run you picked. none · ch. 7
  • Every component of the window has a token count per step and a named owner. Where to look: the trace’s token accounting and the policy block. none · ch. 9; the owner field is this post’s addition
  • The utilization ceiling was measured on our model and our tasks, and the policy records the date. Where to look: the eval report behind the number. none · ch. 7
  • The component caps sum to less than the ceiling, and the policy states how many steps may pass before a trim fires. Where to look: the policy block and the assembly configuration. none · the practice is ch. 7; the arithmetic is this post’s

What does a budget that adds up look like?

A budget adds up when every component has a cap, the caps sum to less than the ceiling, and the gap divided by the growth per step gives the number of steps before a trim must fire. The first group of the context engineering checklist ends on that arithmetic, and the example below is illustrative: every token figure is a round number I invented, including the window size.

Component Cap (tokens) Share of the sum
Standing instructions 2,000 3%
The user’s request 500 1%
Conversation so far 20,000 34%
Memory carried in 1,500 3%
Retrieved passages 12,000 21%
Tool definitions 6,000 10%
Tool results 15,000 26%
Worked examples 700 1%
Output shape 300 1%
Sum 58,000 100%

A cap in this table is the size a trim brings a component back to, and the headroom is how far the pile may overshoot the caps between trims. The sum is 2,000 + 500 + 20,000 + 1,500 + 12,000 + 6,000 + 15,000 + 700 + 300 = 58,000 tokens. It is the planner’s sum of my invented caps, and no chapter prints it.

Take an invented window of 200,000 tokens and a 40% ceiling, and the ceiling is 0.40 × 200,000 = 80,000 tokens. Utilization right after a trim is 58,000 / 200,000 = 29%, and the headroom is 80,000 − 58,000 = 22,000.

If each step adds 2,000 untrimmed tokens, the run reaches the ceiling in 22,000 / 2,000 = 11 steps and the hard limit in ⌊(200,000 − 58,000) / 2,000⌋ = 71. The policy reading is the first number: a trim or a compaction fires by the 11th step. The 40% is the book’s illustration of a practitioners’ figure, which Chapter 7 calls “a practitioner’s rule of thumb for particular models and tasks, offered here as an illustration of the practice, never as a constant”. Replace it with the ceiling you measured.

What does the planner show for these numbers?

The planner below opens on the budget above and shows 29% utilization, 22,000 tokens of headroom, about 11 steps to the ceiling and about 71 to the hard limit. It marks history, tool results and retrieved documents for an operation, since each holds 15% or more of the pile.

With JavaScript on, the Context window budget planner runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Context window budget planner on its own page to share a result by link.

It models growth with no trim applied, so 11 is the longest interval between trims. The neighboring post on context engineering vs prompt engineering runs the same planner on a single recorded call, to show what share of one window a person wrote; here it holds a policy for a whole run, which is why the seed differs. That post’s table of changes by kind of system gives the reason for a practice, and the lines here give its pass condition.

Why cap the component as well as each result?

A cap on the component catches what a limit on each item misses. Suppose each tool result has a hard limit of 5,000 tokens, one that a single result may never exceed. Five results of 4,000 tokens each pass that guard and together hold 5 × 4,000 = 20,000 tokens, which is 5,000 over the 15,000 cap for tool results. That overshoot comes out of the 22,000 headroom until a trim restores the cap.

One open issue on an open-source personal-assistant agent, filed in July 2026, reports this shape: “multiple medium-sized results that collectively overflow the context are not caught”. The context window budget planner takes your own nine counts on its own page.

Is the window assembled in the right order?

The order is right when stable content comes first, variable content comes last, and the instruction follows any long material. Two rules from the book cover it. Chapter 19 gives the caching rule, “stable content first, variable content last”, and Chapter 2, which is free to read, gives the placement rule: “The last thing the model reads should be the thing you most need it to do.”

The two rules fit together. Standing instructions sit at the top and stay byte-identical, and the one rule that carries the load is restated as a fixed line at the tail. One agent product’s team wrote in 2025 that “even a single-token difference can invalidate the cache from that token onward” (Ji, 2025), and their example was a timestamp at the top of the system prompt.

A trim or a compaction rewrites earlier turns, so each one costs a cache rebuild from the point of the rewrite. That is an argument for batching them at boundaries, and the prompt caching savings calculator prices the difference.

  • Nothing that changes on every call sits ahead of content that stays the same. Where to look: diff two consecutive calls with no trim or compaction between them; they should match byte for byte up to the first block the policy lists as variable (the restated rule, fresh passages, the new turn). none · ch. 19
  • When long material is supplied, the instruction follows it, and the one load-bearing rule is restated near the end. Where to look: the ordering rule in the assembly code, then the tail of a late call. none · ch. 2 (free)

What belongs in the standing instructions and the tool list?

Only what almost every task needs belongs there, because both are read on every pass of every run. Chapter 9 calls the always-loaded core the hot tier and gives its rule in four words: “keep it ruthlessly small”. The same chapter treats that file as procedural memory, which is why its edits get review: “A wrong fact costs you an answer. A wrong rule costs you the agent.”

The question of when a tool should become a skill covers which procedures move out of the standing file and which rules must stay. One line here points at it.

  • Each line of the always-loaded instructions has a recorded reason: a mistake its removal would cause. Where to look: the file and its review history. select · ch. 9
  • No standing instruction contradicts another standing instruction or a tool description. Where to look: read the file and the tool descriptions side by side. compress (prune) · ch. 7
  • Every block in the always-loaded file is needed by most tasks or holds a rule that one miss would make costly; everything else loads on demand. Where to look: the file against the on-demand library. select · ch. 7
  • Edits to standing instructions are reviewed and versioned like code, including edits the agent proposes. Where to look: the file’s commit history. write · ch. 9

Which tools should a run carry?

A run should carry the tools its task needs and keep that set for the whole run. Chapter 7 names the first half the tool loadout: “expose only the subset of tools relevant to this task”. The test for ambiguity comes from a 2025 engineering post by Anthropic: “If a human engineer can’t definitively say which tool should be used in a given situation, an AI agent can’t be expected to do better.”

The third line below rests on a source outside the book. The same 2025 product team advises: “unless absolutely necessary, avoid dynamically adding or removing tools mid-iteration”. Tool definitions sit near the front of the window, so a change there invalidates the cache behind it, and earlier steps then refer to tools that are no longer defined.

  • Each task type has a named tool loadout, and its size in tokens is recorded. Where to look: the tool configuration and the trace’s count for definitions. select · ch. 7
  • A person reading any two tool descriptions in a loadout can say which one applies to a given request. Where to look: hand the descriptions to a colleague who has never seen the system. select · ch. 7; the test’s wording is Anthropic’s (2025)
  • The loadout stays the same from the first call of a run to the last. Where to look: compare the tool definitions of an early call and a late call. select · outside the book (Ji, 2025)

Retrieval or a long window: how do you choose?

Paste the corpus whole when it is bounded, fits under your measured ceiling and needs close reading; retrieve when any one of four conditions holds. Chapter 8 lists them: “the corpus outgrows the desk; it changes faster than you care to re-paste; it is private and must remain so; or the answers need receipts”. Only the first is about size.

Three ways to give a model knowledge.
Figure 8.5 Three ways to give a model knowledge. RAG, in accent, pulls fresh or private facts from an external store into the window at question time, leaving the bulk behind; long context places a whole bounded corpus inside one large window for close reading; fine-tuning adjusts the model itself, changing how it behaves and formats rather than what it can look up. RAG is the lowest-risk place to start, which is why it leads. Reuse this diagram
Situation Choice What it costs Source
The fact fits in a paragraph of instructions Put it in the prompt Nothing to build ch. 8
One contract, one handbook or one service’s code, under the ceiling Long window Every request pays for the whole corpus ch. 8
The corpus is large, changing or private, or answers need receipts Retrieval A pipeline to run, latency per question, misses to tune ch. 8
A stable behavior at high volume (a register, a format) Fine-tuning, last A training run for every revision ch. 8

The book’s rule compresses the table: “Retrieval for what the model should know; fine-tuning for how it should respond; the long window for close reading of a bounded set.” Its defense of the long window is one sentence: “Retrieval hands the model fragments; the long window hands it the book.”

What does the evidence say about “RAG is dead”?

The published comparisons support both sides of the argument, and neither supports a slogan. Li and colleagues (2024) compared the two on models of that year: “when resourced sufficiently, LC consistently outperforms RAG in terms of average performance. However, RAG’s significantly lower cost remains a distinct advantage.” LC is their abbreviation for long context.

More passages do not rescue either approach. Jin and colleagues (2024) found that for many long-context models “the quality of generated output initially improves first, but then subsequently declines as the number of retrieved passages increases”. Chapter 8 says it in one sentence: “More retrieved context is not better context.”

So the commenter who wrote in January 2026 “i thought rag/embeddings were dead with the large context windows”, and blamed a chatbot for the idea, had been given the answer to one question of four. A window of any size leaves change, privacy and receipts where they were.

Search the agent runs for itself counts as retrieval in this context engineering checklist, and that includes code search. A system with no index answers the first and the last line, answers the fourth if its answers are meant to come from fetched material, and marks the rest “not applicable, because there is no index”.

  • The policy records the choice between a long window and retrieval, with the four conditions answered. Where to look: the policy block. select · ch. 8
  • A fixed maximum number of passages reaches the model after reranking, and the number was chosen on our eval set. Where to look: the retrieval code and one retrieval call in the trace. select · ch. 8
  • Metadata filters run alongside the similarity search, and a corpus with identifiers is searched by keyword as well as by meaning. Where to look: the query the retrieval code sends. select · ch. 8
  • For questions the corpus should answer, the instructions tell the model to answer only from the supplied passages and to say when they do not contain the answer. Where to look: the standing instructions. none · ch. 8
  • Superseded documents are retired, and the index is rebuilt when documents, the embedding model or the chunking change. Where to look: the index job and the date of its last run. compress (prune at the source) · ch. 8
  • Search the agent directs itself has a hop cap and a spend cap that appear in the trace, and its search tool returns identifiers and snippets (paths, titles, dates) ahead of full text. Where to look: a multi-hop search in the trace. select · ch. 8

Who may write to the agent’s memory?

Each kind of memory gets its own writer: the agent may append to its run log, write facts through a filter, and propose changes to procedures that a person reviews. That split is my reading of Chapter 9, which ranks the three kinds by how dangerous a bad write is and puts the rule for untrusted input plainly: “filtering belongs on the write path, before persistence”.

Agent memory arranged by temperature.
Figure 9.2 Agent memory arranged by temperature. The hot core rides the context window permanently and pays rent on every call, so it is kept ruthlessly small; warm material is loaded when a matching task begins; the cold bulk waits in external storage until retrieval faults it back in. The circulation—write out what must survive, select back in what the step needs—is the virtual-memory pattern: a small desk and disciplined movement, producing the illusion of an unbounded one. Reuse this diagram

The first two lines apply to every long run, with or without a durable store. Chapter 9: “Working state belongs to the run; durable knowledge belongs to the relationship.” If you keep no durable store, mark the other five “not applicable, by decision”. Chapter 7 backs that position: “no paging memory system should exist before an observed failure asks for it”.

The measurements are of models and systems from 2024. Wu and colleagues (LongMemEval) report “a 30% accuracy drop on memorizing information across sustained interactions” for commercial assistants and long-context models. Chen and colleagues (AgentPoison) demonstrated an attack with “an average attack success rate higher than 80%” at “a poison rate less than 0.1%”; that is a result under attack conditions and says nothing about how often poisoning happens.

A later paper (Dong and colleagues, 2025) injects records “by only interacting with the agent via queries and output observations”. The everyday version is duller. One commenter found in April 2026 that some entries an agent had written to its own memory file were wrong, and that managing those files was “more than I can expect people on my team to manage on their own”.

  • Run state and durable knowledge live in different places. Where to look: the path of the plan file and the location of the fact store. write · ch. 9
  • The plan, the progress and the decisions are in a file the agent rereads after any compaction or reset. Where to look: the first calls after a compaction in the trace. write · ch. 7, ch. 9
  • Memory has tiers with a loading rule each: a small core always loaded, task material loaded when the task starts, the rest fetched on demand. Where to look: the assembly code’s memory loader. select · ch. 9
  • Nothing written from untrusted input, including a worker’s summary of it, reaches a durable store without passing a filter. Where to look: the write path in code. write · ch. 9
  • Every stored fact has a timestamp, an expiry rule and a rule for what happens when a new fact contradicts it. Where to look: three entries picked from the store. compress (prune the store) · ch. 9
  • Reads are filtered by owner before any scoring, then scored on relevance, recency and proven usefulness. Where to look: the read query in code. select · ch. 9
  • The store exists because of a failure somebody observed and wrote down. Where to look: the note or commit message that introduced the store. none · ch. 7

What happens to tool results once they have been used?

They are cleared, leaving a one-line fact and a reference. Chapter 7 gives the move: “once a result has been acted on, clear the payload, keep the one-line fact that the call happened, and keep the reference so re-fetching is cheap”. This is the group with the best outside measurement.

Lindenbauer and colleagues (2025) compared two ways of shrinking an agent’s history on one coding benchmark, across five model configurations. They found that “a simple environment observation masking strategy halves cost relative to the raw agent while matching, and sometimes slightly exceeding, the solve rate of LLM summarization”. It is one benchmark of coding tasks.

Sources disagree about resolved errors. The public guide 12-Factor Agents advises: “Consider hiding errors and failed calls from context window once they are resolved.” Chapter 7 agrees, while the 2025 product team quoted earlier says to “leave the wrong turns in the context”.

My reading: keep a failure visible while the same action could still be retried, then replace it with a one-line fact. The line below asks only that your team has decided.

  • Every tool has a hard cap on what it returns; past the cap, the output goes to a file and the tool returns the relevant slice and a pointer. Where to look: the tool code and the largest tool result in the trace. compress and write · the cap is ch. 5, the slice and the pointer are ch. 9
  • A result that has been acted on is replaced by a one-line fact and a reference within the number of steps the policy states, or the policy says “never, for runs under” a stated length and the run was shorter. Where to look: the policy block, then a late call for payloads from early steps. compress (trim) · ch. 7, ch. 9
  • Pruning rules name categories a program can recognize (old tool payloads, resolved errors, superseded drafts), and the policy says what happens to a resolved error. Where to look: the pruning code against the policy block. compress (prune) · ch. 9; the sources disagree on errors

When should the history be compacted, and what must survive?

Compaction should fire below your measured ceiling, at a phase boundary, after the plan is in a file. Chapter 9 gives the reason for firing early: near the limit, “the judgment you are relying on to rescue the run is the judgment the overfull window has already degraded”. The book’s free glossary calls compaction “lossy by design, so what the summary drops is a design decision”.

On what survives, Chapter 9 is specific: “What must survive a compaction hardly varies across reports: the decisions made and why, the problems still open, the current plan, and the handful of artifacts under active work.”

Here the evidence runs out. I found no measurement of how often a compaction drops a constraint, and none of quality by trigger point. What exists is a volume of reports, such as this open issue from September 2026: “requirements I stated at the start of a session, and constraints I set mid-session, stop being followed”. Treat any percentage in your policy as a value to measure on your own tasks.

Harness-provided compaction is a category with a trigger somebody chose for you. Two examples, both read on 2026-10-06: one model provider’s API documentation (OpenAI) describes a compaction item that is “opaque and not intended to be human-interpretable”, and one agent framework’s documentation (Google’s Agent Development Kit) summarizes older history “including instructions, inputs, and model responses”. A summary nobody can read cannot be reviewed, and a summary that covers instructions can paraphrase a rule. That is one of the mechanical reasons why AI agents ignore instructions late in a run.

  • Every compaction in the trace fired at a phase boundary the policy names or at its backstop percentage, at or below the ceiling, and none fired at the hard limit. Where to look: the token count at each compaction event, then the policy block. compress · ch. 9; the trigger value is yours to measure
  • The compaction instruction names what must survive, and it was tuned on our own traces for recall before precision. Where to look: the compaction prompt and its test cases. compress · ch. 9
  • Constraints the user sets during a run are written to the plan file when stated, and after every compaction that file and the standing rules are reinserted word for word, or reread by the agent as its first action. Where to look: the first call after a compaction and the tool call that follows it. write and compress · the principle is ch. 9; the line is this post’s
  • Starting fresh from the notes is a documented move with written triggers. Where to look: the policy block. write and compress · ch. 9

When does work go to a separate window?

Work goes to a subagent when it is bulky and separable, or when the content is untrusted; it stays in the main thread when the pieces share context that keeps changing. Chapter 9 calls the pattern quarantine and prices it: total token spend multiplies, and tightly coupled work split across sealed windows comes back as fragments.

The rule for the boundary is the chapter’s quarantine rule, printed as a display line: “Nothing enters an isolated context except what its brief deliberately admits, and nothing comes back out except a distilled result sized for the desk that will read it—carried as evidence, never as orders.” A system with no delegation marks the first three lines of this group “not applicable”; the fourth still applies to any agent that reads content it does not control.

  • The policy states when work goes to a separate window, and every delegation in the trace fits it. Where to look: the policy block against the delegation calls. isolate · ch. 9
  • Every delegation brief carries the objective, the expected shape of the result, the tools and sources, and the boundaries, with no transcript attached. Where to look: the input of one delegation call. isolate · ch. 9
  • What a subagent returns has a size limit and enters the parent window in the position of a tool result. Where to look: the return of one delegation call. isolate · ch. 9
  • Untrusted content is read by a worker with a minimal toolset and no privileges. Where to look: which window reads uploads, web pages or ticket text, and that reader’s tool configuration. isolate · ch. 9

How do you know a trim did no harm?

You know when a test would have failed had the trim dropped something a later step needed. Chapter 9 explains why reading the output will miss it: “a compression that quietly hurts task success looks identical, from the inside, to one that saved the run”. Replay and step-wise grading are practices this cluster of posts adds to the book’s advice to measure.

The last line borrows the clean-desk test, which belongs to the post on context rot and its four fixes; the five-minute explainer shows the same ground in motion.

  • Recorded windows are replayed after every change to assembly code. Where to look: the test suite’s replay job. none · ch. 7 asks for tests; replay is this cluster’s
  • Every trimming, pruning and compaction rule has a test that fails when the rule drops something a later step needed, graded at an early, a middle and a late step. Where to look: the tests beside each rule. none · ch. 9; grading at three steps is this post’s
  • A failure is reproduced on a clean, minimal context before anyone edits the prompt or blames the model. Where to look: the notes of the last incident. none · ch. 7

What goes in the context policy block?

The policy block holds the choices the context engineering checklist asks about: the ceiling, the caps, the order, the trims, the compaction trigger and survival list, the retrieval decision, the memory writers, the isolation rule and the checks. It is my synthesis of Chapters 7, 8 and 9. Fill it in for one agent and commit it beside the assembly code.

CONTEXT POLICY: [agent or feature] · owner: [name] · reviewed: [date]

Budget
  Utilization ceiling: [ ]% of the window, measured on [task set] on [date]
  Caps (tokens): instructions [ ] · request [ ] · history [ ] · memory [ ] ·
    retrieved passages [ ] · tool definitions [ ] · tool results [ ] ·
    examples [ ] · output shape [ ]
  Sum of caps: [ ] (must be under the ceiling) · growth per step: [ ]
  Steps before a trim or compaction must fire: (ceiling - sum) / growth = [ ]
  Owner per component: [ ]

Order
  Stable first: [what never changes within a run]
  Variable last: [what changes per call]
  Load-bearing rule restated at the tail: [the rule]

Standing instructions and tools
  Always loaded: [file] · loaded on demand: [library]
  Edits reviewed by: [ ] · tool loadout per task type: [name: tools, tokens]

Tool results
  Hard cap per result: [ ] tokens · past the cap: [file + slice + pointer]
  Cleared after: [ ] steps | never, for runs under [ ] steps · leaves: fact + reference
  Resolved errors: [hidden after ... | kept because ...]

History
  Compaction fires at: [phase boundary], and no later than [ ]% · measured on: [task set]
  Survives word for word: constraints, standing rules (from [file]) ·
    put back by: [harness reinserts | agent rereads first]
  Survives in the summary: decisions and why, open problems, current plan,
    artifacts in progress
  Start fresh from the notes when: [signs]

Retrieval
  Decision: [long window | retrieval | both | not applicable]
  Outgrows the window? [ ] · changes? [ ] · private? [ ] · needs receipts? [ ]
  Index: [yes | none, agent search only] · grounding instruction for: [which questions]
  Passages per call after reranking: [ ] · filters: [ ] · keyword search: [ ]
  Agent-directed search: hop cap [ ] · spend cap [ ]

Memory
  Run state lives in: [file] · durable facts live in: [store | none, by decision]
  Who may write: run log [agent] · facts [agent, filtered] · procedures [human review]
  Expiry: [ ] · contradiction rule: [ ] · scoped per: [user | project]

Isolation
  Goes to a separate window when: [bulky and separable | not applicable]
  Untrusted content is read by: [worker, minimal tools | none is read, because ...]
  Brief carries: objective, result shape, tools and sources, boundaries
  Return capped at: [ ] tokens, treated as evidence

Checks
  Replay set: [ ] recorded windows · graded at steps: [early, middle, late]
  Checklist last run on trace: [id] · lines failed: [ ] · tickets: [ ]

What can the checklist not tell you?

A context engineering checklist cannot tell you whether the agent is good at its task, and it cannot supply a number. Every line can pass on a system that still fails, because the list checks that decisions were made and are visible. Whether they were the right decisions is a question for your evals.

Most lines rest on the book’s reasoning and on practice. Measurement exists for input length, position, passage count, tool count and one comparison of masking against summarizing. I found none for compaction triggers, constraint survival or memory expiry rules. The six studies cited used models from 2024 and 2025, I read only their abstracts, and thresholds move with each generation.

Some systems should mark most of the list “not applicable”. A single call has no history to compact and no results to trim, so the first two groups, the standing instructions and, if it retrieves, the retrieval group are the whole audit. A run of a dozen steps that stays far under its ceiling can skip compaction and the first three isolation lines.

Over-curation is the opposite failure, and Chapter 7 names its price: “The trade is always recall against focus”. My harvest of reader questions was mostly from coding agents, so support and research agents may weigh the groups differently.

One recorded run, one hour

The use of a context engineering checklist is to turn “how we manage context” from habit into a page somebody can review. Chapter 9 closes its account of long runs with the line I would put at the top of that page: “a healthy long run is a chain of short runs joined by good notes”. Run the list on one trace this week, fill in the block, and file a ticket for each line that failed.

The order rule is in Chapter 2, free to read, and the glossary is free too. Chapter 7, “Managing the Context Window”, Chapter 8, “Retrieval and Knowledge” and Chapter 9, “Memory: Working State Across Long Runs” are in the full book, as are Chapter 5 and Chapter 19, where the response cap and the caching rule live. The book is written in pseudocode throughout, so nothing in it depends on one harness. The context engineering guide collects the neighboring posts and tools, or you can see the formats.

Questions readers ask

What is a context engineering checklist?
It is a list of statements about what enters a model's context window on each call: which sources, in what order, under what token budget, and what happens when the budget is reached. Each statement is true or false of a real system, and each can be verified in a recorded trace, in the code that assembles the prompt or in the memory store.
When should an agent compact its context?
Below the ceiling you measured and at a phase boundary, after the plan and the decisions are in a file. Chapter 9 of AI Agents, Engineered gives the reason: near the limit the summary is written by the same model reading the same overfull window. No published measurement found for this post fixes the trigger, so treat the threshold as a value to measure on your own tasks.
What must survive a compaction?
The decisions made and why, the problems still open, the current plan and the artifacts under active work, according to Chapter 9. This checklist adds one line of its own: constraints and standing rules are reinserted word for word from a file, because a summary may paraphrase them.
How do I choose between RAG and long context?
Paste the corpus whole when it is bounded, fits under the ceiling you measured and needs close reading. Retrieve when the corpus outgrows the window, changes often, must stay private or the answers need receipts. A 2024 comparison by Li and colleagues found the long window ahead on average performance when given enough resources, and retrieval much cheaper.
How many tokens should each part of the context window get?
No fixed split holds across models and tasks. Measure a utilization ceiling on your own tasks, cap each component, and check that the caps sum to less than the ceiling. The gap divided by the growth per step is the number of steps the run can take before a trim or a compaction has to fire.

Sources

  1. Lindenbauer, Slinko, Felder, Bogomolov, Zharov (2025). The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
  2. Li, Li, Zhang, Mei, Bendersky (2024). Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach (EMNLP 2024 industry track)
  3. Jin, Yoon, Han, Arik (2024). Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG
  4. Wu, Wang, Yu, Zhang, Chang, Yu (2024). LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
  5. Chen, Xiang, Xiao, Song, Li (2024). AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
  6. Dong, Xu, He, Li, Tang, Liu, Liu, Xiang (2025). Memory Injection Attacks on LLM Agents via Query-Only Interaction
  7. Anthropic (2025). Effective context engineering for AI agents
  8. Yichao "Peak" Ji (2025). Context Engineering for AI Agents: Lessons from Building Manus
  9. HumanLayer and contributors (2025). 12-Factor Agents, Factor 3: Own your context window
  10. OpenAI (2026). Compaction (API documentation, read 2026-10-06)
  11. Google (2026). Context compaction (Agent Development Kit documentation, read 2026-10-06)
  12. mooire733 (GitHub) (2026). [FEATURE]: Compaction: deterministic retention of user messages and a configurable compaction prompt (GitHub issue)
  13. ogilvymiles (GitHub) (2026). Context Overflow: large tool outputs exceed context window, compaction can't recover, sessions enter failure loop (GitHub issue)
  14. JohnMakin (Hacker News) (2026). Hacker News comment on auditing an agent's self-written memory
  15. mooball (Hacker News) (2026). Hacker News comment on retrieval and large context windows