Home / Blog / Context engineering and memory / What Is Context Engineering? Curating a Scarce Desk

Context engineering and memory

What Is Context Engineering? Curating a Scarce Desk

What is context engineering? Deciding what a model reads on every call. Learn what is sent each time, three limits on it, and why a model seems to forget.

By Enrique Gutiérrez · Published · 18 min read

What is context engineering? It is the work of deciding what goes into a language model’s input on each call. It starts from one fact: apart from what it learned in training, which no conversation changes, the model sees only what is sent this time. Every call starts from nothing, so software must resend whatever the model should know.

I wrote this for someone who has used a chat window for a year, has perhaps made a first API call, and was told today that the model has no memory. By the end you should know what is sent on each call, which three limits apply to it, and which of three things happened when the model “forgot”. Every count of tokens below (a token is a chunk of text, about a word) is an illustration I made up; none is a fact about any model.

What is context engineering, in one sentence?

The short answer to “what is context engineering?” fits in one sentence: it is deciding, call after call, which text a model gets to read. Chapter 7 of AI Agents, Engineered defines it as “the discipline of deciding what fills the model’s context window at every step of a run.” The context window is, in the book’s words, “the bounded stretch of text a model can consider in a single call”, and the book pictures it as a desk on which everything the model consults must fit at once.

The run in that definition belongs to an agent: a program that calls the model repeatedly and lets it ask for tools, functions the program runs and whose output goes into the next request.

Simon Willison, in a March 2025 essay written before the name caught on: “Most of the craft of getting good results out of an LLM comes down to managing its context—the text that is part of your current conversation.” Anthropic, a model provider, gives the direction in a 2025 engineering essay: good context engineering means finding “the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome.”

Note the direction: “context” sounds like something to add, and the skill is mostly leaving out. Whether context engineering is a new craft or prompt engineering under a new name is a separate question from what the work is.

Does the model remember your conversation?

No. The model keeps nothing between calls, and what looks like memory is the software around it sending the conversation again each time. Chapter 2 puts it this way: “the model does not learn from your conversations; nothing you type changes a single weight [a number inside the model], and what feels like memory within a session is only the growing transcript being re-read on every call.”

What it learned in training is a separate store, which the same chapter describes as “frozen at the cutoff”. The text in front of it is the only part you control.

The technical word is stateless. Two providers’ API documentation, both read in October 2026, say it about their own interfaces: one says “each text generation request is independent and stateless”, and the other that its API “is stateless, which means that you always send the full conversational history to the API.” Those are statements about two products on one date. The book says it of the category, “An LLM API is stateless (each call starts from nothing)”, and says what follows: “the surrounding software lays the entire transcript back on the desk every turn”.

The context window drawn as one bounded frame holding instructions, conversation, tool definitions, tool results, fetched documents and the growing answer, with everything outside it marked as not existing for this call.
Figure 2.2 The context window as a single finite desk: everything the model can use at this moment must physically fit on its surface. Reuse this diagram

Chapter 2 states the consequence for one call: “If a fact is on the desk during this call, the model can use it; if it is anywhere else in the universe, then for the duration of this call it does not exist.”

Why did it seem to remember, then?

It seemed to remember because some software kept the transcript and resent it. Three versions of this differ only in who keeps the transcript.

“It remembered my last message.” In a chat product, the product’s software keeps the transcript and sends it, or a trimmed or summarized version of it, with each new message. In your own program that code is yours, and a second call made without the first exchange knows nothing of it.

“The provider stores the conversation for me.” Some interfaces let a call refer to earlier turns by an identifier in place of resending them. The first provider’s documentation is plain about what that changes: even then, “all previous input tokens for responses in the chain are billed as input tokens in the API.” The storage moved, and the earlier turns are still billed as input on every call.

“The product has a memory feature.” I did not verify how any product implements one, so I say only what the mechanism allows. Nothing you type changes the model while you use it, so whatever is recalled has to reach it as text that software placed in a request. Chapter 7 is blunt: “There is no side channel: no memory the model keeps privately between calls, no rapport accumulating across the session, nothing carried over but what your code lays back on the desk.”

Telling a model to forget something removes nothing, because the text is still in the request. A new conversation is what clears it.

What counts against the context window?

Everything in the request counts, and so does the answer while it is being written. The unit is the token, which Chapter 2 defines as “a chunk of text from a fixed vocabulary—often a word, sometimes a piece of one or a punctuation mark; on average, think three-quarters of an English word.” Each model family splits text its own way, so “a count borrowed from one is only an estimate for another.”

Chapter 2 then lists what lies on the desk “in a typical agent call”: “your standing instructions, the whole conversation so far, the definitions of every tool the agent might use, the full text of every tool result that has come back, whatever documents were fetched for the task—and the answer the model is writing, which lands on the same desk as it grows. One token budget covers them all.” Standing instructions are the text the software puts at the start of every request.

The last item is the one a chat window never shows you: input and output share the budget. The first provider’s documentation, on the same page, says the maximum for its context window “includes input, output, and reasoning tokens”, the last being text some models write for themselves before they answer.

What happens at the edge of the window?

One of three things happens, and you may not be told which. Chapter 2: “the request may be rejected outright with an error; or the input may be truncated—cut down until it fits, usually by dropping the oldest turns, often without telling anyone; or a framework in the middle may summarize the history before sending it.”

The book’s verdict follows: “The outright error is the kind outcome. Silent truncation is the treacherous one”. The second provider’s documentation on context windows (read October 2026) notes that chat interfaces “can also manage the context window on a rolling ‘first in, first out’ basis”; I have not checked how any product trims.

Hence the book’s advice: “When a long session starts ‘forgetting,’ check for overflow before you blame the model.” Many models also cap a single response well below the window, and the same passage says, “An answer that stops mid-sentence usually hit the output cap.” The remedy for that one is to ask the model to continue, or to ask for less per reply.

Why is the window scarce in three different ways?

The window is scarce because three separate things run short as it fills: room, money (with some waiting) and attention. Chapter 2 names room, the bill and the wait in one clause: “nearly every limit you will hit while building—the ceiling on what the model can see, the bill at the end of the month, the pause before an answer begins—is denominated in tokens.” It adds attention later in the same section: “even what fits is used unevenly.” Counting three scarcities, with the wait filed under the bill, is this post’s arrangement and not the book’s.

Scarcity What it is How you notice it What it costs
The hard limit A maximum number of tokens per call, input and output together An error, or nothing at all when something trims silently The oldest material, which may include your first instructions
The bill, and some waiting Every token in the request is read, and billed, on every call The input count reported with each response keeps climbing Money in proportion to the tokens read; a little time
Attention In studies of models from 2023 and 2024, a given fact was used less well as the text around it grew A rule followed early and ignored later, with no error and room to spare Quality, before the limit is reached

The bill is the simple one: you pay for what the model reads, and it reads the whole request every time. Chapter 2 notes that providers soften the cost of re-reading an unchanged pile with caching. That changes the price of those tokens and leaves them in the window.

Attention is the limit nothing reports. Chapter 2, citing a study first posted in 2023 (Liu and colleagues), says models recall material near the beginning and the end of a long input better than material in the middle, and concludes, “A bigger desk therefore solves less than it appears to: the edges stay sharp while the middle grows.”

A long row of tokens on the desk keeps growing to the right, while a fixed attention budget lights only a small pool of them; everything outside the pool is present yet barely seen.
Figure 7.2 The attention budget as a fixed pool of lamplight on an ever-longer desk: the desk can grow without limit, but the circle of light cannot. Everything outside it is present yet barely seen. Reuse this diagram

Does a shorter request get a faster answer?

Yes, but usually by little unless the request is very large, and I found no published controlled measurement. Chapter 2 describes the time before the first token: a bloated context “makes the agent slower to start answering”. One provider’s latency guide (undated, read October 2026) describes the whole response on ordinary prompts, where writing the output dominates: fewer input tokens do lower latency, but “this is not usually a significant factor—cutting 50% of your prompt may only result in a 1–5% latency improvement”, except at “truly massive context sizes”. That figure has no data attached, so do not count on a faster reply.

What does “it forgot” actually mean?

It means one of three different things: the text was never sent, it was cut off, or it was sent and is dim. The three-way split is this post’s arrangement, and it classifies a failure without diagnosing it. An answer that stops mid-sentence is a different case, the output cap described under the edge of the window, and this rule does not cover it.

Take the request of the failing call, as the model received it, and search it for what the model failed to use. Search for the fact and not only the exact sentence, because a summary may have kept it in other words. Then apply the three tests in the table, in order.

Kind The test How to check What to do
sent and dim The fact is in the failing request and was not acted on Search the request as the model received it Rerun the step on a short, clean input holding only the rule and the material it applies to. That rerun separates a context failure from a prompt failure: if it works there, shorten the pile or restate the rule last
cut off The fact is absent, and a limit, a trim rule or a summary removed it, whether from a later request or before its first one (a document truncated on upload, a clipped tool result) The token count is at a limit; your code keeps only the last N messages; a summary step ran; an upload or a tool returned less than the whole Put it where it is resent on every call: the standing instructions, or a file reloaded each time (the file is “write” among the four fixes for context rot). For material clipped on the way in, send less at once
never sent The fact is absent, and no code ever tried to include it Nothing trimmed it: it was said in another conversation, saved but not included, or never fetched Put it in the request. In your own loop, fix the append (the post on what an agent loop is shows where it lives)

“Dim” covers what the post on context rot files under distraction and clash, a rule buried by a long history or contradicted by other text, and says nothing about its other three failure modes. When the contradicting text arrived inside a document or a tool result and carried instructions of its own, that is also a security matter, prompt injection.

When a provider assembles the conversation from a stored identifier, the earlier turns are not in what your program sent. You cannot run the search, and the next section applies.

What if you cannot print the request?

Then two probes stand in for the search, and both are imperfect.

Ask for a quotation. Ask the model to quote, word for word, the sentence it ignored or the passage it skipped; for a document, ask for its last sentence. If the quotation matches your copy, the text is present and the kind is sent and dim. If the model cannot quote it, or invents something, the text is probably absent.

Say it again. Restate the rule, or paste the passage, at the end of your next message. That repairs cut off and dim alike, in a short conversation as well as a long one. A chat gives you no way to delete old text, so when a correction does not stick, start a new conversation that states only the corrected version.

A new conversation needs no probe: nothing from the old one is in it unless you or a memory feature put it there. The kind is never sent and the fix is to paste in what matters. Whether a memory feature dropped the fact or never stored it, you cannot tell from outside, and the fix is the same.

The ten reports below are sorted by kind. Pick a kind to see only its rows.

What you see What happened What to do Kind
Your program’s second call does not know a name the first call was given Each call went out alone; nothing added the first exchange to the second request Keep a list of messages, append both sides of every turn, send the whole list each time never sent
A new conversation knows nothing from last week’s A new conversation starts a new transcript, and last week’s is not in this request Paste in what matters. In your own program, load a notes file at the start never sent
The history is saved in your program, in a list or a database, and the model answers as if it were empty Saving is not sending: the saved history never reached this request Print the request just before it leaves, then fix the code that builds it never sent
A long session loses its first instructions, and the request’s token count is at the window’s size The oldest turns were dropped so the request would fit Keep the instructions in the slot your trimming code leaves alone (check that it does), and shorten the pile cut off
Your program keeps only the last N messages, and the model asks again for something said before them Your own rule trimmed it, well before any hard limit Raise N, or copy what must survive into the standing instructions cut off
After the conversation was summarized (“compacted”), a detail you gave earlier is gone, in any wording The summary replaced the history and left the detail out Write must-keep details to a file that is reloaded after every summary cut off
It summarized a long document and left out the last third, and it cannot quote the document’s last sentence The document was clipped on the way in, by a size limit on the upload or on the tool that read it Send the document in parts, or only the part the question needs cut off
A rule at the top of the request was obeyed for ten steps and then ignored, and the printed request still contains it The rule is present, far from the end, under a growing pile Trim what has been used, restate the rule last, or start fresh with a short summary sent and dim
It asks a question you already answered; the answer is in the printed request and the window is far from full Present and unused; a long pile is the likely cause Start a new conversation with a short statement of where things stand sent and dim
It keeps using an old price after you corrected it, and both the old price and the correction are in the request Two versions sit in one request and nothing marks which one wins Remove the superseded text; do not just add another correction sent and dim

How many tokens does a twelve-call run read?

It reads 156,000 tokens, in a run whose largest request holds 24,000. I constructed the run, and every number in it is illustrative: a small research agent that calls a search tool on each of eleven calls and answers on the twelfth.

Every request starts with the same 2,000 tokens: standing instructions 600, tool definitions 1,200, the user’s question 200. On each call the model writes about 200 tokens and the tool returns about 1,800, so the pile grows by 2,000 per call. Call 1 reads 2,000 tokens, call 2 reads 4,000, and call i reads 2,000 × i.

Call Tokens read on this call Read so far Instructions’ share of the request
1 2,000 2,000 600 / 2,000 = 30%
2 4,000 6,000 15%
3 6,000 12,000 10%
4 8,000 20,000 7.5%
5 10,000 30,000 6%
6 12,000 42,000 5%
7 14,000 56,000 4.3%
8 16,000 72,000 3.75%
9 18,000 90,000 3.3%
10 20,000 110,000 3%
11 22,000 132,000 2.7%
12 24,000 156,000 600 / 24,000 = 2.5%

Add the second column: 2,000 × (1 + 2 + … + 12) = 2,000 × 78 = 156,000. The last request is 24,000 tokens, so the run read 156,000 / 24,000 = 6.5 times its own final size. The total is easy to get wrong because the last request is the only one you see. It is what one call read, and the bill covers all twelve.

In an illustrative 40,000-token window, call 20 would fill it and call 21 (42,000) would not fit. And the 600 tokens of instructions fell from 30% of the request to 2.5%, which measures crowding and not how well the model attends.

This table is my longer answer to “what is context engineering?”: somebody has to decide what each of those twelve requests contains, and if nobody does, the pile decides.

What does the estimator show for this run?

The agent cost-per-task estimator, mounted below on this run, shows the same 156,000 input tokens beside a second figure that also reads 24,000. That figure is a different wrong guess from the one above: it assumes all twelve calls were as small as the first, 12 × 2,000. It matches the size of my last request only because the fixed part and the growth are both 2,000 here.

The tool says “steps” for what I call calls, and “prefix” for the same 2,000 tokens that open every request. For this run, ignore its fields for subagents, hidden reasoning, caching and prices: the seed leaves them off and sets the retry overhead to zero.

With JavaScript on, the Agent cost-per-task estimator runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Agent cost-per-task estimator on its own page to share a result by link.

The line to remember is the last one in its verdict: doubling the run to 24 steps would take the input to 600,000 tokens. Twice the calls, almost four times the reading.

Is a bigger context window the fix?

Only for the first of the three scarcities. A bigger window moves the edge, lets the bill grow further and promises nothing about attention. In the run above, doubling the illustrative window to 80,000 tokens moves the call that fills it from 20 to 40. By then the run has read 2,000 × (1 + 2 + … + 40) = 2,000 × 820 = 1,640,000 tokens, and the instructions are 600 / 80,000 = 0.75% of the request.

The second provider’s documentation on context windows says as much of its own product: “A larger context window allows the model to handle more complex and lengthy prompts, but more context isn’t automatically better. As token count grows, accuracy and recall degrade, a phenomenon known as context rot.”

Hsieh and colleagues (2024; I read the abstract only) tested 17 long-context models of that year on 13 tasks: all claimed at least a certain context size, yet “only half of them can maintain satisfactory performance” at that size. The rest of the measurements, and the remedies, are in the post on context rot, its symptoms and its fixes.

Where do you go next?

Start by looking. Print the full input of one call your program makes, and log the input-token count that comes back with every response.

The book groups the remedies as four operations, write, select, compress, isolate: keep information outside the window, bring in only what this step needs, shrink what stays, and give a messy subtask a window of its own. A short explainer shows all four at work. They also answer how much history to keep: what the next step needs, with whatever must survive moved into something resent on every call.

The first of those decisions most people meet is about documents: retrieving a few passages or putting the whole source in a long window. A whole source may fit and still be read poorly, and it is re-read on every later call.

Where does this explanation stop?

It stops at the single call and at numbers I made up. The twelve-call run has equal steps and no trimming, and I found no public measurement of how fast real agents’ windows grow. The documentation I quote describes two providers’ interfaces as they read in October 2026, not every way a model can be served.

The studies on length are from 2023 and 2024 and describe the models of those years. And the opposite mistake is real: Chapter 7 warns, “starve the model of a fact it needs and you have traded one failure for another.”

The one fact to keep

If a teammate asks you “what is context engineering?”, start where this post did: apart from its training, the model sees only what is sent this time. Memory is a transcript somebody resent, a limit is a count of tokens, and “it forgot” is a question about one request: was the fact in it? Print the request and you can answer.

Chapter 2 is free to read and has tokens, the desk and the edge of the window. Chapter 7, “Managing the Context Window,” (in the full book) has the definition, the attention budget and the four operations. The context engineering guide lists this post’s neighbors, or you can see the formats.

Questions readers ask

What is context engineering in simple terms?
Context engineering is deciding what text goes into each request to a language model. Apart from what it learned in training, which no conversation changes, the model starts every call from nothing, so whatever the software puts into the request is everything else it knows at that moment. In a chat that is the conversation so far; in an agent it also includes tool definitions, tool results and fetched documents.
Does a language model remember previous messages?
No. The model keeps nothing between calls. A chat product seems to remember because its software keeps the transcript and sends it, or a trimmed or summarized version of it, with every new message. In your own program you have to do the same: a second call made without the first exchange knows nothing about it.
Does the context window include the model's answer?
Yes. Input and output share one token budget: the instructions, the conversation, tool definitions, tool results, documents and the answer being written all count against it. Many models also cap the length of a single response separately, which is the usual reason an answer stops in the middle of a sentence.
What happens when the context window is full?
One of three things, depending on the software in between: the request is rejected with an error, the oldest part of the input is dropped until it fits, or the history is replaced by a summary. Truncation is often silent and a summary can drop a detail without saying so, so when a long session starts forgetting, check for overflow first.
Why did the model forget something when the context window was not full?
Either the text was not in the request, or it was there and went unused. Print the failing request and search for the fact. If it is missing, the software never included it, or a trim rule or a summary removed it. If it is there, attention is the likely cause: in the studies published in 2023 and 2024, models used a given fact less well as the text around it grew, well before the window filled.

Sources

  1. Simon Willison (2025). Here's how I use LLMs to help me write code
  2. Anthropic (2025). Effective context engineering for AI agents
  3. OpenAI (2026). Conversation state (API documentation, undated; read October 2026)
  4. Anthropic (2026). Working with the Messages API (API documentation, undated; read October 2026)
  5. Anthropic (2026). Context windows (API documentation, undated; read October 2026)
  6. OpenAI (2026). Latency optimization (API documentation, undated; read October 2026)
  7. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang (2023). Lost in the Middle: How Language Models Use Long Contexts (TACL 2024)
  8. Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, Boris Ginsburg (2024). RULER: What's the Real Context Size of Your Long-Context Language Models? (COLM 2024)