Home / Blog / Context engineering and memory / Context Engineering vs Prompt Engineering: What…

Context engineering and memory

Context Engineering vs Prompt Engineering: What Changed

Context engineering vs prompt engineering: what is a rebrand, what changed with agents, and three tests that separate a prompt failure from a context one.

By Enrique Gutiérrez · Published · 22 min read

Context engineering vs prompt engineering: the name is new and partly a rebrand, by its promoters’ own account, and the work is older than the name. What changed is the system. In an agent nobody types the final prompt, so the text a person writes becomes a small share of what the model reads, and code decides the rest.

Prompt engineering is writing the part of a model’s input that a person writes: instructions, examples, the output contract. Context engineering is deciding, in code and on every call, everything the model reads. I wrote this for an engineer whose agent is fine for ten turns and wrong by turn forty, and who has been told to improve the prompt. By the end you should know what to check first and which artifact to change.

Is context engineering just prompt engineering with a new name?

Partly, yes. The label was promoted as a replacement for a term that had lost its meaning, and the people who promoted it say so. The job the label points at is real, and it changed hands when agents arrived.

Simon Willison, whose June 2025 post helped the term spread, conceded the point in a comment four days later: “They arguably are two names for the same thing. The difference is that ‘prompt engineering’ as a term has failed, because to a lot of people the inferred definition is ‘a laughably pretentious term for typing text into a chatbot’”.

The skeptics deserve a fair hearing, because they are right twice. One commenter wrote in July 2025 that “rebranding ‘prompt engineering’ as ‘context engineering’ and pretending it’s anything different is ignorant at best and destructively dumb at worst.” Another, on 30 June: “I mean yes, duh, relevant context matters. This is why so much effort was put into things like RAG, vector DBs, prompt synthesis, etc. over the years.” At the API boundary there is one input, and retrieval pipelines were shipping long before anyone renamed them.

On substance, the promoters’ claim is about systems. Harrison Chase’s essay of June 2025, on an agent-framework company’s blog, defines the term as “building dynamic systems to provide the right information and tools in the right format such that the LLM can plausibly accomplish the task.”

The term was popularized in mid-2025; who coined it I could not establish. The phrase is in the April 2025 revisions of the public guide 12-Factor Agents, the earliest dated use I verified, which is not proof of coinage. The essay quoted above comes from a company that sells agent software, and this site sells a book with a chapter on the term. So I won’t defend the label; what separates the two crafts is who writes the text and when.

What changed between a single prompt and an agent?

Authorship changed. In a single-call feature a person writes nearly every token the model reads, once, before deployment. In an agent, tools, retrieval and the model’s own earlier output produce most of the tokens at run time, and code decides which of them stay.

Chapter 7 of AI Agents, Engineered puts it in one sentence in its section “Context Engineering versus Prompt Engineering”: “In an agent, no one types the final prompt.” It then freezes a run mid-step and counts what is on the desk, the book’s name for the context window: standing instructions, the user’s request, the conversation so far, memory, retrieved documents, tool definitions, tool results, worked examples and the output shape. Nine components, and then the line that settles the comparison: “Prompt engineering, in this inventory, is the craft of two or three entries.”

The anatomy of an assembled context.
Figure 7.1 The anatomy of an assembled context. What the model reads on a mid-run pass is a single stacked window, but its slices have four different authors: you write the standing instructions and the required output shape, the user writes one request, the model writes the running conversation, and the harness fills the rest—memory, retrieved documents, tool definitions, and the often-dominant tool results. Prompt engineering owns the accented slice; context engineering answers for the whole desk. Reuse this diagram

The figure draws eight of those slices and four authors, and accents one slice, the standing instructions. I count nine components and three hand-written entries because I follow the chapter’s text, which adds worked examples and says “two or three”.

The same section names three differences. A prompt is one component and context is all of them. A prompt is static, while context “is assembled fresh on every call by code that runs before the model does.” And the work never finishes: “A prompt is edited. A context is curated, and curation is a loop.” The short definition of what context engineering is follows from that: deciding what fills the window at every step of a run.

Context engineering vs prompt engineering, side by side

Side by side, context engineering vs prompt engineering differ on seven points: the unit of work, what you control, who writes the text, how failure looks, how you test, what you version and who owns it. The first three rows follow Chapter 7. The last four are this post’s position, and no source I read states them this way.

Prompt engineering Context engineering
Unit of work One message or template: instructions, examples, the output contract Everything in the window on one call, for every call: the book’s nine components
What you control Wording, structure, order, examples, format What enters, what is trimmed or summarized, what is fetched on demand, what moves to another window, where each thing sits
Who writes the text A person, before deployment Four authors: you, the user, the model and the harness (Figure 7.1)
Failure signature Wrong on the first call, on a short clean input, every time Fine at step 3 and wrong at step 30; same prompt, different result by session, user or tool output
How you test it Hold the inputs fixed and vary the prompt Hold the prompt fixed and replay recorded windows; remove or move one component; grade late steps apart from early ones
What you version The template and its eval set The assembly code, its caps, and a set of recorded windows to replay
Who owns it Whoever owns the feature’s behavior Whoever owns the harness code that builds each call

The ownership row is the one teams skip. Assembly code is ordinary code, with review and tests, and Chapter 7 asks for exactly that: “The assembled input is a first-class artifact of your system, worth inspecting, versioning, and testing like code”.

How small is the hand-written share of an agent’s window?

In the window below it is 3.4%. This is an illustrative window computed for this post, not a measurement: one call of a support agent at step 17, split into the book’s nine components. Before reading the numbers, guess what share of the tokens somebody on your team typed.

Standing instructions 1,200 tokens; the user’s request 150; conversation so far 9,000; memory 800; retrieved documents 12,000; tool definitions 6,000; tool results 30,000; worked examples 600; output shape 250. The total is 60,000. The parts a person wrote at design time are the instructions, the examples and the output shape: 1,200 + 600 + 250 = 2,050 tokens, and 2,050 / 60,000 = 3.4%.

Tool results alone are 50% of this window, retrieved documents 20%, history 15% and tool definitions 10%. If your team wrote the tool descriptions too, add their 6,000 tokens and the hand-written share is 8,050 of 60,000, or 13.4%.

With JavaScript on, the Context window budget planner runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Context window budget planner on its own page to share a result by link.

The planner above opens on that window. The 200,000-token size, which the tool labels as yours, is my illustrative figure, and so is the growth of 2,500 tokens per step. It shows 30% utilization, 20,000 tokens of headroom under its illustrative 40% ceiling, about 8 steps to that ceiling and about 56 to the hard limit, and it marks tool results, retrieved documents and history for an operation. It does not print the hand-written share: add the Tokens column for instructions, examples and output shape, then divide by the total.

Now carry the run forward. If the hand-written parts stay fixed and the pile grows by 2,500 tokens a step, it crosses the ceiling at step 25, and step 40 holds 60,000 + 23 × 2,500 = 117,500 tokens, 58.75% of the window. The same 2,050 tokens are 1.7% of that. The prompt you have been told to improve is under 2% of what the model reads at the step where it fails.

A share of tokens is not a share of influence: the standing instructions are read on every call and can outweigh their size. The number says how much of the input a prompt edit leaves untouched, nothing more.

Does prompt wording still matter?

Yes, wherever a person still writes the text, and for a single call that is nearly the whole job. Wording stops being enough when the failure depends on what accumulated in the window: at step 40 of the example, a prompt edit leaves the other 115,450 tokens as they were.

The evidence that the hand-written part matters is measured, and for something smaller than wording: formatting. Sclar and colleagues (2023) varied only the format of few-shot prompts on classification and multiple-choice tasks, such as the separator between a field’s label and its text. They found “performance differences of up to 76 accuracy points” between equivalent formats on one open model, and about 10 points on average across more than 50 tasks and several models of that year.

The authors recommend reporting a range across plausible formats when comparing models, and add that for someone building a system a single format that works well enough is a valid choice. I would still score two or three before trusting one.

Where does wording stop being enough?

Wording runs out at two measured points: the number of instructions a model can follow at once, and the way information arrives. Jaroslawicz and three co-authors, in a 2025 preprint, gave 20 models a business report to write under up to 500 one-word keyword-inclusion instructions. The best models reached 68% accuracy at 500 instructions. The two strongest stayed near-perfect to about 150, smaller and older models degraded early, and the bias toward earlier instructions peaked at 150 to 200.

Those rules are far simpler than a policy rule, so I read the counts as an upper bound: a standing-instructions file is a budget too.

Laban and colleagues, in a 2025 preprint, took tasks specified in one complete message and delivered the same information across several turns to 15 models: scores fell by 39% on average over six generation tasks. Given the same pieces concatenated into one message, the models averaged 95.1% of their single-message scores. The task information was the same and only its delivery differed.

The two crafts also interlock. Shi and colleagues (2023) added one irrelevant sentence to grade-school arithmetic problems and saw accuracy fall on two models of that period; adding the instruction “Feel free to ignore irrelevant information given in the questions” won part of it back. That is a wording fix for a context problem.

One gap needs stating. I found no published comparison of prompt tuning against context changes on the same agent, so “context matters more than wording” is what practitioners say. Chapter 7 words it as a report too (practitioners “report, with striking consistency, that most of them are context failures”), and I could find no dataset behind it.

Isn’t this just RAG?

No. Retrieval is one way of filling one of the nine components. RAG (retrieval-augmented generation) fetches passages from an outside corpus into the window at question time, and teams that built it well were doing context engineering before the term existed. The skeptics are right about that.

What the wider term adds is everything retrieval does not touch: tool results, which are half of the window above; the history; memory carried in from earlier sessions; compaction (summarizing a long history into a fresh window); and isolation, where a subagent works in its own window and returns a short report. The difference between RAG and agents is who decides when to retrieve: a pipeline that retrieves once before the call, or a model that asks for more on step twelve.

Two ideas belong to a neighboring post: context rot, the measured decay in how well a model uses a fact as the window around it fills, and the four operations that remedy it (write, select, compress, isolate). Both are in the post on context rot and its four fixes and in a five-minute explainer.

How do you tell a prompt failure from a context failure?

Read the recorded input, then rerun the failing step on a clean, minimal context built from it. A failure that vanishes there comes from the assembled part of the window. A failure that survives is a gap in the hand-written text, a missing input or a model limit, checked in that order.

The clean-desk test is the book’s; Chapter 7, in “How Context Fails”, calls it “the cheapest diagnostic in this book”. Everything around it (the inspection step, tests 2 and 3, the upstream-bug verdict) is this post’s synthesis. The book reads a failure that persists as fighting the weights or asking for more than the model can deliver. I add two cheaper explanations to rule out first, a gap in the wording and a missing fact, so here “persists” does not yet mean the context is innocent.

“Vanishes” and “persists” are rates: run each desk about five times, and treat a difference of one run as none.

Step 0: inspect. Save the failing call’s exact input as the model received it, with the model identifier, the sampling settings and the tool definitions as sent. Then read the inputs the failing step answered from, in your trace.

  • An input came from code that errored, returned nothing where something exists, or returned the wrong record: upstream bug. Stop.
  • Every input is what its code was designed to produce: go to test 1.

Test 1: clean desk. Open a fresh session holding your hand-written entries as shipped, the tools this step uses, and the step’s input: the message it answers and the one record or passage it must read, copied from the recorded window unchanged. Add nothing the window did not hold. I keep the hand-written entries whole on this pass, as the neighboring post’s version of the test does; test 2 strips them.

  • The behavior is about the history (a repeated call, a self-contradiction): keep the fewest earlier turns it refers to, and note how many.
  • The step’s input is longer than your standing instructions: also run a short input of the same kind. If only the long one fails, count that as vanished.
  • The failure vanishes: go to test 3. It persists: go to test 2.

Test 2: stranger. Hand the clean desk’s text to a colleague who has never seen the system.

  • Instruction gap (an ambiguous or contradictory rule; a missing constraint, fallback or output contract): fix the wording, add two or three examples, rerun. Failure gone: prompt failure.
  • Information gap (a fact, record or tool they would need): if the recorded window held it, add it to the desk and repeat test 1. If not, add it by hand and rerun. Failure gone: ask why it was missing. The code meant to supply it errored or has a bug: upstream bug. That code ran as designed, or no such code exists: context failure of the missing kind.
  • Both gaps had to be closed together: prompt failure, with the missing input logged as a second finding.
  • No gap: strip the hand-written entries to the one governing rule and rerun. Failure gone: the other instructions interfered, a prompt failure. Still failing: model limit.

Test 3: ablation. Return to the recorded window and change one assembled component at a time, rerunning each. Trim acted-on tool results, cut the history to the last few turns, drop the passages and tools the step does not use, move the governing rule to the end, delete a suspect turn.

  • One change restores the behavior: it names the component. Its code ran as designed: context failure. Its code errored or returned the wrong thing: upstream bug.
  • Several changes each restore it: a context failure with several contributors.
  • No single change restores it: change two at a time. Still nothing: unresolved. Escalate with the recorded window attached.

What does each verdict tell you to change?

Each verdict names one artifact to repair: a prompt failure the hand-written text, a context failure the assembly code, a model limit the format, request or model, and an upstream bug the tool or function that produced the bad input.

The last branch of test 2 is the book’s second caution turned into a step: “strip the desk fully, because a contradiction baked into your standing instructions or tool definitions rides into every fresh session and counts as context all the same.” The book counts such a contradiction as context; I file it under prompt failure because the repair is an edit to text a person wrote. The pairs in test 3 follow the chapter’s warning about “the interactions between entries, which is where much of the trouble lives.”

What does the diagnostic say about three constructed failures?

It gives three different verdicts, and only one is fixed by editing the prompt. The cases are constructed for this post on the support agent from the worked example, and every run count in the table is illustrative.

Failure (constructed) Step 0 Test 1: clean desk Next test Verdict and repair
A. Step 38: the agent grants a credit above the limit in its standing instructions. It obeyed the limit at step 6. The request and the account record are what their code was built to return. Recorded window: over the limit in 4 of 5 runs. Hand-written entries, the request and the account record: obeys in 5 of 5. Vanished. Test 3. Trimming acted-on tool results restores it, 5 of 5. So does moving the limit rule after the pile, 5 of 5. Both pieces of code ran as designed. Context failure, two contributors. Assembly code: trim results once used; restate the rule last.
B. Turn 1 of any session: asked to cancel, the agent cancels at once. Policy wants the pause option offered first. The only input is the user’s message. Recorded window: cancels in 5 of 5. Hand-written entries and the message: cancels in 5 of 5. Persists. Test 2. A stranger reading “Handle cancellations helpfully” cannot know about the pause option: an instruction gap. Rewritten with the rule and two examples: offers the pause in 5 of 5. Prompt failure. The standing instructions, plus one new case in the prompt’s eval set.
C. Step 12: “I can’t find an order under that email.” The order exists. The recorded lookup result is an empty list, and the query went out with the wrong account filter. The code returned nothing where a record exists. Not run. None. Upstream bug. The lookup tool.

Case A is the engineer this post is for: the rule works alone, so rewording it changes nothing, and both repairs live in harness code. In case B no amount of context work finds a rule nobody wrote down. Case C is neither craft: the model answered sensibly from a wrong tool result, and reading the recorded input was enough to see it.

For a model limit, Chapter 7 has the example: “a data-extraction pipeline kept failing its tests until the requested output format was switched to the one the model’s training plainly favored”.

How is each one tested?

Prompt tests hold the inputs fixed and vary the wording; context tests hold the wording fixed and vary the assembly. For the hand-written entries, keep a fixed set of cases and rerun it whenever the model changes.

For the assembly, three tests carry most of the weight, beside the ablation of test 3; the selection is my synthesis. Replay stored windows from production after every change to the code that builds them, running the same graded check before and after; a case that passed and now fails is a regression. Grade by step, at an early, a middle and a late step, because an average over steps hides the slope. Assert a budget, failing the build when a fixture’s window passes a token limit you chose.

Both kinds need several trials per case, and the post on how many eval examples you need has the arithmetic. When the output at step 40 is open-ended prose, grading it needs a rubric for an LLM judge that a second person would apply the same way.

What do you do differently on Monday?

It depends on the kind of system you run, and each kind inherits the changes of the simpler ones. A tool loop is any run where the model calls tools and reads the results; a long session is one that runs long enough to be compacted, reset or resumed. Pick yours; the rows that remain are the ones that apply. If your loop does no retrieval, skip the two rows about passages.

Change Why Where it lives System
Version the prompt template and rerun its eval set after every model change Prompts drift when the model behind them changes (ch. 2) Template file, eval set single call, RAG, tool loop, long session, multi-agent
Score two or three plausible formats and phrasings before trusting one Equivalent formats can score far apart (Sclar et al., 2023) Eval report single call, RAG, tool loop, long session, multi-agent
Put the instruction after any long supplied material and restate the load-bearing rule last “The last thing the model reads should be the thing you most need it to do.” (ch. 2) Template file single call, RAG, tool loop, long session, multi-agent
Log the assembled input of every call, with the model and its settings The template alone cannot reproduce a result Calling code RAG, tool loop, long session, multi-agent
Score retrieval apart from the answer: did the passage holding the answer arrive? Separates an information gap from an instruction gap Eval set RAG, tool loop, long session, multi-agent
Cap passages per call and keep near-miss passages in the eval set “material that is present but irrelevant always costs something” (ch. 7) Retrieval code, eval set RAG, tool loop, long session, multi-agent
Replay recorded windows after every change to assembly code A change to trimming or ordering can break a case no prompt test covers CI RAG, tool loop, long session, multi-agent
Trim each tool result to a one-line fact and a reference once it has been acted on “the often-dominant tool results” (Figure 7.1, caption) Harness code tool loop, long session, multi-agent
Give each task only the tools it needs A small model failed a task with 46 tools in view and passed with the 19 relevant ones (ch. 7) Tool configuration tool loop, long session, multi-agent
Grade the same check at an early, a middle and a late step An average over steps hides the slope Eval harness tool loop, long session, multi-agent
Have the agent keep its plan and decisions in a file it rereads A plan that lives only in the transcript is lost when the transcript is compacted or reset (ch. 7) Harness code, one file long session, multi-agent
Assert a token budget per component and fail the build past it Growth is silent until quality drops CI long session, multi-agent
Write each delegation brief from a template that carries the parent’s constraints A subagent knows only what its brief says Brief template multi-agent
Fix the size and shape of what a subagent returns The return sits on the parent’s desk for every later step Subagent output contract multi-agent

Select single call and three rows remain, all of them about the hand-written text and its tests. For a single call with complete inputs, prompt craft is still the work. The two rows that only multi-agent adds are prompt craft as well, since a brief and an output contract are text somebody writes.

What goes in a context manifest?

A context manifest is one entry per component of one recorded call, and each entry answers eight questions: who wrote it, how many tokens, static or assembled at run time, where it sits, which code builds it, who owns it, which test covers it, and what caps it when it grows. Fill one in this week for one mid-run call of your own system.

CONTEXT MANIFEST: [agent or feature], step [N] of run [id], [date]
Model: [id] · sampling: [settings] · tokens counted with: [tokenizer or estimate]
Window: [size] tokens · filled: [total] tokens · utilization: [total / size]

Eight fields per component:
author (you | user | model | harness) · tokens · static or assembled ·
position (first | middle | last) · built by (function or module) ·
owner · test · cap (what trims it when it grows)

1. Standing instructions
   author: [ ] · tokens: [ ] · [static | assembled] · position: [ ]
   built by: [ ] · owner: [ ] · test: [ ] · cap: [ ]
2. User's request
   author: [ ] · tokens: [ ] · [static | assembled] · position: [ ]
   built by: [ ] · owner: [ ] · test: [ ] · cap: [ ]
3. Conversation so far
   author: [ ] · tokens: [ ] · [static | assembled] · position: [ ]
   built by: [ ] · owner: [ ] · test: [ ] · cap: [ ]
4. Memory carried in
   author: [ ] · tokens: [ ] · [static | assembled] · position: [ ]
   built by: [ ] · owner: [ ] · test: [ ] · cap: [ ]
5. Retrieved documents (passages: [ ])
   author: [ ] · tokens: [ ] · [static | assembled] · position: [ ]
   built by: [ ] · owner: [ ] · test: [ ] · cap: [ ]
6. Tool definitions (tools: [ ])
   author: [ ] · tokens: [ ] · [static | assembled] · position: [ ]
   built by: [ ] · owner: [ ] · test: [ ] · cap: [ ]
7. Tool results
   author: [ ] · tokens: [ ] · [static | assembled] · position: [ ]
   built by: [ ] · owner: [ ] · test: [ ] · cap: [ ]
8. Worked examples
   author: [ ] · tokens: [ ] · [static | assembled] · position: [ ]
   built by: [ ] · owner: [ ] · test: [ ] · cap: [ ]
9. Output shape
   author: [ ] · tokens: [ ] · [static | assembled] · position: [ ]
   built by: [ ] · owner: [ ] · test: [ ] · cap: [ ]

Hand-written share: (tokens of every component whose author is "you") / total = [ ]%
Largest component: [ ] at [ ]% of the total · it grows by about [ ] tokens per step
Components with no owner: [ ] · with no test: [ ] · with no cap: [ ]

The author list is the figure’s four, so a tool’s output goes under the harness. Where a component has several authors, as the conversation does, name the one that writes most of it. Count tool descriptions, summarization instructions and subagent briefs as yours when your team wrote them; that is the 13.4% reading of the worked example.

The last three lines are the point of the exercise. Look first at the largest component: if it has no owner, no test and no cap, it is the first thing to fix. If you later work from a context engineering checklist, it is only as good as the inventory under it. The context window budget planner takes the same nine token counts.

Where does this comparison stop being useful?

It stops at the standing instructions, which are both a prompt and a component of the context, and at any system whose proportions differ from my example. A chat feature with no tools is mostly hand-written and user-written text, and the 3.4% is arithmetic on a window I made up.

Running the diagnostic has costs. Count on about five runs of the recorded window, five on the clean desk and five per ablation; with six ablations that is forty runs, most of them re-reading the full window, before any reruns in test 2. The counts are illustrative; yours depend on how often the failure shows. A failure that appears once in fifty runs will not show in five on either desk.

Every measurement here comes from models of 2023 to 2025, and thresholds move with each generation. Curation has its own failure, which Chapter 7 names: “The trade is always recall against focus”. And the claim that context outweighs wording remains unmeasured on any agent I could find. The vocabulary may move again; the inventory and the three tests do not depend on it.

The question to ask before you edit the prompt

Whatever you decide about context engineering vs prompt engineering as labels, keep one habit: before you reword anything, look at what the model read. Chapter 7 gives it as a standing question: “what exactly was on the desk when it failed?” Dump one failing call, read it, run it on a clean desk, and you will know whether the prompt is the first thing to change.

Chapter 7, “Managing the Context Window,” has the nine components, the five ways a context fails and the four operations (in the full book). Chapter 2 is free and holds the prompting fundamentals this post leans on. The context engineering guide places this post beside its neighbors, or you can see the formats.

Questions readers ask

Is context engineering just prompt engineering rebranded?
The name is, in part. Simon Willison, who helped spread it, wrote in 2025 that the two are arguably two names for the same thing and that the older term had come to mean typing into a chatbot. The work differs in one respect: in an agent most of the input is assembled by code on every call, so it has a different owner, a different failure signature and different tests.
Is prompt engineering dead?
No. Instructions, examples and output contracts are still written by hand, and a 2023 study measured gaps of up to 76 accuracy points between equivalent prompt formats on one open model. Every tool description, summarization instruction and subagent brief is also a prompt somebody wrote. Wording stops being enough when the failure depends on what accumulated in the window.
Is context engineering the same as RAG?
No. Retrieval-augmented generation fills one component of the context window, the retrieved documents. Context engineering also decides what happens to tool results, conversation history, memory and tool definitions, and it keeps deciding on every step of a run.
How do I tell a prompt problem from a context problem?
Save the exact input of the failing call and read it; a broken tool result or empty retrieval is an ordinary bug. Otherwise rerun the step several times on a clean, minimal context built only from that recorded input. If the failure vanishes, the assembled part of the window caused it: change one component at a time to find which. If it persists, look for a gap in the wording, then a missing fact, then a model limit.
How do you test context engineering?
Hold the prompt fixed and vary the assembly. Replay recorded windows after every change to the code that builds them, remove or move one component at a time, grade the same check at an early, a middle and a late step, and assert a token budget per component. Prompt tests do the reverse: fixed inputs, varied wording.

Sources

  1. Simon Willison (2025). Context engineering
  2. Simon Willison (2025). Hacker News comment: two names for the same thing
  3. FiniteIntegral (Hacker News) (2025). Hacker News comment on the rebrand
  4. alfalfasprout (Hacker News) (2025). Hacker News comment on retrieval predating the term
  5. HumanLayer and contributors (2025). 12-Factor Agents, Factor 3: Own your context window
  6. Harrison Chase (LangChain) (2025). The rise of "context engineering"
  7. Melanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane Suhr (2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design (ICLR 2024)
  8. Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Chi, Nathanael Schärli, Denny Zhou (2023). Large Language Models Can Be Easily Distracted by Irrelevant Context (ICML 2023)
  9. Daniel Jaroslawicz, Brendan Whiting, Parth Shah, Karime Maamari (2025). How Many Instructions Can LLMs Follow at Once? (preprint)
  10. Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer Neville (2025). LLMs Get Lost In Multi-Turn Conversation (preprint)