The RAG vs long context choice is settled by two questions before accuracy enters: may every asker read every document, and does what they may read fit comfortably in the window? If both answers are yes, the default is to paste the corpus in and read it whole, unless one of five named reasons applies. If either is no, retrieval is in the design, alone or in a hybrid.
I wrote this for a tech lead with a design review this week. Someone has said the window is big enough now, so skip the pipeline; someone else has said “we need RAG.” This post gives an ordered decision table, the arithmetic of a cached corpus against a retrieved slice, and a small test that can overturn the choice. Why a long context degrades at all belongs to the post on what context rot is and how to fix it, and I won’t re-derive it here.
What are you actually choosing between?
You are choosing where the selection happens. With the long window, nobody selects: the whole corpus is carried in the context window on every request and the model finds what it needs. With RAG (retrieval-augmented generation), your code searches the corpus first and puts only the matching passages in front of the model.
Chapter 8 of the book, in the section “RAG versus Long Context versus Fine-Tuning,” sets the two beside a third option: “There are three ways to give a model knowledge it lacks—carry it in the window, fetch it into the window, or train it into the weights”. The book’s figure draws all three.
The third road leaves this comparison quickly. The chapter’s rule is “Retrieval for what the model should know; fine-tuning for how it should respond; the long window for close reading of a bounded set.” Fine-tuning changes behavior well and knowledge badly, since a fact in the weights is frozen at the last training run and cannot be cited. A knowledge question is therefore a choice between the first two roads.
The book is fair to the long window, and so should a design review be. Of pasting a bounded corpus in whole, it says: “for close reading it is the best one,” because “you cannot fail to fetch what is already on the desk.” Then the sentence that carries the trade: “Retrieval hands the model fragments; the long window hands it the book.”
Why isn’t the long window the default for everything?
Because a full read is paid for on every request in three currencies, and none of them is refunded by a bigger window: money, waiting and attention. The book names the first two in one line, “Every request pays for the entire corpus, in money and in waiting,” and points to the previous chapter for the third.
Attention is the one a window size hides. Context rot is the tendency of a model to use any given fact less well as the text around it grows. In a 2025 report, Hong, Troynikov and Huber at Chroma evaluated 18 models with the task held fixed and concluded: “Across all experiments, model performance consistently degrades with increasing input length.”
The makers of long windows say something compatible. One provider’s developer guide to long context (last updated June 2026) notes that its headline recall tests look for one fact, and adds that when there are several pieces of information to find, “the model does not perform with the same accuracy.”
So the book draws the border in a place no vendor’s specification sheet shows: “The long window’s territory is close reading of a small, bounded set, and its border is the soft limit, measured on your task.” The soft limit is the input length at which your own questions start to be answered worse. It is lower than the advertised window, it differs by model and by task, and you find it by testing.
What do the published comparisons of RAG vs long context say?
They disagree on accuracy, and the one that priced both found retrieval far cheaper. Two studies from 2024 found the long window ahead on average; a 2025 benchmark found the sign depends on input length and question type. All of them date quickly, since each tested the models of its year, so read them for the shape of the result and not for the scores.
Li and colleagues (2024) benchmarked both designs and reported that “when resourced sufficiently, LC consistently outperforms RAG in terms of average performance. However, RAG’s significantly lower cost remains a distinct advantage.” Their analysis adds a detail that matters for routing: “for 63% queries, the model predictions are exactly identical.” For most questions in their data, the expensive read bought nothing.
A second evaluation from the same year (Li, Cao, Ma and Sun) agrees on the average and splits it by method: “Summarization-based retrieval performs comparably to LC, while chunk-based retrieval lags behind.” It also found that “RAG has advantages in dialogue-based and general question queries.”
The LaRA benchmark (Li et al., 2025), with 2,326 test cases across 11 models, is subtitled “No Silver Bullet,” and reports a crossover: “With a 32k context length, LC achieved an average accuracy 2.4% higher than RAG across all models. However, with a 128k context length, this trend reversed, with RAG outperforming LC by 3.68%.” By task: “RAG demonstrates similar performance to LC in single-location tasks and offers a significant advantage in identifying hallucinations. In contrast, LC excels in reasoning tasks and comparison tasks.”
I take two things from these. Question type predicts the winner better than any single benchmark average does: lookups are a tie that retrieval wins on price, and comparison across a document favors reading it whole. And since the averages flip with length and model, no published number can stand in for a test on your own questions.
How do you choose? The decision table
For RAG vs long context, apply two gates in order, then read down the table and take the first row that describes your use case. The gates and the row order are this post’s own arrangement; the book supplies the default order and the four conditions that call for retrieval.
Gate A, permissions. If different askers may read different documents, code decides what each one’s request can contain, before any model sees it. The model is never the one enforcing the rule, because anything in the window can end up in the answer.
Gate B, fit. Take what one asker may read, add room for the question, the conversation so far and the answer, and ask whether the total sits under the soft limit you measured. “Fits” in the table means this, never the advertised window.
When both gates pass, the long window is the default. The book’s reason is “The long window while the corpus is bounded, because infrastructure you do not build cannot break.” Rows 6 to 10 are the named reasons to leave it, and row 11 is what remains when none applies.
The book’s own trigger list is shorter: “Retrieval, the road this chapter built, becomes the default the moment any one of four conditions holds: the corpus outgrows the desk; it changes faster than you care to re-paste; it is private and must remain so; or the answers need receipts.” Reading “private” as per-asker permissions is my extension. The buttons filter the table by where each row lands.
| # | Your use case | Design | Why | What to check first | Verdict |
|---|---|---|---|---|---|
| 1 | Askers have different permissions; what any one asker may read fits | Filter by permission in code, then read the remainder whole | The filter is deterministic; the reading keeps the whole-document view | A test asker’s request contains no document they may not open | hybrid |
| 2 | Does not fit; each answer lives in one passage (lookups) | Plain retrieval: search, take the top few passages, answer from them | The price of a question stops depending on corpus size | The needed passage is among those retrieved, on a sample of real questions | RAG |
| 3 | Does not fit; each question concerns one or a few whole documents | Retrieve documents, not fragments, and read each one whole | The search finds the contract; the model compares its clauses | One retrieved document plus the question fits under the soft limit | hybrid |
| 4 | Does not fit; answers join facts from several documents, one leading to the next (multi-hop) | Give an agent the retriever as a tool, with a cap on hops | The second query cannot be written until the first result is read | Cost and latency per question at the hop cap | hybrid |
| 5 | Does not fit; a small stable core is needed by every question, plus a long tail | Core in a cached prefix, tail retrieved | The core is never missed; the tail is paid for only when used | The core stays byte-identical between requests, or the cache misses | hybrid |
| 6 | Fits; you must record which passages each answer was built from | Plain retrieval | The list of retrieved passages is the record; “everything was in the window” is not one | Each answer’s logged passages can be opened by a reviewer | RAG |
| 7 | Fits; a full read misses the latency budget even with a cache | Plain retrieval | A short prompt is read sooner | Time to first token at the 95th percentile, cache cold and warm | RAG |
| 8 | Fits today; growth will break Gate B within your planning horizon | Plain retrieval, built now | Migrating under load is the expensive time to build an index | Corpus size now, growth per quarter, the measured soft limit | RAG |
| 9 | Fits; there are enough lookups for the bill to matter, and the corpus is above the break-even size or traffic is too sparse to keep a cache warm | Plain retrieval | Each question pays for the corpus; see the arithmetic below | F against R ÷ c, and the gap between questions against the cache lifetime | RAG |
| 10 | Fits; lookups and whole-document questions are mixed, and cost matters | Route: answer from retrieved passages first, fall back to the whole read | Most questions get the cheap path; hard ones still get the book | Share of questions the cheap path answers, and how often it answers wrongly without falling back | hybrid |
| 11 | Fits; none of rows 6 to 10 applies | Paste the corpus in, cached, and read it whole | Nothing to index, nothing to miss, nothing to maintain | Accuracy on your own questions at this length | long context |
If Gate A applies to a corpus that also does not fit, use rows 2 to 5 with the permission filter inside the search query, so that a passage the asker may not read is never a candidate. Chapter 8 describes that filter for metadata in general: “The filter prunes the map before the blur gets a vote.”
Freshness, which many comparisons list first, is absent as a row on purpose. A long window assembled from the source at request time is always current. What an edit costs there is a cache miss from the edited position onward, and what it costs retrieval is a re-index of the changed document. Churn therefore shows up in row 9, as a cache that keeps going cold.
Row 4 is where this comparison meets a different one. Putting the retriever inside a loop is agentic retrieval, and whether the loop is worth its cost is the subject of RAG vs agents. The book’s bearing for it is “the simplest retrieval that works.”
What does each design cost per question?
On cost, RAG vs long context comes down to one comparison: a full read costs the size of the corpus on every question, and a cache changes the rate without changing the count. Prompt caching bills a repeated prefix at a fraction of the fresh rate. Call that fraction c, call the corpus F tokens and the retrieved slice R tokens. Per question, in steady state, the cached window costs about c × F and retrieval costs R.
In steady state, the cached window is the cheaper one roughly while F is below R ÷ c.
That inequality is this post’s own approximation for one steady-state question. It ignores the first uncached read, which lowers the break-even in a short sitting (to about 30,000 tokens in the ten-question example below). It ignores the conversation history, which both designs read in equal amounts but which only the cached side gets at a discount, so a very long sitting moves the break-even up again. It also ignores output tokens and the cost of running an index.
A worked example, with every number illustrative: a corpus of 80,000 tokens behind a 500-token system prompt, and a user who asks ten questions in one sitting, each adding 400 tokens of question and answer to the conversation. Retrieval fetches five passages of 800 tokens, so 4,000 per question, and discards them after answering.
| Design | Input tokens read over ten questions | Billed, in fresh-token equivalents | Against retrieval |
|---|---|---|---|
| Long window, no cache | 823,000 | 823,000 | 13.1× |
| Long window, cached at an illustrative 10% | 823,000 (738,900 from cache) | 157,990 | 2.5× |
| Retrieval, 4,000 tokens per question | 63,000 | 63,000 | 1× |
The retrieval total is 10 × 4,500 for the system prompt and passages, plus 18,000 of accumulated conversation. The long-window total is 10 × 80,500 plus the same 18,000. With c at 10% and R at 4,000, the break-even corpus is 4,000 ÷ 0.10 = 40,000 tokens, and this corpus is twice that.
The prompt caching savings calculator is mounted below on the long-window row. It was built for an agent’s run, so read “step” as “question,” and note that the corpus sits in its third prefix block, labeled for instructions; the mapping is mine. It opens on 90% from cache and 157,990 token-equivalents. Change the factor to your provider’s and the corpus to yours.
With JavaScript on, the Prompt caching savings runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Prompt caching savings on its own page to share a result by link.
What happens when the questions are spread out?
The cache lapses and every question pays for the corpus at the fresh rate. The calculator assumes each cached prefix is still stored when the next call arrives, which holds inside a sitting and fails across a quiet afternoon.
Cache lifetimes are short and provider-specific. As dated examples only, one provider’s documentation (read October 2026) says “By default, the cache has a 5-minute lifetime,” and another’s gives 30 minutes as the default for its current models and says that, under one setting for earlier ones, entries “typically remain active for around 5 to 10 minutes of inactivity, up to one hour.” Both document longer retention options.
Take the same ten questions as ten independent 100-token requests, each arriving after the cache has lapsed. The long window reads 10 × 80,600 = 806,000 tokens, all fresh. Retrieval reads 10 × 4,600 = 46,000. The ratio is 17.5 to 1, worse than the uncached sitting, because nothing is shared.
A cached corpus is also exactly as large in the window as an uncached one. The discount applies to the bill and to the wait. The attention problem of the previous section is untouched, which is why row 11 still tells you to check accuracy at that length. For the other levers on the bill, see the guide on how to reduce AI agent costs.
Where does a hybrid win?
A hybrid wins when one design’s failure is the other’s strength, which is the case in five of the eleven rows. The common shape is that code narrows what the model may or must read, and the model then reads that narrower set whole.
Filter, then read whole (row 1). A handbook with manager-only sections is two bounded corpora. Code picks one by role. Nothing needs ranking, and each audience gets the close reading the long window is good at.
Retrieve documents, read them whole (row 3). Two hundred contracts do not fit, and one does. Search for the contract, then hand over all of it. Chapter 8 has the same move at a smaller scale: “Match on the sentence; read the page.”
Route (row 10). Li and colleagues (2024) tested a version they call Self-Route. The model first sees the retrieved passages with the option to decline, through the instruction “Write unanswerable if the query can not be answered based on the provided text,” and only declined questions go to the full read. They report cost “reduced by 65%” for one model and 39% for another, with performance comparable to the long window. This is the routing pattern, staffed by the model’s own judgment of its passages.
Cached core plus retrieved tail (row 5). Put the glossary, the policy summary and the product list that every question needs in the stable prefix. Retrieve everything else. Order matters: the stable part goes first, or each retrieved passage breaks the cache in front of it.
Retriever as a tool (row 4). For multi-hop questions, no single search works. The same 2024 study lists this first among retrieval’s failure reasons: “The query requires multi-step reasoning so the results of previous steps are needed to retrieve information for later steps.”
One hybrid does not win: retrieving far more passages because the window has room. Jin and colleagues (2024) found that “for many long-context LLMs, the quality of generated output initially improves first, but then subsequently declines as the number of retrieved passages increases.” The book’s version is “More retrieved context is not better context.”
Which design is easier to debug?
Retrieval is easier to debug, because its failures separate into stages you can inspect, and that advantage is easy to overstate. A wrong answer from a retrieval system has three possible locations. The passage was never indexed, or it was indexed and not retrieved, or it was retrieved and misread. The first two are visible in a log without reading any model output.
A wrong answer from a long window has one location: the text was present and went unused. The available fixes are few. You can move the relevant material, ask for supporting quotations before the answer, or shorten the corpus, and the last of those is retrieval by another name.
The receipt deserves the same care. The book ranks provenance first among retrieval’s gifts, “a receipt the bare model has no way to print,” and warns about its misuse: “A wrong answer with a citation is more dangerous than a wrong answer without one, because the citation buys trust the answer has not earned.” A logged passage shows what the model was given. It does not show that the answer follows from it.
On a small corpus there is a partial substitute. Ask the long-window answer to quote its support word for word and check each quotation against the corpus with a string match. That proves the sentence exists. It gives you no list of what was considered, which is why row 6 still points to retrieval when an auditor needs that list.
How do you test the choice before you commit?
Settle RAG vs long context for your corpus by running both designs on the same small set of real questions, labeled by type, and comparing by type. The test is cheap when the corpus fits, because the long-window arm needs no infrastructure and the retrieval arm can be the plain pipeline the book says is “an afternoon of work.”
The procedure is my suggestion, and thirty questions is a starting size, not a statistical threshold:
- Collect 30 questions real users asked, or would ask, with the answer and the passage that supports it.
- Label each one: lookup (one passage), whole-document (compares or connects parts), multi-hop (several documents in sequence), or unanswerable from the corpus.
- Run the long window at the real corpus size, and at half of it, to see whether accuracy moves with length.
- Run plain retrieval, and record for every question whether the supporting passage was among those retrieved.
- Score each arm per label. Count a confident answer to an unanswerable question as a failure.
- Record input tokens and time to first token per question for both arms, cache cold and warm.
- Read the misses. A retrieval miss on a lookup is a pipeline fault to fix; a miss on a whole-document question is the design’s limit.
Thirty questions can show a large gap and cannot show a small one. The eval sample size calculator says how many you need before a difference of a given size means anything.
Then write the decision down in a form the next person can overturn:
KNOWLEDGE ACCESS DECISION: <feature> Date: <date>
Corpus: <what it is>, <size in tokens today>, growing <amount per quarter>.
Gate A, permissions: <one audience | N audiences, filtered in code by ...>.
Gate B, fit: largest share one asker may read = <tokens>;
soft limit measured on our questions = <tokens>; fits: <yes | no>.
Question mix (from <n> labeled questions):
lookup <n>, whole-document <n>, multi-hop <n>, unanswerable <n>.
Decision: <long context | RAG | hybrid: which one>, by row <#> of the table.
Cost per question (input tokens): long window <F>, cached at factor <c> = <c x F>;
retrieval <R>. Break-even corpus R / c = <tokens>.
Typical gap between questions <time> against cache lifetime <time>.
Test result by label: long window <score per label>; retrieval <score per label>;
retrieval found the supporting passage in <n of n> questions.
We revisit when: the corpus passes <tokens>; a second audience appears;
the question mix shifts; the provider's cache terms or the model change.
Where does this advice stop?
A RAG vs long context table stops at the edges of what it can know about your system. Four limits are worth stating.
The soft limit is yours to measure, and it moves when the model changes. A 2024 study of 20 models by Leng and colleagues varied context from 2,000 to 128,000 tokens and found that only a handful held consistent accuracy above 64,000. Those models have been superseded, and the finding that the limit differs by model has not.
Vendor thresholds are dated advice about one product. A lab’s 2024 post said that below 200,000 tokens, “about 500 pages of material,” a team “can just include the entire knowledge base in the prompt.” That is a statement about fit in that year. It says nothing about your permissions, your traffic or your question mix.
The cost arithmetic counts input tokens only. It leaves out the engineer who maintains the index, the embedding and storage bills, and output tokens, which are the same in both designs. For a corpus near the break-even size, those omitted terms can decide it, in favor of the design with nothing to operate.
And this post is about a corpus someone else wrote. What an agent learns during its own runs, and how that is stored and fetched, uses the same machinery pointed at a different source; that is the subject of agent memory architecture.
The sentence for the design doc
Decide permissions in code, measure whether the remainder fits, and then let the question mix and the traffic pick between reading it whole and fetching a slice. The RAG vs long context argument in most meetings is about window size, which is the one input that settles none of those.
Whichever row you land on, keep the thirty questions. They are the signal that tells you when the choice has stopped being right.
Chapter 8, “Retrieval and Knowledge,” builds the pipeline, the levers and the three-way decision (in the full book). The context engineering guide places this post beside its neighbors, the explainer on context rot and the four operations animates why the window is scarce, or you can see the formats.
Questions readers ask
- Is RAG obsolete now that context windows are large?
- No. A larger window raises how much text fits. It does not decide who may read which document, it does not make a full read cheap, and it does not stop accuracy from sloping down as input grows. Retrieval stays in the design whenever the corpus does not fit, askers have different permissions, or lookups arrive at a volume where reading everything costs too much.
- When is long context better than RAG?
- When the corpus is bounded, every asker may read all of it, it sits comfortably under the soft limit measured on the task, and the questions need the document read whole: comparing one clause with another, following a cross-reference, noticing that one section amends another. Retrieval hands the model fragments and can miss the one that matters.
- Does prompt caching make long context as cheap as RAG?
- Only for a small corpus under steady traffic. A cached token is still billed at some fraction of a fresh one, so a cached corpus of F tokens costs about that fraction times F per question. With an illustrative factor of 10% and 4,000 retrieved tokens per question, in steady state the cached window is cheaper roughly below 40,000 tokens, and only while questions arrive before the cache lapses.
- Can you combine RAG and long context?
- Yes, in at least five ways: filter by permission in code and read the remainder whole; retrieve whole documents instead of fragments; answer from retrieved passages first and fall back to a whole read when they do not contain the answer; keep a stable core in a cached prefix and retrieve the long tail; or give an agent the retriever as a tool for questions that need several hops.
- Where does fine-tuning fit in the choice?
- Mostly outside it. The book's rule is retrieval for what the model should know, fine-tuning for how it should respond, and the long window for close reading of a bounded set. Facts trained into the weights are frozen at the last training run and cannot be cited, so fine-tuning is for stable behavior at volume, never for a knowledge base.
Sources
- Zhuowan Li, Cheng Li, Mingyang Zhang, Qiaozhu Mei, Michael Bendersky (2024). Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach
- Kuan Li, Liwen Zhang, Yong Jiang, Pengjun Xie, Fei Huang, Shuai Wang, Minhao Cheng (2025). LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs -- No Silver Bullet for LC or RAG Routing
- Xinze Li, Yixin Cao, Yubo Ma, Aixin Sun (2024). Long Context vs. RAG for LLMs: An Evaluation and Revisits
- Bowen Jin, Jinsung Yoon, Jiawei Han, Sercan O. Arik (2024). Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG
- Quinn Leng, Jacob Portes, Sam Havens, Matei Zaharia, Michael Carbin (2024). Long Context RAG Performance of Large Language Models
- Kelly Hong, Anton Troynikov, Jeff Huber (Chroma) (2025). Context Rot: How Increasing Input Tokens Impacts LLM Performance
- Anthropic (2024). Introducing Contextual Retrieval
- Google (2026). Long context (developer guide, last updated June 2026; read October 2026)
- Anthropic (2026). Prompt caching (provider documentation, example; read October 2026)
- OpenAI (2026). Prompt caching (provider documentation, example; read October 2026)