RAG vs agents compares two things that sit at different levels of a product. RAG, short for retrieval-augmented generation, is a way to put the right documents in front of a model before it answers. An agent is a way of deciding what happens next. A product can use one, both or neither, and each combination has its own bill.
This post is for the founder or product manager who heard “we need RAG” and “we need an agent” in the same meeting, offered as alternatives. By the end you can say which of three designs your knowledge feature needs, and what each costs in waiting time, money, testing and failure. You also leave with five questions that make a team or a vendor show which design they built.
RAG vs agents: what does each word mean?
RAG names a source of knowledge, and agent names an owner of decisions. Those are answers to two separate questions about the same product, which is why a debate between them rarely settles.
What is RAG?
RAG is a pattern in which software searches your documents at the moment a question arrives and hands the best passages to the model with the question. The book’s glossary, which is free online, defines it as “The pattern behind most systems that answer from your documents: at request time, retrieve the most relevant chunks from an index and place them in the window beside the question”. A chunk is a passage cut from a document. The window is the context window: everything the model is given to read on one call.
The word has drifted, and the drift feeds the confusion. The paper that coined it, by Lewis and colleagues (2020), described “models which combine pre-trained parametric and non-parametric memory for language generation.” In plain terms, that is a model’s trained knowledge joined to a searchable store outside it. Their store was “a dense vector index of Wikipedia”, a search index built on numerical fingerprints of meaning.
So some engineers reserve RAG for that one search technique, and others use it for any fetch at all. One forum commenter put the result this way: “people seem to be talking about wholly different things?” (schmorptron, Hacker News, 12 October 2025). Another asked, “Isn’t tool use just (more sophisticated) RAG?” and drew two replies that contradicted each other (madeofpalk, 2 October 2025).
In this post, RAG means the pattern: fetched text is placed in the window so the model answers from it. How the fetch is done (keywords, fingerprints of meaning, a database query) is a detail one level down.
What is an AI agent?
An agent is a model placed in a loop with tools, where a tool is a function the model may ask the surrounding software to run. The model looks at the task, picks an action, reads the result and goes again. Chapter 1 of the book states what is new about it: “In an agent, the sequence is decided by the model, at runtime, in response to what it observes.”
Its opposite number is the workflow: several model calls wired together by your code along a path drawn in advance. The glossary gives the dividing property of an agent as “the model, and no flowchart of yours, decides what happens next.” For the full taxonomy written for product teams, including the chatbot, see the comparison of an AI agent vs a chatbot.
Why are RAG and agents answers to different questions?
RAG and agents differ in kind because one is a part and the other is an arrangement of parts. Retrieval is something you attach to a model call. An agent is one way of organizing model calls over time.
The book builds both from a single unit, the augmented LLM (LLM stands for large language model). Chapter 1 defines it: “The augmented LLM is a model call enhanced with three things: retrieval (pulling relevant information in at request time), tools (functions it can request), and memory (some way of carrying forward what matters).” The glossary calls those three “ports”. The idea comes from a 2024 engineering essay, which says “The basic building block of agentic systems is an LLM enhanced with augmentations such as retrieval, tools, and memory” (Anthropic, 19 December 2024).
From that unit the chapter derives the two shapes. A workflow is several units wired by your code along a fixed path. An agent is one unit placed in a loop, choosing its own action on each pass. The chapter’s phrase for this is “Same atom, two arrangements”.
I find it easier to say this to a product team as two layers, and the word “layer” is this post’s own. The book says atom, port and arrangement. The lower layer is the model call and what feeds it, and retrieval lives there. The upper layer is who decides the next call, your code or the model, and “agent” lives there.
Two questions follow, and every product answers both. Where does the knowledge come from: training, a pasted document, or a search? Who decides the next step: code or the model? RAG vs agents sets an answer to the first question against an answer to the second.
A neighboring post, on context engineering vs prompt engineering, puts the same point in one sentence. The difference is who decides when to retrieve: a pipeline that retrieves once before the call, or a model that asks for more partway through.
What are the three designs, and where does the router fit?
A knowledge feature can be built in three ways: with no retrieval, with a fixed retrieve-then-answer pipeline, or with an agent that calls retrieval as a tool. A fourth shape combines the last two behind a router, and it is the one the book recommends for mixed traffic.
Design 1: no retrieval
No retrieval means the model answers from what is already in its window. There are two cases. In the first, the request needs no outside knowledge: rewrite this paragraph, summarize the text I pasted, add these numbers. In the second, the knowledge is small enough to send whole on every call, such as one handbook or one contract.
This is a real option and often the right one. The book’s ladder of designs puts “a single well-crafted model call, perhaps with retrieval and a few examples in the prompt” one rung above plain code. Whether a body of documents should be sent whole or searched is its own decision, covered in the separate comparison of RAG vs long context.
Design 2: the fixed retrieve-then-answer pipeline
The fixed pipeline searches once and answers once, in an order your code set. A question arrives, code turns it into a search, the top passages go into the window, and the model writes an answer from them.
This is plain RAG, and by the book’s definition it is a workflow. Chapter 8 says of it that “the control flow is fixed in advance”, and then names the price: “If the retrieval went wrong, the answer goes wrong, and no step exists at which anything could notice.”
Engineers can tune this design a great deal: better search, a second pass that reorders results, different passage sizes. The chapter observes what all of that tuning preserves. “The pipeline never reads what it fetched and thinks better of it.”
Design 3: an agent with retrieval as a tool
The third design moves the search inside an agent’s loop. In the chapter’s words, the remedy is to “make the retriever a tool, and let the agent decide when and how to call it”.
Four decisions pass from your code to the model. The model decides whether to search at all, what to search for, where to search when there are several collections, and whether what came back is enough. The book’s preferred name for this is agentic retrieval; the trade name is agentic RAG.
The figure comes from Chapter 8, which is in the full book. The chapter describes what separates the two drawings as “one evaluation, and one arrow pointing back”.
Two labeled examples of this category, each in its maker’s words and dated. One lab’s write-up contrasts “static retrieval” with “a multi-step search that dynamically finds relevant information, adapts to new findings, and analyzes results” (Anthropic, 13 June 2025). A search company’s announcement describes “searching, finding interesting pieces of information and then starting a new search based on what it’s learned” over “a few minutes” (Google, 11 December 2024). Neither is a recommendation.
The routed front door
The routed design puts a sorting step in front of the other two. The book describes it this way: “sort incoming questions, send the easy lookups down the static path and only the hard shapes into the loop”.
Routing is a workflow pattern: a cheap decision step sends each input down the path built for its kind. The sorter is usually a classifier, a small model call that assigns a label. So the routed design is a workflow with a loop on one branch. That matches Chapter 1’s forecast: “The dominant production shape is a workflow shell with one or two genuinely agentic steps inside it”.
Which design fits your product?
Three questions asked in order settle a RAG vs agents decision, and the questions your users really ask are the input. Collect the ten most common ones before the meeting. The questions and the table below are this post’s own arrangement; the book gives the reasoning and no table.
- Question 1: does it fit, or is nothing needed? Some requests need no outside knowledge. In other products, everything the feature must know fits in the window with room to spare, and every asker may read all of it. In either case, choose no retrieval.
- Question 2: can every search be written from the question alone? Take each real question and ask whether you could write all its searches before reading any result. If yes for all of them, choose the fixed pipeline.
- Question 3: do plain lookups also arrive? If some questions fail question 2 and lookups still make up part of the traffic, choose the routed design. If nearly every question fails it, choose the agent with retrieval.
Question 2 comes from the book’s example of a question a pipeline cannot answer, where “the second query cannot even be written until the first result has been read”. A question that needs several searches you can write up front is still pipeline work, because code can run them all.
The same holds for a chain whose shape never changes, such as fetching a customer’s record and then searching the manual for the plan it names. Code can carry that field from one search to the next, so the design is still a pipeline, with two fixed steps. The loop is for questions where what to search next cannot be known until a result has been read.
The decision table
The table holds one row per situation, with the design in the last column. Pick your design in the filter to see its rows. Every cost cell is explained in the next section.
| Your situation | Waiting time | Money, by direction | How you test it | How it fails | Design |
|---|---|---|---|---|---|
| Requests need no outside knowledge (rewrite, summarize pasted text) | One model call | One call with a short input | Check answers against the request | Fluent invention when the request did need a fact | no retrieval |
| All the knowledge fits in the window with room to spare, and every asker may read all of it | One call with a long input | The whole document set is read on every question | Nothing can be missed by a search, so check the answers | Attention thins as the input grows; cost per question rises with document size | no retrieval |
| The documents are too many to send whole, or some askers may not read all of them, and every search can be written from the question alone | One search, then one model call | One call per question, with a few passages as input | The path is fixed: test the search and the answer separately | A wrong or stale passage goes in and a confident, cited answer comes out | fixed pipeline |
| Nearly every question needs a search that depends on an earlier result | Calls stacked end to end; plan for a visible wait | Each search adds a model call that rereads what has piled up | The path varies per run: judge outcomes over repeated runs and read the traces | Off-topic chains; a skipped search; the cap is reached with no answer | agent with retrieval |
| Mostly lookups, with some questions whose later search depends on an earlier result | Lookups stay fast; the sorter adds one step; routed questions wait | Loop cost is paid only on routed questions | Give the sorter its own test set, then test each path | A hard question sent to the pipeline, or a lookup sent to the loop | routed |
One situation is missing on purpose. If the product must act (issue the refund, file the ticket) as well as answer, the retrieval choice stays the same and a larger one sits on top of it. The AI agent vs chatbot comparison named above owns that decision, and the Should this be an agent? decision aid walks it question by question.
What does each design cost?
On cost, RAG vs agents comes down to what a second search is worth. Each step from pipeline to loop buys the ability to search again and pays in waiting time, money and predictability. Chapter 1 gives the general rule: “An agent, then, is a cost you pay for adaptability you can name.”
What has been measured?
I found one controlled comparison that covers all three designs and the router, and its numbers need their setup stated beside them. Jeong and colleagues (NAACL 2024) tested open-domain question answering on sets that mix one-step and multi-step questions. They ran three strategies (no retrieval, single-step retrieval, multi-step retrieval) and a router, “a smaller LM trained to predict the complexity level of incoming queries”.
The table gives their averages for two of the three base models they tested, the smallest and the largest. Time is relative to single-step retrieval.
| Strategy in the study | Accuracy score | Searches per question | Time (single-step = 1.00) |
|---|---|---|---|
| Larger base model: no retrieval | 44.27 | 0.00 | 0.71 |
| Larger base model: single-step | 45.27 | 1.00 | 1.00 |
| Larger base model: multi-step | 49.70 | 2.81 | 3.33 |
| Larger base model: routed | 48.97 | 1.03 | 1.46 |
| Smaller base model: no retrieval | 15.97 | 0.00 | 0.11 |
| Smaller base model: single-step | 38.87 | 1.00 | 1.00 |
| Smaller base model: multi-step | 43.70 | 4.69 | 8.81 |
| Smaller base model: routed | 42.10 | 2.17 | 3.60 |
By my subtraction, multi-step retrieval gained 4.43 points over single-step with the larger model and 4.83 with the smaller. It took 3.33 and 8.81 times as long. The router kept 3.70 and 3.23 of those points at 1.46 and 3.60 times the single-step time.
The limits matter as much as the numbers. These are question-answering benchmarks over public knowledge, and your private documents are a different case. With no search at all, the larger model scored one point below single-step retrieval, which I read as a sign that it already knew many of the answers. The study reports time and says nothing about money.
Its multi-step arm is also a fixed iterative method, so it tells you about repeated searching and little about a free-running agent that picks its own tools. The models tested have since been superseded. The authors name the router as the weak point: its training labels were created automatically and “may have the potential to label queries incorrectly”.
Why does the bill move in that direction?
A loop costs more because each search adds a model call, and each call rereads everything gathered so far. Chapter 8 puts it in one clause: “a five-hop investigation pays for five searches and five readings, with the latency stacked end to end and the desk filling as the evidence accumulates”. The desk is the book’s image for the context window, and a hop is one search-and-read step.
An illustrative count shows the shape; these figures are mine and invented for the arithmetic. Suppose the instructions and question take 1,500 tokens (a token is roughly a word-piece), each search returns 2,000, and each query the model writes is 100. The pipeline reads 3,500 tokens once.
A loop with five searches reads 1,500 tokens, then 3,600, 5,700, 7,800 and 9,900, and then 12,000 to write the answer. That is 40,500 tokens of reading across six calls, about 11.6 times the pipeline. The multiple will differ in your system. It agrees in direction with Chapter 1: “A ten-step loop can cost an order of magnitude more than the single call it replaced (the multiplier is illustrative; the direction is what matters).”
Waiting time follows the same logic, since the calls run one after another. A 2025 engineering essay states the trade plainly: “runtime exploration is slower than retrieving pre-computed data” (Anthropic, 29 September 2025). The no-retrieval design has the opposite profile: one call, with a bill that grows with the size of what you paste.
How do testing and failure change?
The pipeline is the easiest design to test because its path never varies. You can check the search by itself (did the right passage come back?) and the answer by itself (did the model use it faithfully?).
A loop takes a different path on different runs. Chapter 1 lists this among the general costs of agents: the same input can take a different route and give a different answer each time. Applying that to retrieval is my step, and it means judging outcomes over repeated runs and reading the trace, the step-by-step record of a run.
The failures differ too. The pipeline fails silently and politely, and the chapter warns about its most persuasive form: “A wrong answer with a citation is more dangerous than a wrong answer without one, because the citation buys trust the answer has not earned.” The loop adds what the book calls the rabbit hole, “the fourth query subtly off-topic, the fifth chasing the fourth”. Its advice is to cap the searches, cap the spend, and show both caps in the trace.
When does “RAG gave wrong answers” mean add a loop?
Add a loop only for wrong answers caused by searching once. Every other kind of wrong answer has a cheaper repair, and a loop leaves some of them exactly as they were.
Sort your failed answers into four piles before anyone proposes a redesign.
- The fact was never in the documents, or the document was stale. Fix the documents. Chapter 8 is blunt that a pipeline serves whatever it indexed. A loop searching the same shelf finds the same stale page.
- The right passage existed and the search missed it. Fix retrieval. This is the tuning work named earlier, and it keeps the cheap design.
- The right passage was fetched and the model misread it. Fix the instructions or the model. Searching again does nothing for a reading error.
- The question needed a second search that depended on the first result. This pile is the case for a loop, or for a router if the pile is small.
The fourth pile has a research pedigree. Trivedi and colleagues (ACL 2023) describe it: “what to retrieve depends on what has already been derived, which in turn may depend on what was previously retrieved.” Their paper also states that the gains from interleaving searches “come with an additional computational cost.”
Field reports on the switch point in opposite directions, and I can only present them as individual reports. One practitioner wrote “My biggest RAG learning is to use agentic RAG”, meaning “you provide the search as a tool to the LLM” (pietz, Hacker News, 21 October 2025). A reply the same day said “the assistant often doesn’t search when it should and very rarely does multiple search rounds” (jokethrowaway, 21 October 2025). Neither posted a measurement.
A third report describes neither extreme. In a thread opened on 27 July 2025, kingkongjaffa wrote of a product “used by thousands of users”. In it, “the user input is just one step in a branching decision tree” of hand-written prompts. In the book’s terms that is a workflow, built by hand around retrieval.
Is agentic RAG vs traditional RAG a clean line?
The line is clean in principle and blurred in the market. In principle, traditional RAG is the fixed pipeline and agentic RAG is retrieval inside a loop the model directs.
In practice the label stretches. The 2025 survey by Singh and colleagues that gathered the family under one name defines it as “embedding autonomous AI agents into the RAG pipeline”, and its taxonomy includes a single-agent router. A router makes one model decision and then follows code. By the book’s definitions that is a workflow.
So “agentic RAG” on a slide can mean the routed design, a capped loop or an open-ended research assistant. Those three have very different bills, and the questions in the next section show which one is on offer.
What should you ask your team or a vendor?
Ask five questions about calls, queries, stops, evidence and question mix. The list is this post’s own, derived from the Chapter 8 passages quoted above, and it works on an internal design review as well as on a sales call.
Five questions about a feature that answers from our documents
1. For one user question, how many model calls and how many searches
happen? Is the number always the same, or does it vary?
A good answer shows: a fixed count (a pipeline), a range with a
stated maximum (a loop), or both with the share of traffic on each
(a router).
2. Who writes the search query: our code, from the user's words, or
the model, after reading earlier results?
A good answer shows: one of the two, said plainly. "The model,
after reading results" means a loop, with a loop's costs.
3. What stops it? Is there a cap on searches and on spend, and does
the cap appear in the log when it is hit?
A good answer shows: two numbers and a log line. "It stops when it
is done" is a missing answer.
4. When an answer is wrong, can you show me which passages the model
was given, and for a loop, each search in order?
A good answer shows: a real record of one failed question, opened
in front of you.
5. Which of our real questions need a second search that depends on
the first result, and what share of traffic are they?
A good answer shows: our questions sorted into piles, with counts.
A guess here means the design was chosen before the questions
were read.
A fixed count in answer to question 1 means a pipeline, whatever the brochure calls it. Question 5 supplies the input to the decision table above.
How does the sort run on a real product?
Take a help-center answer bot over 300 articles, the kind of product a RAG vs agents meeting is held about, and walk it through the three questions. The product and its questions are invented for illustration.
Question 1: does it fit, or is nothing needed? The requests need the articles, so something is needed. Assume the 300 articles are too much to send whole on every question. The answer is no, and the walk continues.
Question 2: can every search be written from the question alone? The ten most common questions are of one kind: how to reset a password, what the refund window is, whether a plan includes exports. Each can be searched from its own words. The answer is yes for all ten, and the design is the fixed pipeline. The table row agrees: too many documents to send whole, with every search writable up front.
Now change the mix. Enterprise customers begin asking the question Chapter 8 uses as its example: “Did last March’s outage affect any customer whose contract guarantees four-nines (99.99 percent) uptime?” The incident report names the failed services, and only then can anyone search for the customers those services touch. Question 2 now fails for that pile, lookups still arrive, and question 3 lands on routed.
I walked two more products the same way. The first is a tool whose questions all join facts across thousands of contracts, where one clause decides what to look for next. It fails questions 1 and 2 and has no lookups left, so it lands on agent with retrieval.
The second is an internal assistant whose whole handbook fits in the window and is open to every employee. It stops at question 1 with no retrieval.
Where does this post stop?
This post sorts RAG vs agents choices by the kind of question and leaves several real inputs out. Five limits deserve a plain statement.
Share matters, and the table ignores it. Suppose dependent questions arrive a few times a year. A pipeline that declines them honestly may then beat building a router. That is a judgment about volume and value, and the three questions do not make it for you.
The measurement is narrow. One study, public question sets, time only, and a scripted multi-step method. I found no controlled comparison of the money cost of agentic retrieval against a pipeline. One practitioner’s suggestion is a better unit: “cost per successful task is more meaningful than cost per query” (reggz, Hacker News, 3 March 2026).
Retrieved text is untrusted. Chapter 8 says so in five words: “retrieved content is untrusted content.” A document can carry instructions planted by anyone able to write to the source, and a loop that acts on what it reads raises the stakes. That subject needs its own treatment.
Retrieval is no cure for invention. The chapter’s verdict is that “retrieval lowers the fabrication rate and raises your odds of catching what remains, and both are worth paying for, and neither is a cure.”
Memory is left out. What an agent stores from its own past runs uses similar machinery pointed at a different source. Retrieval is also one part of a wider discipline, described in the guide to what context engineering is: deciding everything the model is given on each call.
The takeaway
Decide the two layers separately. First ask where the knowledge comes from, then ask who decides the next step, and let your users’ real questions answer both. The RAG vs agents argument fades once the ten questions are on the table and each one has been sorted.
As for the “RAG is dead” headline, one commenter took a middle position: “Personally I see more in a combination of RAG and live querying during the thinking process (e.g. by tools)” (wolvoleo, Hacker News, 7 July 2026). The routed design is one way to build that combination, and the book’s bearing for it is “the simplest retrieval that works”.
Chapter 1, “What Is an Agent?” is free to read online, along with Chapter 2 and the glossary. The pipeline, the loop and the front door are built in Chapter 8, Retrieval and Knowledge (in the full book). The context engineering guide places this post beside its neighbors, or you can see the formats.
Questions readers ask
- Is RAG an agent?
- A plain RAG pipeline is a workflow. Code decides the order: search once, place the passages beside the question, call the model, return the answer. An agent is a model in a loop that decides its own next step. RAG becomes part of an agent only when the search is offered to that loop as a tool.
- What is agentic RAG, and how does it differ from traditional RAG?
- Agentic RAG puts the retriever inside an agent loop as a tool. The model then decides whether to search, what to search for, where to search and whether the results are enough. Traditional RAG searches once, in code, before the model is called. The label is loose: a 2025 survey that named the family also files a single-agent router under it.
- Is agentic RAG better than traditional RAG?
- It depends on the question. Where a second search depends on the first result, iterating helps. In one 2024 study of mixed open-domain questions, multi-step retrieval scored 49.70 against 45.27 for single-step, at 3.33 times the time with the larger base model. On single lookups the book calls the loop pure overhead.
- Is RAG dead?
- The headline usually carries one of two narrower claims. Either a long context window can now hold a bounded set of documents whole, or an agent with search tools can replace a hand-built search step. Both still place fetched text in front of the model, so retrieval remains; what changes is who decides when it runs.
- My RAG feature gives wrong answers. Will an agent fix it?
- A loop fixes the wrong answers caused by searching only once, where the second search depends on what the first one found. It leaves alone the answers that are wrong because a document is stale, missing or contradicted by another document. It also adds a new failure: a model that skips a search it should have run.
Sources
- Patrick Lewis, Ethan Perez, Aleksandra Piktus and colleagues (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv:2005.11401, NeurIPS 2020)
- Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang and Jong C. Park (2024). Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity (arXiv:2403.14403, NAACL 2024)
- Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot and Ashish Sabharwal (2023). Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions (arXiv:2212.10509, ACL 2023)
- Aditi Singh, Abul Ehtesham, Saket Kumar, Tala Talaei Khoei and Athanasios V. Vasilakos (2025). Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG (arXiv:2501.09136, preprint)
- Anthropic (2024). Building effective agents (19 December 2024)
- Anthropic (2025). Effective context engineering for AI agents (29 September 2025)
- Anthropic (2025). How we built our multi-agent research system (13 June 2025)
- Google (2024). Try Deep Research and our new experimental model in Gemini, your AI assistant (11 December 2024)
- Hacker News commenter schmorptron (2025). Hacker News comment on people meaning different things by RAG (12 October 2025)
- Hacker News commenter madeofpalk (2025). Hacker News comment asking whether tool use is a form of RAG (2 October 2025)
- Hacker News commenter pietz (2025). Hacker News comment recommending search offered as a tool (21 October 2025)
- Hacker News commenter jokethrowaway (2025). Hacker News reply reporting that the assistant often does not search (21 October 2025)
- Hacker News commenters yamarldfst and reggz (2026). Ask HN: Agentic search vs. RAG, what's your production experience? (24 February 2026; reply by reggz, 3 March 2026)
- Hacker News commenters TXTOS and kingkongjaffa (2025). Ask HN: Are we pretending RAG is ready, when it's barely out of demo phase? (27 July 2025; reply by kingkongjaffa)
- Hacker News commenter wolvoleo (2026). Hacker News comment on the claim that RAG is dead (7 July 2026)