Figure 8.2 Retrieval-augmented generation. Ahead of time, documents are chunked,
embedded, and written to a vector store (the one thing RAG adds, in accent). At
query time the question retrieves the closest matches, which augment the prompt
the model actually sees; the answer is grounded in what was retrieved.