Skip to main content

How RAG Works, Step by Step

Retrieval-augmented generation end to end: what happens when documents are indexed, what happens when a question arrives, and where each step goes wrong.

BeginnerVerdeshell Team · 6 min read · Last reviewed

RAG answers questions from your own documents by finding the relevant passages at the moment of the question and giving them to the model. It has two halves — preparing the documents, and searching them — and most quality problems start in the first.

Key takeaways

  • Retrieval-augmented generation (RAG) finds relevant passages from your documents at question time and gives them to the model to answer from.
  • It runs in two phases: indexing (split, embed and store documents ahead of time) and querying (retrieve, rerank, prompt, answer).
  • Most RAG failures are retrieval failures — if the right passage is never found, no model can use it.
  • Retrieval must respect permissions: a user should only ever get answers from documents they are allowed to see.
The two phases of retrieval-augmented generationIndexing, done ahead of time: documents are cleaned and split into chunks, each chunk is turned into an embedding, and the chunks are stored in a search index. At question time: the question is used to retrieve the most relevant chunks from the index, the best are reranked, the chunks and the question are assembled into a prompt, and the model writes an answer grounded in them, with citations.INDEXING — DONE AHEAD OF TIMEDocumentspolicies, manuals, ticketsSplitinto chunks + metadataEmbedeach chunk → vectorSearch indexvector + keywordAT QUESTION TIMEQuestionRetrievetop matchesRerankkeep the bestPromptquestion + chunksAnswerwith citationsthe index is searched at question timeMost quality problems start in the top row: a badly split document cannot be retrieved well,and a passage that is never retrieved cannot be used — however good the model is.
Two phases: documents are prepared once, ahead of time; every question then searches what was prepared.

Hover or tap the diagram to replay the animation.

What RAG is

Retrieval-augmented generation is a way of answering questions from information the model was never trained on — your policies, product documentation, contracts, support history. Instead of hoping the model knows, the system looks up the relevant passages when the question arrives and puts them in front of the model alongside the question.

The term comes from a 2020 paper by Patrick Lewis and colleagues at Facebook AI Research, which combined a retriever with a generator so that a model could draw on a document index rather than only on what it had memorised. The idea has since become the standard way to connect language models to private and current information — and the main defence against hallucination on factual questions.

If you are still deciding whether you need it at all, start with prompting, RAG or fine-tuning.

Phase 1 · Indexing: preparing the documents

Indexing happens ahead of time, and again whenever documents change.

Load and clean. Documents arrive as PDFs, web pages, spreadsheets, tickets. Extracting clean text — keeping headings, tables and lists intact, dropping navigation and boilerplate — matters more than it sounds: a table flattened into a jumble of numbers cannot be retrieved or understood.

Split into chunks. Whole documents are too long to retrieve usefully, so they are split into passages, each stored with metadata such as its source, section, date and who may see it. How you split is one of the biggest quality levers — see chunking and embeddings.

Embed and store. Each chunk is converted into an embedding — a list of numbers representing its meaning — and stored in a search index, usually alongside a keyword index so that exact terms can be matched too.

Phase 2 · Answering a question

Retrieve. The question is used to search the index for the most relevant chunks — by meaning, by keywords, or both. Getting this step right is the subject of retrieval: keyword, vector and hybrid search.

Rerank. A first search is fast but approximate, so the top candidates are often re-scored by a slower, more accurate model, and only the best few are kept.

Assemble the prompt. The selected chunks go into the prompt with the question and instructions: answer only from these passages, cite which one supports each claim, and say so if the answer is not there.

Generate and cite. The model writes the answer from the passages. Returning the sources alongside the answer lets users check it — and lets you measure whether the system is working, covered in how to evaluate a RAG system.

Permissions and freshness

Two requirements are easy to miss in a prototype and essential in production.

Access control. If documents have different audiences, retrieval must filter by what the person asking is allowed to see — before the passages reach the model. A model cannot be trusted to withhold text it has been given; the safe design is to never retrieve it.

Freshness. An index is a copy of your documents at a point in time. Decide how changes, deletions and new documents flow into it, and how quickly — an answer built from a superseded policy is wrong in a way that looks right.

Do you need RAG, or a long prompt?

Context windows are now large enough to hold a lot of text, and for a small, stable set of documents you may not need retrieval at all. Anthropic’s 2024 guidance on retrieval suggested that a knowledge base of up to roughly 200,000 tokens — around 500 pages — could simply be included in the prompt.

RAG still wins when the material is large, changes often, has different permissions for different users, or when cost and speed matter: sending only the relevant passages is cheaper and faster than sending everything, and models use information in a long context less reliably than information placed close to the question. See tokens and context windows.

Beyond the basic pipeline

Query transformation. Rewriting the user’s question before searching — expanding abbreviations, splitting a compound question, or generating a hypothetical answer and searching with that — often finds better passages than the raw question.

Agentic RAG. Instead of one fixed search, an AI agent decides when to search, what to search for, and whether the results are good enough or it needs to search again. More capable, and harder to test.

Graph-based RAG. Microsoft Research’s GraphRAG (2024) builds a graph of entities and relationships from the documents, which helps with broad questions about a whole collection — “what are the main themes across these reports?” — that passage-by-passage retrieval answers poorly.

Self-checking. Approaches such as Self-RAG (2023) train the model to decide when retrieval is needed and to critique whether its answer is supported. In practice, most teams get further by first fixing chunking, retrieval and evaluation.

Next in the pathChunking and Embeddings for RAG

Want this built properly?

We design and build AI systems for clients. Tell us the problem and we will tell you honestly whether AI — and which kind — is the right fit for it.