Skip to main content

How to Evaluate a RAG System

Measure retrieval and answers separately, build a test set from real questions, use model-based judges carefully, and know which fix each failure points to.

IntermediateVerdeshell Team · 5 min read · Last reviewed

A RAG system can fail by finding the wrong passages or by misusing the right ones. Measure retrieval and answers separately, on a fixed set of real questions, and every change becomes a decision you can check instead of a guess.

Key takeaways

  • Evaluate retrieval and generation separately: was the right passage found, and was the answer faithful to it?
  • Build a test set of real questions with expected answers and source passages — including questions the documents cannot answer.
  • Faithfulness (is every claim supported by the retrieved text?) is the key answer metric for RAG.
  • Model-based judges scale evaluation, but check them against human judgements before trusting them.
Evaluating a RAG system: retrieval quality and answer qualityA test question flows through retrieval and then generation. Retrieval is measured on its own: did the right passages appear among the results, and how high were they ranked. Generation is measured on the answer: is every claim supported by the retrieved passages, does it answer the question, is it correct, and did the system decline when the documents do not contain the answer.Test questionRetrievepassagesGenerateanswerAnswerRetrievalRight passage found?Ranked near the top?Little irrelevant text?AnswerEvery claim supported by the passages? (faithfulness)Answers the question that was asked?Correct against the expected answer?Declines when the documents are silent?
Measure the two halves separately — a wrong answer means something different when the right passage was never retrieved.

Hover or tap the diagram to replay the animation.

Why RAG needs its own evaluation

When a RAG system gives a wrong answer, there are two very different causes. Either retrieval failed — the passage that contains the answer was never found — or generation failed — the passage was there and the model ignored it, misread it or added to it. The fixes are completely different, so the measurement has to tell them apart.

Evaluation is also what makes every other decision in this topic answerable: chunk size, hybrid search, reranking, which model. Without it, each change is judged on a handful of questions someone happened to try.

Build a test set first

Collect real questions — from support tickets, search logs, or the people who will use the system — rather than inventing them. For each, record the expected answer and the passage or passages that contain it.

Include the hard cases on purpose: questions whose answer spans two documents, questions using different words from the documents, questions about outdated content, and questions the documents do not answer at all. The last group tests whether the system declines instead of hallucinating.

A few dozen good questions are far more useful than a large set nobody has checked. Grow it whenever a real failure is reported.

Measuring retrieval

For each test question, check the retrieved passages against the ones you recorded.

Recall at k: did the right passage appear anywhere in the top k results? If not, nothing downstream can fix it.

Rank: how high did it appear? A passage found at position 40 that the system never sends to the model is as good as missing.

Precision: how much of what was retrieved was relevant? Irrelevant passages cost money and can distract the model.

These numbers point directly at chunking and embeddings and search and reranking — and can be measured without generating a single answer.

Measuring answers

Faithfulness, sometimes called groundedness: is every claim in the answer supported by the retrieved passages? This is the metric that catches a model adding plausible facts of its own.

Relevance: does the answer address the question that was asked?

Correctness: does it match the expected answer you recorded?

Appropriate refusal: for questions the documents cannot answer, did the system say so instead of guessing?

Citation accuracy: if the answer cites passages, does each cited passage actually support the sentence it is attached to?

Using models as judges

Checking faithfulness and relevance by hand does not scale, so teams use a language model as the judge — given the question, the retrieved passages and the answer, it scores each property. The Ragas framework (Es et al., 2023) proposed metrics of this kind for RAG, including faithfulness, answer relevance and context relevance, designed to work without a human-written reference answer for every question.

Judges are useful and imperfect. Before relying on one, have people score a sample of the same answers and check that the judge agrees with them. Keep a person reviewing a sample in production, especially where wrong answers are costly.

From scores to fixes

Low recall: the problem is in indexing or search — chunking, missing context in chunks, the embedding model, or the need for hybrid search or query rewriting.

Good recall, poor faithfulness: the passages are there but the answer drifts — tighten the prompt (answer only from the passages, quote first), reduce the number of passages, or try a different model.

Wrong answers on outdated topics: the index is stale, or date metadata is not being used.

Run the full set on every change — to chunking, prompts, models or documents — and keep the history. A change that improves the question you were looking at and quietly breaks five others is the most common way RAG systems get worse over time.

Next in the path · AI AgentsAI Agents vs Agentic AI vs Generative AI

Want this built properly?

We design and build AI systems for clients. Tell us the problem and we will tell you honestly whether AI — and which kind — is the right fit for it.