Skip to main content

Keyword, Vector and Hybrid Search for RAG

Why keyword and vector search fail in different ways, how hybrid search and reranking combine them, and how query rewriting finds what the raw question misses.

AdvancedVerdeshell Team · 5 min read · Last reviewed

Keyword search finds exact terms and misses synonyms; vector search finds meaning and misses exact identifiers. Running both, merging the results and reranking the best candidates is the most dependable default for retrieval.

Key takeaways

  • Keyword search (such as BM25) matches exact words — strong on names, codes and rare terms, weak on paraphrases.
  • Vector search matches meaning — strong on paraphrases, weak on exact identifiers.
  • Hybrid search runs both and merges the results, often with reciprocal rank fusion.
  • A reranker re-reads the query with each top candidate; it is slower but more accurate, so it is applied only to a shortlist.
Hybrid search: keyword and vector search combined, then rerankedA query is sent to two searches at once. Keyword search, such as BM25, matches exact terms and is good at names, codes and rare words. Vector search matches meaning and is good at paraphrases. Their two ranked lists are merged, for example with reciprocal rank fusion, and a reranker reads the query and each candidate together to pick the few best passages for the prompt.QueryKeyword searchBM25 · exact terms, codesVector searchembeddings · meaningFusemerge rankingsRerankread query + passagetop few → the prompt
Two searches with opposite blind spots, merged by rank, then re-read by a more careful model before anything reaches the prompt.

Hover or tap the diagram to replay the animation.

Keyword search ranks passages by the words they share with the query, weighting rare words more heavily than common ones. The standard method, BM25, comes out of decades of information-retrieval research summarised by Robertson and Zaragoza in 2009, and it remains a strong baseline.

It excels at exact matches: product codes, error messages, policy numbers, names, acronyms. It fails on vocabulary mismatch — a question about “annual leave” will not match a policy that only says “paid time off”.

Vector search compares the embedding of the query with the embeddings of the chunks and returns the closest in meaning. It handles paraphrase and synonyms well: the leave question finds the time-off policy.

Its blind spot is the opposite one. Identifiers, part numbers and rare proper nouns carry little “meaning” for an embedding model, so a query for a specific invoice number may return passages about invoices in general.

Hybrid search: use both

Because the two fail in different places, running both and combining the results usually beats either alone. The common way to merge them is reciprocal rank fusion, introduced by Cormack, Clarke and Büttcher in 2009: each passage scores according to its position in each ranked list, so results that rank well in both rise to the top, and the two methods’ incompatible scores never need to be compared directly.

Anthropic’s 2024 contextual retrieval write-up reached the same practical conclusion — embeddings and BM25 combined, with a reranking step on top.

Reranking

First-stage search is built for speed across millions of chunks, so it compares query and passage only loosely. A reranker — typically a cross-encoder, which reads the query and one passage together — judges relevance much more accurately. The approach was shown for passage ranking with BERT by Nogueira and Cho in 2019.

Rerankers are slow per passage, so the standard pattern is to retrieve a generous shortlist cheaply, rerank it, and keep only the best few for the prompt.

Improving the query

Users ask short, ambiguous questions. Transforming the query before searching often helps more than tuning the search itself.

Rewriting: have a model expand abbreviations, add the likely terms, or resolve a follow-up question (“what about for contractors?”) into a complete one using the conversation so far.

Decomposition: split a compound question into several searches and combine the results.

Hypothetical answers: HyDE (Gao et al., 2022) has a model write a plausible answer first and searches with that, since an answer resembles the target passage more closely than a question does. The hypothetical answer is only used to search — never shown as the answer.

Filtering, and how many results to keep

Apply metadata filters as part of the search — permissions, document type, date, region — so that the best match is the best permitted, current match.

Then decide how many passages go into the prompt. Too few and the answer may be missing; too many and the useful passage competes with noise, costs more, and may sit in the middle of a long context where models use it less reliably. Put the strongest passages first, and measure the effect of the number you choose rather than picking one by feel.

Which of these steps your system actually needs is a question for measurement — the subject of how to evaluate a RAG system.

Next in the pathHow to Evaluate a RAG System

Want this built properly?

We design and build AI systems for clients. Tell us the problem and we will tell you honestly whether AI — and which kind — is the right fit for it.