Skip to main content

Chunking and Embeddings for RAG

How to split documents so the right passage can be found, what embeddings are and how to choose a model, and why metadata matters as much as the text.

IntermediateVerdeshell Team · 6 min read · Last reviewed

A RAG system can only retrieve what its chunks make findable. Split documents along their natural structure, keep enough context in each chunk to stand alone, store the metadata you will need to filter and cite, and test chunk sizes on real questions.

Key takeaways

  • Chunking splits documents into passages small enough to retrieve precisely and large enough to make sense on their own.
  • Split along the document’s structure — sections, headings, paragraphs — rather than at a fixed character count where you can.
  • An embedding is a list of numbers representing meaning; text with similar meaning gets similar numbers.
  • Changing the embedding model means re-embedding every chunk — vectors from different models cannot be compared.

Why documents are split

A retrieval system returns passages, not documents. Return a whole 40-page manual and the relevant paragraph is buried — expensive to send and easy for the model to miss. Return a single sentence and it may have lost the context that makes it meaningful.

Chunking is choosing that unit. It is the step teams most often treat as a default setting, and one of the biggest influences on answer quality, because a passage that was split badly cannot be retrieved well by any search method.

Chunking strategies

Fixed size with overlap. Split every N tokens, with some overlap between neighbouring chunks so that a sentence on a boundary appears in both. Simple and predictable, and blind to structure — it will happily cut a table or a numbered procedure in half.

Structure-aware. Split along the document’s own structure: sections, headings, paragraphs, list items, table rows. Chunks then correspond to units a person would recognise, and headings can be carried into each chunk as context. Where documents have structure, this is usually the better default.

Content-specific. Treat different content differently: keep a table whole, split code at function boundaries, keep a question and its answer together in an FAQ.

Small to retrieve, larger to read. A common pattern is to search over small, precise chunks but pass the surrounding section to the model once a match is found — precise retrieval without losing the context needed to answer.

The context problem, and a fix

A chunk taken out of its document often loses what made it meaningful. Anthropic’s 2024 write-up on what it calls contextual retrieval gives the example of a chunk that says revenue grew over the previous quarter — without saying which company, or which quarter.

The fix it describes is to prepend a short, chunk-specific explanation of where the chunk comes from before it is indexed:

[Context: From the Q2 2026 quarterly report of Acme Logistics, section "Financial performance".]
Revenue grew 3% over the previous quarter, driven by new warehouse contracts in Pune.

Even a simpler version — carrying the document title and section headings into every chunk — makes many chunks retrievable that otherwise would not be.

How big should a chunk be?

There is no universal answer, and any number quoted without your documents and questions behind it is a guess. Too small and chunks lose context; too large and each one mixes several topics, so it matches many questions weakly and costs more to send.

Pick a starting point from your documents’ structure, then test: run a set of real questions, check whether the passage that answers each one is retrieved, and adjust. That test set is described in how to evaluate a RAG system.

Metadata: the part that is not text

Store with each chunk the information you will need later: the source document and section (for citations), dates and version (so outdated content can be excluded), document type, language, and above all who is allowed to see it.

Metadata turns retrieval from “the most similar text anywhere” into “the most similar text in the current HR policy that this employee may read” — which is usually the question actually being asked.

What embeddings are

An embedding is a list of numbers that represents the meaning of a piece of text, produced by an embedding model. Texts with similar meanings produce vectors that sit close together, so “How do I reset my password?” and “I forgot my login” can match even though they share no words.

Sentence-level embeddings became practical for search with work such as Sentence-BERT (2019), and dense passage retrieval (2020) showed that searching by embeddings could beat traditional keyword search on open-domain question answering. Embedding models are separate from the chat models that write answers; you choose them separately.

Choosing an embedding model and an index

Test on your own content. Check quality in every language you serve — including Indian languages if your users write in them — and on your domain’s vocabulary. Consider the size of the vectors, which affects storage and search cost.

Plan for change. Vectors from different embedding models are not comparable, so switching models means re-embedding every chunk. Keep the original text and metadata so that re-indexing is a batch job, not a migration project.

Storage can be a dedicated vector database or a vector extension to a database you already run, such as pgvector for PostgreSQL. Most use approximate nearest-neighbour search — algorithms such as HNSW (2016) — which trades a little accuracy for a great deal of speed on large collections.

Next in the pathKeyword, Vector and Hybrid Search for RAG

Want this built properly?

We design and build AI systems for clients. Tell us the problem and we will tell you honestly whether AI — and which kind — is the right fit for it.