Blog
4 min read

Chunking for RAG: How to Split Documents So Retrieval Works

How you split documents into chunks decides what a RAG system can retrieve. Fixed-size, recursive, structure-aware and semantic chunking, choosing chunk size and overlap, adding context to chunks, parent-child retrieval, chunking code and tables, and how to evaluate it.

In a RAG system, you don't search whole documents — you search chunks: pieces of documents, each embedded and stored. When a question comes in, the closest chunks are retrieved and handed to the model. (What is RAG?, What are embeddings?)

So chunking quietly decides everything downstream. If the answer is split across two chunks, or buried in a chunk about something else, retrieval misses it and the model either guesses or says it doesn't know.

The trade-off

  • Small chunks (a few sentences) — precise matches, but may lack the context needed to understand or answer.
  • Large chunks (pages) — plenty of context, but the embedding becomes a blur of several topics and matches less precisely; they also use more of the model's context.

A common starting point is roughly 200–800 tokens per chunk with 10–20% overlap, then tune with evaluation.

Strategies

1. Fixed-size

Split every N tokens (or characters), with some overlap so sentences cut at a boundary appear in both chunks.

  • ✅ Simple, predictable.
  • ❌ Cuts mid-sentence and mid-thought; ignores structure.

2. Recursive

Try to split on the largest natural boundary first — sections, then paragraphs, then sentences, then words — only going smaller when a piece is still too big. Most libraries' default "text splitter" works this way.

  • ✅ Good general-purpose default.

3. Structure-aware

Use the document's own structure: Markdown headings, HTML sections, PDF headings, slides, FAQ entries, rows. One chunk per section (split further if long).

  • ✅ Chunks align with how humans organised the content — usually the biggest quality win for docs and knowledge bases.

4. Semantic

Embed sentences and split where the topic shifts (where consecutive sentences become less similar).

  • ✅ Topic-coherent chunks for unstructured prose.
  • ❌ More compute; results vary.

Make every chunk self-explanatory

A chunk that says "This setting defaults to 30 days" is useless if you don't know which setting. Add context:

  • Prepend the title and heading path: Billing › Refunds › Time limits: This setting defaults to 30 days…
  • Contextual retrieval: use a model to write a one-sentence summary situating each chunk within its document, and embed that with the chunk. This noticeably improves retrieval in many setups, at extra indexing cost (prompt caching makes it cheaper). (Prompt caching explained)
  • Store metadata: source, URL, section, date, permissions — for filtering and for citations.

Retrieve small, return big

Parent-child retrieval: index small chunks for precise matching, but when one matches, give the model its parent section (or neighbours) for context. You get precise search and enough context to answer.

Special content

  • Code — split by function or class, not by line count. Include the file path and signature.
  • Tables — keep a table together, or convert rows to sentences with column names ("Plan: Pro; Price: $20; Seats: 5").
  • PDFs — extraction quality matters more than the splitter; headers, footers and two-column layouts produce garbage text if not cleaned.
  • Conversations and tickets — chunk by message or thread, with participants and dates.

Don't rely on vectors alone

Combine semantic search with keyword search (hybrid), and consider a reranker — a model that re-scores the top 20–50 candidates for relevance before you pick the final few. (Semantic vs keyword search)

Evaluate, don't guess

Build a set of real questions with the chunk(s) that should answer them. Measure retrieval hit rate (is the right chunk in the top k?) for each chunking setup before judging the final answers. Change one variable at a time. (Evaluating LLM outputs)

Re-indexing

Chunks must be regenerated when documents change. Store a content hash per document and re-chunk only what changed. Changing chunking or embedding model means re-indexing everything — plan for it.

A practical default

  1. Structure-aware splitting on headings, recursive fallback.
  2. ~400–600 tokens, ~15% overlap.
  3. Prepend title + heading path to each chunk.
  4. Hybrid search, top 20 → rerank → top 5.
  5. Evaluate, then tune.

EasySpawn runs your RAG pipeline next to its data: Postgres with pgvector and full-text search, background workers for re-indexing, and Claude Code to tune the chunking. See how it works or join the waitlist.

Related: What Is RAG? · pgvector Tutorial · What Is a Vector Database? · Semantic Search vs Keyword Search

Keep reading