Chunking Strategies for RAG: Fixed-Size, Semantic, and Structure-Aware Splitting Compared
The chunking strategy you pick for a RAG pipeline matters as much as the model you query: split documents too coarsely and answers drown in irrelevant text, split them too finely and you lose context. This guide compares fixed-size, sentence, semantic and structure-aware chunking, explains how overlap and chunk size change answer quality, and gives practical defaults you can apply without running a lab experiment first.
Why chunking is the quiet decision that breaks or makes a RAG system
Retrieval-augmented generation works by searching a vector index for passages related to a question, then handing those passages to a model to compose an answer grounded in the source text. If you are new to the overall mechanism, What Is RAG (Retrieval-Augmented Generation)? A Practical Guide walks through the full loop. Chunking is the step before indexing: deciding where to cut a document into the pieces that will actually get embedded, stored and retrieved.
Every downstream problem traces back to a bad cut. A chunk that ends mid-sentence confuses the embedding model. A chunk that mixes two unrelated topics gets retrieved for the wrong question. A chunk that is too short loses the context a human reader would have gotten from the surrounding paragraph. None of this is visible until you ask a question and the answer is subtly wrong or missing a caveat that was two sentences away in the original file. For the mechanics of how chunks move from upload to a cited answer, see Chunking, indexing, retrieval: how a RAG pipeline works under the hood.
Fixed-size chunking: fast, predictable, and a reasonable default
Fixed-size chunking splits text every N characters or tokens, usually with a sliding window. It is the simplest method to implement and the cheapest to run at scale, since you do not need a model to decide where to cut.
- Pros: predictable chunk count and cost, easy to parallelize, works on any file type including plain text dumps and logs
- Cons: cuts can land mid-sentence or mid-table, unrelated ideas can end up fused together if a paragraph break falls inside the window
- Best for: homogeneous text with short paragraphs, transcripts, or as a fallback when structure-aware parsing fails
In practice, fixed-size chunking rarely means a raw character count. Most implementations snap to the nearest sentence or paragraph boundary near the target size, which removes the worst failure mode (mid-sentence cuts) while keeping the predictability of a fixed budget.
Sentence and paragraph based chunking
Sentence-based chunking groups whole sentences until a size threshold is reached, then starts a new chunk. Paragraph-based chunking does the same at the paragraph level. Both respect natural language boundaries, which keeps embeddings cleaner because the text fed to the embedding model reads as a coherent unit rather than a fragment.
The trade-off is variance: a document with long paragraphs produces large chunks, a document with one-line bullet points produces tiny ones. That variance is usually acceptable for prose-heavy sources like policies, manuals and reports, but it can create uneven retrieval quality on documents that mix dense paragraphs with sparse lists, tables or code. Understanding how text becomes a vector in the first place helps explain why boundaries matter here; see Embeddings Explained for Non-Specialists: How Text Becomes Numbers.
Semantic chunking: cutting where the topic actually changes
Semantic chunking uses embeddings (or a lighter similarity signal) to detect where the topic shifts, rather than relying on a fixed size. The typical approach embeds sentences or small groups of sentences, computes similarity between consecutive segments, and starts a new chunk when similarity drops below a threshold.
This produces chunks that are topically coherent regardless of length, which tends to improve retrieval precision on long, meandering documents like legal contracts or research papers where a single paragraph can quietly pivot from one clause to another. The cost is computational: you need an embedding pass (or a cheaper proxy) just to decide where to cut, before the real indexing embedding pass happens. For most knowledge bases built from well-structured documents, the gain over paragraph-based chunking is modest; for messy, unstructured text it can be the difference between a usable answer and a jumbled one.
Structure-aware chunking: respecting headings, tables and code
Structure-aware chunking parses the document's actual layout (Markdown headings, Word heading styles, PDF bookmarks, HTML tags) and cuts along those boundaries first, only falling back to size-based splitting within a section that is too long. This is usually the strongest default for technical documentation, policies and procedures, because it keeps a section's heading attached to its content, which gives the retrieval step much stronger signal than an arbitrary slice of running text.
- Tables: keep a table (or a logical slice of a very long one) as its own chunk with its caption or preceding heading, rather than letting a fixed-size cut slice rows apart
- Code blocks: never split a function or config block across chunks; treat each block as atomic even if it pushes the chunk over the normal size target
- Headings: prepend the section heading (and ideally its parent heading) to every chunk under it, so the chunk is understandable without the surrounding document
- Lists: keep short lists intact; split only very long lists, and repeat the introductory sentence in each resulting chunk
Match the strategy to the document, not the other way around
A mixed knowledge base (contracts plus spreadsheets plus Slack exports) rarely benefits from one global rule. A pragmatic approach is structure-aware chunking for anything with real headings, paragraph-based chunking as a fallback, and a conservative fixed-size cut as the last resort for unstructured dumps.
Chunk overlap: how much, and why it helps
Overlap means repeating a small amount of text (often 10-20% of the chunk size) at the start of each new chunk, taken from the end of the previous one. The point is to avoid losing context exactly at a boundary: if a key sentence sits right at the cut, overlap gives the retriever a second chance to surface it inside the neighboring chunk.
Overlap is not free. It duplicates content in the index, which increases storage and can slightly inflate retrieval noise if the same sentence gets pulled twice from adjacent chunks. As a rule of thumb, 50 to 100 tokens of overlap on a 300 to 800 token chunk covers most cases; go higher only for dense technical or legal text where a single clause reference can span a boundary, and skip overlap entirely for short, self-contained items like FAQ entries or product descriptions where each chunk is already a full unit.
How chunk size affects answer quality
Chunk size is a direct trade-off between precision and context. Small chunks retrieve more precisely (less irrelevant text mixed in) but can strip out the surrounding explanation a model needs to answer correctly. Large chunks preserve context but dilute the embedding, since a vector representing 2,000 tokens on several sub-topics matches fewer specific questions well.
Chunk size versus typical outcome
| Chunk size | Retrieval precision | Context preserved | Typical use case |
|---|---|---|---|
| 100-200 tokens | High | Low | Short facts, glossary entries, FAQ pairs |
| 300-500 tokens | Balanced | Moderate | General documentation, policies, most knowledge bases |
| 600-1,000 tokens | Lower | High | Narrative content, case studies, legal clauses needing surrounding context |
| 1,000+ tokens | Low | Very high | Rarely recommended; consider structure-aware splitting instead |
When answers come back vague or miss a caveat, the fix is often a smaller chunk size, not a better prompt. When answers come back technically correct but missing the nuance a human reader would catch, the fix is usually the opposite: larger chunks or stronger overlap. If you are troubleshooting answer quality on an existing base, it helps to first confirm what a well-formed retrieval loop looks like end to end in How RAG Works, Step by Step: One Question Followed From Upload to Cited Answer.
Practical defaults you can start with today
If you do not want to run a chunking experiment before launching a knowledge base, these defaults work for the large majority of document sets:
- Start with structure-aware chunking if your source has real headings (Markdown, Word styles, PDF bookmarks); fall back to paragraph-based splitting otherwise
- Target 300 to 500 tokens per chunk for general documentation and policies; drop to 150 to 250 tokens for glossaries, FAQs and short reference entries
- Add 10 to 15% overlap for narrative or legal text; skip overlap for self-contained short entries
- Keep tables and code blocks atomic, never split mid-row or mid-function
- Prepend the nearest heading to each chunk so it reads correctly in isolation
- Re-test with 10 to 15 real user questions after indexing, and adjust size or overlap based on which answers come back thin or noisy
These defaults assume reasonably clean source files. Messy PDFs, scanned documents or exports with broken formatting need cleanup before chunking even starts, which is covered in How to Prepare Documents for AI: Writing a Corpus That RAG Answers Well.
If you would rather skip the pipeline work entirely, Kopik applies structure-aware extraction and hybrid full-text plus semantic indexing automatically when you upload PDF, Word, text or Markdown files, so you get sensible chunk boundaries without tuning a splitter yourself. See How to Build a Knowledge Base From Your Documents (PDF, Word) in Minutes for the upload-to-query flow.
Stop tuning a chunker, start querying
Upload your documents to Kopik and get hybrid search and cited answers out of the box, no chunking pipeline to build or maintain.
When to revisit your chunking strategy
Chunking is not a one-time decision. Revisit it when the document mix changes (adding spreadsheets or code to a base that was all prose), when user questions shift toward short factual lookups versus long explanatory ones, or when you notice a pattern of answers citing the wrong section even though the right section exists in the base. In those cases, re-chunking with a different size or switching from fixed-size to structure-aware splitting is usually faster than trying to fix the problem with prompt engineering alone. You can browse existing public knowledge bases in the catalogue to see how different creators structure similar document types, or query programmatically once your base is tuned using the API and MCP docs.
Frequently asked questions
What is the best chunk size for RAG?
There is no universal best size. 300 to 500 tokens is a solid default for general documentation and policies, 150 to 250 tokens works better for short reference entries like FAQs or glossaries, and 600 to 1,000 tokens suits narrative or legal content where surrounding context changes the meaning of a clause.
Should I always use overlap between chunks?
No. Overlap of 10 to 20% helps dense, narrative or legal text where a key point can sit right at a boundary. For short, self-contained chunks like FAQ pairs or product descriptions, overlap just duplicates content without adding value.
What's the difference between semantic chunking and fixed-size chunking?
Fixed-size chunking cuts text at a predetermined length regardless of topic, which is fast and predictable. Semantic chunking uses embedding similarity to cut where the topic actually shifts, which improves retrieval precision on long, meandering documents but costs an extra embedding pass before indexing.
Does chunk size affect hallucination in RAG answers?
Indirectly, yes. Chunks that are too small can strip out the context a model needs, leading it to fill gaps with assumptions. Chunks that are too large dilute retrieval precision, so the model may receive a passage that is only loosely related to the question. Both situations increase the risk of an answer drifting from the source text.
How should I chunk tables and code blocks for RAG?
Keep them atomic. Never split a table mid-row or a code block mid-function. Attach the table's caption or the code block's preceding heading so the chunk is understandable without the rest of the document, even if that pushes it slightly over your normal size target.
Do I need to build my own chunking pipeline to use RAG?
Not necessarily. Services like Kopik handle extraction, structure-aware chunking and hybrid indexing automatically when you upload a document, so you can query it right away through the website, a REST API, or an MCP client without writing a custom splitter.
Get the Kopik newsletter
New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.
By subscribing you agree to receive our newsletter. We never share your address.