Chunking Strategies for RAG: Fixed-Size, Sentence, Semantic and Structure-Aware
Chunking strategies for RAG decide how your source documents are cut into the passages a vector database actually retrieves, and that single decision affects answer quality more than almost any other setting in the pipeline. Too large and the model drowns in irrelevant text; too small and it loses the context needed to answer properly. This guide compares fixed-size, sentence, semantic and structure-aware chunking, explains how chunk size and overlap interact, and gives practical defaults you can start from today.
What chunking actually does in a RAG pipeline
Retrieval-augmented generation works by searching a store of document fragments for the ones most relevant to a question, then feeding those fragments to a model so it can answer with grounded, citable text. If you want the mechanics behind this, How RAG Works, Step by Step walks through one query from start to finish. Chunking is the step that creates those fragments in the first place: you take a PDF, a Word document or a Markdown file and split it into smaller pieces before each piece is turned into a vector using text embeddings and stored for search.
Get chunking wrong and nothing downstream can fix it. A chunk that splits a clause in half, mixes two unrelated policies together, or runs to twelve pages will either fail to be retrieved when it should be, or get retrieved and confuse the model when it is. This is why most practical improvements to a knowledge base's accuracy come from re-chunking, not from swapping the embedding model or the vector database underneath it.
Four chunking strategies compared
Fixed-size chunking
The simplest approach: cut the text every N tokens or characters, often with a sliding overlap. It is fast, predictable and easy to implement, which is why it remains the default in many open-source RAG tutorials. Its weakness is that it is blind to meaning: a fixed cut at 500 tokens will happily slice a sentence, a table row or a numbered clause in two, leaving both resulting chunks individually useless.
Sentence-based chunking
Here you split on sentence boundaries first, then group consecutive sentences until you hit a target size. This avoids mid-sentence breaks and tends to produce cleaner, more readable passages. It still has no sense of topic, though, so a chunk can start mid-argument if the previous sentence happened to fall just over the boundary.
Semantic chunking
Semantic chunking measures how similar consecutive sentences or paragraphs are (typically using embedding similarity) and places a break wherever the topic shifts noticeably, rather than at a fixed length. This produces chunks that correspond to actual ideas, which tends to improve retrieval precision on long, discursive documents such as reports or meeting transcripts. The cost is extra compute at indexing time and chunks of uneven size, which can complicate downstream token budgeting.
Structure-aware chunking
This approach uses the document's own structure, headings, sections, numbered clauses, table boundaries, list items, as the natural split points. A contract is chunked by clause, a manual by procedure step, a policy by section. It is usually the strongest option for well-formatted business documents because it preserves the unit of meaning the author actually intended, and it is what a well-built ingestion pipeline should default to whenever structure is available.
Chunking strategies at a glance
| Strategy | Best for | Main weakness |
|---|---|---|
| Fixed-size | Quick prototypes, unstructured text | Cuts sentences and clauses arbitrarily |
| Sentence-based | Clean prose, FAQs, articles | Ignores topic shifts within a size window |
| Semantic | Long reports, transcripts, research notes | Slower to index, uneven chunk sizes |
| Structure-aware | Contracts, manuals, policies, technical docs | Needs reliable structure (headings, clauses) to work from |
Chunk size and overlap: the knobs that matter most
Once you have picked a strategy, two parameters do most of the work: chunk size and overlap. Chunk size is usually expressed in tokens, roughly three-quarters of a word in English. Small chunks (150-300 tokens) retrieve with high precision because each one is narrowly focused, but they can lack the surrounding context a model needs to interpret a figure or a clause correctly. Large chunks (800-1,200 tokens) carry more context but dilute the vector representation, so a specific fact can get buried next to unrelated material and the search engine struggles to rank it highly.
Overlap means repeating a small amount of text from the end of one chunk at the start of the next, so that a sentence split across a boundary still appears whole in at least one chunk. It is cheap insurance against the exact failure mode that fixed-size and sentence-based chunking are prone to.
A sensible starting overlap
Set overlap to roughly 10-20% of your chunk size. For 400-token chunks, that is 40-80 tokens of repeated text between neighbours. Go much higher and you waste storage and search time re-indexing the same sentences; go to zero and you reintroduce the boundary problem overlap was meant to solve.
Chunk size also interacts directly with how many passages you retrieve per question (commonly called top-k). Smaller chunks mean you typically need to retrieve more of them (6-10) to cover the same ground that 3-4 larger chunks would; larger chunks need fewer, but each wrong one wastes more of the model's context window. If you are also weighing up where to run this search, Vector Database Cost explains how chunk count and size affect your storage and query bill, since more, smaller chunks mean more vectors to store and search.
Practical defaults by document type
If you do not want to run a full evaluation before launching, these starting points cover most UK business use cases reasonably well:
- Contracts and policies: structure-aware, split by clause or numbered section, 200-500 tokens, overlap 10%, because precision on a specific clause matters more than broad context.
- Technical manuals and SOPs: structure-aware, split by step or sub-section, 300-600 tokens, overlap 15%, keeping numbered steps and their immediate instructions together.
- FAQs and support articles: one chunk per question-and-answer pair regardless of length, no overlap needed since each pair is already a self-contained unit.
- Long reports and research: semantic chunking, target 400-700 tokens, overlap 15-20%, so topic boundaries rather than arbitrary counts drive the split.
- Meeting transcripts and interviews: semantic or speaker-turn based chunking, 300-500 tokens, higher overlap (20%) because conversational context is easily lost.
- Source code and API docs: structure-aware by function, endpoint or heading, with the full signature or endpoint definition kept in a single chunk even if that means exceeding your usual size target.
When you upload documents to build a base on Kopik, the ingestion pipeline extracts the text, applies structure-aware chunking where the document format allows it, and indexes the result for hybrid search that combines full-text matching with semantic keyword expansion. For most PDF, Word and Markdown uploads this removes the need to tune chunking manually before you get a usable knowledge base.
Testing and tuning your chunking strategy
Chunking defaults are a starting point, not a final answer. The only reliable way to know whether your settings work is to test them against real questions and inspect what actually gets retrieved.
- Write down 15-20 questions a real user would ask, including a few that require combining facts from two places in the document.
- Run each question and read the retrieved passages, not just the final answer: are the right paragraphs showing up, and are they complete or truncated mid-thought?
- Check citations against the source document, as described in Document Q&A With Citations, to confirm the model is quoting the passage it was actually given rather than drifting.
- If answers are vague or generic, your chunks are probably too large or too broad; tighten the size or switch to a more structure-aware split.
- If answers miss context or contradict the source, your chunks are probably too small or badly bounded; increase overlap or chunk size, or group related clauses together.
- Re-test after every change. A strategy that works for a 5-page policy will not necessarily work for a 200-page manual from the same organisation.
This evaluation loop matters more than picking the theoretically best strategy on paper. A mediocre strategy you have actually tested against real questions will usually outperform a sophisticated one you have only read about. For a deeper look at why answers go wrong even with good retrieval, see Reducing AI Hallucinations.
Skip the chunking guesswork
Upload your documents to Kopik and let structure-aware chunking and hybrid search handle the splitting for you, then query the result on the website, via the REST API or through an MCP server from Claude, Cursor or ChatGPT.
Where chunking fits alongside the rest of your RAG setup
Chunking is one of several decisions that determine whether a knowledge base gives precise, cited answers or vague ones, alongside embedding quality, retrieval method and prompt design. If you are comparing ready-made options rather than building a pipeline from scratch, RAG as a Service Platforms Compared and RAG as a Service: What It Is and When to Use It cover how much chunking control different providers actually give you, and when it is worth building your own pipeline instead of trusting a provider's defaults. For developers who want to query programmatically once a base is chunked and indexed, the API and MCP documentation covers both REST calls and MCP server setup.
Frequently asked questions
What is the ideal chunk size for RAG?
There is no single ideal size; it depends on the document type. As a starting point, 200-500 tokens works well for dense, clause-based text like contracts, while 400-700 tokens suits narrative reports. Test against real questions and adjust rather than relying on one number for every document.
Should I always use overlap between chunks?
Overlap of roughly 10-20% of chunk size is a sensible default because it protects against sentences or clauses being split across a boundary. The main exception is when you are chunking by natural unit, such as one FAQ pair per chunk, where overlap adds little value.
Does semantic chunking always beat fixed-size chunking?
Not always. Semantic chunking tends to help on long, topic-shifting documents like reports or transcripts, but it adds indexing cost and produces uneven chunk sizes. For short, already well-structured documents such as contracts or manuals, structure-aware chunking by clause or section usually performs at least as well with less overhead.
How does chunk size affect hallucinations?
Chunks that are too large dilute the specific fact a question needs among unrelated text, making the model more likely to guess or blend information incorrectly. Chunks that are too small can strip away the context needed to interpret a figure or clause correctly. Both extremes increase the risk of an answer that sounds plausible but is not properly grounded.
Can I change the chunking strategy after a knowledge base is already indexed?
Yes, but it means re-processing the source documents with the new settings and re-indexing, since the old chunks and their vectors no longer reflect the new boundaries. It is worth testing chunking choices on a small sample of documents before applying them to a full library.
Do I need to chunk documents myself before uploading to a RAG platform?
Not necessarily. Platforms such as Kopik extract and chunk documents automatically on upload, applying structure-aware splitting where the file format allows it. Manual chunking is mainly useful when you have unusual formatting or specific accuracy requirements that automatic defaults do not meet.
Get the Kopik newsletter
New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.
By subscribing you agree to receive our newsletter. We never share your address.