Under the hood

Chunking, indexing, retrieval: how a RAG pipeline works under the hood

The Kopik team10 min read

Retrieval-augmented generation sounds like one feature, but a RAG pipeline is really a chain of small, very concrete steps: extract the text, cut it into passages, index those passages, find the right ones when a question comes in, and hand them to a language model with strict instructions. Each step has knobs, and each knob changes the quality of the final answer. This guide walks through the whole chain, explains the trade-offs at every stage, and shows where most pipelines quietly go wrong.

If you are new to the concept itself, start with our practical guide to RAG. Here we assume you know the idea (give the model the right excerpts instead of hoping it remembers) and want to understand the machinery.

What are the stages of a RAG pipeline?

Every RAG system splits into two phases. The ingestion phase runs once per document, when you add it: it turns files into searchable passages. The query phase runs on every question: it finds relevant passages and generates an answer from them. Keeping the two phases in mind helps a lot when debugging, because a bad answer is almost always caused by one specific stage.

The two phases of a RAG pipeline

StagePhaseWhat it doesTypical failure
ExtractionIngestionTurns PDF, Word, HTML into plain textScanned PDF with no text layer
ChunkingIngestionSplits text into passagesPassages cut mid-idea
IndexingIngestionMakes passages searchableWrong language settings
RetrievalQueryFinds the best passagesRight passage ranked too low
GenerationQueryWrites a grounded answerAnswer goes beyond the sources

Step 1: text extraction, the unglamorous foundation

Nothing downstream can fix text that was never extracted. Formats like .txt, .md, .csv or .json are easy: the text is already there. Word files (.docx) are structured XML and extract cleanly. PDFs are the tricky case: a PDF is a set of drawing instructions, not a document with paragraphs. Extraction tools rebuild the reading order from character positions, which usually works for text-based PDFs but can struggle with multi-column layouts, footnotes or complex tables.

The biggest trap is the scanned PDF. If a document was scanned as an image and never went through OCR (optical character recognition), there is simply no text to extract. The pipeline sees an empty file. A quick test: open the PDF and try to select a sentence with your cursor. If you can't, the file needs OCR before it can go into any knowledge base.

Check your extraction first

When answers look wrong, look at the extracted text before blaming the model. Headers repeated on every page, broken tables and missing sections are extraction problems. Our guide on preparing documents for AI covers how to fix them at the source.

Step 2: RAG chunking, or how to cut documents into passages

Language models have a limited context window, and even when that window is large, stuffing it with whole documents is slow, expensive and dilutes attention. So the text is split into chunks (also called passages): pieces small enough to be retrieved individually, large enough to make sense on their own. Chunking is the stage with the most influence on answer quality, and the one most often configured by default without thought.

Choosing a chunk size

There is no universal best size, but the trade-off is always the same. Small chunks are precise: when retrieved, almost everything in them is relevant. But they lose context, and a rule split across three chunks may only be partially retrieved. Large chunks keep context but bring noise, and fewer of them fit in the prompt. For prose documents such as policies, contracts or guides, passages of roughly one to a few paragraphs (around a thousand characters, give or take) are a common and sensible middle ground.

Why chunk overlap matters

Overlap means each chunk repeats the end of the previous one. It protects against a key sentence falling exactly on a boundary: with overlap, that sentence appears whole in at least one passage. The cost is some duplication in the index. A modest overlap, a sentence or two, is usually enough.

Split on structure, not on character count

Naive chunkers cut every N characters, often mid-word or mid-sentence. Better chunkers are recursive: they try to split on paragraph breaks first, then on sentence boundaries, and only fall back to hard cuts when a single sentence is too long. This keeps each passage a coherent unit of meaning.

  • Fixed-size chunking: simple and predictable, but ignores meaning. Fine for uniform text.
  • Recursive chunking (paragraphs, then sentences): the pragmatic default for most business documents.
  • Structure-aware chunking (by headings, articles, sections): excellent when documents are well structured, such as legal codes or manuals.
  • Semantic chunking (split where the topic shifts, detected with a model): elegant but costlier and harder to predict.

Step 3: indexing, full-text search vs embeddings

Once you have passages, you need a way to find the right ones fast. There are two main families of indexes, and many production systems combine them.

Lexical (full-text) indexing

Full-text search is the technology behind classic search engines. Each passage is broken into words, words are reduced to their stem (so that terminate, terminated and termination match), stop words are dropped, and an inverted index maps each stem to the passages containing it. Ranking functions then score passages by how well they match the query terms. Databases such as PostgreSQL ship this natively, with language-specific dictionaries: see the PostgreSQL full-text search documentation.

Strengths: exact terms, article numbers, product codes and proper names match reliably; results are explainable; no extra model or vector store is required. Weakness: vocabulary mismatch. If the document says dismissal and the user asks about getting fired, a pure lexical search may miss it.

Vector (embedding) indexing

An embedding model converts each passage into a vector of numbers that captures its meaning. Questions are embedded the same way, and the index returns passages whose vectors are closest. This handles synonyms and paraphrases naturally. The trade-offs: you need an embedding model and a vector index, re-embedding is required if you change models, and exact identifiers (a clause number, a SKU) are sometimes matched less reliably than with keywords.

Full-text vs embeddings vs hybrid

ApproachBest atWeak atInfrastructure
Full-text (lexical)Exact terms, codes, namesSynonyms, paraphrasesAny SQL database with FTS
Embeddings (vector)Meaning, paraphrasesExact identifiersEmbedding model + vector index
HybridBoth, merged rankingsMore moving partsBoth of the above

Hybrid search runs both and merges the rankings. It is often the strongest option on paper, but it doubles the components to maintain. A well-tuned lexical index with a smart query step, described below, is a surprisingly strong baseline for professional documents, which tend to use precise, consistent vocabulary.

Step 4: retrieval, from a question to the right passages

Retrieval is where the question meets the index. Users rarely phrase questions the way documents are written, so good pipelines transform the question before searching.

  1. Query expansion: a small, fast model rewrites the question into keywords, synonyms and related terms (getting fired becomes dismissal, termination, notice period). This closes most of the vocabulary gap of lexical search.
  2. Search: the expanded query runs against the index and returns scored candidates.
  3. Top-k selection: only the best k passages are kept. Too few and you miss context; too many and you add noise and cost. Values between 5 and 10 are common for prose.
  4. Optional reranking: a dedicated model re-scores the candidates against the original question for finer ordering.

When retrieval finds nothing

A good pipeline says so. If no passage matches, the honest answer is that the documents don't cover the question, not a guess from the model's general knowledge.

Step 5: generation with citations

The retrieved passages are placed in the prompt, numbered, alongside the question and a system instruction: answer only from these sources, cite them, and say when the information isn't there. The model then writes an answer with references such as [1] or [3] pointing to specific passages. Citations are what make RAG trustworthy: the reader can click through and verify the exact excerpt.

Some consumers don't want a written answer at all. An AI agent that reasons on its own may prefer the raw passages and do the synthesis itself. That is why many RAG services expose a passages-only mode next to the answer mode, which is also what makes it easy to plug a knowledge base into AI agents such as Claude or Cursor with MCP.

How Kopik implements its RAG pipeline

To make this concrete, here is exactly what happens when you create a base on Kopik. Upload PDFs with a text layer (scanned image PDFs can't be read), Word .docx files or text formats (.txt, .md, .csv, .tsv, .json, .html, .xml), up to 4 MB per file, or simply paste text. The text is extracted and split into passages of about 1,200 characters with about 200 characters of overlap, cutting on paragraphs first, then sentences. Passages are indexed as soon as the upload finishes, and the base grows with every new document.

Passages are indexed with PostgreSQL full-text search, tuned to the document language you set for the base (English, French, German, Spanish, Italian, Portuguese, Dutch, or a neutral mode for mixed content): stemming and stop words follow that language. Kopik uses lexical retrieval, not vector embeddings. At question time, a small, fast language model first expands the question into keywords and synonyms in the documents' language, so a question asked in another language still finds the right passages. The top 8 passages are retrieved, and a language model writes the answer from those passages only, with numbered citations such as [1] and [2]. If the base doesn't contain the answer, it says so rather than guessing. A passages mode returns the excerpts without a written answer, and questions that find nothing are neither billed nor counted.

See the pipeline on your own documents

Upload a few files and ask questions: every answer comes with citations to the exact passages used. You can add or remove documents at any time.

Common RAG pipeline pitfalls and how to fix them

SymptomLikely causeFix
Answer says info is missing, but it existsExtraction failed or passage ranked too lowCheck extracted text; improve wording or headings
Answer mixes two different rulesChunks too large or lacking contextAdd clear headings; split long documents by topic
Half of a rule is missingRule split across chunksKeep rules in one paragraph; rely on overlap
Outdated information returnedOld and new versions both indexedRemove superseded documents
Synonym questions failVocabulary mismatchQuery expansion or a glossary in the corpus

Notice how many fixes live in the documents rather than in the code. That's the practical lesson of every RAG project: corpus quality beats clever tuning. If you are deciding whether RAG is even the right approach for your content, our comparison of RAG vs fine-tuning covers the alternatives, and our guide to building a knowledge base from your documents walks through the practical setup.

A checklist to evaluate any RAG pipeline

  • Can you see the extracted text for each document?
  • Is chunking structure-aware (paragraphs, sentences) with some overlap?
  • Does the index handle your documents' language (stemming, stop words)?
  • Is there a query transformation step for vocabulary mismatch?
  • Does every answer cite its sources, and can you open them?
  • Does the system admit when it finds nothing instead of improvising?
  • Can agents get raw passages, not only written answers?

Want to see how other creators structure their content? Browse the public catalogue of knowledge bases and test a few questions, or read the API and MCP documentation to query bases from your own code.

Frequently asked questions

What is a RAG pipeline?

A RAG pipeline is the sequence of steps that lets a language model answer from your documents: text extraction, chunking into passages, indexing, retrieval of the most relevant passages for each question, and generation of an answer grounded in those passages, ideally with citations.

What is the best chunk size for RAG?

There is no universal value. For prose documents, passages of roughly one to a few paragraphs, around a thousand characters, with a small overlap of a sentence or two, are a sensible starting point. Test with real questions and adjust if answers lack context or contain noise.

Do I need embeddings and a vector database for RAG?

No. Embeddings are one way to retrieve passages, but full-text search with stemming and an LLM query expansion step works well, especially for professional documents with precise vocabulary. Hybrid systems combine both at the cost of extra infrastructure.

Why does my RAG system miss information that is in the documents?

The most common causes are failed extraction (for example a scanned PDF without OCR), a key sentence split across passages, or vocabulary mismatch between the question and the document. Check the extracted text first, then the wording and structure of the source.

What does chunk overlap do?

Overlap repeats the end of each passage at the start of the next one. It ensures a sentence that falls on a boundary still appears complete in at least one passage, at the cost of slight duplication in the index.

Which retrieval method does Kopik use?

Kopik uses PostgreSQL full-text search tuned to each base's document language (stemming, stop words), not vector embeddings. A small, fast model expands each question into keywords and synonyms in the documents' language, the top 8 passages are retrieved, and a language model writes the answer from those passages only, with numbered citations, or says so when the base doesn't cover the question.

Get the Kopik newsletter

New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.

By subscribing you agree to receive our newsletter. We never share your address.