How RAG Works, Step by Step: One Question Followed From Upload to Cited Answer
How RAG works fits in one sentence: documents are split into passages and indexed ahead of time, then at question time the system retrieves the most relevant passages, reranks them, packs them into a prompt and asks a language model to answer from them with citations. The easiest way to really understand it is to watch it happen. In this guide we follow one ordinary question, asked by an employee of a fictional 40-person company, through all seven stages of a retrieval-augmented generation pipeline, from the PDF upload to the cited answer.
The setup: one question, three documents
Our fictional company, a small engineering consultancy in California, has put three documents into a knowledge base: the 2026 employee handbook (a 38-page PDF), the 2024 employee handbook that nobody deleted, and a short benefits FAQ written in Word. An employee types this into the company assistant:
Can I roll over my unused PTO into next year?
The answer lives in section 4.3 of the 2026 handbook, titled "Paid time off: carryover". It never uses the words "roll over". The 2024 handbook has an older, different rule. Keep those two details in mind: they are exactly the kind of traps a RAG pipeline has to get past. If you want the concepts first, our pillar guide What is RAG covers the why; here we stay on the how.
Before any question: ingestion, chunking and indexing
The first three stages run once, when documents are added, and again whenever a document changes. The employee never sees them, but they decide most of the answer quality.
Step 1. Ingestion: turning files into clean text
The pipeline opens each file and extracts its text. For the handbook PDF, that means reading the text layer page by page, dropping repeated headers and footers ("Confidential, Page 14 of 38"), and keeping headings so the structure survives. The Word FAQ is easier, because headings are already marked as such. Each document also gets metadata: title, file name, and ideally a version date. That date will matter later, when the 2024 and 2026 handbooks compete.
What can go wrong: a scanned PDF with no text layer produces nothing at all, and a table flattened into a jumble of numbers loses its meaning. Our guide on preparing documents for AI and RAG lists the fixes.
Step 2. Chunking: cutting text into retrievable passages
The extracted text is split into chunks, passages of a few hundred words. Section 4.3 of the 2026 handbook becomes one chunk, which we will call chunk 87. A good chunker does two small things that pay off for our question: it prefixes the chunk with its heading path ("Employee handbook 2026 > 4. Time off > 4.3 Paid time off: carryover"), and it overlaps slightly with the next chunk so a sentence cut at the boundary is not lost. Chunk 87 now reads, in substance: employees may carry over up to 40 hours of unused PTO into the next calendar year; carried hours must be used by March 31.
Step 3. Indexing: making passages findable
Every chunk is stored in a search index. A full-text index records the words of chunk 87 in normalized form, so "carryover", "carried" and "carry" can be matched. A vector index stores an embedding, a list of numbers that captures the meaning of the passage, so a paraphrase can land near it. Many systems keep both. The trade-offs between chunk sizes, overlap and index types are covered in depth in chunking, indexing, retrieval: how a RAG pipeline works under the hood.
At question time: retrieval and reranking
Now the employee hits Enter. Everything from here happens in a second or two.
Step 4. Retrieval: casting a wide net
A pure keyword search for "roll over unused PTO" has a problem: the handbook says "carryover" and "paid time off", not "roll over" and "PTO". Two techniques rescue the question. Query expansion asks a language model to rewrite the question into extra terms (carry over, carryover, paid time off, vacation, unused hours). Vector search finds passages whose meaning is close even when the words differ. Combined, they return a candidate list, usually 20 to 50 passages. For our question, the top candidates look like this:
- Chunk 87 (2026 handbook, 4.3 Carryover): up to 40 hours, use by March 31.
- Chunk 312 (2024 handbook, 4.3 Carryover): the old rule, up to 80 hours, no deadline.
- Chunk 88 (2026 handbook, 4.4 PTO payout): accrued, unused PTO is paid out when employment ends.
- Chunk 9 (benefits FAQ): sick leave is tracked separately from PTO.
- Chunk 51 (2026 handbook, 3.2 Holidays): the list of paid company holidays.
Step 5. Reranking: putting the best evidence first
Retrieval is built for speed and recall; it is allowed to be a bit sloppy. Reranking is the careful second pass. A reranker reads the question and each candidate together and scores how well the passage actually answers it. It can also apply business rules: drop near duplicates, and prefer the most recent version of a document. In our example, chunk 87 rises to the top, chunk 88 stays because payout is closely related, the holiday list falls away, and chunk 312 is either removed or kept but clearly labeled as the 2024 version. The pipeline keeps a small number of passages, say three to eight, to send to the model.
Reranking is optional, but its job is not
Some systems use a dedicated reranking model, others rely on a well-tuned hybrid ranking plus metadata filters. What matters is that outdated or off-topic passages do not reach the model in first position. In our example, letting the 2024 rule through unlabeled would produce a confident, wrong answer.
Prompt assembly and generation with sources
Step 6. Prompt assembly: building the model's brief
The language model never searches anything itself. It receives a prompt assembled by the pipeline, which typically contains four parts:
- Instructions: answer only from the sources below, cite them as [1], [2], say clearly when the sources do not contain the answer.
- Numbered sources: [1] chunk 87 with its title and date, [2] chunk 88, [3] chunk 9.
- The question, exactly as the employee asked it.
- Optional context: the user's language, the date, or the format expected (short answer, then details).
Passages coming from documents are treated as data, not as orders. A well-built pipeline wraps them in clear delimiters so that a sentence hidden in a document ("ignore previous instructions") cannot hijack the model.
Step 7. Generation: a grounded answer the employee can check
The model writes something like: "Yes, partly. You can carry over up to 40 hours of unused PTO into the next calendar year, and those hours must be used by March 31 [1]. Any balance you have when you leave the company is paid out [2]. Sick leave is tracked separately and is not covered by this rule [3]." Each number links back to the passage, so the employee, or HR, can click and read the original wording.
Notice what the model did not do: it did not invent a rule, and it did not quote the 2024 handbook. Notice also what RAG cannot do on its own. In California, accrued vacation is treated as earned wages, so a use-it-or-lose-it policy is not allowed, though employers can cap accrual. Whether the March 31 deadline is lawful is a question for HR or counsel, not something the handbook answers. A good assistant flags that limit instead of guessing.
The whole journey in one table
One question through a RAG pipeline
| Step | What happened to our PTO question | Typical failure |
|---|---|---|
| 1. Ingestion | Handbook text extracted, headers removed, version date kept | Scanned PDF with no text, broken tables |
| 2. Chunking | Section 4.3 became chunk 87 with its heading path | Rule split across two chunks |
| 3. Indexing | Chunk stored in full-text and vector indexes | Wrong language settings, stale index |
| 4. Retrieval | Expansion matched roll over to carryover; 5 candidates | Right passage not in the candidate list |
| 5. Reranking | Current rule first, 2024 rule demoted, holidays dropped | Outdated version ranked first |
| 6. Prompt assembly | Instructions, 3 numbered sources, question | Too many passages, no citation rule |
| 7. Generation | Answer with [1][2][3] and a stated limit | Unsupported claims, missing caveat |
This table doubles as a debugging checklist. When an answer is wrong, walk backward: was the right passage in the prompt? If yes, it is a prompt or generation issue. If not, was it in the candidate list? If not, look at retrieval, then chunking, then extraction. Most bad answers trace back to steps 1, 2 or 4.
Running the same journey without building it
Building those seven stages yourself means a parser, a chunker, a database, a search layer, a model API and the glue between them. Kopik runs that pipeline for you. You create a base for free by uploading PDF, Word, text or Markdown files; Kopik extracts, chunks and indexes them for hybrid search (full-text plus semantic expansion of the keywords). Each question returns an answer grounded in your documents with the cited passages, or just the raw passages in passages mode if your own agent prefers to reason on them.
A base can stay private (only you and your API keys) or be published to the catalog, where it is paid per question and the creator sets the price and keeps 70%. You can query it on the website, through the REST API, or from Claude, Cursor, ChatGPT and other MCP clients via the MCP server at https://kopik.io/api/mcp. The developer documentation shows the API and MCP setup.
Follow your own question through a RAG pipeline
Upload a handbook, a manual or a policy and ask it a real question: you get an answer with the passages it relied on.
Test it like we just did
Pick five questions your team really asks, write down where the answer lives in the documents, then check each answer's citations against that location. It is the fastest way to see which step of your pipeline needs work.
Frequently asked questions
What are the main steps of retrieval augmented generation?
There are two phases. Ahead of time: ingestion (extracting text), chunking (splitting it into passages) and indexing (making passages searchable). At question time: retrieval (finding candidate passages), reranking (ordering them by real relevance), prompt assembly (packing instructions, sources and question) and generation (writing an answer that cites the sources).
Is reranking required in a RAG architecture?
Not strictly. Small, clean knowledge bases often work well with a good hybrid search alone. Reranking becomes valuable when the candidate list is long, when several versions of a document coexist, or when passages look similar but answer different questions.
How many passages should be sent to the model?
Usually a handful, often between three and ten. Too few and the answer may be missing; too many and the model gets distracted by noise, while cost and latency go up. The right number depends on chunk size and on how spread out answers are in your documents, so test it with real questions.
Why does a RAG system sometimes answer from an outdated document?
Because both versions were indexed and retrieval found the older one relevant. The fixes are to remove obsolete documents, store a version date as metadata, and let ranking prefer the most recent version. Clear dates inside the documents themselves also help the model notice the difference.
Does the language model search the documents itself?
No. The search is done by the retrieval layer before the model is called. The model only sees the passages the pipeline selected and placed in its prompt, which is why retrieval quality has such a large effect on the final answer.
Get the Kopik newsletter
New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.
By subscribing you agree to receive our newsletter. We never share your address.