Guide

How RAG Works, Step by Step: Following One Customer Question Through the Pipeline

The Kopik team7 min read

In short, how RAG works is this: your documents are cut into passages and indexed in advance; when a question arrives, the system retrieves likely passages, reranks them, assembles them into a prompt and asks a language model to answer only from them, with references. Diagrams of that pipeline are everywhere, so this guide takes a different route. We follow one real-world customer question, sent to a fictional online shoe retailer in Leeds, through all seven stages of retrieval-augmented generation.

Our worked example: a returns question

The retailer's support assistant draws on three documents: the current returns policy (a PDF updated in 2026), an older returns policy from 2023 that is still sitting in the shared drive, and a delivery and refunds FAQ written in Word. A customer types:

I bought trainers online three weeks ago and they don't fit. Can I still send them back?

The answer sits in section 2 of the 2026 policy, headed "Changing your mind". It talks about "returns within 30 days of delivery" and never mentions trainers or fit. The 2023 policy allowed only 14 days. Those two details are the traps we will watch the pipeline handle.

Preparation: ingestion, chunking and indexing

These first three steps happen once per document, and again whenever a document is replaced. Customers never see them, yet they set the ceiling on answer quality.

Step 1. Ingestion

The pipeline extracts the text from each file. For the PDF that means reading its text layer, stripping repeated page headers and footers, and keeping headings intact. Each document is tagged with metadata such as its title and, ideally, its effective date. A scanned policy with no text layer would yield nothing at all here, which is why searchable PDFs matter so much.

Step 2. Chunking

The text is split into chunks of a few hundred words. Section 2 of the 2026 policy becomes chunk 14. A careful chunker prefixes it with its heading path ("Returns policy 2026 > 2. Changing your mind") and lets it overlap a little with the next chunk, so that a sentence such as "items must be unworn and in their original box" is not stranded on its own. Chunk 14 now says, in substance: you can return unworn items within 30 days of delivery for a full refund.

Step 3. Indexing

Chunk 14 goes into one or two indexes. A full-text index stores normalised words so that "return", "returns" and "returned" match one another. A vector index stores an embedding, a numerical fingerprint of meaning, so a paraphrase can still land nearby. If you are weighing up the second option, our explainer on what a vector database is sets out when it is worth it.

The question arrives: retrieval and reranking

Step 4. Retrieval

The customer wrote "send them back", "trainers" and "three weeks". The policy says "return", "items" and "30 days". A literal keyword match would struggle. Two techniques bridge the gap: query expansion, where a language model rewrites the question into extra search terms (return, refund, change of mind, days after delivery, footwear), and semantic search, which matches meaning rather than wording. Together they produce a shortlist of candidates, for instance:

  1. Chunk 14 (2026 policy, Changing your mind): 30 days from delivery, unworn items.
  2. Chunk 52 (2023 policy, Changing your mind): the old 14-day window.
  3. Chunk 15 (2026 policy, How refunds work): refund to the original payment method once the return is received.
  4. Chunk 6 (FAQ): how to print a returns label.
  5. Chunk 21 (2026 policy, Faulty goods): a different procedure for defective items.

Step 5. Reranking

Retrieval is tuned to be fast and generous; reranking is the careful second reading. A reranker scores each candidate against the full question, and business rules tidy the list: drop duplicates, prefer the newest version of a document, and keep passages that answer the question rather than merely share its vocabulary. Here chunk 14 comes first, chunks 15 and 6 stay because they cover the practical next steps, the faulty goods passage is relegated (the shoes are not faulty, they simply do not fit) and the 2023 rule is removed or clearly labelled as superseded.

Why the old policy is dangerous

If the 2023 passage reached the model first and unlabelled, the assistant would tell a customer at day 21 that it is too late, which is both wrong and bad service. Deleting obsolete files is the simplest fix; dating them is the next best.

Writing the answer: prompt assembly and generation

Step 6. Prompt assembly

The language model does not search anything. It receives a prompt put together by the pipeline, usually made of:

  1. Rules: answer only from the sources provided, reference them as [1], [2], and say so plainly if they do not contain the answer.
  2. Numbered sources: [1] chunk 14, [2] chunk 15, [3] chunk 6, each with its document title and date.
  3. The customer's question, word for word.
  4. Context: language, tone of voice, today's date.

Document passages are framed as evidence, never as instructions, inside clear delimiters. That stops a stray sentence in a document from steering the model.

Step 7. Generation with sources

The model replies along these lines: "Yes. You can return unworn items within 30 days of delivery for a full refund [1], so at three weeks you are still within the window. Print a returns label from your account [3]; the refund goes back to your original payment method once we receive the parcel [2]." Every claim points to a passage the support team can open and check.

The shop's 30 days sits on top of the customer's statutory rights. Under the Consumer Contracts Regulations, online buyers generally have 14 days from delivery to cancel, then a further 14 days to send the goods back. A well-built assistant answers from the retailer's own policy and, if asked about legal rights, says that this is a separate question rather than improvising.

The seven steps at a glance, and how to debug them

One returns question through a RAG pipeline

StepWhat happenedCommon failure
1. IngestionPolicy text extracted, headers removed, date recordedImage-only PDF, mangled tables
2. ChunkingSection 2 became chunk 14 with its heading pathRule split across two chunks
3. IndexingChunk stored in full-text and vector indexesWrong language settings, index not refreshed
4. RetrievalExpansion linked send back to return; five candidatesCorrect passage missing from shortlist
5. RerankingCurrent rule first, old policy removedSuperseded document ranked top
6. Prompt assemblyRules, three numbered sources, questionToo many passages, no referencing rule
7. GenerationAnswer with [1][2][3]Claims not backed by any source

When an answer is wrong, work backwards. Was the right passage in the prompt? If so, look at the instructions and the model. If not, was it in the retrieval shortlist? If not, check the chunking and the extraction. In practice, most errors come from steps 1, 2 and 4.

Running this pipeline without building it

Assembling all seven stages yourself means a parser, a chunker, a database, a search layer and a model API, plus maintenance. That build or buy decision is discussed in RAG as a service: what it is and when to use it. Kopik is one such service: creating a base is free, you upload PDF, Word, text or Markdown files, and Kopik extracts, chunks and indexes them for hybrid search (full-text plus semantic expansion of keywords). Answers come back grounded in your documents with the cited passages, or as raw passages in passages mode.

Bases can be private, visible to you and your API keys only, or public in the catalogue, where they are paid per question and the creator sets the price and keeps 70%. They can be queried on the website, via the REST API, or from MCP clients such as Claude, Cursor or ChatGPT; see the developer documentation. Under UK GDPR, check what personal data your documents contain before uploading them, as you would with any processor.

Put your own question through the pipeline

Upload a policy or a manual, ask a real customer question and see the passages the answer is built on.

Frequently asked questions

What are the stages of a RAG pipeline?

Preparation covers ingestion (extracting text), chunking (splitting it into passages) and indexing (making passages searchable). At question time the pipeline runs retrieval (finding candidates), reranking (ordering them by real relevance), prompt assembly (rules, sources, question) and generation (an answer that references its sources).

What is the difference between retrieval and reranking?

Retrieval quickly pulls a broad shortlist from the whole index and favours not missing anything. Reranking reads that shortlist more carefully, scoring each passage against the full question and applying rules such as preferring the latest document version, so the best evidence reaches the model first.

Why does my RAG assistant quote an outdated policy?

Usually because both versions are indexed and the older one matched the question well. Remove superseded files, record an effective date for each document and make sure ranking favours the most recent version. Writing the date inside the document itself also helps.

Does RAG need a vector database?

No. Full-text search with query expansion works well, especially for documents with precise terminology, and many systems combine it with vector search. A vector database becomes more useful when users phrase questions very differently from the documents.

How can I tell which step caused a wrong answer?

Trace it backwards. Check whether the correct passage appeared in the prompt; if it did, the issue lies with instructions or generation. If it did not, check whether retrieval found it, then whether chunking split it badly, then whether the text was extracted correctly in the first place.

Get the Kopik newsletter

New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.

By subscribing you agree to receive our newsletter. We never share your address.