Guide

What Is RAG (Retrieval-Augmented Generation)? A Practical Guide

The Kopik team11 min read

Large language models are impressive writers, but they only know what they saw during training, and they will happily invent an answer when they don't know. So what is RAG, and why has it become the default way to make AI useful on real documents? Retrieval-augmented generation is a simple idea: before the model answers, you look up the relevant passages in a trusted set of documents and hand them to the model as evidence. The answer is then grounded in your sources, up to date, and verifiable. This guide explains how RAG works, when it is the right tool, where it breaks, and how to build your first knowledge base.

What is RAG, in plain terms?

RAG stands for retrieval-augmented generation. The term was popularized by a 2020 research paper from Facebook AI Research, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, which combined a search component with a text generator. Today the phrase covers any system where an AI model answers a question using information fetched at question time, rather than relying only on what it memorized.

Think of it as an open-book exam. A model without RAG is a student answering from memory: fluent, confident, sometimes wrong. A model with RAG is the same student allowed to open the right textbook to the right page before writing. The student still does the reasoning and the writing, but the facts come from the book, and the student can point to the page.

The set of documents the system searches is usually called a knowledge base. It can be a company handbook, a collection of contracts, product documentation, regulatory texts, course material, research notes: anything written down that you want an AI to answer from.

Why do language models need RAG?

General-purpose models have three structural limits that no amount of clever prompting fully fixes:

  • They don't know your documents. Your internal procedures, client files or niche expertise were never in the training data.
  • Their knowledge is frozen. A model is trained up to a cutoff date. Anything that changed afterwards (a new law, a revised price list, an updated policy) is invisible to it.
  • They can hallucinate. When a model lacks information, it may still produce a plausible-sounding answer. Without sources, you can't tell a correct answer from an invented one.

RAG addresses all three at once. The documents supply the missing knowledge, updating them updates the answers instantly, and because the answer is built from specific passages, the system can cite them so a human can check.

How does RAG work? The three steps

Every RAG system, from a weekend prototype to an enterprise platform, follows the same broad sequence. Part of it happens once, when documents are added; the rest happens every time someone asks a question.

1. Indexing: preparing the documents

Text is extracted from each file (PDF, Word, web pages, and so on), then split into smaller passages called chunks. Chunks are needed because a model can only read a limited amount of text per question, and because a precise passage is a better piece of evidence than a whole 80-page document. Each chunk is then stored in a search index so it can be found quickly later.

2. Retrieval: finding the relevant passages

When a question arrives, the system searches the index and returns a handful of the most relevant chunks. This is the step that makes or breaks a RAG system: if the right passage isn't retrieved, even the best model cannot answer correctly.

3. Generation: writing a grounded answer

The retrieved passages are placed in the model's prompt along with the question and instructions such as "answer only from these sources and cite them". The model writes the answer, ideally with references pointing back to the passages it used.

Each of these steps involves real design choices: chunk size and overlap, the type of search, how many passages to return, how to phrase the instructions. We break them down in detail in how a RAG pipeline works under the hood.

Keyword search, vector search, hybrid: how retrieval finds passages

There are two main families of retrieval, and many production systems combine them.

The main retrieval approaches in RAG

ApproachHow it matchesStrengthsWeak spots
Keyword / full-textShared words, with stemming and rankingExact terms, codes, article numbers; transparentMisses synonyms unless the query is expanded
Vector (embeddings)Similarity of meaning in a numeric spaceParaphrases, fuzzy questionsCan miss exact identifiers; harder to debug
HybridBoth, with merged rankingsCovers both casesMore moving parts to tune

Keyword search is often underrated. Mature full-text engines, such as the one built into PostgreSQL, handle word stems (so "terminate" matches "termination") and rank results by relevance. Its main gap, synonyms, can be closed by asking a language model to rewrite the question into several keyword variants before searching. Vector search shines when users phrase things very differently from the documents. Neither is universally better: the right choice depends on your documents and on how people ask questions.

Key RAG vocabulary

You will meet the same handful of terms in every RAG discussion. Here is what they mean in practice.

RAG glossary

TermMeaning
Chunk / passageA short excerpt of a document, the unit that gets retrieved
OverlapText repeated between consecutive chunks so ideas aren't cut in half
IndexThe searchable store of all chunks
Top-kHow many passages are retrieved per question
Query expansionRewriting the question into extra keywords or synonyms before searching
GroundingConstraining the answer to the retrieved sources
CitationA reference from the answer back to the passage it relies on

Two of these settings deserve early attention. Chunk size is a trade-off: small chunks are precise but can lose context, large chunks keep context but dilute relevance. Top-k is another: retrieving too few passages risks missing the answer, too many adds noise and cost. There is no universal right value; test with your own documents and questions.

What is RAG used for? Common use cases

RAG is worth it whenever people repeatedly ask questions whose answers already exist somewhere in writing. Typical examples:

  • Internal knowledge: HR policies, onboarding guides, IT procedures, answered without pinging a colleague.
  • Customer support: answers drawn from product documentation and help articles, with links to the source.
  • Legal, HR and compliance: navigating regulations, collective agreements or contract clauses, where citations are mandatory. See our guide on RAG for legal, HR and compliance teams.
  • Expert knowledge products: a consultant, trainer or author packaging their know-how as a knowledge base others can query, sometimes for a fee.
  • AI agents: giving an assistant (Claude, ChatGPT, Cursor or your own agent) a reliable source of truth it can consult mid-task, for example through MCP.

RAG vs fine-tuning vs long prompts

RAG isn't the only way to make a model "know" your content. The two main alternatives are fine-tuning (further training the model on your data) and simply pasting documents into a long prompt.

Three ways to bring your knowledge to an AI

CriterionRAGFine-tuningLong prompt
Update contentAdd or remove a fileRetrainEdit the prompt
Cites sourcesYes, nativelyNoPossible, manually
Scales to many documentsYesYesLimited by context size
Best forFacts and referencesStyle, format, behaviorOne-off, small corpora

For factual questions on business documents, RAG is almost always the better starting point: it is cheaper, easier to keep current, and auditable. Fine-tuning is better at teaching a model a style or a task format than at teaching it facts. We compare the two in depth in RAG vs fine-tuning: which one for your business documents.

The limits and pitfalls of RAG

RAG reduces hallucinations; it does not eliminate them. Knowing where it fails helps you design around it.

  • Garbage in, garbage out. Outdated, contradictory or badly structured documents produce weak answers. Scanned PDFs without a text layer can't be read at all.
  • Retrieval misses. If the relevant passage uses different words from the question, or is split awkwardly across chunks, it may never reach the model.
  • Questions that need the whole corpus. "Summarize every contract signed this year" requires reading everything, not retrieving a few passages. RAG is built for targeted questions.
  • Tables and figures. Complex tables, charts and images often lose their structure when text is extracted.
  • Over-trust. Citations make answers look authoritative. Users still need to click through on anything high-stakes.

The cheapest quality upgrade

Most RAG quality problems are document problems. Clear headings, one topic per section, explicit dates and defined terms do more than any model upgrade. Our checklist on how to prepare documents for AI and RAG walks through it.

How to build a RAG knowledge base

You can build a RAG system yourself with open-source components, or use a platform that handles the pipeline. Either way, the steps are the same:

  1. Define the scope. Which questions should the base answer, and for whom? A focused base answers better than a catch-all.
  2. Collect and clean the documents. Remove duplicates and outdated versions; make sure PDFs contain selectable text.
  3. Index them. Extract text, chunk it, and store it in a search index (full-text, vector or both).
  4. Configure the answer step. Choose the model, the number of passages retrieved, and instructions that require citations and allow "I don't know".
  5. Test with real questions. Write 20 to 30 questions your users actually ask, check each answer against the source, and fix the documents where it fails.
  6. Maintain it. Replace documents when they change; remove what is obsolete.

Doing all of this from scratch means gluing together a parser, a chunker, a database, a search layer and a model API. That is a good learning project, but for most people the value is in the documents, not the plumbing. Our step-by-step guide shows how to build a knowledge base from your documents in minutes.

How Kopik puts RAG to work

Kopik is the library of expert knowledge bases for AI, made in France. Creating a base is free. You upload documents (PDF with a text layer, Word .docx, and text formats such as .txt, .md, .csv, .tsv, .json, .html or .xml; up to 4 MB per file) or paste text, and the base is built automatically:

  • Text is extracted and split into overlapping passages of about 1,200 characters, indexed as soon as the file is uploaded.
  • Passages are indexed with PostgreSQL full-text search (no vector database), tuned to the base's document language: English, French, German, Spanish, Italian, Portuguese, Dutch, or a neutral mode for mixed content.
  • At question time, a small, fast model expands the question into keywords and synonyms in the documents' language, so a question asked in another language still finds the right passages. The 8 best passages are retrieved, and a language model writes the answer from those passages only, with numbered citations [1][2]. If the base doesn't contain the answer, it says so.
  • A passages mode returns just the relevant excerpts, without a written answer, for agents that prefer to reason on their own.

Each base can be public (listed in the catalogue of knowledge bases), unlisted (link only) or private. A private base is effectively your own private RAG: you can query it for free from the website, the REST API or MCP with your API key, which makes it a ready-made knowledge backend for your AI agents, with no infrastructure to run. We cover that setup in using a private knowledge base as a RAG for your AI agents. Public bases join the library: subscribers chat with them on the website, and agents pay per request over the API and MCP, from a few cents. Creators receive a fixed share of each subscriber question and 70% of each paid request, which turns specialist documentation into a product. If that interests you, read how to share your expertise as a knowledge base.

Bases can be queried on the website, through a REST API, or from AI agents (Claude, Cursor, ChatGPT or your own) via an MCP server. Setup is covered in connecting a knowledge base to an AI agent with MCP and in the developer documentation.

Turn your documents into a RAG knowledge base

Upload your PDFs or Word files and get a base that answers with citations, ready to use privately, share by link or publish.

Is RAG right for your project? A quick checklist

  • The answers already exist in written documents.
  • People ask targeted questions, not "read everything and summarize".
  • Being able to check the source matters.
  • The content changes over time and must stay current.
  • The documents are text-based (or can be converted to text).

If you tick most of these boxes, RAG is very likely the right approach. If your main goal is to change how a model writes or behaves rather than what it knows, look at fine-tuning or better prompting first.

Where to go next

New to the topic? Start with building a knowledge base from your documents, then dig into the RAG pipeline once you want to understand the trade-offs.

Frequently asked questions

What does RAG stand for in AI?

RAG stands for retrieval-augmented generation. It is a technique where an AI system first retrieves relevant passages from a set of documents, then passes them to a language model that writes an answer based on those passages. The result is grounded in your sources and can cite them.

Does RAG eliminate hallucinations?

No, but it reduces them significantly and makes them easier to detect. Because the answer is built from specific retrieved passages, the model can cite its sources and be instructed to say it doesn't know when nothing relevant is found. Users should still check the cited passages for high-stakes decisions.

Do I need vector embeddings to build a RAG system?

No. Embeddings are one way to retrieve passages, but keyword and full-text search are also valid and work very well for documents with precise terminology. Many systems combine both, and full-text search can be strengthened by having a language model expand the question into synonyms before searching.

Is RAG better than fine-tuning?

For answering factual questions from business documents, RAG is usually the better choice: it is cheaper, easy to update and can cite sources. Fine-tuning is better suited to teaching a model a style, tone or output format. The two can also be combined.

What kinds of documents work best for RAG?

Text-based documents with clear structure work best: headings, one topic per section, explicit definitions and dates. PDFs must contain a real text layer; scanned image-only files need OCR first. Complex tables and diagrams often lose structure during extraction, so key information should also appear in prose.

Can I build a RAG knowledge base without coding?

Yes. Platforms such as Kopik handle extraction, chunking, indexing and answer generation for you: you upload your files and the base is ready to query on the web, by API or from AI agents via MCP. On Kopik, creating a base is free, and querying your own private base is free too. Coding is only needed if you want to build and tune your own pipeline.

Get the Kopik newsletter

New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.

By subscribing you agree to receive our newsletter. We never share your address.