Guide

How to Reduce AI Hallucinations With Grounded Answers

The Kopik team8 min read

The most reliable way to reduce AI hallucinations is to stop asking the model to answer from memory. Retrieve the relevant passages from trusted documents first, require the model to answer only from them with citations, and let it refuse when the sources don't cover the question. Then measure the result with a small set of test questions whose answers you already know. This guide explains why hallucinations happen and walks through each of those layers.

Why do AI models hallucinate?

A language model is trained to predict plausible text. When you ask it a question, it does not look anything up: it generates the sequence of words that best fits the question, based on patterns learned during training. Most of the time the most plausible answer is also the correct one. When it isn't, you get a hallucination: a fluent, confident statement that is simply false.

Several situations make this much more likely:

  • Missing knowledge. Your internal policies, a state regulation updated last month or a niche technical spec were never in the training data, so the model fills the gap with something that sounds right.
  • Precise details. Exact figures, dates, section numbers, case names and product references are where models slip most often, because many similar values look equally plausible.
  • Pressure to answer. Models are tuned to be helpful. Without an explicit way out, they tend to produce an answer rather than admit uncertainty.
  • Ambiguous questions. If a question could mean two things, the model picks one silently and answers it with full confidence.
  • Long, multi-step reasoning. An early small error gets carried forward and amplified in the final answer.

The key point: hallucination is a property of how generation works, not a bug that a newer model version will fully remove. Better models hallucinate less, but none hallucinate never. So the fix has to come from the system around the model.

What are grounded answers?

A grounded answer is one built from specific source material supplied at question time, rather than from the model's memory. The model becomes a reader and a writer, not an oracle. This is the core idea of retrieval-augmented generation, explained in our guide what is RAG.

Grounding changes the failure mode. An ungrounded model that doesn't know something invents it. A grounded model that doesn't find something in its sources can be told to say so. And when it does answer, every claim can be traced back to a passage a human can read.

Ungrounded vs grounded answers

Answer from memoryGrounded answer
Source of factsTraining data, frozen at a cutoff dateDocuments retrieved at question time
When knowledge is missingInvents a plausible answerSays the sources don't cover it
VerifiabilityNone, you must trust itCitations point to exact passages
Keeping it currentWait for a new modelUpdate the documents
Typical residual errorsInvented factsMisread or incomplete passages

Five layers that reduce AI hallucinations

No single trick solves the problem. What works in practice is stacking several layers, each catching what the previous one missed.

1. Curate the sources

Grounding is only as good as what you ground on. Remove outdated versions, contradictory drafts and duplicates. Make sure PDFs contain real text, not scanned images. Put dates and scope at the top of each document so the model knows which version applies. Our checklist on preparing documents for AI covers this in detail.

2. Retrieve the right passages

Most "hallucinations" in RAG systems are actually retrieval misses: the right passage existed but never reached the model, so it answered from memory or from a loosely related passage. Hybrid search (exact keywords plus a broader semantic expansion of the question) helps catch both precise identifiers like "Form 941" and paraphrased questions. Chunk size, overlap and the number of passages retrieved all matter; we break them down in how a RAG pipeline works.

3. Instruct the model to answer only from the sources

The answer prompt should say, plainly: use only the passages provided, cite the passage number after each claim, and do not add facts from general knowledge. It should also tell the model how to behave when passages conflict (report both, with their dates) and when they are only partially relevant (answer the covered part, flag the rest).

4. Require citations

Citations do two jobs. They push the model to stay close to the text, because every sentence must point somewhere. And they let a human check the answer in seconds by opening the cited passage. A good citation points to a specific passage, not just a document title. An answer with no citation for a key claim is a red flag worth flagging in the interface.

5. Allow, and reward, refusal

This is the layer most teams forget. If the retrieved passages don't answer the question, the right response is "the sources don't cover this", ideally with what they do cover. Make this an explicit option in the instructions, and treat a correct refusal as a success in your evaluations, not a failure. A system that always answers will always hallucinate on some share of questions.

A refusal is a feature

Users sometimes complain when an assistant says "I don't know". Explain upfront that the assistant only answers from your documents. In regulated work (HIPAA questions in healthcare, employment rules, contracts), a clear refusal is far cheaper than a confident wrong answer.

How to evaluate hallucinations before you ship

You can't improve what you don't measure, and a few impressive demos prove very little. Build a small evaluation set instead:

  1. Write 30 to 50 real questions that users actually ask, each with the expected answer and the document section that contains it.
  2. Add 10 to 15 questions the sources cannot answer. These test refusal. The correct result is a clear "not covered", never an invented answer.
  3. Check retrieval separately. For each question, did the expected passage appear among the retrieved ones? If not, fix retrieval or the documents before touching the prompt.
  4. Grade each answer on three criteria: correct, fully supported by the cited passages, and correctly refused when out of scope.
  5. Rerun the set after every change to documents, chunking, retrieval settings or prompts, and compare.

Track two numbers above all: the share of answers containing a claim not supported by any cited passage, and the share of out-of-scope questions that got an invented answer instead of a refusal. Those are your hallucination rates. A language model can help grade at scale, but have a human review a sample, especially the failures.

For a broader risk framework, the NIST AI Risk Management Framework treats validity and reliability as core characteristics of trustworthy AI and is a useful reference for documenting your testing.

What grounding does not fix

Grounded answers reduce hallucinations a lot; they don't eliminate them. Keep these limits in mind:

  • Misreading. The model can still misquote a number or merge two passages. Citations make this detectable, not impossible.
  • Wrong sources. If a document is outdated or wrong, the answer will be faithfully wrong.
  • Whole-corpus questions. "List every contract with an auto-renewal clause" requires reading everything, not retrieving a few passages.
  • Over-trust. Citations make answers look authoritative. For high-stakes decisions, someone must open the source.

Getting grounded answers without building the pipeline

Building all of this yourself means a parser, a chunker, a search index, an answer prompt and an evaluation loop. That is a solid engineering project. If you mostly need grounded answers on your own documents, Kopik does the pipeline for you: you upload PDF, Word, text or Markdown files, and Kopik extracts, chunks and indexes them for hybrid search. Answers are written from the documents only, with numbered cited passages, and say so when the base doesn't contain the answer.

You can also skip the written answer and get raw passages (the "passages" mode), which suits AI agents that prefer to reason on the evidence themselves. Bases can stay private or be published in the catalog of knowledge bases, and they are reachable from the site, a REST API or an MCP server usable from Claude, Cursor or ChatGPT. Setup details are in the developer documentation.

Get answers grounded in your own documents

Create a knowledge base for free, upload your files and ask questions that come back with cited passages.

Quick checklist

  • Sources are current, deduplicated and text-based.
  • Retrieval is tested on its own: the right passage shows up for real questions.
  • The prompt says "answer only from these passages" and "cite each claim".
  • Refusal is explicitly allowed and counted as a success when correct.
  • An evaluation set with known answers and out-of-scope questions runs after every change.

Frequently asked questions

Can AI hallucinations be completely eliminated?

No. Hallucination comes from the way language models generate text, so some residual risk always remains. Grounding answers in retrieved sources, requiring citations and allowing refusal reduces it sharply and makes the remaining errors much easier to spot.

Does RAG stop hallucinations?

RAG reduces them significantly but does not stop them on its own. If retrieval misses the right passage, or the prompt doesn't forbid outside knowledge, the model can still invent. RAG works best combined with strict answer instructions, citations, refusal and evaluation.

Why do citations reduce hallucinations?

Requiring a citation for each claim keeps the model close to the retrieved text, and it lets a reader verify any statement in seconds by opening the cited passage. Claims without a citation become easy to spot and challenge.

How do I measure hallucinations in my AI assistant?

Build a test set of real questions with known answers, plus questions your sources cannot answer. Measure how often answers contain claims not supported by the cited passages, and how often out-of-scope questions get an invented answer instead of a refusal. Rerun it after each change.

Is a bigger model enough to fix hallucinations?

Newer, larger models tend to hallucinate less, but they still lack your private documents and anything that changed after their training. Grounding the model in current sources does more for factual accuracy than switching models.

Get the Kopik newsletter

New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.

By subscribing you agree to receive our newsletter. We never share your address.