How to Prepare Documents for AI: Writing a Corpus That RAG Answers Well
When a RAG assistant gives a vague or wrong answer, the first instinct is to blame the model. More often, the cause is upstream: a scanned PDF with no text, a policy that exists in three contradictory versions, or a key fact buried in a table the extractor could not read. Learning to prepare documents for AI is the highest-leverage thing you can do to improve answers. This guide walks through formats, structure, writing style and testing, with checklists you can apply to any corpus.
Why document preparation matters more than the model
A retrieval-augmented generation system works in two steps. First, it searches your documents for the passages most relevant to a question. Then a language model writes an answer based on those passages. If you need a refresher, read what RAG is and how it works.
The model can only be as good as what the search step hands it. If the right passage is never retrieved, because the text was not extracted, the wording does not match the question, or the fact is split across two distant pages, even the best model cannot answer correctly. Good preparation increases the odds that the right passage is found and that it contains everything needed to answer.
The core idea
Write each section so that it could be read on its own, out of context, and still make sense. That is essentially what a RAG system does with your text.
Step 1: Choose formats the AI can actually read
Before anything else, make sure the text can be extracted. This sounds obvious, but it is the most common silent failure.
Common formats and what to watch for
| Format | Works well when | Watch out for |
|---|---|---|
| Exported from a word processor | Scans without OCR have no text | |
| Word (.docx) | Real headings are used | Text inside images or text boxes |
| Markdown / .txt | Clean, structured text | Lost structure if pasted raw |
| HTML | Semantic headings, little clutter | Menus, footers, cookie banners |
| CSV / TSV | One record per row, clear headers | Cryptic column names, codes |
| JSON / XML | Readable field names | Deep nesting with no labels |
- Test a scanned PDF quickly: try to select and copy a sentence. If you cannot, the file is an image and needs OCR before any AI tool can read it.
- Prefer the source file: if you have the original Word document, upload it rather than a PDF export; structure survives better.
- Strip boilerplate from web pages: navigation, footers and legal banners add noise that competes with real content in search results.
- Split very large files: smaller, topic-focused files are easier to maintain and often stay within upload limits.
Step 2: Clean the corpus before you index it
A RAG corpus is not an archive. Every document you add competes for the few passages retrieved per question. Irrelevant or outdated content does not just waste space; it actively pushes good passages out.
- Remove duplicates: the same FAQ pasted into five files means five near-identical passages crowding the results.
- Keep one current version: delete superseded policies, old price lists and drafts. If history matters, keep it in a separate base.
- Date what can change: write the effective date or last update near the top of each document.
- Resolve contradictions: if two documents disagree, fix the source rather than hoping the model picks the right one.
- Remove sensitive data: personal data, credentials and confidential details should not be in a corpus unless strictly necessary and properly protected.
Step 3: Structure documents for chunking
RAG systems split documents into passages, often called chunks, before indexing them. Most splitters prefer natural boundaries: paragraphs first, then sentences. Our article on chunking, indexing and retrieval explains the mechanics. What matters for writers is that clear structure produces clean chunks.
Use descriptive headings
"Section 4" tells a search engine nothing. "Refund policy for annual subscriptions" tells it exactly what the section covers, and it shares vocabulary with the questions people will ask. Headings are cheap and they help both humans and retrieval.
Keep paragraphs focused
One idea per paragraph, and paragraphs of reasonable length. A wall of text mixing three topics produces chunks that match many questions loosely and none precisely. Short, focused paragraphs produce chunks that match one question strongly.
Keep related facts close together
If a rule and its exception sit five pages apart, they will almost certainly end up in different chunks. Put conditions, exceptions and thresholds next to the rule they modify. This is especially important for legal, HR and compliance documents, where exceptions change the answer.
Write self-contained sections
This is the single most useful writing habit for AI-ready documents. When a passage is retrieved, it arrives without the pages around it. Any reference that depends on earlier context loses its meaning.
Rewriting for self-contained passages
| Instead of | Write |
|---|---|
| "As mentioned above, it is 30 days." | "The refund window for annual plans is 30 days." |
| "This does not apply to them." | "This rule does not apply to contractors." |
| "See the previous table." | State the key value in the sentence itself. |
| "The product" | Use the product's actual name. |
| "Contact the team." | "Contact the billing team at the address in the footer of every invoice." |
- Repeat the subject: name the product, policy or population in each section, even if it feels redundant to a human reader.
- Avoid unexplained pronouns: "it", "they", "this" at the start of a paragraph are red flags.
- Spell out acronyms at least once per section, since sections are read in isolation.
- State numbers with units and scope: "5 days per year for full-time employees", not "5 days".
Step 4: Use the words your users use
Many retrieval systems, including keyword-based full-text search, rely heavily on vocabulary overlap between the question and the passage. Even systems based on embeddings benefit when the document uses the same terms as the question. If your documentation says "subscription termination" but customers ask "how do I cancel", the match is weaker than it should be.
- Include common synonyms: "cancel (terminate) your subscription", "sick leave (medical leave)".
- Add a short FAQ section to long documents, phrased the way people actually ask.
- Mirror support tickets: the wording customers use in emails is a goldmine for headings and FAQ questions.
- Define internal jargon: a one-line definition connects an internal code name to the plain-language term.
Quick win
Take your 20 most frequent questions and search your documents for the exact words used in them. Wherever a question's key terms appear nowhere in the corpus, add a sentence or FAQ entry that uses them.
Step 5: Handle tables, lists and numbers carefully
Tables are great for humans and tricky for text extraction. Complex layouts, merged cells or tables exported as images can come out as a jumble of values with no headers. A few habits help:
- Keep tables simple: one header row, no merged cells.
- Add a sentence before the table stating what it contains and for whom it applies.
- For the most important values, repeat them in prose: "The Pro plan costs €20 per month."
- For large datasets, a CSV with explicit column names is often more reliable than a table inside a PDF.
How to test whether your documents are AI-ready
Preparation is not finished until you have tested it. The good news is that testing is simple and requires no technical skills.
- Write a test set of 20 to 50 real questions, each with the expected answer and the document it should come from.
- Ask each question and check two things: is the answer correct, and does the citation point to the right passage?
- Classify failures: text not extracted, wrong document retrieved, right document but incomplete passage, or the answer is simply missing from the corpus.
- Fix the documents: add headings, rewrite for self-containment, add synonyms, remove duplicates.
- Re-run the full set after changes, and again whenever you add or update documents.
Retrieval-only testing is useful too. If your tool can return raw passages without a written answer, you see exactly what the model would receive, which makes diagnosis much faster.
Preparing documents for a Kopik knowledge base
Everything above applies to any RAG tool. Here is how it maps to Kopik specifically, so you can build a knowledge base from your documents with the right expectations.
- Formats: PDF with a text layer, up to 1,500 pages (scanned, image-only PDFs cannot be read, so run OCR first), Word (.docx), and text formats (.txt, .md, .csv, .tsv, .json, .html, .xml), or pasted text. Maximum 4 MB per file and 2 million characters per document; a base holds up to 1,000 documents and 20 million characters, and you can add or remove documents at any time.
- Chunking: text is split into passages of about 1,200 characters with roughly 200 characters of overlap, cutting on paragraphs first, then sentences. Focused paragraphs and clear headings therefore pay off directly.
- Retrieval: Kopik uses PostgreSQL full-text search, not vector search. Before searching, a small language model expands each question into keywords and synonyms in the documents' language, then the top 8 passages are retrieved. Expansion catches many rephrasings, but it is still lexical search, so using your users' vocabulary matters.
- Document language: each base has a document language setting (English, French, German, Spanish, Italian, Portuguese, Dutch, or Mixed/other). It tunes stemming and stop words, so set it to the language your documents are written in, not the language your users ask in: question expansion already bridges that gap. Use Mixed/other for multilingual corpora. Changing it re-indexes the whole base, which is allowed a few times per hour.
- Testing: questions you ask your own bases are free (fair use: 200 a day), even on a private base. The passages mode returns the retrieved excerpts without a written answer, ideal for checking what the search step finds. It is available via the REST API and MCP.
- Real questions: once the base is in use, your dashboard shows the text of the questions people ask (never who asked) and flags the ones that found nothing. That list is your best to-do list: add the missing document, or the term your users actually type.
Put your prepared documents to work
Upload your cleaned corpus, run your test questions and see which passages come back, with citations for every answer.
A well-prepared corpus is also what makes a public base worth consulting. If you plan to share your expertise as a knowledge base, document quality is your product. Browse the catalogue of public bases to see how other creators present theirs.
The AI-ready document checklist
- Text is selectable (no image-only scans).
- One current version per document, with a date.
- No duplicates or contradictory sources.
- Descriptive headings that match how people ask.
- One idea per paragraph; rules and exceptions side by side.
- Each section understandable on its own: subject named, no dangling pronouns.
- Synonyms and plain-language terms next to internal jargon.
- Simple tables with an introductory sentence; key values repeated in prose.
- Sensitive data removed.
- A test set of real questions, passed and re-run after each update.
Frequently asked questions
What is the best file format for RAG?
Clean text formats such as Markdown or well-structured Word documents are usually the most reliable, because headings and paragraphs survive extraction. PDFs work well when they have a real text layer. Scanned PDFs without OCR contain only images and cannot be read by text-based tools.
Do I need to rewrite all my documents for AI?
No. Start with the documents behind your most frequent questions and fix what testing reveals: missing headings, context-dependent sentences, duplicates and outdated versions. Targeted edits on a small part of the corpus often improve answers significantly.
How long should a section or paragraph be for RAG?
There is no universal number, but focused paragraphs that cover one idea work best. Many systems split text into passages of a few hundred to a couple of thousand characters, cutting on paragraph boundaries. Clear headings and self-contained paragraphs matter more than hitting an exact length.
Should I include old versions of documents?
Generally not in the same knowledge base. Old versions compete with current ones during retrieval and can lead to outdated answers. If you need historical versions, keep them in a separate base and label them clearly with their dates.
How do I know if my corpus is good enough?
Build a test set of real questions with expected answers and sources, then check both the answers and the citations. When most questions return the right passage and the remaining failures are understood, the corpus is ready. Re-run the test set whenever documents change.
Get the Kopik newsletter
New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.
By subscribing you agree to receive our newsletter. We never share your address.