How-to

Chat With PDF Files Reliably: A Practical UK Guide to Sourced Answers

The Kopik team7 min read

You can chat with a PDF and get answers you can rely on, provided three conditions are met: the file contains real, extractable text; the tool answers only from that text; and every answer points back to the passage it used. When a "PDF AI" tool disappoints, the culprit is usually the document (a scan, a mangled table) rather than the model. This guide explains how to prepare your PDFs, how to judge the answers, and how to question a whole folder of documents at once without mixing them up.

How asking questions of a PDF actually works

Most tools that let you ask questions of a PDF rely on retrieval-augmented generation, or RAG. The file is not handed to the model in one go. Instead, it is processed once, and then searched every time you ask something.

  1. The text layer is extracted from each page. No text layer, no content.
  2. The text is split into passages of a few paragraphs, small enough to be precise evidence.
  3. The passages are indexed so they can be searched quickly by keyword, by meaning, or both.
  4. Your question triggers a search, which returns the few passages most likely to hold the answer.
  5. A language model drafts the reply from those passages and, in a well-designed tool, cites them.

So the model never sees more than extraction produced and retrieval selected. If the relevant paragraph was garbled on the way in, or simply not retrieved, no amount of clever prompting will rescue the answer. Our walkthrough of how RAG works, step by step follows a single question through this pipeline.

Where extraction goes wrong: scans, tables and layouts

PDF was designed to make documents look the same on every printer, not to make their text easy for software to read. Two PDFs that look identical can behave quite differently. A thirty-second test saves a lot of frustration: open the file, try to highlight a sentence, then search for a word you can see on the page. If neither works, that page is a picture of text, not text.

Typical PDF problems and what to do about them

DocumentProblemRemedy before upload
Scanned letters, signed contractsImage only, so nothing to extractRun OCR, then proofread key pages
Photos of pages taken on a phoneSkew and shadows cause OCR errors, especially in figuresRescan flat, then OCR
Rate tables, price listsRows and columns lose their alignmentWrite the key figures out in a sentence, or add a CSV version
Two-column reports and newslettersLines from both columns can be interleavedUse the original Word or web version if available
Forms with typed-in fieldsField contents may not be extractedFlatten the form or print it to a fresh PDF
Protected PDFsCopying text is disabledLift the restriction if you have the right to

Scanned documents need OCR first

Do not assume a chat tool will read images of text. Put scans through optical character recognition beforehand, using the OCR function in a PDF editor or an open-source tool such as OCRmyPDF, which adds a searchable text layer. Then check figures, names and dates, as these are where recognition errors cluster.

Tables are the classic trap. Fee schedules, VAT examples and staffing ratios are exactly what people want to ask about, and exactly what extraction handles worst. If a table matters, add a short sentence beside it in the source document, or upload the table separately as a plain text or CSV file. A sentence such as "From April 2026 the standard fee is £85 per visit" survives any extraction.

Insist on answers you can check

A well-written answer is not necessarily a correct one. The feature that matters most in any PDF AI tool is a clear trail from each statement back to its source. Without it, you are taking a confident paragraph on trust, which is how invented details end up in board papers and client emails.

When comparing tools, look at how precise the sources are. Some give a page reference, which is handy for long reports. Others quote the passage itself with the document name, so you can check the wording without opening the file. Both are fine; what is not acceptable is a vague note that the answer is "based on your files".

  • Expect "I don't know". A sound tool says the documents do not cover a question rather than filling the gap.
  • Verify anything consequential. Deadlines, amounts, contractual obligations and health information should always be checked against the cited text.
  • Be specific. "What notice does the tenant have to give under clause 4?" retrieves far better than "Summarise the lease".
  • Borrow the document's wording. If the policy says "annual leave", use that phrase rather than "holiday" when precision matters.
  • Beware of duplicate editions. If last year's version is still in the pile, it may be the one quoted.

For more ways to keep answers anchored to the evidence, see our guide to reducing AI hallucinations.

Working with many PDFs at once

Dropping one PDF into a general chatbot is fine for a quick summary. It struggles with a policy library, a folder of tenancy agreements or years of board minutes. Pasting everything into one conversation runs into size limits, and the model starts blending documents. A knowledge base built on RAG copes better, because it searches the whole collection and passes only the relevant passages to the model.

  • Name files meaningfully. "Staff-Handbook-2026.pdf" makes a citation useful; "Scan_0042.pdf" does not.
  • Weed out duplicates and superseded versions before uploading, or put the year in each title.
  • Organise by subject. Separate collections for HR, finance and operations generally answer more cleanly than one enormous pile.
  • Phrase comparisons explicitly. "How do payment terms differ between the 2025 and 2026 supplier contracts?" prompts the search to find both.
  • Respect the limits. "Summarise all 400 files" is not a targeted question, and RAG is built to find specific passages.

Data protection before you upload

PDFs are often where the sensitive material lives: HR files, client contracts, medical letters. If they contain personal data, UK GDPR applies to whatever you upload, and you should be clear about your lawful basis, who can access the documents and where the provider processes them. The Information Commissioner's Office publishes guidance for organisations, and our UK GDPR checklist for document chatbots goes through the practical checks. The simplest safeguard is to strip personal details from documents that do not need them.

Chatting with your PDFs on Kopik

Kopik turns your documents into a knowledge base you can question. Creating a base is free, and a private base is accessible only to you and your API keys. The process looks like this:

  1. Get the files ready. Make sure each PDF has selectable text and run OCR on any scans first. Word, plain text and Markdown files work as well.
  2. Create a base and upload. Kopik extracts the text, splits it into passages and indexes them for hybrid search, combining full-text search with semantic expansion of your keywords so differently worded questions still find the right passage.
  3. Ask away. Answers are drafted from your documents only, with numbered cited passages showing the excerpt and the name of the document.
  4. Check before you reuse. Read the cited passages before copying a figure or clause into anything official.
  5. Connect your other tools. The same base can be queried on the website, through the REST API, or from MCP clients such as Claude, Cursor or ChatGPT. A passages mode returns only the raw excerpts if you would rather read them yourself.

Kopik works from the text contained in your files, which is why the OCR step comes first for scanned material. Questions to your own bases are free within fair use.

Make your PDFs answerable

Upload PDFs, Word or text files and get answers drawn from your own documents, with every passage cited.

Checklist: choosing a tool to chat with PDFs

  • Every answer cites sources precisely enough to verify, by passage or by page.
  • It admits when the documents do not contain the answer.
  • It searches a whole collection of files, not one PDF per conversation.
  • You decide who sees the documents and can delete them at any time.
  • It accepts Word and text formats when the PDF is not the best source.
  • It offers an API or MCP access if you want to plug it into other tools.

Frequently asked questions

Why does the tool say it cannot find anything in my PDF?

Most often the PDF is a scan without a text layer, so there is nothing to search. Try highlighting text in your viewer. If you cannot, run OCR on the file, proofread a few pages and upload the new version.

Can I ask questions across several PDFs at once?

Yes, if the tool builds a searchable knowledge base from all the files. Give the files clear names and remove old versions so citations point to the right document. Requests that require reading every file in full remain a weak point.

How can I tell whether an answer really comes from my document?

Choose a tool that cites its sources, either by page or by quoting the passage with the document name, and read that citation before relying on the answer. An answer without a traceable source should be treated as unverified.

Does UK GDPR apply when I upload PDFs to an AI tool?

If the PDFs contain personal data, yes. You need a lawful basis, appropriate security and clarity about who processes the data and where. Removing personal details that are not needed is the easiest way to reduce risk.

What about tables, charts and images in my PDFs?

Tables often lose their structure during extraction, and charts or images usually contain no extractable text. State important figures in a sentence or upload the table as a separate CSV or text file.

Get the Kopik newsletter

New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.

By subscribing you agree to receive our newsletter. We never share your address.