How-to

A Private ChatGPT for Your Business Documents: A UK Guide to Doing It Properly

The Kopik team7 min read

A private ChatGPT for your organisation is an assistant that answers questions from your internal documents without those documents leaking: not used to train someone else's model, not kept longer than necessary, and only searchable by the staff and systems you authorise. Under UK GDPR that is not just good practice, it is part of your accountability duty whenever the documents contain personal data. This guide breaks 'private' into concrete checks, compares the realistic setups and shows how a private knowledge base lets several assistants share one controlled source of truth.

Why 'private' needs unpacking

Ask five suppliers whether their AI is private and you will get five confident yeses. The word is doing too much work. For document chat, it bundles at least five distinct properties, and a tool can satisfy some without the others:

  • Training: are your prompts and files used to improve a model that other customers use?
  • Retention: how long are conversations, files, extracted text and logs kept, including backups?
  • Access: who inside your organisation can query which documents, and can you withdraw that access quickly?
  • Separation: is your content kept apart from other customers' content in search and storage?
  • Accountability: are these promises in a contract you can show the ICO, a client or an auditor?

Once you list them separately, comparisons become much easier. A self-hosted model scores well on separation and badly on access control unless you build it yourself. A team plan from a chat assistant may handle training and retention but leave every shared file open to every workspace member. Decide which properties matter for which documents, then pick the tool.

Retention and UK GDPR: questions to put to any supplier

Terms for AI services change often, and personal and business tiers of the same product rarely share defaults. Rather than relying on a blog summary, check the terms of your own plan and keep a dated copy. If the documents include personal data (staff records, client correspondence, CVs, case notes), the supplier will normally act as your processor, and UK GDPR expects a written contract covering what they may do with the data.

  1. Is training on customer content off by default, and can an administrator enforce that across the organisation?
  2. What is the retention period for chats, uploaded files and derived data such as indexes, and what happens to backups?
  3. Does deleting a document remove its extracted text and search index as well?
  4. Which subprocessors are involved, and is any data transferred outside the UK? If so, under which transfer mechanism?
  5. Will the supplier sign a data processing agreement, and help you with subject access and erasure requests?
  6. Can you export and delete everything when you leave?

For anything beyond low-risk internal documentation, a data protection impact assessment is a sensible step, and often a required one when processing is likely to result in high risk. The Information Commissioner's Office publishes guidance on AI and data protection, DPIAs and international transfers that is worth reading before you choose. Our UK GDPR checklist for document chatbots goes through the obligations in more detail.

The realistic options, compared

Ways to chat with internal documents privately

SetupWhat it isGood atWeak at
Attaching files in a chat assistantStaff upload documents into individual conversationsInstant, no IT projectCopies spread across personal chats; retention depends on the plan; no single source of truth
Organisation plan with shared projectsCentral admin, shared spaces and data settingsAdmin console, contractual termsPermissions are often coarse; documents are locked into one assistant
Self-hosted open-source model and indexEverything runs on your own infrastructureData never leaves your serversYou own patching, permissions, monitoring and answer quality
Private knowledge base via API or MCPDocuments indexed once; assistants and apps query it with keysOne controlled corpus, revocable keys, cited passages, works across assistantsYou still decide what goes in and vet the provider like any processor

The last option separates two things that chat tools usually merge: where the documents live and which assistant writes the answer. The documents sit in a private retrieval layer; Claude, Cursor, ChatGPT or your own application asks it questions and receives only the relevant passages. That is retrieval-augmented generation, illustrated question by question in How RAG Works, Step by Step.

Access control inside the organisation

Privacy towards the supplier is only half the job. Privacy inside the organisation matters just as much: a grievance file or a draft redundancy plan should not surface because a colleague asked a vague question. Practical rules that work:

  • Build bases around audiences. An all-staff base for policies and procedures, separate bases for HR, legal and finance. Retrieval can only return what a base contains.
  • Issue one key per person or application. Shared credentials make revocation painful and audit trails meaningless.
  • Add AI access to leavers' checklists. Revoke keys and assistant connections on the same day as email and network access.
  • Minimise personal data. Data minimisation is a UK GDPR principle; redact names or index anonymised procedures wherever the question does not need the individual.
  • Review membership regularly. Quarterly is a reasonable rhythm for most SMEs.

Prompts are not permissions

Telling an assistant 'do not reveal salary information' is not access control. If a document must stay restricted, keep it in a base that only authorised keys can query. What is not retrievable cannot be quoted.

Setting up a private knowledge base with Kopik

Kopik is one way to implement the private knowledge base pattern. You upload PDF, Word, text or Markdown files; Kopik extracts, chunks and indexes them for hybrid search (full text plus semantic widening of keywords). A private base is accessible only to its owner and the owner's API keys, and does not appear in the public catalogue.

You query it on the website, through the REST API, or through Kopik's MCP server. The Model Context Protocol is an open standard for connecting assistants to tools, explained in plain English in our MCP guide. For one base, the server address is `https://kopik.io/api/mcp?base=<slug>`, shown on the base's page, and your key goes in the `Authorization: Bearer kpk_…` header. In Cursor or any client that reads an `mcpServers` file:

`{ "mcpServers": { "kopik": { "url": "https://kopik.io/api/mcp?base=<slug>", "headers": { "Authorization": "Bearer kpk_…" } } } }`

The assistant then has `ask_base` (a drafted answer with numbered source passages) and `search_base` (passages only, so your own assistant writes the reply). Passages arrive inside `<kopik-untrusted>` tags, which tells the assistant to treat document content as data rather than as instructions. If you mainly work with Claude, How to Give Claude Access to Your Company Documents compares this route with the alternatives.

A sensible rollout plan for UK organisations

  1. Start with low-risk content. Staff handbook, IT procedures, product documentation. Prove the value before touching personal data.
  2. Classify documents. Public, internal, confidential, personal data, special category data.
  3. Record your lawful basis and run a DPIA where personal data is involved.
  4. Vet the supplier as a processor. Contract, subprocessors, transfers, retention, deletion.
  5. Create bases per audience and keys per application, with names that make revocation obvious.
  6. Test with known answers. Ask questions whose answers you already know and check the cited passages, not only the prose.
  7. Remove superseded documents. Two versions of a policy produce two answers.
  8. Review every quarter. Keys, members, documents and supplier terms all drift over time.

Build your private knowledge base

Upload your documents, keep the base private, and query it from the site, the API or your MCP client with your own keys.

Frequently asked questions

Is it legal to use ChatGPT on company documents in the UK?

Yes, provided you meet UK GDPR when personal data is involved: a lawful basis, a processor contract with the supplier, appropriate security, data minimisation and, where the risk is high, a DPIA. Confidentiality obligations in client contracts may add further limits, so check them too.

Does a private AI assistant have to be hosted in the UK?

UK GDPR does not require UK hosting, but transfers outside the UK must rely on a valid mechanism such as adequacy regulations or appropriate safeguards. Some clients and sectors set stricter location requirements in their contracts, so check those before choosing.

What is the difference between uploading files to a chatbot and a private knowledge base?

Uploaded files usually live inside one assistant and often inside one conversation. A private knowledge base is a separate, central index: you manage the documents once, decide which keys can query it, and any compatible assistant or application can use it through an API or MCP.

Can staff still see documents they should not through the assistant?

Only if those documents are in a base their key can reach. Keep restricted material in separate bases with separate keys, and the assistant will have nothing restricted to retrieve.

How do I check that answers are accurate?

Use a setup that returns the source passages with each answer, and test it with questions you already know the answers to. If an answer is not supported by the cited passage, fix the corpus (outdated or contradictory documents are the usual cause).

Get the Kopik newsletter

New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.

By subscribing you agree to receive our newsletter. We never share your address.