A Private ChatGPT for Your Company Documents: What 'Private' Really Means
A private ChatGPT for your company documents is an AI assistant that answers from your internal files while keeping those files out of other people's reach: not used to train public models, not kept longer than you decide, and not readable by employees or apps who should not see them. You can get there with a business plan of a chat assistant, a self-hosted stack, or a private knowledge base that any assistant queries through an API or MCP. The right choice depends less on the brand of the chatbot than on three questions: who can read the documents, how long anything is kept, and where the answers come from.
What 'private' actually means for an AI assistant
'Private' is used loosely in AI marketing. When a vendor says your data stays private, it can mean very different things. Before you upload a single contract, split the word into the separate promises it can hide:
- No training use: your prompts and files are not used to improve a model that other customers will use.
- Limited retention: conversations, uploaded files and logs are deleted after a defined period, or when you delete them.
- Access control: only the people and applications you authorize can query the documents, and you can revoke that access.
- Isolation: your documents are not mixed into a shared index where another tenant's search could surface them.
- Contractual coverage: the promises above are written into terms you can point to, such as a data processing agreement or, for health data, a business associate agreement.
A setup can be strong on one promise and weak on another. A self-hosted model gives you excellent isolation but no access control unless you build it. A business chat plan may exclude your data from training but still let any colleague in the workspace open a shared file. Treat each promise as a separate checkbox.
Data retention: the questions to ask before uploading anything
Retention is where most surprises happen. Consumer and business tiers of the same assistant often have different defaults, and those defaults change over time. We will not summarize any vendor's current terms here, because they evolve: check the terms of your own plan, in writing, and keep a dated copy. These are the questions worth asking:
- Are prompts, uploaded files and outputs used for model training by default, and can an administrator switch that off for the whole organization?
- How long are conversations kept after a user deletes them, and are there backup or abuse-monitoring copies with a separate clock?
- Are uploaded files stored as files, as extracted text, or as an index, and is each of those deleted when the source file is removed?
- Who on the vendor side can access your content, under what conditions, and is that access logged?
- Which subprocessors handle the data, and in which countries?
- Can you export and delete everything at the end of the contract?
For US companies, the answers feed directly into your compliance work. If the documents contain protected health information, HIPAA requires a business associate agreement with any vendor that creates, receives, maintains or transmits that information on your behalf. If you sell to enterprise customers, their security questionnaires will ask whether your AI vendors hold a SOC 2 report. State privacy laws such as California's CCPA add notice and deletion obligations when personal information is involved. None of this forbids using AI on internal documents, but it does mean the retention answers need to be documented.
Four ways to chat with company documents privately
There is no single 'private ChatGPT' product. In practice, teams pick one of four patterns, sometimes two in parallel.
Private document chat: the main options
| Option | How it works | Strengths | Watch out for |
|---|---|---|---|
| File upload in a chat assistant | Users attach PDFs to a conversation | Zero setup, familiar interface | Retention depends on the plan; knowledge is scattered across personal chats; no shared source of truth |
| Business or enterprise plan with shared workspaces | Admins manage users, files and data settings centrally | Admin controls, contractual terms | Access control is often workspace wide; documents are tied to one assistant |
| Self-hosted open-source stack | You run the model, the index and the interface on your own servers | Maximum isolation | You build and maintain permissions, updates, monitoring and quality |
| Private knowledge base queried by API or MCP | Documents are indexed once in a private base; any assistant or app queries it with a key | One source of truth, revocable keys, cited passages, works with several assistants | You still need to choose what goes in, and check the provider's terms like any other vendor |
The fourth pattern is the least known and often the most practical. Instead of copying the same handbook into every chat tool your teams use, you keep the documents in one private retrieval layer and let assistants ask questions of it. That is retrieval-augmented generation, explained step by step in How RAG Works.
Access control: the part most teams skip
A private assistant that every employee can query about every document is not private inside the company. The HR investigation file, the board deck and the salary grid should not be one prompt away from an intern. Good access control for document chat follows the same principles as file sharing:
- Split by audience, not by topic. One base for the whole company (handbook, policies, product docs), one for HR, one for legal, one for finance. A question can only retrieve what the base contains.
- One key per application or person. If a script, an agent and a colleague all use the same credential, you cannot revoke one without breaking the others.
- Revoke on departure. Add API keys and assistant connections to your offboarding checklist, next to email and SSO.
- Keep personal data out unless it is needed. The safest sensitive document is the one you never indexed. Redact or summarize before uploading.
- Log questions where possible. You want to know which application asked what, especially for regulated content.
A simple rule of thumb
If you would not put a document in a shared drive folder that everyone in a group can open, do not put it in a knowledge base that the same group can query. Retrieval does not invent permissions; it reflects the ones you set when you decide what goes into each base.
Private knowledge bases you query from ChatGPT, Claude or your own code
This is the approach Kopik takes. On Kopik, you create a knowledge base by uploading PDF, Word, text or Markdown files; Kopik extracts the text, splits it into passages and indexes it for hybrid search (full text plus semantic expansion of keywords). A private base is only accessible to its owner and to the owner's API keys. It never appears in the public catalog.
You can then query it three ways: on the website, through a REST API with a key, or through Kopik's MCP server, which MCP clients such as Claude, Cursor or ChatGPT can connect to (check what your workspace plan allows for custom connectors). Each answer is grounded in your documents and comes with numbered source passages, or you can ask for raw passages only and let your own assistant write the reply. The Model Context Protocol is an open standard, so the same base serves several assistants without being copied into each one.
For a single private base, the MCP address is `https://kopik.io/api/mcp?base=<slug>` and the key travels in the `Authorization: Bearer kpk_…` header. In Claude Code, for example:
`claude mcp add --transport http kopik "https://kopik.io/api/mcp?base=<slug>" --header "Authorization: Bearer kpk_…"`
The assistant then gets two tools: `ask_base` for a written answer with sources, and `search_base` for passages only. Content from the base arrives wrapped in `<kopik-untrusted>` tags, so the assistant treats it as data rather than as instructions, a useful protection if a document contains text that looks like a command. The full setup, including Cursor's `.cursor/mcp.json`, is in Connect a knowledge base to AI agents with MCP, and the reasoning behind keeping your own retrieval layer is in A private knowledge base as your own RAG for AI agents.
A rollout checklist for US teams
- Inventory the documents. List what people actually ask about: policies, onboarding, product specs, contract templates, procedures.
- Classify them. Public, internal, confidential, regulated (PHI, financial data, personal information under state law).
- Pick the pattern per class. Internal handbooks can go into a private base quickly; regulated content may need a vendor that signs the agreements you require, or may stay out entirely.
- Read the vendor terms. Training use, retention, subprocessors, deletion. Write the answers down with the date.
- Create bases by audience and issue keys per app. Name keys after their use so revocation is obvious.
- Test with real questions. Ask twenty questions you already know the answers to and check the cited passages, not just the wording.
- Clean the corpus. Remove outdated versions; contradictory documents produce contradictory answers. Our guide on grounded answers covers the common causes.
- Review quarterly. Keys, members, documents and vendor terms all drift.
Privacy for AI on documents is less about the chatbot and more about the plumbing: what is indexed, who holds a key, and how long anything is kept.
Create a private knowledge base
Upload your PDFs and Word files, keep the base private, and query it from the site, the API or any MCP client with your own key.
Frequently asked questions
Is a business plan of ChatGPT private enough for company documents?
It depends on the plan and on your documents. Business tiers of chat assistants usually offer stronger data settings than consumer tiers, but the details (training use, retention, admin controls) vary and change. Check the terms of your own plan in writing and compare them with what your contracts and regulations require.
Can I chat with documents without uploading them to a chatbot?
Yes. With a private knowledge base, documents are indexed once in a separate retrieval layer, and the assistant only receives the passages relevant to each question through an API or MCP. You control which documents are in the base and which keys can query it.
Does a private AI need to run on my own servers?
Not necessarily. Self-hosting gives maximum isolation but makes you responsible for security updates, permissions and answer quality. Many teams prefer a hosted service with clear contractual terms, strict access control and documents they can delete at any time.
Can I use protected health information in a document chatbot?
Only with a vendor setup that meets HIPAA, which includes a business associate agreement with each vendor handling the information. If you cannot obtain one, keep PHI out of the knowledge base and index de-identified procedures instead.
How do I stop an assistant from quoting a document someone should not see?
Do not rely on prompts for that. Put sensitive documents in a separate base and give keys to that base only to the people and applications allowed to read them. The assistant cannot retrieve what is not in the base it queries.
Get the Kopik newsletter
New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.
By subscribing you agree to receive our newsletter. We never share your address.