How-to

Build a GDPR-Compliant Chatbot on Your Documents: A Checklist for US Companies

The Kopik team9 min read

A GDPR chatbot is not a special product: it is a document chatbot whose data flows you have mapped, justified and limited. If you are a US company and your assistant serves customers in the European Union, or if the documents behind it mention EU residents, the GDPR can apply to you even without an office in Europe. This checklist walks through the decisions that matter: lawful basis, vendors acting as processors, retention, data minimization, international transfers, private knowledge bases and the notice your users see. It is a practical guide, not legal advice; run the final setup past your privacy counsel.

Does the GDPR apply to your document chatbot?

Article 3 of the General Data Protection Regulation (EU) 2016/679 reaches companies outside the EU when they offer goods or services to people in the EU, or monitor their behavior there. A support chatbot on a site that sells to German or French customers, in their language and their currency, is a typical case. So is an internal assistant used by your Dublin team.

Personal data shows up in a document chatbot in three places, and you need to look at all three:

  • The documents themselves: contracts with named signatories, HR policies with examples, support tickets, CRM exports, meeting notes.
  • The questions users type: people paste order numbers, emails, health details or complaints into a chat box without being asked.
  • The logs: conversation history, IP addresses, account identifiers and analytics that your stack keeps by default.

If none of these ever contains information about an identifiable person, the GDPR has little to say about the bot. In practice, the questions and logs almost always do, which is why the rest of this checklist matters. Note that US state laws add their own layer: the California Consumer Privacy Act (CCPA, as amended by the CPRA) and similar statutes in other states set notice and opt-out rules for consumer data, so map them alongside the GDPR rather than instead of it.

Step 1: pick a lawful basis and write it down

Every processing activity needs one of the six legal bases in Article 6. For a document chatbot, two usually fit:

Common lawful bases for a document chatbot

Use caseTypical basisWhat you must be able to show
Customer support bot answering account or order questionsPerformance of a contract (Art. 6(1)(b))The processing is needed to deliver the service the customer signed up for
Internal assistant over company policiesLegitimate interests (Art. 6(1)(f))A balancing test showing employees' interests do not override yours
Pre-sales or marketing bot that stores leadsConsent or legitimate interestsClear opt-in for marketing use, easy withdrawal
Using chat logs to improve the botLegitimate interests, rarely consentMinimization, a short retention period and an opt-out

Record the choice in your record of processing activities (Article 30). Avoid stretching one basis over several purposes: answering a customer's question and training a future model on that conversation are two different purposes, and the second needs its own justification. If special category data (health, religion, union membership) can plausibly appear, Article 9 adds stricter conditions; the simplest fix is often to keep those documents out of the bot entirely.

Step 2: sign processor agreements with every vendor in the chain

Your chatbot probably involves a hosting provider, a retrieval or knowledge base service, a language model API and maybe a chat widget vendor. Each one that handles personal data on your behalf is a processor, and Article 28 requires a written contract with specific terms. Before you go live, check for each vendor:

  1. A data processing agreement (DPA) that covers the Article 28 terms: instructions, confidentiality, security, sub-processors, assistance with rights requests, deletion at the end of the contract, audits.
  2. A published list of sub-processors, and a way to be notified when it changes.
  3. A clear statement on whether your inputs and outputs are used to train their models, and how to switch that off.
  4. Their own retention period for prompts and logs, including abuse-monitoring copies.
  5. The countries where data is stored and processed, and the transfer mechanism used.

Controller or processor?

When you deploy the bot for your own customers, you are usually the controller and your vendors are processors. If you sell the chatbot to other businesses, you become their processor and must offer them a DPA yourself.

Step 3: minimize what goes into the knowledge base

Data minimization (Article 5(1)(c)) is where a document chatbot is easiest to get right, because you choose the documents. A bot that answers questions about your return policy does not need last year's ticket export. Practical rules:

  • Index reference documents (policies, product docs, procedures, public regulations), not raw records about individuals.
  • Strip names, emails and signatures from templates and examples before upload; replace them with roles ("the customer", "the HR manager").
  • Split bases by audience: one for customers, one for staff, one for a single team. Retrieval cannot leak a document that is not in the base.
  • Review new uploads like a code change: someone checks that nothing personal slipped in.

Our guide to RAG for legal, HR and compliance teams covers how to scope bases by audience in more depth.

Step 4: set retention, transfers and security

Retention

Storage limitation (Article 5(1)(e)) means you decide in advance how long conversations are kept. Pick a period tied to a purpose, for example 30 days for troubleshooting, then delete or anonymize automatically. Apply the same rule to vendor-side logs, which is why the DPA questions above matter.

International transfers

Sending EU personal data to the United States is a transfer under Chapter V. For US recipients certified under the EU-US Data Privacy Framework, the Commission's adequacy decision covers it; otherwise you typically rely on the Standard Contractual Clauses plus a transfer impact assessment. Hosting in the EU reduces the number of transfers to manage, but it does not remove them if a US vendor can still access the data. The official texts, including the adequacy decisions and the clauses, are collected in the GDPR and international data transfers knowledge base, which you can query with citations.

Security

Article 32 asks for security appropriate to the risk: encryption in transit and at rest, access control on who can upload and who can query, API keys stored as secrets and rotated, and a breach process that can notify the supervisory authority within 72 hours (Article 33). If your team already works to SOC 2, most of these controls exist; map them rather than reinventing them.

Step 5: be transparent with the people using the bot

Articles 13 and 14 require you to tell people what you do with their data. For a chatbot, that means a short notice next to the chat box, linking to the full privacy policy, that states:

  • That they are talking to an automated assistant, not a person. The EU AI Act's transparency rules for systems that interact with people apply from August 2026, so this line is now expected anyway.
  • What happens to their messages, who receives them (vendor categories) and how long they are kept.
  • Not to paste sensitive information, and where to go for a human.
  • How to exercise their rights: access, deletion, objection.

Plan for rights requests before they arrive. If someone asks for their data, you need to find their conversations by account or email and delete them across your logs and your vendors'. If the bot's answers could produce decisions with legal or similarly significant effects (approving a refund claim, screening a candidate), Article 22 restricts fully automated decisions: keep a human in the loop. When the processing is large-scale or involves sensitive data, run a data protection impact assessment (Article 35).

Where private knowledge bases fit

Most GDPR risk in a document chatbot comes from who can reach the documents. On Kopik, a base is either public (listed in the catalog) or private, and a private base can only be queried by its owner and the owner's API keys, from the website, the REST API or the MCP server. That makes it a reasonable back end for an internal assistant: your bot sends the question with your key, gets an answer grounded in your documents with the cited passages, or just the raw passages if you prefer to generate the answer yourself. We describe the pattern in a private knowledge base as your own RAG.

Kopik is still a vendor in your chain, so the same checklist applies to it as to any other: review its terms and privacy information, add it to your records, and keep personal records out of the uploaded documents. Citations help on the accountability side, since every answer points to the passage it came from and reviewers can check what the bot said.

Start with a private base for your documents

Upload your policies and product docs to a private knowledge base, query it with your API key, and get answers with cited passages.

The one-page GDPR chatbot checklist

  1. Map personal data in documents, questions and logs.
  2. Choose and record a lawful basis per purpose.
  3. Sign a DPA with every processor; check training use and sub-processors.
  4. Remove personal records from the knowledge base; split bases by audience.
  5. Set a retention period for chats and apply it to vendors too.
  6. Document the transfer mechanism for any data leaving the EU.
  7. Secure keys, access and breach response.
  8. Publish a short notice at the chat box and a full privacy policy.
  9. Prepare for access and deletion requests; keep humans on significant decisions.
  10. Run a DPIA if the processing is large-scale or sensitive.

Frequently asked questions

Does a US company need to comply with the GDPR for its chatbot?

Yes, if the chatbot offers goods or services to people in the EU or monitors their behavior there, or if you process EU residents' data through an EU establishment. A bot on a site that targets EU customers is the classic case. A purely US-facing bot with no EU users is generally outside its scope, though state privacy laws may still apply.

Do I need consent to run a GDPR-compliant AI assistant?

Not necessarily. Consent is only one of six lawful bases. A support bot often relies on performance of a contract, and an internal assistant on legitimate interests with a documented balancing test. Consent becomes the right basis for optional uses such as marketing follow-ups.

Is the language model provider a processor under the GDPR?

Usually yes, when it processes prompts containing personal data on your instructions. You need an Article 28 agreement, clarity on whether your data is used for training, its retention period and the transfer mechanism if data leaves the EU.

How long can I keep chatbot conversations?

The GDPR sets no fixed number. You must define a period tied to a purpose, such as a few weeks for troubleshooting, document it, and delete or anonymize afterwards. Keeping chats indefinitely just in case is hard to justify.

Does hosting in the EU make a chatbot GDPR compliant?

It helps by reducing transfers, but it is not enough on its own. You still need a lawful basis, processor agreements, minimization, retention rules, security and transparency. And if a non-EU vendor can access the data remotely, that access can still count as a transfer.

Get the Kopik newsletter

New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.

By subscribing you agree to receive our newsletter. We never share your address.