GDPR and AI systems: a builder's guide
What should you know about GDPR AI compliance when building AI systems for EU or UK users?
GDPR-aligned AI needs a lawful basis, DPIAs for high-risk use, processor terms, transfer safeguards and data subject rights support. Most of that is architecture: where data flows, which vendors see it, how it is deleted from every store, and whether a person can be told why the system decided what it did.
GDPR AI compliance means applying a regulation written in 2016 to systems that did not exist then, and the surprising thing is how well it maps. The General Data Protection Regulation regulates processing of personal data of people in the EU, and the UK GDPR does the same for the UK. An LLM application that reads customer emails, a voice agent that transcribes calls, a recommendation model trained on browsing history: each is processing, each needs a lawful basis, and each raises the same questions about processors, transfers, retention and rights that any data system does, plus a few that are specific to AI. This guide is written for the people building the system, not for lawyers, and covers what to decide and build so that the compliance paperwork describes something true.
What GDPR requires, in the order a builder meets it
First, a lawful basis for each processing purpose under Article 6: consent, contract, legal obligation, vital interests, public task or legitimate interests. Second, transparency: people must be told what you do with their data, including any automated decision-making. Third, where processing is likely to be high risk, a Data Protection Impact Assessment under Article 35 before you start. Fourth, contracts under Article 28 with every processor, which includes model API vendors and cloud providers. Fifth, safeguards for any transfer outside the EU or UK. Sixth, support for data subject rights: access, rectification, erasure, restriction, portability, objection, and the Article 22 right not to be subject to solely automated decisions with legal or similarly significant effects. Seventh, security, breach notification to the supervisory authority within 72 hours, and records of processing. The consolidated text at gdpr-info.eu is the easiest reference, and the ICO's guidance on AI is the most practical regulator material.
How each requirement lands on an AI system
| Requirement | AI-specific question | What we build |
|---|---|---|
| Lawful basis | Is the basis for training and evaluation the same as for the live feature? | A data flow register with a basis per flow; consented and legitimate-interest data kept separable |
| DPIA | Does the system profile people, decide about them, or process special-category data at scale? | A DPIA drafted during discovery, updated when scope changes |
| Processor terms | Does the model vendor's contract prohibit training on your data and cover sub-processors? | An AI data processing agreement review before any vendor is wired in |
| Transfers | Which region does the model run in, and does the vendor's support team have access? | Region pinned; transfer mechanism documented; private deployment where none suffices |
| Article 22 | Does the AI make a decision with significant effect without a human? | Human review for consequential decisions; the review must be real, not a rubber stamp |
| Erasure | Where does this person's data sit: index, cache, traces, eval sets, weights? | A deletion job that covers every store and is tested |
| Transparency | Can you explain what the system did with a person's data? | Tracing per request; a plain-language explanation of the logic for automated decisions |
| Security and breach | Does prompt injection or exfiltration through the model count as a breach? | Yes; detection, redaction and an incident runbook that includes the AI components |
GDPR LLM questions: training, prompts and traces
Three places where an LLM application holds personal data get overlooked. The prompt: every retrieved document, customer record or transcript passed to the model is a disclosure to whoever runs the model. The trace: observability tools store prompts and outputs by default, and those stores are usually less protected than the production database. The evaluation set: golden questions built from real conversations are personal data with a long retention life. We handle these by redacting identifiers before external model calls where the task allows, encrypting and expiring traces, and de-identifying evaluation sets. Fine-tuning on personal data adds a fourth: weights cannot be selectively erased, so a person's erasure request cannot be honoured within the model. That is a strong reason to prefer retrieval, discussed in RAG versus fine-tuning, for anything that contains personal data.
The AI data processing agreement
Every model vendor and cloud provider that sees personal data is a processor and needs Article 28 terms. Read the vendor's data processing addendum for four things: a clear statement that your inputs and outputs are not used to train their models; the list of sub-processors and the notice you get when it changes; the deletion commitment, including retention of abuse-monitoring logs; and the transfer mechanism if the vendor is outside the EU or UK. Enterprise tiers from the major vendors generally meet these; consumer or default tiers often do not. Where no acceptable terms exist, or the data is special-category, running the model yourself under our private agentic AI service removes the processor question for the model layer entirely. The decision framework is in self-hosted LLMs: when running your own model beats an API.
DPIAs and Article 22: profiling and automated decisions
A DPIA is required where processing is likely to result in high risk, and the supervisory authorities' lists include systematic profiling, automated decision-making with significant effects, large-scale processing of special-category data and use of new technologies. Most AI systems that decide about people hit at least one. The DPIA is not a form; it is a description of the processing, an assessment of necessity and proportionality, the risks to individuals and the measures that reduce them. Done during discovery it shapes the design: it is where the decision to keep a human in the loop, to pin a region or to redact before external calls gets made. Article 22 then sets the floor for consequential decisions: a loan refusal, a job screening outcome, a claim denial. Either a person genuinely reviews, or one of the narrow exceptions applies with safeguards including the right to contest. Our policy-gated actions and shadow-mode-before-autonomy stance exist for engineering reasons, and they happen to be exactly what Article 22 wants.
EU AI Act basics for builders
The EU AI Act sits alongside GDPR rather than replacing it. It classifies AI systems by risk: prohibited practices, high-risk systems in listed areas such as credit scoring, employment, education and essential services, transparency obligations for systems that interact with people or generate content, and general-purpose model obligations for providers. Application is phased from 2025 onwards, and the exact dates for high-risk obligations have been the subject of proposed adjustments, so check the current timeline before relying on any date. For a company deploying rather than building models, the practical duties are knowing whether your use case is high-risk, keeping the documentation and logs the Act expects, telling people when they are talking to an AI, and ensuring human oversight. The engineering for those overlaps almost entirely with good GDPR practice.
A worked example
A B2B SaaS company with EU customers wanted an in-app copilot that could read a customer's account data and draft actions. The data flow register listed every data type the copilot could see and the basis, which was the contract with the customer for most flows and legitimate interests for product analytics. A DPIA was written during discovery and concluded that the copilot's draft actions needed explicit user confirmation, which became the permission model. Model calls were pinned to an EU region under enterprise terms with a training prohibition, and identifiers were redacted from traces before storage. Erasure was implemented across the account database, the retrieval index and the trace store and tested with a synthetic account. The result is described in the in-app copilot case study; the compliance work added little time because it was designed in rather than bolted on.
Team and timeline
The data flow register, DPIA draft, vendor terms review and erasure design are deliverables of a Sprint Zero discovery sprint at $3,250 / ₹2,00,000, credited to the build. The system follows as an LLM application from $21,000 / ₹13.6L where a public API under acceptable terms fits, or as a private agentic AI build from $31,500 / ₹20.8L plus infrastructure where it does not. Your side needs a data protection lead, or external counsel, to own the basis decisions and sign the DPIA, and an engineering owner for the deletion job. We invoice in USD for EU and UK clients, with studios in London and New York; the bands are on the pricing page.
Before you start: a checklist
- List every AI feature, the personal data it touches, the purpose and the Article 6 basis
- Screen for DPIA triggers: profiling, automated decisions, special-category data, scale
- Review each model vendor's data processing addendum for training use, sub-processors, deletion and transfers
- Decide the region and the transfer mechanism, or decide to self-host
- Identify decisions with significant effects and design real human review
- Design erasure across source, index, cache, traces and evaluation sets
- Write the transparency notice, including how automated decisions are made
- Add AI components to the breach runbook with the 72-hour clock in mind
Glossary
- Controller: the organisation that decides the purposes and means of processing
- Processor: a party processing on the controller's behalf under Article 28 terms, such as a model API vendor
- DPIA: Data Protection Impact Assessment, required before high-risk processing
- Special-category data: health, biometrics, religion, politics, sexual orientation and similar, with stricter rules
- Article 22: the right not to be subject to solely automated decisions with legal or similarly significant effects
- Transfer mechanism: the legal basis for moving data outside the EU or UK, such as an adequacy decision or standard contractual clauses
Related reading
See DPDP Act 2023 and AI for the Indian equivalent, zero data egress: designing AI that never leaves your VPC, and our security page. This article is an engineering guide, not legal advice; the regulation text and your supervisory authority's guidance are the primary sources.
Register the flows, assess the risk before building, read the vendor terms, build erasure everywhere, and keep a real person on consequential decisions.
Frequently asked questions
Can we use OpenAI or Anthropic APIs and stay GDPR compliant?
▾
Yes, under enterprise terms that prohibit training on your data, list sub-processors, commit to deletion and provide a transfer mechanism or an EU region. Check the addendum for your tier; default consumer terms usually do not suffice.
Do we need a DPIA for an AI chatbot?
▾
If it profiles people, makes or supports decisions with significant effects, or processes special-category data at scale, yes. A simple FAQ assistant over public content may not need one, but screening for triggers is always required.
How does the EU AI Act relate to GDPR for AI systems?
▾
They apply together. GDPR governs personal data in the system; the AI Act classifies the system by risk and adds obligations such as documentation, human oversight and transparency for high-risk and interactive uses.