DPDP Act 2023 and AI: what Indian companies must do
What should you know about DPDP Act AI compliance when building AI systems in India?
Under DPDP, AI systems need lawful purpose, consent or legitimate use, retention limits, breach notification and a grievance officer. The Act never mentions AI, but any AI system touching personal data of Indian residents is processing under it, and the choices that make compliance easy are cheapest before the build.
DPDP Act AI compliance is not a separate regime. The Digital Personal Data Protection Act, 2023 regulates the processing of digital personal data, and an AI system that reads customer records, transcribes calls, extracts fields from identity documents or personalises offers is processing personal data like any other software. What changes with AI is the scale, the number of third parties involved (model vendors, cloud regions, labelling contractors) and the difficulty of explaining what the system did with a given person's data. This article sets out what the Act requires, how each requirement lands on an AI build, and the design decisions we make so that compliance is a property of the architecture rather than a document written afterwards.
What the DPDP Act requires of anyone processing personal data
The Act is short by the standards of data protection law. It places obligations on the Data Fiduciary, the organisation that decides why and how personal data is processed, and gives rights to the Data Principal, the individual. Processing must be for a lawful purpose, and either with the individual's consent or under one of the legitimate uses the Act lists, such as data the person voluntarily provided for a specified purpose, employment, or a medical emergency. Consent must be free, specific, informed, unconditional and unambiguous, preceded by a clear notice, and as easy to withdraw as to give. Personal data must be erased when the purpose is served or consent withdrawn, unless retention is required by law. Reasonable security safeguards are mandatory, breaches must be reported to the Data Protection Board and to affected individuals, and every fiduciary must publish a way for individuals to raise grievances. The primary text and the accompanying Rules are on the Ministry of Electronics and IT site, and the Rules phase obligations in over time, so check the current timetable.
How each obligation lands on an AI system
| DPDP obligation | What it means for an AI build | Design response |
|---|---|---|
| Lawful purpose and notice | The purpose you told people about must cover what the model does with their data | Map every AI feature to a stated purpose before building; update notices where a new use appears |
| Consent or legitimate use | Training, fine-tuning and evaluation on customer data need a basis, not just the live prediction | Record the basis per data flow; keep consented and legitimate-use data separable |
| Purpose limitation | Data collected for support cannot silently become marketing personalisation | Tag data by purpose at ingestion; the retrieval layer filters by purpose |
| Erasure and retention limits | Vector indexes, fine-tuned weights, logs and traces all hold personal data | Design deletion across every store, including embeddings and LLM traces, from day one |
| Security safeguards | Prompts, outputs and traces are new places personal data leaks | Redact before logging; encrypt traces; restrict who can read them |
| Data Processors | Model APIs, cloud regions and labelling vendors are processors under contract | Contract terms that cover the AI vendor; residency decision made explicitly |
| Breach notification | A prompt injection that exfiltrates records is a breach | Detection and an incident runbook that includes the AI components |
| Grievance and rights | People can ask what you hold and demand correction or erasure | An access path that covers AI-derived data, and a named grievance contact |
DPDP consent AI teams get wrong
The most common gap is the basis for secondary uses. A company collects a phone number and address to deliver an order, which is straightforward. It then uses the order history to train a recommendation model, transcribes support calls to fine-tune a voice agent, and sends call recordings to a third-party model API. Each of those is a processing activity that needs a purpose the person was told about and a basis. Sometimes the original notice covers it; often it does not. The fix is a data flow inventory listing every AI feature, the personal data it touches, the purpose it serves and the basis relied upon, kept current as features are added. It is a spreadsheet, and it is the single most useful compliance artefact an AI programme produces. For voice specifically, our voice AI compliance in India article covers recording consent in detail.
Erasure is the hard engineering problem
A right-to-erasure request against an ordinary database is a delete statement. Against an AI system it touches the source record, the chunks and embeddings in the vector index, cached retrieval results, LLM request and response traces in the observability tool, evaluation datasets built from real conversations, and possibly a fine-tuned model. The last of these cannot be selectively edited, which is one reason we prefer retrieval over fine-tuning for anything containing personal data, as discussed in RAG versus fine-tuning. For the rest, the requirement is that every store is keyed by the individual so that a deletion job can find and remove their data everywhere, and that the job is tested rather than assumed. Retention limits work the same way: traces and evaluation sets need expiry, not just the production database.
Where the data goes: residency and processors
The Act allows transfer of personal data outside India except to countries the government restricts by notification, and sector regulators such as the RBI impose their own localisation rules for payment data. In practice the question for an AI build is which model vendor and cloud region will see the data, and under what contract. Public model APIs are Data Processors; the terms must prohibit training on your data and set out breach and deletion obligations. Where the data is sensitive, the volume is high or a sector rule applies, a private agentic AI deployment with open-weight models inside your own environment removes the question entirely, and the trade-offs are set out in self-hosted LLMs: when running your own model beats an API.
Significant Data Fiduciaries and children
The government can designate Significant Data Fiduciaries based on the volume and sensitivity of data and the risk of harm. They carry extra duties: a Data Protection Officer based in India, an independent data auditor, and periodic impact assessments. If your AI system could attract designation, build as though it applies now; the impact assessment is cheap at design time and expensive to reconstruct. Children's data has its own rules: verifiable parental consent and a prohibition on tracking, behavioural monitoring and targeted advertising directed at children, which rules out most personalisation for under-eighteen users unless an exemption applies.
Security safeguards specific to AI
- Redact personal identifiers from prompts before they reach an external model where the task allows it
- Treat LLM traces as personal data: encrypt them, restrict access, expire them
- Enforce access rights at the retrieval layer so an internal assistant cannot surface one customer's data to a user without permission, as described in permission-aware retrieval
- Gate every agent action that writes or sends personal data behind a policy check
- Include prompt injection and data exfiltration in the incident runbook and test for them
- Keep evaluation datasets de-identified or under the same controls as production
A worked example
An NBFC building document intelligence for KYC onboarding needed to read identity documents and bank statements, which are about as sensitive as personal data gets. The data flow inventory came first: each document type, the fields extracted, the purpose, the basis, the retention period and every system that would hold a copy. Extraction ran on models inside the NBFC's own cloud tenancy, with no document leaving it. Traces were redacted at the field level before storage and expired on a schedule. Erasure was implemented as a single job that removed a customer's documents, extracted fields, embeddings and traces, and it was tested against a synthetic customer before launch. The grievance contact and the access path were documented as part of the build. The engineering detail is in the KYC document intelligence case study.
Team and timeline
Compliance work belongs inside the build, not alongside it. A Sprint Zero discovery sprint at $3,250 / ₹2,00,000 produces the data flow inventory, the residency decision and the erasure design as deliverables. The build follows as a private agentic AI engagement from $31,500 / ₹20.8L plus infrastructure, or as an LLM application from $21,000 / ₹13.6L where a public API with acceptable terms is appropriate. Your side needs a legal or compliance owner for the basis decisions and the notices, and an engineering owner for the deletion job and incident runbook. We invoice in INR with GST for Indian entities; the bands are on the pricing page.
Before you start: a checklist
- List every AI feature, the personal data it touches, its purpose and its basis
- Check whether existing notices cover training, evaluation and third-party model use
- Decide the residency position: which vendors and regions may see which data
- Review model vendor terms for training prohibitions, deletion and breach clauses
- Design erasure across source, index, cache, traces and evaluation sets
- Set retention periods for traces and evaluation data
- Name the grievance contact and confirm the access request path covers AI-derived data
- Add AI components to the breach detection and notification runbook
Glossary
- Data Fiduciary: the organisation that determines the purpose and means of processing
- Data Principal: the individual the data is about
- Data Processor: a party processing on behalf of the fiduciary, such as a model API vendor
- Legitimate use: the Act's list of grounds that allow processing without fresh consent
- Significant Data Fiduciary: a fiduciary designated for higher obligations because of scale or risk
- Data Protection Board: the adjudicating body that receives breach reports and complaints
Related reading
Continue with GDPR and AI systems: a builder's guide if you serve EU customers, AI governance for mid-size companies, and our security page. The Act, the Rules and any government notifications on MeitY's site are the primary sources; this article is not legal advice.
Inventory the flows, decide where data may go, build erasure to cover every store, and the rest of DPDP compliance is paperwork you can actually fill in.
Frequently asked questions
Does the DPDP Act apply to AI models trained on customer data?
▾
Yes. Training, fine-tuning and evaluation on personal data are processing activities that need a lawful purpose and a basis, the same as the live use. Fine-tuned weights cannot be selectively erased, which is a reason to prefer retrieval for personal data.
Can we send customer data to a public LLM API under DPDP?
▾
Often, if the vendor is a Data Processor under a contract that prohibits training on your data and covers deletion and breaches, and no sector rule requires localisation. Sensitive or regulated data is usually better kept inside your own environment.
What are the penalties under the DPDP Act?
▾
The Schedule sets monetary penalties per instance, with the highest tier, for failing to take reasonable security safeguards, reaching up to ₹250 crore. The Data Protection Board decides amounts based on the nature and gravity of the breach.