azyware
Business

DPDP Act 2023 and AI: what Indian companies must do

EZ
Eazyware
· 7 min read
Quick answer

What should you know about DPDP Act AI compliance when building AI systems in India?

Under DPDP, AI systems need lawful purpose, consent or legitimate use, retention limits, breach notification and a grievance officer. The Act never mentions AI, but any AI system touching personal data of Indian residents is processing under it, and the choices that make compliance easy are cheapest before the build.

DPDP Act AI compliance is not a separate regime. The Digital Personal Data Protection Act, 2023 regulates the processing of digital personal data, and an AI system that reads customer records, transcribes calls, extracts fields from identity documents or personalises offers is processing personal data like any other software. What changes with AI is the scale, the number of third parties involved (model vendors, cloud regions, labelling contractors) and the difficulty of explaining what the system did with a given person's data. This article sets out what the Act requires, how each requirement lands on an AI build, and the design decisions we make so that compliance is a property of the architecture rather than a document written afterwards.

What the DPDP Act requires of anyone processing personal data

The Act is short by the standards of data protection law. It places obligations on the Data Fiduciary, the organisation that decides why and how personal data is processed, and gives rights to the Data Principal, the individual. Processing must be for a lawful purpose, and either with the individual's consent or under one of the legitimate uses the Act lists, such as data the person voluntarily provided for a specified purpose, employment, or a medical emergency. Consent must be free, specific, informed, unconditional and unambiguous, preceded by a clear notice, and as easy to withdraw as to give. Personal data must be erased when the purpose is served or consent withdrawn, unless retention is required by law. Reasonable security safeguards are mandatory, breaches must be reported to the Data Protection Board and to affected individuals, and every fiduciary must publish a way for individuals to raise grievances. The primary text and the accompanying Rules are on the Ministry of Electronics and IT site, and the Rules phase obligations in over time, so check the current timetable.

How each obligation lands on an AI system

DPDP obligationWhat it means for an AI buildDesign response
Lawful purpose and noticeThe purpose you told people about must cover what the model does with their dataMap every AI feature to a stated purpose before building; update notices where a new use appears
Consent or legitimate useTraining, fine-tuning and evaluation on customer data need a basis, not just the live predictionRecord the basis per data flow; keep consented and legitimate-use data separable
Purpose limitationData collected for support cannot silently become marketing personalisationTag data by purpose at ingestion; the retrieval layer filters by purpose
Erasure and retention limitsVector indexes, fine-tuned weights, logs and traces all hold personal dataDesign deletion across every store, including embeddings and LLM traces, from day one
Security safeguardsPrompts, outputs and traces are new places personal data leaksRedact before logging; encrypt traces; restrict who can read them
Data ProcessorsModel APIs, cloud regions and labelling vendors are processors under contractContract terms that cover the AI vendor; residency decision made explicitly
Breach notificationA prompt injection that exfiltrates records is a breachDetection and an incident runbook that includes the AI components
Grievance and rightsPeople can ask what you hold and demand correction or erasureAn access path that covers AI-derived data, and a named grievance contact

The most common gap is the basis for secondary uses. A company collects a phone number and address to deliver an order, which is straightforward. It then uses the order history to train a recommendation model, transcribes support calls to fine-tune a voice agent, and sends call recordings to a third-party model API. Each of those is a processing activity that needs a purpose the person was told about and a basis. Sometimes the original notice covers it; often it does not. The fix is a data flow inventory listing every AI feature, the personal data it touches, the purpose it serves and the basis relied upon, kept current as features are added. It is a spreadsheet, and it is the single most useful compliance artefact an AI programme produces. For voice specifically, our voice AI compliance in India article covers recording consent in detail.

Erasure is the hard engineering problem

A right-to-erasure request against an ordinary database is a delete statement. Against an AI system it touches the source record, the chunks and embeddings in the vector index, cached retrieval results, LLM request and response traces in the observability tool, evaluation datasets built from real conversations, and possibly a fine-tuned model. The last of these cannot be selectively edited, which is one reason we prefer retrieval over fine-tuning for anything containing personal data, as discussed in RAG versus fine-tuning. For the rest, the requirement is that every store is keyed by the individual so that a deletion job can find and remove their data everywhere, and that the job is tested rather than assumed. Retention limits work the same way: traces and evaluation sets need expiry, not just the production database.

Where the data goes: residency and processors

The Act allows transfer of personal data outside India except to countries the government restricts by notification, and sector regulators such as the RBI impose their own localisation rules for payment data. In practice the question for an AI build is which model vendor and cloud region will see the data, and under what contract. Public model APIs are Data Processors; the terms must prohibit training on your data and set out breach and deletion obligations. Where the data is sensitive, the volume is high or a sector rule applies, a private agentic AI deployment with open-weight models inside your own environment removes the question entirely, and the trade-offs are set out in self-hosted LLMs: when running your own model beats an API.

Significant Data Fiduciaries and children

The government can designate Significant Data Fiduciaries based on the volume and sensitivity of data and the risk of harm. They carry extra duties: a Data Protection Officer based in India, an independent data auditor, and periodic impact assessments. If your AI system could attract designation, build as though it applies now; the impact assessment is cheap at design time and expensive to reconstruct. Children's data has its own rules: verifiable parental consent and a prohibition on tracking, behavioural monitoring and targeted advertising directed at children, which rules out most personalisation for under-eighteen users unless an exemption applies.

Security safeguards specific to AI

  • Redact personal identifiers from prompts before they reach an external model where the task allows it
  • Treat LLM traces as personal data: encrypt them, restrict access, expire them
  • Enforce access rights at the retrieval layer so an internal assistant cannot surface one customer's data to a user without permission, as described in permission-aware retrieval
  • Gate every agent action that writes or sends personal data behind a policy check
  • Include prompt injection and data exfiltration in the incident runbook and test for them
  • Keep evaluation datasets de-identified or under the same controls as production

A worked example

An NBFC building document intelligence for KYC onboarding needed to read identity documents and bank statements, which are about as sensitive as personal data gets. The data flow inventory came first: each document type, the fields extracted, the purpose, the basis, the retention period and every system that would hold a copy. Extraction ran on models inside the NBFC's own cloud tenancy, with no document leaving it. Traces were redacted at the field level before storage and expired on a schedule. Erasure was implemented as a single job that removed a customer's documents, extracted fields, embeddings and traces, and it was tested against a synthetic customer before launch. The grievance contact and the access path were documented as part of the build. The engineering detail is in the KYC document intelligence case study.

Team and timeline

Compliance work belongs inside the build, not alongside it. A Sprint Zero discovery sprint at $3,250 / ₹2,00,000 produces the data flow inventory, the residency decision and the erasure design as deliverables. The build follows as a private agentic AI engagement from $31,500 / ₹20.8L plus infrastructure, or as an LLM application from $21,000 / ₹13.6L where a public API with acceptable terms is appropriate. Your side needs a legal or compliance owner for the basis decisions and the notices, and an engineering owner for the deletion job and incident runbook. We invoice in INR with GST for Indian entities; the bands are on the pricing page.

Before you start: a checklist

  • List every AI feature, the personal data it touches, its purpose and its basis
  • Check whether existing notices cover training, evaluation and third-party model use
  • Decide the residency position: which vendors and regions may see which data
  • Review model vendor terms for training prohibitions, deletion and breach clauses
  • Design erasure across source, index, cache, traces and evaluation sets
  • Set retention periods for traces and evaluation data
  • Name the grievance contact and confirm the access request path covers AI-derived data
  • Add AI components to the breach detection and notification runbook

Glossary

  • Data Fiduciary: the organisation that determines the purpose and means of processing
  • Data Principal: the individual the data is about
  • Data Processor: a party processing on behalf of the fiduciary, such as a model API vendor
  • Legitimate use: the Act's list of grounds that allow processing without fresh consent
  • Significant Data Fiduciary: a fiduciary designated for higher obligations because of scale or risk
  • Data Protection Board: the adjudicating body that receives breach reports and complaints

Continue with GDPR and AI systems: a builder's guide if you serve EU customers, AI governance for mid-size companies, and our security page. The Act, the Rules and any government notifications on MeitY's site are the primary sources; this article is not legal advice.

Inventory the flows, decide where data may go, build erasure to cover every store, and the rest of DPDP compliance is paperwork you can actually fill in.

Frequently asked questions

Does the DPDP Act apply to AI models trained on customer data?

▾

Yes. Training, fine-tuning and evaluation on personal data are processing activities that need a lawful purpose and a basis, the same as the live use. Fine-tuned weights cannot be selectively erased, which is a reason to prefer retrieval for personal data.

Can we send customer data to a public LLM API under DPDP?

▾

Often, if the vendor is a Data Processor under a contract that prohibits training on your data and covers deletion and breaches, and no sector rule requires localisation. Sensitive or regulated data is usually better kept inside your own environment.

What are the penalties under the DPDP Act?

▾

The Schedule sets monetary penalties per instance, with the highest tier, for failing to take reasonable security safeguards, reaching up to ₹250 crore. The Data Protection Board decides amounts based on the nature and gravity of the breach.