azyware
Technology

AI MVP Development, security and the DPDP Act: a compliance checklist

EZ
Eazyware
· 7 min read
Quick answer

Is AI MVP development compliant with the DPDP Act?

Not automatically. An AI MVP satisfies India's DPDP Act when it collects personal data against a specific notified purpose with verifiable consent, stores it where you said it would be stored, deletes it when the purpose ends, and logs who saw what. The architecture decides, not the model.

Not automatically. An AI MVP satisfies India's Digital Personal Data Protection Act 2023 when it collects personal data against a specific notified purpose with verifiable consent, keeps it where your notice says it will be kept, deletes it when the purpose ends, and logs who saw what. The architecture decides this, not the model you chose.

What follows is the mapping we use on every AI MVP development engagement: each DPDP obligation, the engineering control that satisfies it, and the evidence a reviewer will ask for. It also covers the two places AI MVPs leak personal data that ordinary web applications do not.

What the DPDP Act requires of an early-stage product

The Digital Personal Data Protection Act 2023 applies to any digital personal data processed in India, and to processing outside India aimed at offering goods or services to people in India. An MVP is not exempt because it is small, internal or a pilot. If ten real customers' records pass through it, the Act applies to those ten records. The Act is published by the Ministry of Electronics and Information Technology, whose official portal hosts the text and the rules made under it.

The obligations that bite hardest on an AI MVP are five: a clear notice tied to a specific purpose, consent that can be withdrawn as easily as it was given, purpose limitation, erasure once the purpose is served or consent is withdrawn, and reasonable security safeguards. The Act attaches financial penalties running to ₹250 crore for a failure of security safeguards, which is why a pilot that quietly copies production data into a spreadsheet is a governance problem rather than a shortcut.

A shorter definition for teams new to this: the DPDP Act 2023 makes you a data fiduciary for the personal data you decide the purpose and means of processing for, and fiduciary duties do not scale down with headcount. Our broader treatment sits in DPDP Act 2023 and AI.

DPDP obligations mapped to MVP engineering controls

This is the table we put in front of a client's legal or risk reviewer during an AI-accelerated MVP. Each row names the evidence, because a control nobody can demonstrate is not a control.

DPDP obligationWhat it means in an AI MVPControl we buildEvidence produced
Notice and purposeThe MVP processes data for one stated purpose onlyPurpose tag on every ingestion job and prompt templatePurpose register mapped to the consent notice
Consent and withdrawalUsers can revoke and the system must stopConsent flag checked at retrieval time, not only at captureWithdrawal test case in the release suite
Purpose limitationPilot data does not become training dataContractual no-training terms plus zero-retention API settingsVendor terms filed alongside the architecture note
Erasure and retentionRecords and their derivatives expireDeletion cascades to embeddings, caches, logs and tracesTimed deletion job with a run log
Security safeguardsEncryption, least privilege, tested restoresRole-based access, encryption in transit and at rest, secrets in a vaultAccess matrix and a restore test report
Breach reportingYou can describe scope within hoursStructured audit log of every record read and every model callQuery that reconstructs who saw what, when

The erasure row is the one teams underestimate. Deleting a row from a database does not delete the vector built from it, the chunk cached in a semantic cache, the trace stored by your observability tool, or the copy in last night's evaluation export. Erasure has to cascade through every derivative, which is an architecture decision made in week one and a rewrite if made in month six.

The consent row deserves a note too. Most teams check consent once, at the moment data is captured, and never again. Under the Act consent can be withdrawn at any time and processing must stop, which means the check belongs at retrieval time as well: before a document is pulled into context, before a record is embedded, before a summary is generated. A consent flag that is only read by the signup form is decorative.

Where AI MVPs leak personal data

Ordinary applications leak through access control. AI MVPs leak through the extra surfaces the model introduces, and these are the ones we check line by line.

  • Prompt payloads. Whole documents get pasted into context. Apply PII redaction before the call, not after the response.
  • Vendor log retention. Some provider defaults retain request payloads for abuse monitoring. Enterprise terms or zero-retention endpoints change this; confirm in writing.
  • Embeddings and vector stores. An embedding derived from a personal record is personal data. It needs the same residency, access and deletion rules as the source.
  • Traces and evaluation sets. Observability tools and labelled test sets are the two most common places live customer data ends up outside the production perimeter.
  • Retrieval without permissions. A knowledge assistant that indexes every folder will answer questions about salaries. Filter by the asking user's rights at query time.
  • Demo environments. Seeding a pilot with a production dump is the single most common DPDP failure we find during onboarding audits.

Two of these surfaces are invisible in a code review because they live in vendor consoles rather than your repository. Provider log retention and observability traces are configuration, not code, so add them to the architecture note and re-check them whenever a provider or tool is swapped.

Does an AI MVP have to keep data in India?

The DPDP Act does not require blanket localisation. It permits cross-border transfer except to territories the central government restricts by notification, which is a narrower rule than the GDPR transfer regime. Sector regulators are stricter: entities under Reserve Bank of India supervision face payment-data storage and outsourcing audit expectations that override the general position, and a fintech MVP should assume Indian residency by default.

In practice we ask one question early: does any personal data have to stay inside your perimeter? If yes, the MVP is designed around a private deployment or a regional endpoint from the start, and the approach in zero data egress applies. If no, a hosted model with contractual no-training terms and a documented region is usually proportionate for a pilot.

What does compliance add to cost and timeline?

Less than teams fear when it is designed in, and a great deal when it is retrofitted. Consent checks, redaction, deletion cascades and audit logging are inside the scope of an AI-accelerated MVP at Eazyware, which runs from $26,500 or ₹17,60,000 to $45,500 or ₹30,40,000. Where the compliance position is genuinely unclear, a ten-day Sprint Zero at $3,250 or ₹2,00,000, credited to the build, settles residency, retention and the sub-processor list before code is written. All figures are published on the pricing page, and our platform controls are described on the trust and security page.

Retrofitting is the expensive path. Adding deletion cascades to a system whose vector store has no link back to source records typically means re-ingesting the corpus and rebuilding the retrieval layer. Budget a fortnight of engineering, or design it in on day one for nothing.

When this checklist is the wrong place to start

If your MVP processes no personal data, do not build a consent framework for it. A pricing copilot over public product documentation, an internal assistant over engineering runbooks, or a forecasting model over aggregated sales figures may touch no personal data at all. Classify the data first; the honest answer is sometimes that DPDP is not engaged and the security work is ordinary application security.

The second wrong start is a full data protection impact assessment before you know whether the feature works. Significant data fiduciary duties, including impact assessments and independent audits, apply to organisations the government notifies as such. A twenty-user pilot inside one department rarely qualifies. Do the proportionate controls now and the formal assessment when the system moves to production scale, an order set out in AI governance for mid-size companies.

What this looked like for a lender

For an NBFC we built document intelligence over KYC and loan onboarding files, where the personal data is unavoidable and the regulator is attentive. The design put extraction inside the client's own environment, kept documents and their derived vectors within the perimeter, and produced an audit record for every field extracted and every human correction. The engagement is described in the KYC document intelligence case study. Nothing in that architecture was exotic; it was decided in week one rather than week twelve.

A pre-build compliance checklist

  • Classify every field the MVP touches as personal, sensitive or neither, in writing
  • Write the purpose statement and check the consent notice already covers it
  • Decide residency per data class before choosing a model provider
  • Confirm no-training and retention terms with every vendor, in the contract
  • Design the deletion cascade across database, vectors, caches, logs and traces
  • Turn on structured audit logging before the first real user, not after
  • Mask personal data in evaluation sets and traces by default
  • Name the person who answers a data principal request and give them a runbook

AI audit trails: what regulators will ask to see covers the logging standard in detail, a security questionnaire for AI vendors gives you the questions to put to providers, and AI MVP development in India sets out the delivery and data rules alongside cost.

Compliance in an MVP is cheap when it is architecture and expensive when it is remediation, and the difference is decided in the first week.

Frequently asked questions

Does the DPDP Act apply to a small AI pilot?

▾

Yes. The Act applies to digital personal data processed in India regardless of the size of the system or the number of users. A pilot with ten real customer records carries the same notice, consent, erasure and security obligations for those records as a production system does.

Can an AI MVP use OpenAI or Anthropic models under the DPDP Act?

▾

Generally yes. The Act permits cross-border transfer except to territories the government restricts by notification. You need contractual no-training and retention terms, a documented processing region, and sector approval where a regulator such as the RBI imposes stricter residency rules.

What does erasure mean for a retrieval system?

▾

Deleting the source record is not enough. Erasure must cascade to embeddings in the vector store, cached chunks, observability traces, evaluation exports and backups. Build the link from every derivative back to its source record at ingestion time, because retrofitting it usually means re-indexing the whole corpus.