azyware
Technology

Machine Learning Development Services, security and the DPDP Act: a compliance checklist

EZ
Eazyware
· 7 min read
Quick answer

Is machine learning development services compliant with the DPDP Act?

No machine learning service is compliant in itself. The DPDP Act binds you as the data fiduciary, so compliance is a property of the system you build: lawful basis for training, a mapped data path, retention that survives erasure requests, and logs an auditor can read.

No machine learning development service is compliant or non-compliant in itself. India's DPDP Act binds you as the data fiduciary, so compliance is a property of the system that gets built: a lawful basis for using personal data in training, a mapped data path, retention rules that survive erasure requests, and audit logs someone outside your team can read.

This is the checklist we work through on regulated builds, organised by obligation rather than by technology. It covers where personal data actually hides in a machine learning system, what residency means once a hosted model is in the loop, the controls that satisfy each duty, and the cases where this level of rigour is genuinely the wrong response.

What the DPDP Act asks of a model

The Digital Personal Data Protection Act, 2023 sets out duties for data fiduciaries handling the personal data of data principals in India. The ones that bite on machine learning are notice and consent, purpose limitation, accuracy, storage limitation, reasonable security safeguards, breach reporting to the Data Protection Board, and the principal's rights to access, correction and erasure. Significant data fiduciaries carry extra duties including a data protection officer and periodic assessments.

Machine learning strains three of those. Purpose limitation is strained because a dataset gathered to deliver a service is convenient to reuse for training. Storage limitation is strained because training corpora are copied, versioned and rarely deleted. Erasure is strained because a record removed from a database may still have influenced model weights that remain in production.

Cross-border transfer is permissible under the Act except to territories the central government restricts, which makes residency a contractual and architectural decision rather than a blanket prohibition. The broader obligations are set out in DPDP Act 2023 and AI.

One practical consequence is worth stating plainly. Consent gathered for delivering a service does not automatically cover training a model on the same records, so either the notice names that purpose or the training set is built from data that no longer identifies anyone. Deciding which of those two routes you are taking, per dataset, is the single most useful hour a machine learning programme can spend with its legal team.

Obligations mapped to controls

This is the table we put in front of a client's legal team, because it converts the statute into things an engineer can build and an auditor can inspect.

ObligationWhat it means for an ML systemControl to implement
Notice and consentTraining on personal data needs a stated purpose the principal agreed toConsent artefact stored with the record, carried into the training manifest
Purpose limitationService data is not automatically training dataSeparate training corpus with a documented lawful basis per field
Data minimisationMost features do not need identifiersPseudonymise at ingest; identifiers stay in a keyed vault
Storage limitationVersioned corpora accumulate silentlyDated corpus versions with an expiry and an owner
Erasure rightsA deleted record may persist in weightsDeletion propagates to the corpus, with a retraining window stated in the policy
Security safeguardsNotebooks and feature stores are the weak pointsRole-based access, encryption at rest and in transit, no production data on laptops
Breach reportingModel endpoints are part of the estateLogging, alerting and a rehearsed reporting path to the Data Protection Board
Auditability"The model decided" is not an answerImmutable inference logs with inputs, version, score and the human action taken

Where personal data hides in a machine learning system

Teams map the database and stop. Personal data in a machine learning system sits in at least five places, and four of them are usually missing from the data inventory.

The training corpus and the feature store

Both are copies, and copies are where governance fails. Give every corpus version a date, an owner, a lawful basis and an expiry. Features derived from personal data are still personal data; a customer-level aggregate is not anonymous simply because it has been averaged.

The inference request path

Every prediction carries a payload. If the model is called through a hosted API, that payload leaves your network, and possibly the country, on every request. Decide deliberately between a region-pinned managed service, a self-hosted open-weight model, and a redaction layer that strips identifiers before the call. Self-hosting is the strongest answer and the most expensive; our self-hosted agentic AI work starts at $31,500 or ₹20.8 lakh plus infrastructure for exactly this reason.

Logs, traces and prompts

Observability tooling captures request bodies by default, which means your tracing system quietly becomes a second copy of the personal data with a weaker access model. Redact at the emitter, set a shorter retention on traces than on the primary store, and put the tracing platform in the same access review as the database.

The model itself

A model trained on personal data can memorise it, particularly at low data volumes or with repeated records. Treat weights as an asset with a residency and access position of their own, not as a neutral artefact, and test for memorisation before exposing a model to users outside the organisation.

The compliance checklist

  • Data map first. List every location personal data reaches: source system, corpus, feature store, inference path, logs, weights, backups.
  • Lawful basis per field, recorded in the training manifest rather than asserted in a slide.
  • Pseudonymise at ingest, keeping identifiers in a separate keyed store with its own access control.
  • Residency decision written down for training, evaluation, inference and logging, each stated separately.
  • Retention and deletion windows that include corpus versions and the retraining cadence that propagates an erasure request.
  • Role-based access with periodic review, covering notebooks, feature stores and tracing tools, not only the production database.
  • Immutable inference logs recording input reference, model version, output and the human decision that followed.
  • Sub-processor register naming every model provider and cloud region, with the contractual terms attached.
  • A rehearsed breach path with named roles and a reporting route to the Data Protection Board.

Row-level controls belong in the same conversation once analysts and models share a warehouse, which is the subject of row-level security for AI analytics.

When this checklist is the wrong response

If your model never touches personal data, most of this does not apply. A demand forecast built on aggregate sales by store and week, a predictive maintenance model on sensor readings, or a document classifier on published regulations carries no DPDP exposure worth a governance programme. Running the full checklist anyway converts a four-week build into a three-month one and teaches the organisation that compliance is theatre.

It is also the wrong response as a substitute for a decision you have not taken. Teams sometimes commission an exhaustive privacy review of a use case nobody has agreed to fund, and the review becomes the project. Decide the use case, scope the data it genuinely needs, then apply the controls that data attracts.

And it is wrong if it becomes a single document signed once. The obligations that matter are continuous: access reviews, corpus expiry, retraining after deletions, and sub-processor changes when a model provider moves a region. AI governance for mid-size companies sets out the minimum routine for organisations with no compliance function.

What compliant delivery costs

Compliance work is not a separate line item for us; it is how a build is scoped. A machine learning engagement starts at $17,500 or ₹11.2 lakh and runs to $70,000 or ₹46.4 lakh, and the data map, residency decision and logging design happen in the first fortnight rather than at the security review. A ten-day Sprint Zero at $3,250 or ₹2,00,000 is often the right place to settle residency before anyone writes code. Post-launch, Care Plans run from $1,000 or ₹68,000 a month to $5,250 or ₹3,40,000 with a named engineer, and starting figures are published on the pricing page. You own the code, the infrastructure, the prompts and the documentation, which matters here: a compliance position you cannot inspect is not a compliance position.

Financial services carry an extra layer of expectations on outsourcing, localisation and audit, covered in RBI guidelines and AI.

Judge the spend by what it removes. A data map and a written residency decision cost a few days and eliminate the rework that follows a failed security review six weeks before launch, which is where compliance genuinely gets expensive.

What this looked like on a live build

An NBFC needed document intelligence for KYC and loan onboarding, which is about as personal as data gets. The architecture decisions were made before the model decisions: private deployment so documents never left controlled infrastructure, identifiers separated from the features used for classification, and an immutable log of every extraction with the reviewer's action attached. The result is described in the KYC document intelligence case study, and the sequence is the transferable part.

A security questionnaire for AI vendors gives you the supplier-side controls list to pair with this. The Ministry of Electronics and Information Technology publishes the data protection framework behind the Act, and the OWASP Machine Learning Security Top 10 catalogues the attack classes, including data poisoning and model inversion, that your safeguards need to cover.

Build the data path you would be willing to show an auditor on the first day, because you will be showing it eventually.

Frequently asked questions

Is machine learning compliant with the DPDP Act?

▾

Machine learning is neither compliant nor non-compliant by itself. The Act places duties on you as the data fiduciary, so compliance depends on the system: a lawful basis for training data, purpose limitation, retention and deletion that reach corpus versions, reasonable security safeguards, and logs that let you answer a data principal's request within the statutory window.

Does training data have to stay in India?

▾

Not as a blanket rule. The DPDP Act permits cross-border transfer except to territories the central government restricts, so residency is an architectural and contractual choice. Sector regulators may be stricter, and the harder question is usually inference rather than training, because a hosted model API can move personal data abroad on every single request.

How do you handle an erasure request for data used to train a model?

▾

Delete the record from the source store, remove it from the versioned training corpus, and retrain within the window stated in your retention policy so the deletion reaches the weights. Write that window down in advance. Keeping the corpus separate from the production database and versioning it by date is what makes this tractable at all.