Machine Learning Development Services, security and the DPDP Act: a compliance checklist
Is machine learning development services compliant with the DPDP Act?
No machine learning service is compliant in itself. The DPDP Act binds you as the data fiduciary, so compliance is a property of the system you build: lawful basis for training, a mapped data path, retention that survives erasure requests, and logs an auditor can read.
No machine learning development service is compliant or non-compliant in itself. India's DPDP Act binds you as the data fiduciary, so compliance is a property of the system that gets built: a lawful basis for using personal data in training, a mapped data path, retention rules that survive erasure requests, and audit logs someone outside your team can read.
This is the checklist we work through on regulated builds, organised by obligation rather than by technology. It covers where personal data actually hides in a machine learning system, what residency means once a hosted model is in the loop, the controls that satisfy each duty, and the cases where this level of rigour is genuinely the wrong response.
What the DPDP Act asks of a model
The Digital Personal Data Protection Act, 2023 sets out duties for data fiduciaries handling the personal data of data principals in India. The ones that bite on machine learning are notice and consent, purpose limitation, accuracy, storage limitation, reasonable security safeguards, breach reporting to the Data Protection Board, and the principal's rights to access, correction and erasure. Significant data fiduciaries carry extra duties including a data protection officer and periodic assessments.
Machine learning strains three of those. Purpose limitation is strained because a dataset gathered to deliver a service is convenient to reuse for training. Storage limitation is strained because training corpora are copied, versioned and rarely deleted. Erasure is strained because a record removed from a database may still have influenced model weights that remain in production.
Cross-border transfer is permissible under the Act except to territories the central government restricts, which makes residency a contractual and architectural decision rather than a blanket prohibition. The broader obligations are set out in DPDP Act 2023 and AI.
One practical consequence is worth stating plainly. Consent gathered for delivering a service does not automatically cover training a model on the same records, so either the notice names that purpose or the training set is built from data that no longer identifies anyone. Deciding which of those two routes you are taking, per dataset, is the single most useful hour a machine learning programme can spend with its legal team.
Obligations mapped to controls
This is the table we put in front of a client's legal team, because it converts the statute into things an engineer can build and an auditor can inspect.
| Obligation | What it means for an ML system | Control to implement |
|---|---|---|
| Notice and consent | Training on personal data needs a stated purpose the principal agreed to | Consent artefact stored with the record, carried into the training manifest |
| Purpose limitation | Service data is not automatically training data | Separate training corpus with a documented lawful basis per field |
| Data minimisation | Most features do not need identifiers | Pseudonymise at ingest; identifiers stay in a keyed vault |
| Storage limitation | Versioned corpora accumulate silently | Dated corpus versions with an expiry and an owner |
| Erasure rights | A deleted record may persist in weights | Deletion propagates to the corpus, with a retraining window stated in the policy |
| Security safeguards | Notebooks and feature stores are the weak points | Role-based access, encryption at rest and in transit, no production data on laptops |
| Breach reporting | Model endpoints are part of the estate | Logging, alerting and a rehearsed reporting path to the Data Protection Board |
| Auditability | "The model decided" is not an answer | Immutable inference logs with inputs, version, score and the human action taken |
Where personal data hides in a machine learning system
Teams map the database and stop. Personal data in a machine learning system sits in at least five places, and four of them are usually missing from the data inventory.
The training corpus and the feature store
Both are copies, and copies are where governance fails. Give every corpus version a date, an owner, a lawful basis and an expiry. Features derived from personal data are still personal data; a customer-level aggregate is not anonymous simply because it has been averaged.
The inference request path
Every prediction carries a payload. If the model is called through a hosted API, that payload leaves your network, and possibly the country, on every request. Decide deliberately between a region-pinned managed service, a self-hosted open-weight model, and a redaction layer that strips identifiers before the call. Self-hosting is the strongest answer and the most expensive; our self-hosted agentic AI work starts at $31,500 or ₹20.8 lakh plus infrastructure for exactly this reason.
Logs, traces and prompts
Observability tooling captures request bodies by default, which means your tracing system quietly becomes a second copy of the personal data with a weaker access model. Redact at the emitter, set a shorter retention on traces than on the primary store, and put the tracing platform in the same access review as the database.
The model itself
A model trained on personal data can memorise it, particularly at low data volumes or with repeated records. Treat weights as an asset with a residency and access position of their own, not as a neutral artefact, and test for memorisation before exposing a model to users outside the organisation.
The compliance checklist
- Data map first. List every location personal data reaches: source system, corpus, feature store, inference path, logs, weights, backups.
- Lawful basis per field, recorded in the training manifest rather than asserted in a slide.
- Pseudonymise at ingest, keeping identifiers in a separate keyed store with its own access control.
- Residency decision written down for training, evaluation, inference and logging, each stated separately.
- Retention and deletion windows that include corpus versions and the retraining cadence that propagates an erasure request.
- Role-based access with periodic review, covering notebooks, feature stores and tracing tools, not only the production database.
- Immutable inference logs recording input reference, model version, output and the human decision that followed.
- Sub-processor register naming every model provider and cloud region, with the contractual terms attached.
- A rehearsed breach path with named roles and a reporting route to the Data Protection Board.
Row-level controls belong in the same conversation once analysts and models share a warehouse, which is the subject of row-level security for AI analytics.
When this checklist is the wrong response
If your model never touches personal data, most of this does not apply. A demand forecast built on aggregate sales by store and week, a predictive maintenance model on sensor readings, or a document classifier on published regulations carries no DPDP exposure worth a governance programme. Running the full checklist anyway converts a four-week build into a three-month one and teaches the organisation that compliance is theatre.
It is also the wrong response as a substitute for a decision you have not taken. Teams sometimes commission an exhaustive privacy review of a use case nobody has agreed to fund, and the review becomes the project. Decide the use case, scope the data it genuinely needs, then apply the controls that data attracts.
And it is wrong if it becomes a single document signed once. The obligations that matter are continuous: access reviews, corpus expiry, retraining after deletions, and sub-processor changes when a model provider moves a region. AI governance for mid-size companies sets out the minimum routine for organisations with no compliance function.
What compliant delivery costs
Compliance work is not a separate line item for us; it is how a build is scoped. A machine learning engagement starts at $17,500 or ₹11.2 lakh and runs to $70,000 or ₹46.4 lakh, and the data map, residency decision and logging design happen in the first fortnight rather than at the security review. A ten-day Sprint Zero at $3,250 or ₹2,00,000 is often the right place to settle residency before anyone writes code. Post-launch, Care Plans run from $1,000 or ₹68,000 a month to $5,250 or ₹3,40,000 with a named engineer, and starting figures are published on the pricing page. You own the code, the infrastructure, the prompts and the documentation, which matters here: a compliance position you cannot inspect is not a compliance position.
Financial services carry an extra layer of expectations on outsourcing, localisation and audit, covered in RBI guidelines and AI.
Judge the spend by what it removes. A data map and a written residency decision cost a few days and eliminate the rework that follows a failed security review six weeks before launch, which is where compliance genuinely gets expensive.
What this looked like on a live build
An NBFC needed document intelligence for KYC and loan onboarding, which is about as personal as data gets. The architecture decisions were made before the model decisions: private deployment so documents never left controlled infrastructure, identifiers separated from the features used for classification, and an immutable log of every extraction with the reviewer's action attached. The result is described in the KYC document intelligence case study, and the sequence is the transferable part.
Related reading
A security questionnaire for AI vendors gives you the supplier-side controls list to pair with this. The Ministry of Electronics and Information Technology publishes the data protection framework behind the Act, and the OWASP Machine Learning Security Top 10 catalogues the attack classes, including data poisoning and model inversion, that your safeguards need to cover.
Build the data path you would be willing to show an auditor on the first day, because you will be showing it eventually.
Frequently asked questions
Is machine learning compliant with the DPDP Act?
▾
Machine learning is neither compliant nor non-compliant by itself. The Act places duties on you as the data fiduciary, so compliance depends on the system: a lawful basis for training data, purpose limitation, retention and deletion that reach corpus versions, reasonable security safeguards, and logs that let you answer a data principal's request within the statutory window.
Does training data have to stay in India?
▾
Not as a blanket rule. The DPDP Act permits cross-border transfer except to territories the central government restricts, so residency is an architectural and contractual choice. Sector regulators may be stricter, and the harder question is usually inference rather than training, because a hosted model API can move personal data abroad on every single request.
How do you handle an erasure request for data used to train a model?
▾
Delete the record from the source store, remove it from the versioned training corpus, and retrain within the window stated in your retention policy so the deletion reaches the weights. Write that window down in advance. Keeping the corpus separate from the production database and versioning it by date is what makes this tractable at all.