AI Discovery Sprint, security and the DPDP Act: a compliance checklist
Is AI discovery sprint compliant with the DPDP Act?
Yes, an AI discovery sprint can be fully DPDP-compliant, because a well-run sprint needs almost no personal data. Work from redacted or synthetic samples, keep processing inside your own cloud account, set a retention date before day one, and log every access.
Yes. An AI discovery sprint is compliant with India's DPDP Act when it is scoped so that personal data barely enters it. Ten days of use-case discovery needs schemas, volumes, sample documents and access patterns, not live customer records. Redact, tokenise or synthesise before the first working session and most of the statutory weight falls away.
This article maps each DPDP obligation that touches an early AI engagement to the control that satisfies it, sets out the architecture we use for AI discovery sprint security, gives the real cost of running one under those constraints, and is honest about the cases where a ten-day sprint is the wrong container for regulated data.
What the DPDP Act actually requires during discovery
The Digital Personal Data Protection Act, 2023 governs digital personal data processed in India, and personal data processed outside India where goods or services are offered to Indian data principals. The Ministry of Electronics and Information Technology publishes the Act and the rules made under it on its data protection framework page. The obligations that bite during a short engagement are narrower than most legal teams expect.
Four apply directly. Purpose limitation: personal data may be processed only for the purpose the data principal agreed to, and "exploring an AI use case" is rarely that purpose. Data minimisation: you take what the purpose needs and nothing more. Storage limitation: you erase once the purpose is served. Security safeguards: the data fiduciary, which is you and not your development partner, must take reasonable measures to prevent a breach and must report one.
Two more shape the contract. Your development partner is a data processor acting on your instructions, so the agreement, not goodwill, defines what it may do with a file. And cross-border transfer is permitted except to countries the central government restricts, which makes residency a commercial and sector-regulator question more than a statutory bar. The definitions are set out in full in our DPDP Act glossary entry.
The practical consequence is simple. If personal data never moves, most obligations never fire. An AI discovery sprint is a decision-making exercise: it needs to know the shape of your data, its volume and its messiness, not the contents of any particular customer record.
DPDP obligations mapped to discovery sprint controls
This is the table we walk through with a client's data protection officer before anything is signed. Each row names the obligation, the control and the artefact that proves the control existed.
| DPDP obligation | What it means in a sprint | Control we apply | Evidence produced |
|---|---|---|---|
| Purpose limitation | No processing beyond the agreed use case | Written scope naming each dataset and its purpose | Signed scope document |
| Data minimisation | Samples, never table dumps | 20 to 200 redacted records per workflow | Sample manifest with counts |
| Storage limitation | Erase when the sprint ends | Retention date fixed on day zero, automated purge | Deletion record |
| Security safeguards | Reasonable technical measures | Processing inside your cloud account, SSO, no local copies | Access log export |
| Processor obligations | Vendor acts only on instruction | NDA and processing addendum before session one | Executed agreements |
| Breach notification | You must notify, quickly | Named incident owner both sides, 24-hour internal SLA | Incident runbook |
| Lawful basis | Something authorises the analysis | Legal review per source system, not per project | Basis register |
The controls we set before day one
Before the first working session of a regulated sprint we agree seven items in writing. It takes about ninety minutes with the right people in the room, and it removes nearly every argument that would otherwise surface in week two.
- A named lawful basis per dataset. Consent, legitimate use or contractual necessity, recorded against each source system rather than against the engagement as a whole.
- Redaction at your boundary, not ours. Names, phone numbers, account numbers and government identifiers are stripped before the file moves, by a script your team runs and keeps.
- One processing location. Everything runs in a single cloud account and region you own, commonly ap-south-1, with no copy on an engineer's laptop and no shared drive.
- A model traffic policy. Which providers may see which classes of text, decided before anyone opens a notebook, using zero-retention endpoints or a self-hosted model where the class demands it.
- A retention date. Written on the scope document as a calendar date, not as "after the project ends".
- Access through your identity provider. Sprint engineers get accounts in your tenant with least-privilege roles, revoked on the final day.
- An exportable audit log. Who read what, when, in a format the person who will audit you has already accepted.
What does a DPDP-safe discovery sprint cost?
The same as any other discovery sprint. Eazyware's AI Discovery Sprint is $3,250 or ₹2,00,000, fixed price, ten working days, and credited against the build that follows. The compliance work sits inside that figure because it is not a separate phase; it is how the sprint is scoped in the first place. Running the environment inside your own account adds your cloud bill, which for ten days of light exploratory workloads is small.
If the sprint concludes that the use case is worth proving, a three-week AI POC sprint follows at $6,250 to $10,500, or ₹4,00,000 to ₹6,80,000. Every starting figure is published on the pricing page. Indian clients are invoiced in INR with GST; international clients in USD.
The cost people forget is internal. A compliant sprint needs roughly four hours from your data protection officer or counsel, and about a day of a data engineer's time to produce and spot-check redacted samples. Budget both before you commit to a start date.
Architecture: where the data actually sits
Residency and the perimeter
We deploy the sprint environment into a project inside your own cloud account. Object storage, notebooks and any temporary index live there. Nothing is copied to an Eazyware account at any point. That makes you the owner of the perimeter and turns data residency into a configuration setting rather than a promise in a proposal. Where a client wants no model traffic to cross the boundary at all, the pattern is the one described in zero data egress.
Redaction and synthetic substitutes
Most discovery questions can be answered from documents in which every identifier has been replaced by a stable token. A loan file still shows its layout, its exception rate and its language mix once the customer name has become CUST_0041. For throughput and volume questions we generate synthetic records that match the real distribution. The rule we apply is blunt: if a sample has to keep a real identifier to be useful, that is a finding about your process, not a reason to copy the field.
Retention and proof of deletion
The sprint environment carries a destroy date in its infrastructure code. On that date the bucket, the index and the notebooks are removed, and we issue a deletion record naming exactly what was destroyed and when. The findings, the architecture diagrams and the evaluation plan survive the sprint; the data does not.
When a discovery sprint is the wrong place for this
There are three situations where we either decline or change the shape of the work, and it is worth knowing them before you write a purchase order.
First, if the only honest way to answer the question is to process live production personal data at volume, ten days is the wrong container. That belongs in a longer programme with a data protection impact assessment, a full processing agreement and your security team present throughout.
Second, if you are an RBI-regulated entity and the use case touches customer financial data, outsourcing rules layer audit rights, access and exit obligations on top of DPDP that a fixed ten-day engagement cannot satisfy by itself. Read RBI guidelines and AI before scoping.
Third, if your organisation has no lawful basis for analysing the dataset at all, a sprint cannot manufacture one. The honest outcome is a narrower sprint on data you can lawfully use, or a consent redesign first and the AI question second.
What this looks like on a real engagement
An NBFC came to us wanting document intelligence for KYC and loan onboarding, which is about as sensitive as Indian data gets. The discovery work ran entirely on redacted material: identity documents with number fields masked, bank statements with account numbers tokenised, and a synthetic set for throughput testing. The production system that followed was deployed inside the client's own perimeter with an audit trail on every extraction. The engagement is written up as private document intelligence for KYC and loan onboarding.
None of that slowed the sprint. Producing and checking two hundred redacted documents took the client's data engineer half a day, and it bought a clean answer on residency that would otherwise have consumed a fortnight of email.
Checklist before the sprint starts
- NDA and data processing addendum signed before the first working session
- Lawful basis confirmed in writing for every dataset in scope
- Redaction script run, and its output sampled by someone who knows the data
- Cloud project created in your account, in your region, with billing alerts on
- Model provider and retention policy agreed for every class of text
- Retention and destroy date written into the scope document
- Named incident owner on both sides, with out-of-hours contact details
- Format of the deletion evidence agreed with whoever will audit you
Related reading
DPDP Act 2023 and AI covers the obligations across a full AI programme rather than a single sprint, AI audit trails describes the evidence a regulator will actually ask to see, and the AI discovery sprint explains what the ten days produce and who needs to be in the room.
Compliance in a discovery sprint is not a legal review bolted on at the end; it is a scoping decision made on day zero, and the cheapest version of it is simply not moving the data.
Frequently asked questions
Does the DPDP Act apply to a short AI discovery engagement?
▾
It applies to any digital personal data processed in India, whatever the length of the engagement. A ten-day discovery sprint escapes most obligations not because it is short but because it can be run on redacted and synthetic samples. The moment real personal data enters the sprint, every obligation applies in full.
Can our data leave India during an AI discovery sprint?
▾
The DPDP Act permits cross-border transfer except to countries the central government restricts, so residency is usually a contractual and sector-regulator question rather than a statutory bar. In practice we run sprints inside the client's own cloud account in an Indian region, which removes the argument and satisfies regulators such as the RBI.
Who is responsible if a breach happens during the sprint?
▾
You remain the data fiduciary and carry the notification duty; your development partner is a processor acting on your written instructions. That is why we sign an NDA and a processing addendum before the first working session, name an incident owner on both sides, and keep an exportable access log from day one.