Self-hosted AI agents, security and the DPDP Act: a compliance checklist
Is self-hosted AI agents compliant with the DPDP Act?
Self-hosting does not make an AI agent compliant with the DPDP Act; it makes compliance achievable. The Act regulates purpose, consent, retention, erasure and breach reporting, none of which are solved by where the GPU sits. Self-hosting removes one hard problem and leaves six for your architecture.
No architecture is compliant by itself, and self-hosting an AI agent does not make it compliant with the DPDP Act. It makes compliance achievable. The Act regulates purpose, notice, consent, retention, erasure, breach reporting and grievance handling, none of which are answered by where the GPU sits. Self-hosting removes one problem and leaves the rest to your design.
This checklist maps each obligation in India's Digital Personal Data Protection Act, 2023 to the specific control that satisfies it in an agentic system, covers the security work that privacy law does not mention but auditors ask about, and is honest about the cases where self-hosting is an expensive answer to a question nobody asked.
What the DPDP Act asks of a system that reasons over personal data
The Digital Personal Data Protection Act, 2023 governs the processing of digital personal data in India. It calls the individual a Data Principal and the organisation deciding the purpose a Data Fiduciary, and it is administered by the Ministry of Electronics and Information Technology, which publishes the text and subsequent rules at meity.gov.in. The glossary entry on the DPDP Act 2023 summarises the definitions.
Three features of the Act bite hardest on AI agents. Purpose limitation means data collected for one stated purpose cannot be quietly repurposed as training or retrieval corpus. Erasure means that when consent is withdrawn, the data must go, including from places engineers forget: vector indexes, prompt logs, evaluation fixtures and trace stores. Accountability means the Fiduciary remains responsible for anything a processor does on its behalf, so an agent calling an external model API is your liability, not the vendor's.
An agent complicates all three because it copies data by design. A single run may pull a customer record into a prompt, embed a fragment into an index, write a summary into a ticket and leave the whole exchange in a trace. Each copy is personal data under the Act, and each needs a purpose, a retention period and a deletion path.
One consequence is worth stating plainly: retention has to be set per store rather than per system. Engineers reasonably assume the database governs the lifecycle, but the trace store usually outlives it, the cache is invisible on the architecture diagram, and the evaluation fixtures were copied into a repository months ago. Erasure that misses any of these is incomplete, and incomplete erasure is the failure most likely to be discovered by an auditor rather than by you.
Obligation by obligation: the control that satisfies each
The following table is the one we work through with a client's legal and security teams before writing code. The right-hand column is architecture, not policy, because a policy without a control is a promise.
| DPDP obligation | What it means for an agent | Control that satisfies it |
|---|---|---|
| Notice and consent | Users are told what the agent does with their data | Consent state read at request time; agent refuses out-of-scope purposes |
| Purpose limitation | Support transcripts cannot silently become training data | Separate stores per purpose; no cross-purpose reads without a new basis |
| Data minimisation | Only the fields needed for the task enter the prompt | Field-level projection in tool contracts; redaction before the model call |
| Accuracy | Wrong data is corrected across copies | Source of truth stays the system of record; index rebuilt on change |
| Storage limitation | Nothing is kept indefinitely by default | TTL on traces, prompt logs, caches and eval fixtures |
| Erasure on withdrawal | Deletion reaches every copy | Deletion job spanning database, vector index, logs and backups |
| Security safeguards | Reasonable measures against breach | Encryption at rest and in transit, RBAC, network isolation, key rotation |
| Breach notification | Board and affected principals informed | Alerting on anomalous access plus a rehearsed incident runbook |
| Grievance redressal | A named route for complaints | Published contact, logged responses, defined turnaround |
Does self-hosting satisfy data residency?
Self-hosting satisfies residency if, and only if, every component of the system sits in the region you claim. The Act permits cross-border transfer except to territories the Central Government restricts, so residency is often a contractual or sectoral requirement rather than a statutory one, which is why an RBI-regulated lender and a D2C brand reach different answers from the same law. The glossary covers data residency in more detail.
Where teams slip is the parts that are not the model. A self-hosted agent on Indian GPUs can still leak data through a hosted embedding API, an overseas error-tracking service, a managed vector database in another region, a translation call, or a webhook to a foreign SaaS tool. Draw the perimeter on a diagram, mark every outbound connection, and check each one against the claim you make in your privacy notice.
Residency also has a time dimension that diagrams miss. Backups, disaster-recovery replicas and log shipping frequently default to a different region from the primary workload, and a nightly snapshot crossing a border is still a transfer. Check the storage class, the replication policy and the retention of every bucket the agent touches, not only the subnet the model runs in.
The security work that privacy law does not name
Prompt injection is an access-control problem
An agent that reads a document, an email or a web page is reading instructions from an untrusted source. OWASP's Top 10 for LLM Applications lists prompt injection and sensitive information disclosure among the leading risks to these systems. The defence is not a cleverer system prompt; it is that the agent's tools cannot do anything the current user could not do themselves, enforced server-side.
Retrieval must respect permissions
A shared index that ignores access rights turns your knowledge agent into an internal data leak with a friendly interface. Filter at query time by the requesting user's entitlements rather than after generation. The pattern is set out in permission-aware retrieval.
Every action needs an audit trail
For each run, record who asked, what the agent was permitted to do, which tools it called with which arguments, what changed, and who approved anything gated. An audit trail is the difference between explaining an incident and guessing about it, and it is the first artefact an assessor asks for.
The compliance checklist
- Map the data. List every field the agent can read, each with a purpose and a lawful basis.
- Redact before the model. Strip identifiers not needed for the task; see PII redaction for what survives.
- Set retention per store. Traces, prompt logs, caches, indexes and eval fixtures each need an explicit TTL.
- Build and test the deletion path. Run a real erasure request end to end and confirm every copy is gone.
- Scope every tool. Narrow contracts, spend limits and approval thresholds, enforced on the server.
- Isolate the perimeter. No outbound calls from the inference subnet except to approved, in-region endpoints.
- Log and review. Audit trail on every action, with a named person reviewing overrides weekly.
- Rehearse the breach. A runbook tested once, not a document written once.
What the compliance work costs, and when to do it
Compliance is cheap when it is designed in and expensive when it is retrofitted, because retrofitting means re-indexing, rewriting tool contracts and backfilling audit data that was never captured. In an agentic AI programme on self-hosted infrastructure, which runs from $31,500 or ₹20,80,000 to $105,000 or ₹72,00,000 plus infrastructure, the privacy and security architecture is part of the build rather than a separate line.
If the obligations are not yet clear, a ten-day Sprint Zero at $3,250 or ₹2,00,000, credited to the build, produces the data map, the perimeter diagram and the retention matrix before a single prompt is written. After launch, the AI system add-on at $750 or ₹40,000 per month sits on top of a Care Plan from $1,000 or ₹68,000 a month and covers the evals and re-indexing that keep deletions honest. Our own posture is documented on the security page, and starting prices are on the pricing page.
Where self-hosting is the wrong answer
Self-hosting is the wrong answer when your data is not personal or sensitive, when volumes are low enough that GPU capacity sits idle, or when the requirement came from a slide rather than from a regulator, a contract or a risk assessment. A public-documentation assistant does not need private infrastructure, and paying for GPUs to reassure a board is an expensive form of comfort.
Run the risk assessment first and let it name the requirement, rather than starting from the architecture and reverse-engineering a justification for it.
It is also the wrong answer when the organisation cannot operate it. Self-hosting transfers patching, key rotation, capacity planning and incident response to you. A team without an on-call rota will run an older, less secure model on an unpatched host, which is worse for security than a well-governed managed service under a solid processing agreement. We say this in first calls regularly, and some prospects find it unwelcome.
What this looks like in practice
An NBFC needed document intelligence for KYC and loan onboarding where customer identity documents could not leave their environment. The controls that mattered were field-level extraction rather than whole-document prompting, confidence thresholds that routed uncertain files to human review, retention tied to the lending record rather than to the AI system, and an audit trail linking every extracted field to the page it came from. The engagement is described in the KYC document intelligence case study, and the wider pattern for regulated lenders appears in our fintech and BFSI work.
Related reading
DPDP Act 2023 and AI: what Indian companies must do covers the obligations beyond agents, Private AI for banks covers perimeter design under sectoral rules, and A security questionnaire for AI vendors gives you the questions to put to anyone selling you a private agent.
Self-hosting buys you the right to make these promises; only the controls above let you keep them.
Frequently asked questions
Does self-hosting an AI agent make it DPDP compliant?
▾
No. Self-hosting addresses where processing happens, which is one requirement among many. The Act also governs notice, consent, purpose limitation, retention, erasure, security safeguards, breach notification and grievance redressal. An agent on your own GPUs that keeps prompt logs forever and cannot delete a customer on request is not compliant.
Where does personal data hide in an AI agent?
▾
In more places than the database: prompt and completion logs, vector indexes, embedding caches, evaluation fixtures, trace stores, ticket summaries the agent wrote, and backups of all of these. Each is a copy subject to retention and erasure obligations, so each needs an explicit TTL and a tested deletion path.
Does the DPDP Act require data to stay in India?
▾
The Act permits cross-border transfer except to territories the Central Government restricts, so a blanket residency rule does not come from the statute itself. Strict residency usually comes from sectoral regulation, contracts or internal risk policy, which is why a regulated lender and a consumer brand reach different conclusions.