azyware
Technology

AI Copilot Development, security and the DPDP Act: a compliance checklist

EZ
Eazyware
· 7 min read
Quick answer

Is AI copilot development compliant with the DPDP Act?

A copilot is compliant with the DPDP Act when it processes personal data for a purpose the user consented to, keeps it no longer than needed, respects access rights inside retrieval, and leaves an audit trail. The Act sets obligations on you as data fiduciary, not on the model vendor.

Yes, an AI copilot can be compliant under India's Digital Personal Data Protection Act, 2023. The Act places obligations on you as data fiduciary: process personal data only for a notified purpose with consent, retain it no longer than necessary, and honour erasure and access requests. AI copilot development security is the work of turning those obligations into architecture.

None of that is automatic. This checklist walks each obligation in turn, states the engineering control that satisfies it, and flags the three places copilots most often leak. It assumes a copilot embedded in a multi-tenant SaaS product handling Indian customer data.

Obligation one: purpose and notice

Under the DPDP Act a data fiduciary must give notice of the purposes for which personal data is processed and obtain consent for those purposes. A copilot introduces a new purpose. Support conversations collected to resolve tickets were not collected to train or ground a copilot, and reusing them without updating the notice is the most common compliance gap we find in existing products.

The engineering control is a purpose tag on every corpus. Each index the copilot retrieves from carries the purpose it was collected for and the consent basis, and retrieval refuses sources whose purpose does not cover the current use. That sounds heavy until the first regulator question, at which point it is the only defensible answer. MeitY, the ministry that administers the Act, publishes the statute and the rules under it, and the notice and consent provisions are where copilot projects most often need legal review before engineering starts.

Obligation two: data residency and where inference happens

The Act does not impose blanket localisation, but sectoral rules do, and RBI's directions on payment data and many enterprise contracts require Indian processing. The practical question for a copilot is simple: when a user types a message containing a customer's name, where does that text travel?

Three architectures satisfy different constraints, and the choice drives cost more than any other decision in the project.

ArchitectureWhere inference runsSuitsTrade-off
Hosted API, default regionVendor's global regionsNon-personal or fully redacted inputsCheapest and fastest; weakest residency story
Hosted API, Indian or contracted regionVendor region with a data processing agreementMost SaaS products with Indian customersModel choice narrows; residency documented
Cloud provider in-regionYour cloud account, Indian regionFinancial services, health, government-adjacentHigher run cost; full control of logs and retention
Self-hosted open-weight modelYour VPC or on-premise GPUsZero egress mandates, sovereign requirementsInfrastructure and evaluation burden sits with you

Whichever you choose, write it into the record of processing and into the customer-facing notice. The data residency entry in our glossary sets out the vocabulary your legal team will expect to see.

Obligation three: minimisation, redaction and retention

Personal data must be limited to what the purpose requires and erased when the purpose ends. Copilots break this quietly, because the prompt is a new copy of the data and the trace is another.

  • Redact before the model call. Strip Aadhaar numbers, PAN, card numbers, phone numbers and email addresses from prompts unless the task genuinely needs them. PII redaction belongs in the request path, not in a nightly job.
  • Set trace retention deliberately. Observability stores full prompts and retrieved chunks. Thirty days is a defensible default for debugging; indefinite is not.
  • Separate the eval set. A golden question set built from real conversations is personal data. De-identify it, store it under the same retention rules, and record it in your processing inventory.
  • Disable vendor training. Confirm in the contract that your inputs are not used to train the provider's models, and record the setting that enforces it.
  • Propagate deletion. When a customer exercises erasure, the record must leave the primary store, the vector index, the trace store and any cached response. Build that path on day one, because retrofitting it means finding every copy.

Deletion deserves a paragraph of its own because it is where architecture and law meet awkwardly. A vector index is not a database row: embeddings derived from a customer record still encode it, and a re-index is often the only honest way to remove them. Decide early whether you will re-index on a schedule or delete by identifier, test the path with a real erasure request during the build, and record the expected completion time so support can answer the customer truthfully.

Obligation four: access control inside retrieval

This is the failure we see most often, and it is a security failure before it is a compliance one. Retrieval that ignores permissions lets a support agent ask a question and receive a passage from a record they could never open in the interface.

The control is filtering at query time by the same identity the application uses, never by post-processing the answer. Every chunk carries its tenant and its object permissions, and the vector query is constrained before ranking. We set out the pattern in permission-aware retrieval, and the tenancy decisions behind it in multi-tenant LLM architecture for SaaS.

Actions need the same treatment in the opposite direction. A copilot that writes must call your existing endpoints as the signed-in user, so that your authorisation layer, rate limits and audit log apply unchanged. Giving the copilot a service account with broad rights is how a support copilot becomes a privilege escalation path.

Obligation five: security safeguards and breach reporting

The Act requires reasonable security safeguards and notification of personal data breaches. For a copilot this means the ordinary controls plus two that are specific to language models.

Prompt injection is an access-control problem

Content the copilot reads can contain instructions. A support ticket, a PDF, a web page or a calendar invite can all carry text telling the model to exfiltrate data or call a tool. OWASP's Top 10 for large language model applications lists prompt injection as the leading risk class, and the mitigation is not a cleverer system prompt: it is scoping tools narrowly, requiring approval for sensitive actions, and never letting retrieved content expand the copilot's permissions.

Audit trails that name the copilot

Every action the copilot takes is logged with the user on whose behalf it acted, the input, the tool called, the arguments and the outcome. If your audit log cannot distinguish an action taken by a human from one proposed by a copilot and approved by a human, you cannot answer the only question an investigation will ask. SSO, RBAC and audit logs covers the foundations this assumes.

Name a data protection point of contact for the copilot as well. Under the Act a data fiduciary must publish the details of a person who can answer questions about processing, and in practice that person needs to understand what the copilot reads, what it writes and how long traces live. If nobody in the engineering team can answer those three questions from memory, the documentation is not finished.

When this checklist is the wrong amount of work

If your copilot answers questions about public product documentation and touches no customer record, most of this is overhead. Confirm no personal data enters the prompt, set a trace retention period, disable vendor training, and ship. Applying a financial-services control set to a documentation assistant delays a useful feature by a quarter for no reduction in risk.

Equally, if you are a Significant Data Fiduciary or handle children's data, this checklist is a floor rather than a ceiling: those categories attract additional obligations including data protection impact assessments and audits, and you need counsel rather than a blog post. The honest position is that compliance scales with the sensitivity of the data, not with the sophistication of the model.

What the compliance work costs

Residency, redaction, permission-aware retrieval and deletion propagation are engineering line items, not paperwork. In our AI copilot development for SaaS programme, which runs from $19,500 or ₹12,80,000 to $63,000 or ₹41,60,000, compliance-driven work typically accounts for a fifth to a third of the build in regulated sectors. All starting prices are published on the pricing page.

Where the constraint is zero data egress rather than documented residency, the answer is usually self-hosted: our agentic AI on your own infrastructure programme starts at $31,500 or ₹20,80,000 plus infrastructure. A ten-day Sprint Zero at $3,250 or ₹2,00,000 is the cheapest way to establish which of these you actually need before anyone writes code.

A comparable engagement is documented in the KYC document intelligence case study, where an NBFC needed document processing that never left its own environment.

The checklist, in order

  • Update the privacy notice to cover copilot processing before launch
  • Tag every retrieval corpus with its purpose and consent basis
  • Decide and document where inference runs, and contract for it
  • Redact personal identifiers in the request path, not after the fact
  • Set explicit retention for traces, caches and evaluation sets
  • Filter retrieval by the signed-in user's permissions at query time
  • Call your own APIs as the user, never through a broad service account
  • Log every copilot action with actor, input, tool and outcome
  • Build and test the erasure path across index, trace store and cache
  • Run a red-team pass for prompt injection before go-live

DPDP Act 2023 and AI: what Indian companies must do gives the wider statutory picture, a security questionnaire for AI vendors is the document to send before you sign anything, and AI governance for mid-size companies shows how to run this without a compliance department.

Compliance for a copilot is not a document you write at the end; it is four or five architectural decisions you make at the start and can evidence a year later.

Frequently asked questions

Does the DPDP Act require AI processing to stay in India?

▾

The Act itself permits cross-border transfer except to countries the government restricts, so it is not a blanket localisation rule. Sectoral regulators and enterprise contracts often are stricter, particularly for payment and health data, so most Indian SaaS copilots end up running inference in an Indian or contracted region.

Can we use customer support conversations to ground a copilot?

▾

Only if the purpose is covered by the notice and consent under which they were collected. Grounding a copilot is a new purpose from resolving a ticket, so update the notice, tag the corpus with its consent basis, and de-identify the conversations you keep in an evaluation set.

What stops a copilot from revealing records a user cannot see?

▾

Filtering retrieval by the signed-in user's permissions at query time, with tenant and object permissions stored on every chunk. Post-processing the answer is not sufficient, because the model has already seen the data. Actions should call your existing APIs as that user, so your authorisation layer applies unchanged.