RAG Development Services, security and the DPDP Act: a compliance checklist
Is RAG development services compliant with the DPDP Act?
No retrieval system is DPDP-compliant as a category; an implementation is compliant or it is not. Under India's Digital Personal Data Protection Act, 2023 a RAG pipeline is processing on behalf of a data fiduciary, so retention, erasure and breach duties apply to the index, not just the source.
No retrieval system is compliant as a category. Under India's Digital Personal Data Protection Act, 2023, a RAG pipeline is processing personal data on behalf of a data fiduciary, so purpose limitation, retention limits, erasure rights and breach notification apply to your vector index exactly as they apply to the database it was built from. RAG development services security is therefore an architecture question, not a certificate.
This checklist walks the obligations that actually bite a retrieval system, the controls that satisfy each, the questions to put to any vendor, and the point at which the honest answer is that the data should never have been indexed.
Why a retrieval index is a compliance object
Teams treat the vector store as a cache. It is not. Embedding a document does not anonymise it; the chunks you store are usually the original text, and the index is a second, searchable copy of personal data sitting outside your system of record. Every obligation attached to the original attaches to the copy.
Three consequences follow immediately. An erasure request must delete the chunk and its embedding, not only the source row. A retention policy must expire index entries on the same clock as the source. And a breach of the vector store is a reportable personal data breach, because the searchable copy is personal data by any reading of the statute.
The Act's text and the rules published alongside it are available from the Ministry of Electronics and Information Technology, and our plain-language summary sits in DPDP Act 2023 and AI: what Indian companies must do. Read both before designing the pipeline, not after the security review.
Obligation by obligation: what it means for a retrieval pipeline
| DPDP obligation | What it means for RAG | Control that satisfies it |
|---|---|---|
| Purpose limitation | Documents indexed for support answers cannot quietly feed an analytics feature | Per-purpose indexes with separate access policies, documented at ingestion |
| Consent and notice | Personal data in the corpus was collected for a stated purpose; answering questions may be a new one | Purpose register mapping each source system to its lawful basis before ingestion |
| Data minimisation | Whole-document ingestion pulls in salary, health and identity fields nobody needs | Field-level redaction at ingestion, with a reviewed allow-list of fields |
| Accuracy | A stale chunk contradicting the live record produces a wrong answer with a citation | Freshness lag monitored as a metric; delta re-indexing on source change |
| Erasure rights | Deleting a customer row leaves their text in chunks and embeddings | Deletion propagates to index and cache by document identifier, with a tested runbook |
| Retention limits | Indexes outlive the policy because nobody set an expiry on them | Time-to-live per index, aligned to the source system's retention schedule |
| Security safeguards | A shared index exposes one tenant's documents to another's query | Permission filters applied at query time, tenant isolation, encryption at rest and in transit |
| Breach notification | You cannot report what you cannot reconstruct | Immutable query and access logs with retrieved document identifiers |
| Processor obligations | A model API is a sub-processor when prompts contain personal data | Contracted sub-processor list, zero-retention terms, or self-hosted inference |
Where does the data actually go?
The first question any reviewer asks is where personal data travels, and in a retrieval system it travels further than teams expect. Text leaves your network three times: to the embedding model at ingestion, to the reranker at query time, and to the generation model in the prompt. Each hop is a processor relationship that must be documented, and each is a place where a hosted API in another region creates a residency question.
Three architectures answer it. Fully hosted, where embeddings and generation run on a provider API under zero-retention terms, is acceptable for most corpora and is the cheapest to operate. Hybrid, where the index and reranker run inside your cloud account and only the final prompt leaves, narrows the exposure sharply. Fully private, where open-weight models run on hardware you control, is the answer when the data cannot leave at all; zero data egress describes how that is built, and self-hosted agentic AI starts at $31,500 or ₹20.8 lakh plus infrastructure.
Regulated lenders have a fourth consideration. RBI outsourcing expectations around localisation, audit access and exit plans apply to an AI vendor as they do to any other, a point covered in RBI guidelines and AI. Treat the retrieval vendor as a material outsourcing arrangement from the first call, and ask for the exit plan in the same conversation as the price.
The access control problem nobody budgets for
The single most common security defect in retrieval systems is a flat index. Documents from seven systems, each with its own permission model, are embedded into one store, and the chat interface cheerfully answers a junior analyst's question using a board paper. Nothing in the pipeline is broken; the permissions were simply never carried across, and nobody noticed because the demo used an administrator account. Access control in retrieval costs real engineering time because it has to be rebuilt per source system, and it is the line most often trimmed when a quote is being squeezed.
- Capture permissions at ingestion. Store the access control list of the source document as metadata on every chunk, and refresh it when the source changes.
- Filter at query time, not after generation. Retrieving forbidden text and then asking a model not to use it is not a control.
- Test with a low-privilege account. Most leaks are found in ten minutes by asking the system a question as a contractor.
- Redact before you embed. Identity numbers, contact details and health fields that no question needs should never enter the index; see PII redaction.
- Isolate tenants physically where the contract demands it. Metadata filtering is adequate for most SaaS; a separate index per tenant is what enterprise buyers will ask for.
- Log the retrieved document identifiers with every answer. Without them you cannot answer a regulator's question about what the system saw.
- Treat prompt injection as an access-control risk. A document containing instructions can steer a system that has tool access; OWASP's top ten risks for LLM applications is the reference list here.
How to implement the first three is set out in permission-aware retrieval, which is the post to hand an engineer who is about to build this.
What compliance work costs
Residency, redaction, permission mirroring and audit logging are not add-ons; they are what moves a build from the bottom to the middle of its range. Retrieval and knowledge engineering runs from $14,000 or ₹8.8 lakh to $49,000 or ₹32 lakh, and a regulated build with per-document permissions and an on-premises index sits in the upper half of that band. Starting prices for every programme are published on the pricing page, and Indian clients are invoiced in INR with GST.
Compliance is also a running cost. Someone must re-run access tests after every source-system change, rotate keys, review logs and re-index when documents move. Care plans start at $1,000 or ₹68,000 a month; the Enterprise tier at $5,250 or ₹3.4 lakh adds a one-hour response and a named engineer, which is what an audit response actually needs.
When the answer is not to index it
Sometimes the honest recommendation is that the corpus should not be indexed at all. Raw call recordings, unredacted medical notes and complete KYC packs frequently fail data minimisation the moment you ask which question needs them. If you cannot name a question that requires the sensitive field, exclude it and lose nothing.
We also advise against indexing when no one will own erasure. An organisation that cannot describe how a deletion request reaches the index within its statutory window should not create a second copy of personal data. Fix the deletion pipeline first, then index. Building on a corpus you cannot delete from is an incident waiting for a date.
A regulated build in practice
An NBFC needed document intelligence over KYC and loan onboarding files that could not leave its own infrastructure. The constraint set the architecture before any model was chosen: private deployment, redaction at ingestion, validation with exceptions routed to humans, and an audit trail covering every extracted field. The engagement is described in the KYC document intelligence case study, and the sequencing is the transferable part: residency and audit decided first, retrieval quality tuned second.
The pre-launch checklist
- Map every source system to a lawful basis and a stated purpose before ingestion
- Confirm and document the region for embedding, reranking and generation
- Get zero-retention terms in writing from every model provider you call
- Test a deletion request end to end, including index, cache and backups
- Set a retention clock on each index that matches its source
- Run an access test with the lowest-privilege account you have
- Enable immutable logging of queries, answers and retrieved document identifiers
- Agree who signs the breach assessment and within what timeframe
- Name the person accountable for quarterly re-testing after launch
Related reading
AI audit trails: what regulators will ask to see sets out the evidence you will be asked to produce, a security questionnaire for AI vendors is the list to send before you sign, and our security page describes how we handle client data during an engagement.
Compliance in a retrieval system is decided at ingestion, because everything you index becomes a second copy you must govern, delete and defend.
Frequently asked questions
Does the DPDP Act apply to a vector database?
▾
Yes, where the chunks contain personal data. Embedding does not anonymise text, so the index is a second searchable copy subject to the same purpose limitation, retention, erasure and breach obligations as the source system. Treat deletion and retention on the index as first-class requirements, not cleanup tasks.
Can we use a hosted model API and still meet DPDP requirements?
▾
Usually yes, provided the provider is documented as a processor, the contract includes zero-retention terms, the region is known and the purpose is recorded. Where data genuinely cannot leave your network, self-hosted open-weight inference is the alternative, starting at $31,500 or ₹20.8 lakh plus infrastructure.
How do we handle an erasure request in a RAG system?
▾
Delete by document identifier across the source, the chunk store, the vector index, any cache and the backups, then re-run the retrieval test for that identifier to confirm nothing surfaces. Build and rehearse this runbook before launch; retrofitting deletion into a live index is expensive and rarely complete.