AI Voice Agent Development, security and the DPDP Act: a compliance checklist
Is AI voice agent development compliant with the DPDP Act?
AI voice agent development is compliant with the DPDP Act only if you build for it. A voice agent is a processor handling personal data, so you need lawful notice and consent before recording, a stated purpose, a retention clock, deletion on request and an audit trail covering every call.
AI voice agent development is compliant with the DPDP Act only if you build for it. Voice calls carry personal data, so the Act applies from the first second of audio. You need notice and consent before recording, a stated purpose, a retention clock, a working deletion path and an audit trail covering every call and every model that touched it.
Below is the obligation-by-obligation checklist we work through on Indian voice projects, the architecture that satisfies each one, the questions auditors ask that teams are not ready for, and the point at which compliance work is genuinely disproportionate.
What the DPDP Act actually requires of a voice agent
The Digital Personal Data Protection Act, 2023, published by the Ministry of Electronics and Information Technology, applies to digital personal data processed in India. A recorded phone call is digital personal data: the voice itself, the phone number, and whatever the caller says about their account, their health or their loan. Nothing about the data being spoken rather than typed changes the obligations. The Act's text and subsequent rules are available from MeitY, and reading the definitions section before you design the pipeline saves rework.
Three roles matter. You, the business running the agent, are the Data Fiduciary. Your development partner and your cloud, speech and model vendors are Data Processors acting on your instructions. The caller is the Data Principal, and they hold rights you must be able to service: access, correction, erasure and grievance redressal.
One consequence catches teams out. Consent must be free, specific, informed and unambiguous, and it must be obtainable for a stated purpose. "We record calls for quality purposes" does not cover using those recordings to fine-tune a speech model. If you intend to train on your audio, say so, separately, and let the caller decline without losing service.
Obligation to control: the mapping
Each row is a requirement on the left and the engineering control that discharges it on the right. This is the table we attach to the architecture document.
| Obligation | What it means on a call | Control that satisfies it |
|---|---|---|
| Notice and consent | Caller is told before recording starts and can decline | Pre-roll consent prompt, consent event written to an immutable log with timestamp and call identifier |
| Purpose limitation | Audio used only for the purpose stated | Separate storage buckets and access policies per purpose; training use behind an explicit second consent |
| Data minimisation | Do not capture more than the task needs | Redact card numbers and identifiers in the transcript pipeline; drop raw audio once the transcript is verified |
| Storage limitation | Audio and transcripts are not kept forever | Lifecycle rules with a retention period per data class, enforced by the storage layer not by a script |
| Data Principal rights | Access, correction and erasure on request | Caller identifier indexed across audio, transcript, logs and vector store so deletion reaches every copy |
| Security safeguards | Reasonable measures against breach | Encryption in transit and at rest, scoped service accounts, no standing human access to raw audio |
| Breach notification | Report to the Board and affected principals | Access logging and alerting good enough to establish what was exposed and to whom |
| Processor accountability | Vendors bound by contract | Data processing agreements with every speech, model and telephony vendor, with sub-processor lists |
Where does the audio actually go?
Answer this first, in writing, before any other compliance work. A typical voice agent sends audio to a telephony provider, a speech to text vendor, a language model API and a text to speech vendor, and stores recordings and transcripts somewhere. That is five places, often in three countries, and most teams have never drawn the diagram.
The DPDP Act permits cross-border transfer except to countries the government restricts, so offshore processing is not automatically unlawful. But sectoral rules can be stricter than the Act. If you are in lending, payments or insurance, RBI guidelines on outsourcing and localisation may require the data to stay in India regardless of what the Act allows. Healthcare adds its own expectations around records.
Where residency is mandatory, the architecture changes: Indian-region telephony, an Indian-region speech vendor or a self-hosted speech model, and either an Indian-region model endpoint or an open-weight model you run yourself. That is what our self-hosted agentic AI work exists for, starting at $31,500 or ₹20,80,000 plus infrastructure. It costs more and it removes the argument.
Consent and recording, done properly
The pre-roll
The agent states, before anything else, who is calling or being called, that the call is recorded, why, and how to opt out. Keep it under eight seconds or callers stop listening. For outbound calls, check the caller against DND and consent registers first; telemarketing regulations sit alongside the DPDP Act, not underneath it.
The consent record
Store consent as an event, not a flag on a customer row. Each event carries the call identifier, timestamp, the exact wording played, the version of that wording, and the outcome. When someone asks in 2028 what a caller agreed to in 2026, a flag cannot answer and an event can. The pattern is covered in voice AI compliance in India.
Withdrawal
Consent can be withdrawn as easily as it was given. In practice that means a documented route, whether by phone, email or in the app, that triggers deletion across audio, transcripts, logs and any retrieval index. If your vector store holds embedded call content and your deletion job does not touch it, the data is still there.
Security risks specific to voice
Voice introduces failure modes text agents do not have. The first is impersonation: a caller who knows a customer name and phone number can sound entirely legitimate, so any agent that can move money, change a delivery address or reveal a balance needs a verification step that does not rely on what the caller volunteers. The second is prompt injection through speech. A caller reading out instructions can attempt to steer the agent, and because the transcript feeds the model as ordinary text, the defence is the same as elsewhere: treat caller speech as untrusted input, never as instruction, and enforce limits in the tool layer rather than in the prompt.
The audit trail auditors actually ask for
By the time someone audits the system, the questions are specific and the answers have to be retrievable per call. Expect to be asked for the consent wording played, the model and version that handled the turn, the tools called with their arguments, the caller identity check performed, and who has accessed the recording since. Building this later is expensive, because the events were never emitted. AI audit trails lists what to log from day one.
Two practices make the difference. Log tool calls with arguments and results rather than a summary, and keep the log immutable and separate from application storage. Both cost little during the build and are close to impossible to retrofit.
A pre-launch security checklist
- Draw the data flow diagram naming every vendor, region and retention period
- Sign data processing agreements with each vendor, including their sub-processors
- Play consent pre-roll on every call and record the consent event immutably
- Redact payment details and government identifiers before transcripts are stored
- Set lifecycle deletion on audio and transcripts, enforced by the storage layer
- Test the erasure path end to end, including the retrieval index and backups
- Remove standing human access to raw audio; require a ticket and log every read
- Run a red-team pass for prompt injection through spoken input and caller impersonation
- Name the grievance officer and publish a working contact route
When this level of compliance work is the wrong choice
If your agent takes no personal data at all, for example a store-hours and directions line with no recording and no caller identification, most of this is disproportionate. Play a short notice, keep no audio, and move on. Compliance effort should track the data you handle, not the anxiety in the room.
Equally, do not self-host purely for comfort. Self-hosted speech and models cost more to run and, unless someone owns patching and monitoring, are less secure than a well-configured managed service. Self-host when a sectoral rule, a contract or a genuine residency requirement forces it. Self-hosted LLMs for BFSI sets out where that line honestly falls.
What it costs to build this in
Compliance is not a separate invoice; it is a constraint on the architecture. Our multilingual voice agent builds start at $17,500 or ₹11,20,000 and run to $56,000 or ₹38,40,000 plus per-minute usage, with consent logging, redaction, retention and audit trails inside that scope. Residency requirements push the work towards self-hosted delivery at higher cost. Ongoing evidence, including eval runs and access reviews, sits in a Care Plan from $1,000 or ₹68,000 a month plus the $750 or ₹40,000 AI add-on. All figures are on the pricing page, and you own the code, prompts and infrastructure at the end either way.
For regulated sectors we usually start with a ten-day Sprint Zero at $3,250 or ₹2,00,000, credited to the build, because the data flow diagram and the residency answer are cheaper to settle before code exists than after. FinTech and BFSI projects almost always need it.
Related reading
DPDP Act 2023 and AI: what Indian companies must do covers the obligations across AI systems generally, and permission-aware retrieval explains how to stop an agent answering from documents the caller is not entitled to.
Treat consent, retention and the audit trail as architecture decisions made in week one, because every one of them is painful to add in week twenty.
Frequently asked questions
Does the DPDP Act allow AI voice agents to record calls?
▾
Yes, provided the caller is given clear notice and gives consent for a stated purpose before recording begins, and the recording is kept only as long as that purpose requires. Consent for quality monitoring does not extend to training a speech model on the audio; that needs its own explicit consent.
Must voice agent data stay in India?
▾
Not under the DPDP Act alone, which permits cross-border transfer except to restricted countries. Sectoral regulators are stricter: RBI-regulated lending and payments work often requires Indian storage. Check your sector's rules, then choose Indian-region telephony, speech and model endpoints where residency is required.
How long should you keep AI voice agent recordings?
▾
Set a retention period per data class tied to the purpose and any sectoral requirement, then enforce it with storage lifecycle rules rather than a cleanup script. Many teams keep verified transcripts longer than raw audio, deleting the audio within days once the transcript and outcome have been confirmed.