Data Analytics Application Development, security and the DPDP Act: a compliance checklist
Is data analytics application development compliant with the DPDP Act?
Not automatically. India's DPDP Act applies to the personal data flowing through an analytics application, not to the category of software, so compliance depends on four architectural choices: what you copy into the warehouse, on what lawful basis, for how long, and who can see which rows.
Not automatically. The DPDP Act applies to the personal data flowing through an analytics application, not to the category of software, so compliance is decided by four architectural choices: what personal data you copy into the warehouse, on what lawful basis, how long derived tables keep it, and who can see which rows. Get those right and the rest is documentation.
This is a working checklist for data analytics application development security under India's Digital Personal Data Protection Act, 2023: the obligations that actually bite in a warehouse, the controls that satisfy each one, what it costs to build them in, and the cases where you are over-engineering.
Why analytics is the awkward case
Most compliance guidance assumes a transactional system with one copy of each record, a clear purpose and an obvious delete button. An analytics application is the opposite by design. It copies data out of source systems, fans it into staging, marts and aggregates, and keeps history precisely so that nothing is lost. Every property that makes a warehouse useful makes erasure, purpose limitation and retention harder.
The Digital Personal Data Protection Act, 2023, published by the Ministry of Electronics and Information Technology, sets duties on a Data Fiduciary covering notice and consent, purpose limitation, accuracy, retention, security safeguards and breach reporting to the Data Protection Board. The Act's text is short and worth reading once in full; our plain-language summary is in DPDP Act 2023 and AI, and the term itself is defined in the DPDP Act glossary entry.
The practical reading for an analytics team is that a derived table built from personal data is still personal data. Aggregation only takes you out of scope when the result cannot be reasonably re-identified, and a cohort of three customers in one pin code is not that.
There is a second awkwardness. In a transactional system the lawful basis sits next to the record, so anyone can check it. In a warehouse the basis was established upstream, three hops and two transformations ago, and by the time a field reaches a dashboard nobody remembers which consent covered it. That lineage has to be carried deliberately as a column, because it will not survive as institutional memory.
Obligations mapped to controls
| Obligation | What it means inside an analytics application | Control that satisfies it |
|---|---|---|
| Purpose limitation | Data pulled for billing cannot be quietly reused for marketing segmentation | Purpose tag on every source-to-mart mapping, reviewed at design time |
| Notice and consent | The lawful basis for each field must be traceable back to a source record | Consent identifier carried as a column through staging into marts |
| Data minimisation | Star schemas tend to copy whole source tables | Explicit column allow-list per ingestion, not SELECT * |
| Retention limits | History tables outlive the purpose that justified them | Per-table retention policy with an automated expiry job and evidence of runs |
| Right to erasure | One deletion request touches staging, marts, extracts and backups | Deletion registry keyed on subject identifier, replayed across every layer |
| Accuracy and correction | A corrected source record does not reach yesterday's mart | Change data capture with reprocessing, not one-way loads |
| Security safeguards | Analysts commonly get broad read access | Row and column level security in the database, plus field-level encryption for sensitive columns |
| Breach reporting | You must know what was exposed, not just that something was | Query-level audit logging with user, rows touched and timestamp |
Does an analytics application have to keep data in India?
The DPDP Act permits cross-border transfer except to territories the central government restricts by notification, so a blanket localisation rule is not the default position under this Act alone. Sector regulation is a different matter: if you are regulated by the RBI, payment and customer data rules are stricter than DPDP, and those rules, not this one, will set your hosting decision. The distinction is explained in data residency.
In practice most Indian enterprise buyers we work with choose an Indian region anyway, because it removes an argument from every procurement review and costs almost nothing on a managed warehouse. Decide it once, write it into the architecture document, and make sure the decision covers backups, logs and any analytics tooling with a hosted control plane, which is where residency commitments usually leak.
The four controls that do most of the work
Row and column level security in the database
Access rules belong where the data lives, not in the reporting tool's configuration. Postgres row level security policies, applied per role and per tenant, mean the same restriction holds whether the query arrives from your application, a scheduled export or an engineer with a psql session. Column level rules handle the narrower case where a role may count rows but not read the email address in them. Row-level security for analytics covers the implementation, and RLS defines the term.
Redaction and tokenisation at ingestion
The cheapest way to protect a field is not to copy it. Where analytics needs to join on an identity but never needs to display it, tokenise at ingestion and keep the mapping in a separate, tightly scoped store. Where free-text fields may contain personal data by accident, such as support notes, run redaction before the text lands in the warehouse. PII redaction sets out the approach.
A deletion registry rather than ad-hoc deletes
Erasure requests are the obligation analytics teams underestimate. Handle them with a registry: every request is recorded with a subject identifier and a timestamp, and every layer of the pipeline consumes the registry on a schedule, removing or tombstoning matching rows and recording that it did so. Ad-hoc deletion works once and fails the second time a mart is rebuilt from staging that was never cleaned.
Audit logging at query level
You need to answer, months later, who read which rows on which day. Application-level logs that record the user, the report, the filters and the row count are usually enough, and they are what an investigation or a customer security review will ask for. AI audit trails describes what regulators actually request, and the controls Eazyware runs internally are on our security page.
The checklist
- Inventory personal data fields by source, including free-text columns that may contain it accidentally
- Record a purpose and a lawful basis per field, not per system, and review them when a new mart is added
- Set a retention period per table with an automated expiry job and a log of its runs
- Build the deletion registry before the first mart, because retrofitting it means reprocessing history
- Apply row and column level security in the database, then verify it with a test that runs as each role
- Decide hosting region and write it down, covering backups, logs and any hosted control plane
- Turn on query-level audit logging with user, report, filters and row counts retained for an agreed period
- Name the Data Protection Officer or equivalent contact and the breach escalation path, with a rehearsal
- Restrict non-production environments, because masked test data is a control and copied production data is a breach waiting to happen
What compliance work adds to the build
Done during the build, these controls are part of the architecture rather than a phase. Our data and analytics application programmes run from $14,000 or ₹8,80,000 to $56,000 or ₹36,80,000, and a build with a full consent lineage, deletion registry and role-tested access model sits in the upper half of that band rather than outside it. Published bands are on the pricing page, and Indian clients are invoiced in INR with GST.
Retrofitting the same controls into a live warehouse costs more, because history has to be reprocessed and every existing extract audited. If you already have a warehouse and want to know where you stand, a ten-day Sprint Zero at $3,250 or ₹2,00,000, credited against the build, produces the data inventory, the gap list and a costed remediation plan.
After launch, compliance is a maintenance activity: retention jobs fail, new sources arrive, roles drift. A Care Plan from $1,000 or ₹68,000 a month on Essential up to $5,250 or ₹3,40,000 on Enterprise with a named engineer covers that work.
One planning note: the controls above are not independent. The deletion registry depends on a stable subject identifier, which depends on tokenisation at ingestion, which depends on the field inventory. Sequence them in that order and each piece is a day of work. Build them out of order and you will rebuild the pipeline twice.
Where this is over-engineering
An analytics application over machine telemetry, stock levels or financial aggregates with no personal data in scope does not need a consent lineage. Adding one because a checklist said so wastes weeks and teaches the team that compliance is theatre.
Equally, a ten-person company with one internal reporting tool does not need a formal Data Protection Officer function. It needs a named owner, a retention policy and access that is actually restricted. Scale the process to the data, and be honest when the risk is low.
The genuine mistake in the other direction is treating aggregation as an escape hatch. Small cohorts, rare attributes and joinable identifiers re-identify easily, and "it is only a summary table" is not a defence anyone has to accept.
Related reading
Data analytics application development in India covers costs and delivery models alongside the data rules, data analytics application development: a practical implementation guide shows where these controls land in the delivery phases, and tenant isolation explains what enterprise buyers ask about shared infrastructure.
Compliance in analytics is decided by what you copy and who can see it, so make both decisions at design time rather than at the security review.
Frequently asked questions
Does the DPDP Act apply to aggregated analytics data?
▾
It applies whenever the data can reasonably identify a person. Aggregation removes scope only if re-identification is genuinely infeasible, which small cohorts, rare attributes and joinable identifiers defeat. Treat derived tables built from personal data as personal data until a specific analysis shows otherwise, and document that analysis.
Do analytics applications have to store Indian data in India?
▾
The DPDP Act permits cross-border transfer except to territories the central government restricts by notification, so it does not impose blanket localisation by itself. Sector rules are stricter: RBI-regulated entities have their own requirements. Most Indian enterprise buyers choose an Indian region regardless, and the decision must cover backups and logs.
How do you handle a deletion request in a data warehouse?
▾
Use a deletion registry rather than ad-hoc deletes. Record each request against a subject identifier, then have every pipeline layer consume the registry on a schedule and remove or tombstone matching rows, logging that it did so. Otherwise the next rebuild from uncleaned staging restores the data you deleted.