azyware
Technology

Text to SQL Solution, security and the DPDP Act: a compliance checklist

EZ
Eazyware
· 7 min read
Quick answer

Is text to SQL solution compliant with the DPDP Act?

A text to SQL solution can be DPDP-compliant, but only by design. The Act cares about four things your architecture decides: where processing happens, which rows the asker may see, what the logs keep, and whether the purpose was stated. Get those right and the risk is no worse than your existing BI tool.

Yes, a text to SQL solution can meet India's Digital Personal Data Protection Act, but nothing in the technology guarantees it. Text to SQL solution security rests on four choices: where queries and results are processed, which rows the asker may see, what the logs retain, and whether the purpose was disclosed to the person whose data it is.

What follows is the checklist we work through before a natural language querying layer goes anywhere near a production database, written as obligations mapped to controls rather than as legal commentary. It assumes you already know roughly what the system does; if you do not, start with the plain definition of text-to-SQL and come back.

What the DPDP Act actually asks of a text to SQL solution

The DPDP Act 2023 regulates the processing of digital personal data about identifiable individuals in India. It applies to you as a Data Fiduciary, which means you decide the purpose and means of processing, and it does not care whether the query that surfaced a customer's mobile number was typed in SQL or asked in English. The obligations attach to the data, not to the interface.

Four duties drive almost all of the engineering work. You must process personal data only for the purpose the person was told about and consented to. You must apply reasonable security safeguards. You must not keep personal data longer than the purpose requires. And you must be able to answer a request from the person about what you hold and to erase it. The Ministry of Electronics and Information Technology, which administers the framework, publishes the Act and the rules made under it at meity.gov.in.

Nothing in that list bans natural language querying. What it does is make several habits expensive: a shared service account that reads every table, chat transcripts kept forever in a vendor's cloud, result sets exported to a spreadsheet nobody tracks, and a prompt log that quietly accumulates customer names. Our wider position on the Act is set out in DPDP Act 2023 and AI, and the wider engineering stance in that piece applies here without change.

Where personal data enters the pipeline

A text to SQL solution has six stages, and personal data can appear at five of them. Mapping the stages is the fastest way to find the gap your policy document has not noticed.

StagePersonal data exposureControl that satisfies the Act
Question captureThe user types a name, phone number or PAN into the questionInput redaction before the prompt leaves your network; block patterns you never need
Schema and semantic layerColumn names and sample values sent to the modelSend schema and synthetic examples only; never real rows as few-shot context
SQL generationModel provider processes the promptRegional endpoint or self-hosted model, zero-retention contract, no training on your data
Query executionThe SQL reads rows the asker is not entitled toRow-level security and a per-user database role, not a shared service account
Result renderingRaw personal data appears in a chat window and gets pasted elsewhereAggregate-only defaults, masking, export controls, watermarked downloads
Logging and cachingQuestions, SQL, results and cached answers persist indefinitelySeparate retention clocks: SQL forever, results short, personal data never

The last row is where most projects fail an audit. Teams remember to secure the database and forget that the observability stack is now a second copy of it.

The text to SQL solution compliance checklist

Work through these eight items before launch. Each one is a yes or no with an owner, not a discussion.

  • Identity flows end to end. The person's SSO identity reaches the database as a session context, and the query executes under their permissions. If your architecture cannot do this, you do not have a compliant system, you have a shared account with a language model in front of it.
  • Row-level security is enforced in the database. Filtering in application code is bypassed the first time someone calls the API directly. PostgreSQL enforces this in the engine through row security policies, which apply to every query regardless of how it arrived.
  • The semantic layer excludes sensitive columns by default. Columns are opted in, never opted out. A table exposed wholesale is a breach waiting for a curious analyst.
  • Generation is read-only. The database role has SELECT and nothing else, statement timeouts are set, and data-modifying keywords are rejected before execution, not after.
  • Purpose is recorded with the query. Each workspace is bound to a stated processing purpose, so an audit can show why a marketing analyst read a customer table.
  • Retention is split by artefact. Keep the question and generated SQL for audit; keep result sets for hours, not months; never persist personal data in prompt logs or caches.
  • Erasure reaches the derived stores. When a person exercises their right to erasure, the deletion must sweep caches, exports, embeddings of schema samples and the analytics warehouse, not just the source row.
  • Every answer is traceable. One immutable record per question: who asked, what SQL ran, which rows returned, how long it was kept. Our view on what that record should contain is in AI audit trails.

Does your data leave India? Text to SQL solution data residency

Under the DPDP framework, transfers outside India are permitted unless the government restricts a particular country, which is a lighter rule than GDPR. Sector regulators are stricter. If you are a bank, an NBFC or a payment system participant, RBI's data storage directions already require that the payment data stays in India, and your contract with an American model provider will not change that.

In practice, text to SQL solution data residency has a clean answer because the model never needs the data. It needs the schema, the semantic definitions and the question. If the question is redacted and the results never travel, a hosted model in another region is processing metadata, not personal data. When even that is unacceptable, an open-weight model on your own GPU inside your VPC keeps the whole path inside your boundary, at a higher running cost and a modest accuracy penalty on complex joins.

Write the choice down. Auditors accept a documented decision with a rationale far more readily than a system nobody can explain.

What does a DPDP-ready text to SQL solution cost?

Compliance is not a separate line item; it is a scope decision made at the start. Our natural language data querying programmes run from $12,500 or ₹8,00,000 for a governed layer over one warehouse and a defined question set, up to $38,500 or ₹25,60,000 for several data domains with per-tenant isolation, redaction and a full audit store. Retro-fitting identity propagation into a system built on a shared account usually costs more than building it correctly the first time.

If the data model is contested or the sensitive-column inventory does not exist, a ten-day Sprint Zero at $3,250 or ₹2,00,000, credited to the build, produces the semantic definitions, the access matrix and the retention policy before anyone writes code. After launch, a Care Plan from $1,000 or ₹68,000 a month covers patching and access reviews, and the AI add-on at $750 or ₹40,000 covers evals and prompt regression when models change. Current figures sit on the pricing page.

Where a text to SQL solution is the wrong choice

Three situations. First, when the underlying data model has no owner. If nobody can say authoritatively what "active customer" means, a language model will invent a definition and your governance problem becomes a correctness problem too. Build the semantic layer first.

Second, when the questions are regulatory filings or anything where a wrong number has legal consequence. Generated SQL is probabilistic; a statutory return is not. Use a fixed, reviewed report and keep the natural language layer for exploration. The honest accuracy picture is in text-to-SQL accuracy: what 95% really means.

Third, when the only people asking questions are three analysts who already write SQL fluently. You would be buying a governance obligation to save a skill they already have.

What a compliant deployment looks like in practice

The pattern we ship most often has four parts. A semantic layer defines the entities and metrics that may be queried, with sensitive columns excluded and joins pre-declared. The database enforces access through per-user roles and row-level policies, so the language model cannot widen anyone's permissions even with a perfect prompt. A redaction filter strips identifiers from the question before it reaches the model, and a policy engine decides whether a result set may be shown raw, masked or only in aggregate. Finally, an append-only log records the asker, the question, the SQL and the row count, with personal data excluded by construction.

Layered this way, the compliance argument becomes short: the model never sees personal data, the database never returns rows the asker is not entitled to, and every answer has a record. Teams running this pattern alongside a governed analytics product such as Eazy Insights AI tend to pass security review in one pass rather than three. The same permission discipline applied to documents is what we call permission-aware retrieval.

Row-level security for AI analytics covers the database mechanics in detail, GDPR is the stricter comparison for European operations, and text to SQL solution cost in 2026 breaks down the budget once the compliance scope is known. If you want the controls reviewed against your own schema, talk to us.

Treat text to SQL solution security as an access-control problem with a language model attached, and the DPDP questions answer themselves.

Frequently asked questions

Is a text to SQL solution allowed under the DPDP Act?

▾

Yes. The Act regulates the processing of personal data, not the interface used to query it. A text to SQL solution is permitted where processing serves a disclosed purpose, access is restricted to entitled users, security safeguards are reasonable, and retention is limited. The interface changes nothing about those duties.

Does a text to SQL solution send our customer data to OpenAI or Anthropic?

▾

It should not. A well-built system sends the schema, semantic definitions and the redacted question to the model, and executes the returned SQL inside your own database. Results never leave your boundary. If your vendor sends sample rows as prompt context, that is a design choice you can refuse.

How do we stop a text to SQL solution showing salaries to the wrong person?

▾

Propagate the user's identity to the database and enforce row-level security policies there, so the query runs with that person's permissions. Exclude sensitive columns from the semantic layer by default. Application-side filtering is not sufficient, because any direct API call bypasses it.

What logs should a compliant text to SQL solution keep?

▾

Keep the question, the generated SQL, the asker's identity, the timestamp and the row count in an append-only store for audit. Keep result sets only for the hours a session needs them. Never persist personal data in prompt logs, caches or traces, and make erasure requests sweep every derived store.