What to put in a text to SQL solution RFP
What should a text to SQL solution RFP include?
A text to SQL solution RFP needs six things: a ranked inventory of real questions, an honest description of your warehouse, acceptance criteria with numeric thresholds, the permission and DPDP requirements, a commercial model that states IP ownership, and support terms. Feature lists make bids incomparable.
A text to SQL solution RFP should contain six things: a ranked inventory of real questions, an honest description of your warehouse, acceptance criteria with numeric thresholds, the permission and DPDP requirements, the commercial model with IP ownership stated, and the support terms. Feature lists make bids incomparable; question inventories make them comparable.
Having bid on a good many of these, the difference between an RFP that produces useful quotes and one that produces a spread of four to one is almost entirely in the first two sections. This article gives you the section-by-section content, the thresholds worth naming, what to leave out, and the situation where an RFP is the wrong instrument altogether.
Why most of these RFPs produce incomparable bids
The typical document asks for natural language querying, dashboard generation, multi-database support and enterprise security, then asks for a fixed price. Every bidder prices a different project, because nothing in the document constrains scope. The cheap bid assumed twenty questions on one schema; the expensive one assumed a governed layer across the whole warehouse. You then compare them as though they were the same thing.
The fix is not a longer document. It is a narrower one that specifies the work in terms a supplier can size: which questions, against which tables, for which users, judged against which threshold. Everything else is commentary.
Section one: a ranked question inventory
This is the most valuable page in the document and the one most RFPs omit. Export twelve months of ad hoc data requests from your ticket queue or Slack channel, deduplicate them into distinct question shapes, and rank the top forty to eighty by frequency. Include the exact wording people used, not a tidied version.
Attach the twenty hardest as a separate annexe and ask each bidder to state which they expect to answer in the first release and which they would exclude. A supplier who marks five as out of scope and explains why is telling you more than one who claims all eighty. This inventory is also the seed of your golden question set, so the work is not wasted whichever way the tender goes.
Add one more column to the inventory: who asks. A question asked weekly by eleven people in operations is worth more than one asked monthly by the finance director, and bidders cannot infer that from the text of the question. The distribution of askers is also what tells you whether the project clears its own running cost, since a system used by a dozen people rarely does.
Section two: your data estate as it actually is
Bidders price uncertainty. Reduce it by stating the warehouse platform and version, the number of tables and subject areas in scope, whether a modelling layer already exists, how documented the schema is, and how often it changes. Say plainly if column naming is inconsistent or if there are three revenue columns with no agreed definition. Concealing that does not make it cheaper; it makes the quote wrong and the change request inevitable.
State which metrics already have a governed definition and which do not. A semantic layer that exists shortens delivery by two to four weeks, and one that does not is the single largest variable in any bid you receive. The semantic layer: why text-to-SQL needs one is worth circulating internally before you write this section.
Section three: acceptance criteria with numbers in them
Acceptance criteria are what turn a proposal into a contract. Name the metric, the measurement method and the threshold. Vague criteria such as high accuracy are unenforceable and invite optimistic bids.
| Criterion | How it is measured | Threshold worth naming |
|---|---|---|
| Execution accuracy | Golden question set of 100+ questions, results compared to analyst-verified answers | 85% or better on the agreed set |
| Refusal behaviour | Out-of-scope questions in the test set | Declines rather than guesses, with a stated reason |
| Permission enforcement | Question set run as each user role | Zero rows returned outside the role's entitlement |
| Query cost control | Scanned bytes and statement duration per query | Hard ceiling enforced in the database, not the application |
| Latency | Median and 95th percentile, measured end to end | Under 10 seconds median for the agreed question set |
| Transparency | Manual inspection of the interface | Generated SQL visible to the user on every answer |
| Regression safety | Golden set run in CI on schema and model changes | Build fails on any regression beyond an agreed tolerance |
Set the accuracy threshold against a question set you own, not a public benchmark. Ask each bidder how they would measure it and be sceptical of anyone quoting a headline percentage without naming the data it was measured on. Text-to-SQL accuracy: what 95 percent really means explains the ways that number gets inflated.
Section four: permissions, residency and DPDP
State the permission model explicitly: which user groups exist, what each may see, and whether restrictions are row-level, column-level or both. Require enforcement in the database rather than in prompt text, because a control that a well-phrased question can argue away is not a control. Row-level security for AI analytics sets out what a compliant implementation looks like.
If the warehouse holds personal data, the Digital Personal Data Protection Act applies to processing carried out in India and to processing of Indian data principals' data abroad, and the text of the Act is published by MeitY. Name your obligations in the RFP rather than discovering them at security review: purpose limitation, retention of query logs, breach notification, and whether any data may leave your cloud region. Also state whether model inference may use a hosted API or must run inside your perimeter, because that single line moves both price and timeline.
Section five: commercial model and ownership
Ask for a fixed price against the named scope with change control priced by the day, not a time and materials estimate. Require the quote to separate build, semantic layer work, evaluation suite construction and post-launch support, so you can compare like with like across bidders.
State ownership plainly: you should own the code, the prompts, the semantic model, the evaluation suite and the deployment configuration. At Eazyware the client owns all of it, along with model choices and documentation, and any supplier unwilling to write that into the contract is selling you a dependency. Ask also who pays for model API usage; ours is billed through the client's own accounts, with budgets and dashboards set so it stays predictable.
Section six: support after launch
A text to SQL system has two dependencies you do not control: your schema and the model provider's versions. Both change. Require the bid to price ongoing evaluation runs, re-indexing and prompt regression, not just incident response. Eazyware's care plans run from $1,000 or ₹68,000 a month for business-hours cover with a ten-hour allowance, to $5,250 or ₹3,40,000 a month for 24x7 with a one-hour response and a named engineer, with an AI add-on at $750 or ₹40,000 covering evaluations, cost monitoring and prompt regression.
What to leave out
- Model names. Specifying a provider dates the document and removes the supplier's ability to route between models by task.
- Vector database requirements. Text to SQL is a schema and semantics problem; a vector store may or may not be part of the answer.
- Dashboard feature parity. If you want your existing dashboards rebuilt, that is a different procurement.
- Unbounded multi-database support. Name the databases in scope; every additional source is real integration work.
- Conversation length or memory requirements. These are design decisions, not procurement criteria.
- Accuracy figures copied from vendor marketing. Set your own threshold against your own question set.
When an RFP is the wrong instrument
If you cannot yet answer section one or section two, an RFP will waste a quarter and produce bids you cannot compare. Run a short paid discovery instead. A ten-day Sprint Zero at $3,250 or ₹2,00,000, credited against a subsequent build, produces the question inventory, the metric gap list and the permission model, which is precisely the RFP content you are missing. A three-week ProofRun at $6,250 or ₹4,00,000 goes further and returns a measured accuracy figure on your own schema.
An RFP is also wrong when the real demand is a small, stable set of questions. That is a dashboard, and a tender for a natural language system will get you a more expensive answer to a question you had already solved. We would tell you that rather than bid.
One more case: a tender written to satisfy a procurement policy when the decision has already been made. Everyone recognises it, the serious bidders decline, and you end up choosing between three suppliers who had nothing better to do that fortnight. If you already know who you want, run a paid discovery with them and put the output through procurement as a scoped statement of work.
Budget guidance worth including
Publishing a range gets you better bids, because suppliers stop guessing at your ambition. A scoped natural language data querying build starts at $12,500 or ₹8 lakh and runs to $38,500 or ₹25,60,000 depending on schema breadth, governed metric count and integration surface, over eight to twelve weeks. Our published figures are on the pricing page, and the 2026 cost breakdown explains what drives movement inside that range.
Related reading
Questions to ask a text to SQL solution vendor before you sign covers the shortlist conversation that follows the RFP, five ways text to SQL solution projects fail is a useful risk annexe, and how long does text to SQL solution take helps you set a delivery window bidders can meet.
Specify the questions and the thresholds, and the bids will compare themselves.
Frequently asked questions
What should a text to SQL solution RFP include?
▾
A ranked inventory of real questions taken from your request log, an honest description of the warehouse and its metric definitions, acceptance criteria with numeric thresholds, the permission and DPDP requirements, a commercial model stating IP ownership and who pays for model usage, and priced post-launch support.
What acceptance criteria should I set for text to SQL?
▾
Execution accuracy of 85 per cent or better on a golden question set you own, correct refusal on out-of-scope questions, zero rows returned outside a user role's entitlement, enforced query cost ceilings, median latency under ten seconds, visible generated SQL, and a CI run that fails on regression.
Should the RFP specify which AI model to use?
▾
No. Naming a model dates the document and removes the supplier's ability to route simple questions to a cheaper model and hard ones to a stronger one. Specify the outcome instead: accuracy threshold, latency, whether inference may leave your perimeter, and who is billed for usage.