Questions to ask a text to SQL solution vendor before you sign
What should you ask a text to SQL solution vendor?
Ask a text to SQL solution vendor how they measure accuracy, who writes and owns the semantic layer, what happens when your schema changes, who pays for model and warehouse usage, and what you keep if you leave. Vague answers to those five are the reliable signal.
Ask a text to SQL solution vendor five things: how they measure accuracy on your data, who writes and owns the semantic layer, what happens when your schema changes, who pays for model and warehouse usage, and what you keep if the relationship ends. A vendor who has shipped one answers all five in specifics within a minute.
This is an interview script rather than a checklist. It is grouped into three rounds, with the answers that should reassure you and the ones that should end the meeting, plus the questions that sound rigorous but reveal nothing.
The opening question that sorts the room
Start here: "Show me the golden question set from your last deployment, with the questions you failed." A text to SQL solution is a system that converts a natural-language question into SQL, runs it under the asker's permissions and returns the result with the query visible. Every team that has built one has a list of questions it could not answer, because ambiguity in business language is the hard part, not SQL syntax.
A vendor who produces that list, redacted, and talks through why each question failed has run a real evaluation. A vendor who offers a live demo on a sample database instead has built a demo. The distinction matters more here than in most AI work, because a demo on a clean schema tells you almost nothing about your warehouse. Our position on this is set out in evals over demos.
Round one: how do they prove accuracy?
The compressed answer is that credible vendors measure execution accuracy against known-correct results on your own schema, and publish the number with its denominator. Anything else is theatre.
Press on the definition. "95 per cent accurate" is meaningless without knowing whether that is exact string match on generated SQL, execution result match, or a human judging plausibility, and without knowing how many of the questions were simple single-table lookups. Public research makes the point: the BIRD benchmark evaluates text-to-SQL on large, messy, real-world databases and reports execution accuracy well below what clean academic schemas suggest. Ask which of those conditions your evaluation will resemble.
Then ask who writes the golden set. If the vendor writes it alone, the questions will be the ones the system can answer. The right arrangement is that your analysts write the questions and supply the correct answers, and the vendor is measured against them. We unpack the arithmetic in what 95 per cent accuracy really means.
One more accuracy question is worth asking and rarely is: what does the system do when it is not confident? The useful behaviours are refusing, asking a clarifying question, or answering with the assumption stated on screen. A system that always produces a confident number is not more capable than one that sometimes declines; it is simply hiding the cases where it guessed, and those are the cases that end up in a board pack.
Round two: the questions and the answers that should worry you
| Ask this | A good answer sounds like | Warning sign |
|---|---|---|
| How do you handle an ambiguous question? | The system asks a clarifying question or returns the assumption it made, in the interface | It picks the most likely interpretation silently |
| Who maintains the semantic layer after launch? | Your team, with our handover and documentation; we train two people | We manage it for you as part of the subscription |
| What stops a generated query from scanning the whole warehouse? | Row limits, statement timeouts, a cost estimate before execution | The model is good at writing efficient SQL |
| How does the system know what this user may see? | Row-level security tied to the person's identity, tested per role | We filter results after the query runs |
| What happens when a column is renamed? | The eval suite fails, we are alerted, the definition is updated | The model adapts automatically |
| Can we see the SQL the model wrote? | Always, next to every answer, with the rows returned | It is available in the logs on request |
| Who pays for model and warehouse usage? | You do, on your own accounts, with budgets and dashboards we set up | It is bundled, so you do not have to think about it |
The bundled-usage answer is the one people most often mistake for generosity. Bundled usage means the vendor now has an incentive to make the system ask fewer and cheaper questions than your business needs, and you lose the ability to see what anything costs.
Round three: ownership, exit and the contract
The short answer is that you should own the semantic layer, the prompts, the eval suite and the generated code, and be able to run the system without the vendor within a week of a handover. That is our standard position, described in you own everything, and it is a reasonable thing to require of anyone.
Take these clauses into the commercial conversation rather than the technical one, because they are where a good technical partner and a bad contract most often meet.
- Ownership of artefacts: code, prompts, semantic layer definitions, eval sets and documentation transfer to you on payment, not on completion of an exit process.
- Model independence: no clause that ties you to one provider, and a written route to swap models with the eval suite as the gate.
- Accuracy gate for acceptance: a named percentage on a named question set, agreed before the build, not after the demo.
- Usage on your accounts: your OpenAI, Anthropic or cloud keys, your warehouse, your budget alerts.
- Data handling: where your schema and sample rows are processed, whether they leave the country, and what is retained by any model provider.
- Change process: what a new data domain costs, quoted as a rate card rather than negotiated under pressure later.
- Support terms: response versus resolution times, named engineer or pool, and what the plan covers when the model provider deprecates a version.
- Exit: a documented handover, a knowledge transfer session, and a fixed period of paid support afterwards.
For the full version of this in procurement language, what to put in a text to SQL solution RFP covers the document itself, and a security questionnaire for AI vendors covers the security annexe.
What should a credible price answer sound like?
It should be a range with the drivers named. Our natural language data querying programme runs from $12,500 or ₹8 lakh to $38,500 or ₹25.6 lakh, and the drivers are the number of data domains, the state of the warehouse and how much semantic modelling already exists. A vendor quoting a single number without asking about your schema has priced a template.
Ask what a discovery step costs and whether it is credited. A ten-day Sprint Zero at $3,250 or ₹2 lakh, credited against the build, is a fair structure because it lets both sides walk away cheaply. Post-launch, ask for the care plan tiers in writing: ours start at $1,000 or ₹68,000 a month with a $750 or ₹40,000 add-on for evals, regression and cost monitoring, all listed on the pricing page.
Questions that sound tough but tell you nothing
"Which model do you use?" invites a brand name and proves nothing; the useful version is "how do you decide, and how would we switch?". "How many of these have you built?" rewards volume over fit; ask instead which one most resembled your schema and what went wrong on it. "Can it handle complex queries?" will always get a yes; ask for a failed one. "Do you use RAG?" is a category error in this context, because the work is schema grounding and query validation rather than document retrieval, and a vendor who accepts the framing without correcting it has not thought hard about your problem.
"Is it secure?" is the weakest question in the set. Replace it with three specific ones: does the query run as the user or as a service account, is the connection read-only by construction, and what appears in the audit log when someone asks about salaries.
When you should not be hiring a vendor yet
If your warehouse does not have agreed definitions for your top twenty metrics, no vendor can help you, because the system will be asked to be certain about things your business is not certain about. Fix the definitions first, ideally in a semantic layer you own, and the vendor conversation gets shorter and cheaper. In practice this means a finance lead and a data lead sitting in a room and agreeing what revenue means, which is an organisational task no procurement process can outsource.
Equally, if fewer than twenty people will ever use it and their questions are stable, buy a BI tool with a natural-language add-on and spend the difference elsewhere. And if your data is known to be unreliable, a querying layer will simply distribute the unreliability faster. The build-or-buy trade-off is laid out in build or buy for text to SQL solutions.
Related reading
How to measure whether text to SQL solution is working gives you the metrics to write into the contract, and the semantic layer post explains the artefact you are insisting on owning. If you want the shortlist conversation to start with a scope rather than a sales call, the contact page is the fastest route.
The vendor worth signing is the one who shows you the questions their last system failed, and tells you what they changed afterwards.
Frequently asked questions
What is the single most useful question to ask a text to SQL vendor?
▾
Ask to see the golden question set from their last deployment, including the questions the system failed. Teams that have shipped a text to SQL solution keep that list, because ambiguous business language is the hard part. A vendor who offers a demo instead of a failure list has not run a real evaluation.
Should the vendor pay for model and warehouse usage?
▾
No. Usage should run on your own provider accounts with budgets and dashboards set up for you. Bundled usage hides the true running cost and gives the vendor an incentive to limit query volume. Eazyware sets budgets, routing and monitoring, but the accounts and the spending stay yours.
What should the contract say about ownership?
▾
Code, prompts, semantic layer definitions, evaluation sets and documentation should transfer to you on payment, with no clause tying you to one model provider. You should be able to run and change the system without the vendor after a documented handover and one knowledge transfer session.