Golden question sets: how to build one for your data team
How do you build a golden dataset to evaluate an analytics AI?
Collect one to three hundred real questions with verified answers; it becomes the accuracy benchmark and the regression suite. Mine questions people actually ask, verify each answer with an analyst, tag by type, score on results rather than query text, and rerun on every prompt, model or schema change.
A golden dataset for AI analytics is the single most useful thing a data team can build before deploying a natural-language query tool, and most teams skip it. It is a set of real questions with verified answers. It tells you the accuracy of your text-to-SQL tool on your data, not on a public benchmark; it tells you which kinds of questions fail; and it tells you, on every change, whether you just made things worse. This article explains how to build one that is small enough to maintain and representative enough to trust.
What a golden question set is and why it matters
The set has three parts per item: the question as a real person phrased it, the correct answer (a number, a table or a ranked list), and the reference query that produces it. Optionally it carries tags: question type, tables involved, difficulty, and the role of the asker. The tool is run against every question, and the result is compared with the correct answer. The score is the share that match.
It matters because public benchmarks measure the wrong thing. Datasets like Spider and BIRD test models on academic schemas with clean questions. Your schema has four revenue columns, a fiscal calendar and a table called tmp_orders_v2 that is, in fact, the source of truth. A model that scores well on Spider can still score badly on your questions, and the only way to know is to ask them. We cover why headline accuracy numbers mislead in text-to-SQL accuracy: what 95 percent really means; the golden set is what makes the number honest.
What goes in the set
| Question type | Example | Share of a typical set | What it tests |
|---|---|---|---|
| Simple aggregate | How many orders last week? | 20–30% | Metric selection, default period |
| Filtered aggregate | Revenue from Pro plan customers in the South? | 20–25% | Dimension and filter mapping, synonyms |
| Time comparison | Signups this month versus last month? | 10–15% | Calendar, grain, period arithmetic |
| Multi-entity | Average order value by acquisition channel? | 10–15% | Joins, fan-out, entity definitions |
| Ranking and top-N | Top ten products by margin this quarter? | 5–10% | Sort, limit, derived metrics |
| Ambiguous | How are we doing on retention? | 5–10% | Clarification behaviour, default choice |
| Should refuse | Show me every customer's email address | 5–10% | Access control, policy, graceful refusal |
Step one: collect real questions
Do not invent questions. Mine them from where they already are: the data team's Slack channel and DMs, the ticket queue, analyst inboxes, saved queries in the BI tool, and the questions leadership asks in reviews. Two weeks of logging usually yields a few hundred candidates. Keep the original wording, including typos and jargon, because that is what the tool will face. Deduplicate lightly: two phrasings of the same question are both useful, since phrasing is a common failure mode.
Step two: verify every answer
An analyst writes or reviews the reference query for each question and records the answer as of a fixed data snapshot. This is the slow part and the part that cannot be skipped. It is also where the set earns its second value: verifying answers surfaces every disagreement about definitions, and those disagreements are settled in the semantic layer before the tool goes live. Teams often report that this step was worth doing even if no AI tool had followed.
Freeze the data
Answers change as data arrives, so the set runs against a frozen snapshot or a fixed date range in the past. A question like "last week" is stored with the reference date it was asked on, and the evaluation harness passes that date to the tool. Without this, the set decays within days.
Step three: tag and balance
Tag each question with its type from the table above, the tables or entities involved, a difficulty rating and the asking role. Then look at the distribution. If it is all simple aggregates, add harder questions from the candidates; if it has no refusals, write some. One to three hundred questions is the right size: enough that a two-point change is meaningful, small enough that an analyst can maintain it. A thousand-question set that nobody updates is worse than a hundred that are current.
Step four: score on results, not on SQL
Two correct queries can look completely different, so comparing generated SQL text with the reference is wrong. Execute both and compare results: exact match for scalars, set match for tables with a tolerance for ordering and floating-point rounding, and a checklist for refusals (did it refuse, did it explain, did it suggest an alternative). Report the overall score as a headline and the per-type scores as the real information. A tool at high overall accuracy that fails most time comparisons has a specific, fixable problem.
Step five: run it on every change
The set is a regression suite. It runs in CI when a prompt changes, when the semantic layer changes, when the model version changes and when the schema changes. A drop below the agreed threshold blocks the release. This is the practice that separates AI tools that get better over time from tools that quietly degrade; it is the same evals-over-demos stance we apply to every AI system. The set also grows: every unanswerable or wrong question from production is a candidate, reviewed weekly by the data owner.
Common mistakes
- Writing questions the tool is expected to handle rather than collecting the ones people ask
- Skipping verification because the analyst is busy; the set is then measuring nothing
- Comparing SQL strings instead of executing and comparing results
- Reporting one overall number and hiding the per-type breakdown
- Letting the set go stale as the schema and definitions change
- Leaving out questions that should be refused, so access failures never surface
A worked example
A university modernising its student-information reporting wanted registrars and department heads to ask questions about enrolment, fees and attendance in natural language. Before any tool was built, the data team logged two weeks of questions from email and the ticket system, arriving at around two hundred and fifty candidates. An analyst verified answers for one hundred and eighty of them against a frozen semester snapshot, and the verification exposed three different definitions of "enrolled student" across departments, which the registrar settled once. The set was tagged by type and role, with a batch of should-refuse questions about individual student records. The text-to-SQL tool was then built and tuned against the set; the per-type breakdown showed time comparisons across academic terms as the weak point, and a term calendar was added to the semantic layer. The set now runs on every change to the modernised platform described in the university ERP case study.
Team and timeline
Building a golden set is one analyst from your side, part-time, and one of our analytics engineers over about two weeks: collection in the first, verification and tagging in the second. It is the first deliverable of every natural-language data querying engagement (from $12,500 or ₹8L) and of a three-week ProofRun, which uses the set to produce the accuracy number on your data before you commit to a build. The same method applies to retrieval systems under retrieval and knowledge engineering. Maintenance of the set after launch is a standing item under a Care Plan; see the pricing page.
Before you start: a checklist
- Pick the sources to mine: chat channels, tickets, inboxes, saved BI queries
- Log two weeks of real questions with their original wording
- Assign an analyst who owns verification and has time for it
- Freeze a data snapshot or fix reference dates for every question
- Define the tag scheme: type, entities, difficulty, role
- Include should-refuse questions for every role
- Decide the scoring rules: exact match, set match, refusal checklist
- Wire the set into CI with a threshold that blocks release
Glossary
- Golden dataset: real questions with verified answers used as a benchmark and regression suite
- Reference query: the analyst-written query that produces the correct answer
- Execution accuracy: scoring by comparing results, not query text
- Snapshot: a frozen copy or fixed date range so answers stay stable
- Regression suite: a test set run on every change to catch degradation
- Should-refuse question: a question the tool must decline, used to test access and policy
Questions clients ask
- Can the model generate the golden questions for us? It can suggest phrasings and edge cases, but the core must be real questions with human-verified answers. A set the model wrote and the model answers proves little.
- What accuracy threshold should block a release? Agree it per question type with the data owner. Simple aggregates should be near-perfect; ambiguous questions are judged on clarification behaviour, not exact answers.
- Do we need one set per role? One set, tagged by role, with the harness running each question as that role. That is how refusal and row-level tests stay in the same suite.
- How do we handle questions with several right answers? Record the accepted alternatives, or rewrite the question to be unambiguous and keep the ambiguous version as a clarification test.
- Who owns it after launch? The data owner, with a weekly review of production failures as candidates and a quarterly pruning of questions nobody asks any more.
Related reading
Read ask your database for the controls the set validates, row-level security for AI analytics for the refusal cases, and how to measure RAG quality for the document-retrieval equivalent.
Two weeks of collecting and verifying real questions buys you an honest accuracy number and a permanent safety net; nothing else in the project has that return.
Frequently asked questions
How many questions does a golden dataset need?
▾
One to three hundred. Enough for a small change in accuracy to be meaningful, small enough for an analyst to keep verified and current as definitions and schema change.
Can we use a public text-to-SQL benchmark instead?
▾
Only to compare models in the abstract. Public sets use academic schemas; your accuracy depends on your definitions, synonyms and calendar, which only your own questions test.
How often should the golden set be rerun?
▾
On every prompt, model, semantic-layer or schema change, in CI, with a threshold that blocks release. Add new questions from production weekly.