azyware
Technology

Natural-language reporting: turning questions into dashboards

EZ
Eazyware
· 7 min read
Quick answer

What should you know about natural language reporting inside a SaaS product?

Natural-language reporting maps a question onto a semantic layer, runs a validated read-only query and returns a chart with definitions. Without the semantic layer it is text-to-SQL guessing at what "active customer" means. This article covers the architecture, the guardrails, accuracy and the build.

Natural language reporting lets a user type "overdue invoices by region this quarter" and get a chart, a table and a plain statement of how each term was defined. It is the most requested copilot job in any product with a data model, and the one most often built badly, because teams wire a model straight to SQL and hope. The version that works maps the question onto a semantic layer first, runs a validated read-only query, and shows its definitions with the result. It is a standard job in the SaaS copilots we build, and this article explains how it fits together.

We cover why the semantic layer is the whole product, the query path and its guardrails, how to present results so they can be trusted, how accuracy is measured, and the team and budget the build needs.

What natural language reporting is and why the semantic layer matters

Every product has terms that sound simple and are not. "Active customer" might mean logged in this month, has a live subscription, or placed an order in ninety days. "Revenue" might be booked, recognised or invoiced. Two analysts in the same company will define them differently, and a model asked to write SQL will define them differently each time it is asked.

A semantic layer fixes the definitions once. It is a catalogue of metrics (revenue, churn, overdue balance), dimensions (region, plan, month), the joins between them and the filters each supports, expressed in a form both people and the model can read. The model's job shrinks from "write SQL for this question" to "pick the metric, dimensions and filters that match this question", which it does far more reliably. The semantic layer article covers the design; the dbt semantic layer documentation is a primary reference for one way of expressing it.

Three approaches compared

ApproachHow it worksAccuracy on real questionsSafetyFit
Raw text-to-SQLModel sees the schema and writes SQLUnpredictable; definitions vary per queryNeeds heavy sandboxingPrototypes and internal analysts who read SQL
Semantic layer plus generated queryModel selects metrics, dimensions, filters; query is generated from definitionsHigh on covered metrics; refuses or asks when not coveredRead-only by construction; row-level security appliedCustomer-facing reporting in a product
Pre-built report pickerModel maps question to one of N saved reportsPerfect within N; useless outsideTrivialNarrow products or a first step

Most products start with the third approach for a handful of reports, then move to the second as the semantic layer grows. The first is where teams begin when they have not read this article.

The query path, step by step

Understand the question

The model receives the question, the user's tenant and role, and the semantic catalogue. It returns a structured selection: metric, dimensions, filters, time range, and any term it could not map. Ambiguity is surfaced, not guessed: "By 'active' do you mean has a subscription or logged in this month?" is a better answer than a chart built on an assumption.

Generate and validate the query

The query is generated from the selection by code, not by the model, using the definitions in the catalogue. It is checked before it runs: read-only, against allowed tables only, with the tenant filter and row-level security applied automatically, within a row and time limit. The model never sees the database credentials and never writes SQL that runs unreviewed. Our ask your database and row-level security for AI analytics articles go into the guardrails.

Run, chart and explain

The query runs on a read replica or warehouse with a timeout. The result comes back as a table and a chart, with the chart type chosen from the shape of the data: a line for a metric over time, bars for a metric by category, a single number for a scalar. Beside it, the definitions used: "Overdue: invoice due date passed and balance greater than zero. Region: billing address. Quarter: calendar Q3 2026." The definitions are what make the result trustworthy, and they are what a user shares with a colleague.

Guardrails that are not optional

  • Read-only database role for the reporting path, enforced at the connection, not the prompt
  • Tenant and row-level filters injected by the query generator, never supplied by the model
  • Allow-list of tables and columns exposed through the semantic layer; everything else invisible
  • Row limits, time-range limits and query timeouts to protect the database from an expensive question
  • Refusal or clarification when a term is not in the catalogue, instead of a guess
  • Every query logged with the question, the selection, the generated SQL and the user

From one answer to a dashboard

A single answer is useful; a saved one is a report; a set of saved ones is a dashboard. Let the user pin an answer, schedule it, and share it with the definitions attached. Because each pinned answer is a structured selection rather than a text prompt, it re-runs deterministically and survives model changes. That is the difference between conversational BI in a product and a chat window that happens to draw charts.

Follow-up questions work the same way: "now split by plan" modifies the selection rather than starting again, and the chart updates. Narrative explanations of what the chart shows are a separate feature, covered in AI narrative insights.

Measuring accuracy honestly

Accuracy for natural-language reporting is measured on a golden set of real questions with verified answers, scored on whether the selection was correct, whether the query returned the expected result, and whether ambiguous questions were surfaced rather than guessed. Public benchmarks such as BIRD show what text-to-SQL achieves on academic schemas; your own set on your own schema is what matters, and building it is covered in golden question sets. Our text-to-SQL accuracy article explains why a headline percentage hides more than it shows.

A useful target is coverage: the share of real questions the catalogue can answer. Every question that falls outside it is logged, and the clusters tell the data team which metric to define next. Coverage grows weekly; accuracy on covered questions should be high from the start.

A worked example

A B2B SaaS company for field-service teams had a report builder that customers found hard to use and a support queue full of "can you send me a report of..." requests. They wanted customers to ask in plain language. A first attempt with raw text-to-SQL produced plausible charts with inconsistent definitions of "completed job" and "first-time fix", and was pulled after a customer questioned a number in a review meeting.

The rebuild started with a semantic catalogue of the twenty metrics and eight dimensions that covered most support requests, each with a definition the customer success team signed off. The model maps questions to selections; code generates read-only queries with the tenant filter; results show the chart, the table and the definitions. Questions outside the catalogue are logged and reviewed weekly, and the catalogue has grown since. The report-request queue shrank, and the copilot's reporting job became its most used. The in-app copilot case study describes the product.

Team and timeline

Natural-language reporting needs a data engineer to build the semantic catalogue, an AI engineer for the mapping, validation and eval suite, and a front-end engineer for charts, definitions and pinning. A first catalogue of twenty to thirty metrics with the full query path takes six to eight weeks. It is offered as natural language data querying from $12,500 / ₹8L, or as one job inside a SaaS copilot build from $19,500 / ₹12.8L. If your metrics are not yet defined anywhere, a ten-day Sprint Zero produces the first catalogue and the golden set, credited to the build. Growing the catalogue after launch sits in a Care Plan. Prices are on the pricing page.

Before you start: a checklist

  • List the twenty to thirty metrics and dimensions behind most report requests
  • Get a signed-off definition for each from the people who answer for the numbers
  • Set up a read-only role on a replica or warehouse with tenant and row-level filters
  • Decide the model's output: a structured selection, never raw SQL
  • Build a golden set of one hundred real questions with verified answers
  • Design the result view with chart, table and definitions together
  • Log every question outside the catalogue and review weekly
  • Plan pinning, scheduling and sharing so answers become dashboards

Glossary

  • Semantic layer: a catalogue of metrics, dimensions, joins and filters with agreed definitions
  • Selection: the structured output of the model: metric, dimensions, filters and time range
  • Row-level security: filters that restrict rows to what the user is entitled to see
  • Coverage: the share of real questions the catalogue can answer
  • Golden set: real questions with verified answers used to score accuracy
  • Pinned answer: a saved selection that re-runs deterministically as a report

The ten copilot jobs users actually ask for, conversational analytics in Slack and Teams and text-to-SQL for MongoDB extend the topic. The dbt semantic layer docs and the BIRD benchmark are primary references.

Define the terms once, let the model choose among them rather than invent them, and every chart it returns will carry the one thing a dashboard usually lacks: an explanation of what it means.

Frequently asked questions

How does natural language reporting differ from text-to-SQL?

▾

Text-to-SQL has the model write queries from the schema, so definitions vary per question. Natural-language reporting has the model pick metrics and dimensions from a semantic layer, and code generates a validated read-only query from agreed definitions.

Is it safe to let customers query data in plain language?

▾

Yes, when the query path is read-only at the connection, tenant and row-level filters are injected by code, only catalogued tables are visible, and limits and timeouts protect the database. The model never writes SQL that runs unreviewed.

How accurate is natural language reporting?

▾

High on questions the catalogue covers, when measured on a golden set of your own real questions. The practical metric is coverage, which grows as questions outside the catalogue are logged and new metrics defined.