AI narrative insights: from chart to explanation
How do AI data insights turn a chart into an explanation people can trust?
Narrative insights turn a chart into a short explanation of what moved and why, grounded in the same metric definitions. The reliable version computes the change and its drivers with deterministic queries and uses the model only to write the sentences; it never guesses a number, and every claim links to its query.
AI data insights are the feature every dashboard vendor now advertises: a paragraph beside the chart saying revenue fell, why, and what to look at. Done well, it is the most-read part of a report. Done badly, it is a model looking at a picture and making things up. The difference is architecture. This article explains how narrative analytics should be built so the numbers are computed, not imagined, what a good narrative contains, where it goes wrong, and what a build involves.
What narrative insights are and why they matter
A narrative insight is a short, generated explanation attached to a metric: "Net revenue was down versus last month. The decline is concentrated in the South region and the Starter plan; Enterprise grew. Refunds rose in the same period." It answers the question a reader asks of every chart, "so what?", without waiting for an analyst. It matters because most people do not read dashboards; they read sentences. A weekly post in a channel with three sentences of explanation gets read; a dashboard with twelve tiles gets opened once. The narrative also forces a discipline on the data team: if the explanation cannot be generated from governed metrics, the metric is not well enough defined.
Two ways to build it, one of which works
| Approach | How it works | Risk | Verdict |
|---|---|---|---|
| Model reads the chart or the raw rows | Send the image or a data table to a language model and ask for insights | Model invents numbers, misreads axes, picks spurious drivers, cannot cite | Demo only |
| Compute, then narrate | Deterministic queries compute change, contribution and anomalies from the semantic layer; the model writes prose from those facts | Prose can over-interpret; mitigated by templates and citations | Production |
Compute what moved
The first stage is arithmetic, not language. For each metric on the report, the system computes the change against the relevant comparison period (last week, last month, same period last year, the plan) using the same definitions the dashboard uses, from the same semantic layer. Then it computes contribution: which dimension values account for the change. A standard decomposition over the declared dimensions (region, plan, channel, product line) ranks the drivers by their share of the movement. Anomaly detection flags whether the change is unusual against the metric's own history. All of this is ordinary analytics engineering with well-understood methods, and every number in it is reproducible by a query.
Choose the comparison carefully
"Down 8% versus last week" and "up 3% versus the same week last year" can both be true, and the wrong one tells the wrong story. The semantic layer declares each metric's default comparison and seasonality, and the narrative states which one it used. Where both are informative, it says both.
Then write the sentence
The language model receives a structured fact set: metric, period, change, comparison basis, top three drivers with their contributions, anomaly flags, and any related metrics that moved together. It is asked to write two to four sentences in the house style, using only those facts, with each factual claim mapped to a fact id. The output is checked: every number in the prose must appear in the fact set, and any sentence that cannot be traced is dropped. Templates cover the common shapes ("X was [up/down] by Y versus Z, driven mainly by A and B"), and the model's job is to make the template read naturally, choose emphasis and connect related movements. That is a job models do well and a job with a small blast radius.
Correlation is not cause
"Refunds rose in the same period" is a fact. "Refunds caused the decline" is a guess, and the narrative should not make it. The prompt forbids causal language unless the semantic layer declares a known relationship, and the reviewer checks for it during shadow mode. Readers will draw their own conclusions; the narrative's job is to make the facts easy to find.
What a good narrative contains
- The headline change, with the comparison basis stated
- The two or three drivers that explain most of it, with their share
- Whether the change is unusual against history
- Related metrics that moved in the same period, without causal claims
- A link from every claim to the query that produced it
- A question the reader might ask next, which the query tool can answer
Where it goes wrong
The commonest failure is the picture-reading approach: a model looking at a chart image and describing what it thinks it sees. It will read a dip that is a data-loading delay as a collapse in sales. The second is drivers without a proper decomposition, where the narrative names whichever segment the model noticed rather than the one that mathematically explains the change. The third is stale definitions: a narrative computed from a different revenue definition than the dashboard beside it, so the two disagree in front of the CFO. All three are avoided by computing from the semantic layer and narrating from facts. Guidance on presenting data honestly, including the Nielsen Norman Group's work on dashboards and data visualisation, is a useful reference for the design side.
Where narratives live
Beside the chart in the dashboard, as a paragraph that updates with the data. In the weekly channel post from a Slack or Teams analytics bot, where "why?" in a thread triggers the same decomposition. In the PDF board pack, where the narrative replaces the analyst's Sunday evening. And inside a product, where customers see an explanation of their own usage or spend, which is a SaaS copilot feature that the same engine can serve.
A worked example
A hospital network's operations team received a weekly dashboard of appointment volumes, no-shows and call-centre metrics across sites, and nobody read it. The narrative feature we built computed each metric's change against the previous week and the same week last year, decomposed the movement across site, department and booking channel, and flagged anomalies against twelve weeks of history. The model wrote three sentences per metric from those facts, in the network's plain style, with every number traceable. Two weeks of shadow mode, in which the operations lead compared the narratives with her own notes, caught two problems: a comparison basis that ignored public holidays, fixed in the semantic layer, and a tendency to describe the multilingual voice agent's booking volume, which was growing as the voice agent rolled out, as causing the drop in call-centre calls, which was removed as a causal claim. The Monday morning narrative post is now the report; the dashboard is where people go when a sentence prompts a question.
Team and timeline
Narrative insights on top of an existing semantic layer are an analytics engineer and an AI engineer over three to four weeks: the decomposition and anomaly logic in the first two, the narration, checks and delivery surface in the last two, with shadow mode overlapping. Without a semantic layer, it is part of a four-to-eight-week natural-language data querying build from $12,500 or ₹8L, where the query tool and the narratives share definitions. If the narratives are a product feature for your customers, the data and analytics applications service covers the reporting layer. Programs and Care Plans are on the pricing page.
Before you start: a checklist
- Confirm the metrics have governed definitions the narrative can compute from
- Declare each metric's default comparison basis and seasonality
- List the dimensions a decomposition may use, in priority order
- Decide the house style: length, tone, what is never said
- Forbid causal language unless a relationship is declared
- Build the fact-to-sentence check so every number is traceable
- Choose the surfaces: dashboard, chat post, board pack, in-product
- Run shadow mode with the person who currently writes the commentary
Glossary
- Narrative insight: a generated explanation of what a metric did and which segments drove it
- Comparison basis: the period a change is measured against, such as prior month or same period last year
- Contribution analysis: decomposing a change across dimension values to rank drivers
- Anomaly flag: a marker that a change is unusual against the metric's history
- Fact set: the structured numbers the model is allowed to write from
- Traceability check: verifying every number in the prose appears in the fact set
Questions clients ask
- Can it write in our CFO's style? Yes. A few examples of past commentary set the tone, and the fact-set constraint keeps the content honest regardless of style.
- What if a metric moved for a reason not in the data? The narrative says what moved and stops. A human can add a note ("price change on the 12th"), and declared events can be included as facts in later releases.
- Does every metric need a narrative? No. Start with the five the leadership team asks about and add on request; a page of narratives is as unread as a page of tiles.
- Can users ask follow-up questions? In chat, yes: "why?" triggers the decomposition, and "show South by week" hands off to the query tool over the same definitions.
- How do we know it is right? Every number is traceable to a query, and shadow mode with the person who writes the commentary today is the acceptance test.
Related reading
See natural-language reporting for building reports from questions, golden question sets for evaluating the underlying answers, and how to handle hallucinations in production for the general principle of computing before narrating.
Compute the change, decompose the drivers, and let the model write only what the numbers already say; that is the whole recipe for AI data insights people will trust.
Frequently asked questions
Can an AI explain a dashboard accurately?
▾
Yes, if the explanation is computed first with deterministic queries over governed metrics and the model only writes the sentences. A model reading a chart image and guessing is not accurate enough for business use.
Will narrative insights say why a metric moved?
▾
They say which segments drove the change and which related metrics moved together. They should not claim causes unless a relationship is declared, because correlation in a report is not evidence.
How long does it take to add narrative analytics?
▾
Three to four weeks on top of an existing semantic layer, longer if definitions must be built first. It is usually delivered alongside a natural-language query tool sharing the same definitions.