Reporting on AI cost per account
How should a SaaS company report AI cost per account?
Track inference and retrieval cost per account against plan price so the AI tier stays inside its margin. AI cost per account shows whether a premium tier is a business or a subsidy: attribute every model call, embedding and index query to a tenant, price it at provider rates, and report it weekly beside plan revenue.
AI cost per account is the only number that tells a SaaS company whether its AI tier makes money. Aggregate model spend hides the truth: a handful of heavy accounts can consume most of the bill while paying a mid-tier price, and the average looks fine until the finance team asks why gross margin has moved. This article explains what to attribute, how to attribute it, what the report should show, and what to do when an account's cost exceeds what it pays.
Why cost per account, not total spend
Software gross margin has traditionally been high because the marginal cost of serving one more account was close to zero. AI features change that: every request costs money in tokens, embeddings and retrieval, and the cost scales with use rather than with seats. An account that runs reports all day costs more to serve than one that runs them weekly, and both may be on the same plan.
SaaS AI margins therefore depend on knowing the distribution, not the mean. Cost per account, tracked from the first day the feature is live, lets you set allowances, price overage, route heavy tasks to cheaper models and have a fact-based conversation with the accounts that cost the most. The SaaS industry page describes how we build this into products from the start.
What to attribute to an account
| Cost component | How it is measured | Attribution key | Notes |
|---|---|---|---|
| Model inference | Input and output tokens per request, priced at the provider's published rate for that model | Tenant and user on every request | Track cached-input tokens separately where the provider discounts them |
| Embeddings | Tokens embedded at index time and query time | Tenant of the document or the query | Index-time cost spikes on onboarding; amortise or show separately |
| Vector and search queries | Queries and storage per namespace or index | Tenant of the namespace | Managed vector databases price by storage and reads; map both |
| Tool and action calls | External API calls the copilot makes on the account's behalf | Tenant | Often small, occasionally the largest line for integration-heavy features |
| Voice or speech minutes | Minutes of speech-to-text and text-to-speech | Tenant | Only for voice features; priced per minute by providers |
| Tracing and storage | Trace volume and retention per tenant | Tenant | Usually allocated rather than measured; keep it visible |
| Private or reserved capacity | GPU hours or reserved throughput | Allocated by share of tokens | Fixed cost; allocate by use so heavy accounts carry it |
Attribution: the tenant on every request
AI usage cost tracking is only as good as the identifier carried through the request pipeline. Every model call, embedding job, retrieval query and tool call must carry the tenant identity and, where possible, the user, the feature and the prompt version. This is the same trace metadata you need for observability and for the tenant-isolation questions enterprise buyers ask, so it should be one design rather than three. We covered the observability layer in LLM observability: tracing every request from prompt to cost.
Pricing the usage is straightforward once tokens are counted: multiply by the provider's rate for the model used. Rates change, and models are swapped by routing, so store the rate card with a version and date rather than hard-coding it. OpenAI's pricing page and equivalents from other providers publish per-token rates; your own rate table should mirror them per model and date.
What the report should show
The cost report is a weekly view for product and finance, not a monthly accounting artefact. It should answer four questions: which accounts cost most, which accounts are outside margin, which features drive cost, and whether cost per unit of work is falling as routing and prompts improve.
- Cost per account this week and trailing four weeks, beside plan price and, if metered, overage billed
- Margin per account on the AI tier: plan price plus overage minus attributed AI cost, flagged when below the agreed floor
- Cost by feature and by model route, so the team knows which capability and which provider drive the bill
- Cost per completed task: per report generated, per draft sent, per bulk change applied, which is the unit economics number
- Distribution: the top ten accounts by cost, and the share of total cost they represent
- Index-time versus query-time cost, so onboarding spikes are not mistaken for a trend
The distribution view is the one that changes decisions. It is common for a small number of accounts to dominate cost, and each is a conversation: a higher plan, an overage arrangement, or a cheaper model route for their workload.
Unit economics of AI features
Cost per account tells you about margin; cost per task tells you whether the engineering is improving. A report that once cost a certain amount in tokens should cost less six months later, through prompt tightening, caching of repeated context, routing routine requests to smaller models and retrieval that returns fewer, better chunks. Track cost per task by feature and prompt version so each change can be seen. Model-agnostic routing is the biggest lever: when a cheaper model passes the evaluation suite for a task, the route flag moves and the cost per task drops without a product change. We covered the techniques in cutting inference costs by a third.
Keeping the tier inside its margin
Once the report exists, the controls follow. Set a fair-use allowance per plan, expressed in tasks rather than tokens so customers can understand it, and meter overage above it. Set a margin floor per account and alert when an account breaches it for two consecutive weeks. Give the account team a script for that conversation that is about value rather than tokens. And keep a hard cap per account per day as a safety measure against runaway automation or abuse, with a clear message to the user when it is reached. The billing mechanics are in usage metering and AI billing for SaaS.
Common mistakes
- Reporting total provider spend monthly and assuming the average account is typical
- Hard-coding token rates, so a model swap silently misprices cost
- Attributing by user only, when the plan and the renewal are per account
- Ignoring embedding and retrieval cost because inference is the visible line
- Pricing the tier before the first month of cost data exists
- Treating onboarding index cost as a recurring trend and over-reacting
A worked example
A field-service B2B SaaS launched an in-app copilot with drafting, bulk actions and reporting in a premium tier. Tenant, user, feature and prompt version were attached to every request from the first beta week, and a rate table with dated provider prices converted tokens to cost nightly. The weekly report showed cost per account beside plan price, cost by feature and model route, and cost per completed task.
Within the first months the distribution view showed a few accounts consuming a disproportionate share of cost through scheduled reports over very large datasets. The response was not a price rise: routine report generation was moved to a smaller model that passed the evaluation suite, repeated report context was cached, and those accounts were offered a higher allowance on the next plan. Cost per report fell and margin per account moved back inside the floor. Finance then had a number they trusted for the tier's contribution.
Team and timeline
Cost attribution and reporting is a two-to-three-week piece of work when built alongside an AI feature: a backend engineer for the request metadata, rate table and nightly aggregation, a data engineer or analyst for the report, and a finance owner from your side who sets the margin floor and allowance. It fits inside a Launch 6 build at $26,500–45,500 or from ₹17,60,000, or as a standalone data and analytics application from $14,000 or ₹8.8L when the feature already exists. Rate-card updates and route changes after launch are handled under a Care Plan; the pricing page lists them, and the SaaS copilot service includes attribution by default.
Before you start: a checklist
- Attach tenant, user, feature and prompt version to every model, embedding, retrieval and tool call
- Build a dated rate table per provider and model, and version it
- Decide how index-time cost is shown separately from query-time cost
- Agree the margin floor per account with finance
- Define the fair-use allowance per plan in tasks, not tokens
- Set a daily hard cap per account and the message users see
- Schedule the weekly report to product, finance and account management
- Plan the first routing review once four weeks of cost per task exist
Glossary
- Cost per account: all attributable AI cost for one tenant over a period
- Margin per account: plan price plus overage minus attributed AI cost
- Cost per task: attributed cost divided by completed units of work, such as reports generated
- Rate table: the dated list of per-token or per-minute prices by provider and model
- Allowance: the included usage per plan before overage applies
- Model route: the configured choice of provider and model per task type
Related reading
See LLM inference costs: how to forecast your monthly bill, AI features that justify a higher SaaS tier and how to price an AI feature in your SaaS product for the surrounding decisions.
Attribute every call to a tenant, price it from a dated rate table, report it weekly beside plan revenue, and the AI tier stays a business rather than a subsidy.
Frequently asked questions
How do you calculate AI cost per account in a SaaS product?
▾
Attach the tenant identity to every model, embedding, retrieval and tool call, count tokens and queries, price them from a dated rate table per provider and model, and aggregate nightly. Report the result weekly beside the account's plan price.
What margin should an AI tier target?
▾
That is a finance decision, but it should be set per account as a floor, with alerts when an account breaches it for two consecutive weeks, rather than judged from the aggregate provider bill.
How do we reduce AI cost per account without raising prices?
▾
Route routine tasks to smaller models that pass evals, cache repeated context, tighten prompts and retrieval, and set a fair-use allowance with overage. See cutting inference costs.