Multi-tenant LLM architecture for SaaS
How should a SaaS product architect a multi-tenant LLM backend?
Multi-tenant LLM backends isolate prompts, retrieval and logs per tenant and meter usage so features can be billed. The tenant ID travels with every request from the API to the vector store to the trace, and nothing crosses it. Here is the architecture, the isolation points, the metering model and the cost to build.
A multi-tenant LLM backend is the part of a SaaS AI architecture that lets one deployment serve many customers without any customer's data, prompts or usage leaking into another's. The rule is simple to state and demanding to implement: a tenant ID is attached to every request at the edge and enforced at every layer it touches, including retrieval, prompt assembly, tool calls, caching, traces and billing. Get this right and an AI feature can be sold per seat or per usage with a clean audit trail; get it wrong and one customer's contract appears in another's answer. This article sets out the architecture we build under the LLM applications service for SaaS products, and it assumes the wider tenancy decisions covered in multi-tenant SaaS architecture have been made.
Why tenant isolation for AI is harder than for ordinary SaaS
Ordinary multi-tenant code isolates rows in a database. An LLM feature adds new places where data mixes. Retrieval pulls passages from a shared index; if the filter is wrong, another tenant's passage lands in the prompt. Prompt assembly composes instructions, tenant configuration, retrieved context and history; a bug in any of those carries one tenant's data into another's request. Caches keyed on the question alone return one tenant's answer to another. Traces collect everything. And the provider bill arrives as one number that must be attributed to the customers who caused it. Each is a point of isolation to design and test.
The isolation points, layer by layer
| Layer | What must be isolated | How | Test |
|---|---|---|---|
| API edge | Tenant identity | Resolve tenant from the authenticated session; never from a request body field | Forged tenant ID in body is ignored |
| Prompt assembly | System prompt, tenant configuration, few-shot examples | Load per-tenant overrides from a tenant-scoped store; validate against a schema | Tenant A's custom instructions never appear in tenant B's rendered prompt |
| Retrieval | Documents and chunks | Tenant ID in every chunk's metadata; filter before ranking; or a collection per tenant | Query as tenant A for a passage only tenant B holds returns nothing |
| Tool calls | Records fetched or changed | Tools receive the tenant context and enforce row-level rules | Tool called with tenant A context cannot read tenant B's order |
| Caching | Stored responses | Cache key includes tenant ID and prompt version | Identical question from two tenants never shares a cached answer |
| Traces and logs | Inputs, outputs, context | Tenant tag on every trace; access controls on the trace store | Support engineer scoped to tenant A cannot open tenant B's traces |
| Metering and billing | Token and cost attribution | Tokens and cost per step tagged with tenant and feature; aggregated per period | Sum of per-tenant cost equals the provider bill within rounding |
Tenant identity travels with the request
The tenant ID is resolved once, from the authenticated session or API key, and passed explicitly through every function that touches data or the model. It is never inferred from user input or read from a field the client can set. In practice this is a request context object that the retrieval client, prompt builder, tool executor, cache and tracer all require as an argument, so an omission is a test-time error rather than a production incident.
Retrieval: shared index with filters, or a collection per tenant
Two patterns work. A shared index with a tenant ID on every chunk and a mandatory filter applied before ranking is simpler to operate and suits products with many small tenants. A collection or namespace per tenant gives stronger isolation, simpler deletion when a customer leaves, and independent tuning, at the cost of more objects to manage; it suits products with fewer, larger or more regulated tenants. Both require the filter or collection choice to be enforced by the retrieval client from the request context, not passed in by the caller. The vector store's own support for tenancy is one of the criteria in pgvector vs Pinecone vs Qdrant; with Postgres, row-level security on the chunk table is a sound belt-and-braces measure, documented in the PostgreSQL row security policies.
Within a tenant, user-level permissions still apply: a tenant's finance documents should not be visible to every user of that tenant. The same filter mechanism carries both tenant and user entitlements, as described in permission-aware retrieval.
Per-tenant configuration without per-tenant prompts
Customers will ask for their own tone, their own terminology, their own guardrails. Resist maintaining a separate prompt per tenant; it does not scale and cannot be evaluated. Instead, keep one versioned prompt template per feature and a tenant configuration record with a small, schema-validated set of fields: product name, glossary terms, tone setting, allowed topics, escalation contact. The template renders these fields in fixed places. Every tenant runs the same prompt version, so an eval on the golden set covers everyone, and a template change ships to all tenants with one review, as set out in prompt versioning and evaluation.
Per-tenant metering: the basis for billing
Every model call records tokens in and out, the model used and its unit price at the time, tagged with tenant, feature, user and prompt version. Aggregating this per tenant per period gives cost of goods per customer, which is what pricing an AI feature requires: a flat fee needs a forecast of typical usage; a usage-based fee needs a meter the customer can see; a per-seat fee needs the distribution across users. The same data drives fair-use limits, quotas and alerts when a tenant's usage departs from its plan. The trace pipeline described in LLM observability is the source; the billing system reads an aggregated table, not raw traces. Pricing strategy is covered in how to price an AI feature in your SaaS product.
Rate limits, quotas and noisy neighbours
One tenant's bulk job must not slow every other tenant's chat. Apply per-tenant rate limits at the edge, queue batch work separately from interactive requests, and enforce plan allowances from the metering table with a soft warning and a hard cap. Batch work can also go to a cheaper model or a self-hosted endpoint without affecting the interactive path, as discussed in multi-model routing.
Data residency and private tenants
Some customers will require that their data never leaves a region or a VPC, or is never sent to a third-party model provider. Make the model endpoint a per-tenant routing decision from the start: most tenants go to the shared provider pool; a regulated tenant routes to a self-hosted open-weight model in their region, with the retrieval and trace stores following the same rule. This is the pattern behind the private agentic AI service.
Testing tenant isolation
Isolation is proven by negative tests run on every change: seed two tenants with distinct documents, configuration and users; run every feature as tenant A and assert that nothing of tenant B's appears in retrieved context, rendered prompts, tool results, cached responses or traces visible to A's support scope. Add a reconciliation test that per-tenant metered cost sums to the provider's reported usage. These tests are cheap and catch the one class of bug that ends a customer relationship.
A worked example
A field-service SaaS company added an in-app copilot that answered questions from each customer's job history and manuals. The first version used a shared vector index with the tenant filter applied in the application code after retrieval, a cache keyed on the question text, and a single provider bill nobody could split.
The review found no leaks in production, but the negative tests written during it found two paths where a missing filter argument would have retrieved across tenants, and the cache would have served one tenant's answer to another for an identical question. The rebuild made the tenant context a required argument for the retrieval client, moved the filter before ranking, added row-level security on the chunk table, put tenant and prompt version in the cache key, and tagged every trace and token count with tenant and feature. A metering table let the company launch a usage-based tier, and one regulated customer was routed to a self-hosted model in their region without changes to the feature. The build is the in-app copilot case study, and the pattern is standard in the SaaS copilots we ship.
Team and timeline
For a new AI feature in an existing multi-tenant SaaS, the tenancy work is typically an AI engineer and a backend engineer for two to three weeks within a six-to-twelve-week build: request context and retrieval isolation in the first week, configuration, caching and metering in the second, negative tests and residency routing in the third. Retrofitting to a feature that shipped without it takes longer because the filter and cache paths must be found and tested. The work is part of every LLM applications and SaaS copilot engagement, from $19,500 / ₹12.8L for a copilot, and the six-week Launch 6 program ships an MVP with tenancy and metering in place. Current figures are on the pricing page; the client owns the code, tests and metering pipeline.
Before you start: a checklist
- Resolve the tenant ID from the session, never from a client-supplied field
- Make tenant context a required argument for retrieval, prompt assembly, tools, cache and tracing
- Choose shared-index-with-filter or collection-per-tenant and enforce it in the client
- Keep one versioned prompt per feature with schema-validated tenant configuration
- Include tenant ID and prompt version in every cache key
- Tag tokens and cost per step with tenant and feature; reconcile against the provider bill
- Apply per-tenant rate limits and separate batch from interactive traffic
- Write negative isolation tests for two seeded tenants and run them on every change
Glossary
- Tenant: one customer organisation sharing a deployment with others.
- Request context: the object carrying tenant, user and permissions through every layer of a request.
- Row-level security: database policies that restrict which rows a query can see, enforced by the database itself.
- Metering: recording usage (tokens, calls, cost) per tenant and feature for billing and quotas.
- Noisy neighbour: one tenant's load degrading service for others.
Related reading
See multi-tenant SaaS architecture for the platform decisions this builds on, why AI copilots inside SaaS beat standalone chatbots for the product case, and what makes an LLM application production-ready for the wider engineering practices.
Attach the tenant to every request, enforce it at every layer, meter every token, and prove it with tests; then the AI feature is something you can sell rather than something you hope holds.
Frequently asked questions
Should each tenant have its own vector index?
▾
Not necessarily. A shared index with a tenant ID on every chunk and a mandatory filter before ranking suits many small tenants; a collection per tenant suits fewer, larger or regulated ones. Either way the retrieval client must enforce the choice from the request context.
How do you bill an AI feature per tenant?
▾
Tag every model call with tenant, feature, tokens and unit price, aggregate per period into a metering table, reconcile it against the provider bill, and let the billing system read the aggregate. That supports flat, usage-based or per-seat pricing with a meter the customer can see.
Can one tenant use a private model while others use a shared provider?
▾
Yes, if the model endpoint is a per-tenant routing decision from the first version, with the retrieval and trace stores following the same rule. We build this into LLM applications for SaaS where a regulated customer is expected.