Shipping AI features behind flags to a beta cohort
How should a SaaS company roll out an AI feature?
Ship to a chosen cohort, measure adoption and evals, iterate twice, then release generally with the metric that proves it. AI feature rollout works when the flag controls the model and prompt version as well as the UI, the cohort is chosen for feedback quality, and general release waits for a number, not a date.
AI feature rollout SaaS teams get wrong in a predictable way: the feature is built, demoed to leadership, and released to everyone with a banner. Two weeks later usage has collapsed, support has a queue of confused tickets, and nobody knows whether the problem is the model, the prompt, the UI or the idea. The alternative is unglamorous and reliable: put the feature behind a flag, ship it to a cohort chosen for feedback, measure adoption and evaluation results together, iterate at least twice, and release generally only when a specific metric says so. This article sets out how to do that.
Why AI features need a different rollout from ordinary features
A normal feature either works or it does not, and a bug report tells you which. An AI feature works most of the time, fails in ways users describe vaguely, and changes behaviour when the model provider ships an update you did not ask for. The failure modes are probabilistic, so you need volume to see them and an evaluation suite to distinguish a model regression from a UX problem. A beta cohort gives you the volume under control; the flag gives you the ability to change the model, prompt or retrieval configuration for that cohort without a deploy. The SaaS industry page describes the products we have taken through this.
What the flag should control
Most feature-flag AI setups only gate the UI: the button appears or it does not. That is necessary but not sufficient. The flag, or a small set of flags, should control every variable you might want to change during the beta.
| Flag or variant | What it controls | Why you need it during a beta |
|---|---|---|
| Feature visibility | Whether the entry point appears for this account and user | Cohort selection; instant kill switch |
| Prompt version | Which versioned prompt template the feature uses | Iterate on wording without a deploy; compare versions in evals |
| Model route | Which provider and model handle each task type | Swap when a provider regresses or a cheaper model proves good enough |
| Retrieval configuration | Index, chunking and top-k settings | Tune groundedness without touching code |
| Action permissions | Whether the feature can take actions or only draft | Start in draft mode, enable actions once evals pass |
| Usage allowance | Requests per account per day during beta | Contain cost; spot heavy users for interviews |
| Telemetry level | How much of the request is traced | Detailed traces for beta, reduced for general release |
Flags that control prompt and model versions should reference an artefact with its own version history and evaluation results, not free text in a config UI. We covered why in prompt versioning and evaluation: treat prompts like code.
Choosing the beta cohort
The cohort for a beta launch AI feature is chosen for feedback quality, not for size or for how friendly the account is. A good cohort has three properties: it contains the kind of data the feature will meet in general release, including messy data; it includes users who will complain in detail, because vague praise teaches nothing; and it is small enough that a bad week does not become a churn event. Ten to thirty accounts across two or three segments is a common shape, though the right number depends on volume per account.
- Include at least one account with old, inconsistent data, because that is where retrieval and reporting break
- Include an admin-heavy account if the feature takes actions, because admins find permission gaps
- Exclude accounts in a renewal window or an active escalation
- Name a contact per account who has agreed to a fifteen-minute call every fortnight
- Tell the cohort what the feature is for, what it is not yet good at, and how to report a bad answer in one click
Measuring adoption and evals together
Adoption numbers alone mislead. If usage falls in week two, you cannot tell whether the answers got worse or the novelty wore off. Evaluation numbers alone mislead the other way: a feature can pass every golden test and still be ignored because it appears in the wrong place. The rollout dashboard must show both for the cohort, week by week.
- Weekly active use per account and the tasks completed, not prompts typed
- Thumbs-down and edit rate: how often users reject or heavily rewrite the output
- Eval pass rate per prompt and model version on the golden set, re-run on every change
- Groundedness and citation accuracy for retrieval features, sampled from real cohort traffic and reviewed by a person
- Latency and cost per request per model route, so the general-release budget is known
- Support tickets tagged to the feature with the reason clustered weekly
Real cohort traffic also feeds the evaluation set. Every rejected or edited output is a candidate golden case, which is how the suite grows from the launch set into something that reflects actual use. See evals: the practice that separates AI demos from AI products for the mechanics.
Iterate twice before you decide
One iteration is not enough to know whether a feature can be fixed. The first round of cohort feedback usually reveals a placement or wording problem that is easy to change and produces a visible bump. The second round tells you whether the underlying task is one the feature can do reliably. A rule we apply: no general-release decision until the cohort has seen two versions, each with its own evaluation run and at least two weeks of use. If the second version does not move the metric, the feature is reworked or dropped, and that is a cheap outcome compared with a failed general launch.
Progressive AI rollout and the metric that proves it
General release should be gated on a number agreed before the beta starts. Examples: persistence of weekly use in the cohort from week one to week six; edit rate below an agreed threshold; eval pass rate stable across two model versions; cost per account inside the tier's budget. Choose one primary metric and two guard rails, and write them down with the product lead and the finance owner. When the beta hits them, ramp the flag by segment, plan or region rather than to everyone at once, and keep the kill switch and the model route flag live in production permanently. Standard feature-flag practice, described well in Martin Fowler's writing on feature toggles, applies directly; the AI-specific addition is that the toggle also selects the prompt and model.
A worked example
A field-service B2B SaaS built an in-app copilot that drafted job summaries and took scheduling actions. It went behind flags for visibility, prompt version, model route and action permission. The cohort was twenty accounts across two segments, including two with years of inconsistent job data and three admin-heavy accounts, each with a named contact.
The first fortnight showed strong initial use, a high edit rate on summaries and several permission gaps found by admins. The second version changed the summary prompt, added a preview step for actions and closed the permission gaps; evals were re-run and the edit rate fell. A third version enabled actions for the whole cohort. General release was gated on persistence of weekly use across six weeks and cost per account inside the tier budget, and the ramp went by plan over a month with the model-route flag left live. The support team reported that the general launch produced almost no tickets, because the cohort had already found the problems.
Team and timeline
A beta rollout runs alongside the feature build rather than after it. The flags, telemetry and evaluation harness are set up in the first fortnight of a six-week Launch 6 programme at $26,500–45,500 or from ₹17,60,000, with an AI engineer owning evals and prompt versions, a backend engineer owning flags and telemetry, and a product lead from your side owning the cohort and the release metric. A shorter ProofRun at $6,250–10,500 can take a feature from working prototype to cohort-ready in three weeks. Ongoing model and prompt changes after general release are covered by a Care Plan; the pricing page lists them. The SaaS copilot service includes this rollout method by default.
Before you start: a checklist
- Put visibility, prompt version, model route and action permission behind flags
- Build the golden evaluation set before the cohort sees the feature
- Choose ten to thirty accounts across segments, including messy-data and admin-heavy ones
- Name a contact per cohort account and agree a fortnightly call
- Add a one-click bad-answer report inside the feature
- Agree the general-release metric and two guard rails with product and finance
- Set a usage allowance per account for the beta
- Plan the ramp by segment and keep the kill switch live after release
Questions clients ask
- Can we use our existing feature-flag tool? Yes; the AI-specific flags are ordinary flags whose values point at versioned prompts and model routes.
- How long should the beta run? Long enough for two iterations with two weeks of use each, so six to eight weeks is typical.
- What if the cohort loves it in week one? Wait for week six. Launch enthusiasm is the least reliable signal in AI features.
- Should beta users pay? Usually not, but tell them the feature will be part of a paid tier, so feedback reflects real value.
- What if a model provider changes behaviour mid-beta? The eval run catches it; the model-route flag lets you switch without a deploy.
Related reading
See feature flags and beta cohorts for safe SaaS releases for the general practice, shadow mode: the right way to launch AI agents for features that act, and copilot adoption: why most AI features die in a month for what the metrics should look like.
Flag everything that can change, pick a cohort that will tell you the truth, iterate twice, and release on a number.
Frequently asked questions
What should a feature flag control for an AI feature?
▾
Visibility, the prompt version, the model route, retrieval configuration, whether actions are enabled and the usage allowance. Gating only the UI leaves you unable to fix the feature without a deploy.
How big should an AI beta cohort be?
▾
Ten to thirty accounts across two or three segments, chosen for feedback quality rather than size, including accounts with messy data and admin-heavy accounts, each with a named contact.
When is an AI feature ready for general release?
▾
When a metric agreed before the beta, such as persistence of weekly use across six weeks, is met after at least two iterations, with eval pass rate stable and cost per account inside budget.