azyware
Technology

Feature flags and beta cohorts for safe SaaS releases

EZ
Eazyware
· 7 min read
Quick answer

How should a SaaS team use feature flags and beta cohorts for safe releases?

Flags let you ship to a beta cohort, measure and roll back without a release; every AI feature should launch behind one. The discipline is separating deploy from release, choosing cohorts deliberately, deciding the metric before the flag turns on, and deleting the flag when the rollout is done.

Feature flags in SaaS separate two events that most teams still treat as one: deploying code and releasing a feature. With a flag, code ships to production switched off, then turns on for a beta cohort, then for a percentage, then for everyone, with measurement at each step and a rollback that takes seconds and no deploy. Every AI feature should launch behind one, because AI features fail in ways that unit tests do not catch and that only real users on real data expose. This article covers how flags work, how to choose beta cohorts, what to measure, and the housekeeping that stops flags becoming their own source of risk.

What progressive delivery changes

Without flags, a release is a bet on the whole user base at once, and a bad release is a rollback deploy under pressure. With flags, a release is a sequence of small, reversible decisions: the code is already in production, and the question at each step is whether to widen the audience.

QuestionRelease by deployRelease by flag
How do we test with real users?Ship to everyone and watch supportEnable for a named cohort and watch their metrics
How fast can we roll back?Minutes to hours; needs a deploySeconds; flip the flag
Can we ship half-finished work?No; the branch waitsYes, dark; integrates continuously
Can we compare with and without?Only before versus afterConcurrent holdout, same period
Can sales promise a feature to one customer?Only by shipping to allEnable per tenant
What does a bad AI answer cost?Everyone sees itThe beta cohort sees it, and told you it might
Who can turn it on?Engineering, via deployProduct, via the flag console, within policy

Kinds of flags and how long they should live

Not all flags are the same, and mixing them is how flag systems become unmanageable. Release flags gate a new feature during rollout and should be deleted when it reaches everyone. Experiment flags split users for a controlled test and end with the test. Operational flags act as kill switches for risky dependencies, such as a model provider or a payment gateway, and live as long as the dependency. Entitlement flags reflect what a tenant has paid for and belong in the billing and permissions model, not in the flag system at all.

The rule that keeps the system healthy: every release and experiment flag has an owner and an expiry date at creation. A flag past its date is a bug. Teams that skip this end up with hundreds of flags nobody dares to remove, and combinations nobody has tested.

Choosing a beta cohort

A beta rollout in SaaS is only as useful as the cohort. The wrong cohort gives you either silence or a skewed picture. The right one is chosen on purpose, with a hypothesis about what it will reveal.

  • Internal users first: your own team on production data, to catch the obvious
  • Friendly tenants: customers who have agreed to give feedback and tolerate rough edges, ideally across sizes and use patterns
  • Representative by data shape: for AI features, tenants whose data covers the variety the model will see, not only the cleanest
  • Excluding the fragile: tenants in a renewal negotiation, an audit, or a peak season are not beta candidates
  • Opt-in where possible: a beta programme that users join sets expectations and produces better feedback than a silent enable

For AI features, add a stage before the beta cohort sees anything: shadow mode. The feature runs for the cohort, its outputs are logged and evaluated, and nobody sees them. When the evaluation says the outputs are good enough, the flag turns the interface on. Shadow mode is the same idea applied to agents that take actions, and the flag is the mechanism that makes it practical.

Decide the metric before the flag turns on

A beta without a metric is a demo with extra steps. Before enabling, write down what the feature is supposed to change and how you will know: task completion, time to first value, retention of the cohort, support tickets mentioning the feature, or for AI features, acceptance rate of suggestions, thumbs-down rate, escalation rate and cost per use. Then instrument the feature so the flag state is attached to every event, and compare the cohort to a concurrent holdout, not to last month.

The concurrent holdout matters more for AI features than for most, because their quality drifts with model updates and data changes. The comparison has to be same-period. Personalisation lift needs a controlled test makes the case for one category; the principle is general.

Rollback criteria written in advance

Alongside the success metric, write the condition under which the flag turns off without a meeting: an error rate, a latency threshold, a thumbs-down rate, a cost per use above forecast. Automate the check where you can. A rollback decided in advance is calm; one decided during an incident is not.

Flags and AI features: the specific cases

AI features benefit from flags at more than one layer. A feature-level flag controls whether users see the copilot. A model-level flag controls which model or prompt version serves a request, so a provider update can be tested on a cohort before it reaches everyone and reverted if the evaluation suite shows a regression. A tool-level flag controls which actions an agent may take, so autonomy can be widened intent by intent. Together they let a team practise the policy-gated actions approach without redeploying.

One caution: flags that change model or prompt must be recorded in the evaluation and audit trail. If a user reports a bad answer, you need to know which flag state produced it. Tie flag evaluations into your LLM observability so every trace carries the flag context.

Implementation choices

Flag evaluation should be fast, local and consistent: a user should not see a feature flicker between requests. Most teams use a managed or open-source flag service with an SDK that caches rules and evaluates them in-process, targeting by tenant, user, percentage and attribute. Keep flag definitions in an exportable format, make changes auditable, and give product managers a console with guard rails: they can widen a rollout within policy, but a full enable for a high-risk feature needs a second person.

Multi-tenant products need tenant-level targeting as a first-class concept. That interacts with the entitlement model in the enterprise SaaS features article: flags decide what is rolled out, entitlements decide what is paid for, and the two should not be confused.

A worked example

A field-service B2B SaaS company launched an in-app copilot that drafted job summaries and suggested next actions for dispatchers. The copilot was deployed dark, then run in shadow mode for three friendly tenants chosen for the variety of their job data, with drafts logged and scored against the evaluation set. When the scores were acceptable, the interface flag turned on for those tenants only, with a concurrent holdout among similar tenants. Acceptance rate of drafts and time to close a job were the agreed metrics; a thumbs-down rate threshold was the automatic rollback. One tenant's data exposed a failure mode with a job type the evaluation set had underrepresented; the flag stayed on for the other two while the prompt was fixed and re-evaluated. The rollout widened by percentage over the following weeks. The in-app copilot case study describes the outcome; the flag discipline is why the rough edges were seen by three tenants who expected them rather than by everyone.

Team and timeline

Flag infrastructure and a beta programme are part of the first release in a SaaS development engagement, from $31,500 or ₹20.8L, and every AI feature we build, whether a SaaS copilot from $19,500 or an agent, ships behind a flag with shadow mode and a defined metric as standard. Adding flags to an existing product is typically one to two weeks for an engineer, plus the product work of defining cohorts and metrics. A Launch 6 MVP, six fixed weeks from $26,500 or ₹17,60,000, includes the rollout plan as a deliverable. Figures are on the pricing page; to discuss a safe launch for an AI feature, contact us.

Before you start: a checklist

  • Separate deploy from release: code ships dark by default
  • Classify each flag as release, experiment, operational or entitlement, and keep entitlements out of the flag system
  • Give every release and experiment flag an owner and an expiry date
  • Choose beta cohorts deliberately, with internal users first and fragile tenants excluded
  • Write the success metric and the automatic rollback condition before enabling
  • Compare the cohort with a concurrent holdout, not with last month
  • Attach flag state to every event and trace, especially for AI features
  • Schedule flag deletion as part of the rollout, not as later cleanup

Glossary

  • Feature flag: a runtime switch that controls whether a code path is active for a given user, tenant or request
  • Progressive delivery: releasing to widening audiences with measurement at each step
  • Beta cohort: a deliberately chosen group of users who see a feature before general release
  • Holdout: a comparable group that does not get the feature, measured over the same period
  • Kill switch: an operational flag that disables a risky dependency or feature immediately
  • Shadow mode: running a feature for a cohort without showing its output, to evaluate before release
  • Flag debt: expired flags left in the codebase, creating untested combinations

See copilot adoption: why most AI features die in a month, prompt versioning and evaluation and our SaaS industry page. Martin Fowler's site hosts the standard reference article on feature toggles, including the categorisation of flag types used above.

Ship dark, enable for a cohort you chose on purpose, measure against a holdout, roll back by rule, and delete the flag when you are done.

Frequently asked questions

Why should every AI feature launch behind a feature flag?

▾

Because AI features fail on real data in ways tests do not catch. A flag allows shadow mode, a chosen beta cohort, a concurrent holdout, model and prompt switching, and a rollback in seconds without a deploy.

How do I choose a beta cohort for a SaaS feature?

▾

Start with internal users, then friendly tenants who agreed to give feedback, chosen to cover the variety of data the feature will see. Exclude tenants in renewals, audits or peak periods, and prefer opt-in programmes.

How do we stop feature flags piling up?

▾

Give every release and experiment flag an owner and an expiry date at creation, treat an expired flag as a bug, keep operational kill switches separate, and keep entitlements in the billing and permissions model rather than the flag system.