Zero-downtime cutovers: how to plan them
How do you plan a zero downtime migration and cutover for a production system?
Zero downtime needs feature switches, rehearsed rollbacks, business sign-off per slice and cutovers outside critical windows. Plan it as many small reversible switches rather than one event, keep old and new running side by side with data in step, and have a business owner decide go or no-go on written criteria.
A zero downtime migration is not a heroic weekend. It is a sequence of small switches, each of which can be reversed in minutes, taken outside the windows when the business cannot afford surprises, with a named person deciding go or no-go against criteria written in advance. The technical parts, feature switches, parallel running, data synchronisation and rehearsed rollback, are well understood. What separates the migrations that go quietly from the ones that make the news is cutover planning: who decides, on what evidence, and what happens in the first hour if the evidence turns bad. This article sets out the plan we use for production migration in modernisation work, and what it takes in team and time.
Why zero downtime is a planning outcome, not a technology
Downtime during a cutover comes from three sources: a change that cannot be reversed quickly, data that is out of step between old and new, and a decision that nobody was empowered to make at 3 a.m. Each is a planning failure. A switch that is configuration rather than deployment reverses in seconds. Data kept in sync by replication and reconciled nightly is never out of step by more than the last few minutes. And a written runbook with a named decision-maker turns the 3 a.m. question into a lookup. The tools matter less than the discipline; Google's Site Reliability Engineering book remains the primary reference for release and rollback practice.
Cutover strategies compared
| Strategy | How it works | Rollback | Best for |
|---|---|---|---|
| Big-bang | Stop the old system, migrate, start the new | Restore from backup; slow and lossy | Small systems with a genuine maintenance window |
| Blue-green | Two full environments; switch traffic at the load balancer | Switch back; data written to green must be reconciled | Stateless services and full-platform moves |
| Canary / staged | New version serves a small share of traffic, then more | Reduce the share to zero | Most application changes |
| Strangler slice | One capability routed to the new implementation behind a façade | Re-route the slice | Legacy modernisation, module by module |
| Parallel run | Old and new both process everything; only the old result is used until sign-off | Nothing to roll back until the switch | Financial, academic and regulated processes |
| Dual write with reconciliation | Both data stores written; nightly comparison | Stop writing to the new store | Database migrations behind live services |
The plan, step by step
Define the slices
Break the migration into units that can be switched independently: a module, a report, a job, a customer segment, a region. Each slice gets its own switch, its own success criteria and its own window. A migration with one slice is a big-bang with better branding; aim for many. The strangler pattern is the usual way to find the slices in a legacy system.
Build the switches
Every slice is controlled by a feature switch that lives in configuration and takes effect without a deployment: a routing rule at the façade, a flag read by the application, a percentage at the load balancer. The switch must support partial traffic, so a slice can go to 5% of users before 100%. And it must be observable: the dashboard shows which slice is on which implementation at any moment.
Keep the data in step
If old and new have separate data stores, replication or dual write keeps them aligned during the transition, and a reconciliation job compares them every night and reports every difference. Reconciliation is the control that lets you roll back a week after a switch without losing what happened in between. Where data cannot be kept in step, the slice is not ready to switch, whatever the calendar says.
Rehearse the rollback
Before the first real switch, the team switches a slice on and off in a staging environment that mirrors production, then does the same in production during a quiet period with a non-critical slice. The rehearsal times the reversal, checks that the dashboards show it, and confirms that data written during the on period survives the off. A migration rollback plan that has never been run is a hope, not a plan.
Write the go/no-go criteria
For each slice, before the window opens: the metrics that must stay within range, the comparison results that must be clean, the period they must hold for, and the person who decides. The person is from the business, not engineering, because the criteria are about business outcomes. Agreeing this in a calm meeting is what prevents the debate at midnight.
Choose the window
Cutovers happen outside critical periods: no month-end close, no fee deadline, no exam week, no sales peak, no payroll run. They happen when the team that can respond is awake and when the slice's first full cycle of use falls inside a period where someone is watching. A quiet Tuesday morning with the whole team present beats a heroic Saturday night with two engineers.
The cutover runbook
- Pre-checks: reconciliation clean, dashboards green, rollback rehearsed within the last month, decision-maker present
- Switch to a small share of traffic; watch error rates, latency and the slice's business metric for the agreed period
- Increase in steps; at each step the same checks, with the decision-maker confirming
- Full switch; the old implementation stays live and receiving synchronised data
- Observation period of one full cycle of the slice's work, with reconciliation each night
- Sign-off by the business owner; only then is the old path retired, and even then read-only first
- A written post-cutover note recording what was observed, so the next slice's plan improves
Where cutovers go wrong
The failures we see repeat. A switch that turns out to need a deployment, because a cache or a queue was not behind the flag. A schema change that cannot be reversed, made in the same step as the traffic switch rather than in a separate, earlier, backward-compatible step. Reconciliation that runs but nobody reads. A go/no-go held by an engineer who is also the one who wrote the code. And the window chosen by the project deadline rather than the business calendar. Each has a simple prevention that belongs in the plan from the start, and each is a reason the legacy-to-AI modernization programme treats cutover planning as a deliverable in its own right rather than a task at the end. For infrastructure moves, cloud migration for legacy apps covers the platform side.
A worked example
A last-mile logistics operator moved dispatch from a monolith to a new platform while hundreds of drivers used the app every day. The migration was cut into slices by function and then by depot. The first slice was shipment tracking, a read-only path, switched for one depot on a weekday morning with the operations lead watching the delivery-status dashboard. Reconciliation between the old and new stores ran nightly and showed clean results for a week before the second depot switched. Dispatch itself moved depot by depot, each with a rehearsed reversal and a sign-off from the operations lead. The old dispatch path stayed available in read-only mode for a quarter, then was retired. No driver saw a maintenance page. The dispatch platform case study tells the wider story, and the same discipline shaped the university ERP modernisation, where the windows were set by the academic calendar.
Team and timeline
Cutover planning for a modernisation is owned by the architect, with a DevOps engineer for switches, environments and dashboards, a QA engineer for reconciliation and comparison tooling, and a client-side business owner per slice who holds the go/no-go. The plan, runbooks and first rehearsal are produced in the opening weeks of a ReCore programme, which runs 8–16 weeks from $31,500 / ₹22,40,000, with the ranges on the pricing page. During and after the switches the system runs under a Care Plan; the Standard plan's 24×5 cover and four-hour response is the usual fit for cutover weeks, and the Enterprise plan's 24×7 cover is used where the business runs around the clock. Everything, including the runbooks and switch configuration, is documented and yours.
Before you start: a checklist
- The migration broken into independently switchable slices
- A feature switch per slice that works without a deployment and supports partial traffic
- Data synchronisation between old and new, with a nightly reconciliation report someone reads
- Schema changes made backward-compatible and released before the traffic switch, never with it
- A rollback rehearsed in production on a non-critical slice, timed and recorded
- Written go/no-go criteria and a named business decision-maker per slice
- Business calendar blackout windows in the plan, agreed by the offices they protect
- A retirement plan for the old path, with a read-only period before switch-off
Glossary
- Feature switch: a configuration value that routes traffic or enables behaviour without a deployment
- Canary: a small share of traffic sent to a new version to test it under real load
- Parallel run: both implementations process everything; only one result is used
- Reconciliation: an automated comparison of two data stores that reports every difference
- Expand-contract: changing a schema in two backward-compatible steps so either version of the code works
- Go/no-go: the written decision point, with criteria and an owner, before each switch
Related reading
See characterisation tests for the comparison that gates each slice, migrating PHP 5 to PHP 8 for a runtime cutover, and phased ERP delivery for going live module by module.
Many small switches, data kept in step, a rehearsed way back and a business owner holding the decision: that is what zero downtime is made of.
Frequently asked questions
Is zero downtime migration really possible for a legacy system?
▾
Yes, when the migration is cut into slices behind switches, old and new data are kept in sync and reconciled, and each switch has a rehearsed rollback. The remaining risk is in planning, not technology.
What should a migration rollback plan include?
▾
The switch that reverses the slice, the time it takes, what happens to data written while the slice was live, the dashboards that confirm the reversal, and a record of the last rehearsal. A rollback that has never been run in production is not a plan.
Who should decide go or no-go at a cutover?
▾
A named business owner for the slice, using criteria written before the window opened. Engineers report the evidence; the business decides. Our legacy-to-AI modernization service builds that into every phase.