The strangler pattern: modernizing without stopping the business
What is the strangler pattern and how does it modernise a legacy system?
The strangler pattern wraps the legacy system with a façade and replaces modules one at a time behind it, each with rollback. Traffic for each module is switched to the new implementation only after it has run in parallel and matched the old one, so the business never stops and no single cutover can take it down.
The strangler pattern is the safest known way to replace a legacy system that the business cannot afford to stop. You put a façade in front of the old system, route every request through it, and then move one capability at a time to a new implementation behind the façade. Each move has its own test, its own switch and its own rollback. The old system shrinks until nothing calls it, and then it is switched off. The name comes from the strangler fig, which grows around a host tree until the host is gone. This article explains how the pattern works in practice, where it goes wrong, and what a strangler fig migration costs in team and time.
Why the strangler pattern beats a rewrite
A rewrite asks the business to wait for a finished replacement, then bet everything on a single cutover. The replacement always takes longer than planned because the old system's real behaviour is only discovered as the new one is built, and the cutover is the riskiest day of the year. The strangler pattern replaces that one bet with dozens of small ones. Each slice is delivered, tested against the old behaviour, switched on for a share of traffic, and either kept or rolled back. Value arrives from the first slice, and the risk of any single step is bounded. Martin Fowler's original description of the pattern remains the best short primary source.
| Approach | Risk profile | When value arrives | Best fit |
|---|---|---|---|
| Big-bang rewrite | One large cutover; discovery of legacy behaviour late | At the end, if it ships | Small systems with complete documentation |
| Strangler fig migration | Many small cutovers, each reversible | From the first slice | Live business systems that cannot stop |
| Lift and shift | Low, but nothing improves | Infrastructure savings only | When the code is fine and the hosting is not |
| Wrap and extend | Low; legacy stays as-is under an API | Immediately for new features | When replacement is not the goal yet |
How a strangler fig migration works
Step one: the façade
The façade is a routing layer that sits between callers and the legacy system. It can be an API gateway, a reverse proxy, a thin service or, for a monolith with a database everyone writes to, an interception layer around the data access. Its only job at first is to pass everything through unchanged while recording what passes. That recording is the map of the system: which endpoints exist, who calls them, how often, and with what data. Most teams find capabilities in that log that nobody knew were still in use.
Step two: choose the first slice
The first slice should be a capability that is valuable, well-bounded and low-risk. Read-only reporting, a lookup service or a notification pipeline are typical starts. It should not be the billing engine. The purpose of the first slice is to prove the mechanism, the parallel run, the switch and the rollback, on something where a mistake costs an afternoon rather than a quarter.
Step three: parallel run and comparison
The new implementation runs alongside the old one, receiving the same requests, with its outputs compared and its differences logged. This is the same shadow-mode discipline we apply to AI agents before they act. A slice is promoted only when the comparison shows it matching the legacy behaviour on real traffic, including the odd cases that no specification ever mentioned. Characterisation tests, built from the old system's actual outputs, are the gate; we describe them in characterisation tests for legacy code.
Step four: switch, watch, roll back if needed
The façade switches the slice's traffic to the new implementation, first for a share of users, then for all. Business metrics and error rates are watched for an agreed period. If anything moves the wrong way, the switch is flipped back and the slice returns to the parallel run. Rollback is a configuration change, not a deployment, and it is rehearsed before the first real switch. The watch period is agreed with the business in advance, typically one full cycle of whatever the slice does, so that a weekly job is observed for at least a week.
Step five: repeat, then retire
Slices are taken in dependency order, with the data model migrated behind the façade as capabilities move. When the last caller of the legacy system is gone, it is put into read-only mode for a period, then archived. The retirement is planned from the first day, because a legacy system that is 90% replaced and never switched off is the most expensive outcome of all.
Where the strangler pattern goes wrong
- The façade is built but slices never follow, because nobody owns the sequence and the budget stops after the gateway
- Slices are chosen by developer interest rather than by business value and dependency order
- The shared database is left untouched, so the new services are coupled to the old schema and cannot evolve
- Parallel runs are skipped under time pressure and the switch becomes a hope
- Rollback is assumed rather than rehearsed, so the first incident becomes a firefight
- Retirement is never scheduled, and the business pays for two systems for years
The data question
Code is easier to strangle than data. Most legacy monoliths have one database with every module reading and writing every table. The workable approach is to give each migrated slice ownership of its tables, with the façade or a synchronisation job keeping the legacy copy consistent until the last legacy reader is gone. That synchronisation is where most of the engineering care goes, and where dual-write bugs hide. Reconciliation reports that compare both stores nightly are not optional. The same discipline applies whether the destination is a modern relational store or a set of services, and it is the part of the work our legacy-to-AI modernization programme spends the most time planning.
Incremental modernization and AI
The façade has a second benefit beyond safe replacement: it exposes the legacy system as an API. Once requests and data flow through a layer you control, you can add document extraction, search, copilots and reporting on top without touching the legacy code. Clients often get more early value from that than from the replacement itself, and it changes the business case: the modernisation pays for itself through new capability while the risky replacement proceeds at its own pace. We cover the mechanics in adding an API layer to a legacy monolith.
A worked example
A university ran admissions, fees, timetabling and results on a single ageing system with no test suite and a vendor that had moved on. A rewrite had been quoted twice and declined twice. The strangler approach began with a façade that recorded every request for a term, which revealed which reports were still used and which were not. The first slice was the results-publication service, replaced during a quiet period and run in parallel through one exam cycle before switch-over. Fees followed, timed outside the collection window, then admissions with a full parallel run through an intake. The old system was retired after the last module moved. Staff never saw a downtime notice. The university ERP case study tells the fuller story.
Team and timeline
A strangler fig migration is led by an architect who owns the slice sequence and the data plan, with two to four engineers on slices, a QA engineer who owns the characterisation suite and comparison tooling, and a client-side product owner who signs off each switch. The first phase, façade plus first two or three slices, fits the ReCore programme at 8–16 weeks from $31,500 / ₹22,40,000, with the full range on the pricing page. Subsequent slices are planned in phases of the same shape, and the operating system runs under a Care Plan between phases. Clients own the façade, the tests and every slice from the first commit.
Before you start: a checklist
- A recorded map of every request into the legacy system over at least one business cycle
- A slice sequence ordered by business value and dependency, agreed with the product owner
- Characterisation tests for the first slice, generated from real legacy outputs
- A façade with per-slice switches that can be flipped without a deployment
- A data ownership plan and a nightly reconciliation report
- A rehearsed rollback for the first slice before its first real switch
- Business calendar blackout windows written into the plan
- A retirement date for the legacy system, even if provisional
Glossary
- Façade: the routing layer in front of the legacy system that directs each request to old or new
- Slice: one capability moved from legacy to the new implementation as a unit
- Parallel run: both implementations receive the same traffic; only the old one's output is used
- Characterisation test: a test that captures what the system actually does today
- Dual write: writing to old and new data stores during migration; the main source of subtle bugs
- Retirement: the planned switch-off of the legacy system after its last caller is gone
Related reading
Read application modernization vs rewrite for the decision itself, zero-downtime cutovers for the switch-over mechanics, and custom enterprise software for what typically replaces the last slices.
Wrap it, move one slice at a time with a rehearsed way back, and retire the old system on a date you chose rather than one it chose for you.
Frequently asked questions
How long does a strangler fig migration take?
▾
The façade and first slices usually fit an 8–16 week phase. The whole migration depends on how many capabilities the legacy system has and how tangled its data is; most run as a series of phases over one to two years, delivering value from the first.
Does the strangler pattern work with a shared legacy database?
▾
Yes, but the database is the hard part. Each slice takes ownership of its tables, a synchronisation job keeps the legacy copy consistent, and nightly reconciliation catches dual-write errors until the last legacy reader is retired.
What is the difference between the strangler pattern and a rewrite?
▾
A rewrite replaces everything in one cutover after a long build. The strangler pattern replaces one capability at a time behind a façade, each with its own test and rollback, so the business keeps running and value arrives early. See our modernization service.