azyware
Technology

Characterisation tests: the safety net for legacy code

EZ
Eazyware
· 7 min read
Quick answer

What are characterisation tests and why do they matter for legacy code?

Characterisation tests capture what the system actually does before you change it, so refactoring is checked against reality. You record real inputs and the outputs the legacy code produces today, including its quirks, and turn them into automated tests that fail the moment a change alters behaviour someone depends on.

Characterisation tests are tests written to describe what a piece of code does, not what it should do. You feed the existing system real inputs, record exactly what comes out, and lock that behaviour in as an automated check. They are the first thing we build in any modernisation, because legacy code testing is not about proving the old system correct; it is about making sure that whatever it does, the new version does the same until you decide otherwise. This article explains how to build them, where the technique is known as golden master testing, how they enable refactoring safely, and what they cost.

Why characterisation tests matter in a modernisation

Legacy systems have behaviour nobody wrote down. A rounding rule in the fee calculator, a date edge case in payroll, an order of operations in a report that finance has quietly relied on for a decade. None of it is in the specification, if there is one, and some of it is a bug that the business has built a process around. If you change the code without capturing that behaviour first, you find out what mattered when someone downstream complains, usually at month end.

Characterisation tests move that discovery to the start. They do not tell you whether the behaviour is right. They tell you the moment it changes. That is exactly the guarantee you need to refactor, upgrade a runtime, replace a module behind a façade or swap a database, because each of those is meant to change structure without changing outcomes. Michael Feathers introduced the term in his work on working effectively with legacy code, and the practice has not needed much revision since.

Types of test for legacy code

Test typeWhat it checksCost to buildRole in modernisation
Characterisation / golden masterOutput today equals recorded outputLow: generated from real runsThe safety net for every change
Unit testsA small piece behaves to a specificationHigh for legacy code: needs seamsWritten for new code, rarely for old
Integration testsModules work together across a boundaryMediumGuards the façade and data sync
Parallel-run comparisonNew implementation matches old on live trafficMedium: needs routing and diffingThe final gate before a switch-over
Approval testsA human approves a recorded output once; then it is lockedLowGood for reports, documents and UI snapshots

How to build a characterisation test suite

Find the seams

A seam is a place where you can observe or substitute behaviour without editing the code. Common seams in a legacy system are the HTTP boundary, the command line, a database table that a job writes to, a generated file, an email or a queue message. You want the seam furthest out that still isolates the behaviour you care about, because tests at the outer boundary survive internal refactoring. Tests that reach into internal functions break every time you tidy the code, which defeats the purpose.

Collect real inputs

Synthetic inputs miss the odd cases. Pull real ones: a month of requests from logs, a sample of records from production stratified by the fields that drive branching, the actual files that came in last quarter. Anonymise personal data before it goes into the test repository, and keep a record of the anonymisation rules so the sample stays reproducible. A few hundred well-chosen inputs beat ten thousand random ones. Add to the sample every time a production incident reveals a case it missed, so the suite grows with the system's known history rather than staying frozen at the first collection.

Record the golden master

Run each input through the current system and store the exact output: response body, file, database delta, rendered document. That stored output is the golden master. Normalise anything that legitimately varies, such as timestamps and generated identifiers, so the comparison does not fail on noise. The first run will show you things you did not expect; note them but do not fix them yet.

Automate the comparison

The test harness re-runs every input and diffs the result against the golden master. It runs in the pipeline on every commit, and a difference fails the build. The harness should print a readable diff, because a failing golden master test is only useful if an engineer can see in seconds whether the change is intended. Approval-test tooling handles most of this out of the box, and a simple harness is a few days' work for anything it does not cover.

Golden master testing and intended change

The suite will fail for two reasons: an accidental regression, or a deliberate improvement. The second needs a process. When a change is intended, the engineer reviews the diff with the business owner, the new output is approved and the golden master is updated in the same commit. That review is where the decade-old rounding bug finally gets a decision: keep it for compatibility or fix it now, with everyone who depends on it informed. Either way the choice is explicit and recorded.

Refactoring safely behind the net

With the suite green, refactoring becomes ordinary work. Extract a module, rename, split a class, upgrade the language runtime, move a job to a new host: after each step, the suite runs and either stays green or points at the exact input that changed. This is what makes a PHP 5 to PHP 8 migration tractable, and it is the gate for every slice in a strangler pattern migration. It also makes AI-assisted refactoring safe to use at scale: a model can propose large mechanical changes, and the characterisation suite decides whether they are accepted. Without the suite, AI-generated changes to legacy code are a liability; with it, they are a productivity gain.

What characterisation tests do not do

  • They do not prove correctness; a captured bug is preserved until someone decides otherwise
  • They do not cover inputs you never collected, so the sample has to be deliberate
  • They do not replace a parallel run on live traffic before a real switch-over
  • They do not test performance; add timing checks separately if latency matters
  • They do not survive being ignored; a suite that is red for a week is no longer a safety net

A worked example

A lending business ran its repayment schedule engine on code that predated everyone on the current team. The plan was to move it behind an API and then replace it. Before touching anything, the team pulled a stratified sample of loan records, ran the existing engine on each, and stored every generated schedule as the golden master. The first run surfaced a leap-year quirk and a rounding rule that collections had built a manual step around. Both were documented and deliberately preserved. Over the following months the engine was refactored, moved to a new runtime and finally replaced, with the suite green at every step and the two quirks fixed on an agreed date with collections in the room. That approach is the backbone of our legacy-to-AI modernization work and of the core lending modernisation pattern.

Team and timeline

A first characterisation suite for one business-critical module takes a QA engineer and a backend engineer two to four weeks: finding seams, collecting and anonymising inputs, recording the golden master and wiring the harness into the pipeline. Your side supplies a domain owner who can answer "is that difference a bug or a feature" within a day. The work is scoped inside a Sprint Zero when the modernisation is still being planned, or as the opening weeks of a ReCore programme, which starts at $31,500 / ₹22,40,000 with the full range on the pricing page. The suite stays with you, and keeping it green is part of any Care Plan.

Before you start: a checklist

  • The modules that will change first, and the outer seam for each
  • A stratified sample of real inputs, anonymised, with the rules written down
  • A normalisation list: timestamps, identifiers and anything else that varies legitimately
  • A harness that prints readable diffs and fails the pipeline
  • A named domain owner who adjudicates intended versus accidental differences
  • An agreed process for updating the golden master when a change is approved
  • A rule that the suite must be green before any refactor is merged
  • Storage for large golden outputs, such as generated documents, outside the code repository

Glossary

  • Characterisation test: a test that records current behaviour rather than specified behaviour
  • Golden master: the stored output that future runs are compared against
  • Seam: a boundary where behaviour can be observed or substituted without editing code
  • Approval test: a golden master test where a human approves the first output
  • Normalisation: removing legitimate variation from outputs before comparison
  • Parallel run: sending live traffic to old and new implementations and diffing the results

See application modernization vs rewrite for when modernisation is the right choice, zero-downtime cutovers for what happens after the suite is green, and evals: the practice that separates demos from products for the same idea applied to AI systems. Feathers' original treatment is summarised in Martin Fowler's notes on legacy seams.

Record what the system does today, lock it in, and every change after that is checked against reality instead of memory.

Frequently asked questions

What is the difference between characterisation tests and unit tests?

▾

Unit tests check code against a specification of what it should do. Characterisation tests check code against a recording of what it currently does, quirks included. For legacy code with no specification, characterisation tests are the practical starting point.

How many inputs does a golden master suite need?

▾

A few hundred deliberately chosen real inputs, stratified by the fields that drive branching, usually cover more behaviour than thousands of random ones. Add inputs whenever a production incident reveals a case the sample missed.

Can AI help write characterisation tests?

▾

Yes, for the harness, the normalisation rules and for proposing stratification of inputs. The recorded outputs must come from the real system, and a domain owner must adjudicate differences. See our modernization service for how we combine the two.