How to measure whether legacy application modernization services is working
How do you measure whether a legacy modernisation programme is working?
Measure a modernisation on four layers: business outcomes it was funded for, operational load on the people using it, system health against stated service levels, and output quality for any AI in the loop. Capture every baseline before the first module moves, because you cannot reconstruct it afterwards.
Measure legacy application modernization services on four layers: the business outcomes the programme was funded to move, the operational load on the people using the system, system health against stated service levels, and output quality for any AI in the loop. Capture every baseline before the first module moves, because once the old screens are gone you cannot reconstruct it.
Most programmes are measured on delivery instead: modules shipped, tickets closed, percentage complete. Those numbers can all be green while the business case quietly fails. This article sets out what to instrument, when to read it, which numbers mislead, and how to gate releases on evidence rather than on a demo.
Capture the baseline before you change anything
The single most common measurement failure on a modernisation programme is that nobody wrote down what the old system did. Six months later the new one is faster, cheaper or more accurate according to everyone's recollection, and the finance director quite reasonably asks for the comparison. There is none.
Baseline capture takes about a week and belongs in discovery. You need the volumes that run through each workflow, the time the workflow takes end to end, the error and rework rate, the cost of running the incumbent including licences and hosting, the support ticket profile by category, and the service levels the old system actually delivered rather than the ones in its documentation. Take these from the systems, not from interviews, and take them for a period long enough to cover a normal cycle.
One discipline is worth borrowing from site reliability engineering: define the indicators and the targets before you build, in the language Google's service level objectives chapter uses. An indicator you agree on in advance is evidence. One you choose after seeing the data is an argument.
The four layers, and what belongs in each
| Layer | What to measure | Read it | Who owns it |
|---|---|---|---|
| Business outcome | Cycle time, cost per transaction, revenue leakage, compliance exceptions | Monthly, against baseline | The programme sponsor |
| Operational load | Task completion time, rework rate, support tickets by category, adoption per role | Weekly during rollout | Head of the affected function |
| System health | Availability, p95 latency, error rate, batch completion, incident count | Continuous, reviewed weekly | Engineering lead |
| AI output quality | Groundedness, extraction accuracy, escalation rate, cost per task | Every release, plus every model change | AI engineer and a business reviewer |
| Migration integrity | Records reconciled, exceptions open, ageing of exceptions | Daily through cutover | Finance or operations owner |
The layers are ordered deliberately. A programme that is green on system health and red on business outcome has built the wrong thing well. A programme green on business outcome and red on operational load has moved cost onto the people using it, which shows up as attrition rather than as a metric, and always eventually shows up.
Two practical notes on the table. Adoption per role is worth breaking out rather than reporting as a single percentage, because an average hides the one department that has quietly reverted to the old process. And cost per transaction should include the AI usage line if there is a model in the loop, since that cost scales with the adoption you are also trying to increase, and the two numbers only make sense read together.
The metrics that mislead
Several numbers look like progress and are not. Watch for these in status reports and ask what sits underneath them.
- Percentage complete. It measures the plan, not the system. A programme can be ninety per cent complete for a quarter.
- Lines of legacy code retired. Retiring the easy half of a monolith is not half the work, and the difficult modules are the ones holding the business.
- Story points delivered. Velocity measures a team's estimation habits, not the value shipped to a user.
- Aggregate uptime. Ninety-nine point nine per cent availability means nothing if the outage falls on a payroll run or a filing deadline. Measure availability during the windows that matter.
- Support ticket volume falling. It can fall because users gave up and went back to spreadsheets. Pair it with adoption per role.
- Demo quality for AI features. A model that answers ten questions well in a meeting tells you nothing about two hundred real ones.
- Average response time. Averages hide the tail. Track p95 and p99, because your loudest users live there.
Gating releases on evidence
Measurement earns its keep when it can stop a release. Three mechanisms do that work on a modernisation programme, and each answers a different question.
Characterisation tests for parity
A characterisation test records what the legacy system actually does, including its quirks and its bugs, and fails when the new module behaves differently. It answers the question "did we change behaviour we did not mean to change?" This is the safety net that makes incremental replacement survivable, and it is covered in characterisation tests for legacy code.
Shadow running for confidence
Before a module takes live traffic alone, it processes the same traffic as the incumbent and the two outputs are compared. Disagreements are triaged rather than assumed to be bugs, because sometimes the new system is right. A module moves to live when the disagreement rate is low and every remaining disagreement is understood. The rollout discipline is described in shadow mode.
Evaluation suites for AI modules
Any AI component in the modernised system needs a fixed set of cases with known correct answers, scored automatically, with a pass threshold agreed in advance by someone from the business. Run it on every release and on every model change, because providers deprecate and replace models on their own schedule and your prompt is not portable by default. Evals: the practice that separates AI demos from AI products sets out how the suite is built, and LLM observability covers tracing cost and latency per request once it is live.
All three mechanisms share one property that makes them useful: they produce a pass or fail that a person agreed to in advance. A metric nobody committed to is a talking point, and talking points do not stop a bad release on a Friday afternoon. Write the thresholds into the delivery plan, name who can waive them, and record every waiver, because a pattern of waivers is itself a signal worth reading.
How long before the numbers move?
Expect system health metrics to improve within days of a module going live, operational load to get worse before it gets better, and business outcomes to lag by a quarter. The productivity dip is real: people who have used the same screen for a decade are slower on a new one for six to ten weeks. A programme judged at week four will be judged a failure. Set the review point at the end of the first full business cycle after the module is live, and say so in writing before you start.
Budget for the measurement itself. Instrumentation, dashboards and the evaluation suite are work, typically a small share of a Legacy-to-AI Modernization Program that runs from $31,500 or ₹22,40,000 to $105,000 or ₹72,00,000 and above. Keeping the suite current afterwards sits in the Care Plan, from $1,000 or ₹68,000 a month at Essential to $5,250 or ₹3,40,000 at Enterprise with 24x7 cover and a named engineer, with the $750 or ₹40,000 AI add-on covering evals, cost monitoring, prompt regression and re-indexing. Everything is published on the pricing page.
When measurement becomes theatre
There is a point where more instrumentation stops helping. A dashboard with forty tiles that nobody opens is worse than five numbers in a weekly email, because it creates the impression of oversight without the substance. If no metric on your dashboard has ever caused someone to change a decision, the dashboard is decoration.
Two more honest limits. If your legacy system has no usable telemetry and your workflows are undocumented, the baseline you construct will be partly estimated, and you should label it as such rather than present it as measurement. And if the programme's real driver is that the incumbent vendor is exiting support, the business case is risk, not efficiency, and measuring efficiency gains will make a sound programme look weak. Pick metrics that match the reason you are doing this.
What this looks like on a live programme
On a university ERP programme covering admissions, fees and examinations, the metric that governed sequencing was not technical. It was the academic calendar: results and fee records could not move mid-term, so each module's go-live was gated on a window, and migration integrity, meaning records reconciled and exceptions still open, was read daily through each cutover rather than weekly. The programme is described in the university ERP modernisation case study. The transferable lesson is that the metric which decides your go or no-go is usually specific to your operating cycle, and it will not be on a generic template.
A measurement plan in eight lines
- Capture the baseline from the systems, over a full cycle, before the first module changes
- Agree the indicators and their targets in writing with the sponsor and the affected function
- Instrument the four layers, and name an owner for each
- Write characterisation tests against the legacy system's real behaviour, bugs included
- Run every module in shadow against the incumbent before it takes traffic alone
- Build the evaluation suite for AI modules before the modules, with a business-agreed threshold
- Set the first outcome review at the end of a full business cycle, not at week four
- Review the dashboard quarterly and delete any metric that has never changed a decision
Related reading
The ROI of legacy application modernization services covers the business case these metrics have to defend, Legacy application modernization services: a practical implementation guide covers the delivery sequence they sit inside, and the AI agent ROI calculator is a starting point for modelling the outcome layer before you commit.
A modernisation is working when a number you agreed on beforehand has moved, and everything else is opinion.
Frequently asked questions
What should you measure before a modernisation programme starts?
▾
Workflow volumes, end-to-end cycle times, error and rework rates, the full running cost of the incumbent, the support ticket profile by category, and the service levels the old system actually delivered. Take these from the systems rather than from interviews, and cover a full business cycle so seasonal variation does not distort the comparison.
How soon should a modernisation show results?
▾
System health improves within days of a module going live. Operational load usually worsens for six to ten weeks while people learn the new system. Business outcomes lag by roughly a quarter. Set the first outcome review at the end of a full business cycle after go-live, and agree that review point in writing before the programme starts.
How do you test AI features added during modernisation?
▾
With an evaluation suite: a fixed set of cases with known correct answers, scored automatically, with a pass threshold agreed in advance by a business reviewer. Run it on every release and on every model change, because providers deprecate models on their own schedule and prompt behaviour is not portable between them.