azyware
Technology

How to measure whether custom enterprise software development is working

EZ
Eazyware
· 7 min read
Quick answer

How do you measure custom enterprise software development?

You measure custom enterprise software development on four layers: delivery health, adoption, business outcome and running cost. Pick two or three numbers per layer before the first sprint, baseline them against the process the software replaces, then review them monthly. A build with no baseline can only be defended.

You measure custom enterprise software development on four layers: delivery health, adoption, business outcome and running cost. Choose two or three numbers on each layer before the first sprint starts, baseline them against the manual process the software replaces, and review them monthly with the people who own the work. A build with no baseline can only be defended, never judged.

What follows is the metric ledger we put in front of steering committees: which custom enterprise software development KPIs earn their place on each layer, which numbers flatter a project without proving anything, what the measurement work itself costs, and the point at which more instrumentation stops paying for itself.

What "working" actually means for an internal system

An internal enterprise system is working when the people who do the job today finish it faster, with fewer errors or fewer handoffs, and when a number the finance team already trusts moves in the right direction. Everything else is proxy evidence.

That definition has an awkward consequence. Most of the value of a custom build shows up in someone else's ledger: fewer credit notes, a shorter order-to-cash cycle, less overtime at month end. If you do not agree in writing which of those ledgers the project may claim, the evaluation argument happens at the end of the programme instead of the start, and by then nobody has the before picture.

So the first measurement task is not instrumentation. It is a fortnight of watching the current process and recording how long each step takes, how often it is redone, and how many systems a person opens per case. That baseline is worth more than any dashboard you build later, and it expires the moment the new software ships.

The four layers of custom enterprise software development metrics

Layer one: delivery health

Delivery health tells you whether the build will still be maintainable in year three. Track lead time from accepted story to production, change failure rate, and the proportion of releases that needed a rollback. These are process metrics, not value metrics; a team can score well on all three and still build the wrong system, which is why they never travel alone.

Layer two: adoption

Adoption kills more enterprise projects than any other layer, and it is usually measured badly. Logins are not adoption. The number that matters is task completion inside the new system, expressed as a share of the total volume of that task across the organisation. If half the invoices still arrive by email to a shared inbox, the software is at fifty per cent regardless of how good the demo was. The pattern is common enough that we wrote about why enterprise software implementations fail on adoption separately.

Layer three: business outcome

Business outcome is the layer your sponsor cares about: cycle time, error rate, cost per transaction, revenue leakage recovered. Each of these needs a source system that is not the new software, because a system measuring its own success is not evidence. Pull the number from the ERP, the general ledger or the data warehouse.

Layer four: running cost

Running cost covers hosting, licences the build depends on, and the support hours the system consumes each month. A build that saves four hundred hours a quarter but needs thirty hours of engineering attention a month has a thinner margin than it looks. Track support hours against the allowance in your care plan, and treat a rising trend as a defect, not as demand.

A metric ledger you can take to a steering committee

LayerMetricWhat it provesHow it misleads
DeliveryLead time, story accepted to productionThe team can ship small changes safelyFalls when the team splits work into meaningless tickets
DeliveryChange failure rateReleases are tested against real casesLooks perfect when nothing risky ships
AdoptionShare of total task volume completed in the systemPeople have actually switchedHides pockets of the business still on spreadsheets unless volume is org-wide
AdoptionMedian time to complete one caseThe interface fits the workImproves if the hard cases quietly route around the system
OutcomeCycle time from source systemThe process is genuinely fasterMoves for seasonal reasons; compare like periods
CostSupport hours consumed per monthThe system is stableStays low while an internal hero patches things unofficially

How do you choose which KPIs to track?

Choose the smallest set that a sceptical CFO would accept as evidence. Six to eight numbers is usually right, and each one should pass these tests before it goes on the page.

  • It has a baseline. You know the value before the software existed, measured the same way. Without that, the metric is decoration.
  • It comes from a system the business already trusts. Finance data beats application telemetry in any argument about value.
  • It has an owner who is not on the build team. The person accountable for the outcome reports the number.
  • It cannot be gamed by shipping less. If the fastest route to a good score is to avoid difficult work, replace the metric.
  • It survives a bad quarter. A metric that only moves when demand is high tells you about the market, not the software.
  • It maps to a decision. Say in advance what you will do if it goes the wrong way for two consecutive months.

Google's site reliability engineering team makes the same argument about service metrics: pick a small number of indicators that reflect what users actually experience, because a wall of dashboards produces no decisions. Their chapter on service level objectives is the clearest primary source on choosing few indicators and setting explicit targets against them.

What does the measurement work cost?

Budget five to eight per cent of the build for instrumentation, baselining and the reporting layer. On a custom and enterprise software development programme, which runs from $24,500 or ₹16,00,000 to $175,000 or ₹1.2 Cr depending on scope, that is a real line item and it should appear in the quote rather than being absorbed. All of our starting prices are published on the pricing page.

Baselining is cheaper than people expect. A ten-day Sprint Zero at $3,250 or ₹2,00,000, credited against the build, is usually enough to map the current process, record the before numbers and agree the ledger with the sponsor. Doing it after go-live costs the same and buys nothing, because the process you wanted to measure no longer exists.

After launch, the review cadence is support work. Care Plans run from $1,000 or ₹68,000 a month for Essential, with business-hours cover in IST and an eight-hour response, to $5,250 or ₹3,40,000 a month for Enterprise, with round-the-clock cover, a one-hour response and a named engineer. The sixty hours a month on Enterprise are enough to run the review and act on it; the ten hours on Essential are not, and pretending otherwise is how measurement quietly stops. Both tiers sit under software maintenance and support.

The numbers that mislead

Four metrics show up in almost every enterprise status pack and almost none of them survive contact with a hard question.

Story points delivered measures how the team estimates, not what it built. Logins measure curiosity in week one and habit never. Percentage of scope complete describes a plan written before anyone learned anything. Satisfaction surveys taken in the first fortnight measure novelty; run them at ninety days and they start telling the truth.

There is a subtler trap in the outcome layer. When a new system makes errors visible for the first time, the error rate goes up, and a naive reading says the software made things worse. Record in advance whether each metric should rise, fall or stay flat in the first quarter, so a predicted rise is not read as failure.

When heavy measurement is the wrong choice

Not every build deserves this apparatus. If the system serves fewer than about twenty people and costs less than a quarter of an engineer's annual salary, a monthly conversation with the three people who use it will tell you more than a dashboard will. Instrumentation has a running cost and a maintenance burden of its own.

Heavy measurement is also wrong when the decision it informs has already been made. If the regulator requires the system, or the incumbent vendor exits support next year, the metric ledger is useful for tuning the build, not for justifying it. Say so plainly rather than manufacturing a return.

What a real review looks like

When we modernised a fifteen-year-old university ERP, the measurable question was not whether the new modules were better in the abstract. It was whether admissions, fees and results could run through the new path during a live academic cycle without the old path as a safety net. The monthly review counted cases completed, exceptions that fell back, and support hours consumed, and each module moved on only when those three numbers held for a full cycle. The university ERP modernisation case study describes the sequencing.

That is the shape of an honest review: three numbers, one decision per module, and a written rule for what happens when a number goes the wrong way. It fits on one page and takes forty minutes.

Checklist before the first sprint

  • Write down the process baseline: time per case, rework rate, systems opened, volume per month
  • Agree which ledger the project is allowed to claim savings in, and who signs that claim off
  • Name an owner outside the build team for each outcome metric
  • Choose the source system for every number, and confirm someone can export it without heroics
  • Set the review date and the escalation rule before go-live, not after
  • Budget the instrumentation and the post-launch review hours explicitly in the quote

Our practical implementation guide to custom enterprise software development covers the build sequence these metrics sit on top of, and the ROI of custom enterprise software development turns the outcome layer into a business case a finance team will accept. If adoption is the layer you are worried about, the adoption metrics glossary entry defines the measures we use and what each one omits.

Measure the four layers, publish the baseline before you write code, and accept that a project which refuses to name a number in advance has already told you what it expects to find.

Frequently asked questions

What are the most important metrics for custom enterprise software?

▾

Task completion share inside the system, median time to complete one case, a cycle-time or error-rate figure pulled from a source system outside the build, and monthly support hours consumed. Those four cover adoption, usability, business outcome and running cost, and each needs a baseline recorded before the software ships.

How soon after go-live can you tell if an enterprise build is working?

▾

Adoption and delivery signals are readable within four to six weeks. Business outcome usually needs one full business cycle, which for month-end processes means a quarter and for annual cycles such as academic admissions can mean longer. Judging outcome metrics before a complete cycle produces confident conclusions from incomplete data.

Who should own the metrics for an internal software project?

▾

The outcome metrics belong to the business owner of the process, not the delivery team, because a team reporting on its own success is not evidence. The delivery team owns lead time, change failure rate and support hours. Both sets appear in the same monthly review, with one named person per number.