How to measure whether software product development company is working
How do you measure software product development company?
You measure a software product development company on three things: delivery predictability, product outcomes and the state of the codebase it hands back. Track promised-versus-shipped scope per milestone, one business metric per release, and defect escape rate, and review all three on a fixed cadence.
You measure a software product development company on three things: delivery predictability, product outcomes and the state of the codebase it hands back. Track promised-versus-shipped scope per milestone, one business metric per release, and defect escape rate. If a partner cannot show you all three by the end of week four, the engagement is already drifting.
This article sets out the software product development company metrics we publish to our own clients, the numbers that look reassuring and tell you nothing, the cadence that makes the scorecard useful rather than ceremonial, and what instrumenting a product properly actually costs.
What does a working engagement look like?
A working engagement is one where the forecast you were given in week one still predicts week twelve. That is the whole test, and everything else is a leading indicator of it. Software product development company evaluation goes wrong when buyers judge a partner on activity, because activity is easy to produce and easy to misread.
Three separate questions sit underneath it. Is the team shipping what it said it would ship, in the order it said? Is the shipped software changing a number the business already cared about before the project started? And is the system getting easier or harder to change as it grows? A partner can pass the first and fail the other two for months before anyone notices.
Set the baseline before the build starts. If you cannot state today's figure for the metric you intend to move, you will not be able to prove movement later, and a reasonable partner will insist on capturing it during discovery rather than reconstructing it afterwards.
The three layers of software product development company KPIs
Group the numbers by who acts on them. Engineering metrics are for the delivery lead, product metrics are for the product owner, and health metrics are for whoever will still own this system in three years. Mixing them into one dashboard is how scorecards die.
| Layer | Metrics | Cadence | Acts on it |
|---|---|---|---|
| Delivery | Scope promised vs shipped per milestone, cycle time from ticket to production, deployment frequency | Weekly | Delivery lead |
| Quality | Defect escape rate, crash-free sessions, error budget burn against the agreed SLO | Weekly | Engineering owner |
| Product outcome | One primary business metric per release, activation rate, task completion in the core flow | Per release | Product owner |
| Codebase health | Test coverage on changed code, build time, count of modules with a single expert | Monthly | CTO or architect |
| Commercial | Burn against fixed price, change requests raised and accepted, hosting and licence run rate | Monthly | Sponsor |
Five rows is enough. The moment a scorecard has twenty numbers on it, nobody can say which one triggers a decision, and a partner with something to hide will happily fill the space.
Which metrics mislead?
Most vanity metrics in software delivery share a property: the partner controls the number directly, so it can be improved without the product improving. Watch for these.
- Story points delivered. The team estimates and the team delivers, so the unit is self-referential. Velocity tells you about planning stability, never about value.
- Lines of code or commit count. More code is usually a cost, not an achievement, and both numbers rise fastest when a team is duplicating rather than designing.
- Tickets closed. Closing rate is a throughput measure that says nothing about whether the right tickets existed.
- Uptime alone. A product nobody uses has excellent uptime. Pair availability with an adoption number or it flatters everyone.
- Demo quality. A weekly demo shows the happy path. Ask for the same feature exercised with real production data and the picture changes.
- Hours logged. On a fixed-price programme the hours are the supplier's problem; on time and materials they are a cost, never an outcome.
None of these is useless internally. They mislead when they are used to answer the question of whether the money is working.
How do you build the scorecard in the first two weeks?
Write the metric definitions before the first sprint, agree who publishes them, and put the publication in the contract. In practice that means a short instrumentation workstream running alongside early build: an event schema for the core flows, a dashboard the sponsor can open without asking anyone, and an alert on the one quality number that should stop a release.
Define a service level objective for each user-facing surface and an error budget that goes with it, so that quality has a number rather than an opinion. Google's SRE guidance on service level objectives makes the case for choosing a small number of indicators that represent the user experience rather than instrumenting everything you can measure, which is exactly the discipline a scorecard needs.
Adoption instrumentation deserves the same care as the feature itself. We cover the practical side in measuring adoption after go-live, and for mobile surfaces crash-free sessions is the one quality figure that predicts store ratings better than anything else.
What does the measurement work cost?
Instrumentation is not a separate purchase; it is a line inside the build. Our product and platform development programmes start at $42,000 or ₹28 lakh and run to $175,000 or ₹1.2 crore, and the analytics event layer, dashboards and release gates sit inside that scope rather than arriving as a change request in month four. If the product is a single web application rather than a platform, full stack web application development starts at $14,000 or ₹8.8 lakh.
Where the measurement question comes first, a ten-day Sprint Zero at $3,250 or ₹2,00,000, credited against the build, produces the baseline numbers, the metric definitions and the instrumentation plan. After launch, a Care Plan keeps the dashboards honest: Essential at $1,000 or ₹68,000 a month, Standard at $2,500 or ₹1,60,000, Enterprise at $5,250 or ₹3,40,000 with a named engineer and one-hour response. Every starting figure is on the pricing page, and Indian clients are invoiced in INR with GST.
Mechanics: making the numbers arrive on their own
Instrument the flow, not the page
Events named after user intent survive redesigns. An event called checkout_completed still means something after three UI revisions; an event called button_click_3 does not. Agree the event schema with the product owner and version it in the repository alongside the code.
Gate the release, do not report on it
A metric that only appears in a report is advisory. A metric wired into the deployment pipeline is a rule. Failing evaluation suites, a coverage drop on changed files or an error budget already spent should block a release automatically, and the exception should require a named person to override it.
Review on a cadence nobody can skip
Thirty minutes a week on delivery and quality, one hour a month on outcomes and codebase health, with the same five rows every time. Changing the metric set mid-programme is the most common way a scorecard stops being evidence and starts being narrative.
When measurement is the wrong focus
There are engagements where building this apparatus is a waste of the budget. A four-week proof of concept that exists to answer one technical question needs a pass or fail, not a dashboard. A product with no users yet has no adoption metric worth tracking, and instrumenting it early produces confident charts of noise.
Measurement also fails when the organisation has no appetite to act on a bad number. If the sponsor will not pause a launch when the quality gate goes red, the gate is decoration, and the honest move is to fix the decision rights before buying the tooling. We would rather say that out loud than sell you an observability workstream that changes nothing.
Finally, a scorecard cannot substitute for a relationship that has already broken down. If you are measuring a partner to build a case for terminating them, stop measuring and have the conversation.
What this looks like on a real programme
On the dispatch platform and driver apps we built for a last-mile operator, the metrics that mattered were not engineering ones. The delivery layer tracked milestone scope and cycle time as usual, but the numbers the sponsor watched were operational: how many jobs were dispatched without manual intervention, how often the offline-capable driver app resynchronised cleanly after a dead zone, and how much of the day a dispatcher spent on the phone.
Those figures existed because we agreed them before the first sprint and instrumented the flows during the build, not because someone assembled them at the end to justify the invoice. The engineering metrics kept the programme predictable; the operational metrics told everyone whether the predictability was worth anything.
A checklist before you sign
- Write down today's baseline for the one business metric the product must move
- Agree the five scorecard rows and their definitions in the statement of work
- Name who publishes each number and where the sponsor reads it without asking
- Set the service level objective and the error budget for each user-facing surface
- Decide which metric can block a release and who is allowed to override it
- Book the weekly delivery review and the monthly outcome review before kickoff
- Confirm that dashboards, event schemas and alert rules are handed over with the code
- Check that the contract names the codebase health measures you will inspect at handover
Related reading
How to measure whether AI MVP development is working applies the same logic to a shorter engagement, fixed price, fixed date: how we make it work explains the commercial structure that makes promised-versus-shipped measurable at all, and the adoption metrics glossary entry defines the terms your product owner will need. If the partner refuses to publish numbers, that is itself a result.
Pick five numbers, publish them weekly, and let the trend rather than the demo tell you whether your money is working.
Frequently asked questions
What KPIs should I use to evaluate a software development partner?
▾
Use five: scope promised versus shipped per milestone, cycle time from ticket to production, defect escape rate, one business outcome metric per release, and a codebase health measure such as test coverage on changed files. Any more and nobody can say which number triggers a decision.
How soon should a development partner show measurable results?
▾
Delivery and quality numbers should exist from the first shipped milestone, typically week three or four. Product outcome numbers need real users, so expect them from the first production release onwards. A partner who cannot show delivery metrics by week four is not tracking them.
Is velocity a useful measure of a development company?
▾
Velocity is useful for the team's own planning and useless as a buyer's metric, because the same team both estimates and delivers the points. Judge planning stability by whether milestone scope arrives intact, and judge value by a business metric the product was meant to move.