azyware
Business

How to measure whether full stack development company is working

EZ
Eazyware
· 7 min read
Quick answer

How do you measure full stack development company?

Measure a full stack development company on four layers: delivery predictability, reliability, user experience and business outcome. That means scope delivered against the committed date, change failure rate, time to restore, Core Web Vitals, task completion, and the one business number the application exists to move.

Measure a full stack development company on four layers: delivery predictability, reliability, user experience and business outcome. Concretely that means scope delivered against the committed date, change failure rate and mean time to restore, Core Web Vitals and task completion, and the one business number the application exists to move. Story points measure nothing.

This article gives you the four layers with the specific metrics under each, the instrumentation that produces them, the six numbers that look like progress and are not, and a test suite design that gates a release rather than describing one after it has shipped.

Why the usual dashboard measures the wrong thing

Most project dashboards report activity. Tickets closed, velocity, commits, hours burned. Every one of those numbers rises when a team works harder and also when a team works on the wrong thing, which makes them useless for the question you are actually asking: is this programme going to produce a working application that moves a business number on a date I can plan around.

Worse, activity metrics are trivially gamed without anybody intending to game them. Velocity rises if estimates inflate. Ticket counts rise if work is split more finely. Neither movement tells you anything about the product. A good measurement set has the property that you cannot improve the number without improving the thing, and that property is what to test each metric against before you adopt it.

The second failure is measuring only one layer. Teams that watch delivery alone ship on time and then spend six months on defects. Teams that watch reliability alone run a beautifully stable application nobody uses. The four layers below exist because each one catches a failure the others miss, and the list of ways these programmes go wrong is set out in five ways full stack projects fail.

The four layers

LayerMetricWhere it comes fromA healthy signal
DeliveryCommitted scope delivered by the committed dateThe increment plan and the release logConsistent, with variances explained in the same week they appear
DeliveryChange requests raised against locked scopeThe change registerSteady and small after week four; a rising trend means discovery was thin
ReliabilityChange failure rate and mean time to restoreDeployment and incident recordsFailures happen and are restored in minutes, not hours
ReliabilityError rate and p95 latency per endpointApplication monitoringFlat under load, with alerts that fire before users notice
ExperienceCore Web Vitals on real user trafficField data, not a lab scoreLCP under 2.5 seconds, INP under 200 milliseconds, CLS under 0.1
ExperienceTask completion rate and time on the primary workflowProduct analytics on a named funnelCompletion rising and time falling release over release
OutcomeThe one business number the application exists to moveYour own systems, agreed before the buildMoves within a quarter of launch, measured against a pre-launch baseline

Layer by layer

Delivery: predictability, not speed

The metric that matters is whether what was committed arrives when it was committed. Track it per increment rather than per project, so a slip is visible in week four instead of week fourteen. Track change requests alongside it, because a programme that hits every date while absorbing forty changes did not hit its dates; it renegotiated them quietly. Expectations for a normal schedule are in how long a full stack build takes.

Reliability: failure and recovery

Two numbers carry most of the signal: what proportion of deployments cause a problem, and how long it takes to restore service when one does. A team that deploys weekly with a two per cent failure rate and a ten-minute restore is in far better shape than one that deploys quarterly and has never measured either. Pair these with a written definition of response versus resolution, because most arguments about reliability are really arguments about which of the two was promised.

Experience: field data, not lab scores

Measure the experience real users get on real devices and networks. Google's Core Web Vitals define the thresholds a good experience should meet: Largest Contentful Paint within 2.5 seconds, Interaction to Next Paint under 200 milliseconds, and Cumulative Layout Shift below 0.1, assessed at the 75th percentile of page loads. A lab score on a developer laptop will tell you none of this. Add task completion on your primary workflow, because a fast application that nobody finishes a task in is still failing.

Outcome: one number, agreed in advance

Name the business number before the build starts, with a baseline recorded in the same system you will use to measure it afterwards. Orders processed per operator per day, days sales outstanding, time from enquiry to quote, support tickets per hundred customers. One number, chosen by the sponsor, is worth more than a dozen chosen by the team. The way to construct and defend that number is covered in the ROI post for full stack programmes.

Six numbers that look like progress

  • Velocity. Measures estimation habits, not output. It rises whenever estimates inflate and falls whenever a team estimates honestly.
  • Lines of code or commit count. A proxy for typing. The best weeks on a mature codebase are frequently net negative.
  • Test count. Two thousand tests that never fail are slower feedback, not better safety. Coverage of the critical paths is the metric worth reporting.
  • Uptime alone. Ninety-nine point nine per cent uptime is compatible with a checkout that has been broken for a fortnight. Measure the journeys, not the host.
  • Page views. Rising page views on an internal tool often mean people cannot find what they need.
  • Lab performance scores. A synthetic score from a fast machine on a fast network describes an experience none of your users are having.

Building an evaluation suite that gates a release

A measurement set only changes behaviour if something refuses to ship when the numbers are wrong. That is what an evaluation suite is: a named set of checks that must pass before a build reaches production, run automatically on every merge rather than reviewed in a meeting. For a web application it usually includes the critical-path end-to-end journeys, a performance budget enforced against field thresholds, an accessibility check on the primary templates, a schema migration rehearsal against a copy of production data, and a security scan of dependencies.

Put the suite in before the second increment, not before launch. The value is in catching a regression the week it appears, and a suite written in the final fortnight only documents what already broke. Pair it with feature flags and beta cohorts so a release that passes the gates can still be exposed to ten per cent of users first and pulled without a deploy.

What does the measurement work cost?

Instrumentation is part of the build rather than an extra. A full stack web application at Eazyware starts at $14,000 or ₹8,80,000 and runs to $63,000 or ₹41,60,000, and monitoring, the test suite and the analytics events for the primary funnel sit inside that scope because a delivery you cannot measure is not finished. All starting figures are on the pricing page.

Where cost does appear separately is in reporting depth. If you want dashboards that join application events to commercial data across several systems, that is a data and analytics application from $14,000 or ₹8,80,000 in its own right. After launch, a care plan from $1,000 or ₹68,000 a month keeps the alerts, the suite and the dependency scans maintained, which matters because an unattended test suite decays into noise within two quarters.

When measuring harder is the wrong move

There is a point where instrumentation becomes a substitute for judgement. An internal tool with forty users does not need a real-user monitoring pipeline, a seven-metric dashboard and a weekly metrics review; it needs someone to ask the forty users whether it is any good. Small audiences are better measured by conversation, and the money saved on telemetry is better spent on the product.

Measuring too early is the other error. Before a workflow has real traffic, funnel numbers are noise, and teams that optimise against noise chase movements that were never signal. Wait until the baseline is stable, then act. A related mistake is adding metrics without removing any, until a dashboard nobody reads is quietly reporting a broken checkout to an empty room.

And a metric with no owner is decoration. Every number in your set should have a named person who is expected to explain it when it moves. If you cannot name that person, delete the metric rather than displaying it.

What this looks like in practice

Our in-app copilot for a field-service SaaS is a useful illustration because the success criterion was not a technical one. The question was whether the product became the tool users opened first, which is an adoption measure taken from the client's own usage data rather than from a delivery report. Programmes framed that way tend to make better product decisions, because the team can see which release moved the number and which one did not.

The hidden costs quotes leave out covers the running lines that show up in the same reviews as these metrics, and questions to ask a vendor before you sign tells you how to get measurement written into the contract rather than added later. If you want help choosing the one outcome number for your programme, start with the workflow.

Pick four numbers, give each an owner, and let one of them be allowed to stop a release.

Frequently asked questions

What are the most important metrics for a web application project?

▾

Four: committed scope delivered by the committed date, change failure rate with mean time to restore, Core Web Vitals measured on real user traffic, and one business outcome agreed before the build starts. Each layer catches a failure the others miss, which is why a single-layer dashboard misleads so reliably.

Why is velocity a poor measure of a development team?

▾

Velocity measures estimation habits rather than output. It rises when estimates inflate and falls when a team estimates honestly, and it can be improved without improving the product at all. A useful metric is one you cannot move without moving the underlying thing, which velocity, commit counts and test counts all fail.

When should measurement be built into a project?

▾

From the second increment, not before launch. Monitoring, the critical-path test suite and analytics events for the primary funnel belong inside the build scope at Eazyware, because a delivery you cannot measure is not finished. A suite written in the final fortnight only documents regressions that have already shipped.