How to measure whether API development services is working
How do you measure API development services?
Measure three layers separately: delivery, whether the signed contract shipped on time and unchanged; operations, meaning availability, 95th and 99th percentile latency and error rate; and adoption, meaning integrations live and tickets per integration. Two strong layers and one weak one is not a working programme.
Measure three layers separately. Delivery: did the signed contract ship on time and unchanged. Operations: availability, latency at the 95th and 99th percentiles, and error rate split by cause. Adoption: integrations live, calls per consumer, and support tickets per integration. A programme scoring well on two layers and badly on the third is not working.
This article sets out what to instrument in each layer, what a defensible target looks like, which popular numbers actively mislead, and how to turn the whole thing into a gate that a release has to pass rather than a dashboard nobody opens.
Why API programmes get measured badly
The usual failure is measuring the thing that is easy to count. Endpoint counts, story points burned and lines of documentation all go up reliably while an API gets worse. None of them tells you whether a partner can integrate without calling your team, which is the only question that matters commercially. Counting output is comfortable precisely because it never returns bad news.
The second failure is measuring averages. Mean response time is close to useless for an API, because the tail is what breaks consumers. An endpoint averaging 120 milliseconds while five per cent of calls take four seconds will produce timeouts, retries and duplicate writes in every client that talks to it, and the average will not move enough to notice.
The third failure is measuring only after launch. If the first measurement happens at go-live, you have no baseline, no gate and no way to attribute a regression to a change. Eazyware builds the measurement into API development and integrations work from the specification onwards, because a contract you cannot test against is a contract you cannot enforce.
The three layers
Delivery: did you get what you signed
Delivery metrics compare the signed OpenAPI specification to what actually shipped. Count specification drift, meaning endpoints, fields or status codes that differ from the agreed document. Count change requests and separate the ones caused by genuine new requirements from the ones caused by discovery that should have happened earlier. Track date variance against the original plan and, more usefully, how early a slip was reported. Two weeks of warning is a scheduling problem; two days is a trust problem. A team that reports a two-week slip in week three is managing the project; one that reports it in the final week is not.
Operations: does it hold up
Operational metrics describe the API as consumers experience it. Availability measured from outside your network, not from a health check inside it. Latency at the 95th and 99th percentiles per endpoint, because that is where timeouts live. Error rate split into client errors, server errors and upstream failures, since the three imply completely different fixes. Webhook delivery success on first attempt, and how many succeed only after retries. Queue depth and consumer lag where work is asynchronous.
Adoption: is anyone using it
Adoption is the layer most programmes skip and the one that predicts commercial return. Count integrations live against integrations planned. Count time from credential issue to a consumer's first successful call, which is the single best proxy for developer experience. Count support tickets per integration and read the topics: repeated questions about the same field mean the documentation is wrong, not that the integrators are careless. Adoption numbers also tell you when to stop building. If three of twelve planned integrations carry ninety per cent of the traffic six months in, the remaining nine were a plan, not a requirement, and the maintenance they would have cost is money saved.
What to instrument, and where to start
| Metric | Layer | How to instrument | A sane starting target |
|---|---|---|---|
| Specification drift | Delivery | Contract tests run against the signed OpenAPI document in CI | Zero undeclared differences at release |
| Availability | Operations | External synthetic probes per critical endpoint | 99.9 per cent monthly for internal APIs, higher for partner-facing |
| p95 and p99 latency | Operations | Server-side traces exported per route and per consumer | p95 under 300 ms for reads, with writes agreed per endpoint |
| Error rate by class | Operations | Structured logs tagged 4xx, 5xx and upstream | Server errors under 0.1 per cent of calls |
| Webhook first-attempt delivery | Operations | Delivery log with attempt count and final status | Above 99 per cent on first attempt |
| Time to first successful call | Adoption | Timestamp from credential issue to first 2xx per consumer | Under one working day with sandbox access |
| Tickets per live integration | Adoption | Support system tagged by consumer and topic | Falling month on month after the first quarter |
The numbers that mislead
Each of these appears in status reports and each can rise while the API deteriorates.
- Total API calls. Volume grows when clients retry. A rising call count alongside a rising error rate is a symptom, not a success.
- Mean latency. Hides the tail that actually breaks consumers. Report percentiles or report nothing.
- Endpoints shipped. Twelve endpoints nobody integrates with are worse than three that carry the business, because you now maintain twelve.
- Uptime measured from inside. A health check that passes while DNS, TLS or the gateway fails is measuring the wrong thing entirely.
- Documentation page views. High views with high ticket volume means the documentation is being read and is not answering the question.
- Test coverage percentage. Coverage says lines ran, not that the contract holds. Contract tests are the ones that catch breakage.
How do you set a target you can defend?
Pick targets from consumer impact rather than from what the system currently does. The discipline is straightforward: choose a service level indicator that reflects what users feel, set a service level objective slightly tighter than the point at which consumers complain, and treat the gap between the objective and one hundred per cent as an error budget you are allowed to spend on releases. Google's engineering teams document this approach in the chapter on service level objectives, and the error budget idea is what stops the argument between shipping and stability becoming a personality contest.
Two practical rules follow. Set different objectives for different endpoints, because a bulk export and a payment authorisation do not deserve the same number. And publish the objective to consumers of a partner-facing API, because an undocumented expectation is not a commitment. Review the numbers quarterly against real complaints: an objective nobody has ever missed is set too loose to be informative, and one missed every month is a budget for arguments rather than releases. Response and resolution commitments are a separate matter from availability, and SLAs that mean something covers how those are written.
Make the measurement a gate, not a dashboard
Dashboards get looked at after an incident. Gates prevent the incident. The version that works: contract tests run in continuous integration against the signed specification, and a breaking change fails the build unless a version bump accompanies it. A load test runs against staging with production-shaped traffic and fails if p99 regresses beyond an agreed margin. A synthetic consumer, one small client that calls the API the way a real integrator does, runs on a schedule and pages when it cannot complete its journey. Each of those is cheap to build during the project and disproportionately expensive to add after the first partner is live, because by then a failing gate blocks someone else's roadmap as well as yours.
Roll releases out behind flags to one consumer before all of them, so that the first evidence of a regression comes from a controlled cohort rather than from your largest partner. Feature flags and beta cohorts describes the mechanics, and the same staged approach applies to deprecations.
What this costs and when it happens
Instrumentation is not a separate project. Contract tests, structured logging, traces and a synthetic consumer are part of a competent build and are included in Eazyware's API engagements, which run from $7,000 or ₹4,40,000 to $35,000 or ₹23,20,000 depending on the number of systems in scope. Ongoing monitoring, target reviews and the on-call cover behind them sit in a Care Plan: $1,000 or ₹68,000 a month for business-hours IST, $2,500 or ₹1,60,000 for 24x5 with a four-hour response, and $5,250 or ₹3,40,000 for 24x7 with a one-hour response and a named engineer. Current ranges are on the pricing page, and the cover levels are set out under maintenance and support.
When measurement is the wrong focus
If the API has two internal consumers and both teams sit in the same building, a full measurement programme is overhead. Instrument errors and p99 latency, keep contract tests in CI, and skip the rest until a third consumer appears. Measurement should be proportional to the number of people who will be surprised when something breaks. A three-person team running a full error-budget process on an internal API is spending review meetings it does not have.
Measurement is also the wrong focus when the real problem is that nobody agreed what the API was for. A precise dashboard over a contested contract produces confident reporting about the wrong thing. Fix the specification and the ownership question first; numbers cannot arbitrate a disagreement about purpose.
Related reading
API-first SaaS explains why the signed contract is what makes any of this measurable, five ways API development projects fail covers the failure modes these metrics surface first, and the ROI of API development services turns the adoption layer into a business case.
Instrument the contract, report percentiles, and let the numbers block a release rather than explain one.
Frequently asked questions
What are the most important API metrics to track?
▾
Availability measured externally, latency at the 95th and 99th percentiles per endpoint, error rate split into client, server and upstream causes, webhook first-attempt delivery, and time from credential issue to a consumer's first successful call. The last one is the best single proxy for whether integrators can work without your help.
Why is average API response time a bad metric?
▾
Because consumers break on the tail, not the mean. An endpoint averaging 120 milliseconds while five per cent of calls take four seconds produces timeouts, retries and duplicate writes in every client, and the average barely moves. Report the 95th and 99th percentiles per endpoint instead, and set different targets for different routes.
How do you stop a breaking API change reaching consumers?
▾
Run contract tests in continuous integration against the signed OpenAPI document so that any undeclared difference fails the build unless a version bump accompanies it. Add a scheduled synthetic consumer that calls the API the way a real integrator does, and release behind flags to one consumer before all of them.