SLAs that mean something: response, resolution and cover
What should a software support SLA define to be meaningful?
A useful SLA states response and resolution by severity, cover hours, escalation path and what happens when it is missed. It also defines the severities in plain language, says whose clock is used, and reports performance against itself every month, so both sides can see whether the promise is being kept.
A support SLA definition worth signing has five parts: what counts as which severity, how quickly someone responds to each, how quickly each is resolved or mitigated, during which hours those clocks run, and what happens when a target is missed. Most software support SLAs we review have one or two of these and a lot of adjectives. "Prompt" is not a response time. "Commercially reasonable" is not a resolution target. This article sets out the structure we use in our own Care Plans, explains the trade-offs behind each choice, and shows how to read a vendor's SLA so you know what you are actually buying.
Why most support SLAs are decorative
An SLA fails in one of three ways. It defines response but not resolution, so the vendor can acknowledge a ticket in an hour and fix it in a month while staying compliant. It defines times but not severities, so every ticket is argued into the lowest tier. Or it defines everything except the consequence, so a missed target is a conversation rather than a credit. A meaningful SLA closes all three gaps, and it is short enough that the person raising the ticket can remember it.
| Severity | Definition | Response target | Resolution or mitigation target |
|---|---|---|---|
| S1 Critical | System down, data at risk, or a security incident | Per tier: 8 h, 4 h or 1 h | Mitigation within the same cover window; root cause fix scheduled |
| S2 High | Core function broken for many users, no workaround | Same cover day | Fix or workaround within a few working days |
| S3 Medium | Function degraded, workaround exists | Next cover day | Scheduled in the next release |
| S4 Low | Cosmetic, minor, or a question | Within two cover days | Backlog, reviewed monthly |
Severity definitions in plain language
Severity is where SLA arguments start, so the definitions must be concrete. We write them as situations, with examples from the specific system: "customers cannot check out" is S1; "the export report times out for large accounts" is S2; "the dashboard chart renders slowly" is S3. The client's named contact assigns severity when raising the ticket; the vendor can propose a change with a reason, but not silently downgrade. A security incident is always S1, whatever its visible impact, because the clock on disclosure obligations starts at detection.
Response time: when the clock starts and what it means
Response means a qualified engineer has acknowledged the ticket, confirmed the severity and started work. It does not mean an auto-reply. The clock starts when the ticket is logged through the agreed channel, not when someone reads it, and it runs only within cover hours unless the tier is 24×7. State the channel: a ticket system with timestamps, not a phone call to someone's mobile. Incident response times are only auditable if the timestamps exist.
Resolution: the target vendors leave out
Resolution is harder to promise than response because some problems are genuinely hard. The honest way to handle this is to separate mitigation from fix. For S1, the promise is mitigation within the cover window: service restored, even if by rollback, failover or a temporary workaround. The root-cause fix is then scheduled with a date. For S2, a fix or workaround within a stated number of working days. For S3 and S4, a scheduling commitment rather than a time. This structure lets the vendor make a promise it can keep and gives the client the thing they actually need, which is the service back.
Cover hours and 24×7 support tiers
Cover hours are the multiplier on everything else. An eight-hour response during business hours in one timezone can mean a Friday evening outage is first looked at on Monday. That is acceptable for an internal tool and unacceptable for a payments flow. We offer three tiers so the price matches the need: Essential covers business hours IST with an eight-hour critical response; Standard covers 24×5 with a four-hour response; Enterprise covers 24×7 with a one-hour response and a named engineer who knows the system. The question to ask is not "do we want 24×7" but "what does an hour of outage cost at 2 a.m. on Sunday, and how often does it happen". The cost side is in what a Care Plan should cost.
Named engineer versus rota
A rota gives cover; a named engineer gives context. For systems with deep domain logic, the first hour of an incident is often spent understanding the system, and a named engineer removes that hour. It costs more because it constrains the vendor's staffing. It is worth paying for on systems where S1 incidents are expensive and rare rather than cheap and frequent.
Escalation path and communication
The SLA should name who is contacted if the response target is missed, on both sides, with a second level above them. It should also set a communication cadence during an S1: an update at a stated interval until mitigation, then a written incident report within a few working days covering timeline, cause, fix and what will prevent recurrence. Google's SRE book has a useful treatment of service level objectives and why the target should be what users need rather than the best the team can do; the same logic applies to support SLAs.
What happens when the SLA is missed
A consequence makes the SLA real. Common forms are a service credit against the next month's fee for each missed S1 or S2 target, a formal review after repeated misses, and a termination right after a stated number of misses in a period. Credits should be automatic, triggered by the monthly report rather than by the client having to claim them. We prefer credits to penalties because they keep the relationship commercial rather than adversarial, but the point is that missing the target costs something.
Reporting against the SLA
Every month, the report lists each ticket with severity, time logged, time responded, time mitigated and time resolved, alongside the targets. Misses are highlighted, with the reason. Over a quarter this shows whether the severities are being assigned honestly, whether the cover tier is right, and whether the system itself is generating more incidents than it should, which is a maintenance question rather than a support one. The report belongs to the client and lives in the client's systems. For AI components, we add evaluation results and cost per feature, since a model regression is an incident in everything but name.
A worked example
A hospital network running a multilingual voice agent for appointment booking had an SLA from its telephony vendor that promised "24×7 support" with no severities and no resolution terms. When the speech provider had a regional outage on a weekend, the ticket was acknowledged within minutes and not actioned until Monday, entirely within the letter of the contract. We rewrote the SLA for the whole stack around four severities, made a failed call-completion rate above a threshold an automatic S1 with a paged engineer, set the mitigation promise as failover to the secondary speech provider, and added the monthly report. The next provider outage was mitigated by routing before most patients noticed. The build is described in multilingual voice agent for a hospital network, and the voice agents page covers the stack.
Team and timeline
Writing the SLA is a day of work with the client's operations lead and our support lead, and it is done before the Care Plan starts rather than after the first incident. Essential is $1,000 a month, Standard $2,500 and Enterprise $5,250, with INR pricing on the pricing page; the tier sets the cover hours and critical response, and the severity structure is the same across all three. For systems built elsewhere, a short takeover audit precedes the SLA so the targets are set against a known baseline; see taking over a system you didn't build.
Before you start: a checklist
- Four severities defined as situations, with examples from your system
- Response, mitigation and resolution targets per severity
- Cover hours and timezone stated, and whether clocks run outside them
- A ticket channel with timestamps as the only official route
- Escalation contacts named on both sides, two levels deep
- Communication cadence during S1 and a written incident report afterwards
- Service credit or other consequence for missed targets, applied automatically
- Monthly report format agreed, with a named reader on your side
Questions clients ask
- Can we have a one-hour response on business-hours cover? Yes; response and cover hours are independent. One hour during business hours is a reasonable middle option.
- What counts as a security incident? Any confirmed or suspected unauthorised access, data exposure or exploited vulnerability. It is S1 regardless of visible impact.
- Who decides severity? Your named contact when raising the ticket. We may propose a change with a reason; we do not downgrade silently.
- Does the SLA cover third-party outages? Response and mitigation yes; resolution of the third party's problem no. Mitigation usually means failover or a workaround.
- Is 24×7 worth it for a small team? Only if the system earns money or serves users overnight. Many clients start on Standard and move up when usage justifies it.
Related reading
Application maintenance contracts: what should be in an AMC, Security patching cadence for production applications and How to measure an AI agent cover the surrounding commitments.
An SLA is a set of clocks with consequences; if you cannot say when one started and what happens when it runs out, it is a brochure.
Frequently asked questions
What is the difference between response time and resolution time?
▾
Response is when a qualified engineer acknowledges the ticket and starts work. Resolution is when the problem is fixed or, for critical incidents, mitigated so service is restored. A meaningful SLA states both, by severity.
What response time should we expect for a critical incident?
▾
It depends on the tier you pay for: eight hours in business hours, four hours on 24×5 or one hour on 24×7 in our Care Plans. The right tier depends on what an hour of outage costs you.
Should SLA credits be automatic?
▾
Yes. Credits triggered by the monthly report, without the client having to claim them, keep the SLA honest and the relationship commercial. A credit you have to argue for is not much of a consequence.