azyware
Technology

Five ways application maintenance and support services projects fail, and how to avoid each

EZ
Eazyware
· 7 min read
Quick answer

Why do application maintenance and support services projects fail?

Application maintenance and support fails for five recurring reasons: a transition with no verified runbooks, an SLA measuring acknowledgement not resolution, preventive work crowded out by tickets, ambiguous ownership, and a contract you cannot exit. Each has an early warning signal and a fix.

Application maintenance and support services fail for five recurring reasons: a transition that produced no verified runbooks, an SLA that measures acknowledgement rather than resolution, preventive work crowded out by ticket volume, ambiguous ownership at system boundaries, and a contract you cannot leave. None of these is about engineering skill, and every one of them is visible in the first ninety days.

This post takes each pattern in turn: what it looks like from the inside, the signal that appears before anyone admits there is a problem, and the specific decision that prevents it.

The five patterns, and their early warning signals

Failure patternEarly warning signalThe decision that prevents it
No verified runbooksFirst night incident escalates to the previous teamEnd discovery on a tested restore and rollback, not a document
SLA measures the wrong thingResponse targets always met, users still unhappySeparate response and resolution clocks by severity
Preventive work squeezed outDependency versions unchanged for two quartersRing-fence a fixed share of hours for patching
Ambiguous ownershipTickets bounce between teams for daysName a triage owner with authority over both sides
Lock-in by designRunbooks and monitoring live in the vendor's accountsRequire artefacts in your repository from week one

One pattern deserves naming before the others, because it hides them all. A monthly report built from ticket counts will show every one of these five failures as a healthy green number, since closures rise when quality falls and stay flat when preventive work is skipped. Review the trend in repeat incidents, the age of the oldest unpatched dependency and the time a ticket waits for an owner, and the picture changes immediately.

The table is the short version. The detail matters because each pattern has a plausible-sounding excuse attached to it, and the excuse is usually what keeps it alive for a year.

Pattern one: a transition that produced documents nobody tested

The commonest failure is a takeover that ended on paper. A handover pack was produced, a slide said transition complete, and the first genuine incident at 2am escalated straight back to the team you were replacing.

It happens because discovery is scoped as reading rather than as verification. Reading produces a document. Verification produces knowledge that a restore actually completes, that the last release can actually be rolled back, and that the alert path actually reaches a human. In most takeovers at least one of those three fails, which is exactly why they are worth testing.

The fix is to make discovery end on evidence. A transition is complete when an engineer who has never seen the system can work a severity-one incident from the runbook, and when the restore and rollback have been performed with the timings recorded. The phasing that supports this is set out in our implementation guide to application maintenance and support services.

Pattern two: an SLA that measures acknowledgement

The second pattern shows up as a contradiction in the monthly report: every target is green and the business is complaining. It nearly always means the SLA measures response, and only response.

A response clock stops when someone acknowledges the ticket. A resolution clock stops when the user can work again. A vendor can hit a one-hour response target indefinitely while a severity-two defect sits open for three weeks, and nothing in the contract is breached. The excuse is reasonable on its face, because resolution time genuinely depends on the defect, which is why the answer is not a single resolution number but a resolution target per severity with an escalation trigger when it is missed.

Define severity by business impact rather than by volume of complaint: payments failing is severity one, a slow export is not. Then publish both clocks. SLAs that mean something covers the wording, and SLA response vs resolution covers the distinction in one page.

Pattern three: preventive work crowded out by tickets

This is the slowest failure and the most expensive. Ticket volume is visible and urgent; patching is invisible until it is an incident. Month after month the hours pool goes to corrective work, and two quarters later the dependency tree has not moved, the runtime is approaching end of support and a routine upgrade has become a project.

The OWASP Top 10 lists vulnerable and outdated components as a top-ranked web application security risk, and notes that organisations frequently do not know the versions of the components they run. A support contract that never reserves hours for upgrades is how a codebase arrives in that state while appearing to be professionally maintained.

The fix is arithmetic rather than intent. Ring-fence a share of the monthly hours for preventive work and report it separately, so that spending it on tickets requires a decision rather than happening by drift. Publish the cadence, including the threshold at which a critical advisory breaks the release freeze, as described in security patching cadence for production applications.

Pattern four: ambiguous ownership at the boundaries

Tickets that bounce are a symptom of a boundary that was never drawn. The classic version is an application maintained by one team, an integration owned by another and infrastructure owned by a third, with a fault whose origin is genuinely unclear.

The unhelpful response is a longer responsibility matrix. Matrices describe steady state well and ambiguity badly, and every real dispute is ambiguity. The useful response is to name a single triage owner with the authority to pull in either side without a meeting, and to write into the contract that the triage owner holds the ticket until the origin is established. Whoever is wrong about the origin is not penalised, because penalising misdiagnosis teaches people to hand tickets away rather than investigate.

This is sharpest during a phased migration, when old and new paths run at once and an incident can legitimately belong to either. That was the live constraint in this university ERP modernisation programme, where modules moved one at a time around the academic calendar and the support boundary moved with them.

Pattern five: a contract you cannot leave

The last pattern is quiet. Nothing is broken, the reports are fine, and then you decide to change vendor and discover the runbooks live in their wiki, the monitoring is configured in their cloud account, the alerting integrations are on their credentials and the deployment pipeline runs on a machine you cannot see.

Nobody plans this; it accumulates. Prevent it with a rule that costs nothing at the start and is very difficult to retrofit: every artefact produced under the contract lands in your repository, in your accounts, under your identity provider. Runbooks, infrastructure definitions, monitoring configuration, dashboards, scripts. If it cannot be handed over in a week, it was never yours.

Write the exit clause on the way in, with a notice period, a named list of handover artefacts and a support window that overlaps with the next provider. Our position on ownership is unambiguous: you own the code, the infrastructure, the prompts and the documentation, and a support contract should not create an exception to that.

The warning signs, in one list

  • Every SLA target is green and the business is unhappy. You are measuring acknowledgement, not outcomes.
  • The first serious incident escalated to the previous team. Discovery ended on documents rather than verification.
  • Dependency versions have not changed in two quarters. Preventive work is being silently deprioritised.
  • The same tickets reopen. Symptoms are being fixed rather than causes; watch the reopen rate rather than the close rate.
  • Tickets take days to find an owner. Nobody holds triage authority across the boundary.
  • You cannot name where the runbooks live. You have a dependency, not a contract.
  • Monthly reports arrive as a ticket count. Counting closures says nothing about whether the system is getting healthier.

When the contract is not the problem

One honest caveat. If most of your incidents come from the same subsystem, no support arrangement will fix it, and changing vendor will not either. Faster response to a recurring structural failure is a treadmill. The money belongs in remediation or in a legacy modernisation programme, which starts at $31,500 or ₹22,40,000, and the incremental route is described in the strangler pattern.

Equally, if you have bought business-hours cover and are unhappy about weekend response, the contract is working as written. The remedy is a different tier rather than a dispute: cover runs from $1,000 or ₹68,000 a month for business hours with an eight-hour response, through $2,500 or ₹1,60,000 for 24 x 5, to $5,250 or ₹3,40,000 for 24 x 7 with a one-hour response and a named engineer. Systems with models in production add $750 or ₹40,000 for evaluation runs and cost monitoring. The inclusions sit on the maintenance and support service page and the numbers on the pricing page.

Application maintenance contracts: what should be in an AMC covers the clauses that prevent three of these five patterns, how to measure whether application maintenance and support services is working covers the metrics that expose the other two, and what a care plan should cost and what it should include covers the tier you should actually be on. If a current arrangement shows two or more of these signals, tell us what you are seeing.

Support fails at the boundaries and in the invisible work, so put your attention there rather than on the ticket count.

Frequently asked questions

Why do maintenance vendors hit every SLA while users stay frustrated?

▾

Because the SLA measures response rather than resolution. A response clock stops when someone acknowledges the ticket; a resolution clock stops when the user can work again. Fix it by publishing a resolution target for each severity level, with an escalation trigger that fires automatically when the target is missed.

How do you stop security patching being deprioritised in a support contract?

▾

Ring-fence a fixed share of the monthly hours for preventive work and report it separately from corrective work, so spending it on tickets becomes a deliberate decision. Publish the patch cadence and the severity threshold at which a critical advisory breaks a release freeze, and review dependency versions monthly.

What makes a support contract hard to exit?

▾

Artefacts that live in the vendor's estate: runbooks in their wiki, monitoring in their cloud account, alerting on their credentials, pipelines on machines you cannot see. Require from week one that everything produced lands in your repository and your accounts, and write a notice period and handover list into the contract.