How to rank AI use cases by ROI, not excitement
How should a company prioritise AI use cases by return on investment?
Score each use case on business value, data readiness, model risk, build effort and monthly running cost, in the open with stakeholders. The session takes an afternoon, the result is a ranked list everyone has seen, and the first build is the one that scores well on all five, not the one with the loudest sponsor.
AI use case prioritisation works when it is a scoring exercise done in the open. Put every candidate on one sheet, score each on business value, data readiness, model risk, build effort and monthly running cost, and do it in a room with the people who own the process, the data and the budget. The ranked list that comes out is rarely the one the executive sponsor expected, and that is the point: the exercise replaces enthusiasm with evidence before any money is spent. This article describes the framework we use in AI product strategy engagements, the questions behind each score, and the traps that make a ranking wrong.
Why AI use case prioritisation goes wrong
Most companies arrive with a list assembled from vendor demos, conference talks and one executive's pet idea. The list is ranked by excitement: the most visible or futuristic item goes first. The consequences are predictable. The chosen use case turns out to need data nobody has cleaned, or to involve decisions a regulator will question, or to cost more to run each month than it saves. Meanwhile a dull document-processing task with clean data and a clear owner, which would have paid for itself in a quarter, sits at the bottom of the list. Ranking by ROI does not mean picking the boring option every time; it means seeing the whole picture before choosing.
The five scores and what each one asks
| Dimension | The question | What a high score looks like |
|---|---|---|
| Business value | What changes if this works: hours returned, revenue protected, risk removed, and for whom? | A named owner can describe the change in their own numbers and will be measured on it |
| Data readiness | Does the data exist, is it accessible, is it clean enough, and are we allowed to use it? | Data is in a system with an API or export, an owner can vouch for its quality, and its use is permitted |
| Model risk | What happens when the model is wrong, and can a person catch it before harm? | Errors are visible, reversible and cheap; a human review step is natural |
| Build effort | How many integrations, how much new UI, how much process change? | One or two integrations, an existing interface, a process that mostly stays the same |
| Running cost | What does it cost per month to run at real volume, including inference, retrieval, monitoring and people? | Cost is predictable and clearly below the value it produces |
Each dimension is scored from one to five with a written reason. The reason matters more than the number; it is what makes the ranking defensible six months later when someone asks why the chatbot was not first.
An AI ROI framework that finance will accept
The business value score needs to survive a conversation with the finance team, so the value must be expressed in units they use: hours of a known role at a known cost, contacts avoided at a known cost per contact, revenue at risk from a known churn cause, or penalties avoided. Where the value is speculative, say so and score it lower rather than inventing a figure. Running cost is the other half of the equation and is the one most often left out; a use case that saves a team's time but costs more in inference and maintenance than that time is worth has negative ROI regardless of how impressive the demo was. Our AI total cost of ownership article lists what belongs in that line.
Value is per use case, not per technology
"Deploy a copilot" is not a use case. "Let field engineers ask the manual a question in their own words and get an answer with the page reference" is. Break each technology idea into the specific jobs it would do, and score the jobs. Several will score very differently from each other, and the ranking becomes much more useful.
AI use case selection: weighting and vetoes
The five scores can be summed with equal weight for a first pass, but two of them act as vetoes. A data readiness score of one means the use case cannot be built until the data work is done, whatever its value; it moves to a foundations list with the data task as the actual next step. A model risk score of one, meaning errors would be harmful and hard to catch, means the use case needs a human-in-the-loop design and a compliance conversation before it is scored again. After vetoes, the weights can shift to match the company's situation: a business under cost pressure weights running cost higher; one with a regulatory deadline weights model risk higher. Write the weights down and keep them the same across the portfolio.
Running the session in the open
The scoring session works best with the process owner, the data owner, someone from finance, someone from IT or security, and the executive sponsor, in the same room for an afternoon. Each use case gets ten minutes: the owner describes the job, the group scores each dimension aloud, disagreements are recorded rather than resolved on the spot, and the reasons are written on the sheet. The point of doing it in the open is that everyone sees why the ranking came out as it did, so the result is not relitigated in corridors. It also exposes the gaps: use cases nobody owns, data nobody can locate, and value nobody can describe. We run this session as part of the AI discovery sprint and also as a standalone half-day.
What a good first pick looks like
The best first use case scores at least four on data readiness and model risk, three or more on value and effort, and has a running cost that is obviously small relative to the value. It usually looks unglamorous: extracting fields from a document type the company handles daily, answering staff questions from a manual that already exists, drafting a routine reply that a person still sends, or classifying incoming requests so they route correctly. It has a named owner who wants it, and it can be measured in a quarter. Winning with that first pick earns the credibility to attempt the riskier, higher-value ones next, and the evaluation and platform work carries over.
A worked example
A field-service software company came with a list headed by an autonomous scheduling agent that would reassign jobs in real time. Scored in the open, it rated high on value but low on data readiness, because job history was inconsistent across customers, and low on model risk, because a bad reassignment was expensive and hard to reverse. Further down the list, an in-app assistant that let technicians ask about a product's documentation and their own past jobs scored well on every dimension: the documentation existed, errors were visible to the technician, one integration was needed, and the running cost was modest. It was built first, adopted quickly, and the data cleanup it required became the foundation for the scheduling work later. The build is described in our in-app copilot case study.
Team and timeline
A prioritisation exercise on its own is a half-day session plus a day of preparation and a written ranking, run by a strategist from our AI product strategy practice. Within a Sprint Zero of ten working days at $3,250 or ₹2,00,000, credited to the next build, the ranking is joined by a data readiness check, an architecture outline and a costed plan for the top use case. Your side provides the process owners, the data owner and a finance contact for the session. The top-ranked use case then typically proceeds to a ProofRun or straight to a Launch 6 build, both priced on the pricing page.
Before you start: a checklist
- A list of candidate use cases written as specific jobs, not technologies
- A named owner for each, who would be measured on the result
- A finance contact who can confirm the units of value
- A data owner who can say where each dataset lives and how clean it is
- Written weights for the five dimensions, agreed before scoring
- Vetoes agreed: what data readiness or model risk score stops a use case
- An afternoon with the right people in the room
- A place to record reasons, not just scores
Glossary
- Use case: a specific job the system would do for a specific person, with a measurable outcome
- Data readiness: whether the required data exists, is accessible, is clean enough and is permitted for the purpose
- Model risk: the cost and detectability of a wrong output
- Running cost: the monthly cost to operate at real volume, including people
- Veto dimension: a score that removes a use case from ranking until a precondition is met
- Foundations list: data, integration or policy work that must precede a use case
Related reading
AI readiness assessment: the ten questions before you build, building an AI roadmap that survives the first quarter, and why AI pilots never reach production. McKinsey's research on AI adoption is a useful external reference on where value has actually been captured.
Score every candidate on the same five dimensions, in the open, and let the ranking rather than the loudest voice decide what gets built first.
Frequently asked questions
What criteria should be used to prioritise AI use cases?
▾
Business value, data readiness, model risk, build effort and monthly running cost, each scored one to five with a written reason. Data readiness and model risk act as vetoes: a low score there sends the use case to a foundations list.
How long does AI use case prioritisation take?
▾
A day of preparation and an afternoon session with the process owners, data owner, finance and IT in the room. Within a ten-day discovery sprint it is combined with a readiness check and a costed plan.
Should the first AI use case be the highest-value one?
▾
Not necessarily. The first should score well on every dimension, especially data readiness and model risk, so it succeeds quickly and builds the platform and credibility for higher-value, riskier use cases next.