Reading an AI case study: what to look for and what to ignore
How should you evaluate an AI vendor's case study?
Look for the measurement method, the control, the ownership of the system and whether it still runs; ignore round-number uplift claims. A case study that names its metric, its baseline and its failure modes tells you how the vendor works; one that leads with a percentage tells you how the vendor sells.
AI case study evaluation comes down to four questions. How was the result measured? What was it compared against? Who owns the system now? And is it still running? A case study that answers those tells you how the vendor works. One that leads with a round-number uplift and skips the rest tells you how the vendor sells. This article explains what each question exposes, the patterns that should raise doubt, and how to use case studies as one input into vendor due diligence rather than as proof.
We write case studies ourselves, so this is partly a description of the standard we try to hold. Our own are anonymised and qualitative, which has its own limits; the section on what to ignore applies to them as much as to anyone's.
What a credible AI case study contains
| Element | What to look for | What its absence suggests |
|---|---|---|
| The problem | A specific workflow, user and failure mode before the system existed | Generic problem statement written to fit the solution |
| The metric | Named, defined, and the one the client's business actually uses | Vanity metric chosen because it moved |
| The baseline | What the number was before, and how it was measured | The uplift is unanchored and cannot be checked |
| The control | A holdout, an A/B test or a shadow-mode comparison | The change may be seasonal, coincidental or self-reported |
| What did not work | A failure mode, a deferred scope item, a limitation | Nothing went wrong, which never happens |
| Ownership | Client owns code, prompts, models and infrastructure | Client is renting the system from the vendor |
| Status | Still running; how long; who maintains it | A pilot that was switched off after the write-up |
| Timeline and team | Weeks, people, client effort | Effort hidden, so the price cannot be inferred |
Look for the measurement method
A number without a method is a claim, not a result. "Resolution rate improved" means something only if you know how resolution was defined: closed with no reopen within a week? Closed by the bot? Closed by anyone? The definition changes the number by a wide margin, and a vendor who chose the metric after seeing the data will have chosen the one that looked best.
Look for the metric to be the one the client's business already used, defined before the build, and measured by the client's own systems rather than the vendor's dashboard. Why ticket deflection is the wrong metric is a worked example of how a plausible-sounding metric can hide a poor outcome.
Look for the control
AI systems are usually launched during other changes: a new pricing plan, a busy season, a hiring round. Without a control, any improvement may belong to something else. The credible controls are a holdout group that did not get the system, an A/B test, or a shadow-mode period where the system's outputs were compared against what humans did. If a case study reports a before-and-after with no control, treat the number as suggestive at best. Personalisation lift needs a controlled test makes the argument for one common category of claim.
Time windows matter too
A result measured over two weeks after launch is a launch effect, not an outcome. Ask over what period the number was measured and whether it held. A case study written a year after go-live is worth more than one written the week of the launch.
Look for ownership and status
Two questions that case studies rarely answer unprompted: who owns the system, and is it still running? Ownership tells you whether the client could leave the vendor without losing the system. Status tells you whether the result was a pilot that made a good story or a production system that a team maintains. A vendor with confidence in their work will answer both; ask them directly. If the answer to "is it still running?" is vague, the case study is describing a demo.
Related to status: who maintains it, and under what arrangement? A system on a care plan with named response times is a different thing from a system that was handed over and never touched. The contract terms that matter cover what ownership should include.
What to ignore in AI success stories
- Round numbers: "40% faster", "3x productivity" and similar are almost always rounded from something less impressive or measured loosely
- Uplift without a baseline: an improvement from an unknown starting point is unverifiable
- Vanity metrics: messages sent, queries answered, documents processed; volume is not value
- Named logos with no detail: the client's brand is doing the work the evidence should do
- Timelines that skip effort: "deployed in two weeks" often means the vendor's two weeks, after months of client preparation
- Quotes that could apply to any project: "the team was responsive and professional" says nothing about the system
- Screenshots of chat transcripts: a good conversation proves the demo worked once
None of these prove a case study is false. They mean it is not evidence, and you should ask for the evidence separately.
How to verify AI claims in a conversation
Case studies are a starting point for questions, not an end. In the vendor conversation, pick one case study and ask: how was the metric defined, what was the control, what went wrong, who owns it, is it still running, and can we speak to the client? The quality of the answers will tell you more than the document did. A vendor who built the system will answer with specifics, including the parts that did not work. A vendor whose case study was written by marketing will struggle with the second question.
Then ask to see the evaluation approach for your own project, because that is what will determine your result. The case study tells you what happened to someone else; the evals tell you what will happen to you. Our 12-point checklist for choosing an AI development company puts case studies in their place among the other checks.
Why anonymised case studies are still useful
Many clients, especially in regulated sectors, will not be named. Anonymised case studies lose the reassurance of a logo but can be more honest, because the vendor is not constrained by the client's marketing team. Judge them by the same standard: the specificity of the problem, the metric, the control and the failure modes. An anonymised case study that names its limitations is worth more than a named one that does not.
A worked example
A university that had a legacy ERP modernised without a rewrite is one of our own case studies, and it illustrates the standard. The credible version says what the system was, why a rewrite was rejected, how characterisation tests were used as the safety net, which modules were delivered in which phases, what was deferred, and that the institution owns the result and runs it. It does not give a percentage cost saving, because the counterfactual rewrite never happened and any number would be invented. The legacy ERP modernisation case study is written that way deliberately. A reader who wants a percentage should ask instead for the phased plan and the test approach, because those are what they will get on their own project.
Team and timeline
Case-study review is part of the diligence you do before a first program, and it costs nothing but an hour with the vendor's engineers. The lowest-risk way to test what you read is Sprint Zero: ten working days at $3,250 or ₹2,00,000, credited to the build, during which the vendor's actual team works on your actual data and produces a written feasibility answer and an evaluation plan. If the case study described how they work, the discovery will confirm it. Program and care plan prices are on the pricing page, the AI strategy service covers discovery in more depth, and you can contact us to ask any of the questions in this article about our own case studies.
Before you start: a checklist
- For each case study, write down the metric, baseline, control and time window; note any that are missing
- Ask whether the system is still running and who maintains it
- Ask who owns the code, prompts, models and infrastructure
- Ask what went wrong or was deferred; a blank answer is a warning
- Request a reference call with the client where the vendor permits
- Compare the case study's problem with yours: same data shape, same actions, same risk?
- Ask for the evaluation approach for your own project, in writing
- Weight case studies below engineering practice in your overall scoring
Questions clients ask
- Should we only trust named case studies? No. Judge by specificity and method. Anonymised studies in regulated sectors are often more candid.
- Is one strong case study enough? It shows the vendor has done the work once. Ask how many times they have run the same shape of program.
- What if the vendor cannot give any numbers? That is acceptable if they explain why and give the method they would use on your project. It is not acceptable if they cannot describe the metric at all.
- Can we ask for the evaluation set from a case study? It belongs to the client, so usually no, but the vendor should describe how it was built.
Related reading
See questions to ask before hiring an AI agency, how to measure an AI agent: the six metrics that matter and our work page for case studies written to the standard above. For the general principles of running a controlled comparison before claiming an effect, the Nielsen Norman Group's articles on quantitative research are a clear, non-vendor reference.
Read a case study for its method, not its number; the method is the part you will actually get.
Frequently asked questions
What should an AI case study include to be credible?
▾
A specific problem, a named metric defined before the build, a baseline, a control such as a holdout or shadow-mode comparison, what did not work, who owns the system and whether it is still running.
Why should I ignore percentage uplift claims?
▾
Because without the metric definition, baseline, control and time window, a percentage cannot be checked. Round numbers are usually rounded from looser measurements, and launch-week results rarely hold.
How do I verify a vendor's case study?
▾
Ask the engineers who built it how the metric was defined, what the control was, what went wrong, who owns it and whether it still runs. Request a reference call, then ask for the evaluation approach they would use on your project.