Signs your AI vendor is out of their depth
How do you tell if your AI vendor is out of their depth?
Your AI vendor is out of their depth when demos replace measurement, when nobody can show you a request trace, when every quality problem is answered with a prompt edit, and when there is no plan for a wrong action. These signs appear mid-build, not at pitch.
Your AI vendor is out of their depth when they show demos instead of evaluation scores, cannot produce a trace of a single production request, answer every quality complaint by editing the prompt, and have no plan for what happens when the system acts wrongly. Those four AI vendor red flags appear weeks after signing, not during the pitch.
Pitch-stage warning signs are covered elsewhere. This piece is for the harder position: you have signed, money is moving, and something feels wrong but you cannot name it. Each sign below comes with what it usually means, the specific question that confirms it, and what to do about it that is short of cancelling.
Why capable-looking vendors run out of road
Building something that works once is genuinely easy now. A competent web team can wire a model to a document store in a fortnight and show you a convincing conversation. The skills that separate a working demo from a production system are different skills: evaluation design, retrieval quality, permission modelling, cost control, failure handling and the patience to run something in shadow mode while it is embarrassing.
Most vendors who get into trouble are not dishonest. They are a good software team who priced an AI project as if it were a web project and discovered in week five that the last stretch of accuracy is where nearly all the work lives. The signs you are looking for are signs of that discovery, and they surface on a predictable schedule.
There is a second, less comfortable possibility: the brief was wrong, the data is worse than anyone admitted, or the success criterion was never agreed. Separating vendor incapacity from client ambiguity is most of the diagnosis, which is why every sign below has a confirming question rather than a verdict.
The signs, and when they appear
Read the table as a timeline. Early signs are recoverable with a conversation; late signs usually mean the engagement needs restructuring.
| Roughly when | What you see | What it usually means | The question that confirms it |
|---|---|---|---|
| Weeks 1 to 2 | No written success criterion | Nobody agreed what finished looks like | What number, measured how, means we are done? |
| Weeks 2 to 3 | Progress shown only as live demos | There is no evaluation set behind the demo | Show me the scores on the same fifty cases from last week |
| Weeks 3 to 4 | Answers improve after each meeting | The prompt is being tuned to the questions you ask | Which cases regressed when that change landed? |
| Weeks 4 to 6 | No request trace available | No observability, so no debugging and no cost visibility | Pull up yesterday's worst answer and show me every step |
| Weeks 5 to 8 | Retrieval problems described as model problems | The team does not understand where the error came from | Was the right document retrieved, yes or no? |
| Weeks 6 to 10 | No answer on what happens after a wrong action | Permissions and approval gates were never designed | What is the maximum damage one bad call can do? |
| Any time | Cost per request unknown | Nobody instrumented the spend | What did last week cost, split by intent? |
| Any time | Reluctance about code and prompt handover | Lock-in is part of the commercial model | Can we take a full export next Friday? |
The four that matter most
Demos where measurement should be
A demo proves the system can be right once. An evaluation suite proves how often it is right across cases it has not been tuned on. If, six weeks in, your progress reports are still screen recordings, the team has no way to know whether last week's change helped or hurt, which means neither do you. Our position on this is set out in Evals over demos, and it is the single question we would ask if we could only ask one.
The prompt as the only tool
Prompt editing is real engineering, but it is one tool among many. When retrieval is returning the wrong documents, no prompt fixes it. When a schema is being violated, structured output constraints fix it, not a plea in the system message. A team that responds to every failure with a prompt change is treating symptoms, and the giveaway is that old problems return. Ask which cases regressed; if there is no answer, prompts are not versioned or tested.
No trace, no truth
You should be able to pick any request and see the retrieved chunks, the exact prompt sent, the model and version, the tool calls, the tokens and the cost. Without that, nobody can tell you why an answer was wrong, and every incident becomes a guess. Tracing is not advanced practice; it is the baseline described in LLM observability.
No answer for the bad action
If the system can do anything more than answer questions, it can do the wrong thing to a real customer. A team in control of this can tell you the blast radius of each tool, which actions are gated for human approval, and how the system runs in shadow mode before it acts alone. A team that has not thought about it will talk about how accurate the model is, which is not the same subject. Prompt injection and insecure output handling are documented failure classes in the OWASP Top 10 for LLM applications, and a vendor who has never read it is telling you something.
How to run the conversation without blowing up the project
Most of these situations are recoverable, and replacing a vendor mid-build is expensive in time as well as money. Work through this sequence before you escalate.
- Ask for artefacts, not reassurance. Request the eval set, last week's scores, a trace export and the cost report. Three working days is a fair deadline for things that should already exist.
- Separate your ambiguity from theirs. If nobody wrote down the success criterion, that is shared. Write it now, in numbers, and reset from there.
- Narrow the scope rather than extending the date. One intent working at production quality is worth more than five at demo quality, and it is the honest test of whether the team can finish anything.
- Insist on a full export this month. Code, prompts, eval data, infrastructure definitions and documentation. If that is difficult, you have learned the most important thing.
- Bring in a second opinion for a fixed, small engagement. A technical review of the architecture, evals and traces costs far less than a restart.
- Agree a decision date. Name the day you will decide whether to continue, and what evidence you need by then.
Ownership is the part to be firm about. At Eazyware the client owns the code, prompts, infrastructure and model choices from the first commit, for exactly this reason; the argument is in You own everything. A contract that makes leaving painful was designed by someone who expected you to want to.
When the vendor is not the problem
Three situations look identical from the outside and are not the vendor's fault.
The data was not what you described. If the knowledge base is out of date, contradictory or locked in scanned PDFs nobody told them about, no team ships good answers on top of it. The right response is a data remediation workstream, not a new supplier.
The success criterion keeps moving. If the definition of a good answer has changed three times, the team is chasing a target you are holding. Freeze it in writing for eight weeks and judge them against that.
The use case was wrong. Some workflows are not model problems at all; they need a rules engine, better search, or a fixed process. A vendor who tells you this early is demonstrating depth, not avoiding work, and a second opinion should be willing to say the same.
What good looks like, concretely
A team that is not out of their depth will, without being asked, show you an evaluation report with scores per intent and a note on what regressed, a trace for the worst answer of the week, a cost figure per completed task, a list of actions that are gated and why, and a shadow-mode acceptance rate that is going up. They will also tell you when something is not working, before you notice. The first thirty days of a healthy engagement are described in what to expect in the first 30 days.
If you are weighing a supplier against building the capability yourself, the trade-offs are laid out in Eazyware versus an in-house AI team, and the broader options in in-house team vs agency vs freelancers. Neither route is inherently safer; both fail the same way when nobody is measuring.
If you decide to restart
Restarting is cheaper than most people fear when the scope is cut properly first. A ten-day Sprint Zero at $3,250 or ₹2,00,000, credited to the build, establishes the success criterion, the intent list and the eval plan. A three-week ProofRun at $6,250 or ₹4,00,000 proves the hardest intent end to end before anyone commits to a full programme. Published starting prices for every service are on the pricing page, and we will say on a first call if we think the honest answer is that you do not need us. The one thing not to do is re-award the same unmeasurable brief to a different supplier.
Related reading
Red flags when hiring an AI development partner covers the signals you can see before signing, and Questions to ask before hiring an AI agency turns them into a shortlist conversation. If you are at the point of a formal reset, start with what a fixed-price AI quote should contain.
Depth shows up as evidence produced without being asked; its absence is the only red flag you really need to watch.
Frequently asked questions
What is the clearest sign an AI vendor is struggling?
▾
They cannot show you scores on a fixed set of test cases from one week to the next. Demos prove a system can be right once; evaluation scores prove how often it is right on cases it was not tuned against. A team without that measurement cannot tell whether their own changes helped.
Should we replace an AI vendor mid-project?
▾
Usually not before you have asked for artefacts and narrowed the scope. Request the eval set, a request trace and a cost report, freeze the success criterion in writing, and set a decision date. Replacing a vendor costs months, and some of these situations turn out to be brief problems rather than capability problems.
What should we own at the end of an AI engagement?
▾
Source code, prompts and prompt history, evaluation datasets and results, infrastructure definitions, model choices and documentation, all in your own repositories and cloud accounts. If a full export next Friday would be difficult for your supplier, that difficulty is deliberate and worth resolving in the contract now.