Red flags when hiring an AI development partner
What are the red flags to watch for when hiring an AI development partner?
Vendor demos on curated data, no eval suite, single-model lock-in, no running-cost estimate and no security review are the warning signs. Each predicts a failure after launch: accuracy that collapses on real data, regressions nobody catches, a provider change that breaks the product, a surprise bill or audit findings.
AI vendor red flags are easy to spot once you know what each one predicts. A demo on the vendor's own curated data predicts accuracy that collapses on yours. No evaluation suite predicts regressions nobody catches. A single-model architecture predicts a broken product when the provider changes something. No running-cost estimate predicts a surprise bill. No security review predicts an audit finding. This article walks through the warning signs we see most often when clients bring us projects a previous vendor started, what each one leads to, and how to check for it before you sign.
Why AI outsourcing risks are different
Conventional software either works or does not, and you can usually tell in acceptance testing. AI systems work partially and probabilistically, which means a vendor can deliver something that looks finished and passes a demo while failing a third of real cases. The gap is invisible unless someone measures it, and a vendor who has not built the measurement has no incentive to. The bad AI agency signs below are all variations on one theme: the vendor is selling a demo and calling it a product. The interview questions that surface these are in questions to ask before hiring an AI agency.
The red flags and what they predict
| Red flag | What it predicts | How to check |
|---|---|---|
| Demo on curated data | Accuracy collapses on real inputs | Bring your own messy data to the demo and watch |
| No evaluation suite | Silent regressions after every change | Ask to see an eval report from a past project |
| Single-model lock-in | A provider update or price change breaks the product or the budget | Ask where the model switch lives in the code |
| No running-cost estimate | A surprise inference bill in month two | Ask for cost per request and the method behind it |
| No security review | Data exposure, prompt injection, audit findings | Ask for the review checklist and where data flows |
| Big-bang go-live | Customer-facing errors on day one | Ask about shadow mode and policy gates |
| Vendor keeps prompts or platform | Lock-in; you cannot leave or maintain it | Ask for the handover list in writing |
| Accuracy promised before measurement | A padded price or a dispute at acceptance | Ask how the number was arrived at |
Red flag: the demo is on their data
Every AI demo works on the data it was built for. The only demo that tells you anything runs on a sample of your documents, tickets or calls, chosen by you, including the ugly ones. A vendor who resists that, or who asks for time to "prepare" your sample, is telling you the system is tuned to the demo. Insist on a live run with your data, and count the failures yourself. The difference between a demo and a proof of concept is spelled out in AI proof of concept vs demo.
Red flag: there is no evaluation suite
Ask any vendor how they will know the system got worse after a change. If the answer involves a person trying a few questions, there is no eval suite. A real one is a graded set of real examples, a harness that runs them automatically, and a report by category with history. It should exist by the second week of a build and be handed over at the end. Without it, every prompt tweak, model update and content change is a gamble, and the person who discovers the loss is a customer.
Red flag: one model, hard-wired
A system that calls one provider's model directly from every part of the code is a system that breaks when that provider retires the model, changes its behaviour or doubles the price. Ask to see the routing layer: the place where each task is mapped to a model and provider by configuration. Ask what happens if you want to move a task to a cheaper model or a self-hosted one. A vendor who says a single model is "the best" for everything has not measured, and a benchmark-first method is described in model and vendor selection. Where data residency matters, ask whether the same architecture can run on private, self-hosted infrastructure.
Red flag: no running-cost estimate
Build price is only part of what you are buying. A vendor who cannot estimate cost per request, per conversation or per user has not traced a real request through the system and does not know how many model calls it makes. Ask for the estimate, the method and the controls: daily budgets, per-tenant caps, output limits, caching. A good vendor will refine the estimate in shadow mode and hand you a measured number before launch. Our build and care prices are on the pricing page, and every quote with a model in it carries a running-cost line.
Red flag: security is a paragraph in the proposal
AI systems have their own attack surface: prompt injection through documents or messages, data leaking across tenants through a shared index, tool calls that can be coerced into unintended actions, and secrets in prompts. The OWASP Top 10 for LLM applications lists the common ones. Ask the vendor which of those they test for and how, where data is stored and processed, which providers see it and under what terms, and whether a security review is part of the scope or an extra. Our practices are on the security page; ask every vendor for the equivalent.
Red flag: launch day is autonomy day
An agent that takes actions, whether refunds, bookings or record updates, should run in shadow mode first, drafting while a human approves, and then go live with policy gates that limit what it may do without a person. A vendor who plans a big-bang launch has not run an agent in production. Ask for the go-live plan and the list of actions that stay human regardless of accuracy.
Red flag: you will not own it
Read the handover terms. If prompts are proprietary, if the system runs on the vendor's platform, if the evaluation set stays with them, or if infrastructure sits in their accounts, you are renting. That is a legitimate business model for a product company; it is a red flag for a development partner, because it means you cannot maintain, move or audit what you paid for. Our position is that clients own all code, prompts, models, infrastructure and documentation, and the contract should say so in plain words.
A worked example
A field-service SaaS company came to us with a copilot a previous vendor had built. It demonstrated well on the vendor's sample manuals and failed on the client's real ones, which were scanned and inconsistently formatted. There was no evaluation set, so nobody could say how often it was wrong. The model was hard-coded to one provider, and a price change had already pushed the running cost past what the feature earned. The rebuild started with an evaluation set graded by the client's support engineers, a retrieval layer designed for the real documents, routing across providers by task, tracing for cost, and shadow mode before release. The in-app copilot case study describes the result. Every red flag above had been present in the first build.
Softer warning signs
- The proposal has no questions for you; a good vendor asks about your data before quoting
- Everything is possible and nothing is hard
- The team in the pitch is not the team in the contract
- No mention of what will stay human or what the system will refuse to do
- Timelines with no discovery stage
- References are all recent, and none has a system still running
- The price is a single number with no scope behind it
Team and timeline
Checking for these red flags takes a week: two meetings per vendor, one reference call each, and a live run on your data. If you are already in a project showing these signs, a Sprint Zero discovery at $3,250 / ₹2,00,000 over ten working days assesses what can be kept, builds the missing evaluation set and produces a fixed-price plan to finish or rebuild; the fee is credited to the work that follows. Every build we deliver includes evals, routing, cost tracing, a security review and shadow mode as standard, with a Care Plan from $1,000 a month after launch. To talk through a project, contact us with what you have.
Before you start: a checklist
- Bring your own sample data to every demo and count the failures
- Ask for an anonymised evaluation report from a previous project
- Ask where the model routing lives and whether a provider can be swapped
- Get a running-cost estimate and the method behind it
- Ask for the security review checklist and the data-flow description
- Confirm shadow mode and policy gates are in the go-live plan
- Read the handover list and the platform terms before signing
- Insist on a discovery stage before any fixed build price
Glossary
- Curated demo: a demonstration on data chosen to make the system look good
- Evaluation suite: graded real examples plus a harness that runs them on every change
- Model lock-in: a system that depends on one provider's model with no way to switch
- Prompt injection: malicious instructions hidden in inputs to redirect the model
- Shadow mode: running on real traffic without acting, to measure before autonomy
- Policy gate: rules that limit which actions the system may take without a person
- Handover list: everything transferred to you at the end of a build
Related reading
See questions to ask before hiring an AI agency, how to choose an AI development company and AI proof of concept vs demo. The OWASP Top 10 for LLM applications is the reference for the security questions.
Each red flag predicts a specific failure after launch; check for all five before you sign, and walk away from any vendor who cannot show you their evals.
Frequently asked questions
What is the biggest red flag when hiring an AI vendor?
▾
No evaluation suite. Without graded real examples run on every change, nobody knows the system's accuracy or when it degrades, and every other problem stays hidden until customers find it.
Is single-model lock-in really a problem?
▾
Yes. Providers retire models, change behaviour and change prices. A routing layer that lets each task move between providers by configuration protects both quality and budget.
Can a project started by a bad vendor be rescued?
▾
Often. A discovery sprint assesses what can be kept, builds the missing evaluation set and produces a fixed-price plan to finish or rebuild. Prices are on the pricing page.