azyware
Business

Questions to ask before hiring an AI agency

EZ
Eazyware
· 7 min read
Quick answer

What questions should you ask an AI agency before hiring them?

Ask how they measure accuracy, who owns the code, what happens when a model changes, and to see a production system they still run. The answers separate teams that build evaluated, owned, maintained systems from teams that build demos. Ask the same questions of every agency and compare the answers side by side.

The questions to ask an AI agency are the ones that reveal whether they have run something in production and lived with it: how do you measure accuracy, who owns the code and prompts, what happens when the model provider changes something, and can we see a system you built that is still running. Four questions, and most agencies fail at least one. This article gives the full list we suggest buyers use in an AI vendor interview, explains what a good answer sounds like, and shows how to compare agencies fairly.

Why the usual vendor questions are not enough

Buyers know how to evaluate a software vendor: references, portfolio, team, price. AI adds properties that those questions miss. A feature can look finished and be wrong a third of the time. A system can work on launch day and degrade quietly when a provider updates a model. A build can be delivered with prompts the client cannot read and evals that do not exist. AI partner due diligence has to probe measurement, ownership and operation, not just delivery. Our 12-point checklist for choosing an AI development company covers the general criteria; this article is the interview script.

The questions, and what good answers sound like

QuestionGood answerWeak answer
How do you measure accuracy?A graded test set from your data, run on every change, with per-intent or per-category results"We test thoroughly"; a demo on their sample data
Who owns the code, prompts and evals?You do, on handover, with documentation and infrastructure accessLicensed platform; prompts are proprietary
What happens when a model changes?Evals re-run, routing adjusted, rollback if needed, under the care plan"We use the latest model"; no process
Show us a system you still runA live production system with metrics and the client's permission to discuss itSlides; a demo built for the meeting
What will it cost to run?A per-request estimate from traced calls, refined in shadow modeNo estimate, or "it depends" with no method
How do you go live?Shadow mode, then policy-gated actions, then expanded autonomyBig-bang launch
Which models do you use?Chosen per task by benchmark on your data; routing across providersOne provider, always
How do you handle our data?Named storage locations, access controls, a security review, contractual termsVague assurances

On measurement

Ask to see an evaluation report from a past project with the client's details removed. It should show a test set built from real data, a metric that matches the business outcome, results broken down by category, and a history across versions. Ask who graded the set and how disagreements were resolved. Ask what accuracy was on the first run, because a team that only quotes the final number has not been honest about the journey. The approach we use is described in model and vendor selection: a benchmark-first approach.

On ownership

Ask for the handover list. It should include source code, infrastructure as code, prompts, the evaluation set and harness, model configuration and routing, documentation, and credentials transferred to your accounts. Ask whether anything runs on the agency's platform after handover and what happens if you stop paying. A system you cannot move is not yours, whatever the contract says about intellectual property.

On model change

Providers retire models, change behaviour between versions and change prices. Ask the agency to describe the last time this happened to a system they run and what they did. A good answer has specifics: the model that changed, the eval that caught it, the fix. A weak answer treats model updates as free improvements. Ask whether the system is model-agnostic, meaning a different provider could be swapped in per task, and ask them to show where in the code that switch lives.

On production experience

Ask to see a running system, not a recording. Ask what its metrics were last month, what broke in the last quarter, and what the care plan for it covers. Ask to speak to the client who runs it. Agencies that have operated a system for a year talk differently about it than agencies that shipped and left; they talk about escalation reviews, cost per resolution and the intent that never quite worked. Our case studies are written that way for that reason.

On cost and contract

Ask for a fixed price and a fixed date for a scoped build, and for the method they use to arrive at the scope. Ask what a discovery sprint costs and whether it is credited to the build. Ask for the running-cost estimate and the care plan price alongside the build price, because the total is what you are buying. Our numbers are on the pricing page; ask every agency to put theirs beside them.

On security and data

Ask where data will be stored and processed, which models will see it and under what terms, who at the agency has access, and whether a security review is part of the build. Ask about applicable regulation for your sector and whether they have worked under it. The NIST AI Risk Management Framework is a useful common vocabulary for this conversation, and our own practices are on the security page.

On the team you will actually get

Ask who will do the work by name, what they have shipped, and how much of their time you get. Agencies often pitch with seniors and deliver with juniors; ask whether the people in the room will be the people on the project and get it into the contract. Ask how decisions will be recorded, how often you will see working software rather than slides, and who on your side they expect to have available. A good agency will tell you what they need from you, including a product owner who answers within a day and a subject-matter expert to grade the evaluation set, because they know a build without those people fails regardless of the engineering.

A worked example

A non-banking financial company interviewing agencies for a KYC document intelligence build asked each the same eight questions. One agency demonstrated a polished extraction demo on its own sample documents and could not say how accuracy would be measured on the client's photographed proofs. Another proposed a licensed platform where prompts and models stayed with the vendor. The agency selected proposed a discovery sprint to inventory real document types, an evaluation set graded by the client's operations team, a review queue for low-confidence cases, and full handover of code and evals. The build was fixed price against a measured target and the client runs it under a care plan. The KYC document intelligence case study describes the outcome.

How to run the comparison

  • Send every agency the same one-page brief and the same question list
  • Score each answer on the good and weak column above, with notes
  • Weight measurement, ownership and production experience highest
  • Ask for one reference call per agency with a client whose system is still running
  • Compare total cost: discovery, build, running cost and care, over two years
  • Prefer the agency that told you something you did not want to hear

Team and timeline

A proper AI vendor interview takes two meetings per agency and a week of reference calls; three agencies is enough. The cheapest way to evaluate AI consultants is to buy a discovery sprint from your shortlisted choice: our Sprint Zero costs $3,250 / ₹2,00,000 over ten working days, produces a scope, an evaluation plan and a fixed quote, and is credited to the build. If the sprint output is weak, you have spent little and learned a lot. To start that conversation, contact us with the one-page brief you sent everyone else.

Before you start: a checklist

  • Write a one-page brief with the outcome, the data, the systems and the users
  • Prepare the question list and send it in advance
  • Ask each agency for an anonymised evaluation report from a past project
  • Ask for the handover list in writing
  • Request one reference with a system still in production
  • Ask for build, running and care costs together
  • Check security and data-handling answers against your own requirements
  • Decide who on your side will grade the evaluation set

Glossary

  • Evaluation set: real examples with graded correct outputs, used to test every change
  • Shadow mode: the system runs on real traffic without acting, so accuracy is measured before autonomy
  • Model-agnostic: built so the model behind each task can be swapped by configuration
  • Policy gate: rules limiting what an AI system may do without a human
  • Handover list: everything transferred to you at the end of a build
  • Care plan: a monthly retainer for monitoring, fixes and evaluation re-runs
  • Discovery sprint: a short fixed-price engagement that produces the scope and quote

See red flags when hiring an AI development partner, how to choose an AI development company and AI proof of concept vs demo. The NIST AI RMF is a neutral reference for the governance questions.

Ask about measurement, ownership, model change and a system still running; the agency that answers all four with specifics is the one to hire.

Frequently asked questions

What is the single most important question to ask an AI agency?

▾

How do you measure accuracy on our data? A graded evaluation set run on every change is the difference between engineering and demos. Everything else follows from that answer.

Should we ask for a free proof of concept?

▾

No. A free POC is a demo on curated data. Pay for a short discovery or POC sprint that measures accuracy on your data and produces a fixed quote; ours is credited to the build. See the pricing page.

How many agencies should we interview?

▾

Three, with the same brief and the same questions. Score answers side by side, take one reference call each, and choose on measurement, ownership and production experience rather than price alone.