azyware
Business

Questions to ask an AI customer service agent vendor before you sign

EZ
Eazyware
· 7 min read
Quick answer

What should you ask an AI customer service agent vendor?

Ask for the evaluation suite first. A vendor who has shipped will show a golden question set built from real conversations, name the metrics they report, state who owns the code and prompts, explain how running costs are billed, and describe what happens when a model is deprecated.

Ask for the evaluation suite first. A vendor who has shipped will show you a golden question set built from real conversations, name the metrics they report, state who owns the code and prompts afterwards, explain how running costs are billed, and describe exactly what happens when the model underneath the agent is deprecated.

What follows is a question bank you can take into a vendor meeting, grouped by what each group is actually testing. For every question there is a version of the answer that means the vendor has done this before and a version that means you will be their learning experience. Use the second column as a filter, not a script.

What you are really testing

Most AI customer service agent vendor conversations are demos, and a demo tells you almost nothing. Any competent team can produce twenty impressive exchanges. The questions worth asking are the ones a demo cannot answer: how do you know it is right, what happens when it is wrong, who owns it afterwards, and what does month thirteen look like.

There is a reason evaluation is the first filter. An evaluation suite is a set of scenarios with known correct outcomes, run automatically on every prompt, retrieval or model change, producing a score you can compare over time. A vendor without one cannot tell you whether their last change helped, which means they cannot safely change anything once you are live. Our position on this is set out in evals over demos.

The second reason to lead with evidence is commercial. Vendors who cannot measure quality compete on demo polish and price, and both of those are cheap to fake.

The question bank

Ask thisA vendor who has shipped saysWalk away if you hear
Show me an eval suite from a past buildScenarios from real logs, scored per intent, run on every changeWe test manually, our QA team checks samples
What will you report in month three?Auto-resolution rate, reopen rate, cost per resolved conversationContainment and deflection percentage
Who owns the code, prompts and eval set?You do, transferred with documentation at handoverThe prompts are our intellectual property
How is model usage billed?Through your own provider accounts, with budgets and dashboardsIt is bundled, do not worry about it
What happens when a model is deprecated?We re-run the evals on candidates and give you a comparisonWe upgrade automatically to the latest model
How do escalations reach a human?With full context, transcript and proposed action attachedIt hands over to your existing queue
What does shadow mode look like here?Three to four weeks, acceptance rate per intent, written exit criteriaWe can go live immediately
What happens if we leave after a year?Repository, prompts, indexes, evals and runbooks are already yoursWe will scope a transition project

Questions about evidence

Ask to see the scenarios, not the demo

Request a redacted evaluation set from a previous engagement and look at three things: whether the scenarios came from real conversations, whether they include the awkward cases rather than only the clean ones, and whether results are broken down per intent. Averages hide the intent that is quietly wrong most of the time. If the vendor cannot show a redacted set for confidentiality reasons, ask them to build twenty scenarios from your own logs during the sales process and score them in front of you.

Ask which metric goes on the board slide

This question separates the field faster than anything technical. If the answer is containment or deflection, the incentive is to stop customers reaching a human, and you will be paying for a queue that empties into frustration. The metrics that mean something are auto-resolution rate, reopen rate within seven days, escalation quality and cost per resolved conversation. The argument in full is in AI ticket deflection is the wrong metric.

Questions about ownership and exit

Ask who owns the repository, the prompts, the retrieval index, the evaluation set and the deployment configuration once the invoice is paid. Ask it as five separate items, because a vendor can say "you own the code" while keeping the prompts and the eval set that make the code useful. Ask where the system runs and whether you hold the credentials for your own provider accounts.

Then ask the exit question directly: if we end this engagement in month thirteen, what do we have and what do we lose? The answer should be nothing beyond the vendor's attention. Our stance is that you own the code, infrastructure, prompts, model choices and documentation, which we explain in you own everything. The contract language that makes it real is covered in who owns the code, prompts and models.

Questions about money

Ask for the price as a fixed figure with a written scope, then ask what is outside it. There are only a few honest answers: model usage, third-party services, your team's reviewer time and the post-launch support arrangement. A vendor who cannot separate build from run has not run one.

On usage, ask who holds the provider account. You should, so you see the bill and can move providers. Ask what the vendor does to control it: model routing by intent complexity, caching of repeated context, prompt length discipline and a monthly budget with alerting. Ask them to size your monthly cost per resolved conversation rather than per message; if they cannot, they have not thought about the architecture carefully. Ask what support costs after launch, because what a care plan should cost is a question with a real answer.

Questions about life after launch

Ask what happens in month six when your refund policy changes. The answer should involve a prompt change under version control, updated eval scenarios and a regression run, not a support ticket that disappears. Ask who reads the escalation log weekly and whether knowledge gaps found by the agent feed back into your help centre, because an agent that silently fails on the same question for six months is a system nobody owns.

Ask about security and data specifically: where customer data is processed, what is redacted before it reaches a model, how long transcripts are retained, and which sub-processors are involved. In India these answers have to satisfy the DPDP Act 2023, and in regulated sectors they have to satisfy your auditor. A security questionnaire for AI vendors is the list we are asked to complete most often, and a vendor who has shipped will have answers ready rather than promises.

Eight questions to take into the meeting

  • Show me an evaluation suite from a build you have shipped. Per intent, from real logs, run on every change.
  • Which metric will you put in front of our board? Resolution and reopen rate, never deflection.
  • Who owns the prompts and the eval set? Ask about each artefact separately.
  • Whose provider account pays for inference? It should be yours, with a budget and a dashboard.
  • What is your process when a model is deprecated? Evals on candidates, then a recommendation.
  • How long is shadow mode and what ends it? Written exit criteria, not a feeling.
  • What does an escalation carry with it? Transcript, account context and the proposed action.
  • What do we keep if we stop working with you? Everything, or the price was not what you thought.

Questions that sound rigorous but tell you nothing

Some due-diligence questions have become rituals. "Which model do you use?" invites a brand answer when the useful version is "how do you choose, and how would you know if a different one were better?" Most production systems route between two or three models anyway, and we pick whichever benchmarks best for your task across OpenAI, Anthropic, Google, Meta, Mistral and open-weight options.

"How accurate is it?" is similarly hollow without a defined test set; any number offered in response to it is marketing. "Do you have experience in our industry?" matters less than whether they can build a golden set from your conversations in a week. And a request for a free proof of concept usually costs you more than it saves, because unpaid work is scoped to impress rather than to inform, which is why we run a paid three-week ProofRun instead.

How we answer these questions

For the record, so you can compare. An AI customer service agent build is $12,500 to $42,000, or ₹8 lakh to ₹28 lakh, fixed price with a written scope, including the evaluation suite and shadow mode. You pay model usage through your own accounts and we set the budgets and dashboards. You own the code, prompts, indexes and evals at handover. A ten-day Sprint Zero through the AI Discovery Sprint at $3,250 or ₹2,00,000 is credited to the build, and a three-week ProofRun via the AI POC Sprint from $6,250 or ₹4,00,000 proves the hardest intent first. Starting prices are on the pricing page. If you want us to answer this list against your specific workflow, get in touch.

Build or buy: the honest case for each in AI customer service agent covers the decision before vendor selection, and the hidden costs of AI customer service agent lists what the quote will not include. If a vendor claims evaluation is impossible for conversational systems, the open-source Ragas documentation describes reference-free metrics such as faithfulness and context precision, which shows the measurement problem is solved well enough to be a standard practice.

A vendor who leads with evidence will bore you for the first twenty minutes and save you the next twelve months.

Frequently asked questions

What is the single most important question to ask an AI support agent vendor?

▾

Ask to see an evaluation suite from a system they have already shipped, broken down per intent and built from real conversation logs. A vendor without one cannot prove quality, cannot safely change anything once you are live, and cannot tell you whether a new model release has made your agent worse.

Should the vendor or the client pay for model usage?

▾

The client should hold the provider accounts and pay usage directly, with the vendor setting budgets, routing rules and dashboards so the number stays predictable. Bundled usage hides the unit economics, removes your ability to switch providers, and makes cost per resolved conversation impossible to verify independently.

What contract terms matter most for an AI customer service agent?

▾

Ownership of code, prompts, retrieval indexes, evaluation sets and deployment configuration, named separately rather than as one clause. Then a fixed price against a written scope, a defined change process, and an exit position where ending the engagement leaves you with a working system you can run or hand to someone else.