Questions to ask a self-hosted AI agents vendor before you sign
What should you ask a self-hosted AI agents vendor?
Ask how they will prove the agent works, who owns the code, prompts and weights, what the first twelve months cost including infrastructure, what happens when an agent acts wrongly, and how you leave. Vendors who have shipped self-hosted agents answer all five without preparation.
Ask a self-hosted AI agents vendor five things: how they will prove the agent works on your data, who owns the code, prompts and any fine-tuned weights, what the full first twelve months cost including infrastructure, what happens when an agent takes a wrong action, and how you exit. Teams who have shipped this answer all five without preparation.
The rest of this article is the long form of those five: twenty-two questions grouped by theme, with the answer that should reassure you and the answer that should worry you. Take it into the second meeting, after the demo and before the contract.
Why demos are not evidence
Every vendor can show you an agent doing something impressive. A demo is a curated path through a system on data the vendor chose, and the gap between that and a production agent acting on your records unattended is where most of the money goes. The questions worth asking are therefore not about capability. They are about what happens on the bad day: when retrieval returns the wrong document, when the model is superseded, when a tool call fails halfway, when someone asks in nine months why the agent approved a particular case.
Self-hosting adds a second axis. The vendor is also asking you to take on infrastructure and operational responsibility, so the diligence has to cover their engineering practice and their handover discipline. A firm that builds well and documents badly leaves you with an asset you cannot maintain, and you will not discover that until the engagement has ended and the people who knew the system have moved on.
The five questions that separate the shortlist
If you only get twenty minutes, ask these. The right-hand column is what you should hear if the answer is genuine.
| Question | Answer that should reassure | Answer that should worry |
|---|---|---|
| How will you prove this works? | A named metric on a scenario set built from your data, measured by you | "We will demo it at the end of each sprint" |
| Who owns the output? | You own code, prompts, configuration, eval sets, weights and documentation | "Our platform, licensed to you annually" |
| What does year one cost in total? | Build, infrastructure at forecast utilisation, and a named care tier | A single build number with infrastructure "to be confirmed" |
| What happens when the agent acts wrongly? | Reversible actions, thresholds, approval gates, an audit trail, a rollback path | "Our accuracy is very high" |
| How do we leave? | Documented handover, runs without their licence, a stated support window | Discomfort, or a reference to a renewal conversation |
Questions about evidence and evaluation
This is the cluster that predicts delivery quality better than any other. A vendor who cannot describe their evaluation practice in specifics has probably shipped prototypes.
- How do you build an evaluation set, and who curates it? The answer should involve your real cases and your domain experts, not synthetic examples generated by a model.
- Which metrics will we agree on, and what is the pass threshold? Task completion, escalation precision, groundedness and latency at concurrency are the usual four. Vague answers here become vague acceptance later.
- When does the eval suite run? On every prompt change, model change and retrieval change, in a pipeline, not on request.
- Show me an eval report from another engagement, redacted. A firm that has this can produce it in a day. A firm that does not will offer a case study instead.
- What was the last thing your evals caught before production? The answer reveals whether the practice is real or presentational.
- How will the agent be launched? The answer you want is shadow mode first, with the agent proposing and humans approving until acceptance rates justify autonomy.
Our own position is on the record in evals over demos, and the metric definitions we use are conventional: open frameworks such as Ragas document measures like faithfulness and context precision for retrieval-grounded systems, so there is no excuse for a vendor inventing private vocabulary for the numbers you will be judged against.
Questions about ownership, security and exit
Self-hosting is chosen for control, so a contract that hands control back to the vendor defeats the purpose. Ask each of these in the room and watch for hesitation rather than for the words.
Ownership
Who owns the source code, the prompts, the retrieval configuration, the evaluation sets, the documentation and any weights fine-tuned under the contract? Is there any component you may not modify, and does anything call home to a vendor service? Our answer is that the client owns all of it, which is what code and IP ownership means in practice, and it is worth making a competing vendor say the same sentence out loud.
Security and data
Which of your engineers will have access to our environment, under what identity, with what logging? Where does data sit at rest and who holds the keys? What is your policy on using our data or prompts for anything outside our engagement? Are NDAs signed before the first working session? A security questionnaire for AI vendors lists the rest, and your security team will have its own.
Exit
If we end the relationship at month nine, what do we have? Ask for the specific artefacts: repository, deployment manifests, runbooks, eval suite, prompt history, architecture decision records. Ask how long they will support a handover and at what rate, and ask whether the system runs indefinitely without any key, licence or hosted component of theirs. The test that settles it is blunt: could a competent platform engineer who has never met the vendor redeploy this system from the repository and the runbook alone? If the honest answer is no, you have bought a dependency rather than an asset, whatever the contract says about ownership.
Questions about cost and commercial shape
Ask for the twelve-month number, not the build number. A self-hosted agent programme has three commercial lines: build, infrastructure and care. Our own bands are published rather than negotiated per deal: self-hosted agentic AI runs from $31,500 or ₹20,80,000 to $105,000 or ₹72,00,000 plus infrastructure, care plans run from $1,000 or ₹68,000 to $5,250 or ₹3,40,000 a month with a $750 or ₹40,000 AI add-on for evals and cost monitoring, and everything sits on the pricing page. Ask a vendor to show you the equivalent page.
Then ask three follow-ups. Is this fixed price against a locked scope, and what triggers a change request? Who pays for model and GPU usage during development, and through whose accounts? What is your assumption about the number of integrations, and what happens if a target system turns out to be read-only? The hidden costs of self-hosted AI agents covers the lines that most often surface later.
Questions about the team and the handover
- Who exactly is on this project, and are they the people in this room? Named engineers, not a capability slide.
- Can our engineers pair with yours? About half our work is paired with internal teams, and a vendor who resists this is protecting something.
- What does your documentation look like at the end? Ask to see a redacted runbook from a finished engagement.
- Who do we call at two in the morning in month seven? The answer depends on the care tier and should be stated as a response time, not a sentiment.
- What have you told a client not to build? A vendor who has never talked a client out of an AI project is selling, not advising.
When the vendor is not the problem
Diligence has a failure mode of its own. If you cannot describe the workflow precisely, no answer to any of these questions will protect you, because the scope will move and every vendor will be right to reprice. If your data has never been assembled, the honest answer from a good vendor is that they cannot quote yet, and treating that as weakness will select for the vendor willing to guess.
The same applies to the build-or-buy question underneath all this. If the workflow is generic and a product already does it, a self-hosted custom agent is an expensive way to reach the same outcome. We say so when it is true, and Eazyware versus building an in-house AI team sets out the third option honestly.
What a good answer sounds like in practice
On a multilingual voice programme for a hospital network, the questions that mattered were not about model quality. They were about consent and recording, about what happened when a caller was misunderstood, about how a call was handed to a human without the patient repeating themselves, and about who could read a transcript afterwards and for how long it was kept. The hospital voice agent case study describes the result, and the useful signal during selection was that the constraints were discussed before the capabilities.
Related reading
What to put in a self-hosted AI agents RFP covers the document that precedes this conversation, Five ways self-hosted AI agents projects fail explains what you are diligencing against, and How to measure an AI agent gives you the metrics to insist on. If you want these questions answered by us in writing before a call, get in touch.
The vendor you want is the one who answers the awkward questions faster than the flattering ones.
Frequently asked questions
What should you ask a self-hosted AI agents vendor?
▾
How they will prove the agent works on your data with a named metric and threshold, who owns the code, prompts, eval sets and any fine-tuned weights, what the full first twelve months cost including infrastructure and care, what happens when an agent acts wrongly, and exactly what you receive if you leave at month nine.
How can you tell if an AI vendor has shipped before?
▾
Ask for a redacted evaluation report and a redacted runbook from a finished engagement. Teams who have run systems in production produce both within a day. Teams who have not will substitute a case study or a demo. Also ask what their evals last caught before production, which is hard to answer without the practice.
What should be in the contract for a self-hosted AI agent?
▾
Client ownership of code, prompts, configuration, evaluation sets, documentation and fine-tuned weights; a fixed scope with a named change-request trigger; separate lines for build, infrastructure and monthly care; an audit and logging commitment; and an exit clause stating handover artefacts, support window and that the system runs without vendor licences.