azyware
Business

Questions to ask a LLM application development vendor before you sign

EZ
Eazyware
· 7 min read
Quick answer

What should you ask a LLM application development vendor?

Ask a LLM application development vendor five things: how they measure accuracy, who owns the code and prompts, what the system costs to run, what happens when the model is wrong, and how you leave. Vendors who have shipped answer all five with numbers and artefacts.

Ask a LLM application development vendor five things: how they will measure accuracy before launch, who owns the code, prompts and infrastructure, what the system costs to run each month, what happens when the model is wrong, and how you exit. Vendors who have shipped production systems answer all five with numbers and artefacts. Vendors who have not, answer with adjectives.

Below is the script we would use if we were the buyer, grouped into five rounds, with the answer that should reassure you and the answer that should worry you. Take it into a call and write the responses down; the pattern across five vendors is more informative than any single reply.

Round one: how do you know it works?

This is the question that separates demonstration from delivery. An LLM application that has not been measured is an opinion. Ask the vendor to describe, for a named past project, the evaluation set they built, who wrote the correct answers, how many cases it contained and what the pass rate was at launch.

A good answer contains a number and a method: a golden question set of a few hundred cases, written with the client's subject-matter expert, scored for retrieval recall, factual groundedness and task completion, run in continuous integration on every prompt change. A weak answer is "we test it thoroughly" or a demonstration of the model answering questions the vendor chose. We set out our position in evals over demos, and the practice itself in how to measure RAG quality.

Follow up with a harder one: what was the failure rate, and what did it look like? A vendor who cannot describe how their system fails has not looked. Ask who wrote the correct answers for the evaluation set, because if the vendor wrote them, the system was graded by the people who built it. The right answer involves your subject-matter experts spending real hours, and a good vendor will have insisted on that time before quoting.

Round two: who owns what when we stop paying?

LLM applications accumulate assets that are easy to lose: prompts, evaluation sets, fine-tuning data, retrieval indexes, orchestration code and the configuration that ties them to your infrastructure. Some vendors hand over application code but keep the prompt library on their own platform, which means you own a car without an engine.

Insist that the contract names every artefact: source code, prompts and prompt history, evaluation sets and results, index build scripts, model choices and the reasoning behind them, deployment configuration and documentation. Ask whose cloud accounts and API keys the system runs on. Our own stance, and the clauses we use, are in who owns the code, prompts and models.

Round three: what will this cost to run?

Build price is the number on the proposal. Running cost is the number that decides whether the application survives. Ask for a modelled monthly figure at your stated volume, with input and output tokens shown separately, the percentage attributed to evaluation and retry traffic, and the model tier each step routes to.

Ask who pays for usage. At Eazyware the client pays providers directly through their own accounts, so there is no margin on tokens, and we configure budgets, routing and dashboards during the build. A vendor who resells tokens at an undisclosed rate has an incentive you do not want. The lines that rarely appear in a proposal are collected in the hidden costs of LLM application development.

Then ask the awkward version: what happens to this number when usage triples, and what would you do about it? The answer should mention routing to smaller models, prompt caching and reranking, not "we would discuss it".

The five rounds, and what the answers should sound like

What you askA reassuring answerA warning sign
How do you measure accuracy?A named golden set, a size, a pass rate, run in CIA live demonstration and the word robust
Who owns prompts and evals?Client owns all artefacts, listed in the contractPrompts live on the vendor's platform
What does it cost to run?A modelled monthly figure with token split and routing planA single per-user price with no working shown
What happens when it is wrong?Confidence thresholds, escalation, human approval, audit logGuardrails, described only as a feature name
How do we exit?Runbook, handover sessions, no proprietary runtimeA managed platform with no export path
Who is on the team?Named engineers, with a stated allocation per weekA pool, with people to be assigned later
What is out of scope?A written exclusion list before signatureNothing is out of scope

Round four: what happens when the model is wrong?

Every LLM application is wrong sometimes. The question is what the system does about it. Ask how the vendor detects low confidence, what the application does at that point, whether actions are gated by a human approval, and where the audit trail lives.

Press specifically on prompt injection, because an application that reads untrusted text and has tool access can be steered by that text. OWASP maintains a Top 10 for Large Language Model Applications that lists prompt injection as the leading risk class, alongside insecure output handling and excessive agency. A vendor who has shipped will recognise those terms immediately and tell you how they scope tool permissions. A vendor who has not will talk about content filters. The wider checklist is in a security questionnaire for AI vendors.

Ask one more thing in this round: what does the application refuse to do? A production system has a defined boundary, and the vendor should be able to name the questions it declines, the actions that always require a person, and the point at which it hands off to a human with the full context attached. A system with no refusal behaviour has not met real users yet.

Round five: how do we leave?

  • Ask for the exit clause in writing. What is handed over, in what format, within how many days of notice.
  • Ask what runs on vendor-proprietary infrastructure. Anything that does is a lock-in point, and it should be a deliberate choice rather than a surprise.
  • Ask for a handover session in the contract. Two or three working sessions with your engineers, recorded, not a document dump.
  • Ask who can rebuild the retrieval index. If only the vendor can, you do not control your own knowledge layer.
  • Ask how model changes are handled after handover. A model deprecation twelve months out is certain, and somebody needs the runbook for it.
  • Ask for a reference who left. A vendor confident in their handover will give you one.

What should a LLM application development quote contain?

A quote should contain a fixed scope, a fixed price, a named exclusion list and a dated delivery plan. At Eazyware, LLM application development starts at $21,000 or ₹13,60,000 and runs to $84,000 or ₹56,00,000, with all starting figures published on the pricing page rather than quoted on request.

If a vendor cannot fix a price, ask why. Usually the honest reason is that the scope is not yet knowable, in which case buy discovery first: a ten-day Sprint Zero at $3,250 or ₹2,00,000, credited to the build, or a three-week ProofRun at $6,250 or ₹4,00,000 that tests the hardest step against your real data. What a proper quote includes is broken down in what a fixed-price AI quote should contain.

When due diligence is the wrong use of your time

This script is built for a build worth twenty thousand dollars or more, where the application will touch customer data and stay in production for years. If you are buying a two-week experiment to find out whether an idea has legs, running a seven-round vendor interview costs more than the experiment. Buy the experiment, judge the vendor on how they handle it, and run the full diligence before the real build.

Diligence also cannot rescue an undefined problem. If you cannot state what the application should do, what a correct answer looks like and who will use it daily, no vendor answer will help, because the best vendors will give you the same answer: define that first. We decline work at this stage more often than any other, and the clients who come back six weeks later with a defined workflow get a better system for less money.

Before the first call

  • Write the workflow the application should support, as it runs today
  • Collect fifty real examples with the outcome you would accept
  • List the systems it must read from and write to, with owners
  • Decide your accuracy threshold and who signs off against it
  • Set the monthly running-cost ceiling you are willing to carry
  • Agree internally who owns the application after handover

What to put in a LLM application development RFP turns these questions into a document you can send, fixed price vs time and materials for AI projects covers the commercial model, and model and vendor selection: a benchmark-first approach explains how to judge the model layer separately from the partner. When you are ready to compare answers, talk to us and use the same script on us.

Judge a vendor on the questions they answer without flinching, because those are the ones they have already lived through.

Frequently asked questions

What is the single most important question to ask an LLM vendor?

▾

How will you measure that it works, before launch? A credible vendor describes a golden question set of a few hundred cases, written with your subject-matter expert, scored for retrieval recall, groundedness and task completion, and run automatically on every prompt change. A demonstration is not a measurement.

Should the client own the prompts?

▾

Yes, along with the code, evaluation sets, index build scripts, model choices, deployment configuration and documentation. Prompts are the accumulated knowledge of the project. A contract that hands over application code but keeps prompts on a vendor platform leaves you unable to change, audit or move the system later.

How do I check a vendor has really shipped LLM applications?

▾

Ask them to describe a failure in production: what went wrong, how it was detected, what it cost and what changed afterwards. Teams who have shipped answer specifically within seconds. Ask also for a reference from a client who left, which only a vendor confident in their handover will provide.