azyware
Business

Questions to ask a RAG development services vendor before you sign

EZ
Eazyware
· 7 min read
Quick answer

What should you ask a RAG development services vendor?

Ask a RAG development services vendor how they measure retrieval quality, how they enforce who may see what, what the system costs to run at your query volume, what happens when a model is deprecated, and what you receive on the day the contract ends.

Ask a RAG development services vendor how they measure retrieval quality, how they enforce who may see what, what the system costs to run at your query volume, what happens when a model is deprecated, and what you receive on the day the contract ends. The answers separate shipped systems from slide decks.

Every vendor demonstrates well, because a demo is a curated question against curated content. What follows is a set of questions designed so that a vendor who has only built demos cannot answer them fluently, along with what a good answer sounds like and where the conversation usually goes wrong.

Ask questions that retire a risk, not questions that invite a pitch

A useful due-diligence question has a specific failure in mind. "Tell us about your AI experience" retires nothing. "Show us a retrieval evaluation report from a system you shipped, with the numbers that failed first" retires the risk that the vendor has never had to defend an accuracy figure to a customer.

There are five risks in a retrieval project: the system retrieves the wrong passages, it invents claims the passages do not support, it shows someone content they are not entitled to see, it costs more to run than the problem is worth, and you cannot leave. Every question below maps to one of those.

The five questions and the answers you want to hear

QuestionA weak answerA strong answer
How will you measure retrieval quality on our content?We use the latest embedding model and it performs very wellWe build a golden set with your experts, report recall at k and groundedness weekly, and hold thirty per cent back for acceptance
What happens when the system does not know?The model is trained to be helpfulIt refuses, cites what it did find, and routes to a named human, and we measure the refusal rate on out-of-scope questions
How do you stop a user seeing a document they cannot access?We can add permissions laterAccess metadata is attached at indexing and applied as a pre-filter before ranking, tested with per-role fixtures
What will this cost us to run at 5,000 questions a month?Inference is very cheap nowHere is the token model per answer, the re-indexing volume, and the routing and caching design that keeps it flat
What do we get if we terminate?All the code is in the repositoryCode, prompts, index configuration, evaluation harness, golden set and a runbook, all in your accounts from day one

Questions about your corpus, not their product

The best signal in an early conversation is how much the vendor wants to know about your content. A vendor who asks nothing about document formats, update frequency or who owns the wiki is planning to find out during the build, at your expense.

Ask what they would do with a corpus that is forty per cent scanned PDFs of variable quality. Ask how they would handle three versions of the same policy with no clear authority marker. Ask what chunking approach they would propose for contracts, and why it differs from what they would do for support articles. A practitioner answers these quickly and with caveats. A salesperson answers them smoothly and identically.

Ask also what they will refuse to do. Retrieval is the wrong tool when the questions require calculation over structured data rather than passages from documents, and the honest answer there is a query layer over your database instead. A vendor who says every problem is a retrieval problem is describing their inventory, not your requirement.

What should you ask about evaluation?

Ask to see a real evaluation report, redacted if necessary, from a system in production. You are looking for three things: metrics defined precisely enough to reproduce, a segmented breakdown rather than a single headline number, and evidence that some number failed and was fixed. Reports where everything passed on the first run are marketing.

Then ask who writes the reference answers. The correct answer is your subject-matter experts, with the vendor providing structure. If the vendor writes both the questions and the answers, they are grading their own homework. Our position on this is set out in evals over demos, and the metrics themselves are explained in how to measure RAG quality.

Finally, ask what happens to the eval suite after launch. It should be runnable by your team, versioned with the prompts, and re-run on every model change. If it lives on the vendor's laptop, you have bought a report rather than a capability.

Questions about permissions, residency and security

Ask how access control is enforced and at which point in the pipeline. Filtering after ranking is a common shortcut that leaks through result counts and latency; filtering before ranking is the correct design. Ask for the test fixtures that prove it, role by role.

Ask about prompt injection specifically. Content in your corpus can carry instructions, and a vendor should be able to describe how retrieved text is isolated from instructions and what is logged when something looks like an attempt. The OWASP project maintains a Top 10 list of risks for large language model applications that names prompt injection and sensitive information disclosure among the leading categories, and a serious vendor will already be working against it.

Then the boring but decisive questions: where does data sit at rest, which model providers see your content, what the retention setting is on those provider accounts, and whether a self-hosted option exists if a regulator later insists. Our own posture is documented on the security page, and a security questionnaire for AI vendors gives a fuller list to send out.

What should you ask about cost and contract?

Ask for three numbers in writing: the build, the expected monthly running cost at your stated volume, and the cost of a model migration. Ask whose accounts the model usage is billed to; the right answer is yours, so the bill is visible without a markup. Ask what a scope change costs and how one is agreed.

On ownership, insist the contract names code, prompts, index configuration, evaluation data and documentation as yours, which is the argument made in who owns the code, prompts and models. Ask what the handover contains and how long it takes. For reference, our retrieval and knowledge engineering programmes run from $14,000 or ₹8.8 lakh to $49,000 or ₹32 lakh, and all starting figures are published on the pricing page rather than discovered in a negotiation.

One more contractual question is worth asking out loud: who is actually on the team. Ask for the names of the engineers who will do the work, not the ones presenting, and ask what proportion of their week you are buying. Ask what happens if that person leaves mid-project. Retrieval work depends heavily on judgement built from previous corpora, so a substitution part-way through is a real risk rather than a staffing detail.

The questions that are not worth asking

Which vector database do you use is a poor question, because the answer rarely predicts quality and the honest reply is that several would work. Which model do you use is similarly weak: most production systems route between two or three based on benchmarks against your own question set, as argued in model and vendor selection. How many AI engineers do you have tells you about the vendor's size, not about whether your project will be staffed by the people in the room.

Asking for a free proof of concept is also a mistake, and not only because good vendors decline. Unpaid work is scoped to impress rather than to inform, which is the opposite of what you need. Pay for a small, honest one instead: a three-week ProofRun from $6,250 or ₹4,00,000 tests the hardest subset of your corpus against thresholds you set, and the output is evidence either way.

When the vendor is not the problem

If three credible vendors all decline your scope or price it far above your budget, the scope is the issue. The commonest version is a corpus nobody has curated, where the real first project is content work rather than software. The second is a question set that requires arithmetic over transactions, which belongs in a natural language data querying layer rather than a retrieval pipeline.

It is also worth asking whether you should build this internally at all. If you have platform engineers with spare capacity and a long-term product reason to own the pipeline, the calculus differs; Eazyware vs an in-house team sets out the trade honestly.

A shortlist scoring checklist

  • Did they ask more questions about our content than we asked about their product
  • Can they show a production evaluation report with a failure in it
  • Do they propose a held-out acceptance set they have not seen
  • Is access filtering applied before ranking, with per-role test fixtures
  • Is model usage billed to our accounts, with dashboards and budgets
  • Are code, prompts, index configuration and eval data ours in the contract
  • Do they name who on our side must be available, and in which weeks
  • Did they tell us one thing we were planning that they think is wrong

Five ways RAG development services projects fail covers the failure modes these questions are designed to detect, how to measure whether RAG development services is working sets the post-launch metrics, and why basic RAG fails in production explains the technical traps behind the weak answers above.

The vendor worth signing is the one who tells you which part of your plan will not work.

Frequently asked questions

What is the single most revealing question to ask a RAG vendor?

▾

Ask to see an evaluation report from a system they shipped, including a metric that failed and what they changed. Vendors who have only built demos cannot produce one, because a demo is never measured against a held-out question set written by somebody other than the person who built it.

Should a RAG vendor provide a free proof of concept?

▾

No, and you should not want one. Free work is scoped to impress rather than to test the hard cases in your corpus. A short paid engagement against thresholds you set produces evidence you can act on, and our three-week ProofRun starts at $6,250 or ₹4,00,000.

What contract terms matter most in a RAG project?

▾

Ownership of code, prompts, index configuration, evaluation data and documentation; model API usage billed to your own provider accounts; a defined exit package with a delivery deadline; numeric acceptance criteria measured on a held-out set; and a written scope-change process with prices attached.