What to put in a LLM application development RFP
What should a LLM application development RFP include?
A LLM application development RFP should name one workflow, the measured baseline, the accuracy and latency thresholds that define acceptance, the systems to integrate, the data residency rule, the ownership terms and the support expectation. Without thresholds, bids are not comparable and the quotes will not hold.
A LLM application development RFP should name one workflow, state the measured baseline, define acceptance as numeric thresholds, list the systems to be integrated, set the data residency rule, and spell out ownership and support terms. Bids become comparable the moment acceptance is a number. Without that, every quote prices a different imagined project.
What follows is the section structure we would want to receive, the acceptance language that makes quotes hold, the commercial framing that stops bidders guessing, and the cases where issuing an RFP at all is the wrong move.
What an RFP for an LLM application is actually for
An RFP is a comparison instrument. Its job is to make three proposals differ only in the things you want to choose between: approach, evidence and price. Anything left undefined becomes a place where bidders differ for reasons you cannot see, and the cheapest bid usually wins by having assumed the least.
That matters more for language-model work than for ordinary software, because the deliverable is probabilistic. Two vendors can both deliver "an assistant that answers questions from our policy documents" and one of them can be right sixty per cent of the time. If the document does not say how correctness will be measured and on what data, you have not bought quality, you have bought a demo.
So the centre of gravity of a good RFP is not the feature list. It is the acceptance section, and everything else exists to make that section achievable.
The nine sections to include
| Section | What to specify | If you leave it out |
|---|---|---|
| Business context | The workflow, its volume, its current owner | Bidders scope to the wrong user |
| Baseline | Current handling time, error rate, rework cost | No agreed definition of improvement |
| Scope boundary | What is in phase one and explicitly what is not | Change requests from week three |
| Data and content | Sources, formats, volumes, who maintains them | Retrieval quality is unquotable |
| Integrations | Each system, its API status, sandbox availability | Timeline risk lands on you |
| Acceptance criteria | Numeric thresholds and the test set they run on | Bids are not comparable |
| Security and residency | Where inference and data may run, audit needs | Architecture reopens after award |
| Ownership and exit | Code, prompts, infrastructure, model choices, handover | Vendor lock-in discovered late |
| Support expectation | Response times, eval cadence, re-indexing, escalation | Support gets cut to win the bid |
Nine sections is enough. RFPs that run to sixty pages of boilerplate get skimmed, and the sections that decide the outcome are the ones vendors read carefully.
Write acceptance as thresholds, not adjectives
Replace every instance of "high accuracy", "fast" and "robust" with a number and the data it is measured on. These are the thresholds worth naming.
- Task-level accuracy on a held-out set. State that acceptance is measured on a golden set of your real inputs with human-labelled correct outputs, and say who labels it.
- Groundedness. For anything answering from documents, require every claim to be traceable to a retrieved passage, and require citations in the interface.
- Refusal behaviour. Specify the cases the system must decline rather than attempt, and test them deliberately. A system that never refuses is not safe, it is untested.
- Latency at the ninety-fifth percentile. Averages hide the requests that make users abandon the feature.
- Cost per completed task. Ask for a modelled figure at your stated volume, not a cost per thousand tokens.
- Human correction rate in shadow mode. Require a shadow period and set the acceptance gate on accepted-versus-corrected output rather than on a demo.
- Regression policy. Require the eval suite to run in continuous integration and on every model or prompt change after launch.
One sentence in the acceptance section does more work than a page anywhere else: state that no threshold is accepted on vendor-supplied examples, only on a test set your team assembled.
The commercial section: asking for comparable prices
Ask for a fixed price against a locked scope, with change control defined, rather than a day rate against an estimate. Ask bidders to price discovery separately so you can stop after it if the answer is no. Ask explicitly whether model API usage is billed through your accounts or theirs, because a retainer that hides inference makes the two-year total impossible to compare.
Require a two-year running-cost model alongside the build price: inference at your volume, retrieval infrastructure, observability, support. And require the support quote in the same document as the build quote, or it will be cut to win the bid and reappear as a surprise in month four. The hidden costs that quotes leave out lists what to ask for line by line.
For reference, our LLM application development engagements are fixed price from $21,000 or ₹13,60,000 to $84,000 or ₹56,00,000, discovery is a ten-day Sprint Zero at $3,250 or ₹2,00,000 credited to the build, a three-week ProofRun is $6,250 or ₹4,00,000, and care plans start at $1,000 or ₹68,000 a month with a $750 or ₹40,000 AI add-on for evals and re-indexing. Published figures are on the pricing page, and quoting them here is the point: a bidder who will not publish a range is asking you to negotiate blind.
Governance, security and residency clauses
State where inference may run and where data may rest. For Indian organisations, obligations under the DPDP Act 2023 usually decide between a hosted API, a deployment inside your own cloud account and self-hosted weights, and that decision changes both architecture and price, so it belongs in the RFP rather than in a post-award workshop.
Ask bidders to map their controls to a recognised framework rather than to describe them freely. The NIST AI Risk Management Framework organises the work into govern, map, measure and manage functions, and asking for responses in those terms makes three proposals comparable on safety in the way thresholds make them comparable on quality. Add the practical clauses too: audit logging of every action, PII redaction in traces, retention periods, and a named security contact.
Ownership deserves its own paragraph. Ask for code, prompts, infrastructure definitions, model choices, eval sets and documentation to transfer to you, and ask what happens to the system if the engagement ends at month three. Our compliance checklist for LLM applications covers the clauses in more detail.
What to leave out
Do not name the model. Specifying a provider in the requirements removes the one lever that most reliably improves cost and quality, and any serious bidder will route between two or three models by task anyway. Do not prescribe the vector database, the framework or the orchestration library; ask for the reasoning instead and judge it. The same applies to team composition: asking for a fixed number of named engineers optimises for staffing rather than for outcome, and it is the acceptance thresholds, not the headcount, that determine whether you get what you paid for.
Do not ask for a free proof of concept as part of the bid. It selects for vendors with capacity to give work away and produces demos tuned to whatever sample you hand over, which is the exact failure the acceptance section is designed to prevent. Pay for a short, scoped proof from a shortlist of two instead.
How to score the responses you get back
Publish the scoring weights in the RFP itself. We would suggest roughly forty per cent for evidence, meaning comparable work described with numbers rather than logos, thirty per cent for the proposed approach to evaluation and rollout, twenty per cent for total two-year cost, and ten per cent for team and ways of working. Price alone is a poor discriminator when the underlying scopes are identical, which is the whole purpose of the acceptance section.
Read the risk section of each response first. A proposal that names the things that could go wrong on your specific data, and says what it would do about each, is worth more than one that reads as though delivery were certain. Then check one thing in every response: does the plan contain a shadow-mode period before real users are affected, and is the go-live gate expressed as a number? A proposal without both has quietly moved the risk onto you.
When an RFP is the wrong instrument
If you cannot yet name the workflow or state a baseline, an RFP will return nine proposals for nine different projects and the evaluation will be a beauty contest. Run a paid discovery engagement first, then write the RFP with its output. If the whole budget is under roughly $10,000, the procurement overhead consumes too much of it; approach two vendors directly.
And if a packaged product would do the job, say so in the document and invite bidders to argue for buying rather than building. A vendor who recommends a product over their own build has told you something useful about how they will behave in month six. Build or buy sets out the honest case for each.
Related reading
Questions to ask a vendor before you sign covers the shortlist conversation, LLM application development cost in 2026 explains what drives the number you will receive, and a practical implementation guide shows the delivery sequence a good response should describe. If you would like your draft reviewed before it goes out, send it to us.
An RFP that states its acceptance thresholds in numbers will get three comparable bids; one that asks for high accuracy will get three different projects at three different prices.
Frequently asked questions
What should an LLM application RFP include?
▾
Business context, a measured baseline, an explicit scope boundary, the data sources and who maintains them, each integration and its API status, numeric acceptance thresholds with the test set they run on, security and data residency rules, ownership and exit terms, and the support expectation priced in the same document.
Should an RFP specify which model to use?
▾
No. Naming a provider removes the main lever for improving cost and quality, and most production systems route between two or three models by task. Ask instead how the bidder will select and benchmark models on your data, and how they will handle a provider deprecating a model you depend on.
How do you compare LLM development proposals fairly?
▾
Fix the acceptance thresholds and the test set in the RFP so every bidder prices the same outcome, require a fixed price against a locked scope, and require a two-year running-cost model including inference, infrastructure and support. Then differences in price reflect approach and evidence rather than differing assumptions.