azyware
Business

What to put in an AI copilot development RFP

EZ
Eazyware
· 7 min read
Quick answer

What should an AI copilot development RFP include?

An AI copilot development RFP needs nine things: the named jobs, the systems and API surface, the data and permission model, the action list with approval thresholds, acceptance criteria as numbers, the evaluation method, security and residency terms, ownership of code and prompts, and the support model after launch.

An AI copilot development RFP needs nine things: the named copilot jobs, the systems and API surface, the data and permission model, the write actions with approval thresholds, acceptance criteria expressed as numbers, the evaluation method, security and residency terms, ownership of code and prompts, and the support model after launch. Anything vaguer than that produces bids you cannot compare.

This is a section-by-section walkthrough of the document itself: what each section must state, what a weak response to it looks like, and the specific sentences that make a fixed-price quote hold rather than reopen in month three.

What an AI copilot development RFP is actually for

An RFP for a copilot is a specification of a product surface, not a procurement form for a licence. The deliverable is an assistant embedded in your own SaaS application that reads your customers' account data under their permissions, drafts work and takes scoped actions through your API. Every part of that sentence is a requirement someone has to price.

Most copilot RFPs fail at the same point. They describe an outcome ("an AI assistant that helps users work faster") and leave the jobs, the data access and the write actions to the bidder's imagination. Five vendors then imagine five different products, quote five incomparable numbers, and the evaluation turns into a comparison of slide decks. The cure is to over-specify the boundary and under-specify the implementation: say exactly what the copilot must do and what it must not touch, and let bidders propose how.

The second failure is asking for a demo instead of evidence. A copilot demo is cheap to produce and tells you almost nothing about behaviour on a messy real account. Ask instead for the evaluation method, a sample of graded scenarios and a reference system that has been in production for at least six months.

The nine sections and what each must contain

SectionWhat to specifyA weak response looks like
JobsThree to seven named jobs with current click counts or ticket volumesA list of capabilities: summarise, search, generate
Systems and APIEvery system the copilot reads or writes, and whether an API existsWe will integrate with your stack
Data and permissionsTenancy model, roles, what each role may retrieveRole-based access will be respected
ActionsEach write action, its authorisation and its approval thresholdThe copilot can take actions on the user's behalf
Acceptance criteriaNumbers: completion rate, groundedness, latency, cost per accountHigh accuracy and fast responses
EvaluationSize of the graded set, who writes it, when it runsThorough testing before launch
SecurityResidency, retention, subprocessors, penetration test expectationsEnterprise-grade security
OwnershipCode, prompts, eval sets, infrastructure and model choicesLicence to use the delivered solution
SupportResponse times, eval reruns on model change, hours per monthOngoing support available

Score the responses against those nine rows rather than against overall impression. Give each row a weight before you read a single bid, and mark the sections independently so a strong security answer cannot carry a vague actions answer. In practice the ownership and support rows predict long-term cost better than the price row does, because a copilot you cannot move to another team is a copilot you will keep paying a premium to maintain.

How do you write acceptance criteria a vendor can price?

Write them as thresholds on a fixed test set, not as adjectives. "The copilot resolves at least 70 per cent of the two hundred graded scenarios without human correction, with groundedness above 95 per cent on cited answers, median first-token latency under two seconds, and inference cost under a stated ceiling per account per month." That sentence is priceable. "Accurate and responsive" is not.

Two details make the difference. First, say who owns the graded set: if the vendor writes the test and also marks it, the number means nothing, so the set should be written jointly and held by you. Second, say when the test runs: on every prompt change, on every model version change, and as a condition of each payment milestone. The practice behind this is set out in evals: the practice that separates AI demos from AI products.

The thresholds to state as numbers

  • Task completion rate. The share of graded scenarios finished correctly without human correction, measured on a set you hold.
  • Groundedness. The share of cited answers where every claim traces to a retrieved source, tested adversarially as well as normally.
  • Tenant isolation. Zero cross-tenant retrievals under a defined adversarial prompt suite, proven in an automated test.
  • Latency. Median and 95th percentile time to first token, measured from inside your application rather than a vendor sandbox.
  • Cost per account. A monthly inference ceiling per active account, with routing and alerting to hold it.
  • Correction rate. The share of copilot proposals that users edit before accepting, reported weekly after launch.
  • Escalation quality. The proportion of low-confidence cases that hand off with full context rather than a dead end.

What should the RFP say about budget and timeline?

State a band rather than a number, and state the shape of the engagement you want. Our own AI copilot development for SaaS programmes start at $19,500 or ₹12,80,000 and reach $63,000 or ₹41,60,000 where there are more jobs, deeper integration or heavier governance, and the published bands are on the pricing page. Quoting a band tells bidders which product you are buying and stops the two failure responses: a token quote that reopens later, and a six-figure proposal for a three-job copilot.

On timeline, ask for a date rather than a duration, and ask what the bidder needs from you to hold it. A copilot build of three to five jobs is typically eight to sixteen weeks including shadow mode. If the scope is still moving, say so and ask bidders to price a paid discovery stage first: ours is a ten-day Sprint Zero at $3,250 or ₹2,00,000, credited to the build, and a three-week ProofRun at $6,250 or ₹4,00,000 that proves the hardest job before the full commitment. The mechanics of holding a fixed price are covered in what a fixed-price AI quote should contain.

The clauses that stop a fixed price from reopening

Three clauses do most of the work. The first is scope lock: the jobs, actions and integrations are listed in an annexe, and anything outside it is a change request with its own price, which protects both sides. The second is ownership, stated explicitly across code, prompts, eval sets, infrastructure definitions and model selection, because prompts and eval sets are the assets that make switching vendors possible. The third is a model-change clause: when a provider deprecates or updates a model, whose responsibility is it to rerun the evals and fix regressions, and within what window?

A fourth clause is worth adding where the copilot touches regulated data: a residency and retention statement covering where account data is processed, how long prompts and completions are retained by each provider, and which subprocessors sit in the path. Providers differ on retention defaults for their enterprise tiers, so the answer belongs in the contract rather than in a slide.

Ask about the AI API bill separately from the build fee. In our contracts the client pays for model usage through their own provider accounts, and we set budgets, routing and dashboards so it stays predictable. An RFP that does not separate build fee from running cost invites a bid that hides one inside the other. The trade-offs between engagement shapes are set out in fixed price vs time and materials for AI projects.

When an RFP is the wrong instrument

If you cannot yet name the three jobs, an RFP will not find them for you. Running a competitive process on an undefined product wastes a quarter of your time and produces proposals that are really discovery documents with a price attached. Buy a short paid discovery from one or two vendors instead, compare the artefacts they produce, and write the RFP afterwards with real numbers in it.

An RFP is also the wrong instrument when the copilot is a two-week experiment inside a product that has not found its usage pattern yet, or when your API has no write endpoints and no roadmap to build them. In the second case the honest scope is API work first and a copilot second, and a good bidder will say so rather than quoting for the copilot alone. If you want that conversation before you write anything, talk to us.

A worked scope, for reference

A field-service SaaS company scoped its copilot around three jobs drawn from support ticket volumes: reassign a job, find the nearest available engineer, and close a batch of work orders. Each job mapped to endpoints that already existed, each write action got an approval threshold, and the acceptance test was a graded scenario set run against anonymised production accounts. The finished system is described in the in-app copilot case study. The scope document was four pages, and none of it was about the model.

A security questionnaire for AI vendors gives you the security section almost verbatim, and copilot actions through your existing API explains the permission model your actions section should demand. NIST's AI Risk Management Framework organises AI delivery into govern, map, measure and manage functions; asking each bidder which of the four they will own, and how, separates teams who have delivered under governance from teams who have only demonstrated.

A copilot RFP is good when a competent stranger could read it and build roughly the right thing without calling you.

Frequently asked questions

How long should an AI copilot development RFP be?

▾

Four to eight pages is usually enough. The length should sit in the annexes: the job list with volumes, the systems and endpoints, the action list with approval thresholds, and the acceptance thresholds. Long narrative sections about vision add nothing a bidder can price and dilute the parts that matter.

Should the RFP name a model or a framework?

▾

No. Name the acceptance thresholds and let bidders choose. Most production copilots route between two or three models from different providers, and pinning a model in the RFP freezes a decision that should be made on benchmark evidence for your specific jobs and revisited when providers ship new versions.

What ownership terms should an AI copilot RFP require?

▾

Require ownership of source code, prompts, evaluation sets, infrastructure definitions, documentation and model configuration. Prompts and eval sets matter most: without them a new vendor restarts the quality work from zero. State that deliverables transfer on payment of each milestone rather than at final acceptance.