Model and vendor selection: a benchmark-first approach
What should you know about AI vendor selection and choosing a model provider?
Select by benchmarking candidates on your tasks for accuracy, latency and cost, then design routing so no single vendor locks you in. Public leaderboards rank models on other people's problems; a few hundred of your own labelled examples tell you which model clears your bar at a price you can sustain.
AI vendor selection goes wrong when it is done the way enterprise software has always been bought: a shortlist from analyst reports, a scripted demo, a procurement scorecard, a three-year contract. Language models do not reward that process. They change every few months, they behave differently on different data, and the best one for your invoice extraction may be the wrong one for your support agent. The only reliable method is to benchmark candidates on your own tasks, choose on evidence, and build so the choice can be changed later without a rewrite.
This article sets out the selection criteria that matter, how to run a benchmark that a sceptic would accept, and how routing turns a one-off decision into an ongoing advantage.
Why LLM vendor evaluation is different
Three things distinguish choosing an AI provider from choosing a database or a CRM. First, quality is task-specific: leaderboard rank predicts little about performance on your documents in your language with your edge cases. Second, price and capability move quickly, so the best deal today is unlikely to be the best deal in a year. Third, switching cost is mostly under your control. If prompts, evals and integration sit behind a thin routing layer, changing providers is a configuration change and a regression run; if they are scattered through application code, it is a project.
The practical consequence is that the goal of selection is not to pick the perfect vendor. It is to pick a good one now, on evidence, and to keep the cost of changing your mind low.
AI model selection criteria that survive contact with production
| Criterion | How to measure it | Common mistake |
|---|---|---|
| Task accuracy | Score on your labelled sample, per category of input | Trusting a public benchmark or a vendor demo |
| Latency | p50 and p95 for your prompt sizes, measured from your region | Quoting a vendor's headline figure measured elsewhere |
| Cost per task | Tokens in and out per real task multiplied by list price, plus retries | Comparing per-token prices without measuring tokens per task |
| Data handling | Data residency, retention, training-use terms in the actual contract | Assuming the consumer terms apply to the API |
| Reliability | Error rate and rate-limit behaviour over a week of realistic load | A single afternoon of testing |
| Capability fit | Context length, structured output, tool calling, languages you need | Choosing the largest model when a smaller one fits |
| Exit cost | Hours to switch, measured by actually switching once during the benchmark | Never testing the switch until it is forced |
Running a benchmark a sceptic would accept
Build the sample
Two to five hundred real examples, drawn from production, covering the messy cases as well as the clean ones. Each is paired with the output a competent person would produce. If the task is subjective, two people label independently and disagreements are reviewed; the agreement rate becomes the ceiling any model can reach. This sample is the most reusable asset you will create, because it becomes the regression suite for the whole life of the system. We describe how it feeds later work in Why basic RAG fails in production.
Choose the candidates
Four or five models: the strongest from two or three of OpenAI, Anthropic and Google, one open-weight model that could be self-hosted, and one smaller, cheaper model from any vendor. The cheap one is not a token gesture. On many bounded tasks it clears the bar, and knowing that is worth a great deal of money over a year.
Hold the prompt fair
Use the same prompt and the same examples for every candidate, with only the formatting each vendor requires. Do not tune the prompt for one model and then compare. If a model clearly needs different handling, tune for all of them and report both runs.
Report the whole table
Accuracy by category, latency at p50 and p95, cost per task, failures. The recommendation reads the table, not one column. A model that is three points behind on accuracy and a quarter of the price is often the right answer if the three points fall where a human reviews anyway.
Choosing an AI provider: the contract questions
Once the benchmark narrows the field, the contract decides the rest. Ask whether API inputs are used for training, what the retention period is and whether zero retention is available, where data is processed and whether a regional endpoint exists, what the rate limits are at your expected volume and how they are raised, and what the deprecation policy is for a model version you have tested. Vendors publish most of this; OpenAI's enterprise privacy documentation and Anthropic's documentation are the places to start, and the terms in your signed agreement are the ones that count.
For regulated buyers, the residency question can settle the choice before the benchmark does. If data cannot leave the country or the building, the shortlist is open-weight models on your own infrastructure, and the benchmark tells you which of them is good enough. The trade-offs are covered in Self-hosted LLMs for BFSI: a practical checklist.
Routing: how you avoid lock-in
Routing is a thin layer between your application and the model providers. The application asks for a capability ("classify this ticket", "extract these fields"); the router decides which model handles it, with what prompt version, and records the result. Three things follow. Cheap tasks go to cheap models and hard tasks to strong ones, which lowers the bill. A vendor outage or price change becomes a configuration change. And every request is logged with its cost, so the benchmark never really ends.
The design cost is small if it is done from the start and large if retrofitted. Prompts live in versioned files, not in application code. Evals run against the router, so any model swap is a regression run away from being safe. Provider-specific features are used behind an adapter, never directly. The detailed pattern is in Multi-model routing: cutting LLM costs without cutting quality.
A worked example
A B2B SaaS company adding an in-app copilot began with a strong preference for a single well-known provider, largely because their engineers had used it. The benchmark sample was three hundred real user requests across the copilot's four intents, labelled by two product managers. Results showed the preferred model best on the hardest intent, a smaller model from a different vendor equal on the other three at a much lower cost, and an open-weight model close behind on all four. The router sends the hard intent to the strong model and the rest to the cheap one; the open-weight model is kept in the eval suite as the fallback. When one vendor later raised prices, the change was a configuration edit and a regression run. The product is described in the in-app copilot case study.
Team and timeline
A benchmark-first selection takes one to two weeks: two or three days to assemble and label the sample, three to five days to run and analyse candidates, and a short write-up. We run it as part of Sprint Zero, the ten-day AI discovery sprint, which sits inside our AI product strategy service, and the results become the eval suite for the build that follows. Routing is built into every LLM application we deliver; we are model-agnostic across OpenAI, Anthropic, Google and open-weight models, and the client owns the prompts, evals and routing configuration. Fixed prices are on the pricing page.
Before you start: a checklist
- Define the task in one sentence and the metric that counts as success
- Assemble 200–500 real examples with correct outputs, including messy ones
- Shortlist four or five candidates including one cheap model and one open-weight model
- Confirm data residency and training-use terms for each candidate before sending data
- Measure latency from where your users are, not from a laptop next to the vendor
- Count tokens per real task before comparing per-token prices
- Decide who signs off the recommendation and what would change their mind
- Plan for the router from day one so the choice stays reversible
Questions clients ask
- Should we just pick the model at the top of the leaderboard? No. Leaderboards measure general tasks; your task is specific. Run the sample and let the table decide.
- Is multi-vendor worth the complexity? For anything beyond a prototype, yes. The router is a small amount of code and it converts vendor risk into a configuration change.
- How often should we re-benchmark? On every prompt change, on every model version change, and quarterly against new releases. The sample makes it cheap.
- Can we use a vendor's fine-tuning and stay portable? Partly. Keep the training data and evals in your own repository so a fine-tune can be reproduced elsewhere.
- What if our data cannot go to any external vendor? Then the shortlist is open-weight models on your infrastructure, and the benchmark decides among them.
Related reading
RAG vs fine-tuning: which does your product need? covers the design decision that usually follows model selection, and How to choose an AI development company: a 12-point checklist applies similar evidence-first thinking to choosing a build partner.
Benchmark on your data, read the whole table, sign a contract you have actually read, and route so you can change your mind: that is vendor selection that still looks right a year later.
Frequently asked questions
How many examples do we need to benchmark models properly?
▾
Two to five hundred real examples with correct outputs is enough for most bounded tasks. Fewer than a hundred gives noisy results; more than a thousand rarely changes the ranking but makes a better regression suite.
Is the most expensive model usually the best choice?
▾
Often not. On bounded tasks a smaller model frequently clears the bar at a fraction of the cost, and routing lets you reserve the strongest model for the requests that need it.
How do we avoid vendor lock-in with LLMs?
▾
Keep prompts in versioned files, run evals against a routing layer rather than a vendor SDK, wrap provider-specific features in adapters, and test a switch once during the benchmark so the exit cost is known.