Open-weight models vs GPT-class APIs: a benchmark approach
What should you know about open source LLMs vs OpenAI before choosing a model for your product?
Benchmark on your tasks; open-weight models often match on retrieval and extraction and lag on hard reasoning, which routing can cover. Public leaderboards tell you which model is best at leaderboards. A golden set of a few hundred of your own tasks, scored the same way for every candidate, tells you what to ship.
The open source LLM vs OpenAI question is usually asked as if it had a general answer, and it does not. Which model is better depends on the task, the language, the length of the inputs, the latency you can tolerate and what a wrong answer costs. What we can offer instead is a method that produces the right answer for your product in about two weeks, and that keeps producing it as new models arrive every few months. This article describes that method, what we have generally found when applying it, and how to design the system so the model choice stays reversible.
Why open source LLM vs OpenAI has no general answer
Public benchmarks measure broad capability on academic tasks: exams, maths, coding puzzles, general knowledge. They are useful for ruling models out and useless for ruling one in, because your product does not ask exam questions. It asks the model to pull six fields from an invoice in a fixed layout, to answer a customer's question from a specific set of policy documents, to classify a ticket into one of your forty categories, or to hold a booking conversation in Hindi and English mixed. Performance on those tasks varies between models in ways that leaderboards do not predict. Models also change under you: an API model is deprecated on the vendor's schedule, and an open-weight family ships a new size or version every few months. The only durable asset is your own benchmark.
Where open-weight and API models tend to differ
| Task family | What we usually find | Implication |
|---|---|---|
| Retrieval-grounded question answering | Well-chosen open-weight models match frontier APIs when the retrieval is good | Invest in retrieval; the model matters less than people assume |
| Structured extraction from documents | Comparable, and open-weight models are cheaper at volume | Strong self-host candidate |
| Classification and routing | Comparable; small open-weight models often suffice | Use the smallest model that clears the bar |
| Summarisation and drafting | Close; style preferences dominate quality differences | Decide with human ratings, not automatic scores |
| Multi-step reasoning and planning | Frontier APIs lead, sometimes by a wide margin | Route these requests to the API |
| Tool calling and agent loops | Frontier APIs more reliable on long chains; open-weight improving quickly | Benchmark per agent; gate actions regardless |
| Indian languages and code-mixed text | Varies widely by model family; some open-weight models are strong | Test in the actual languages your users write |
| Long context | APIs generally handle very long inputs more robustly | Chunk and retrieve rather than rely on context length |
Building the golden set
A golden set is a few hundred real tasks drawn from your product's actual traffic or documents, each with an expected output that a person has checked. For extraction, the expected output is the correct field values. For question answering, it is a reference answer and the source passage. For classification, the correct label. For drafting, a rubric rather than a single answer. The set must cover the distribution you care about: easy and hard, short and long, English and other languages, clean and messy inputs. Building it takes a domain expert a few days, and it is the most valuable artefact the project produces, because it outlives every model. We describe the construction in evals: the practice that separates AI demos from AI products.
Scoring
Exact match for structured fields, normalised match for names and amounts, a rubric-based judge model with human calibration for free text. The judge model must be checked against human ratings on a sample before its scores are trusted, and it should be a different model from any candidate to avoid it preferring its own style. Every candidate is scored on the same set with the same prompt template, adjusted only for each model's formatting conventions.
Running the open weight model comparison
We typically benchmark three or four open-weight candidates and two API models. Open-weight candidates come from the current families on Hugging Face, chosen by size band to match the hardware you could plausibly run. Each is served with the same engine and settings you would use in production, because quantisation and serving configuration affect quality, and a model tested at full precision but deployed quantised is a different model. Beyond the accuracy score, we record latency at the expected concurrency, cost per thousand tasks including hardware idle time for self-hosted models, and failure modes: refusals, format errors, hallucinated fields. The output is a table your team can read in five minutes, and a recommendation of the smallest model that clears your quality bar for each task family.
Llama vs GPT, Mistral vs OpenAI: what the comparison usually shows
Across the projects where we have run this, a consistent shape appears. For tasks grounded in your own documents, the quality of retrieval decides more than the model, and once retrieval is fixed, the gap between a mid-size open-weight model and a frontier API is small enough that cost and residency decide. For extraction and classification, the same holds, and the open-weight model often wins on cost per task by a wide margin at volume. For hard reasoning, long agent chains and unusual tasks, the frontier API is better and sometimes the only option. That is not a reason to choose one side; it is the reason to route. Specific model names change too fast to be worth printing; the method does not.
Routing: the answer is usually both
Once you have per-task benchmark results, the architecture follows. A router sends each request to the cheapest model that met the bar for its task family, with sensitive data constrained to the private model regardless of score. The application sees one interface. When a new model arrives, it is benchmarked on the golden set and, if it wins on a task family, the routing table changes without touching application code. This is what model-agnostic routing across OpenAI, Anthropic, Google and open-weight models means in practice, and it is how multi-model routing cuts cost without cutting quality. For the deployment side of the private model, see the private agentic AI service.
Mistakes that make benchmarks lie
- Testing on a handful of hand-picked examples rather than a representative set
- Using a different prompt for each model and then comparing the results
- Scoring free text with a judge model that was never checked against humans
- Testing a model at full precision and deploying it quantised
- Ignoring latency and concurrency, then discovering the winning model is too slow
- Letting golden-set examples leak into few-shot prompts or fine-tuning data, which inflates scores as described in data leakage
- Running the benchmark once and never again
A worked example
A D2C brand needed a WhatsApp agent for order queries and personalised recommendations, in English, Hindi and code-mixed messages. The golden set was three hundred real conversations with expected responses checked by the support team, split by language and intent. Four open-weight models and two APIs were benchmarked with the same prompts and retrieval. On order-status and policy questions in English, the mid-size open-weight models matched the APIs; on code-mixed Hindi, one open-weight family was clearly stronger than the others and close to the best API; on multi-step recommendation reasoning, the API led. The routing table sent status and policy intents to the self-hosted model, code-mixed messages to the strongest open-weight model for that language, and recommendation conversations to the API with customer identifiers removed. The agent and the personalisation engine behind it are described in the WhatsApp personalisation case study.
Team and timeline
A model benchmark on your own tasks is the core of a Sprint Zero discovery sprint: ten working days, $3,250 / ₹2,00,000, credited to the build, delivering the golden set, the comparison table across open-weight and API candidates, and a routing recommendation. It needs an ML engineer from us and a domain expert from you for a few days of checking expected outputs. The system that follows is an LLM application from $21,000 / ₹13.6L, or a private agentic AI build from $31,500 / ₹20.8L plus infrastructure if a self-hosted model is part of the answer. Re-running the benchmark when a new model ships is a standard task under a Care Plan. Bands are on the pricing page.
Before you start: a checklist
- Define the task families your product actually performs and the quality bar for each
- Collect a few hundred real examples across languages, lengths and difficulty
- Have a domain expert write or check the expected output for each
- Decide the hardware you could plausibly run, to set the size band for open-weight candidates
- Fix one prompt template and one serving configuration per candidate
- Record latency, cost per task and failure modes alongside accuracy
- Classify which data may reach an API and which must stay private
- Plan to re-run the benchmark on a schedule
Glossary
- Open-weight model: a model whose parameters are published under a licence that permits use, often called open source though licences vary
- Golden set: a fixed set of real tasks with checked expected outputs used to compare models
- Judge model: an LLM used to score free-text outputs against a rubric, calibrated against human ratings
- Quantisation: reducing weight precision to fit hardware, with a measurable quality effect
- Routing: choosing a model per request based on task family, data sensitivity and cost
- Frontier model: the most capable current generation from a major API vendor
Related reading
See how we choose between OpenAI, Anthropic, Gemini and open-weight models per task, self-hosted LLMs: when running your own model beats an API, and model and vendor selection: a benchmark-first approach.
Build the golden set once, benchmark every candidate the same way, route by task, and the open-versus-API argument becomes a table you update when the next model ships.
Frequently asked questions
Are open source LLMs as good as OpenAI's models?
▾
On retrieval-grounded answering, extraction and classification, well-chosen open-weight models are usually comparable and cheaper at volume. On hard multi-step reasoning and long agent chains, frontier APIs still lead. The honest answer comes from your own golden set.
How many examples do we need to benchmark models properly?
▾
A few hundred real tasks with checked expected outputs, covering the languages, lengths and difficulty your product sees. Fewer than a hundred gives noisy results; more than a thousand rarely changes the ranking.
Should we pick Llama, Mistral or another open-weight family?
▾
Pick by benchmark, not by name. Families differ by task and language, and each ships new versions every few months. Design for routing so the winner on each task family can change without a rewrite.