azyware
Technology

Self-hosted LLMs: when running your own model beats an API

EZ
Eazyware
· 7 min read
Quick answer

What should you know about self-hosted LLMs before choosing them over a public API?

Self-host when residency, contracts or volume demand it; open-weight models match public APIs for most retrieval and extraction tasks. If none of those three pressures applies, an API is cheaper and simpler. If one does, plan for GPUs, a serving stack, an evaluation suite and the people to run them.

A self-hosted LLM is an open-weight model running on hardware you control: your own servers, a rented GPU box in a data centre of your choosing, or a private cloud tenancy where prompts and outputs never leave your network. The public APIs from OpenAI, Anthropic and Google are easier, and for most companies they are the right first choice. But there are three situations in which running your own model is the better answer, and this article sets out how to recognise them, what it actually takes, and how to decide with evidence rather than instinct.

When a self-hosted LLM is the right call

The first reason is data residency and contracts. If a regulator, a client agreement or your own policy says customer data cannot be sent to a third-party processor outside a defined boundary, an API is out unless the vendor offers an in-region, contractually acceptable deployment. Indian BFSI clients and healthcare providers meet this most often. The second is volume. At a few million tokens a day an API bill is a rounding error; at hundreds of millions, a dedicated GPU running a well-chosen open-weight model can cost less per token, provided the machine is kept busy. The third is control: you want a model whose behaviour does not change under you when the vendor retires a version, and you want to fine-tune it on your data without sending that data anywhere.

API versus self-hosted: the honest comparison

FactorPublic APISelf-hosted open-weight model
Time to first working prototypeHoursDays to weeks, once hardware is available
Quality on hard reasoning tasksFrontier models leadBest open-weight models trail but close the gap each release
Quality on retrieval, extraction, classificationExcellentComparable for well-chosen models; measure on your tasks
Cost at low volumeLow; pay per tokenHigh; idle GPUs still cost money
Cost at high, steady volumeGrows linearlyFlat once the box is saturated; often lower per token
Data residency and egressDepends on vendor terms and regionFully under your control
Model stabilityVersions deprecate on the vendor's scheduleYou decide when to upgrade
Operational burdenNoneServing stack, GPU drivers, monitoring, on-call
Fine-tuning on private dataLimited, and data goes to the vendorUnrestricted, in your environment

What you are actually running

A private LLM deployment is more than a model file. The serving layer, typically vLLM or a similar engine, batches requests, manages GPU memory and exposes an API compatible with the public ones so application code does not change. Around it sit a gateway for authentication and rate limits, an observability layer that traces every request and its cost in GPU-seconds, an evaluation harness that runs your golden set on every model change, and the retrieval and tool-calling infrastructure that the application needs regardless of where the model runs. Our private agentic AI service delivers exactly this stack, and the zero-data-egress design article explains the network boundary.

Hardware, briefly

Model size determines GPU memory, and GPU memory determines the hardware bill. A mid-size open-weight model quantised sensibly fits on a single modern data-centre GPU and serves a modest workload well. Larger models need several GPUs with fast interconnect. The right starting point for most mid-size companies is one or two GPUs sized to the model you have actually chosen after benchmarking, not a cluster bought on the assumption you will need the biggest model. We cover sizing in GPU sizing for private AI.

Open-weight models enterprise teams can rely on

The open-weight landscape moves quickly, which is a reason to design for model swaps rather than to pick a favourite. Model families from Meta, Mistral, Google, Alibaba and others are released with permissive or commercially usable licences, and within each family there are sizes from small enough for a laptop to large enough to rival frontier APIs on many tasks. The right way to choose is to benchmark three or four candidates on your own tasks, with your own golden set, and pick the smallest that clears the quality bar, because the smallest is the cheapest to run and the fastest to respond. That method is described in open-weight models versus GPT-class APIs.

Routing: you do not have to choose only one

The most common production pattern we build is a hybrid. The self-hosted model handles everything that touches sensitive data or runs at volume: retrieval over internal documents, extraction from forms, classification of tickets, first-draft generation. A public API handles the rare hard reasoning task on data that has been redacted or that was never sensitive. A router decides per request based on task type and data classification, and the application code sees one interface. This keeps the GPU busy with the work it is good at and buys frontier quality only where it is needed. Model-agnostic routing across OpenAI, Anthropic, Google and open-weight models is part of every stack we ship, precisely so this decision can change without a rewrite.

Signs you should stay on an API

Self-hosting is sometimes chosen for reasons that do not survive scrutiny. A general unease about cloud vendors is not a residency requirement; a contract clause is. A wish to "own the AI" is satisfied by owning the prompts, evaluation sets, retrieval indexes and application code, all of which you own regardless of where the model runs. Low volume with occasional spikes is the worst possible profile for dedicated GPUs. And a team with no platform engineer and no appetite for on-call should not take on a serving stack, however good the per-token maths looks. If none of the three reasons above applies clearly, use an API in a region acceptable to you, with vendor terms that exclude training on your data, and revisit when volume or requirements change.

The costs people forget

  • GPU utilisation: a box running at a fraction of capacity costs the same as one running flat out; batch work and background jobs are how you fill it
  • Upgrades: a new open-weight release every few months means a benchmark run and a controlled swap, not a free improvement
  • On-call: when the model server falls over at midnight, someone has to bring it back; a Care Plan or an internal owner is required
  • Evaluation: without a golden set you cannot tell whether the model you self-hosted is as good as the API you left
  • Security patching of the serving stack, drivers and base images, which is ordinary infrastructure work that is easy to neglect

A worked example

A hospital network needed an assistant for clinical documentation and a multilingual voice front-end, and its data governance policy did not permit patient records to leave its own environment. A public API was ruled out for anything touching records. We benchmarked three open-weight models on the network's own de-identified transcripts and summaries, chose the smallest that met the clinicians' accuracy bar, and deployed it on a pair of GPUs in the network's data centre behind vLLM, with a gateway, tracing and a golden-set evaluation that runs on every change. Non-sensitive tasks such as general medical terminology lookups route to a public API with no patient data in the prompt. The voice side of that programme is described in the multilingual voice agent case study.

Team and timeline

A self-hosted deployment needs an ML engineer for model selection and evaluation, a platform engineer for the serving stack and GPUs, and someone on your side who owns infrastructure and security sign-off. We start with a Sprint Zero discovery sprint at $3,250 / ₹2,00,000 to confirm the residency requirement, benchmark candidates and size the hardware. The build is a private agentic AI engagement from $31,500 / ₹20.8L plus infrastructure, typically eight to twelve weeks to a production stack with evaluation and monitoring. Running it afterwards fits a Standard or Enterprise Care Plan, depending on the response time you need. The pricing page lists the bands, and everything we deploy, including the model weights, prompts and infrastructure code, is owned by you.

Before you start: a checklist

  • Write down the actual residency or contractual requirement, with the clause, not the assumption
  • Estimate daily token volume for the next year, honestly
  • Build a golden set of a few hundred real tasks with expected outputs before benchmarking anything
  • Confirm who will own the serving stack after launch and what their on-call arrangement is
  • Decide which tasks may route to a public API and how data is classified for that decision
  • Check GPU availability and lead time in your region or cloud account
  • Agree how model upgrades will be evaluated and approved

Glossary

  • Open-weight model: a model whose trained parameters are published for download and use under a licence
  • Quantisation: reducing the numeric precision of weights to fit a model in less GPU memory with a small quality cost
  • Serving engine: software such as vLLM that runs a model efficiently and exposes an API
  • Golden set: a fixed collection of real tasks with expected outputs, used to compare models
  • Routing: sending each request to the model best suited to it, by task type and data sensitivity
  • Egress: data leaving your controlled network boundary

See the self-hosted LLM checklist for BFSI, LLM inference costs: how to forecast your monthly bill, and the security page for how we handle client environments.

Self-host because a requirement demands it or the volume justifies it, benchmark before you commit, and keep a router so the decision stays reversible.

Frequently asked questions

Is a self-hosted LLM cheaper than an API?

▾

Only at high, steady volume where the GPU stays busy. At low volume, idle hardware costs more than per-token pricing. Model the year's token volume before deciding; the crossover point depends on the model size you actually need.

Are open-weight models as good as GPT-class APIs?

▾

For retrieval, extraction, classification and drafting, well-chosen open-weight models are comparable on most business tasks. On hard multi-step reasoning, frontier APIs still lead. Benchmark on your own golden set rather than on published leaderboards.

Can we self-host and still use a public API?

▾

Yes, and it is the pattern we recommend: route sensitive or high-volume work to the private model and redacted, hard-reasoning tasks to an API, behind one interface so the split can change later.