azyware
Business

Build or buy: the honest case for each in self-hosted AI agents

EZ
Eazyware
· 7 min read
Quick answer

Should you build or buy self-hosted AI agents?

Buy the commodity layers of a self-hosted AI agent and build the layers that encode your business. In practice you buy the model runtime, vector store and observability, and build the tool contracts, policy gates and evaluation suites, because nobody sells your process back to you.

Buy the commodity layers of a self-hosted AI agent and build the layers that encode your business. Buying means the model runtime, the vector store, the tracing stack and the queue. Building means the tool contracts, the policy gates and the evaluation suites. Nobody sells those, because they are your process rather than software.

That one-line rule is right about eighty per cent of the time. This article covers the other twenty: the layer-by-layer split, the five questions that decide each layer, what both routes actually cost at Eazyware prices, and the three ways a self-hosted AI agents build vs buy decision goes wrong even when the reasoning looked sound on the slide.

What build and buy even mean once the agent runs on your own infrastructure

A self-hosted AI agent is an agent whose model weights, retrieval index and execution loop all run inside infrastructure you control, usually your own VPC or your own rack, with no prompt or document leaving the perimeter. The concept is covered in the glossary entry on zero data egress.

Self-hosting changes the build versus buy question rather than answering it. Once you have taken on the infrastructure, you still choose for every layer above it whether to adopt an existing component or write one. Teams who treat "self-hosted" and "custom-built" as the same decision end up writing an inference server, which is six weeks of work that vLLM already did better.

The useful framing is ownership of behaviour. Anything that behaves the same way for you as it does for a logistics firm in Pune is commodity: token streaming, batching, embedding storage, span collection, retries. Anything that would be wrong if you copied it from another company is yours: what counts as an eligible refund, which approval threshold applies to a credit note, what a correct answer looks like.

Where the line falls: a layer-by-layer comparison

The following split is what we recommend on most private agentic AI programmes. The third column is the default we start from, and we move a layer only when a question in the next section pushes it.

LayerOff-the-shelf optionCustom optionOur default
Model weightsOpen-weight models from Meta, Mistral, QwenFine-tune or continue pre-trainingBuy; fine-tune only after evals prove prompting has plateaued
Inference servingvLLM, TGI, Ollama on your GPUsIn-house serving layerBuy, always
Retrieval indexpgvector, Qdrant, OpenSearchBespoke index and rankerBuy the store, build the chunking and reranking policy
OrchestrationLangGraph, Temporal, a workflow engineHand-rolled state machineBuy for durable, long-running flows; build for three-step agents
Tool contractsGeneric connectors and MCP serversNarrow, scoped tools over your APIsBuild; this is where correctness lives
Policy gatesPlatform approval widgetsThresholds tied to your authority matrixBuild; a platform cannot know your delegation rules
Evaluation suiteRagas, generic benchmark packsScenario suites from your own transcriptsBuild; buy only the runner
ObservabilityLangfuse, OpenTelemetry, GrafanaCustom tracingBuy, always

Five questions that decide each layer

Run these against any layer you are unsure about. Three or more answers pointing the same way settles it.

  • Would a competitor's version of this work for us unchanged? If yes, buy. Nobody has ever won a deal because their span collector was proprietary.
  • Does it encode a rule that a regulator, an auditor or a board member would ask us to justify? If yes, build. You cannot defend a threshold you did not set.
  • How often will it change? Layers that change monthly, such as prompts, tool scopes and eval cases, belong in your repository. Layers that change yearly can be a dependency.
  • What happens when the vendor deprecates it? Self-hosted AI agents software with a single commercial vendor behind it reintroduces exactly the dependency risk you self-hosted to avoid. Prefer open-weight and open-source components with a real contributor base.
  • Can we staff it? A custom orchestrator needs someone who understands distributed retries at two in the morning. If that person does not exist on your team or ours, buy the engine.

The pattern that falls out is consistent: buy downwards, build upwards. Everything below the agent loop is infrastructure and should be boring. Everything at or above the loop is product.

What does each route cost?

A mostly-bought stack with custom tools and gates is the cheaper starting point and the one we quote most often. Agentic AI solutions on self-hosted infrastructure run from $31,500 or ₹20,80,000 up to $105,000 or ₹72,00,000, plus infrastructure, and that range already assumes open-weight models on vLLM rather than a bespoke serving layer. A multi-agent orchestration programme, where a planner coordinates several specialised workers, is priced separately from $24,500 or ₹16,00,000 to $84,000 or ₹56,00,000.

Building a layer we would normally buy adds calendar time rather than licence cost, which is why it is easy to underestimate. A custom orchestration engine is typically four to six additional weeks and permanent maintenance that never appears in the build quote. GPU capacity is a separate line you carry monthly, sized by concurrency rather than by user count. Full starting figures are on the pricing page, and post-launch the AI system add-on at $750 or ₹40,000 per month covers evals, cost monitoring and prompt regression on whichever route you chose.

If the split is genuinely unclear, a ten-day Sprint Zero at $3,250 or ₹2,00,000, credited to the build, produces the layer decisions, the tool inventory and a GPU sizing estimate before anyone commits budget.

When buying is the right answer

Buy when the workflow is common, the volume is modest, and nothing about your data forces the model onto your own hardware. A support agent answering policy questions over public documentation has no self-hosting requirement at all, and a hosted product such as Eazy Chat AI will be live in a fraction of the time. Self-hosting for its own sake is the most expensive form of reassurance we see teams buy.

Buy also when you are still learning. If you cannot yet write twenty scenarios with known correct outcomes, you do not know the problem well enough to build for it. Run something off the shelf for a quarter, read the logs, and let the failures tell you which layer deserves custom work.

When building pays back

Building the upper layers pays back when the agent touches money, health records or regulated decisions, when one workflow carries enough volume that a two per cent accuracy gain is worth a quarter of engineering, or when your rules simply have no equivalent in any product. An NBFC checking identity documents against internal watchlists is not buying that logic from anyone. Custom versus platform stops being a preference at the point where a wrong action costs more than the build.

The self-hosted AI agents off the shelf route also fails a specific test in regulated Indian financial services: the RBI outsourcing guidelines expect you to demonstrate control over the processing environment, and a hosted multi-tenant agent platform makes that argument harder than an open-weight model on your own GPUs.

Where this decision goes wrong

Three failures recur. The first is building the fashionable layer instead of the load-bearing one: teams write an orchestrator because it is interesting and then wire it to generic connectors that cannot enforce a spend limit. The second is buying a platform on the strength of a demo and discovering that its evaluation story is a dashboard rather than a scenario suite, which leaves you unable to tell whether last week's model upgrade made things worse.

The third is the expensive one: treating build versus buy as permanent. It is a per-layer decision with a review date. We have replaced a bought orchestrator with a custom one eighteen months in, and we have deleted custom retrieval code once pgvector caught up. Write down why you chose each layer so the next team can revisit the reasoning rather than the code.

Honestly, if your volumes are low and your data is not sensitive, the wrong choice here is self-hosting at all. We say so in first calls more often than we recommend it.

What a mixed engagement looks like in practice

An NBFC came to us needing document intelligence for KYC and loan onboarding where customer documents could not leave their environment. The serving stack, the vector store and the tracing were all bought and assembled. What we built was the extraction schema, the validation rules, the confidence thresholds that decide which files a human reviews, and the eval set drawn from real historical applications. The result is described in the KYC document intelligence case study. Roughly a fifth of the engineering went into custom work, and that fifth is the part that made the system usable.

A checklist before you decide

  • List every layer from GPU to interface and mark each buy, build or undecided
  • For each undecided layer, answer the five questions above and record the reason
  • Check the licence of every open-weight model and component against your commercial use
  • Size GPU capacity against peak concurrency, not headcount
  • Confirm who maintains each bought component after launch, and under which Care Plan tier
  • Write twenty evaluation scenarios before committing to any custom build
  • Set a review date, twelve months out, for the three most expensive layer decisions

Build vs buy vs integrate: an AI decision framework generalises this logic beyond agents, Self-hosted LLMs: when running your own model beats an API covers the infrastructure economics, and GPU sizing for private AI turns concurrency into hardware. The vLLM documentation shows why the serving layer is a buy decision: it exposes an OpenAI-compatible server for open-weight models with continuous batching, which is precisely the work you should not repeat.

Decide layer by layer, write down the reason, and revisit it on a date you set now rather than on the day something breaks.

Frequently asked questions

Is it cheaper to build or buy self-hosted AI agents?

▾

Buying the commodity layers is almost always cheaper to start. A mostly-bought stack with custom tools and gates runs from $31,500 or ₹20,80,000 plus infrastructure at Eazyware. Building layers such as inference serving or orchestration adds four to six weeks of build time and permanent maintenance that rarely appears in the original quote.

Which parts of a self-hosted AI agent should never be bought?

▾

Tool contracts, policy gates and evaluation suites. These encode your authority matrix, your definition of a correct outcome and your regulatory obligations. No vendor can supply them accurately, and a platform's generic approval widget will not survive an auditor asking who set a particular threshold and why.

Can you start with a platform and move to a custom build later?

▾

Yes, and it is often the sensible order. Run an off-the-shelf agent for a quarter, collect real transcripts, and use the failures to decide which layer earns custom work. Your evaluation scenarios and tool definitions transfer; the platform's internal orchestration does not.