LLMOps for small teams: the minimum viable operations
What is the minimum LLMOps a small team needs to run an LLM application?
Minimum LLMOps is tracing, an eval suite on every change, cost dashboards, model pinning and a rollback for prompts. A team of two or three engineers can run all five with open-source tooling and a few days of setup, and without them an LLM application in production is being operated on hope rather than evidence.
Minimum viable LLMOps for a small team is five things: tracing of every request from prompt to response to cost, an evaluation suite that runs on every change, a cost dashboard that someone looks at weekly, pinned model versions behind a routing layer, and a versioned prompt store with a rollback. Nothing else is required to run an LLM application responsibly, and nothing less is enough. A team of two or three engineers can set all five up in a week with open-source tools and keep them running in a few hours a month. This article is the checklist we install on every build and maintain under our maintenance and support service, written for teams who do not have a platform group and do not want to become one.
The temptation is to buy an LLMOps platform and assume the problem is solved, or to do nothing until the first incident. Both are expensive. The middle path is a small set of practices that fit in a normal engineering workflow.
What LLMOps is, for a team without a platform group
LLM operations is the set of practices that keep an LLM application behaving as accepted, costing what was budgeted and recoverable when something goes wrong. It borrows from ordinary software operations (observability, CI, versioning, rollback) and adds what is specific to language models: outputs that are not deterministic, behaviour that can change upstream, and quality that has to be measured rather than asserted. For a small team the aim is not a mature platform; it is enough evidence to make each change safely and enough visibility to notice when the world moves.
The five practices, what they cost and what they catch
| Practice | Minimum implementation | Setup effort | What it catches |
|---|---|---|---|
| Tracing | Every call logged with prompt version, model, tokens, latency, cost, user and outcome; searchable | One to two days with an open-source tracing tool | Slow paths, expensive users, the exact input behind a bad output |
| Eval suite on every change | Golden set of real cases with expected outcomes; scored in CI; regression blocks merge | Two to three days for the first hundred cases | Prompt regressions, retrieval changes, provider model updates |
| Cost dashboard | Spend by feature, model and tenant, daily, with an alert on a threshold | Half a day on top of tracing | Runaway loops, a tenant abusing a feature, a model switch that doubled cost |
| Model pinning and routing | Explicit model versions in config; one routing layer; a fallback per task | One day | Silent provider updates, outages, deprecations |
| Prompt versioning and rollback | Prompts in the repository, versioned, referenced by id in traces; deploy and revert like code | One day | The change that broke production, and the way back |
Tracing: the foundation everything else stands on
Without tracing, every incident starts with "can you reproduce it?" and the answer is no. A trace records the full chain for each request: the prompt version and rendered prompt, retrieved context, model and version, tool calls, the response, tokens, latency, cost, and whatever outcome signal you have (thumbs, escalation, task completed). Open-source tools such as Langfuse or an OpenTelemetry pipeline do this with a few lines of instrumentation. The rule is that a trace must let an engineer see exactly what the model saw. LLM observability: tracing every request goes into the schema.
Evals on every change
The eval suite is the CI test for behaviour. Start with a hundred real cases pulled from traces: the questions users actually asked, with the answers a business owner agrees are right, plus the failures you have already seen. Score them automatically where the answer is checkable (a field extracted, a classification, a tool called) and with an LLM judge or a person where it is not. Run the suite on every prompt change, retrieval change, code change and provider release, and block the merge on a regression. The suite grows with every production failure, so month six is better protected than month one. Evals: the practice that separates AI demos from AI products covers how to build and score one.
A cost dashboard someone looks at
LLM costs are usage-based and can move faster than a monthly invoice reveals. The dashboard is a view over the traces: spend by feature, by model and by tenant or customer, daily, with a threshold alert. The questions it answers are the ones finance will ask: which feature costs the most, which customer is expensive, what did the last model change do to the bill. Once a week someone opens it. Cutting inference costs by a third describes what to do when the numbers are wrong, and LLM inference costs: how to forecast your monthly bill covers budgeting.
Pin models and route through one layer
Every model call goes through one routing layer that maps a task name to a pinned model version and a fallback. The application never names a model directly. This gives you three things for one day of work: no silent provider updates, a switch to a fallback during an outage, and a migration path when a version is deprecated that is a config change plus an eval run rather than a code hunt. It also makes a model-agnostic stance practical: the same task can be evaluated on OpenAI, Anthropic, Google and open-weight models and routed to whichever wins on quality and cost. Handling model deprecations shows the migration in detail.
Prompts as code, with a way back
Prompts live in the repository, not in a dashboard or a database someone edits by hand. Each version has an identifier that appears in every trace, so a bad output can be tied to the exact prompt that produced it. Deploys go through the same pipeline as code, with the eval suite as the gate, and a revert is a normal revert. Teams that skip this discover that the prompt in production is not the one in the document, and that nobody knows when it changed. Prompt versioning and evaluation sets out the practice.
What you can leave out at first
A small team does not need a feature store, a fine-tuning pipeline, a custom judge model, automated red-teaming or a multi-region deployment on day one. It does not need a commercial LLMOps platform if the open-source tracing tool and CI are in place. It does need the five practices above, and it needs one person who owns the weekly ritual: read the cost dashboard, review a sample of traces, add failures to the eval set, check provider deprecation notices. That ritual is the whole of LLM operations for most applications.
A worked example
A B2B SaaS company shipped an in-app copilot with a team of three engineers and no operations tooling beyond application logs. The first month went well; the second brought a cost surprise from one customer's heavy use and a quality complaint nobody could reproduce. Installing tracing answered both within a day: the traces showed the customer's usage pattern and the exact prompt and context behind the complaint. A hundred-case eval set built from the traces caught a regression on the next prompt change before it shipped. Model pinning behind a routing layer meant the provider's next release was evaluated first and adopted deliberately. The in-app copilot case study describes the product; the operations were what made month three uneventful.
Team and timeline
Setting up the five practices takes about a week for one engineer on an existing application, and it is included in every Eazyware build so that a system is handed over with tracing, an eval suite, dashboards, routing and prompt versioning in place, all owned by the client. For teams running the system themselves, the weekly ritual is a few hours; for teams who want it done, a Care Plan covers it: Essential at $1,000 per month (₹68,000) for business hours and 10 hours, Standard at $2,500 (₹1,60,000) for 24×5 and 25 hours, Enterprise at $5,250 (₹3,40,000) for 24×7 with a named engineer. Applications built elsewhere without these practices are usually brought up to this baseline in the first month under the plan. The pricing page has the details, and LLM applications describes the build service that ships with them.
Before you start: a checklist
- Instrument every model call with prompt version, model, tokens, latency, cost and outcome
- Pull a hundred real cases from traces and agree expected outcomes with a business owner
- Wire the eval suite into CI and block merges on regression
- Build the cost view by feature, model and tenant with a threshold alert
- Move every model name into a routing config with pinned versions and a fallback
- Put prompts in the repository with ids that appear in traces
- Name the owner of the weekly ritual and put it in the calendar
- Subscribe to deprecation notices from every provider you route to
Glossary
- Trace: the full record of one request, from rendered prompt and context through model call to response and cost
- Golden set: curated real cases with agreed expected outcomes, used to score every change
- LLM judge: a model used to score outputs against a rubric where exact matching is impossible
- Routing layer: the single component that maps a task to a pinned model version and a fallback
- Prompt id: the version identifier for a prompt, recorded in traces and deployed like code
- Regression: an eval score below the last accepted run, which blocks release
Related reading
Why AI systems degrade after launch explains what these practices are defending against, and What makes an LLM application production-ready covers the build-side checklist. For a concrete open-source starting point, the Langfuse documentation covers tracing, prompt management and evaluation in one tool.
Five practices, a week of setup and a weekly ritual are all the LLMOps most small teams need; skip them and every incident starts from nothing.
Frequently asked questions
What is the minimum LLMOps for a small team?
▾
Tracing of every request, an evaluation suite run on every change, a cost dashboard by feature and tenant, pinned model versions behind a routing layer, and versioned prompts with a rollback. Together they take about a week to set up.
Do we need a commercial LLMOps platform?
▾
Not at first. Open-source tracing tools, CI for the eval suite and a routing config cover the five practices. A platform becomes worth it when multiple teams or many applications need shared tooling.
Who should own LLM operations in a small company?
▾
One named engineer who runs a weekly ritual: read the cost dashboard, review sampled traces, add failures to the eval set and check deprecation notices. Teams who want it done for them can use an Eazyware Care Plan.