azyware
Technology

Prompt versioning and evaluation: treat prompts like code

EZ
Eazyware
· 7 min read
Quick answer

How should you version and evaluate prompts in a production LLM application?

Version prompts, run them against a golden set on every change and block releases on regression, exactly as you would with code. Prompts edited in place with no test are the most common reason an AI feature that worked in April is worse in June. Here is the workflow, tooling and eval design that make changes safe.

Prompt versioning is the practice of treating every prompt as a versioned artefact: stored in the repository, changed through review, tied to an evaluation result and rolled back when a release goes wrong. Paired with prompt evaluation against a golden set, it is what lets a team change a prompt on Tuesday and know by Tuesday afternoon whether the change helped. Without it, prompts drift through admin panels and chat windows, nobody can say which version answered a complaint, and a fix for one case silently breaks fifty others. This article sets out the workflow we use on every build under the LLM applications service.

Why prompts need versioning more than ordinary code

A code change affects the inputs it was written for. A prompt change affects every input the model will ever see, because the model interprets the whole prompt afresh each time. Adding one sentence to handle a refund edge case can change how the model formats dates, how often it asks clarifying questions, or whether it still refuses out-of-scope requests. The blast radius is wide and invisible until a user reports it. The only defence is to run every version against a representative set of inputs before it ships and to keep every version so the previous behaviour can be restored.

There is a second reason. Providers update models under stable names and deprecate older ones. A prompt that was tuned against one model snapshot can behave differently against the next. Versioned prompts with a recorded eval score let you re-run the same set against a new model and see the difference as a number rather than a feeling.

Prompt engineering best practices, before and after versioning

PracticeWithout versioningWith versioning and evaluation
Where prompts liveDatabase field, admin panel, hard-coded stringRepository or registry, one file per prompt, semantic version
How a change shipsEdit and hopePull request, review, eval run attached, merge
Knowing what is liveAsk whoever edited it lastVersion ID logged on every request trace
Reacting to a complaintGuess which change caused itReproduce against the logged version and inputs
Rolling backRetype the old text from memoryRevert the version; traffic moves in minutes
Changing modelRe-tune blindRe-run the golden set, compare scores per version

The versioning workflow

Store prompts as files

Each prompt is a file with a name, a version, the template text, the variables it expects, the model and parameters it was tuned for, and a short changelog. Templates use a simple substitution syntax so the same file is used in the application and in the eval harness. Prompts that are composed from parts (a system preamble, a task section, a few-shot block) keep each part versioned so a change to the shared preamble is visible everywhere it applies.

Change through review

A prompt change is a pull request. The reviewer reads the diff, reads the eval comparison the harness attaches, and looks at a sample of outputs that changed. The rule is simple: no eval run, no merge. This is also where a second person catches the sentence that solves the reporter's case and breaks the format for everyone else.

Log the version on every request

The trace for each request records the prompt version, the model, the rendered inputs and the output. This is the link between versioning and observability: when a user flags an answer, the team can pull the exact prompt and inputs, reproduce the result, and add the case to the golden set. Tools such as Langfuse support prompt management and tracing together, and the OpenTelemetry semantic conventions for generative AI define the attributes to record if you prefer your own stack.

Prompt evaluation: building the golden set

A golden set is one to three hundred real inputs with an expected output or a grading rubric for each. Real inputs matter: invented cases test what the author imagined, production cases test what users actually do. Start from traces, sample across the input types the feature sees, and include the awkward ones: ambiguous requests, out-of-scope questions, inputs in the wrong language, and the cases that generated complaints. Each case records who verified the expected output and when.

Scoring depends on the output. Structured outputs are checked exactly: did the classification match, did the extracted fields match, was the JSON valid. Prose outputs are graded against a rubric, usually by a stronger model acting as judge with the rubric in its prompt, and a sample is checked by a person each cycle so the judge's own drift is caught. Groundedness checks confirm the output stays within the provided context. The overall score is reported per category, because an average hides a category that fell off a cliff. The wider method is in evals: the practice that separates AI demos from AI products.

Prompt regression testing in the pipeline

The eval harness runs in continuous integration like any other test suite. On a prompt or model change it runs the full golden set, compares scores per category with the current version, and fails the build if any category drops below a threshold agreed in advance. Cost and latency are recorded alongside quality, so a change that improves accuracy by making the prompt three times longer is visible as a trade-off rather than a free win. Because LLM outputs vary, the harness runs each case more than once where the budget allows and reports the spread; a change inside the noise band is neither a regression nor an improvement.

Two practical points. Keep a small smoke subset (twenty or thirty cases) that runs on every commit in a minute, and the full set on merge. And version the golden set itself, so a score from March can be compared with one from September even though the set has grown.

What versioning does not fix

Versioning makes change safe; it does not make a bad prompt good. Prompt design still matters: clear task statements, explicit output formats, examples for ambiguous cases, instructions on what to do when the input is out of scope. Structured outputs remove a whole class of format problems, covered in structured outputs and function calling. And a prompt cannot compensate for retrieval that returns the wrong passage; if the eval fails on a category because the context was wrong, the fix is in retrieval, not the prompt.

A worked example

A fintech company had a support assistant whose prompt had been edited in an admin panel around forty times over six months by three different people. Each edit fixed a reported case. By month six the prompt was several pages long, contradicted itself in two places, and the assistant had become cautious to the point of refusing routine balance questions. Nobody could say which edit had caused the refusals, and the previous versions existed only in a chat thread.

The team moved the prompt into the repository, reconstructed a history from the thread, and built a golden set of two hundred cases from traces, categorised by intent. Running the current prompt against it showed the refusal problem concentrated in one category. A rewrite that was a third of the length scored higher in every category, and the two contradictory instructions were found because the diff made them adjacent. From then on every change went through review with the eval attached. The next model upgrade was tested against the same set in an afternoon and shipped with a documented score. The same discipline underpins the assistant in the in-app copilot case study, where prompts change weekly without incident.

Team and timeline

Setting up prompt versioning and an eval harness is one AI engineer for one to two weeks at the start of a build, or two to three weeks when retrofitting to an existing feature because the golden set has to be assembled from traces and verified. Ongoing, an engineer spends a few hours a week extending the set from production cases. It is included in every LLM applications build, which starts at $21,000 / ₹13.6L, and in the six-week Launch 6 MVP program; a Care Plan covers the ongoing eval maintenance after launch. Current figures are on the pricing page, and the golden set, harness and prompt history belong to the client.

Before you start: a checklist

  • Move every prompt into the repository with a version and a changelog
  • Record the prompt version on every request trace
  • Sample 100–300 real inputs across categories and verify expected outputs
  • Decide scoring per output type: exact match, rubric grading, groundedness
  • Set per-category regression thresholds before the first run
  • Add the harness to CI with a fast smoke subset and a full run on merge
  • Version the golden set alongside the prompts
  • Agree who reviews prompt changes and that no eval run means no merge

Glossary

  • Golden set: a curated collection of real inputs with verified expected outputs used to score every change.
  • Regression: a drop in eval score for any category compared with the current version.
  • LLM-as-judge: using a stronger model with a rubric to grade prose outputs, spot-checked by people.
  • Prompt registry: a store that holds prompt versions and serves the current one to the application.
  • Smoke subset: a small, fast slice of the golden set run on every commit.

See what makes an LLM application production-ready for where versioning sits among the six requirements, LLM observability for the tracing side, and how to handle hallucinations for what the eval should catch.

A prompt is code in English; give it a version, a test and a reviewer, and it will stop surprising you.

Frequently asked questions

Where should prompts be stored?

▾

In the repository or a prompt registry, as versioned files with a changelog, never in a database field edited in production. The version must be logged on every request so any answer can be traced back to the prompt that produced it.

How big does a golden set need to be?

▾

One to three hundred real cases sampled across input categories, with expected outputs verified by a person. Grow it from production traces, especially from complaints, and version the set so scores remain comparable over time.

How do you test prompts when the model's output is not deterministic?

▾

Run each case more than once where the budget allows, report the spread, and treat changes inside the noise band as neither regression nor improvement. Structured outputs and exact checks reduce the variance for data-producing prompts.