By Mert Dönmezler11 min read

Evaluating Language-Model Systems in Production: Beyond Benchmark Scores

How to evaluate a language-model feature inside a real product: a frozen task set from real traffic, measured rater agreement, controlled model judges and release gates.

  • llm-evaluation
  • ai-engineering
  • research
  • metrics
  • llm-as-judge
  • production-ai
  • prompt-engineering

Public benchmark scores say very little about whether a language-model feature will work inside a particular product. What you deploy is a whole system (prompt, retrieval, tools, validators, fallbacks), and it runs against a task distribution that no benchmark samples. So useful evaluation has to be built locally. It starts with a frozen task set drawn from real traffic, and with rubrics whose inter-rater agreement has been measured before any judge is trusted. Model-based judging can then be used with its known biases controlled, and the online metrics get rotated so that Goodhart drift doesn't set in. The article goes through assembling such a set, computing agreement on it, deciding when a judge can be used, and specifying a regression gate that is allowed to block a release.

What you deploy is a system

A benchmark measures a model on its own, on a distribution its authors chose. A production feature is built from many parts, and the model is only one of them. There is a retrieval step over a private corpus, in the pattern established by retrieval-augmented generation (Lewis et al., 2020), whose components carry their own architectural commitments. There is a prompt scaffold that may impose intermediate reasoning steps (Wei et al., 2022), a tool schema with its own ways of failing, decoding parameters, a validator, and a fallback for when the validator rejects twice. In our experience those components account for more of the variance than the choice between two comparable base models, and the difference between base models is the only thing a leaderboard reports.

Few-shot prompting made the prompt the main interface to these models (Brown et al., 2020), and it also made the prompt part of the thing under test. A prompt tuned against one model does not keep its measured quality on another, so swapping the model changes the system. Version the object of evaluation as a tuple: model identifier, prompt template, retrieval index snapshot, tool schema, decoding parameters, post-processing, judge version. If any element changes, the earlier numbers describe a different system.

Four AI agents with four different evaluation problems

BrewX, a venue operating system we have in development, contains four AI agents, and each one poses a different evaluation problem. Churn prediction and fraud detection make predictions whose ground truth arrives later, with class imbalance and asymmetric error costs. Those are supervised problems, and precision, recall and cost-weighted thresholds can answer them. The promotion and hype agent writes copy, and copy has no ground truth at all. All you have is rubric judgement plus a downstream behavioural outcome. Atmosphere scoring is harder still. An atmosphere score means whatever the people defining it decide it means, so the definition has to be agreed before anything is measured against it.

RelationCRM drafts messages in six tones across ten languages. Quality there means tone fidelity, grounding in the relationship record and fluency in each language, and a parser can decide none of these. It also shows a constraint that evaluation has to live with. Contacts are anonymised to identifiers of the form Contact-001 before any model call, in line with our GDPR and KVKK posture and the wider case for privacy by design in on-device AI. A failure that depends on a real name therefore can never show up in our sets.

A frozen task set drawn from real traffic

Build the set from logged traffic. Examples written by the team encode the author's idea of the task, which is the same idea that produced the prompt, so they leave out the inputs nobody anticipated, and those are the ones that fail. Between 150 and 300 real items with careful rubric labels are worth more than ten thousand unlabelled ones.

Sample by stratum rather than uniformly, for example by intent, language, input length, whether retrieval returned a relevant document, and known failure classes. Rare but costly cases have to be oversampled on purpose, and that changes the arithmetic, because the unweighted mean over an oversampled set no longer estimates production quality. Report a frequency-weighted score S = Σ wᵢ sᵢ, where sᵢ is the mean in stratum i and wᵢ that stratum's share of production traffic. Always publish the per-stratum scores next to it, because a weighted mean is exactly the number that hides a collapse in a small stratum.

Sampling biases inherited from logged traffic

Real traffic only samples what the current system already tolerates, and it is biased in several ways. Users whose first attempt failed often stopped asking, so the hardest intents are underrepresented in proportion to how badly they were served. That is survivorship. Requests that crashed or timed out before logging are missing entirely, which is a selection effect of the instrumentation. Distributions also drift over time. GUZZL rotates content weekly, so a set frozen before a rotation describes a distribution that no longer exists. Finally there is feedback: the deployed system trains its users, so the shape of the traffic is partly produced by the prompt under test.

The cost of freezing an evaluation set

A frozen set lets you compare versions, and it costs you three things. It goes stale as the distribution moves. The team gradually overfits to it, because they read its failures and tune against them. And once items leak into few-shot examples or training data, it is contaminated. To limit this, keep a sealed subset that is opened only for release decisions, rotate a fixed fraction on a schedule, and version the sets so they are appended to and never edited.

Rubrics and inter-rater agreement

A usable rubric is a set of dimensions that can fail independently: instruction compliance, factual grounding, tone fidelity, safety, format validity. Each gets an ordinal scale whose levels are defined by observable anchors, not by adjectives. A single overall five-point score cannot tell these failures apart. One rule saves a surprising amount of work: if a parser can decide it, write a test for it and keep it out of the evaluation.

Next, measure agreement between two human raters on the same items, corrected for chance. Cohen's kappa (Cohen, 1960) is the standard instrument for nominal and dichotomised scales: κ = (p_o − p_e) / (1 − p_e), with p_o the observed agreement and p_e the agreement expected from the raters' marginals alone.

Do the arithmetic once by hand. The numbers here are assumed for illustration and were not observed. Suppose 120 items scored pass or fail by two raters who agree on 96, so p_o = 0.80, and both pass 70% of items overall. Then p_e = (0.7 × 0.7) + (0.3 × 0.3) = 0.58 and κ ≈ 0.52. Raw agreement of 80% looks reassuring, but once chance is removed it is only moderate. Low kappa usually points to an underspecified rubric more than to careless raters, so the fix is sharper anchors. Sterner instructions to the raters change little. The loop goes like this: label fifty items independently, compute kappa per dimension, adjudicate every disagreement, rewrite the anchors that produced it, and repeat until agreement stabilises.

LLM-as-judge and its documented biases

Model-based judging is the only scoring method that scales to every commit. It is usable only under certain conditions, because judges have biases that are well documented and easy to reproduce.

The first is position bias. In pairwise comparisons the verdict depends on which candidate comes first, so score both orders, count only the verdicts that hold in both, and report the inconsistency rate as a property of the judge. Verbosity bias makes longer, hedged answers score higher regardless of content. Use length-controlled comparisons, and probe with a padded version of a response you have already scored. If padding raises the score, the judge is measuring length. Judges also prefer text from their own model family, so judge with a different family, or run two judges and send their disagreements to human review. That review step needs its own defences against automation bias in human-in-the-loop designs. On top of this, judges compress toward the middle of a scale and are sensitive to how the rubric is phrased.

A few design rules follow. Judge one dimension per call, since an overall score quietly averages dimensions with unknown weights. Make the judge name the anchor it applied, so its errors can be audited. Force a discrete label, and give the judge reference material instead of asking it to recall facts. Most important, evaluate the judge like any other system. The human-adjudicated items are its frozen set, and its kappa against the human gold label is its score. If a judge agrees with humans clearly less often than humans agree with each other, it is a noisier instrument than the raters it replaces. Put both numbers in the same report.

The offline–online gap and shadow evaluation

An offline set answers a narrow question: did this change alter behaviour on this distribution? It can't tell you whether users are better off. Shadow evaluation closes part of that gap. The candidate runs over live traffic without its output being served, and the results are logged for comparison with the incumbent. It costs inference budget and carries no risk that users can see. It is also the only cheap way to observe what a frozen set cannot contain: the real input distribution with its tail, real tool and retrieval failure rates, and real latency at real payload sizes.

Online comparison then measures behaviour instead of scores. The best proxies are the ones that show the user finishing the job: acceptance rate of a generated draft, edit distance between the drafted and the sent message in RelationCRM, retry rate, escalation rate to a human. Next to them sit guardrails that must not regress at all, such as p95 latency, cost per resolved task, safety violation rate and format failure rate. Forsgren, Humble and Kim (2018) argue for the same structure in delivery measurement: a few outcome metrics that deliberately balance each other, so that improving one at another's expense shows up and doesn't get rewarded.

Goodhart drift and metric rotation

Once a measure becomes a target, it stops being a good measure. The regularity is associated with Goodhart (1975), and Campbell (1979) stated it for social indicators: the more an indicator is used for decision-making, the more it will be corrupted. Language-model systems drift quickly, because the optimisation loop is a person editing a prompt with the metric on screen. Three mechanisms keep coming up. Prompts get tuned to the wording of the judge's rubric instead of to the task. The team memorises the frozen set and sometimes leaks it into the prompt as few-shot examples. And the output style drifts toward whatever the judge rewards, which is usually length.

The countermeasures need to be structural. Keep a sealed holdout for release decisions only, rotate a defined fraction of the set each quarter, keep a small adversarial subset written after the prompt was frozen, and change the judge's model family from time to time. Then watch a second-order number, which is how well the offline score predicts the online outcome. If offline scores keep improving and online results don't follow, that points to gaming, and a metric that no longer tracks the outcome should be retired.

Decision rules and regression gates

Five rules cover most cases. Don't bring in a judge until the humans agree, because weak human-human kappa means the rubric isn't finished. Anything a parser can decide goes into the tests. Gate on the worst stratum, since the mean can hide it. Treat the prompt as code and release it through the application's own pipeline. And fail closed on format, because a schema violation is a system error and shouldn't be scored as a weak answer.

for each candidate system version:
  run deterministic checks on frozen set   # schema, refusal, PII leakage
  if any hard check fails: block
  run judge per dimension, per stratum
  block if mean(dimension) < baseline − delta
        or min over strata < stratum_baseline − delta_s
        or p95_latency > budget or cost_per_task > budget
  else promote to shadow evaluation

The gate also has to respect the statistics. With 150 items, a two-point movement in a mean is noise, so state the smallest detectable effect before you choose delta, and prefer a paired comparison on identical items to two independent means. The other thing you can do is shrink the layer that needs judging. Our automated quality-control pipeline for a collectible card producer is the deterministic version of this. It has 90+ validation checkpoints covering roughly 95% of the QC process and cut an eight-hour manual pass to about fifteen minutes, and it works because every check is decidable. Make the deterministic layer as large as it can honestly be, and use judgement only where you can't avoid it.

Limitations

This is a design argument based on building these systems, and it is not an empirical study. We have not run a controlled comparison of judge and human agreement across our products, so reporting judge kappa next to human kappa is a method we recommend, without a measured result behind it. The kappa figures are arithmetic on assumed inputs, and any thresholds they suggest for kappa or gate deltas are conventions that don't come from our data. The bias mitigations come from published accounts of judge behaviour and our own inspection. We have not run ablations on our systems. BrewX and RelationCRM are in development, so there are no post-launch results behind the sections on shadow evaluation and drift. Finally, the argument is about bounded deployments shaped around a workflow. It does not cover open-ended assistants or frontier safety evaluation, where the distribution you care about can't be sampled from a product's traffic at all.

References

  • Brown, T. (2020). Language Models are Few-Shot Learners.
  • Campbell, D. (1979). Assessing the Impact of Planned Social Change.
  • Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales.
  • Forsgren, N., Humble, J., & Kim, G. (2018). Accelerate: The Science of Lean Software and DevOps.
  • Goodhart, C. (1975). Problems of Monetary Management: The U.K. Experience.
  • Lewis, P. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.
  • Wei, J. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.