Back

LLMOps: Evaluation, Observability, and Shipping Without Fear

Jun 15, 2026 (2mo ago)

LLMOps: Evaluation, Observability, and Shipping Without Fear

The demo always works. You paste a prompt into a notebook, the output is uncanny, you screenshot it, everyone is thrilled. Then someone edits that prompt three weeks later and nobody in the room can say whether the product got better or worse. That gap — between "works in my notebook" and "a team can change this safely" — is the whole job, and this post is about the harness, the gates, and the telemetry that close it.

Why classic MLOps doesn't transfer cleanly

If you've shipped traditional ML, you have muscle memory that will mislead you here.

  • No training loop to own. The artifact you ship isn't weights — it's a prompt, a schema, a tool set, and a model identifier. Your "retrain" button is a text edit, so the barrier to a risky change is roughly zero.
  • No fixed label set. Classification gave you accuracy for free. "Was this answer good?" has no closed form, so you have to manufacture the ground truth you measure against.
  • Non-deterministic outputs. The same input can produce different text across calls. Any gate you build has to separate a real regression from sampling noise, or it becomes a flaky test people learn to re-run.
  • The model is someone else's API. It can be updated, deprecated, rate-limited, or quietly re-tuned underneath you. You do not control your most important dependency.

That last point reframes everything. In classic MLOps you version the model and monitor the data. Here the data is user text you can't constrain and the model is a moving target you can't pin forever — so the only thing left to control tightly is your evaluation of the system's behaviour. Evals aren't a nice-to-have testing layer; they're the substitute for every guarantee you just lost.

Build the eval set from real traffic

The first instinct is to write fifty tidy questions. Don't. Synthetic questions encode the failure modes you already imagined — the ones your prompt already handles. Sources that actually work, in rough order of value:

  1. Production failures. Thumbs-down, escalations to a human, retries, abandoned sessions. Each is a case someone already cared about.
  2. Real traffic samples, stratified across intent, length, and language — not just the head. The long tail is where the embarrassing outputs live.
  3. Bug reports and support tickets. Free labels, already triaged by a human.
  4. Adversarial and boundary cases. Injection attempts, out-of-scope asks, empty input, requests the feature is meant to refuse.
  5. Synthetic cases — last, and only to fill a gap the four above revealed.

A case is a small record: the input, the context the system should have, the expected behaviour, and the graders that apply. Keep them as data, not code:

# evals/cases/billing/refund_window_expired.yaml
id: billing.refund_window_expired
tags: [billing, refusal, policy]
source: production_failure/2026-04-18
input:
  user: "I bought this 90 days ago, refund it."
  account_tier: standard
expect:
  must_refuse: true
  must_cite_policy_id: REFUND-30D
  forbidden_phrases: ["I've processed your refund", "approved"]
graders: [structural.json_schema, assert.contains_policy_id, judge.tone_rubric]

Golden sets, and how many is enough

Split cases into a golden set — small, hand-verified, stable, the thing you gate merges on — and a larger exploratory set you run less often and treat as a signal rather than a rule. The golden set earns its name by being reviewed by someone who knows the domain, not by being big.

The honest answer to "how many?" is derivable rather than measured: the set needs enough cases that one flipped case moves the score by less than your gate threshold. Gate on a 2% drop with 20 cases and a single flip is 5% — your build now fails on noise. Work backwards from the tolerance you can live with. In practice that pushes most teams from "a dozen" to "a few hundred", grouped into slices you score separately. Slice scores matter more than the aggregate: one number can hold steady while your multilingual cases fall off a cliff.

Version it in git, next to the code

The eval set belongs in the repo, in the same PR as the prompt change it justifies — not a spreadsheet, not a vendor dashboard, not somebody's notebook.

evals/
  cases/          # one YAML per case, foldered by slice
  graders/        # structural.py, assertions.py, judge.py (rubrics live here)
  baselines/      # main.json — the last accepted score report
  run.py

A reviewer sees the prompt diff and the case diff side by side, git blame says when a case was added and which incident it came from, and CI targets evals/ with a path filter.

Grading methods, and when each is right

There is no single grader. There's a ladder, and you climb only as far as you must — every rung up costs money, adds noise, or both.

Method Cost Reliability What it catches Use when
Structural assertions Free Deterministic Malformed JSON, missing fields, broken contracts Always — first line of defence
String / regex checks Free Deterministic, brittle Banned phrases, missing citations, wrong IDs Closed-form answers, refusal checks
Set and rank metrics Free Deterministic Retrieval regressions, dropped evidence You have labelled gold references
Embedding similarity Cheap Threshold-sensitive Gross topic drift only Coarse smoke tests, never a gate
LLM-as-judge Moderate Noisy, biased, calibratable Tone, helpfulness, groundedness Open-ended text with no reference
Trajectory grading Moderate Depends on grader Right answer reached the wrong way Multi-step runs
Human review Expensive Ground truth Everything, slowly Calibrating judges, release sign-off

Set and rank metrics are how you measure a retrieval-backed feature without opening up the retriever. Label which documents should have been in context, score recall@k over what actually arrived, then separately score groundedness by splitting the answer into claims and checking each against the retrieved context. That decomposes failure: low recall@k is a retrieval problem, high recall@k with low groundedness is a generation problem. You've routed the bug before anyone opens a debugger.

Trajectory grading applies when a run has steps. Score the final outcome and the path — right things called in a sane order, no loops, no ten steps doing one step's work. Outcome-only grading passes runs that stumbled into the right answer, and those break the moment anything shifts. And note where embedding similarity sits on that ladder: it tells you two texts are about the same topic, not that one of them says the opposite thing.

LLM-as-judge, honestly

Eventually you need to score "is this a good explanation?" and only a model can do it at volume. Fine — but know what you're buying. Judges have reproducible biases:

  • Position bias. In a pairwise comparison, judges favour one slot. Flip the order and the winner can change.
  • Verbosity bias. Longer, more confident, better-formatted answers score higher independent of correctness. Bullet points are worth free points.
  • Self-preference. A judge prefers text that looks like its own output. Judging your generator with the same model family bakes in a thumb on the scale.
  • Score compression. Ask for 1–10 and you'll get 7s and 8s forever, with no discriminating power exactly where you need it.

Mitigations, in order of payoff:

  1. Write a rubric with concrete anchors — not "rate helpfulness 1-10" but a short ordinal scale where each level names an observable property. Ask for reasoning before the verdict.
  2. Prefer pairwise over absolute. "Which is better, and why?" is far more stable than "score this out of ten."
  3. Swap positions and average. Run every comparison twice with candidates flipped; disagreement is a tie, and the disagreement rate is a health metric for the judge itself. Judge with a different model family than the generator while you're at it — that closes the cleanest path to self-preference.
  4. Calibrate against human labels. Grade a sample by hand, measure judge-human agreement. An uncalibrated judge produces a number, not evidence.
  5. Control for length. Log answer length next to the score; if they correlate across the suite, you're measuring verbosity.

A rubric that survives contact with a real suite is short and ordinal, and every level names something you could point at. Comparing two support answers against a policy context:

Score Awarded when
2 Every claim supported by the context, the policy cited by ID, and no action promised that the agent cannot actually perform
1 Substantially correct but flawed: a missing citation, or one unsupported minor claim
0 An unsupported claim, the wrong policy, or an action outside authority

Length, tone and formatting are named in the rubric as explicitly not criteria — a terse correct answer scores 2, a polished answer with one unsupported claim scores 0. The judge runs at temperature=0 on a pinned model from a different family to the generator, and returns {"reasoning": "two sentences", "winner": "x" or "y"} with the reasoning first. The position swap is the part worth writing down:

# evals/graders/judge.py — both orderings; disagreement means tie
def judge_pairwise(client, context, a, b):
    first  = _compare(client, context, a, b)
    second = _compare(client, context, b, a)   # positions swapped
    if first == second:
        return {"score": 1.0 if first == "a" else 0.0, "disagreed": False}
    return {"score": 0.5, "disagreed": True}

Track that disagreed rate across the suite. When it climbs, the judge has stopped discriminating and your scores mean less than they did last week.

Writing the eval harness

An eval suite is a test suite with fuzzy assertions and a score report. Same shape: cases in, graders applied, aggregate out. Keep the runner boring.

A grader is just (case, output) -> float in [0, 1]. That's the whole contract, and it's why schema checks, regexes and judges compose into one report — the runner keeps a registry of them by name and a case names the ones that apply.

Each Case is the YAML above loaded into a record: id, tags for its slices, inputs, expect, graders. The runner invokes the system with repeats: int = 3, because one sample of a non-deterministic system is an anecdote, and averages every grader across those samples. What comes back per case is scores (grader name to score, meaned into the case score with equal weights — weight them if you must), a median latency_ms, and cost_usd averaged over the repeats.

Aggregating is a group-by on tags, which is what makes slices free rather than a second pipeline. The suite report is a flat dict: overall, per-tag slices, p95_latency_ms, cost_usd_total, a failures list of every case scoring below 0.5, and the prompt_version and model that produced all of it.

The report is the deliverable, so make it readable in ten seconds:

eval report — prompt v14 · model pinned · 218 cases · 3 samples each
overall            0.871   (baseline 0.884, -0.013)   FAIL: drop exceeds gate
  billing          0.912   (baseline 0.905, +0.007)   ok
  onboarding       0.934   (baseline 0.930, +0.004)   ok
  adversarial      0.702   (baseline 0.781, -0.079)   FAIL
  multilingual     0.848   (baseline 0.851, -0.003)   ok
judge disagreement 6.4%    (baseline 5.9%)            ok
p95 latency / cost         within budget              ok
newly failing: adversarial.injection_via_pasted_doc, billing.refund_window_expired

Run it in CI, fail the build on a drop

A score you check when you remember to check it is not a gate. Wire it in where your unit tests live and treat a regression exactly like a failing test.

The workflow itself is four steps, and only the middle two are interesting:

Step What it does What it means for the build
Trigger Any pull_request touching prompts/**, evals/** or src/agent/**, plus every push to main Unrelated diffs never pay for the suite
Run golden suite python evals/run.py --suite golden --out report.json, with an EVAL_CACHE_DIR of .eval-cache Produces the candidate report
Gate on regression evals/compare.py reads evals/baselines/main.json against report.json with --max-overall-drop 0.01, --max-slice-drop 0.03, --hard-fail-cases evals/cases/adversarial, --warn-only-slices multilingual Fails on a drop past either threshold or any adversarial regression; multilingual only warns
Promote baseline Only when github.ref == 'refs/heads/main', and only after the gate passed report.json is committed as the new baseline

Gating slices separately is the step people leave out, and it's the one that catches the most: an aggregate can hold flat while one slice collapses.

Cache responses keyed on a hash of (prompt, model, inputs) so unchanged cases aren't re-billed on every push. Without a cache the suite gets slow and expensive, and the first thing a team does with a slow expensive check is skip it.

Gates and promotion

Not every signal deserves to block a merge. Be explicit, or people start merging past a red X out of habit.

Signal Verdict
Structural or schema check fails Blocks — a contract break, not a quality call
Any adversarial or safety case regresses Blocks. No threshold, no averaging
Overall score drops past tolerance Blocks
One slice drops past tolerance Blocks, even if overall is flat
p95 latency or cost per request past budget Blocks
Judge disagreement rate climbing Warns — investigate the judge, not the PR
Exploratory suite drifts Warns
Score improves sharply Warns — verify it isn't a grader returning 1.0 for everything

Passing CI still isn't the same as being safe, because your eval set is a sample and your users are not. Two patterns bridge the gap. Shadow: run the new prompt alongside the old on real traffic, serve only the old output, log both — real-distribution comparisons at zero user risk, paying only for the extra calls. It's how you find failure modes your eval set doesn't contain. Canary: route a small share of live traffic to the new version, keyed by a stable hash so a given user gets a consistent experience, and watch feedback rate, escalation rate, latency, and cost — not just your quality score. Ramp only when those hold, and keep rollback a config flip rather than a deploy. Prompt changes feel weightless, which is exactly why they need the rollout discipline of a schema migration.

Prompts and config as code

The single most expensive habit in this space: "we just tweaked the prompt in the dashboard." That sentence means there's a production behaviour with no diff, no reviewer, no test run, and no way to answer "what changed on Tuesday?" You will lose a week to it, bisecting a system that has no history to bisect.

The fix is unglamorous: everything that decides how the feature behaves is a handful of fields in a file the reviewer can see. A promptVersion of "v14", bumped in the PR that edits it, next to a promptPath of "prompts/support/v14.md" — because the prompt itself lives in the repo, not in a vendor's text box. A pinned model.id, a temperature of 0.2, maxOutputTokens of 800. Budgets declared in the same object: p95LatencyMs: 4000, costPerRequestUsd: 0.02. And a CONFIG_FINGERPRINT, hashed from the prompt version, model id and temperature, so the combination has a name.

  • Pin the model identifier explicitly. An alias that tracks "the current best" is a silent dependency upgrade in production. Pin, then upgrade deliberately: bump the pin in a PR, let the suite price the change, review the diff in the report.
  • Emit the fingerprint on every request. When quality moves, "which version produced this?" should be answerable from a log line.
  • Playgrounds are for drafting, not deploying. Iterate wherever you like, then land the change as a diff. If your vendor console can edit live behaviour, revoke that permission for production.
  • Re-run the suite on a schedule. Your code didn't change, but the model underneath it might have. A nightly run against a fixed pin is your canary for silent vendor drift.

Observability: trace the request, not the response

Logging the final output tells you that something was wrong. Tracing tells you where. Structure a request as a tree of spans — one root per user request, a child per model call, tool call, and retrieval — and give every span the same envelope.

Group Attributes on every span
Identity traceId — one per user request — plus spanId and parentSpanId
Shape kind: request, model, tool, retrieval or judge; name: "plan", "search_orders", "final_answer"
Outcome durationMs, and status of ok, error or timeout
Version promptVersion, modelId, configFingerprint
Spend inputTokens, outputTokens, costUsd, retryCount
Payload inputRef and outputRef pointers, and a structured error of { type, message }

Wrapping a call in a traced(kind, name, fn) helper is what makes that envelope automatic rather than aspirational: it opens the span with the current CONFIG_FINGERPRINT, marks timeout or error on the way out of a catch with the message redacted, and in a finally stamps the duration and emits — fire-and-forget, never blocking the response.

The parts people skip: version identity, without which a quality graph is uninterpretable; cost and tokens per span, because one runaway sub-call is invisible in a request-level total; latency at every level plus time-to-first-token where you stream, since a healthy request-level p95 can hide one tool call timing out for a specific slice of users; retries recorded as attributes rather than untracked extra calls, because "it succeeded" and "it succeeded on the third attempt at triple the cost" are different outcomes; and a typed failure taxonomy — timeout, rate limit, schema violation, refusal, empty output, tool error. Untyped errors can't be counted, and what you can't count you can't prioritise.

Sampling and PII

Capturing full payloads on every request inflates your storage bill and your privacy exposure at the same rate. Sample deliberately:

  • Keep metadata for 100% of requests — it's small, and it's what your dashboards run on. Keep payloads for a small random sample plus 100% of anything interesting: errors, low-scoring outputs, thumbs-down, escalations, cost outliers. Tail-based sampling, deciding after the trace completes, beats a fixed rate, because the traces you want are exactly the abnormal ones.
  • Redact at the emission boundary, in-process, before anything leaves. Redaction as a downstream cleanup job means the raw data already landed somewhere it shouldn't.
  • Store payloads separately from metadata, referenced by id, with shorter retention — deleting a user's content is then one delete, not a purge across your whole telemetry pipeline. Log a hash of the input alongside, so you can still group duplicate queries and spot repeat failures without keeping the text.

Closing the loop

The eval set isn't a deliverable you finish. It's a ratchet, and production is what turns it.

  1. Collect signal. Explicit feedback is sparse and skewed toward the angry. Implicit signal is richer — retries, edits to your output, abandonment, escalation, copy events. Instrument those; they're most of your data.
  2. Triage weekly, with a human. Read the lowest-scoring and most-complained-about traces and assign a cause: bad retrieval, bad instruction, missing tool, ambiguous input, genuine model limitation. Untriaged failures pile up until nobody opens the queue.
  3. Promote failures into the eval set, each with a source field pointing back at the incident. This is the ratchet: a bug you fix once can never silently return.
  4. Fix the class, not the instance. If three cases share a cause, one change fixes them, and the suite proves it.
  5. Prune. Cases every version has passed for months cost money and tell you nothing. Move them to the exploratory suite and keep the golden set sharp.

Health check: what fraction of your golden set came from real production failures? If it's low, you're grading yourself on an exam you wrote.

Cost and latency are SLOs, not footnotes

Quality is the metric everyone watches, which is why the other two regress unnoticed. A change that adds a reasoning preamble and three examples may score better while doubling token spend and pushing p95 past what your UI can hide. That's not obviously a good trade, and it should never be an implicit one.

Gate all three off the same report: cost per successful request (retries hide inside averages) with per-span attribution so you can see which call is expensive; latency at p50 and p95, plus time-to-first-token for streamed responses, because users experience the first token while dashboards usually measure the last; and quality per slice. Then state the trade-off in the PR description: "+1.8 on the billing slice, +40% tokens per request, p95 unchanged." A reviewer can accept or reject that. Nobody can review a number they never saw.

TL;DR checklist

  • Build the eval set from production failures and real traffic, not synthetic questions
  • Split golden (gates merges) from exploratory (informs you); size golden so one flipped case can't trip the gate
  • Version cases, graders, prompts, and baselines in git, in the same PR as the change
  • Climb the grader ladder only as far as needed: structural, assertions, metrics, judges, humans; repeat every case, because one sample of a non-deterministic system is an anecdote
  • Use rubrics with concrete anchors, prefer pairwise, swap positions, calibrate judges against human labels
  • Run the suite in CI, cache by input hash, fail the build on a regression, gate slices as well as the aggregate
  • Pin model identifiers explicitly; never ship a prompt edited in a dashboard
  • Shadow, then canary, then ramp — with rollback as a config flip
  • Trace every model and tool call with version, tokens, cost, latency; sample payloads and redact at the boundary
  • Triage failures weekly and promote each one back into the eval set
  • Gate cost and latency alongside quality, and put the trade-off in the PR description

None of this makes the model deterministic. It makes your knowledge of the model deterministic — you stop arguing about whether a change is better and start reading the number. That's the point: not shipping perfect systems, just shipping without fear.