Lesson 20 / 25

Evaluation as a Release Gate

Block changes that regress key metrics.

Thresholds per metric, run on every change

Run the evaluation set automatically on every change to chunking, embeddings, retrieval settings, prompts or models, and compare with the current production baseline. Define rules per metric: improvements welcome, small regressions within a tolerance, larger regressions block the release. Include quality (recall, faithfulness), cost and latency, and floors for important slices. Store results with versions so you can see trends. A gate turns evaluation from an occasional report into a habit.

Gate, observe, refresh

Turn evaluation into a release gate, then keep measuring with live signals.

Three ideas: gate, online signals, drift.
Figure 7.1 — Gate, online signals and drift.

A release gate comparing candidate and baseline, run

I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. The candidate improves recall@5 (0.82 to 0.86) and latency, and raises cost within tolerance, but faithfulness drops from 0.91 to 0.88, beyond the allowed 0.01, so the gate blocks the release. All numbers are example values.

baseline = {"recall@5": 0.82, "faithfulness": 0.91, "p95_latency_ms": 1800, "cost_per_query": 0.0040}
candidate = {"recall@5": 0.86, "faithfulness": 0.88, "p95_latency_ms": 1650, "cost_per_query": 0.0043}
rules = {  # metric: (direction, allowed change)
    "recall@5": ("higher", -0.01), "faithfulness": ("higher", -0.01),
    "p95_latency_ms": ("lower", 200), "cost_per_query": ("lower", 0.0005)}
ok = True
for m, (direction, tol) in rules.items():
    delta = candidate[m] - baseline[m]
    passed = delta >= tol if direction == "higher" else delta <= tol
    ok &= passed
    print(f"{m:<15} {baseline[m]:>8} -> {candidate[m]:<8} {'PASS' if passed else 'FAIL'}")
print("release gate:", "PASS" if ok else "BLOCKED")

Output:

recall@5            0.82 -> 0.86     PASS
faithfulness        0.91 -> 0.88     FAIL
p95_latency_ms      1800 -> 1650     PASS
cost_per_query     0.004 -> 0.0043   PASS
release gate: BLOCKED

Keep tolerances above the noise

Set allowed changes larger than the run-to-run variation, or the gate will fail randomly and people will learn to ignore it.

Quick check: Why did the example gate block the release despite better recall?

  • Recall improved too much
  • Faithfulness regressed beyond its allowed tolerance
  • Latency got worse
  • Cost fell
Answer

Faithfulness regressed beyond its allowed tolerance — Guardrail metrics matter as much as the headline.