Lesson 20 / 25
Evaluation as a Release Gate
Block changes that regress key metrics.
Thresholds per metric, run on every change
Run the evaluation set automatically on every change to chunking, embeddings, retrieval settings, prompts or models, and compare with the current production baseline. Define rules per metric: improvements welcome, small regressions within a tolerance, larger regressions block the release. Include quality (recall, faithfulness), cost and latency, and floors for important slices. Store results with versions so you can see trends. A gate turns evaluation from an occasional report into a habit.
Gate, observe, refresh
Turn evaluation into a release gate, then keep measuring with live signals.
A release gate comparing candidate and baseline, run
I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. The candidate improves recall@5 (0.82 to 0.86) and latency, and raises cost within tolerance, but faithfulness drops from 0.91 to 0.88, beyond the allowed 0.01, so the gate blocks the release. All numbers are example values.
baseline = {"recall@5": 0.82, "faithfulness": 0.91, "p95_latency_ms": 1800, "cost_per_query": 0.0040}
candidate = {"recall@5": 0.86, "faithfulness": 0.88, "p95_latency_ms": 1650, "cost_per_query": 0.0043}
rules = { # metric: (direction, allowed change)
"recall@5": ("higher", -0.01), "faithfulness": ("higher", -0.01),
"p95_latency_ms": ("lower", 200), "cost_per_query": ("lower", 0.0005)}
ok = True
for m, (direction, tol) in rules.items():
delta = candidate[m] - baseline[m]
passed = delta >= tol if direction == "higher" else delta <= tol
ok &= passed
print(f"{m:<15} {baseline[m]:>8} -> {candidate[m]:<8} {'PASS' if passed else 'FAIL'}")
print("release gate:", "PASS" if ok else "BLOCKED")
Output:
recall@5 0.82 -> 0.86 PASS faithfulness 0.91 -> 0.88 FAIL p95_latency_ms 1800 -> 1650 PASS cost_per_query 0.004 -> 0.0043 PASS release gate: BLOCKED
Keep tolerances above the noise
Set allowed changes larger than the run-to-run variation, or the gate will fail randomly and people will learn to ignore it.
Quick check: Why did the example gate block the release despite better recall?
- Recall improved too much
- Faithfulness regressed beyond its allowed tolerance
- Latency got worse
- Cost fell
Answer
Faithfulness regressed beyond its allowed tolerance — Guardrail metrics matter as much as the headline.