पाठ 4 / 25
Offline Evaluation and the Launch Bar
Overall and per slice, with uncertainty.
A bar for every important slice
Build an evaluation set of realistic inputs with expected behaviour (answers, labels or grading rubrics) covering important slices: user segments, languages, input types, and cases that should be declined. Set a launch bar for the overall pass rate and a floor for each slice, because a good average can hide a failing group. Report uncertainty: with small slices, a confidence interval shows how much the true rate could differ. Re-run the same set for every prompt or model change.
Earn the right to ship
Before any user sees the feature, evaluation, safety testing and budgets must clear their bars.
Checking an overall bar and slice floors, run
I ran this with Python 3 (scipy 1.18.1 where imported) on example numbers, not data from a real product. The overall pass rate is 0.868, above the 0.85 bar, but the Hindi slice is at 0.762, below its 0.80 floor, with a 95% lower bound of 0.659 on only 80 cases. The launch should wait for that slice to improve, or exclude Hindi users from the first stages.
import math
results = { # eval set results: (passed, total) per slice
"billing questions": (184, 200), "technical questions": (171, 200),
"Hindi questions": (61, 80), "unanswerable (should decline)": (44, 50)}
bar = {"overall": 0.85, "per_slice_floor": 0.80}
def wilson_low(p, n, z=1.96): # lower end of a 95% interval for a pass rate
return (p + z * z / (2 * n) - z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n))) / (1 + z * z / n)
passed = sum(p for p, _ in results.values()); total = sum(n for _, n in results.values())
print(f"overall pass rate {passed / total:.3f} (bar {bar['overall']}) -> {'OK' if passed / total >= bar['overall'] else 'BELOW'}")
for name, (p, n) in results.items():
rate = p / n; low = wilson_low(rate, n)
flag = "OK" if rate >= bar["per_slice_floor"] else "BELOW FLOOR"
print(f" {name:<30} {rate:.3f} (95% low {low:.3f}) {flag}")
Output:
overall pass rate 0.868 (bar 0.85) -> OK billing questions 0.920 (95% low 0.874) OK technical questions 0.855 (95% low 0.800) OK Hindi questions 0.762 (95% low 0.659) BELOW FLOOR unanswerable (should decline) 0.880 (95% low 0.762) OK
Launch narrower if a slice fails
If one slice misses its floor, you can launch only to slices that pass while you improve the rest.
त्वरित जाँच: Why set a floor per slice and not only an overall bar?
- Slices always score higher
- A good overall average can hide a slice that performs badly
- Overall rates cannot be computed
- It reduces evaluation cost to zero
Answer
A good overall average can hide a slice that performs badly — Protect every important group.