# Offline Evaluation and the Launch Bar — Safe Rollout Plans for AI Features

Source: https://www.skillbyai.com/en/ai-feature-rollouts/p-eval

> Overall and per slice, with uncertainty.

## A bar for every important slice

Build an **evaluation set** of realistic inputs with expected behaviour (answers, labels or grading rubrics) covering important **slices**: user segments, languages, input types, and cases that should be declined. Set a launch bar for the overall pass rate and a **floor for each slice**, because a good average can hide a failing group. Report uncertainty: with small slices, a confidence interval shows how much the true rate could differ. Re-run the same set for every prompt or model change.

## Earn the right to ship

Before any user sees the feature, evaluation, safety testing and budgets must clear their bars.

![Three ideas: quality bar, safety bar, cost and latency budget.](assets/figures/ai-feature-rollouts/section-2-map.svg) — Figure 2.1 — Quality, safety and budget gates.

## Checking an overall bar and slice floors, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example numbers, not data from a real product. The overall pass rate is 0.868, above the 0.85 bar, but the Hindi slice is at 0.762, below its 0.80 floor, with a 95% lower bound of 0.659 on only 80 cases. The launch should wait for that slice to improve, or exclude Hindi users from the first stages.

```python
import math
results = {  # eval set results: (passed, total) per slice
    "billing questions": (184, 200), "technical questions": (171, 200),
    "Hindi questions": (61, 80), "unanswerable (should decline)": (44, 50)}
bar = {"overall": 0.85, "per_slice_floor": 0.80}
def wilson_low(p, n, z=1.96):   # lower end of a 95% interval for a pass rate
    return (p + z * z / (2 * n) - z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n))) / (1 + z * z / n)
passed = sum(p for p, _ in results.values()); total = sum(n for _, n in results.values())
print(f"overall pass rate {passed / total:.3f} (bar {bar['overall']}) -> {'OK' if passed / total >= bar['overall'] else 'BELOW'}")
for name, (p, n) in results.items():
    rate = p / n; low = wilson_low(rate, n)
    flag = "OK" if rate >= bar["per_slice_floor"] else "BELOW FLOOR"
    print(f"  {name:<30} {rate:.3f} (95% low {low:.3f}) {flag}")
```

Output:

```
overall pass rate 0.868 (bar 0.85) -> OK
  billing questions              0.920 (95% low 0.874) OK
  technical questions            0.855 (95% low 0.800) OK
  Hindi questions                0.762 (95% low 0.659) BELOW FLOOR
  unanswerable (should decline)  0.880 (95% low 0.762) OK
```

## Launch narrower if a slice fails

If one slice misses its floor, you can launch only to slices that pass while you improve the rest.

**Quiz:** Why set a floor per slice and not only an overall bar?

- [ ] Slices always score higher
- [x] A good overall average can hide a slice that performs badly
- [ ] Overall rates cannot be computed
- [ ] It reduces evaluation cost to zero

*Answer:* A good overall average can hide a slice that performs badly. Protect every important group.
