# Human Review of Sampled Outputs — Safe Rollout Plans for AI Features

Source: https://www.skillbyai.com/en/ai-feature-rollouts/m-review

> Read what users actually see.

## Targeted plus random samples

Metrics cannot tell you everything about AI output quality, so set up a regular **human review** of sampled outputs. Combine **targeted** samples (thumbs-down, escalations, check failures, new slices) with **random** samples, because complaints over-represent some problems and miss others users did not notice. Use a rubric, track ratings over time, and feed findings into the evaluation set and prompt fixes. Respect privacy: restrict who can see user content and redact where possible.

## Building a review queue, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example numbers, not data from a real product. From 5,000 logged responses, where only 0.6% have a thumbs-down, the queue takes 20 thumbs-down items and 30 random items, so reviewers see both known complaints and unreported problems.

```python
import random
rng = random.Random(1)
logs = ([{"id": i, "feedback": "down"} for i in range(30)] + [{"id": i, "feedback": "up"} for i in range(30, 400)]
        + [{"id": i, "feedback": None} for i in range(400, 5000)])
def review_sample(logs, n_down=20, n_random=30):
    downs = [l for l in logs if l["feedback"] == "down"]
    rest = [l for l in logs if l["feedback"] != "down"]
    return rng.sample(downs, min(n_down, len(downs))) + rng.sample(rest, n_random)
s = review_sample(logs)
print("review queue size:", len(s))
print("thumbs-down items:", sum(l["feedback"] == "down" for l in s), "| random items:", sum(l["feedback"] != "down" for l in s))
print("share of thumbs-down in all traffic:", f"{30 / len(logs):.1%}")
```

Output:

```
review queue size: 50
thumbs-down items: 20 | random items: 30
share of thumbs-down in all traffic: 0.6%
```

## Turn findings into eval cases

Every confirmed bad output reviewed by humans should become a test case in the offline evaluation set.

**Quiz:** Why include random samples, not only thumbs-down items, in review?

- [ ] It reduces privacy
- [ ] Random samples are always bad
- [ ] Thumbs-down items are always wrong
- [x] Many problems are never reported by users

*Answer:* Many problems are never reported by users. Complaints are a biased sample.
