पाठ 19 / 25
Human Review of Sampled Outputs
Read what users actually see.
Targeted plus random samples
Metrics cannot tell you everything about AI output quality, so set up a regular human review of sampled outputs. Combine targeted samples (thumbs-down, escalations, check failures, new slices) with random samples, because complaints over-represent some problems and miss others users did not notice. Use a rubric, track ratings over time, and feed findings into the evaluation set and prompt fixes. Respect privacy: restrict who can see user content and redact where possible.
Building a review queue, run
I ran this with Python 3 (scipy 1.18.1 where imported) on example numbers, not data from a real product. From 5,000 logged responses, where only 0.6% have a thumbs-down, the queue takes 20 thumbs-down items and 30 random items, so reviewers see both known complaints and unreported problems.
import random
rng = random.Random(1)
logs = ([{"id": i, "feedback": "down"} for i in range(30)] + [{"id": i, "feedback": "up"} for i in range(30, 400)]
+ [{"id": i, "feedback": None} for i in range(400, 5000)])
def review_sample(logs, n_down=20, n_random=30):
downs = [l for l in logs if l["feedback"] == "down"]
rest = [l for l in logs if l["feedback"] != "down"]
return rng.sample(downs, min(n_down, len(downs))) + rng.sample(rest, n_random)
s = review_sample(logs)
print("review queue size:", len(s))
print("thumbs-down items:", sum(l["feedback"] == "down" for l in s), "| random items:", sum(l["feedback"] != "down" for l in s))
print("share of thumbs-down in all traffic:", f"{30 / len(logs):.1%}")
Output:
review queue size: 50 thumbs-down items: 20 | random items: 30 share of thumbs-down in all traffic: 0.6%
Turn findings into eval cases
Every confirmed bad output reviewed by humans should become a test case in the offline evaluation set.
त्वरित जाँच: Why include random samples, not only thumbs-down items, in review?
- It reduces privacy
- Random samples are always bad
- Thumbs-down items are always wrong
- Many problems are never reported by users
Answer
Many problems are never reported by users — Complaints are a biased sample.