SkillByAIOpen interactive version →

Lesson 19 / 25

Human Review of Sampled Outputs

Read what users actually see.

Targeted plus random samples

Metrics cannot tell you everything about AI output quality, so set up a regular human review of sampled outputs. Combine targeted samples (thumbs-down, escalations, check failures, new slices) with random samples, because complaints over-represent some problems and miss others users did not notice. Use a rubric, track ratings over time, and feed findings into the evaluation set and prompt fixes. Respect privacy: restrict who can see user content and redact where possible.

Building a review queue, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example numbers, not data from a real product. From 5,000 logged responses, where only 0.6% have a thumbs-down, the queue takes 20 thumbs-down items and 30 random items, so reviewers see both known complaints and unreported problems.

import random
rng = random.Random(1)
logs = ([{"id": i, "feedback": "down"} for i in range(30)] + [{"id": i, "feedback": "up"} for i in range(30, 400)]
        + [{"id": i, "feedback": None} for i in range(400, 5000)])
def review_sample(logs, n_down=20, n_random=30):
    downs = [l for l in logs if l["feedback"] == "down"]
    rest = [l for l in logs if l["feedback"] != "down"]
    return rng.sample(downs, min(n_down, len(downs))) + rng.sample(rest, n_random)
s = review_sample(logs)
print("review queue size:", len(s))
print("thumbs-down items:", sum(l["feedback"] == "down" for l in s), "| random items:", sum(l["feedback"] != "down" for l in s))
print("share of thumbs-down in all traffic:", f"{30 / len(logs):.1%}")

Output:

review queue size: 50
thumbs-down items: 20 | random items: 30
share of thumbs-down in all traffic: 0.6%

Turn findings into eval cases

Every confirmed bad output reviewed by humans should become a test case in the offline evaluation set.

Quick check: Why include random samples, not only thumbs-down items, in review?

  • It reduces privacy
  • Random samples are always bad
  • Thumbs-down items are always wrong
  • Many problems are never reported by users
Answer

Many problems are never reported by users — Complaints are a biased sample.