Lesson 10 / 25
Slices: Averages Hide Weak Spots
Break results down by segment.
Report per segment
An overall metric can look healthy while one group of users is badly served. Tag every query with slices such as topic, language, query length, user type or document source, and report metrics per slice with its sample size. Small slices have wide intervals, so treat their numbers as warnings to investigate, not precise scores. Typical discoveries: non-English queries, very short queries, questions about tables or PDFs, and new product areas perform much worse than the average suggests.
An overall score hiding a weak slice, run
I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. Over 100 made-up queries the overall hit@5 is 0.88, but the 10 Hindi queries score 0.40 while billing and account queries score 0.92 and 0.95.
results = ( # (slice, hit@5) per query
[("billing", 1)] * 46 + [("billing", 0)] * 4 +
[("account", 1)] * 38 + [("account", 0)] * 2 +
[("hindi queries", 1)] * 4 + [("hindi queries", 0)] * 6)
total = sum(h for _, h in results) / len(results)
print(f"overall hit@5 = {total:.2f} over {len(results)} queries")
for s in ["billing", "account", "hindi queries"]:
rows = [h for name, h in results if name == s]
print(f" {s:<14} {sum(rows) / len(rows):.2f} (n={len(rows)})")
Output:
overall hit@5 = 0.88 over 100 queries billing 0.92 (n=50) account 0.95 (n=40) hindi queries 0.40 (n=10)
Set a floor per important slice
In release rules, require each key slice to stay above a minimum, not only the overall average.
Quick check: Why report metrics per slice?
- An average can hide a segment that performs badly
- Slices always have more queries
- It makes the overall score higher
- Slices remove sampling error
Answer
An average can hide a segment that performs badly — Look at the distribution, not only the mean.