Lesson 8 / 25
How Many Queries Do You Need?
The width of the uncertainty shrinks slowly with more queries.
Uncertainty shrinks with the square root
A metric measured on a sample of queries has sampling error. For a rate such as hit@5 near 0.75, the 95% interval half-width is roughly 1.96 times the square root of p(1-p)/n: about plus or minus 0.17 with 25 queries and 0.085 with 100. Quadrupling the queries only halves the uncertainty. Practical consequence: with 50 queries, differences of a few points are invisible; either grow the set, use paired comparisons (next topic), or focus on large effects.
Uncertainty, pairing and slices
Small evaluation sets give noisy numbers; quantify uncertainty before declaring a winner.
Interval width versus number of queries, run
I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. Using the normal approximation for a rate near 0.75, the 95% half-width falls from 0.170 at 25 queries to 0.085 at 100 and 0.030 at 800.
import math
p = 0.75 # observed hit rate
print("queries | 95% interval half-width for a rate near 0.75")
for n in [25, 50, 100, 200, 400, 800]:
print(f"{n:>7} | +/- {1.96 * math.sqrt(p * (1 - p) / n):.3f}")
Output:
queries | 95% interval half-width for a rate near 0.75
25 | +/- 0.170
50 | +/- 0.120
100 | +/- 0.085
200 | +/- 0.060
400 | +/- 0.042
800 | +/- 0.030Always print n
Show the number of queries next to every metric; 0.80 on 20 queries and 0.80 on 2,000 queries are very different claims.
Quick check: Roughly how much does quadrupling the number of queries shrink the interval?
- It does not change it
- It divides it by four
- It halves it
- It doubles it
Answer
It halves it — Width scales with one over the square root of n.