Lesson 8 / 25

How Many Queries Do You Need?

The width of the uncertainty shrinks slowly with more queries.

Uncertainty shrinks with the square root

A metric measured on a sample of queries has sampling error. For a rate such as hit@5 near 0.75, the 95% interval half-width is roughly 1.96 times the square root of p(1-p)/n: about plus or minus 0.17 with 25 queries and 0.085 with 100. Quadrupling the queries only halves the uncertainty. Practical consequence: with 50 queries, differences of a few points are invisible; either grow the set, use paired comparisons (next topic), or focus on large effects.

Uncertainty, pairing and slices

Small evaluation sets give noisy numbers; quantify uncertainty before declaring a winner.

Three ideas: sample size, paired comparison, slices.
Figure 3.1 — Sample size, pairing and slices.

Interval width versus number of queries, run

I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. Using the normal approximation for a rate near 0.75, the 95% half-width falls from 0.170 at 25 queries to 0.085 at 100 and 0.030 at 800.

import math
p = 0.75  # observed hit rate
print("queries | 95% interval half-width for a rate near 0.75")
for n in [25, 50, 100, 200, 400, 800]:
    print(f"{n:>7} | +/- {1.96 * math.sqrt(p * (1 - p) / n):.3f}")

Output:

queries | 95% interval half-width for a rate near 0.75
     25 | +/- 0.170
     50 | +/- 0.120
    100 | +/- 0.085
    200 | +/- 0.060
    400 | +/- 0.042
    800 | +/- 0.030

Always print n

Show the number of queries next to every metric; 0.80 on 20 queries and 0.80 on 2,000 queries are very different claims.

Quick check: Roughly how much does quadrupling the number of queries shrink the interval?

  • It does not change it
  • It divides it by four
  • It halves it
  • It doubles it
Answer

It halves it — Width scales with one over the square root of n.