Lesson 9 / 25

Paired Comparisons and the Bootstrap

Compare two systems on the same queries.

Pair queries, resample, read the interval

When comparing systems A and B, run both on the same queries and look at the per-query differences. Many queries are easy or hard for both, so pairing removes that shared noise and detects smaller differences than comparing two separate averages. A simple, general tool is the paired bootstrap: resample the queries with replacement thousands of times, compute the mean difference each time, and take the 2.5th and 97.5th percentiles as a 95% interval. If the interval excludes zero, the difference is unlikely to be sampling noise.

A paired bootstrap on 120 queries, run

I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. On simulated per-query hits for 120 queries, system B scores 0.692 versus A 0.758. The 95% paired bootstrap interval for B minus A is about -0.133 to -0.008, which excludes zero, so B is very likely worse. The data is simulated.

import random
random.seed(7)
n = 120  # queries
a = [1 if random.random() < 0.70 else 0 for _ in range(n)]           # system A hit@5 per query
b = [x if random.random() < 0.9 else 1 - x for x in a]                # system B: mostly same, some flips
diff = sum(b) / n - sum(a) / n
boots = []
for _ in range(5000):
    idx = [random.randrange(n) for _ in range(n)]
    boots.append(sum(b[i] - a[i] for i in idx) / n)
boots.sort()
lo, hi = boots[int(0.025 * 5000)], boots[int(0.975 * 5000)]
print(f"hit@5  A={sum(a)/n:.3f}  B={sum(b)/n:.3f}  difference={diff:+.3f}")
print(f"95% paired bootstrap interval for B-A: [{lo:+.3f}, {hi:+.3f}]")
print("interval includes 0 -> not a reliable difference" if lo <= 0 <= hi else "interval excludes 0 -> likely real")

Output:

hit@5  A=0.758  B=0.692  difference=-0.067
95% paired bootstrap interval for B-A: [-0.133, -0.008]
interval excludes 0 -> likely real

Keep the random seed and the query list

Store both with results so the comparison can be reproduced exactly.

Quick check: Why compare systems on the same queries (paired)?

  • It doubles the dataset size
  • It makes both systems faster
  • Shared difficulty cancels out, so smaller real differences become detectable
  • It removes the need for labels
Answer

Shared difficulty cancels out, so smaller real differences become detectable — Pairing reduces noise.