# Paired Comparisons and the Bootstrap — RAG Retrieval & Evaluation

Source: https://www.skillbyai.com/en/rag-evaluation/s-boot

> Compare two systems on the same queries.

## Pair queries, resample, read the interval

When comparing systems A and B, run both on the **same queries** and look at the **per-query differences**. Many queries are easy or hard for both, so pairing removes that shared noise and detects smaller differences than comparing two separate averages. A simple, general tool is the **paired bootstrap**: resample the queries with replacement thousands of times, compute the mean difference each time, and take the 2.5th and 97.5th percentiles as a 95% interval. If the interval excludes zero, the difference is unlikely to be sampling noise.

## A paired bootstrap on 120 queries, run

I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. On simulated per-query hits for 120 queries, system B scores 0.692 versus A 0.758. The 95% paired bootstrap interval for B minus A is about -0.133 to -0.008, which excludes zero, so B is very likely worse. The data is simulated.

```python
import random
random.seed(7)
n = 120  # queries
a = [1 if random.random() < 0.70 else 0 for _ in range(n)]           # system A hit@5 per query
b = [x if random.random() < 0.9 else 1 - x for x in a]                # system B: mostly same, some flips
diff = sum(b) / n - sum(a) / n
boots = []
for _ in range(5000):
    idx = [random.randrange(n) for _ in range(n)]
    boots.append(sum(b[i] - a[i] for i in idx) / n)
boots.sort()
lo, hi = boots[int(0.025 * 5000)], boots[int(0.975 * 5000)]
print(f"hit@5  A={sum(a)/n:.3f}  B={sum(b)/n:.3f}  difference={diff:+.3f}")
print(f"95% paired bootstrap interval for B-A: [{lo:+.3f}, {hi:+.3f}]")
print("interval includes 0 -> not a reliable difference" if lo <= 0 <= hi else "interval excludes 0 -> likely real")
```

Output:

```
hit@5  A=0.758  B=0.692  difference=-0.067
95% paired bootstrap interval for B-A: [-0.133, -0.008]
interval excludes 0 -> likely real
```

## Keep the random seed and the query list

Store both with results so the comparison can be reproduced exactly.

**Quiz:** Why compare systems on the same queries (paired)?

- [ ] It doubles the dataset size
- [ ] It makes both systems faster
- [x] Shared difficulty cancels out, so smaller real differences become detectable
- [ ] It removes the need for labels

*Answer:* Shared difficulty cancels out, so smaller real differences become detectable. Pairing reduces noise.
