# Pairwise Comparisons and Ratings — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/v-pairwise

> Rank versions by head-to-head wins.

## Relative judgements are easier

People and judges are often better at saying which of two outputs is better than at giving absolute scores. Run **pairwise comparisons** between prompt versions on the same inputs, then turn win counts into **ratings** with an Elo-style or Bradley-Terry model, the method behind public model leaderboards. Ratings rank versions and estimate win probabilities. Combine with absolute checks so a ranking of three poor prompts is not mistaken for quality.

## Is the new prompt really better?

Rankings, paired tests and repeated sampling separate real improvements from noise.

![Three ideas: pairwise ratings, paired significance, variance.](assets/figures/prompt-quality-scoring/section-6-map.svg) — Figure 6.1 — Ratings, significance and variance.

## Elo-style ratings from head-to-head wins, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example or seeded simulated data, not results from a real product. From example comparison counts, v3 rates highest (1142), then v2 (1042), then v1 (816). v1 wins only 16% of its comparisons against v3, while v2 wins 42% against v3, a closer contest.

```python
import math
from itertools import combinations
wins = {("v1", "v2"): (12, 38), ("v1", "v3"): (8, 42), ("v2", "v3"): (21, 29)}   # (a wins, b wins)
ratings = {"v1": 1000.0, "v2": 1000.0, "v3": 1000.0}
games = []
for (a, b), (wa, wb) in wins.items():
    games += [(a, b)] * wa + [(b, a)] * wb
for _ in range(30):                                   # several passes for stable Elo-style estimates
    for w, l in games:
        e = 1 / (1 + 10 ** ((ratings[l] - ratings[w]) / 400))
        ratings[w] += 4 * (1 - e); ratings[l] -= 4 * (1 - e)
for name, r in sorted(ratings.items(), key=lambda x: -x[1]):
    print(f"{name}: rating {r:.0f}")
for (a, b), (wa, wb) in wins.items():
    print(f"{a} vs {b}: {a} wins {wa / (wa + wb):.0%}")
```

Output:

```
v3: rating 1142
v2: rating 1042
v1: rating 816
v1 vs v2: v1 wins 24%
v1 vs v3: v1 wins 16%
v2 vs v3: v2 wins 42%
```

## Compare on the same inputs

Pairwise judgements are only fair when both versions answered the same input.

**Quiz:** Why are pairwise comparisons often used for open-ended outputs?

- [ ] They remove all bias automatically
- [ ] They need no inputs
- [x] Choosing the better of two is easier and more consistent than absolute scoring
- [ ] They measure cost

*Answer:* Choosing the better of two is easier and more consistent than absolute scoring. Relative judgements are more reliable.
