SkillByAIOpen interactive version →

Lesson 17 / 25

Pairwise Comparisons and Ratings

Rank versions by head-to-head wins.

Relative judgements are easier

People and judges are often better at saying which of two outputs is better than at giving absolute scores. Run pairwise comparisons between prompt versions on the same inputs, then turn win counts into ratings with an Elo-style or Bradley-Terry model, the method behind public model leaderboards. Ratings rank versions and estimate win probabilities. Combine with absolute checks so a ranking of three poor prompts is not mistaken for quality.

Is the new prompt really better?

Rankings, paired tests and repeated sampling separate real improvements from noise.

Figure 6.1 — Ratings, significance and variance.

Elo-style ratings from head-to-head wins, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example or seeded simulated data, not results from a real product. From example comparison counts, v3 rates highest (1142), then v2 (1042), then v1 (816). v1 wins only 16% of its comparisons against v3, while v2 wins 42% against v3, a closer contest.

import math
from itertools import combinations
wins = {("v1", "v2"): (12, 38), ("v1", "v3"): (8, 42), ("v2", "v3"): (21, 29)}   # (a wins, b wins)
ratings = {"v1": 1000.0, "v2": 1000.0, "v3": 1000.0}
games = []
for (a, b), (wa, wb) in wins.items():
    games += [(a, b)] * wa + [(b, a)] * wb
for _ in range(30):                                   # several passes for stable Elo-style estimates
    for w, l in games:
        e = 1 / (1 + 10 ** ((ratings[l] - ratings[w]) / 400))
        ratings[w] += 4 * (1 - e); ratings[l] -= 4 * (1 - e)
for name, r in sorted(ratings.items(), key=lambda x: -x[1]):
    print(f"{name}: rating {r:.0f}")
for (a, b), (wa, wb) in wins.items():
    print(f"{a} vs {b}: {a} wins {wa / (wa + wb):.0%}")

Output:

v3: rating 1142
v2: rating 1042
v1: rating 816
v1 vs v2: v1 wins 24%
v1 vs v3: v1 wins 16%
v2 vs v3: v2 wins 42%

Compare on the same inputs

Pairwise judgements are only fair when both versions answered the same input.

Quick check: Why are pairwise comparisons often used for open-ended outputs?

  • They remove all bias automatically
  • They need no inputs
  • Choosing the better of two is easier and more consistent than absolute scoring
  • They measure cost
Answer

Choosing the better of two is easier and more consistent than absolute scoring — Relative judgements are more reliable.