Lesson 14 / 25

Calibrating a Judge Against Humans

Agreement, systematic bias and weighted kappa.

Measure before trusting

Have humans grade a sample (often 100 or more outputs), run the judge on the same outputs, and compare. For ordinal scales, weighted kappa credits near-misses (4 versus 5) more than big misses (1 versus 5) and corrects for chance. Also check systematic bias: a judge may be consistently more generous, which matters when comparing against absolute thresholds even if rankings agree. Recalibrate when the judge model, prompt or task changes.

Exact agreement versus weighted kappa, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example or seeded simulated data, not results from a real product. On 20 example ratings the judge matches humans exactly only half the time, but its quadratic weighted kappa is 0.84 because disagreements are mostly by one point. It is on average 0.50 points more generous, so absolute thresholds need adjusting.

from collections import Counter
human = [5, 4, 4, 2, 1, 3, 5, 4, 2, 3, 4, 5, 1, 2, 4, 3, 5, 4, 3, 2]
judge = [5, 5, 4, 3, 2, 3, 5, 5, 2, 4, 4, 5, 2, 2, 5, 4, 5, 5, 3, 3]
def weighted_kappa(a, b, k=5):
    n = len(a); w = lambda i, j: (i - j) ** 2 / (k - 1) ** 2
    obs = sum(w(x, y) for x, y in zip(a, b)) / n
    ca, cb = Counter(a), Counter(b)
    exp = sum(ca[i] * cb[j] * w(i, j) for i in range(1, k + 1) for j in range(1, k + 1)) / (n * n)
    return 1 - obs / exp
exact = sum(x == y for x, y in zip(human, judge)) / len(human)
bias = sum(y - x for x, y in zip(human, judge)) / len(human)
print(f"exact agreement {exact:.2f} | quadratic weighted kappa {weighted_kappa(human, judge):.2f}")
print(f"judge is on average {bias:+.2f} points more generous than humans")

Output:

exact agreement 0.50 | quadratic weighted kappa 0.84
judge is on average +0.50 points more generous than humans

Check agreement on failures specifically

A judge can agree well overall yet miss many of the bad outputs; measure its recall on the outputs humans failed.

Quick check: What does quadratic weighted kappa reward compared with exact agreement?

  • It measures speed
  • Only identical ratings count
  • It ignores chance agreement
  • Near-miss ratings count as partial agreement, corrected for chance
Answer

Near-miss ratings count as partial agreement, corrected for chance — Weighted agreement suits ordinal scales.