पाठ 18 / 25
Calibrating a Judge Against Humans
Agreement percentages can be misleading.
Measure agreement beyond chance
Before trusting a judge, have humans label a sample (100 or more) and compare. Raw agreement can look high just because most answers are good: a judge that says good to everything agrees 80% of the time when 80% are good. Cohen's kappa corrects for chance agreement, and the judge's recall on bad answers shows whether it catches the failures you care about. Iterate on the rubric until agreement is acceptable, then re-check periodically and whenever the judge model changes.
Raw agreement versus kappa, run
I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. On 100 made-up labels where humans mark 20 answers bad, the judge agrees 84% of the time, but chance agreement is 75%, so Cohen's kappa is only 0.35, and the judge catches just 30% of the bad answers.
human = ["good"] * 80 + ["bad"] * 20
judge = ["good"] * 78 + ["bad"] * 2 + ["good"] * 14 + ["bad"] * 6
n = len(human)
agree = sum(h == j for h, j in zip(human, judge)) / n
p_h = human.count("good") / n; p_j = judge.count("good") / n
chance = p_h * p_j + (1 - p_h) * (1 - p_j)
kappa = (agree - chance) / (1 - chance)
caught = sum(h == "bad" and j == "bad" for h, j in zip(human, judge)) / human.count("bad")
print(f"raw agreement {agree:.2f} | chance agreement {chance:.2f} | Cohen kappa {kappa:.2f}")
print(f"judge catches {caught:.0%} of answers humans marked bad")
Output:
raw agreement 0.84 | chance agreement 0.75 | Cohen kappa 0.35 judge catches 30% of answers humans marked bad
Focus on the failure class
Measure judge recall on the bad answers specifically; that is what a quality gate depends on.
त्वरित जाँच: Why can 84% raw agreement be unimpressive?
- When most answers are good, a lenient judge agrees often by chance
- 84% is always perfect
- Agreement cannot be measured
- Kappa is always higher than agreement
Answer
When most answers are good, a lenient judge agrees often by chance — Kappa and failure recall tell the real story.