# Calibrating a Judge Against Humans — RAG Retrieval & Evaluation

Source: https://www.skillbyai.com/en/rag-evaluation/j-calib

> Agreement percentages can be misleading.

## Measure agreement beyond chance

Before trusting a judge, have humans label a sample (100 or more) and compare. **Raw agreement** can look high just because most answers are good: a judge that says good to everything agrees 80% of the time when 80% are good. **Cohen's kappa** corrects for chance agreement, and the judge's **recall on bad answers** shows whether it catches the failures you care about. Iterate on the rubric until agreement is acceptable, then re-check periodically and whenever the judge model changes.

## Raw agreement versus kappa, run

I ran this with Python 3 (numpy 2.5.3 and scikit-learn 1.9.1 where imported) on small made-up data, with fixed seeds where random. On 100 made-up labels where humans mark 20 answers bad, the judge agrees 84% of the time, but chance agreement is 75%, so Cohen's kappa is only 0.35, and the judge catches just 30% of the bad answers.

```python
human = ["good"] * 80 + ["bad"] * 20
judge = ["good"] * 78 + ["bad"] * 2 + ["good"] * 14 + ["bad"] * 6
n = len(human)
agree = sum(h == j for h, j in zip(human, judge)) / n
p_h = human.count("good") / n; p_j = judge.count("good") / n
chance = p_h * p_j + (1 - p_h) * (1 - p_j)
kappa = (agree - chance) / (1 - chance)
caught = sum(h == "bad" and j == "bad" for h, j in zip(human, judge)) / human.count("bad")
print(f"raw agreement {agree:.2f} | chance agreement {chance:.2f} | Cohen kappa {kappa:.2f}")
print(f"judge catches {caught:.0%} of answers humans marked bad")
```

Output:

```
raw agreement 0.84 | chance agreement 0.75 | Cohen kappa 0.35
judge catches 30% of answers humans marked bad
```

## Focus on the failure class

Measure judge recall on the bad answers specifically; that is what a quality gate depends on.

**Quiz:** Why can 84% raw agreement be unimpressive?

- [x] When most answers are good, a lenient judge agrees often by chance
- [ ] 84% is always perfect
- [ ] Agreement cannot be measured
- [ ] Kappa is always higher than agreement

*Answer:* When most answers are good, a lenient judge agrees often by chance. Kappa and failure recall tell the real story.
