Lesson 19 / 25
Judge Biases and Mitigations
Known failure modes of model graders.
Position, length, self-preference, leniency
LLM judges show systematic biases: position bias in pairwise comparisons (preferring the first or second answer), length bias (preferring longer answers), self-preference (favouring outputs from the same model family), and leniency (rarely failing anything). Mitigations: swap the order of pairwise answers and count only consistent verdicts, use rubrics that penalise unnecessary length, use a different model family as judge where possible, ask narrow binary questions, and keep human spot-checks on a sample of every evaluation run.
Order-swapping for pairwise judging (sketch)
Count a win only if it survives both orders. Not run here.
def compare(question, a, b):
v1 = judge(question, first=a, second=b) # returns "first" or "second"
v2 = judge(question, first=b, second=a)
if v1 == "first" and v2 == "second": return "A"
if v1 == "second" and v2 == "first": return "B"
return "tie" # inconsistent across orders -> position biasSpot-check every run
Read 20 judged examples per evaluation run; drifting judges are found by reading, not by averages.
Quick check: How do you reduce position bias in pairwise judging?
- Remove the rubric
- Always put your system first
- Use longer answers
- Judge both orders and count only consistent verdicts
Answer
Judge both orders and count only consistent verdicts — Order-swapping exposes and cancels position effects.