Lesson 19 / 25

Judge Biases and Mitigations

Known failure modes of model graders.

Position, length, self-preference, leniency

LLM judges show systematic biases: position bias in pairwise comparisons (preferring the first or second answer), length bias (preferring longer answers), self-preference (favouring outputs from the same model family), and leniency (rarely failing anything). Mitigations: swap the order of pairwise answers and count only consistent verdicts, use rubrics that penalise unnecessary length, use a different model family as judge where possible, ask narrow binary questions, and keep human spot-checks on a sample of every evaluation run.

Order-swapping for pairwise judging (sketch)

Count a win only if it survives both orders. Not run here.

def compare(question, a, b):
    v1 = judge(question, first=a, second=b)   # returns "first" or "second"
    v2 = judge(question, first=b, second=a)
    if v1 == "first" and v2 == "second": return "A"
    if v1 == "second" and v2 == "first": return "B"
    return "tie"   # inconsistent across orders -> position bias

Spot-check every run

Read 20 judged examples per evaluation run; drifting judges are found by reading, not by averages.

Quick check: How do you reduce position bias in pairwise judging?

  • Remove the rubric
  • Always put your system first
  • Use longer answers
  • Judge both orders and count only consistent verdicts
Answer

Judge both orders and count only consistent verdicts — Order-swapping exposes and cancels position effects.