# Judge Biases and Mitigations — RAG Retrieval & Evaluation

Source: https://www.skillbyai.com/en/rag-evaluation/j-bias

> Known failure modes of model graders.

## Position, length, self-preference, leniency

LLM judges show systematic **biases**: **position bias** in pairwise comparisons (preferring the first or second answer), **length bias** (preferring longer answers), **self-preference** (favouring outputs from the same model family), and **leniency** (rarely failing anything). Mitigations: swap the order of pairwise answers and count only consistent verdicts, use rubrics that penalise unnecessary length, use a different model family as judge where possible, ask narrow binary questions, and keep human spot-checks on a sample of every evaluation run.

## Order-swapping for pairwise judging (sketch)

Count a win only if it survives both orders. Not run here.

```python
def compare(question, a, b):
    v1 = judge(question, first=a, second=b)   # returns "first" or "second"
    v2 = judge(question, first=b, second=a)
    if v1 == "first" and v2 == "second": return "A"
    if v1 == "second" and v2 == "first": return "B"
    return "tie"   # inconsistent across orders -> position bias
```

## Spot-check every run

Read 20 judged examples per evaluation run; drifting judges are found by reading, not by averages.

**Quiz:** How do you reduce position bias in pairwise judging?

- [ ] Remove the rubric
- [ ] Always put your system first
- [ ] Use longer answers
- [x] Judge both orders and count only consistent verdicts

*Answer:* Judge both orders and count only consistent verdicts. Order-swapping exposes and cancels position effects.
