Lesson 15 / 25
Length Bias
Judges often like longer answers.
Check whether length predicts the score
LLM judges (and people) tend to rate longer, more elaborate answers higher, even when length does not add quality. If you then optimise prompts against such a judge, outputs grow longer and more expensive without getting better. Check by correlating judge scores with output length and comparing with the correlation between length and human-judged quality. Mitigations: include concision in the rubric, instruct the judge not to reward length, and compare length-matched outputs.
Detecting a length-biased judge, run
I ran this with Python 3 (scipy 1.18.1 where imported) on example or seeded simulated data, not results from a real product. In 200 simulated outputs, length is only weakly related to true quality (+0.17), but strongly related to the judge's score (+0.45). The judge still tracks quality (+0.90), yet it also rewards length, which would push prompts toward longer answers.
import random, statistics
rng = random.Random(2)
rows = []
for _ in range(200):
quality = rng.randint(1, 5) # true quality from expert labels
length = rng.randint(50, 400) # answer length in words, unrelated to quality here
judge = min(5, max(1, round(quality + (length - 225) / 150 + rng.gauss(0, 0.5))))
rows.append((quality, length, judge))
def corr(x, y):
mx, my = statistics.mean(x), statistics.mean(y)
return sum((a - mx) * (b - my) for a, b in zip(x, y)) / (len(x) * statistics.pstdev(x) * statistics.pstdev(y))
q, l, j = zip(*rows)
print(f"correlation(length, true quality) = {corr(l, q):+.2f}")
print(f"correlation(length, judge score) = {corr(l, j):+.2f} <- the judge rewards length")
print(f"correlation(judge, true quality) = {corr(j, q):+.2f}")
Output:
correlation(length, true quality) = +0.17 correlation(length, judge score) = +0.45 <- the judge rewards length correlation(judge, true quality) = +0.90
Track output length alongside scores
Plot average length per prompt version; rising scores with rising length deserve a closer look.
Quick check: What happens if you optimise prompts against a length-biased judge?
- Outputs get longer and costlier without real quality gains
- Outputs always get shorter
- Nothing changes
- Costs drop automatically
Answer
Outputs get longer and costlier without real quality gains — Biased judges steer optimisation.