SkillByAIOpen interactive version →

Lesson 15 / 25

Length Bias

Judges often like longer answers.

Check whether length predicts the score

LLM judges (and people) tend to rate longer, more elaborate answers higher, even when length does not add quality. If you then optimise prompts against such a judge, outputs grow longer and more expensive without getting better. Check by correlating judge scores with output length and comparing with the correlation between length and human-judged quality. Mitigations: include concision in the rubric, instruct the judge not to reward length, and compare length-matched outputs.

Detecting a length-biased judge, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example or seeded simulated data, not results from a real product. In 200 simulated outputs, length is only weakly related to true quality (+0.17), but strongly related to the judge's score (+0.45). The judge still tracks quality (+0.90), yet it also rewards length, which would push prompts toward longer answers.

import random, statistics
rng = random.Random(2)
rows = []
for _ in range(200):
    quality = rng.randint(1, 5)                       # true quality from expert labels
    length = rng.randint(50, 400)                     # answer length in words, unrelated to quality here
    judge = min(5, max(1, round(quality + (length - 225) / 150 + rng.gauss(0, 0.5))))
    rows.append((quality, length, judge))
def corr(x, y):
    mx, my = statistics.mean(x), statistics.mean(y)
    return sum((a - mx) * (b - my) for a, b in zip(x, y)) / (len(x) * statistics.pstdev(x) * statistics.pstdev(y))
q, l, j = zip(*rows)
print(f"correlation(length, true quality)  = {corr(l, q):+.2f}")
print(f"correlation(length, judge score)   = {corr(l, j):+.2f}  <- the judge rewards length")
print(f"correlation(judge, true quality)   = {corr(j, q):+.2f}")

Output:

correlation(length, true quality)  = +0.17
correlation(length, judge score)   = +0.45  <- the judge rewards length
correlation(judge, true quality)   = +0.90

Track output length alongside scores

Plot average length per prompt version; rising scores with rising length deserve a closer look.

Quick check: What happens if you optimise prompts against a length-biased judge?

  • Outputs get longer and costlier without real quality gains
  • Outputs always get shorter
  • Nothing changes
  • Costs drop automatically
Answer

Outputs get longer and costlier without real quality gains — Biased judges steer optimisation.