# Goodhart's Law: Overfitting to the Score — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/t-goodhart

> When a measure becomes a target.

## Keep the score honest

"When a measure becomes a target, it ceases to be a good measure." Repeatedly tweaking prompts to raise one test set's score overfits to that set: special-case wording for known inputs, verbose answers that please a length-biased judge, refusals that avoid failing strict checks. Protect against it: keep a **held-out** test set used rarely, refresh test sets with new production cases, combine several metrics (quality, length, refusals, cost), and confirm improvements with human review and live metrics.

## Teaching to the test

Students drilled on last year's exam questions can score well without learning the subject; a new exam reveals the gap.

## Keep a held-out set

Score candidate prompts on a hidden set only at release time; if gains disappear there, the prompt overfit.

**Quiz:** Which practice helps avoid overfitting prompts to a test set?

- [ ] Never reviewing outputs
- [ ] Tuning on the same 20 cases forever
- [ ] Using only one metric
- [x] Keeping a held-out test set and refreshing tests with new production cases

*Answer:* Keeping a held-out test set and refreshing tests with new production cases. Fresh and held-out data keep scores honest.
