# Reference-Based Metrics — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/c-ref

> When a correct answer is known.

## Exact match, F1, numeric tolerance

For tasks with a known answer (classification labels, extracted fields, short facts), compare outputs with references: **exact match** after normalisation, **token F1** for short phrases, **numeric tolerance** for numbers, and **set overlap** for lists of items. For long free text, word-overlap metrics such as BLEU or ROUGE correlate poorly with quality, so prefer rubric or judge-based scoring there. Keep references reviewed; wrong references make good prompts look bad.

## Choosing a reference metric

Match the metric to the output type.

```text
output type                 metric
class label                 exact match, per-class precision/recall
extracted fields            exact match per field, normalised (case, spaces, dates)
short factual answer        exact match or token F1
numbers                     absolute/relative tolerance
list of items               precision/recall over the set
long free text              rubric or calibrated judge (not BLEU/ROUGE alone)
```

## Normalise before comparing

Lower-case, trim, standardise dates and numbers; otherwise formatting differences count as errors.

**Quiz:** Which metric suits extracted invoice totals?

- [ ] Word count
- [ ] BLEU score
- [x] Numeric comparison with a tolerance
- [ ] Human vibes

*Answer:* Numeric comparison with a tolerance. Compare numbers as numbers.
