पाठ 8 / 25
Reference-Based Metrics
When a correct answer is known.
Exact match, F1, numeric tolerance
For tasks with a known answer (classification labels, extracted fields, short facts), compare outputs with references: exact match after normalisation, token F1 for short phrases, numeric tolerance for numbers, and set overlap for lists of items. For long free text, word-overlap metrics such as BLEU or ROUGE correlate poorly with quality, so prefer rubric or judge-based scoring there. Keep references reviewed; wrong references make good prompts look bad.
Choosing a reference metric
Match the metric to the output type.
output type metric
class label exact match, per-class precision/recall
extracted fields exact match per field, normalised (case, spaces, dates)
short factual answer exact match or token F1
numbers absolute/relative tolerance
list of items precision/recall over the set
long free text rubric or calibrated judge (not BLEU/ROUGE alone)Normalise before comparing
Lower-case, trim, standardise dates and numbers; otherwise formatting differences count as errors.
त्वरित जाँच: Which metric suits extracted invoice totals?
- Word count
- BLEU score
- Numeric comparison with a tolerance
- Human vibes
Answer
Numeric comparison with a tolerance — Compare numbers as numbers.