पाठ 8 / 25

Reference-Based Metrics

When a correct answer is known.

Exact match, F1, numeric tolerance

For tasks with a known answer (classification labels, extracted fields, short facts), compare outputs with references: exact match after normalisation, token F1 for short phrases, numeric tolerance for numbers, and set overlap for lists of items. For long free text, word-overlap metrics such as BLEU or ROUGE correlate poorly with quality, so prefer rubric or judge-based scoring there. Keep references reviewed; wrong references make good prompts look bad.

Choosing a reference metric

Match the metric to the output type.

output type                 metric
class label                 exact match, per-class precision/recall
extracted fields            exact match per field, normalised (case, spaces, dates)
short factual answer        exact match or token F1
numbers                     absolute/relative tolerance
list of items               precision/recall over the set
long free text              rubric or calibrated judge (not BLEU/ROUGE alone)

Normalise before comparing

Lower-case, trim, standardise dates and numbers; otherwise formatting differences count as errors.

त्वरित जाँच: Which metric suits extracted invoice totals?

  • Word count
  • BLEU score
  • Numeric comparison with a tolerance
  • Human vibes
Answer

Numeric comparison with a tolerance — Compare numbers as numbers.