Lesson 21 / 27

Choosing Metrics for the Task

Match the measure to what the output must do.

Exact, structural, judged

Pick the cheapest metric that reflects success. Exact or structured checks suit classification and extraction (accuracy, precision/recall, field-level match, schema validity). Programmatic checks suit code (tests pass) and formats (parses). Human or model judges suit open-ended writing, but judges have biases (position, length, self-preference), so calibrate them against human labels on a sample. Track format validity, refusal rate, length, cost and latency beside quality. A single headline number hides regressions.

Metric by task

A starting menu; adapt it.

classification      accuracy, per-class precision/recall, confusion matrix
extraction          field-level exact match, schema validity, missing/extra fields
code generation     unit tests pass, lints clean, runs within time
summarisation       human or calibrated-judge rating, factual-consistency checks
chat / support      resolution rate, escalation rate, sampled human review
all of them         format validity, refusal rate, length, cost, p95 latency

Calibrate the judge

If a model grades outputs, compare its grades with 50 human-labelled cases and swap answer order to expose position bias.

Quick check: Why track format validity and cost beside quality?

  • They replace quality
  • A single headline score can hide regressions in other dimensions
  • They are required by LoRA
  • They improve the loss
Answer

A single headline score can hide regressions in other dimensions — Look at several dimensions together.