Lesson 21 / 27
Choosing Metrics for the Task
Match the measure to what the output must do.
Exact, structural, judged
Pick the cheapest metric that reflects success. Exact or structured checks suit classification and extraction (accuracy, precision/recall, field-level match, schema validity). Programmatic checks suit code (tests pass) and formats (parses). Human or model judges suit open-ended writing, but judges have biases (position, length, self-preference), so calibrate them against human labels on a sample. Track format validity, refusal rate, length, cost and latency beside quality. A single headline number hides regressions.
Metric by task
A starting menu; adapt it.
classification accuracy, per-class precision/recall, confusion matrix
extraction field-level exact match, schema validity, missing/extra fields
code generation unit tests pass, lints clean, runs within time
summarisation human or calibrated-judge rating, factual-consistency checks
chat / support resolution rate, escalation rate, sampled human review
all of them format validity, refusal rate, length, cost, p95 latencyCalibrate the judge
If a model grades outputs, compare its grades with 50 human-labelled cases and swap answer order to expose position bias.
Quick check: Why track format validity and cost beside quality?
- They replace quality
- A single headline score can hide regressions in other dimensions
- They are required by LoRA
- They improve the loss
Answer
A single headline score can hide regressions in other dimensions — Look at several dimensions together.