# Choosing Metrics for the Task — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/o-metrics

> Match the measure to what the output must do.

## Exact, structural, judged

Pick the cheapest metric that reflects success. **Exact or structured checks** suit classification and extraction (accuracy, precision/recall, field-level match, schema validity). **Programmatic checks** suit code (tests pass) and formats (parses). **Human or model judges** suit open-ended writing, but judges have biases (position, length, self-preference), so calibrate them against human labels on a sample. Track **format validity**, **refusal rate**, **length**, **cost** and **latency** beside quality. A single headline number hides regressions.

## Metric by task

A starting menu; adapt it.

```text
classification      accuracy, per-class precision/recall, confusion matrix
extraction          field-level exact match, schema validity, missing/extra fields
code generation     unit tests pass, lints clean, runs within time
summarisation       human or calibrated-judge rating, factual-consistency checks
chat / support      resolution rate, escalation rate, sampled human review
all of them         format validity, refusal rate, length, cost, p95 latency
```

## Calibrate the judge

If a model grades outputs, compare its grades with 50 human-labelled cases and swap answer order to expose position bias.

**Quiz:** Why track format validity and cost beside quality?

- [ ] They replace quality
- [x] A single headline score can hide regressions in other dimensions
- [ ] They are required by LoRA
- [ ] They improve the loss

*Answer:* A single headline score can hide regressions in other dimensions. Look at several dimensions together.
