Lesson 3 / 27

When Fine-Tuning Is the Right Tool

Recognise the situations where training pays off.

Style, format, narrow skills, cost and latency

Fine-tuning pays off for consistent style or format that a prompt follows only most of the time, narrow repeated tasks (classification, extraction, routing) where a small model can match a large prompted one, shorter prompts (behaviour learned in training no longer needs long instructions), lower latency from a smaller model, and distilling a large model into a small one. It works poorly for learning new facts reliably, fast-changing information, tasks with few or inconsistent examples, and anything you cannot evaluate.

Good and poor candidates

Quick screening before you invest.

GOOD candidates                                          POOR candidates
classify tickets into 12 categories (5k labelled)         answer questions about last week's policy change
extract invoice fields into a strict JSON schema          "be smarter" in general
write replies in our brand voice (1k approved replies)    tasks with 20 inconsistent examples
cut a 3,500-token few-shot prompt to 400 tokens           anything with no evaluation set
distill an expensive model for one task                   a knowledge base that changes daily (use RAG)

Count your examples honestly

Hundreds of consistent, reviewed examples beat thousands of messy ones; if you cannot produce them, stay on the ladder.

Quick check: Which task is a good fit for fine-tuning?

  • A task with no way to measure success
  • Answering questions about this morning's news
  • Learning facts that change daily
  • Extracting invoice fields into a strict schema with thousands of labelled examples
Answer

Extracting invoice fields into a strict schema with thousands of labelled examples — Narrow, repeated, measurable tasks with plentiful examples suit training.