Lesson 26 / 27

Common Pitfalls

Learn the mistakes before you make them.

Eight mistakes that waste weeks

(1) Fine-tuning before trying a good prompt. (2) Fine-tuning to add fast-changing facts. (3) No evaluation set, or one contaminated by training data. (4) Noisy, inconsistent or tiny datasets. (5) Comparing against a lazy baseline. (6) Tuning too long or too aggressively, causing overfitting and forgetting. (7) Ignoring safety, privacy and licence checks on the data. (8) Treating the result as finished: no versioning, no monitoring, no plan for base-model retirement. Each has a cheap preventive habit, all of which appear earlier in this course.

Avoid the common mistakes, then decide

Most failed fine-tuning projects repeat the same few mistakes.

Two ideas: pitfalls, checklist.
Figure 8.1 — Pitfalls and checklist.

Pitfall to fix

Map each mistake to its habit.

no good prompt first          -> climb the ladder, build the eval set first
facts that change             -> retrieval, not weights
contaminated test set         -> group split, freeze the test set
noisy data                    -> read 50 rows, write rules, measure agreement
lazy baseline                 -> best prompt + caching + retrieval comparison
overfitting / forgetting      -> few epochs, low rate, LoRA, replay, regression suite
data risk                     -> scan, licence, provider terms, safety re-test
no lifecycle                  -> model card, monitoring, scheduled re-evaluation

Keep a mistakes log

Add every failed run and its cause to a shared log; the same cause repeats across teams.

Quick check: Which is a common fine-tuning mistake?

  • Skipping a strong prompted baseline and an evaluation set
  • Reading sample rows of the data
  • Saving model versions
  • Planning for rollback
Answer

Skipping a strong prompted baseline and an evaluation set — Evidence against a fair baseline is the core discipline.