Lesson 27 / 27
A Decision Checklist
A one-page way to decide and to review a proposal.
Ask these in order
Is the failure about facts, instructions or behaviour? Is there an evaluation set and a target? Has the best prompt, with caching and example selection, been measured? Would a stronger model or retrieval close the gap? Do you have enough clean, consistent, legally usable examples, split without leakage? Does the learning curve suggest more data will help? Does the break-even volume make sense with maintenance included? Have you planned LoRA or a hosted option, early stopping, a regression suite, safety re-tests, versioning and a fallback? If any answer is no, fix that before training.
The checklist
Use it in design reviews.
[ ] gap named: facts / instructions / behaviour
[ ] eval set + target written before any training
[ ] best prompt (+ caching, + selected examples) measured
[ ] stronger model and retrieval tried where relevant
[ ] clean, consistent, licensed data; group split; test frozen
[ ] learning curve still rising with data
[ ] break-even volume computed with upkeep
[ ] LoRA / hosted option, early stopping, low learning rate
[ ] regression + safety suite; privacy scan done
[ ] model card, monitoring, fallback, re-evaluation dateMake no a valid answer
A review that ends with stay on prompting for now is a success; the checklist exists to prevent unneeded training.
Quick check: What should you do if the checklist shows no evaluation set exists?
- Skip evaluation to save time
- Train first and decide later
- Use the training set as the test set
- Build one before training anything
Answer
Build one before training anything — Without a measure you cannot know whether tuning helped.