Lesson 2 / 27
The Decision Ladder: Cheapest Effective Step First
Escalate only when an evaluation shows the cheaper step is not enough.
Climb one rung at a time, with numbers
A sensible order: (1) define success and build a small evaluation set; (2) improve the prompt and add examples; (3) try a stronger model; (4) add retrieval if facts are the gap; (5) only then consider fine-tuning, ideally a smaller model, with enough high-quality examples. Each rung costs more effort and maintenance, so move up only when the evaluation shows the current one misses the target. Often the remaining errors after careful prompting are data or specification problems that fine-tuning would not fix.
The ladder
Cost and maintenance rise as you go down.
1. write an evaluation set + success metric (hours) free to change
2. improve prompt / examples / structure (hours) instant, reversible
3. try a stronger or different model (hours) per-token cost may rise
4. add retrieval (RAG) (days) index to build and keep fresh
5. fine-tune (usually a smaller model) (days-weeks) data, training, evaluation, re-training
Stop at the first rung that meets the target.Keep the evaluation set
The same evaluation set that justifies fine-tuning later proves it worked; build it on rung one.
Quick check: Why try prompting before fine-tuning?
- It is faster, cheaper and reversible, and often enough
- Fine-tuning is not allowed
- Prompts train the model
- There is no difference in effort
Answer
It is faster, cheaper and reversible, and often enough — Use the cheapest step that meets the target.