# The Decision Ladder: Cheapest Effective Step First — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/d-ladder

> Escalate only when an evaluation shows the cheaper step is not enough.

## Climb one rung at a time, with numbers

A sensible order: (1) define success and build a small **evaluation set**; (2) improve the **prompt** and add examples; (3) try a **stronger model**; (4) add **retrieval** if facts are the gap; (5) only then consider **fine-tuning**, ideally a smaller model, with enough high-quality examples. Each rung costs more effort and maintenance, so move up only when the evaluation shows the current one misses the target. Often the remaining errors after careful prompting are data or specification problems that fine-tuning would not fix.

## The ladder

Cost and maintenance rise as you go down.

```text
1. write an evaluation set + success metric       (hours)       free to change
2. improve prompt / examples / structure           (hours)       instant, reversible
3. try a stronger or different model                (hours)       per-token cost may rise
4. add retrieval (RAG)                              (days)        index to build and keep fresh
5. fine-tune (usually a smaller model)              (days-weeks)  data, training, evaluation, re-training

Stop at the first rung that meets the target.
```

## Keep the evaluation set

The same evaluation set that justifies fine-tuning later proves it worked; build it on rung one.

**Quiz:** Why try prompting before fine-tuning?

- [x] It is faster, cheaper and reversible, and often enough
- [ ] Fine-tuning is not allowed
- [ ] Prompts train the model
- [ ] There is no difference in effort

*Answer:* It is faster, cheaper and reversible, and often enough. Use the cheapest step that meets the target.
