Lesson 20 / 27
Always Compare Against a Strong Prompted Baseline
The question is not whether tuning improved the base model, but whether it beats your best prompt.
Fair comparison or no comparison
A tuned model beating an untuned model with a lazy prompt proves little. Compare on the same held-out test set: (a) base model with your best prompt and examples, (b) base model plus retrieval if relevant, (c) the tuned model with a short prompt. Use the same scoring for all, include cost and latency, and report uncertainty (small test sets give noisy scores). Check subgroups and hard cases, not only the average, and read failures by hand. Ship the tuned model only when the gain is real and worth its maintenance.
Prove it, ship it, keep it
A tuned model is only worth shipping if it beats a well-prompted baseline on your own held-out data.
A comparison table template
Fill it with your own measured numbers; the values here are placeholders.
System quality cost/1k calls p95 latency notes
A base + best prompt + examples ___ ___ ___
B base + prompt + retrieval ___ ___ ___
C tuned + short prompt ___ ___ ___
D tuned + retrieval ___ ___ ___
same test set - same scorer - 95% interval reported - failures read by handReport the interval
With 100 test cases, a 3-point gain is often noise; report a confidence interval or run a paired comparison.
Quick check: What is the fair baseline for a tuned model?
- An untuned model with a one-word prompt
- The base model with your best prompt and examples on the same test set
- The training set accuracy
- A different task entirely
Answer
The base model with your best prompt and examples on the same test set — Beat your best alternative, not a strawman.