SkillByAIOpen interactive version →

Lesson 20 / 27

Always Compare Against a Strong Prompted Baseline

The question is not whether tuning improved the base model, but whether it beats your best prompt.

Fair comparison or no comparison

A tuned model beating an untuned model with a lazy prompt proves little. Compare on the same held-out test set: (a) base model with your best prompt and examples, (b) base model plus retrieval if relevant, (c) the tuned model with a short prompt. Use the same scoring for all, include cost and latency, and report uncertainty (small test sets give noisy scores). Check subgroups and hard cases, not only the average, and read failures by hand. Ship the tuned model only when the gain is real and worth its maintenance.

Prove it, ship it, keep it

A tuned model is only worth shipping if it beats a well-prompted baseline on your own held-out data.

Figure 6.1 — Baseline, metrics, serving and versions.

A comparison table template

Fill it with your own measured numbers; the values here are placeholders.

System                              quality   cost/1k calls   p95 latency   notes
A base + best prompt + examples      ___       ___             ___
B base + prompt + retrieval          ___       ___             ___
C tuned + short prompt               ___       ___             ___
D tuned + retrieval                  ___       ___             ___

same test set - same scorer - 95% interval reported - failures read by hand

Report the interval

With 100 test cases, a 3-point gain is often noise; report a confidence interval or run a paired comparison.

Quick check: What is the fair baseline for a tuned model?

  • An untuned model with a one-word prompt
  • The base model with your best prompt and examples on the same test set
  • The training set accuracy
  • A different task entirely
Answer

The base model with your best prompt and examples on the same test set — Beat your best alternative, not a strawman.