Lesson 23 / 25

Model Upgrades and Prompt Changes

Every change is a mini-launch.

Re-run the gates

Providers release new models and retire old ones; teams tweak prompts constantly. Treat each change as a small launch: run the evaluation and red-team suites, compare with the current version on quality and on refusal rate, output length, latency and cost, then roll out through flags and stages. A newer model can score higher on your eval but refuse more, write longer (more expensive) answers or change tone. Plan ahead for retirement dates so you are not forced into an untested switch.

Launch is the beginning

Model updates and prompt changes need the same gates as the first launch.

Three ideas: upgrades, case study, checklist.
Figure 8.1 — Upgrades, case study and checklist.

Gating a model upgrade, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example numbers, not data from a real product. A candidate model improves the pass rate (0.874 to 0.889) and latency, but its refusal rate rises to 5.8% (limit 4%) and its answers are about 53% longer, beyond the allowed 30% increase. The switch is blocked until the prompt is adjusted and the comparison re-run.

current = {"pass_rate": 0.874, "refusal_rate": 0.031, "avg_output_tokens": 340, "p95_ms": 3100}
candidate = {"pass_rate": 0.889, "refusal_rate": 0.058, "avg_output_tokens": 520, "p95_ms": 2700}
rules = {"pass_rate": ("min_delta", -0.005), "refusal_rate": ("max_abs", 0.04),
         "avg_output_tokens": ("max_ratio", 1.3), "p95_ms": ("max_ratio", 1.2)}
ok = True
for m, (kind, v) in rules.items():
    c, n = current[m], candidate[m]
    good = (n - c >= v) if kind == "min_delta" else (n <= v) if kind == "max_abs" else (n / c <= v)
    ok &= good
    print(f"{m:<18} {c} -> {n}  {'OK' if good else 'FAIL'}")
print("switch model version:", ok)

Output:

pass_rate          0.874 -> 0.889  OK
refusal_rate       0.031 -> 0.058  FAIL
avg_output_tokens  340 -> 520  FAIL
p95_ms             3100 -> 2700  OK
switch model version: False

Track provider deprecation dates

Put model retirement dates on the team calendar months ahead, so upgrades go through the normal gates.

Quick check: Why can a model with a higher eval score still be blocked?

  • Other guardrails, such as refusal rate or cost, got worse
  • Higher scores are always fake
  • New models cannot be used
  • Eval scores do not matter
Answer

Other guardrails, such as refusal rate or cost, got worse — Check every guardrail, not only quality.