पाठ 23 / 25
Model Upgrades and Prompt Changes
Every change is a mini-launch.
Re-run the gates
Providers release new models and retire old ones; teams tweak prompts constantly. Treat each change as a small launch: run the evaluation and red-team suites, compare with the current version on quality and on refusal rate, output length, latency and cost, then roll out through flags and stages. A newer model can score higher on your eval but refuse more, write longer (more expensive) answers or change tone. Plan ahead for retirement dates so you are not forced into an untested switch.
Launch is the beginning
Model updates and prompt changes need the same gates as the first launch.
Gating a model upgrade, run
I ran this with Python 3 (scipy 1.18.1 where imported) on example numbers, not data from a real product. A candidate model improves the pass rate (0.874 to 0.889) and latency, but its refusal rate rises to 5.8% (limit 4%) and its answers are about 53% longer, beyond the allowed 30% increase. The switch is blocked until the prompt is adjusted and the comparison re-run.
current = {"pass_rate": 0.874, "refusal_rate": 0.031, "avg_output_tokens": 340, "p95_ms": 3100}
candidate = {"pass_rate": 0.889, "refusal_rate": 0.058, "avg_output_tokens": 520, "p95_ms": 2700}
rules = {"pass_rate": ("min_delta", -0.005), "refusal_rate": ("max_abs", 0.04),
"avg_output_tokens": ("max_ratio", 1.3), "p95_ms": ("max_ratio", 1.2)}
ok = True
for m, (kind, v) in rules.items():
c, n = current[m], candidate[m]
good = (n - c >= v) if kind == "min_delta" else (n <= v) if kind == "max_abs" else (n / c <= v)
ok &= good
print(f"{m:<18} {c} -> {n} {'OK' if good else 'FAIL'}")
print("switch model version:", ok)
Output:
pass_rate 0.874 -> 0.889 OK refusal_rate 0.031 -> 0.058 FAIL avg_output_tokens 340 -> 520 FAIL p95_ms 3100 -> 2700 OK switch model version: False
Track provider deprecation dates
Put model retirement dates on the team calendar months ahead, so upgrades go through the normal gates.
त्वरित जाँच: Why can a model with a higher eval score still be blocked?
- Other guardrails, such as refusal rate or cost, got worse
- Higher scores are always fake
- New models cannot be used
- Eval scores do not matter
Answer
Other guardrails, such as refusal rate or cost, got worse — Check every guardrail, not only quality.