Lesson 25 / 27
Versioning, Drift and Retirement
A tuned model is a product with a lifecycle, not a one-off file.
Version everything, plan the re-train
Record, for every tuned model: the base model and version, dataset version and hash, hyperparameters and seed, code commit, evaluation results and the date, so you can reproduce or roll back. Expect to re-train: providers retire base models, your data and users drift, and new base models may beat your tuned one with a plain prompt. Keep the evaluation set and pipeline automated so a re-train costs hours, not weeks. Monitor live quality, cost and format validity, sample outputs for review, and keep a documented rollback to the previous version or to the prompted fallback.
A model card record
Store this beside every released model.
model: tickets-router-v3
base: <base model id and version>
data: tickets-2026-09 (sha256 abc123..., 8,000 rows, group split)
method: LoRA r=8, lr=2e-4, 2 epochs, seed 42
code: <commit>
eval: held-out 600 tickets, report with 95% interval
safety: refusal + regression suite passed on <date>
fallback: prompted model <id>
owner/review: <name>, re-evaluate every quarter or on base-model retirementSchedule the re-evaluation
Put a quarterly re-run of the evaluation, including the newest untuned base models, on the calendar; sometimes retiring your tuned model is the right call.
Quick check: Why record dataset version and base model for each tuned model?
- Because regulators forbid logging
- To make the model smaller
- So the model can be reproduced, compared and rolled back
- It has no practical use
Answer
So the model can be reproduced, compared and rolled back — Reproducibility is what makes re-training safe.