Lesson 22 / 25

Rollback and Incident Response

Prepare the response before the incident.

Detect, mitigate, learn

When a model misbehaves (alerts fire, complaints spike), the first goal is mitigation: roll back the alias to the previous version, switch to a rule-based fallback, or disable the feature. Then diagnose using prediction logs, data checks and the change history (new model? new data? new upstream code?). Afterwards, write a blameless review and add tests or monitors so the same failure is caught earlier. Practise rollback before you need it.

A model incident runbook

Keep it next to the dashboard.

1 confirm: which metric, since when, which model version, which slices?
2 mitigate: move champion alias to previous version  OR  enable fallback rule
3 communicate: owners, affected teams, status updates
4 diagnose: data validation results, drift report, recent deploys, upstream changes
5 fix + verify on holdout + shadow
6 review: timeline, cause, new tests/monitors, runbook updates

Always have a fallback

Keep a simple rule or the previous model ready so the product works, if less well, while you fix the model.

Quick check: What is the first priority in a model incident?

  • Wait for the weekly review
  • Retrain a bigger model immediately
  • Delete the logs
  • Mitigate impact, for example by rolling back or using a fallback
Answer

Mitigate impact, for example by rolling back or using a fallback — Stop the harm first, then diagnose.