पाठ 22 / 25
Rollback and Incident Response
Prepare the response before the incident.
Detect, mitigate, learn
When a model misbehaves (alerts fire, complaints spike), the first goal is mitigation: roll back the alias to the previous version, switch to a rule-based fallback, or disable the feature. Then diagnose using prediction logs, data checks and the change history (new model? new data? new upstream code?). Afterwards, write a blameless review and add tests or monitors so the same failure is caught earlier. Practise rollback before you need it.
A model incident runbook
Keep it next to the dashboard.
1 confirm: which metric, since when, which model version, which slices?
2 mitigate: move champion alias to previous version OR enable fallback rule
3 communicate: owners, affected teams, status updates
4 diagnose: data validation results, drift report, recent deploys, upstream changes
5 fix + verify on holdout + shadow
6 review: timeline, cause, new tests/monitors, runbook updatesAlways have a fallback
Keep a simple rule or the previous model ready so the product works, if less well, while you fix the model.
त्वरित जाँच: What is the first priority in a model incident?
- Wait for the weekly review
- Retrain a bigger model immediately
- Delete the logs
- Mitigate impact, for example by rolling back or using a fallback
Answer
Mitigate impact, for example by rolling back or using a fallback — Stop the harm first, then diagnose.