# Rollback and Incident Response — MLOps

Source: https://www.skillbyai.com/en/mlops/t-incident

> Prepare the response before the incident.

## Detect, mitigate, learn

When a model misbehaves (alerts fire, complaints spike), the first goal is **mitigation**: roll back the alias to the previous version, switch to a rule-based fallback, or disable the feature. Then diagnose using prediction logs, data checks and the change history (new model? new data? new upstream code?). Afterwards, write a blameless review and add tests or monitors so the same failure is caught earlier. Practise rollback before you need it.

## A model incident runbook

Keep it next to the dashboard.

```text
1 confirm: which metric, since when, which model version, which slices?
2 mitigate: move champion alias to previous version  OR  enable fallback rule
3 communicate: owners, affected teams, status updates
4 diagnose: data validation results, drift report, recent deploys, upstream changes
5 fix + verify on holdout + shadow
6 review: timeline, cause, new tests/monitors, runbook updates
```

## Always have a fallback

Keep a simple rule or the previous model ready so the product works, if less well, while you fix the model.

**Quiz:** What is the first priority in a model incident?

- [ ] Wait for the weekly review
- [ ] Retrain a bigger model immediately
- [ ] Delete the logs
- [x] Mitigate impact, for example by rolling back or using a fallback

*Answer:* Mitigate impact, for example by rolling back or using a fallback. Stop the harm first, then diagnose.
