SkillByAIOpen interactive version →

Lesson 1 / 25

How ML Systems Fail in Production

A good notebook score is not a working product.

Failure modes ordinary software does not have

An ML system can fail without any code error: the data changes (new customer behaviour, a new product, a changed upstream export), features are computed differently in production than in training, labels arrive late so nobody notices accuracy falling, the model cannot be reproduced because nobody recorded the data or settings, or a new version silently performs worse on an important group. MLOps applies DevOps ideas (automation, versioning, testing, monitoring) plus ML-specific practices (data validation, experiment tracking, model registries, drift monitoring, retraining) to make ML systems dependable.

From notebook to dependable system

Most ML effort fails not in modelling but in getting models into production and keeping them working.

Figure 1.1 — Failure modes, lifecycle and maturity.

Common production failures and their guards

Each failure has a matching practice in this course.

failure                                         guard
"which data trained this model?"                 experiment tracking + data fingerprints
features computed differently in serving         shared feature code, skew checks, feature store
broken upstream batch (nulls, units, columns)    data validation before training and scoring
new model worse for one segment                  evaluation gates with slices
accuracy decays as the world changes             drift + performance monitoring, retraining
bad release                                      shadow, canary, fast rollback

List your silent failures

For every model, write down how it could become wrong without raising an error; each item needs a check or a monitor.

Quick check: Which is a failure mode specific to ML systems?

  • A missing semicolon
  • A syntax error in Python
  • Input data changes so the model becomes less accurate with no code error
  • A server running out of disk
Answer

Input data changes so the model becomes less accurate with no code error — ML systems can fail silently through data.