पाठ 1 / 25
How ML Systems Fail in Production
A good notebook score is not a working product.
Failure modes ordinary software does not have
An ML system can fail without any code error: the data changes (new customer behaviour, a new product, a changed upstream export), features are computed differently in production than in training, labels arrive late so nobody notices accuracy falling, the model cannot be reproduced because nobody recorded the data or settings, or a new version silently performs worse on an important group. MLOps applies DevOps ideas (automation, versioning, testing, monitoring) plus ML-specific practices (data validation, experiment tracking, model registries, drift monitoring, retraining) to make ML systems dependable.
From notebook to dependable system
Most ML effort fails not in modelling but in getting models into production and keeping them working.
Common production failures and their guards
Each failure has a matching practice in this course.
failure guard
"which data trained this model?" experiment tracking + data fingerprints
features computed differently in serving shared feature code, skew checks, feature store
broken upstream batch (nulls, units, columns) data validation before training and scoring
new model worse for one segment evaluation gates with slices
accuracy decays as the world changes drift + performance monitoring, retraining
bad release shadow, canary, fast rollbackList your silent failures
For every model, write down how it could become wrong without raising an error; each item needs a check or a monitor.
त्वरित जाँच: Which is a failure mode specific to ML systems?
- A missing semicolon
- A syntax error in Python
- Input data changes so the model becomes less accurate with no code error
- A server running out of disk
Answer
Input data changes so the model becomes less accurate with no code error — ML systems can fail silently through data.