पाठ 17 / 25
What to Monitor
Three layers of signals.
Service, data and model health
Monitor three layers. Service: request rate, error rate, latency percentiles, resource use. Data: schema violations, missing-value rates, feature distributions versus training, volume and freshness of batches. Model: prediction distribution (share of positives, average score), confidence, and, when labels arrive, real quality metrics overall and per slice. Alert on clear breaches, review trends on dashboards, and link each model's monitors to its registry version.
System, data, model
Monitor services like any software, plus the data going in and the quality coming out.
A monitoring plan for one model
Example thresholds; tune them to your system.
layer signal alert when
service p99 latency > 150 ms for 10 min
service error rate > 1% for 5 min
data null rate in income > 2x training rate
data PSI per key feature > 0.2
model share predicted fraud outside 0.5% - 3%
model precision@top-100 (weekly labels) < 0.7
model per-region recall any region < 0.6Alert on few, review many
Page people only for breaches needing action; put the rest on dashboards reviewed weekly.
त्वरित जाँच: Which is a model-layer signal?
- Disk space
- CPU usage
- The share of positive predictions over time
- Number of deploys
Answer
The share of positive predictions over time — Prediction distributions reveal model changes early.