पाठ 17 / 25

What to Monitor

Three layers of signals.

Service, data and model health

Monitor three layers. Service: request rate, error rate, latency percentiles, resource use. Data: schema violations, missing-value rates, feature distributions versus training, volume and freshness of batches. Model: prediction distribution (share of positives, average score), confidence, and, when labels arrive, real quality metrics overall and per slice. Alert on clear breaches, review trends on dashboards, and link each model's monitors to its registry version.

System, data, model

Monitor services like any software, plus the data going in and the quality coming out.

Three ideas: what to monitor, drift, delayed labels.
Figure 6.1 — Signals, drift and delayed labels.

A monitoring plan for one model

Example thresholds; tune them to your system.

layer    signal                              alert when
service  p99 latency                         > 150 ms for 10 min
service  error rate                          > 1% for 5 min
data     null rate in income                 > 2x training rate
data     PSI per key feature                 > 0.2
model    share predicted fraud               outside 0.5% - 3%
model    precision@top-100 (weekly labels)   < 0.7
model    per-region recall                   any region < 0.6

Alert on few, review many

Page people only for breaches needing action; put the rest on dashboards reviewed weekly.

त्वरित जाँच: Which is a model-layer signal?

  • Disk space
  • CPU usage
  • The share of positive predictions over time
  • Number of deploys
Answer

The share of positive predictions over time — Prediction distributions reveal model changes early.