पाठ 19 / 25
Delayed Labels and Proxy Metrics
When you cannot see accuracy right away.
Measure what you can, when you can
Often the true outcome arrives late (did the loan default? did the customer churn in 90 days?) or never (fraud that was blocked). Until labels arrive, monitor proxies: prediction distributions, input drift, and early business signals (chargebacks, complaints, manual review overturn rates). Join labels back to stored predictions as they arrive and compute real metrics per cohort. For blocked actions, keep a small randomised holdout or human review sample so you can still estimate quality.
A prediction log that makes later evaluation possible
Store this for every prediction.
prediction_id, timestamp, model_version, input_features (or a reference),
score, decision, threshold, request_context (channel, region)
-- later --
label, label_timestamp (joined by prediction_id)
=> metrics by week, version and slice once labels arriveLog the model version with every prediction
Without it you cannot tell whether a drop in quality came from a new model or from new data.
त्वरित जाँच: What can you monitor before true labels arrive?
- The training loss
- Final accuracy only
- Nothing at all
- Input drift, prediction distributions and early business signals
Answer
Input drift, prediction distributions and early business signals — Proxies bridge the label delay.