SkillByAIOpen interactive version →

Lesson 19 / 25

Delayed Labels and Proxy Metrics

When you cannot see accuracy right away.

Measure what you can, when you can

Often the true outcome arrives late (did the loan default? did the customer churn in 90 days?) or never (fraud that was blocked). Until labels arrive, monitor proxies: prediction distributions, input drift, and early business signals (chargebacks, complaints, manual review overturn rates). Join labels back to stored predictions as they arrive and compute real metrics per cohort. For blocked actions, keep a small randomised holdout or human review sample so you can still estimate quality.

A prediction log that makes later evaluation possible

Store this for every prediction.

prediction_id, timestamp, model_version, input_features (or a reference),
score, decision, threshold, request_context (channel, region)
-- later --
label, label_timestamp  (joined by prediction_id)
=> metrics by week, version and slice once labels arrive

Log the model version with every prediction

Without it you cannot tell whether a drop in quality came from a new model or from new data.

Quick check: What can you monitor before true labels arrive?

  • The training loss
  • Final accuracy only
  • Nothing at all
  • Input drift, prediction distributions and early business signals
Answer

Input drift, prediction distributions and early business signals — Proxies bridge the label delay.