Lesson 24 / 25

Responsible Machine Learning

Models affect people; check how.

Fairness, transparency and monitoring

Models learn from historical data, including its biases. Before deploying a model that affects people (loans, hiring, healthcare, pricing), check performance per group (for example error rates by region, age band or gender where lawful and appropriate), question whether features act as proxies for protected attributes, document the intended use and limits, keep a human in the loop for high-stakes decisions, and comply with applicable law. After deployment, monitor: input data drifts, the world changes, and accuracy decays, so plan retraining and re-evaluation.

A model card outline

Document every model that affects people.

model: loan-default-v3          owner: risk-ml team
intended use: rank applications for human review (not automatic rejection)
data: 2022-2025 applications, region coverage noted, known gaps listed
metrics: overall recall@precision 0.8; per-group error rates reported
limitations: not validated for business loans; drifts with interest rates
monitoring: monthly drift report; retrain quarterly or on alert

Compare error rates by group

An average that looks fine can hide a much higher error rate for one group of people.

Quick check: Why monitor a model after deployment?

  • Monitoring trains the model
  • Models improve automatically
  • Data and the world change, so performance can decay
  • It is only needed for regression
Answer

Data and the world change, so performance can decay — Deployment is the start, not the end.