पाठ 12 / 25

Confusion Matrix, Precision and Recall

Why accuracy fails on imbalanced data.

Four kinds of outcome

A confusion matrix counts true negatives, false positives, false negatives and true positives. From it: precision = of the cases predicted positive, how many were positive (cost of false alarms); recall = of the actual positives, how many were caught (cost of misses). With imbalanced classes (fraud, disease, churn), accuracy misleads: always predicting "no" can be 95% accurate and useless. Choose metrics from the business cost of each error type, and report precision and recall (or F1) for the class you care about.

Accuracy versus recall on 5% positives, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. On a synthetic dataset with 5% positives, always predicting negative scores 0.947 accuracy with zero recall. Logistic regression reaches 0.99 accuracy, precision 0.978 and recall 0.83, with 9 positives still missed.

from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import confusion_matrix, precision_score, recall_score, accuracy_score
from sklearn.model_selection import train_test_split
X, y = make_classification(n_samples=4000, weights=[0.95], random_state=1)   # 5% positives
X_tr, X_te, y_tr, y_te = train_test_split(X, y, random_state=0, stratify=y)
always_no = [0] * len(y_te)
print("always predict negative: accuracy", round(accuracy_score(y_te, always_no), 3), "recall 0.0")
p = LogisticRegression().fit(X_tr, y_tr).predict(X_te)
print("logistic regression    : accuracy", round(accuracy_score(y_te, p), 3),
      "precision", round(precision_score(y_te, p), 3), "recall", round(recall_score(y_te, p), 3))
tn, fp, fn, tp = confusion_matrix(y_te, p).ravel()
print(f"confusion: TN={tn} FP={fp} FN={fn} TP={tp}")

Output:

always predict negative: accuracy 0.947 recall 0.0
logistic regression    : accuracy 0.99 precision 0.978 recall 0.83
confusion: TN=946 FP=1 FN=9 TP=44

Name the costly error

Decide with stakeholders whether a missed positive or a false alarm is worse before choosing the metric.

त्वरित जाँच: Recall answers which question?

  • Of predicted positives, how many were right?
  • Of all actual positives, how many did the model catch?
  • How fast is the model?
  • How many features are used?
Answer

Of all actual positives, how many did the model catch? — Recall measures misses.