Lesson 18 / 25

Classification Metrics

Accuracy is not enough.

Confusion matrix, precision, recall, F1

The confusion matrix counts correct and wrong predictions per class. Precision asks "of the cases predicted positive, how many were?", recall asks "of the actual positives, how many did we find?", and F1 balances them. With imbalanced classes or unequal costs (missing a malignant tumour is worse than a false alarm), choose metrics and thresholds that reflect those costs rather than accuracy alone.

Evaluating a breast cancer classifier, run

I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). On 171 test cases, 5 malignant tumours were predicted benign and 2 benign as malignant; malignant recall is 0.922 even though overall accuracy is 0.959.

from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report, confusion_matrix
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

X, y = load_breast_cancer(return_X_y=True)   # 0 = malignant, 1 = benign
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=1, stratify=y)
clf = make_pipeline(StandardScaler(), LogisticRegression()).fit(X_tr, y_tr)
pred = clf.predict(X_te)
print(confusion_matrix(y_te, pred))
print(classification_report(y_te, pred, target_names=["malignant", "benign"], digits=3))

Output:

[[ 59   5]
 [  2 105]]
              precision    recall  f1-score   support

   malignant      0.967     0.922     0.944        64
      benign      0.955     0.981     0.968       107

    accuracy                          0.959       171
   macro avg      0.961     0.952     0.956       171
weighted avg      0.959     0.959     0.959       171

Tune the decision threshold

Use predict_proba and choose a threshold that meets the recall or precision your use case needs instead of the default 0.5.

Quick check: Which metric measures how many actual positives the model found?

  • Recall
  • Precision
  • Accuracy
  • Support
Answer

Recall — Recall = found positives / all positives.