Lesson 18 / 25
Classification Metrics
Accuracy is not enough.
Confusion matrix, precision, recall, F1
The confusion matrix counts correct and wrong predictions per class. Precision asks "of the cases predicted positive, how many were?", recall asks "of the actual positives, how many did we find?", and F1 balances them. With imbalanced classes or unequal costs (missing a malignant tumour is worse than a false alarm), choose metrics and thresholds that reflect those costs rather than accuracy alone.
Evaluating a breast cancer classifier, run
I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). On 171 test cases, 5 malignant tumours were predicted benign and 2 benign as malignant; malignant recall is 0.922 even though overall accuracy is 0.959.
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report, confusion_matrix
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = load_breast_cancer(return_X_y=True) # 0 = malignant, 1 = benign
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=1, stratify=y)
clf = make_pipeline(StandardScaler(), LogisticRegression()).fit(X_tr, y_tr)
pred = clf.predict(X_te)
print(confusion_matrix(y_te, pred))
print(classification_report(y_te, pred, target_names=["malignant", "benign"], digits=3))
Output:
[[ 59 5]
[ 2 105]]
precision recall f1-score support
malignant 0.967 0.922 0.944 64
benign 0.955 0.981 0.968 107
accuracy 0.959 171
macro avg 0.961 0.952 0.956 171
weighted avg 0.959 0.959 0.959 171Tune the decision threshold
Use predict_proba and choose a threshold that meets the recall or precision your use case needs instead of the default 0.5.
Quick check: Which metric measures how many actual positives the model found?
- Recall
- Precision
- Accuracy
- Support
Answer
Recall — Recall = found positives / all positives.