# Classification Metrics — NumPy / Pandas / scikit-learn

Source: https://www.skillbyai.com/en/numpy-pandas-sklearn/s-metrics

> Accuracy is not enough.

## Confusion matrix, precision, recall, F1

The **confusion matrix** counts correct and wrong predictions per class. **Precision** asks "of the cases predicted positive, how many were?", **recall** asks "of the actual positives, how many did we find?", and **F1** balances them. With imbalanced classes or unequal costs (missing a malignant tumour is worse than a false alarm), choose metrics and thresholds that reflect those costs rather than accuracy alone.

## Evaluating a breast cancer classifier, run

I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). On 171 test cases, 5 malignant tumours were predicted benign and 2 benign as malignant; malignant recall is 0.922 even though overall accuracy is 0.959.

```python
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report, confusion_matrix
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

X, y = load_breast_cancer(return_X_y=True)   # 0 = malignant, 1 = benign
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=1, stratify=y)
clf = make_pipeline(StandardScaler(), LogisticRegression()).fit(X_tr, y_tr)
pred = clf.predict(X_te)
print(confusion_matrix(y_te, pred))
print(classification_report(y_te, pred, target_names=["malignant", "benign"], digits=3))
```

Output:

```
[[ 59   5]
 [  2 105]]
              precision    recall  f1-score   support

   malignant      0.967     0.922     0.944        64
      benign      0.955     0.981     0.968       107

    accuracy                          0.959       171
   macro avg      0.961     0.952     0.956       171
weighted avg      0.959     0.959     0.959       171
```

## Tune the decision threshold

Use predict_proba and choose a threshold that meets the recall or precision your use case needs instead of the default 0.5.

**Quiz:** Which metric measures how many actual positives the model found?

- [x] Recall
- [ ] Precision
- [ ] Accuracy
- [ ] Support

*Answer:* Recall. Recall = found positives / all positives.
