# Logistic Regression and Probabilities — Machine Learning Basics

Source: https://www.skillbyai.com/en/machine-learning/c-logreg

> A linear model that outputs probabilities.

## From scores to probabilities

**Logistic regression** computes a weighted sum of features and squeezes it through the logistic (sigmoid) function to give a **probability** between 0 and 1 for the positive class; predicting the class means comparing that probability with a **threshold** (0.5 by default). Despite its name it is a classification method. It is fast, works well with scaled features, gives interpretable weights, and its probabilities are often reasonably calibrated, making it an excellent first classifier.

## Models and the metrics that matter

Classifiers output classes or probabilities; accuracy alone can mislead.

![Four ideas: logistic regression, trees and neighbours, confusion matrix, thresholds.](assets/figures/machine-learning/section-4-map.svg) — Figure 4.1 — Models, confusion matrix and thresholds.

## Scaled logistic regression with probabilities, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. On the breast cancer dataset (where class 1 means benign), a scaled logistic regression reaches 0.958 test accuracy; the first three test rows get probabilities of 0.996, 0.000 and 0.000, which become classes 1, 0 and 0 at the 0.5 threshold.

```python
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, random_state=0, stratify=y)
clf = make_pipeline(StandardScaler(), LogisticRegression()).fit(X_tr, y_tr)
print("test accuracy:", round(clf.score(X_te, y_te), 3))
probs = clf.predict_proba(X_te[:3])[:, 1]
for p in probs:
    print(f"P(benign) = {p:.3f} -> predicted class {int(p >= 0.5)}")
```

Output:

```
test accuracy: 0.958
P(benign) = 0.996 -> predicted class 1
P(benign) = 0.000 -> predicted class 0
P(benign) = 0.000 -> predicted class 0
```

## Use predict_proba

Keep the probabilities, not just the class; they let you choose thresholds and rank cases by risk.

**Quiz:** What does logistic regression output before thresholding?

- [x] A probability for the positive class
- [ ] A cluster id
- [ ] A continuous price
- [ ] A decision tree

*Answer:* A probability for the positive class. The threshold turns a probability into a class.
