# Regression and Baselines — NumPy / Pandas / scikit-learn

Source: https://www.skillbyai.com/en/numpy-pandas-sklearn/s-regression

> Always compare to something simple.

## MAE, R² and a dummy baseline

Regression predicts numbers. **MAE** (mean absolute error) is in the target's units; **R²** compares the model with always predicting the mean (0 means no better, 1 is perfect). Always include a **baseline** such as `DummyRegressor` or `DummyClassifier`: a model is only useful if it clearly beats it. A more complex model is not automatically better, especially on small datasets.

## Baseline, ridge and random forest on the diabetes data, run

I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). On this small dataset (442 patients), simple ridge regression beat a 200-tree random forest, and both beat the mean baseline.

```python
from sklearn.datasets import load_diabetes
from sklearn.dummy import DummyRegressor
from sklearn.ensemble import RandomForestRegressor
from sklearn.linear_model import Ridge
from sklearn.metrics import mean_absolute_error, r2_score
from sklearn.model_selection import train_test_split

X, y = load_diabetes(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=0)
for name, m in [("baseline (mean)", DummyRegressor()),
                ("ridge", Ridge(alpha=0.1)),
                ("random forest", RandomForestRegressor(n_estimators=200, random_state=0))]:
    p = m.fit(X_tr, y_tr).predict(X_te)
    print(f"{name:16} MAE={mean_absolute_error(y_te, p):6.1f}  R2={r2_score(y_te, p):6.3f}")
```

Output:

```
baseline (mean)  MAE=  58.3  R2=-0.000
ridge            MAE=  44.6  R2= 0.369
random forest    MAE=  48.1  R2= 0.250
```

## Start simple

Linear models are fast, interpretable baselines; reach for ensembles or boosting when they clearly win on validation data.

**Quiz:** What does an R² of 0 mean?

- [ ] The data has no variance
- [ ] The model is perfect
- [ ] The model predicts zero
- [x] The model is no better than predicting the mean

*Answer:* The model is no better than predicting the mean. R² compares to the mean baseline.
