Lesson 19 / 25
Regression and Baselines
Always compare to something simple.
MAE, R² and a dummy baseline
Regression predicts numbers. MAE (mean absolute error) is in the target's units; R² compares the model with always predicting the mean (0 means no better, 1 is perfect). Always include a baseline such as DummyRegressor or DummyClassifier: a model is only useful if it clearly beats it. A more complex model is not automatically better, especially on small datasets.
Baseline, ridge and random forest on the diabetes data, run
I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). On this small dataset (442 patients), simple ridge regression beat a 200-tree random forest, and both beat the mean baseline.
from sklearn.datasets import load_diabetes
from sklearn.dummy import DummyRegressor
from sklearn.ensemble import RandomForestRegressor
from sklearn.linear_model import Ridge
from sklearn.metrics import mean_absolute_error, r2_score
from sklearn.model_selection import train_test_split
X, y = load_diabetes(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=0)
for name, m in [("baseline (mean)", DummyRegressor()),
("ridge", Ridge(alpha=0.1)),
("random forest", RandomForestRegressor(n_estimators=200, random_state=0))]:
p = m.fit(X_tr, y_tr).predict(X_te)
print(f"{name:16} MAE={mean_absolute_error(y_te, p):6.1f} R2={r2_score(y_te, p):6.3f}")
Output:
baseline (mean) MAE= 58.3 R2=-0.000 ridge MAE= 44.6 R2= 0.369 random forest MAE= 48.1 R2= 0.250
Start simple
Linear models are fast, interpretable baselines; reach for ensembles or boosting when they clearly win on validation data.
Quick check: What does an R² of 0 mean?
- The data has no variance
- The model is perfect
- The model predicts zero
- The model is no better than predicting the mean
Answer
The model is no better than predicting the mean — R² compares to the mean baseline.