पाठ 19 / 25

Regression and Baselines

Always compare to something simple.

MAE, R² and a dummy baseline

Regression predicts numbers. MAE (mean absolute error) is in the target's units; R² compares the model with always predicting the mean (0 means no better, 1 is perfect). Always include a baseline such as DummyRegressor or DummyClassifier: a model is only useful if it clearly beats it. A more complex model is not automatically better, especially on small datasets.

Baseline, ridge and random forest on the diabetes data, run

I ran this with Python 3.12.3 and scikit-learn 1.9.1 on a dataset bundled with scikit-learn (no download). On this small dataset (442 patients), simple ridge regression beat a 200-tree random forest, and both beat the mean baseline.

from sklearn.datasets import load_diabetes
from sklearn.dummy import DummyRegressor
from sklearn.ensemble import RandomForestRegressor
from sklearn.linear_model import Ridge
from sklearn.metrics import mean_absolute_error, r2_score
from sklearn.model_selection import train_test_split

X, y = load_diabetes(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=0)
for name, m in [("baseline (mean)", DummyRegressor()),
                ("ridge", Ridge(alpha=0.1)),
                ("random forest", RandomForestRegressor(n_estimators=200, random_state=0))]:
    p = m.fit(X_tr, y_tr).predict(X_te)
    print(f"{name:16} MAE={mean_absolute_error(y_te, p):6.1f}  R2={r2_score(y_te, p):6.3f}")

Output:

baseline (mean)  MAE=  58.3  R2=-0.000
ridge            MAE=  44.6  R2= 0.369
random forest    MAE=  48.1  R2= 0.250

Start simple

Linear models are fast, interpretable baselines; reach for ensembles or boosting when they clearly win on validation data.

त्वरित जाँच: What does an R² of 0 mean?

  • The data has no variance
  • The model is perfect
  • The model predicts zero
  • The model is no better than predicting the mean
Answer

The model is no better than predicting the mean — R² compares to the mean baseline.