Lesson 16 / 25
Regularisation
Penalise complexity to generalise better.
Shrink the weights
Regularisation adds a penalty for large model weights, pushing the model towards simpler solutions. Ridge (L2) shrinks all weights; Lasso (L1) can set some exactly to zero, acting as feature selection. The strength (alpha in scikit-learn's Ridge and Lasso; C, its inverse, in logistic regression) is a hyperparameter chosen by cross-validation: too weak and the model overfits, too strong and it underfits. Regularisation matters most when you have many features relative to rows.
Ridge strength with 40 features and 60 rows, run
I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. With only 60 noisy rows and 40 features (5 of them useful), weak regularisation overfits badly (cross-validated R2 of -2.466 at alpha 0.001). Alpha 100 gives the best score, 0.083, and alpha 1000 is too strong (-0.202). Even the best score is low because the data is so limited.
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.model_selection import cross_val_score
rng = np.random.default_rng(0)
X = rng.normal(size=(60, 40)); w = np.zeros(40); w[:5] = [3, -2, 2, 1, -1]
y = X @ w + rng.normal(scale=2, size=60)
print("alpha | mean CV R2")
for a in [0.001, 1, 10, 100, 1000]:
print(f"{a:>6} | {cross_val_score(Ridge(alpha=a), X, y, cv=5, scoring='r2').mean():.3f}")
Output:
alpha | mean CV R2
0.001 | -2.466
1 | -0.791
10 | -0.038
100 | 0.083
1000 | -0.202Scale before regularising
Penalties treat all weights equally, so features must be on comparable scales for regularisation to be fair.
Quick check: What happens when regularisation is far too strong?
- Training stops working entirely
- The model overfits more
- The model underfits
- The data gets larger
Answer
The model underfits — Too much penalty makes the model too simple.