# Regularisation — Machine Learning Basics

Source: https://www.skillbyai.com/en/machine-learning/g-reg

> Penalise complexity to generalise better.

## Shrink the weights

**Regularisation** adds a penalty for large model weights, pushing the model towards simpler solutions. **Ridge** (L2) shrinks all weights; **Lasso** (L1) can set some exactly to zero, acting as feature selection. The strength (alpha in scikit-learn's Ridge and Lasso; C, its inverse, in logistic regression) is a **hyperparameter** chosen by cross-validation: too weak and the model overfits, too strong and it underfits. Regularisation matters most when you have many features relative to rows.

## Ridge strength with 40 features and 60 rows, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. With only 60 noisy rows and 40 features (5 of them useful), weak regularisation overfits badly (cross-validated R2 of -2.466 at alpha 0.001). Alpha 100 gives the best score, 0.083, and alpha 1000 is too strong (-0.202). Even the best score is low because the data is so limited.

```python
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.model_selection import cross_val_score
rng = np.random.default_rng(0)
X = rng.normal(size=(60, 40)); w = np.zeros(40); w[:5] = [3, -2, 2, 1, -1]
y = X @ w + rng.normal(scale=2, size=60)
print("alpha  | mean CV R2")
for a in [0.001, 1, 10, 100, 1000]:
    print(f"{a:>6} | {cross_val_score(Ridge(alpha=a), X, y, cv=5, scoring='r2').mean():.3f}")
```

Output:

```
alpha  | mean CV R2
 0.001 | -2.466
     1 | -0.791
    10 | -0.038
   100 | 0.083
  1000 | -0.202
```

## Scale before regularising

Penalties treat all weights equally, so features must be on comparable scales for regularisation to be fair.

**Quiz:** What happens when regularisation is far too strong?

- [ ] Training stops working entirely
- [ ] The model overfits more
- [x] The model underfits
- [ ] The data gets larger

*Answer:* The model underfits. Too much penalty makes the model too simple.
