पाठ 14 / 25

Underfitting and Overfitting

Too simple misses the pattern; too complex learns the noise.

The bias-variance trade-off

An underfitting model is too simple to capture the pattern: it does badly on both training and test data. An overfitting model is so flexible it fits noise in the training data: training error is low but test error is higher. Between them lies a sweet spot. You move along this scale with model complexity (polynomial degree, tree depth, number of neighbours) and with the amount of data: more data lets more complex models generalise. Diagnose by comparing training and validation errors.

Fit the signal, not the noise

The goal is good performance on new data; validation techniques tell you whether you have it.

Three ideas: over- and underfitting, cross-validation, regularisation.
Figure 5.1 — Fit, validation and regularisation.

Polynomial degree and test error, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. Fitting a noisy sine curve with 30 points: degree 1 underfits (test error 0.267), degree 3 is best (0.050), and degrees 5 and 15 fit the training points more closely (0.018) but do worse on test data (0.067).

import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures
from sklearn.metrics import mean_squared_error
rng = np.random.default_rng(0)
x = rng.uniform(0, 1, 30).reshape(-1, 1); y = np.sin(2 * np.pi * x).ravel() + rng.normal(0, 0.2, 30)
xt = rng.uniform(0, 1, 200).reshape(-1, 1); yt = np.sin(2 * np.pi * xt).ravel() + rng.normal(0, 0.2, 200)
print("degree | train error | test error")
for d in [1, 3, 5, 15]:
    m = make_pipeline(PolynomialFeatures(d), LinearRegression()).fit(x, y)
    print(f"{d:>6} | {mean_squared_error(y, m.predict(x)):>11.3f} | {mean_squared_error(yt, m.predict(xt)):.3f}")

Output:

degree | train error | test error
     1 |       0.303 | 0.267
     3 |       0.051 | 0.050
     5 |       0.018 | 0.067
    15 |       0.018 | 0.067

Plot learning curves

Plot training and validation scores as data grows; a persistent gap means overfitting, two low curves mean underfitting.

त्वरित जाँच: Low training error but much higher test error indicates what?

  • Data leakage is impossible
  • Underfitting
  • A perfect model
  • Overfitting
Answer

Overfitting — The model learned noise specific to the training set.