SkillByAIOpen interactive version →

Lesson 11 / 27

Overfitting, Epochs and Early Stopping

Stop training when validation quality stops improving.

Memorising the training set is not learning

A model too flexible for the data memorises training examples, noise included, and does worse on new inputs: overfitting. Detect it with a held-out validation set: training accuracy keeps rising while validation stalls or falls. Remedies: fewer epochs and early stopping (keep the best-validation checkpoint), more and more varied data, regularisation (weight decay, dropout, lower LoRA rank) and a lower learning rate. On small datasets one to three epochs are common; overtrained LLMs parrot training answers and lose general ability.

Train versus validation accuracy, run

I ran this with Python, numpy 2.5.3 and scikit-learn 1.9.1, with fixed random seeds. It trains small classical models, not a language model: the mechanics (gradient descent, learning rate, overfitting, forgetting, low-rank updates) are the same ideas that apply to fine-tuning an LLM, but the numbers are not LLM results. With 60 training examples and 60 features, the most regularised model underfits (about 0.53 train, 0.55 validation), moderate regularisation (C = 0.01) gives the best validation accuracy, 0.667, and weak regularisation fits training perfectly (1.000) while validation falls to 0.58 to 0.63.

import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(2)
X = rng.normal(size=(120, 60))                                    # many features, few examples, weak real signal
y = (X[:, 0] + X[:, 1] + rng.normal(size=120) * 1.5 > 0).astype(int)
Xtr, Xva, ytr, yva = train_test_split(X, y, test_size=0.5, random_state=0)

print("regularisation C (higher = fewer limits) | train acc | validation acc")
for C in (0.001, 0.01, 0.1, 1, 100):
    m = LogisticRegression(C=C, max_iter=5000).fit(Xtr, ytr)
    print(f"{C:>40} | {m.score(Xtr, ytr):9.3f} | {m.score(Xva, yva):.3f}")

Output:

regularisation C (higher = fewer limits) | train acc | validation acc
                                   0.001 |     0.533 | 0.550
                                    0.01 |     0.850 | 0.667
                                     0.1 |     1.000 | 0.633
                                       1 |     1.000 | 0.600
                                     100 |     1.000 | 0.583

Save checkpoints and compare them

Evaluate several checkpoints on held-out data and keep the best, not simply the last.

Quick check: What is the classic sign of overfitting?

  • Loss is exactly zero on step one
  • Both curves are flat from the start
  • Validation is always better than training
  • Training performance keeps improving while validation performance stalls or worsens
Answer

Training performance keeps improving while validation performance stalls or worsens — Hold out data and watch the gap.