# Overfitting, Epochs and Early Stopping — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/h-overfit

> Stop training when validation quality stops improving.

## Memorising the training set is not learning

A model too flexible for the data **memorises** training examples, noise included, and does worse on new inputs: **overfitting**. Detect it with a held-out **validation set**: training accuracy keeps rising while validation stalls or falls. Remedies: **fewer epochs** and **early stopping** (keep the best-validation checkpoint), **more and more varied data**, **regularisation** (weight decay, dropout, lower LoRA rank) and a **lower learning rate**. On small datasets one to three epochs are common; overtrained LLMs parrot training answers and lose general ability.

## Train versus validation accuracy, run

I ran this with Python, numpy 2.5.3 and scikit-learn 1.9.1, with fixed random seeds. It trains small classical models, not a language model: the mechanics (gradient descent, learning rate, overfitting, forgetting, low-rank updates) are the same ideas that apply to fine-tuning an LLM, but the numbers are not LLM results. With 60 training examples and 60 features, the most regularised model underfits (about 0.53 train, 0.55 validation), moderate regularisation (C = 0.01) gives the best validation accuracy, 0.667, and weak regularisation fits training perfectly (1.000) while validation falls to 0.58 to 0.63.

```python
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(2)
X = rng.normal(size=(120, 60))                                    # many features, few examples, weak real signal
y = (X[:, 0] + X[:, 1] + rng.normal(size=120) * 1.5 > 0).astype(int)
Xtr, Xva, ytr, yva = train_test_split(X, y, test_size=0.5, random_state=0)

print("regularisation C (higher = fewer limits) | train acc | validation acc")
for C in (0.001, 0.01, 0.1, 1, 100):
    m = LogisticRegression(C=C, max_iter=5000).fit(Xtr, ytr)
    print(f"{C:>40} | {m.score(Xtr, ytr):9.3f} | {m.score(Xva, yva):.3f}")

```

Output:

```
regularisation C (higher = fewer limits) | train acc | validation acc
                                   0.001 |     0.533 | 0.550
                                    0.01 |     0.850 | 0.667
                                     0.1 |     1.000 | 0.633
                                       1 |     1.000 | 0.600
                                     100 |     1.000 | 0.583
```

## Save checkpoints and compare them

Evaluate several checkpoints on held-out data and keep the best, not simply the last.

**Quiz:** What is the classic sign of overfitting?

- [ ] Loss is exactly zero on step one
- [ ] Both curves are flat from the start
- [ ] Validation is always better than training
- [x] Training performance keeps improving while validation performance stalls or worsens

*Answer:* Training performance keeps improving while validation performance stalls or worsens. Hold out data and watch the gap.
