Lesson 9 / 27
Learning Rate: The Most Sensitive Knob
Step size decides whether training converges or explodes.
Too small is slow, too large diverges
The learning rate sets the size of each weight update. Too small and training barely moves; too large and updates overshoot, so loss bounces or explodes. Fine-tuning usually needs a much smaller rate than training from scratch because you want to refine, not overwrite. Typical full fine-tuning rates are small (around 1e-5 to 5e-5) and LoRA adapters often use larger ones (around 1e-4 to 3e-4), but follow current guidance for your model and tune on a validation set. Use warm-up and decay, and always plot training and validation loss.
Four learning rates on the same problem, run
I ran this with Python, numpy 2.5.3 and scikit-learn 1.9.1, with fixed random seeds. It trains small classical models, not a language model: the mechanics (gradient descent, learning rate, overfitting, forgetting, low-rank updates) are the same ideas that apply to fine-tuning an LLM, but the numbers are not LLM results. The same least-squares problem is solved with four step sizes: 0.001 learns slowly (loss 11.8 to 9.7 in 60 steps), 0.05 and 0.4 reach near 0.01, and 1.2 diverges to about 7e31.
import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(200, 5)); w_true = np.array([1.0, -2.0, 0.5, 0.0, 3.0]); y = X @ w_true + rng.normal(size=200) * 0.1
def train(lr, steps=60):
w = np.zeros(5); losses = []
for _ in range(steps):
err = X @ w - y; losses.append(float(np.mean(err ** 2)))
w -= lr * (2 / len(X)) * X.T @ err
return losses
for lr in (0.001, 0.05, 0.4, 1.2):
l = train(lr)
note = "diverges (loss explodes)" if l[-1] > l[0] else "learns"
print(f"learning rate {lr:<6} loss start {l[0]:9.3f} -> end {l[-1]:12.4g} {note}")
Output:
learning rate 0.001 loss start 11.813 -> end 9.696 learns learning rate 0.05 loss start 11.813 -> end 0.01021 learns learning rate 0.4 loss start 11.813 -> end 0.009382 learns learning rate 1.2 loss start 11.813 -> end 7.28e+31 diverges (loss explodes)
Sweep three rates
Try your rate, half and double it on a small run; pick the best validation loss, not the lowest training loss.
Quick check: A loss that suddenly shoots up usually suggests what?
- The model is perfect
- Training is finished
- The learning rate is too high (or the data/batch is bad)
- The GPU is faster
Answer
The learning rate is too high (or the data/batch is bad) — Lower the learning rate and inspect the data and logs.