पाठ 13 / 25
Learning-Rate Schedules and Early Stopping
Change the step size over time and stop at the right moment.
Warm up, decay, watch validation
A learning rate that is good early in training is often too big later. Schedules change it over time: warm-up (start small to avoid early instability, common for transformers), step or cosine decay (reduce gradually), or reduce on plateau (cut when validation stops improving). Pair this with early stopping: evaluate every epoch, save the best checkpoint, and stop after several epochs without improvement. Plot training and validation loss together to choose these settings.
A cosine schedule with early stopping (sketch)
Standard PyTorch pieces. Not run here.
opt = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=0.01)
sched = torch.optim.lr_scheduler.CosineAnnealingLR(opt, T_max=epochs)
best, patience = float("inf"), 0
for epoch in range(epochs):
train_one_epoch(model, opt)
val = evaluate(model)
sched.step()
if val < best:
best, patience = val, 0
torch.save(model.state_dict(), "best.pt")
else:
patience += 1
if patience >= 5: breakFind a learning rate quickly
Run a short sweep of learning rates (for example 1e-4 to 1e-1) for a few hundred steps and pick the largest that still lowers the loss smoothly.
त्वरित जाँच: What does early stopping keep?
- The last checkpoint always
- The checkpoint with the best validation score
- The first checkpoint
- Only the optimiser state
Answer
The checkpoint with the best validation score — Stop when validation stops improving and keep the best.