पाठ 13 / 25

Learning-Rate Schedules and Early Stopping

Change the step size over time and stop at the right moment.

Warm up, decay, watch validation

A learning rate that is good early in training is often too big later. Schedules change it over time: warm-up (start small to avoid early instability, common for transformers), step or cosine decay (reduce gradually), or reduce on plateau (cut when validation stops improving). Pair this with early stopping: evaluate every epoch, save the best checkpoint, and stop after several epochs without improvement. Plot training and validation loss together to choose these settings.

A cosine schedule with early stopping (sketch)

Standard PyTorch pieces. Not run here.

opt = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=0.01)
sched = torch.optim.lr_scheduler.CosineAnnealingLR(opt, T_max=epochs)
best, patience = float("inf"), 0
for epoch in range(epochs):
    train_one_epoch(model, opt)
    val = evaluate(model)
    sched.step()
    if val < best:
        best, patience = val, 0
        torch.save(model.state_dict(), "best.pt")
    else:
        patience += 1
        if patience >= 5: break

Find a learning rate quickly

Run a short sweep of learning rates (for example 1e-4 to 1e-1) for a few hundred steps and pick the largest that still lowers the loss smoothly.

त्वरित जाँच: What does early stopping keep?

  • The last checkpoint always
  • The checkpoint with the best validation score
  • The first checkpoint
  • Only the optimiser state
Answer

The checkpoint with the best validation score — Stop when validation stops improving and keep the best.