Lesson 7 / 25

Optimisers: SGD, Momentum and Adam

Different rules for turning gradients into steps.

No optimiser wins everywhere

SGD takes plain gradient steps. Momentum keeps a running direction, smoothing noisy gradients and speeding progress along consistent slopes. Adam adapts the step size per parameter from running averages of gradients and squared gradients; it is a popular default because it is forgiving across many problems (AdamW, with decoupled weight decay, is common for transformers). Results depend on the learning rate and problem: always compare on validation data rather than trusting a default.

Three optimisers on one regression problem, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. After 100 steps on a 10-feature linear regression, SGD with momentum reaches loss 0.0099, plain SGD 0.1181, and Adam at learning rate 0.01 is at 0.1931. On this simple problem momentum wins; on deep networks Adam is often the easier default. The lesson is to compare, not assume.

import torch
torch.manual_seed(0); torch.set_num_threads(1)
X = torch.randn(512, 10)
true_w = torch.randn(10, 1)
y = X @ true_w + 0.1 * torch.randn(512, 1)
for name, make in [("SGD lr=0.01", lambda p: torch.optim.SGD(p, lr=0.01)),
                   ("SGD+momentum", lambda p: torch.optim.SGD(p, lr=0.01, momentum=0.9)),
                   ("Adam lr=0.01", lambda p: torch.optim.Adam(p, lr=0.01))]:
    torch.manual_seed(1)
    model = torch.nn.Linear(10, 1); opt = make(model.parameters())
    for step in range(100):
        opt.zero_grad(); loss = torch.nn.functional.mse_loss(model(X), y); loss.backward(); opt.step()
    print(f"{name:<14} loss after 100 steps {loss.item():.4f}")

Output:

SGD lr=0.01    loss after 100 steps 0.1181
SGD+momentum   loss after 100 steps 0.0099
Adam lr=0.01   loss after 100 steps 0.1931

Tune the learning rate first

The learning rate usually matters more than the optimiser choice; try a few values on a log scale.

Quick check: What does momentum add to SGD?

  • A running direction that smooths noisy gradients
  • A larger dataset
  • Random weight resets
  • A second loss function
Answer

A running direction that smooths noisy gradients — Momentum accelerates along consistent directions.