Lesson 7 / 25
Optimisers: SGD, Momentum and Adam
Different rules for turning gradients into steps.
No optimiser wins everywhere
SGD takes plain gradient steps. Momentum keeps a running direction, smoothing noisy gradients and speeding progress along consistent slopes. Adam adapts the step size per parameter from running averages of gradients and squared gradients; it is a popular default because it is forgiving across many problems (AdamW, with decoupled weight decay, is common for transformers). Results depend on the learning rate and problem: always compare on validation data rather than trusting a default.
Three optimisers on one regression problem, run
I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. After 100 steps on a 10-feature linear regression, SGD with momentum reaches loss 0.0099, plain SGD 0.1181, and Adam at learning rate 0.01 is at 0.1931. On this simple problem momentum wins; on deep networks Adam is often the easier default. The lesson is to compare, not assume.
import torch
torch.manual_seed(0); torch.set_num_threads(1)
X = torch.randn(512, 10)
true_w = torch.randn(10, 1)
y = X @ true_w + 0.1 * torch.randn(512, 1)
for name, make in [("SGD lr=0.01", lambda p: torch.optim.SGD(p, lr=0.01)),
("SGD+momentum", lambda p: torch.optim.SGD(p, lr=0.01, momentum=0.9)),
("Adam lr=0.01", lambda p: torch.optim.Adam(p, lr=0.01))]:
torch.manual_seed(1)
model = torch.nn.Linear(10, 1); opt = make(model.parameters())
for step in range(100):
opt.zero_grad(); loss = torch.nn.functional.mse_loss(model(X), y); loss.backward(); opt.step()
print(f"{name:<14} loss after 100 steps {loss.item():.4f}")
Output:
SGD lr=0.01 loss after 100 steps 0.1181 SGD+momentum loss after 100 steps 0.0099 Adam lr=0.01 loss after 100 steps 0.1931
Tune the learning rate first
The learning rate usually matters more than the optimiser choice; try a few values on a log scale.
Quick check: What does momentum add to SGD?
- A running direction that smooths noisy gradients
- A larger dataset
- Random weight resets
- A second loss function
Answer
A running direction that smooths noisy gradients — Momentum accelerates along consistent directions.