# Optimisers: SGD, Momentum and Adam — Deep Learning & Neural Networks

Source: https://www.skillbyai.com/en/deep-learning/t-optim

> Different rules for turning gradients into steps.

## No optimiser wins everywhere

**SGD** takes plain gradient steps. **Momentum** keeps a running direction, smoothing noisy gradients and speeding progress along consistent slopes. **Adam** adapts the step size per parameter from running averages of gradients and squared gradients; it is a popular default because it is forgiving across many problems (AdamW, with decoupled weight decay, is common for transformers). Results depend on the learning rate and problem: always compare on validation data rather than trusting a default.

## Three optimisers on one regression problem, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. After 100 steps on a 10-feature linear regression, SGD with momentum reaches loss 0.0099, plain SGD 0.1181, and Adam at learning rate 0.01 is at 0.1931. On this simple problem momentum wins; on deep networks Adam is often the easier default. The lesson is to compare, not assume.

```python
import torch
torch.manual_seed(0); torch.set_num_threads(1)
X = torch.randn(512, 10)
true_w = torch.randn(10, 1)
y = X @ true_w + 0.1 * torch.randn(512, 1)
for name, make in [("SGD lr=0.01", lambda p: torch.optim.SGD(p, lr=0.01)),
                   ("SGD+momentum", lambda p: torch.optim.SGD(p, lr=0.01, momentum=0.9)),
                   ("Adam lr=0.01", lambda p: torch.optim.Adam(p, lr=0.01))]:
    torch.manual_seed(1)
    model = torch.nn.Linear(10, 1); opt = make(model.parameters())
    for step in range(100):
        opt.zero_grad(); loss = torch.nn.functional.mse_loss(model(X), y); loss.backward(); opt.step()
    print(f"{name:<14} loss after 100 steps {loss.item():.4f}")
```

Output:

```
SGD lr=0.01    loss after 100 steps 0.1181
SGD+momentum   loss after 100 steps 0.0099
Adam lr=0.01   loss after 100 steps 0.1931
```

## Tune the learning rate first

The learning rate usually matters more than the optimiser choice; try a few values on a log scale.

**Quiz:** What does momentum add to SGD?

- [x] A running direction that smooths noisy gradients
- [ ] A larger dataset
- [ ] Random weight resets
- [ ] A second loss function

*Answer:* A running direction that smooths noisy gradients. Momentum accelerates along consistent directions.
