Lesson 6 / 25
Automatic Differentiation in PyTorch
Let the framework apply the chain rule.
requires_grad, backward, .grad
PyTorch records operations on tensors that have requires_grad=True and builds a computation graph. Calling loss.backward() walks that graph backwards and stores each gradient in the tensor's .grad. Optimisers then read .grad to update parameters. Two habits matter: call optimizer.zero_grad() before each backward pass (gradients accumulate by default), and wrap inference in torch.no_grad() so no graph is built when you are not training.
Autograd reproduces the hand gradient, run
I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. The same neuron and loss in PyTorch: loss.backward() produces dL/dw = -0.300732 and dL/db = -0.200488, identical to the hand calculation.
import torch
w = torch.tensor(0.3, requires_grad=True)
b = torch.tensor(-0.1, requires_grad=True)
x, y = torch.tensor(1.5), torch.tensor(1.0)
loss = (torch.sigmoid(w * x + b) - y) ** 2
loss.backward() # PyTorch applies the chain rule for us
print(f"autograd dL/dw = {w.grad.item():.6f}")
print(f"autograd dL/db = {b.grad.item():.6f}")
Output:
autograd dL/dw = -0.300732 autograd dL/db = -0.200488
Zero gradients every step
Forgetting zero_grad silently adds gradients from previous batches and makes training erratic.
Quick check: Why call optimizer.zero_grad() each step?
- It loads new data
- It saves the model
- It sets the learning rate
- PyTorch accumulates gradients by default
Answer
PyTorch accumulates gradients by default — Clear old gradients before computing new ones.