SkillByAIOpen interactive version →

Lesson 6 / 25

Automatic Differentiation in PyTorch

Let the framework apply the chain rule.

requires_grad, backward, .grad

PyTorch records operations on tensors that have requires_grad=True and builds a computation graph. Calling loss.backward() walks that graph backwards and stores each gradient in the tensor's .grad. Optimisers then read .grad to update parameters. Two habits matter: call optimizer.zero_grad() before each backward pass (gradients accumulate by default), and wrap inference in torch.no_grad() so no graph is built when you are not training.

Autograd reproduces the hand gradient, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. The same neuron and loss in PyTorch: loss.backward() produces dL/dw = -0.300732 and dL/db = -0.200488, identical to the hand calculation.

import torch
w = torch.tensor(0.3, requires_grad=True)
b = torch.tensor(-0.1, requires_grad=True)
x, y = torch.tensor(1.5), torch.tensor(1.0)
loss = (torch.sigmoid(w * x + b) - y) ** 2
loss.backward()                     # PyTorch applies the chain rule for us
print(f"autograd dL/dw = {w.grad.item():.6f}")
print(f"autograd dL/db = {b.grad.item():.6f}")

Output:

autograd dL/dw = -0.300732
autograd dL/db = -0.200488

Zero gradients every step

Forgetting zero_grad silently adds gradients from previous batches and makes training erratic.

Quick check: Why call optimizer.zero_grad() each step?

  • It loads new data
  • It saves the model
  • It sets the learning rate
  • PyTorch accumulates gradients by default
Answer

PyTorch accumulates gradients by default — Clear old gradients before computing new ones.