# Automatic Differentiation in PyTorch — Deep Learning & Neural Networks

Source: https://www.skillbyai.com/en/deep-learning/t-autograd

> Let the framework apply the chain rule.

## requires_grad, backward, .grad

PyTorch records operations on tensors that have `requires_grad=True` and builds a computation graph. Calling `loss.backward()` walks that graph backwards and stores each gradient in the tensor's `.grad`. Optimisers then read `.grad` to update parameters. Two habits matter: call `optimizer.zero_grad()` before each backward pass (gradients accumulate by default), and wrap inference in `torch.no_grad()` so no graph is built when you are not training.

## Autograd reproduces the hand gradient, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. The same neuron and loss in PyTorch: loss.backward() produces dL/dw = -0.300732 and dL/db = -0.200488, identical to the hand calculation.

```python
import torch
w = torch.tensor(0.3, requires_grad=True)
b = torch.tensor(-0.1, requires_grad=True)
x, y = torch.tensor(1.5), torch.tensor(1.0)
loss = (torch.sigmoid(w * x + b) - y) ** 2
loss.backward()                     # PyTorch applies the chain rule for us
print(f"autograd dL/dw = {w.grad.item():.6f}")
print(f"autograd dL/db = {b.grad.item():.6f}")
```

Output:

```
autograd dL/dw = -0.300732
autograd dL/db = -0.200488
```

## Zero gradients every step

Forgetting zero_grad silently adds gradients from previous batches and makes training erratic.

**Quiz:** Why call optimizer.zero_grad() each step?

- [ ] It loads new data
- [ ] It saves the model
- [ ] It sets the learning rate
- [x] PyTorch accumulates gradients by default

*Answer:* PyTorch accumulates gradients by default. Clear old gradients before computing new ones.
