# Gradient Descent and Backpropagation — Deep Learning & Neural Networks

Source: https://www.skillbyai.com/en/deep-learning/t-backprop

> The chain rule, applied layer by layer.

## Which way to move each weight

**Gradient descent** updates each weight a small step in the direction that reduces the loss: new weight = old weight minus learning rate times the gradient. **Backpropagation** computes those gradients efficiently by applying the **chain rule** backwards from the loss through every layer, reusing intermediate results. You can check a hand-derived gradient numerically by nudging a weight slightly and measuring the change in loss; frameworks do the derivation for you, but the idea explains many training problems.

## Backprop by hand with a numerical check, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. For one sigmoid neuron with a squared loss, the chain-rule gradient for w is -0.300732, matching a numerical estimate exactly. One step with learning rate 0.5 lowers the loss from 0.1709 to 0.1137.

```python
import numpy as np
# one neuron, squared loss: L = (sigmoid(w*x + b) - y)^2
x, y, w, b = 1.5, 1.0, 0.3, -0.1
def loss(w, b):
    p = 1 / (1 + np.exp(-(w * x + b)))
    return (p - y) ** 2
p = 1 / (1 + np.exp(-(w * x + b)))
dL_dp = 2 * (p - y)                 # chain rule, step by step
dp_dz = p * (1 - p)
grad_w = dL_dp * dp_dz * x
grad_b = dL_dp * dp_dz * 1
eps = 1e-6
num_w = (loss(w + eps, b) - loss(w - eps, b)) / (2 * eps)
print(f"backprop gradient dL/dw = {grad_w:.6f} | numerical check {num_w:.6f}")
print(f"backprop gradient dL/db = {grad_b:.6f}")
w2 = w - 0.5 * grad_w; b2 = b - 0.5 * grad_b
print(f"loss before step {loss(w, b):.4f} -> after one step {loss(w2, b2):.4f}")
```

Output:

```
backprop gradient dL/dw = -0.300732 | numerical check -0.300732
backprop gradient dL/db = -0.200488
loss before step 0.1709 -> after one step 0.1137
```

## Gradient-check custom code

If you write a custom layer or loss, compare its gradients with a numerical estimate on a tiny input.

**Quiz:** What does backpropagation compute?

- [ ] The final accuracy
- [x] The gradient of the loss with respect to every weight, via the chain rule
- [ ] The dataset size
- [ ] The best learning rate

*Answer:* The gradient of the loss with respect to every weight, via the chain rule. Gradients tell each weight which way to move.
