Lesson 5 / 25
Gradient Descent and Backpropagation
The chain rule, applied layer by layer.
Which way to move each weight
Gradient descent updates each weight a small step in the direction that reduces the loss: new weight = old weight minus learning rate times the gradient. Backpropagation computes those gradients efficiently by applying the chain rule backwards from the loss through every layer, reusing intermediate results. You can check a hand-derived gradient numerically by nudging a weight slightly and measuring the change in loss; frameworks do the derivation for you, but the idea explains many training problems.
Backprop by hand with a numerical check, run
I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. For one sigmoid neuron with a squared loss, the chain-rule gradient for w is -0.300732, matching a numerical estimate exactly. One step with learning rate 0.5 lowers the loss from 0.1709 to 0.1137.
import numpy as np
# one neuron, squared loss: L = (sigmoid(w*x + b) - y)^2
x, y, w, b = 1.5, 1.0, 0.3, -0.1
def loss(w, b):
p = 1 / (1 + np.exp(-(w * x + b)))
return (p - y) ** 2
p = 1 / (1 + np.exp(-(w * x + b)))
dL_dp = 2 * (p - y) # chain rule, step by step
dp_dz = p * (1 - p)
grad_w = dL_dp * dp_dz * x
grad_b = dL_dp * dp_dz * 1
eps = 1e-6
num_w = (loss(w + eps, b) - loss(w - eps, b)) / (2 * eps)
print(f"backprop gradient dL/dw = {grad_w:.6f} | numerical check {num_w:.6f}")
print(f"backprop gradient dL/db = {grad_b:.6f}")
w2 = w - 0.5 * grad_w; b2 = b - 0.5 * grad_b
print(f"loss before step {loss(w, b):.4f} -> after one step {loss(w2, b2):.4f}")
Output:
backprop gradient dL/dw = -0.300732 | numerical check -0.300732 backprop gradient dL/db = -0.200488 loss before step 0.1709 -> after one step 0.1137
Gradient-check custom code
If you write a custom layer or loss, compare its gradients with a numerical estimate on a tiny input.
Quick check: What does backpropagation compute?
- The final accuracy
- The gradient of the loss with respect to every weight, via the chain rule
- The dataset size
- The best learning rate
Answer
The gradient of the loss with respect to every weight, via the chain rule — Gradients tell each weight which way to move.