SkillByAIOpen interactive version →

Lesson 23 / 25

Debugging Neural Networks

Start tiny and verify each piece.

A routine that finds most bugs

When a network does not learn, work in order: check data and labels by eye; check shapes; check the loss at initialisation (for 10 balanced classes, cross-entropy should start near ln(10) = 2.3); overfit a single small batch (if the model cannot drive the loss near zero on 16 examples, there is a bug in the model, loss or optimiser step); then train on the full data and compare training and validation curves; then tune. Change one thing at a time and keep a log of runs.

Make it work, ship it, review it

A methodical debugging routine and careful deployment matter as much as architecture.

Figure 8.1 — Debugging, serving and checklist.

Overfitting one batch as a sanity check, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. A small network drives the loss on a single batch of 16 random examples from 1.1210 to 0.0004 in 50 steps and 0.0001 by step 200, showing the model, loss and optimiser are wired correctly.

import torch
torch.manual_seed(0); torch.set_num_threads(1)
xb = torch.randn(16, 20); yb = torch.randint(0, 3, (16,))
model = torch.nn.Sequential(torch.nn.Linear(20, 64), torch.nn.ReLU(), torch.nn.Linear(64, 3))
opt = torch.optim.Adam(model.parameters(), lr=1e-2)
for step in range(201):
    opt.zero_grad(); loss = torch.nn.functional.cross_entropy(model(xb), yb); loss.backward(); opt.step()
    if step in (0, 50, 200):
        print(f"step {step:>3} loss {loss.item():.4f}")
print("can memorise one batch:", loss.item() < 0.01)

Output:

step   0 loss 1.1210
step  50 loss 0.0004
step 200 loss 0.0001
can memorise one batch: True

Check the initial loss

An initial loss far from the expected value (like 2.3 for 10 classes) often means wrong labels, a double softmax or broken scaling.

Quick check: What does failing to overfit a single small batch suggest?

  • A bug in the model, loss or training step
  • The dataset is too large
  • The model is perfect
  • You need more epochs on all data
Answer

A bug in the model, loss or training step — A working setup can memorise a few examples.