# Building a Small CNN — Deep Learning & Neural Networks

Source: https://www.skillbyai.com/en/deep-learning/c-cnn

> Convolutions, pooling, then a classifier.

## Conv, activation, pool, repeat

A typical CNN stacks blocks of **convolution, activation and pooling**. **Max pooling** keeps the strongest response in each small window, shrinking the spatial size and adding some tolerance to small shifts. Channels usually increase as resolution decreases. Finally the feature maps are flattened (or globally averaged) and passed to a linear classifier. Modern image models add batch normalisation and residual connections; ready-made architectures (ResNet, EfficientNet, ConvNeXt) are usually fine-tuned rather than built from scratch.

## A two-block CNN on 8x8 digits, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. Two convolution and pooling blocks followed by a linear layer (6,090 parameters) reach 0.971 test accuracy on the digits dataset after 15 epochs, better than the fully connected network earlier.

```python
import torch
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
torch.manual_seed(0); torch.set_num_threads(1)
X, y = load_digits(return_X_y=True)
X = torch.tensor(X / 16.0, dtype=torch.float32).view(-1, 1, 8, 8); y = torch.tensor(y)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=0, stratify=y)
cnn = torch.nn.Sequential(
    torch.nn.Conv2d(1, 16, 3, padding=1), torch.nn.ReLU(), torch.nn.MaxPool2d(2),   # 16 x 4 x 4
    torch.nn.Conv2d(16, 32, 3, padding=1), torch.nn.ReLU(), torch.nn.MaxPool2d(2),  # 32 x 2 x 2
    torch.nn.Flatten(), torch.nn.Linear(32 * 2 * 2, 10))
opt = torch.optim.Adam(cnn.parameters(), lr=2e-3)
loader = torch.utils.data.DataLoader(torch.utils.data.TensorDataset(X_tr, y_tr), batch_size=64, shuffle=True)
for epoch in range(15):
    cnn.train()
    for xb, yb in loader:
        opt.zero_grad(); torch.nn.functional.cross_entropy(cnn(xb), yb).backward(); opt.step()
cnn.eval()
with torch.no_grad():
    print("CNN test accuracy:", round((cnn(X_te).argmax(1) == y_te).float().mean().item(), 3))
print("parameters:", sum(p.numel() for p in cnn.parameters()))
```

Output:

```
CNN test accuracy: 0.971
parameters: 6090
```

## Write the shape after each block

Comment the output shape of each block (16 x 4 x 4, then 32 x 2 x 2); it makes the final Linear size obvious.

**Quiz:** What does max pooling do?

- [x] Keeps the strongest value in each window, shrinking the feature map
- [ ] Adds more pixels
- [ ] Normalises the labels
- [ ] Computes the loss

*Answer:* Keeps the strongest value in each window, shrinking the feature map. Pooling downsamples and adds shift tolerance.
