Lesson 16 / 25
Why Convolutions Are Efficient
Weight sharing slashes parameter counts.
Same filter everywhere
A fully connected layer connects every input pixel to every unit, so parameters explode with image size and the layer cannot reuse what it learns at one position elsewhere. A convolution shares the same small filter across all positions: parameters depend on kernel size and channels, not image size, and a pattern learned in one corner is recognised everywhere (translation equivariance). This built-in assumption (local, repeated patterns) is why CNNs learn images with far less data and computation.
Dense layer versus convolution parameter counts, run
I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. Connecting a flattened 224x224 RGB image to just 64 units needs 9,633,856 parameters; a 3x3 convolution producing 64 channels needs 1,792.
import torch
dense = torch.nn.Linear(224 * 224 * 3, 64)
conv = torch.nn.Conv2d(3, 64, kernel_size=3)
count = lambda m: sum(p.numel() for p in m.parameters())
print(f"dense layer on a 224x224 RGB image -> 64 units: {count(dense):,} parameters")
print(f"3x3 conv layer, 3 -> 64 channels : {count(conv):,} parameters")
Output:
dense layer on a 224x224 RGB image -> 64 units: 9,633,856 parameters 3x3 conv layer, 3 -> 64 channels : 1,792 parameters
Count parameters as a sanity check
Print sum(p.numel()) after building a model; an unexpected count often reveals a wrong layer size.
Quick check: Why does a convolution layer have few parameters?
- It only processes one pixel
- It has no weights
- The same small filter is shared across all image positions
- It stores the dataset
Answer
The same small filter is shared across all image positions — Weight sharing is the key.