Lesson 11 / 26
CNN Architectures in Brief
From LeNet to ResNet and beyond.
Deeper, wider, smarter
Convolutional neural networks stack convolution, normalisation, activation and pooling layers that learn edges, textures, parts and objects. Landmarks: LeNet (digits), AlexNet (2012 ImageNet breakthrough), VGG (deep, simple), ResNet (residual connections make very deep networks trainable), MobileNet and EfficientNet (efficient for mobile and edge), ConvNeXt (modernised CNN design). Vision Transformers (ViT) split images into patches and apply transformer layers; with enough data or pretraining they match or beat CNNs. In practice you start from a pretrained model of a suitable size.
Pretrained networks do most of the work
Convolutional and transformer networks pretrained on large datasets classify images and transfer to new tasks.
Choosing a backbone
Size versus accuracy versus speed.
family strengths typical use
ResNet-18/50 reliable, well understood, many weights baselines, transfer learning
MobileNetV3 tiny and fast phones, embedded devices
EfficientNet accuracy per compute balanced deployments
ConvNeXt modern CNN, strong accuracy server-side vision
ViT / DINO strong pretrained features, scales well large data, foundation modelsStart small
Try a small pretrained backbone first; move to larger ones only if validation results justify the extra latency and cost.
Quick check: What made very deep CNNs like ResNet trainable?
- Training without labels
- Removing all convolutions
- Using grayscale only
- Residual (skip) connections
Answer
Residual (skip) connections — Skip connections keep gradients flowing.