SkillByAIOpen interactive version →

Lesson 11 / 26

CNN Architectures in Brief

From LeNet to ResNet and beyond.

Deeper, wider, smarter

Convolutional neural networks stack convolution, normalisation, activation and pooling layers that learn edges, textures, parts and objects. Landmarks: LeNet (digits), AlexNet (2012 ImageNet breakthrough), VGG (deep, simple), ResNet (residual connections make very deep networks trainable), MobileNet and EfficientNet (efficient for mobile and edge), ConvNeXt (modernised CNN design). Vision Transformers (ViT) split images into patches and apply transformer layers; with enough data or pretraining they match or beat CNNs. In practice you start from a pretrained model of a suitable size.

Pretrained networks do most of the work

Convolutional and transformer networks pretrained on large datasets classify images and transfer to new tasks.

Figure 4.1 — Architectures, pretrained inference and transfer learning.

Choosing a backbone

Size versus accuracy versus speed.

family          strengths                                 typical use
ResNet-18/50    reliable, well understood, many weights   baselines, transfer learning
MobileNetV3     tiny and fast                              phones, embedded devices
EfficientNet    accuracy per compute                       balanced deployments
ConvNeXt        modern CNN, strong accuracy                server-side vision
ViT / DINO      strong pretrained features, scales well    large data, foundation models

Start small

Try a small pretrained backbone first; move to larger ones only if validation results justify the extra latency and cost.

Quick check: What made very deep CNNs like ResNet trainable?

  • Training without labels
  • Removing all convolutions
  • Using grayscale only
  • Residual (skip) connections
Answer

Residual (skip) connections — Skip connections keep gradients flowing.