Lesson 21 / 25
Transfer Learning and Fine-Tuning
Start from a model that already knows a lot.
Reuse, then adapt
Transfer learning starts from a model pretrained on a large dataset (images, text, audio) and adapts it to your task. The cheapest form freezes the pretrained backbone and trains only a new output head; fine-tuning then unfreezes some or all layers with a small learning rate. This needs far less data and compute than training from scratch and is the default approach in practice. Parameter-efficient methods such as LoRA update small added matrices instead of all weights.
Freezing a backbone and training a new head, run
I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. With the backbone frozen, only the new 3-class head is trainable: 1,539 of 526,851 parameters (0.29%), and the optimiser receives just those two tensors.
import torch
backbone = torch.nn.Sequential(torch.nn.Linear(512, 512), torch.nn.ReLU(), torch.nn.Linear(512, 512), torch.nn.ReLU())
head = torch.nn.Linear(512, 3) # new task: 3 classes
for p in backbone.parameters():
p.requires_grad = False # freeze the pretrained part
model = torch.nn.Sequential(backbone, head)
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total = sum(p.numel() for p in model.parameters())
print(f"trainable {trainable:,} of {total:,} parameters ({trainable / total:.2%})")
opt = torch.optim.Adam([p for p in model.parameters() if p.requires_grad], lr=1e-3)
print("optimizer parameter tensors:", len(opt.param_groups[0]["params"]))
Output:
trainable 1,539 of 526,851 parameters (0.29%) optimizer parameter tensors: 2
Match preprocessing to the pretrained model
Use the same image normalisation or tokenizer the model was pretrained with; mismatches silently hurt accuracy.
Quick check: Why is transfer learning usually preferred over training from scratch?
- It never needs a validation set
- It needs much less data and compute because the model already has useful features
- It makes models smaller automatically
- Pretrained models cannot be changed
Answer
It needs much less data and compute because the model already has useful features — Start from what the model already knows.