SkillByAIOpen interactive version →

Lesson 21 / 25

Transfer Learning and Fine-Tuning

Start from a model that already knows a lot.

Reuse, then adapt

Transfer learning starts from a model pretrained on a large dataset (images, text, audio) and adapts it to your task. The cheapest form freezes the pretrained backbone and trains only a new output head; fine-tuning then unfreezes some or all layers with a small learning rate. This needs far less data and compute than training from scratch and is the default approach in practice. Parameter-efficient methods such as LoRA update small added matrices instead of all weights.

Freezing a backbone and training a new head, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. With the backbone frozen, only the new 3-class head is trainable: 1,539 of 526,851 parameters (0.29%), and the optimiser receives just those two tensors.

import torch
backbone = torch.nn.Sequential(torch.nn.Linear(512, 512), torch.nn.ReLU(), torch.nn.Linear(512, 512), torch.nn.ReLU())
head = torch.nn.Linear(512, 3)                   # new task: 3 classes
for p in backbone.parameters():
    p.requires_grad = False                      # freeze the pretrained part
model = torch.nn.Sequential(backbone, head)
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total = sum(p.numel() for p in model.parameters())
print(f"trainable {trainable:,} of {total:,} parameters ({trainable / total:.2%})")
opt = torch.optim.Adam([p for p in model.parameters() if p.requires_grad], lr=1e-3)
print("optimizer parameter tensors:", len(opt.param_groups[0]["params"]))

Output:

trainable 1,539 of 526,851 parameters (0.29%)
optimizer parameter tensors: 2

Match preprocessing to the pretrained model

Use the same image normalisation or tokenizer the model was pretrained with; mismatches silently hurt accuracy.

Quick check: Why is transfer learning usually preferred over training from scratch?

  • It never needs a validation set
  • It needs much less data and compute because the model already has useful features
  • It makes models smaller automatically
  • Pretrained models cannot be changed
Answer

It needs much less data and compute because the model already has useful features — Start from what the model already knows.