# Transfer Learning and Fine-Tuning — Deep Learning & Neural Networks

Source: https://www.skillbyai.com/en/deep-learning/m-transfer

> Start from a model that already knows a lot.

## Reuse, then adapt

**Transfer learning** starts from a model **pretrained** on a large dataset (images, text, audio) and adapts it to your task. The cheapest form **freezes** the pretrained backbone and trains only a new output head; **fine-tuning** then unfreezes some or all layers with a small learning rate. This needs far less data and compute than training from scratch and is the default approach in practice. Parameter-efficient methods such as LoRA update small added matrices instead of all weights.

## Freezing a backbone and training a new head, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. With the backbone frozen, only the new 3-class head is trainable: 1,539 of 526,851 parameters (0.29%), and the optimiser receives just those two tensors.

```python
import torch
backbone = torch.nn.Sequential(torch.nn.Linear(512, 512), torch.nn.ReLU(), torch.nn.Linear(512, 512), torch.nn.ReLU())
head = torch.nn.Linear(512, 3)                   # new task: 3 classes
for p in backbone.parameters():
    p.requires_grad = False                      # freeze the pretrained part
model = torch.nn.Sequential(backbone, head)
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total = sum(p.numel() for p in model.parameters())
print(f"trainable {trainable:,} of {total:,} parameters ({trainable / total:.2%})")
opt = torch.optim.Adam([p for p in model.parameters() if p.requires_grad], lr=1e-3)
print("optimizer parameter tensors:", len(opt.param_groups[0]["params"]))
```

Output:

```
trainable 1,539 of 526,851 parameters (0.29%)
optimizer parameter tensors: 2
```

## Match preprocessing to the pretrained model

Use the same image normalisation or tokenizer the model was pretrained with; mismatches silently hurt accuracy.

**Quiz:** Why is transfer learning usually preferred over training from scratch?

- [ ] It never needs a validation set
- [x] It needs much less data and compute because the model already has useful features
- [ ] It makes models smaller automatically
- [ ] Pretrained models cannot be changed

*Answer:* It needs much less data and compute because the model already has useful features. Start from what the model already knows.
