Lesson 13 / 26
Transfer Learning: Linear Probes and Fine-Tuning
Reuse features for your own classes.
Features first, fine-tune if needed
For your own classes (defect types, plant diseases), remove the pretrained classifier head and use the network as a feature extractor. A linear probe trains only a simple classifier on those features and often works well with a few hundred images. If that is not enough, fine-tune: unfreeze some or all layers and train with a small learning rate, using augmentation. Pretrained features are much more informative than raw pixels because they already encode shapes and textures.
Raw pixels versus pretrained features, run
I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. Eighty random grayscale crops from the cat and coffee photos are classified with logistic regression. Subsampled raw pixels reach 0.912 in 5-fold cross-validation; frozen ResNet-18 features reach 1.000. A toy task, but it shows how much the pretrained features help.
import torch, numpy as np
from torchvision.models import resnet18, ResNet18_Weights
from torchvision.transforms import v2
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
from skimage import data
torch.manual_seed(0)
weights = ResNet18_Weights.DEFAULT
backbone = resnet18(weights=weights).eval(); backbone.fc = torch.nn.Identity() # 512-d features
crop = v2.Compose([v2.RandomResizedCrop(224, scale=(0.08, 0.3), antialias=True), v2.Grayscale(num_output_channels=3),
v2.ToDtype(torch.float32, scale=True),
v2.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])])
X, y = [], []
for label, img in enumerate([data.chelsea(), data.coffee()]):
t = torch.tensor(img).permute(2, 0, 1)
for _ in range(40):
X.append(crop(t)); y.append(label)
with torch.no_grad():
feats = backbone(torch.stack(X)).numpy()
pixels = torch.stack(X).flatten(1).numpy()[:, ::50]
for name, F in [("raw pixels (subsampled)", pixels), ("pretrained ResNet features", feats)]:
print(f"{name:<27} 5-fold accuracy {cross_val_score(LogisticRegression(max_iter=2000), F, y, cv=5).mean():.3f}")
Output:
raw pixels (subsampled) 5-fold accuracy 0.912 pretrained ResNet features 5-fold accuracy 1.000
Try a linear probe first
A linear probe trains in seconds and gives a strong baseline before any fine-tuning.
Quick check: What does a linear probe train?
- Every layer of the network from scratch
- Only a simple classifier on top of frozen pretrained features
- The image sensor
- A new tokenizer
Answer
Only a simple classifier on top of frozen pretrained features — Cheap, fast and often strong.