SkillByAIOpen interactive version →

Lesson 13 / 26

Transfer Learning: Linear Probes and Fine-Tuning

Reuse features for your own classes.

Features first, fine-tune if needed

For your own classes (defect types, plant diseases), remove the pretrained classifier head and use the network as a feature extractor. A linear probe trains only a simple classifier on those features and often works well with a few hundred images. If that is not enough, fine-tune: unfreeze some or all layers and train with a small learning rate, using augmentation. Pretrained features are much more informative than raw pixels because they already encode shapes and textures.

Raw pixels versus pretrained features, run

I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. Eighty random grayscale crops from the cat and coffee photos are classified with logistic regression. Subsampled raw pixels reach 0.912 in 5-fold cross-validation; frozen ResNet-18 features reach 1.000. A toy task, but it shows how much the pretrained features help.

import torch, numpy as np
from torchvision.models import resnet18, ResNet18_Weights
from torchvision.transforms import v2
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
from skimage import data
torch.manual_seed(0)
weights = ResNet18_Weights.DEFAULT
backbone = resnet18(weights=weights).eval(); backbone.fc = torch.nn.Identity()   # 512-d features
crop = v2.Compose([v2.RandomResizedCrop(224, scale=(0.08, 0.3), antialias=True), v2.Grayscale(num_output_channels=3),
                   v2.ToDtype(torch.float32, scale=True),
                   v2.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])])
X, y = [], []
for label, img in enumerate([data.chelsea(), data.coffee()]):
    t = torch.tensor(img).permute(2, 0, 1)
    for _ in range(40):
        X.append(crop(t)); y.append(label)
with torch.no_grad():
    feats = backbone(torch.stack(X)).numpy()
pixels = torch.stack(X).flatten(1).numpy()[:, ::50]
for name, F in [("raw pixels (subsampled)", pixels), ("pretrained ResNet features", feats)]:
    print(f"{name:<27} 5-fold accuracy {cross_val_score(LogisticRegression(max_iter=2000), F, y, cv=5).mean():.3f}")

Output:

raw pixels (subsampled)     5-fold accuracy 0.912
pretrained ResNet features  5-fold accuracy 1.000

Try a linear probe first

A linear probe trains in seconds and gives a strong baseline before any fine-tuning.

Quick check: What does a linear probe train?

  • Every layer of the network from scratch
  • Only a simple classifier on top of frozen pretrained features
  • The image sensor
  • A new tokenizer
Answer

Only a simple classifier on top of frozen pretrained features — Cheap, fast and often strong.