पाठ 12 / 26
Classifying With a Pretrained Model
Load weights, preprocess exactly, read the top predictions.
Use the model's own preprocessing
Libraries such as torchvision and timm provide models with pretrained weights (often trained on ImageNet's 1,000 classes). To use them correctly, apply the same preprocessing the weights expect (resize, centre crop, scaling, mean and standard deviation normalisation), switch to eval() mode, run under no_grad, and read the softmax probabilities. Look at the top few classes: fine-grained categories (several cat breeds) often share probability, which is expected rather than an error.
ResNet-18 on two bundled photos, run
I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. The cat photo is spread across Egyptian cat (0.391), tiger cat (0.387) and tabby (0.218), all correct as "a cat". The coffee photo is espresso with 0.852 confidence.
import torch
from torchvision.models import resnet18, ResNet18_Weights
from skimage import data
weights = ResNet18_Weights.DEFAULT
model = resnet18(weights=weights).eval()
prep = weights.transforms()
for name, img in [("chelsea", data.chelsea()), ("coffee", data.coffee())]:
x = prep(torch.tensor(img).permute(2, 0, 1))[None]
with torch.no_grad():
probs = model(x).softmax(1)[0]
top = probs.topk(3)
print(name, "->", [(weights.meta["categories"][i], round(p.item(), 3)) for p, i in zip(top.values, top.indices)])
Output:
chelsea -> [('Egyptian cat', 0.391), ('tiger cat', 0.387), ('tabby', 0.218)]
coffee -> [('espresso', 0.852), ('chocolate sauce', 0.041), ('consomme', 0.038)]Use weights.transforms()
torchvision bundles the correct preprocessing with each weights object; use it instead of guessing the normalisation.
त्वरित जाँच: Why must you use the same preprocessing as the pretrained weights?
- It makes images larger
- The model learned from inputs scaled and normalised that way
- Any preprocessing works the same
- It is only needed for training
Answer
The model learned from inputs scaled and normalised that way — Mismatched inputs degrade predictions.