Lesson 21 / 26
Robustness to Real-World Conditions
Test with the images you will actually get.
Rotation, blur, noise, lighting
Models trained on clean, well-lit images can degrade with motion blur, sensor noise, unusual angles, low light, compression artefacts or new camera types. Build a robustness test set with realistic perturbations and new conditions, measure performance per condition, and add matching augmentation or data. A model can keep its top label while losing confidence, so track probabilities too. Also watch for distribution shift after deployment, such as a new camera or season.
Real-world conditions
Models must cope with real images, consistent preprocessing and hardware limits.
One photo under five conditions, run
I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. ResNet-18 keeps a cat label for the original, rotated, heavily blurred and darkened versions (combined cat probability 0.96 to 1.00), but with strong noise the cat probability drops to 0.45 even though the top label is still a cat breed.
import torch, cv2, numpy as np
from torchvision.models import resnet18, ResNet18_Weights
from skimage import data
weights = ResNet18_Weights.DEFAULT
model = resnet18(weights=weights).eval(); prep = weights.transforms(); cats = weights.meta["categories"]
base = data.chelsea()
rng = np.random.default_rng(0)
variants = {
"original": base,
"rotated 90": np.ascontiguousarray(np.rot90(base)),
"heavy blur": cv2.GaussianBlur(base, (31, 31), 0),
"strong noise": np.clip(base + rng.normal(0, 60, base.shape), 0, 255).astype(np.uint8),
"dark (x0.2)": (base * 0.2).astype(np.uint8),
}
for name, img in variants.items():
with torch.no_grad():
p = model(prep(torch.tensor(img).permute(2, 0, 1))[None]).softmax(1)[0]
cat_mass = p[281:286].sum().item() # ImageNet indices 281-285 are domestic cat classes
print(f"{name:<13} top-1 {cats[p.argmax()]:<16} prob of any cat class {cat_mass:.2f}")
Output:
original top-1 Egyptian cat prob of any cat class 1.00 rotated 90 top-1 Egyptian cat prob of any cat class 0.96 heavy blur top-1 tiger cat prob of any cat class 0.98 strong noise top-1 Persian cat prob of any cat class 0.45 dark (x0.2) top-1 Egyptian cat prob of any cat class 0.98
Collect images from the real device
A few hundred images from the actual cameras and locations are worth more than thousands of web images.
Quick check: Why track probabilities, not only the top label, in robustness tests?
- Probabilities never change
- Confidence can collapse even when the top label stays the same
- Top labels are always wrong
- It makes models faster
Answer
Confidence can collapse even when the top label stays the same — Falling confidence is an early warning.