पाठ 21 / 26

Robustness to Real-World Conditions

Test with the images you will actually get.

Rotation, blur, noise, lighting

Models trained on clean, well-lit images can degrade with motion blur, sensor noise, unusual angles, low light, compression artefacts or new camera types. Build a robustness test set with realistic perturbations and new conditions, measure performance per condition, and add matching augmentation or data. A model can keep its top label while losing confidence, so track probabilities too. Also watch for distribution shift after deployment, such as a new camera or season.

Real-world conditions

Models must cope with real images, consistent preprocessing and hardware limits.

Three ideas: robustness, preprocessing consistency, deployment.
Figure 7.1 — Robustness, consistency and deployment.

One photo under five conditions, run

I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. ResNet-18 keeps a cat label for the original, rotated, heavily blurred and darkened versions (combined cat probability 0.96 to 1.00), but with strong noise the cat probability drops to 0.45 even though the top label is still a cat breed.

import torch, cv2, numpy as np
from torchvision.models import resnet18, ResNet18_Weights
from skimage import data
weights = ResNet18_Weights.DEFAULT
model = resnet18(weights=weights).eval(); prep = weights.transforms(); cats = weights.meta["categories"]
base = data.chelsea()
rng = np.random.default_rng(0)
variants = {
    "original": base,
    "rotated 90": np.ascontiguousarray(np.rot90(base)),
    "heavy blur": cv2.GaussianBlur(base, (31, 31), 0),
    "strong noise": np.clip(base + rng.normal(0, 60, base.shape), 0, 255).astype(np.uint8),
    "dark (x0.2)": (base * 0.2).astype(np.uint8),
}
for name, img in variants.items():
    with torch.no_grad():
        p = model(prep(torch.tensor(img).permute(2, 0, 1))[None]).softmax(1)[0]
    cat_mass = p[281:286].sum().item()          # ImageNet indices 281-285 are domestic cat classes
    print(f"{name:<13} top-1 {cats[p.argmax()]:<16} prob of any cat class {cat_mass:.2f}")

Output:

original      top-1 Egyptian cat     prob of any cat class 1.00
rotated 90    top-1 Egyptian cat     prob of any cat class 0.96
heavy blur    top-1 tiger cat        prob of any cat class 0.98
strong noise  top-1 Persian cat      prob of any cat class 0.45
dark (x0.2)   top-1 Egyptian cat     prob of any cat class 0.98

Collect images from the real device

A few hundred images from the actual cameras and locations are worth more than thousands of web images.

त्वरित जाँच: Why track probabilities, not only the top label, in robustness tests?

  • Probabilities never change
  • Confidence can collapse even when the top label stays the same
  • Top labels are always wrong
  • It makes models faster
Answer

Confidence can collapse even when the top label stays the same — Falling confidence is an early warning.