# Data Augmentation — Computer Vision

Source: https://www.skillbyai.com/en/computer-vision/g-augment

> More varied training images for free.

## Random but label-preserving changes

**Augmentation** applies random, label-preserving changes to training images: crops, flips, small rotations, colour jitter, blur, noise, cutout. Each epoch the model sees new variations, which reduces overfitting and builds robustness to conditions it will meet. Evaluation uses fixed, deterministic preprocessing so scores are comparable. Choose augmentations that reflect reality: horizontal flips suit natural photos, not text or some medical images, and strong colour changes can harm tasks where colour matters.

## Random training views versus fixed test preprocessing, run

I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. With torchvision transforms v2, two training passes over the same cat photo produce different 224 x 224 views, while the test pipeline (resize then centre crop) produces identical outputs every time.

```python
import torch
from torchvision.transforms import v2
from skimage import data
torch.manual_seed(0)
img = torch.tensor(data.chelsea()).permute(2, 0, 1)        # (C, H, W) uint8
train_tf = v2.Compose([v2.RandomResizedCrop(224, scale=(0.5, 1.0), antialias=True), v2.RandomHorizontalFlip(),
                       v2.ColorJitter(0.3, 0.3, 0.3), v2.ToDtype(torch.float32, scale=True)])
test_tf = v2.Compose([v2.Resize(256, antialias=True), v2.CenterCrop(224), v2.ToDtype(torch.float32, scale=True)])
a, b = train_tf(img), train_tf(img)
print("original:", tuple(img.shape), "-> augmented:", tuple(a.shape))
print("two training views identical:", torch.equal(a, b))
print("two test views identical    :", torch.equal(test_tf(img), test_tf(img)))
```

Output:

```
original: (3, 300, 451) -> augmented: (3, 224, 224)
two training views identical: False
two test views identical    : True
```

## Visualise augmented batches

Look at a grid of augmented training images before training; it quickly shows augmentations that destroy the label.

**Quiz:** Why is test-time preprocessing kept deterministic?

- [ ] Tests have no images
- [ ] Augmentation is illegal at test time
- [x] So evaluation scores are stable and comparable
- [ ] It makes training faster

*Answer:* So evaluation scores are stable and comparable. Random changes belong in training.
