# Geometric Transforms — Computer Vision

Source: https://www.skillbyai.com/en/computer-vision/g-affine

> Resize, rotate, warp.

## Matrices move pixels

Resizing, rotating, translating and shearing are **affine transforms**, described by a 2x3 matrix applied to pixel coordinates; **perspective (homography)** transforms use a 3x3 matrix and can, for example, flatten a photographed document. When resizing, choose interpolation deliberately (area for shrinking, linear or cubic for enlarging) and decide whether to keep the aspect ratio (letterboxing) or stretch. Remember that boxes, masks and keypoints must be transformed with the image.

## Transform images, multiply data, label well

Geometric transforms reshape images; augmentation and good labels decide how well models learn.

![Three ideas: transforms, augmentation, datasets.](assets/figures/computer-vision/section-3-map.svg) — Figure 3.1 — Transforms, augmentation and datasets.

## A rotation-and-scale matrix applied to points, run

I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. A 90-degree rotation with scale 2 about the origin maps (100, 0) to (0, -200) and (0, 50) to (100, 0). Note that OpenCV's resize takes (width, height) but the resulting array shape is (height, width): (32, 64).

```python
import cv2, numpy as np
pts = np.float32([[0, 0], [100, 0], [0, 50]])          # three corners of a shape
M = cv2.getRotationMatrix2D(center=(0, 0), angle=90, scale=2.0)
out = cv2.transform(pts[None], M)[0]
print("rotation+scale matrix:\n", M.round(3))
for p, q in zip(pts, out):
    print(p.tolist(), "->", q.round(1).tolist())
img = np.zeros((100, 200), np.uint8)
print("resize 200x100 -> 64x32 shape:", cv2.resize(img, (64, 32)).shape)
```

Output:

```
rotation+scale matrix:
 [[ 0.  2.  0.]
 [-2.  0.  0.]]
[0.0, 0.0] -> [0.0, 0.0]
[100.0, 0.0] -> [0.0, -200.0]
[0.0, 50.0] -> [100.0, 0.0]
resize 200x100 -> 64x32 shape: (32, 64)
```

## Transform labels with images

Apply the same geometric transform to boxes, masks and keypoints, or your training labels will no longer match the pixels.

**Quiz:** What must happen to bounding boxes when an image is rotated for training?

- [ ] They should be deleted
- [ ] Nothing
- [x] They must be transformed with the same rotation
- [ ] They should be doubled in size

*Answer:* They must be transformed with the same rotation. Labels must follow the pixels.
