# Bounding Boxes and IoU — Computer Vision

Source: https://www.skillbyai.com/en/computer-vision/d-iou

> How much do two boxes overlap?

## Intersection over union

Detectors describe objects with **bounding boxes**, usually as (x1, y1, x2, y2) corners or (centre x, centre y, width, height); always check which format a library uses. **Intersection over Union (IoU)** measures overlap: the area of the intersection divided by the area of the union, from 0 (no overlap) to 1 (identical). IoU decides whether a prediction matches a ground-truth object (commonly at 0.5 or higher), which predictions are duplicates, and how accurate localisation is.

## Where and what

Detectors output boxes, labels and scores; IoU, NMS and average precision make them usable and measurable.

![Four ideas: IoU, NMS, running a detector, average precision.](assets/figures/computer-vision/section-5-map.svg) — Figure 5.1 — IoU, NMS, detection and AP.

## IoU for five predicted boxes, run

I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. Against a 100 x 100 ground-truth box, a perfect prediction scores 1.000, a 20-pixel shift 0.667, a box 1.5 times too large 0.444, a half overlap 0.333 and a miss 0.000.

```python
def iou(a, b):   # boxes as (x1, y1, x2, y2)
    ix1, iy1, ix2, iy2 = max(a[0], b[0]), max(a[1], b[1]), min(a[2], b[2]), min(a[3], b[3])
    inter = max(0, ix2 - ix1) * max(0, iy2 - iy1)
    area = lambda r: (r[2] - r[0]) * (r[3] - r[1])
    return inter / (area(a) + area(b) - inter)
truth = (50, 50, 150, 150)
for name, pred in [("perfect", (50, 50, 150, 150)), ("shifted 20px", (70, 50, 170, 150)),
                   ("too big", (25, 25, 175, 175)), ("half overlap", (100, 50, 200, 150)), ("miss", (200, 200, 260, 260))]:
    print(f"{name:<13} IoU {iou(truth, pred):.3f}")
```

Output:

```
perfect       IoU 1.000
shifted 20px  IoU 0.667
too big       IoU 0.444
half overlap  IoU 0.333
miss          IoU 0.000
```

## Check the box format

Mixing (x, y, w, h) with (x1, y1, x2, y2) silently produces nonsense IoU values.

**Quiz:** A prediction shifted by 20 pixels on a 100-pixel box has IoU of about...

- [ ] 2.0
- [ ] 1.0
- [ ] 0.0
- [x] 0.67

*Answer:* 0.67. Overlap shrinks quickly with shifts.
