# Evaluating Detectors: Precision, Recall and AP — Computer Vision

Source: https://www.skillbyai.com/en/computer-vision/d-ap

> Match predictions to ground truth, then rank.

## Average precision and mAP

To evaluate a detector, sort predictions by score, match each to an unmatched ground-truth box of the same class with IoU at least a threshold (true positive), and count the rest as false positives; unmatched ground truths are misses. Computing **precision and recall** as you go down the ranking gives a precision-recall curve; **average precision (AP)** summarises it. **mAP** averages AP over classes, and COCO-style mAP also averages over IoU thresholds from 0.5 to 0.95. Use standard tools (pycocotools, torchmetrics) for official numbers; the simple version below shows the idea.

## A hand-computed AP at IoU 0.5, run

I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. Five predictions for three objects: ranks 1, 2 and 4 are correct, rank 3 is a false positive and rank 5 duplicates an object already found, so it counts as a false positive. Precision ends at 0.60 with recall 1.00, and the simple, non-interpolated AP is 0.917.

```python
def iou(a, b):
    ix = max(0, min(a[2], b[2]) - max(a[0], b[0])); iy = max(0, min(a[3], b[3]) - max(a[1], b[1]))
    inter = ix * iy; area = lambda r: (r[2] - r[0]) * (r[3] - r[1])
    return inter / (area(a) + area(b) - inter)
gts = [(10, 10, 60, 60), (100, 100, 160, 170), (200, 40, 260, 90)]
preds = [((12, 11, 61, 58), 0.95), ((105, 98, 158, 168), 0.90), ((300, 300, 340, 340), 0.85),
         ((198, 45, 255, 95), 0.60), ((11, 12, 59, 61), 0.55)]      # last one duplicates a found object
matched, tp_flags = set(), []
for box, score in sorted(preds, key=lambda p: -p[1]):
    best = max(range(len(gts)), key=lambda g: iou(box, gts[g]))
    ok = iou(box, gts[best]) >= 0.5 and best not in matched
    if ok: matched.add(best)
    tp_flags.append(ok)
tp = cum = 0; ap = 0.0; prev_recall = 0.0
print("rank score  TP  precision recall")
for i, (flag, (_, score)) in enumerate(zip(tp_flags, sorted(preds, key=lambda p: -p[1])), 1):
    tp += flag; precision, recall = tp / i, tp / len(gts)
    if flag: ap += precision * (recall - prev_recall); prev_recall = recall
    print(f"{i:>4} {score:.2f}  {'Y' if flag else 'N'}  {precision:9.2f} {recall:6.2f}")
print(f"average precision at IoU 0.5 (simple, non-interpolated): {ap:.3f}")
```

Output:

```
rank score  TP  precision recall
   1 0.95  Y       1.00   0.33
   2 0.90  Y       1.00   0.67
   3 0.85  N       0.67   0.67
   4 0.60  Y       0.75   1.00
   5 0.55  N       0.60   1.00
average precision at IoU 0.5 (simple, non-interpolated): 0.917
```

## Report per-class AP

mAP can hide a class that is never detected; inspect AP per class and per object size.

**Quiz:** Why is a second box on an already-matched object counted as a false positive?

- [ ] It has a low IoU
- [ ] Boxes cannot overlap
- [x] Each ground-truth object can be matched only once
- [ ] All second boxes are correct

*Answer:* Each ground-truth object can be matched only once. Duplicates are penalised.
