Lesson 17 / 26
Evaluating Detectors: Precision, Recall and AP
Match predictions to ground truth, then rank.
Average precision and mAP
To evaluate a detector, sort predictions by score, match each to an unmatched ground-truth box of the same class with IoU at least a threshold (true positive), and count the rest as false positives; unmatched ground truths are misses. Computing precision and recall as you go down the ranking gives a precision-recall curve; average precision (AP) summarises it. mAP averages AP over classes, and COCO-style mAP also averages over IoU thresholds from 0.5 to 0.95. Use standard tools (pycocotools, torchmetrics) for official numbers; the simple version below shows the idea.
A hand-computed AP at IoU 0.5, run
I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. Five predictions for three objects: ranks 1, 2 and 4 are correct, rank 3 is a false positive and rank 5 duplicates an object already found, so it counts as a false positive. Precision ends at 0.60 with recall 1.00, and the simple, non-interpolated AP is 0.917.
def iou(a, b):
ix = max(0, min(a[2], b[2]) - max(a[0], b[0])); iy = max(0, min(a[3], b[3]) - max(a[1], b[1]))
inter = ix * iy; area = lambda r: (r[2] - r[0]) * (r[3] - r[1])
return inter / (area(a) + area(b) - inter)
gts = [(10, 10, 60, 60), (100, 100, 160, 170), (200, 40, 260, 90)]
preds = [((12, 11, 61, 58), 0.95), ((105, 98, 158, 168), 0.90), ((300, 300, 340, 340), 0.85),
((198, 45, 255, 95), 0.60), ((11, 12, 59, 61), 0.55)] # last one duplicates a found object
matched, tp_flags = set(), []
for box, score in sorted(preds, key=lambda p: -p[1]):
best = max(range(len(gts)), key=lambda g: iou(box, gts[g]))
ok = iou(box, gts[best]) >= 0.5 and best not in matched
if ok: matched.add(best)
tp_flags.append(ok)
tp = cum = 0; ap = 0.0; prev_recall = 0.0
print("rank score TP precision recall")
for i, (flag, (_, score)) in enumerate(zip(tp_flags, sorted(preds, key=lambda p: -p[1])), 1):
tp += flag; precision, recall = tp / i, tp / len(gts)
if flag: ap += precision * (recall - prev_recall); prev_recall = recall
print(f"{i:>4} {score:.2f} {'Y' if flag else 'N'} {precision:9.2f} {recall:6.2f}")
print(f"average precision at IoU 0.5 (simple, non-interpolated): {ap:.3f}")
Output:
rank score TP precision recall 1 0.95 Y 1.00 0.33 2 0.90 Y 1.00 0.67 3 0.85 N 0.67 0.67 4 0.60 Y 0.75 1.00 5 0.55 N 0.60 1.00 average precision at IoU 0.5 (simple, non-interpolated): 0.917
Report per-class AP
mAP can hide a class that is never detected; inspect AP per class and per object size.
Quick check: Why is a second box on an already-matched object counted as a false positive?
- It has a low IoU
- Boxes cannot overlap
- Each ground-truth object can be matched only once
- All second boxes are correct
Answer
Each ground-truth object can be matched only once — Duplicates are penalised.