पाठ 17 / 26

Evaluating Detectors: Precision, Recall and AP

Match predictions to ground truth, then rank.

Average precision and mAP

To evaluate a detector, sort predictions by score, match each to an unmatched ground-truth box of the same class with IoU at least a threshold (true positive), and count the rest as false positives; unmatched ground truths are misses. Computing precision and recall as you go down the ranking gives a precision-recall curve; average precision (AP) summarises it. mAP averages AP over classes, and COCO-style mAP also averages over IoU thresholds from 0.5 to 0.95. Use standard tools (pycocotools, torchmetrics) for official numbers; the simple version below shows the idea.

A hand-computed AP at IoU 0.5, run

I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. Five predictions for three objects: ranks 1, 2 and 4 are correct, rank 3 is a false positive and rank 5 duplicates an object already found, so it counts as a false positive. Precision ends at 0.60 with recall 1.00, and the simple, non-interpolated AP is 0.917.

def iou(a, b):
    ix = max(0, min(a[2], b[2]) - max(a[0], b[0])); iy = max(0, min(a[3], b[3]) - max(a[1], b[1]))
    inter = ix * iy; area = lambda r: (r[2] - r[0]) * (r[3] - r[1])
    return inter / (area(a) + area(b) - inter)
gts = [(10, 10, 60, 60), (100, 100, 160, 170), (200, 40, 260, 90)]
preds = [((12, 11, 61, 58), 0.95), ((105, 98, 158, 168), 0.90), ((300, 300, 340, 340), 0.85),
         ((198, 45, 255, 95), 0.60), ((11, 12, 59, 61), 0.55)]      # last one duplicates a found object
matched, tp_flags = set(), []
for box, score in sorted(preds, key=lambda p: -p[1]):
    best = max(range(len(gts)), key=lambda g: iou(box, gts[g]))
    ok = iou(box, gts[best]) >= 0.5 and best not in matched
    if ok: matched.add(best)
    tp_flags.append(ok)
tp = cum = 0; ap = 0.0; prev_recall = 0.0
print("rank score  TP  precision recall")
for i, (flag, (_, score)) in enumerate(zip(tp_flags, sorted(preds, key=lambda p: -p[1])), 1):
    tp += flag; precision, recall = tp / i, tp / len(gts)
    if flag: ap += precision * (recall - prev_recall); prev_recall = recall
    print(f"{i:>4} {score:.2f}  {'Y' if flag else 'N'}  {precision:9.2f} {recall:6.2f}")
print(f"average precision at IoU 0.5 (simple, non-interpolated): {ap:.3f}")

Output:

rank score  TP  precision recall
   1 0.95  Y       1.00   0.33
   2 0.90  Y       1.00   0.67
   3 0.85  N       0.67   0.67
   4 0.60  Y       0.75   1.00
   5 0.55  N       0.60   1.00
average precision at IoU 0.5 (simple, non-interpolated): 0.917

Report per-class AP

mAP can hide a class that is never detected; inspect AP per class and per object size.

त्वरित जाँच: Why is a second box on an already-matched object counted as a false positive?

  • It has a low IoU
  • Boxes cannot overlap
  • Each ground-truth object can be matched only once
  • All second boxes are correct
Answer

Each ground-truth object can be matched only once — Duplicates are penalised.