पाठ 16 / 26

Running a Pretrained Detector

Boxes, labels and scores from one call.

Detector families

Two-stage detectors (Faster R-CNN) propose regions then classify them: accurate, slower. One-stage detectors (SSD, RetinaNet, the YOLO family) predict boxes directly: fast, widely used in real time. Transformer-based detectors (DETR and successors) predict a set of objects without hand-crafted anchors. Pretrained COCO models detect 80 everyday categories; for your own objects you fine-tune on labelled boxes. Always apply a score threshold suited to your use case.

Faster R-CNN (MobileNetV3) on the astronaut photo, run

I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. The lightweight Faster R-CNN returns 17 raw detections; only one passes the 0.5 score threshold: a person at 0.99 confidence with a box covering most of the photo.

import torch
from torchvision.models.detection import fasterrcnn_mobilenet_v3_large_320_fpn, FasterRCNN_MobileNet_V3_Large_320_FPN_Weights
from skimage import data
weights = FasterRCNN_MobileNet_V3_Large_320_FPN_Weights.DEFAULT
model = fasterrcnn_mobilenet_v3_large_320_fpn(weights=weights).eval()
img = torch.tensor(data.astronaut()).permute(2, 0, 1).float() / 255
with torch.no_grad():
    out = model([img])[0]
print("raw detections:", len(out["boxes"]))
for box, label, score in zip(out["boxes"], out["labels"], out["scores"]):
    if score >= 0.5:
        print(weights.meta["categories"][label], round(score.item(), 2), "box", box.round().int().tolist())

Output:

raw detections: 17
person 0.99 box [18, 16, 335, 482]

Choose the threshold from costs

Lower thresholds catch more objects but add false alarms; set them per class from validation precision and recall.

त्वरित जाँच: Why apply a score threshold to detector outputs?

  • Thresholds make boxes larger
  • Scores are always 1
  • Detectors return many low-confidence boxes that are usually wrong
  • It is needed to load the model
Answer

Detectors return many low-confidence boxes that are usually wrong — Raw outputs include many weak guesses.