पाठ 16 / 26
Running a Pretrained Detector
Boxes, labels and scores from one call.
Detector families
Two-stage detectors (Faster R-CNN) propose regions then classify them: accurate, slower. One-stage detectors (SSD, RetinaNet, the YOLO family) predict boxes directly: fast, widely used in real time. Transformer-based detectors (DETR and successors) predict a set of objects without hand-crafted anchors. Pretrained COCO models detect 80 everyday categories; for your own objects you fine-tune on labelled boxes. Always apply a score threshold suited to your use case.
Faster R-CNN (MobileNetV3) on the astronaut photo, run
I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. The lightweight Faster R-CNN returns 17 raw detections; only one passes the 0.5 score threshold: a person at 0.99 confidence with a box covering most of the photo.
import torch
from torchvision.models.detection import fasterrcnn_mobilenet_v3_large_320_fpn, FasterRCNN_MobileNet_V3_Large_320_FPN_Weights
from skimage import data
weights = FasterRCNN_MobileNet_V3_Large_320_FPN_Weights.DEFAULT
model = fasterrcnn_mobilenet_v3_large_320_fpn(weights=weights).eval()
img = torch.tensor(data.astronaut()).permute(2, 0, 1).float() / 255
with torch.no_grad():
out = model([img])[0]
print("raw detections:", len(out["boxes"]))
for box, label, score in zip(out["boxes"], out["labels"], out["scores"]):
if score >= 0.5:
print(weights.meta["categories"][label], round(score.item(), 2), "box", box.round().int().tolist())
Output:
raw detections: 17 person 0.99 box [18, 16, 335, 482]
Choose the threshold from costs
Lower thresholds catch more objects but add false alarms; set them per class from validation precision and recall.
त्वरित जाँच: Why apply a score threshold to detector outputs?
- Thresholds make boxes larger
- Scores are always 1
- Detectors return many low-confidence boxes that are usually wrong
- It is needed to load the model
Answer
Detectors return many low-confidence boxes that are usually wrong — Raw outputs include many weak guesses.