# Running a Pretrained Detector — Computer Vision

Source: https://www.skillbyai.com/en/computer-vision/d-run

> Boxes, labels and scores from one call.

## Detector families

**Two-stage** detectors (Faster R-CNN) propose regions then classify them: accurate, slower. **One-stage** detectors (SSD, RetinaNet, the YOLO family) predict boxes directly: fast, widely used in real time. **Transformer-based** detectors (DETR and successors) predict a set of objects without hand-crafted anchors. Pretrained COCO models detect 80 everyday categories; for your own objects you fine-tune on labelled boxes. Always apply a **score threshold** suited to your use case.

## Faster R-CNN (MobileNetV3) on the astronaut photo, run

I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. The lightweight Faster R-CNN returns 17 raw detections; only one passes the 0.5 score threshold: a person at 0.99 confidence with a box covering most of the photo.

```python
import torch
from torchvision.models.detection import fasterrcnn_mobilenet_v3_large_320_fpn, FasterRCNN_MobileNet_V3_Large_320_FPN_Weights
from skimage import data
weights = FasterRCNN_MobileNet_V3_Large_320_FPN_Weights.DEFAULT
model = fasterrcnn_mobilenet_v3_large_320_fpn(weights=weights).eval()
img = torch.tensor(data.astronaut()).permute(2, 0, 1).float() / 255
with torch.no_grad():
    out = model([img])[0]
print("raw detections:", len(out["boxes"]))
for box, label, score in zip(out["boxes"], out["labels"], out["scores"]):
    if score >= 0.5:
        print(weights.meta["categories"][label], round(score.item(), 2), "box", box.round().int().tolist())
```

Output:

```
raw detections: 17
person 0.99 box [18, 16, 335, 482]
```

## Choose the threshold from costs

Lower thresholds catch more objects but add false alarms; set them per class from validation precision and recall.

**Quiz:** Why apply a score threshold to detector outputs?

- [ ] Thresholds make boxes larger
- [ ] Scores are always 1
- [x] Detectors return many low-confidence boxes that are usually wrong
- [ ] It is needed to load the model

*Answer:* Detectors return many low-confidence boxes that are usually wrong. Raw outputs include many weak guesses.
