Lesson 1 / 26

What Computer Vision Does

The main tasks and where they are used.

Classify, detect, segment, track, measure

Computer vision extracts information from images and video. Core tasks: classification (what is in this image?), object detection (what objects, and where, as boxes), segmentation (which pixels belong to which object or class), keypoint and pose estimation, tracking across video frames, optical character recognition, and 3D reconstruction and measurement. Applications range from medical imaging and manufacturing inspection to retail, agriculture, document processing and driver assistance. Classical image processing and modern deep learning both play a role.

From light to numbers

Computer vision turns grids of pixel values into useful decisions.

Three ideas: tasks, pixels, colour spaces.
Figure 1.1 — Tasks, pixels and colour spaces.

Tasks and their outputs

Choose the task from the output you need.

task                    output                              example use
classification          label (+ probability)               is this X-ray normal?
object detection        boxes + labels + scores             count people in a store
semantic segmentation   class per pixel                     road vs sidewalk
instance segmentation   mask per object                     separate touching cells
keypoints / pose        point coordinates                   exercise form feedback
OCR                     text + positions                    read invoices
tracking                object identities over time         traffic flow analysis

Pick the cheapest sufficient output

If you only need to know whether a defect exists, classification is cheaper to label and run than segmentation.

Quick check: Which task outputs bounding boxes with labels?

  • Semantic segmentation
  • Image classification
  • Object detection
  • OCR only
Answer

Object detection — Detection localises objects with boxes.