Lesson 18 / 26

Semantic, Instance and Panoptic Segmentation

Which pixels, which objects.

Three granularities

Semantic segmentation assigns a class to every pixel (road, car, person) without separating individual objects. Instance segmentation gives each object its own mask (person 1, person 2). Panoptic segmentation combines both: every pixel gets a class, and countable objects also get instance ids. Typical architectures: U-Net (encoder-decoder with skip connections, popular in medical imaging), DeepLab (atrous convolutions), Mask R-CNN (instance masks on top of detection) and transformer models such as Mask2Former. Promptable models like Segment Anything produce masks from clicks or boxes.

Pixel-level understanding

Segmentation labels every pixel, giving precise shapes and areas.

Three ideas: kinds of segmentation, running a model, metrics.
Figure 6.1 — Kinds, models and metrics.

Choosing a segmentation type

Match the output to the question.

question                                   type          output
what share of the field is crop vs weed?   semantic      class map
how many cells, and how big is each?       instance      one mask per object
full scene understanding for a robot       panoptic      class map + instance ids
cut out the object the user clicked        promptable    mask for that prompt

Label efficiently

Pixel labels are expensive; use model-assisted labelling (predict, then correct) and polygon tools to save time.

Quick check: Which segmentation type separates two touching people into different masks?

  • Instance segmentation
  • Semantic segmentation only
  • Image classification
  • Edge detection
Answer

Instance segmentation — Instances are individual objects.