Lesson 18 / 26
Semantic, Instance and Panoptic Segmentation
Which pixels, which objects.
Three granularities
Semantic segmentation assigns a class to every pixel (road, car, person) without separating individual objects. Instance segmentation gives each object its own mask (person 1, person 2). Panoptic segmentation combines both: every pixel gets a class, and countable objects also get instance ids. Typical architectures: U-Net (encoder-decoder with skip connections, popular in medical imaging), DeepLab (atrous convolutions), Mask R-CNN (instance masks on top of detection) and transformer models such as Mask2Former. Promptable models like Segment Anything produce masks from clicks or boxes.
Pixel-level understanding
Segmentation labels every pixel, giving precise shapes and areas.
Choosing a segmentation type
Match the output to the question.
question type output
what share of the field is crop vs weed? semantic class map
how many cells, and how big is each? instance one mask per object
full scene understanding for a robot panoptic class map + instance ids
cut out the object the user clicked promptable mask for that promptLabel efficiently
Pixel labels are expensive; use model-assisted labelling (predict, then correct) and polygon tools to save time.
Quick check: Which segmentation type separates two touching people into different masks?
- Instance segmentation
- Semantic segmentation only
- Image classification
- Edge detection
Answer
Instance segmentation — Instances are individual objects.