Lesson 10 / 26

Datasets and Labelling

Good labels beat clever models.

Coverage, consistency, splits

Vision models learn what the data shows. Collect images that cover real conditions (lighting, angles, devices, backgrounds, rare defects), write labelling guidelines with examples (how tight should boxes be? what about partly visible objects?), measure agreement between labellers, and review a sample of labels. Split data so near-identical images (frames from the same video, photos of the same item) do not land in both train and test. Public datasets such as ImageNet and COCO are useful for pretraining and benchmarks, but check their licences and how well they match your domain.

A labelling guideline excerpt

Concrete rules make labels consistent.

object: "damaged package"
- box covers the whole package, tight to its visible edges
- label if any tear, crush or water stain is visible
- partly visible packages (>= 30% visible): label; otherwise skip
- blurry images: mark "unusable", do not guess
- ambiguous cases: flag for review, do not choose arbitrarily
quality: 10% of images double-labelled; target box IoU agreement >= 0.8

Split by source, not by frame

Group frames by video or by physical item before splitting, or test scores will be inflated by near-duplicates.

Quick check: Why should frames from the same video not be split across train and test?

  • Near-duplicate frames inflate test scores
  • Videos cannot be labelled
  • Frames are always blank
  • It reduces training speed
Answer

Near-duplicate frames inflate test scores — Keep related images on one side of the split.