Lesson 10 / 26
Datasets and Labelling
Good labels beat clever models.
Coverage, consistency, splits
Vision models learn what the data shows. Collect images that cover real conditions (lighting, angles, devices, backgrounds, rare defects), write labelling guidelines with examples (how tight should boxes be? what about partly visible objects?), measure agreement between labellers, and review a sample of labels. Split data so near-identical images (frames from the same video, photos of the same item) do not land in both train and test. Public datasets such as ImageNet and COCO are useful for pretraining and benchmarks, but check their licences and how well they match your domain.
A labelling guideline excerpt
Concrete rules make labels consistent.
object: "damaged package"
- box covers the whole package, tight to its visible edges
- label if any tear, crush or water stain is visible
- partly visible packages (>= 30% visible): label; otherwise skip
- blurry images: mark "unusable", do not guess
- ambiguous cases: flag for review, do not choose arbitrarily
quality: 10% of images double-labelled; target box IoU agreement >= 0.8Split by source, not by frame
Group frames by video or by physical item before splitting, or test scores will be inflated by near-duplicates.
Quick check: Why should frames from the same video not be split across train and test?
- Near-duplicate frames inflate test scores
- Videos cannot be labelled
- Frames are always blank
- It reduces training speed
Answer
Near-duplicate frames inflate test scores — Keep related images on one side of the split.