Lesson 24 / 26
Vision Foundation Models
One pretrained model, many tasks.
Zero-shot, promptable, multimodal
Large models pretrained on huge image or image-text collections are reshaping vision. CLIP-style models embed images and text in the same space, enabling zero-shot classification (compare an image with text prompts like "a photo of a damaged box") and image search. Self-supervised models such as DINO provide strong general features. Promptable segmentation (Segment Anything) creates masks from clicks or boxes. Vision-language models answer questions about images and read documents. They speed up prototyping and labelling, but still need evaluation on your data, and can be large and costly.
New capabilities, real responsibilities
Large pretrained vision and vision-language models change workflows; privacy and fairness need deliberate attention.
Zero-shot classification with a CLIP-style model (sketch)
Shows the idea with the open_clip package; not run here because it needs extra model downloads.
import open_clip, torch
model, _, preprocess = open_clip.create_model_and_transforms("ViT-B-32", pretrained="laion2b_s34b_b79k")
tokenizer = open_clip.get_tokenizer("ViT-B-32")
labels = ["a photo of an intact box", "a photo of a damaged box"]
with torch.no_grad():
img = model.encode_image(preprocess(image)[None])
txt = model.encode_text(tokenizer(labels))
img, txt = img / img.norm(dim=-1, keepdim=True), txt / txt.norm(dim=-1, keepdim=True)
probs = (100 * img @ txt.T).softmax(-1)
print(dict(zip(labels, probs[0].tolist())))Use foundation models to bootstrap labels
Let a promptable or zero-shot model pre-label images, then have people correct them; it cuts labelling time sharply.
Quick check: What does zero-shot classification with CLIP-style models need?
- A segmentation mask
- Thousands of labelled images per class
- Text descriptions of the classes, not labelled training images
- A new GPU architecture
Answer
Text descriptions of the classes, not labelled training images — Images are compared with text prompts.