SkillByAIOpen interactive version →

Lesson 24 / 26

Vision Foundation Models

One pretrained model, many tasks.

Zero-shot, promptable, multimodal

Large models pretrained on huge image or image-text collections are reshaping vision. CLIP-style models embed images and text in the same space, enabling zero-shot classification (compare an image with text prompts like "a photo of a damaged box") and image search. Self-supervised models such as DINO provide strong general features. Promptable segmentation (Segment Anything) creates masks from clicks or boxes. Vision-language models answer questions about images and read documents. They speed up prototyping and labelling, but still need evaluation on your data, and can be large and costly.

New capabilities, real responsibilities

Large pretrained vision and vision-language models change workflows; privacy and fairness need deliberate attention.

Figure 8.1 — Foundation models, ethics and checklist.

Zero-shot classification with a CLIP-style model (sketch)

Shows the idea with the open_clip package; not run here because it needs extra model downloads.

import open_clip, torch
model, _, preprocess = open_clip.create_model_and_transforms("ViT-B-32", pretrained="laion2b_s34b_b79k")
tokenizer = open_clip.get_tokenizer("ViT-B-32")
labels = ["a photo of an intact box", "a photo of a damaged box"]
with torch.no_grad():
    img = model.encode_image(preprocess(image)[None])
    txt = model.encode_text(tokenizer(labels))
    img, txt = img / img.norm(dim=-1, keepdim=True), txt / txt.norm(dim=-1, keepdim=True)
    probs = (100 * img @ txt.T).softmax(-1)
print(dict(zip(labels, probs[0].tolist())))

Use foundation models to bootstrap labels

Let a promptable or zero-shot model pre-label images, then have people correct them; it cuts labelling time sharply.

Quick check: What does zero-shot classification with CLIP-style models need?

  • A segmentation mask
  • Thousands of labelled images per class
  • Text descriptions of the classes, not labelled training images
  • A new GPU architecture
Answer

Text descriptions of the classes, not labelled training images — Images are compared with text prompts.