# Vision Foundation Models — Computer Vision

Source: https://www.skillbyai.com/en/computer-vision/f-foundation

> One pretrained model, many tasks.

## Zero-shot, promptable, multimodal

Large models pretrained on huge image or image-text collections are reshaping vision. **CLIP**-style models embed images and text in the same space, enabling **zero-shot classification** (compare an image with text prompts like "a photo of a damaged box") and image search. **Self-supervised** models such as DINO provide strong general features. **Promptable segmentation** (Segment Anything) creates masks from clicks or boxes. **Vision-language models** answer questions about images and read documents. They speed up prototyping and labelling, but still need evaluation on your data, and can be large and costly.

## New capabilities, real responsibilities

Large pretrained vision and vision-language models change workflows; privacy and fairness need deliberate attention.

![Three ideas: foundation models, ethics, checklist.](assets/figures/computer-vision/section-8-map.svg) — Figure 8.1 — Foundation models, ethics and checklist.

## Zero-shot classification with a CLIP-style model (sketch)

Shows the idea with the open_clip package; not run here because it needs extra model downloads.

```python
import open_clip, torch
model, _, preprocess = open_clip.create_model_and_transforms("ViT-B-32", pretrained="laion2b_s34b_b79k")
tokenizer = open_clip.get_tokenizer("ViT-B-32")
labels = ["a photo of an intact box", "a photo of a damaged box"]
with torch.no_grad():
    img = model.encode_image(preprocess(image)[None])
    txt = model.encode_text(tokenizer(labels))
    img, txt = img / img.norm(dim=-1, keepdim=True), txt / txt.norm(dim=-1, keepdim=True)
    probs = (100 * img @ txt.T).softmax(-1)
print(dict(zip(labels, probs[0].tolist())))
```

## Use foundation models to bootstrap labels

Let a promptable or zero-shot model pre-label images, then have people correct them; it cuts labelling time sharply.

**Quiz:** What does zero-shot classification with CLIP-style models need?

- [ ] A segmentation mask
- [ ] Thousands of labelled images per class
- [x] Text descriptions of the classes, not labelled training images
- [ ] A new GPU architecture

*Answer:* Text descriptions of the classes, not labelled training images. Images are compared with text prompts.
