पाठ 21 / 25

Deep Learning at Scale

Why recent progress was so fast.

Data, compute and architectures

Recent AI progress came from combining large datasets, massive compute (GPUs and specialised accelerators) and architectures that scale well, especially the transformer. Researchers observed fairly predictable improvements as models, data and compute grew together ("scaling laws"), then added instruction tuning and reinforcement learning from human feedback to make models helpful and safer. The same recipe extends beyond text to images, audio, video, code and multimodal models that combine them. Costs, energy use and data rights are now major considerations.

Ingredients of a modern foundation model

A simplified pipeline.

1 pretraining      next-token prediction on very large text/code/image corpora
2 instruction tune  examples of following instructions and answering helpfully
3 preference tune  human or AI preference data (RLHF, DPO and similar)
4 safety work      red-teaming, refusals for harmful requests, evaluations
5 deployment       APIs and apps, with tools, retrieval and monitoring

Use, adapt, or train

Most organisations use or adapt existing foundation models; training from scratch requires resources few can justify.

त्वरित जाँच: Which combination drove recent AI progress?

  • Hand-written expert rules only
  • Large datasets, massive compute and scalable architectures like the transformer
  • Smaller models and less data
  • Removing evaluation
Answer

Large datasets, massive compute and scalable architectures like the transformer — Scale plus architecture plus tuning.