पाठ 21 / 25
Deep Learning at Scale
Why recent progress was so fast.
Data, compute and architectures
Recent AI progress came from combining large datasets, massive compute (GPUs and specialised accelerators) and architectures that scale well, especially the transformer. Researchers observed fairly predictable improvements as models, data and compute grew together ("scaling laws"), then added instruction tuning and reinforcement learning from human feedback to make models helpful and safer. The same recipe extends beyond text to images, audio, video, code and multimodal models that combine them. Costs, energy use and data rights are now major considerations.
Ingredients of a modern foundation model
A simplified pipeline.
1 pretraining next-token prediction on very large text/code/image corpora
2 instruction tune examples of following instructions and answering helpfully
3 preference tune human or AI preference data (RLHF, DPO and similar)
4 safety work red-teaming, refusals for harmful requests, evaluations
5 deployment APIs and apps, with tools, retrieval and monitoringUse, adapt, or train
Most organisations use or adapt existing foundation models; training from scratch requires resources few can justify.
त्वरित जाँच: Which combination drove recent AI progress?
- Hand-written expert rules only
- Large datasets, massive compute and scalable architectures like the transformer
- Smaller models and less data
- Removing evaluation
Answer
Large datasets, massive compute and scalable architectures like the transformer — Scale plus architecture plus tuning.