Large Language Models
Understand how LLMs work from tokens and embeddings to attention, training, prompting, evaluation, safety and deployment, with small runnable models you can verify.
Syllabus
What an LLM Is
- What Is a Large Language Model
- The Life of a Model: Pretraining to Chat
- Tokens and Byte-Pair Encoding
- What LLMs Are Good and Bad At
From Text to Probabilities
- Embeddings and Similarity
- Logits, Softmax and Temperature
- Sampling: Greedy, Top-k and Top-p
- A Tiny Language Model You Can Run
- Loss and Perplexity
The Transformer
- Self-Attention: Queries, Keys, Values
- Causal Masking and Generation
- Layers, Feed-Forward Blocks and Positions
- Context Window and the KV Cache
Training and Adapting Models
- Scaling Laws and Data
- Fine-Tuning, LoRA and When Not to Fine-Tune
- Alignment: RLHF and Preference Optimisation
Building with LLMs
- Prompting Basics
- Retrieval-Augmented Generation (RAG)
- Tool Use and Structured Output
- Cost, Latency and Caching