Lesson 20 / 25
The Transformer Block
Attention plus a feed-forward network, with residuals and normalisation.
A repeatable unit
A transformer is a stack of identical blocks. Each block has multi-head self-attention (tokens exchange information) and a feed-forward network applied to each token, both wrapped with residual connections and layer normalisation. Because attention has no built-in notion of order, positional information is added (learned or sinusoidal positions, or rotary position embeddings). Encoder stacks (like BERT) read whole inputs; decoder stacks (like GPT-style models) generate text one token at a time using a causal mask so tokens cannot see the future.
Modern building blocks
Most current systems stack transformer blocks and start from pretrained models.
A two-layer transformer encoder in PyTorch, run
I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. A 2-layer encoder with 64-dimensional tokens and 4 heads keeps the shape (8, 20, 64): every token gets a new, context-aware vector. This small stack has 66,944 parameters; large language models have billions.
import torch
torch.manual_seed(0)
layer = torch.nn.TransformerEncoderLayer(d_model=64, nhead=4, dim_feedforward=128, batch_first=True)
encoder = torch.nn.TransformerEncoder(layer, num_layers=2)
x = torch.randn(8, 20, 64) # batch of 8 sequences, 20 tokens, 64-d each
print("input :", tuple(x.shape))
print("output:", tuple(encoder(x).shape))
print("parameters:", sum(p.numel() for p in encoder.parameters()))
Output:
input : (8, 20, 64) output: (8, 20, 64) parameters: 66944
Use maintained implementations
For real projects, start from established libraries and pretrained checkpoints rather than writing attention from scratch.
Quick check: Why do transformers need positional information?
- Attention only works on images
- Self-attention by itself does not know token order
- To reduce parameters
- Because embeddings are random
Answer
Self-attention by itself does not know token order — Positions are added to the token vectors.