Lesson 20 / 25

The Transformer Block

Attention plus a feed-forward network, with residuals and normalisation.

A repeatable unit

A transformer is a stack of identical blocks. Each block has multi-head self-attention (tokens exchange information) and a feed-forward network applied to each token, both wrapped with residual connections and layer normalisation. Because attention has no built-in notion of order, positional information is added (learned or sinusoidal positions, or rotary position embeddings). Encoder stacks (like BERT) read whole inputs; decoder stacks (like GPT-style models) generate text one token at a time using a causal mask so tokens cannot see the future.

Modern building blocks

Most current systems stack transformer blocks and start from pretrained models.

Three ideas: transformer blocks, transfer learning, generative models.
Figure 7.1 — Transformers, transfer and generation.

A two-layer transformer encoder in PyTorch, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. A 2-layer encoder with 64-dimensional tokens and 4 heads keeps the shape (8, 20, 64): every token gets a new, context-aware vector. This small stack has 66,944 parameters; large language models have billions.

import torch
torch.manual_seed(0)
layer = torch.nn.TransformerEncoderLayer(d_model=64, nhead=4, dim_feedforward=128, batch_first=True)
encoder = torch.nn.TransformerEncoder(layer, num_layers=2)
x = torch.randn(8, 20, 64)                       # batch of 8 sequences, 20 tokens, 64-d each
print("input :", tuple(x.shape))
print("output:", tuple(encoder(x).shape))
print("parameters:", sum(p.numel() for p in encoder.parameters()))

Output:

input : (8, 20, 64)
output: (8, 20, 64)
parameters: 66944

Use maintained implementations

For real projects, start from established libraries and pretrained checkpoints rather than writing attention from scratch.

Quick check: Why do transformers need positional information?

  • Attention only works on images
  • Self-attention by itself does not know token order
  • To reduce parameters
  • Because embeddings are random
Answer

Self-attention by itself does not know token order — Positions are added to the token vectors.