# The Transformer Block — Deep Learning & Neural Networks

Source: https://www.skillbyai.com/en/deep-learning/m-transformer

> Attention plus a feed-forward network, with residuals and normalisation.

## A repeatable unit

A **transformer** is a stack of identical blocks. Each block has **multi-head self-attention** (tokens exchange information) and a **feed-forward network** applied to each token, both wrapped with **residual connections** and **layer normalisation**. Because attention has no built-in notion of order, **positional information** is added (learned or sinusoidal positions, or rotary position embeddings). Encoder stacks (like BERT) read whole inputs; decoder stacks (like GPT-style models) generate text one token at a time using a causal mask so tokens cannot see the future.

## Modern building blocks

Most current systems stack transformer blocks and start from pretrained models.

![Three ideas: transformer blocks, transfer learning, generative models.](assets/figures/deep-learning/section-7-map.svg) — Figure 7.1 — Transformers, transfer and generation.

## A two-layer transformer encoder in PyTorch, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. A 2-layer encoder with 64-dimensional tokens and 4 heads keeps the shape (8, 20, 64): every token gets a new, context-aware vector. This small stack has 66,944 parameters; large language models have billions.

```python
import torch
torch.manual_seed(0)
layer = torch.nn.TransformerEncoderLayer(d_model=64, nhead=4, dim_feedforward=128, batch_first=True)
encoder = torch.nn.TransformerEncoder(layer, num_layers=2)
x = torch.randn(8, 20, 64)                       # batch of 8 sequences, 20 tokens, 64-d each
print("input :", tuple(x.shape))
print("output:", tuple(encoder(x).shape))
print("parameters:", sum(p.numel() for p in encoder.parameters()))
```

Output:

```
input : (8, 20, 64)
output: (8, 20, 64)
parameters: 66944
```

## Use maintained implementations

For real projects, start from established libraries and pretrained checkpoints rather than writing attention from scratch.

**Quiz:** Why do transformers need positional information?

- [ ] Attention only works on images
- [x] Self-attention by itself does not know token order
- [ ] To reduce parameters
- [ ] Because embeddings are random

*Answer:* Self-attention by itself does not know token order. Positions are added to the token vectors.
