पाठ 19 / 25

Self-Attention

Every token looks at every other token.

Queries, keys and values

In self-attention, each token's vector is projected into a query, a key and a value. The query of one token is compared with the keys of all tokens (dot products, scaled by the square root of the dimension); a softmax turns the scores into attention weights that sum to 1; the output is the weighted sum of the values. Each token's new representation therefore mixes in information from the tokens most relevant to it, so the word "bank" can take meaning from "river". Multi-head attention runs several of these in parallel.

Scaled dot-product attention in numpy, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. Three tokens with random (untrained) projections: each row of attention weights sums to 1, and river attends almost entirely to itself (0.939). Trained models learn projections that make these weights meaningful.

import numpy as np
np.set_printoptions(precision=3, suppress=True)
rng = np.random.default_rng(0)
tokens = ["the", "bank", "river"]
X = rng.normal(size=(3, 4))                       # one 4-d vector per token
Wq, Wk, Wv = (rng.normal(size=(4, 4)) for _ in range(3))
Q, K, V = X @ Wq, X @ Wk, X @ Wv
scores = Q @ K.T / np.sqrt(K.shape[1])            # scaled dot-product
weights = np.exp(scores) / np.exp(scores).sum(axis=1, keepdims=True)   # softmax per row
out = weights @ V
print("attention weights (row = token attending):")
for t, row in zip(tokens, weights):
    print(f"  {t:<6}", row, "sum", round(row.sum(), 3))
print("output shape:", out.shape)

Output:

attention weights (row = token attending):
  the    [0.384 0.3   0.315] sum 1.0
  bank   [0.321 0.206 0.473] sum 1.0
  river  [0.053 0.008 0.939] sum 1.0
output shape: (3, 4)

Mind the quadratic cost

Attention compares every token with every other, so cost grows with the square of sequence length; long inputs need efficient attention variants or chunking.

त्वरित जाँच: What do attention weights in each row sum to?

  • The model size
  • 0
  • The sequence length
  • 1, because of the softmax
Answer

1, because of the softmax — They form a weighted average over values.