Lesson 17 / 25

Embeddings: Tokens as Vectors

Turn ids into learnable dense vectors.

A lookup table that learns

Networks need numbers, so text is split into tokens and each token id is mapped to a dense vector by an embedding layer, a learnable lookup table of shape (vocabulary size, dimension). During training, tokens used in similar contexts end up with similar vectors. Sequences in a batch are padded to the same length with a padding id whose vector stays zero and is masked out. The same idea embeds users, products or categories in recommendation and tabular models.

From tokens to context-aware vectors

Text and other sequences are handled by embedding tokens and letting them exchange information.

Three ideas: embeddings, recurrent networks, attention.
Figure 6.1 — Embeddings, recurrence and attention.

An embedding layer on padded sequences, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. A five-word vocabulary with 4-dimensional embeddings maps a (2, 3) batch of ids to a (2, 3, 4) tensor of vectors. The refund vector is a learnable random start; the padding vector stays zero.

import torch
torch.manual_seed(0)
vocab = {"<pad>": 0, "refund": 1, "invoice": 2, "crash": 3, "login": 4}
emb = torch.nn.Embedding(num_embeddings=len(vocab), embedding_dim=4, padding_idx=0)
ids = torch.tensor([[1, 2, 0], [3, 4, 0]])           # two padded sentences of token ids
vectors = emb(ids)
print("token ids shape:", tuple(ids.shape), "-> embeddings shape:", tuple(vectors.shape))
print("vector for 'refund':", emb.weight[1].detach().numpy().round(3))
print("padding vector stays zero:", emb.weight[0].detach().numpy())

Output:

token ids shape: (2, 3) -> embeddings shape: (2, 3, 4)
vector for 'refund': [ 0.599 -1.555 -0.341  1.853]
padding vector stays zero: [0. 0. 0. 0.]

Mask the padding

Make sure losses and attention ignore padding positions; otherwise the model learns from filler.

Quick check: What is an embedding layer?

  • A data loader
  • A loss function
  • A learnable lookup table from ids to dense vectors
  • A type of pooling
Answer

A learnable lookup table from ids to dense vectors — Embeddings turn discrete ids into learnable vectors.