Lesson 17 / 25
Embeddings: Tokens as Vectors
Turn ids into learnable dense vectors.
A lookup table that learns
Networks need numbers, so text is split into tokens and each token id is mapped to a dense vector by an embedding layer, a learnable lookup table of shape (vocabulary size, dimension). During training, tokens used in similar contexts end up with similar vectors. Sequences in a batch are padded to the same length with a padding id whose vector stays zero and is masked out. The same idea embeds users, products or categories in recommendation and tabular models.
From tokens to context-aware vectors
Text and other sequences are handled by embedding tokens and letting them exchange information.
An embedding layer on padded sequences, run
I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. A five-word vocabulary with 4-dimensional embeddings maps a (2, 3) batch of ids to a (2, 3, 4) tensor of vectors. The refund vector is a learnable random start; the padding vector stays zero.
import torch
torch.manual_seed(0)
vocab = {"<pad>": 0, "refund": 1, "invoice": 2, "crash": 3, "login": 4}
emb = torch.nn.Embedding(num_embeddings=len(vocab), embedding_dim=4, padding_idx=0)
ids = torch.tensor([[1, 2, 0], [3, 4, 0]]) # two padded sentences of token ids
vectors = emb(ids)
print("token ids shape:", tuple(ids.shape), "-> embeddings shape:", tuple(vectors.shape))
print("vector for 'refund':", emb.weight[1].detach().numpy().round(3))
print("padding vector stays zero:", emb.weight[0].detach().numpy())
Output:
token ids shape: (2, 3) -> embeddings shape: (2, 3, 4) vector for 'refund': [ 0.599 -1.555 -0.341 1.853] padding vector stays zero: [0. 0. 0. 0.]
Mask the padding
Make sure losses and attention ignore padding positions; otherwise the model learns from filler.
Quick check: What is an embedding layer?
- A data loader
- A loss function
- A learnable lookup table from ids to dense vectors
- A type of pooling
Answer
A learnable lookup table from ids to dense vectors — Embeddings turn discrete ids into learnable vectors.