पाठ 20 / 25

Language Models: Predicting the Next Word

From counting to neural networks.

Next-token prediction

A language model assigns probabilities to what comes next in text. The simplest is a bigram model that counts which word follows which. Sampling from it generates text that is locally plausible but quickly loses track of meaning, because it sees only one previous word. Modern large language models (LLMs) apply the same objective, predicting the next token, with transformer networks trained on vast text collections, capturing long-range context, grammar, facts and reasoning patterns, and are then tuned to follow instructions.

Predict, generate, act

Today's headline systems learn language from huge datasets and increasingly take actions with tools.

Three ideas: language models, deep learning at scale, AI agents.
Figure 7.1 — Language models, scale and agents.

A bigram model from a tiny corpus, run

I ran this with plain Python 3 (standard library only), with fixed random seeds where randomness is used. From four short sentences, the model learns that after the come cat or dog (0.33 each), mat (0.22) or rug (0.11), and that sat is always followed by on. Sampling generates the cat saw the dog . the cat on, plausible pairs without overall meaning.

import random
from collections import Counter, defaultdict
text = ("the cat sat on the mat . the dog sat on the rug . the cat saw the dog . "
        "the dog saw the cat on the mat .").split()
counts = defaultdict(Counter)
for a, b in zip(text, text[1:]):
    counts[a][b] += 1
def next_probs(word):
    total = sum(counts[word].values())
    return {w: round(c / total, 2) for w, c in counts[word].most_common()}
print("after 'the':", next_probs("the"))
print("after 'sat':", next_probs("sat"))
rng = random.Random(3); word, out = "the", ["the"]
for _ in range(8):
    options = counts[word]
    word = rng.choices(list(options), weights=options.values())[0]; out.append(word)
print("generated:", " ".join(out))

Output:

after 'the': {'cat': 0.33, 'dog': 0.33, 'mat': 0.22, 'rug': 0.11}
after 'sat': {'on': 1.0}
generated: the cat saw the dog . the cat on

Remember what LLMs optimise

LLMs produce likely-sounding text; fluency is not proof of truth, so verify important facts.

त्वरित जाँच: What do language models fundamentally learn to do?

  • Store a database of all facts
  • Predict the next token given previous text
  • Search game trees
  • Colour maps
Answer

Predict the next token given previous text — Next-token prediction at scale.