Lesson 12 / 27
Preference Tuning: RLHF and DPO
Train on comparisons instead of single answers.
Teaching what people prefer
SFT shows one good answer; it is often easier to say which of two answers is better. RLHF trains a reward model on human rankings, then optimises the language model against it. DPO skips the reward model and trains directly on (chosen, rejected) pairs, raising the chosen answer's relative probability against a frozen reference model. For a pair, the DPO loss is log 2 ≈ 0.693 when policy equals reference, falls when the policy favours the chosen answer more than the reference does, and rises when it favours the rejected one. Over-optimising can cause verbosity or flattery, so evaluate on held-out comparisons and real tasks.
The DPO loss on one pair, run
I ran this with plain Python 3 (standard library only), using example numbers. When policy equals reference the loss is 0.6931 (log 2); a policy raising the chosen answer and lowering the rejected one reaches 0.5130; the opposite rises to 0.9130. The log-probabilities are example numbers for one pair.
import math
def dpo_loss(policy_chosen, policy_rejected, ref_chosen, ref_rejected, beta=0.1):
"""Direct Preference Optimisation loss for one preference pair, from sequence log-probabilities."""
margin = beta * ((policy_chosen - ref_chosen) - (policy_rejected - ref_rejected))
return math.log(1 + math.exp(-margin)) # = -log(sigmoid(margin))
ref = (-12.0, -11.0) # reference model log-probs of (chosen, rejected)
for name, pol in (("policy = reference", (-12.0, -11.0)),
("prefers chosen more", (-10.0, -13.0)),
("prefers rejected more", (-14.0, -9.0))):
print(f"{name:22} loss {dpo_loss(pol[0], pol[1], ref[0], ref[1]):.4f}")
Output:
policy = reference loss 0.6931 prefers chosen more loss 0.5130 prefers rejected more loss 0.9130
Label consistently
Give raters written criteria and measure their agreement; noisy preferences teach noisy behaviour.
Quick check: What data does preference tuning use?
- Only raw web pages
- Pairs of answers labelled chosen versus rejected
- Only images
- GPU temperature logs
Answer
Pairs of answers labelled chosen versus rejected — Comparisons teach which behaviours are favoured.