SkillByAIOpen interactive version →

Lesson 12 / 27

Preference Tuning: RLHF and DPO

Train on comparisons instead of single answers.

Teaching what people prefer

SFT shows one good answer; it is often easier to say which of two answers is better. RLHF trains a reward model on human rankings, then optimises the language model against it. DPO skips the reward model and trains directly on (chosen, rejected) pairs, raising the chosen answer's relative probability against a frozen reference model. For a pair, the DPO loss is log 2 ≈ 0.693 when policy equals reference, falls when the policy favours the chosen answer more than the reference does, and rises when it favours the rejected one. Over-optimising can cause verbosity or flattery, so evaluate on held-out comparisons and real tasks.

The DPO loss on one pair, run

I ran this with plain Python 3 (standard library only), using example numbers. When policy equals reference the loss is 0.6931 (log 2); a policy raising the chosen answer and lowering the rejected one reaches 0.5130; the opposite rises to 0.9130. The log-probabilities are example numbers for one pair.

import math

def dpo_loss(policy_chosen, policy_rejected, ref_chosen, ref_rejected, beta=0.1):
    """Direct Preference Optimisation loss for one preference pair, from sequence log-probabilities."""
    margin = beta * ((policy_chosen - ref_chosen) - (policy_rejected - ref_rejected))
    return math.log(1 + math.exp(-margin))                      # = -log(sigmoid(margin))

ref = (-12.0, -11.0)                                            # reference model log-probs of (chosen, rejected)
for name, pol in (("policy = reference", (-12.0, -11.0)),
                  ("prefers chosen more", (-10.0, -13.0)),
                  ("prefers rejected more", (-14.0, -9.0))):
    print(f"{name:22} loss {dpo_loss(pol[0], pol[1], ref[0], ref[1]):.4f}")

Output:

policy = reference     loss 0.6931
prefers chosen more    loss 0.5130
prefers rejected more  loss 0.9130

Label consistently

Give raters written criteria and measure their agreement; noisy preferences teach noisy behaviour.

Quick check: What data does preference tuning use?

  • Only raw web pages
  • Pairs of answers labelled chosen versus rejected
  • Only images
  • GPU temperature logs
Answer

Pairs of answers labelled chosen versus rejected — Comparisons teach which behaviours are favoured.