# Preference Tuning: RLHF and DPO — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/h-pref

> Train on comparisons instead of single answers.

## Teaching what people prefer

SFT shows one good answer; it is often easier to say which of two answers is **better**. **RLHF** trains a reward model on human rankings, then optimises the language model against it. **DPO** skips the reward model and trains directly on (chosen, rejected) pairs, raising the chosen answer's relative probability against a frozen **reference model**. For a pair, the DPO loss is `log 2 ≈ 0.693` when policy equals reference, falls when the policy favours the chosen answer more than the reference does, and rises when it favours the rejected one. Over-optimising can cause verbosity or flattery, so evaluate on held-out comparisons and real tasks.

## The DPO loss on one pair, run

I ran this with plain Python 3 (standard library only), using example numbers. When policy equals reference the loss is 0.6931 (log 2); a policy raising the chosen answer and lowering the rejected one reaches 0.5130; the opposite rises to 0.9130. The log-probabilities are example numbers for one pair.

```python
import math

def dpo_loss(policy_chosen, policy_rejected, ref_chosen, ref_rejected, beta=0.1):
    """Direct Preference Optimisation loss for one preference pair, from sequence log-probabilities."""
    margin = beta * ((policy_chosen - ref_chosen) - (policy_rejected - ref_rejected))
    return math.log(1 + math.exp(-margin))                      # = -log(sigmoid(margin))

ref = (-12.0, -11.0)                                            # reference model log-probs of (chosen, rejected)
for name, pol in (("policy = reference", (-12.0, -11.0)),
                  ("prefers chosen more", (-10.0, -13.0)),
                  ("prefers rejected more", (-14.0, -9.0))):
    print(f"{name:22} loss {dpo_loss(pol[0], pol[1], ref[0], ref[1]):.4f}")

```

Output:

```
policy = reference     loss 0.6931
prefers chosen more    loss 0.5130
prefers rejected more  loss 0.9130
```

## Label consistently

Give raters written criteria and measure their agreement; noisy preferences teach noisy behaviour.

**Quiz:** What data does preference tuning use?

- [ ] Only raw web pages
- [x] Pairs of answers labelled chosen versus rejected
- [ ] Only images
- [ ] GPU temperature logs

*Answer:* Pairs of answers labelled chosen versus rejected. Comparisons teach which behaviours are favoured.
