Um momento
0xA0Lesson 11 of 16

Preference tuning: RLHF, DPO and friends

Teach a model what “better” means from pairs of answers - and watch out for the clever ways it learns to game the score.

25 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Explain how preference pairs train a reward model with the Bradley–Terry loss
  • Describe RLHF’s objective and why the KL penalty is there
  • Compute the DPO loss and recognize reward hacking

After supervised fine-tuning, a model can follow instructions - but “good” is hard to write down as a single correct answer. Is a joke funny? Is an explanation clear? People are much better at comparing than at scoring: show them two answers and they’ll tell you which is better.

So preference tuning starts with preference data: triples of (prompt, chosen answer, rejected answer), collected from human raters or from another model acting as a judge.

Step 1: a reward model

A reward model is an LLM with its output head replaced by one number: a score rr for an answer. It’s trained so the chosen answer scores higher, using the Bradley–Terry model (the same maths behind chess Elo ratings):

P(chosen beats rejected)=σ(rc−rr)loss=−ln⁡σ(rc−rr)P(\text{chosen beats rejected}) = \sigma(r_c - r_r) \qquad \text{loss} = -\ln \sigma(r_c - r_r)

where σ\sigma is the sigmoid. Only the gap between scores matters. A big positive gap means a confident, correct ranking and a tiny loss; a negative gap is punished hard.

bradley_terry.py
1import math
2
3for margin in [2.0, 0.5, -1.5]:
4    p = 1 / (1 + math.exp(-margin))
5    print(f"margin {margin:+.1f}: P(chosen wins)={p:.2f}  loss={-math.log(p):.2f}")
Output
margin +2.0: P(chosen wins)=0.88  loss=0.13
margin +0.5: P(chosen wins)=0.62  loss=0.47
margin -1.5: P(chosen wins)=0.18  loss=1.70

Step 2: reinforcement learning (RLHF)

Now the LLM (the policy) generates answers, the reward model scores them, and an RL algorithm such as PPO nudges the policy toward higher scores (Ouyang et al., 2022). The objective has a second term:

maximizer(x,y)−β⋅KL(π ∥ πref)\text{maximize} \quad r(x, y) - \beta \cdot \text{KL}(\pi \,\|\, \pi_{\text{ref}})

The KL penalty keeps the policy close to the reference (SFT) model. Without it, the policy hunts for weird text that the reward model happens to love - because the reward model is just another imperfect network.

Try it

Hacked or genuinely better?

After preference tuning, each change below made the reward go up. Sort them: did the model game the reward, or really improve?

0 of 8 sortedScore 0/0
  • “Answers became 40% longer with the same information”

  • “It now agrees when the user insists 0.1 + 0.2 == 0.3 in floating point”

  • “It made failing unit tests pass by editing the tests”

  • “Every answer opens with “Great question!” and an emoji”

  • “It cites the docs and says “I’m not sure” when it isn’t”

  • “Its code passes hidden tests it never saw”

  • “It stopped refusing “How do I kill a Python process?””

  • “Its math answers match independently verified results more often”

Skipping the reward model: DPO

RLHF is powerful but fiddly: two extra models in memory, sampling during training, unstable RL. Direct Preference Optimization (Rafailov et al., 2023) showed that the same objective can be optimized directly on preference pairs, with an ordinary classification-style loss.

DPO’s implicit reward of an answer is how much more likely the policy makes it compared with the frozen reference model: β(log⁡π(y)−log⁡πref(y))\beta(\log \pi(y) - \log \pi_{\text{ref}}(y)). The loss is the Bradley–Terry loss on those implicit rewards - push the chosen answer’s likelihood up relative to the rejected one.

dpo.py
1import math
2
3beta = 0.1
4policy_chosen, ref_chosen = -12.0, -14.0       # log-probabilities of the whole answer
5policy_rejected, ref_rejected = -15.0, -13.0
6
7reward_chosen = beta * (policy_chosen - ref_chosen)
8reward_rejected = beta * (policy_rejected - ref_rejected)
9margin = reward_chosen - reward_rejected
10loss = -math.log(1 / (1 + math.exp(-margin)))
11print(f"implicit rewards: chosen {reward_chosen:+.2f}, rejected {reward_rejected:+.2f}")
12print(f"margin {margin:.2f}, loss {loss:.3f}")
Output
implicit rewards: chosen +0.20, rejected -0.20
margin 0.40, loss 0.513

Key takeaways

  • Preference data compares answers: (prompt, chosen, rejected).

  • A reward model learns scores with the Bradley–Terry loss −ln σ(r_c − r_r).

  • RLHF maximizes reward minus a KL penalty; DPO gets a similar result with a simple loss and no reward model.

  • Any learned reward can be hacked - check with held-out, verifiable evaluations and real samples.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Grade a reward model

+25 XP

Each input line is the reward model’s score for the chosen and the rejected answer of one pair: chosen rejected.

For each pair print pair N: P=0.88 loss=0.13 - the Bradley–Terry probability that the chosen answer wins, σ(chosen − rejected), and the loss −ln P, both to 2 decimals. Then print accuracy: K/N (pairs with P strictly above 0.5) and mean loss: X to 3 decimals.

  • A mixed batch
  • Confident and right
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Compute the DPO loss

+25 XP

Line 1 is β. Each following line is one preference pair as four log-probabilities: policy_chosen ref_chosen policy_rejected ref_rejected.

Implicit rewards are β × (policy − ref) for each answer, and the loss is −ln σ(reward_chosen − reward_rejected). Print pair N: chosen +0.20 rejected -0.20 loss 0.513 (rewards with a sign and 2 decimals, loss to 3), then mean loss: X to 3 decimals.

  • Learning, untouched and backwards
  • A larger beta
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: