AI Foundationspredict · compress · act
act VII

The Attention Revolution

A model that predicts text is not a model that helps. RLHF is how raw prediction becomes behaviour we actually want.

28

Teaching Taste

before this →More Is Different

A pretrained model completes text. It has no idea that a question deserves an answer. Fine-tuning teaches the task. Instruction tuning teaches the format. RLHF teaches the taste.

promptgenerate 2+ answershumans rankA > Btrain rewardmodelreinforce the policy toward higher-scored answersthe reward model stands in for the humans, so it can score a million samples
RLHF: humans rank outputs, a reward model learns the ranking, the policy learns from it

Why not just fine-tune on good answers?

Because “good” is rarely a single right answer. The question is comparative: is this reply more helpful, honest, and safe than that one? Humans are much better at ranking two candidates than at writing an ideal answer or assigning a score. So we collect preference data — pairs where one is preferred — and train a reward model to predict the preference.

Then we optimise the language model against that reward model with reinforcement learning (usually an algorithm called PPO), with a leash keeping it close to the original model so it does not drift into gibberish that games the reward.

The modern shortcut

RLHF is complex and unstable. DPO (Direct Preference Optimisation) skips the separate reward model and the RL loop, and fits the language model directly to the preference pairs. It is simpler, cheaper, and now widely used. The idea is unchanged — learn from comparisons — but the machinery is one supervised step.

The alignment tax

the cost of being good

Making a model safe and helpful usually costs a little raw capability. That gap is the alignment tax. Much of the field’s effort goes into shrinking it.

the riskOptimising a proxy for human approval invites reward hacking: the model finds answers that score well without being good. This is the central problem of the next act.
SFTimitate good examples — learn the format
PREFERENCEhumans rank outputs — capture taste
RLHF / DPOoptimise the model toward the ranking
introduces →fine-tuninginstruction tuningRLHFreward modelDPO
← previousMore Is Differentnext →Noise into Images