AI Foundationspredict · compress · act
act VIII

The Attention Revolution

A model that predicts text is not a model that helps. RLHF is how raw prediction becomes behaviour we actually want.

32

Teaching Taste

before this →More Is Different

A pretrained model completes text. It has no idea that a question deserves an answer. Fine-tuning teaches the task. Instruction tuning teaches the format. RLHF teaches the taste.

promptgenerate 2+ answershumans rankA > Btrain rewardmodelreinforce the policy toward higher-scored answersthe reward model stands in for the humans, so it can score a million samples
RLHF: humans rank outputs, a reward model learns the ranking, the policy learns from it

Why not just fine-tune on good answers?

Because “good” is rarely a single right answer. The question is comparative: is this reply more helpful, honest, and safe than that one? Humans are much better at ranking two candidates than at writing an ideal answer or assigning a score. So we collect preference data — pairs where one is preferred — and train a reward model to predict the preference.

Then we optimise the language model against that reward model with reinforcement learning (usually an algorithm called PPO), with a leash keeping it close to the original model so it does not drift into gibberish that games the reward.

The modern shortcut

RLHF is complex and unstable. DPO (Direct Preference Optimisation) skips the separate reward model and the RL loop. It fits the language model directly to the preference pairs. It is simpler, cheaper, and now widely used. The idea is unchanged — learn from comparisons — but the machinery is one supervised step.

The thing fine-tuning cannot do

Fine-tuning changes the model, and it does not know how to change only part of it. Updating the weights on new data can quietly degrade the old behaviour: the network that improved at your task gets worse at everything else, the failure known as catastrophic forgetting. The usual fix is not to learn continuously at all. Keep the base model frozen and learn a small adapter, or retrain the whole model from scratch with the new data mixed in. Both are expensive and both are clumsy, which is why continual learning — taking in a stream of new things without erasing the old — is one of the guide’s open problems rather than a shipped feature.

The alignment tax

the cost of being good

Making a model safe and helpful usually costs a little raw capability. That gap is the alignment tax. Much of the field’s effort goes into shrinking it.

the riskOptimising a proxy for human approval invites reward hacking: the model finds answers that score well without being good. This is the central problem of the next act.
SFTimitate good examples — learn the format
PREFERENCEhumans rank outputs — capture taste
RLHF / DPOoptimise the model toward the ranking
introduces →fine-tuninginstruction tuningRLHFreward modelDPOcontinual learning
← previousMore Is Differentnext →Noise into Images