Teaching Taste
A pretrained model completes text. It has no idea that a question deserves an answer. Fine-tuning teaches the task. Instruction tuning teaches the format. RLHF teaches the taste.
Why not just fine-tune on good answers?
Because “good” is rarely a single right answer. The question is comparative: is this reply more helpful, honest, and safe than that one? Humans are much better at ranking two candidates than at writing an ideal answer or assigning a score. So we collect preference data — pairs where one is preferred — and train a reward model to predict the preference.
Then we optimise the language model against that reward model with reinforcement learning (usually an algorithm called PPO), with a leash keeping it close to the original model so it does not drift into gibberish that games the reward.
The modern shortcut
RLHF is complex and unstable. DPO (Direct Preference Optimisation) skips the separate reward model and the RL loop, and fits the language model directly to the preference pairs. It is simpler, cheaper, and now widely used. The idea is unchanged — learn from comparisons — but the machinery is one supervised step.
The alignment tax
Making a model safe and helpful usually costs a little raw capability. That gap is the alignment tax. Much of the field’s effort goes into shrinking it.