Teaching Taste
A pretrained model completes text. It has no idea that a question deserves an answer. Fine-tuning teaches the task. Instruction tuning teaches the format. RLHF teaches the taste.
Why not just fine-tune on good answers?
Because “good” is rarely a single right answer. The question is comparative: is this reply more helpful, honest, and safe than that one? Humans are much better at ranking two candidates than at writing an ideal answer or assigning a score. So we collect preference data — pairs where one is preferred — and train a reward model to predict the preference.
Then we optimise the language model against that reward model with reinforcement learning (usually an algorithm called PPO), with a leash keeping it close to the original model so it does not drift into gibberish that games the reward.
The modern shortcut
RLHF is complex and unstable. DPO (Direct Preference Optimisation) skips the separate reward model and the RL loop. It fits the language model directly to the preference pairs. It is simpler, cheaper, and now widely used. The idea is unchanged — learn from comparisons — but the machinery is one supervised step.
The thing fine-tuning cannot do
Fine-tuning changes the model, and it does not know how to change only part of it. Updating the weights on new data can quietly degrade the old behaviour: the network that improved at your task gets worse at everything else, the failure known as catastrophic forgetting. The usual fix is not to learn continuously at all. Keep the base model frozen and learn a small adapter, or retrain the whole model from scratch with the new data mixed in. Both are expensive and both are clumsy, which is why continual learning — taking in a stream of new things without erasing the old — is one of the guide’s open problems rather than a shipped feature.
The alignment tax
Making a model safe and helpful usually costs a little raw capability. That gap is the alignment tax. Much of the field’s effort goes into shrinking it.