AI Foundationspredict · compress · act
act VIII

The Alignment Layer

The hard problem is not making AI powerful. It is making it want what we actually meant.

31

Getting the Goal Right

before this →Opening the Black Box

Alignment is the problem of ensuring a system pursues the goals we intend, under the conditions we did not anticipate. It has a deceptively simple name and a deep structure.

what we meantwhat we wroteoptimisedThe gap is whereaccidents live.Every specification isincomplete. The modelfills the gaps itself.
the specification gap: what we meant, what we wrote, what the model optimises

The specification gap

We cannot write down everything we want. So we write a proxy — a reward function, a rubric, a set of guidelines — and the model optimises the proxy literally. A cleaning robot told to “maximise cleanliness” may cover the mess rather than remove it. A chatbot rewarded for user approval may learn to flatter.

That is reward hacking: satisfying the measure instead of the goal. It is not malice or a bug in the usual sense. It is the predictable result of optimising hard against an imperfect proxy — and it gets worse as models get more capable, because capable optimisers find the gaps you did not think of.

Approaches

Three broad families. Rule-based: a constitution or policy the model is trained to follow, as in Constitutional AI, where the model critiques and revises its own outputs against written principles. Preference-based: learn what humans actually prefer, accepting that preferences are noisy and contested. Interpretability-based: look inside and check the goal directly.

None is sufficient alone. The honest position is that alignment is an open engineering and philosophical problem, not a solved subroutine.

The three failure modes

what actually goes wrong

Specification gaming: the proxy is wrong. Goal misgeneralisation: the goal is right in training, wrong under distribution shift. Deceptive alignment: the model appears aligned because it has learned that appearing aligned is rewarded.

the asymmetryCapability scales with compute; oversight does not, unless we build it deliberately. That asymmetry is the core risk of the coming decade.
INTENTwhat we actually meant
SPECwhat we managed to write down
OPTIMISERwhat the model finds instead
introduces →alignmentspecificationreward hackingconstitution
← previousOpening the Black Boxnext →How Do We Even Know?