AI Foundationspredict · compress · act
act VIII

The Alignment Layer

A model can be right a million times and we still may not know why. Interpretability is the attempt to read the machine.

30

Opening the Black Box

before this →Noise into Images

A neural network gives us numbers. We want reasons. Mechanistic interpretability tries to reverse-engineer the computation inside a trained model — to find the actual circuits, not to tell a plausible story about them.

formal toneFrenchcodesarcasmlegalesethousands of featuresshare a few dimensionsthat is superposition:efficient, and hard to read
features are directions in activation space; superposition packs many into few dimensions

Features and directions

A feature is a property the model represents — “this is French,” “this is legalese,” “this is a cat.” Researchers have found that features correspond to directions in the activation space. Add that direction and the model’s behaviour shifts toward that concept. That is real, causal evidence, not narrative.

The catch is superposition. A model has far more concepts than dimensions, so it stores many features in overlapping directions. Reading one feature at a time accidentally reads several at once. This is why the insides look like noise at first.

Sparse autoencoders

The current best tool for the job is the sparse autoencoder (SAE): a wide layer trained to reconstruct activations using only a few active features at a time. Forcing sparsity disentangles the crowded directions into cleaner, more monosemantic ones. The result is a dictionary of interpretable features that can be steered, ablated, and audited.

This is not academic. If we can find the features for deception, sycophancy, or “I am being tested,” we can check for them directly. Interpretability aims to be the inspection regime that safety otherwise lacks.

Two kinds of explanation

plausible story vs. verified circuit

A model can produce a convincing rationale that has nothing to do with its actual computation. Interpretability’s discipline is to demand causal evidence: intervene, and watch the behaviour change.

the stakesAlignment by measurement is weaker than alignment by understanding. You cannot reliably supervise what you cannot see.
PROBEfind a direction that encodes a concept
ABLATEremove it — does behaviour change?
STEERadd it — can we control the model?
introduces →mechanistic interpretabilityfeaturesuperpositionsparse autoencoder
← previousNoise into Imagesnext →Getting the Goal Right