Opening the Black Box
A neural network gives us numbers. We want reasons. Mechanistic interpretability tries to reverse-engineer the computation inside a trained model — to find the actual circuits, not to tell a plausible story about them.
Features and directions
A feature is a property the model represents — “this is French,” “this is legalese,” “this is a cat.” Researchers have found that features correspond to directions in the activation space. Add that direction and the model’s behaviour shifts toward that concept. That is real, causal evidence, not narrative.
The catch is superposition. A model has far more concepts than dimensions, so it stores many features in overlapping directions. Reading one feature at a time accidentally reads several at once. This is why the insides look like noise at first.
Sparse autoencoders
The current best tool for the job is the sparse autoencoder (SAE): a wide layer trained to reconstruct activations using only a few active features at a time. Forcing sparsity disentangles the crowded directions into cleaner, more monosemantic ones. The result is a dictionary of interpretable features that can be steered, ablated, and audited.
This is not academic. If we can find the features for deception, sycophancy, or “I am being tested,” we can check for them directly. Interpretability aims to be the inspection regime that safety otherwise lacks.
Two kinds of explanation
A model can produce a convincing rationale that has nothing to do with its actual computation. Interpretability’s discipline is to demand causal evidence: intervene, and watch the behaviour change.