Opening the Black Box
A neural network gives us numbers. We want reasons. Mechanistic interpretability tries to reverse-engineer the computation inside a trained model — to find the actual circuits, not to tell a plausible story about them.
Features and directions
A feature is a property the model represents — “this is French,” “this is legalese,” “this is a cat.” Researchers have found that features correspond to directions in the activation space. Add that direction and the model’s behaviour shifts toward that concept. That is real, causal evidence, not narrative.
The catch is superposition. A model has far more concepts than dimensions, so it stores many features in overlapping directions. Reading one feature at a time accidentally reads several at once. This is why the insides look like noise at first.
Sparse autoencoders
The current best tool for the job is the sparse autoencoder (SAE): a wide layer trained to reconstruct activations using only a few active features at a time. Forcing sparsity disentangles the crowded directions into cleaner, more monosemantic ones. The result is a dictionary of interpretable features that can be steered, ablated, and audited.
This is not academic. If we can find the features for deception, sycophancy, or “I am being tested,” we can check for them directly. Interpretability aims to be the inspection regime that safety otherwise lacks.
The limit of features
A feature is a direction that correlates with a concept. That is not the same as a cause. A model may represent “this is a photo of a cow” and “this is grass” as two directions without ever representing that the grass is why the cow is there. A system that knows only correlations fails when the correlation inverts. Causal representation learning is the harder target: recover the variables the world is actually made of, and the causal relations between them, so the model can answer questions about interventions (“what happens if I change this?”) rather than merely predict what co-occurs. It remains open because raw data never hands you the variables; disentangling them is part of the problem, not a precondition for it.
Two kinds of explanation
A model can produce a convincing rationale that has nothing to do with its actual computation. Interpretability’s discipline is to demand causal evidence: intervene, and watch the behaviour change.