← all weeks
MLn CLUB · WEEK 12

Towards Monosemanticity

Reading

Paper audio

About this week

Can we identify meaningful, human-interpretable features inside a language model when individual neurons are polysemantic and represent many unrelated concepts at once? Can sparse autoencoders recover the hidden features represented in superposition—and give us a better unit for mechanistic interpretability than neurons themselves?

Anthropic’s Towards Monosemanticity investigates whether sparse autoencoders can decompose a language model’s activations into interpretable features hidden by superposition. Rather than treating individual neurons as the fundamental units of computation, the authors train overcomplete sparse autoencoders to reconstruct a transformer’s MLP activations as sparse combinations of learned feature directions—revealing concepts such as Arabic script, DNA sequences, Base64, and Hebrew that may be distributed across many neurons.

The resulting features are substantially more interpretable than neurons and can also behave as causal units: artificially activating features can steer the model toward the corresponding behavior, such as generating Base64 or Arabic text. The paper also uncovers deeper structure, including feature splitting as dictionary size increases, similar features appearing across independently trained models, and groups of features interacting in ways resembling finite-state automata. Together, these results provide an early proof of concept for sparse autoencoders as a tool for mechanistic interpretability, suggesting that models may be better understood in terms of learned features rather than their raw neuron basis.

Join us at CASI for discussion at 8 pm, and (optional) quiet reading from 7 pm.