Adam Jermyn · Anthropic · 2022-11-21
Before Towards Monosemanticity became Anthropic's flagship interpretability result, Jermyn was poking at the core question on his personal blog: can you actually engineer a network where each neuron means one thing? He builds tiny models, perturbs them, and reports what makes monosemanticity emerge versus collapse.
This is interpretability research caught in the lab-notebook stage — messy, exploratory, and honest about what the toy results do and don't tell you. Read it next to the polished Anthropic papers to see how a research direction is actually shaped: not by clean ideas, but by stubborn iteration on small examples until something gives.