Engineering Monosemanticity in Toy Models

Adam Jermyn · Anthropic · 2022-11-21

Read on adamjermyn.com

Before Towards Monosemanticity became Anthropic's flagship interpretability result, Jermyn was poking at the core question on his personal blog: can you actually engineer a network where each neuron means one thing? He builds tiny models, perturbs them, and reports what makes monosemanticity emerge versus collapse.

This is interpretability research caught in the lab-notebook stage — messy, exploratory, and honest about what the toy results do and don't tell you. Read it next to the polished Anthropic papers to see how a research direction is actually shaped: not by clean ideas, but by stubborn iteration on small examples until something gives.