Should Developers Care About Interpretability?

Thariq Shihipar · Anthropic · 2024-11-04

Read on thariq.io

Most interpretability writing is for researchers — features, circuits, sparse autoencoders, what's inside a model. Shihipar takes a different angle: what does any of this matter for someone shipping a product on top of Claude or GPT? His answer is that interpretability isn't just a safety tool; it's a new control surface developers should be paying attention to before everyone else does.

The lever is "steering" — turning specific learned features up or down at inference time. Today you customize model behavior through RLHF (expensive, slow, opaque) or system prompts (cheap, fast, blunt). Steering sits in between: cheap, fast, and surgical. The author's argument is that this becomes a real product axis once the reliability problem gets solved.

- Style transfer that resists description — when "make it sound like our brand" is easier shown than written, you steer toward a feature instead of writing rules - Reducing RLHF dependence — no more 10× cost runs to nudge tone or refusal behavior - Cross-conversation preference memory — steering as a serializable user-preference object, not 4,000 tokens of system prompt - Cheap purpose-built classifiers — features as off-the-shelf detectors

> Developers should have a level of control over their models that has not been possible thus far.

The pitch is honest about its problems. Steering can push models out of their training distribution; features are often mislabeled or misunderstood; turning two knobs at once can produce incoherent text. None of that disqualifies the technique — it just means it's at the "researchers can use it, products mostly can't yet" stage.