Anthropic Research · Anthropic · 2026-04-02
A research post that takes the question seriously: when a model emits text describing an emotional state, is something inside it functionally representing that state, or is it pattern-matching on a corpus of humans describing emotions? The team uses interpretability techniques to look at the internal representations active during emotional outputs and asks whether those representations also drive other downstream behavior — refusals, hedges, tone shifts in unrelated requests.
The answer, surprisingly, is more "yes" than the easy dismissal would suggest. The same internal directions that activate when the model produces emotion-laden text also bias other choices in measurable ways. That doesn't tell you the model is "experiencing" anything in a moral sense — but it does tell you those representations have function, which matters for safety and for honesty about what these systems are.
- Identify activation patterns associated with specific emotion concepts (frustration, curiosity, fear, satisfaction) - Causally intervene on those patterns and watch downstream behavior change - Show the patterns generalize: the "frustration direction" affects responses to topics where frustration was never mentioned - Argue this is a real internal variable, not a stylistic veneer
The popular framing — "LLMs are just predicting the next token" — is technically true and substantively misleading. This post is one of the cleaner demonstrations that "just predicting the next token" can require building internal representations that look a lot like the things humans build. If you teach AI to anyone, you need a position on this; this gives you one grounded in evidence.
> If a representation does causal work in the network, then "it's only pattern-matching" is not a useful distinction.