Constitutional classifiers — defending against universal jailbreaks

Anthropic Research · Anthropic · 2025-02-03

Read on anthropic.com

A practical safety paper: Anthropic builds classifiers that filter inputs and outputs against a written constitution of what counts as harmful. Unlike RLHF, which bakes preferences into the model's weights, this is an external layer — you can update it without retraining, and you can audit what it's checking against.

The clever bit is the "universal jailbreak" framing. A specific jailbreak ("how do I make a Molotov cocktail") is easy to filter. A universal one — a prompt that breaks the model across many topics simultaneously — has historically been very hard. The classifiers are trained against a continuously-updated set of these universal attacks, and the paper claims meaningful improvements at low cost to legitimate-use latency and pass rates.

- Safety becomes layered: model alignment + classifier filter + monitoring, rather than betting everything on the model - Updating the constitution is editorial work, not ML work — closer to content moderation than reinforcement learning - Independent red-team evaluation is the way you trust this; the post documents that process - The trade-off is real: false-positive filters frustrate legitimate users, and the paper is honest about that

If you're building on top of an LLM in a regulated context (healthcare, education, legal), the question isn't "is the underlying model safe" — it's "what's the system I deploy, and what can I audit." Constitutional classifiers are the cleanest production example of an audit-able external safety layer. Worth reading if you're going to ship anything with real consequences.

> The constitution is a living document. The classifier is a compiled artifact of it.