Why we still don't know how to write ASL-4

AIAcademy · AIAcademy · 2026-05-16

Read Anthropic's Responsible Scaling Policy

Anthropic's Responsible Scaling Policy v3.0 (February 2026) and the v3.2 revision (April 2026) did something the earlier RSPs avoided: they walked away from the original promise of pre-defined, future-capability tripwires. The clean ladder of ASL thresholds — ASL-3 here, ASL-4 there, each with a published model-eval that would trigger it — has quietly been replaced with something looser. ASL-3 protections (weight-exfiltration hardening, uplift-resistant deployment) now ride along with every Opus release as the operating floor. ASL-4 is the level that has not been written down, and the silence is the point.

The disagreement is not political. It is empirical. To write a useful ASL-4 threshold you need a measure for "the capability beyond which our current defensive posture stops working" — and nobody has one. The evals that worked for ASL-3 (CBRN-uplift, autonomy at limited horizons, weights-security adversarial probes) saturate or stop discriminating as you cross into the regime ASL-4 is supposed to cover. Anthropic's own automated alignment researchers work hints at the underlying bet: the only way to measure a research-grade AI is with another research-grade AI, and that machinery is still being built.

The other frontier labs have not solved this either. Google DeepMind's Frontier Safety Framework update ran into the same wall and resolved it the same way — denser process commitments around critical capability levels rather than fixed numeric tripwires. OpenAI's Preparedness Framework has drifted toward category-based rather than score-based thresholds. The convergence is not coordination; it is everyone discovering the same gap independently.