Has any model been blocked from release yet?

AIAcademy · AIAcademy · 2026-05-16

Mythos Preview red-team writeup

For two years the Responsible Scaling Policy lived in a strange middle ground. ASL-3 protections had been on every Claude Opus since May 2025 — uplift mitigations for CBRN, classifier defences, abuse monitoring. But no model had actually been withheld from general release on capability grounds. The policy was real; the bite had not been demonstrated.

Mythos Preview is the demonstration.

Anthropic's public red-team writeup describes a model that beats every prior frontier system on long-horizon autonomy (METR puts its 50%-horizon at "at least 16 hours" with confidence interval up to 55), beats human baseline on OSWorld computer use, and exhibits enough new-capability surface that the company chose not to ship a public API. Instead, Mythos is deployed only to ~50 critical-infrastructure organizations through Project Glasswing, with monitored access, in-environment safety classifiers, and a stated obligation to publish incident reports.

The framing matters. This is not a model that failed a capability test. It is a model that passed too many of them, in a regime the existing ASL-3 mitigations were not designed for. By withholding general deployment, Anthropic has effectively published a draft ASL-4 standard through behaviour rather than policy text — "we built it, we won't deploy it generally" is a far more legible signal than a paragraph in a governance document.

Three things to watch in 2026–27. First, whether the unwritten standard becomes a written one before the next frontier model crosses the same line. Second, whether OpenAI, Google, and the Chinese labs respond with comparable disclosure — the precedent is set, but it is not yet a norm. Third, whether Glasswing-style restricted deployments become the default for capability-leading models, or whether competitive pressure forces an early walk-back.

The answer to the question in the title is now yes. The interesting questions are what comes next.