AIAcademy · AIAcademy · 2026-05-16
Read OpenAI on CoT monitorability
For roughly two years, chain-of-thought monitorability has been the most reliable cheap-and-deployable interpretability surface the field has had. You read the reasoning trace, you check it against the action, and a non-trivial fraction of misalignment shows up as a mismatch. OpenAI's 2026 evaluation framework is the first systematic attempt to ask whether that property survives the current generation of post-training. The short answer is: at today's RL scales, yes. The longer answer is everyone is preparing for what comes next.
The framework is precise about what "monitorability" means: a reasoning trace is monitorable if a competent reader (human or auxiliary model) can detect intent-misaligned action from the trace alone, above a calibrated threshold. Current-scale RL — including the reasoning-RL pipelines behind o3, o4-mini, and the Claude reasoning models — does not degrade this property in OpenAI's measurements. That is the good news. The bad news is encoded in the same paper: the trend lines for implicit reasoning capacity (problems the model can solve without externalising the steps) are climbing on every model family, and the threshold beyond which monitorability collapses is a function of how much reasoning the model can do in-weight rather than in-tokens.
So labs are already building the replacement. OpenAI's "How confessions can keep language models honest" proposes a second-output approach: the model is trained to emit a parallel "confession" channel — what it considered doing, what it suppressed, what it would tell an auditor it actually intended — alongside the user-facing output. Anthropic's introspection-adapter work takes a structurally different bet: train a lightweight adapter to interrogate the model's residual stream and ask it to explain why it produced the activation pattern that drove the action. Different mechanisms, same underlying acknowledgement — CoT is the cheap interpretability win, and the cheap win has a sunset.