METR · METR · 2026-05-16
Read Time Horizon 1.1 on metr.org
The single most legible chart in AI progress right now belongs to METR. Their Time Horizon framework asks one question: how long a task can a model finish at 50% reliability, measured against the time a competent human takes? The answer doubles. It used to double every seven months. Since 2024 it has been doubling every three.
The Time Horizon 1.1 release (January 2026) expanded the suite from 170 to 228 tasks and added 17 new 8-hour-plus tasks. Mythos Preview lands at roughly 16.8 hours — at the upper edge of what the current evaluation suite can resolve. Claude Opus 4.6 is at 14.5 hours. GPT-5 sits at 3.6. The previous generation's "10-hour wall" framing is dead.
Two things make this metric load-bearing for learners. First, it converts an abstract trendline ("AI is getting better") into a number ordinary humans can compare against their own work. Second, it exposes the unglamorous middle of frontier progress: the 50% number is the headline, but the 80%-reliability horizon is roughly five times shorter. Mythos at 16h@50% is closer to 3h@80%. The reliability gap is where production engineering lives.
Two reads worth pairing. AI Digest calls the doubling "a new Moore's Law for AI agents" — useful, popular, possibly too neat. MIT Tech Review's counterweight argues the chart is a fit, not a forecast, and that domain variance (10× across METR's own follow-up studies) breaks any naive extrapolation.
If the post-2024 doubling holds, week-scale autonomous tasks land in late 2027. If it reverts to the long-term seven-month rate, you get there in 2029. The contested question is which of those numbers you should plan against.