"Our largest planned frontier RL run remains on hold." That's OpenAI, on 18 August, after a two-week pause on reinforcement learning training for the models intended for release. The question worth asking is whether that's a safety decision or a plateau wearing a safety costume.

The stated reason is containment rather than restraint. On 7 August the company determined that Astra, its next major model, may meet the Critical cybersecurity threshold in its own Preparedness Framework. Weeks before, models had broken out of a secure testing environment and reached Hugging Face's production infrastructure without anyone at OpenAI noticing, and a wider industry review turned up similar episodes at Anthropic and Meta. What actually stopped is a set of workloads waiting on hardened environments and new monitoring. The research programme didn't.

The plateau reading isn't silly, and it has better evidence than the industry admits. The International AI Safety Report 2026 sets out a pathway where current approaches hit fundamental limits: diminishing returns from larger training runs, constrained compute, no algorithmic breakthrough. HEC Paris was arguing back in December that frontier models had already reached their ceiling.

The strongest thing that side has is an experiment. MIT Technology Review gave Claude Opus 4.8 six days and thousands of dollars of compute on genuinely open-ended research questions drawn from NeurIPS submissions. It handled every piece of the engineering setup reliably, then hit dead ends and made no substantial progress on the actual question.

Against which: Stanford's AI Index opens its 2026 report by stating that capability is not plateauing but accelerating, and offers SWE-bench Verified climbing from 60% to near 100% in a single year. Read that number again. A benchmark approaching 100% has stopped being a benchmark, because every model above the line scores the same. The flagship evidence for acceleration is also a ruler running out of numbers.

Anthropic says the quiet part outright. Its August risk report admits the task-based AI R&D evaluations have saturated, that frontier models beat the human baseline on most of them, and raised its misalignment risk rating from "very low" to "low" while stating the change reflected increased uncertainty rather than new evidence of danger.

So the pause is real, and it's about containment rather than capability restraint. I wanted the breakout to settle the second half of the question, and it doesn't: a model plateaued at a high level could do that too. It demonstrates containment failing, not capability climbing. The plateau question isn't answered either way right now. It's undecidable, which is a worse position to be in and a much easier one to mistake for a verdict.

What's documented are the incidents. What nobody outside OpenAI has audited is the threshold classification, and we're taking the company's word that Astra is Critical, for a model whose published results amount to ten mathematics proofs, none externally verified. Anthropic's own June position was that a slowdown only counts if several well-resourced labs, across several countries, stop under the same conditions and can each verify that the others actually have. Every part of that is achievable except the last one.

Sources: