OpenAI paused its largest frontier RL run after Astra's cyber evaluations couldn't rule out 'Critical' hacking capability. Here's what's confirmed — and what isn't.
OpenAI has paused a deployment-focused reinforcement-learning phase for two weeks and is holding back its largest planned frontier RL run — after evaluations of its upcoming Astra system couldn't rule out "Critical" cybersecurity capabilities under the company's own risk framework. OpenAI says it's using the time to strengthen security, monitoring, and alignment safeguards.
This is a real, consequential pause — but a narrow one. OpenAI hasn't stopped all training, cancelled Astra, or set a restart date. What it has done is let an internal capability evaluation directly delay a training schedule, rather than just shape release documentation after the fact. Axios independently confirmed the hold and reported OpenAI is also reconsidering parts of its preparedness framework as a result.
The trigger is specific: Astra's cyber-capability evaluations reached a point where OpenAI couldn't confidently rule out Critical-tier hacking ability. In response, the company moved related work into more secure research environments, restricted tool and network access, and added monitoring — meaning the risk assessment changed how researchers are allowed to work with the system, not just what gets published about it later.
What's still missing is the substance behind the headline. OpenAI hasn't disclosed the actual benchmark results, which specific cyber tasks triggered the concern, or the threshold Astra needs to clear before the held run resumes — and there's no public timeline for that decision. That means outsiders can't yet independently judge whether OpenAI's "Critical" classification is well-calibrated, or whether the new safeguards are actually sufficient.
Bottom line: This is a genuinely notable step in AI safety governance — a capability evaluation now has the power to gate a frontier training run, not just annotate a model card after launch. OpenAI deserves credit for disclosing it publicly. But the real test comes next: what evidence OpenAI accepts before resuming, and whether that bar can be scrutinized by anyone outside the company.
OpenAI says it's using the time to strengthen security, monitoring, and alignment safeguards.
OpenAI hasn't stopped all training, cancelled Astra, or set a restart date.
What it has done is let an internal capability evaluation directly delay a training schedule, rather than just shape release documentation after the fact.
Continue reading