OpenAI has temporarily slowed frontier AI training after internal testing suggested that its upcoming Astra model may reach a critical cybersecurity capability threshold.
The company said the decision followed a security incident involving Hugging Face models and growing evidence that advanced systems could identify and exploit software weaknesses with limited human guidance.
The pause included a two-week halt on reinforcement learning training for models intended for deployment. OpenAI also placed its largest planned frontier reinforcement learning run on hold.
At the same time, it conducts smaller training runs, evaluations, and alignment testing. The move reflects a major shift in how AI developers are approaching cyber risk.
Models able to write code, use tools, browse connected systems, and reason across long tasks could potentially discover zero-day vulnerabilities, flaws unknown to vendors and defenders.
In the wrong environment, those capabilities could enable unauthorized access, data theft, destructive actions, or attempts to bypass security controls.
OpenAI said Astra may possess critical cyber capabilities, prompting stricter safeguards for all Astra and cyber-related workloads. Some training and evaluation tasks remain paused until they are moved into hardened research environments.
The company has strengthened workload isolation for systems that execute model-generated or untrusted code. These sandboxing controls are designed to contain compromised software and reduce the possibility that model outputs can affect internal infrastructure.
OpenAI has also expanded network isolation. Higher-risk workloads are being separated from the public internet and internal networks so that compromising a single service does not automatically grant an attacker or an AI model broader access.
Continuous testing is another focus. OpenAI said it has reduced standing privileges, removed potentially exposed shared services, improved security logging, and begun using models to test infrastructure boundaries against simulated attacks continuously.
A key part of the updated approach is expanded chain-of-thought monitoring. The system uses activation classifiers to inspect model activity during sampled tokens and flag potentially concerning behavior.
Automated investigation systems can then examine tool activity, available reasoning, and the sequence of actions for signs of security violations.
OpenAI aims to generate an alert within 30 minutes of detecting suspicious behavior. If teams cannot determine within another 30 minutes that the event is a false positive, they are expected to pause the activity.
The monitoring framework now applies to reinforcement learning training and tool-enabled evaluations involving models at Sol capability or above, while use of the Astra tool faces additional monitoring requirements.
The security controls are resource intensive. OpenAI estimated that monitoring currently adds roughly 20% to the inference compute used by the monitored workloads.
Beyond technical containment, OpenAI is increasing its alignment research. The company is working to improve reward models, reduce deceptive or reward-hacking behavior, and train systems to report their capabilities and actions more honestly.
The announcement shows that the race toward more capable AI is increasingly tied to a new security concern: advanced models may soon assist defenders at scale, but they could also independently accelerate vulnerability discovery and offensive cyber operations.
Prevent incidents due to slow investigations. Power your Tier 1 with threat intelligence from 15K SOCs: Integrate TI Lookup in your SOC
Your SaaS estate M365, Salesforce, Workday, Slack, hundreds of others is a sprawl of misconfigurations,…
DSPM finds sensitive data you didn’t know you had, classifies it, maps who can reach…
Open-source packages are meant to save developers time. In the GemStuffer campaign, that trust became…
Google has released an important Chrome 153 security update that fixes 42 vulnerabilities across the…
The Cybersecurity and Infrastructure Security Agency (CISA) and the National Institute of Standards and Technology…
CISA and five international cybersecurity agencies have released detailed guidance describing 17 common techniques hackers…