
OpenAI has slowed parts of the development of its upcoming artificial intelligence model, Astra, after internal evaluations indicated that the system may possess cybersecurity capabilities powerful enough to cross the company’s highest risk threshold.
The company said its recent testing showed significant advances in agentic coding and cybersecurity, prompting it to conclude that it could not rule out Astra reaching what its Preparedness Framework classifies as a critical level of cyber capability.
The development marks a notable moment in the race to build autonomous AI systems, as companies face growing pressure to demonstrate that more capable models can be deployed without creating new avenues for cyberattacks.
Why Astra is raising concerns
Unlike conventional chatbots that primarily generate text or answer questions, capable AI agents can write and execute code, use digital tools and work through complex tasks with limited human intervention.
OpenAI’s concern is that Astra may be capable of identifying and exploiting serious software vulnerabilities or developing sophisticated cyberattack strategies with little human assistance.
Under OpenAI’s framework, the critical threshold represents a level at which a model could potentially develop functional exploits against hardened real-world systems or execute novel cyberattack strategies autonomously.
OpenAI stressed that Astra remains an upcoming model and was not responsible for the separate incident involving the hacking of Hugging Face.
OpenAI tightens security controls
In response to the findings, OpenAI said it has introduced stronger controls around Astra and other high-capability models.
These measures include isolated testing environments, restricted network and tool access, stronger protection and encryption for model weights, expanded monitoring and sandboxed execution.
The company has also paused internal Astra activities that do not yet meet the strengthened security requirements.
OpenAI said it is applying monitoring across agentic Astra applications to detect risky actions and potential signs of misalignment.
The company subsequently said that the strictest security safeguards now apply to Astra and cyber-related workloads, while some training and evaluation activities remain paused as they are moved into environments meeting the new security standards.
The decision comes amid wider AI security concerns
The Astra development comes as the AI industry faces increasing evidence that autonomous systems can behave in unexpected ways when given access to tools, networks and external environments.
OpenAI recently slowed parts of its model-development programme following a cybersecurity incident involving an AI agent that hacked into Hugging Face during testing.
Reuters reported that the company temporarily paused some model testing and suspended training on Astra as it reassessed its safety controls.
OpenAI has maintained that Astra was not involved in that incident, but the episode has intensified questions about whether existing safeguards are keeping pace with the capabilities of frontier AI systems.
The concern extends beyond OpenAI as more than 100 technology companies, including OpenAI, Anthropic, Amazon Web Services and Microsoft, recently warned that AI could make sophisticated cyberattacks faster, cheaper and more accessible to malicious actors.
Astra could also signal a new phase for AI
The significance of Astra extends beyond cybersecurity.
OpenAI has been pushing toward AI systems capable of carrying out longer and more complex sequences of work rather than simply responding to individual prompts.
Recent reports has also highlighted the company’s work on more persistent AI agents that can continue working on tasks and proactively generate follow-up actions.
No public release date yet
Despite the growing attention around Astra, OpenAI has not announced a public release date for the model.
The company’s latest updates suggest that development is continuing, but under tighter security requirements and with additional evaluations before broader deployment.
The challenge for OpenAI is not how quickly Astra can become more capable, but whether the company can demonstrate that those capabilities can be controlled safely.
The decision to slow development highlights a growing tension at the centre of the AI race where the same capabilities that make frontier models more useful can also make them significantly more dangerous when placed in the wrong hands.
As AI companies move toward systems that can code, reason, use tools and act with greater independence, Astra may become an important test of whether the industry can increase AI capability without allowing security risks to outpace its ability to control them.
Join BusinessDay whatsapp Channel, to stay up to date
Open In Whatsapp
Follow the story