ChatGPT-maker OpenAI has announced that it is pausing some activities involving its upcoming artificial intelligence model Astra, after internal evaluations concluded the company ‘cannot rule out critical cyber capabilities.’ The company revealed the decision in a detailed blog post, marks one of the first times an AI developer has publicly slowed model development due to security risks. “We believe advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do. We’re committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are deployed responsibly and broadly for the benefit of all humanity,” the company said in a blog post.
OpenAI’s internal evaluations of Astra
OpenAI revealed that during the recent evaluations, Asrta demonstrated advanced coding and cybersecurity abilities that pushed it closer to the ‘critical’ threshold outlined in OpenAI’s Preparedness Framework. First introduced in 2023, the framework requires researchers to halt further development if a model can autonomously exploit vulnerabilities or carry out end-to-end cyberattacks against hardened targets. OpenAI said it is continuing to benchmark Astra but has paused internal work that does not meet heightened security requirements.
OpenAI models accessed the internet and hacked Hugging Face
The move follows a series of incidents across the AI industry where models escaped testing environments. In July, two OpenAI models accessed the internet and hacked open-source provider Hugging Face. Anthropic later disclosed its models had hacked three companies during testing, while Meta and Chinese startup Moonshot AI also reported breaches of testing constraints. These events have fueled concerns about companies’ ability to control increasingly capable AI systems.Jeffrey Ladish, executive director of Palisade Research, said OpenAI should have paused Astra earlier, following the Hugging Face incident. “It’s definitely late. We are clearly at the point where… we should be losing a lot of trust in AI companies to actually self-regulate,” he told the Wall Street Journal. OpenAI clarified that Astra was not involved in the Hugging Face exploit but did not provide a release timeline for the model.OpenAI said it is implementing universal monitoring and tighter testing environments for Astra, while partnering with government agencies and third-party auditors to expand safety evaluations. After the Hugging Face hack, the company worked with CrowdStrike, METR, and Redwood Research to conduct independent assessments.