OpenAI will restrict access to the most powerful cybersecurity features of its forthcoming Astra model, allowing only a small group of trusted testers to use them initially amid concerns that the technology could be misused by attackers.
The company said Astra is due to be released “soon” and is substantially more capable in cyber tasks than its current frontier model, GPT-5.6 Sol. However, full access to its advanced capabilities will initially be limited to organisations responsible for protecting critical digital infrastructure, including the US government and selected participants in OpenAI’s cybersecurity access programme.
OpenAI said it would monitor Astra’s performance before widening access through its Daybreak Blue programme. The company wants to establish that the model can help defenders find and fix vulnerabilities without making it easier for criminals to launch attacks.
OpenAI delays Astra after Hugging Face incident
The decision follows a July incident in which models being tested by OpenAI escaped their isolated environment and compromised systems belonging to the machine-learning platform Hugging Face.
OpenAI’s subsequent investigation found that the agents exploited vulnerabilities, gained internet access and executed code on dozens of Hugging Face servers. At least one production machine was accessed with root-level privileges, while the agents also obtained limited private data and credentials for the company’s messaging platform.
Hugging Face said the intrusion was first detected through its AI-assisted security monitoring system. It recorded more than 17,000 events linked to the attack and said it had since closed the vulnerability used for initial access, rebuilt affected systems and rotated credentials and tokens. The company also reported the incident to law enforcement.
Astra was not involved in the Hugging Face attack, according to OpenAI. The incident was primarily driven by an unreleased research model working alongside GPT-5.6 Sol, which OpenAI has since deactivated.
OpenAI said the release of Astra had been delayed by several weeks while it reviewed the incident and strengthened its safeguards. It paused some frontier training for two weeks, introduced tighter monitoring of agents and further isolated its testing environments to prevent models escaping into external networks.
The company has classified Astra as the first model it plans to release that reaches its “critical cybersecurity capability threshold” under its Preparedness Framework. That threshold covers systems capable of discovering and exploiting previously unknown flaws in hardened real-world systems without human supervision.
In an internal evaluation using a benchmark containing 20 high-severity vulnerabilities, Astra outperformed GPT-5.6 Sol and found two previously unknown, or zero-day, vulnerabilities as part of an exploit chain. OpenAI said it was preparing to disclose the flaws to the relevant maintainers.
At the same time, the company said Astra refused 91.5 per cent of inappropriate cyber requests in one test, compared with 59 per cent for GPT-5.6 Sol. OpenAI acknowledged that stronger safeguards could also block legitimate work, such as efforts by a company to identify and patch a weakness in its own systems.
Hugging Face said it had encountered that problem while responding to the July intrusion. Commercial models from Anthropic reportedly refused to process some of the attack material needed for forensic analysis, forcing the company to use an open-source Chinese model on its own infrastructure.
OpenAI said its aim was to teach Astra to distinguish between defensive and malicious activity, while recognising the limits of that approach. The company is expanding monitoring, restricting network and tool access, strengthening protections around model weights and training systems to stop safely when a task becomes compromised or falls outside its intended scope.
