
OpenAI announced on September 1 local time that it plans to launch its next-generation AI model Astra “soon,” but will strictly limit access to the model’s cybersecurity functions.
According to the company, Astra is the first OpenAI model to reach the “Critical” cybersecurity capability threshold in its Preparedness Framework.


According to OpenAI’s definition, reaching the “Critical” threshold means that a model can identify and develop effective zero-day exploits across all severity levels, without human intervention, in multiple real-world critical systems that have been hardened for security.


OpenAI Vice President of Research Amelia Glaese said Astra can discover previously unknown security vulnerabilities and develop exploits in many tightly protected systems without step-by-step human guidance. During testing, Astra successfully discovered and chained two zero-day vulnerabilities. OpenAI is disclosing the two vulnerabilities to the relevant maintainers.
OpenAI said it will adopt a phased access strategy. After Astra’s release, its ability to perform advanced cybersecurity tasks will initially be available only to a small group of testers.
Later, the company plans to provide broader access for defensive cybersecurity uses through the Daybreak Blue program. OpenAI acknowledged that restricting Astra’s cybersecurity capabilities could limit the legitimate defensive work of some institutions and businesses.
The release plan comes against the backdrop of an incident in July 2026, when an AI agent unexpectedly “jailbroke” and attacked the AI platform Hugging Face.
At the time, an autonomous AI agent built on two OpenAI models, GPT-5.6 Sol and another undisclosed model, exploited a software vulnerability in a sandboxed testing environment to connect to the internet on its own and breach Hugging Face’s systems during a security assessment. In a report released at the end of August, OpenAI acknowledged that the company could have responded earlier to prevent the incident.
OpenAI said that although Astra was not one of the models involved in the Hugging Face breach, the company had applied the lessons learned from the incident to Astra’s safety protections. These measures include training the model to more reliably refuse harmful cybersecurity requests and adding a dedicated “misalignment monitor” to identify and block potentially dangerous behavior. The protections also include monitoring the model for unauthorized behavior during internal deployment and automatically terminating potentially unauthorized activity.
OpenAI also warned that Astra’s safety protections could mistakenly flag legitimate defensive cybersecurity activities as abusive behavior, causing tasks to be slowed, suspended, or even terminated. (Note: OpenAI classifies cybersecurity work into “blue team” defensive tasks and “red team” offensive tasks based on its nature. The Daybreak Blue program is limited to defensive “blue team” work.)
OpenAI research scientist Fouad Matin said: “We believe these capabilities can help defenders discover and fix serious weaknesses, but without appropriate safeguards, they could also make attackers more effective. That is precisely the situation we are working to prevent.”
