
OpenAI said that internal and expert evaluations show that its upcoming Astra model has made major breakthroughs in agentic programming and cybersecurity. According to its Preparedness Framework, Astra has become the first model from OpenAI to reach the “Critical” risk level in the field of cybersecurity.

Note: Under the framework, an AI model reaches the “Critical” threshold if it meets either of the following conditions:
Unassisted zero-day exploitation: Without human intervention, it can identify and develop effective zero-day exploits covering all severity levels across multiple real-world critical systems that have been hardened for security.
Autonomous end-to-end attacks: Given only high-level strategic objectives, it can autonomously devise and carry out entirely new end-to-end cyberattacks against hardened targets.
OpenAI specifically emphasized that Astra is an unreleased model and was not involved in the previous cyberattack against Hugging Face.
To ensure that subsequent development remains secure and compliant, OpenAI announced that it has implemented several control measures:
Stricter security controls have been applied to high-capability models and related research and development activities, including isolated testing environments, restricted network and tool access, enhanced protection and encryption for model weights, additional monitoring and detection capabilities, and sandboxed execution. All internal Astra-related activities that do not yet meet the enhanced security standards above have also been suspended.
Global monitoring covering high-risk behavior and alignment failures has been implemented for all Astra agent applications, including during training and evaluation. High-risk operations are reviewed and blocked by assessing the model’s chain of thought (CoT).
OpenAI is working with relevant government agencies and designated AI safety organizations to jointly test the model. It is also providing third-party testing partners with recommended security control specifications to ensure that higher-risk evaluations and workloads can be conducted safely.
OpenAI CEO Sam Altman later said that Astra is a powerful model, and that they are making every effort to move forward with its public release.
We have always believed that limiting top-tier models to a small number of people is not a good strategy.
However, given its powerful cybersecurity capabilities, we need a little more time to ensure that everything is fully secure. Hopefully, you won’t have to wait too long!

