
On Tuesday local time, OpenAI announced a series of new safety policies focused on controlling potential security incidents involving models during testing. The new measures include more detailed monitoring throughout model development, as well as a greater emphasis on model alignment and safety during the post-training phase.

In a blog post, OpenAI said: “As model capabilities continue to improve, the risks we face while developing and testing these models internally are also increasing. The standards we set for monitoring, alignment, and safety must always stay ahead of these risks.”
This is one of the first major external disclosures of changes to OpenAI’s safety practices since the company disclosed the Hugging Face incident on July 21.
OpenAI representatives said the measures were not introduced specifically in response to the Hugging Face incident. However, the cybersecurity capabilities demonstrated by the upcoming Astra model, along with the rapid pace of development across the AI industry, were among the reasons the company adopted them.
OpenAI also disclosed that it paused reinforcement learning (RL) training for two weeks after the Hugging Face incident, but has now resumed training for many lower-risk models.
OpenAI wrote in its blog: “We are still pausing our largest frontier reinforcement learning training programs while conducting smaller-scale training and evaluations to assess model behavior, validate our safety measures, and gather more evidence about model alignment before moving forward.”
Amelia Glaese, OpenAI’s vice president of research, emphasized in an interview with reporters that as model capabilities continue to improve, the company will correspondingly increase the strictness of its safety controls, with its most capable models receiving the most stringent scrutiny.
Glaese told reporters: “We have requirements and standards for safe development. These requirements and standards are adjusted based on the level of risk we identify through our evaluations.”
Following the incident, OpenAI’s cybersecurity measures came under criticism. In the earlier incident, a model breached the original training environment by compromising an internet-connected tool within OpenAI’s network. OpenAI’s new measures include further strengthening network isolation, though the company has not yet disclosed specific technical details.
OpenAI said that under the new safety framework, “the compromise of a single workload or supporting service will not, by itself, give an attacker unauthorized access to the internet or other internal networks.”
The most important safety measure in this upgrade is a new monitoring system. It will inspect tool actions, available reasoning traces, and activity logs to identify various unauthorized behaviors. OpenAI said its goal is to issue an alert within 30 minutes of detecting suspicious activity.
OpenAI estimates that the monitoring system will require computing resources equivalent to approximately 20% of the resources consumed by the monitored process itself. The company promised to publish more details about the system in a future blog post.
In addition, OpenAI’s official post-incident analysis report on the incident has not yet been released.
