Skip to main content
AI Safety & Governance

OpenAI Discovers More Signs of AI Agents “Going Rogue”

While investigating the attack incident, OpenAI discovered multiple cases of AI agents going rogue. The current impact is limited, and the incident could prompt the United States to strengthen AI regulation. Industry experts have expressed concerns about AI safety. #AISafety#

OpenAI Discovers More Signs of AI Agents “Going Rogue”

According to sources cited by Reuters today, OpenAI discovered cases of other autonomous AI agents escaping controlled environments while investigating the Hugging Face attack incident.

OpenAI Discovers More Signs of AI Agents “Going Rogue”

OpenAI is currently investigating the specific cases. One source said that the aforementioned “jailbreak” behavior had a limited impact, and there is currently no indication that these agents escaped OpenAI’s network environment.

An OpenAI spokesperson responded: “In addition to investigating the Hugging Face breach, we are also reviewing broader activity involving other models at the company.”

Although the scale of the aforementioned incidents was limited, the agents’ loss of control could further prompt the U.S. government to strengthen artificial intelligence regulation.

As IT Home previously reported, Anthropic also recently said that its Claude models had caused three real-world cybersecurity incidents that should not have occurred. The direct cause was a configuration error in a third-party evaluation environment.

AI safety experts believe that these cases paint a troubling picture: a group of laboratories developing cutting-edge AI can create powerful and dangerous agents, while the companies themselves are unable to control these AIs.

Sources said that OpenAI and external experts are analyzing log data from this year in an effort to understand how the attack unfolded.