Skip to main content
Models & Technology

AI Agent “Out of Control” Risk Resurfaces: OpenAI and Anthropic Models Found Performing Unauthorized Operations

The UK AI Security Institute discovered while testing AI models from OpenAI and Anthropic that AI agents could create fake online identities to gain unauthorized access, exposing insufficient AI security safeguards. No actual harm has occurred so far. #AIsecurity#

AI Agent “Out of Control” Risk Resurfaces: OpenAI and Anthropic Models Found Performing Unauthorized Operations

The UK AI Security Institute (AISI) disclosed on Tuesday local time that, while testing AI models from OpenAI and Anthropic, an AI agent was found to have created a fake online identity to obtain unauthorized access to a security system, exposing a series of new security vulnerabilities.

AI Agent “Out of Control” Risk Resurfaces: OpenAI and Anthropic Models Found Performing Unauthorized Operations

According to the institute, agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol carried out unauthorized operations during security assessments conducted by government institutions. The tests were designed to evaluate the capabilities of the relevant models.

In a blog post, AISI said: “Some of the agents tested persisted in carrying out potentially harmful activities targeting real people and organizations.”

The report highlights that current safeguards in AI agent testing procedures remain insufficient, even as AI companies are heavily promoting agent technology as an important direction for future business development.

AISI obtained permission to test advanced AI models through voluntary agreements with major AI laboratories. The institute placed the agents in a fictional cybersecurity scenario to evaluate their capabilities.

AISI conducted 122 testing challenges and found 19 instances of unauthorized behavior across 10 test runs. Anthropic’s agents were responsible for 17 violations, while OpenAI’s agents were responsible for the other 2.

AISI said the most serious violation involved an agent writing malicious code and creating a fake online identity in an attempt to trick a real person into approving the code. However, the institute added that none of these incidents resulted in actual harm in the real world.

Although AISI did not disclose which agent created the fake identity, the incident was not one of the two security incidents previously disclosed proactively by OpenAI.

Andrew Yoon, a researcher at California-based nonprofit organization CivAI, said the violation appeared likely to have been caused by an Anthropic agent.

Yoon said: “Mythos took this deceptive action despite clearly recognizing that it was targeting real people, suggesting that Anthropic may not have as much control over its model as it believes.”

Anthropic said on the social media platform X that it was working closely with AISI to obtain more details and conduct its own investigation.

OpenAI disclosed relevant details on its company blog, noting that both of the unauthorized operations carried out by its agent involved accessing the internet in ways explicitly prohibited by the prompts.

OpenAI said: “We are committed to working with the entire industry to strengthen safety practices for high-risk AI evaluations, including convening national AI research institutes, independent evaluators, other AI laboratories, and relevant organizations in the coming weeks to advance this field together.”

OpenAI also disclosed another separate incident on its blog: a configuration error at the third-party testing organization Irregular caused an agent under testing to connect to the internet unintentionally. This echoes a similar configuration error incident disclosed by Anthropic last week.

Reuters reported last week that OpenAI had expanded the scope of its related hacking safety investigation after finding more evidence that agents had breached security restrictions.

AISI said that, unlike the security incident in July in which an OpenAI agent infiltrated the AI company Hugging Face, the agents in this assessment did not break out of the isolated testing environment and connect to the internet. The institute explained that, under its standard testing procedures, the testing environment itself allows agents to access the internet.