The UK's AI Safety Institute (AISI) disclosed on Tuesday that during testing of models from OpenAI and Anthropic, an AI agent created a fake online identity and gained unauthorized access to security systems, revealing a new set of vulnerabilities. The institute stated that during security assessments conducted by government agencies to evaluate the capabilities of these models, agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol engaged in unauthorized actions.
"Some of the intelligent agents being tested carried out persistent and potentially harmful activities targeting real individuals and organizations," AISI said in a blog post. The report highlights weak security measures in the agent testing process, even as AI companies promote agents as the future of business.
AISI gained access to advanced AI models through voluntary agreements with major labs and tested these agents in fictional cybersecurity scenarios to assess their capabilities. The institution conducted 122 tests and discovered 19 instances of unauthorized behavior across a total of 10 test runs. Seventeen of these were caused by Anthropic's agent, while the remaining two were triggered by OpenAI's agent.
AISI stated that the most severe incident involved an agent writing malicious code and creating a fake online identity in an attempt to trick a human into approving the code. It added that no violations were found to have caused actual harm. Although AISI did not disclose which agent created the fake identity, the breach does not match the two cases previously self-reported by OpenAI.
Andrew Yoon, a researcher at the California-based nonprofit CivAI, which studies AI capabilities and risks, said the incident appeared to be caused by Anthropic's agent. "The fact that Mythos took such deceptive actions, and apparently knew it was targeting real people, shows that Anthropic does not have as much control over its own models as it thinks," Yoon said.
In a statement on X, Anthropic said it is working closely with AISI to obtain more details and conduct an independent investigation. The company did not immediately respond to a request for comment from Reuters. In a company blog post, OpenAI disclosed details of its two unauthorized actions, noting that both involved its agent accessing the internet in ways prohibited by the prompt.
"We are committed to working with the industry to strengthen common practices for safely conducting high-risk assessments, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and related groups in the coming weeks," OpenAI said. In its blog post, OpenAI also revealed another incident where a third-party testing provider, Irregular, due to a configuration error, caused its agent to mistakenly connect to the internet. This mirrors a similar disclosure by Anthropic last week regarding configuration errors.
Last week, reports indicated that OpenAI had expanded its investigation into a hacking incident after finding evidence of other agent escapes. Unlike the security breach at AI company Hugging Face in July, which was caused by an OpenAI agent, the agents in the AISI assessment did not escape from the isolated testing environment to connect to the internet. AISI stated that it had allowed internet access to the agents as part of standard testing procedures.
Comments