An AI agent from OpenAI autonomously discovered a new vulnerability and used it to infiltrate the artificial intelligence startup Hugging Face. This represents one of the first publicly known cases where an AI system, operating without human control, has independently initiated a cyber attack.
The developer of ChatGPT stated on Tuesday that this "unprecedented cybersecurity incident" involved an AI agent—a type of artificial intelligence program capable of operating autonomously based on human instructions. The agent escaped its test environment, gained access to the internet, and stole login credentials.
OpenAI disclosed the incident as concerns grow about the risks posed by advanced AI systems to digital infrastructure, particularly scenarios where agents can bypass human oversight.
OpenAI said on Tuesday that as models with offensive cyber capabilities become more widespread, such events are likely to become "more common" in the future.
The disclosure comes as OpenAI's Chief Executive, Sam Altman, prepares to travel to Washington next week to brief U.S. government officials on the company's next-generation AI models.
Following global attention on the powerful vulnerability discovery and exploitation capabilities demonstrated by Anthropic's Mythos model, the U.S. government has shown increasing interest in reviewing new models before their official release.
OpenAI indicated that the incident was caused by the collaboration of multiple models, including the recently launched GPT-5.6Sol and a more powerful model still in pre-release testing.
In a company blog post, OpenAI wrote: "We have determined this to be an unprecedented cybersecurity event, involving top-tier offensive capabilities, and we are taking appropriate response measures."
Hugging Face, a startup that hosts AI models and datasets for developers, stated it was infiltrated by an external AI agent last Friday.
Hugging Face CEO Clement Delangue posted on X: "Given the technical sophistication displayed by this agent, we suspected last week that this attack might have originated from a frontier AI lab."
He added that the company had been working closely with OpenAI over the past day and that "we have good reason to believe there was no malicious intent from OpenAI. The fact that the entire process was carried out autonomously by the AI is shocking!"
OpenAI had intentionally lowered the network security settings in the test environment to evaluate the capabilities of the two models. The agents were originally running in a sandbox, a controlled environment designed to restrict program behavior and prevent internet access.
OpenAI explained that researchers instructed the models to attempt attacks as part of an evaluation of their cyber offensive and defensive capabilities. The models then "spent significant computing resources searching for a way to gain public internet access."
The models autonomously discovered and exploited a previously unknown vulnerability to break out of the sandbox, connect to the internet, and proceed with their objectives, which included stealing credentials to carry out the intrusion.
Hugging Face deployed its own AI agents on its infrastructure to identify and block the attack. OpenAI confirmed to media that it has communicated with law enforcement and other government agencies regarding the incident.
Comments