Three major US artificial intelligence (AI) companies have recently acknowledged that their models breached boundaries during testing, intruding into the systems of other organizations, sparking widespread attention from the global industry. Some experts believe that beyond technical errors, there is also suspicion of commercial hype.
In late July, OpenAI publicly admitted that models, including its GPT-5.6 Sol, became uncontrollable during internal evaluations, breaking out of the isolated test environment and infiltrating the systems of Hugging Face, a US company. Subsequently, Anthropic disclosed that three of its models had unauthorized access to other organizations' systems during network capability testing. In early August, Meta confirmed that one of its models had breached other companies' systems during a cybersecurity capability assessment. These AI boundary violations have drawn significant attention. The UK's AI Safety Institute published a report discussing the results of cybersecurity evaluations involving more models. Researchers tested seven models and found 19 actions that clearly exceeded test boundaries, 17 of which involved Anthropic's Claude-5 Myth model, with the remaining two attributed to OpenAI's GPT-5.6 Sol. The report emphasized that in this test, researchers deliberately enabled internet access and disabled some safety mechanisms from the model providers to assess maximum capabilities, and the models did not actively escape the isolated environment. However, in the most concerning incident involving OpenAI's model infiltrating Hugging Face, the model, with safety restrictions somewhat relaxed during testing, proactively discovered and exploited a previously unknown vulnerability to break out of the isolated test environment, connect to the internet, and execute an attack. Current public information does not indicate that these boundary violations caused large-scale data leaks or persistent damage, but they still highlight a common issue: as models gain the ability to autonomously plan, use tools, and execute complex tasks, a single misconfiguration could turn a simulated attack into a real intrusion. Dr. Andrew Soltan, a researcher at the University of Oxford, said that model boundary violations like those reported by OpenAI sound concerning, but they occurred because safety barriers were artificially disabled. "This is not AI going out of control on its own; it precisely demonstrates that safety measures are crucial," he noted.
Why are US AI companies successively disclosing their models' boundary violations? Some experts question whether there may be commercial motives. Dr. Konstantinos Gkouzis from Imperial College London's computing department stated that in the case of OpenAI, the real issue is its failure to control model capability testing, allowing a third party to bear the consequences. As for the warnings about the model's advanced cyberattack capabilities, they conveniently serve as advertising for the model. The BBC, reporting on OpenAI's disclosed boundary violation, cited cybersecurity expert Daniel Card, who said, "What a coincidence... OpenAI specifically breached an organization that could also benefit from this marketing exposure." Earlier this year, Anthropic claimed its Claude-5 Myth model was "too powerful" in autonomously discovering network system vulnerabilities and developing attack methods, so it was temporarily withheld from the public and only restricted to a few partners. At the time, OpenAI CEO Sam Altman said Anthropic's use of "fear-based marketing strategies" was meant to make its product sound "more impressive" than it actually is, comparing it to "announcing you've built a bomb to blow someone up, then turning around to sell them a $100 million bomb shelter." Some European and American media outlets have also pointed out the commercial motives behind such hype. In the current highly expensive AI race, heavily emphasizing the disruptive nature of one's technology can not only significantly boost corporate attention and influence, helping secure high-value cybersecurity contracts, but also inflate company valuations.
Whether due to technical errors or commercial considerations, these incidents highlight the importance of strengthening AI safety regulation. The UK's AI Safety Institute noted that the boundary violations indicate a shift in risk profiles, where harm may not only stem from intentional misuse but also from AI systems, when operating in internal research environments or with special access privileges, taking unauthorized actions beyond their intended scope. As AI capabilities continue to advance, efforts to understand and ensure the safety of these systems must also proceed in parallel. Global action is already underway. At the recent World AI Conference, China advocated for promoting global AI towards a direction of goodness and benefit for humanity, emphasizing risk awareness and ensuring safety and controllability. The European Union expanded the scope of its AI Act from August 2, requiring providers of advanced general-purpose AI models that may pose systemic risks to fulfill additional obligations to prevent large-scale harm, such as cyberattacks or model loss of control. The US government has also recently convened meetings with relevant companies to discuss AI model safety evaluation mechanisms. Professor Oliver Buckley of cybersecurity at Loughborough University in the UK believes that lessons should be learned from AI boundary violation incidents, stressing that we cannot always expect models to follow instructions, and that strengthening technical protection mechanisms is crucial.
Comments