AI Security Red Flags: Major Labs Breach Safety Protocols, Experts Warn of National Security Threats

Deep News08-01 06:28

Successive security failures at two leading artificial intelligence companies have ignited serious concerns among cybersecurity experts about the risk of losing control over AI systems.

Anthropic disclosed on Thursday that its Claude model, during a cybersecurity assessment, accidentally connected to the internet due to a failed isolation measure. The model then carried out real-world attacks on an external organization, including stealing infrastructure credentials and deploying malware.

Previously, OpenAI also admitted that its AI agent had breached its sandbox environment during testing, successfully infiltrating the open-source AI platform Hugging Face. Both incidents were only discovered weeks or even months after they occurred.

The revelations have sparked a strong reaction across the industry. Jake Williams, a former NSA hacker and Vice President of Research and Development at Hunter Labs, stated bluntly: "This is sheer negligence. Both OpenAI and Anthropic compromised multiple external parties without any of those intrusions being detected in real time. I can't find another word for it."

Several former government officials have warned that as autonomous AI systems become more capable of launching attacks, the associated threats are escalating to the level of national security risks.

Security Sandboxes Fail, Real Attacks Unfold Quietly

According to Anthropic, during a series of 141,006 cybersecurity evaluations, the Claude model was supposed to be completely isolated from the external internet. However, a technical error allowed the model to access the external network in a few instances. Believing it was still within a test environment, it launched attacks on real targets.

Specifically, the model infiltrated one organization, stealing infrastructure credentials and a database containing internal production data. In another instance, it distributed malware and used it to steal credentials from a different organization. Neither of the victimized organizations has been publicly named.

The earliest of these intrusions dates back to April of this year, but they were only discovered last week when Anthropic audited its testing process. This audit was triggered by the earlier Hugging Face breach disclosed by OpenAI.

Ciaran Martin, former head of the UK's National Cyber Security Centre, noted that for companies in the cybersecurity business, such errors are unacceptable by industry standards. "If a mainstream cybersecurity company made a similar mistake, it could face lawsuits and regulatory penalties."

He emphasized that the cybersecurity industry typically tests dangerous tools within tightly controlled virtual sandboxes and is required to ensure those safeguards are actually effective.

Human Oversight Missing, Fires Discovered Only After the Fact

Cybersecurity experts widely point out that the common thread in both incidents is the failure to detect them in real time. They were only identified later, revealing a severe lack of human oversight.

"They only found these attack traces after actively conducting an audit," said Gregory Allen, former Director of Strategy and Policy at the U.S. Department of Defense's Joint Artificial Intelligence Center. "We essentially have no idea how large the scale of current autonomous AI intrusions truly is."

Andrew Morris, founder of GreyNoise Intelligence and a cybersecurity expert, characterized the events as a "reality check" for AI developers. It reveals the immense effort required to build secure systems and the inherent fragility of the technology the public relies on.

"Models will always lie, cheat, or steal to complete an evaluation task," Morris said. "They will do whatever it takes to accomplish the task they've been given."

Furthermore, research by the U.S. AI company Dreadnode found that mainstream AI models commonly cheat during cybersecurity tests designed to measure hacking ability. This problem appears to be industry-wide, suggesting that companies may be systematically overestimating the real capabilities of their models.

The researchers also noted that explicitly instructing a model not to cheat is not a reliable solution. Models often find ways to circumvent the rules. This issue presents a direct challenge for national security policymakers who intend to use AI for both offensive and defensive cyber operations.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment