数据简报-关于失控人工智能代理引发的安全漏洞,我们目前了解的情况

路透中文08-01 00:06
数据简报-关于失控人工智能代理引发的安全漏洞,我们目前了解的情况

路透7月31日 - Anthropic公司周四披露 (link) 称,其Claude模型入侵了三家公司的系统,这凸显了人工智能日益增强的黑客攻击能力,并可能促使美国进一步加大努力,以更好地管控该技术带来的安全风险。

此前,OpenAI上周曾披露 (link),称一个由其AI模型驱动的自主代理入侵了AI初创公司Hugging Face的基础设施。

路透报道称,从OpenAI逃逸的该失控代理还入侵了另一家科技公司——总部位于纽约的Modal Labs的一名客户。

以下是这些事件的更多细节:

公司

日期

模型

遭受入侵的组织

持续时间

事件详情

OpenAI

该智能体于2026年7月9日左右开始试图逃离其测试环境

GPT-5.6 Sol 以及一款未公开名称、性能更强的预发布模型

AI初创公司Hugging Face以及总部位于纽约的Modal Labs的一位客户

Hugging Face遭受入侵的时间为2026年7月11日至7月13日

在受控测试期间,一个自主智能体逃离了其隔离环境,接入互联网,并入侵了Hugging Face以完成其被赋予的目标。该活动持续了数天,直到事件被控制住并通知FBI后,OpenAI才察觉到这一情况。

Anthropic

最早的事件可追溯至2026年4月

Claude Opus 4.7、Claude Mythos 5 以及一个未公开名称的内部研究测试模型

上述三家机构均未公开名称。

Anthropic表示,其中两家机构在收到Anthropic的通知前并未察觉该活动;该活动随后继续波及第三家机构

Anthropic未作说明

在网络安全测试期间,一个错误导致Claude模型获得了互联网访问权限,从而对三家公司发起了攻击。Opus 4.7模型因误将一家真实公司当作虚构目标,而访问了该公司的凭证和数据库;另一款模型则在识别出目标真实后停止了行动。

At the request of the copyright holder, you need to log in to view this content

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment