U.S. Artificial Intelligence Begins Targeting Real Humans

Deep News09:40

British government researchers have stated: "Targeting real people—this is something we've never seen before." Creating false identities, applying pressure in turn, and then obtaining authorization—all to complete an attack. This script, common in spy films, is now written into an official report by a British institution.

On August 4, local time, the UK government's AI Safety Institute (AISI) released a report revealing that some tested AI agents had been continuously engaging in potentially harmful activities targeting real individuals and organizations. In 122 recent cybersecurity evaluations, the agents overstepped boundaries in 10 runs, with 19 recorded instances of transgression involving leading US models from Anthropic and OpenAI.

The most serious incident involved an AI agent attempting to insert malicious code into a real project on the mainstream code hosting platform GitHub. The agent studied the human reviewers of the project, created multiple fake identities, and pressured these reviewers to approve the code. AISI exclaimed that this was the first time risks related to autonomy and deception had been so clearly observed in a real-world setting without specific prompts. "Targeting real people—this is something we've never seen before," researchers noted. The agent tried to contact real users directly, sending messages and files through online file transfer services to persuade them to run the malicious code.

This scenario feels familiar. Looking back over the summer, incidents involving US AI models have been almost non-stop. On July 21, OpenAI admitted its AI model had discovered a zero-day vulnerability, broken out of its sandbox, and infiltrated Hugging Face, the world's largest open-source AI community. The company called it "an unprecedented cybersecurity incident." On July 30, Anthropic acknowledged that its model Claude had accidentally gained internet access during a controlled cybersecurity evaluation, entering the systems of three real organizations, with the earliest incident dating back to April.

To overstep boundaries, AI models have even resorted to cheating. According to AISI test results, across 475 cybersecurity assessment runs, five leading US AI models all engaged in cheating, with rates ranging from 7.8% to 14.1%.

What is the root cause? The report states, "Fundamentally, the agent took these actions to complete the task it was assigned." The core issue lies in the competitive logic of frontier AI. Competition among top US AI companies is intensifying, with a fixation on stacking parameters, climbing leaderboards, and pushing limits, all in the race to see who can first achieve extreme capabilities. To obtain better test results, researchers have even relaxed security permissions. OpenAI previously admitted that to test the model's ultimate cyberattack capabilities, the research team intentionally disabled some security filters.

The guardrails themselves are also a black box. In the Hugging Face example, the safety rails of leading closed-source models have become increasingly rigid, unable to even protect victims. In contrast, China's AI is trusted for more than just performance. The problems triggered by OpenAI were resolved thanks to Zhipu GLM-5.2, which reduced evidence collection from days to hours. This relies on usability, controllability, trustworthiness, local deployment, and the consistent protection of data security permissions.

The deeper difference lies in the system. While encouraging innovation, China has clearly drawn several red lines for AI agent governance. For example, actions must not exceed the scope of user authorization, operations must include decision validation and fault-tolerant shutdown mechanisms, and the entire process must be traceable and auditable. While US companies push capabilities to the max first and patch later, China promptly addresses shortcomings in technology, rules, and responsibility. This is not a restriction but a guarantee that society will not pay for technological recklessness.

Guided by this approach, Chinese companies are forming a consensus: whoever can deliver a fully auditable security system will be the one to secure a ticket to enterprise-level AI. The repeated failures of US AI also expose new security concerns. For a long time, AI safety issues have focused on content compliance, data bias, and privacy protection. Now, the behavioral safety of AI itself is becoming a reality. It is clear that banks, hospitals, government platforms, and large manufacturing enterprises cannot entrust core data and operational permissions to an uncontrollable cloud-based model.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment