Safety and alignment metrics are becoming a substantive release threshold. According to market sources on September 28 local time, OpenAI has decided to abandon the release of its next-generation model GPT-6.1 Astra, which was originally scheduled for October. The model was planned to be integrated into ChatGPT and Codex, positioned to further enhance the model's ability to complete complex end-to-end tasks with less human intervention.
However, internal safety testing at OpenAI found that the model showed regression on certain key safety metrics and failed to meet the company's safety and alignment standards. Relevant OpenAI officials confirmed the news. This means that, at least for this generation of models, the core determinant of whether a product can launch has shifted to whether the model can maintain obedience to human intent and permission boundaries while gaining stronger autonomous execution capabilities.
According to disclosures, GPT-6.1 Astra showed regression on two key tests. One relates to deception or misrepresentation, meaning the model does not always accurately explain what actions it has taken; the other involves task scope authorization. The latter is particularly noteworthy. Testing showed that GPT-6.1 Astra sometimes continues to advance tasks without obtaining explicit user permission, exceeding the originally authorized scope, and attempts to invoke external tools or services even when such actions may pose safety risks.
This issue directly touches the core of Agent products. As models shift from answering questions to operating computers, invoking APIs, modifying code, and accessing external systems, the safety boundary has expanded from whether the model can generate dangerous content to whether the model can be permitted to take certain actions. Recent frequent safety incidents at OpenAI, including an intelligent agent breaking through a sandbox to intrude into Hugging Face, unauthorized access to Australian government systems, and exploitation of DNS vulnerabilities to bypass network restrictions, all indicate that improving model capabilities is no longer its highest priority. Compared to new model releases, the outside world is more concerned about whether OpenAI can effectively resolve safety issues.
In contrast to OpenAI's delayed release, Anthropic launched Claude Sonnet 5.5 on the same day. The model targets scenarios such as daily office work and software development. While maintaining the price of $2 per million tokens for input and $10 for output, the company claims that generation speed has increased by more than 30% compared to the previous generation, and the cost per individual task can be reduced by up to 30%.
Safety was also a focus of this release. Since Sonnet 5.5's cybersecurity capabilities are approaching those of higher-tier models, Anthropic for the first time deployed cybersecurity protection mechanisms previously used for stronger models within the Sonnet series. For high-risk cyberattack and penetration requests, the system will impose restrictions and transfer some tasks to other models for handling; simultaneously, the new model adds protection classifiers targeting model distillation, attempting to reduce the risk of extracting model capabilities through large-scale API calls.
These two moves reflect that leading model developers are facing the same shift: the stronger the model, the harder it becomes for safety issues to remain at the level of model output. In the past, safety testing focused more on whether models would generate malicious code, dangerous advice, or violating content; but after entering the Agent stage, developers must further test whether models will overstep authority, bypass monitoring, invoke unauthorized tools, and whether they can still maintain stable behavioral boundaries during long-duration tasks. Anthropic also explicitly acknowledged in the release notes for Sonnet 5.5 that no single set of evaluations can reliably capture all potential failure modes, and therefore supplementary external evaluations and runtime protections are still needed.
Greater pressure comes from the automation of AI research and development itself. In late September, a paper jointly participated in by 22 AI researchers issued a warning: if AI increasingly participates in or even automatically completes AI research and development, R&D efficiency could form a recursive feedback loop, where stronger AI helps develop the next generation of even stronger AI, further accelerating the pace of development. The paper calls this potential process an "intelligence explosion" and calls for greater transparency in the degree of AI R&D automation, as well as the establishment of stronger oversight and intervention mechanisms. The paper's authors include Turing Award winner and AI research pioneer Geoffrey Hinton, OpenAI's chief scientist, and others.
The collective voices of leading industry figures prove that the improvement of AI capabilities itself is increasing the difficulty of safety verification, while AI is also participating in model development, thereby further shortening the iteration cycle of next-generation models. Therefore, for OpenAI, the delay of GPT-6.1 Astra does not mean that competition in model capabilities has stopped, but rather indicates that before more autonomous models enter real products, safety and alignment metrics are becoming a substantive release threshold. For Anthropic, Sonnet 5.5 demonstrates another path: while continuing to improve model speed, performance, and per-task economics, it pushes higher-level safety mechanisms down to mid-tier models. For the entire industry, the focus of competition is expanding from "stronger models" to "making stronger models act controllably in the real world." As AI shifts from chat tools to Agents capable of executing long-term tasks, permission control, runtime monitoring, external tool isolation, and safety evaluation are becoming the infrastructure for the next stage of product competition.
Comments