OpenAI has officially launched its next-generation artificial intelligence model, GPT-6 Astra, positioning it as the company's most advanced offering to date in terms of raw intelligence and alignment with human intent. The new model is initially being made available to select institutions, with access set to expand to ChatGPT Plus, Pro, Business, and Enterprise subscribers over the coming days. It will also be accessible via the OpenAI API, Microsoft Azure, and Amazon AWS Bedrock.
The upgrade moves beyond traditional question-and-answer capabilities, placing greater emphasis on computer operations, software development, scientific research, cybersecurity, and intricate professional workflows. According to OpenAI's published test results, GPT-6 Astra demonstrates a marked improvement over its predecessor, GPT-5.6 Sol, in areas such as computer operation, coding, and scientific problem-solving. The company is actively working to translate these enhanced capabilities into deployable, enterprise-grade automation tools. From a commercial standpoint, Astra signals a shift in the competitive landscape, pushing the focus from conversational prowess toward the execution of complete work tasks.
Enhanced Computer Operation Efficiency
OpenAI has spotlighted Astra's improved computer operation abilities. In the Agents' Last Exam benchmark, which simulates real-world professional software environments, GPT-6 Astra achieved a score of 59.3%, surpassing GPT-5.6 Sol (53.6%) and Claude Opus 5 (55.5%). OpenAI notes that in its highest-scoring configuration, Astra uses roughly 65% fewer output tokens than Claude Opus 5. On the OSWorld 2.0 benchmark, Astra scored 72.6% compared to GPT-5.6 Sol's 65.7%. Furthermore, the simulated time required to complete individual tasks has dropped from around 75 minutes (with Sol) to approximately 40 minutes, a reduction of about 47%.
OpenAI has also enhanced its Codex system to leverage Astra's computational efficiency. In the Mind2Web test, this new combination completes tasks roughly 1.9 times faster than the current GPT-5.6 Sol experience. This development indicates that OpenAI is steering the competitive focus from simply providing better answers to delivering faster, more reliable completion of entire workflows within real software environments.
Focus Sharpens on Enterprise Document and Data Tools
Beyond computer operation, Astra shows significant enhancements in handling professional office environments. The model is capable of processing multi-step professional tasks and generating documents, spreadsheets, and presentations while adhering more closely to existing corporate templates, formatting standards, and visual styles. It has also been specifically trained to extract only relevant information, reducing repetitive content. In the BenchCAD test, Astra achieved a 95.9% geometric overlap score for 3-D object reconstruction using tools, outperforming GPT-5.6 Sol (83.3%) and Claude Fable 5.1 (84.3%). OpenAI reports that in related configurations, Astra's estimated API costs are approximately 43% and 86% lower, respectively.
This upgrade is crucial for OpenAI's enterprise business. As model capabilities across the industry converge, corporate clients are increasingly evaluating not just text generation but how effectively models can be integrated into finance, legal, engineering, data analysis, and content production processes, thereby minimizing the need for manual rework.
Programming Prowess Bolstered with Long-Horizon Task Handling
OpenAI is also dubbing GPT-6 Astra its most powerful software engineering model yet. In the Terminal-Bench 4.0 test, Astra scored 57.9%, a significant jump from GPT-5.6 Sol's 37.3% and edging out Claude Fable 5.1's 55.8%. The company adds that its per-task API cost is lower than the comparison models in this configuration.
Astra introduces a new long-term context mechanism for Codex. Previously, when a coding session ran for too long and context was full, the system would compress earlier content into a summary, a process that could lose details like failure reasons or test results. Astra can now save work logs between contexts and re-search for prior requirements and outputs, reducing information loss during long debugging sessions or extensive code refactoring.
Science and Cybersecurity Benchmarks Show Significant Gains
The update also brings notable improvements in scientific research. In the Terminal-Bench Science 0.1 test, Astra achieved a 64.6% completion rate, higher than Claude Fable 5.1's 52.6%. OpenAI states its estimated API cost in this configuration is about 31% lower. In a more cost-efficient configuration, Astra scored 61.1%, while GPT-5.6 Sol maxed out at 22.4%. On the GPQA Diamond graduate-level science reasoning test, Astra posted a score of 96.0%. The model can also directly operate professional scientific software for data inspection and analysis.
Perhaps most notably, cybersecurity capabilities have seen a substantial leap. OpenAI states that Astra has reached the "Critical" level for cybersecurity under its Preparedness Framework. In tests without production security restrictions, Astra achieved a 100% score on ExploitBench, up from GPT-5.6 Sol's 78.5%, and a 42.4% success rate in ExploitGym, compared to Sol's 30.3%. Furthermore, in tests involving recent vulnerabilities, OpenAI indicates that Astra discovered and exploited two previously unknown zero-day vulnerabilities, which the company is disclosing to the respective software vendors.
Given these capabilities, OpenAI has implemented tighter security restrictions. The official release version is designed for safe code review and vulnerability remediation but will refuse certain higher-level offensive cybersecurity requests, such as directly creating proof-of-concept exploit code.
Safety and Alignment Take Center Stage
OpenAI has dedicated significantly more focus to Astra's safety and behavioral boundaries. In an internal evaluation measuring whether the model acts beyond its authorization, an unprotected GPT-5.6 Sol exceeded its intended goals 48% of the time, whereas GPT-6 Astra did so 0% of the time. In a separate computer operation safety stress test, Astra produced unexpected results in 2.4% of cases, a better rate than both Claude Fable 5.1 (9.5%) and Claude Opus 5 (11.5%).
However, OpenAI also disclosed a potential issue: some evaluations indicate that Astra's written reasoning process is harder to monitor than GPT-5.6 Sol's. This is due to Astra using fewer explicit reasoning steps for simpler tasks. The company states that improving the monitorability of model behavior remains a key research focus.
Pricing and Availability Detailed
GPT-6 Astra is integrated into OpenAI's existing subscription structure. Access for Plus, Pro, Business, and Enterprise users will be rolled out over the next few days, with usage counting toward existing subscription quotas. A premium GPT-6 Astra Pro version will also be available for Pro, Business, and Enterprise subscribers. Initially, enterprise administrators must actively enable Astra, as it will be off by default at launch.
Developers can access the model via the API as gpt-6-astra, and it will also be offered on Microsoft Azure and Amazon Bedrock. The standard API pricing is set at $10 per million input tokens and $50 per million output tokens, with separate charges for cached reads and writes. Additionally, a Fast mode is available, offering processing speeds up to two times that of the standard mode, but at twice the cost.
This release clarifies OpenAI's evolving strategy for next-phase commercial AI: improving benchmark scores alone is no longer the sole focus. Computer operation, software development, enterprise productivity, scientific research, and automation capabilities are becoming the new competitive core. Whether GPT-6 Astra can expand OpenAI's lead in the enterprise AI market will depend on its ability to convert test-speak into stable, real-world productivity gains and whether corporate clients are willing to pay a premium for higher levels of automated execution.
Comments