GPT-5.6's Premier Sol Model Set to Debut Tomorrow, Introducing Ultra Multi-Agent Capability

Deep News07-08

OpenAI is on the verge of launching its newest premier model, signifying a new phase in AI capability and commercial deployment.

OpenAI CEO Sam Altman announced on social media Tuesday that GPT-5.6 Sol will be officially released this Thursday. As the flagship model in the GPT-5.6 series, it features a new ultra multi-agent mode and max reasoning strength, setting new best-in-class records across core benchmarks including coding, biology, and cybersecurity.

Pricing and Availability

In terms of pricing, Sol is positioned at the top end of the three GPT-5.6 series products at $5 per million input tokens and $30 per million output tokens. OpenAI also announced that Sol will be launched on Cerebras hardware in July, achieving inference speeds of up to 750 tokens per second.

The release will follow a phased strategy, with initial API and Codex access granted only to a select group of trusted partners. OpenAI plans to roll out the full GPT-5.6 series to a broader user base over the coming weeks.

Benchmark Leadership

On the coding benchmark Terminal-Bench 2.1, GPT-5.6 Sol Ultra leads the industry with a score of 91.9%, followed by GPT-5.6 Sol at 88.8%. The competitor Claude Mythos 5 ranks third with 88.0%, a gap of about 0.8 percentage points. Gemini 3.1 Pro Preview trails at 70.7%, showing a significant distance from the top tier.

From a tier perspective, GPT-5.6 Sol Ultra and Sol constitute the first tier. Claude Mythos 5, GPT-5.6 Terra, and Claude Fable 5, all scoring above 84%, form the second tier. GPT-5.5 and GPT-5.6 Luna belong to the third tier.

In terms of cost efficiency, GPT-5.6 Sol consistently achieves the highest scores for the same API cost, offering the best value in the series. GPT-5.5 and GPT-5.6 Luna show a distinct "cost bottleneck," where increased investment yields very limited performance gains.

Regarding reasoning capability, as the number of output tokens increases, GPT-5.6 Sol's score improvement slope is the steepest, indicating it can most effectively utilize complex reasoning processes to enhance output quality. In contrast, Luna's curve is relatively flat, showing limited quality improvement even with increased output.

Architectural Advancements

GPT-5.6 introduces two key technical upgrades. The first is max reasoning strength, granting Sol ample time for deep reasoning. The second is the ultra mode, which calls upon sub-agents to work collaboratively, breaking through the capability limits of a single agent and is specifically designed to accelerate complex tasks.

In the field of biology, Sol achieved better results than GPT-5.5 on the GeneBench v1 benchmark for evaluating long-cycle genomics and quantitative biology analysis, using fewer tokens and demonstrating higher computational efficiency.

In cybersecurity, on the ExploitBench benchmark, Sol required only about one-third of the output tokens to compete with Claude Mythos Preview. On the ExploitGym benchmark, created by UC Berkeley researchers in collaboration with OpenAI and other leading labs, all three models—GPT-5.6 Sol, Terra, and Luna—showed significant growth in cybersecurity capabilities as reasoning strength increased.

OpenAI stated that Sol is better at helping users discover and fix vulnerabilities than reliably executing end-to-end attacks, emphasizing that its priority is ensuring these capabilities benefit defenders.

Product Tiers and Pricing Strategy

The GPT-5.6 series adopts a new naming system using numbers to denote generations, with Sol, Terra, and Luna representing three independently evolving capability tiers, aiming to provide users and developers with clearer choices regarding intelligence, speed, and cost.

Regarding pricing, Sol costs $5 per million input tokens and $30 per million output tokens. Terra is priced at $2.50 for input and $15 for output, approximately half of Sol's cost. Luna is the most cost-effective option in the series at $1 for input and $6 for output. OpenAI noted that Terra's performance is competitive with GPT-5.5 at half the cost, while Luna provides basic capabilities at the lowest cost.

For caching mechanisms, GPT-5.6 introduces more predictable prompt caching, supporting explicit cache breakpoints and a minimum cache lifetime of 30 minutes. Cache writes are billed at 1.25 times the uncached input rate, while cache reads enjoy a 90% discount.

Additionally, OpenAI plans to launch GPT-5.6 Sol on Cerebras in July, with inference speeds reaching up to 750 tokens per second. Initial access will be limited to select customers, expanding gradually as capacity increases.

Security and Staged Rollout

Given Sol's powerful capabilities in cybersecurity, OpenAI has equipped the GPT-5.6 series with its most robust safety protection system to date and is employing a staged release strategy.

The protection system uses a multi-layered architecture, including model-level refusal mechanisms during training, real-time cybersecurity and biology misuse classifiers during generation, account-level behavior review, and differentiated access control. For high-risk requests, the system can pause output during generation, have a larger reasoning model review the context, and intercept the content before it reaches the user.

For stress testing, OpenAI dedicated over 700,000 A100-equivalent GPU-hours to automated red team testing, focusing on discovering universal jailbreak attacks effective across various prompts and contexts, supplemented by third-party human expert red team testing.

According to OpenAI's preparedness framework, GPT-5.6 Sol does not cross the cybersecurity "critical" threshold. In evaluations involving Chromium and Firefox, Sol identified vulnerabilities and exploit primitives but, under test conditions, could not autonomously generate a usable, complete attack chain exploit.

OpenAI stated that this initial limited preview for trusted partners is part of ongoing discussions with the U.S. government. However, it also clarified that it "does not believe this government access process should be the long-term default" and will work with the government to develop a repeatable process for future model releases.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment