OpenAI has kicked off a "promotional" war this summer.
On July 31, Beijing time, OpenAI CEO Sam Altman announced a new pricing matrix for the entire GPT-5.6 product line. The entry-level model, Luna, saw its price drop by as much as 80%, further pushing down the price floor in the AI API market.
This adjustment covers three models. The Luna model saw the most aggressive price reduction, the Terra model was cut by 20%, and the flagship Sol model introduced a "reverse premium" model for the first time, adding a Fast mode that increases speed, priced at twice the standard rate, while intelligence levels remain unchanged.
Altman stated on social media that OpenAI's goal is to "offer the best price-intelligence ratio at every tier." This statement clearly outlines the company's competitive strategy in the AI API market: using tiered pricing to cover a full spectrum of customers, from cost-sensitive developers to high-performance users.
The Luna Model's Major Price Cut Reshapes the Entry-Level Pricing Benchmark
The Luna model saw the most significant price reduction across the GPT-5.6 product line. Its input price dropped to $0.20 per million tokens, and its output price to $1.20 per million tokens, an 80% decrease from previous levels. This directly drives down the cost of using lightweight reasoning to extremely low levels.
This pricing gives Luna a clear cost advantage in high-frequency, large-batch call scenarios, making it especially suitable for developers and enterprise clients who need to process large-scale text classification, content generation, or simple Q&A. For users with limited budgets and moderate performance requirements, Luna's new pricing significantly lowers the barrier to accessing GPT-5.6 capabilities.
The Terra Model's Moderate Price Cut Solidifies Its Mid-Range Position
Compared to Luna's aggressive adjustment, the Terra model's price reduction was 20%, with new pricing set at $2 per million tokens for input and $12 per million tokens for output.
This more restrained adjustment reflects Terra's mid-range role within the product line: balancing performance and cost for application scenarios that require a certain level of model capability but do not yet need flagship-level computing power. The 20% reduction maintains the relative value positioning of this tier while also exerting pricing pressure on competitors.
Sol Fast Mode: Trading Price for Speed for Latency-Sensitive Scenarios
The Sol model did not see a price cut this time but instead introduced a new Fast mode. This mode offers up to 2.5 times faster response speed at the API level, priced at twice the standard mode, with no change in model intelligence.
The design logic is clear: for real-time conversations, streaming generation, or production environments highly sensitive to latency, developers can pay a premium for faster throughput without compromising performance.
By introducing the Fast mode, OpenAI is essentially incorporating "speed" as an independent paid dimension into its pricing system, offering a new choice for enterprise-level users with high-concurrency, low-latency needs.
The Price War Deepens, Intensifying Competitive Pressure
This round of price cuts is OpenAI's latest move to exert pressure in the AI API market. Altman has clearly set "the best price-intelligence ratio at every tier" as the company's goal, indicating that price reductions are not a one-time event but a component of a systematic competitive strategy.
The 80% price cut for the Luna model is particularly noteworthy. This magnitude is enough to directly impact competing products in a similar market position and may force other AI model providers to adjust their pricing in response.
For developers and enterprise users, the continued decline in API call costs means the marginal economics of AI applications are steadily improving, helping to accelerate commercial deployment across more scenarios. Overall, this adjustment further strengthens OpenAI's pricing dominance in the AI infrastructure layer.
Comments