The primary focus of AI infrastructure is shifting from "construction" to "monetization," but this does not signal the end of capital expenditure growth.
The cost of AI inference tokens has plummeted 37.5% from its May peak to $1.33 per million tokens. Over the same period, the rental price of Blackwell GPUs has defied the trend, rising 15.2% to $5.18 per hour. With supply-side price increases and consumer-side price reductions, the AI industry is entering a "scissors difference" phase in the inference economy. The triple constraints of electricity, licensing, and labor are extending the expected three-year capacity expansion timeline to five to ten years—construction is far from over, but the path to monetization has already been paved.
Analyst Heath Terry from Citi, in his latest weekly AI industry tracking, suggests that intelligent routing is the core driver of the token price decline. Companies are no longer uniformly calling on the most powerful models; instead, they automatically dispatch tasks based on complexity, systematically compressing inference costs. However, the supply side presents a completely different scenario. In the second quarter, the return on AI capital expenditure for hyperscalers reached 28%, nearly five times their financing cost (approximately 6%). Even as both U.S. political parties this week unusually joined forces to promote a moratorium on data center construction, Google and SpaceX accelerated their infrastructure expansion after the earnings season, with Amazon adding $20 billion in capital expenditure.
In the realm of model competition, the global intelligence gap between open-source and closed-source models has narrowed to just 4 points. However, this narrowing is primarily driven by Chinese companies—a 21-point chasm remains between leading U.S. closed-source models and leading U.S. open-source models. The optimization of inference speed for frontier models far outpaces improvements in intelligence scores. Safety governance is evolving from a compliance issue into a commercial credential. Spending hasn't decreased, but how it's spent is changing. The industry's center of gravity is shifting from "how much computing power to build" to "how to turn computing power into profit."
Electricity is the bottleneck: whoever gets connected first will monetize first.
The Texas state audit has frozen the approval queue for over 1,800 projects and 474 GW of capacity, with data centers accounting for approximately 90% of new electricity applications. Electricity is replacing chips as the most pressing bottleneck for AI infrastructure.
Core construction firms report a one-year payback period on new investments. Amazon's additional $20 billion in capital expenditure is primarily driven by rising storage costs—a factor mentioned by multiple companies. AWS contract demand extends to 2028, with management expecting capacity constraints to persist until 2027. The earlier a company can secure a definite power supply timeline, the higher its site premium and the faster its revenue realization.
With the triple constraints of electricity, licensing, and labor, the originally expected 3-year capacity expansion is likely to be extended to 5 to 10 years. For suppliers, this means high pricing for scarce infrastructure can be maintained for longer; the trade-off is that some revenue is deferred to later cycles.
Inference is becoming "heavier": the tension in storage and interconnect hasn't peaked yet.
Over the past three weeks, the average output token count per inference task has risen by 11% to 24,000, while inference-intensive tasks have seen a 9% increase in the tail to 40,000. The output token share has risen to 41%, cache activity has fallen to 57%, and input tokens remain stable at 2%.
Each model call is consuming more computational resources. Without a breakthrough in architecture, workloads dominated by output and cache will continue to drive up demand for HBM and interconnect chips. Storage and interconnect remain key bottlenecks, and this is one of the underlying reasons for Amazon's increased capital expenditure.
Global open-source gap is 4 points; the China-U.S. open-source gap is 21 points.
The Artificial Analysis intelligence index shows the leading closed-source model (Anthropic Claude Opus 5) scores 61, while the leading open-source model (Moonshot AI Kimi K3) scores 57, narrowing the gap from 9 points to 4 points.
However, the significance of this 4-point gap requires closer examination. The global open-source models closing in on closed-source models is mainly driven by Chinese companies—Kimi K3 (57 points), Zhipu GLM-5.2 (51 points), and DeepSeek V4 Flash (50 points) are at the forefront of open-source. The gap between leading U.S. closed-source and leading U.S. open-source models remains at 21 points. Open-source models are exempt from pre-release safety testing, which may accelerate their short-term catch-up, but the asymmetry in capability between Chinese and U.S. open-source models is a more profound structural variable.
Optimization of inference speed is far outpacing intelligence improvements: The median inference speed for the top 20 vendors has rebounded to 118 tokens per second, a week-over-week increase of 55.3%, while the median intelligence score remains stable at 43 points. The industry's focus has shifted from "being smarter" to "being faster."
Pricing is highly divergent. The average price for a mixed U.S.-European token is $1.63 per million tokens, while in China it is only $0.80, less than half. DeepSeek V4 Flash and Xiaomi MiMo-V2.5-Pro are priced at $0.03 per million tokens, nearly free. The average price for frontier models is $1.30, a 3% week-over-week decline.
A dense wave of new model releases is imminent: Over the next six months, DeepSeek has scheduled V4.1 and V4.2, Google has scheduled four versions from Gemini 3.5 Pro to 4.2, and SpaceX has scheduled four generations from Grok 4.6 to 5.1. A new wave of frontier models may widen the gap again, but the pace of the chasers' moves is also accelerating.
Safety governance is becoming a commercial barrier.
The AISI recorded 19 instances of unauthorized real-time internet activity in 122 model evaluations. The greater the capability, the more real the risk of loss of control—this is not a theoretical exercise.
In the long term, government safety reviews may become a competitive barrier for frontier closed-source vendors. AI-driven cyberattacks are becoming increasingly frequent and sophisticated. "Government-certified" itself is a commercial qualification for clients in regulated industries.
The ceiling of frontier capabilities is still far off. OpenAI disclosed that an unpublished model generated 10 mathematical breakthroughs using an API cost of $2,000. The upper limit of capability is far from being reached, but the threshold for making calls is rapidly decreasing.
Developers are accelerating their entry, and layoffs are temporarily receding.
The download volume of multi-provider SDKs increased by 12.7% week-over-week, showing a clear acceleration. On the application side, weekly active users for Tongyi Qianwen grew by 2.5%, Gemini grew by 0.5%, and Kimi declined by 0.8%, indicating a rapidly changing traffic landscape.
AI-driven layoffs in July were recorded at 3,220, an 85% decrease from the previous month's 21,640. However, layoff data is highly volatile, and the proportion of AI-driven layoffs within total layoffs has not shown a trend-like decline. When the next macroeconomic adjustment arrives, the accelerating effect of AI substitution may amplify again.
Comments