• Like
  • Comment
  • Favorite

Chinese AI Models Seize Nearly Half of US Enterprise Token Usage, Transforming the Cost Calculus

Deep News07-21 19:04

The conversation in Silicon Valley a year ago centered on whether Chinese AI models could compete. This year, that question has been decisively answered by a compelling set of data.

OpenRouter data reveals that the share of weekly token usage by US enterprises attributed to Chinese models has surged from 4.5% in the first half of 2025 to a consistent weekly rate above 30% since early 2026, with a peak reaching 46%. This is not a fleeting spike caused by a single viral model but represents a growth trajectory that has not reversed since February 9th. Intriguingly, the shift is not limited to Silicon Valley startups but includes major players like Microsoft, a key ally of OpenAI. Price, bills, and a collective transatlantic pivot are transforming Chinese models from a technical discussion into an unavoidable cost-benefit calculation for foreign companies.

The Unforeseen Trajectory from 4.5% to 46%

If the prevailing consensus over the past two years was that US models won on quality while Chinese models won on price, 2026 is rewriting that consensus to show Chinese models are now winning on volume. AI product lead Zhu Yijun contextualizes this growth curve, stating the biggest change this year is not a comprehensive Chinese model superiority but the full-scale entry of large models into a commoditization phase. He notes the industry's focus has shifted from debating which model is smarter or has stronger reasoning capabilities to prioritizing how to complete the vast majority of tasks at low cost.

According to the latest OpenRouter statistics, the weekly token share for US enterprise calls to Chinese models was a mere 4.5% in H1 2025, with a 12-month average of just 11%. However, starting the week of February 8, 2026, this figure has not fallen below 30%, reaching 46% mid-year. Notably, this growth occurred against a backdrop of explosive platform expansion. OpenRouter's own weekly processing volume grew from about 5 trillion tokens in April 2025 to over 20 trillion by April 2026. In essence, Chinese models are not taking share from a shrinking pie but claiming an increasingly large slice of a rapidly expanding one.

Zhu Yijun also highlights an easily overlooked transmission path. As a developer platform, OpenRouter sees developers migrate first to models with price advantages, with enterprises following suit. This suggests the current curve may be a prelude to changes in B2B enterprise procurement structures.

Concurrently, data from Vercel AI Gateway confirms Chinese models are accelerating their entry into the global developer ecosystem. The platform reports that DeepSeek's token usage share surged from under 1% to 17% within May 2026, becoming its third-largest token source. Meanwhile, Zhipu's GLM-5.2, launched in June, saw its daily token volume grow approximately 50-fold in just two weeks, rapidly ascending the Vercel AI Gateway model call rankings.

This shift has captured attention even within US AI investment circles. A16Z co-founder Marc Andreessen noted on X that multiple AI practitioners judge GLM-5.2 to be the first Chinese AI model capable of matching or even beating top US lab public models on multiple tasks without significant compromise.

From Silicon Valley to Europe, numerous overseas enterprises are adopting Chinese AI models to reduce reliance on a single US provider. Companies like DoorDash, Siemens, and Airbnb have reportedly deployed Chinese AI model products, attracted by lower prices, steadily improving performance, and support for on-premise or hybrid cloud deployment. A Siemens representative stated the company internally uses a mix of several AI models, including those from Chinese firms like DeepSeek and Zhipu.

Zhu Yijun points to a changing procurement logic, with more enterprises favoring private or hybrid deployments. Some European firms, highly sensitive to data security, privacy, and avoiding vendor lock-in with US or European providers, are naturally prioritizing Chinese AI. A recent Financial Times report indicated that US export control policies targeting models like Anthropic's have prompted many European firms to reduce dependence on US AI tools, a concern that persists even if supply restrictions are lifted.

The Breaking Point of US Corporate Bills Benefits Chinese AI

When asked about the ultimate reason for this collective shift, several interviewees pointed unanimously to one factor: price. The answer lies in the ever-thickening bills from US AI model agent workflows.

An April paper from institutions including the Stanford Digital Economy Lab revealed stark data: for the same coding task, an AI Agent automated workflow consumed an average of 4.17 million tokens, while a standard code Q&A consumed only 3,390 tokens—a difference exceeding 1,000-fold. The cost stems from the Agent's need for continuous context reading, tool calling, and iterative planning; the expense is not in answering a question but in keeping the AI working.

As of July 2026, DeepSeekV4Flash's input token price is $0.14 per million, compared to $5 for OpenAI's GPT-5.5, representing a price differential fluctuating between 4x and 100x. OpenRouter states that open-source Chinese models are generally 60% to 90% cheaper than flagship products from Anthropic and OpenAI. When such price differences are applied to tasks requiring hundreds of model calls, the magnitude of the bill is amplified dramatically.

Under this billing pressure, Microsoft confirmed to Axios it is exploring using a fine-tuned, Azure-deployed version of DeepSeek V4 as a lower-cost model option for Copilot Cowork, to replace some workloads currently handled by OpenAI and Anthropic models, while shifting the product from fixed to usage-based pricing.

Zhu Yijun interprets this as Microsoft entering a "multi-model strategy era," where its relationship with OpenAI is becoming non-exclusive. He argues Microsoft now sells Copilot, Microsoft 365, and GitHub as its core offerings, with models merely providing underlying output capabilities. Given that some Chinese large models are sufficiently capable, stable, and cost-effective, they become viable options, reflecting a shift in corporate mindset.

Investor Shen Wei sees another facet of the same trend in Microsoft's moves: future enterprise clients will inevitably use a hybrid of models, combining cheaper and premium options. He observes this "mixed use" is evolving from a temporary expedient for individual companies into a default configuration.

Smaller companies were quicker to vote with their feet. San Francisco startup Lindy migrated all its AI Agents from Anthropic to DeepSeek V4, a move the founder claims saved millions of dollars and improved many core performance metrics. Airbnb CEO Brian Chesky has publicly stated the company heavily relies on Alibaba's Qwen model for large-scale production scenarios due to its sufficient effectiveness, faster speed, and lower cost.

However, Shen Wei cautions that the flip side of this shift is a price war from which no one emerges unscathed. The global large model market has entered a period of significantly intensified, even deteriorating, competition, with price wars markedly escalating in recent months. He notes this is not a unilateral price cut by Chinese models but a collective卷入 by major players. OpenAI led, Anthropic followed, both recently offering incremental benefits to heavy users—effectively disguised price cuts—while Meta has also joined the low-price, high-volume battle. This is commercial competition: when one side cuts prices, others must follow.

Shen Wei's view is blunt: this fierce battle is intensifying irreversibly, with no player's business model appearing truly healthy. In other words, the "breaking" of US model bills is less a unilateral victory for Chinese models and more the entire industry being dragged into a price competition where Chinese vendors happen to hold a more advantageous position.

Which Use Cases Are Chinese Models Capturing First?

While Chinese models are catching up in usage volume, an undeniable fact is they are primarily capturing the "middle layer," not the "pinnacle," according to AI sector analyst Sun Hao. He notes that although the performance gap in programming agent benchmarks between Chinese and top US AI models has narrowed to a single percentage point, disparities remain in areas like multimodal understanding, ultra-long context consistency, and highly sensitive scenarios involving strict regulation.

Chinese models are capturing high-volume, low-sensitivity tasks such as batch classification, content drafting, and code scaffolding, where the cost of potential data leakage is not catastrophic. This represents a substitution of "good enough" for "optimal," not a comprehensive frontal defeat of competitors.

How far this path extends depends on a sharper question: can the market share gained through low prices be converted into genuine profit? Sun Hao believes Chinese models are currently in a微妙 phase, winning on token volume but not seeing proportional growth in revenue. The low-price strategy has helped them rapidly enter developer and enterprise workflows, but the commercial value behind massive token call volumes has not been fully unlocked.

For the paradoxical situation of exploding token consumption without corresponding revenue growth, Zhu Yijun offers a detailed "pathological analysis": First, model pricing is too low, fundamentally a multiplication problem. Second, open-source models dilute inference revenue, with large model companies earning limited licensing income. Third, there is a lack of upper-layer applications; using GPT as an example, the model is merely the underlying base, with real revenue creation still coming from upper-layer apps. Fourth, enterprise software revenue is insufficient; Chinese large model companies primarily rely on API call fees, capturing only a thin sliver of value in the industry chain.

Investor Shen Wei frames this issue within a larger context: this is not a problem unique to Chinese models. The current business of selling tokens for large models itself has not yet established a complete commercial闭环, and industry business models are still under exploration. He posits that AI is more akin to electricity than the internet. Revenue is layered: model API and token billing is the first layer, AI subscriptions the second, enterprise custom software markups the third, Agent automated workflows the fourth, and vertical industry solutions the fifth. How AI makes money does not have a single answer but rather a layered profit portfolio.

Shen Wei also suggests several key indicators to watch for the direction of this industry battle: whether Agents can move from demo prototypes to real business applications; whether enterprise clients can progress from pilots to renewals and expansion; whether token growth corresponds to real revenue growth; and the dynamic changes in inference costs and gross margins. Only tokens consumed in high-value business scenarios possess true commercial worth.

A year ago, the core question for Chinese models was whether the performance gap could be closed. A year later, as the performance gap continues to narrow and price advantages are fully realized, the industry's core question has transformed into who will dominate the establishment of a new AI commercial order. The watershed moment for the AI industry may not be about who has the most tokens, but about who can create the highest industrial value from massive token consumption.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Report

Comment

empty
No comments yet
 
 
 
 

Most Discussed

 
 
 
 
 

7x24