The mismatch between surging demand for large language model computing power and sluggish revenue generation is reshaping the profit distribution landscape across the global AI industry chain. Over the past two years, daily token call volumes in the Chinese market have skyrocketed by more than a thousandfold, yet the full-year public cloud MaaS (Model as a Service) revenue in 2025 is expected to remain at only a few billion yuan. This massive consumption has not translated into equivalent book income, and China and the US are now pursuing distinctly different approaches to computing bottlenecks and commercialization paths.
Analyst Song Xinzhu from Northeast Securities, in an analysis of the token economy industry chain, proposed that AI profit retention is driven by four mechanisms: scarcity premium, generational premium, integrated internal settlement gains, and migration cost premium.
Currently, profits are flowing into financial statements in a top-down order: the upstream computing base layer is the first to cash in on scarcity dividends; the midstream model layer is mired in deflation caused by the commoditization of same-generation capabilities; and the downstream application layer is benefiting from lower computing costs, building long-term moats through accumulated "migration costs" over time.
Regarding the ultimate destination of premium flows, due to differences in the paying habits of the two markets, the incremental value of AI in the US is settling into high-priced software subscription systems. In contrast, China's low-priced token dividends are spilling directly into the application layer, awaiting a value reassessment after a comprehensive shift in pricing models.
Computing investment nears cash flow limits, with thousand-fold growth in calls yielding only a 3 billion yuan market.
The token economy remains in a heavy-asset construction phase. On the demand side, daily token call volumes in China have surged from approximately 100 billion in early 2024 to 100 trillion by the end of 2025. However, most token consumption occurs within the proprietary scenarios of major tech firms, not forming external transactions. The portion that does go through external transactions has its transaction prices severely compressed. Additionally, the application layer's billing has not yet fully transitioned to token-based pricing, meaning the thousand-fold increase in usage has only generated a 3.07 billion yuan public cloud MaaS market size.
Corresponding to the meager API revenue is extreme capital expenditure on the computing side. The intensity of capital spending is approaching the coverage limits of operating cash flow. As of the second quarter of 2026, the ratio of TTM (trailing twelve months) capital expenditure to operating cash flow for the four major US cloud vendors ranged from 0.63 to 1.05. Alphabet saw its first quarter of negative free cash flow, while Meta's free cash flow plummeted by 91% year-on-year. The funding for the construction phase has overflowed from operating cash flow to the capital markets.
The pace of investment in China is distinctly different. Alibaba has seen the highest capital expenditure intensity among Chinese tech firms, while Baidu is the only one among the top eight computing buyers in China to experience both a decline in revenue and an increase in capital spending.
Upstream captures scarcity dividends, as computing bottlenecks diverge between China and the US.
The upstream sector is the only link currently firmly booking profits. The "scarcity premium" based on supply gaps directly contributed to Nvidia's Fiscal Year 2026 data center revenue of $193.7 billion. Facing the same computing shortage, China and the US, under the same regulatory constraints, have formed completely different ways of clearing markets and different industrial bottlenecks.
In the US, the industry's bottleneck lies in power access. The queue for interconnection approval at ERCOT (Electric Reliability Council of Texas) includes over 1,800 projects, about 90% of which are data centers, representing a cumulative power demand of approximately 474 GW. The extended approval and power interconnection timelines have pushed North American data center vacancy rates to historic lows. In the US, scarcity is cleared through pricing, with the benefits of price increases accruing to leading companies like Nvidia.
In China, the industry bottleneck directly involves AI computing chips. Under export controls, the Chinese market clears according to the quota allocation imposed by regulations, with institutional forces driving demand toward domestic substitution. By 2025, domestic manufacturers already accounted for over 40% of the AI accelerator card market. New computing capacity is being concentrated in the "Eastern Data, Western Computing" hub nodes, with construction led by a combination of public entities, telecom operators, and private capital, forming a computing system dominated by the public sector.
Open weights break down generational barriers, turning the midstream model layer into standardized capacity.
Tokens at the same capability level are highly substitutable, and open-weight models are the primary force flattening the price gap. The cost for buyers to switch suppliers is extremely low, so competition directly comes down to listed prices. Estimates show that the cost of calling a model with capabilities equivalent to GPT-4 has dropped to about one-fortieth of its original price each year.
Prices for same-generation capabilities are rapidly converging globally. At the capability level of the AA Intelligence Index around 51 points, the blended prices of four major models from China and the US (GPT-5.6 Luna, GLM-5.2, MuseSpark 1.1, Gemini 3.6 Flash) all fall within a very narrow range of 14 to 22 yuan per million tokens. The lowest price in this tier does not come from a Chinese vendor but from Meta, which entered the API market. Once a capability tier is matched by open-weight models, the token becomes a commodity, and its price only changes with usage volume and cost.
The midstream sector is thus squeezed from both ends. On the selling price side, the actual transaction price is often about an order of magnitude lower than the listed price. On the cost side, depreciation and amortization account for over 70% of unit inference costs, with electricity costs making up less than 10%. The core space for cost reduction is not in electricity prices, but in depreciation periods and computing utilization rates.
In China, the commoditization of tokens is being pushed to the infrastructure level. Through direct subsidies to buyers via computing vouchers and the establishment of a unified platform for pricing and comparison, the room for intermediaries to add markup is being squeezed out. The distribution link is inherently destined for volume growth but thin profits. The only ends that can sustainably make money are the production end, which relies on extreme utilization for cost advantages, and the scenario end, which retains profits through high customer switching costs.
Profit destinations diverge: settling in subscriptions in the US, spilling over to applications in China.
The deflationary price of tokens releases the dividends of model upgrades to the downstream. The way China and the US capture these dividends differs due to their respective paying habits.
In the US, the incremental value of AI is being absorbed by existing subscription systems. American users are accustomed to paying high prices for software subscriptions. Frontier model vendors and existing software giants compete for the same ecosystem niche. Microsoft bundles AI features into its high-priced subscriptions, turning a temporary capability advantage into sustained customer payments.
The Chinese market is constrained by a lower willingness to pay for software subscriptions. For the same basic office software, the price in the Chinese market is often only one-fifth or even less than in the US. Consequently, midstream vendors in China typically price based on computing costs, and the value of low-priced tokens directly spills over to the application layer.
For application-layer companies in China, the shift in pricing model determines where profits will go. If subscription pricing is maintained, the dividends from lower token prices remain on the cost side, manifesting as improvements in gross margin and operating leverage. If the transition is toward token-based or outcome-based pricing, the dividends directly enter the revenue side.
However, a change in pricing model does not automatically mean profits are secured. Whether the downstream can retain profits depends on the "migration costs" established within specific scenarios. This is the only weapon to prevent buyers from clawing back the dividends through price bargaining.
The quality of a scenario is determined by the attributability of output value, the exclusivity of scenario assets, the customer structure, and the consumption intensity of the task. The moat within a scenario is maintained by three types of carriers: first, the entry qualifications and procurement systems granted by external regulations (e.g., in government and regulated industries); second, the data integration and standard interfaces accumulated over time (e.g., in financial data management and intelligent operations); and third, the replacement costs arising from system-level integration.
As the old seat-based billing system collapses, the token-based and outcome-based pricing models will significantly shorten the verification cycle for customer stickiness. Customer retention, which previously required a year-end renewal to assess, can now be clearly observed through quarterly token usage. Ultimately, the players that can retain profits will be those downstream companies that bind customers into their own business processes through extremely high switching costs.
Comments