Open-Weight AI Drives A Model Routing Revolution As Token Costs Plummet, Amplifying Demand For Computing Power

Stock News07-27 20:21

The recent trajectory of AI large model technology and market penetration has surprised institutional investors focused on the valuation prospects of closed-source AI application leaders like OpenAI and Anthropic. This is especially true as global tech powerhouses, including Nvidia, Microsoft, Meta, IBM, and venture capital giant Andreessen Horowitz, urge the U.S. government to support the development of open-weight AI large models.

More than 20 of the world's top tech companies, including Nvidia and Microsoft, jointly signed an open letter supporting open-weight AI models. OpenAI and Anthropic did not sign the letter, but world's richest man and SpaceX founder Elon Musk, along with Microsoft CEO, have publicly expressed their support. In the AI developer ecosystem, open-weight AI typically refers to models where developers can download the trained parameters and deploy, fine-tune, and run inference on their own servers or cloud environments, without necessarily making the training data, processing methods, code, or full architecture details public.

This open letter, released last Friday, arrives as the White House considers whether to prohibit Chinese open-weight models over national security concerns. A total of 20 companies signed the letter, including A16z, Dell, Microsoft, Meta, Nvidia, and Palantir. For global investors focused on the AI computing power supply chain and the AI super-bull market, open-weight models, model routing, and architectural efficiency improvements will lower the unit cost of intelligence, but may amplify total computing power demand through the Jevons paradox. As the cost of simple calls decreases, enterprises will deploy more continuously running intelligent agents, parallel sub-agents, long-context analysis, code automation, and real-time multimodal services.

A single task may consume less computing power, but the number of tasks, inference steps, and deployment nodes could grow faster. Even a large MoE model like KIMI K3, which activates only a few experts, still requires massive weights to be distributed across high-capacity memory and high-speed interconnect clusters. Therefore, the proliferation of open models will shift computing power demand from a few closed-source labs to cloud service providers, sovereign clouds, enterprise AI data centers, and local inference clusters.

Why U.S. Tech Giants Suddenly Support Open-Weight AI?

Microsoft and other signatories stated in the open letter: "Open-weight AI large models expand the opportunities for global enterprises to participate in the economic prosperity of the AI era." The letter noted that organizations can develop on top of advanced AI models without needing to train models from scratch or pay the expensive token prices of the most cutting-edge models. The tech giant added that open weights can promote competition, ensuring the benefits of AI are shared more broadly, rather than being concentrated in the hands of a few. "Open weights allow every organization to match the right model to the right task at the right cost, reserving frontier-scale capabilities for truly frontier problems while integrating, running, and deploying efficient, specialized models in all other scenarios."

Nvidia CEO Jensen Huang has been a strong advocate for open-weight AI. When sharing the joint letter, he wrote: "I share a letter signed by Nvidia explaining why open AI large models are critical. AI will ultimately transform every industry, empower every enterprise, and be built and led by every country." Unlike closed-source AI systems, open-weight models allow developers and enterprises to download trained parameters and run them on their own dedicated AI computing infrastructure. This concept is similar to the open-source software movement that transformed the global computer industry decades ago.

The open letter argues that open-weight models can largely solve the problem of high training costs, allowing developers to build on existing AI models without starting from scratch. For smaller enterprises, this significantly lowers the barrier to entry in the AI product race. The letter states that open-weight models not only promote competition among AI developers but also drive healthy competition among cloud service providers, chip companies, software providers, and AI application developers. "This competition stimulates innovation, lowers costs, and widely distributes the benefits of AI across the entire economy."

In contrast, developers of closed-source AI models, such as OpenAI, Anthropic, and Alphabet's Google, clearly have strong economic incentives to maintain exclusive control of their top technologies. So, how do these tech giants, including Nvidia, view the security risks of open-weight AI large models? The open letter adds that cybersecurity defense needs priority access to these advanced AI capabilities to counter increasingly complex cyber attacks. Researchers in the letter believe that open-weight AI large models allow more researchers to identify vulnerabilities, improve security systems, and conduct large-scale, independent testing without relying solely on the original model developers.

Kimi Ignites the "Open-Weight + Low-Cost Token" Wave, Model Routing Takes Over the AI Gateway

The real impact of Kimi is not just "lower prices," but the combination of high performance-to-cost ratio, open weights, long context, and development interface compatibility that collectively lowers the barrier to model migration. Kimi K2.6's official price is approximately $0.95 and $4 per million input and output tokens, with cache input costing only $0.16. Kimi K3 offers a 1-million-token context and is priced at $3 and $15 on OpenRouter. Therefore, "low price" is relative to its capability level. More critically, open weights allow enterprises to deploy on their own infrastructure, regional clouds, or third-party inference platforms, weakening dependence on a single API provider.

The concept of "model routing" means that in the future, enterprises will no longer bind all tasks to OpenAI, Anthropic, or a single model. Instead, when a request arrives, they will dynamically select a model based on quality, cost, latency, compliance, context length, and availability. Simple classification, summarization, and customer service can be handled by low-cost open models, while complex reasoning, critical code, and high-risk decisions are directed to expensive frontier models. Microsoft's Model Router already supports three routing modes: cost, quality, and balance, with automatic failover. Microsoft believes the router can direct 60% to 80% of traffic to cheaper models without a measurable drop in quality. AWS internal tests show that smart routing can save about 16% to 56% in costs across different model families, reaching 63.6% in some RAG tests.

This means the core entry point for AI applications is shifting from "model brand" to the routing and orchestration control layer, which controls traffic distribution, model pricing, and actual token consumption. This is where the Jevons paradox may replay in the AI industry. Model routing, caching, quantization, and open weights reduce the marginal cost of each inference, but cheap intelligence will create new demand that was previously economically unviable. As long as the elasticity of token demand to price is greater than 1, a 50% drop in unit price leading to more than double the call volume will cause total computing power consumption and total inference spending to rise.

Therefore, low-cost models may trigger short-term concerns of "peak computing power demand," but in the long term, they are more likely to spread AI from a few high-value tasks to billions of daily workflows, transforming the staged capital expenditure of the training era into continuous, distributed, and high-frequency computing power consumption in the inference era.

AI Computing Power Demand is Far from Over

For the AI computing power supply chain, this is not about the disappearance of AI computing power demand, but a shift in its structure from centralized training to a balance of training and widespread inference. Frontier training will still require the highest-end GPUs, advanced processes, HBM, and high-speed interconnects. Open models and model routing will expand inference deployment in enterprise private clouds, sovereign AI, regional clouds, and local data centers, boosting demand for inference GPUs, specialized accelerators, server CPUs, high-end DRAM/NAND storage, Ethernet infrastructure, optical modules, and liquid cooling systems. The pricing anchor will also shift from purely peak FLOPS to per-watt tokens, per-dollar tokens, memory bandwidth, KV cache efficiency, and cluster utilization.

OpenAI and Anthropic will not lose all growth space, but the "toll station" model of calling closed-source models at high prices will face structural erosion. Anthropic's Claude Opus 5 is priced at $5 and $25 per million input and output tokens. Facing open models like Kimi and automatic routing, closed-source labs must prove that their frontier capabilities, reliability, security compliance, tool ecosystem, and end-to-end agents can create business value significantly exceeding the price difference.

According to Goldman Sachs, the AI super-bull market is far from over, but is entering a second phase from an "AI chip buying frenzy" to "large-scale AI factory construction." The next wave of excess alpha will no longer solely belong to the strongest names in AI GPU/AI ASIC, but will systematically spread to the entire AI computing power infrastructure stack. This includes high-performance CPUs, DRAM/NAND/HBM storage, AI PCBs, liquid cooling systems, optical interconnects, ABF substrates, and broader wafer foundries.

Morgan Stanley analysts recently significantly raised their 2027/2028 capital expenditure forecasts for the world's five largest hyperscale cloud and computing vendors to approximately $1.2 trillion and $1.4 trillion, respectively. The firm's 2026 capital expenditure expectations for large U.S. tech companies were revised sharply upward from $433 billion a year ago to $805 billion. Morgan Stanley stated that the capital expenditure super-cycle is not over, but 2026 and 2027 may be the steepest growth years. After 2028, stock prices will be determined not by "who spends the most," but by "who can fastest convert AI computing resources into revenue, profit, and free cash flow."

Citigroup's latest report indicates that the AI arms race is shifting from "whose model is most intelligent" to "who can continuously produce intelligence at the lowest cost and highest efficiency under physical constraints." Open-weight models like Kimi K3 are rapidly approaching closed-source frontiers, meaning model capabilities are becoming commoditized faster. However, the synchronous expansion of parameter scale, long context, and multi-step agent reasoning shifts the bottleneck from pure FLOPs to HBM capacity and bandwidth, GPU high-speed interconnect, cluster scheduling, and power access.

If the AI super-bull market continues, future investment trends in the AI computing power supply chain may focus on AI chips, storage, optical interconnects, AI control planes, and vertical applications. This is why Wall Street giants like Morgan Stanley and Bank of America remain bullish on leaders in the AI computing power chain, such as Nvidia, Micron, SK Hynix, and Intel. The most caution is needed for software companies that lack proprietary data, workflow barriers, and distribution channels, and merely wrap a single closed-source API—the more mature model routing becomes, the more easily their products can be replaced and their margins consumed by underlying price wars.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment