The Shared Infrastructure Choice of Top Chinese AI Model Makers Signals a New Era's Foundation

Deep News07-20

An interesting observation has emerged.

At today's WAIC forum, both MiniMax and StepFun were present at the same AI infrastructure company's event. The former participated in a strategic cooperation signing, while the latter delivered a keynote speech.

The signs were visible even earlier. Four months prior at the Zhongguancun Forum, Zhipu AI's Zhang Peng and Kimi's Yang Zhilin shared a stage with this company's co-founder and CEO for a panel discussion, where it was highlighted that the company already provides services for Kimi and Zhipu AI.

Four leading domestic foundational model companies have now entered into deep collaborations with the same AI infrastructure provider. In a playful sense, gathering all four might just summon something extraordinary.

This AI infrastructure company is Infinigence-AI. Its position is somewhat analogous to "the CATL of the battery world" – while automakers can pursue their own paths, the foundational battery layer is an unavoidable necessity.

Infinigence-AI aims to become the common choice for the large model era. This raises a compelling question: what exactly are these model companies seeing in it?

The answer lies on both sides of the equation. On the demand side, by 2026, inference is projected to surpass training as the primary consumer of AI computing power. While inference costs have dropped 280-fold in two years, total enterprise AI spending has not decreased but risen. Daily Token calls in China have already exceeded 140 trillion, a 40% increase in a year, with demand surging exponentially.

The supply side, however, presents a completely different picture. The expansion of physical computing power remains linear – building more data centers and buying more chips – yet the supply gap is projected to remain unfilled for the next three to five years.

With exponential growth on one side and linear scaling on the other, the chasm in the middle is precisely where Infinigence-AI seeks to position itself.

The greater challenge is that the quality of model deployment is difficult for outsiders to judge. A request might receive a normal-looking model response, but the output accuracy could have silently degraded by 30%, an issue conventional monitoring cannot detect. By the time users notice something amiss in their business applications, the model's reputation may already be tarnished. This invisible threshold is what truly determines which suppliers remain in the game.

Key Reasons for Preference

Since the full-scale explosion of the Token economy, inference demand has accelerated, and the computing power gap continues to widen. However, the barriers to the MaaS (Model-as-a-Service) business are much higher than commonly perceived.

The decisions by Kimi, Zhipu AI, MiniMax, and StepFun naturally involve their own strategic calculations, but they fundamentally represent a bet on three critical factors simultaneously: whether model performance degrades, if costs are sustainable, and whether the system remains stable under pressure.

First, performance. The aforementioned hidden performance degradation is a major industry pain point. There are cases where third-party deployments have led to a 30% drop in accuracy compared to the original model provider's service, a change the client didn't notice until key business metrics began to decline.

Infinigence-AI's approach to this is to establish a set of准入测试标准 (admission testing standards). This standard set verifies everything from tool calling consistency to inference mode accuracy alignment. Every new model must pass this checkpoint before being listed. The result is that clients experience virtually no difference whether they use Infinigence-AI's service or the original model provider's API.

Second, cost. At this year's WAIC, Infinigence-AI announced a self-developed hardware technology: cross-cluster heterogeneous P-D separation. In large model inference, the Prefill and Decode stages have completely different workloads and hardware requirements. Deploying them separately allows different chip types to do what they do best.

However, separating them across different clusters introduces a new problem. After P-D separation, transmitting the KV Cache between heterogeneous chips over a wide-area Ethernet network faces issues of low bandwidth and high latency. It's like two relay runners performing excellently individually but fumbling the baton handoff.

To address this, Infinigence-AI first innovatively migrated its Radix Cache technology, designed for Decode instances, into this cross-cluster architecture, reducing the volume of transmitted data by an order of magnitude.

Next, they pioneered a PDD architecture, splitting the traditional P-D chain into three levels: P, RelayDecode, and MainDecode. For requests with high transmission latency, RelayDecode steps in first to deliver Tokens to the user, who remains unaware of the actual latency lasting tens of seconds. Once data transmission is complete, it seamlessly hands off to MainDecode. Tests show this architecture reduces Time-to-First-Token (TTFT) by 51.5% while lowering per-Token cost by 37.5%.

Third, stability. At large cluster scales, faults are often deeply hidden. Large model inference requires managing tens to hundreds of clusters, handling nationwide daily traffic distribution on a terabyte scale. If a server fails, relying on human monitoring is insufficient for timely response.

Infinigence-AI has now released an "Intelligent Computing Cluster Operations and Maintenance Agent System." This system provides end-to-end solutions for operational challenges in real production environments, offering 7x24 all-weather monitoring. It transforms cluster operations from "people finding problems" to "problems finding people" and even "problems solving themselves," achieving over a 5x improvement in operational personnel efficiency and a 6x increase in critical fault resolution efficiency.

Building Systemic Competitiveness

Solving these three issues merely forms the foundation of the "Token factory" layer. The complete solution Infinigence-AI aims to deliver is integrating computing power, Tokens, and productivity into a complete system, corresponding to its own proposed formula: AI Productivity = Intelligent Resource Scale × Token Conversion Efficiency × AI Productivity Conversion Efficiency.

The "central hub" refers to the "Computing Power Distribution Center," namely the Agentic Infra autonomous infrastructure platform. The domestic chip ecosystem is inherently fragmented. Infinigence-AI aggregates scattered computing resources, enables elastic scheduling, and facilitates on-demand utilization, building a sufficient, stable, and scalable computing foundation for the model and application layers. The core goal is clear: maximizing intelligent resource scale.

This distribution center has already deployed and accessed 37,000 Petaflops of computing power, covering 16 mainstream chip types. The challenge of cross-cluster reinforcement learning is also tackled at this layer. Infinigence-AI believes that with the ongoing development of Post-training Scaling Laws, reinforcement learning has become key to unlocking intelligence. Its hardware requirements are more complex and its scale more massive. Based on the dual necessities of heterogeneity and超大算力规模 (ultra-large computing scale), cross-cluster reinforcement learning has become a new anchor point for intelligent scaling and continuous evolution. This is a core scenario Infinigence-AI's Agentic Infra platform focuses on.

By optimizing across the network, platform, and framework layers, Infinigence-AI has successfully achieved stable, uninterrupted cross-domain reinforcement learning training for a continuous week, making large-scale cross-domain reinforcement learning not just feasible, but also fast and stable.

In the future, the company will continue to deepen core technologies like parallel strategies, communication fusion, intelligent operators, and extreme fault tolerance, expanding the support scale for cross-domain computing resources to over 100,000 cards.

The "back-end factory" is the "Token Factory," i.e., the Agentic MaaS large model service platform. This is where the aforementioned advantages in performance, cost, and stability are truly realized. The core concept is "deriving production capacity from efficiency on top of scale." It is a complete service technology stack from gateway and routing to underlying inference instances, where "every layer is optimizable, every point offers incremental gains."

Within this "factory," besides the new "cross-cluster heterogeneous P-D separation architecture," Infinigence-AI also collaborates deeply with several leading large model companies, continuously refining inference efficiency and service stability in real business scenarios. The growing business scale, in turn, accelerates technological iteration. The company announced that as of July, the daily Token call volume on its Agentic MaaS platform has increased 40-fold compared to December last year.

It has also co-created the "Tianwen" model service portal with Shanghai Mobile and developed "TokenDance" for developers with Guanzha, positioning it as a counterpart to OpenRouter.

The "front-end store" is the "AI Productivity Store" that extends accumulated capabilities across countless industries, i.e., the Agentic Infra industry solutions. This solution set already covers sectors like entertainment/gaming, healthcare, legal terminals, and energy/power, transforming technological potential into practical, usable, and replicable value for various industries.

For example, it has developed a 3D gaming solution with VAST, explored healthcare large model implementation scenarios with Shanghai Ruijin Hospital, created AI partner products for several law firms, and is working with China Southern Power Grid on energy consumption prediction for intelligent computing centers.

While supporting the intelligent upgrade of myriad industries, Infinigence-AI also continuously applies agent capabilities to AI Infra itself. The company has built an infrastructure agent swarm, which currently includes a platform管家智能体系统 (steward agent system) for users that independently resolves 80% of daily issues; the aforementioned operations and maintenance agent system for clusters; and an operator generation agent system for chips that continuously self-iterates through real compilation trial-and-error and knowledge沉淀 (precipitation), stably delivering high-quality, production-ready industrial-grade operators.

In tests with the GLM 5.2 model, this system achieved a 7.3% end-to-end performance improvement on NVIDIA flagship cards compared to industry baselines, and a more significant 39% overall performance leap on AMD flagship cards.

These three pieces fit together. The experience gained from the "front-end store" across industries feeds back to optimize the inference strategies of the "back-end factory" and the scheduling of the "central hub," creating a cycle of mutual reinforcement.

The Broader Strategic Vision

If Infinigence-AI is viewed merely as a company selling Tokens, its complete business model is missed. The "Token Factory" is just the entry point. What it is truly building is a comprehensive Agentic Infra strategic layout of "front-end store, back-end factory, central hub," designed for the Agent era. These combined directions tell the full story it aims to convey.

The persuasiveness of this narrative is most directly evidenced by the deep collaboration partners mentioned earlier. Looking overseas, Infinigence-AI's counterparts include Baseten, Together AI, and Fireworks AI, with valuations in the $13-15 billion range. However, these three are built on the单一生态 (single ecosystem) of NVIDIA, while Infinigence-AI operates within a far more fragmented domestic chip ecosystem. Consequently, it has turned多元异构 (multi-vendor heterogeneity) and a more complete architectural layout into its own moat.

Returning to the initial question – why have these leading model companies all agreed on this infrastructure provider – the answer may not be so complex. Mastering computing power, understanding model architecture, excelling at training/inference optimization, and comprehending application scenarios: excelling at all four simultaneously is the truly稀缺的能力 (scarce capability) in this field.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment