Moonshot AI seeks up to 30% revenue share from major cloud providers for its Kimi models

Deep News08-31 19:26

Chinese AI startup Moonshot AI is in discussions with Microsoft, Amazon.com, and Alphabet to place its Kimi K3 large language model on their Azure, AWS, and Google Cloud platforms, according to recent reports. The company is proposing to take up to 30% of the revenue generated from selling K3-related services on these cloud platforms. If successful, this would mark one of the first major model revenue-sharing agreements between a Chinese AI firm and major US cloud providers.

However, the negotiations are still in their early stages. Key issues remain unresolved, including the exact basis for calculating revenue, the scope of usage data Moonshot would receive, and how token sales by cloud vendors would be audited. All three cloud companies and Moonshot have declined to comment on the talks.

Meanwhile, Moonshot has already signed similar agreements with several smaller cloud platforms. Chinese IT services provider Chinasoft International has publicly announced it will share token revenue and other derivative income from Kimi large models with Moonshot.

Over the past year, the primary goal for Chinese open-source model companies expanding overseas was to get developers to adopt their models first. This involved opening weights, lowering API prices, listing on platforms like Hugging Face and OpenRouter, and getting models onto as many inference platforms as possible to build global influence.

But the question remains: how can open-source models sustainably generate revenue from the vast overseas ecosystem?

Based on this latest development, Moonshot is attempting to turn overseas usage into a new form of licensing revenue. Cloud vendors can purchase their own GPUs, deploy Kimi, and sell inference services to customers. Once this business reaches a significant scale, the open-source model company can also take a share of the profits.

Starting with smaller platforms

In addition to the talks with the three major cloud providers, Moonshot's developer website now features a dedicated "Access Kimi models anywhere" entry point, explicitly stating that K3 can be accessed through third-party inference partners.

Based on public information from various platforms, K3 is already available on inference service providers including Together AI, Fireworks, DigitalOcean, Modal, Baseten, and DeepInfra. When counting other smaller inference platforms offering K3 APIs, the total number of providers exceeds ten.

There is an important distinction that is often overlooked: these companies are not all API resellers.

An API reseller is essentially just a channel intermediary. When users purchase Kimi services on Platform A, that platform does not actually run Kimi itself. Instead, it forwards the request to Moonshot or another upstream provider that truly operates the model. The reseller earns the difference between the API purchase price and the retail price, but has no involvement in where the model actually runs or how inference is optimized.

True model hosting is completely different. Providers directly obtain K3's open weights, deploy this 2.8-trillion-parameter model on their own GPU clusters, and handle model parallelism, inference engines, memory management, scaling, and stability themselves. They then sell tokens to customers through their own APIs. In this scenario, the model company, like Moonshot, no longer bears the inference cost for each call—the cloud vendor becomes the true computing infrastructure provider.

It can now be confirmed that a significant portion of K3 providers fall into the latter category.

Together AI announced a strategic partnership with Moonshot in late July, explicitly using its own US infrastructure to "natively serve" Kimi models. Beyond K3, future open-weight models released by Moonshot will also be available on Together's platform. Modal has similarly stated it worked with Moonshot and vLLM to achieve day-zero support for K3, offering both shared API and dedicated deployment options.

DigitalOcean has disclosed that to launch K3, it needed to select its own hardware, adjust its inference stack, complete model optimization, and validate results against Moonshot's benchmarks. This means when users access K3 through DigitalOcean, the underlying GPU and inference infrastructure belong to DigitalOcean, not Moonshot.

It is now clear that Moonshot has established around ten or more third-party inference channels, many of which are true K3 hosting providers bearing their own computing costs. Additionally, some mid-sized cloud platforms have already agreed to revenue-sharing commercial agreements.

An overseas channel system is taking shape

These partnerships reveal that Kimi is attempting a lighter-weight approach to international expansion compared to selling its own APIs. Previously, model companies had to rent GPUs, build inference clusters, and sell tokens to developers themselves. Now, Moonshot can leverage cloud vendors' existing computing power and enterprise channels to reach overseas markets without deploying an expensive global inference network.

The key to making this model work is that Kimi has added a commercial threshold on top of its open weights.

Kimi K3 uses a custom license created by Moonshot. The first part of the agreement follows a permissive MIT-style authorization, allowing use, modification, deployment, and commercialization. However, it adds additional commercial conditions for large MaaS (Model as a Service) enterprises. If a company engaged in MaaS and its affiliates have combined revenue exceeding $20 million over twelve consecutive months, they must sign a new commercial agreement with Moonshot.

The $20 million figure is only a threshold to trigger negotiations, not an automatic 30% deduction. The "up to 30%" reported by foreign media is one of the specific commercial terms Moonshot is currently proposing to larger clients.

Another detail about K3 deserves attention. Multiple independent inference platforms are selling K3 at strikingly similar prices: Moonshot's official API charges approximately $3 per million input tokens and $15 per million output tokens. Together AI and Modal are in roughly the same price range, while some providers like DigitalOcean are slightly lower.

Different inference companies have different GPU procurement costs, scales, and inference frameworks. For an open-weight model, they could theoretically price independently. This price convergence is not a natural outcome for open-source models.

While there is no evidence yet that Moonshot has imposed uniform pricing on partners, it is clear that K3 is not being independently downloaded, deployed, and sold by each cloud vendor. Instead, Moonshot is involved in technical adaptation and coordinated operations when the model enters these platforms.

A channel system is emerging: Moonshot trains the model, maintains the model brand, and handles technical integration with providers. Inference platforms handle purchasing computing power, deploying the model, and finding customers. Finally, some partners share commercial revenue with the model company according to their agreements.

If AWS, Azure, and Google Cloud eventually join, Kimi would gain more than just three additional API access points. It would obtain a distribution network extending from specialized mid-sized inference platforms to the global large-enterprise cloud market.

Challenges with the big three cloud providers

Compared with the mid-sized inference platforms that quickly launched K3, negotiations with Microsoft, Amazon.com, and Alphabet remain in the discussion phase. The reported disagreements center on revenue distribution, data access, and token usage auditing.

These matters become significantly more complex on large cloud platforms. Inference platforms like Together AI primarily sell model services on a token basis, making it relatively straightforward to calculate K3-generated revenue.

However, AWS and Azure deal with annual contracts, enterprise discounts, reserved capacity, and bundled AI product offerings when serving large enterprises. If K3 is just one model within a multi-million-dollar cloud contract, determining how much revenue should be attributed to Kimi directly affects Moonshot's share.

Token auditing also raises data boundary issues. Moonshot needs to verify how much K3 service a cloud platform actually sells to complete settlements, while cloud vendors must protect enterprise customer data. These issues will be far more complex to resolve than with smaller cloud providers.

Alibaba exploring similar territory

Moonshot is not the only Chinese model company redesigning its open-source business model. Reports indicate that Alibaba is also considering revenue-sharing mechanisms for large commercial users of Qwen models. The new Qwen license has begun setting separate commercial authorization thresholds for large MaaS and AI Work Assistant businesses.

Alibaba's definition of AI Work Assistant includes AI coding and office productivity standalone products. This suggests future monetization targets may not be limited to cloud platforms like AWS or Together AI, but could also include numerous coding tools, office agents, and other AI applications built on open models. MiniMax has adopted a similar approach, requiring commercial users that reach certain revenue thresholds to obtain written authorization again.

However, not all Chinese model companies are moving toward "scale-based charging." Tencent's latest Hy4 preview and StepFun's Step 3.7 Flash both use the standard Apache 2.0 license without extra revenue thresholds for large MaaS operations. DeepSeek V4 Pro continues to use the MIT license, with no requirement for third-party businesses to re-authorize or share revenue as they grow.

This means an overseas cloud vendor could deploy these models themselves and build a large-scale token business without owing any additional payment to the model company under the license terms alone.

Zhipu's latest flagship GLM-5.3 has switched from the MIT license used by some of its previous models to a custom GLM-5.3 license. The new threshold appears designed not for revenue collection: only companies operating MaaS businesses with combined revenue exceeding $10 billion over twelve consecutive months are required to pass Z.AI's security review before commercial use. This extremely high threshold effectively targets only the world's largest cloud platforms and reflects model governance and safety control rather than revenue sharing.

Chinese open models are thus diverging into two distinct paths. One continues to use free and permissive licensing as its primary competitive weapon, pursuing deployment volume and market share at the lowest possible entry threshold. The other is experimenting with "free diffusion, scale-based charging"—first getting models onto as many cloud platforms and AI applications as possible, then collecting fees from participants that have created substantial commercial value.

Can a model royalty system ultimately work?

From a revenue perspective, this model is attractive. If global third-party cloud platforms sell $1 billion worth of Kimi services annually, and Moonshot receives an average 20% revenue share, that means $200 million in income. At 30%, it would be $300 million. More importantly, this revenue does not require Moonshot to bear all the investment in global inference infrastructure alone.

But the limitations are equally clear: open models can be replaced very quickly. Today, Kimi leads in performance, and cloud vendors may accept higher commercial licensing fees. But in a few months, if DeepSeek, Hunyuan, Qwen, or another open model achieves comparable performance with lower inference costs and more permissive licensing, Kimi's bargaining power will decline.

Therefore, 30% is more like the highest offer a model can make during a strong phase, rather than an established long-term "model tax." Whether model companies can create sustained revenue streams similar to software licensing or chip IP ultimately depends on how difficult the model is to replace and how high the switching costs are for developers and enterprises.

If this revenue-sharing model is eventually established with the big three cloud providers, the commercial logic of Chinese open-source model expansion would shift significantly. Previously, model companies mainly sold tokens through their own APIs. In the future, they may hand models directly to global cloud vendors for deployment and continuously share in the revenue these channels generate.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment