On July 27, Moonshot AI uploaded the complete model weights for its Kimi K3 to Hugging Face, accompanied by a technical report on GitHub and a custom license called the Kimi K3 License. This marks the first open-weight model globally to enter the 3 trillion parameter range, with a total of 2.8 trillion parameters, 104 billion activated, and a 1 million token context window. The release came exactly 11 days after K3's initial debut as a hosted service on July 16.
During those 11 days, independent evaluators ranked K3 third globally in intelligence index, trailing only behind Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol. The impact on financial markets was immediate: Z.ai's Hong Kong-listed shares fell 30% intraday, while MiniMax dropped 16%. Moonshot's daily revenue surged at least sixfold, and its valuation anchor shifted from $20 billion to $50 billion. The White House's Office of Science and Technology Policy director publicly accused the company of distilling technology from Anthropic, and the Treasury Secretary threatened sanctions. In response, 25 U.S. companies including NVIDIA, Microsoft, and Meta jointly signed a letter opposing a ban on open-source models, with OpenAI, Anthropic, and Google notably absent from the initial signatory list.
From "Open Source Claim" to "File Release"
Between July 16 and July 26, K3 existed in the English-speaking tech community as an "API model with an open-source press release." The Hugging Face repository returned a 404 error, and developers could only find K2.7 Code. Moonshot's official documentation promised to release full weights before July 27, and several overseas tech media outlets ran countdown pages for ten consecutive days. Today's weight release fulfills that promise, coinciding with the peak of Washington's heated debate over whether to block Chinese open-source models.
Core Parameters (According to Official Model Card and Technical Report)
Total parameters / activated parameters: 2.8 trillion / 104 billion. Architecture: MoE, 896 experts, 16 tokens activated per token. Context window: 1 million tokens. Multimodal: Native visual understanding, vision encoder MoonViT-V2. Attention mechanism: Kimi Delta Attention (KDA) covering 69 of 93 layers, plus 24 layers of gated latent attention. Key innovations: Attention Residuals, Stable LatentMoE. MoE efficiency: Officially stated to be approximately 2.5 times the scaling efficiency improvement over K2. Moonshot claims it is the "world's first 3 trillion parameter level open model."
The License is Today's Most Underrated Information
The true new variable is not the weights themselves, but the Kimi K3 License released alongside them—a custom agreement that "looks like MIT but has additional clauses above a revenue threshold." Unite.AI's detailed analysis notes that the main body of the agreement is similar to MIT: anyone can freely use, copy, modify, distribute, sublicense, and sell the model, and can deploy, fine-tune, and build derivative models. However, there are two additional obligations. First, the Model as a Service (MaaS) clause. If a licensee provides inference or fine-tuning access to a third party, and that third party has substantial control over inputs, parameters, or training data, when that entity and its affiliates exceed $20 million in revenue in any consecutive 12-month period, they must sign a separate agreement with Moonshot. Second, the attribution clause. Commercial products with over 100 million monthly active users or over $20 million in monthly revenue must prominently display "Kimi K3" in the interface. Pure internal use, or access through Moonshot's own products or its certified inference partners, is not subject to these constraints. The practical effect of this license is to segment users by business model, not by user. A team fine-tuning K3 on its own infrastructure has zero obligations; a cloud provider reselling K3 inference at scale must first obtain a contract. This is the first time a leading Chinese AI company has explicitly turned "cloud provider free-riding" into a clearly priced commercial term in an open-source license.
The Value of Third Place Has Significantly Increased
Independent evaluator Artificial Analysis gave K3 an Intelligence Index score of 57 when it was only available as an API, ranking it third globally, alongside Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5. It is notably behind Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol. What is truly striking is the gap it opens over other open-source models: GLM-5.2 scored 51, and DeepSeek V4 Pro scored 44. Its performance in the agentic dimension is even stronger. GDPval-AA v2 scored 1668 Elo, compared to K2.6's 1190. AutomationBench-AA ranked first globally at 53%. The single-task cost is $0.94, close to GPT-5.6 Sol and about half of Opus 4.8. Frontend coding has been its most prominent arena. The Arena's Frontend Code Arena uses a human blind test voting mechanism. K3 scored 1679 on its launch day, topping the charts, leading in 6 out of 7 frontend sub-fields—a jump of 17 positions from K2.6's 18th place. Moonshot's official benchmarks report GPQA Diamond at 93.5% and BrowseComp at 91.2%, but footnotes acknowledge that different models used different agent harnesses. The official release also states that overall performance still lags behind Fable 5 and GPT-5.6 Sol, and explicitly points out gaps in user experience. The significance of open weights lies in this: from today, anyone can run these numbers themselves, no longer needing to accept conclusions under the manufacturer's harness.
Pricing: Tearing a Hole in the Logic of Closed-Source Premiums
K3's official API pricing is $3 per million input tokens, $0.3 for cached input, and $15 for output, with uniform pricing across the entire context window. For comparison, Anthropic's Claude Fable 5 is priced at $10 and $50. The Week, citing the Financial Times, reports that K3's operating cost is about one-third of Opus 4.8; Axios's figure is a 40% reduction. However, two countervailing variables must be clarified. First, K3 is "token-hungry." Multiple testers have reported that it consumes more tokens than Fable 5 for the same task. Combined with the default max reasoning mode and the $15 output unit price, the actual single-task cost may be higher than the per-token comparison suggests. Second, self-hosting is not a cost-saving button. Moonshot recommends deployment on supernodes with 64 or more accelerators, supporting vLLM, SGLang, and TokenSpeed. Because KDA disrupts conventional prefix caching, Moonshot has contributed a proprietary caching implementation to the vLLM project. The technical report also highlights a potential pitfall: K3 is trained in a "preserved thinking history" mode, meaning the harness must return the complete previous assistant message along with reasoning content and tool calls. Discarding it midway or switching from another model will cause output instability. In short, this is a three-fold event combining Hugging Face, cluster, and legal aspects, not a "download-and-use" app.
Capital Markets: First Kill Rivals, Then Kill Valuation Anchors
In the week of K3's initial launch, the Hong Kong-listed and US-listed Chinese AI sectors experienced textbook-style peer competition. Z.ai's Hong Kong shares fell 30% intraday, MiniMax fell 16%, and Alibaba fell 4%. On the US market, several media outlets mentioned that chip stocks were affected. Moonshot's own financial trajectory is steep and rare. After K3's release, daily revenue grew at least sixfold (Bloomberg); June ARR reached $300 million, up from $200 million in April. In May, the company completed a $2 billion funding round at a $20 billion valuation, led by Meituan's Longzhu Investment. It is now seeking a new funding round at a $50 billion valuation and preparing for a potential Hong Kong IPO within the year. In two months, the valuation anchor shifted from $20 billion to $50 billion. K3 is the pricing document.
Washington: From "How to Respond" to "Infighting"
At the allegation level, White House Office of Science and Technology Policy Director Michael Kratsios publicly accused Moonshot of developing K3 through distilling technology from Anthropic's Fable 5. He distinguished between two scenarios: legitimate AI distillation plays an important role in open innovation ecosystems, but "large-scale, covert, industrial-grade distillation aimed at stealing US proprietary technology" is unacceptable. Treasury Secretary Scott Bessent subsequently told CNBC that the government would investigate whether Chinese companies stole US intellectual property, stating, "We have the ability to impose sanctions on them for this." White House officials and Anthropic characterized it as "IP theft" and "industrial espionage." The Chinese side has denied the allegations, calling them unfounded. At the policy toolkit level, Axios reported that the Commerce Department last year considered adding several Chinese AI labs to the Entity List. The NSA and the National Cyber Director's Office weighed issuing advice discouraging US companies from using Chinese AI. The White House considered an executive order that would only allow US companies that can guarantee security and bear responsibility for leaks to host Chinese models. These measures were not implemented at the time. K3 has made all of them viable options again. At the infighting level, Fast Company's observation is most pointed: K3 has narrowed the seemingly comfortable lead of US frontier labs, and Washington's and Silicon Valley's first reaction was to open fire on each other. Dean Ball, OpenAI's Strategic Future Lead and former White House AI advisor, called China's open-source approach "thorough AI communism" and predicted the Trump administration would eventually realize that "the best strategy is to create a lot of regulatory risk around the use of Chinese open-source models." Notably, he also acknowledged that K3 is "a very good model" and its performance cannot be explained by distillation. Former White House AI and Crypto Lead David Sacks publicly countered: "The panic over Kimi must stop. Leading US models are still stronger." Samuel Hammond, AI Policy Director at the Innovation Foundation, said Washington had previously assumed China was months behind the frontier, adding, "One of the US intelligence community's responsibilities is not to be surprised." Researcher Kristy Loke from the MATS program noted that people close to Washington "do have a strong interest in banning these models," but she cautioned that the vast majority of US companies are not OpenAI or Anthropic. "Most companies are happy to have the option of open-source models," and a ban would protect frontier lab revenue but raise costs for US businesses and startups. The industry's counterattack came on July 24, when 25 US technology and investment institutions signed a joint open letter titled "Open Weights and American AI Leadership." Signatories included NVIDIA, Microsoft, Meta, IBM, Dell, Palantir, a16z, Hugging Face, Y Combinator, Mozilla, Mistral, Replit, Perplexity, and the Linux Foundation. The letter argued that relying solely on closed-source models is not inherently secure; they can also be broken, misused, or fail in ways invisible to the outside. Concerns about illegal distillation should be addressed through "targeted legal and commercial frameworks," not "comprehensive restrictions on a technical approach that plays an important role in innovation." NVIDIA CEO Jensen Huang promoted the letter with his first post on X, garnering over 11 million views within hours. Who did not sign is more telling than who did: OpenAI, Anthropic, and Google were not on the initial list. These three are precisely the biggest beneficiaries of Washington's ban on Chinese open-source models. Anthropic did not sign and did not provide a specific reason for this letter, but its consistent stance is that publicly releasing powerful model weights poses a security risk because access, once granted, is difficult to revoke. According to MLQ, the coalition later expanded to about 50 companies, including Google, OpenAI, Microsoft, and AMD. This is a subsequent development with a different list from the initial one, so the timing must be noted when citing it. A hedging data point: According to a joint UK-US assessment cited by Arabiya Business, K3's cybersecurity capability score is 32.2%, while the average score of leading US models is 76.2%, "far below" the latter. However, BigGo's report also mentions that K3's discovery of a zero-day vulnerability on a Reddit server within 27 minutes simultaneously shocked developers and national security officials. These two pieces of information are not entirely contradictory, but they are sufficient to show that there is no internal consensus within the US on how significant a security risk K3 poses.
The Most Critical Point: Anyone Can Download and Use It
AI Weekly's summary hits the deadlock of the entire situation: The signal beneath this week's Kimi headlines is policy, not rankings. And policy faces a physically unavoidable fact: once the weights are on the internet, they are on the internet. A more subtle concurrent event further complicates the picture. According to a report cited by The Hill, federal officials previously asked Anthropic to remove its latest Claude model from the global market for over two weeks after Amazon raised cybersecurity concerns. On one hand, there is unprecedented US regulatory intensity on domestic labs; on the other hand, anyone with a server can download and run K3 without ever passing through any US cloud. This constitutes the most awkward comparison in AI governance for 2026. Another data point changing the balance of power: According to byteiota, citing OpenRouter data, Chinese open-source models now account for 46.4% of token usage routed through the multi-model API gateway, compared to 35.7% for US models. This is the quantitative answer to the question, "Is it technically too late to ban them?"
Post-Market Observations
For domestic investors and industry observers, there are four verifiable observation points in the next 30 days. First, the Entity List: Whether Moonshot is added to it. If so, it directly creates procurement compliance risks for US companies and will reprice the valuation range of its potential Hong Kong IPO. Second, the Executive Order: Whether the White House signs an executive order targeting the output responsibility of open-source or Chinese models. As of July 25, no executive order or legislation has been passed, and self-hosting K3, Llama, Mistral, or GLM is currently not restricted by any existing law. Third, the inference service price war: The day the weights landed is the starting gun. All inference service providers will compete on the same weights, focusing on price, latency, quantization quality, and reliability. "Same model, different services" becomes a real procurement decision variable for the first time. This is a demand increment for NVIDIA and a price pressure for closed-source APIs. Fourth, community replication: The technical report is now public. Whether KDA, Attention Residuals, and Stable LatentMoE can be independently verified will determine the credibility of the "distillation theory." This is both a technical issue and a geopolitical one. Finally, a reminder of a definitional issue easily overlooked in the Chinese-language discourse: K3 is open-weight, not open-source. You receive the weights, license, and technical report, but the training data and complete training code remain the property of Moonshot. The English-speaking tech community draws this line very clearly, and mixing up the terms in Chinese reports will directly impact credibility.
Comments