China Securities Co., Ltd. has released a research report stating that the release of the Kimi K3 marks the entry of domestic open-weight models into the global competition for capability boundaries and commercial pricing power. The focus moving forward will be on observing community deployment scale, API call volumes, and paid conversion rates following the full release of the model weights.
Taiwan Semiconductor Manufacturing Company (TSM.US) reported earnings that exceeded expectations and raised its full-year revenue guidance. This follows similarly impressive preliminary earnings reports from downstream chip, server, and switch-related companies, indicating sustained momentum transmission through the AI hardware chain. This is ultimately reflected in continuous iteration on the model side and application implementation. The firm remains optimistic about segments including GPU, CPU, storage, high-speed networking, and compute leasing services. The key points from China Securities Co., Ltd. are as follows:
Kimi K3 Officially Launched
The launch of the Kimi K3 signifies that domestic models are further closing in on the global frontier in terms of parameter scale, long-context agent capabilities, and knowledge work proficiency. Moonshot AI released Kimi K3 on July 16, 2026. The model boasts a total parameter scale of 2.8 trillion, employs a Mixture of Experts (MoE) architecture, natively supports visual understanding and a 1 million token context window, and is primarily oriented towards software engineering, deep research, complex knowledge work, and multimodal content generation. K3 is currently accessible via Kimi, Kimi Work, Kimi Code, and official APIs. The company plans to release the complete model weights before July 27. If completed on schedule, this would make it the world's largest open-weight model by parameter scale. The release of K3 indicates that competition among domestic models is shifting from cost advantage and catching up in specific capabilities to a comprehensive competition encompassing model scale, architectural innovation, and complex task delivery capability.
Architectural Innovation and Scale Expansion
K3 supports its trillion-level parameters through extremely sparse MoE and hybrid linear attention mechanisms. It utilizes a Stable LatentMoE architecture where each token activates only 16 out of 896 experts. The model introduces two core technologies: Kimi Delta Attention and Attention Residuals. The former reduces computational overhead for long sequences via hybrid linear attention, while the latter allows the model to selectively invoke historical representations across different network depths, improving information transmission efficiency in ultra-deep networks. Building on this, techniques like Quantile Balancing, Per-Head Muon, SiTU, and Gated MLA optimize expert load balancing, attention head training, and activation control. Official disclosures indicate K3 achieves approximately a 2.5x overall scaling efficiency improvement compared to K2. On the training side, quantization-aware training is introduced from the SFT stage, employing MXFP4 for weights and MXFP8 for activations, balancing training stability for ultra-large models with subsequent hardware adaptation efficiency.
Front-End Programming and Long-Context Agent Performance
K3's capability improvements have extended from single-turn code generation to complex engineering delivery. K3 scored 1679 points in the Frontend Code Arena, ranking first globally, and secured the top spot in six out of seven subcategories including brand marketing, reference design, data analysis, consumer products, simulation, and content creation tools. An evaluation by Artificial Analysis on July 17 showed K3 scored 57 points on the Intelligence Index, ranking third, behind only Claude Fable 5 and GPT-5.6 Sol, and above Claude Opus 4.8 and GPT-5.5. In long-context knowledge work (AA-Briefcase), it achieved an Elo rating of 1547, ranking second. In terms of engineering cases, K3 can continuously perform GPU kernel optimization, build a MiniTriton compiler from scratch, and autonomously complete chip design, optimization, and verification using open-source EDA tools over a 48-hour run, reflecting significantly enhanced capabilities in tool calling, code execution, visual feedback, and long-duration task planning.
Independent Testing and Cost Considerations
Independent testing by tech media validates front-end and coding capabilities, though inference speed, stability, and usage costs still have room for optimization. According to evaluations, K3 can form a closed loop of "code generation - runtime screenshot - visual recognition - code modification" in front-end development and can handle large code repositories and cross-file development tasks using its 1 million token context (QuantumBit). In a 3D front-end test reusing the same prompts, K3 was deemed superior to the GPT-5.6 series in page completion, visual effects, and interaction design, while also showing strength in identifying conflicting requirements, locating structural bugs, and long-context research. However, some complex front-end tasks took nearly an hour, and initial runs, mobile adaptation, and visualization projects may still require multiple rounds of fixes. The model also exhibited an issue of being overly proactive when intent is unclear (Photon Planet). Overall, K3 has entered the global top tier, but its commercial ceiling will still be determined by product experience, output speed, and success rates on complex tasks.
Pricing Strategy Shift
Performance leaps are driving a significant upward shift in pricing, with K3 beginning to move from low-cost competition to pricing based on frontier capabilities. The official API pricing for K3 is $0.30 per million tokens for cache-hit input, $3 for regular input, and $15 for output, with corresponding domestic prices of approximately 2 yuan, 20 yuan, and 100 yuan. Compared to K2.7 Code, regular input and output prices have increased by about 2.2x and 2.75x, respectively, reaching roughly 3.2x and 3.75x the previous generation's levels. Artificial Analysis estimates the average cost for K3 to complete a single Intelligence Index task is about $0.94, close to GPT-5.6 Sol's $1.04, but only about 52% of Claude Opus 4.8's cost. This price increase reflects both the significantly higher resource consumption of long-context reasoning and agent tasks, and indicates that domestic model vendors are beginning to attempt to capture commercial premiums based on capability differentiation, rather than relying solely on low-cost APIs to expand call volume.
Infrastructure Demand and Cost Dynamics
While per-token costs continue to decline, total compute consumption driven by complex tasks is still expected to expand. Moonshot AI recommends deploying K3 on super-nodes composed of 64 or more accelerator cards. The demands of a 1 million token context, 896-expert MoE, visual closed loops, and temporal agent operation place higher requirements on GPU compute, HBM capacity, high-speed interconnects, CPU scheduling, KV Cache, and storage systems. Techniques like KDA, quantization-aware training, and prefix caching can lower per-token inference costs. However, with simultaneous increases in inference depth, output length, tool-calling rounds, and task duration, the tokens and infrastructure resources consumed per task are still likely to rise.
TSMC Exceeds Expectations Again
TSMC's earnings call once again exceeded expectations, leading to an upward revision of its full-year revenue guidance. On July 16, TSMC released its Q2 2026 financial results. Quarterly revenue reached NT$1.27 trillion, a year-over-year increase of 36% and a sequential increase of 12%. Revenue in USD terms was $40.2 billion, hitting the upper end of prior guidance. Gross margin was 67.7% and operating margin was 60.3%, both exceeding the high end of guidance. Net profit attributable to shareholders was NT$706.6 billion, a year-over-year surge of 77.4%, significantly surpassing market expectations. Structurally, by application, HPC revenue grew 20% sequentially, increasing its revenue share to 66%, while smartphone business revenue declined 4% sequentially. By process node, 2nm contributed approximately 3% of wafer revenue for the first time, while 3nm, 5nm, and 7nm contributed 30%, 33%, and 11% respectively. A steep ramp for the 2nm process is expected starting in Q3. Benefiting from strong demand from CSP customers, TSMC raised its 2026 USD revenue growth guidance from "above 30%" to "slightly above 40%." Concurrently, it significantly increased its full-year capital expenditure guidance from $52-56 billion to $60-64 billion, with 70%-80% allocated to advanced processes and 10%-20% to advanced packaging, testing, and masks. The company explicitly stated that AI demand is expanding from GPUs to data center CPUs, custom ASICs, and networking chips, and that advanced packaging capacity remains tight. For Q3, revenue guidance is $44.6-$45.8 billion, implying sequential growth of about 12% and year-over-year growth of about 37%. On the profitability side, due to the steep 2nm ramp and overseas capacity expansion, gross margin is expected to be diluted by 3-4 percentage points in the second half, with a Q3 gross margin guidance of 65%-67%.
Sustained Momentum in AI Supply Chain
The transmission of positive momentum through the AI industry chain continues, with domestic chip and server manufacturers reporting high growth in preliminary earnings. Recently, some companies in the AI compute supply chain released preliminary first-half results, all showing rapid growth. On the chip side, Hygon Information estimated H1 2026 revenue of RMB 8.5-9.3 billion, a year-over-year increase of 55.56%-70.20%, with net profit attributable to shareholders of RMB 1.7-1.83 billion, up 41.50%-52.32% year-over-year, or 74.27%-84.71% after excluding share-based payment effects. This indicates that domestic CPU+DCU solutions are benefiting from the push for self-reliance in domestic compute and AI infrastructure construction. On the server and interconnect side, Inspur Information estimated H1 2026 net profit attributable to shareholders of RMB 2.6-3.1 billion, a year-over-year increase of 226%-288%, with Q2 single-quarter net profit of RMB 1.995-2.495 billion, surging 494%-643% year-over-year. This shows that the volume ramp of AI servers and super-nodes has begun to significantly boost profitability. Foxconn Industrial Internet estimated H1 2026 net profit attributable to shareholders of RMB 23.4-24.4 billion, up 93%-101% year-over-year, with AI server revenue from cloud service providers growing over 230% year-over-year, and shipments of data center switches above 800G increasing 1.4 times year-over-year. This indicates that AI compute investment is spilling over from pure server procurement to high-speed interconnect infrastructure.
Key Takeaways
The momentum transmission through the "foundry - advanced packaging - domestic chips - AI servers - network interconnect" AI hardware chain remains strong.
In summary, the release of Kimi K3 marks the beginning of domestic open-weight models competing for global capability boundaries and commercial pricing power. Subsequent focus will be on community deployment scale, API call volume, and paid conversion following the full weight release. TSMC's earnings beat and raised guidance, coupled with similarly strong preliminary results from downstream chip, server, and switch companies, underscore the sustained positive momentum transmission through the AI hardware chain. This ultimately manifests in continuous model iteration and application deployment. The firm maintains a positive outlook on segments including GPU, CPU, storage, high-speed networking, and compute leasing services.
Risk warnings include accounts receivable bad debt risk, intensifying industry competition, and impacts from changes in the international environment.
Comments