Goldman Sachs has released a research report maintaining a "Buy" rating on Xiaomi-W (01810) with a 12-month target price of HK$39. The report highlights that Xiaomi's MiMo model began publicly livestreaming its MiMo-V2.6 RL (reinforcement learning) process several days ago, with both the V2.6-Flash and V2.6-Pro variants having stopped after 30 training steps.
Goldman Sachs notes this marks the first time a major AI laboratory has publicly broadcast real-time post-training telemetry data, and the RL training has led to significant improvements in long-horizon agent performance. Within 30 training steps, MiMo-V2.6-Flash/Pro achieved top scores of 67.86/72.57 points (avg@3) on the DeepSWE benchmark.
The bank points out that MiMo-V2.6-Flash appears to offer more favorable RL investment returns, achieving a DeepSWE score improvement of over 19 points with a total expenditure of US$850,000 (US$10 per million tokens or US$2.85 per second) over 3.5 days. In comparison, MiMo-V2.6-Pro spent US$2.62 million (US$35 per million tokens or US$5.71 per second, 2-3 times higher than Flash) to secure a DeepSWE score improvement of over 14 points.
Goldman Sachs estimates that the RL training for MiMo-V2.6-Flash/Pro is based on clusters of approximately 4,000/8,000 H200-equivalent GPUs. The mid-course interventions shown in the livestream, including infrastructure error reruns, bad-pattern datasets, zero-gradient tasks, and GPU out-of-memory issues, suggest that RL bottlenecks have shifted toward engineering optimization.
The bank expects that pricing for the MiMo-V2.6 series upon its official release will attract attention. V2.6 will solidify Xiaomi's focus on agentic coding and multimodal integration, offering higher API pricing potential, although Xiaomi may continue to disrupt market pricing structures alongside DeepSeek.
Goldman Sachs also anticipates that MiMo-V3, which adopts the new HySparse architecture, could launch within a few months, potentially around early 2027. Its sparse ratio may increase from the 7:1 seen in V2/V2.5 to a more aggressive 11:1, further advancing the efficiency frontier and reducing inference costs.
Comments