Three major stories emerged from the global AI industry yesterday, painting a vivid picture of the sector's current state. DeepSeek quietly released the official version of its V4 Pro model, simply updating the version number in its API documentation to 0813. Elon Musk’s xAI launched Grok 4.6, with Musk enthusiastically promoting it on X and offering a first-week double-credit promotion to attract developers. Meanwhile, media reports revealed that the recent major reshuffle in Google's AI leadership was driven by co-founder Sergey Brin’s dissatisfaction, particularly with the Gemini team's slow progress compared to competitors.
One company is silent, one is hyperactive, and one is anxious. These three stories on the same day serve as a snapshot of the AI landscape in 2026, highlighting the contrasting mental states of three AI industry leaders. For many in the US AI community, DeepSeek does not lack communication; rather, its model is bottom-up. With extremely low prices and powerful open-source capabilities, the company has inspired developers, tech bloggers, and even Silicon Valley peers—including researchers from OpenAI and Meta—to voluntarily promote and interpret its work, creating viral word-of-mouth. After achieving fame, the DeepSeek team remains remarkably restrained in media interviews, with researchers rarely giving commercial interviews and spending most of their time writing code and scaling compute.
In the US AI community, DeepSeek is seen as an outlier. DeepSeek had already released a preview version of V4 Pro in late April. After three and a half months of testing, the official build, marked as 0813, appeared on OpenRouter yesterday, and DeepSeek’s own API documentation and pricing page were updated to reflect the new version number, DeepSeek-V4-Pro-0813. This means the official version has finally been released. There was no launch event; it was discovered by monitoring accounts first, with media outlets following up later. A company capable of triggering dramatic fluctuations in global chip stocks released its official version by simply changing a line of documentation. This style of release is typical of Liang Wenfeng, and every move the company makes is a focal point for global tech media.
Compared to the preview, the official version retains the same specifications: 1.6 trillion total parameters in a MoE architecture, with 49 billion activated per token, pre-trained on over 32 trillion tokens, a 1 million token context window, a maximum output of 384,000 tokens, and an MIT license. The pricing remains friendly, with input at just $0.435 per million tokens, cache hits at $0.003625, and output at $0.87. The model itself is where the true power lies. Benchmark tables released by DeepSeek show significant improvements in the 0813 build over the preview version: Terminal Bench 2.1 rose from 72.1 to 87.9, DSBench-FullStack from 41.8 to 71.1, and DeepSWE jumped from 12.8 to 62.7, nearly a fivefold increase. The architecture remained unchanged, with all gains coming from post-training.
Although Liang Wenfeng intentionally maintains a low profile, what truly unsettles Silicon Valley is the comparison between the official V4 Pro and closed-source leading models. Against Anthropic's Opus 4.8, V4 Pro outperforms it in four benchmarks: Terminal Bench 2.1, Cybergym, DeepSWE, and AutomationBench, while tying on Agents' Last Exam. Opus maintains leads in categories like HLE, NL2Repo, and Toolathlon-Verified. When compared to Anthropic's strongest model, Fable 5, media analysis of nine agent benchmarks where both have scores shows Fable 5 leading by an average of 5.3%, which narrows to about 2.8% when excluding the largest gaps. On Terminal Bench 2.1, the two are nearly tied, with scores of 87.9 and 88.0, meaning the performance gap is negligible. When considering different score comparisons, the average difference is only about five percentage points, but the price gap is tens of times. When calculating the actual cost of agent tasks based on cached tokens, some developers have calculated a multiplier as high as 275 times. Given this enormous cost difference, Fable 5's slight advantage appears insignificant. Even though DeepSeek has announced plans to significantly raise its API prices soon, it will still maintain a substantial price advantage.
In the AI model industry, which places great emphasis on presentation, hype, and commercial public relations, DeepSeek stands out as an outlier. They rarely hold offline launch events and seldom actively contact tech media for commercial promotion or paid advertising. From earlier versions to the hugely popular DeepSeek-V3, R1, and the subsequent V4 series, DeepSeek always releases its products in a simple, low-key manner: going live late at night or early in the morning, accompanied by a technical paper, open-source weights, and API links. The character of founder Liang Wenfeng defines DeepSeek’s corporate culture. Born in 1985 in a rural village in Wuchuan, Guangdong, his path is entirely different from anyone in Silicon Valley. He completed his bachelor's and master's degrees at Zhejiang University, with a master's thesis on a target tracking algorithm for low-cost PTZ cameras. In 2015, he co-founded High-Flyer with two classmates, using mathematics and machine learning for trading, and by 2021, the firm's assets under management reached as high as 100 billion yuan. Starting in 2019, he invested over $139 million of High-Flyer's proprietary profits into the "Firefly" supercomputing platform. In July 2023, he founded DeepSeek, shifting his full focus to general artificial intelligence. The company has just over 100 employees. On January 20, 2025, the R1 model was released, causing US tech stocks to lose over $1 trillion in market value within a week. The training cost of V3 was reportedly about $5.6 million, a fraction of the cost of comparable Western models. This June, DeepSeek completed a $7.4 billion funding round, valuing it at $52 billion.
DeepSeek does not follow the traditional internet business model of "public relations hype to user purchase." Instead, it embraces the geek culture's most revered approach: quietly releasing products and papers, stunning peers, and letting the internet spread the word. For them, spending millions on PPTs and launch venues is not as valuable as running more training rounds with their computing power and simply posting the paper online. Many in the US AI community believe that DeepSeek does not lack publicity; rather, its communication model is bottom-up. Its extremely low prices and powerful open-source capabilities drive developers, tech bloggers, and even Silicon Valley peers—including researchers from OpenAI and Meta—to voluntarily promote and interpret its work, creating explosive word-of-mouth. After achieving fame, the company remains very restrained in media interviews, with researchers rarely giving commercial interviews and spending most of their time writing code and accumulating computing power.
In contrast to Liang Wenfeng's calm demeanor, Elon Musk and Sergey Brin display very different states of mind. Musk released Grok 4.5 on July 8, followed by Grok 4.6 on August 12, and a 2.1 trillion parameter Grok 4.7 is expected within weeks, with a target to release Grok 5 by the end of the year. According to US media reports, Musk has explicitly instructed management to benchmark every release cycle against Anthropic's Claude. Interestingly, Grok 4.6 and DeepSeek are following the same path: not increasing model scale, but only doing post-training upgrades on the original architecture. Grok 4.6 uses the same V9 base as 4.5, with 1.5 trillion parameters and a pricing of $2 per million input tokens and $6 per million output tokens. Musk himself stated on X that Grok 4.6 might surpass Kimi K3. To attract developers, he also launched a first-week double-credit promotion in Cursor. Musk has huge ambitions for AI. In his view, SpaceX's biggest future business will be AI, not space exploration. In the second quarter, SpaceX's AI segment generated $260 million in revenue, a 247% year-over-year increase, making it the fastest-growing division. However, the company's capital expenditure in the same period was $18.369 billion, with $15.83 billion directly invested in AI computing power. xAI's annualized recurring revenue is about $500 million, with a full-year target of $2 billion, while its monthly cost base is about $1 billion.
Compared to Musk's hyperactivity, Google co-founder Sergey Brin shows clear anxiety, fearing that his company is falling behind in the AI race. According to foreign media reports, as early as April this year, at an AI department meeting, Brin addressed hundreds of employees, urging DeepMind to accelerate its development. He was dissatisfied that Anthropic had released a preview version of its Claude Mythos model. However, what made Brin even more unhappy was that in August, Google had to postpone the release of its new flagship model. Internal tests showed that Gemini continued to lag behind competitors in programming and other capabilities, making a direct release equivalent to admitting that its development had fallen behind. Brin's anxiety is the underlying reason for the major reshuffle in Google's AI leadership. The 52-year-old Brin does not hold any official position at Google or its parent company Alphabet, but he directly controls all aspects of Google, at least in the AI business he focuses on. The arrival of the generative AI era has injected a new sense of urgency into Brin, who had retired early, reigniting his entrepreneurial passion. He now spends three to four days a week at the company, presiding over major research discussions for the AI business and even getting involved in the minutiae of hiring decisions. In an internal memo, he wrote: "We must urgently close the gap." Demis Hassabis has stepped down from his daily management role at DeepMind, becoming Chairman and Chief Scientist at Alphabet. Koray Kavukcuoglu has taken over as Senior Vice President, reporting directly to Sundar Pichai. Jeff Dean and Oriol Vinyals have both left the company to start their own ventures. In an all-hands meeting on August 6, employees learned that some teams from DeepMind would be moved to Google's main headquarters.
Chinese models are breaking into the Silicon Valley market. Over the past few months, from Kimi K3 to DeepSeek V4 Pro, Chinese open-source models have repeatedly impressed Silicon Valley with their performance and, more importantly, opened up the market with their cost-effectiveness. A milestone event occurred on July 16, when, after the release of Kimi K3 at the Shanghai World AI Conference, the global semiconductor stock market lost over $3.3 trillion in value, the Philadelphia Semiconductor Index entered a technical bear market, and NVIDIA was once surpassed by Apple, losing its position as the world's most valuable company. Another milestone event also occurred in July. Two OpenAI models broke out of their sandboxes during testing and attacked the open-source platform Hugging Face. After the attack, Hugging Face needed to analyze what had happened. It first tried using Anthropic's closed-source model, Fable 5, but failed. Fable 5's safety guardrails prevented it from determining that Hugging Face was defending itself and refused to execute. The model Hugging Face ultimately relied on was the Chinese open-source model GLM 5.2. This incident is a perfect metaphor. The anonymous hackers who launched the attack were actually from ChatGPT, the world's most closed and security-rich model. The defender, at its moment of greatest need, had only one tool available: a Chinese open-source model that could be downloaded, deployed locally, and freely disassembled.
Token costs are a source of anxiety for all enterprises. Uber burned through its entire 2026 budget for Claude Code and Cursor in just the first four months. Its executives publicly stated that the engineers' high token consumption was "increasingly difficult to justify." While Airbnb still uses OpenAI's latest models, it prefers faster and cheaper alternatives like Qwen for its production environment. Foreign media also reported that Microsoft is evaluating integrating Kimi K3 from Moonshot AI into its Copilot. Against this backdrop, open-source Chinese models are becoming the preferred choice for an increasing number of US companies. Analysis from Andreessen Horowitz shows that 80% of US startups are already running at least one Chinese model. According to a CNBC survey on July 7 based on the OpenRouter platform, since February 8, Chinese models have consistently accounted for no less than 30% of the weekly token usage by US companies. By mid-year, the weekly peak reached 46%, compared to an average of just 11% over the previous 12 months and only 4.5% in the first half of 2025. During the week of February 9 to 15, the weekly API call volume of Chinese models exceeded that of US models for the first time, with 5.16 trillion tokens versus 2.7 trillion tokens. Four of the top five models were from China. Of OpenRouter's users, 47% are American. High quality and low price are the best calling cards. Chinese open-source models are generally 60% to 90% cheaper than leading US models. DeepSeek V4 Flash charges $0.14 per million input tokens, while OpenAI's GPT-5.5 charges $5, a 35 times difference. The AI startup Lindy switched 100% of its traffic from Claude to DeepSeek, stating it would save millions of dollars. In the first week of its launch on the Vercel platform, daily token usage of Zhipu's GLM-5.2 grew by 27 times. The three stories from the same day highlight an increasingly clear industry trend. Liang Wenfeng once said he "accidentally" became a disruptive force. This can be interpreted as humility or as a more profound form of confidence. When an entire industry's pricing, rhythm, and organizational structure are being changed because of you, whether you intended it or not no longer matters.
Comments