According to data compiled from OpenRouter, global AI large model call volume for the week ending August 9 reached 69 trillion tokens, a 21.48% increase from the prior week. Among the models on the leaderboard, Chinese AI large models collectively recorded 34.25 trillion tokens in weekly call volume, up 21.76% week-over-week. Over the same period, U.S. AI large models saw a weekly call volume of 9.17 trillion tokens, a 109.36% increase. Chinese large models have now outpaced their U.S. counterparts in weekly call volume for fifteen consecutive weeks, maintaining the top global position.
The top four spots on the global call volume leaderboard were all held by Chinese AI models. Topping the list was DeepSeek-V4-Flash-0731 (the official version of DeepSeek-V4-Flash), which achieved a weekly call volume of 8.83 trillion tokens, an impressive 570% increase from the previous week. Tencent Hy3 ranked second with a weekly call volume of 8.05 trillion tokens, up 67% week-over-week. DeepSeek-V4-Flash-0423 (the preview version of DeepSeek-V4-Flash) fell to third place, with 5.88 trillion tokens in weekly call volume, a 19% decline. Xiaomi MiMo-V2.5 dropped to fourth, registering 5.39 trillion tokens in weekly call volume, a 14% decrease.
On July 31, DeepSeek announced the public beta launch of the API for the official version of DeepSeek-V4-Flash. The official version maintains the same model structure and size as the preview version but has undergone a retraining process. According to the official update log, the official V4-Flash version demonstrates significantly enhanced Agent capabilities, with benchmark test results far exceeding the preview version. Data from Artificial Analysis shows that the average cost per single task completed by the official V4-Flash version is approximately $0.03, which is only 1/100th of the cost of the flagship model Claude Fable 5 from Anthropic, making it the most cost-effective model among current mainstream global models.
OpenAI's GPT-5.6 Luna entered the top five, with a weekly call volume of 4.43 trillion tokens, a 128% increase from the prior week. On July 10, OpenAI officially launched the GPT-5.6 series models publicly. GPT-5.6 Luna is the fastest and most cost-effective model in this series, designed for large-scale, high-frequency task processing, supporting tool calls and multi-step workflows. Less than three weeks after the series launch, OpenAI announced pricing and performance optimizations for the GPT-5.6 series to enhance the cost-effectiveness of AI applications for enterprise users. The price of the GPT-5.6 Luna model was reduced by 80%.
Google Gemini 3.6 Flash entered the leaderboard for the first time, ranking ninth with a weekly call volume of 2.33 trillion tokens, a 446% increase week-over-week. Gemini 3.6 Flash was released on July 22. Data from the third-party evaluation agency Artificial Analysis Index indicates that Gemini 3.6 Flash consumes an average of 17% fewer tokens during output compared to its predecessor, Gemini 3.5 Flash. Its price per million input tokens remains at $1.50, while the output price has been reduced from $9.00 in the previous generation to $7.50, a decrease of approximately 16.7%.
Notably, MiniMax M3, which ranked seventh the previous week, and Step 3.7 Flash from Stepfun, which ranked ninth, have both dropped off the leaderboard.
Disclaimer: The content and data in this article are for reference only and do not constitute investment advice. Please verify before use. Trading based on this information is at your own risk.
Comments