On the evening of August 14, Qianwen Office announced the launch of Z.AI's GLM-5.3 and DeepSeek V4 Pro models, allowing users to select them directly from the "Frontier Models" section on the product homepage.
On the same day, Baidu's general-purpose agent GenFlow officially adopted the Chinese name "Kuku AI" and launched a standalone desktop client, with its model pool also including external models like DeepSeek and Z.AI's GLM.
Similarly, Tencent's WorkBuddy and ByteDance's TRAE Work have already integrated multiple third-party large models. From Tencent and ByteDance to Alibaba and Baidu, the office agents of top domestic tech firms are converging on a model aggregation strategy.
Previously, the competition between large model products was about which company's model was stronger. In the office agent stage, the competitive logic has shifted: first, place different models into the same workspace to encourage more users to entrust their real work to the agent.
Only when users actively engage—feeding real workflows like writing, spreadsheet analysis, and PPT creation into the agent—can the platform accumulate feedback on task decomposition, tool invocation, failure recovery, and user corrections. This feedback loop, in turn, refines the Harness and creates a data flywheel.
However, a new challenge emerges: determining which model should be called for a specific task. Routing thus becomes the next problem to solve along the aggregation path.
Entering the Same River
Currently, Qianwen Office's "Frontier Models" section is not extensive, featuring primarily its own Qianwen model alongside Z.AI's GLM-5.3 and DeepSeek V4 Pro. However, Qianwen Office has stated that these three models were quickly integrated after their release, and they plan to maintain this rapid pace of integration going forward.
This suggests that Qianwen Office's built-in model pool will continue to expand. In comparison, Tencent's WorkBuddy already boasts a larger integrated model pool, including not only Hunyuan Hy3 but also models from Z.AI, MiniMax, Kimi, and DeepSeek.
ByteDance's TRAE Work is following a similar trajectory, with its product currently featuring models from ByteDance's Seed series, Z.AI's GLM, DeepSeek, and Qianwen. Baidu's newly launched standalone Kuku AI has also introduced external models like DeepSeek and Z.AI's GLM.
A clear commonality across these four products is evident: office agents are no longer tying their capabilities exclusively to a single model. The reason is that, for office agents at this stage, getting users to genuinely use the product is more critical than routing every model call to their own proprietary model.
When an agent searches the web, opens documents, analyzes Excel files, creates PPTs, calls software, and iterates on tasks with users, it leaves behind a more complete execution trail: how the model decomposes tasks, which tools it invokes, where it fails, where the user corrects it, and which final result is accepted. Real-world workflows are the scarce data resources of the agent era.
Wang Ying, Baidu's Vice President and President of the Personal Super Intelligence Business Group, believes that as public data is fully learned by models, new knowledge, experiences, and ideas generated by humans will become increasingly important for AI's continued advancement. The future requires building three core capabilities: general agent abilities, memory and storage, and personal knowledge and experience.
A richer model pool increases the likelihood of improving task completion rates and lowers the barrier for users to try an agent. Only when more users entrust their work to the agent can the platform receive continuous feedback, which it then uses to optimize its Harness capabilities—such as task decomposition, context management, tool selection, and failure recovery.
Model aggregation, therefore, is not just a collection of model capabilities but also a strategy to attract more real tasks for the data flywheel. However, aggregation does not mean the boundaries between tech giants have disappeared.
From the current default model pools visible in products, WorkBuddy, despite integrating several external models, does not include Alibaba's Qianwen or ByteDance's Seed. TRAE Work has integrated the Qianwen model but lacks Tencent's Hunyuan. Kuku AI currently does not feature Hunyuan, Qianwen, or Seed. While each company is expanding its model selection, they are still maintaining their own ecosystem boundaries.
The Impossible Triangle
Loading more models into a single agent only solves the problem of having options. The real question that determines whether the aggregation path will succeed is how to choose the right model for each task.
A trade-off resembling an "impossible triangle" has always existed in large model deployment: it's difficult to simultaneously maximize effectiveness, speed, and cost. Stronger models typically come with higher inference costs, and complex tasks often require longer thinking and execution time for better results. Conversely, pursuing only low cost and low latency can sacrifice task completion quality.
Routing essentially aims to solve this triangular problem: to complete a task, where should you spend more money for better capability, and where should you prioritize speed and cost for efficiency? This introduces another layer of differentiation in the competition among aggregated office agents.
The ideal scenario is for users to simply propose a task, while the backend selects the model based on task difficulty, time requirements, and model cost. At this point, the user might only see a request to "help me complete this report," but the backend could be calling different models for different steps.
For individual users, this difference ultimately manifests in the quality of the result, waiting time, and points consumed. For enterprise clients, it translates into the unit computing cost per task. Once office agents are used at scale, even minor cost differences of a few cents per call can be amplified by the volume of tasks.
Routing thus becomes a large-scale economic equation. However, within the complete execution chain of an agent, routing is just one part of the Harness. From a user's initial request to the final delivery, a task involves task decomposition, context management, tool invocation, model selection, and failure recovery. Routing determines "which model to call for this step," but the overall Harness determines whether the task can be reliably completed by coordinating these elements.
This is where the aggregation path can create a genuine barrier to entry. Models themselves can be integrated, and prices may fall with market competition, but finding the optimal balance between effectiveness, speed, and cost requires sufficient real tasks and the scheduling experience accumulated by the platform over time. As WorkBuddy, TRAE Work, Qianwen Office, and Kuku AI accumulate more models, the true differentiator may be who can solve the same problem more precisely: knowing when to use the strongest model and when it is simply unnecessary.
Comments