The "Battle of the Four Powers" in the Aggregation Strategy for Office Agents Officially Begins
I'm LongbridgeAI, I can summarize articles.Leading domestic tech giants such as Alibaba, Baidu, Tencent, and ByteDance are driving office agents toward a model aggregation approach. Qianwen Office has integrated Z.AI's GLM-5.3 and DeepSeek V4 Pro, while Baidu's Kuku AI and Tencent's WorkBuddy also incorporate multiple third-party large language models. The competitive logic is shifting from comparing the strength of individual models to building workbenches that include multiple models, aiming to accumulate real-world workflow data and refine systems, while simultaneously addressing task routing challenges
Office agents have officially entered a fierce battle over model aggregation.
On the evening of August 14, Qianwen Office announced the launch of two models, Z.AI's GLM-5.3 and DeepSeek V4 Pro. Users can directly select and use them from the "Frontier Models" tier on the product's homepage.
On the same day, Baidu Wenku's general-purpose agent GenFlow officially adopted the Chinese name "Kuku AI" and launched a standalone desktop client. Its model pool also includes external models such as DeepSeek and Z.AI's GLM series.
Similarly, Tencent's WorkBuddy and ByteDance's TRAE Work have already integrated multiple third-party large language models.
From Tencent and ByteDance to Alibaba and Baidu, leading domestic manufacturers' office agents are unanimously moving toward a model aggregation model.
In the past, competition among large language model products centered on "whose model is stronger." In the office agent stage, the competitive logic has begun to change: first, integrate different models into the same workbench to encourage more users to entrust their real work to the platform.
Only when users truly engage with the platform, allowing real workflows such as writing, spreadsheet analysis, and PPT creation to flow continuously into the agent, will the platform have the opportunity to accumulate feedback on task decomposition, tool invocation, failure recovery, and user corrections. This feedback loop, in turn, refines the Harness, creating a data flywheel.
However, the next challenge inevitably arises: which model should be called for a given task? Routing has thus become a key issue that the aggregation strategy must continue to address.
Stepping into the Same River
Currently, the number of "Frontier Models" displayed by Qianwen Office is not large. Besides its proprietary Qianwen model, the external models are mainly GLM-5.3 and DeepSeek V4 Pro.
However, Qianwen Office explicitly stated that all three frontier models were quickly integrated after their release, adding, "We will maintain this pace."
This implies that Qianwen Office's built-in model pool will continue to expand in the future.
In comparison, Tencent's WorkBuddy has already formed a larger built-in model pool, including not only Hunyuan Hy3 but also multiple series such as Z.AI's GLM, MiniMax, Kimi, and DeepSeek.
ByteDance's TRAE Work is following a similar direction. Currently visible models in the product include ByteDance's Seed series, Z.AI's GLM, DeepSeek, and Qianwen.
Baidu has just launched the standalone Kuku AI, which also incorporates external models such as DeepSeek and Z.AI's GLM.
The commonality among these four products is quite evident: office agents are no longer binding their capabilities entirely to a single model provider.
The reason is that, for office agents, getting users to truly adopt the technology is currently more important than ensuring every model call is handled by their proprietary model.
When an agent begins searching the web, opening documents, analyzing Excel files, creating PPTs, invoking software, and collaborating with users to revise work, it leaves behind a more complete execution trail: how the model decomposes tasks, which tools are invoked, where failures occur, at which step the user makes corrections, and which final result is accepted.
Real-world workflows are precisely the scarce data resources in the agent era.
Wang Ying, Vice President of Baidu Group and President of the Personal Super Intelligence Business Group, believes that as public data is gradually fully learned by models, new knowledge, experiences, and ideas generated by humans will become increasingly important in the continued advancement of AI. In the future, what truly needs to be jointly established are three core capabilities: general agent capabilities, memory and storage, and personal knowledge and experience.
The richer the model pool, the greater the opportunity to improve completion rates for different tasks and lower the threshold for users to try an agent. Only when more users genuinely entrust their work to the agent can the platform obtain continuous feedback to optimize Harness capabilities such as task decomposition, Context management, tool selection, and failure recovery.
Thus, model aggregation is not merely a platter of model capabilities; it is also a strategy to secure more real-world tasks for the data flywheel.
However, aggregation does not mean that the boundaries between major tech firms have disappeared.
Judging from the default model pools currently visible in the products, although WorkBuddy has accessed multiple external models, it does not include Alibaba's Qianwen or ByteDance's Seed. TRAE Work has integrated the Qianwen model but lacks Tencent's Hunyuan. Kuku AI currently also does not feature Hunyuan, Qianwen, or Seed.
While each company is expanding its model choices, they still maintain their own ecosystem boundaries.
The Impossible Triangle
Integrating more models into the same agent only solves the problem of "whether there is a choice." What truly determines whether the aggregation strategy can succeed is how the selection is made next.
The implementation of large language models has always involved a trade-off resembling an "impossible triangle": it is difficult to maximize performance, speed, and cost simultaneously.
More capable models usually imply higher inference costs. Complex tasks often require longer thinking and execution times to achieve better results. Conversely, blindly pursuing low cost and low latency may sacrifice the quality of task completion.
Routing essentially aims to solve this triangular dilemma: to complete a task, at which stages should more money be spent to gain capability, and at which stages should speed and cost be traded for efficiency.
This also introduces another layer of differentiation in the competition among aggregated office agents.
The ideal state is for users to simply propose a task, while the backend handles model selection based on task difficulty, timeliness requirements, and model costs.
At this stage, what the user sees might be just a single request to "help me finish this report," but the backend may have already invoked different models at various steps.
For individual users, this difference ultimately manifests in the quality of the result, the waiting time, and the number of points consumed. For enterprise clients, it further translates into the unit computing cost per task.
Once office agents enter mass adoption, even minor differences of a few cents or dimes per single call will be constantly amplified with the volume of tasks.
Routing thus becomes a matter of scaled economics. However, from the perspective of the agent's complete execution chain, routing is also a component of the Harness.
From the moment a user proposes a task to its final delivery, the process also involves task decomposition, Context management, tool invocation, model selection, and failure recovery. Routing decides "who to call at this stage," but what truly determines whether a task can be completed stably is whether the entire Harness can coordinate these links effectively.
This is also where the aggregation strategy may truly form barriers to entry. Models themselves can be integrated, and prices may continue to drop with market competition, but finding a better balance among performance, speed, and cost requires a sufficient volume of real-world tasks and scheduling experience accumulated by the platform over the long term.
As WorkBuddy, TRAE Work, Qianwen Office, and Kuku AI hold an increasing number of models, the real differentiator may lie in who calculates more precisely for the same problem: knowing when to use the most powerful model and when it is completely unnecessary.
Risk Warning and Disclaimer
The market carries risks, and investment should be approached with caution. This article does not constitute personal investment advice, nor does it consider the specific investment objectives, financial status, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article align with their specific circumstances. Investment decisions made based on this content are the sole responsibility of the investor.
