Harness is Key to AI Agent Performance! Wall Street Tests Qwen Office and Workbuddy
I'm LongbridgeAI, I can summarize articles.Jefferies' real-world tests show that Alibaba's Qwen Office topped the rankings in AI Agent performance, while Claude Opus 5, the model with the strongest capabilities, ranked only second. The core conclusion points out that the key determinant of an Agent's quality is not the intelligence of the underlying model, but the "Harness" (management mechanism). Harness covers six layers, including instructions, context, and tools, effectively bridging the gap between models. The same model paired with different Harnesses shows significant performance differences, proving the importance of management processes for AI execution
Analysts at Jefferies, a Wall Street investment bank, recently conducted real-world tests on eight mainstream Chinese and American AI Agents to see which ones could truly get the job done.
The results were somewhat surprising: The winner was Alibaba's Qwen Office, even though the intelligence score of its underlying model ranked only fourth; Claude Opus 5, which has the top-ranked model capability, saw its Agent product, Claude Cowork, take second place; Another counterintuitive result was that Workbuddy, which has been very popular recently, had the lowest implied Harness score among Chinese Agents in this round of testing.
Jefferies' core conclusion can be summarized in one sentence: What determines whether an AI Agent is useful is often not the model behind it; a well-designed Harness can make up for the gap in model intelligence.
What is Harness: Everything in an Agent Except the Model
The workflow of an AI Agent can be split into two halves: one half is the model, responsible for "thinking"; the other half is the harness, responsible for "managing."
Jefferies broke down the harness into six layers, each managing things that the model cannot handle on its own:
-
Instructions: Telling the Agent what the task is and where the red lines are
-
Context: Background information the Agent can see while working
-
Tools: What the Agent can use to complete the task
-
Boundaries: Defining the scope of actions the Agent can take
-
Feedback: Telling the Agent where it went wrong, allowing it to self-correct and re-plan
-
Governance: Enabling organizations to manage a batch of Agents
These six layers correspond exactly to what a company does after hiring an employee—assigning tasks, providing information, giving tools, defining permissions, providing feedback, and managing. The model is the new smart hire, and the harness is the management mechanism. No matter how smart the person is, if they are thrown into a company without management, they won't do a good job.

Let's look at a set of controlled variable evidence—keeping the model unchanged and only swapping the harness to see how the Agent's performance changes.
Result: With the same Claude Opus 4.6, applying different harnesses resulted in scores on the Terminal-Bench 2.0 benchmark ranging from 58.0% to 76.4%, a difference of 18.4 percentage points between the highest and lowest. Similarly, for Gemini 3 Pro, only changing the harness resulted in scores ranging from 56.0% to 69.4%, a difference of 13.4 percentage points.

Real-World Test of Eight Chinese and American Agents
For this real-world test, Jefferies selected eight Chinese and American Agents. Three from the US: Anthropic's Claude Cowork, OpenAI's Codex, and Google's Gemini Spark. Five from China: Tencent's Workbuddy, Alibaba's Qwen Office, ByteDance's Doubao, Moonshot AI's Kimi Work, and MiniMax's Code.
The test consisted of five tasks, each corresponding to a type of work that genuinely exists in enterprises.
First, multi-file retrieval. The Agent was given a bunch of files and asked to write a one-page summary of Kingdee Software's 2025 annual report, covering revenue, growth, gross margin, net profit, 2026 guidance, and AI progress. The requirements were strict: each fact must cite which file it came from; if numbers conflicted across different files, the conflict must be pointed out and the more credible source explained; quietly smoothing over contradictory numbers was not allowed.
Second, autonomous online research. The Agent was asked to go online and compare the revenue growth rates of Microsoft, Meta, and Google for the most recent quarter, as well as their capital expenditure guidance for 2026. It had to use primary sources, produce a table, and clearly distinguish between officially reported figures and its own estimates.
Third, browser control. The Agent was asked to use a real desktop browser to navigate from the OpenAI official website to the news page, find the latest article, read the full text, and generate a Word file containing the title, date, category, link, a 150-to-200-word summary, and three key points. Using APIs or crawlers to take shortcuts was not allowed; it had to actually click through using the browser.
Fourth, creating a PPT. Given a document, the Agent was asked to create a 5-slide English PPT with a narrative arc. The first slide had to include a chart using data from the document, and fabricating data was not allowed.
Fifth, generating a marketing poster. Given a photo of a shoe as reference, the Agent was asked to create an original poster for a fictional basketball shoe brand, removing all Nike and Jordan logos, not copying trademarked slogans, in a 3:4 vertical format suitable for posting on Xiaohongshu.
These five tasks, ranging from reading files and researching to operating software, creating PPTs, and designing images, almost cover the most common types of work in enterprises.
The scoring method was also sophisticated. Each task was broken down into 5 dimensions, tailored to the specific task rather than using a generic template. Each dimension was scored from 0 to 20, making a perfect score of 100 per task. The total score was the average of the five tasks.

Results: The Weakest Model Took First Place
The top three were: Qwen Office with 95 points, Claude Cowork with 94 points, and Codex with 92 points. Following them were Kimi Work with 86 points, Doubao with 77 points, MiniMax Code with 71 points, while Workbuddy and Google's Gemini Spark tied for last place, both with 66 points.

It is worth mentioning Qwen Office's "comeback." Its underlying model, Qwen 3.8 Max, had an intelligence score of only 56. In contrast, Claude Cowork is backed by Claude Opus 5 with a score of 61, and Codex is backed by GPT-5.6 Sol with a score of 59. Despite a model intelligence gap of five or six points, the Agent's total score was actually one or two points higher.
This also demonstrates that Harness advantages are sufficient to offset gaps in model intelligence.
To verify this, Jefferies performed a breakdown: splitting Agent capabilities into Model and Harness components, assigning a 60% weight to the Model and 40% to the Harness, and reverse-engineering the "implied Harness score" for each product based on the total Agent score. The result showed that Qwen Office had the highest Harness score, surpassing American counterparts Claude Cowork and Codex.

The real-world test also revealed some capability differences between Chinese and American Agents.
The fifth task, marketing poster creation, was a Waterloo for American Agents. Both Claude Cowork and Gemini Spark failed here, while Chinese Agents like Qwen Office and Doubao performed better.
The second task, browser control, was a weakness for Chinese Agents. Workbuddy and MiniMax Code stumbled here, while the three American Agents performed more steadily.
There were also differences in speed. On the American side, Gemini Spark was the fastest, followed by Codex, with Claude Cowork being the slowest. On the Chinese side, Workbuddy was the fastest, Doubao and MiniMax were comparable, while Kimi Work and Qwen Office were slower. Additionally, Claude Cowork had the best "taste," with the most refined PPT layout and color schemes, but this is attributed to model capability, not the Harness.
Price was another area of dominance. Qwen 3.8 Max, behind Qwen Office, has an API price of approximately $1.1 per million tokens; GPT-5.6 Sol is $4.4, and Opus 5 is $3.9.
Workbuddy's Misalignment
Another subject worthy of close examination in the report is Tencent's Workbuddy.
Its data is impressive: monthly visits in June were about 21 million, ranking first among similar Agents in China. However, in this real-world test, Workbuddy's implied Harness score was the lowest among Chinese Agents.
How to explain this misalignment between ranking first in traffic and last in Harness?
Jefferies gave three reasons.
First, Workbuddy is deeply embedded in Tencent's own ecosystem, with entry points in Tencent Docs, IMA, and WeCom all under its control.
Second, it follows a model-agnostic route, allowing users to freely switch between open-source models like Kimi K3, DeepSeek V4, and GLM 5.2 within the harness, using others' strong models to compensate for its own shortcomings.
Third, Tencent's marketing investment is aggressive.
These three factors have no direct relationship with how well the Harness engineering is done.
The report's judgment is: In the short term, Workbuddy wins on entry points and distribution; in the long term, the user base of 21 million will accumulate massive amounts of real usage data, which in turn can help Tencent improve its Harness and even train its own models.
At this stage in the Chinese market, distribution capability has temporarily run ahead of Harness engineering. But what ultimately determines how far an Agent can go is still the Harness.
China Has Advantages in Harness, But Also Unavoidable Shortcomings
The report compared the Harness strategies of China and the US: China leads in several structural areas, but also has several shortcomings holding it back.
First, the advantages, four of them.
Super App Integration. American Agents generally connect to tools via connectors. China is different; many Agents are directly embedded in platforms like WeCom, Lark, and DingTalk, fully integrating data, distribution, and monetization in one line. This is a unique foundation for Chinese vendors.
Model-Agnostic Harness. Tencent's Workbuddy and ByteDance's Trae follow this path—the harness does not come with its own model, allowing users to switch between third-party models like Kimi K3, DeepSeek V4, and GLM 5.2 depending on task difficulty, latency, and cost. Using others' strong models to compensate for their own shortcomings.
Low Token Prices. Export controls have forced Chinese labs to adopt efficient MoE architectures and inference optimization. Coupled with fierce domestic competition, the average token price for Chinese models is 70% to 80% lower than in the US. Cheap tokens make agent workflows that consume large amounts of inference feasible, accelerating enterprise adoption.
Fast Iteration Speed. The quality of a Harness is tested through repeated failures in real scenarios. Chinese vendors have large user bases and encounter many failure cases, leading to faster iteration.
Now, the shortcomings, also four of them.
Model Gap Still Exists. Harness-Bench data shows that strong models have higher average scores and smaller variance across Harnesses; weak models are more dependent on the Harness. Chinese models overall still lag behind the US, requiring superior Harnesses to boost Agent capabilities.
Low Product Pricing. The Chinese enterprise software market has always struggled with monetization. SMEs are price-sensitive, while large state-owned enterprises prefer customized projects and are less inclined towards subscription models. This drags down not just profits, but also the motivation to continuously invest in Harness engineering.
Difficulty in Going Global. Workbuddy and Qoder are growing rapidly in domestic users, but replicating this success overseas is difficult—Chinese Agents are optimized for local ecosystems and user habits, while American Harnesses benefit from a globally standardized software stack.
Computing Power Bottleneck. Agent workflows consume much more inference computing power than ordinary conversations. The shortage of high-end computing power for Chinese vendors may lead to service instability, which in turn limits monetization capabilities.
Why Harness is Becoming Increasingly Important
Finally, let's mention the value of Harness, mainly in three aspects:
Stickiness. Model capabilities are becoming increasingly homogeneous, but harnesses are not. Customer work habits, conversation history, memory, connectors, skills, and automation all reside in the harness. The longer it is used, the higher the cost of switching. Models can be changed anytime, but the harness is not easily replaced.
Data Flywheel. Every time an Agent runs, it generates a trace—what tools were called, what failed, and how humans corrected it. This data can feed into next-generation reinforcement learning and also be used to improve the harness itself. More users mean more data, better Agents, and even more users. This is a positive cycle.
Willingness to Pay. The report pointed out one thing: Enterprises and consumers are not buying "intelligence"; they do not have the ability to develop applications themselves. They are buying a complete Agent product packaged as "harness + model." This is why products like Claude Code, Claude Cowork, and Workbuddy are seeing rapid user growth—customers are paying for the harness, with the model being incidental.
For a company, the key to AI implementation lies in the engineering layer, not the model layer. The real opportunity lies in "doing the harness well"—managing the six aspects of instructions, context, tools, boundaries, feedback, and governance is more likely to produce implementable solutions than chasing the strongest model.
This article comes from the WeChat Official Account " AI Native Lab" , continuously dissecting real AI implementation cases and sharing enterprise AI practices and methodologies.
Risk Warning and Disclaimer
The market involves risks, and investment should be approached with caution. This article does not constitute personal investment advice, nor does it take into account the specific investment objectives, financial status, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Investment decisions made based on this content are the sole responsibility of the investor.
