AI Agent Efficiency Rests on Harness, Not Just the Brain, Wall Street Tests Reveal Surprising Winners

Deep News
Aug 18

Jefferies analysts recently conducted a hands-on evaluation of eight leading AI agents from the US and China to determine which ones can genuinely deliver on complex tasks. The results overturned expectations, with Alibaba's Qwen Office taking the top spot despite its underlying model ranking only fourth in raw intelligence. Meanwhile, the model with the highest intelligence score, Claude Opus 5, powered its agent, Claude Cowork, to second place. In another counter-intuitive twist, the popular Workbuddy scored the lowest implied Harness rating among the Chinese agents in this test.

The core conclusion from Jefferies can be summarized in one sentence: the factor determining whether an AI agent performs well is often not the model behind it, and a well-executed Harness can effectively bridge the gap in model intelligence.

What is a Harness: Everything in an Agent Except the Model

The workflow of an AI agent can be divided into two halves. One half is the model, which is responsible for thinking; the other half is the harness, which is responsible for management. Jefferies broke down the harness into six distinct layers, each managing aspects the model cannot control on its own. These layers include instructions, which tell the agent what the task is and where the boundaries lie; context, which provides the background information the agent can see; tools, which define what the agent can use to complete the task; boundaries, which set the limits of the agent's action scope; feedback, which tells the agent where it went wrong so it can self-correct and re-plan; and governance, which allows an organization to manage a fleet of agents.

These six layers correspond to the steps a company takes when hiring an employee, which are assigning tasks, providing information, giving tools, setting permissions, offering feedback, and managing performance. The model is the smart newcomer, and the harness is the management system. No matter how capable the smart person is, they cannot perform well if they are thrown into a company without management.

Consider a set of controlled experiments where the model remains unchanged and only the harness is swapped to see how agent performance shifts. The results show that the same Claude Opus 4.6, when paired with different harnesses, scored between 58.0% and 76.4% on the Terminal-Bench 2.0 benchmark, a difference of 18.4 percentage points between the highest and lowest. Similarly, Gemini 3 Pro, with only a harness change, saw scores range from 56.0% to 69.4%, a difference of 13.4 percentage points.

Testing Eight Agents from the US and China

For this evaluation, Jefferies selected eight agents from the US and China. From the US, there were three: Anthropic's Claude Cowork, OpenAI's Codex, and Google's Gemini Spark. From China, there were five: Tencent's Workbuddy, Alibaba's Qwen Office, ByteDance's Doubao, Moonshot AI's Kimi Work, and MiniMax's Code. The test consisted of five tasks, each representing a real type of work found in enterprises.

The first task involved multi-file retrieval. Agents were given a set of documents and asked to write a one-page summary of Kingdee Software's 2025 annual report, covering revenue, growth, gross margin, net profit, 2026 guidance, and AI progress. The requirements were strict: every fact had to be cited to its source file, conflicting numbers across different files had to be flagged with an explanation of which was more credible, and agents were prohibited from silently smoothing over contradictory data.

The second task was autonomous online research. Agents had to browse the internet to compare the latest quarterly revenue growth rates of Microsoft, Meta, and Google, along with their 2026 capital expenditure guidance. They were required to use primary sources, produce a table, and clearly distinguish between official numbers and their own estimates. The third task was browser control. Agents had to use a real desktop browser to navigate from the OpenAI homepage to the news section, find the latest article, read it in full, and generate a Word file containing the title, date, category, link, a 150 to 200-word summary, and three key points. They were not allowed to use APIs or scrapers to cut corners and had to genuinely click through the browser.

The fourth task was creating a PowerPoint presentation. Agents were given a file and asked to make a 5-page English PPT with a narrative arc, using the file's data to create a chart on the first page without fabricating any numbers. The fifth task was generating a marketing poster. Agents were given a photo of a shoe as a reference and had to create an original poster for a fictional basketball shoe brand, removing all Nike and Jordan logos, not copying trademarks or slogans, and adhering to a 3:4 vertical format suitable for posting on Xiaohongshu. These five tasks cover the most common types of work in enterprises, from reading files and researching information to operating software, creating presentations, and designing graphics.

The scoring method was also meticulous. Each task was broken down into five dimensions, which were tailored to the specific task rather than a generic template. Each dimension was scored from 0 to 20, with a total of 100 per task, and the final score was the average of all five tasks.

Results: The Agent with the Weakest Model Took First Place

The top three were Qwen Office with 95 points, Claude Cowork with 94 points, and Codex with 92 points. Following them were Kimi Work with 86 points, Doubao with 77 points, MiniMax Code with 71 points, and Workbuddy and Google's Gemini Spark tied for last place with 66 points. Notably, Qwen Office's rise was a surprise. The model behind it, Qwen 3.8 Max, had an intelligence score of only 56. In contrast, Claude Cowork was powered by Claude Opus 5 with a score of 61, and Codex was backed by GPT-5.6 Sol with a score of 59. Despite a five to six-point gap in model scores, Qwen Office's agent score was one to two points higher, demonstrating that a Harness advantage can fully offset a model intelligence gap.

To validate this, Jefferies performed a breakdown, splitting agent capability into model and Harness components with a 60% weight for the model and a 40% weight for the Harness. By using the total agent score, they reverse-engineered each product's implied Harness score. The result showed that Qwen Office had the highest Harness score, surpassing US agents like Claude Cowork and Codex.

The tests also revealed some differences in capabilities between US and Chinese agents. The fifth task, marketing poster creation, was a Waterloo for US agents. Both Claude Cowork and Gemini Spark stumbled here, while Chinese agents like Qwen Office and Doubao performed better. The second task, browser control, was a weakness for Chinese agents. Workbuddy and MiniMax Code struggled, while the three US agents were more stable. There were also differences in speed. Among US agents, Gemini Spark was the fastest, Codex came second, and Claude Cowork was the slowest. Among Chinese agents, Workbuddy was the fastest, Doubao and MiniMax were comparable, and Kimi Work and Qwen Office were slower. Additionally, Claude Cowork had the best taste in terms of PPT layout and color scheme, but this was attributed to model capability rather than the Harness.

Price was another area of stark contrast. The API price for Qwen 3.8 Max, behind Qwen Office, was approximately $1.1 per million tokens, while GPT-5.6 Sol was $4.4 and Opus 5 was $3.9.

The Discrepancy in Workbuddy's Performance

Another focal point of the report was Tencent's Workbuddy. Its data was impressive, with about 21 million monthly visits in June, ranking first among similar Chinese agents. However, in this test, Workbuddy's implied Harness score was the lowest among Chinese agents. How can this discrepancy between being first in traffic and last in Harness be explained? Jefferies offered three reasons.

First, Workbuddy is deeply integrated into Tencent's own ecosystem, with entry points in Tencent Docs, IMA, and WeCom. Second, it follows a model-agnostic approach, allowing users to freely switch between open-source models like Kimi K3, DeepSeek V4, and GLM 5.2 within the harness, using others' strong models to compensate for its own weaknesses. Third, Tencent has invested heavily in marketing. These three factors have no direct bearing on how well the Harness engineering is executed. The report's assessment is that in the short term, Workbuddy wins on entry points and distribution. In the long term, the 21 million user base will accumulate vast amounts of real-world usage data, which can, in turn, help Tencent improve its Harness and even train its own models. At this stage of the Chinese market, distribution capability has temporarily outpaced Harness engineering, but it is the Harness that will ultimately determine how far an agent can go.

China's Advantages and Unavoidable Shortcomings in Harness

The report compared the Harness approaches of China and the US, noting that China leads in several structural areas but also faces several persistent weaknesses. Starting with the advantages, there are four. First is super-app integration. US agents typically connect to tools via connectors, whereas in China, many agents are directly embedded in platforms like WeCom, Feishu, and DingTalk, seamlessly integrating data, distribution, and monetization. This is a unique foundation for Chinese vendors. Second is model-agnostic Harness. Tencent's Workbuddy and ByteDance's Trae follow this path, where the harness itself does not come with a model, and users can switch between third-party models like Kimi K3, DeepSeek V4, and GLM 5.2 based on task difficulty, latency, and cost. This allows them to leverage others' strong models to address their own shortcomings. Third is low token prices. Export controls have forced Chinese labs to pursue efficient MoE architectures and inference optimization, and combined with intense domestic competition, Chinese model token prices are on average 70% to 80% lower than in the US. Cheaper tokens make agent workflows that consume significant inference resources viable, accelerating enterprise adoption. Fourth is fast iteration speed. The quality of a Harness is honed through repeated failures in real-world scenarios. Chinese vendors have large user bases and encounter more failure cases, which paradoxically speeds up their iteration cycles.

Turning to the weaknesses, there are also four. The first is the ongoing model gap. Data from Harness-Bench shows that stronger models have higher average scores and lower variance across different harnesses, while weaker models are more dependent on the Harness. Chinese models still lag behind their US counterparts overall, requiring a more outstanding Harness to prop up agent capabilities. The second is low product pricing. The Chinese enterprise software market has always struggled with monetization; small and medium-sized enterprises are price-sensitive, and large state-owned enterprises prefer custom projects over subscription models. This not only drags down profits but also dampens the motivation for sustained investment in Harness engineering. The third is difficulty in going global. Products like Workbuddy and Qoder are growing quickly among domestic users, but replicating this success overseas is challenging because Chinese agents are optimized for the local ecosystem and user habits, whereas US harnesses benefit from a globally standardized software stack. The fourth is compute bottlenecks. Agent workflows consume far more inference compute than ordinary conversations. The shortage of high-end compute for Chinese vendors could lead to service instability and, in turn, limit their monetization capabilities.

Why the Harness is Becoming Increasingly Important

Finally, the value of the Harness can be summarized in three main aspects. The first is stickiness. Model capabilities are becoming increasingly similar, but the Harness is different. Customer work habits, conversation history, memory, connectors, skills, and automations all reside in the Harness. The longer it is used, the higher the switching cost. Models can be swapped out at any time, but the Harness is not easily replaced. The second is the data flywheel. Every time an agent runs, it generates a trace, recording which tools were called, what failed, and how humans corrected it. This data can feed into the next generation of reinforcement learning and also be used to improve the Harness itself. More users lead to more data, which leads to a better agent, which attracts more users. This is a positive cycle. The third is willingness to pay. The report highlighted a key point: enterprises and consumers are not buying intelligence itself, as they lack the capability to develop their own applications. What they are buying is a complete agent product that packages the Harness and the model together. This is why products like Claude Code, Claude Cowork, and Workbuddy are seeing rapid user growth. Customers pay for the Harness, and the model is merely a byproduct.

For a company, the key to AI implementation lies in the engineering layer, not the model layer. The real opportunity is in perfecting the Harness, which means managing the six aspects of instructions, context, tools, boundaries, feedback, and governance. This is more likely to produce practical, deployable solutions than chasing the most powerful model.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10