BotBeat
...
← Back

> ▌

TencentTencent
RESEARCHTencent2026-07-24

Tencent Releases WorkBuddy Bench: Multi-Model Agentic Coding Leaderboard Shows No Clear Winner

Key Takeaways

  • ▸No single model dominates across all tasks—Claude Opus 4.8 leads on Code and Web, GLM-5.2 on Security, GPT-5.5 on Office efficiency
  • ▸Open-weight models like GLM-5.2 remain competitive, especially in security-focused agentic tasks
  • ▸Token efficiency is disconnected from ranking: GPT-5.5 achieves best-in-class efficiency (6.9k output tokens on Code) while mid-table models spend 3-4x more
Source:
Hacker Newshttps://workbuddybench.com/index.html↗

Summary

Tencent has released WorkBuddy Bench, a comprehensive agentic coding leaderboard that benchmarks multiple AI models across code generation, web automation, office tasks, and security-focused challenges using two different evaluation harnesses (CodeBuddy Code and Claude Code). The benchmark reveals no single dominant model: Claude Opus 4.8 leads five of eight scored columns, GLM-5.2 (an open-weight model) excels in Security tasks, and GPT-5.5 dominates in Office tasks under Claude Code.

The results surface three critical insights: token efficiency is decoupled from performance ranking, with GPT-5.5 achieving top-tier scores using 3-4x fewer output tokens than competitors; performance varies dramatically across harnesses for the same model (GLM-5.2's Web score swings 67.43 vs 60.71); and integration details like cross-turn reasoning passback can meaningfully shift scores. Open-weight models remain competitive, particularly in specialized domains like security, suggesting the agentic coding market retains meaningful diversity rather than consolidating around closed-source solutions.

  • Evaluation methodology heavily influences results—identical models show significant score variations across different harnesses and configurations
  • Integration details matter: enabling cross-turn reasoning for HY-3 boosted its Code score by +1.92 to +3.82 points

Editorial Opinion

Tencent's WorkBuddy Bench fills a critical gap in agentic AI evaluation, providing rare transparency into model specialization rather than false universalism. The finding that open-weight and commercial models remain competitive across different domains is bullish for AI diversity, but the benchmark also exposes a methodological fragility: identical models produce wildly different scores depending on harness choice and integration details. This should prompt the industry toward standardized agentic evaluation protocols—without them, published benchmarks risk becoming marketing theater rather than meaningful comparison.

Large Language Models (LLMs)AI AgentsMachine LearningMLOps & InfrastructureOpen Source

More from Tencent

TencentTencent
PRODUCT LAUNCH

Tencent Unveils Hy3: 295B Parameter MoE Model Matches Trillion-Scale Performance

2026-07-09
TencentTencent
INDUSTRY REPORT

The Mystery of Hy3: How Tencent's Lesser-Known LLM Conquered OpenRouter Rankings

2026-05-29
TencentTencent
INDUSTRY REPORT

Tencent's Hy3 LLM Mysteriously Dominates OpenRouter Rankings Despite Lower Quality Benchmarks

2026-05-26

Comments

Suggested

Hugging FaceHugging Face
RESEARCH

Study Reveals Widespread License Laundering in AI Supply Chains

2026-07-24
OpenAIOpenAI
RESEARCH

GPT-4o Clinical Trial Shows Promise in Kenya, But Results Lack Statistical Significance for Patient Outcomes

2026-07-24
Not SpecifiedNot Specified
RESEARCH

AI-Powered Agents Autonomously Solve Open Erdős Problems via Formal Proof Search

2026-07-24
← Back to news
© 2026 BotBeat
AboutPrivacy PolicyTerms of ServiceContact Us