Tencent Releases WorkBuddy Bench: Multi-Model Agentic Coding Leaderboard Shows No Clear Winner
Key Takeaways
- ▸No single model dominates across all tasks—Claude Opus 4.8 leads on Code and Web, GLM-5.2 on Security, GPT-5.5 on Office efficiency
- ▸Open-weight models like GLM-5.2 remain competitive, especially in security-focused agentic tasks
- ▸Token efficiency is disconnected from ranking: GPT-5.5 achieves best-in-class efficiency (6.9k output tokens on Code) while mid-table models spend 3-4x more
Summary
Tencent has released WorkBuddy Bench, a comprehensive agentic coding leaderboard that benchmarks multiple AI models across code generation, web automation, office tasks, and security-focused challenges using two different evaluation harnesses (CodeBuddy Code and Claude Code). The benchmark reveals no single dominant model: Claude Opus 4.8 leads five of eight scored columns, GLM-5.2 (an open-weight model) excels in Security tasks, and GPT-5.5 dominates in Office tasks under Claude Code.
The results surface three critical insights: token efficiency is decoupled from performance ranking, with GPT-5.5 achieving top-tier scores using 3-4x fewer output tokens than competitors; performance varies dramatically across harnesses for the same model (GLM-5.2's Web score swings 67.43 vs 60.71); and integration details like cross-turn reasoning passback can meaningfully shift scores. Open-weight models remain competitive, particularly in specialized domains like security, suggesting the agentic coding market retains meaningful diversity rather than consolidating around closed-source solutions.
- Evaluation methodology heavily influences results—identical models show significant score variations across different harnesses and configurations
- Integration details matter: enabling cross-turn reasoning for HY-3 boosted its Code score by +1.92 to +3.82 points
Editorial Opinion
Tencent's WorkBuddy Bench fills a critical gap in agentic AI evaluation, providing rare transparency into model specialization rather than false universalism. The finding that open-weight and commercial models remain competitive across different domains is bullish for AI diversity, but the benchmark also exposes a methodological fragility: identical models produce wildly different scores depending on harness choice and integration details. This should prompt the industry toward standardized agentic evaluation protocols—without them, published benchmarks risk becoming marketing theater rather than meaningful comparison.



