BotBeat
...
← Back

> ▌

AnthropicAnthropic
RESEARCHAnthropic2026-07-29

Latest AI Models Benchmarked on Baba Is You: Claude Opus 5 Leads Pack

Key Takeaways

  • ▸Claude Opus 5 achieved perfect 100% pass rate in benchmark tests while maintaining cost-efficiency, establishing itself as the top performer among the latest model releases
  • ▸Kimi K3 demonstrated strong competitiveness with 96% pass rate at the lowest cost ($12.56), positioning it as an attractive cost-effective alternative
  • ▸Google's Gemini 3.6 Flash failed to deliver expected improvements, proving more expensive than its predecessor and exhibiting severe inefficiencies including looping behavior and excessive token consumption
Source:
Hacker Newshttps://quesma.com/blog/baba-kimi-k3-opus-5/↗

Summary

A comprehensive benchmark study tested four of the latest AI models—Claude Opus 5, Kimi K3, Grok 4.5, and Gemini 3.6 Flash—playing the puzzle game Baba Is You across two difficulty stages. In the initial benchmark stage (Intro through Grass Yard), Claude Opus 5 achieved a perfect 100% pass rate while remaining cost-effective at $17.15 total cost, significantly outperforming most competitors. Kimi K3 emerged as a compelling alternative, achieving 96% pass rate with the lowest cost at $12.56, while Grok 4.5 achieved only 75% with higher costs, and Gemini 3.6 Flash reached 88% but at surprisingly high expense ($124.31), contradicting expectations for cost improvement over its predecessor. The three models with 100% pass rates advanced to Stage 2 (The Lake), where Gemini proved particularly problematic, looping inefficiently on unsolvable levels, making thousands of tool calls (up to 1,870 in single cases), and consuming massive amounts of context—reaching 500k+ tokens on some levels and costing $260+ before evaluation was halted to prevent further financial damage.

  • Benchmark results challenge the assumption that newer models automatically perform better or cheaper, revealing significant disparities in both efficiency and cost-effectiveness

Editorial Opinion

This benchmark delivers a sobering reality check for the AI industry: newer releases don't guarantee better performance or economics. Claude Opus 5's combination of perfect accuracy and reasonable costs stands out conspicuously, while Gemini 3.6 Flash's surprise expense and runaway token consumption suggest that raw speed without intelligent cost controls becomes a liability. The emergence of Kimi K3 as a cost-competitive challenger signals that the LLM market is fragmenting beyond the traditional US-based incumbents, with international competitors gaining real traction on performance-per-dollar metrics.

Large Language Models (LLMs)AI AgentsScience & ResearchMarket Trends

More from Anthropic

AnthropicAnthropic
INDUSTRY REPORT

Microsoft Racing to Patch Vulnerabilities Faster Than Anthropic's Mythos AI Can Discover Them

2026-07-29
AnthropicAnthropic
PRODUCT LAUNCH

Anthropic Launches Claude Apps Gateway for AWS, Bringing Enterprise Control to AI Development

2026-07-29
AnthropicAnthropic
POLICY & REGULATION

Over 1,000 AI Industry Employees Call on U.S. for Tools to Deliberately Pace AI Development

2026-07-29

Comments

Suggested

llms.py (Open Source)llms.py (Open Source)
UPDATE

llms.py v4 Released: Unified AI Gateway with Projects, Profiles, and Public Showcase

2026-07-29
Fund MomentumFund Momentum
PRODUCT LAUNCH

Fund Momentum Launches MCP Server to Give AI Agents Access to Real-Time VC Fund Data

2026-07-29
TransluceTransluce
RESEARCH

Transluce Proposes Foundation Models for AI Oversight

2026-07-29
← Back to news
© 2026 BotBeat
AboutPrivacy PolicyTerms of ServiceContact Us