Latest AI Models Benchmarked on Baba Is You: Claude Opus 5 Leads Pack
Key Takeaways
- ▸Claude Opus 5 achieved perfect 100% pass rate in benchmark tests while maintaining cost-efficiency, establishing itself as the top performer among the latest model releases
- ▸Kimi K3 demonstrated strong competitiveness with 96% pass rate at the lowest cost ($12.56), positioning it as an attractive cost-effective alternative
- ▸Google's Gemini 3.6 Flash failed to deliver expected improvements, proving more expensive than its predecessor and exhibiting severe inefficiencies including looping behavior and excessive token consumption
Summary
A comprehensive benchmark study tested four of the latest AI models—Claude Opus 5, Kimi K3, Grok 4.5, and Gemini 3.6 Flash—playing the puzzle game Baba Is You across two difficulty stages. In the initial benchmark stage (Intro through Grass Yard), Claude Opus 5 achieved a perfect 100% pass rate while remaining cost-effective at $17.15 total cost, significantly outperforming most competitors. Kimi K3 emerged as a compelling alternative, achieving 96% pass rate with the lowest cost at $12.56, while Grok 4.5 achieved only 75% with higher costs, and Gemini 3.6 Flash reached 88% but at surprisingly high expense ($124.31), contradicting expectations for cost improvement over its predecessor. The three models with 100% pass rates advanced to Stage 2 (The Lake), where Gemini proved particularly problematic, looping inefficiently on unsolvable levels, making thousands of tool calls (up to 1,870 in single cases), and consuming massive amounts of context—reaching 500k+ tokens on some levels and costing $260+ before evaluation was halted to prevent further financial damage.
- Benchmark results challenge the assumption that newer models automatically perform better or cheaper, revealing significant disparities in both efficiency and cost-effectiveness
Editorial Opinion
This benchmark delivers a sobering reality check for the AI industry: newer releases don't guarantee better performance or economics. Claude Opus 5's combination of perfect accuracy and reasonable costs stands out conspicuously, while Gemini 3.6 Flash's surprise expense and runaway token consumption suggest that raw speed without intelligent cost controls becomes a liability. The emergence of Kimi K3 as a cost-competitive challenger signals that the LLM market is fragmenting beyond the traditional US-based incumbents, with international competitors gaining real traction on performance-per-dollar metrics.



