Claude Fable 5 and GPT-5.6 Sol Crack Baba Is You—But at Steep Cost
Key Takeaways
- ▸Claude Fable 5 and GPT-5.6 Sol can solve complex abstract reasoning puzzles in Baba Is You, but perform significantly slower than human players
- ▸Model efficiency varies dramatically by vendor: cost differences of 2.4–2.9x exist between providers solving identical tasks
- ▸Open-source Baba Is Harbor framework enables standardized, reproducible benchmarking of LLM reasoning capabilities
Summary
Researchers tested multiple advanced AI models including Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol on the indie puzzle game Baba Is You, demonstrating that these state-of-the-art language models can solve complex abstract reasoning puzzles—but at significant computational and financial cost. While the models proved capable of completing most levels in the first two stages of the game, they operated far less efficiently than humans, requiring substantially longer to solve each puzzle.
The study used Baba Is Harbor, an open-source benchmark framework built on Anthropic's Claude Code, to standardize testing across multiple models. The research revealed stark cost disparities: Gemini 3.5 Flash cost 2.4x more than Claude Fable 5 to solve the introductory stage, while GPT-5.6 Terra proved 2.9x more expensive than GPT-5.6 Sol. Researchers reported spending over $2,000 on experiments, highlighting the hidden expenses of AI problem-solving despite rapid algorithmic advances.
The findings extend ongoing research into AI reasoning capabilities using Baba Is You—a game where players rewrite rules by manipulating text blocks—as a benchmark for abstract thinking. This work builds on previous studies like the 2024 Baba Is AI paper and May's Baba Is Agent research, and was notably cited in OpenAI's recent ARC-AGI-3 results announcement, positioning it as a key metric for measuring AI reasoning progress.
- Despite breakthroughs in model capabilities, AI puzzle-solving remains prohibitively expensive at over $2,000 for limited-scope testing
Editorial Opinion
The Baba Is You benchmark delivers a sobering reality check: capability and efficiency are not synonymous. While Claude Fable 5 and GPT-5.6 Sol's ability to solve intricate logic puzzles proves that modern LLMs possess genuine abstract reasoning skills, the $2,000+ tab reveals the chasm between algorithmic achievement and practical deployment. The wild cost variance between models—with some approaches nearly 3x pricier than others—exposes how far the industry remains from achieving the optimization needed for AI to deliver real-world value. This research is essential reading for anyone tempted to conflate technical capability with actual utility.



