New Benchmark: Claude Fable 5 and Other AI Models Solve Complex Puzzle Game 'Baba Is You'—But at Hefty Cost
Key Takeaways
- ▸Claude Fable 5 demonstrated superior cost-efficiency compared to competing models on complex reasoning tasks
- ▸AI models can solve abstract reasoning puzzles but perform slower and more expensively than human players
- ▸New open-source 'Baba Is Harbor' framework enables reproducible benchmarking of AI reasoning across multiple models
Summary
Researchers evaluated multiple AI models, including Claude Fable 5, on the indie puzzle game 'Baba Is You'—a game requiring abstract reasoning and rule manipulation. The models successfully completed introductory levels but demonstrated inefficiency compared to humans, taking significantly longer and consuming over $2000 in computational costs across experiments.
To enable systematic benchmark evaluation, the researchers created 'Baba Is Harbor,' an open-source framework that extracts the game's logic into a reproducible evaluation harness compatible with various AI models. This approach allows precise measurement of model performance on complex spatial reasoning tasks alongside detailed cost analysis.
The cost analysis reveals significant disparities across models: Gemini 3.5 Flash was 2.4x more expensive than Claude Fable 5 for solving the introductory stage, while the supposedly budget-friendly GPT-5.6 Terra proved 2.9x more costly than GPT-5.6 Sol. These findings highlight a critical gap between raw capability and practical cost-effectiveness in AI deployment.
- Cost-effectiveness varies dramatically between models—a factor often overlooked in capability comparisons
Editorial Opinion
The benchmark exposes an uncomfortable reality: current AI models can match human performance on complex reasoning tasks, but not their efficiency. While the $2000+ price tag and longer solve times are presented matter-of-factly, they demand serious consideration for real-world deployment. The emphasis on cost-effectiveness—not just capability—should become standard practice in AI evaluation, since raw intelligence without efficiency is a luxury most practical applications cannot afford.



