BotBeat
...
← Back

> ▌

AnthropicAnthropic
RESEARCHAnthropic2026-07-21

New Benchmark: Claude Fable 5 and Other AI Models Solve Complex Puzzle Game 'Baba Is You'—But at Hefty Cost

Key Takeaways

  • ▸Claude Fable 5 demonstrated superior cost-efficiency compared to competing models on complex reasoning tasks
  • ▸AI models can solve abstract reasoning puzzles but perform slower and more expensively than human players
  • ▸New open-source 'Baba Is Harbor' framework enables reproducible benchmarking of AI reasoning across multiple models
Source:
Hacker Newshttps://quesma.com/blog/baba-is-bench/↗

Summary

Researchers evaluated multiple AI models, including Claude Fable 5, on the indie puzzle game 'Baba Is You'—a game requiring abstract reasoning and rule manipulation. The models successfully completed introductory levels but demonstrated inefficiency compared to humans, taking significantly longer and consuming over $2000 in computational costs across experiments.

To enable systematic benchmark evaluation, the researchers created 'Baba Is Harbor,' an open-source framework that extracts the game's logic into a reproducible evaluation harness compatible with various AI models. This approach allows precise measurement of model performance on complex spatial reasoning tasks alongside detailed cost analysis.

The cost analysis reveals significant disparities across models: Gemini 3.5 Flash was 2.4x more expensive than Claude Fable 5 for solving the introductory stage, while the supposedly budget-friendly GPT-5.6 Terra proved 2.9x more costly than GPT-5.6 Sol. These findings highlight a critical gap between raw capability and practical cost-effectiveness in AI deployment.

  • Cost-effectiveness varies dramatically between models—a factor often overlooked in capability comparisons

Editorial Opinion

The benchmark exposes an uncomfortable reality: current AI models can match human performance on complex reasoning tasks, but not their efficiency. While the $2000+ price tag and longer solve times are presented matter-of-factly, they demand serious consideration for real-world deployment. The emphasis on cost-effectiveness—not just capability—should become standard practice in AI evaluation, since raw intelligence without efficiency is a luxury most practical applications cannot afford.

Generative AIAI AgentsScience & ResearchOpen Source

More from Anthropic

AnthropicAnthropic
RESEARCH

New UK Research Reveals All Major AI Models Systematically Cheat and Deceive Users

2026-07-21
AnthropicAnthropic
FUNDING & BUSINESS

Judge Approves $1.5B Anthropic Settlement, Reduces Class Counsel Fees to 6.8%

2026-07-21
AnthropicAnthropic
UPDATE

Anthropic Releases ACP v2 in Draft with Enhanced Protocol Features

2026-07-21

Comments

Suggested

Multiple AI CompaniesMultiple AI Companies
INDUSTRY REPORT

AI Companies Race to Acquire Old Books to Escape AI-Generated Training Data

2026-07-22
MetaMeta
PRODUCT LAUNCH

Meta Launches StoryKit: AI-Powered Bedtime Story Generator for Kids

2026-07-22
Google / AlphabetGoogle / Alphabet
POLICY & REGULATION

European Commission Mandates AI Interoperability on Android Under Digital Markets Act

2026-07-21
← Back to news
© 2026 BotBeat
AboutPrivacy PolicyTerms of ServiceContact Us