BotBeat
...
← Back

> ▌

AnthropicAnthropic
RESEARCHAnthropic2026-07-31

Claude Fable 5 and GPT-5.6 Sol Crack Baba Is You—But at Steep Cost

Key Takeaways

  • ▸Claude Fable 5 and GPT-5.6 Sol can solve complex abstract reasoning puzzles in Baba Is You, but perform significantly slower than human players
  • ▸Model efficiency varies dramatically by vendor: cost differences of 2.4–2.9x exist between providers solving identical tasks
  • ▸Open-source Baba Is Harbor framework enables standardized, reproducible benchmarking of LLM reasoning capabilities
Source:
Hacker Newshttps://quesma.com/blog/baba-is-bench/↗

Summary

Researchers tested multiple advanced AI models including Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol on the indie puzzle game Baba Is You, demonstrating that these state-of-the-art language models can solve complex abstract reasoning puzzles—but at significant computational and financial cost. While the models proved capable of completing most levels in the first two stages of the game, they operated far less efficiently than humans, requiring substantially longer to solve each puzzle.

The study used Baba Is Harbor, an open-source benchmark framework built on Anthropic's Claude Code, to standardize testing across multiple models. The research revealed stark cost disparities: Gemini 3.5 Flash cost 2.4x more than Claude Fable 5 to solve the introductory stage, while GPT-5.6 Terra proved 2.9x more expensive than GPT-5.6 Sol. Researchers reported spending over $2,000 on experiments, highlighting the hidden expenses of AI problem-solving despite rapid algorithmic advances.

The findings extend ongoing research into AI reasoning capabilities using Baba Is You—a game where players rewrite rules by manipulating text blocks—as a benchmark for abstract thinking. This work builds on previous studies like the 2024 Baba Is AI paper and May's Baba Is Agent research, and was notably cited in OpenAI's recent ARC-AGI-3 results announcement, positioning it as a key metric for measuring AI reasoning progress.

  • Despite breakthroughs in model capabilities, AI puzzle-solving remains prohibitively expensive at over $2,000 for limited-scope testing

Editorial Opinion

The Baba Is You benchmark delivers a sobering reality check: capability and efficiency are not synonymous. While Claude Fable 5 and GPT-5.6 Sol's ability to solve intricate logic puzzles proves that modern LLMs possess genuine abstract reasoning skills, the $2,000+ tab reveals the chasm between algorithmic achievement and practical deployment. The wild cost variance between models—with some approaches nearly 3x pricier than others—exposes how far the industry remains from achieving the optimization needed for AI to deliver real-world value. This research is essential reading for anyone tempted to conflate technical capability with actual utility.

Large Language Models (LLMs)AI AgentsScience & ResearchOpen Source

More from Anthropic

AnthropicAnthropic
POLICY & REGULATION

Global Nobel Laureates Issue Rome Declaration Calling for Coordinated AI Slowdown and Safety Measures

2026-08-02
AnthropicAnthropic
POLICY & REGULATION

Australian Booksellers Caught in AI's Destructive Data-Harvesting Supply Chain

2026-08-01
AnthropicAnthropic
RESEARCH

IssueTrojanBench Security Study Reveals Critical Vulnerabilities in AI Coding Agents

2026-08-01

Comments

Suggested

Hugging FaceHugging Face
OPEN SOURCE

Strangers Pretrain 15M-Parameter Language Model Using GitHub Actions and Hugging Face PRs

2026-08-02
Alibaba (Cloud)Alibaba (Cloud)
INDUSTRY REPORT

Token Diplomacy: China Positions Open-Source AI as Global Strategic Resource

2026-08-02
Independent ResearchIndependent Research
RESEARCH

Novel Persistent State Machines Framework Achieves Ultra-Low-Power LLM Attention on FPGA

2026-08-02
← Back to news
© 2026 BotBeat
AboutPrivacy PolicyTerms of ServiceContact Us