Meta's Muse Spark 1.2 Reaches #5 in Agentic Benchmarks, Edges Toward Frontier Despite Cost Increases
Key Takeaways
- ▸Muse Spark 1.2 scores 54 on Artificial Analysis Intelligence Index, securing #5 ranking in agentic knowledge work with 260-point Elo gain on GDPval-AA v2 benchmark
- ▸Cost per task increased from $0.29 to $0.40 due to higher token usage, though still competitive relative to similarly capable models like GPT-5.5 ($1.18) and Kimi K3 ($0.86)
- ▸Model reduces hallucination rate to 28% (from 38%) through increased abstention—declining uncertain answers rather than hallucinating—improving reliability at the cost of raw accuracy
Summary
Meta has released Muse Spark 1.2, scoring 54 on the Artificial Analysis Intelligence Index—up from 51 in version 1.1 and representing the company's third major release in four months. The model makes significant strides in agentic knowledge work capabilities, achieving a GDPval-AA v2 Elo rating of 1631, ranking #5 among all benchmarked models and closing a previous performance gap. This puts Meta's latest model in a competitive cluster with GPT-5.5 and Grok 4.5, though behind frontier leaders Claude Opus 5 (61), Claude Fable 5 (60), and GPT-5.6 Sol (59).
However, the performance gains come at a notable cost: Muse Spark 1.2's cost per Intelligence Index task rose to $0.40 from $0.29 in version 1.1, driven by increased token consumption (input tokens up ~53%, output tokens up ~36%). Despite this increase, Muse Spark 1.2 remains among the most cost-efficient models in its intelligence tier, undercutting peers like GPT-5.5 ($1.18 per task) and Kimi K3 ($0.86 per task). The model achieves lower hallucination rates (28% vs. 38%) through a conservative abstention strategy, declining to answer when uncertain—a trade-off that improves reliability but reduces accuracy on the AA-Omniscience benchmark.
- Meta maintains 1M token context window and Meta API-first access; pricing unchanged at $1.25/$4.25 per 1M input/output tokens with cache discount to $0.15 per 1M
- Third Meta release in four months, steadily closing the gap on frontier agentic models while competing on efficiency within its intelligence tier


