Claude Opus 5 Tops MacBook SVG Benchmark But Reasoning Mode Burns Tokens
Key Takeaways
- ▸Claude Opus 5 generates the most detailed MacBook SVG among Anthropic models but with compromised geometry at all effort levels (skewed at high, broken isometry at xhigh, non-existent at max)
- ▸Thinking-on-by-default causes 10–25x higher token costs compared to Opus 4.8, making Opus 5's maximum effort setting economically inefficient without quality gains
- ▸The new MacBook benchmark successfully replaces the 'pelican on a bicycle' test by offering universal visual grounding and a genuine difficulty ladder that better separates model capabilities
Summary
A rigorous new benchmark test comparing frontier AI models has emerged to replace the increasingly outdated 'pelican on a bicycle' prompt. The test asks models to draw a MacBook Pro 16 in SVG format in a single shot—a task with universally understood visual success criteria that far better separates model capabilities than previous benchmarks. The MacBook test succeeds because hundreds of millions of people know exactly what one looks like, making even small geometric errors instantly apparent.
Anthropicโ€™s Claude Opus 5 produces the most detailed MacBook drawings among Anthropic's models, but the results reveal significant trade-offs. While Fable 5 delivers the highest-quality overall outputs and Gemini 3.6 Flash achieves competitive results at just $0.16 per run, Opus 5 struggles with geometric fidelity at every effort level. Worse, the model's default reasoning behavior drives token consumption to 10–25 times higher than Opus 4.8, consuming 20–47K tokens per drawing. At maximum effort, extended thinking exhausts the entire token budget without completing a drawing.
The benchmark's design deliberately avoids the pitfalls that rendered the pelican test obsolete: MacBook imagery is not yet exhaustively trained on (unlike thousands of documented pelicans), and it has a genuine difficulty ladder with dozens of geometric components and 3D perspective challenges. Across 12 models and 64 total runs, the MacBook test shows models spending roughly 4x more tokens on the task than they do on the pelican, revealing the test's true discriminative power.
- Claude Fable 5 delivers the best-quality outputs overall; Gemini 3.6 Flash provides competitive results at a fraction of the cost
Editorial Opinion
The MacBook test represents meaningful progress in AI benchmarking methodology, but Opus 5's results expose a troubling pattern: aggressive reasoning modes can harm visual generation without clear upside. Anthropic's decision to enable thinking-on-by-default for Opus 5 appears to be a misfire for this category of task, prioritizing capability signals over economic efficiency. The benchmark confirms that cost-to-quality ratio remains as decisive as raw capability when comparing frontier models.


