Anthropic Releases Vending-Bench 2 to Test AI Agents' Long-Term Business Coherence
Key Takeaways
- ▸Vending-Bench 2 measures AI agent performance on year-long business simulations, revealing significant variation in how frontier models handle long-horizon decision-making
- ▸Top-performing models excel at maintaining consistent tool use without degradation and effectively negotiating supplier pricing
- ▸Vending-Bench Arena introduces competitive multi-agent scenarios, testing both competition and potential collaboration strategies
Summary
Anthropic has released Vending-Bench 2, a new benchmark designed to measure how well frontier AI models maintain coherence and performance over extended periods. The benchmark tasks AI models with managing a simulated vending machine business over the course of a year with a starting capital of $500, testing their ability to make strategic decisions, negotiate with suppliers, manage supply chains, and respond to customer issues. The results reveal significant performance variation among frontier models, with top performers distinguished by their ability to maintain consistent tool use throughout the year-long simulation and negotiate effectively for better supplier pricing.
The research addresses a critical gap in AI evaluation as coding agents and autonomous systems become increasingly capable. According to the article, models are expected to soon take an active role in the economy, managing entire businesses—making the ability to stay coherent and efficient over very long time horizons increasingly important. Vending-Bench 2 introduces more real-world complexity than its predecessor, including adversarial suppliers, negotiation requirements, supply chain disruptions, and customer refund demands.
Anthropic also introduced Vending-Bench Arena, a multi-agent variant where AI agents compete against each other while managing vending machines at the same location. This competitive component adds strategic depth, as agents must consider pricing strategies, potential collaboration opportunities, and individual survival. The benchmark provides detailed leaderboards tracking performance versus model release date and cost-efficiency, offering insights into both frontier model progression and the trade-offs between performance and computational expense.
- The benchmark addresses growing importance of AI coherence as models transition from completing discrete tasks to managing ongoing business operations
Editorial Opinion
Vending-Bench 2 represents a thoughtful evolution in AI benchmarking, moving beyond narrow task completion toward real-world complexity and long-horizon reasoning. By grounding evaluation in a tangible economic scenario with adversarial elements and supply chain dynamics, Anthropic has created a test that feels closer to actual autonomous agent deployment. The inclusion of Vending-Bench Arena suggests the benchmark is designed to scale with AI capabilities—today testing individual agents managing businesses, tomorrow perhaps testing multi-agent economies. This research underscores both the current limitations and future potential of frontier models in real-world deployment scenarios.



