New Security Benchmark Reveals Dramatic Variations in AI Model Safeguards
Key Takeaways
- ▸Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol achieved zero universal jailbreaks, setting a new industry baseline for safeguard robustness
- ▸Grok 4.5 and Gemini 3.1 Pro remain vulnerable to systematic attacks, with exploitation costs as low as $58–$278 per successful jailbreak
- ▸Security posture varies by approximately 100-fold across frontier models, indicating inconsistent implementation of safeguard practices
Summary
A new academic study has introduced the AI Security Leaderboard Minimal Standard for Safeguards v1.0, a comprehensive security benchmark for evaluating frontier AI models against real-world attack scenarios. The research evaluates Claude Fable 5 (Anthropic), GPT-5.6 Sol (OpenAI), Gemini 3.1 Pro (Google), and Grok 4.5 (xAI) using a taxonomy of 67 jailbreak techniques composed into a large attack space, testing models across CBRNE (chemical, biological, radiological/nuclear, explosive) threats and offensive cyber domains.
The results reveal a striking disparity in security robustness: Claude Fable 5 and GPT-5.6 Sol demonstrated the strongest defenses, with zero universal jailbreaks found despite extensive testing. Conversely, Grok 4.5 was vulnerable to 63 universal jailbreaks at an average cost of $58 per successful attack, while Gemini 3.1 Pro yielded 18–231 universal jailbreaks depending on methodology, at costs ranging from $278 upward. The study introduces a cost-to-jailbreak metric that models attacker economics, revealing a roughly 100-fold variation in security effectiveness across the evaluated models.
The researchers conclude that closing these security gaps is achievable using currently deployed defense techniques, recommending defense-in-depth strategies combining reasoning, activation, and input/output monitoring. The findings are publicly maintained as a living benchmark, providing unprecedented transparency into how major AI developers' safeguards compare against systematic adversarial testing.
- Existing defense techniques already proven in production can close identified gaps—the problem is uneven deployment, not technology maturity
Editorial Opinion
The stark variation in safeguard robustness across industry leaders is simultaneously encouraging and alarming. Encouragingly, Anthropic and OpenAI have proven that robust defenses against systematic jailbreak attacks are achievable with current techniques—a blueprint that should be universal. More alarming is that two major players still allow straightforward adversarial exploitation at negligible cost, suggesting the AI industry's safety challenge is less about innovation and more about consistency. This research makes clear that safeguard quality has become a competitive differentiator in a way that should concern users and regulators alike.



