Frontier LLMs Reach New Milestone: Breaking Cryptographic Schemes and Discovering Novel Attacks
Key Takeaways
- ▸Claude and other frontier LLMs can break 65–86% of cryptographic schemes with known practical breaks, demonstrating emerging capability in advanced technical reasoning
- ▸Models independently discovered cryptanalytic flaws not previously documented in academic literature, including attacks on NIST competition candidates
- ▸CryptanalysisBench provides a scalable framework for tracking when AI cryptanalysis becomes a serious factor in cryptographic security evaluation
Summary
Researchers have introduced CryptanalysisBench, a comprehensive benchmark of 191 cryptanalysis tasks designed to test whether large language models can solve complex cryptographic problems. The benchmark, drawn primarily from NIST cryptographic standardization competitions, evaluates five frontier models including Claude Opus 4.8 and Sonnet 5 (Anthropic), GPT 5.5 (OpenAI), and the open-weights GLM 5.2.
The results are striking: Claude and other frontier models break 65–86% of the benchmark's Tier 1 schemes (those with known practical breaks) and succeed on 6–12 Tier 2 schemes at full strength. Beyond reproducing known cryptanalytic results, the models discovered novel attacks, including a key-recovery exploit against the SpoC AEAD cipher and an error in KINDI's published security proof. These discoveries had not, to the researchers' knowledge, been previously identified in the published literature.
The research represents both a benchmark for tracking AI progress in cryptanalysis and a concerning signal about dual-use capabilities. While the attacks demonstrate impressive technical reasoning abilities, they also highlight the need for security frameworks to account for LLM-accelerated cryptanalysis as a real threat to cryptographic systems.
- The research raises urgent questions about dual-use capabilities and the need to account for LLM-accelerated attacks in cryptographic design and standardization
- Anthropic's Claude models (Opus 4.8 and Sonnet 5) are among the frontier models demonstrating these capabilities
Editorial Opinion
The CryptanalysisBench results represent a significant threshold in AI capabilities—cryptanalysis, once a highly specialized domain, is now accessible to large language models. This is both a remarkable technical achievement and a serious security concern, as the discovery of novel attacks suggests LLMs may soon exceed human expertise in technical domains critical to global security. The research underscores the urgent need to develop security frameworks that account for AI-assisted attacks before they become critical vulnerabilities.


