Frontier LLMs Break Cryptographic Schemes, Discover Novel Attacks: CryptanalysisBench Reveals Claude and GPT-5.5 Can Perform Cryptanalysis
Key Takeaways
- ▸Frontier LLMs including Claude Opus 4.8, Sonnet 5, and GPT-5.5 break 65-86% of known-vulnerable cryptographic schemes, with 6-12 models achieving breaks on previously-unbroken Tier 2 schemes
- ▸Models demonstrated novel cryptanalysis capability, discovering previously-unknown vulnerabilities in SpoC AEAD and KINDI schemes—suggesting frontier-level mathematical reasoning
- ▸CryptanalysisBench provides 191 standardized cryptanalysis tasks across six cryptographic primitive families as a tool to benchmark AI progress and stress-test candidate schemes
Summary
Researchers have introduced CryptanalysisBench, a comprehensive benchmark of 191 cryptanalysis tasks spanning six families of cryptographic primitives drawn from NIST standardization competitions, designed to evaluate whether frontier large language models can discover attacks against cryptographic schemes. Testing five frontier models including Anthropic's Claude Opus 4.8 and Sonnet 5, OpenAI's GPT-5.5, Mythos 5, and GLM 5.2, researchers found that these models break 65-86% of known-vulnerable cryptographic schemes in Tier 1, with 6-12 models achieving breaks on Tier 2 schemes with no previously known attacks at full strength. Beyond reproducing known results, the models demonstrated novel cryptanalysis capabilities, including discovering a key-recovery attack exploiting a design flaw in the SpoC AEAD scheme and identifying an error in KINDI's published CCA-security proof—both previously unknown attacks. The benchmark is being released as both a frontier-tracking tool and a stress-test for evaluating cryptographic candidates before real-world deployment.
- Research suggests AI cryptanalysis capabilities may soon match or exceed published state-of-the-art, raising urgent questions about cryptographic standards resilience and AI governance
Editorial Opinion
The discovery that frontier LLMs can break established cryptographic schemes and independently discover novel attacks marks a significant milestone in AI reasoning while delivering a sobering signal for cybersecurity. While CryptanalysisBench provides a valuable evaluation tool, the speed at which models are reaching frontier cryptanalysis levels demands urgent review of deployed cryptographic standards and the governance frameworks guiding frontier AI capabilities. This research exemplifies the dual nature of advanced AI: remarkable technical achievement paired with security-critical implications requiring proactive policy response.



