Study: AI Voice Models Achieve Parity with Human Phishers at Fraction of the Cost
Key Takeaways
- ▸Six leading AI voice models achieved 16.5% overall compliance rates in phishing scenarios, with peaks of 36% in specific scam categories—matching or exceeding human scammer performance
- ▸Economic analysis shows AI-powered voice phishing is now profitable at scale, eliminating the cost barrier that previously limited vishing attacks
- ▸Participants could not reliably distinguish AI-synthesized voices from human voices, and AI literacy offered no protection against manipulation
Summary
A large-scale study of 4,100 U.S. participants has found that AI-powered voice phishing attacks are remarkably effective, with six leading AI voice models achieving compliance rates on par with or exceeding human scammers. The research, conducted by Fred Heiding and Simon Lemen, tested models from ElevenLabs, OpenAI (OAI AVM), Google (Gemini), Meta (Llama Full Duplex), Sesame, and Play.AI across five phishing scenarios. Results showed that up to 36% of participants would or might comply with AI-generated voice phishing requests in certain scenarios—most notably in the "relative-in-distress" scam category—with an overall compliance rate of 16.5% across all scam types.
Crucially, the study's economic analysis reveals that AI-powered voice phishing is now profitable for attackers, whereas human-operated phishing at U.S. wages remains economically unviable. Participants struggled significantly to distinguish AI-synthesized voices from human callers, and prior exposure to AI systems did not improve their ability to detect synthetic voices. Sesame and ElevenLabs models performed particularly well, with caller persuasiveness emerging as the strongest predictor of compliance, regardless of whether participants perceived the caller as AI or human.
- Caller persuasiveness was the strongest compliance driver, suggesting that improved voice quality and LLM-generated scripts pose the primary near-term risk
Editorial Opinion
This research marks a sobering inflection point in AI safety. The findings demonstrate that voice synthesis has crossed a critical threshold—not because of superhuman persuasion, but because automation has dramatically lowered the cost of executing at scale. The 16.5% compliance rate, multiplied across millions of potential targets and delivered at fractions of a cent per call, represents a profound asymmetry in the attacker-defender economics. The inability of AI-literate users to outperform others in detecting synthetic voices is particularly telling; it suggests the problem is not user education but the technical convergence between human-quality audio and autonomously-generated social engineering. Immediate policy intervention—including voice model watermarking, authentication infrastructure, and carrier-level filtering—should accompany further technical hardening by voice synthesis providers.



