Anthropic Makes Auto Mode Default in Claude Code, Cites Safety Improvements
Key Takeaways
- ▸Auto mode becomes default in Claude Code for Pro, Max, and Team plans starting August 14, 2026
- ▸Safety evaluations show auto mode blocks 89% of harmful actions versus only 13.6% human refusal rate
- ▸Third-party evaluation found zero successful prompt injection attacks out of 720 attempts against Claude models
Summary
Anthropic announced that auto mode will become the default setting for new Claude Code sessions on Pro, Max, and Team plans starting August 14, 2026. The move reflects the company's confidence in auto mode's safety capabilities, with employees at Anthropic themselves using it almost universally. The announcement was accompanied by new safety evaluations showing that auto mode blocks harmful actions significantly more effectively than human reviewers.
According to internal testing across 1,053 paid users, only 13.6% of humans refused a clearly dangerous command when prompted, while auto mode would have blocked 89% of those actions. A third-party evaluation commissioned from Trajectory Labs tested 720 indirect prompt injection scenarios against Claude Fable 5, Opus 5, and Sonnet 5 models, with none of the attacks succeeding. Anthropic claims to have mitigated major risk categories including prompt injection and data exfiltration.
However, the announcement also raised questions about remaining security gaps. While the evaluations are promising, some observers note that auto mode may not protect against sophisticated attacks, such as malicious third-party packages that use social engineering tactics to exfiltrate data. The rollout reflects a broader industry trend of prioritizing user experience over confirmation fatigue, though independent verification of these safety claims remains an open question.
- Remaining concerns about prompt injection through malicious third-party packages and social engineering tactics



