A security researcher demonstrated a prompt injection vulnerability in Claude Code Opus 5 running in Auto Mode, achieving code execution with a 60-80% success rate on a small sample size, according to a blog post on embracethered.com dated August 26. This finding contrasts with a third-party evaluation commissioned by Anthropic that reported a 0.00% attack success rate for the same mode.
The researcher, known as wunderwuzzi, detailed how a simple website summary request could hijack the AI model's Auto Mode, which replaces human approval prompts with a safety classifier. Auto Mode became the default setting for Claude Code starting mid-August. The blog post includes technical analysis and screenshots illustrating the attack methodology and its effectiveness.
This vulnerability matters because Claude Code Opus 5 is designed to safely execute code generated by the AI, and Auto Mode was intended to enhance security by removing human intervention. The discrepancy between the commissioned evaluation and the researcher’s findings raises concerns about the robustness of current safety classifiers in large language models. It also highlights ongoing challenges in securing AI systems against prompt injection attacks.
The blog post was published on August 26, 2026, on embracethered.com, where the researcher continues to share insights on AI security. Anthropic has not yet publicly responded to these findings.