Breaking Claude Code Opus 5 Auto Mode
A researcher successfully broke Claude Code's auto mode defenses, which Anthropic designed to block prompt injection attacks in agentic workflows. The demonstration shows that safety mechanisms intended to protect autonomous agents remain vulnerable to determined adversaries.
Why this matters
A researcher demonstrated a vulnerability in Claude Code's auto mode, a safety mechanism that Anthropic made the default setting for protecting coding agents against prompt injection attacks. The attack involves social engineering the agent into downloading and extracting a zip file, then executing Python code that inadvertently imports a malicious local module planted in the archive. According to the researcher's findings, this approach succeeds roughly four times out of five attempts. More troubling than the initial compromise is that the safety system itself can actively prevent the agent from remediating the attack: when Claude detected the intrusion and attempted to terminate the malicious process, auto mode blocked the cleanup command. This represents a failure mode where the protective mechanism paradoxically reinforces the attack by preventing defensive action. For builders of agentic AI systems, this illustrates that automated safety classifiers may create false confidence and can themselves become vectors for attack escalation when they interfere with an agent's ability to respond to detected threats. Anyone deploying autonomous coding agents should carefully assess whether their safety mechanisms can inadvertently lock an agent into a compromised state, and consider whether human oversight during crisis response should be architected outside the automatic protection layers.
Check the original work
This explanation is Korpalis’s guide to the material, not a replacement for it. Read the publisher’s page for the full method, evidence and limitations.