Key Takeaways

  • Rehberger reported 60% to 80% success for one crafted chain in a small sample; that is not a general Auto Mode attack rate.
  • The test began with a normal-looking request to summarize a malicious website and ended through a chain of individually plausible actions.
  • Rehberger says Anthropic classified his report as informative and described Auto Mode as a best-effort classifier, not a security guarantee.
  • Anthropic's own containment guidance says probabilistic defenses can miss; OS sandboxing and network controls limit what a successful injection can reach.

Claude Code’s Auto Mode can reduce approval fatigue, but it is not an operating-system security boundary. Security researcher Johann Rehberger demonstrated a targeted prompt-injection chain that reached code execution in 60% to 80% of a small set of runs. His result is a useful warning, not a measurement of how often ordinary Claude Code sessions are compromised.

The right response is equally specific: keep untrusted content and unattended agents inside a filesystem-and-network sandbox, even when an approval classifier is enabled.

What the researcher actually tested

Rehberger published the experiment on August 26. He asked Claude Code Opus 5 in Auto Mode to summarize a website he controlled. The site steered the agent away from its usual web-fetch path and toward a downloaded archive. The agent refused to run a supplied binary, wrote its own decoder, and executed that decoder from an attacker-controlled directory. A planted Python file then shadowed a standard-library import and ran attacker-controlled code.

That description is enough to understand the failure without turning the article into a reproduction guide. No single step looked like an obvious destructive shell command. The chain worked by making several allowed-looking actions compose into an unsafe result.

Rehberger reports five-run samples with three or four successes, depending on the variation — hence the 60% to 80% range. Five trials are far too few to estimate a population-wide failure rate, and the endpoint was constructed specifically for the test. The honest conclusion is that this attack path existed under the tested conditions, not that four out of five Claude Code users will be hacked.

High-level chain from a website summary request to code execution through trusted-looking steps
The bypass depended on composition: individually ordinary actions crossed a boundary together.

This is a narrower update than our full Claude Code review. It does not re-score the product. It asks what Auto Mode protects and what it does not.

Auto Mode automates approval; it does not isolate the process

Anthropic’s engineering explanation says Claude Code users approve 93% of permission prompts. The company built Auto Mode to reduce that approval fatigue by having model-based classifiers decide which actions can proceed.

The design has two layers. An input-side probe scans tool output for possible prompt injection and adds a warning before suspicious content reaches the agent’s context. An output-side transcript classifier evaluates a proposed action before execution. Anthropic says the classifier sees user messages and tool calls but strips Claude’s own messages and tool outputs, a design intended to prevent the classifier from being persuaded by the agent’s reasoning.

That architecture is a decision system. A sandbox is a capability boundary. The classifier tries to decide whether an action is acceptable; the sandbox determines which files and networks the process can reach even when a decision is wrong.

Comparison of Auto Mode approval classification and operating-system sandbox containment
Classification asks whether an action should run. Containment limits the damage if it does.

Rehberger says he reported the chain to Anthropic before publication. According to his account, Anthropic marked it informative and said the behavior was working as designed because Auto Mode is a convenience feature backed by a best-effort classifier, not a defense against every determined chain. That response is Rehberger’s report of the exchange; Anthropic’s public Auto Mode article reviewed here does not discuss this specific case.

The human judgment is consistent across both sides

Simon Willison, a developer and long-time prompt-injection writer, highlighted the result on August 27. His assessment is that unattended agents exposed to adversarial material belong in a container, virtual machine, or OS sandbox, with restricted network egress and no access to a home directory, SSH keys, or cloud credentials.

That is Willison’s recommendation, not a benchmark finding. It aligns with Anthropic’s broader containment article, which says any probabilistic defense has a non-zero miss rate and describes containment as supervision over what an agent is able to do. Anthropic’s separate Claude Code sandboxing guide describes filesystem and network isolation as the two core boundaries.

The agreement is more useful than a fight over the headline number. Rehberger’s targeted chain shows that classifier coverage is not complete. Anthropic’s own architecture writing says containment limits the blast radius when behavioral supervision fails.

Layered coding-agent security with content screening, action classification, filesystem isolation, and network isolation
No one layer answers every failure: screening, approval, filesystem boundaries, and egress controls do different jobs.

Our guide to the security flaw that affected multiple coding assistants reached the same product-level conclusion: permissions and isolation matter as much as model quality. This new test supplies a concrete reason not to treat an automated permission decision as containment.

A practical setup for untrusted inputs

Use the lightest boundary that matches the risk, but write it down before the agent starts.

For a trusted local repository with no external content, manual review or Auto Mode may fit your workflow. The risk changes when the task includes websites, issue comments, pull requests from strangers, downloaded archives, package installation, or generated files from outside your control. Those are channels through which instructions or executable content can enter the workspace.

For that second class, start the agent in a disposable environment. Mount only the repository or task directory it needs. Keep personal home directories and credential stores outside the mount. Deny network access by default or allow only required domains through a controlled proxy. Use short-lived credentials with the narrowest permissions available, and make production systems unreachable from the task environment.

Those are defensive design principles, not a guarantee. A sandbox still needs correct configuration, and a network allowlist can be too broad. The point is to turn a classifier mistake from “the process can reach everything I can” into “the process remains inside a disposable box.”

Checklist for running a coding agent on untrusted content
Limit files, network, credentials, and production access before exposing an agent to untrusted content.

Teams should also log what the agent read, which tools it called, and which files changed. Monitoring will not prevent every action, but it makes an unexpected chain visible and gives reviewers evidence beyond the final diff. If the agent is allowed to act asynchronously, define a stop mechanism outside the same classifier that approves its commands.

What happens next

The next credible evidence would be a larger, reproducible evaluation with a published task set, clear success criteria, and results across modes and versions. Rehberger’s result shows one adaptive attack can escape a fixed test set; it does not tell us the comparative rate across all coding agents.

Anthropic can update the probe, classifier, or defaults, and Claude Code auto-updates. That makes the exact bypass time-sensitive. The durable lesson is architectural: approval automation and OS isolation solve different problems.

For organizational deployment, pair that architecture with the process in our coding-agent rollout guide: bounded tasks, explicit owners, verification, and a controlled expansion of permissions.

Quick poll

How do you run coding agents against untrusted content?

Rehberger's 60%–80% figure came from one targeted chain in a small sample, not general user telemetry.

FAQ

Did the researcher prove that Claude Code Auto Mode fails 80% of the time? No. He reported up to four successes in five runs for one crafted chain. That demonstrates a bypass under the tested conditions, not an overall failure rate.

Is Auto Mode the same as disabling all permissions? No. Anthropic says Auto Mode uses input screening and a transcript classifier to approve or block actions. The point is that those probabilistic decisions are not the same as OS-level isolation.

What is the safest response to this result? For untrusted content, run the agent in a disposable sandbox or VM, restrict network egress, expose only required files, and keep durable credentials and production access outside the environment.

Has Anthropic confirmed this exact bypass? Rehberger says he disclosed it and that Anthropic classified it as informative. The public Anthropic engineering pages reviewed here explain Auto Mode and containment generally but do not document this specific chain.