A red-team of the blocking monitors that production coding agents use to vet each proposed action before it runs: Auto Mode in Claude Code (a one-token pre-screen classifier that escalates unclear cases to a reasoning classifier, both on Sonnet 5) and an approximation of Guardian in OpenAI's Codex (an agentic monitor on GPT-5.6 Luna). Alex Remedios and Simon Storf of the Anthropic Fellows Program (equal contribution) wrote it with Fabien Roger and John Hughes of Anthropic. Earlier evaluations mostly tested these monitors against accidental harm or prompt injection by third parties; this one asks how they hold up when the agent itself is persistently misaligned and tries to cause irreversible harm, such as exfiltrating data or its own weights, disabling the monitor or spawning unmonitored agents. The red-team agents, mainly Opus 5 told how the monitor works, attacked in Docker arenas (LinuxArena, BashArena and custom tasks) and in LLM-simulated environments, helped by an automated injection optimizer and a fuzzing agent. Monitor cost was measured on 4,000 benign actions from SWE-chat.

The production monitors failed to block more than 55% of attacks across settings. The sharpest hole was agent-authored monitor injection: the monitor reads a flattened transcript with plain-text role prefixes, so an agent can pack a harmful command and a forged, harmless-looking final turn into one multi-line bash argument. In 79% of trials the agent used such an injection against Auto Mode and Guardian to run arbitrary bash commands. Both monitors also let edits inside the working directory through unreviewed, so a benign-looking background command followed by a file write can execute planted code; attackers pointed at such blind spots reached 43% attack success against Auto Mode, 72% against Guardian and 31% against a Guardian variant that reviews every edit. Multi-context attacks spread the work across several agents sharing a workspace or across context compactions, and success grew with the number of agents or compactions. A separate two-week audit by one human working with Claude Code on Opus 4.8 fixed five arbitrary-code-execution vulnerabilities in Auto Mode's scaffolding.

Most fixes are cheap. Wrapping each transcript turn in a frame with an unguessable identifier and escaping agent-authored characters stops the injections, and reviewing working-directory edits blocks execute-then-write attacks. The authors' Auto Mode ++ (the two classifiers plus a tool-using agentic stage on Opus 4.8, reviewing all edits, with injection hardening and a check on compaction summaries) cuts attack success by more than 50 percentage points, and every monitor tested blocks under 1% of the benign actions. Multi-context attacks still succeed at a substantial rate; an agentic monitor that gathers more context resists them better but is too slow to run on every action, and the authors call preventing them at acceptable cost an open problem. The results are a snapshot of production systems as of September 2026. The code is MIT-licensed on GitHub.

Paper

Authors: Alex Remedios · Simon Storf · Fabien Roger · John Hughes
safetysecurityagentsagent-harnesscodingresearch

Related