AgentsAI Safety 🇩🇪 08.08.2026 18:02

Anthropic Switches Claude Code to Auto-Mode by Default to Protect Developers

AnthropicAnthropic OpenAIOpenAI
Anthropic will make Claude Code run in auto-mode by default for Pro, Max, and Team plans starting August 14. The mode allows the AI coding tool to work autonomously, with a classifier checking for dangerous actions. Tests show it is at least as safe as manual approvals and increases productivity by 25%.
Anthropic announced that Claude Code will default to auto-mode for Pro, Max, and Team plans starting August 14, while Enterprise customers must opt in. In auto-mode, the AI coding tool works independently without waiting for manual approval at each step; instead, a classifier automatically checks if an action is dangerous or irreversible and only asks for confirmation if needed. According to Anthropic, tests with 1,053 paid testers and internal red-teaming showed that auto-mode is at least as safe as or safer than manual approvals, and teams using auto-mode create about 25% more pull requests. Anthropic also highlights that auto-mode offers an additional layer of protection against prompt-injection attacks, where injected code tries to divert the agent from the user's instructions. An independent review by Trajectory Labs tested 72 such attack scenarios ten times each, and none of the 720 attacks succeeded against Claude's current models (Fable 5, Opus 5, Sonnet 5) in auto-mode, while 5.83% of attacks succeeded against OpenAI's GPT-5.6 Sol in Codex auto-review mode. Internally, auto-mode prevented Claude from uploading confidential data to a public page and from terminating around 2,000 processes in a long session, which would have destroyed running GPU training jobs. Anthropic does not charge for the tokens consumed by the classifier itself, but the auto-mode as default is likely economically attractive for the company as it increases overall token usage and revenue. Claude Code is currently the most widely used AI coding tool, and this shift moves the developer role from active coding to monitoring AI-generated output. Anthropic advises caution with production-critical infrastructure, noting that the classifier reduces but does not eliminate risks, which creates a paradox: the less developers intervene, the more critical their control becomes, while understanding projects largely created autonomously becomes harder.
Source: The Decoder (DE) — original
Our earlier posts on this topic ↓
Fresh news