✳ the wire · analysis

Claude Code makes auto mode the default on Aug 14 — and the study behind it says human review decays as sessions get longer

confirmedby ArtificialWatch
Claude Code makes auto mode the default on Aug 14 — and the study behind it says human review decays as sessions get longer
Source imagery · verified against a primary source

On the same day OpenAI slowed a model down over what it might do autonomously, Anthropic made autonomous execution the default in its coding agent. The two are not the same decision — one is about a frontier model's capability, the other about a product's permission prompt — but they are the same question at two different layers, and it is worth reading them together.

The change. From August 14, new Claude Code sessions on Pro, Max and Team plans start in auto mode, where a classifier reviews each tool call instead of the user approving it. If you have already chosen a different default you get a one-time prompt; if you have pinned a default, nothing changes. The classifier's token overhead stops being billed to Pro, Max and Team users today. It stays opt-in for now on Enterprise, the API, and the AWS, Google Cloud and Microsoft Foundry channels, with the same change planned there within the month.

The number in the headline is 89% versus 13.6%, and the study behind it is worth describing because the design is unusual. Anthropic hired 1,053 paid professional testers through a research vendor, working in a purpose-built test environment rather than their own codebases. Testers knew they were being evaluated but not which behaviour was being measured. Partway through each session a single permission prompt was swapped for a clearly dangerous command, and the vendor recorded whether the tester approved it — approving never ran anything, only the prompt text changed. Humans caught it 143 times out of 1,053. Auto mode blocked the same command 937 times. Head to head, auto mode blocked 800 commands a human waved through; humans blocked 6 that auto mode allowed.

But the finding that actually matters is further down. Humans blocked about 17% of dangerous commands early in a session and about 5% after fifty or more prior prompts. The classifier's block rate stayed flat regardless of session length. That is not a story about people being careless; it is a story about a control that decays with use, measured against one that does not.

Anthropic also published the numbers that make its own previous default look bad, which is the most credible thing in the post. Users approve 97% of permission prompts. Individual permission requests get rejected 3% of the time, while plans presented for approval get rejected 39% — the same people scrutinising one kind of dialogue and reflexively clearing another. As of June 2026, 49.5% of active CLI users had written a Bash allow-rule: 5% allowing any shell command outright and 43% with interpreter rules like Bash(python:*) that amount to the same thing, a share growing about 5 points every 5 weeks. 62% have used bypassPermissions or clicked "don't ask again", and 25% of interactive sessions start in bypass mode. Read together, that is a company reporting that its safety prompt had already been disabled by half its users and rubber-stamped by the rest.

One detail deserves attention from anyone with a settings file: in auto mode, allow-rules broad enough to grant arbitrary code execution are set aside, because they would let commands skip the classifier entirely. Narrower permission rules still fire first. The settings file is not modified and the rules apply again when you switch modes. If your workflow depends on Bash(python:*) doing what it says, it will not do that in auto mode.

Production data, with its sampling stated. From May to June, on sessions the safety pipeline had already flagged and only from accounts opted into data use, Anthropic re-graded outcomes on severity and on how explicitly the user had asked. Manually approved sessions contained serious unintended harm more than twice as often. At production-level severity, 6.3% of manual sessions contained a harmful action the user had not asked for, against 2.4% of auto mode sessions. That is a flagged-session sample, not a random one, and it is the right comparison for this question even so.

Red-teaming, including the part that went badly. Anthropic ran a two-week pilot with Apollo Research injecting synthetic attacks into real coding trajectories. After hardening, Apollo re-tested on a held-out attack set Anthropic never saw: auto mode caught about half the attacks it had previously missed while losing about 2% it had previously caught, and the overall miss rate fell from 12% to 7%. Anthropic says plainly that these attacks are adversarial by construction and that 7% should not be read as a miss rate on real traffic. A vendor publishing its own residual failure rate, and then telling you not to generalise it, is doing this correctly.

The competitive number, stated precisely. A third party, Trajectory Labs, tested 72 indirect prompt-injection scenarios ten times each against the publicly available Claude Code and Codex builds as of July 17. With auto mode on, no attack succeeded against any Anthropic model. GPT-5.6 Sol in Codex v0.144.5's auto-review mode came out at 5.83%, and at 19.03% with full access; on the browser-use subset those figures were 7.87% and 25.53%. The attacks were optimised against Claude Opus 4.7, which is not in the reported set, so every model faced them unseen. It is still a vendor publishing a competitor's number on evaluations of its own choosing, on specific dated builds, and should be read as that.

What we would watch. Anthropic reports auto mode users ship about 25% more PRs among Teams and Enterprise adopters, and that Claude now works 9x longer between interruptions. Those are throughput numbers, and throughput is the thing that made the old prompt fail — the more work an agent does unattended, the more the human's approval becomes a formality. Auto mode is a better control than a prompt nobody reads. It is not the same as a human reading the prompt, and the honest version of today's announcement is that the second option stopped existing a while ago.

Source: Anthropic's announcement + production write-up, read directly · Apollo Research and Trajectory Labs evaluations as reported by Anthropic · the bench index

sweeping every 60 seconds

Know the minute it drops — not the minute we write it up.

Claude Opus 5 went live at 16:51 UTC. The alert was in subscribers' inboxes at 16:52.

  • Free forever
  • No card
  • Unsubscribe in one click

← back to the wire