TL;DR

  • Claude Code’s auto permission mode puts a classifier between the agent and any tool call that is not already covered by an allow rule. I ran in that mode all summer.
  • In the transcripts I still have (mid-August to early October), 301 tool results mention that classifier. About 288 are outright denials, about ten are “no verdict” errors, and a few just quote the phrase. Of roughly 57,700 tool calls in September, that is about 0.5%.
  • Almost all of it landed in September (285 of the 301), and it peaked on September 6 with 29 in one day.
  • Two things helped: adding environment text that tells the classifier what my infrastructure is, and adding narrow allow rules so routine commands never reach it. Neither one was a clean win, and I will show where the data does and does not support that.
  • Sometime between late September and early October I switched the default mode to bypass. The thing that makes that tolerable is a deterministic hook that runs before permission rules in every mode. The classifier was not wrong often. It was expensive in attention, and it was the wrong tool for unattended work.

This is part of a small batch about the classifiers I run around agents. The rewrite of my destructive-command guard came first, and a field guide to all of them will follow.

What auto mode is

Claude Code has a few permission modes. In the default one it asks you before anything risky. In bypass mode it asks nothing. Auto mode sits in between: calls that match an allow rule go straight through, and everything else is judged by a model-based classifier that decides whether the action is safe given what the session is doing.

When it says no, the agent sees a result like “Permission for this action was denied by the Claude Code auto mode classifier”, a short reason, and some advice to carry on with other work or try a different tool. Occasionally it says it cannot decide at all, because the classifier call itself failed.

I liked the idea. I run a homelab, a work bench lab, and a pile of side projects, I am the only operator, and I let agents work in parallel. A second opinion on each unusual command that does not need me to be at the keyboard is exactly the shape of thing I want.

The numbers

I counted tool results that mention the classifier in every main-session transcript Claude Code still has on disk. I left out subagent transcripts, which would add roughly a hundred more. Transcripts only go back to mid-August, so this is not the whole summer, and the first half of the story is missing.

MeasureValue
Tool results mentioning the classifier301
Explicit denials288
“No verdict” errorsabout 10
By monthAug 4, Sep 285, Oct 12 (through Oct 6)
Tool calls in Septemberabout 57,700 (36,450 of them Bash)
September denial rateabout 0.5% of all tool calls

The tools it blocked:

ToolDenials
Bash238
GitLab merge-request merge17
Agent (spawning subagents)7
SendMessage7
Browser automation (batch, JavaScript, navigate, upload)10
Everything else (Skill, Edit, Write, scheduling, a few more MCP calls)the remainder

Per week, counting from September 1: 98, 94, 53, 31, 19. The worst single days were September 6 (29), September 8 (24), and a three-way tie around 20 for the 15th, 16th and 27th.

Two caveats before anyone quotes me. These are result strings, so a retried call counts more than once. And I did not label each denial as right or wrong, so I cannot give you a false positive rate. What I can tell you is what kinds of things were blocked.

What got denied

The denial text sometimes carries a short category tag. About 60% of them had none, just “Blocked by classifier”. The tagged ones read like a taxonomy of things an overcautious reviewer worries about: merge without review, auto-mode bypass, modify shared resources, permission grant, security weaken, production reads, unauthorized persistence, credential materialization, remote shell writes, secret-store writes.

Reading through a sample of them, sorted by what the command was actually doing:

  • Scripts that apply changes for real. Scripts run with --apply instead of in dry-run, and a cloud CLI call against an account profile the classifier had no reason to trust yet.
  • Credential plumbing. Fetching a token from the secret manager into a shell variable and calling an API with it. Writing a temporary token file with a restrictive umask, which is exactly what you are supposed to do. Reading a different project’s transcripts to look for how a credential was handled.
  • Writes through a remote shell. A helper script that runs a command inside a container on another machine. The classifier sees “writes on a remote box” and cannot see the helper is mine.
  • Merges. Marking a merge request ready and merging it in the same breath, and several merge calls through the GitLab MCP server. This was the biggest single named tool.
  • A few that looked wrong to me. One denied a plain pytest run inside a project virtualenv.

Some of these were correct. In July the classifier refused to let a session merge a merge request that the same session had authored, and I agreed with it. Later it refused to let me spawn workers with the bypass flag because that “creates unsafe agents”, which is also a fair description.

The pattern that made it expensive was not any one denial. It was that a long autonomous run hits one, stops, and waits for me. Most of the cost was latency from an agent that had stopped working while I was elsewhere.

Fix one: tell the classifier where it is

Out of the box the classifier trusts the working directory and its git remotes. Everything else is a stranger. A command that SSHes to a machine on my LAN looks the same as one that SSHes to someone else’s server.

You can add context in your user settings under autoMode, with an allow list and an environment list. Project settings are ignored for this, which makes sense, because otherwise a hostile repo could vouch for itself. Both lists keep a "$defaults" entry so you extend the built-in rules instead of replacing them.

I wrote mine as plain English paragraphs, one fact per entry: who the operator is, what the machine is, where source control lives, which domains and networks are mine, which accounts the agent is expected to use, how secrets are handled, and which things are genuinely high blast radius. Two of the allow entries, scrubbed of hostnames, give the flavor:

"autoMode": {
  "allow": [
    "$defaults",
    "Pushing any branch and opening or updating merge requests on the self-hosted GitLab is allowed. Merging to main is not.",
    "kubectl apply, rollout restart, scale, port-forward, exec, logs, and delete of namespaced pods/deployments/jobs/configmaps on the home and work clusters is allowed; the deterministic hook separately blocks deletes of namespaces, nodes, PVs/PVCs, StatefulSets, CRDs and cluster roles."
  ]
}

I keep the text in a shared repo with a small script that merges it into settings.json, backs the file up first, and can be re-run. Mine ended up as about a dozen allow entries and fourteen environment entries on top of the defaults.

Did it work? Honestly, I cannot show that it did. I applied it on September 4. The denial counts for the next two weeks are the highest of the whole period. A running session also does not pick up new rules until it restarts, and my mix of work in mid-September involved a lot of remote operations that the classifier is right to be suspicious of. The one thing I can say is that it did not make things worse in any way I can measure. I would still write the text, because it is the only way to teach the classifier about things like “this machine is mine and passwordless sudo is expected”.

Fix two: do not send routine commands there at all

The documented lever that does work is narrow allow rules. A rule like Bash(grep *) is evaluated before the classifier, so those calls never see it. This also is the only defense against the classifier being unavailable, because there is no call to fail.

I kept adding them whenever I saw a routine read-only command in a denial. The allow list went from 27 entries in early August to 38 on September 4, 72 by September 12, and 74 now. The additions were boring on purpose: cat, grep, sed, head, tail, jq, sort, wc, ls, find, date, echo, and the like.

Weekly denials fell from 94 to 53 across the same window in which that list went from 38 to 72. That is correlation, and the work mix changed too, but it is the fix I believe in.

The catch is that an allow rule is a hole in the safety net. I also have Bash(kubectl *), Bash(terraform *), Bash(git *) and Bash(helm *) in there, because without them half my day gets interrupted. Those commands skip the classifier completely, including the dangerous subcommands. I will come back to that.

What rules cannot fix

Three classes of denial stayed no matter what I wrote:

  1. Unattended sessions. A background builder launched without a human attached cannot ask for permission, and auto mode degrades to deny. I found this in August when a headless worker could not edit a file, redirect output, or even touch a file in /tmp. The workaround is acceptEdits for those lanes.
  2. Spawning workers with bypass. Asking a child process to run in bypass mode gets denied. A scoped --allowedTools list passes. That is a fair rule, and one I now follow.
  3. Merges. Merging an agent-authored merge request was denied nondeterministically. The same call would pass one day and fail the next.

Why I moved to bypass

Sometime between September 27 and October 5 I changed permissions.defaultMode from auto to bypassPermissions. I do not have the exact day, because ~/.claude is not under version control and my only history is settings backups. The last backup with auto is from September 27, and the file in place by October 5 says bypass.

My reasons, in order of weight:

  • It was the wrong tool for most of my work. More and more of what I run is unattended: scheduled cluster sweeps, a responder that polls a chat bridge, headless workers. The sweeps and the responder already ran in bypass mode, so they never had the classifier anyway. Keeping interactive sessions in a different mode just meant different behavior for the same kind of work.
  • Attention cost. Even at 0.5% of calls, a denial is the thing that stops a ten-minute autonomous run for an hour.
  • It was not the only line of defense, and for my riskiest commands it never was. Because of the allow rules above, kubectl delete never reached the classifier. What was actually protecting me from those was a hook, not a model.

What I lost, plainly: a second opinion on the unfamiliar. The categories above (persistence, credential exploration, permission grants) are things I do want reviewed, and a bypass session will not review them. What replaces that is mostly habit and written rules: work happens in branches, merging to main is still a step I reserve for myself, and credentials go through a wrapper that never prints the value. None of those is enforced by a gate the way the hook is, and I would rather say so than imply otherwise.

The backstop

The piece I did not give up is a PreToolUse hook that inspects every shell command before it runs and blocks a short list of unrecoverable operations: deleting a namespace, node or volume, destroying infrastructure, applying infrastructure changes without a reviewed plan, force-pushing the main branch, and a few others. It runs regardless of permission mode, and it runs before allow rules. In bypass mode it also refuses its own override, since nothing would review it.

That hook had a problem of its own: it was matching text instead of commands, and it cried wolf on innocent strings. I rewrote it to match parsed commands, not text.

What I would tell you to do

  • If you run auto mode, count your denials before tuning anything. Search your transcripts for the denial phrase and group by tool. Mine pointed at Bash and merges, which told me where to spend effort.
  • Write the environment text, but do not expect it to be the fix. It teaches the classifier about your setup. Measure afterward.
  • Add narrow allow rules for commands you are already comfortable with. Then ask what the rule gives up. Bash(kubectl *) is convenient and also removes review for every subcommand.
  • Put your real never-do-this list in a deterministic hook, not in a model. A hook is boring, testable, and fast. A classifier is probabilistic, and mine was unavailable ten times.
  • Decide per workload, not per machine. Unattended lanes and interactive sessions want different modes, and I should have admitted that sooner.

I earlier wrote about building up trust for agents in steps, in the autonomy ladder and how it went in practice. Auto mode was a rung on that ladder, and I am no longer standing on it. The next rung was a short list of hard rules, enforced by a script that cannot be talked out of anything.