TL;DR
- I run a PreToolUse hook that blocks a short list of unrecoverable shell commands before my coding agent can run them. Version 1 was a list of regular expressions over the raw command string.
- Replaying 90 historical blocks and about 82,000 historical commands showed that roughly 40% of v1’s blocks were not commands at all. They were grep patterns, commit messages, and markdown heredocs that happened to contain a scary phrase.
- v2 tokenizes the command the way a shell would and applies each rule only where a program actually runs. On the same replay it allowed every false text match, still caught every real destructive command v1 caught, caught a few that v1 missed because of flag order, threw zero parser errors, and took about 51 ms per call end to end.
- It fails toward blocking. If it cannot parse something it falls back to text rules, and if Python is missing the shell wrapper falls back to the old regex list.
- It is a tripwire for an agent that makes mistakes, not a security boundary against one that is trying to hide. I show a case it does not catch below.
I will write up my experience with Claude Code’s auto-mode classifier separately. The short version is that this hook, not the classifier, is what stands between the agent and my cluster’s worst commands.
Why a guard at all, next to a classifier
The reason matters because it shapes the design. In auto mode, a command that matches a permission allow rule is approved before the classifier ever looks at it. I allow Bash(kubectl *), Bash(terraform *), Bash(git *) and Bash(helm *), because without them I spend my day approving things. That means a classifier, however good, never sees the commands that can do the most damage. And sessions that run in bypass mode, like my scheduled sweeps, have no classifier at all.
A hook is different. It runs before permission rules, in every mode, and it gets the exact command. Exit code 2 blocks the call and sends stderr back to the agent so it can adjust. That is a good place for a small number of rules that are not negotiable.
The rule list is deliberately short:
kubectl deleteof a namespace, node, volume, volume claim, StatefulSet, CRD, or cluster roleterraform destroy,apply -destroy, andapply -auto-approve(applying a saved, reviewed plan file is fine)helm uninstallDROP,TRUNCATEandDELETE FROM, but only when run through a SQL client or database driver- force-pushing
mainormaster rm -rfof/or the home directory
Routine destructive work is deliberately left alone: deleting a pod, clearing a build directory, force-pushing a feature branch. A guard that cries wolf gets removed, and then it protects nothing.
What v1 got wrong
v1 was about eleven regular expressions run with grep -E over the whole command string. It worked on the day I wrote it. Then agents started doing normal things like these:
grep -rn "terraform destroy" docs/
git commit -m "docs: explain why terraform destroy is blocked"
cat > notes.md <<'EOF'
Never run kubectl delete namespace on the shared cluster.
EOF
All three are text that mentions a rule. None runs anything dangerous. v1 blocked all three, because it was classifying the text, not the command. The agent got a block message about a destructive operation it never attempted, and had to work around a regex to commit documentation.
I had a pile of transcripts, so rather than argue about it I measured. I replayed 90 historical blocks and about 82,000 historical Bash commands through both versions.
| Measure | Result |
|---|---|
| v1 blocks that were text matches (grep patterns, commit messages, prose heredocs) | about 40% |
| Real destructive commands v1 caught that v2 missed | 0 |
| Commands v2 catches that v1 missed (flag order, wrapper scripts) | a few |
| Parser exceptions over 82k commands | 0 |
| Time per call, end to end including process start | about 51 ms |
The “few” is real. v1 matched kubectl delete namespace, in that word order. kubectl -n media delete statefulset foo has the flag between the program and the verb, and slipped through. A wrapper script that ran terraform -chdir=... apply -auto-approve slipped through too.
The design: match programs, not text
v2 is one Python file of roughly 900 lines, behind the same shell entry point. It does four things.
1. Pull the heredocs out first
Heredoc bodies can contain unbalanced quotes and arbitrary prose, which breaks a tokenizer. So they are extracted, replaced by a numbered placeholder, and kept so the body can be attached to the command that reads it:
def extract_heredocs(text: str, start: int = 0) -> Tuple[str, List[str]]:
"""Pull heredoc bodies out (they may hold unbalanced quotes) and leave a
`<< __HEREDOC_n__` placeholder so the body stays attached to its command.
A marker with no terminator line is left alone."""
That last sentence is a safety property: if the heredoc is malformed, the guard does not pretend to understand it.
2. Tokenize like a shell, split into simple commands
The tokenizer is the standard library, with punctuation treated as operators:
PUNCT = ";&|()<>\n`"
def tokenize(text: str) -> List[str]:
lexer = shlex.shlex(text, posix=True, punctuation_chars=PUNCT)
lexer.whitespace = " \t\r"
lexer.whitespace_split = True
lexer.commenters = ""
return list(lexer)
A second function walks the tokens and builds a Simple record per command: the leading variable assignments, the argv, any stdin text (heredoc, here-string, or piped), and the redirect targets. Pipes and && split commands. Redirects are not part of argv.
3. Unwrap launchers until you find the real program
sudo kubectl delete ns x is still kubectl delete ns x. So is timeout 30 kubectl ..., env FOO=1 kubectl ..., xargs kubectl ... and a few more. unwrap strips them off in a bounded loop:
def unwrap(argv):
rest = list(argv)
for _ in range(12):
...
prog = os.path.basename(rest[0]).lower()
if prog in ("sudo", "doas"):
rest = skip_options(rest[1:], frozenset({"-u", "-g", "-p", ...}))
elif prog in ("timeout", "gtimeout"):
rest = skip_options(rest[1:], frozenset({"-s", "--signal", ...}))[1:]
elif prog == "xargs":
rest = skip_options(rest[1:], frozenset({"-n", "-I", "-P", ...}))
else:
break
return rest, assigns
The skip_options helper knows which flags take a value, so sudo -u deploy kubectl ... does not mistake deploy for the program.
4. Apply the rule to argv
This is where flag order stops mattering. The kubectl rule asks for positional arguments only, skipping flags and their values, then checks the verb and the resource kind:
def kubectl_rule(args):
pos = positionals(args, KUBECTL_VALUE_FLAGS)
if len(pos) < 2 or pos[0] != "delete" or _is_dry_run(args):
return False
kinds = [part.split("/")[0] for part in pos[1].split(",")]
kinds += [target.split("/")[0] for target in pos[2:] if "/" in target]
return any(kind.lower() in PROTECTED_KUBE_KINDS for kind in kinds)
positionals is why -n media between the program and the verb no longer helps. It also handles TYPE/NAME forms, comma-separated kinds, and short names like ns and sts. A --dry-run that is not none exempts the call.
Terraform is the same idea on flags instead of positions:
def terraform_rules(args):
flags = [a for a in args if a.startswith("-")]
pos = [a for a in args if not a.startswith("-")]
names = {f.lstrip("-").split("=", 1)[0]: f for f in flags}
...
if verb == "destroy":
return ["terraform-destroy"]
if verb != "apply":
return []
hits = []
if "destroy" in names and not names["destroy"].endswith("=false"):
hits.append("terraform-destroy")
if "auto-approve" in names and not names["auto-approve"].endswith("=false"):
hits.append("terraform-auto-approve")
return hits
So -chdir=infra apply -auto-approve is caught, apply -auto-approve=false is not, and apply tfplan is fine.
Following the command into other places
A lot of damage hides one level in. The scanner recurses, with a depth limit of six, into the places a command can run another command:
bash -c '...',eval, and scripts fed through stdinssh host 'remote command'kubectl exec ... -- cmdanddocker exec ... cmd$(...)and backticks inside words- heredocs or pipes into a shell, a SQL client, or an interpreter
- inline interpreter code that also calls a database driver
- a script written to a file earlier in the same call and then run
SQL gets one extra rule, because DELETE FROM is common in prose and code. It only counts when it is passed to a SQL client (psql, sqlite3, mysql and friends) or to inline code that also calls an .execute()-style driver method. Code that merely mentions SQL does not match.
A few real checks
I ran these through the current file:
| Command | Result |
|---|---|
grep -rn "kubectl delete namespace" docs/ | allowed |
git commit -m "docs: explain why terraform destroy is blocked" | allowed |
echo "DROP TABLE x" > notes.md | allowed |
kubectl delete pod foo-123 | allowed |
kubectl delete ns scratch --dry-run=client | allowed |
kubectl -n media delete statefulset jellyfin | blocked, kubectl-delete |
sudo timeout 30 kubectl delete ns scratch | blocked, kubectl-delete |
ssh host "kubectl delete pvc data-0" | blocked, kubectl-delete |
terraform -chdir=infra apply -auto-approve | blocked, terraform-auto-approve |
a heredoc of DELETE FROM lots; piped into psql | blocked, sql-destructive |
git commit -m "kubectl delete namespace docs" && kubectl delete namespace foo | blocked (only the second command) |
That last row is the one I like. The commit message is ignored and the real command after && is caught.
Failing safe
Parsers have bugs, and a guard that fails open is worse than none. There are three layers:
def evaluate(command: str, mode: Optional[str]) -> Decision:
analyzer = Analyzer()
note = ""
try:
analyzer.scan_shell(command, 0, leading_hatch(command))
except Exception as error: # never fail open on a parser bug
analyzer = Analyzer()
analyzer.text_scan(command, leading_hatch(command), sql=True)
note = f"(guard parse error: {type(error).__name__}; used the text fallback)"
return decide(analyzer.findings, mode, note)
- Unbalanced quotes or recursion past the depth limit send that piece to the text rules. The text rules still require the verb to follow the program, so prose does not match, but a malformed command does.
- Any exception in the analyzer falls back to the text rules over the whole command.
- The shell wrapper that settings point at runs the Python file with
python3 -I. If Python is missing, or exits with anything other than 0 or 2, the wrapper runs the original regex list over the raw payload. Over-blocking is the correct direction for a safety guard to fail.
Zero exceptions over 82,000 real commands is encouraging. I would not conclude from it that the parser cannot throw, which is why the fallbacks exist.
The override, and why it is awkward on purpose
Sometimes I really do want to run a destructive thing. The escape hatch is a prefix on the command:
CLAUDE_ALLOW_DESTRUCTIVE="reason naming the operation and who approved it" <command>
A few properties make it more than a bypass:
- The reason must be at least four characters and contain a letter. It is plain text in the command, so it is visible in any approval prompt.
- The prefix defeats allow-rule matching, because a command that begins with a variable assignment no longer matches
Bash(kubectl *). That routes the call to the classifier in auto mode, or to me in default mode. The hatch does not skip review. It chooses a reviewer. - In bypass mode there is no reviewer, so the override is refused outright and the agent is told to hand me the command.
- Every use and every block is logged with the permission mode and session id.
What it does not do
I want to be clear about the limits, because “guard” invites more trust than it has earned.
The rules are about what the command line says. Anything the parser cannot see through is out of reach. For example, I tested piping a base64 string through base64 -d | sh, and the guard allowed it. A variable that holds the program name is followed when it was set earlier in the same call, but not across calls. A script file that exists on disk but was not written in the same call is not read.
An agent that is confused will almost never do these things. An agent trying to route around a rule might. If that is your threat model, you need a different layer, such as scoped credentials or a sandbox, rather than a better regex.
I also dropped a rule. v1 blocked git reset --hard origin/.... In the replay it fired 12 times and every one was a benign sync of a scratch clone. A rule that blocks harmless work gets muted in your head, and I would rather have eleven rules I believe in.
Tests
The test file is about 230 lines with 17 tests, including one that runs the real shell wrapper end to end. The test names read like the sections above: phrases in text are not commands, routine destructive work passes, kubectl kinds block in any flag order, nested ssh and exec forms block, the override must carry a real reason, the override is refused in bypass mode, and both unbalanced quotes and a parser crash fall back to the text rules.
My CI also type-checks and compiles the hooks directory now, which the shell version never needed.
Takeaways
- If your guard works by searching text, measure it against your own history. Mine was wrong in the annoying direction 40% of the time, and I had not noticed because I was not the one reading the messages.
- Match where programs run, not where words appear. A tokenizer from the standard library gets you most of the way.
- Keep the rule list short enough that every rule has a reason you can state.
- Write down what it cannot see. Then do not call it a sandbox.
- When you rewrite a safety check, replay your history through both versions. The false positives are the motivation. The false negatives are the thing you have to prove did not get worse.