Rewriting my agent's destructive-command guard to parse, not grep
TL;DR I run a PreToolUse hook that blocks a short list of unrecoverable shell commands before my coding agent can run them. Version 1 was a list of regular expressions over the raw command string. Replaying 90 historical blocks and about 82,000 historical commands showed that roughly 40% of v1’s blocks were not commands at all. They were grep patterns, commit messages, and markdown heredocs that happened to contain a scary phrase. v2 tokenizes the command the way a shell would and applies each rule only where a program actually runs. On the same replay it allowed every false text match, still caught every real destructive command v1 caught, caught a few that v1 missed because of flag order, threw zero parser errors, and took about 51 ms per call end to end. It fails toward blocking. If it cannot parse something it falls back to text rules, and if Python is missing the shell wrapper falls back to the old regex list. It is a tripwire for an agent that makes mistakes, not a security boundary against one that is trying to hide. I show a case it does not catch below. I will write up my experience with Claude Code’s auto-mode classifier separately. The short version is that this hook, not the classifier, is what stands between the agent and my cluster’s worst commands. ...