A checklist-driven auditor catching defects in a supervisor's review A checklist-driven auditor catching defects in a supervisor's review

The cheaper model audited the expensive one and found two bugs

TL;DR On July 31 a fleet of long-running Claude manager sessions that I was using to supervise parallel work collapsed badly enough that I had to cold-start the whole thing. My first hypothesis was that the supervisor model was not smart enough. That hypothesis was checkable, and false. The transcripts showed the managers really were running Opus 5. They still failed. The dominant failure markers in their ledgers were stale and duplicate, which are bookkeeping problems that a script gets right every time and a model gets right most of the time. The same week, a Sonnet auditor at medium effort, given an explicit list of gates to check, found two real defects that the Opus supervisor had missed. The Opus supervisor was better at something else: writing the gate list and noticing what was not on it. I turned that into a table that allocates model and effort per type of turn, with scripts doing the work that is decidable. I am going to be careful about how much that one afternoon proves, which is not much. It is one observation, and I did not run a benchmark. I wrote earlier about sending finished code to a rival model for review. This is the same instinct from the other direction: reviewer and author should not share a blind spot, but they also do not need to share a price tier. ...

September 4, 2026 · 9 min · zolty

Affiliate Disclosure: Some links on this site are affiliate links (Amazon Associates, DigitalOcean referral). As an Amazon Associate, I earn from qualifying purchases. This does not affect the price you pay or my editorial independence — I only recommend products and services I personally use and trust.