TL;DR
- On July 31 a fleet of long-running Claude manager sessions that I was using to supervise parallel work collapsed badly enough that I had to cold-start the whole thing. My first hypothesis was that the supervisor model was not smart enough.
- That hypothesis was checkable, and false. The transcripts showed the managers really were running Opus 5. They still failed. The dominant failure markers in their ledgers were
staleandduplicate, which are bookkeeping problems that a script gets right every time and a model gets right most of the time. - The same week, a Sonnet auditor at medium effort, given an explicit list of gates to check, found two real defects that the Opus supervisor had missed. The Opus supervisor was better at something else: writing the gate list and noticing what was not on it.
- I turned that into a table that allocates model and effort per type of turn, with scripts doing the work that is decidable. I am going to be careful about how much that one afternoon proves, which is not much. It is one observation, and I did not run a benchmark.
I wrote earlier about sending finished code to a rival model for review. This is the same instinct from the other direction: reviewer and author should not share a blind spot, but they also do not need to share a price tier.
What I was running
I supervise a lot of parallel Claude Code sessions. A “manager” in my setup is a long-lived session that owns a lane of work: it writes briefs for worker sessions, launches them, watches them, and records what happened in a ledger file. A fleet is several managers at once, each with its own ledger.
On July 31 the fleet fell apart. The ledgers filled with stale and duplicate rows, two managers had churned through to their thirteenth and fourteenth generations, and what was left could not be resumed. The recovery was a cold start.
Was the model the problem?
The obvious story is that the supervisor was not capable enough, and the obvious fix is a bigger model. I did not want to buy that without checking, because there was an easy way to check.
The manager’s manifest says what model it was supposed to be running, but a manifest line is a self-report, and I already had a reason not to trust self-reports. So I went to the authoritative source, which is the message.model field on each assistant turn in the session transcript.
| Manager session | Assistant messages | What actually ran |
|---|---|---|
| Manager A | 60 | 59 on Opus 5 |
| Manager B | 43 | 42 on Opus 5 |
| Manager C | 134 | 133 on Opus 5 |
| Manager D | 243 | 174 on Opus 5, 67 on Sonnet 5 |
| Manager E | 1 | a synthetic placeholder only, never ran |
| Manager F | 1 | a synthetic placeholder only, never ran |
(Counts are as I recorded them. A message or two per session carried no resolved model.)
Three results came out of that, and only the first one was what I was looking for.
- The managers were Opus 5. Four of the six really ran. Capability was not the cause of a collapse that happened under a top-tier model.
- Two of my six “managers” never ran at all. They had the same display name, one synthetic message each, and sat in the registry as blocked rows. “It was backgrounded” is not evidence that anything started.
- One session drifted. Manager D ran 174 turns on Opus and 67 on Sonnet inside a single session. A supervisor that records “Opus, high” once at launch is not describing what ran.
What actually failed
I counted outcome markers across all 16 manager ledgers from that fleet:
| Marker | Files | Hits |
|---|---|---|
BLOCKED | 14 | 80 |
stale | 15 | 61 |
ESCALATE | 12 | 56 |
duplicate | 10 | 47 |
FAILED | 12 | 30 |
These are keyword counts, not an incident taxonomy, and a keyword in a ledger is not always a failure. Still, the shape was clear. stale and duplicate were the two I could attribute to a specific kind of mistake: is this row live, is this lane owned twice, is this commit current. Those are registry and ledger hygiene. They are decidable facts, and they were being decided by Opus sessions.
That is a different problem from the one I thought I had. A model can look at a registry row and judge it plausible. A script can check that the process id exists and that no two lanes share a name. At fleet scale, “gets it right most of the time” is the same as “never”.
So the first change had nothing to do with models. I already had a snapshot script that computes an attention flag for the fleet. The gap was that nothing forced a manager to run it. Now every watch pass begins with the snapshot, a row with no live process id is historical whatever its state column says, and liveness is a four-part check against the transcript rather than a message that says “backgrounded”. Duplicate display names are, by the rule I wrote down, an error instead of a judgment call.
If you take one thing from this post, take that. When your agents fail at bookkeeping, take the bookkeeping away from the agent.
The audit
The second finding is the one in the title. The same week I had a plan to review and two reviewers available: the Opus supervisor, who was running the session, and a Sonnet auditor at medium effort, launched fresh with an explicit list of gates to check.
What the Opus supervisor contributed:
- It noticed that a required gate was missing from the plan. Not wrong, absent. A checklist verifies what is on it. Only judgment catches what is not.
- It inferred that removing one reviewer from the process would not hand that reviewer’s authority to the supervising model. The authority collapses back to me. Nothing in the text said so.
- It caught the auditor claiming to be a different model than its transcript showed.
- It declined to dispatch a writer into an infrastructure lane without a decision from me.
What the Sonnet auditor, with a gate list, contributed:
- It discharged seven gates, each with file and line evidence.
- It found two real defects that the supervisor had missed: a document that contradicted itself about which service owned a component, and a deploy job that was missing a preflight check.
- It graded each gate pass or fix and honestly flagged where its own verification was limited.
- It got its own model identity confidently wrong. In its final report it said it was a different, older model. Its transcript said Sonnet 5 on 32 of 32 messages.
The split was clean enough to turn into a rule. The Opus supervisor’s edge was writing the gate list and noticing absences. The Sonnet auditor’s execution against that list was more productive than the supervisor’s own review of the same artifact. Specification was the scarce capability. Execution against a good specification was not.
What this does not prove
This is one afternoon, one plan, one auditor, and a bounded gate list. The comparison was not controlled: the two reviews were not held to identical conditions, and I did not repeat the run. A different artifact could have gone the other way. The claim I am willing to make is narrow: an explicit gate list with file-and-line evidence requirements is a good fit for a mid-tier model, and the expensive model’s attention is better spent on the list than on re-doing the checking.
Effort is not a free dial either
I also looked at how much effort buys. I did not run my own study. In the published judge-model studies I read, for model families that are not Claude, accuracy as a judge climbed in a consistent shape as reasoning effort went up: one small reasoning model went from 70.6% at low to 76.6% at medium to 80.9% at high, and a small open model went from 79.9% to 86.0% to 89.3%. The step from low to medium bought about six points. Medium to high bought three or four.
I would not quote those numbers as Claude numbers, and I do not. I take the shape: the first step up buys more than the second, and the headroom shrinks as the base model gets stronger. That points two ways.
- Never run a supervisory turn at low effort. That is the expensive cut.
- Medium is defensible on a strong model for routine watching, and high is for the named judgment moments.
The table
This is what I ended up with. It is in my supervision skill, and I use it as the default per-turn allocation.
| Type of turn | Model and effort |
|---|---|
| Write the gates, a worker contract, or review a plan | Opus, high |
| Boundary judgment: authority, merge or deploy, escalation | Opus, high |
| Routine watch pass or heartbeat | Opus, medium |
| Bounded audit against an explicit gate list | Sonnet, medium |
| Mechanical extraction or status | Haiku or Sonnet, low, or a cheap external model |
| A named ambiguity that remains after a high pass | a premium-tier advisor, once, then come back down |
| Liveness, duplicates, staleness | a script, no model |
The rules that go with it:
- Escalate for a named uncertainty, not because the task feels important.
- Never run supervisory turns at low effort.
- Change setting only at a safe boundary: before dispatch, after a checkpoint, or after pausing a worker. Never mid-edit, mid-migration, or mid-test.
- Record the model, the effort, the reason, and when I will reassess.
- If the surface cannot switch a live session, start the next bounded session instead. Never claim a switch took effect when it did not.
- When Opus quota is tight, move routine work down first. Do not downgrade the final gate.
- Track what ran from the transcript, not what was declared.
How those settings live in my config, and the traps in what transcripts report, are a separate post on upgrading through the Claude 5 line.
What I would do again, and what I would not
I would do again:
- Verify the model from the transcript before forming a theory about why a model failed.
- Move decidable checks out of the model entirely.
- Hand a mid-tier model a gate list and require evidence per gate.
- Spend the top tier on writing the list and on noticing what is missing from it.
I would not claim:
- That a cheaper model is generally better at auditing. One auditor found two defects in one review.
- That the effort curve I quoted applies to Claude. It is a shape from other model families.
- That scripts are enough. The Opus supervisor caught things no checklist contained, and that is exactly the job I want it to keep.
The pattern I trust is boring. A script decides what can be decided. A mid-tier model checks a specific list and shows its work. The expensive model writes the list and judges the boundary. Capability is rarely the bottleneck. Specification and bookkeeping are.