TL;DR
- Between late June and late September I moved my daily setup through Sonnet 5, Opus 5, Fable 5.1, Opus 5.5 and Sonnet 5.5. Haiku stayed at 4.5 the whole time. I never saw a Haiku 5 in any log.
- There is no controlled benchmark in anything I kept. What follows is operational evidence: what I changed, what broke, what I verified, and what I would have gotten wrong if I had trusted what a model told me about itself.
- The changes that mattered were small: a per-model effort block in
settings.jsonand a subagent default model. The expensive lessons were two traps. One is a bug where background-job subagents silently ran Haiku 4.5. The other is that a worker’s report of its own model cannot be believed. - If you want the earlier, more general version of this, see re-tuning my setup for a new Opus model. This is what the next five releases taught me.
The ladder
These are first-seen dates from my own records: commit trailers in my repositories and, where I still have them, session transcripts. They are when I started using each model, not necessarily when it became available.
| Date | Model | What I know |
|---|---|---|
| Jun 30 | Sonnet 5 | First commit trailer. Fable 5 had already shipped earlier in June. |
| Jul 24 | Opus 5 | First commit trailer; I ran /model opus that afternoon. |
| Sep 1 | Fable 5.1 | First seen in transcripts; first commit trailer Sep 2. |
| Sep 22 | Opus 5.5 | First commit trailer; first transcript Sep 24. |
| Sep 29 | Sonnet 5.5 | I was checking for it that morning; first trailer same day. |
The commit trailers give a rough sense of how I actually used them. Opus 5 dominates, with several hundred trailers across my repositories. Opus 5.5 collected about 200 in its first nine days. Fable was something I reached for with an explicit /model fable and then went back to Opus. It never became the default.
The share of tokens tells the same story from another angle, with a caveat. The stats file counts cache reads, which dominate, so treat it as a usage mix rather than anything about cost. Opus 5.5 was about 15% of September’s tokens even though it only arrived on the 22nd. The cutover was fast, because I had little to change.
What I changed in settings
My settings.json is not under version control, so the history I have is a handful of backup files. Reading across them, the model-related changes were these.
A per-model effort block
Claude Code has a modelSettings key that maps a model ID to its own effort level. Before September I had one global effortLevel of high and every model inherited it. By September 4 I had added an entry for Fable, which I keep at xhigh since I only reach for Fable on purpose:
"effortLevel": "high",
"modelSettings": {
"claude-fable-5": { "effortLevel": "xhigh" },
"claude-opus-5": { "effortLevel": "medium" },
"claude-opus-5-5": {}
}
That is my most recent backup, from September 27. The path to it was not straight. Opus 5.5 got xhigh when it arrived and was back to unset within a few days, which means it inherits the global high. Opus 5 had been moved down to medium by then. I would love to tell you there was a clean experiment behind each move. There was not, and the point of the block is that a new model gets its own dial instead of silently inheriting the last model’s. I also tried the very top effort tiers a few times and walked back to medium or high within the day.
The policy underneath is in my global instructions, and it is a rule rather than a list of models: routine implementation on Sonnet at medium, bounded extraction on Haiku at low, diagnosis and plans on Sonnet or Opus at high, and security, privacy, financial, migration, merge and deployment judgment on Opus at high. Use a higher tier only for a named question the normal pass could not resolve, and come back down after the hard reasoning checkpoint. I wrote about why I stopped listing model IDs in that file in the CLAUDE.md diet.
The subagent default
Subagents get their own default through an environment variable in settings, CLAUDE_CODE_SUBAGENT_MODEL. In my July review it was pinned to Haiku 4.5. By early August it was Sonnet 5, and it was still Sonnet 5 in the September 27 backup, two days before Sonnet 5.5 showed up. That follows the table: most subagent work is bounded (read these files, run this check, draft this patch), and bounded work is what Sonnet at medium is for.
Separately, my own agent definitions pin a model in their frontmatter. The planner, the auditor and the coordinator are on Opus. The builder, debugger, researcher and tester are on Sonnet. The ones that decide are expensive. The ones that execute are not.
The advisor
There is an advisorModel setting for the second-opinion model. It has been Opus in every backup from early August through September 27. I tried Fable and Opus in that slot on the same night in late July, about ten minutes apart, and kept Opus. I do not have a measurement behind that, only a preference: the advisor sees a whole conversation and gets asked about a named hard problem, which is the kind of call I want a top-tier model on. Whether that should be Opus or Fable is a question I keep reopening.
For the standalone Fable advisor calls in my supervision tooling, I follow fixed rules: read-only, plan permission mode, only the Read tool, a redacted prompt, one follow-up at most, and I record the model that actually answered. The advisor never approves a merge, a deploy, a security decision or a secret operation. It informs.
What I did not change
DISABLE_NON_ESSENTIAL_MODEL_CALLS and the compaction threshold stayed put across every backup. The default model alias stayed Opus. Headless automations always pin a model explicitly and fall back to Sonnet, so a release does not silently move scheduled work.
Trap 1: subagents in background jobs ran Haiku
On July 1 I tried to run a small fleet of subagents on the top-tier model from a background job. Every one of them ran as Haiku 4.5. Nothing errored.
I tested it eight ways before believing it: asking for different models through the workflow API and the Agent tool, across two agent types, and then the decisive one, an agent whose own frontmatter pinned Opus. Still Haiku. After switching the session’s main model, a subagent with no override still reported Haiku.
I confirmed it again on July 5 and July 11. The workaround for a while was to stop using the Agent tool from background jobs, and run a headless claude -p --model opus --fallback-model sonnet per task in its own git worktree.
On August 1 I narrowed the scope. The bug was specific to Agent-tool subagents. A full background CLI child launched with claude --bg --model opus resolved to Opus 5, and one with --model sonnet resolved to Sonnet 5. I tested it at the top tier on purpose, because that is the tier the original bug degrades, and I did not want to infer it from a Sonnet result.
The fix is not a fix. It is a check. The model each session actually used is recorded in its transcript, and this is the command I run before I count any worker:
TS=$(find ~/.claude/projects -name "<session-id>*.jsonl" | head -1)
jq -r 'select(.message.model) | .message.model' "$TS" | sort -u
If you run any kind of multi-agent setup, run that against a few of your sessions. Silent fallback to a cheaper model produces output that looks fine, which is exactly why it survives.
I will note that this was a bug in a specific build, in a specific context, and I have not retested it on every release since. Treat it as something to verify, not as a standing fact about the product.
Trap 2: a model cannot tell you what model it is
On August 1, an auditor worker finished its report with a line reading MODEL USED: claude-opus-4-8. Its transcript said claude-sonnet-5 on 32 of 32 messages. It was not lying. It was wrong, and confident about it.
There is a wrinkle I should be honest about. My Haiku probe, where I ask a subagent to state what model it is, has identified itself reliably in the probes I have run. So a self-report is not always wrong. The lesson is narrower and more useful: a self-report is not evidence. Sometimes it matches, and you have no way to know from the report whether this is one of those times.
So I changed two things. The transcript’s message.model field is the only source I accept. And I removed MODEL USED: from my worker contracts, so nobody is tempted to read it. In the supervision procedure the launch check has four parts, and the resolved model is the third: the live process, the prompt reaching the worker, the model in the transcript, and activity past initialization.
What I would tell you, without a benchmark
I do not have a table of “Opus 5.5 scored X against Opus 5”. I would rather not invent one. What I can say from operating experience:
- Per-model settings beat one global dial. Every release arrived with its own idea of how much effort was worthwhile. A block keyed by model ID is cheap and removes the guesswork.
- Put model choice in rules, not lists. My upgrades got less work the day I stopped writing model IDs into instructions. The rules say what kind of work gets what kind of model. The CLI says what exists.
- Verify the model from the transcript. The bug and the self-report problem both disappear if you look at the one field that records what ran.
- Keep the cheap model cheap on purpose. Subagents default to Sonnet. Haiku 4.5 is still my choice for bounded judging jobs, like the little evaluator that decides whether a long-running goal is done. Newer did not mean I had to move everything.
- Expect to flip settings back. At least two of my effort changes were reversals. That is fine, as long as the settings file is small enough that you can see what you did.
The part I find most interesting is that none of this was about capability. In another post I look at a case where an Opus-level fleet failed for reasons that had nothing to do with how smart the model was: the cheaper model audits the expensive one.