TL;DR
- On September 6 I audited every file Claude Code loads at the start of a session on my main machine. The global
CLAUDE.mdwent from 19,286 bytes to 7,607 (down 60.6%). The four biggest project memory indexes went from 64,900 bytes combined to 6,513 (down 90.0%). - The file had more than doubled in under three months. A backup from June 18 was 9,242 bytes. Every addition had seemed reasonable when I made it.
- The biggest single category of cuts was model information: release lists, retirement dates, introductory pricing, spend estimates. I replaced all of it with a rule, “use only the models and efforts the active CLI offers”, and kept the judgment about which tier fits which task.
- The honest caveat: this is a document check. I verified that files parse, links resolve, and the constraints I care about survived. I did not benchmark whether Claude behaves better or worse, and I did not take a token count from a live session. Bytes are not tokens.
If you have read my earlier post on re-tuning my setup for a new model, this is the sequel with a different lesson. That time I learned not to trust a model’s suggestions about what to delete. This time I put a validator and a rollback around the edits.
How a config file doubles
Nothing about a global instruction file pushes back when you add to it. A new model ships, so you add a section about it. An incident happens, so you add a rule so it cannot happen again. A project gets complicated, so its status goes into the memory index “so the next session has context”. You never delete anything, because deleting feels like losing information, and the cost of an extra paragraph is invisible.
The cost is real, though, and it has two parts. Everything in a startup file competes for attention with the task in front of the model. And everything in it gets stale at a different rate. A rule like “never print secrets” is true for years. “As of June, this model costs X” is wrong within a quarter, and it stays there looking authoritative.
What I had by September was a file full of the second kind.
What Claude actually loads
Before cutting anything, I needed to know what counts. Claude Code concatenates the global file, any ancestor directory files, and the project file. The per-project auto memory loads only the first 200 lines or 25 KB of its MEMORY.md index at the start. The topic notes the index points to are read on demand, not at startup. Plain Markdown links to other files do not import them. Only an @ import pulls the target in eagerly.
That gave me a design with three layers:
- Global
CLAUDE.md: durable cross-project behavior only. - Memory indexes: a short retrieval table that points at notes. Not a log.
- Everything else: topic notes, references, archives. Read when relevant.
I also added no new @ imports, because an import is a way of loading optional material at startup while pretending you did not.
The numbers
Sizes are UTF-8 bytes of the edited files. They are not token counts.
| Entry point | Before | After | Change |
|---|---|---|---|
Global CLAUDE.md | 19,286 | 7,607 | -60.6% |
| Workspace root memory index | 17,268 | 1,625 | -90.6% |
| Resale project memory index | 18,419 | 1,513 | -91.8% |
| Cluster project memory index | 12,248 | 1,617 | -86.8% |
| Fourth project memory index | 16,965 | 1,758 | -89.6% |
| Four indexes combined | 64,900 | 6,513 | -90.0% |
The worst offenders by size were the topic notes the indexes pointed to. The largest was 109,663 bytes. Its default read is now 1,551 bytes: current constraints and a link to the archived history. Three others went from roughly 31 to 80 KB down to about 1.4 to 1.6 KB each. The full old content is still in an _archive folder next to each note.
Global instructions plus the largest applicable memory index shrank by roughly 71% to 76% together. That applies to those files only. The parent workspace file, repo instructions, tool definitions, plugin text, and the conversation itself are all still context on top of that.
What I cut
Model information. The old “Model Strategy” section had a dated list of which models existed and which were retired, an introductory pricing note that had since expired, and fixed cloud-spend estimates. All of it was a snapshot of one week. It is gone, and I am deliberately not reproducing the numbers here either. What replaced it is a rule and a small policy:
Use only models/efforts surfaced by the active CLI. Reassess at start, checkpoints, scope/risk/usage changes, and repeated failure: Sonnet/medium for routine implementation; Haiku/low for bounded extraction; Sonnet/high or Opus/high for diagnosis/plans…
The policy says which kind of work gets which tier. It does not say which model IDs exist today. The CLI knows that. The line-up has already changed more than once since I started. Because the file names no model IDs, a new release does not require editing it.
Claims about the world that expire. A note that a symlinked skill checkout was “230 commits behind” on some branch. Local inspection showed a different branch with a very different divergence, so the old number was obsolete. I replaced the number with a verification procedure: resolve the symlink, look at its branch and diff, compare with the canonical main before declaring code missing. “Docker is always running” became “check availability; ask me to start Docker Desktop if needed”.
Optional personal and business detail. Family context, legal identifiers, and business profile moved to a reference file that loads on demand. The global file keeps the privacy rules and a pointer. This is the same principle: the rule is cross-project, the detail is not.
Contradictory guidance. An old memory note said to wait for a context failure before starting a fresh chat. My instructions said to compact at safe checkpoints and keep the conversation continuous. I kept the second, which is what I actually want.
Stale status. A July incident still read as blocked in one note, even though a later note recorded the August resolution. The pointer now goes to the resolution.
Credential-shaped literals. While scanning, I found recognizable token and key strings in live memory and in archives. I removed them and kept the instruction for where such values live. This was a pattern scan, not an audit. The tokens were not rotated by this exercise, and the private backup I made before editing still contains the originals, so it stays local.
What I kept
The rule I used was: if removing a line would change what Claude does tomorrow, it stays. If it only records what was true yesterday, it moves or goes.
- Safety and secrets rules, the deployment verification habits, and the quality gates.
- Model and effort checkpoints, and the data boundaries for auxiliary providers.
- Pointers to where details live.
- One more convention that I like: when CI starts enforcing a rule, the prose shrinks to one line that names the enforcer. A rule that a pipeline checks does not need three paragraphs explaining why.
The resulting file has five sections: working relationship, source of truth and verification, safety and secrets, build quirks, and model and context budget.
How I checked I had not broken anything
This is the part I would do again. I did not hand the files to a model and ask it to trim them. The edits came with a change plan, a validator, and a rollback script.
- Every planned write was checked against the original file’s hash, then written atomically. If something else had edited the file in the meantime, the write would stop.
- All 470 original references from the four large indexes still resolve by walking links from the new indexes: 110, 129, 107 and 124 per index.
- Authored Markdown links all resolve, and every edited index sits inside the documented startup limit.
- A list of constraints I care about, covering secrets, deployment, effort and model checkpoints, containers, and the SDLC gates, was checked for presence after the edit.
- Settings parsed as JSON and their hash matched the baseline, because I wanted no settings changes mixed into this.
- The rollback script ran in dry-run mode and verified hashes for every applied file and backup. It refuses to overwrite later edits.
Total changes: 67 Markdown writes and the removal of one already-broken symlink. The earlier lesson about a model deleting load-bearing constraints is why the “constraints preserved” check exists. I wanted a mechanical test for it, not my own reading.
The caveat that matters
I want to state this plainly, because the numbers above look more authoritative than they are.
This was a document check, not a behavioral benchmark.
- I did not measure token counts of a live session before and after. Bytes are a proxy.
- I did not run the same tasks under the old and new files and compare results. I have no evidence that the agent got better, and I have no evidence that it got worse.
- Sessions that were already open kept the text they had loaded.
- The savings apply when the new entry points load. Whether they change outcomes is something I only learn by watching ordinary sessions for regressions.
What I do believe, without a benchmark: a smaller file with fewer stale claims has fewer ways to be wrong, and the model has fewer things to reconcile when two statements disagree. That is a design argument, not a measurement. If you want to test it, the cheapest experiment is a handful of representative tasks, run before and after, with a note of anything that changed.
What I would do differently
I would run the audit on a schedule. A file that doubles in three months will do it again, and the right time to catch it is the quarter it starts. I would also put a size ceiling in a check somewhere, the way I already check other things, so the argument for adding a paragraph has to include what comes out.
For now the practical habit is simple. Before adding anything to a startup file, ask whether it is a rule or a fact about today. Rules stay. Facts about today go in a note that gets read on demand and dated, so the next reader knows how much to trust it.