7 Claude Code Tools To Stop Wasting Tokens
A 28-minute practical breakdown of seven tools that attack token waste at session startup, during input, and in model output.
May 27thA creator distills Anthropic's own guidance on running efficient Claude Code sessions into six habits across three buckets.
Six habits split across context management, resource selection, and noise filtering determine whether a Claude Code session burns tokens efficiently or quietly racks up avoidable cost.
Anthropic published guidance on running efficient Claude Code sessions, and this video distills it into six habits across three buckets. Context management means clearing between unrelated tasks, only running /compact within the first hour before the prompt cache expires, and auditing total context with /context, which can reveal tens of thousands of tokens loaded before a single message from forgotten MCP servers or skills. Resource efficiency means locking the model and effort level at session start, since switching either mid-session invalidates the whole cache, and @-mentioning files directly instead of letting Claude search for them. Noise reduction means filtering noisy commands and routing anything likely to produce heavy output through a cheap sub-agent.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →
Sean frames the video: Anthropic published an article on maximizing Claude Code session value, and he's pulling out the six biggest takeaways, grouped into three buckets.

The first and most underused habit: run /clear whenever you're switching to unrelated work so stale context doesn't ride along. Spec-driven workflows help because every stage gets saved to an artifact, so clearing is always safe.

Compact only pays off inside the first hour of a session. Wait longer than that and the prompt cache has already expired, so /compact forces a full reread instead of a cheap summarize.

Running /context shows exactly what's loaded before a single message is sent, 47,000 tokens on an Opus 5 window in this example, with 11,000 of that tied to memory. The /memory command breaks down whether that's user-level or project-level bloat.

Installed MCP servers, custom agents, and viral skills (he name-checks Graphify) often get forgotten and keep adding tokens to every session's startup cost, sometimes pushing pre-message load past 100,000 tokens.

The three context-management habits, clear, time your compacts, and audit with /context, combine into one weekly ritual: check your context total and prune anything unrelated.

The model and effort-level pickers look like harmless settings, but switching either mid-session invalidates the entire prompt cache, every prior message has to be reread and repaid for on the next turn. Swap models by spawning a sub-agent instead.

Directly @-mentioning a file attaches it to the request outright, cutting the read calls and search operations Claude would otherwise burn tokens on hunting for it across directories.

Letting Claude run an unfiltered git status (or similar) on a messy repo dumps everything into the context window and the cache, muddying both and costing tokens for output that was never useful.

His actual habit: kick off a Sonnet or Haiku sub-agent to run the noisy command, isolate just the useful result (clean branch, last five commits, nothing staged), and hand only that back to the main orchestrator window.

Subscription plans hide per-token cost, but that discipline becomes mandatory the moment you're paying metered API rates or building on open-source models, habits worth building now before the flat-fee subsidy goes away. He closes by linking Anthropic's original article.
The gap between a Claude Code session that burns tokens fast and one that doesn't comes down to three buckets: managing context, picking resources deliberately, and filtering noise before it hits the model.
“You should be clearing out your context entirely.”
“You've basically lost the entire cache.”
“47,000 tokens are being loaded up before I have even sent anything.”
“When you change this, if you change it mid session, it is going to completely invalidate the cash that you have.”
“We don't really care how many tokens it takes to get something done, and that doesn't actually work when you are using open source models.”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
Anthropic published its own guidance on running efficient Claude Code sessions. Sean Kochel pulls out the six biggest takeaways and sorts them into three buckets: what to do with your context, which resources to lock in early, and how to keep noisy command output from bloating every turn.
Anthropic's official guidance boils down to six habits split across three buckets that together determine how many tokens, and how much cache, a Claude Code session burns.
“I'll link to this article below because there's actually a lot of really good information on here.”
Soft, single-mention CTA pointing to the original Anthropic blog post, delivered in the closing seconds rather than pitched mid-video.
00:00
00:13
00:23
00:32
00:41
00:50
01:00
01:09
01:18
01:27
01:37
01:46
01:55
02:04
02:14
02:23
02:32
02:41
02:49
03:00
03:09
03:18
03:28
03:37
03:46
03:55
04:05
04:12
04:23
04:32
04:42
04:51
05:00
05:09
05:19
05:28
05:37
05:46
05:56
06:05
06:14
06:23
06:33
06:42
06:51
07:00
07:10
07:19
07:24
07:42
07:47
07:56
08:05
08:14
08:24
08:33
08:42
08:51
09:01
09:10
09:17
09:28
09:38
09:47
09:52
10:05
10:15
10:24
10:35
10:42
10:52
11:01
11:10
11:19
11:29
11:38
11:47
11:56
12:06
12:15Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
A 28-minute practical breakdown of seven tools that attack token waste at session startup, during input, and in model output.
May 27thSix composable patterns that turn Claude Code into a real multi-agent orchestrator — with two live workflow demos and a token-budget survival guide.
June 4thSean Kochel wires Mobbin's screen library into Claude Design through a community MCP server, then uses it to generate three UX directions, an onboarding flow copied structurally from a named competitor, and a single mocked-up UI component — all for a lactation-support app he invents on the spot.
August 5thA solo builder scopes, researches, designs, and epics-out a real fitness-tracking app, then hands the implementation to an unattended overnight Claude Code loop with a verification sub-agent watching every phase.
July 24thSean Kochel installs an open-source Claude Code skill, answers a five-question brand interview, and watches it build — and bill him $20 for — a scrollable 3D website.
July 15thA one-hour engineering checklist for builders who can ship prototypes but keep breaking production.
June 26th