A $200/month Claude Max power user runs the real numbers on his own token usage, then rebuilds his workflow around open-weight models and tiered model routing instead of switching to a different single vendor.
Posted
yesterday
Duration
Format
Essay
educational
Views
846
67 likes
57 · 43
Big Idea
The argument in one line.
The $200 Claude Max plan hides thousands of dollars of real API cost per heavy-use day, and that subsidy is what creates vendor lock-in, not the tool itself.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You're a heavy Claude Code or Claude Max user who has never checked what your daily token usage would cost at raw API rates.
You've built real workflows, a business, or side projects entirely inside one AI subscription and haven't stress-tested what happens if the plan or pricing changes.
You want a concrete model-routing setup (which model plans, which codes, which handles small tasks) instead of vague 'try open source' advice.
You care about which company is quietly building a profile of your creative process through your coding assistant.
SKIP IF…
You're a light or occasional Claude user for whom $20-200/month already comfortably covers your usage; the economics argument here doesn't apply to you.
You want a step-by-step OpenCode installation tutorial; this video explains the reasoning and shows the routing rules, not a full setup walkthrough.
TL;DR
The full version, fast.
A $200/month Claude Max plan let this creator burn over 700 million tokens on peak days, work that would cost roughly $1,200-$1,900 even after cache discounts, and tens of thousands at raw API rates. That subsidy works like a hook: the cheaper Claude Code feels, the more a builder's tools and business get built around one vendor with no guaranteed price floor. His fix wasn't switching to a different single AI company. He moved to OpenCode, a model-agnostic harness, and wrote routing rules into AGENTS.md so an expensive planner model only plans, a mid-tier model codes, and a near-free model handles simple tasks, cutting one measured task from 726,470 Claude Code tokens to 418,190.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
He describes being a $200/month Claude Max power user pushing hundreds of millions of tokens a day, and states upfront that this dependency is exactly why he cancelled.
01:56 – 03:03
02 · The Token Math ($50,000 For Two Days)
Runs the numbers on two peak days of usage: tens of thousands of dollars at raw API pricing, still $1,200-$1,900 after cache discounts, against a $200 subscription.
03:03 – 04:36
03 · Subsidized Convenience Is The Lock-In
Argues the subsidized inference is what created his dependency: he stopped thinking about cost or model choice and just kept working inside Claude Code.
04:36 – 05:36
04 · The Audit: How Much Of My Life Ran On Claude
Lists what he actually built with Claude Code, landing pages, ad campaigns, video assets, and a side project book/game/website built with his son.
05:36 – 07:01
05 · Why 'Replace Claude' Is The Wrong Question
Realizes looking for one single model to replace Claude was the wrong frame, since different tasks (planning, coding, simple work) don't need the same model.
07:01 – 08:09
06 · The New Stack: Open-Weight Models + OpenCode
Introduces the open-weight models he now uses (Kimi K3, GLM 5.2, DeepSeek V4 Flash) and OpenCode, the model-agnostic harness that replaced Claude Code.
08:09 – 09:38
07 · Privacy And Profiling (And My Home Network)
Raises privacy as a second reason to move off proprietary tools: every prompt profiles how you think, and he doesn't want Anthropic on his home network.
09:38 – 11:53
08 · Smart Model Routing In The Workspace
Shows the actual AGENTS.md routing rules, planner/implementer/simple/review tiers, and a measured comparison: 726,470 Claude Code tokens vs 418,190 in a tight OpenCode workspace.
11:53 – 12:17
09 · The Generation Getting Handicapped
Warns that builders who never learn model routing will be caught out when subsidized plans change, and pitches his channel for workspace-building content.
12:17 – 13:29
10 · The Experiment: Venice Max And 300+ Models
Details his replacement plan: Venice Max at $200/month with $225 in credits across 300+ models, positioned as the actual test of whether this can replace Claude.
13:29 – 14:05
11 · Own The Workspaces, Not The GPUs
Argues the durable asset is the workspace and instructions, not the specific model or hardware, since a portable workspace survives a vendor swap.
14:05 – 14:40
12 · The Honest Exit: Will I Resubscribe?
Frames this as an open experiment rather than a permanent stance, and commits to posting honest updates including the possibility he resubscribes to Claude.
14:40 – 15:24
13 · AI Captains Academy (Come Say Hi)
Closing pitch for his paid community, the AI Captains Academy, and free email courses on AI foundations and command-line basics.
Atomic Insights
Lines worth screenshotting.
Two days of one Claude Max power user's actual usage would cost $1,200-$1,900 even with Anthropic's cache discounts, against a $200/month subscription price.
A $200 AI subscription that lets you use tens of thousands of dollars of raw API tokens isn't a bargain, it's a hook that gets you dependent before the pricing changes.
Anthropic already switched enterprise customers to usage-based billing, so a consumer flat-rate plan holding its price forever isn't guaranteed.
The real lock-in isn't the AI model, it's the workflows and unquestioned defaults a builder stacks on top of one subscription over months.
Asking 'what replaces Claude Code' is the wrong question, because no single model needs to plan, code, and handle small tasks equally well.
OpenCode is a model-agnostic harness: the same agentic workflow that runs on Claude can run on Kimi K3, GLM 5.2, or DeepSeek V4 Flash instead.
Every prompt sent to a proprietary coding assistant also profiles the way you think, not just the code it outputs.
Writing model-routing rules directly into AGENTS.md let one measured task drop from 726,470 Claude Code tokens to 418,190 tokens in OpenCode.
A tight workspace with skills, instructions, and sub-agent profiles burns fewer tokens than a fresh, unconfigured session solving the same problem from scratch.
Splitting AI work by tier, an expensive model to plan, a mid-tier model to implement, a near-free model for simple tasks, is what actually controls cost, not switching to a cheaper single vendor.
Zero data retention agreements matter because they mean no profile is built from your inference at all, not just that the raw output isn't stored.
Owning your workspace and instructions matters more than owning the GPU, because a portable AGENTS.md setup survives a vendor swap and a fixed local model doesn't.
Takeaway
How To Stop Renting One AI Vendor
VENDOR INDEPENDENCE
Before you assume your flat-rate AI plan is a deal, price out your actual usage, then split your workflow across tiered models instead of renting everything from one company.
01The $200 Power User Confession
Check what your actual AI usage would cost at raw per-token API rates before assuming your flat-rate plan is simply a good deal.
A subscription price that stays flat while your usage scales unpredictably is a signal to ask who is actually absorbing the cost difference.
02The Token Math ($50,000 For Two Days)
Cache discounts genuinely lower real API cost, but the discounted total can still dwarf a flat subscription fee for a heavy user.
Public API pricing pages are the fastest way to sanity-check whether a subscription's economics make sense for how much you actually use it.
03Subsidized Convenience Is The Lock-In
The moment you stop thinking about the cost of a request is the moment a subsidized tool has started shaping your workflow instead of the other way around.
A plan that looks generous today can tighten the moment the company behind it needs to hit different financial targets, including an IPO.
04The Audit: How Much Of My Life Ran On Claude
List every real project, business task, and side project currently running through a single AI tool before deciding how replaceable that tool actually is.
Side projects with no guaranteed return are often the ones most dependent on subsidized inference, because they wouldn't get built at full price.
05Why 'Replace Claude' Is The Wrong Question
Looking for one single model to replace another single model skips the real fix: no one model needs to be equally good at planning, coding, and busywork.
Vendor lock-in isn't really about the AI company, it's about how much of your own workflow you built with no exit plan.
06The New Stack: Open-Weight Models + OpenCode
Open-weight models now benchmark close enough to frontier models to run real production workflows, not just toy demos.
A model-agnostic harness lets you keep the same agentic workflow you already built while swapping which company's model powers it.
07Privacy And Profiling (And My Home Network)
Every prompt sent to a proprietary coding assistant is also training data about how you think, not just a request for code.
Zero data retention agreements are worth asking for by name, since they mean no profile is built from your inference at all.
08Smart Model Routing In The Workspace
Write your model-routing rules directly into your agent's instructions file so the routing happens automatically instead of by memory each session.
A well-instructed workspace with skills and sub-agent profiles uses meaningfully fewer tokens than a fresh, uninstructed session solving the same problem from scratch.
09The Generation Getting Handicapped
Builders who never learn to route between models risk being stuck when a subsidized plan's economics change and the convenience disappears.
10The Experiment: Venice Max And 300+ Models
A single alternative provider plan can bundle access to hundreds of open-weight models, which makes tiered routing practical without juggling separate API keys for each one.
11Own The Workspaces, Not The GPUs
The durable asset isn't the specific model you use today, it's the workspace and instructions you've built, because those transfer to whatever model comes next.
You don't need to run your own hardware to be independent of one vendor, you need a workspace that isn't hard-wired to only one provider.
12The Honest Exit: Will I Resubscribe?
Treat a vendor-independence experiment as reversible and time-boxed rather than a permanent stance, so you can honestly report back if it doesn't work.
Glossary
Terms worth knowing.
Token
The unit AI providers bill by, roughly a few characters of text. Both what you type and what the model generates count toward the price.
Cache reads/writes
A discounted rate a provider charges for reusing recently-sent context instead of resending it fresh, which is why real usage often costs less than raw per-token math suggests.
Open-weight model
A model whose trained parameters are published for anyone to run or rent, as opposed to a closed model only accessible through one company's paid API.
Model-agnostic harness
A coding agent tool, like OpenCode, that can connect to any AI provider's model instead of being locked to one company's API.
AGENTS.md
A plain-text instructions file an AI coding agent reads at the start of a session, used here to hard-code which model handles which type of task.
Model routing
Sending different parts of a job to different AI models based on how demanding the task is, instead of using one model for everything.
Zero data retention
A provider agreement stating that none of a user's prompts or outputs are stored or used to train future models.
“418,190 tokens versus 726,470 tokens for Claude Code.”
a hard number comparison that reads as a graphic on its own→ TikTok hook↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphorstory
I was a CloudMax subscriber. I paid $200 a month. I loved it.
I was a beast power user using hundreds of millions of tokens per day sometimes. I can easily say that just over the last 12 months, CloudCode went from pretty neat to OMG can't live without it. Something as a builder and solopreneur I genuinely couldn't do without.
And that is exactly why I canceled my subscription. Because I realized there is something seriously wrong with the dependency on Anthropic. Let me explain.
You see, over the last 30 days of my subscription, on two of them I surpassed 700 million tokens per day, and other days were well on their way. If we do all the math here with the general API cost, those two days alone would have cost over $50 ,000, around $2 ,000 if you factor in all the token caching, but we'll get to that later.
But even with the discounted cash tokens, we're still looking at approximately $1 ,800 just for two days of usage. And my subscription only costs $200. Something is seriously wrong with the economics here, and I don't want to be hooked on Claude when the other shoe drops.
How was I getting all of this for $200 a month? Someone has to foot the energy bill here. AI is not free.
Yet the more I use Claude Code, the more my work and business depended on it. Until I finally started thinking about what would happen if that $200 plan disappeared next week. What if they changed the limits or the product itself?
At that point, how easy would it be for me to leave? And over time, that question started bothering me. A lot.
Because a large part of my life was revolving around Claude. Fortunately, there's a way out of this. Today, we have open weight models that are not only surprisingly good, but they offer results comparable to Opus and Fable.
Meanwhile, there are more and more open source agent frameworks and harnesses that are model agnostic, which allow you to get creative past the handful of models offered by Claude, or OpenAI, or any American frontier family of models. So I decided to try something. I want to see whether I can replace my $200 Claude subscription with open models and a smarter workflow without giving up the things that made Claude so damn useful to me in the first place.
Let's start with the token inference. As I said, those two days alone would have cost approximately $1 ,800. But that's only two days of my 30 -day subscription that I pay, in total, $200.
So I started digging into that because it seemed strange. The API prices are public. If you took the amount of input and output tokens from those days and priced them as normal API usage, well, you see these high numbers in the tens of thousands.
So either their API margins are obscene, or they're burning VC money to subsidize user inference, get us hooked, and hiding the actual cost of delivery here. It's probably a little bit of both. They have a good product.
People want to pay for it. Cloud Code has become one of the most useful products some of us have ever paid for in our lives. We use it constantly.
Someone like me, I could sit down on a weekend, start building something and not stop for 20 hours straight. The more you use it, the more you got from that easy $200. That's an incredible deal for a heavy user.
But there is a flip side. Once you get used to working that way, you build your entire workflow around it and you become dependent. You stop thinking about the cost of every request.
You stop worrying about which model you're using. You just open Cloud Code and keep working. And I had definitely gotten to that point.
Lazy with my AI usage because the Cloud Code harness itself with the subsidized inference allowed me to just work and not have to really think about anything under the hood. So when I started looking at how much of my actual work was being done with Cloud, I started wondering what happens if the deal changes? Because we've seen AI products or anything in tech changed their limits before.
A plan that looks generous today can look very different once the economic reality of running a business kicks in, which is often what happens when a company is about to IPO. But that's another story. So I had already seen how quickly my usage could explode with over 700 million tokens in a single day.
That's just what my workflow started to look like. So I had to ask myself the uncomfortable question. If Anthropic changed the price tomorrow, or tightened the limits, or changed the way Cloud Code worked, how much of my workflow would I have to rebuild or start from scratch?
And that question bothered me a lot more than the $200. So I used Cloud Code to do a lot. built landing pages, sales pages, automated meta ad campaigns, generate visual assets for videos or presentations with Remotion and Hyperframes, and handle parts of my content workflow, including repurposing, metadata generation, et cetera.
I did a lot. And then there were my side projects. I had a project I was working on with my son where he created a world, built a story, and we turned it into a book and a YouTube channel and an educational learn -to -read game on its own website.
We made our own vibe -coded Patreon. And Claude was sitting right in the middle of that. If I didn't have the subsidized free inference, essentially, to build that, I wouldn't.
Because I wasn't guaranteed to make any money. It was just a fun project. And it was with this project with my son I realized I would open Claude code before I'd even think about whether there was another way to get the job done.
I already knew how to use it. The tools were there. The context and memory was there.
I built up all my workflows around it. So starting something new was easy. And that convenience adds up as an ignorance debt that eventually we must pay back.
Because you've stopped questioning the tool because it just works. And after enough time, you inevitably feel locked in whether you realize it or not. The cost of switching just becomes too high.
And that's what Anthropic wants. That's what any business wants. It's good business.
Because it turned into even the thought of me canceling my Cloud subscription was more than just cutting another subscription from my life. I'm not just going to stop building with AI. I have to figure out an alternative.
There's a cost to switching. Some of my workflows would be easy to move, but some of it would not. And truthfully, I didn't like how little I'd actually thought about that before.
I'd spent all this time making Claude more useful to me, while quietly making my own workflow harder to separate from it. And that's a weird position to end up in. You start with a tool because it's convenient, then you build around it, and eventually, that convenience becomes the reason you don't want to leave.
It's like the relationship you know you need to end, but it's just too difficult. So eventually two things were really bothering me. The future economics of the subscription and the company, the vendor lock -in.
The fact that I had built so much around a system that I didn't really control. So at this point, I asked myself, I don't want to just be stuck with Claude. And I don't want to just resign to that.
Especially with all the geopolitical drama Anthropic has been cooking up. So, I started looking for a way out. Now the question arises, what the hell do we use instead of Claude?
or open ai because i'm talking about claude here but the truth is this applies to any ai provider especially one that doesn't respect your privacy and is giving you something for free remember what they say there is no such thing as a free lunch and if lunches were tokens well i've gotten a lot of free lunches with Claude Code.
But the answer to this question is not to replace Claude or Codex with one thing. And that's how I was thinking about it the wrong way at the beginning. I was looking for a single tool that could replace Claude.
Oh, is Codex better? Is Gemini CLI better? But that doesn't really make sense because why should one model or one family model have to provide everything?
If I'm writing code, maybe I want just one model. If I'm planning a project, I might want another. And this is where things got interesting.
because the options had changed a lot since I first started using Cloud Code. Open weight models now can do some serious work, from Kimi K3 to GLM 5 .2, GLM 5 .3, DeepSeq V4, Quen 3 .8, and you can connect all these to open source agent harnesses. And a tool like OpenCode, which has quickly become my Cloud Code replacement, is what they call model agnostic.
I can connect any provider, any model into there. And it's completely compatible with my Cloud Code agentic workflows. So once I started thinking about it that way, the whole setup changed.
I was able to start using Kimi K3 as more of a planner, an orchestrator. It's more expensive, but it's super intelligent, and it's actually a lot friendlier and easier to work with than Fable. Meanwhile, GLM 5 .2 is cheaper, could handle a lot of the coding, and it's a very good creative writer as well.
And for even simpler jobs, DeepSeek V4 Flash may as well be free. It's fast, intelligent, and while I wouldn't want it planning my projects, it can handle the manual tasks perfectly. And getting comfortable with this meant that I could move away from the belief that I needed cloud code to achieve the productivity that I desired.
But there's another advantage I want to double click on, and that is actually choosing where the models come from. Now, as I was talking about, when I'm creating something for my business content, like what does the privacy really matter? Well, in a lot of ways, the output itself, whatever.
Privacy, no privacy, whatever. But there's the actual creative process that comes from me, my unique way of thinking as a human being that is being profiled by tools. like Claude Code, Codex, Gemini.
They're learning my psychology and making a profile about me. And once you start going deep into that, it's kind of freaky. So there's another reason to shift from these proprietary frontier services.
And that begins to align with the inevitable next step if you're a builder with AI is what you can start using these harnesses for inside your own house. I started messing around with home security systems with Raspberry Pis in the living room hosting video game emulators. And once I hit the stage, I realized I don't want Anthropic in my home network.
No, thank you. And the models I can host locally just aren't there yet. So that's why I use Venice as a provider of all of my favorite open weight models.
because they have zero data retention agreements with the providers, meaning nothing is being stored by the provider of my inference. No profile is being built. And even if there's a nefarious entity inside the GPUs that my provider doesn't even know about, they're not going to know what's coming from me.
So this started creating a comprehensive view of how I want to approach my AI setups now. Instead of opening Cloud Code and asking it to do every little thing, I can now build workflows that use different models, different tools, different providers for different parts of the job.
And that brings us into a topic that most cloud code users have probably not thought about much, and that is smart model routing. I had no choice but to start digging into this. I'd gotten into the habit of treating cloud like the default answer to every problem.
So once I stopped doing that, the economics actually started to look very different. I could have one model think through what actually needs to happen and then hand off the plan to another model for implementation, a cheaper model. And it's not a coding thing, I could use even simpler models.
And then when I need a serious final review or audit, then I bring the big boys back in. So the most interesting part of something like smart model routing is that it can be handled in the workspace itself. So what might not be obvious as a cloud code user is that your AI agent can make calls to other models with instructions, with a system prompt, whatever.
And what's really interesting about this, if you've only used Cloud Code or Codex and you've never built your own workspaces, is that the workspace itself can handle a lot of this smart model routing. The better your instructions are in AgentsMD, the smarter your agent can be in sending certain tasks to certain models. This will not only give you better performance and results, but it will save money, which is important when we're no longer using the subsidized inference by...
Anthropic. A well -structured workspace means my agent is going to be smart about the tasks it does versus the tasks it sends sub -agents to do. So as an example, I could use cloud code for a task, I could use open code for a task, and they might use a completely different amount of tokens, even if I'm using the same model.
So let me explain. If I throw everything at Cloud Code and say, hey, figure this thing out, just do it. It'll probably use a lot of tokens in a fresh workspace with no instructions.
It'll make its own instructions. It's so smart. It'll figure it out.
It'll probably do a good job. But it's going to burn through way more tokens than if you use the same model, Opus, Fable, in a tight workspace with instructions, with skills, with sub -agent profiles, et cetera, because it has context to work with. Not only that, but within the model workspace, you can integrate model routing.
So if something needs to be done by a different agent, in your instructions you simply say when we do this thing send it to that model from that provider here's the api key done you could do this with cloud code but why would you even think to do that when cloud code is just so good and you're not paying for the actual inference And a whole generation of AI users are going to be handicapped when reality comes knocking.
And I don't wish that for you. So I'm glad you're watching this. And now is probably a good time to subscribe to my channel if you do want to learn how to make workspaces that are not dependent on cod code, because that's what I explore in my videos on the channel.
So thanks for hitting subscribe. So once I started understanding the different economics of running different models in a workspace, I wanted to know how far I could take it. Could I actually rebuild enough of my workflows that I wouldn't miss Claude?
The answer was way more obvious than I could have expected. So I decided to actually test it out. I've canceled my Claude subscription.
It's time to move on. I'm running a Venice Max account, which is also $200 a month, and it gives me $225 worth of credits with their service. But there's over 300 different models that I can choose from, from the top of the line open weight models like Himi K3 to GLM 5 .2 to DeepSeek V4 Flash, etc.
There's a lot to play with here that will allow me to save money and retain performance. It just means sacrificing the convenience. My goal is not to recreate Claude perfectly, but it is to get similar results without having to spend a lot more money.
The whole point of my experiment is to find out where the gaps in my own workspaces and workflows are. I have to be present with every single word and every single prompt because all of those are costing me tokens, which cost money. If a task doesn't need the strongest model, I shouldn't send the task there.
If a bunch of context is going to be used, can it be cached? If I'm recycling instructions, I need to make skills and put them in my behavioral context files. This is the stuff I'm most interested in because this is the true architecture of owning your own AI.
Sure, I might not be running the actual intelligence on a machine I own, but theoretically in the not so distant future, I could. And more importantly, as long as there are vendors out there willing to provide for me the models at a cost that I want to use, then the most important aspect for me is owning the workspaces themselves because they can be mine and no one can take those away from me.
No one can shut those off. I'll just find another model provider. We live in a world of markets.
Find the service provider that works for you. This way, if a particular tool or model or provider stops working for me, the rest of my workflow, the rest of my life with AI doesn't have to come down with it. I don't know yet how far this will go, but that's what the next few weeks of my life are gonna all be about, not using cloud code and moving into agentic workspaces that I control with inference that I'm paying for every single prompt accounted for.
A responsibility that I have chosen to stop outsourcing to the corporate provider of my artificial intelligence. Can open weight models, open source agents, and smart model routing replace what I was getting with cloud code? Or am I going to end up resubscribing to cloud, tail between my legs in a few weeks from now?
now. That's the experiment. Stay tuned because I will be posting updates, letting you know how it goes.
Finally, if you're interested in learning how to build these agentic workspaces yourself, I invite you to check out the AI Captains Academy. There is a whole curriculum inside for understanding from the ground up how AI agents work, how to select between models and providers and agent frameworks, etc. How to approach building your agents, including knowledge infrastructure and context engineering.
We do two calls a week. One is a workshop, one is a questions and answer session.
Come say hi, be in a community of builders that use AI, that love AI like you and me, and we share our solutions to these problems and challenges we face, such as replacing cloud code. So if that sounds interesting to you, check out the link below and I hope to see you there. Thanks for watching and I hope this has inspired you to reconsider your own cloud or open AI subscription.
The Hook
The bait, then the rug-pull.
He was a $200-a-month power user pushing hundreds of millions of tokens a day, and the video opens with the math nobody else bothers to run: what does my actual usage cost at real API rates?
Frameworks
Named ideas worth stealing.
10:29list
Tiered model routing (planner / implementer / simple / review)
A four-tier routing table written directly into AGENTS.md so the agent automatically sends each task to the cheapest model capable of doing it well, instead of using one model for everything.
Steal forany AI coding or content workflow currently doing every task at the same subscription price
CTA Breakdown
How they asked for the click.
VERBAL ASK
15:00product
“check out the AI Captains Academy”
soft plug woven into the closing minutes after the experiment framing lands, points to a paid community plus two free 5-day email courses rather than a hard sell
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
A three-level walkthrough of prompting, saving a design system, and layering in AI-rendered images to turn one topic into a branded Instagram carousel.
A fifteen-year video editor spent three months vibe-coding the ingest, sync, and AI-assistant tool he'd wanted for eight years, then shipped it as a one-time purchase instead of a subscription.
A creator feeds a public archive of 2,000+ Alex Hormozi coaching transcripts into a Claude Code skill that interviews any business and spits out an AI-written constraint diagnosis — then uses the free download as the top of a paid-consulting funnel.
A Claude Code creator mines his own chat history into a personal test pack, then builds a slash-command benchmark that tells him in one run whether a new model release is actually worth switching to.