Modern Creator
Theo - t3․gg · YouTube

If you have a Claude sub, watch this

A detailed, ethically loose walkthrough of squeezing thousands of dollars in token inference out of a $200 Claude or Codex subscription, then running a dozen AI coding agents at once without losing track of them.

Posted
6 days ago
Duration
Format
Essay
comedic-rant
Views
260.8K
4.7K likes
Big Idea

The argument in one line.

A $200 Claude or Codex subscription converts into thousands of dollars of inference that resets whether you use it or not, so the real skill in 2026 isn't writing code faster, it's running enough parallel agent threads, backed by the right accounts, proxy, and always-on hardware, to actually spend what you're already paying for.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You already pay for a $100-$200/month Claude or Codex subscription and want to know whether you're using anywhere near its token allowance.
  • You run, or want to run, several AI coding agents in parallel and need a system for tracking which threads actually need your attention.
  • You're comfortable setting up a proxy, Tailscale, or a spare Linux box and want a cheaper way to keep agents running around the clock.
  • You manage engineers and are curious how a heavy agent-user structures a workday around dozens of concurrent threads.
SKIP IF…
  • You're looking for API-based production advice. This is specifically about personal subscription token limits, not serving inference to users.
  • You want a first introduction to AI coding agents. The video assumes you already run Claude Code or Codex regularly.
  • You're not willing to risk a subscription ban. The presenter is explicit that several tactics described may violate terms of service.
TL;DR

The full version, fast.

A $200/month Claude subscription converts to roughly $8,000 of token inference and a $200 Codex subscription to about $12,000, both heavily subsidized compared to API pricing, and both reset on a timer whether or not you used them. The system: route several accounts through one residential IP via a local proxy, prioritize burning whichever account is closest to reset, and run agents from an always-on home Linux box reachable over Tailscale rather than a flagged VPS. On the usage side, the real shift is treating agent threads like an email inbox, bringing the agent in at the problem stage rather than the already-solved stage, and accepting that humans stay single-threaded while agents don't, so the job becomes triaging finished work instead of babysitting one thread.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:00 – 02:00

01 · The premise: are you even running agents right now?

Theo opens with a question, how many agent threads are you running right now, then frames the rest of the video as material Anthropic won't love: a disclosure that across roughly 30 shared Claude accounts, only two temporary bans have happened, both tied to a VPN sign-in and a mismatched billing name.

02:00 – 03:15

02 · Sponsor break: Parallel

A sponsored segment for Parallel, an AI-focused web search and extraction API, demonstrated head-to-head against OpenAI's search for speed.

03:15 – 05:28

03 · The roadmap, and getting ahead of the haters

Theo lays out the five-part structure (getting, accessing, using well, using poorly, using while asleep) and pre-empts the expected Twitter backlash, insisting the tactics scale from side projects to large companies and that the goal is inspiration, not a copy-paste formula.

05:28 – 11:04

04 · The subscription math: how $200 becomes $8,000-$12,000 in tokens

The core economics: a $200 Claude subscription converts to about $8,000 of inference (only $4,000 to the flagship model), a $200 Codex subscription to about $12,000, the flagship model's pricing history (a leaked $25/$125 per-million enterprise rate, now roughly 70% cheaper), and a ballpark 95% inference margin that explains the subsidy.

11:04 – 15:22

05 · The ethical line: data sharing, and never resell your sub to users

Two ground rules: disable 'improve the model for everyone' in account data controls so a personal sub's terms match a team account's, and never route public, user-facing traffic through a personal subscription instead of the API, which the presenter calls an actual bannable offense.

15:22 – 22:13

06 · Accessing your tokens: CLI Proxy, quota dashboards, and burn order

Walkthrough of CLI Proxy (run as a fork called Vibe Proxy) for holding multiple accounts' OAuth sessions behind one API endpoint, the quota-management dashboard, prioritizing whichever account is closest to reset, the five-hour-vs-seven-day limit math, an Opus 5.5 aside, and session/account affinity to protect prompt caching.

22:13 – 28:57

07 · The fleet: Tailscale, a residential IP, and the reset-mechanics gotcha

Why signing in from a VPS or cloud IP is risky, how one home proxy box reached over Tailscale from every other machine keeps traffic looking residential, the 'how to set up a box' fleet-management runbook, and the difference between Codex's full-clock resets and Claude's percentage-only resets.

28:57 – 38:32

08 · Using tokens well: de-risk the merge, give agents problems not solutions

The shift from coding to verifying (roughly 10-15% vs 80-90% of token spend), the 15-second revert threshold for when autonomous coding is actually safe, and the habit of bringing an agent in at the problem stage instead of handing it an already-decided solution.

38:32 – 49:45

09 · Inside T3 Code: threads as to-dos, not histories

A live tour of T3 Code's thread sidebar: treating threads as disposable to-dos, the 'settle' action for inbox-zero, picking which machine a thread runs on, sending a problem plus permission to build directly or mock it up first, and the background-send shortcut.

49:45 – 58:43

10 · Agents are multi-threaded, humans aren't: running 10+ threads as an inbox

The stated thesis, agents can be multi-threaded, humans are single-threaded, demonstrated with a live sidebar of 10+ concurrent threads, a dimmed-UI trick for in-progress work, and the case against watching an agent work in real time.

58:43 – 1:06:36

11 · Using tokens while sleeping: cheap always-on boxes over VPS rentals

The closing pitch for moving dev work onto an always-on, low-spec home box over Tailscale instead of a rented VPS, with cost comparisons against Hetzner-style cloud rentals, a live look at an underused 32-core server, and a final riff on why this has made engineering fun again.

Atomic Insights

Lines worth screenshotting.

  • A $200 Claude subscription converts to roughly $8,000 of total inference, but only $4,000 of that can go to the flagship model.
  • A $200 Codex subscription converts to about $12,000 of inference with no such split between flagship and cheaper models.
  • Anthropic's flagship-model pricing is reportedly about 70% cheaper than the rate it first quoted enterprise early-access customers.
  • Industry napkin math puts AI inference margins at roughly 95%, meaning a $100 charge covers about $5 of real compute and electricity cost.
  • Codex subscriptions reportedly gave out 14 resets in 30 days, which turns a weekly token limit into something closer to a three-day limit.
  • Claude's usage reset only restores your percentage, not your clock, so unspent tokens before a scheduled reset simply vanish rather than roll over.
  • Multiple AI coding accounts should route through a single residential IP, since requests from hosting providers like AWS or Hetzner draw extra scrutiny.
  • Switching accounts mid-thread forces the API to rebuild cached context, so keeping one account per thread avoids paying repeatedly for the same cache writes.
  • On a typical coding task, maybe 10-15% of token spend goes to writing code and 80-90% goes to verifying it's actually correct.
  • If undoing a bad agent-made change takes more than 15 seconds, the surrounding systems aren't ready for autonomous coding yet.
  • Handing an agent a problem instead of a pre-decided solution avoids wasting tokens defending a fix that may not have been the right one.
  • A secondhand PC with as little as 4-8GB of RAM and a few cores can run multiple coding agents at once, since the work is lightly threaded, not GPU-bound.
  • Renting a 32-core, 32GB cloud server for about $275/month costs more within roughly two and a half months than buying equivalent used hardware outright.
  • Agents can run many threads in parallel; a human working on them can only focus on one at a time, which makes attention, not code generation, the real bottleneck.
Takeaway

Spending your tokens is now the real skill.

TOKEN ECONOMICS

A paid Claude or Codex subscription is a use-it-or-lose-it token budget that resets on a clock, and the real craft in 2026 is running enough parallel agent threads to actually spend it before it disappears.

01The premise: are you even running agents right now?
  • The real benchmark for running AI coding agents in 2026 isn't how well you prompt, it's how many agent threads you already have running at any given moment.
  • Across roughly 30 shared Claude accounts, only two bans occurred, both temporary, tied to a VPN sign-in and a mismatched billing name rather than to the usage pattern itself.
03The roadmap, and getting ahead of the haters
  • The tactics described scale from solo side projects up to large-company teams, apart from outright subscription resale to paying customers, which the presenter calls out as the genuinely risky version of this.
  • The goal isn't to copy the exact setup shown, it's to absorb enough examples that you start generating your own workarounds for problems the presenter hasn't thought of yet.
04The subscription math: how $200 becomes $8,000-$12,000 in tokens
  • A $200 Claude subscription converts to about $8,000 of inference (only $4,000 of which can go to the flagship model); a $200 Codex subscription converts to about $12,000 with no such split.
  • Anthropic's flagship-model pricing is reportedly about 70% cheaper than the rate it first quoted enterprise early-access customers, a move aimed at undercutting competitors rather than reflecting falling compute cost.
  • Industry margin estimates put AI inference profit around 95%, which is the subsidy that makes burning a subscription to zero before its reset the financially rational move.
05The ethical line: data sharing, and never resell your sub to users
  • Turning off 'improve the model for everyone' in a personal account's data controls makes its terms of service effectively match a paid team account, so the privacy argument for paying extra to use the API instead doesn't hold up.
  • Routing public, user-facing traffic through a personal subscription instead of the API is explicitly against the terms and a real ban risk, unlike the account-stacking tactics covered elsewhere in the video.
06Accessing your tokens: CLI Proxy, quota dashboards, and burn order
  • A local proxy (CLI Proxy, run as a fork called Vibe Proxy in the presenter's setup) holds OAuth sessions for multiple Claude and Codex accounts and exposes them as one API endpoint that Claude Code or Codex can point to.
  • Routing traffic evenly across accounts wastes the ones closest to reset; prioritizing the soonest-to-expire account first avoids losing unspent tokens for nothing.
  • Keeping a prompt thread on the same account it started on avoids forcing the API to rebuild cached context, which otherwise gets billed at full price again.
07The fleet: Tailscale, a residential IP, and the reset-mechanics gotcha
  • Signing into Claude Code or Codex from a VPS or cloud IP draws extra scrutiny, since hosting-provider IPs are associated with reselling subscription access to other people's traffic.
  • Running one proxy box on a home network and reaching it from every other machine over Tailscale keeps all traffic looking like it comes from a single residential IP, regardless of which machine triggered it.
  • Codex resets restart the entire weekly clock, while Claude resets only restore the usage percentage and leave the clock untouched, so unspent Claude tokens before a reset are gone for good.
08Using tokens well: de-risk the merge, give agents problems not solutions
  • Roughly 10-15% of token spend goes into writing code, and the other 80-90% goes into verifying it's actually correct before merging.
  • If undoing a bad agent-made change takes longer than about 15 seconds, the surrounding revert and rollback systems need fixing before autonomous coding is safe to lean on harder.
  • Bringing an agent in at the problem stage, before you've decided on a fix, costs fewer tokens than handing it your own solution and then needing extra passes when that solution turns out wrong.
09Inside T3 Code: threads as to-dos, not histories
  • Treating each agent thread as a disposable to-do rather than a saved conversation history changes how much backlog accumulates in a sidebar.
  • A 'settle' action that clears a finished thread, plus auto-settling anything untouched for several days, turns a thread sidebar into a working inbox-zero system.
  • Sending a problem description plus explicit permission, go do it directly if you have a solution, otherwise mock it up, lets the agent choose its own confidence level instead of guessing what's expected.
10Agents are multi-threaded, humans aren't: running 10+ threads as an inbox
  • Agents can run many threads in parallel; a human working on them can only focus on one thread at a time, so the real bottleneck shifts from writing code to triaging finished work.
  • Visually dimming threads that are still working, while keeping 'done' or 'needs input' threads fully visible, reduces the urge to babysit work that isn't ready yet.
  • Watching an agent work in real time is one of the lowest-value uses of your attention: if it succeeds, merge it; if it fails, read the reasoning trace or ask why instead of hovering.
11Using tokens while sleeping: cheap always-on boxes over VPS rentals
  • A roughly $275/month rented 32-core cloud server costs more within about two and a half months than buying comparable used hardware outright.
  • A secondhand PC with as little as 4-8GB of RAM and a few cores is usually enough to run multiple coding agents, since the work is lightly threaded, not GPU-heavy inference.
  • Moving agent work onto an always-on home box reachable over Tailscale removes the 'I have to close my laptop' excuse for not firing off one more thread before stepping away.
Glossary

Terms worth knowing.

Token maxxing
Squeezing the maximum usable AI inference out of a flat-rate coding subscription before its usage limit resets, rather than letting unused capacity expire.
CLI Proxy
A local proxy server that holds the login sessions for multiple AI coding accounts and exposes them through one API endpoint a coding tool can point to.
Session affinity
Keeping every prompt in an ongoing agent thread routed through the same account, so the provider doesn't have to rebuild (and rebill) cached context when the account changes.
Banked reset
An optional usage-limit reset a provider can grant on demand, as opposed to the automatic reset that happens on a fixed schedule.
Residential IP
A home internet connection's IP address, used here to route multiple accounts' traffic so it doesn't look like it's coming from a server or hosting provider.
Tailscale / tailnet
A private mesh VPN that lets multiple computers reach each other, and a shared proxy, securely without exposing anything to the public internet.
Settle
Marking a finished agent thread as done so it disappears from the active sidebar, the core mechanic behind treating a thread list like an email inbox.
Fable
The flagship Claude model referenced throughout the video, priced and rate-limited separately from the rest of a Claude subscription's token pool.
Astra
The flagship Codex model referenced as Fable's rival, with no separate limit carved out of the broader Codex subscription pool.
T3 Code
The presenter's own agentic coding tool, used throughout the video's second half to demo thread management, cross-machine routing, and load balancing.
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

Before we get started with this one, I have a question for you. How many agent threads do you have running right now? I don't mean, have you run today?
I don't mean, are you gonna run later? I mean, in the exact moment where you clicked this video, did you already have things running in the background? If the answer is less than five, you really, really need to watch this video because you are not taking advantage of the genuinely awesome and unique opportunity we live in today.
And if the answer is more than five, you should also probably watch this because you'll get a lot of nice edges and solutions to problems you might not have even known about. And you'll have a great resource to share with your friends who are not taking advantage of where we're at today. I'll be real with you guys, though.
This video is not my norm. And as such, this doesn't feel quite right. I don't think we need caffeine for this one.
We're going full rodeo.
Much better. This is the video that Anthropic does not want you to watch and I'm going to make it very hard for them to do anything about it. So before you get too worried about bans, I promise I've done everything in my power to prevent it.
And of the bans that have occurred across me and my friends who do all of these things, of the 30 plus Claude accounts we're hopping between, I've had one friend get banned twice because he was on a VPN when he registered an account with a credit card for a business that was not his, it was mine. And those two accounts got temporarily banned and reinstated almost immediately after.
While this is probably not something the companies want you to do, it is a thing that they are not stopping us from right now. In fact, the only thing that can really stop us is a quick break for today's sponsor. I have two quick questions.
First, have you ever used an agent without search? If you have, you know how painful and miserable it is. It basically can't do anything.
My second question is the opposite. Have you ever used an agent with insanely fast and accurate search? I personally hadn't until I started using today's sponsor Parallel because they have the best and fastest search results for pretty much every single thing you can measure.
Historically, I've really liked the search built into OpenAI, but when you compare it to Parallel, it's just night and day. Watch and see just how fast Parallel can get you good results. Going now.
all real time of course it took just over a second for parallel and openai still going still going meanwhile openai's endpoint took over six seconds long if search was all they did they'd be one of the best options available but they also have everything else you would need on the web from a proper monitor that will send your agents info when things change on different pages to a traditional response api for when you want to actually get synthesized results from your search queries to their extract endpoints that lets you send a url and get back the data that your agents actually want and need even parsing js heavy pages, by the way.
There's even MCP so you can expose Parallel to your existing agents in whatever tool you're using now. Every month you'll get 5 ,000 requests for free. And on top of that, if you sign up today, you'll get $80 in credit.
What are you waiting for? Join now at soydiv .link slash Parallel. This is a real, real rough order of events that I have planned here.
We're going to start with how do you actually get enough tokens to do the terrible things I'm about to show you guys. Right after that, we're going to talk about actually accessing them once you've taken advantage of the ways to get more. Then I'm going to talk about how you actually use them practically well.
And of course, right after poorly, which I'd argue is more fun and much more important. And then at the end, the thing that we all want, how do you use your tokens when both you and your laptop are asleep? While I will do my best to make this a top level guide on like what things you should do.
The thing that will matter more is the little pieces that I show. Not because you should copy all of them, but they should help you understand how I think and operate with agents and how I've been able to do like 100 plus PRs a day for a bit now, part time. The point here is not to give you the exact formula to copy.
It's to get these ideas in your head and these examples floating around in your brain so that you'll apply some of these lessons in hopefully your own ways. The best possible outcome of this video would be you guys going off, building more cool shit with agents, and then coming to me with solutions for problems I hadn't even thought about yet that make my life and my agent maxing better.
I can already tell this video is going to cause huge blowback on Twitter. So I'm going to just list the questions they're all going to bring up to try and dunk on me for making this ahead of time. So let's go through the common questions and asshole comments that I know are going to happen the moment this video goes live.
This only works for side projects, not real apps. I have a good feeling if you're saying that my apps are more real than yours. I also happen to know that a lot of the strategies I'm showing here have worked at companies scaling as big as, I don't know, AWS itself, where I have friends who are learning these lessons from me, as well as at Microsoft, where I have friends who are learning these lessons from me, and even some Fortune 500 companies that are doing some of the subscription abuse stuff that I'm going to show y 'all.
We'll talk about that in a bit. Not all of us are rich. Cool.
I'm sorry. This is a great way to save a shitload of money, though. These strategies work at small scale and large scale.
And the crazy numbers you've seen me spending in the 300k plus range are not spent at all. I'm putting in like $1 ,000 for every 50 ,000 of tokens I'm taking out, like worst case. And if you're doing worse than that, you're paying API prices.
We'll get there in a bit. Of course, a JS dev thinks this. Oh yeah, the devs who understand the importance of having software people can use instead of software that people can gawk at and not use.
Yeah, you're right. I do know what people want, which means I am pre -qualified to talk about what people want in building real software. Wow, OpenAI has paid you off.
I have paid OpenAI more money than you're worth. Wow, Anthropic has paid you off. I've paid Anthropic more money than you're worth.
Now that we've got this out of the way, let's actually be useful. First and foremost, getting more tokens. Let's do some math here.
Who here knows how much money and tokens you get when you put in $200 on the Cloud API? I'll give you a hint. It's $200.
What about the Codex API? I'll give you a hint. The answer is the same.
Which is why what we will be doing today does not involve the API. You might think this is silly or dumb or not viable for real businesses. If I was to list the big companies I know...
that have their employees subscribing to personal to your accounts and just turning off the setting for data sharing, they'd all be really mad at me because these companies are very private and they don't want these things known, but they're all fucking doing it. You have no idea how many companies are just paying for the subscriptions.
And when you do the very basic math, you'll understand why. Because a $200 Claude sub gets you around $8 ,000 in tokens. That's 200 bucks a month for 8K in tokens.
There is a catch for this one though. Only 4 ,000 of that is Fable because they have the 50 % limit on Fable separately. That still means you're putting in $200 and taking out 4 ,000 to Fable.
That's a very good deal. The $100 tier is weird because you do actually get half of this total amount, but you get a fourth of the hourly limits. So the five -hour limit is much more strict, which means it's harder to burn that all on the $100 account.
So personally, I don't really recommend the $100 account. I think you should save a little more, do the $200, burn it, find a way to make money off how you burnt it, and then reinvest and keep going. And the $200 Codex sub is even more interesting because it's roughly $12 ,000 of inference.
And there is no distinction between what you're doing with Astra and with other cheaper models like Sol. They don't split it that way. You just get the tokens.
And if you guys think these subsidies are brutal, I got an even crazier one. Y 'all know how much Fable costs? It's $10 per mil in and $50 per mil out.
That sounds expensive, especially when you compare to other cheaper models, especially in the open weight world, where there are things that are doing real work for a tenth or less that price. What if I told you this number was also more than 50 % off? Fun fact, when Anthropic first announced Claude Mythos Preview in April, they actually quoted a price.
Even though the model wasn't out, this was just for internal use at companies for Project Glasswing for securing their stuff. Those companies were offered to continue using the model after at $25 per million in and $125 per million out. That means the price we're paying is like almost 70 % cheaper than what they intended to charge.
Why would they ever make it so much cheaper? It's because their margins are insane. prior rumored margins for how much money they made selling tokens at api price compared to like the energy costs and the eventual usage of the compute causing it to fail over time roughly calculated from various napkin math aficionados i know to around a 95 profit margin so if they charge you 100 bucks they made 95 and they spent five on replacing dead compute and paying for electricity or in the case of anthropic renting the compute from spacex so why would they knock this price down so much Well, Anthropic knocked the price down because they know everyone has the margins, but they were already in the lead and they wanted to maintain their lead.
So they decided to make the price lower in order to make it harder and harder to justify using competitors. I would argue that at 25 in, 125 out, this model was, especially at the time, worth it. But at $50 out, it's a bargain.
200 to 8 ,000. So the 1 to 40 ratio of subsidization means you're getting some really good deals. So you can kind of mentally double the numbers for Claude as well.
And also Codex, because there is no world in which OpenAI actually plan to charge the literal exact same rate as Mythos. OpenAI does numbers very simply. They keep them the same, they lower them 20 to 80%, or they double them.
Those are the only things OpenAI knows how to do, except for after Fable destroys them. And after that, all of a sudden, they're going to nudge the price a little bit. So the reason that Astra is as cheap as it is, is because of Fable.
And the reason that Fable is so cheap is they decided to eat their margins a little bit in order to guarantee a win, at least at the time. And this is all without even talking about the resets. Last I saw someone do the numbers, Tebow gave out 14 resets in 30 days.
that means you're averaging a reset every two to three days with your codex ups that means that the weekly number which is four grand is more like a three day number which means you can double the codex sub here with 24 grand if you're really maxing it out it's so bad right now that they've temporarily paused new subscriptions on the 200 tier for codex which sucks due to the nature of this video and i have a bad feeling this video is going to make the problem worse quickly sorry in advance I feel a little bad exposing these things to the world, but my job and my loyalty is as a journalist and informant, not as somebody who's trying to protect and shelter the few hundred people who know how abusable this shit is.
I think everyone should know, and I'm going to share this info with the world, and how things proceed after is how things proceed after. Don't be mad that I'm sharing. Be mad that I'm right.
Okay. When you set up your account, the first thing you should do in Codex... is hop over to settings, go to data controls, go to improve the model for everyone, and turn it off.
As soon as you do that, the terms of service for your personal sub are effectively identical to the terms of service for a team account. There is no real difference. So if you are paying for API prices because you think you're not going to have your data trained on, you're not saving anything.
You're just wasting money. Every dollar you spend on API is 80 to 90 cents you've thrown away. And if you're not thinking of it that way, that's on you.
There's a similar setting in the cloud code settings. I don't feel like going to find it. You get what I'm talking about.
So now with all of this established, the best way to get more tokens is obviously to get more accounts. And since we've established the accounts aren't meaningfully different from API pricing, you get a good deal. You should still use the API for a couple of things though.
One is very important. Anything facing users. If you are rating your prompts or your agents are rating your prompts, or the prompts are running on your system on your code base from things that are in your github issues or whatever all that is fine that is awesome but if you are trying to serve public traffic where users can submit a request and it runs on your inference that is entirely against the rules with the accounts if you're using this as an alternative to the api to serve user traffic you deserve your ban that's not what these accounts are for at all they are being very generous letting us experiment with coding with these huge subsidies You should abuse the generosity by building cool shit, not by reselling things to squeeze margin that they aren't.
I'm giving this advice up front because the things I'm about to show would make it hypothetically easier to do that. I think it's time to get there. It's time to talk about how we actually access the tokens.
In my first token maxing video, I gave the example of using two clod code subs. And I showed that if you have clod code running in the terminal, and it's completing a task, and you use another tab to change the auth for your Cloud Code instance.
It'll correctly recover and keep going. That is great when you have one or two computers and one or two threads with two accounts. That does not work when you are doing what we are going to be doing here.
You need systems that manage this, not just for your local Cloud Code with your one or two accounts. You need a solution that will handle the... insanity that is how often the cloud code auth breaks in a single shared bucket on a residential IP address that you can route all your other traffic through.
So let's go through some pro tips here. So tip number one, I love VPSs. We'll talk a lot about VPSs.
I highly advise you do not sign in to cloud code or codex, but especially cloud code on a VPS. anthropic in particular is very very nervous around people reselling their cloud code subs in vps's serving traffic to other people with cloud code because it lets them sell it for cheap and also collect a bunch of data they can use for distillation so they're being extra aggressive with sign -ins and often requests coming from ip addresses that are known to be servers like anything on aws or hetsner or any of these other providers so i highly recommend you do not do this over a vpn or on a cloud that you're serving the traffic from i would also advise against having that traffic come from multiple different computers at the same time because that's how they know you're doing some sketchy stuff so how do we avoid that how can we do our best to make sure most traffic comes from one place ideally one residential IP address, regardless of what you're doing, what account you're using, and where you're using it from.
Especially with something like Cloud Code, where the auth breaks every two to three days. Video from the future, because I forgot an important disclosure. Before I go any further, this might get you banned.
This might be against terms of service. I am just speaking from my experience. Do not hold me accountable.
I am literally drinking and filming as I talk about all these nerdy things. If you get banned and I didn't, that's unfortunate and I'm sorry. I can't be held liable.
You'll probably be fine though. Just don't do this with your Google accounts because when your Google account gets banned, you're fucked. Thankfully, there's no use for Gemini models anyways, so there's no reason to deal with that.
Cheers.
As we were saying, it's time to talk about the proxy. CLI proxy API is a thing that's come up in a lot of my videos, but not really been showcased in most of them. The way that this works is you sign into it using your auth for Codex and Cloud.
And once you have signed in, it will maintain the auth for you and give you an API that you can use in things like Cloud Code and Codex, where you point it to this instead of pointing it to the traditional OAuth servers. And now everything is routed accordingly. It is very, very convenient.
we'll get back to this in a sec because i realized i forgot one other thing we need to talk about in the getting more token section so one last thing here so i can actually wrap it up you might have noticed the subscriptions i was talking about i was talking about clod subs and codex subs i didn't include other subscriptions here why did i not include other things like the open code sub or the subscription to t3 chat or the subscription to cursor even the subsidies are comically Comically less.
Cursor does probably get a deal with Anthropic. The deal is probably 30 to 50 at best percent off. That is very different from 95 to 98 percent off.
You get a better deal than Cursor does. You get a better deal than OpenCode does. You get a better deal than pretty much everyone does.
Did a reset just get announced for Codex? Really? Kind of crazy that even in a compute crunch like they have right now at OpenAI, where they don't have enough compute for all the asterisks, they cancel the subs, they're still giving us resets.
I will point out it is currently Friday evening, which makes it a lot easier for them to do this because they are hoping the power users like us, yes, I'm including you now, viewer, you're joining us on this journey. The power users like us will burn it all to the ground before Monday when the enterprises wants to pay API prices again.
So yeah, the only subscriptions that you can really massively benefit from in terms of you put in money, you get more tokens out are Clod and Codex. If Grok 4 .7 comes out good, it might make sense. But I just, it is hard to token max with Grok because the Grok models just cannot stay coherent for as long yet.
So back to where we were with accessing the tokens. I gave you the prereqs. You want it on your home network.
You want it in one place, ideally. And you want everything to route through. CLI Proxy is a very cool place to start.
But CLI Proxy, not necessarily the best UI. The guy who got me into CLI Proxy didn't know it had a dashboard because he has avoided it to the best of his ability. Technically, I think the version he uses and that I started from was a fork called Vibe Proxy.
But I have extensively forked this project. I have made a lot of changes. And when you set it up and look at the dashboard, it won't look anything like this.
And if you want to change that, then congrats. You have your first token maxing task. Give it a screenshot of this and tell it you want it to look like this.
And it'll probably figure it out fine. Especially if you use Fable instead of Astra. There is one other change you almost have to make.
And I would argue these changes combined would be worthwhile if somebody creating a like real fork with an easier setup like Flow. Because the other change you need to make is account prioritization. By default, the way CLI proxy works is it routes traffic.
to all of your accounts evenly and it distributes across them which sounds great until you look at my usage here where you see i have one account that expires at 1 pm tomorrow and i have another that expires in four days obviously i want to burn as much as possible of the one that expires tomorrow and not touch the ones that have another few days the only reason you wouldn't want to do this is because you're constantly hitting five hour limits even with multiple accounts I can't believe I'm saying this in the token maxing video.
What I would say there is slow down a little because a five hour limit gets you through 40 % of your seven day fable. I've done the math. I know this one for a fucking fact.
So please do not give me shit. I know for sure that 100 % of the five hour on the 20X plan is exactly 40 % on the seven day. It's only 20 % of your full seven day limit.
But there's a reason I have this column grayed out. This is the opus bucket from hell. ignore that right column it doesn't matter we only care about the five hour and the seven day and we really only care about the fable seven day i often accidentally call my cloud accounts my fable accounts because opus is a trash model and you should avoid it to the best of your ability sup nerds less drunk future theo here wanted to call out something that didn't exist when i filmed this a few days ago opus 5 .5 i'll be clear all the advice in this video still applies and opus 5 5 behaves very similarly to fable 5 1 so anything i say about fable for the most part applies here too The big difference is you might not need five to 10 clod subs anymore because Opus 5 .5 on just one sub, I have found pretty hard to like hit your limits with.
So take that as you will. Maybe you don't need as many subs to take advantage of these patterns now. You might not even need to set up something like the proxy I'm about to teach you about.
But yeah, thought it was worth calling this out because 5 .5 is really solid overall and has changed my way of using things. Opus 5 was a useless dumpster fire that I had to work around a lot. 5 .5.
pretty good model back to whatever drunk theo was rambling about once you've gotten all of these signed in and you've had one of your models go through and make these changes changing the routing to optimize for which accounts have the next reset up soonest and the other one i think is important is by default your codec subs won't go through websocket so there's a few changes you can make to fix that which hugely increases performance just like the time passing messages back and forth goes down a ton.
So make sure you have the website stuff working. It's a little annoying, but you can. One last piece that's important, and thankfully my agents have been smart enough to do it right, is called session affinity and account affinity.
What this means is if I have a thread going, ideally the next prompt in that thread goes to the same account because caches are account specific. So if you are 800K tokens into a thread and you switch to a different account quietly behind the scenes with CLI proxy, Cloud Code doesn't give a fuck.
It doesn't know the difference. But the API then has to go recreate those cache entities. You only have to do it once per switch.
But if your stuff is switching per prompt or worse per tool call, then you're going to eat a lot of cache rates that you probably don't need to. Thankfully, the affinity stuff works very well by default in CLI proxy, and the caches only last five minutes anyways. So it's not too, too bad.
Just thought you should know because... If you're using a dumber model to set this up and it gets it wrong, like if you use Opus, it might get it wrong. So keep an eye on that.
Make sure that you are looking at how much cache reading and writing you're doing and that those numbers aren't getting crazy. You'll get a good feel for it soon. You can always just ask the agent, hey, how are we handling session affinity?
Are we writing cache more than we should be? And it can probably get an answer. For this next piece, I will request that you look up.
The top of the screen, you'll see the URL I'm on. BB1, which is my framework, my desktop in the other room. dot porpoise micro ts .net this is my tail scale this is my tail net 4318 is the default and then i'm on the management html page this url is being accessed this way because i am using it over tail scale tail scale is very very good to have for a setup like this because it makes it easy for other machines to access this proxy endpoint and it means you don't even need auth i was experimenting with some things yesterday and i set up an api key for a different thing and it broke my inference on everything because by default it just works with no api key this sounds horrible until you realize it's only exposed to other things on my tail net so if you don't have my google account that i use to sign into tail scale you cannot hit this endpoint so it doesn't matter so you're fine and this also makes it really really really convenient to set up other machines for this next part i'm gonna have to do a thing you probably didn't expect for me in 2026.
We're going to have to look at a real repo. This is a project that I have on GitHub and on my machine. And despite the fact that most of my projects are on most of my machines, this one's only on this machine.
That's because this is my fleet management repo. This is the repo that explains how all of the other computers I do my work on function, including the one that hosts my proxy and manages the tail scale for it so that everything else can connect. one of the files in here is how to set up a box this describes all the things i want set up in my boxes and all the random tools and things like i want rgfd jq tmx btop node python and build tools all of this codecs and cloud code installed i want them to use the proxy and all these other things all the details all the machines are in here gets updated whenever i add a new one so when i get a new machine step one i set up ssh step two i go to this repo I open up Clodder Codex, or in my case, usually D3 code.
And I say, look, here's a new box. Here's the SSH key. Go set it up the way I like.
And in not very much time, I'll all of a sudden have a page open in Helium for a tail scale approval. And like, oh, that was probably my agent running. I blindly hit accept.
And now I have the machine set up and it can do inference on these other boxes easily. So now I have what I needed. I have one machine on my home network going through a residential IP address.
that all five of my cloud accounts and all four of my codex accounts go through and all my machines all around the world connect to that over tail scale and it is impossible for anthropic or open ai to know that i'm using them for servers and something has been cloud code and codex is trivial it's so trivial that i'm not going to tell you how i'm going to tell you to tell your agent to do it it will swap two variables in the configs and you're done Thankfully, both Cloud Code and Codex have to work with API endpoints because they're being used so heavily with companies that are doing all their inference through Bedrock on AWS, which means they need to put a URL in.
So this will work indefinitely. It would be bad for them if it didn't. In fact, the Codex desktop app handled this poorly and enough people on AWS complained when Bedrock happened that they fixed it.
And now if you're using Codex with this setup, not only does it work great, it'll even show you in the corner the name of your proxy. It's pretty legit. And if you're skeptical or you really, really want to have your normal install with a normal auth, you can still do that.
And you can make another instance of Codex or Cloud Code with a different home directory that has this configured. So you have a different command for each. The other real cool benefit of this is if you do set it up, you get access to all the models that are in that setup inside of your Cloud Code.
Notice what I said there inside of Cloud Code. Do not, under any circumstance, Use this to use your cloud code sub in something other than cloud code unless you really, really like dealing with anthropic support and getting accounts banned all the time.
You will be banned. You will be upset. I highly recommend you only ever use cloud models through your cloud subs in cloud code over the proxy.
One last thing I think is worth knowing about the setup because it catches me off guard every time. Anthropic resets work different than codex ones. The difference is the dates.
When Codex does a reset, it resets your weekly entirely. So if your weekly was going to renew in a day or 10 minutes or whatever else anyways, your reset is now seven days off. So when the reset Tebow just promised hits at midnight, all my accounts that are currently three days away from a reset will now be seven days away from a reset.
And this means you end up with a weird pacing issue where you can quickly get to a point where you're six days away from a reset across all your accounts, which really sucks. It gets kind of balanced out with the banked resets, which across my accounts, I have three, six, seven, eight. I have two more of my other.
So I have 10 resets across my codec subs. And the two types of resets are banked and immediate. Banked resets are the ones you can click whenever and those cost them more money because the reason they can do these resets so freely is they do them in hours where they're not competing for the inference with their enterprise customers.
So they can freely give them on Fridays and Saturdays, especially Friday nights. They can't give them so freely on a Tuesday morning when they're about to have all their customers trying to use Astra at work that are paying API prices. So the banked resets cost them a lot more literally because it's opportunity cost.
I do see a future where they start pushing us to only use these subsidized subscriptions at off hours. I'm legitimately at a point where I would shift my sleep schedule in order to maintain this level of subsidization because it's so good and I'm fucking addicted. Let's be real.
And we all will be soon. So now that this is all established, the difference with Claude is there's only one type of reset. It's a limit reset and not a timer for the limit, just the percentage for the limit.
So if Claude was to do a reset right now, The only thing that would change is all of these numbers become 100 % again. The dates and times all stay the same.
So if I was to get a reset from Claude right now and my next actual one was in a day, that means I need to turn on the furnace immediately. I need to burn all the tokens I can because they'll vanish at that point. And this is the mindset shift I need y 'all to get in.
When your normal reset hits. So at 1 p .m. tomorrow for me with this account.
Whatever I have left here is money I lost. Don't think of this as you spent $200 and you got more. Think of this as you've been generously gifted $4 ,000, but every week where you don't spend a thousand of it, it disappears forever.
This is the opposite of the mindset you're in when you're paying API prices, where you're trying to get as much as possible for as few tokens. You need to have the depressed mindset where you feel yourself losing money because you are, and we don't want to be stuck at the permanent underclass. Burn your tokens.
Cool. That's most of what I need to show in this dashboard. I'll probably be back here.
I know I will be back here. Let's be real. I guess this is most of the accessing your tokens bit, which means next, you talk about using the tokens well.
This one is going to be full of all sorts of different layers. What I'll say for the core of this one is I have another video that's probably already out by now. that's about how I code without one of my hands.
It's a video all about my life after accepting that I can't really type even after my surgery and I get my hand back. I might not be able to type very well, which means I have to use my computer different, not just like voice to text, but I don't want to switch between apps as much as I can't command tab. I don't want to be staring at the thread as it generates because I have other things to do.
I don't want to be at my computer that much. I am sitting at my computer less and I'm coding more. and these are the strategies that will get you there that video has a lot of the good examples but i do want to give a couple important points that i don't necessarily think were in that and then also show you some dumb examples that weren't there that could be useful the question i want you to start getting into your head is how can i use more tokens to care less about this when you have a problem that you're trying to solve how early can you pull in the agent to start solving the problem And how long can you have it go until you need to take another look?
This is a twofold thing. You can even think of this in terms of a specific important button merge. You want to de -risk both sides of the merge button.
You want it to be more likely that by the time you go to GitHub and you're looking at the merge button, that it is safe to hit it because the agent has addressed as many of the potential problems as possible. It's kind of crazy to think of it this way, but I would honestly guess. that of my token use, maybe 10 or 15 % is actually coding and the other 80 to 85 % and the other 85 to 90 % is being spent verifying the code.
So by the time I am actually looking, the work's done and not like it's kind of done, but it has these bugs. If you can knock down the likelihood that the PR sucks or has some small issue from 5 % to 1%. you'll be able to do way way more because that means you can focus less on each pr so you want to de -risk merge on that side and you also want to after merge it needs to be easier to revert the things that are failing you need systems that catch things before they hit your users or systems that will revert things after they hit their users and you've noticed the problem if undoing a bad change takes more than 15 seconds you probably shouldn't be vibe coding at all yet You probably need to fix your systems.
So undoing something broken is way cheaper and faster. Then you can go a little harder. Holy shit.
I think I just saw the actual worst take of all time in my chat. The argument that I'm making is essentially you need to watch as much Netflix as you can because you need to maximize your subscription. I'm sorry.
We might be using a different Netflix. I just I personally don't know how I would use my unlimited movie watching to build real businesses. I don't see how I would use that to save 40x plus.
I don't see how Netflix can help you escape the permanent underclass. I've actually never, I hope you're rage baiting because this is legitimately the worst take I think I've ever seen. And it's my job to read shitty takes, which means congrats.
You can probably provide a lot of value to the world through Netflix. Because you've successfully baited me in the middle of what will be a very good video by sending the actual stupidest possible fucking message. So seriously, congrats.
Salute. Hats off. Fantastic work.
Back to real world. So we've talked about de -risking before merge by making the code more likely to be good. And we talked about de -risking after as well.
So if the code is bad, it's easy to fix. Whether or not you are using agents, these are things worth doing. Make it easier to verify code is good before you look at the PR and make it easy to revert if the code is bad.
But there is one other phase here, which is not really a button, which means this isn't the right UI, but whatever, you get the idea. Knowing what you want. There's a very good chance if you're watching this video and you're using agents for coding.
that by the time you have sent the prompt to the thread, you probably know what you want. I'm saying this because this is the case for me even just a few weeks ago. That is not the case anymore.
I have finally rewired my brain where I don't bring in the agent when I'm done thinking. I bring in the agent when I start thinking so we can talk it out. Maybe we find a shortcut that I hadn't thought about before.
I've had times where I thought about a problem, not like fully, but like pretty actively for three weeks. And then I went to an agent to build it. And I just asked, is there a stupid, simple solution I'm not thinking of here?
And it gave me one. And I realized I had just wasted three weeks of my time. I want to be really clear here because after the video about knowing your code base, I realized you guys don't listen very well.
So I'm putting this in largely to have a quote that I can grab when people misquote me from this video. I am not saying you should stop thinking. I am saying you should figure out if something's worth thinking about before you think about it.
Because you don't know until you put the time in thinking about it. But if the agent can solve it before you have to think about it, you both just saved a bunch of time. And tokens too.
Because if I ask the agent about a problem, and the agent has a good solution, then I don't have to trick it with my bad solution and run in circles a whole bunch with it. You save your tokens and your brain if you give the agent the problem instead of the solution. And this was a hard habit for me to break.
When people would DM me a bug in T3 code, for example, I would think through the bug, I would think through where the problem was, and I would go to my agent with a solution. I would tell it, I want you to change these things in this way, and then it would change it, and I would look at it, and I would realize it doesn't actually quite solve the problem I wanted to.
So now I just handed the screenshot to the agent and say, fix it. And it often does. And if it fails to, I now know this problem is too hard for the agent to do itself.
At which point, I know it's time to think more about it. I think it's been really cool to learn that a lot of the problems I would have used a bunch of my mental energy on didn't need it. And also, and this is even cooler, some of the problems I thought were really simple weren't.
So the agent outright failed. You ready for the spicy take here? This is similar to how it felt to be a manager.
When I realized that my team doesn't need to be handed solutions, it needs to be handed interesting problems, and they would usually come back with solutions. And if the first few times they come back with a solution is not quite right, then I know this problem is novel and difficult in some weird way, and I have to dive in and help steer it a bit.
So instead of diving in once you know the solution, dive in when you discover the problem. See if it can solve it autonomously. And if it can't, whatever.
Buy another cloud account. Or get your boss to. Seriously though, if you're not employed as a dev, Don't stack subs until you're making money like I know this shit's expensive But devs make a lot of money and the companies hiring devs also make a lot of money We're in the money when you have it burn somebody else's when you don't so the core thing I'm trying to say here is that at each of these stages from Problem to knowing what you want to do to the code being written and merged too many of y 'all live Not even this whole range, but like between a small set here where you are just letting the agent operate between knowing what you want and pr filed i want you to do this go all the way from where you first hear about the problem to when you hit merge and maybe even let the agent merge itself once you build more confidence and now your job is talking to users and of course sre for when things do inevitably fail the time thinking about this changes you in important ways so again dan i
Respect you heavily. You were fantastic to work with. But I'll push back on this specifically.
Because I, again, I even gave myself this quote earlier to make sure I have my get out of jail free card. I'm not saying think less. And I personally believe it is a better use of our time to think about problems agents can't solve than the ones agents can.
You will learn more and better yourself more thinking about the problems that your agents can't solve. or reading through the solutions that they can solve, then you would get thinking about a problem for days that an agent can solve in minutes. I'm not saying that you should replace your thinking with the agent.
I'm saying that you should optimize for thinking about things that matter more. Don't outsource thinking, scale it. Yes, exactly.
Pull in the model to figure out if you need to think about the thing. And if you do, that's where your brain energy goes. And we're going to get to this tip in a bit.
But part of the skill here isn't to give it the task, let it go do the thing while you go and play a video game. Since the agent's doing the task, it's time for you to do the next task, and then the next one, and then the next one, and then you see the first one done, so you go back and check it. And that's where this starts to get really cool.
If you let the agent run for this whole path, this can take hours. it is possible that from when you show it the problem to when the pr is ready to go the agent has to run for a couple hours maybe it runs for 30 minutes and it tests the thing quick it gets up a pull request it waits 15 minutes for getting i don't know like a review from any of our awesome ai code review sponsors and it spends 10 minutes fixing the changes puts it up again wasting their 15 for a follow -up review addresses those and now it's good to go that's an hour plus that you could be spending sending more prompts to other threads to do more things.
Here's where that ends, because I'm going to be so fucking real. I have been very kind and polite to the alternatives to T3 code, but they fucking suck when you have more than three threads going. And here is where we need to talk about a very good question we just got from Strawman Twitch.
How do you keep them all straight in your head? I'm going to be so real with you. If you're not using T3 code, I do not know how you do it.
I hate that this is the case because the competitors have been trying to copy us and they are not succeeding. And it is really genuinely frustrating. I want the Codex activity thing to be good.
I have offered to go there for free and write the code for them because I want to use these things well more than I want us to win with our open source project that makes no money. So what the hell am I talking about this arrogantly? I'm not actually arrogant about this.
I'm pissed off about this because the solution was really simple. It was admittedly one that I thought about for far too long because I was trying to figure out how do I deal with the fact that I have, let's be real, far too many threads at any given time. I realized that I'm not treating threads properly.
Threads are not histories that you actively go back to all the time. Threads are tasks. They're to -dos.
And when they are not working, they should not matter. I have a bunch of stuff here because I've been streaming, so all of these things are done. Also, I was at Demo Day yesterday, but normally at any given time, I got 10 plus things running and 5 plus in the done state.
And this is where things get really, really cool with T3 code specifically. When you are done with a thread, you check settle, and now it is gone. I know, I know.
Theo thinks he's so cool for giving a new word to archive. This is not that. Settle is different.
The goal here is to turn your sidebar into inbox zero. You should try to end your day with all your threads gone or running. So when you go to bed and wake up the next day, you have some cool things to look at.
But that doesn't mean we have cured the context switching costs. We haven't. We have just made it easy to see which context needs your attention.
And this is where you have to start rewiring a bit. This is where if you have ADHD, you have a solid advantage. Because once you get in the mindset of, Thread is up.
Not my problem now. You're not going to get out of it. What's something I want to add to T3 code?
I'm going to be so real. T3 code has progressed so much and added so many of the things I wanted. I don't really think about it that much anymore.
Like I don't, I'm running out of things to add, which is great. It means we're doing very well. Once orchestrator V2 is in, that will change.
But let's say I want some things and I do. I have a couple in mind. Here's the first one.
I would really like a simple and minimal queuing system where I can send a message and it will show as pending and it will go up when the next tool call is completed, or I can click steer and immediately send it. The attached screenshot shows an example of what I'm thinking of here from another app. Since I don't have a good way to get the screenshot because my codecs don't reset for another three hours or so, you should be able to find a good example of this in Codecs.
Look for screenshots online if necessary. I have two tips I'm about to give you that are important. They're going to be rapid fire, so pay attention right now.
Tip one, look at the computer I have selected on the bottom left. This is one of the things I think T3 Code does exceptionally. Right now, it's Theo's MacBook Pro.
That is the computer I'm currently using. This computer is going to have to be closed when I go downstairs later. I don't know if this thread will be done by then.
So I am not going to run this on this computer because I don't run anything on this computer other than my fleet management because I'm doing that in the loop. So we're going to click here. There's going to be any repo.
And as long as the origin for Git is the same on the different machines, they will all be bunched under here. So I can pick any of my servers and pick where I want this to run. Usually I pick one of these three because they're my Linux boxes.
But if I really want this to work on mobile or do computer use, which right now is better in macOS, I can pick my two Macs at the bottom here to do it. I don't care for this one, so I'm just going to throw it on a random cloud server. Cool.
Now it's going. And I just realized I missed the second tip. So I will give that after showing another thing that I want to fix.
I hop in here to connections and scroll down. You'll see all the devices that currently have connected. I hate the UI here because you have update or disconnect and no remove for my T3 connect machines.
And disconnect doesn't seem like it is or isn't permanent. So it's not very clear. So here's what I'm going to do.
I'm going to grab a screenshot of the whole UI here. Screenshot grabbed. We're going to go back.
Command shift O, enter, new thread, cool, paste. I really want to rethink the UI for the devices connected in T3 Connect and settings on the desktop app. There's a couple issues I have.
First, I feel like disconnect doesn't seem as temporary as it is. It would be nice if that was a toggle instead. I also don't like that when update is an option, remove isn't.
I should always be able to remove and it should be very clear that it's a permanent destructive action. If you think this is simple enough to do directly, go do it and send me a screenshot. If you feel like you aren't quite sure, make me a few mocks using my HTML skill and we can decide between them.
Cool. There's a couple of things I did here that might be useful. First, I gave it the exact issue I have with not a prescribed solution, but the problem and then some ideas of solutions.
I then told it that it can go do it directly if it has a solution it's happy with. And I also told it that if it doesn't, to give me mocks using a skill I built so that we can make a better decision together. And now if it does make one solution, I don't like it.
I can tell it again, like go make the mocks and it knows what I mean. And here is where one of my favorite T3 code pro tips comes in. I'm not going to press enter.
I'm going to press command enter, which sends off the thread in the background and leaves me exactly where I am. so i can now send off another prompt without having to do anything it's very nice and this helps you get into the mindset of oh that's a problem i will go fire it off and i will look when i am done i do not check my threads until they say done or input in the corner and if you're too lazy to even pick which server to run maria has built an awesome feature for you in settings here load balancing You can now set up auto load balancing across all of the servers that you have connected in T3 code.
So it will fire off the threads across your different machines. And once you work this way, your job becomes different. You're no longer sitting there carefully babysitting the exact change.
You are firing off different things you want to have done. And then you go through them when they are ready for your attention. And here is the harshest reality I need you guys to get through your thick fucking skulls.
And this was so hard for me. That's why I'm being mean. It took me forever to accept this.
Agents can be multi -threaded. Humans are single -threaded. You cannot focus on two things at once.
You cannot do it. This means our job's a bit different now. We are the bottleneck.
So how can you get the agents to unblock themselves as much as possible so you only have to come in when they need you? And how do you make it so they need you less and that you come in later? So look at this one that I filed because people were mad about how I think about streaming.
I wanted to move to a chunked by paragraph solution potentially. So I sent a prompt asking specifically how hard would it be to do this? If I really was committed and wanted this, I would have just told it to go do it.
But I wanted to see if there's any difficulty here. Before doing it, because if it's any friction, I just don't want this feature. So I asked, and it said not hard, and then it hallucinated.
It said about half a day, all in one. No, I don't fucking care about what you think half a day of work is. It's kind of funny that these models still don't know how long work takes.
Yeah, eventually it'll be fixed. Regardless, it said it would be pretty easy. It said a bunch of variable names that seemed fine.
I then said, can you build it for me? Once you get it working, file a PR and spit up an environment with tail scale for me to try. Now, and this is very important.
Imagine I didn't do the second sentence here. Imagine I just said, can you build it for me? And then in five minutes it comes back.
Okay, I built it. And then I'm like, okay, can you spit it up on tail scale so I can try it quick? I leave and then I get the ding.
I go back and then I click the link and I try. I'm like, okay, this is great. Can you file the PR?
And then it does. And I look, I'm like, oh, cool. You should babysit this too.
I probably should have included babysit here as well. So I'll do that now. Can you babysit this PR and make sure everything is good?
Cool. Now it will continuously monitor that PR and it's not my problem again. So your goal here is to make sure everything you need is there the next time you click.
Ideally, you want to maximize the chance that the next time you check that thread, you're ready to merge. And really think about that. What can you add?
What context can you give? What tools can you let your agent use to make it more likely that by the time you go back to that thread, you can merge the PR? and this is where one of my favorite t3 code features comes in when you merge the pr the thread disappears which means if you tell the agent you can merge the pr if it passes these conditions or meets these requirements then you can send the prompt and never see the thread again i would honestly guess that around half my threads are archived without me ever seeing their final message because it doesn't matter i see people confused about the agents stepping on each other's toes while they're working i thought we were developers guys We're not vibe coders.
We made the right default in T3 code for this for a reason. By default, when you start a new thread, it starts in a new work tree. Not going to pretend work trees are perfect, but they are good enough.
And now that the models are smart enough to deal with the weird bullshit that is Git, they will fix your work tree issues for you. For example, what I run into a lot is I have a work tree with a branch. that makes a pull request.
And then I want another agent, usually another model, to review it and give thoughts. And if those both are on the same machine and they're both on work trees, they can't both have the same branch and then Git freaks out. Previously, models were dumb enough they would get stuck there.
Now they can figure it out. They have workarounds. They'll make a new branch.
They'll pull it down in some other way. They'll make a clone. They'll do whatever they have to.
It doesn't matter. I don't look. I don't care.
Models are smart enough that if I give it a pull request with a branch that it already has in a work tree, it'll figure out how to read it. don't spend your time thinking about those things anymore. They don't matter anymore.
This is another habit I've gotten into. By default, T3 code auto settles threads that you have not touched for at least three days, which I think is the right call. I pumped it to seven because I usually have better discipline about clearing out my inbox, but I haven't been good enough at it here, which is why my sidebar is a little chaotic.
As exploring thread pop -outs, this one I actually do want to play with later, but not now. So I'm going to use the snooze feature to make this my problem later. You'll start to see my mindset as I go through this.
I want to make it so I don't have as much stuff trying to take my focus. This is why I've done some strategic things with the UI and the sidebar. Notice that the ones that are working are semi -transparent.
I wanted to hide working threads, but nobody would let me. So I made them more transparent so you don't look as closely at them. I like this a lot.
It has made it much easier for me to only prioritize things that need me, things that are done or things that are waiting for input. i also really really really want to emphasize there are few worse uses of your time than watching an agent as it works if it works and it succeeds awesome merge the code if it works and it fails look maybe read the reasoning trace or even better ask the agent why this is a good opportunity to pivot into the next section here which is using your tokens poorly i know it sounds silly but i'm actually gonna be framing these things in ways that could genuinely be helpful god i One last thing I want to crash out at a little bit.
You can censor the chatter if you prefer, Jeff, in the edit. I see a lot of sentiment like this still. Like, if it works, that's fine, but what if it installs some malware on the side?
I need to be so real with y 'all. If you think your agent is more likely to install malware than you are, then you are dumber than your agent. If you are smart enough to realize that won't happen, then you're both smart enough to not install malware.
if you think the agent is more likely than you are you have installed malware recently and you probably have someone on your machine you should consider a new windows install because i know you're also using windows let's be real here anyways we're going to go back to talking to real engineers because that's what we're here for let's talk more about using tokens poorly one of the things i want you to think about is when you face a problem or have a question or uncertainty about something how can you use your tokens instead of your brain When I have a pull request that is hard to parse that a teammate filed, how can I use tokens to figure out what it is?
If I'm curious what's changed in the Orchestrator V2 rewrite that Julius is working on, I could ask him, but he's busy, he's prompting. I could rather just ask my agent to get me the info on what he has been working on. If I've lost track of all my pull requests and I don't know what I should focus on today, I ask my agent to go through them and find something that I should be more focused on.
If I can't find an email from one of my five email inboxes, I open ChatGBT and tell it to go find it, and it does. If I don't want to sit there and download all my medical records for my surgeries, I don't ask my assistant to do it like I used to. I ask Codex to go through my inbox and go through all of my dashboards for my medical shit on my browser and get it all for me.
And what's even more fun is I would have my computer going through and spending 30 plus minutes collecting all my medical records. And while it does that, I go back to T3 code and I fire off four more prompts for things that I notice that are annoying me or for a feature I want to iterate on. And then every couple of minutes I go back, I take a quick look and I see, oh, these things are done.
These things are still working. Cool. I really need to finish the durable objects.
I was going to save us a shitload of money if I can get it right. This is how I think about it. The issue with the other tools that you can use for this is that the sidebar doesn't help you prioritize the work that you're doing and the work that is done.
It doesn't make completed work go away. The search isn't trustworthy enough to get old work back. And you can't use the same sidebar across multiple machines.
Like of the threads here, this first one is on one of my MacBooks. This one is on my server Alvin. This one's also on Alvin because I just spawned it.
This one was too because I just spawned it. This one's on BB1. This one's on my MacBook.
This one's on this MacBook. This one's on the other MacBook. This one's on my main cloud server.
They're all in different boxes and it doesn't matter. They're all in the same UI and they're all on my phone app too. I will say that right now, sadly, if you're not using Tailscale, you'll only be able to connect three devices at once with T3 code for free by default.
We're working on it. I want to bump the number a bunch, but since you're already going to be using CLI proxy, that means you're also going to be using Tailscale. That means you don't even need to use T3 Connect.
You can just connect directly. This all works without using our servers at all. T3 Code is fully free and open source.
It fully supports your subscriptions. It works incredibly with CLI proxy. So if you have a couple boxes that have Cloud Code and Codex, and you have them CLI proxied, you go to one of those boxes.
I'll show you just how hard it is to set up T3 Code on a new box. I don't have a new box to set it up on. All my boxes are configured.
Let's say they weren't. You go to the box, you SSH in, you're on NPX, T3, connect. And then you click the link.
Oh, fuck, email leak, great. I should work on that, but yeah. You click the link that comes up, you sign in, and now as long as you're signed in on the client, on the website, the mobile app, and then once you're signed in on the mobile app, the website, or the desktop app, you can now control that machine.
You don't have to install anything. You just need CloudCoder Codex, ideally with a sub or a proxy. You run the T3 connect command, and now you can connect through our layer.
or you can do npx t3 serve and it will instead let you use tail scale we even have a dash dash tail scale built in directly or crazy thought i don't know kind of risky here but we are in the use your tokens poorly section you can just tell your agent to go set it up you have a new server and you're already building a repo that manages your servers you can tell quad or codex when you set up that server with quad and codex you should also probably set it up with t3 code so i can connect remotely over tail scale and it will just do it and it will just work and it will be just great i'm gonna frame the stupid prompts a little silly have you ever found yourself google searching where did i leave my keys because you got so in the habit of google searching things that you just default to that if you don't find yourself doing that with agents you're not offloading to them enough you need to get to the point where you ask a thing to agents that it obviously can't know and you feel silly for asking
If you don't find the things you're asking the agents a little bit silly, then you're not asking it silly enough shit and you're not pushing the limits here yet. Here's a fun example. I have a lot of side projects and I had a bit of inference to burn the day I sent this.
So I asked, I want you to go through all my GitHub projects as well as unfinished work on this machine using lots of sub -agents to help me figure out what I should be putting more time into. What are some of my ideas and side projects that seem like they're more in demand now that I might have forgone or that are worth bringing back and finishing?
They might be half baked or projects that have been abandoned for a while and deserve another pass. Go through all my work on this machine and GitHub and find things I should potentially revive. This is a shitty prompt for a shitty problem.
None of this matters. But I feel bad when my limits reset and they weren't at zero. So I was looking for things to burn them on so that I could have them all dead and not feel bad when they reset.
And it found a handful of my things that it thought I should prioritize more. I think that is fun and cool. And I have already decided what I want to prioritize.
So now the next step. snooze well subtle i don't need to snooze because i don't want it back context users breakdown i decided what i want to do on this one and what i want to do is kill it so we're going to go here and i'm going to close it and i'm going to archive it because i don't care anymore show works recreation progress i don't know how far we got with this how far along is this work is the tail scale dev server still up spin it up if you can i'd also love for you to file a pr and babysit it to make sure everything is good to go and i'll come back later Here's one that I've forgotten about because it was a while ago.
It says three days, but I'm pretty sure I started this way before then. So I'm just going to ask. I'll be so real.
I haven't kept up with this workflow in a bit. I have no idea what the state of this is or what the value is. Can you give me a rough idea of where this PR is at and why we should merge it?
This was going to fail because I'm still mostly out of Astra. It might route correctly. I have a weird routing issue right now.
We'll figure it out. These two are still going. That means that they are legitimate prompts doing legitimate stuff.
Awesome. I will come back to them later when they say done. I would not have ever clicked these two if I wasn't making content because I don't care when they are still working.
Here's a real example of what my sidebar looks like when I'm actually working. Do you see how many threads I have here? There were even more below.
1, 2, 3, 4, 5, 6, 7, 8, 9, 10. Then a monitor. There's at least like two or three more underneath that.
And across my five clod code subs. i was able to do all of this and still have some usage left over not bad and after this i went to bed and i woke up the next day i went through them one at a time made a decision do i want to merge this do i want to do a follow -up do i want to get rid of it what do i want to do i went through one at a time did all that and after i noticed some other things i wanted to fix so i spun up more threads to fix them and i went back and looked to see if everything was done and went through and clicked all the ones that were done and did it I really have been treating this like an email inbox now, and it has made me so much more productive.
So I just said they're confused because those threads have only been running or those threads only ran for 20 minutes. I had spun them all up within the last 20 minutes. Some of them took 30 minutes.
Some of them took 18 hours. That was me trying to get everything going before bed. That wasn't work going.
Those all said working. I had. 12 plus threads that had been working for at least 20 minutes.
Most of them kept working after. All of them actually did. I don't think any of them were done even with the next 10 minutes.
So as silly as the using your tokens poorly section may have been, I do want you to get into that mindset, especially when you notice a limit's coming up and you haven't burned it. Find silly things to have your agents do. you'll learn things from that you'll learn a lot more than you expect from that i am still amazed at how good agents are at triaging large numbers of prs at finding things that i should bring in in merge that i might have missed or forgotten about otherwise you'll be amazed at how much random you can find and here is where we get into the last section i've touched on a good bit of this here but i want to really really emphasize this part using your tokens while sleeping i mean this both literally and metaphorically If you've ever had the feeling of, I want to fix this right now, but I have a meeting coming up or I have to leave the office soon.
So I'm not going to send this message to this thread because I have to close my laptop. I get it. I was like this not long ago.
You need to move your dev work remote. You need to pick up a crappy little Linux box. You need to take some old computer that you haven't used in a while, install Ubuntu on it, set it up with Tailscale, install CLI proxy on it so that it's always coming from your residential IP.
and then use it as a box that your agents can work with. Now, connect it with something like T3 Code, and you can fire shit off on that box and not have to think about it. I see people asking, if they don't have a computer, what's the most budget option?
The most budget option by far is to talk to your friends and family and find somebody with an old laptop or desktop that they'll give you for free that has at least eight gigs of RAM and at least four cores, and you'll be able to do some real work there. And I'm already seeing some very, very stupid replies. If you think a VPS is the solution here, I might have to get into selling VPSs because those sound much more lucrative than bridges nowadays.
Need 32 gigs of RAM and 16 threads? You can get it on Hetzner. It's only $275 a month.
That's a great deal. Or you can waste all your money spending $700. on a box that is 32 gigs of RAM and a one terabyte drive, that would be terrible.
That's such a waste of money. That's like two and a half whole months of renting a worse computer on an IP address that'll get you banned from Claude. Why would you ever buy hardware that you can run on an IP address that won't get you banned when you can rent something for two months that will get you banned?
Hopefully you understand sarcasm because there is pretty much no reason to do cloud servers unless you're grandfathered into a good deal. Theo, you're wrong here because insert something stupid. It's actually very funny someone said that in chat right after somebody complained about electricity costs for a computer that pulls conservatively 60 to 80 watts of power at worst.
And remember, it's not doing inference. It's doing code. It's going to be using one or two threads most of the time.
Let's go hop into my massively over -specced server to take a look. I have six threads or so running on Alvin right now. And let's take a look at Btop quick.
Huh. You can't see because my face is covering it. This is a box I have at least six threads running on right now.
It's got 32 cores because I got it for a good deal. You might have noticed it's not using them very much. It's rounding to 0%.
So once again, you don't need a lot. You just need enough. Sadly, that box I shared earlier with the GMK tech has gone up in price because my tweet caused them to sell out.
You can still hunt, find something, but I would highly recommend you do not do a Mac for this because Mac OS is a shit show for parallel work. If you are looking at the numbers I'm sharing here and saying that's not right, I run three agents in Codex and my laptop's overheating. You're right, your MacBook is overheating because your MacBook has a shit file system and an even shittier security policy that is hard bottlenecking how much shit you can run at once.
Move to a real OS like Linux and you won't have these problems. You don't need a lot. And I bet my ass if you can find your mom or uncle's old PC that's got four to eight gigs of RAM in it and an OK -ish chip and you flash Ubuntu on it and you plug it into your router somewhere.
You're going to have a better experience than you have doing dev work on a Mac, like immediately. And if you live in an area that isn't San Francisco, well, I'll be real. If you live in San Francisco, hopefully you can afford compute.
And if you can't get out, the city's too expensive. So if you live anywhere else, Facebook Marketplace will be your friend. I bet you can find a surprisingly good deal.
And any old laptop with any real RAM, you're going to be fine. And once you make that change. Suddenly, you're not bottlenecked the same way.
There will obviously be edges, like if your agents run CI on the machine locally, and it's a Rust project that saturates all your cores when it compiles, you'll run into problems. You're an engineer, though, and you have tokens. Burn your brain and your tokens to solve the problem.
Maybe you move CI to only run on GitHub, or you use one of our awesome partners like Blacksmith or Depot. Maybe you have a different server that runs the CI. Maybe you set up a queuing system where the CI gets triggered one after another so you don't have five Rust compiles destroying your machine at once.
You're engineers. You can solve those problems. I don't want to just tell you all the solutions to those things because your problems will be different from mine.
Apparently, I succeeded with my goal of moving Maria to the Blacksmith CLI in order to let her run the CI via CLI instead of it running on her machines directly because she was a bit RAM constrained. And now she's way less RAM constrained. And that's how we end up with posts like this from Steven that Maria shared.
Request. Can you make T3 code less productive and experienced so that he burns fewer tokens? This is where you want to be.
And I have one last thing I want to lean into with this bit. This shit's so fun. I know a lot of engineers have felt like the thing they love died.
And that everything has changed too much. And now engineering isn't fun anymore because you send off a prompt and then you sit there and watch it make a bunch of mistakes and then get annoyed that the code doesn't work and you could have wrote it faster yourself. I still have that experience.
The only difference is I don't watch the thread. I go do something else. And then another thing.
And then another thing. Maybe I spin up three threads and then I go to dinner with my friends. Maybe I spin up seven threads while I am waiting in the lobby to queue in a game.
Maybe I go skate while I have some things running. And when I'm sitting and recharging, I pull out my phone and check on two of them quick and kick them off to go do a bit more work. And the fun isn't in those parts.
I'll be clear. The fun isn't that I'm checking my phone and sending prompts. The fun is this new type of engineering problem.
I can now justify spending more time micro -optimizing how and when I trigger CI. I'm thinking more about my cores and how they're being used. I'm thinking more about how I can work on the same thing eight times on one box and not have to think as much after.
And I love doing this. I find it so genuinely fun. And I hope this video helps inspire some more of this fun for y 'all.
Because that's my real goal here. All of this chaos has made engineering the most fun I've ever had with it. And if you're not having fun yet, I hope this can help you get there.
And maybe just maybe... you have a reason to go spin up a couple more accounts and burn a few more tokens. And if you do want to throw some of those tokens our way to make some real improvements to T3 code, we'll probably ignore them, but we might merge them.
So consider it. Oh yeah, I did have this DM back and forth with Maria. I could probably scroll and find it, but I don't want to scroll through our DMs publicly.
I set her up to do remote stuff and she got a Linux box. You could do it at her place. She hits me up in all caps.
Remote work is the best. This is so fun. and i said i'm sorry you're about to burn so many tokens and if you've been struggling to hit your limits i really hope you don't struggle anymore because this is so so fun i think i've said all i have to on this one i'm gonna go kick off some more threads and grab another drink enjoy the new world of agent maxing and get as many of these tokens out as you can before the subsidization ends if it ever does let me know if you want a video about that too because i have a lot of thoughts about how this goes long term But this wasn't about that.
This is about actually using them, and I hope it was helpful. Let me know, and until next time, peace nerds.
The Hook

The bait, then the rug-pull.

Theo opens by asking how many agent threads you have running right now, then spends the next hour making the case that if the answer is under five, you're leaving thousands of dollars of paid-for inference on the table every single week.

CTA Breakdown

How they asked for the click.

FROM THE DESCRIPTION
PRIMARY CTAWhere the creator wants you to go next.
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.

41:32
Theo - t3․gg · Tutorial

So I was using Fable wrong...

Theo spends forty minutes inside Anthropic's own Fable 5.1 prompting guide, rebuilding his habits around effort levels, finishing the whole task, and trusting the model's defaults instead of babysitting them.

September 22nd
33:48
Theo - t3․gg · Essay

He's right.

Boris Cherny said coding is solved. Matt Pocock called it VC-funded bullshit. Theo argues they're both right, because they're using the word coding to mean two different things.

August 24th