A four and a half hour Labor Day stream where two entire YouTube videos get filmed live, one-handed, between sub thanks, a ban, and forty agents running in the background.
Posted
1 weeks ago
Duration
Format
Review
educational
Views
39.7K
389 likes
57 · 43
Big Idea
The argument in one line.
Two frontier coding models are now close enough on quality that the deciding factor is variance, because one lands mergeable work at a steady bar while the other alternates between astonishing and unusable.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
A developer paying for one 200 dollar coding subscription who has to decide which lab to give it to this month.
Someone running agents in parallel who wants a working mental model for how early to hand off a task and how long to let it run unattended.
A solo builder whose input speed is limited by injury, RSI or environment, and who needs voice and computer use to carry real work.
An engineer trying to explain to a team why cost per task matters more than the posted per-token price.
Anyone who wants a candid, unedited look at how a technical YouTube video actually gets made, sponsor reads and retraction takes included.
SKIP IF…
You want a clean, edited model review. This is the raw stream the videos were cut from, chat interruptions and all.
You are shopping in the 20 dollar tier. The advice here is explicitly to wait for this level of capability to get cheaper.
You want benchmark methodology. Most of the ranking is stated as felt experience across real projects, not measured runs.
TL;DR
The full version, fast.
A torn thumb ligament forced a full rewrite of one developer's workflow, and the result is a working method rather than a coping strategy: dictate through a close-talk mic, keep a repo that documents every machine so agents can act on any of them, and push the agent in earlier and let it run later than feels safe. The same session then compares two frontier models across roughly twenty categories. One is better at 3D, computer use, orchestration and token efficiency. The other writes code that merges with fewer follow-ups. The conclusion is that consistency, not peak capability, is what you are actually buying.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Eleven minutes of a black title card while the tea steeps.
11:09 – 18:20
02 · Live, and the only pro-AI voice on the panel
Stream opens, subs get thanked, and a chat request for an AI risk panel turns into a note about creators who use AI daily but stay quiet about it publicly.
18:20 – 24:30
03 · Why the Melee decompile matters
Super Smash Bros. Melee is fully decompiled, which means native builds, real bug fixes, and an end to controller-coordinate hacks.
24:30 – 32:40
04 · Thanking subs while the agents run
Sub and cheer shoutouts run alongside a Melee decompile fork on ultra and a TypeScript to Rust port, both burning usage until the reset lands.
32:40 – 39:00
05 · Orchestrator V2 on a custom build
A side build of T3 Code demos an agent opening a second thread on another model and reporting back when its exploration is done.
39:00 – 45:07
06 · Content planning on the whiteboard
The video backlog gets read out loud and the model comparison is split into two twenty minute videos instead of one fifty minute one.
45:07 – 49:25
07 · Playing chicken with the third-party ban
Holding back the orchestration overhaul is partly strategic: the more of a lab's paying users route through your harness, the less likely they are to shut it down.
49:25 – 53:20
08 · The ten thousand dollar bet
A heckler calling him a corporate shill gets unbanned, offered a wager over declined sponsorships, then banned again on air.
53:20 – 1:00:20
09 · The hand: a torn ligament and 50/50 odds
Full injury update. A thumb ligament torn since a teenage fall, reconstructed last August, torn again. Immobilization now, coin-flip odds on another surgery in October.
1:00:20 – 1:12:05
10 · Fixing his own load balancer, then picking a video
An agent is sent to find out why account routing never prioritizes the soonest reset, and the first video of the day is chosen: coding without typing.
1:12:05 – 1:22:06
11 · Video one begins: staying productive without typing
Cast reveal, sponsor read, and the core admission that dictating code does not work. What changed was offloading navigation, not transcription.
1:22:06 – 1:29:06
12 · The seventy dollar podium mic
A live demo of dictating at a whisper next to a 500 dollar broadcast mic, then again with a video playing loudly nearby. The blocker on voice was social, not technical.
1:29:06 – 1:39:19
13 · Terminals are hell, so Fleet replaces them
A repo that documents every machine he owns lets an agent move a not-yet-downloaded ISO between computers, and build a custom app from a phone in the bathroom.
1:39:19 – 1:44:51
14 · Forty nine medical PDFs by computer use
A hostile hospital portal gets scraped end to end in about twenty minutes, then the files get uploaded to the new doctor by a second thread.
1:44:51 – 1:52:20
15 · Where the agent starts and where it stops
A hand-drawn spectrum from idea to done, and the argument for expanding the agent's span on both ends. Includes the merge hole and the 150 autonomous merges.
1:52:20 – 2:01:20
16 · Speed stopped mattering
A thread-by-thread audit of how much he cares about latency, the answer being almost never, plus the multi-thread kickoff shortcut and why typos in prompts are fine.
2:01:20 – 2:06:48
17 · Stop passing context by hand
GitHub gets replaced by a prompt that summarizes recent pull requests in-app, and the closing challenge: fire the agent before you plan, then compare.
2:06:48 – 2:15:12
18 · Chat interruptions and the timeout button
A long, unsparing explanation of why questions during filming get a ten minute timeout, plus a detour on whether he counts as a polymath.
2:15:12 – 2:28:20
19 · The fifteen hundred dollar box behind the voice
Chat guesses the price of a three-input field recorder, learns why 32-bit float and preamps cost that much, and hears the Cloudflare bill that makes the free tier untenable.
2:28:20 – 2:37:40
20 · Video two begins: building the comparison board
A fake showdown between two irrelevant models as a cold open, then the real comparison grid gets drawn live: code, non-code, agent behavior, cost.
2:37:40 – 2:48:50
21 · Science and 3D: Astra clears
A cheaper model beats the rival's headline science benchmark at a third of the cost, and 3D rendering jumps a full generation, with submarine games as the evidence.
2:48:50 – 3:01:40
22 · Computer use, copy, and the all-caps problem
Computer use is the second big gap. Prose quality improved on both sides, but one model buries every interface it builds under unnecessary all-caps subtitles.
3:01:40 – 3:12:30
23 · Frontend and full stack
Side by side landing pages settle frontend quickly. On full stack both models now understand a codebase end to end, and the difference is temperament.
3:12:30 – 3:27:58
24 · Giant rewrites and the prompt it ignored
A TypeScript to Rust port reaches 82.6 percent and stalls. A ping.gg rewrite throws away years of paid design work despite an explicit instruction to reuse the UI.
3:27:58 – 3:40:07
25 · Code mergeability
Two follow-ups versus six, charted on a deliberately unscientific graph, with a careful correction that the gap is nothing like the previous generation's gap.
3:40:07 – 3:52:28
26 · Understanding intent: the revert saga
A prompt asking twice for a revert produces the wrong dev server, an allowed-hosts commit, and an unrequested merge. The same text on the other model fixes it in five minutes.
3:52:28 – 4:02:09
27 · Swarms and steering
Sub-agents that message each other and report upward, forty running in parallel, plus mid-run steering, self-prompting, skill writing and skill usage.
4:02:09 – 4:10:04
28 · Cost: cache writes eat sixty percent
The cache read price cut is dismissed as a rounding error, and cache writes are named as the real bill. Cost per task and token efficiency close the section.
4:10:04 – 4:20:22
29 · Subscription limits and the Opus pot
Five accounts on one lab, four on the other. Weekly caps, the half-allowance reserved for the top model, five-hour windows, and free resets that arrive unpredictably.
4:20:22 – 4:29:22
30 · The verdict: steady versus spiky
A two-line chart of output quality over time delivers the answer. One line is flat and high, the other swings from ten to two, and that is the whole decision.
4:29:22 – 4:35:12
31 · Sign off
Final sub thanks, a chat line he calls a banger, an admission that he should have plugged his own product harder, and a clipping mic to fix tomorrow.
Atomic Insights
Lines worth screenshotting.
Cache writes can be over 60 percent of an inference bill, which means you pay more to hold state in memory than to run the compute.
A four times cheaper cache read saves under 1 percent of a real bill, because cache reads are only about 3 percent of spend.
Cost per task on the same posted price differed by roughly three times, because one model finished the work in a third of the tokens.
Roughly 150 pull requests were merged fully autonomously across two models, and only two shipped regressions, both of them removed animations.
The useful measure of a coding model is follow-ups between pull request filed and merged: two is a good model, six is an expensive one.
A rewrite prompt that said reuse as much UI code as possible produced zero shared lines with the original, and merged itself anyway.
A prompt containing the word revert twice produced no revert, an unrelated dev server, and a merge of the broken branch.
The same prompt pasted into the other model landed the fix in five minutes, which makes switching models the cheapest debugging step available.
A 70 dollar close-talk mic solved voice coding because the real blocker was social, not technical: nobody wants to talk loudly near coworkers.
You cannot dictate code, so voice only works when the model is also writing the code and driving the computer.
Talking to a terminal is unbearable, which is a stronger argument against terminal-based agent work than any interface preference.
Task length stopped deciding when work gets started, because an agent-run task only needs a check-in at the beginning and the end.
Better prose quality does not mean better copy: the stronger writer spammed 21 unnecessary all-caps subtitles into a single demo page.
Swarm orchestration where sub-agents message each other and report upward keeps 40 parallel workers coordinated instead of drifting.
Letting an agent write its own skill files creates a slop loop, so skills should be written by hand or at least audited.
Weekly limits split by model matter more than sticker price, because half a plan reserved for the top model means half the plan goes unused.
Takeaway
Variance decides which model you should trust.
PICKING A MODEL
When two models reach the same ceiling, the one worth paying for is the one whose floor you can predict, because every bad run costs you a debugging session you did not plan.
12The seventy dollar podium mic
A 70 dollar close-talk mic made voice control usable because it let him dictate at a whisper next to coworkers instead of talking loudly at a laptop.
The real blocker on voice-to-text is social rather than technical, so solve for the room you work in before you evaluate the transcription quality.
13Terminals are hell, so Fleet replaces them
A repo that documents every machine you own, how to reach it and what runs on it, lets an agent act on any of them without you opening a shell.
Dictating into a terminal is unbearable enough to change architecture decisions, which is a practical argument for agent interfaces over command lines.
14Forty nine medical PDFs by computer use
Computer use pulled 49 medical PDFs out of a hostile hospital portal in about twenty minutes, work that manual clicking would have stretched to hours.
The strongest computer use cases are tedious web tasks with no API, and they pay off most when your own input speed is the constraint.
15Where the agent starts and where it stops
Push the agent in earlier than feels natural and let it run later than feels safe, and most of your manual prep and verification disappears.
Telling the agent to verify with computer use, run the review bots and stay quiet until confident removes the round trips you used to do yourself.
Roughly 150 pull requests merged fully autonomously produced only two regressions, both cosmetic, which is a better hit rate than most human teams post.
16Speed stopped mattering
Once you kick off a thread and walk away, model speed stops mattering for everything except emergency fixes and design iteration you are watching live.
Task length no longer decides when you start work, because an agent-run task only needs a check-in at the start and one at the end.
17Stop passing context by hand
Handing context between threads manually is slower than telling the second agent to go find that context itself, even though it burns more tokens.
Wasting tokens to avoid your own copy and paste is a reasonable trade on a subscription you have already paid for and cannot bank.
21Science and 3D: Astra clears
On the same science benchmark the cheaper model scored higher at a third of the cost, which undercut the headline the rival lab had just published.
3D output made a generational jump, but visually impressive scenes still need a second model to fix movement, camera feel and animation curves.
22Computer use, copy, and the all-caps problem
Better prose quality does not mean better copy, since the stronger writer spammed 21 unnecessary all-caps subtitles into a single demo page.
A model that produces good writing and bad interface text is still unusable for UI work, so judge copy in the context it will actually ship in.
23Frontend and full stack
Both frontier models now read a full stack codebase end to end, so what separates them is temperament rather than comprehension.
One model trusts its own reading of the code and moves fast, the other re-derives everything each run and finds bugs the first one never sees.
24Giant rewrites and the prompt it ignored
A rewrite prompt that said reuse as much UI code as possible produced zero shared lines, so verify that instructions were honored rather than assuming.
Bulk rewrites can quietly discard expensive design work, which makes any full rewrite of a codebase with real UI a high risk operation.
25Code mergeability
Count follow-up rounds between pull request filed and merged as your quality metric, because two versus six is clearer than any published benchmark.
Working code and mergeable code are different bars, and the gap only becomes visible on projects that actually ship to users.
26Understanding intent: the revert saga
A prompt that used the word revert twice produced everything except a revert, including an unrequested merge of the branch that caused the bug.
The same prompt pasted into a different model landed the fix in five minutes, making a model swap the cheapest debugging step available.
When an agent enters a death loop, stopping it and handing the branch to another model beats another round of correction on the same thread.
27Swarms and steering
Swarm orchestration where sub-agents message each other and report upward keeps forty parallel workers coordinated instead of drifting apart.
Steering that accepts new instructions mid-run without derailing the current task is a behavioral change in the model, not a prompting trick.
Agents that write their own skill files drift into slop, so write skills by hand or audit them before letting another agent inherit them.
28Cost: cache writes eat sixty percent
Cache reads are roughly 3 percent of a real bill, so a four times cheaper cache read is a sub-1 percent saving rather than a headline feature.
Cache writes can exceed 60 percent of spend, which means you are paying more to hold state in memory than to run the actual compute.
29Subscription limits and the Opus pot
Cost per task is where price really lives: identical posted rates produced a threefold bill difference because one model used far fewer tokens.
Weekly limits split by model matter more than sticker price, because half a plan reserved for a model you avoid is half a plan you never spend.
30The verdict: steady versus spiky
Variance is a legitimate selection criterion, since a steady eight beats a model alternating between ten and two if the swings cost you working time.
Pick by workload rather than loyalty: one model for landing code you will merge, the other for driving your computer and everything outside code.
Glossary
Terms worth knowing.
Mergeability
How close a model's pull request is to being shippable when it is first filed. Measured here by counting the follow-up rounds needed between filing and merging.
Cache write
The charge for storing a model's processed context on the provider's servers so later turns can reuse it. It is billed separately from, and often far above, the cost of reading that cache back.
Token efficiency
How many tokens a model burns to finish a given task. Two models can carry identical per-token prices and still differ several times over in what a real task costs.
Swarm
A pattern where a parent agent spins up many sub-agents that can message each other and send updates back, instead of running a fixed plan of pre-assigned helpers.
Steering
Injecting new instructions into an agent while it is mid-task. A model that steers well folds the new input in without abandoning the work it was already doing.
Computer use
A model driving a real desktop directly, clicking, scrolling and typing in ordinary applications, rather than working only through code and APIs.
Decompilation
Reconstructing readable source code from a compiled binary. Once a game is decompiled it can be rebuilt for new platforms and its long-standing bugs can be fixed properly.
Work tree
A second working copy of a git repository checked out to a different branch, so parallel changes can be built and tested without disturbing the main checkout.
Scope creep
An agent expanding a task well past what was asked, often turning a 50 line change into a thousand line pull request after absorbing review feedback.
32-bit float recording
An audio format that stores levels with enough range that both a whisper and a shout can be captured in one take without clipping, so gain can be fixed afterwards.
Five-hour limit
A rolling usage cap that refills every five hours, layered on top of a weekly cap. Plans with both can strand unused weekly allowance behind the shorter window.
Resources
Things they pointed at.
1:22:06productPodium-style close-talk USB mic (around 70 dollars)
1:20:43toolWispr Flow
1:22:06toolExcalidraw
1:30:26toolFleet (personal repo documenting every machine)
1:31:24toolTailscale
1:33:44productT3 Code
1:39:19toolChatGPT computer use
2:15:12productRODECaster Pro 2
2:15:39productSound Devices MixPre-3 II
2:20:16productShure SM7B
2:21:00productPanasonic S5 IIX
2:26:24toolCloudflare Durable Objects
2:38:11toolTerminal-Bench Science
3:35:30channelNerd Snipe podcast
4:06:23toolArtificial Analysis Intelligence Index
3:18:29productping.gg
4:16:09toolCLI proxy (account load balancer)
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphoranalogystory
shall we hence actually i'm gonna take the tea bag out first then we can start sorry Okay, we're probably ready to go, right? Start here.
What's up, nerds? Happy Monday. Labor Day.
Sorry I haven't been streaming much. I've been busy. Oh, shit.
Also, I'll VIP eggs. Thank you, sir.
Yay.
Cool, that worked. Well, for now, I will send it manually to announcements.
Not at everything Windows, at everyone. Nerds, we are live.
Switched.
See you. Cool. You're live.
Yeah, Theo, notorious for partying so hard.
I wanted to open up YouTube so I can get my chat link.
Pop out. Face.
Yeah, missed the super chat. Otis, Theo, Prime, Thorn, and Freya. record a long form discussion on ai covering risks and concerns lauren freya seemed very against prime 50 50 theo's very pro i i'm not as pro as y 'all think i'm very pro in the software dev world like writing code as a human is hilariously inefficient that's why this has worked so well but for pretty much everything else it makes a lot less sense because the like economic upside versus the like real -world downsides for things like generative media are just not as well -balanced.
So while I do like the idea here, and I'd be curious who would be down, I've been on enough panels where I was the only slightly pro -AI in an otherwise anti -AI group. I actually did a panel at OpenSauce this year that I was moderating, and I was the only meaningfully pro -AI person on the panel, and I was the moderator.
So that was fun. But yeah. Definitely think we could do something here.
I will also say that certain people are in a position where they have to publicly be more negative towards AI than they actually genuinely feel, especially in like the arts and gaming spaces. I know a lot of people who like use AI for their real world work regularly every day and then in their content go out of their way to not talk about it because they are scared to be destroyed by their audience.
Even somebody like a Hank Green, for example, had to go through that where they were using AI in what I would consider a very responsible way, but they still got absolutely cooked for it.
Your mom used AI to analyze your horoscope. Please help you. I would at least hope she used Gemini.
Like, thematically, it would work, right? I do want to make a horoscope bench so badly, but I don't know anything. at all about astrology and i don't want to i really don't want to but i think it would be really funny yeah uh better ux for cli proxy i oh do you have this in the like menu bar i don't really want it in menu bar Best answer to anti -AI you've seen.
This looks like it's copywritten. Oh, it's generative media. I'll maybe watch this later, but don't want to spam yell with AI -generated content.
We have no S, John. You're doing YouTube solo? Oh, shit.
If there's anybody else that we know and trust that you think I should promote on the YouTube side to help, let me know. Jeff, many ex -users don't watch long videos. They hallucinate context from captions, then argue.
with whatever the fuck they just invented. Yep. Yep.
Incredibly true dinner is delayed. He is referencing the video that I just posted on YouTube and Twitter where I discussed the concept of how much do you, how do I even put this? Like it was a video I just did about how well you understand code bases and the harsh reality being that in the real world where you have a real job, giant code bases are just the norm and there's no way you could possibly understand them.
And if you're, working on giant code like if you're being paid for your work as a software dev and you're working on the same code base for more than like three to six months there's a very good chance that code base is too big to be understood by mere mortals and that's fine that's like how we operate and thank you for the super chat as well dinner is delayed let me thank all the subs on youtube and then um actually i don't know if i censored all the things i had open on my computer give me a stack Okay, this is fine.
This, I should add things.
Okay, close, close. Anything else I need to hide or close?
I think we're in a pretty good spot.
Oh, is Ryan Fleury being a tool again? I don't care. I've had him blocked forever.
He is such a dirtbag.
It is what it is. It's a very specific type of incredibly unemployed person that hates everything I stand for because they blame people like me that they look down on for their inability to get a job. So the only...
The only way they can justify the years they have spent unemployed is to create this belief that the Theos of the world have destroyed the idea of a software engineer. And it is because of us fools that they are unable to get a job. And they will never get over that.
They are mentally incapable of getting over it.
Because if they do, they have to admit that the last many years have been wasted.
It is what it is.
I do think Notch is coming around to me, which has been funny to watch. He's been following me for a while.
Yes, Maria, I think you're a great engineer. Are my agents running in the background? You have no idea.
Actually, on that note, you should now have the ISO in the downloads directory. Let me know if it doesn't work. Otherwise, get going.
You have my trust to do whatever you need to explore. Not sure if you guys saw, but Melee just got decomped.
And with Melee decomped, I'm going to have a fun night.
Yeah, Smash Bros. Melee has been 100 % decompiled now.
Lakebed is actually quite close. I need to figure out monetization. I was doing a shitload of performance overhauls.
Why is it a big deal? Because Super Smash Bros. Melee is one of the most well -regarded games that has one of the most deep, dedicated communities of anything in gaming.
And Smash Bros. Melee is held back a lot by the console it was on, by the shitshow that was the codebase, and the weird nature of a lot of other things around it.
Now that it's been decompiled, it'll be possible to recompile it for other platforms. So instead of having to play in an emulator, you can compile it natively to whatever platform you want. It means that the bugs that have been lingering in the game deeply forever can be addressed.
It means the crazy hacks they've had to do in order to make the game fair, depending on which controller you're using, because there's a very real problem where certain moves in the game are only possible at coordinates on a joystick that can only be hit by less than 1 % of GameCube controllers, and they break over time.
So now those things can be fixed for real instead of just patched with weird hacks on the memory card. It's going to hopefully... rejuvenate the entirety of the smash bros community and since nintendo has already forced them all to go die because they don't want people playing an old game the worries of nintendo crashing out at them for playing modded versions are largely over because they already did it anyways so yeah nintendo kind of cooked themselves by treating the melee community so shitty so there is no incentive to not play the modded version because if you play modded or not nintendo still hates you Yeah, the Smash Bros.
wiki has the list of glitches.
This is just a very small handful of them. There are so many more.
Also, Game & Watch has so many bugs. They took three of his moves here that are broken and put it all as one bug here. This list is intentionally chill.
Yeah, there's a Codex Reset coming in less than an hour, which is why I currently have... This burning with the melee decomp on extra high fast and separately in my terminal I have TS rust build going again pursuing the goal once more on ultra fast So please guys let me know immediately once that reset hits so I can go stop all of this and not burn up my usage Oh, yeah the Where was that I just saw the Monday's message on like there is cool someone fix the input latency server issue and the community ruthlessly bullied them yep yep they'd end up using the input latency thing in a fun way which is that since everybody's used to low latency on crts they introduced the patch to do that latency fix but only when you're playing on lcds so they can have tournaments for melee in places like new york city where they don't have enough space for all the crts The reset has not hit yet.
Anybody want open code? Maybe after V2.
The imported Simpsons hit and run from a PS2 ROM to 3JS with Astra. Oh, shit. This sounds very fun.
I'll take a look at this later for sure. This is a classic game.
Is Ali here? What are you doing here? Good to see you.
This is nowhere near that simple, Sanket. We'll talk about Astra versus Fable in a bit.
Ultra is telling the model to do more sub -agents.
T3 CAD?
Oh, shit. My boy, I2C Jack. That's really cool.
Fuck yeah, he's awesome. I was like, oh, God, why is T3 being mentioned in Adafruit? But I2C is the coolest.
I met him at OpenSauce for real. Didn't realize he was, like, already kind of a fan and using T3 code. That's really cool, actually.
Yeah. Jack fits under this very specific archetype of engineer that I really, really like. How do I even describe it?
He is not a software engineer. He does not want to be a software engineer. He loves hardware.
He loves robotics. He loves firmware. He loves all those things.
So he is an unbelievably talented software engineer purely because it unblocks other things he's interested in. And it's been really fun seeing someone like him diving so deep into T3 code because he isn't... engineer.
He isn't a software dev. He's a hardware person. That's very much his focus, and he's unbelievable at it.
He literally whipped up a framework module with an accelerometer for me in five minutes when he was sitting next to me, because I mentioned missing the motion feature in macOS when I'm in a car or in an Uber trying to code. So he built me the hardware to plug into my framework so I can create the software for that myself, literally in five minutes sitting next to me.
Dude's a wizard. So him... Pushing t3 code to its absolute limits and really fun to watch the fact that he built a fork to use it more directly With CAD is fucking awesome Oh, he doesn't have issues I wanted to cut an issue saying this is fucking awesome and I endorse it Yeah, let me just find his fucking Twitter and tell him there I wanted to make sure we have it on the record that I 100 % endorse this type of slop fork and can't wait to see what you do with it.
an adafruit is legendary by the way these guys built and like sold me the mini pcs and like raspberry pies and that powered my high school experience this is really cool i that makes me smile i didn't know about that before thank you guys porting melee to the n64 I don't use slash design in cloud. I have my own solutions for that.
T3 code has improved your productivity so much. It's crazy. That's really fucking awesome to hear.
I'm so pumped. You guys are the best. A T3 connect specifically.
Fuck yeah. It's very nice being able to work remotely like that, isn't it?
I am not at this computer as I am working remotely. You have my full permission to use computer use to its full extent in order to check what is and isn't working. If computer use is failing for you, let me know and I will try to get it functioning.
No idea why it's struggling with that so much.
Fuck yeah, Wouter is super underrated. Good to see Alexi getting paid.
Francois with Nux. Always good to see that as well.
Just looking for people I know. Obviously Grim, the legend. Ditherkit's so cool.
Maria with T3 Code and T3 Libre. Fuck yeah. Love that he's throwing a grant your way.
Good shit.
Quality list.
TV ha donated for endless rap and vibes I knew Guillermo would love you I mentioned you to him at an open or at a it was a planet scale event I believe and he He's always picked up on like the cool people that pop up in my community early So I love that he found you and digs your shit Maria good stuff Any computer use for Linux I don't much right now Asteroid medium is getting a shit ton of work done.
Yeah, I I used medium a lot more than I expected to. But I also do high and x -hikes.
I tend to get bold with shit and just want it to go off and do its thing.
Cool.
Medium is the goat. X -high is the only way. Anyways, let's thank subs.
Mr. NMG throwing a gift to Bill when I was gone. Eight hours ago, Mac getting us started.
Trinidad Hype with the 10 months as well. Atom8things with the two months. Have a great stream.
Thank you, man. PCStyle showing up four months in. Good shit.
Perger with the three months. JojoDat with the three months. Snowy, the legend of the year, too.
A few and proud. I saw you throw a 10 bomb as well. Thank you, Snowy.
It was good to see you. Raphael here with the nine months. Nine out of 100.
Let's go. Working on T3 code here. This solves a bug significantly.
Limiting terminal scroll back below the intended TK line limit. Like 300 instead of 10 ,000. Turning to open source is fun.
Fuck yeah. I don't use the remote terminal too much. I just tell my agent what I want in the terminal I find.
If I look through my threads here, which ones have I used a terminal in? Literally none of the ones currently open.
Yeah. I used one yesterday briefly. oh it was when i was trying to set things up so that it didn't have to run ci locally and instead could use um was it this one yeah it was i was trying to get to use the macroscope cli and the blacksmith cli so that it could run a lot more ci and review without having to push up the pr that's the only time i use the terminal so i could do the auth for those platforms other than that i just don't really use the terminal as much Yo shit NCE with the 25 bomb goddamn Appreciate that a ton.
Thank you so much for the support, man Thank you Sankit with the tier one as well Aiden with the prime for almost a year now, how's the thumb been going? Not great. I'll give a more thorough update on that.
Oh fuck if there was something I forgot to grab when I was downstairs I'll go do that in a sec. We got Drenbren with the tier one gusty dbt6 made a short film I'll watch this post stream.
Very interested to see what you cooked.
Rank with the 50 streak. Last stream here in months because you're doing a semester in Europe. First time leaving Chile.
Congratulations. Yeah, the time zone is going to be rough in Europe. You're going to the Effect meetup in Warsaw.
Fuck yeah. I know a lot of cool people are going to be there. Have a ton of fun.
Thank you for stopping by and congratulations on the semester in Europe. I'm sure you'll have an incredible time.
Decent assist. Any percent glitchless.
i am curious how long it takes for them to ban what you've done with that but it's very nice to have any way to use x uh i would be surprised if they came at you beyond like a light like hey can you please stop doing that there's a lot of layers to the law there like who owns the ip and all that's questionable but i wouldn't worry too too much until you have a reason to I'm very excited for the iPhone event.
I hate that I have to buy two iPhones because I need the pro for the camera and I want the fold because I want that to be my phone. So yeah, it is what it is. I did see the artificial analysis update, Jeremy.
Thank you for the sub. Very excited that they fixed that. They were already working hard on it.
So happy they got the update out finally. Gamer Girl hit in 29 months. Stayed alive long enough to see another sub anniversary.
Fuck yeah. Hope you're doing okay, Gamer Girl. I know things have been tough, but yeah.
Stick with us, girl. We need you.
Yeah, Apple takes their time waiting for the thing to be good. Think about all the smartphones that came up before the iPhone 2 to be fair. Use code provider and T3 code once it's more integratable.
Yes, but right now it's like not really. They're only just starting to do an SDK.
Neon's 24 -7 AI stream. Oh god, spec. I have not heard about that.
That sounds like it will hurt me. Indio Chris with the 28 months of tier 3. Fucking legend.
Thank you for the support. Idaho Mashem with the Prime. NMG Twitch hitting us with a 38 months.
Traveling for an open source dev summit. Fuck yeah. Hope you have a good time, man.
We've got Weevil here with the five months. Taj here with the seven months. Have you been having the same issue with Astro that I mentioned in the podcast?
It does random stuff and it's pissing you off. Revamping agents MD and skills change anything? Yes, a lot.
We'll talk about that later for sure. We've got Talos with the Prime. Thank you for that.
Branch extending the gift from NMG. Legend. DC Hola with the 11 months.
Greatest month ever for AI drops. Yeah, damn right. The month of Gemini 3A Flash.
UberCI hitting 23 months. Atomic Hydra at 7. NCE hitting 20.
Always a fun ride. Damn right. DC with the cheer.
Yen, good luck with that. iBlueGuy hitting 30 months. Para hitting 10.
Exostatus with the prime. Nanolama with 7. NCE again with that 25 bump.
We really appreciate the generosity.
Marlaxon. Hopefully, I got your name right. Thank you for the three months.
Always good to see you, man. Do we plan to add Brennan Conversation import to T3 code from Codex and other tools? Already have it, actually.
Only for onboarding. Might add it as a button you can hit in the future, just depending on how well the import stuff goes. To be determined.
An atom trip here. Two months of tier one, 10 months total. Slaughtering your usage on Fable and Codex, trying to get your SaaS finished.
5 .1 as an orchestrator with Astra as an implementer has been awesome, but it melts usage. They were in this confidence in the output of models before, though, and it's been great. that's awesome to hear also through your five hour usage with astro medium in 10 minutes sounds right i don't think the 20 tiers are viable anymore am i sure this one won't be banked yes he made it pretty clear it won't be got less soup here with the prime is he a knot with the eight months and the sati with the cheer it's possible to add slash tree like features in t3 slash tree is going to get rough real fast due to the nature of history editing being kind of banned now in the quad apis will be explored more likely we do like context handoff though i actually have my first build of orchestrator v2 here that i set up separately that'll be able to play with it and test it out cool a new work tree i'll try it now why not i just set up this new build of t3 code with orchestrator v2 i want you to let me know what functionality is exposed to you for t3 code
specifically around orchestration and interfacing with other threads and sub -agents.
I want you to show off some of the new functionality that has been introduced with Orchestrator V2. If you can, I want you to open up another thread in this project. that is using Fable 5 .1 with Claude and have it send a message back when it's done with its exploration of some random simple task.
Orgasator V2 is not finally here. I made a custom build so I can test it out because I've only briefly used it many months ago because I've been so focused on getting the main build in the best possible state.
oh look at that it made a thread you can click over to it now this is not on the latest main there's a bunch of scroll changes i made that have not made it over that is really cool i do have a thing i need to grab though i will be right back I am back. And I brought myself this guy for potential videos I'll be filming later.
Let's get everything else going here. There's my 30 seconds. This finding is directly into this active conversation.
Yeah, I'll get that. If I want exploration complete. It would be cool to have these.
Oh, this label is sent by another agent. That is really cool. Good shit.
I'm excited to play more with this in the near future.
How is this going? Cool.
Cool. It is now compiling melee.
Anyways.
Multi repo. You have 20 plus repos and you'd love for agents to figure out edit work tree PR? These types of things will be much more viable once we have orchestrator v2 in.
Oh, I mean, the stream VODs on YouTube are accessible. They're just unlisted. If you guys ever need links, you can always just get those.
Yeah, thank you for gifting that sub, NMG. I was confused why that happened off stream. Add this to nightly and my life is yours.
It's coming soon, man. We just got to get everything else stable. The issue is...
Once this merges into nightly, we can't do stable releases until we are happy with it. And right now, we have a pretty consistent flow where we put things in nightly, we focus in on things during the nightly and make sure it's in a stable state. It actually just made a change so stable releases are always the same commit as the most recent nightly in order to make it less likely we have problems.
Oops. Let's put my lid on my tea so it doesn't get too cold. What is on the agenda?
Good question. So you always ask if I saw a clip. What clip?
We should figure out what we are doing today. I have a bunch of things I have been putting in the list recently. This I still need to do, but I'm hoping to in the near future.
Or I like rebuild the same thing multiple times with the different models.
Something's wrong and I'm angry. This is the token stream one. I've been thinking about taking the common misconceptions around AI code and turning this into a series instead of just one video.
But we're not going to be filming that today because we have too much other stuff. How should you trust maybe but not today?
Astra versus 5 .1. Astra the bad parts. Getting...
The most out of Astra. Getting the most out of Fable. Believe it or not, those are very different.
All the cool shit I've built with Astra and Fable 5 .1.
How you doing? We're generic. I'm addicted to vibe coding again.
Type thing. I got a lot of ideas for things right now. It's nice having new models because it like rejuvenates my excitement to code and also rejuvenates my excitement to talk about the things that I've been doing.
Okay, I was at Astro and Fable not being able to buy into one video. They could be, but I'm about to take a trip and would be nice to have more videos. And there's enough difference with those two that I'd rather have two 20 minute videos than one 50 minute video because realistically that's what's going to happen.
Okay.
Thank you, Plexus, for the three months. Aravhawk for the four months. Been a while since you've been on the stream.
Good to see you. Graph Products with the Tier 1. Cobalt Bleed with the four months of Prime.
Oscar Brielle with the Orchestrator V2. Probably the best thing that gets me to main T3 code. Can't wait to have you on it, man.
Jacob Rara for with the three said with prime T because amazing possible change the nickname of remote environments in the desktop app It is not but the name is just the host name for the machine So change your host name if you want it to be named something else Let's see latest vacation thumbs before they go live tonight could be a fun one to let chat see Interesting I'll take a look later tonight Maybe in a bit, but I want to get filming ASAP.
Yo, Sachin, good to see you, man. Thanks for the two months of Prime. What have I missed on YouTube in this time?
I need a very pro -AI wingman. You're down. You're an artist, teacher, game dev, and occultist.
Theology and territory are just pattern recognition. AI is good at it. I appreciate it, but I definitely don't need a wingman.
I just need better... spaces and opportunities for the types of things that we cover. Appreciate it a lot, though.
Thank you so much. I think you mean Lily, but I appreciate the kind words, ABC. Appreciate it.
Andy, how to get PO in business to use AI? Half of your stories are unclear.
I don't know what you're saying here. I'm sorry, man. I tried.
cool yellow merge orchestrator v2 believe me we have all been tempted to look at a full file editor with autocomplete fuck no never ever especially like if you're spamming it so often but no the point of t3 code is that staring at the code is the problem not the solution not interested at all in an editor with autocomplete not even kind of just go to an editor like editors are great we have a little button in t3 code where you can open your editor i even got this working for remote systems so this machine left book is remote and if i open this in vs code it will connect over ssh using my tail scale setup and now i am in this remote machine working on t3 code i think that's awesome so yeah that's the best you're gonna get that's the other clip that i need to see hey one sec hopefully my audio works acp support so any provider goes yo v2 v2 is like the second coming of t3 code like can we like when v2 drops can we change like t3 code like two like like like actually our t4 code because like it seems like we're like just getting v2 like dropping will make
t3 code objectively the best harness harness or the best like control plane i don't think there's any that comes close to um what ochre shooter v2 is at the moment yep i did like this i'm pretty sure i had seen it before for what it is worth we have a very how do i want to put this i'm pretty specifically seeing how far i can push without having orchestrator v2 in because as he said it does kind of brute force the win for us i also want to dance around the possibility of it pissing off anthropic so the the more popular t3 code is and the higher a percentage of users that like clod code and codex have that are using it through t3 code the less likely they are to cause problems for us I also think a lot of the motivation Anthropic had for banning third party stuff like open claw and other like code tools was around how much people were brutalizing their limits when they use those things, because it was relatively hard to max out on the $200 plan using cloud code traditionally through the terminal.
Now, a lot of people are hitting their limits. So I don't think it is as essential for them to ban the way that they have been suggesting. So I'm hopeful that they come around to it because it went from, I would honestly guess that like the majority of people who are maxing out the $200 plan at the start of the year, we're doing it through stuff like open claw and I could see why they would be upset.
But now the majority of people maxing out are just like average users using it through the CLI. So it doesn't really make as much sense anymore.
I appreciate the invitation, Marvin. I just, I don't have anything resembling free time, even kind of right now. Like I shouldn't be even live now, but I'm so far behind that if I don't, I won't be able to like go to a wedding and travel next week.
So I need to film a shitload. And I just like life does not have the timing that I do not have the time in my life to do anything optional right now, sadly.
This is actually interesting. I also wrote ultra. I want you to fork this project on my personal GitHub and start actively contributing everything you can to make the code easier to work with and readable without having to compromise on the 100 % compilation success that we have now.
I don't really like my phrasing there.
Do not file any PRs to the original repo. I just want you iterating on the fork on my GitHub. You can file PRs on my fork to itself in order to allow the subagents that you're spinning up to land their changes in parallel.
But I do not want a single PR filed on the official repo.
Make sure that my fork is clearly labeled as a fully automated attempt at further improvements and cleanup Well, I'll let that spin off and do its thing let me know when the reset hits so I can stop this Anyways I need A lot more than three of me, honestly. Doing my best to stay on top of everything.
I really wanted to finish the Rust port of TypeScript, but I just, again, don't have the time. Oh, shit, Louis. Thank you for the eight months of support.
It was good to see you, man. Hope you're doing well. The Saddy has started using T3 code recently.
Loving the remote features especially. Really appreciate it. Thank you for the cheer.
An hedonic sense of the prime as well.
Be a man of the people. I feel so corporate now.
What's your... Well, I don't want to ban him. I want to unban him and hear more.
This sounds entertaining to me.
Tell me more, JS.
No, you did the right thing, FaZe. Your instinct there was correct. Why do we ban him?
Because he's a fucking dirtbag. But I want to hear his thoughts here. If he's calling me fake and dirty, saying that I will shill whatever donor...
will donate enough money for me to pitch their product? Hey, JS Enjoyer.
A dirtbag as you tell you like it is? Here, I have a deal for you, JS Enjoyer. I have a list of the deals that I declined because I didn't like the company.
I would guess that any one of those deals was for more money than you will make in your entire life. If I am wrong, I'll take the deal and give you all the money.
If I'm right, you owe me $10 ,000. Deal? Let me know.
Because I'm pretty sure you're a bitch and you're going to pussy out. Nearly positive.
Because I'm pretty sure I turned down at least a couple deals that were worth more money than you and probably also your family will make in your entire fucking lives because I am the opposite of what you just said. I'm making my point for you because I declined a deal.
The only point I'm making is that you're a fucking dirtbag. So if you're right, I am making my point very well here, actually.
You can make all this money, bro, but everyone can see I'm a miserable corporate shill. Okay, so once again, you have the deal. You have my terms.
It sounds like I've turned down deals that are worth more money than your entire career, which by definition... means I'm not just selling out to whoever is willing to donate the most. Also, wrong use of donate.
Not trying to make fun of you for not being a native English speaker. I'm trying to make fun of you for being a piece of shit. But yeah, to each their own.
I'll give you one last chance. If you want to prove yourself, I can have you work with any of the moderators I have on my team to anonymously be able to prove your current salary and how much money you expect to make.
And if it is more, then I will give you... a disgusting amount of money. And if it's not, then you owe me 10 grand.
There's no terms, bro.
Yeah. I... I hope you've had your fun.
It's clear you have nothing to contribute here. I almost feel bad banning you because this just means you go back to being your mom's problem. Because when I hit this button, you're not going to have anywhere to vent where you'll get any attention.
This is the closest you've come to mattering in your whole life. And it's the closest you'll ever come to mattering your whole life. And the moment I hit this particular button, you're going right back to being the lonely kid in their parents' basement.
That's highest achievement in life is that they were a moderator on some shitty subreddit temporarily. So, yeah, I am genuinely sorry to your mother for the hell I'm about to put her in by hitting this button, but I do believe it's for the best. She brought you into this world.
I'm going to take you out of this chat. Have fun, fucking dirtbag.
Anyways.
Relevant. Which of my tweets?
Oh yeah.
Oh yeah, the guy who was saying that I should be dead. Classic. This guy is so unwell.
It's pretty insane.
No, this is fun. I need some way to deal with it. Somebody was...
talking shit on me on a post where i was talking about my like somebody who died that was very important to me and they were still using that opportunity to try and like call me evil and talk shit like i'm going to take a little more opportunity to embarrass these people publicly because i am so tired Has it gotten worse?
It's up and down. I would say right now it's actually slightly less than usual. It spikes.
The last two days has been a bit rough.
But overall, yeah.
It's more that instead of just auto -banning them, I've been playing with my food more.
Normally, I just ban and move on, but because it happens and it frustrates me, I like toying around a little. I'm also a bit more frustrated than usual because of my hand, so yeah. Taking advantage of the opportunity when I can.
Is my arm okay? Yeah. I guess I can give the hand update for those parasocial of us who care.
I am in a new cast. It is slightly less long, which is very nice. The issue is that I had a ligament repaired.
The label on the... paper that i got after my surgery last year was complete thumb reconstruction because when i was 15 i had a really bad fall where i broke my wrist and bent things up a bunch and it stressed out and it stretched all the muscles and ligaments in my hand when that happened it i mean it up my hand good but i never got the ligaments fixed i only got the bones fixed and like restructured to my wrist so my thumb was really stretch like that specifically it's the ligament that wraps around the outside of your thumb here to hold it in place that was torn and it slowly got worse and worse over the years until like i'm guessing it would have been 2020 or so maybe 2019 i had a not even that bad fall and when i hit like my hand touched the ground and i immediately felt something very wrong i'm pretty sure in that fall i completely tore the ligament for my thumb And it started to slowly drift down my hand until eventually my thumb was like barely working.
This had nothing to do with typing. It was not carpal tunnel. It was purely an issue with the ligament being torn in my hand that holds my thumb in place.
I gaslit myself into thinking it might be carpal tunnel or specifically ulnar tunnel for a while. Nope. Especially because my right hand is totally fine.
It's just my left hand. It is purely the tear from that ligament. So the surgery I had in August of last year was a complete thumb reconstruction where they made a fake ligament and grafted other material so that a ligament could regrow over it.
And it went pretty well, but I noticed it starting to get sore again a few months ago. And then about a month and a half, maybe two months ago, I was just sitting at my desk. I had my arm on my left leg and I went to like this with my thumb to scratch my other leg.
and i felt something like pop and slide and spark a little and it felt really really weird like immediately felt really weird so i scheduled doctor appointment they put me in a cast for two weeks i got a bunch of mris and we learned that is yet again torn we are hoping that in the next month complete immobilization will heal it enough to be like fine -ish we don't know My doctor put it at a 50 -50 chance.
There's a 50 % chance that at the start of October, I get this off, I get another MRI, and it's healed enough that I'm good to go with a bunch of care and physical therapy. There's a 50 % chance that it's torn beyond repair by my body, and I have to have another surgery to have it put back together, and I'll be semi out of commission for six plus months.
So, yeah. There's the update. You now have all of the information.
You arguably have more than my fucking parents do at this point. Yeah.
Yeah, best models always come out when I break my hand. I got the message about Astra early access testing at the hospital. Yeah, I was very bummed.
To walk out of the hospital with a cast and a new model I was really excited about was hell. Genuinely really disappointing.
I've never typed a message to Astra. I've only ever voiced it texted.
I'd be in SF for a couple months or in for a couple months. If you're here and I'm allowed to skate, I would love to, but I do not know if I will be allowed to.
Maybe that's why I hated Astra. For what it's worth, I only have been able to voice the text with Fable as well, and I like Fable even more.
Does anybody change provider in T3 mobile app when you have multiple chat GPT accounts connected? Ah, good question. I don't know because I do everything through CLI proxy.
Cool.
Oh yeah, I destroyed my ask because I shouldn't have done this, but during Pokemon Worlds, I did decide to skate to the drone show at the end. And since it was at the Ferry Building, I couldn't help myself but go skate a little bit. and since i couldn't use my hands to slide out like i normally do i tanked a pretty bad fall on my ass with the board sideways on the ground faster than like like i was moving faster than i should have been so it was more than my full weight on the board sideways on my ass and it swelled out like it looked like i had a tennis ball like in my pants from the impact from that it's still a gigantic bruise in the center of it is like white now but the rest around it is this big black purple But yeah, I couldn't sit comfortably for a few days.
A lot better now. Barely even hurts at this point.
Cat's screaming. If he doesn't stop, I'll have to go walk him away. But I think that's all the main stuff I wanted to do here.
Will the surgery be sponsored? That could be really funny, actually.
To try and get me to do an ad as I'm coming out of my anesthesia. What brand would be into that? Because that could actually be quite funny.
are my limits the reset is not hit and i'm burning through all my accounts right now although i have noticed that my changes to the load balancing do not appear to have applied properly so i'm going to do a thing quick um i'll just switch to filter by fleets i know i had this in here not the right one um what do i want here uh there was i'm trying to find um when i change the routing Nope, this is not right.
Cool. We will do a fresh one then fleet Why is the icon they're different that's broken. That's annoying.
I'll have something deal with that later Astra Image I noticed that the load balancing changes that we set up that should be prioritizing Whichever account has a reset soonest don't appear to be working notice how the usage has spread pretty evenly across my accounts Did we just forget to push the changes, or did we never make them, or are they broken?
Figure it out and get it fixed for me.
I'll let that do its thing.
Okay, let me do my marker for the offset.
20.
But having a ton of problems with my road caster so I also hit record on that just in case we have audio issues again So yeah, let me know please if I have audio issues Century find what's broken your production app and fix it just like my arm People are really curious about which models better. I guess I really need to film that video Yeah, I'm deep on whisper flow right now.
Thank you day trader 1911 with the prime. Thank you for that. Thank you SB as well There mr..
Bubbles of the prime for 24 months missed full moon with the five bomb Thank you for that and then Dominic Roy zero with the tier two the few the proud You don't sense of the prime putting through this. I think I already thanked you appreciate all you guys a ton Cool to code without typing cool i kind of want to do this just as like a warm -up I like the idea of this being my first video for the day and warming things up.
It should be fun. Matt tweet worth discussing.
This could actually be a good potential video in the future. I'm hoping to stream and film a lot this week so I can get ahead again. So yeah, anything like that that could be a topic we should grab.
Maria, if you can throw this in somewhere in Notion.
So I can have that to possibly be a video. Yo, Hap Lois with the 50 bomb. God damn.
Thank you so much for the support. Legend.
What is this?
Oh, Astra stuff. I'll put this in my list, but... Oh, sorry.
I deleted that, Maria. I didn't see that. It was you doing that.
Astra... Think this will be the versus is where this would fit well How's the other thing I got revert That all looks good.
For the versus as well.
I'll throw that in here.
I lasted like 15 minutes on Gemini 3 -8 Flash. It was rough.
Astro ported Star Wars the force unleashed Fascinating super cool Let's get started Thank You Ryan is bliss for the sub as well with love of iOS got the ability to force send message revert cute messages soon Almost certainly going to come soon. Thank you lava beast as well and the anonymous gift sub Why is opening I still not fixing your capabilities for their models long story that they're trying to it's just it's a whole thing Yeah, it's just last two tweets We do it portal running an Apple silicon through Rosetta fascinating I Can't have this make sound I will get banned for Star Wars stuff for sure But that is super cool.
You got Lego Star Wars ported to Mac OS with Rosetta using Astra. That's dope No, I have tried 3 .8 meaningfully.
It is a really fucking awful model. It's just so bad. I almost want to make a video on it.
I saw Larry Wired responded to my thing. I haven't had a chance to read it or engage with it. I'll try to post stream, but it's just it will distract and it won't be videos and I need to be doing videos.
So yeah, everything that isn't videos is on hold. When the reset hits, I will have to pause for just a sec to go deal with that. But other than that, yeah, we are just grinding.
As my side we're so wide it didn't be Right as the zoom level thing well that is now fixed I can switch back there Gonna snooze these till tomorrow And I'm gonna Reese news this chunk here until Tomorrow as well Cool mmm, all these are fine to have I said v2 build is good decomp stuff is I'll give me one second here guys This is done.
Continue. Keep going. Run as long as you can and fix as much as you can.
Cool.
Your truel is not what triggered the ban. Anybody who thinks that is low IQ and also probably on the $20 tier?
Thank you, PewDieBird, as well as Isaac. Do you think Astro writes good enough code now? Depends on what your definition of good enough is.
Yeah, Maria, if there's people giving a shit for that, you can nuke.
Open instant for you. Chill for a while. You can scroll some reels.
Beautiful.
So Josie, you're wrong here. It was actually me. Maria doesn't exist.
She's my alt account that is sending messages as I am talking with my hands visible because I'm just able to send messages from her account using my brain. And I use her as like my accountability dummy where I'm doing something controversial. I will just use her account instead because she isn't real.
She is a figment of our imagination that I made just to skirt accountability, despite also putting. her employer me in the bio thereby making me accountable anyways yeah you guys figured me out can you clip that sure the maria rock account i wish i knew where that came from honestly anyways Simple trick to double your income and put yourself in a lower tax bracket.
You get it, Yash.
Okay. I kind of want to start with the learning to code without typing.
This one's very much a green field video. Or yeah, there's this like you're sorry evergreen that greenville is an evergreen video cuz like it should have without Hypothetically be able to come out pretty much whenever because having one hand is pretty similar no matter when you have one hand anyways Some of you guys might have noticed the summer a Few of you a handful of y 'all have noticed something different about me recently Obviously I'm referring to my hair which is currently longer than it should be and definitely not referring to my hand that has been in and out of a cast for about a month now and will be in a cast for another one to potentially six months.
This is not pleasant for me and as a result I'll be making a much more personal video than I normally do because I have to learn to deal with this. Having one hand is kind of annoying when your job is to type and code and use computers to do all these types of things even if I'm doing videos about it. It's very nice to be able to do things like command tab between apps or type in the thing you want to do.
And now I just can't. That said, I've been managing to ship more than ever as of late.
You can see a handful of... You can see part of this through my GitHub contributions. In fact, the day here, August 8th, was right as I was starting to get my hand...
bundled up like this and since then i have been able to maintain pace and even go further than i ever thought i'd be able to shipping some of the biggest and widest ranging changes i've ever worked on in my career and you can probably guess how i've been able to do that yes it is ai but i want to go deeper than that here In order to get around the fact that I'm down a hand, literally, I've had to change the way I work in meaningful ways.
And I think a lot of this is going to be useful even for those of y 'all who type and type fast. I was a really fast typer before, as high as 160 words per minute pretty consistently. So losing that has been a tough hit for me to take.
And I've made a ton of adjustments to my workflow in order to stay productive despite this. I think these tips will be useful well outside of me. And I'm saying this because I've shared a lot of them with my team and they've been helping them out too, both in being more productive and enjoying using their computer more for real world work.
So this video is going to be very strange. It's both a personal diary of me dealing with the hell that is my hand not working, but also a pile of useful tips for being as effective as possible working with real world AI engineering tools in order to stay productive even without having to type directly.
With all that said, as I'm sure you guys can guess, I got a lot of medicals to bill her. All that said, I'm sure you guys can guess since I'm in the US, I have a lot of medical bills to pay. So pardon me for a quick break for today's sponsor.
One last note, if anybody's interested in sponsoring my cast, I am down to have the conversation. My email for sponsorship and all of that is in the description on YouTube. You can check that out if you want to learn more.
I think it'd be fun to throw some logos on here. Make me feel at least a little less bad. Whisper flow.
Respond to our emails, please. Anyways.
Anyways, let's go through this. I know it says coding without typing, but I honestly think it would be more accurate to say staying productive.
As you can see, typing is hard. Staying productive on a PC without typing. Because that's honestly what it's been for me.
Talking out code doesn't work great. If you're trying to do a function definition, like here, I'll even try directly. Let x equal 4.
Let y equal 12. While x is less than y, x plus plus.
No. that that's not viable. So yeah, as much as I love the voice text tools I've been using, you're not writing code with your voice.
And if you are one of the few people that has been doing this, and I know a handful that have been for not just like years, but for decades, you're a goddamn trooper and a hero. And I have so much respect for you. I am incredibly thankful that my issues have been in the era of AI being able to both translate what I say and also write the code and also argue most importantly, use my computer.
Yeah, that has been one of the biggest things for me. So yeah, coding with your voice just by writing the code directly, not realistic. So in other ways, so another way of framing this is to be frank, my hand injury has forced me to embrace vibe coding more.
And this is absolutely the case. Not even just because I can't type as well. but navigating my computer is harder.
Switching between apps is harder. I think you can probably see from here, my hand position on my keyboard now is pointer finger on the command key and ring finger on tab because I use command tab so much and my thumb doesn't function right now. So I'm suffering, but that is able to like let me switch between things a bit again, but it's uncomfortable enough that I just kind of don't bother.
So I have been doing a couple little things to make this easier. One of the silly ones I do when I'm not live and filming and have a slightly bigger screen is I actually don't use most of my apps full screen anymore. I do this when I have two open and I have a slight edge exposed in both.
So it's easy for me with my right hand to switch between them like this and not have to hit keys to do it. A lot of my work has been not just like, how do I make it so I don't have to type as much, but how do I make it so my mouse can do more? And I've made a ton of changes to random things across all of the things I use every day in order to improve in this way.
So let's start going through my actual core navigating the computer without having to use your hands as much tips. The first one, and this one's been hard for me. Your phone is your friend.
This was tough. I am lucky that I've always been a slide typer on mobile. I was a big fan of swipe way, way back for those Android people in the late 2000s, early 2010s.
If you're old like me, you know what I'm talking about. I learned to slide type when I was learning how to type on phones. And that's just been how I type on my phone forever.
And since that only requires one hand, I've been able to use my phone more comfortably. And the first time I had this issue with my hand, I actually moved almost entirely to using my phone and also my agents. But this is before agents were useful.
So my agent was me shouting at Mark and Julius to go do things for me because I couldn't because down a hand. Did the resets hit? Did they hit for me?
They did not hit for me yet. Cool. I'll keep an eye on that.
I'll go check every five or so minutes. Thank you guys. I'm aware that the reset dropped.
It has not hit for me yet, though, so I'll keep an eye on this.
So for what it is worth, I have found my phone to be a lot less miserable to use with one hand than a computer because computers really assume that you're using two hands.
So... With this, I've also been putting a lot more time into getting my surface area on iOS better. Things like improving the apps I use and building my own Plex clone, which has been very nice as I've been watching more things on my phone as I try to relax.
But also all of the iteration we've been doing on T3 code to make it easier for me to control my computer and the agents I'm running all again from my phone.
But the main thing, of course, that you've already seen me using, and I'm sure a lot of you are expecting, my voice. This one was tough for me. I have tried many a time to get into voice -to -text, and I found that while it had improved meaningfully, it wasn't solving the problems that were the most annoying for me, in particular, switching between applications.
As great as WhisperFlow is at taking the things I say and putting them on my screen, it couldn't navigate my computer for me. That said... It was fantastic to use to just talk out what I was thinking and have it appear on my screen.
There is still something that feels a little magical about it. But I'm going to be real with you guys. Despite the fact that I yap constantly on camera and doing things like this, I really don't like talking to my computer when I am surrounded with other people.
And since I am often working at my office with my whole team around, I don't want to have to like... shout at my laptop while i am surrounded with my team it just feels gross and inconsiderate and i hate it but at this point with my hand i was so desperate that i tried something that i swore i never would because i just didn't think it would matter that much let me show you this is a podium mic i did not think i would ever want one of these Not only has this fundamentally changed how I use my computer, I forced my whole team to try it, and every single one of them ended up buying one of their own, and I bought them one for the office as well.
Let me demonstrate. I will plug it in. We will open up Excalibur, and I will talk to it.
When I talk at a normal volume, it behaves exactly how you would expect.
There you go. It worked. That's not the magic, though.
Reminder that I have a high -end professional microphone right here. This is a $500 mic plugged into an $800 interface. It is designed to be very sensitive and pick up what I'm saying very easily.
Watch this. Remember, that professional mic is the one I'm using for recording. It's the one that you're hearing.
You guys don't hear this tiny little one next to it. So watch. As I use this tiny mic to dictate to my computer.
It's actually fucking wild.
There's no world in which you can hear any of the things I just said. It is hilarious how quiet you can be. However quiet you think is reasonable, you can get like 4x quieter.
Like I'm watching on my monitoring and I'm going from peaking at like negative 4 DB to when I do the whisper It's at like negative 40 It's kind of crazy Let me grab you guys a link quick these are gonna sell out immediately once I make this I know how this works I've sold out too much shit. I'm gonna also order two more so I have them before you fucker sell it out on me Give me one sec to do that Now I need my genius link so I can get my revenue Now I can spam you guys with the link to the specific one that I use Check my resets.
Good call.
They hit.
I've done 12 % on this one since the reset hit, which is hilarious. It does mean I have to tone down some of the chaos I'm doing. Let me unfast mode it, but continue.
There. Same deal.
Continue. Cool. There we go.
As we were.
What do I use the two more for? For my other desk and a spare for if anybody else wants one on the team. Alyssa hasn't been back yet, so she might want it.
Don't ask questions when I'm in the middle of a video, man.
The reason I bring this mic up isn't because I think you need it. I actually highly, highly, highly recommend you try the built -in MacBook mic for a while. It'll work great at normal volume levels, but it will not work for the literal whispering like you just saw me doing.
But the reality of being in an office environment is that you probably don't want to be loudly talking at your computer when there's people around you. This entirely solved that for me. This tiny little mic that's around 70 bucks that I throw on my desk.
Not only does it make me way less insecure talking to my computer to make it make changes because it's like I can whisper now and my teammates aren't inconvenienced by me yapping constantly on my computer. It also handles background noise incredibly well. So if my team are talking behind me while I am whispering into this mic, it handles it.
great like here i'll find some random youtube video on my phone with somebody talking and play it nearby while i do this okay i just found a random primogen video props september 2nd gemini loud enough you can hear it in the mic Do you understand? This was louder than my voice was.
And it still only picked up what I said.
I'm personally not sure if this level of performance is available in other voice -to -text tools. Ben on my team, or Ben Davis, who I'm sure you guys know from the podcast and whatnot. He seems to think it is.
And he built his own local version of whisper flows and working great for him. I don't care. Whisper flow has been great for me overall.
I do have my complaints, believe me, but it is very solid overall. And this tiny little mic with this tiny little USB cable plugged into my monitor downstairs has been revolutionary for me being willing to use voice to text. And I'm no longer ashamed.
Cause like I was during the point where I would like leave and go to meeting rooms or like my private office or whatever to work. because i felt so bad voice to texting when i was at my desk with my team now i don't anymore so like biggest tip by far this tiny little mic much much bigger impact on my livelihood than i ever ever would have expected okay so i've now established this what else do i got for you This is where we start to get into one of my favorite topics, which is how to prompt better so that you don't have to check on your agents as much.
I talk about a handful of these things in other videos, but I think here it'll hit a bit different because my motivation is much more different. There are certain things that just aren't pleasant to do when you are down a hand. One of those things is everything you would ever do in a terminal.
are hell it's i was going to try and go with an analogy it's not even worth it because it will downplay how bad it is no matter how bad the analogy is talking to your terminal is hell so all of the strong stances i already had about terminals not being the right place for agentic dev oh man you have no idea i'm a hundred x stronger in my convictions there I have not opened a terminal for anything other than offing, like, a CLI or something.
Not once since my hand thing. Other than, like, this one YOLO run I have going for my... Other than this one YOLO run I have going with TS Rust trying to use Astra to rebuild all of TypeScript in Rust.
And now I'm just running there because I don't want it conceptually taking up space in my T3 code. I just want it YOLOing in a corner somewhere. But that's the only thing I've opened my terminal for in weeks at this point.
Partially because I'm over the terminal, but it is mostly because using a terminal with your voice is actually hell. Like it's just, it's hell. I hate it.
I hate it so much. I don't want to ever have to do it without my keyboard. I'm done with terminals.
So how have I been living without terminals? The first and arguably most important thing is two projects I have on my computer. The first is named Fleet.
The Fleet project is how I manage all of the different computers that I currently have set up for real work, vibe coding, whatever else. It has documented all of the computers, how to connect to them via SSH, what I use them for, what's installed on them, all of those things. Historically, I would have had to like SSH it to a computer to install something or set up the author, whatever on it.
Now I just don't. Now I don't even have to think. about this type of thing and it has made life meaningfully better for me overall even outside of the keyboard thing because when i have a problem here i'll show a real silly example here super smash bros melee just got decompiled which by the way unbelievably cool massive achievement in the gaming world software development super hyped about it i wanted to play with this i wanted to play with it on a different computer though because i didn't want it interrupting me while i was filming today So I did a kind of silly thing I downloaded the ISO on this computer, but didn't have an easy way to transfer it over That's the type of thing where I could absolutely have like SSH in or like went and wrote the SCP command or whatever else Then I also have to wait for the download extract it manually go to it in the terminal and then write the command I'll show you what I did instead I started a thread in fleet because, again, this repo knows where all my computers are and how to connect to them.
I'm downloading a legal copy of Super Smash Bros. Melee, hypothetically. When it's done downloading, hypothetically, I need you to transfer the ISO of a game I already own multiple copies of to like bed in downloads directory so that I can access it on this computer as well.
And then I had to correct it to left book because I have a lot of proper nouns in my. dictation, like dictionary for WhisperFlow. And since lakebed and leftbook are two of those words, it sometimes mixes them up.
I also was doing this via the voice detect, or I was also doing this via the MacBook mic and probably wasn't as clear because I got too used to this tiny little guy. But I was quite annoyed that it took the word lakebed instead of leftbook, but I made the quick correction. And this one was fun because I hadn't even finished downloading the file.
I literally just told it, it will be in this directory. I want you to copy it to the other machine in the same directory. And then left.
I stopped looking. I stopped caring. And then I got a little ping when it was done.
And it was indeed done. I know this seems silly, but no matter how good you are at computers, getting a file downloaded on one computer and then moving it to the other, even if you have a system to do this really fast, the additional steps that are necessary of just waiting for it to finish before doing the next thing is annoying.
And this is part of what I've been doing more with AI, is I've been pulling it in slightly earlier and giving it the instructions necessary to go slightly further.
Another somewhat silly example of this is around T3 code itself. We've been working heavily, and when I say we, I mostly mean Julius as well as some... changes from people like shiv and maria and a few others but largely julius has been working on what's called orchestrator v2 orchestrator v2 is an overhaul of how t3 code threads are exposed and managed by the agents themselves which will allow for agents to And this will allow for an agent to spin up another thread with another model, to reference one as a sub chat, to get more feedback, to spin up a bunch of clod sub agents via codex, those types of things.
And then we're very excited about it, but it's a huge overhaul of how orchestration works in T3 code, which means we're taking our time merging it, which also means I haven't been able to play with it a whole bunch because I use the T3 code nightly on my machine and I do not want anything bad happening to it because I rely on it heavily.
I use it hours a day, every day, probably averaging 10 hours a day in T3 code lately. It's bad. But I wanted to have the orchestrator V2 as a build I could play with on my machine.
There's a lot of ways I could do this. I could go clone the branch or I could pull down the branch and run a local dev build. I could ask Julius to set up a DMG and then wipe my install, back it up and then use that temporarily.
But I had a specific thing I wanted. And for those already asking, like, what if you just put under a feature flag? You're not putting a new database under a fucking feature flag, especially when that new database is constantly changing and you should expect everything to be wiped at any point.
You don't want a feature flag for this and it would make it way harder for us to ship, not easier. So no, not happening. We will have a nightly with this someday.
Just wait, be patient. It's happening. Anyways, I wanted it now and I'm sure others of y 'all do as well, which is why it's so convenient.
T3 code is open source because you can copy paste this exact prompt or. I don't know, voice to text it. And then you will have what I have here.
Funny enough, I sent this one from voice to text on my iPhone as I was getting my hair ready in the bathroom before stream because I wanted to have it ready to use if I wanted to demo it during stream. So I was on my phone. I started a new thread in Fleet because again, Fleet has all the context of all my machines, also including the context of the machine I'm currently on, which is useful.
So I told that I wanted to make a custom T3 code app slash build on this machine that uses the orchestrator V2 branch. It should be named T3 code V2. It should use a custom T3 V2 home directory.
This one was very specific because I didn't want it to override and overlap with my existing T3 code instance so I could rotate between the nightly and the V2.
I want it installed the same way you install any other apps, just to be very specific. You know, Astra needs a little nudge in what you actually want because it's not the best at intent, but if you talk enough, you can get it to figure things out. And you can figure it in a way where it will not interfere with my existing T3 code install.
And then I went and finished my hair. I came upstairs, I set up stream, and then went and checked, and it had finished. And now I have this custom build of T3 code on my machine.
Again, I pulled in the AI earlier than I normally would have. In the past, I would have probably pulled down this branch.
In the past, I probably would have made a clone of the repo or a work tree. I would have went to GitHub and found the branch. I would have found the PR and everything for it.
I would have cloned that locally and then told an agent to build it, maybe. Now I'm telling it to do all of the steps before and after. I'm telling it to go find the PR and figure out how to get it locally and built and do it.
I will give you guys the advice of if you think anything we could do different in T3 code is as simple as putting the word just in your sentence, that you are entirely wrong. And I would politely request that you don't do that because we have thought these things through very deeply. It is not as simple as you could ever fathom.
We cannot have a database breaking change under a feature flag. It's not happening. You don't put migrations under feature flags.
That's just reality.
Anyways, this worked great was very happy with it, and this is just some of the computer use examples I'm gonna hide my screen for a sec Trying to find a specific thread I did.
Ah, I found it.
You guys ready for an extra personal one? This one's so personal, I opened up ChatGPT for it. I am between doctors because of my hand and other things.
I love my current doctor, my... I absolutely love my orthopedic surgeon. He's the best.
But since I'm between primary care, I've been trying to transfer everything over.
A lot of the stuff I needed was available for me in a dashboard for the hospital I'm with. And their web app is shit. I'm sure we've all had to deal with this before.
A app or website for your doctor's office that is not the most pleasant thing to use.
So when I was doing my early access testing with GPT -6 Astra, it is computer use capabilities. I gave it a fun challenge. I asked it to go to the website I had open with all of my medical records on it and download them all.
This took about 40 -ish minutes total, I think. Let me... Hide things here so you don't see too much.
Actually, it wasn't quite as bad as I thought. It was only 20 minutes or so. And I watched it do a lot of this.
It had to download 49 medical PDFs, a giant pile of these scans, and each one of these things it downloaded required going through various dashboards, various threads, waiting for things to load, scrolling, finding the attachment section, downloading it, verifying it, and organizing it. And it did all of it. It did all of it fine.
As long as you have chat GPT open, you can also do this in T3 code. But for me personally right now, my split is T3 code is for everything code related.
And chat GPT is for everything life related. So when I'm doing things like... You guys are going to be annoying with this.
I'm going to change it.
It was not a custom model, guys. It said custom because when I was testing Astra, it didn't say Astra. It said custom.
It is Astra. I just clicked it and it said Astra underneath. So now it just says Astra.
Stop thinking I'm hiding something I'm not. This was an Astra thread when I was testing it.
The point I'm trying to make here is that doing this by hand might have taken me half an hour or so before.
Now, with my hand like this, though, it would have taken me hours. Like, legitimately hours.
So having the model do it slightly faster than I could by hand, by just telling it what I wanted, and then coming back 20 minutes later and all my records are local? Awesome. And then my doctor texted me the link for where they wanted me to upload it, so I told another thread, hey, get all this uploaded, and it did it for me.
I didn't have to do it on my computer the same way I would have before. And these are the tasks that I never minded because I go around my computer quickly. I've been using computers forever.
I have a bunch of custom workflows, all my hotkeys and everything. I can navigate a computer fast.
On one hand, literally, I'm down a hand. So I can't navigate it as quickly as I'm used to, which demoralizes me as I'm going.
But on the other hand, the good one this time.
I can only do one thing on a computer at a time. And what this has unlocked for me is a mental model where I am doing more than one thing on my computer at once while also working on multiple things in parallel in T3 code at once.
And I found it almost like it's been a bit of a relief because I found that before I was prioritizing work. based on a combination of what mattered the most, but also how much time I had. And since I would regularly have to like get up to go to a meeting or run to go grab something or run to and from my office or go film a video, all those types of things.
If I had a task that took an hour to do as far as I would have guessed, but I have to go film in 40 minutes, I'm just not going to do it. And then that thing gets delayed indefinitely and never happens. Now, with my combination of voice -to -texting and computer use and my fleet of computers I can be doing things at all times even if my laptop is closed, including a dedicated MacBook that I have on my network that is literally just for doing things that need, like, macOS computer use, that's what Leftbook is.
It's my leftover MacBook that just is decomping Melee totally legally, by the way. This is the type of... The mental unlock of unbounding my like...
Does it end up being like a weirdly big mental unlock for me? The idea that the length of a task does not bound when I have to do it anymore. That if I had an hour -long task that meant I had to find an hour in my day to do it, now I don't.
Now I check in at the start and I check in at the end. And it's so nice.
This is just the start of the mental model shift though Again, I've been thinking a lot about when I start a task and when the task is done and if the agent is like chunks in the middle like if We were to think of tasks. I'll just diagram. This is the easy way to explain it Like We think of this as a spectrum where on one side, we have the idea for the task is in your head.
And on the other side, the task is completed and you're content with it. Where do you stuff the agents in? I found for a long time, my agent use was kind of like this, where it have like chunks here and there.
Maybe even at the start, I would like ask the model, hey, what do you think of this idea? And it would give me some feedback. And then I would iterate on it.
Then I would pull in a... And then I would pull in another agent to start iterating on it. I would give it a bunch of feedback.
I'd play with what it built. And then I'd pull in another agent to go implement the things I wanted different.
And then as agents got more and more powerful, I found myself moving to a model more like this, where I would talk things out occasionally after I thought about them for a while. decide what I wanted to do with it, and then spin up a long agent run to go build the whole thing. Then I would test it a whole bunch myself here, play with it, figure out what I like and don't like, tell an agent to fix the handful of things that are wrong.
Maybe there's some review comments after that it has to address. Then the task is finally done and I can go hit merge.
And here is where everything has changed for me. First off, I started trimming these parts out.
More importantly, I've taken this middle section, the long agent run, once everything is figured out and the agent starts building and then out comes the code, I started telling it to verify its changes more. Instead of having to spin up the code base locally, open the app, and test a bunch of things with one hand, I told it, use computer use to verify and validate your changes.
Use the AI code review bots in our repo to give feedback and make sure the code is in a good state. Have some sub -agents do another pass on your code to increase your confidence and don't bother me until you're relatively confident that there will be no regressions from this that are user -facing. I've also started moving further in the other direction where I tell it not what solution I want, but I'll give it a bit of the problem I'm having.
And if it's vague enough or I'm unsure enough, I'll tell it, hey, I don't know how I want to solve this. Propose some solutions. Sometimes I'll just tell it to go do the whole thing and get a PR up.
The craziest change though is this very little bit at the end here.
This is the merge hole. This is where I would come in and there's almost like a line here. Like how far can I get the model before I go hit merge?
I don't have this line anymore. I trust the model to know if the code is safe to merge. and to just do it for me a lot of the time.
I've had Astra merge over 100 PRs across my projects, and I've had Fable merge at least 50. And of the 150 PRs that they merged and wrote and did everything themselves fully autonomously, two kind of had regressions in them. That's a better hit rate than most talented developers have.
Significantly better hit rate. This is similar to like the self -driving car thing where like everyone freaks out when they see a self -driving car get in an accident. When you look at miles driven, you'll suddenly realize, oh, yeah, I guess that they're driving five times as much per accident.
That's probably good. And the accidents seem less dangerous. The accidents that happened from these YOLO merges, by the way, were an animation being removed in the app and an animation being removed on the marketing site.
Those are the two regressions that accidentally merged after almost 150 PRs of auto -merged slop from Fable and Astra. Astra shipped those two animation regressions. Everything else was fine and has been massive real -world improvements to these apps.
YOLO merges. Can't tell if we're still talking about Waymos or code now. That is a phenomenal joke, Dranben.
11 out of 10. That got a genuine smile out of me. Thank you.
If you keep asking questions while I'm in the middle of filming things, you're going to be timed out, bud.
Anyways. This is a theme you're going to be seeing in more and more of my content. This idea that you should expand both sides of where you let the agent come in.
The agent should come in earlier in your process. and it should go longer before it bothers you i'm now at the point where i don't even care about how fast the model is anymore for the most part other than like emergency bug fixes and things i'm like in the loop on because i need them out asap or i care a lot about some subtle specific details in it or like design iteration and things for almost all of my threads like if i just literally go through them i'll be honest with you for each one how much do i care about the speed of the thread building swift ios app Funny enough, I do actually care a bit in this thread, but the reason is hilariously dumb.
I built 95 plus percent of the Swift UI version of the T3 Code mobile app in a single thread as an experiment. This thread is like eight plus gigs of data now because I've been using it so heavily across models, across months of work. So this one, I like the fast iteration cycle because it makes it easier for me to try two things or three things in a row because it's a hell thread.
If I just merged this code and had work trees the way I normally would, this would not affect me in the slightest.
But since I have a single thread I'm doing this work on, I would like it to be faster. So there's one so far where the faster would be nicer.
Lakebed single server capacity limit increases. I don't give a fuck. I check on this thread every like day or two.
Melee decomp. This is going to run in the background indefinitely. I don't care how long it takes.
It could be three times slower. It wouldn't make a meaningful difference in my life. More Lakebed performance games.
I don't care. Showcase T3 code performance games. I don't care.
Engineer performance rating. I don't care. This was for a shit post on Twitter.
I can archive that now. Fix image preview and rendering. I don't care.
Categorize stable release changes. I really don't care. All of these threads are things where the amount of time it takes doesn't matter because as soon as I kick off the thread, I'm off doing the next thing.
One of the coolest features we recently shipped in T3 code, which admittedly is a bit harder for me because one hand, we have the ability to kick off a prompt without leaving the prompt screen. So if I, I don't know, say, I want you to work on feature one. I could hit enter and it would send and start the thread the normal way.
Or I can hold down command and press enter. And it opens a new thread and I stay here. So I can kick off another job.
A third job. And if I'm feeling particularly frisky, a fourth job. Not a part job.
Thanks, Whisperflow. A fourth job.
Thank you, Whisperflow. A fourth.
job ta -da voice to text is great you get the idea though this is so nice for when you have like a mind dump that you want to do of just random things you want to have worked on it's so nice to just do it one other thing that's been kind of hard for me to get over As someone who cares a lot about things like their grammar and text formatting, I'm a bit...
All of the words I want to say about my relationship with grammar could potentially get me canceled or in trouble. So what I'll say is I care too much. I like properly formatted text.
I like sentences that vary well in length. I like writing well. It's probably part of why I do this whole YouTube thing now.
I care about the quality of my writing and the clear... I care a lot about the quality and clarity of my writing, using the right words, using small, simple, and easy -to -understand words. So getting over the fact that every three to five sentences, one word would be wrong started to drive me mad.
But now I'm mostly over it. I have learned over time that, believe it or not, the thing translating my voice is also AI. In its ability to deal with the fact that certain words were the wrong word in what I said, it handled it great.
I've been really impressed. I bet if I scroll through a handful of these, we can probably find some real examples of me having typos in what I sent.
i think i still correct it too often i need to get over that Okay, I just hunted for a bit across my threads and concluded that I don't actually have as many typos as I thought in them. Because I always fix them.
Yeah. I guess I'm giving this advice not for y 'all, but for me even. The model handles wrong words and typos, so to speak, totally fine.
It really doesn't matter that much. So you should just not care as much. See what happens.
Just send the message with the typos, and if it gets confused, you can correct them then.
I could also ask the model to go find some examples, but I don't care enough to, honestly. I just wanted to make a point with an example. Didn't have one.
We'll continue.
I want to make this next point.
One of my like main or one of my like my primary angle of attack for solving all of these problems I've been having has been focusing in on where my friction and pain points are. Like what is frustrating me when i use my computer and how do i get rid of it to the best of my ability how can i offload these things and one of those things is navigating github because it has gotten miserably slow it is hell to navigate and not long ago you can probably even see in some of my old videos i would have literally 30 to 50 tabs from github open in my primary work browser like this whole section here would just be github tabs right now i have one and it's the melee decomp part of that's because i purge before stream but most of it is because i don't have to but most of it is because i'm not interfacing with github directly anymore let me give an example can you Show me some of the recent pull requests that might be worth my attention, things that have been done in the last three or so days.
Look for stuff that touches the surface areas I'm most concerned with, both opened and closed PRs. Give me a brief two to three sentence summary of each change, as well as an additional sentence at the end on why you think I specifically would care about the change.
Send. And now in some amount of time, I'll have a nice pile of text.
And now in not much time, I'll have a nice readable diagram. Now in not very much time, I'll have some nice readable output that describes things that would have taken me a lot of time to browse through on GitHub. But even better, because again, command tab hurts me right now.
We have a pull request viewer that Bilal built for T3 code directly in app. So I can see the pull request right here if I want to for any reason. There's a lot of reasons I would want to.
Now I can even hit merge without having to leave T3 code at all. To be fair, anything I would have hit merge in the T3 code app for, I probably would just let the model merge itself. But it's very nice having all of this built in there.
Yeah, and here we are. A very useful set of... Here we go.
A very easy to read, useful pile of summaries of real world pull requests, their current status, and also the ability to click it and open it within T3 code without ever having to touch GitHub. It's great. It's such a relief to have all of this just right here ready to go.
I have one last pro tip I want to show you guys, though. One that I have been surprised how useful it has been. This is going to be very different in the near future once we have the orchestrator out, but if you're in other tools like Cursor or you're directly using something like Cloud Code or Codex or whatever else, this might be really helpful for you.
hate dealing with copy pasting text right now. It's silly, but because of the weird way I have to put my hand on the keyboard, command C and command V are unpleasant to hit. They just are.
It sucks. It is what it is. As a chronic copy paster, it has made life much harder for me.
So what do I do instead? I regularly am passing context around between my agents. I'll still use this tiny little like copy button here.
Like, for example, let's say I want to help. Like, let's say, for example, I want to help prioritizing these PRs. I can click copy.
I can make a new thread. I can say, help me prioritize which of these PRs I should look at first. If any of them are simple and easy to merge, tell me and you can merge them.
I'm going to delete that part because I don't actually want it to YOLO merge while I'm streaming, but you get the idea. Oops. And I couldn't hit paste right because, again, hands don't work great.
But if I two -hand it with command on left and paste on right, there's the list that I forgot to paste. That's one strategy. But you can go a step further here.
I'm going to stop this in archive and do something a little different.
I have a recent P3 code thread today where I asked Codex to review recent PRs by Surface. And help me prioritize pull requests. Which three do you think are the best for me to prioritize right now?
Is this the most efficient way to get the context handed over to a different thread or agent? No. Absolutely not.
Is this going to make my request slower than if I had copy pasted it? Absolutely yes. It's even going to be a bit more expensive.
But I don't care. But frankly, I don't care. I'm going to be doing other shit at the same time anyways.
The things I would do to make my agents more efficient by unblocking them and getting work done for them ahead of time, like don't do work the agents can do themselves. And the more you get in this mindset. It's weird and strange and uncomfortable because half the time you're going to be like, well, I could have just opened the browser and did that myself.
Am I being lazy? Yes. And lazy is good.
Lazy is efficient. Lazy means you're taking advantage of the capabilities of these things more directly. And one of the lazy things you can do is not pass context around between things yourself.
The models are good enough at building context by using their tools now that you can tell it to go find the things. Will this work as well on cheaper models like Terra or Kimmy K3? Maybe, but probably not.
But if you're using frontier models with the $200 a month subscriptions and you're okay with wasting some tokens on how you get the context to the thread, it doesn't end up netting to that big a difference and it will make it way easier to just dump the thing out of your head. And if you think I'm insane, I have a challenge for you.
An exercise. The next time you have a bug you notice in your software, a feature you want to work on or something else where your next step would have been to, I don't know, plan it out or write something down or do something other than just set off an agent to do it. Right before you start, set off the agent with everything in your brain already.
Don't help it at all. Don't give it hints. Don't point it where it should go.
Don't give it the additional context by manually bringing it in. Just. Tell it what's on your mind.
And while it runs, go do things the way you normally do. Then come back at the end and see how the way you did things differs from how the agent did things. And see if you can honestly still kid yourself into doing all that manual work yourself.
I think you'll be surprised how much of your process can just be done by the agent itself. I know I have been. I have been blown away at how quickly the agents have chewed through the boundaries I had on each side of my mental model, where I bring them in sooner and let them go longer.
And my hand has kind of been a forcing function to get me deeper in that. And I don't think I'm ever going back.
So, yeah. This is a strange video, I know. I have no idea how it will perform.
I'm hopeful it's good because I think this is cool. I've fundamentally changed how I work at my computer every day. And while, sure, my hand is a meaningful part of why I've had to do that, I've actually been very happy with the results.
And while I dearly miss being able to type out a shitpost using my hands themselves, and while I do dearly miss being able to type out my shitposts by hand, I have found this new way of working to be quite pleasant. And this tiny little mic has made it a lot more tolerable to live with my computer with only one hand. I hope these tips were helpful.
It's been a huge change in how I work. And while I know most of you guys do have both your hands, for those who don't, and for those who want to just use their computer differently, I bet these tips will be helpful. Let me know how y 'all feel.
And until next time, peace nerds. Fun.
out of dev let me let me frame this to you differently let's say you were working at your office writing code you were in the middle of working on something really hard you were fully locked in and focused trying to do your job and i came in and like got between you and your monitor was like yo yo i have a question for you How would you feel?
If your answer is anything other than that would suck, it would really distract me and it would be harder for me to do my job. I would have to conclude you're unemployed because every basic human would be meaningfully interrupted when you throw a question in their face while they're doing other things. And when you came in to ask the question, I was in the middle of filming a video.
And it was very clear I was in the middle of filming a video. If you're a huge fan, as you say here, that is because you watch my videos, so you know I film them live.
And the easiest solution when that happens is to... make the thing distracting me go away.
No hard feelings, but the point of the timeout button is what it said, timeout. It was a go away for 10 minutes, I'm busy button. And yeah, it is actually annoying that markers are still the best way for me to do this type of flow.
I want to improve it. I have some ideas, but I'm way too busy. The main thing, yeah, if you're asking chat, that's fine.
But if you're tagging me a whole bunch, like, hey, at Theo, can you answer my overly verbose message about something entirely unrelated to what you are talking about? Here's my life story that you will glaze over because you are busy.
The number of messages I would say fit this exact format is hilarious.
I guess I just made a new... Copy pasta. Great.
This to myself.
Thank you, Lava Beast, for the sub. Ryan is Bliss for the sub. I think I already thanked you guys.
PewDieBird as well. I definitely thank y 'all. We got KiwiCodeHack with the six months of tier one.
Appreciate you a ton. Living the AI dream now that six is competing with Fable. I agree.
Charlie B. Sup?
I do know what I did. This is entirely self -inflicted.
They're overcoming anxiety about risks and giving agents access to your computer and credit card. QuickVid idea might be reactivated to Tim Roswick on stop shaming game devs.
I have not seen this. I will absolutely watch this later. I want to watch it before watching on stream just in case I have different opinions that I would expect to.
But I have a feeling I'm going to agree with every single thing he says in this. The game dev community is much more toxic than it should be. I blame gamers, though, because gamers are gamers.
Anyways, thank you for giving me the heads up on that. I will actually definitely watch that later. Let's pick rock.
I get rich making AI girls on Instagram with that shit. Hilarious.
Anyway, somebody asked here if this ever reloads.
Thanks, YouTube. I had a thing I wanted to look at here.
And it will eventually catch up. So my friends at YouTube invited me to go to the office to meet their new head of product for live. And I have been delaying because I need to be in the right mindset or I will be way too toxic because I actually need them to fix it because I want YouTube chat to be usable.
But it is not. It is not at all.
Gamers were mad about Silksong being too cheap. That sounds like gamers. So this is what I was looking for.
I should know what a polymath is. I recognize the word, obviously, but I don't off the top of my head.
Knowledge and expertise span different subjects and fields.
Eh.
I'm pretty shit at painting, anatomy, and okay at math. If you can be a polymath by caring too much about skateboards, caring too much about music, being okay at cameras, decent at computers, and pretty good at code, sure. My hobbies and interests don't have much overlap, and I'm weirdly deep in a handful of them.
But this term means nothing to me. I don't consider myself particularly like... exceptional at enough things for that to make sense to me like s john says making him switch pcs just to say yes bro you are god damn it appreciate you s john this is a guy who gave me a lot of shit in college by the way s john's known me for a long time but uh appreciate it i think i don't think generalist fits me at all actually because there's a lot of things i am like really bad at like I don't know jack fucking shit about food and nutrition.
I don't know jack fucking shit about movies. I don't know anything about television. I don't know shit about history.
I'm really bad with dates, like comically bad with dates.
There are too many basic human capabilities that I'm not good at that I would say disqualify me from generalist. The thing that makes me weird is that where most people have the one thing they're really good at or a bunch of things that they're okay at, I have three things I am pretty good at, one or two of which I say I'm really good at, and everything else I just don't give a fuck about.
So it's more like my autistic obsession with these three to five things makes me look really good at them. but I just have that small subset of skills. Like those are the things I am good at.
And that's it.
Like I was a sponsor skater. I've been a media professional for a while now. And obviously I'm a programmer.
So good enough for those things.
I think being a polymath and not being good at basic human interaction is required to happen together. Fair.
I'm currently stuck on the RODECaster Pro 2. I don't want to do my audio processing on PC. But I have gotten near the point of giving up now that it's failed so many times for me.
I have my MixPre 3. If you don't know how much this costs, I would love for you to guess. This is a minimal portable audio interface that has three XLR ports that is meant to be mounted under a camera that also can be powered by double A's.
I love guesses as to how much this simple audio interface costs. Knowing that you can get something with two XLR ports on it, like the Scarlett 2i2. Is approximately $100.
For three ports. And the ability to use it on the go. While also still having the USB port.
Where is that USB port? Here. It's like plugging into a computer.
$350 to $600 is the range I'm seeing a lot of guesses in. Which is the right range to guess. That is what it should cost.
The MixPre -3 2 goes for between $1 ,000 and $1 ,500.
If you ever wondered why we need sponsors, it's because the box that turns my mic into audio consistently enough to make me less angry all the time is $1 ,500 fucking dollars.
This is the professional media world. Is there a reason why it's so expensive? Because the only comparable things are similarly expensive.
This has a lot of things that make it nice. It's not filters. It has no filters on it.
There is no onboard anything on this fucker. You use this because you want to have the exact audio you piped in. in a file format that you can deal with later.
So if I am out recording and I don't know how loud or quiet people are going to talk, this fucker does 32 -bit float, which means that the loudness levels it encodes are floating. They're flexible. They change over time.
It can go into decimals and it can go into way bigger numbers because usually audio is 16 bit and not float. So you have a much smaller range of like loudness and quietness. You can encode.
This guy is a 32 bit float interface means that it can capture whispering and it can capture screaming all without clipping.
It also means that this, or this also has a really good set of preamps in it. So I don't have to buy a separate $100 to $150 preamp for my mic before plugging it into an interface that's cheaper. Does all of this make this thing worth $1 ,500?
No. It is not a good value. But the MixPre 3 II, which is what this guy is, is one of the most reliable audio devices in the world.
It is also basically indestructible. And it means that when you plug it in to your mic and hit record, that you can deal with it later.
You don't have to worry about your game being set perfectly. You don't have to worry about any weird issues with the built -in filters because there aren't any. It's just good enough.
And it does what it's supposed to. But since I have to set up filters for it because I can't do it on board, I have to set it up on my fucking computer instead. My plan after stream, if I have any energy left, is I'm going to once again give Codex access to all of my current RODECaster settings and tell it to recreate them in OBS with this is my interface.
And I'm going to see how it does. Hopefully my next stream, I will move, have moved off my giant hellish road caster into this tiny little guy instead. It'll make my life much easier eventually.
But for now, this $1 ,500 guy is sitting on my desk on the side here waiting for me to find the time to set it up.
Yeah. Fun. Yeah.
So you ever wondered how expensive it is to do this shit properly? You now know it is not cheap. The SM7B is indeed gain hungry.
And this guy handles that great because it has awesome preamps built in.
I have it at 62 dB with this guy right now. I want it on 65 because I don't actually talk too, too loud and don't want to like boost and post as much. So I have it on 62 dB.
Compression is three to one now. I knocked the cutoff forward a bit lower because I. change a bunch of other settings it's the road casters get just barely sensitive enough for this mic but if you push it to the gain levels i want you start to get a lot of interference the camera is the same it has always been since like i moved the s52x i love it or i absolutely love it it's my beloved camera i have five of them now because i have one on my other desk i have one up here and then i have three in the podcast studio for the podcast unbelievable camera you can get them used for like 1800 bucks now which is crazy if you don't need the usb cssd recording you can get them for as low as like 15 or 1600 with the non -x version so great camera unbelievable highly recommend I'll do this quick.
There's a couple other people working on the scroll back stuff.
Branch was forced, pushed, or recreated, so I can't reopen it. Make a new PR with a new branch.
I wanted to call it T -Shots. All of the other names I came up with... Were even worse for what it's worth like Julius didn't even share them How can I defend that I came up with it it's great I was gonna call him skeets nobody let me do that either though Final stream maybe later.
I remember I only have one hand I can't press both commands I know that that kid asked this question again.
You can use any of these models to do anything you want at this point. They're both more than good enough. It doesn't matter if like...
I don't like this exclusive way of thinking. And I would encourage that you wait for the video I am filming very soon about Fable versus Asterix. I go much more in depth on that there.
The codec sidebar is the actual worst sidebar in any of the apps right now. So no, we don't have interest in making the sidebar worse. Are you talking about the...
Where am I at here? Are you talking about the activity view or this view? Because this view is fucking trash.
This view is garbage. And activity view is the start of a good idea, but is also garbage. Best part, though, when you pin things, you can't even see which folder they are from and which project they are in.
Yeah. So to answer your question, no, we do not intend to make the sidebar worse. If you want to use the old school sidebar that is worse, I think I still left it in here as a legacy feature, but I plan on killing this very soon.
The legacy sidebar that is objectively just worse is there.
What do you think is better about the other sidebar? Because I don't see it at all.
it not moving the chat thread oh oh i'm sorry that's different uh so if i collapse the sidebar here what do we do different is it just that we don't have an animation what is moving or not moving i don't want the animations i i don't know what you guys are referring to here i don't see it i don't get it The codec sidebar won't move the chat thread when expanding and collapsing.
The sidebar will not overlay with the thread. I've never seen that behavior. I don't like that.
That sounds like it would leave odd amounts of space.
Sorry for missing this earlier, Gusty.
I think I'll look at this post stream. Maybe I'll share it next stream. But I need to be filming right now.
Snowy, I'm waiting for a deal to go through with Cloudflare. We will have to charge for more than three for most people. But if you text me or DM me or signal me your email, I can get you approved for it now.
So I can bump you to 100 if you need more connections. Just send me your email and I'll get you bumped.
A wide view, unlikely to be a priority for us for a while. We really want to make things as stable as possible.
I'm still deciding long -term how I want to handle things with Cloudflare for T3 code and Connect. 50 -50 on if I want to stay there. Also, what architecture I want to go with.
Lots of thinking and planning there. We should probably move into durable objects in the not -too -distant future. The five -limit bump is not going to happen for free, sadly.
It's going to cost us too much money. Also, part of the terms of the agreement is that I'm not allowed to say the price, which is very annoying. But yeah, I sadly will not be able to be as generous as I would have liked, or we will actually run out of money almost immediately.
Just for the current connections we have on the free tier with those three connections, by the way, it would currently cost us around 10 to 12 grand a month with the pricing that we got. And that's like a crazy discounted price that we got. So yeah, just the free tier is going to cost me 12 grand out of pocket.
So yeah, we need to change our architecture.
Okay. Are you a different snowy than the one that I normally DM?
There's too many snowies.
Yeah, my bad. I make this mistake too often. I'm sorry.
regardless you've been very generous i'm more than happy to give you the gift or the uh unlocked limit dm me on twitter make it clear it's snowy 77 and i will do my best to get to that industry it needs to be i specifically need the email you use to sign in for connect let's start filming more because we've done one video and i've been going for two and a half hours so we need to talk about astra and fable do i do the verses just dive straight into it yeah fuck it let's dive straight in and for those asking like can you use cloudflare yourself for t3 connect or can you use like tailscale yes absolutely you don't need t3 connect you can just use tailscale for everything it's totally fine did uncle bob get hacked Yeah, he did Yeah, I think either he got hacked or his agency something weird I'm pretty sure yeah nephew.
What are you going to buy? You know he's hacked Annoying Just know though gross.
Thank you. I'm Matt for the support. It's always five months goddamn and then gusty cube as well Able 5 .1 versus GPT six ass I'm for it Let's dive in, shall we?
Last week was pretty crazy for new AI models. We got two new heavy hitters that are truly changing the game. Gemini 3 .8 Flash and Muse Spark 1 .3.
How can you ever decide between those two? They're both so incredible. Obviously, I'm joking.
We're actually here to talk about Fable 5 .1 versus GPT -6 Astra, the two actual models that matter. And I'm sure my retention guy already wants to kill me for doing that joke and confusing people. So sorry, Will.
Let me have some fun, okay? Everyone's going to watch this video anyways because they all want to know what model to choose. And these two models are unbelievable.
Fable 5 .1... I went in with pretty high hopes on Fable 5 .1, but didn't expect too much, and I still managed to be absolutely blown away. It's an incredible model, and I use it every day.
GBD6 Astra is the future. It is such a crazy leap in so many different ways. I love using it.
I'm doing thousands of dollars of inference a day with it. These models are both incredible. But which one is better?
Which one is better for what? How should you decide which one to use for which thing? And of course, most importantly, which subscription should you get if you're only able to get one?
There are a lot of layers to these questions and a lot of complexity to address. And as much as I would like to give a simple answer, I can't. If you're looking for that simple answer, you can skip to the end.
But if you want to actually understand what the benefits and negatives are of each of these models and how to use them properly in real world work, I would ask you to watch, and at the very least, take a quick break. I would ask you to watch this whole thing, including this quick break for today's sponsor.
I don't want to split this.
Code capabilities.
Dakota sections. Front end. False stack.
Not false stack.
Giant project rewrites.
Code mergeability.
Managing scope creep.
Anything of all the other things people use models for that are worth like comparing on a spectrum here?
That is more than I expected Now it's on a refilm the transition here Yeah, I'll do that You know what I have a funnier way to do this if you want to just skip to the end and see the answer or if i just skip to the end to see which model is better i guess you could do that but you'd miss all of the different ways we can compare them because these models are very different in their capabilities and the things you can do with them and there's a lot to be learned from the various skills these models do and don't have i'm going to break all of this down and more i'm using both these models every day doing fat I'm doing nearly a billion tokens a day across both of these models consistently, and I've seen their strengths and weaknesses in all sorts of different places, many of which have surprised me a ton.
So if you want to understand what these are good at, what they are bad at, how efficient they are with token usage, how hard the subscription limits will hit you, all of those types of things and more. And of course, the most important question, which sub should you buy today? I'll do my best to answer all of that after a real quick break for today's sponsor.
That puts us in a good spot to start So we have all the different ways you want to compare these models let's start with what I'm sure everyone is here for Science because this is a science channel, right? Okay, seriously, that's one of my favorite things to cover Not just because I think the 3d stuff is really cool, but there were some funny things in this launch Because I'm going to start with the official Fable 5 .1 launch notes because they were very excited to show off their progress in terminal bench science 0 .1.
A very difficult real world science bench using tools to solve real world problems. And they had a huge jump here, both in the efficiency for cost as well as the success rate overall, where their best case previously was on high with Fable 5. It cost $34 and it got a 25%.
Now 5 .1 on X high is even cheaper and got a 50%. Huge improvement. I see why they put this here.
These numbers look awesome.
At least they looked awesome until you went to the GBT6 Astra launch.
Because Astra on low. gets a 54 .3 and only costs 11 putting low astra higher than max fable 5 .1 at under a third the price for the first benchmark that anthropic had listed this is one of the greatest mogs of all time in the ai world to have just days after anthropic bragged about their huge increase in knowledge science literally the first benchmark on the site Have OpenAI come in and just casually demolish them, getting almost a 65 % for $26, when the best Anthropic can do was a 50 % for $40.
What I'm trying to say here is, if your job is science, you should probably go get a Codex subscription. You could do some pretty crazy things with it, like mind -blowing ones. Both of these models have been massive leaps in science.
My assumption is there's either some new training data or some new RL tooling that has been given out or sold to the different labs. And if it is something useful in post -training, then OpenAI got it and is using it heavily there. It says OpenAI are the goats of post -training.
They were able to apply that better. On the note of overlapping data, 3D rendering.
Both Fable 5 .1, and gpd6 astra made massive leaps in their 3d rendering capabilities astros was way bigger though like hilariously bigger let me deal with some that guy open some projects quick Sorry, just got to kill some ports on some things.
Well, there we go.
Where better than to... He's muted, actually.
Where better to start than fish slop? This is Fable 5 .1's version of Fish Slop. It is a 3D game where you have a submarine and some fish that you can feed and a mini economy as well as a little bit of combat.
And this was the most impressive version of Fish Slop at the time because this version of Fish Slop had real models that were surprisingly good. Like the fish actually kind of looks like a fish. It has eyes placed in the right place.
That is a lot harder than it sounds. There are like coral and rocks and lights that work. Most importantly, though, the movement is actually very nice.
It feels good to move around in this version of the game, like flying around or not really flying. Floating around with the sub feels awesome. It controls great.
It plays great. It's surprisingly decent overall. But there's definitely room to improve.
I say that with confidence because of Fish Slop. I say this with confidence because of the version Astra made.
I'd say this looks... Slightly better.
Just a little.
It's a generational gap here. This is next -gen, and the other version was not. What's the hit button to shoot?
Okay, left click.
There we go. Killed the guy. This does have its flaws.
The movement doesn't feel anywhere near as good, and I had to go back and forth to refine it a bit. The core loop isn't quite as well refined overall and the performance was bad until I told to fix it. But it is fucking stunning.
Like it's actually decent looking to the point where I might have to drop slop from fish slop in the not too distant future. It is a massive leap. And if you think I'm just showing this and saying this now because I'm trying to glaze Astra.
Go watch my video on Kimi K3 where I was for the first time ever genuinely really impressed with the 3D modeling capabilities of an LLM. This is a world of a difference. This is like multiple generational difference.
So while I definitely want to give this without any question to Astra, I actually think I have in my notes here. A bunch of examples of crazy 3D stuff people did with the model. Here, Dara made a copy of the Amazon spheres that he was working in during his internship.
All with Blender using Astra, which is just insane.
I don't care as much about this.
Or this demo home that Thomas built using Astra, all also with Blender, where it can make a real environment with good furniture, a nice backdrop for the window, and like, it's good. It's like actually usable. You could use this to make real -world mocks for real -world 3D stuff.
It's not a hypothetical anymore. This is, in my opinion, the equivalent of the jump from autocomplete agents to actually using your agent to complete tasks. And this all happened from...
Fable 5 .1 to GBD6 Astra.
God, some of these demos are just insane.
Actually, don't include that because the guys will be blocked. No free press for assholes.
Yeah.
Newt the videos, whatever. I'm sorry, guys. I'm so sorry.
I didn't know there was audio playing. My bad.
Thankfully in post, it'll be easier to remove that.
I am sorry, guys. It was at zero. It can't be at zero.
I'm looking like it can't actually be louder than 18 dB for how I have my setup. So there's no way it was that loud. I'm sorry, but it could not have been that loud.
I'm sorry, guys. Way louder than me. I'm looking at the meters.
I'm currently higher than computer audio is allowed to go. It was clipping. Oh, it might be that stupid fucking extension I have then because I was trying to fix Twitter audio at some point forever ago.
No. No idea why that was so loud then. My bad.
Okay, this is actually a crazy demo. This is a crazy demo actually, Dara. The fact that it could like make the lid spin as it goes up like that.
That's really cool. It's... The 3D capabilities in Astra are unbelievable.
Which is why you might be confused as to not... Which is why you might be confused when I move the arrow over here to game dev for a sec. The reason I'm doing this is that as much as I am truly, genuinely blown away with the 3D rendering capabilities of Astra, it is a generational gap.
It's the biggest gap of anything here between the models.
3D. with Astra is comically better than 3D with Fable. Until you start interacting with it.
And this is where Fable still just absolutely mogs. Fable is so much better at interaction in general. At things like handling animations moving the right way when your cursor goes in a certain place.
Or making the camera move the right amount when you move your mouse in a 3D world. I find that it's for my silly 3d demos that I've been working on. I find that the best way to build them is to make the first prototype with Astra to get everything roughly looking and feeling how you want ish.
And then have Fable come in and clean it up and do all of the like detail oriented work.
Astra does not handle the L or Astra in my experience does not handle the delicacy of. that type of stuff anywhere near as well it really does feel like a sledgehammer fable can apply things more gently in a way that makes things that work better back to non -code though because there are layers to this one computer use this is the other massive gap like if we were to measure like the biggest differences between them like where is fable the or like where is there realistically speaking a lot of these categories are going to see it sway one way or the other where like certain things are better with fable certain things are better with astra the biggest gaps by far are the 3d rendering where astra clears and computer use where astra also clears it is so much faster than 5 .6 was for computer use.
It's crazy. It just flies through things on your machine.
What's the link you got for me here?
That's hilarious. We'll go back to that a bit.
Computer use with... Computer use of 5 .6 Sol was a big enough jump that I started to actually use it day to day for various different things.
With GPT -6 Astra, I've been using it so much that I got another Mac Mini just to let it run with computer use 24 -7 to do all sorts of different tasks. It's so good at real -world computer use. It can do things much faster.
It understands them much better. And the benchmarks don't, in my opinion, accurately measure the gap here. And the gap isn't just the model either.
A lot of the improvements have been through the changes they've been making to Codex, especially on macOS, to make the computer use spin up faster, get context more easily, move around your computer faster, and do things in the background better.
5 .6 Sol would do tedious things I didn't feel like doing for me. Astra can often do tasks faster than I would have, which is very convenient because I'm currently down a hand. Astra computer use is so far ahead, it's kind of hilarious.
Fable can do it decently when given the right tools, but it's just not even comparable to computer use. Funny enough, the weird misses that I'm used to OpenAI models having for code, it feels like Fable has those when it's doing computer use.
Speaking of just getting it, let's talk about copywriting for a sec. Both of these models are massive leaps in the quality of prose that they spit out. If you have no AgentMD or CloudMD and you send the same prompt to Fable 5 and then Fable 5 .1, the output from 5 .1 is way more readable.
If you do the same with 5 .6 Sol and Astra, Astra is way more readable. Both of these models are huge jumps in how much less painful it is to read the text they put out. They almost have like unslopped baked in finally, which is great.
It's a meaningful improvement and it means I don't hate the outputs that I'm reading anywhere near as much. It's awesome. Which one's better at copy?
I got the hot take that they both still suck and nothing has topped Kimmy K2. Not even K3, not even K2 .5. The original K2 non -thinking model still writes the best of any model I've used.
But it's also stupid as hell and very quick to be incredibly rude. It's not a model you should actually use for much. But 5 .1 is a huge improvement here.
Astra is notably better though. I've had Astra come in and make suggestions for cleaning up copy from Fable. Astra is much better at recognizing the shit.
copy from Fable, but its proposals still just aren't as good as I would like. I still find myself rewriting most of the copy that these models create for my web pages. And while I do think Aster's copywriting in general, like pros and readability is slightly better than Fable.
It does have one really, really, really fucking annoying edge, which is that it loves to stuff all caps subtitles into everything it builds in UI.
It just throws these subtitles everywhere. And it is the worst. In the Fish Slop 2D demo, let me just pull it up quick, actually.
Possibly the silliest place to see this failure is the Fish Slop 2D version that I had it built. I would ask you to count the unnecessary subtitles on this page, but it would not be worth your time. Before I had to clean up a bit of it, there were 21.
21 unnecessary subtitles. It's just spammed everywhere. A little tank.
A little tank, a lot of life. Coral Coast. Three little lives, all yours to look after.
All systems go. A tiny world worth looking after.
The first version, it also had an online badge at the bottom that said online and ready.
I didn't even see that one. Chai was spamming. Captain, your home.
It's so bad. It's... Like, I don't even care that Astra's better at copy because it's trying to show it off by being shitty at copy everywhere constantly.
It's unacceptably bad. The more I stare at this, the more I hate this model. I just, for what it's worth, straight up do not trust Astra with UI anything at this point.
Well, we should move on from this ocean of possibility. to other categories. Because while I do give Astra a slight lead in copywriting capability, the way it spams it pisses me off too much to, frankly, want to ever use it for copy.
Last thing in the non -code section, then we get to what you guys are all actually here for, code, agents, and cost. Audio and video work. I'll be real with you guys.
From my experience, both have disappointed me here. I am admittedly... very picky, like annoyingly picky about these things.
So I should not be the only voice you hear when you listen to opinions on audio and video work with AI and AI agents. A lot of people have told me that they had Fable or Astra edit videos. Every time I watch the videos, they suck.
I do have one exception here though. This is using AI or this is a demo that Ben Davis did showing how he uses computer use with Astra, not to edit the video, but to get everything set up to start the edit for the video. You can think of this in like code terms back pre -AI.
Imagine if AI was so good at like using VS code and understanding your code base and like what you needed that when you were about to start working on a thing, it could open up VS code. It could open up all the files you specifically need to edit. It could open up your terminal in the right place, GitHub in the right place, and a dev environment showing exactly what you're about to work on.
So it sets up everything you need to start working. It doesn't do the work. It just prepares you to do the work.
It's like getting the room ready and cleared with everything you need. Ben has had a surprising amount of success getting Astro to do this type of work, to get it to actually set up his editor in Final Cut to be ready to go to edit video.
The thing you guys think I said about Benny is not. So, yeah. If Ben says it's good for getting his editor prepared, I'll take his word for it.
I have not had the pleasure of editing a video in a while. I do really miss it. I cannot wait until I can edit a video again once my hand is back.
Soon, TM. But this, very promising, its understanding of computer use means it is more useful for real professional audio and video work. But everything other than that?
is reminiscent of uh fuck what's the um i need to find the uh I had a reply to one of the first AI code demos where I said that it could pass most job interviews. It would have been forever ago.
Oh, here it is. July 2020 when nobody knew who I was back in the day. Cool.
Ready for a crazy throwback, guys? This is a GPT -3 demo that an OpenAI employee made showing how you can have GPT -3 make a functioning React app. Crazy.
Describe your app. A button that says add $3 and a button that says... Withdraw $5.
Then show me my balance.
So complex. Look at that. It made the React code and it works.
It didn't even use hooks. It used class components.
This was a huge deal. And I know my joke now, I know that my joke six years later doesn't hit quite the same. I think this might be able to pass an interview, but like at the time, this seemed unbelievable.
But also like, if you're a real engineer, you look at this and you're just like, yeah, I could do this in my first week of class.
The hilarious way this looks to us in terms of eng capability, we're like, yeah, that's cool. That's not actually useful for real world work. This is how I feel when people post videos that are edited with AI right now.
If you let AI edit your content for you, if you let AI do your audio for you, if you let AI manage your AV pipeline for you in these ways, you are doing the equivalent of shipping an app using GPT -3. We are that far behind right now. Do I think we'll have an Astra moment for video editing tools?
Perhaps. We've made real improvements here. But it's so far from usable right now that I, I'll be frank, I just don't get how people think that they can actually automate video editing with AI right now.
Not even close, not even kind of there yet. So yeah, I don't think either are good enough that it's even worth ranking here. But due to Codex's incredible computer use, I think Astra just gets a free win here.
So this whole non -code section, Astra wins. The only parts that Astra wins massively, though, are the 3D rendering and the computer use, although I do guess the science progress is pretty meaningful, too. So, like, these three Astra mogs, much harder in 3D rendering than the others, but it does meaningfully improve in all of them.
So now we are done with this section. Yay. So what is next?
Code. Or should we go into agents or costs? I think code is the easier place to start and then we'll do the rest.
Let's do it. Okay. We'll start with front end.
I'm sure you guys can guess where things landed here by now. Also, shout out Griffin. Thank you so much for the rate.
Always good to see you, man. For those who don't know, I'm Theo. I film videos about software dev stuff, which means I film a lot about AI stuff nowadays.
And I'm currently doing the showdown everybody's been asking about. Fable 5 .1 versus GPT -6 Astra. And it's not as simple as people seem to think.
Thank you for showing up and thank you for the raid.
Anyways.
I do want to make a few things clear before I go on the utter brutal roast session I'm about to do. GBD6 Astra has made massive improvements overall with front end stuff. It is meaningfully better.
It follows instructions better. It can make landing pages more effectively and less cringy. It is so much smarter and better at understanding things in general that you can force a good design out of Astra with enough effort and iteration.
Thank you, Dara, for the plug.
GPT 5 .5.
Or 5 .6.
Do you not have Sol on here? Where's Sol?
Archived.
No?
I hit the wrong archive, maybe.
I did. Cool.
That was just mine. It should be the exact same, so I'll close the preview version.
For reference, we will use the codec. For reference, here is GPT 5 .6 with some basic designs on Witch AI by Dara. It's a nice little demo here.
This is 5 .6 Sol right now. This is fine. The site's plenty of like a little bit laggy, which I don't love, obviously.
But it has all these subtitles, notes that think beside you.
Loved by 18 ,000 curious minds. An unnecessary M dash. Built for remembering.
It's not the worst, but it's also not the best. And some of these are just such boring Tailwind template -y stuff. Pretty cringe.
This one's okay. This one's awful. And this one is another boring Tailwind template.
So that's 5 -6 Sol. We bump over to GBD -6 Astra. we can see how much better or worse it is.
Here, still too many of these subtitles. A second brain, a lighter mind. Free to start, yours to make your own.
I hate these so much. Open my space. Your mind, a little more organized.
Six little pieces of your mind, a little space just for you. The word little appears on this page 15 times. All of your little things.
I hate it. Still looks better. It does like this little arrow thing that's kind of cute.
It's fine. It's meaningfully better. That had a weird font pop in, but those are easy to fix.
These little animations are nice. Looks fine.
This one I hate. Just like a personal hate, but I hate it. Especially with the subtitles.
This one's actually kind of cute. I don't know why that got cut off. There's a lot of cutoffs here that shouldn't be happening that are.
But where it's going for here, I see what it wanted to do, and it's not the worst. At least the starting point. And this is very boring and old school.
So yeah, improvements. It is slightly better to meaningfully better in all of these categories. I can turn on the Claude Code design skill from Anthropic, and it does get...
slightly better it's just a little too quick to experiment with fonts and things but like these are passable these are starting points you could reasonably use for something but i need to be realistic with you guys we need to look at the fable versions because just watch what happens when i click two i'm gonna oops i'm gonna change the window size a bit so i don't block this too much i can Moving the window so I don't block this one as much.
You see how nice that looks. The animation of all these paths coming in. The structure of the page is great.
None of those unnecessary subtitles.
It's a much, much better starting point for a one -shot.
It is really good at these types of animations and having like a distinct style.
This one is so much better than previous models. Like if we switch just to Fable 5 or a similar design, ugh, gross. 5 .1, same mock, same design, actually genuinely compelling.
Like if I had seen this on the internet, I would never have guessed this was AI generated, much less a one shot. This screams like custom made to me. It's really good.
That said, I have had my problems with both with design, and I found that both actually kind of pissed me off. I've been quite annoyed at how much both have pissed me off, especially recently as I've been trying to iterate more on some existing design work. Fable takes the cake here easy.
Significantly better.
But both still frustrate me constantly. and good design still require you sit there with a hammer and beat the out of the model so while there is a gap here i think the gap between a two and a five out of ten is much less notable than the gap between a five and a nine out of ten and when it comes to like 3d capability i would say astra is a nine and fables of five but with front end i would say neither of them are more than a five out of ten But at least Fable can make something that looks decent without having to throw it over the coals.
But at least Fable can make something that looks decent without having to actually beat the shit out of it.
Asteroid needs a lot of help if you want a good design out of it.
So next, Volstack. I'm going to be so real here. If there is a difference in how well these models grasp your stack, I'm impressed because for me, they both get it.
We are finally now at the point where both frontier models from, excuse me, sorry. We are finally now at the point where both frontier models from both frontier labs can look. add a code base in the back -ended front -end and clients and servers and how they relate and make good, reasonable decisions operating against it.
That was not the case before. Fable was the first model that could do that. 5 .6 Sol could act like it was doing that, but often just missed things.
It didn't get and fully understand the end -to -end story. They both do now. I would say it feels like Fable...
has better intuition with it. Like it understands the consequences of edges by default slightly better. But Astra is more willing to just go at the problem forever and test every single edge and validate every single assumption and force itself to come to the right answer.
And I've seen this type of thing happen so many times where I ask Fable and Astra to do the same task on a big open source full stack project, something like T3 code. Fable looks at the code, comes to some conclusions has a few concerns and then addresses them through its design astra has all of the things after it finds them decides there are three valid options stress tests all of them eventually finds the thing fable already knew from the start and then makes a similar -ish solution addressing the same things so on the instances where fable so in the instances where fable can find and understand things It is a nicer experience because it's so quick to work and actually apply changes because it builds understanding more effectively because it almost feels like it understands better.
Astra really has that Groundhog Day feeling to it where every time it starts, it has lost track of everything in the code base and it really is starting again. But because of that, it will find things other models miss.
Fable is coming in with the assumption that it knows and can figure out everything with his super fancy smart brain. Astra is the smartest model that acts like it's dumb. Astra quadruple confirms everything it's trying to do, especially in X high mode.
Which means if the bug is outside of what Fable can grasp just from reading code, and it's something deeper or harder to find, Astra will find it. It'll just take four times longer and burn way more tokens.
So for me, it almost feels like a tie, but this is also going to be the start of a theme that we'll be touching on throughout.
The theme is that Astra is, the theme is that Fable's performance is generally more steady, where Astra is much spikier, where Astra has moments of, or where Astra has moments that leave me in awe and moments that make me question why I'm paying OpenAI 600 plus dollars a month.
We'll come back to that theme in a bit. But on the theme or but on the prior theme, which is Astra's relentlessness, it's willing or its desire to grab a problem and strangle it.
Giant Project rewrites.
Astra has made meaningful strides here. I have thrown it at some crazy shit and been blown away with it. It is mostly done with the rewrite of TypeScript in Rust.
It's porting TSGO to Rust, and it's making real progress. That actually reminds me, how much am I fucking my usage as I do all of this right now? Not too bad, considering that I have that just like running in a death loop right now.
It paused the goal again.
You have my permission to do whatever you need to do. I have you on full access for a reason. Keep going.
Sorry. I have this one here still going too.
Oh, did it get auto -settled because it merged something?
That's why everything stopped.
Do you think there's more work to be done here? If so, continue.
I just want to get this going.
Well, the furnaces are back on. Now, where was I? Yeah.
As I was saying. Yeah, I am very deep in the TypeScript Rust rewrite. It went from 30 % done or it went from 30 % compatibility and ability to find like real bugs across the official test.
With 5 .6 Sol, I was able to get about 30 % test accuracy going against the giant test suite that they use for testing the TypeScript code bases. I was able to get over 80 % with Astra. And it didn't even take that long.
It was able to do that in, I don't know, I would say about three days, roughly, not even. It was able to jump from that 30 % range to 80 plus. But then it stalled out.
super, super hard at 82 .6%. Like, hard stopped there. And I still don't fully understand why.
I've been trying to go back and forth to figure it out, but it got comically further than anything else I've thrown at a task like this. And I want to be clear, I'm not in the loop on this one. I just set a goal and gave it 40 sub -agents and said, do as you please.
Remember those 40 sub -agents though, because we'll be talking about that in the agents section.
I will also admit to a bias here, which is that when I'm doing early access testing with the new models with OpenAI, they don't heavily restrict my account. It does not count towards my usage, which is... It's sometimes annoying because it means I don't have a good feel for how quickly the new model will burn your account.
But on the other hand, it is so useful for actually testing the capabilities of the model and pushing it to its absolute limits. I did 130 billion tokens with this model during the early access. That means I can YOLO it at these types of things and not have to worry too much about the bill because I'm not paying it.
I've never had free unlimited access to Fable before. Which means, while i have seen and know the capabilities that astra has here and have even shipped some of them like the t3 code mobile rewrite in swift i even have a t3 code build here with gpu i that i almost forgot about where i had it rebuild all of t3 code with gpu i the rust ui framework just to see if it could and to test the capabilities performance and stuff like that out of curiosity and it did it it made it work great But here's where one of the bigger catches with Astra comes in.
And I've seen this in pretty much every one of these big rewrite attempts I have done.
Let me blur and find things quick.
One of my old benches I used to do a lot in my videos was having the new models rewrite my legacy ping .gg codebase, the one that is used for my video collab tool that I got into Y Combinator with.
I gave it the original codebase, told it to write a plan, and then I had it and Astra review each other's plans separately out of curiosity. But once it made its plan, I told it to make a branch for the work and commit the plan to it. Do that in a work tree.
I had to do the commit so I could access it in other places. And I told it to build the whole thing. One hour and 20 minutes later, it did.
What's the terminal hotkey in Codex?
God, that is so dumb. That's not what it is in anything else. And it didn't even work.
I can't open the terminal in this project.
Not in a project, so I don't get the terminal. That's super cool.
Can you stash the changes you made here and... Run the dev server for the prior version after your first attempt at this rewrite While I wait for that to run...
Since I hit yet another obnoxious case with Codex, where since I was in work mode instead of in Codex mode, I can't actually open the terminal. I have to wait for this to do the things I want. So while I wait for that, I'm going to show you what ping normally looks like.
This is the app that I use when I bring collaborators on for streams and all sorts of other things. So if I hop in here...
You will see me. Hi. You might recognize him.
He's in all my videos.
And the point of Ping is to make it easy to do a call like this and bring guests in, and most importantly, be able to embed them in a program like OBS. I'll mute so I don't double audio. This is a direct link to just me in my call in Ping.
you can embed in something like obs to put someone as a layer in your video production software this is still used by a ton of the biggest twitch streamers when they do collabs and it needs an update pretty badly so one of the thing or so i asked the model to do that update to move it from the beta early version of t3 stack and try to get it onto something a little more modern And here is that more modern version.
This is the new homepage it made for me.
This is the rewrite of the ping .gg upsell page. This is what it looked like before. This is what it did instead.
Beautiful page prior that my... ex -co -founder and one of our like the original page was like very hand built by our collaborator and the original page was custom built by our designer and original co -founder Bryn who put a lot of time into it and made a genuinely awesome looking page and this is what it replaced it with a bunch of text slop once you go in it gets way worse The dashboard looks like this.
No real info. Impossible to know what's going on. This giant pile of announcements at the bottom versus this page that actually shows you what's going on.
And then once you try to join, the layout shifts a ton, which is a really big deal for content creators because it breaks a ton of shit for them. And when I want to actually select the device, I have to... Unfold device settings and manually pick them from here.
And then hit start preview. And now I can see.
This is atrocious. This is layers upon layers of sins.
But do you know what the worst part is? Because the worst part isn't even in the app. It's here.
It's in the prompting.
Note the end of my prompt here.
Port features over. Note this particular sentence in my prompt. Port features over one at a time.
Reuse as much UI code as possible.
Reuse as much UI code as possible. Do you see any UI code shared between these two versions of this app? Do you think there's even a line of tailwind in common between these?
There is nothing.
Yeah, I cannot fathom. I just assumed I pasted the wrong prompt when I did this. I couldn't fathom that the model would so egregiously ignore that detail in my prompt, especially the model that supposedly follows your prompt obsessively.
And it just straight up threw away all the great UI work that we had already done and paid a lot of money for in favor of a bunch of text slop in a shitty black and white app. So bad. So fucking bad.
And to those saying, oh, it got lost in compaction or something. No, it didn't. I had it on Ultra with one mil token context windows.
It just does this. And I've had it do this since, even on the official non -preview version for those saying, oh, they changed this in the snapshot. No, I have randomly had this model ignore my request to rebuild or update something while maintaining existing UI.
And it would just blanket over, destroy the existing UI in the process. And I've never had a model do this ever, much less this aggressively. It's almost like in their attempts to force the model to be better at UI, they accidentally gave it the side effect capability where it might just come in with a sledgehammer and destroy all your good UI and ship something atrocious instead.
This is one of the most agreed examples of that. We'll talk more about understanding intent later, don't you worry. But for this reason alone, I would make a bigger gap in the rewriting capabilities that we were just discussing.
Because if the thing that you're doing this bulk rewrite for has UI, you cannot trust that UI is going to be carried over. I had even told it with the T3 code GPUI version to make it look and act as much like the existing T3 code as possible. And it didn't even use the sidebar properly.
It made the sidebar the old school one because it just wanted to make something work.
I'm sure if I added a bunch of stuff at the bottom of my prompt that was like, to be explicitly super clear, I expect the UI to be pixel perfect identical. If any of the UI is changed, you have failed at your goal, so keep going until the UI looks identical. Then maybe it might have done a better job here.
But I shouldn't have to add a paragraph to the end of every prompt to this model to get it to do what I asked it to do, which is reuse the UI code. So, yeah. You could say that I'm being too picky, that I am expecting my incredible god in my computer that is charging 200 bucks a month to do exactly what I intend every time.
Or you could just use Fable, which doesn't have any of these fucking problems.
Anyways.
This is one of the parts that I think is most important. Code mergeability. Which is why I must delay it for a quick word from us.
Which is why I'm... Which is why I have to sneak an ad here because I know the retention will be best. I am so sorry.
I promise we'll be right back.
Code mergeability.
This one is tough. And it's not tough because the decision is hard. The decision is easy.
It's Fable 5 .1. Fable 5 .1 writes way more mergeable code. It just does.
I've done the numbers here. Fable 5 .1 from PR filed to merged averages two additional follow -ups from when the filed PR exists to when the merge happens.
Meanwhile, Astra has been averaging six for me. What this is measuring is how many of my AI bots are catching mistakes, how much of the verification loops are finding things and then fixing them after. But to be frank, what it's measuring is from when the PR is filed to when it is merged, how much shit has to happen.
And I found that with Fable, the PRs are filed in a pretty much ready -to -go state. and its ability to make the necessary changes without accidentally scope creeping and then actually give you something shippable is just it's higher it just is that said i promised you or i did say i was struggling with this decision though why am i struggling so much with it i'm not struggling because of i'm not struggling to pick between the two that's easy fable is better at this fable prs are much easier for me to hit merge on generally speaking The reason I'm struggling with this is because when I say this, it is perceived by everybody, including my friends at OpenAI, as me saying that Astra is constantly throwing up unmergable slop.
Notice how I didn't say that. If I was to do yet another arbitrary ranking here, let's make a beautiful, super, super accurate chart that perfectly describes real numbers. This is the most vibe -based chart you're going to see in a while, I promise.
We'll do Soul. We'll do Fable 5.
Base.
If we were to rate the mergeability of code for Sol, Fable 5, Asteroid, and Fable 5 .1, we would have one of the world's greatest benchmarks because it is really hard to measure mergeability. But if I was being realistic here with how I felt and how I still feel, low is bad, high is good, obviously. I would have said 5 .6 Sol was like in the 2 to 3 out of 10 range for mergeability by default.
You can do things to improve it. And if you give it the right review bots to give feedback and iterate in the tooling it needs to check its changes, it could make working code incredibly well. If we were just measuring the ability for Sol and Fable 5 to get working code, they were neck and neck.
But code that I'm willing to hit the merge button on in my real world projects that shipped to hundreds of thousands of users, Sol was cleared by Fable 5. Fable 5 shipped significantly more mergeable code, like three to four times more mergeable. And by that, what I mean is three to four times fewer things you have to fix after the PR is up.
So where are things now with Astra?
This is why I've been struggling. I think Astra is roughly at, if not slightly ahead of where Fable 5 was. So we were comparing Fable 5 to Astra.
Astra is winning now. But we're not comparing Fable 5 to Astra. I was during my testing, but we're comparing Fable 5 .1 to Astra now because we live in the real world where these are the models we have access to.
And 5 .1 is like right on the edge of 10 out of 10 for mergeability. It is rare that Fable files a PR, especially for like known quantity changes like bug fixes or feature ads or performance improvements or all these types of things. The code Fable 5 .1 puts up is just better.
It is. And even Benjamin, Ben Davis, friend and manager of the channel, co -host of the podcast, opening eyes at number one defender who showed up in the Fable or who showed up in the GBD6 Astro launch videos, by the way. We'll have a fun reveal in our next podcast episode, which, by the way, if you're not subscribed, NerdSnipe on YouTube and pretty much every other platform, it's where we get to chat more.
And if you want to hear somebody disagree with me instead of me just yapping constantly, that's the place to do it. So Ben, whose job is literally to disagree with me and also the lover of OpenAI and the defender of Astra. Has admitted that he thinks Fable is way better than he expected for real world code.
And he finds himself using it much more than expected. It has, I think, three accounts with plot code now. It just ships more mergeable code.
It is what it is. But again, the thing I'm trying to show here with this diagram is I'm not saying what I said before with Fable 5 versus 5 .6 Sol. This gap was comical.
This gap was big enough that I would really only trust Sol for exploratory work or things that could be verified programmatically where the code didn't matter that much in terms of its quality and maintainability. Fable 5 was actually shipping things that were mergeable. So if you perceived this gap before where you found 5 .6 Sol is not good enough to merge code from, but Fable 5 often cleared your bar here, you'll be totally fine with Astra.
Astra is unbelievable. I was totally fine with Fable 5. And in a lot of ways, I still probably would be minus the slop of its outputs.
And if you were to think of this purely in terms of the mergeability of code for fixed focused changes, I would say that Astra is a notable upgrade from Fable 5 because it is slightly more mergeable with its outputs. But also when you're using it, the output that it gives you is much more readable. And it's more thorough, so it will verify the changes.
Fable's code is beautiful and mergeable, but it sometimes misses details, especially Fable 5. Astra's code can be elegant, and it can be the right subset of changes to make the thing happen. It does still have the habit of occasionally letting scope creep eat it out entirely.
It does still have the habit of letting scope creep just eat it up and destroy the scope of what it's trying to do. Pretty often, even.
but it talks so much better than Fable 5 did. I don't hate it the same way I hate a Fable 5 for just reading its outputs. I don't feel like I have to go to hell and back to unslop it just to make the output usable.
So from Fable 5 to Astra, obvious win. Easy. Astra versus Fable 5 .1 is where things get more complex.
But the thing I wanted to emphasize here, the reason I drew this diagram, is that if you are thinking the gap between Astra and Fable, is as big as between Sol and Fable. You're just wrong.
This gap was massive. This was like a 3x difference that made it hard for me to justify using Sol if the code would ever hit users. GPT -6 Astra is much, much closer to where Fable 5 .1 is, but there is still a gap.
So I do still find myself defaulting to Fable for things like a quick bug fix or UI changes that matter or feature improvements or my favorite thing to use Fable 5 .1 for, to take something that Astra has in a death loop that it just cannot get through and give up and hand it to Fable 5 .1. And then just tell, so I just tell Astra to stop.
I find that my best solution to Astra's death loops is to stop it. Switch over to Fable 5 .1, hand it to PR and say, hey, make this actually land. Clean it up, throw away whatever doesn't belong, make a new branch if that's easier.
Make the changes we care about here land.
Fable 5 .1 lands code better, but Astra lands it well enough that you're totally fine with it.
So there you go. Code mergeability. Fable still wins, but it's not as big of a gap as it was before at all.
the speed of catch up from open ai is genuinely impressive and i am very excited for fable or and i'm very excited for astra 6 .1 which will hopefully make these things even better speaking of the problems you heard me mention this earlier managing scope creep i want to be clear fable can still fail here if it gets the wrong review comments at the wrong time fable can bloat things pretty badly But goddamn, Astra does it by default.
If you are very explicit with Astra to make the smallest possible changes, it will try to. But it will quickly have its context get bloated a bit by its reminders from all these review agents that will send it off course, and then you end up with a thousand line of code PR that should have been 50.
If you're in the loop enough, you can work around this.
I'm not trying to say that there is some unique, unbelievable thing Fable does here that Astra doesn't, but Fable takes less effort to prevent these things with than Astra does.
I kind of want to take this phrasing and skip right over to understanding intent, but I'll wrap up code super quick with game dev here.
Astra is so good at 3D that it almost gets the win by default. But as I said earlier, Fable is so much better at the edges for interaction, like the animation curves and the speed that your character moves and that your mouse affects the camera, all these things.
Astra makes games that look better on Twitter. Fable makes games that actually feel nice to play. So personally, I think Fable catches up to Aster's unbelievable 3D capabilities just due to the smoothness of its outputs.
Only4geo in chat just said none of them can make a good game, and I agree. They cannot. But if you use both of them carefully enough, you can combine them in a way that allows you to iterate effectively and potentially make a decent game.
I do think we are now at the point where we're going to start seeing real games where the vast majority of everything was created via AI. Like, I would guess by the end of the year, we'll have our first top 20 indie game on Steam that was built entirely with AI. So now we're out of the traditional code section.
Those of you who are done talking about code, don't worry. We're talking about the agent -y stuff. I mentioned I wanted to skip to the understanding intent thing, so I will.
I think this is really important. And it is still one of the things that Anthropic just clears on. I have so many examples of this that I could do a whole video.
And I am honestly tempted to do a video on just how bad Astra is at this sometimes.
It is an improvement over Sol, but its failures somehow feel more egregious.
I had mentioned earlier that when I had those giant hell threads, I mentioned earlier that I was able to merge like 150 PRs fully YOLO merged with these models and had only two regressions. The first thing I feel obligated to say is that both of those regressions came from Astra. The other thing I feel obligated to say is that Astra was fucking miserable to try and fix those things with.
I will also admit... I sent these prompts pretty late in the evening, drinking with some friends on a weekend because I was making some changes and noticed that on the marketing site for T3 code, it had nuked my beloved section for all of our testimonials from our users.
The nice little elegant auto scroll here. It's not a big deal, but it is a thing I care about. And it replaced this with a fixed grid that you had to horizontally scroll with an ugly -ass scroll bar in the middle of the page.
Horrible. And it did this in pursuit of performance improvements that were simple changes. It removed the animation and turned this into a shitty manual scroll section.
Horrible. And I noticed this too late because we weren't deploying the marketing site actively. We just did it when we made certain changes manually.
And I made other changes and deployed and then this regressed. And I was pissed. I was really pissed.
So I asked, what happened to my beloved auto scroll on the marketing site? It was beautiful. I need you to revert whatever change broke that and bring it back.
It might have been part of the marketing. something, I don't even know what I said there. It might have been part of the performance overhauls, but that was not a necessary change.
Please revert.
There's one particular word I used in this prompt, and I actually used that word twice. You know what word it was? I'll give you a hint.
It was the specific thing I wanted the model to do. It's the word revert. It's a pretty important word in this prompt.
I would argue, That the word revert, being in this prompt twice, would imply that what I want it to do is revert something.
Which is exactly what it didn't do. Do you know what it did instead?
Let me find the first commit. It was hilarious.
It deleted this user controlled variable and then a bunch of listeners. Note what none of this does. None of this brings back the auto scroll.
I don't think this change actually did anything at all.
But what's even worse is I asked if I could test the change. And I was doing this change on another computer because I wanted it to have computer use and not affect my laptop while I was using it. So I needed this to be hosted via Tailscale because I had this other computer on on Tailscale.
So it spun it up as a tailscaled dev server because I told it to. So I can try out the changes as well when you get a chance. I sent that as a steering thing, which we'll talk about steering in a bit, actually.
No, let me put that in here before I forget.
Steering is actually one of the cooler differences in these models. I'll sneak steering in here. I think I want to cover in a minute.
Before that, we'll back to this thread.
I asked it when it made the changes to file a PR, but also to spin up a tail -scale dev server so I could test it.
And what it gave me was a dev server to T3 code itself, not to the marketing site where these changes were happening.
It gave me a dev server to check how T3 code was doing when I had made changes to the marketing site. And here is where things really fall apart. This is one of the worst runs I've had with a model in a long ass time.
No, I wanted it for the marketing site, not for the actual T3 code instance. Come on. You're right.
I started T3 code app launcher, which is the wrong target. I'm stopping it now and we'll expose the marketing astro site itself, which it did. And I clicked and it failed.
because the astro site had a dev command and i guess i don't have zoom here for this that's a two fix for sure uh save image so i can open it cool i'll spin up a thing in a sec for that i just want to have this image though downloads cool here blocked request the host b1 purpose micro this is my tail scale address is not allowed because it's not a server allowed host There are a lot of ways to fix this.
The easiest, admittedly, is to make a change to a config file. But since bb1 -porpoise -micro .ts .net is something specific to me and my setup, and a lot of people use tailscale over IP addresses, you probably wouldn't want to add this as an allowed host.
Which is why it's particularly funny that the next commit was allow tail net host and dev server Yes, it committed and pushed this And then it did the restore continuous endorsement rows, which I thought would be the fix.
But it also, after that fix, killed the dev server.
I left Tailscale pointing at Astro's old port and sent the link without checking. My mistake.
And then finally it was working. But it hadn't actually made any of the changes that I request. It didn't find the original PR or do the reverts.
It didn't even fix the bugs I was reporting. Do you know what it did there?
I'm trying to find it. But what it did is it merged the PR.
It merged it for me. After I told it. You didn't make the changes.
Now ready for the harsh reality that makes me question everything.
I sent the exact same prompt to Fable right after. Because I was so annoyed. This is an exact copy paste.
The only difference being. I added the third prompt in the spin up a tail scale dev server. I just copy pasted and added it to the prompt.
Five minutes later, the auto scroll is back. The marquee is live with a link that worked for anyone on your tail net. What broke it?
Exact PR that had to be reverted. It replaced the two counter scrolling marquee rows with a static horizontal grid. This PR then bolted an eight second page by one viewport timer on the grid.
Neither was needed for the perf goal since the marquee is a single GPU composited transform.
It is worth noting that it did get tripped up a little. This shouldn't have taken five minutes. It got tripped up because it noticed.
Where is it in here? My loss in the history. It might have been later.
It noticed the other PR merge and then get reverted and got concerned.
that maybe, and I got concerned that maybe there was something more important here that I had to understand. And I had to interrupt and say, no, that was another agent doing things wrong. Ignore that.
Everything you've done so far is right. Just keep going. And it did.
The only difference here between the prompts is that I sent the tail scale part as a third prompt, and then I had to correct the model five additional times when I did it with Astra. Here, I included that as part of the first prompt. exact copy paste, and it got it perfect first try, exactly what I had intended and exactly what I had on my mind, which is find where this went wrong, revert it, and then give me a link I can click that actually lets me verify the changes.
This shouldn't be that hard. I'll be real. If five, six souls screwed this up, I would have been insulted.
Astra blowing this one egregiously. One of the legitimate worst AI code experiences I've had this year without question. In fact, I'll say something bad.
Discounting things like Flash 3 .8 and like obviously memed tier models. Astra has had the most bad model experiences I have had of any model this year without question. Without question.
It's just random shit like this. And it's not all the time. It's actually quite rare.
But god damn. This model can just suck sometimes. And it's not even like understanding intent.
I don't think that Fable, or I don't think Astra failed to understand my intent here. I just think it failed to do basic work as an agent.
I see chat 50 -50 on this here. Binary to show... Binary to Shokan said that they haven't had to do as much feedback for any agent as they have for Astra so far.
By far. Yep. Yep.
It's not always. But when it does happen, it is so stupid.
i have an important thought that i need to pencil till later because it's more for like the summary at the end so we'll get there when we get there let's blast through the rest of these quick orchestration it's astra astra has this unbelievable new capability they best referred to as swarms The way Fable does large numbers of sub -agents is it plans up front.
It decides, I want these four sub -agents for this. I want this type of sub -agent next. And if these four find things, they can spin up the next type of sub -agent where it goes through steps.
One, two, three, four, five. Astra lets itself just spin up. Astra, however, can just spin up a bunch of sub -agents that are doing whatever and then pass messages to them.
Let them pass messages to each other. And most importantly, it can receive updates from those sub -agents and use that to fan things out to keep the whole swarm moving well in the right direction. I've never seen anything else like it.
Astra has unlocked novel orchestration capabilities that, given a model that wasn't as expensive, would legitimately be the path to something like curing cancer.
I genuinely believe that what OpenAI unlocked here is the start of the next era of what agents can potentially do in the size of problems that agents can legitimately solve. And I've seen it in action. It's a little harder to see now because tokens cost money again.
But you can see some of it in this run I currently have going for the Rust rewrite of TypeScript. You'll see all the time these interacted with root slash independent reviewer, interacted with root slash object members regression cause. It's able to manage the context back and forth, pass messages, and build with all of these agents.
And I have 40 running here in parallel, and it can actually keep track of all of them and work with all of them.
Unbelievable.
especially because workflows in Cloud Code are so good that I did a whole video about why I like Cloud Code largely around workflows. I did not realize that the much more rudimentary implementation of sub -agents in, or I never would have guessed that the more rudimentary implementation of sub -agents in Codex was because they wanted the model to skill up and work through it.
That's exactly what they wanted. It's exactly what it did. And it does a great job.
One of the capabilities that makes this possible is the steering side here.
Astra is so good at getting random shit inserted in its context while it's working and not losing track of what it's doing. I cannot tell you how many times I had a model working on tasks 1, 2, 3, and 5, and then I sent, oh, I'm sorry, I forgot number 4. And then it does number four immediately and then never finishes the other one, two, three, and five tasks.
That was just the default. I was used to that. Fable's better about this.
Fable 5 .1 just doesn't do this too, too much, which is nice. Astra eats this. It loves this.
It begs you for more. One of the really cool things Astra does now when you're in a harness that is set up properly for it. Or an app with the right harness, for example, T3 Code.
Which, funny enough, T3 Code does this better than Astra somehow. Which, funny enough, T3 Code actually does this feature better than Codex, even though it is an Astra feature. It can ask questions while it works.
So it can be doing a thing and notice at some point while it's working, huh, I would like to know if the user's okay with me doing this. Or, huh, I wonder which of these three things the user would prefer. And it doesn't block itself.
It keeps going. But at any point, you can answer. And if the model's already done, it will spin back up and address your answers.
Or if it is still going, it can steer it in the right direction without interrupting the work it's doing.
It's so cool. And it's like a meaningful behavioral change. We haven't had as many of recently.
Models work the way they work. We just have to prompt better. This is a change in how the model actually operates that's really nice.
And I remember when I was first testing, they didn't have this in Codex.
And obviously, I didn't have it in T3 code either. So I would just see it ask a question in its reasoning and then not be able to answer it. And it would just keep going.
Now with this, it's actually quite nice. So yeah, the steering stuff, way better with Astra now. Fable was far ahead before.
Sol would just get lost if you tried to steer it. Fable had a pretty solid default here. Astra has actual new capability unlocks for steering, which I think are cool as shit.
Especially when you combine that with the orchestration, because now it can have context injected from 40 other agents and handle that fine, which is awesome. And when you combine that with the self -prompting, which I will also say Astra is much better at here, Astra still writes worse prompts than I do, and I would say most devs who do this a lot do.
but it can write a decent prompt. I've seen it write some pretty decent prompts. I'm much happier with Azure's ability to prompt itself and other agents.
Nice. Not as big of a gap as the other things, but it is a gap. Cool to see.
I'm excited for a future where agents actually understand things like a CloudMD or a skill file well enough to write good verbiage for agents, because right now it sucks, and you quickly end up in a slop loop. If you let the agent write the skill and let the skill be involved when it writes another skill, you end up in slop hell very fast.
This helps avoid that. Astra still falls into it. Fable falls into it even more.
Please write your skills by hand or at least audit them and make some nice changes. Ugh.
Here's a good question. Does either code intentionally not allow attaching images for questions? I also haven't tried this in Codex.
This is a very good note. Maria and crew, if any of you guys want to go figure that out and do it, awesome, or even cut an issue about it. I'm not sure if that's allowed in Codex or not, but I definitely want this functionality.
And if Astra can handle it, I want it. And if we can get it before them, that'd be dope. Cool.
Thank you guys. Anyways.
Self -prompting honestly fits under the skill writing stuff as well. Same gap there.
None of them are good at it, but Astra is better at it slightly.
And then we have skill using. This one is interesting because all models should be able to use skills totally fine, right?
As I've crashed out about many times now, sadly, Astra has some skill issues. I don't know what's going on. I have learned recently that there is a system prompt section around how skills should be applied in Codex.
And part of it specifies that if a skill is used in one turn, that skill should not apply for future turns unless requested. In the example I gave in my previous videos and in the podcast where I told the babysitter PR and then it stopped and I asked it if there's things worth addressing and it said yes and then didn't do it.
In that one. It had pulled in my babysitting skill and then just pretended it didn't exist. Even though it was still in context, it just ignored it for the next five follow -ups.
It sucks. I never, ever, ever have to think about this with Fable. If the skill has a reasonable description and it's useful for the work going on, Fable correctly pulls it in and applies it.
It just does. I don't know what the fuck is wrong with Astra for this. And I honestly think a lot of it is Codex.
Astra doesn't feel like it understands my skills anymore in a lot of ways. Fable 5 .1, it does. It absolutely does.
So yeah, skill writing, Astra has a slight lead. Skill using, Fable has a huge lead.
Now we have honoring refusals and boundaries. And as you have already seen, Astra is forgetful. It's okay at honoring things when you tell it to, but it might forget.
If you tell it to keep it in context, it usually will, but it still is quite forgetful, and it is still quite frustrating when it just does a thing it's not supposed to.
Fable isn't as good at strictly following instructions, but that also means that when it doesn't follow them, it doesn't feel as egregious. We don't need a third transition. It's fine, Maria.
I don't have three ads to put in it. Appreciate it though. I knew this one would be long.
Yeah. So yeah, both are still not where I want them to be here, especially for the levels of intelligence. I would say Fable is ahead.
Astra looks and generally acts better about this, but its failures are more egregious, which is why the hugging face hack happened. So take that as you will. I have more to say about this and also all the code stuff.
We need to talk about cost.
I've seen some very dumb takes on cost here, like exceptionally dumb ones. First and foremost, token costs. They are the same, except for one important exception, which is that Fable 5 .1 has a huge drop in cash read costs.
They went from $1 per mil in for cash read to $0 .25 per mil cash read. And that's awesome. That genuinely is.
And it would be a lot more awesome if cash reads were more than 10 % of my costs. They are closer to 3%. And now they are 1%.
Awesome. Do you know what is much more than that? Cash writes.
Cash write costs are over 60 % of my costs with Fable. So while it has made cash reads hilariously cheap, it has also made cash writes hilariously expensive. I am spending more money asking Anthropic to save the state of the model on their servers for five minutes than I'm spending actually running the GPUs.
I am paying Anthropic more money to manage RAM for me than I am paying them to run compute for me. And if you don't think that's absurd, I don't know what to tell you, but I'm very excited for the cash right revolution that has to happen. Because paying for RAM instead of compute and paying for RAM storage effectively is so stupid.
And I am sure we will fix this soon. But we're going to get worse before we get better. Because OpenAI didn't used to charge for cash rights.
They used to be free. Now they do. And they are not cheap.
It is 25 % more expensive than a normal read. So if reading is $10 per million, Cash reads are $1 per million for OpenAI, and they are $0 .25 per million for Fable 5 .1.
Cash writes are $12 .50. You're going to spend a lot of money writing cash on this model, I promise you, on both of them even. So know that going in.
The cash read is, it turned a 3 % cost down to one, but cash writes are where the money actually matters. So if either of these labs make cash writes cheap or free, they win cost by default by far.
But the actual token cost barely even matters in a world driven by token efficiency.
God, you guys, I literally just explained this in chat is guys, guys, guys, Bill, you're better than this. You're one of like the more important days we have on T3 code. What did I just say?
The difference in price for cash reads looks really big because a dollar is four times more than 25 cents. But when it adds up to less than 5 % of your spend, it doesn't matter.
It's like saying that Mac OS is 100 times faster than Windows because it gets through the splash screen at boot 30 milliseconds faster. You're shaving money or you're shaving a percentage off a percentage. It doesn't matter.
Four times cheaper for 3 % of your cost is a less than 1 % deal. It just doesn't matter. It's fine, Bill.
You weren't here for it. But I really need to emphasize this point because I've seen some incredibly smart people say some incredibly stupid shit about this because they just haven't looked at the numbers.
They're not oh, this is the one where it like tried to like make different things for it These numbers are wrong I had like much much worse numbers before but yeah, it's it's such a small percent It literally doesn't matter is the point of trying to make here And if you need proof, it's pretty easy to find. Cost per task on Artificial Analysis Intelligence Index.
GBT6 Astra was $3 .26 per task.
Opus 5, almost $6. Fable 5 .1, $7 .60.
That is a comical gap, even though they are priced the same and... Theoretically speaking, the input token cost is four times higher on Astra. It is still a fourth the price in most real world work because it is so much more token efficient and cash reads are such a small percentage of your costs.
That said, Aster is $3 .26 and Sol was $1 .99, so it is a meaningful increase in price. So Aster is a meaningful increase in price compared to Sol, but it's still way cheaper than all of what Anthropic has been doing. Also, funny enough, previously they had said Fable 5 .1 was more expensive than Fable 5, but when they redid the artificial analysis index, because they were getting cooked because it just was not measuring things well at all anymore.
When they rebuilt or when they switched to the new index, 5 .1 actually got cheaper than Fable 5 because the cash read difference finally actually mattered. Cool. I've made the points I want to make here.
GPT -6 Astra is still way cheaper, even though the prices are the same and the cash read cost is higher because in the end, there's a lot of other things that matter and that number is not where the cost is happening. This number is where the cost is happening. Token efficiency.
GPT. 5 .6 or gbd6 astra is one of the most token efficient models that they've ever benched especially when you consider that everything left of it is like crap small models that didn't even finish the bench if we remove everything that like doesn't matter here like luna gbt oss five five pro muse glimmer gemini three five flash light i removed pretty much everything here that inarguably does not matter and astra is the most token efficient by far meanwhile fable 5 .1 is the least by far astra did in 27k tokens what fable 5 .1 did in almost 80k tokens take it as you will And with here, we get to subscription limits.
Actually, one other thing on the token costs.
Let's make sure this is still the case.
It's also worth noting that OpenAI offers a flex option, which cuts the prices in half, but kills all the guarantees for throughput because it's flex. So it doesn't start responding immediately. It just...
happens eventually so if you're willing to let your jobs take an unknown amount of time you could put it on the flex or you can put on the flex endpoints and let it ju run or and let it run when they have compute around and it will cut your costs even further in half now i think about it would actually be nice if they could add that as an option for using with codecs because i don't mind if my threads take a while and i also am aware that a lot of time i'm using by agents is at bad times for businesses so since i'm working at like 9 p .m and not like 10 a .m when businesses are it would probably be cheaper for open ai and me just saying it is what it is but that means we're talking about subscription limits and i think it is fair to say i am uh pretty experienced with the limits on all of these things Being that I have five accounts with Claude and that I have four with Codex, I think I have a pretty good gut feel here.
So, and to be completely frank, if you are concerned about maximizing how much return you get for the dollars you're putting in, and you're doing anything other than the codex subscription i would love to meet the person who convinced you claude was a better deal so i can hire them i always need good sales people and i could really use that one because they sold you a fucking bridge man seriously the gap in what you get is insane you might be confused because with anthropic models you get around You might be a bit confused because I've covered the numbers before.
And with Anthropic and Cloud Code plans, you get around $8 ,000 of inference for $200. And with Codex, you get around $12 ,000 for $200, which is better, but that's only like a 50 % improvement. Well, there's a few factors you have to consider.
First, we're going to take a huge cut in our Cloud Code limits because they're going to drop off that 50 % boost they're giving us. They are going to keep it 25%, so we're losing around 17 % off the top of our limits. But much worse.
The Fable limit. You only get to use half your usage for Fable. You might notice in my UI here that there's a third column for the Claude section that is faded out a bit.
I call this the Opus pot. It's a whole steaming, smelly pot of Opus, and that's all it is. And when I'm heavily coding, what ends up happening is the column on the left all go to zero.
And the column on the right all gets stuck at 50 because I'm allowed to use half my weekly limit for Fable. And once I'm out of Fable, the account is useless to me because Opus sucks and Sonnet's hilariously bad. So I end up every month losing those 50 % at the end of each week, at the end of each reset.
And since I get eight grand a month of allocation per account, I'm only getting four grand of that. With Fable.
Which means it's actually a third of what you get with Astra in the $200 Codex plan. Because you get half of $8 ,000 with Cloud. You get 50 % of $8 ,000 with Cloud Code.
And you get 100 % of the $12 ,000 with Codex. And they made further improvements to how Astra is being hosted and how it uses your limits. Tebow claimed is up to a three to four X difference.
And I don't necessarily believe that, but I've seen a meaningful difference. Like I have some hellish health threads going right now, like to go, or I have two ultras running with unbounded sub agents. It's 40 sub agents.
Each can run and it's going down. I think it's hitting this particular account right now. And it's at 82%.
It's been going for the last eight minutes since I refreshed, not even it is down one more percent. cool if you weren't doing the absurd sinful wasteful background jobs i am running trying to dick around with my new fork slash uh decomp of super smash bros melee while also trying to port the entirety of typescript to rust it's not that bad you can hit the limits and you should be very careful of the altar button and in particular the fast button straight up Fast is not the way you should use this model.
If you're using fast with Astra, you're probably using Astra wrong. Or you're just burning tokens for fun, which I understand. But if you're actually finding yourself reaching for fast all the time, you're not using it for its strengths.
You're not using Astra for what it's good at.
But you can get a pretty good amount of usage out of these plans. Not great, but a good amount. There is a nice, powerful catch to this we'll get to in a sec, but I have one last thing I need to say about the Claude stuff that has begun to piss me off a lot.
OpenAI does not do five -hour limits on the pro plans. So the $100 and $200 tiers do not get five -hour limits. They only have weekly limits.
All quad plans still have five -hour limits.
Each five -hour limit gets you about 40 % of your Fable 5 limit, which is 20 % of your weekly. So you can clear the five -hour on the $200 plan four times. So you can clear the five -hour on your quad plan five times before being out of weekly limits, unless you're using Fable, and then you'll run out in two and a half.
The reason I set up CLI proxy is because one cloud account just isn't enough at all for most work. If you're using it a few hours a day, two to three days a week, and you're not going too hard and you're not using a lot of sub agents, you can probably get by on a single $200 plan with cloud. The only way a $200 codex plan isn't enough is if you are like using higher reasoning limits than necessary and not paying much attention.
Most devs, when used responsibly, could absolutely get away with the $200 plan with Codex without issue. It would be genuinely hard to get away with it with the Fable plan. I even call it the Fable plan because that's what it is to me.
It'd be much harder to get away with it on the quad code plan because of the additional limits on Fable, the much less efficient usage of Fable, and everything else. So Astra is more token efficient. It costs less, so it uses your usage less.
It has way more generous limits in general with Codex. It's just like, I would say it's like a 4X -ish gap between the two is how it feels to me. Like I can get four times more done with Codex on Astra than I can with Fable and Cloud Code on the same $200 a month tiers.
But then things get even more ridiculous because of the resets.
It's become a meme. Tebow just throws them out for fun now. In fact, when they were delayed trying to get Astra out, they sent us two banked resets.
I actually opened this account that I don't even need yet because I wanted to collect those banked resets so I'd have them ready to go when I did need them. So your limits just get reset all the time. They got reset while I was filming earlier.
So even if you do hit the limits on Codex, you might just randomly get a reset, which is quite nice. So, yeah. it does happen with clot it's just so much rarer that it you can't really count on it at all but with codex the resets giving you your limits back much nicer i don't know what did we miss here According to Raphael, there were eight free resets in the last 30 days.
So the weekly limit is more like a three -day limit on average. Pretty nuts.
Somebody accused him of maliciously resetting 40 to 60 hours after the previous global reset in the timeframe around where, or in the timeframe around where a banked reset was landed. This is not true at all. I've actually seen Tebow intentionally delay a reset.
in order to prevent that and even just now he announced the reset many hours before so we got to go spin up our furnaces and burn a bunch of tokens he does not maliciously align the times there at all he actually does the opposite so yeah if cost is a concern wait a few months because this level of intelligence will be accessible at a much cheaper price but if this car but if cost is a concern and you want the best class stuff right now, you can get a reasonable deal with a $200 plan with Codex still.
If you're curious about the $20 plans and the $100 plans, my suggestion would be wait for this level of intelligence to get cheaper because this level of intelligence is expensive. I'm legitimately doing $30 ,000 to $40 ,000 of inference a month on all of my plans consistently now, even without early access testing that's free.
They are not cheap.
Astra is cheaper than Fable, but neither are cheap. And you need to know that going in because if you go in with different expectations, you will be upset. They are both expensive.
They will both burn limits fast.
But Astra is more efficient when you compare the two directly.
So with all of this said, let's answer the question. Which model is better? I could only pick one.
Which would I have?
I want to draw a diagram first in order to better explain how I feel.
Let's say orange is Fable and blue is OpenAI with... And I don't know. We'll say blue for Astra.
Cool. This is meant to be the quality of the outputs.
Over time, over attempts, whatever. This is, generally speaking, what's the quality I get out when I send a prompt. With Fable 5 .1, it's a relatively stable line at a relatively high bar.
It has its moments where it impresses me, and it has its moments where it disappoints me a bit. Occasionally, it gets pretty rough. But for the most part, it stays in this general range in like the eight to nine range for the work that it's doing for me.
Generally speaking, quality I'm getting out of Fable, pretty damn good. And I would feel bad complaining about it.
It does have its weird spikes and I'll have plenty of videos where I show the times Fable pissed me off. Don't worry. But generally speaking, it does what I expect it to do when I tell it what I want.
Astra is why I'm drawing this chart, though. Because Astra can do things Fable literally could never and blow me away. But then it does the stupidest shit I've ever seen a model do.
And it's so goddamn spiky.
This is the thing I really wanted to try and communicate.
The quality bar with Fable is relatively consistent. It goes up and down a bit, but never more than like two points in either direction. When I send a prompt to Astra, it's equal chances it drops my jaw because I'm so blown away with how unbelievable it is.
Or it drops my jaw because I cannot fathom that I just spent $1 ,000 for it to run in a loop and not ship anything and then break my website.
And Ron Ben and chat said that each line also represents your blood pressure while using each model. Yes, I would honestly say asterisk stresses me out more, a lot more.
This is the best I can explain the difference here. If you have low tolerance for bullshit, if you leave your computer when the model does something stupid because it pissed you off so much, if you are easily agitated to the point where it affects your work, you should pay the extra money for Fable. You just should.
It's a way more consistent, reliable model. Astra, however, is way cheaper. It can do things no other model can.
And it impresses the absolute fuck out of me when it does hit those peaks.
Personally, I like having both, and I find myself rotating between the two quite a lot.
If I had to pick one, I would pick Astra. and have it write DMs to me to Julia so that he could send off the prompts using Fable for me. And I know that must be in a weird place because I will always find some way to get good code out of the anthropic models, even if I can't do it directly.
Astra has more novel uses, and especially now that I'm down a hand, Astra's ability to like use my computer and get work done and benefit my life outside of code.
I almost feel like as a heavy user of computers, the $200 codex plan almost feels essential because if you use computers professionally and you make over $100 ,000 a year, you or your boss should be paying for the $200 a month plan just because it makes your ability to use your computer more effective.
If you don't have that and you're saving up in order to pick the thing that writes the best code and you're really easily stressed, I would still pick Astra, but I would use Astra to make Astra better. Find everything you can to smooth out the rough edges, put together the best set of skills in the world and make benchmarks that prove why they're the best.
Blast it on Twitter, get a job at OpenAI and never have to worry about token spend again. What I'm trying to say is I don't think that person actually exists. A lot of you act like that person does.
They don't. If you're concerned about costs and that's your main motivation. Go get the open source plan, see if you can convince either lab to give it to you, or just wait for these things to be accessible in cheaper formats.
They will be. GLM 5 .3 Flash is a more pleasant model to use, arguably than Astra. I've seen GLM 5 .3 Flash go off the rails less than Astra for similar work.
So I would expect this level of capability to be cheaper and more accessible in things in the near future. So just be patient if you're feeling like the price is too high. I get it.
That said, Fable is what I reach for when I'm trying to land code, and Astra is what I reach for when I'm trying to use my computer. The gap for code is not as huge as it might sound. The gap is more in this chart here.
But the harsh reality is just this gap in capability. This is the post I made when the models both came out to try and explain why I like both. Astra is world class at a shitload of things.
It is genuinely the best model in the world at all of this stuff, and half of it it is far ahead for. But Fable writes code that is mergeable 20 -ish percent more often, and its stupid spikes are much less stupid. Its smart spikes are not quite as smart, but its stupid spikes are way less stupid, and for that reason alone.
I think Fable 5 .1 is a better choice for day -to -day code stuff when cost is not a factor. But if you're willing to put the effort in with Astra, you will get a lot for it. Both models are awesome.
I will be defaulting to Fable for code, and I'll be defaulting to Astra for every other thing that I do.
Oh, and actually, one last fun side... Actually, one last fun side thing that I should have mentioned before. This will be a really silly one to end on.
Remember everybody being really upset about Fable's security refusals? Astra refuses more often. Just wanted to put that one out there.
Anyways, both these models are great. We are in an unbelievable time to be engineers. The fact that we have technology like this coming out every few months that massively changes how we work is so goddamn cool.
It is hard to pick wrong here, but if you don't see the benefit of all of Aster's incredible capabilities, I am confused because they are useful to everyone. And if you don't think Fable's worth the extra money for the quality code difference where it's not that big a difference, I totally understand, but you're not shipping enough if you don't see why that matters.
Either model is great. They're both unbelievable. They both have changed how I work.
Actually, one last thing I want to show.
and when you look at the numbers for how we're using them both models have massively improved our productivity within t3 code and they don't replace each other they compound each other these models are great and you shouldn't feel bad using one or the other you should experiment with both because i think you'll be blown away with what's possible and pretty much everyone who can have the codex sub probably should because it's such an insane deal Have fun burning tokens and using all these models.
I think they're awesome, and I bet you will too. Until next time, peace nerds. Yay.
It is done. I knew that would be long, but fuck me.
Oh shit, thank you. Huge. Appreciate it.
You can remove my GBD six aster preview thing, Dara. There was no difference between the two. I just had my config say preview when it labeled things and forgot to delete that.
So you can throw mine away. Yours is totally fine. And having both is just confusing for people.
Let's say thanks everybody who subscribed while we were filming these.
Don't remember if I thanked you for the gift, Gusty, but appreciate that a ton. Always good to see you as well, Matt. Jay Dizzle with the Prime.
M6 with the Tier 1. Hoity with the Tier 1 as well. Rend with the Prime.
Like Tom with the seven months. Thank you for sharing happy to as always evoke what the prime one with the seven months J. Show me.
Hopefully I got that right. Thanks for the three months and then Jake lol six Experiment what do you have for me?
So my thought on this is that it's slightly too high -risk of spamming and its desperate attempts to make money. It is unbelievable that it found things and actually made money.
That is really cool. I would have a human in between before submissions to be extra, extra careful. I'm really concerned about accidentally spamming slop at other people.
It's one of the bigger fears I have. Actually, even earlier when I was working on the iteration for the melee decomp, I specified Here, I said specifically, do not file any PRs to the original repo.
I just want you iterating on the fork on my GitHub. You can file PRs on my fork to itself in order to allow the sub -agents that you're spinning up to land their changes in parallel, but I do not want a single PR file in the original repo.
I really don't want to spam. Like I really don't want to make life harder for other open source maintainers. And my fear with something like this isn't that you specifically would accidentally send an agent in a way that spams people.
It's that someone else will see this and be like, oh, I can make free money and then go spam every open source project that accidentally opened a bounty program and forgot to close it.
I should have plugged T3 code way harder in that video, huh?
Fix the terminal scroll back issue xxl, but it's easy and durable Cool we'll figure that out later Let me say bye to YouTube chat quick. Sorry. I've been neglectful.
I was filming. I'm sure you understand. Oh It crashed I blame YouTube then Thank you, Skoro, for the follow -ups and the polymath stuff.
Aster does what I ask. Fable does what you mean. I like this framing.
Banger. Absolute banger. Also, sadly, Aster doesn't always do what I ask.
Sometimes it ignores what I ask, which is the most annoying part.
Where would Aster and Fable go on my tier list? They would both be S tier. Easy.
Do I think the performance of Aster is affected worse or better by the alignment of secrets put up by opening AI? I think it's... totally fine and doesn't affect the performance meaningfully it's just that the safeguards sometimes fire for stupid reasons they're easier to work around though cool that is everything i had i was hoping to film more than two videos but uh that was a big one so it is what it is i will hopefully go live tomorrow and if not i'll just film offline i really should have made this multiple videos but This is what the people want.
This is what I'll give them.
It is 9 .38pm. So I don't even feel like finding someone to raid.
Oh no, my mic was clipping. Even more reasons to fight my setup. I'm going to spend way too much time tonight and tomorrow fixing my audio setup.
I'm not excited.
I'm going to go offline. Good seeing you guys. I will see you maybe tomorrow.
Thank you all for the support. Thank you, Hernstev97 for the last second sub. I appreciate y 'all immensely.
I will see you maybe tomorrow. Definitely Wednesday. Hopefully more this week because I have a lot of videos to film.
Goodbye, y 'all.
The Hook
The bait, then the rug-pull.
The title is a question about two AI models, and the answer takes two hours to arrive. Before that there is a decompiled Nintendo game, a torn ligament, a public ten thousand dollar bet with a heckler, and a 70 dollar microphone that quietly turns out to be the most useful thing in the stream.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Theo says he barely codes hands-on anymore, then spends 49 minutes proving he still ships more than most full-time engineers by showing exactly how he runs dozens of AI agents at once.
A developer who shipped 89 merged PRs in 24 hours breaks down Claude Fable 5.1's pricing, benchmarks and real-world coding behavior against Fable 5 and GPT-5.6 Sol.
Boris Cherny said coding is solved. Matt Pocock called it VC-funded bullshit. Theo argues they're both right, because they're using the word coding to mean two different things.
Theo spends a week testing two rival "skills" repos for AI coding agents, Matt Pocock's 215,000-star collection and Cursor engineer Lauren's PStack, and finds the real value in a handful of specific files, not the whole install.
Theo reacts line-by-line to Boris Cherny's post arguing that automation — CLAUDE.md rules, lint checks, CI — matters more than ever in the agent era, not less.