A real-time test of OpenAI's priciest inference tier: $600 to review two pull requests, and the one workflow that might justify it anyway.
Posted
2 days ago
Duration
Format
Review
sincere
Views
82.2K
1.5K likes
57 · 43
Big Idea
The argument in one line.
OpenAI's Ultrafast inference tier for GPT-6 Astra costs six to ten times standard pricing, and the only workflow that justifies it is staying in an uninterrupted, real-time feedback loop rather than saving money.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You're using OpenAI's Codex or API products and want to know whether the $500/month Ultrafast tier is worth paying for.
You want to understand how raw token-generation speed changes the way you prompt and iterate, not just how it changes the price.
You're comparing GPT-6 Astra against other coding models on real cost-per-task before picking a default.
SKIP IF…
You don't use OpenAI's Codex or API products at all; the specific pricing here doesn't transfer to other providers.
You want a step-by-step build tutorial; this is a cost and workflow review, not a how-to for the apps shown.
TL;DR
The full version, fast.
OpenAI's Ultrafast speed tier for GPT-6 Astra prices output at $300 per million tokens, six times the standard rate, and $450 per million on long-context work. Reviewing two small pull requests cost $600, and a single live-build session hit $306 in one thread. The $500/month plan's weekly allowance can burn down in about two hours of real use, and even OpenAI's own employees reportedly need approval to use it. The honest math says it almost never pays for itself. The real argument for it is that staying in an unbroken, real-time back-and-forth with the model produces work you wouldn't have finished, or wouldn't have designed the same way, if you'd had to wait minutes between every change.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
A sponsored segment for Greptile's T-Rex, an AI review agent that runs code on a real machine instead of just reading it.
03:11 – 04:22
04 · The reveal: $600 to review two small pull requests
Two PR reviews under 100 lines each burned $600 in Ultrafast usage, and the token math shows why the multiplier is so brutal.
04:22 – 05:31
05 · Mapping the $500/month plan's burn rate
The $500 Ultrafast plan grants about $1,000 of weekly usage, which disappears in roughly two hours of real use.
05:31 – 09:48
06 · Live demo: building a kanban app in real time
A toy kanban-and-chat app gets rebuilt live on screen, prompt by prompt, with changes appearing in seconds instead of minutes.
09:48 – 12:02
07 · Watching the usage meter drain live
A handful of small prompts knock the $500 plan's weekly balance down several points in minutes, and OpenAI reportedly throttles even its own employees on Ultrafast.
12:02 – 21:27
08 · Building Slopalytics: Astra Ultrafast vs a slower model
A real AI-model benchmark dashboard gets built from scratch, comparing a 90-second Astra Ultrafast draft against a 15-20 minute draft from a slower, higher-quality model, then iterated in a fast back-and-forth loop.
21:27 – 24:39
09 · The real case for Ultrafast: flow state, not savings
On a pure cost basis Ultrafast makes no sense, but staying in the loop instead of losing context between steps is what actually got the project finished.
24:39 – 27:53
10 · Verdict: use it only when you don't care about losing it
The final take: don't upgrade to the $500 tier for Ultrafast alone, reserve it for throwaway or truly urgent work, and watch for a cheaper model to get the same speed option.
27:53 – 28:17
11 · Sign-off
A quick tease for a follow-up video on whether the $200/month plan is worth using.
Atomic Insights
Lines worth screenshotting.
OpenAI's Ultrafast mode for GPT-6 Astra charges $300 per million output tokens, six times the standard $50, and jumps to $450 per million on long-context runs.
Reviewing two pull requests under 100 lines each cost $600 in Ultrafast usage, roughly $300 per review.
The same review work priced at a comparable model's standard speed would have cost about $1.10, a roughly 40x difference.
The $500/month Ultrafast plan grants about $1,000 of weekly usage, which a heavy user can burn through in about 2.1 hours.
Cached input reads cost $6 per million tokens on Ultrafast, and cache writes hit $75 per million, both far above standard-tier pricing.
OpenAI employees reportedly now need approval per task to use Ultrafast internally because even they hit compute restrictions with it.
A working first draft took under two minutes on Astra Ultrafast versus 15-20 minutes on a slower, higher-quality model, though the fast draft looked visibly rougher.
One well-known AI coding streamer burned over half his $500 Ultrafast allotment in a single hour of casual use.
Measured real-world throughput ran roughly 320-340 tokens per second on Ultrafast versus about 30 tokens per second at regular speed.
When responses return in seconds instead of minutes, the bottleneck shifts from waiting on the model to waiting on tool calls like builds and git commands.
A full build-and-iterate session on a benchmark dashboard cost $306 for the main thread and $250 for follow-up work, versus an estimated $90 on standard pricing or $12 on a cheaper comparable model.
The real value of the fast tier isn't speed for its own sake, it's staying in flow, since waiting between steps means losing track of what you wanted to change next.
The only defensible everyday use case named for the fast tier is urgent incident response, and realistically only for people who don't personally pay the bill.
A future fast tier on a cheaper model, priced with the same multiplier, could end up cheaper than the current expensive model's own standard pricing.
Takeaway
What Ultrafast inference actually buys you
SPEED PREMIUM
Inference speed is a separate dial from model quality, and paying six to ten times more for it only pays off when staying in an uninterrupted feedback loop is what gets the project finished.
01Cold open: live-updating code feels trippy
Watching code edit itself in real time changes how you work with a model, shifting from fire-and-forget prompts to an active back-and-forth.
Raw generation speed becomes a distinct product feature, separate from how smart or accurate the underlying model actually is.
02The catch: $300-450 per million tokens
A single speed tier can carry a 6x price multiplier over the same model's standard tier, with no change in output quality.
Long-context work compounds the penalty further, pushing output pricing even higher on the same speed tier.
04The reveal: $600 to review two small pull requests
Reviewing two pull requests under 100 lines each cost $600 in Ultrafast usage, roughly $300 per review.
Token spend on a task has almost no relationship to the size of the code being reviewed once the agent starts reasoning at length.
05Mapping the $500/month plan's burn rate
A $500/month plan advertised as $1,000 of weekly usage can be burned through in about two hours of real use.
Two small prompts, under a minute of generation combined, dropped a usage balance by a full percentage point.
06Live demo: building a kanban app in real time
Splitting your screen between the app and the model, instead of looking away and coming back, keeps you from losing track of what you wanted changed.
Explicitly telling the model to skip slow tool calls like deploys and previews matters more once generation itself takes seconds instead of minutes.
When inference drops from 10 minutes to 30 seconds, the 30 seconds you spend waiting on builds or git commands becomes the dominant cost, not an afterthought.
07Watching the usage meter drain live
A handful of casual prompts can measurably drain a weekly usage balance in minutes, not hours.
Even a company's own employees reportedly need approval per task to use a tier this expensive, because the compute cost is real even when the bill isn't.
08Building Slopalytics: Astra Ultrafast vs a slower model
A fast model can produce a working first draft in under two minutes versus 15-20 minutes for a slower, higher-quality model, trading initial polish for iteration speed.
Sending dozens of small corrections in one thread, more than fifty messages, produced a better end result than a handful of large batched requests.
Running two models on the same task in parallel is a cheap way to find out which one actually earns its price for a given job.
The same build cost hundreds of dollars more on the fast tier than it would have on standard pricing or a cheaper comparable model.
09The real case for Ultrafast: flow state, not savings
On pure dollar-for-time math, paying for speed almost never beats being patient and waiting for a cheaper model to finish.
The real cost of waiting isn't the clock time, it's losing the mental context of what you wanted changed and having to rebuild it.
Staying in the loop produces design decisions you wouldn't think to make from a single batched request, because you're reacting to what you see instead of planning ahead.
10Verdict: use it only when you don't care about losing it
The one defensible use case for a premium-speed tier is genuinely urgent work, like incident response, not everyday building.
It rarely makes sense to upgrade a plan tier specifically for a feature this expensive until a cheaper model offers the same speed option.
When a cheaper model eventually gets the same speed multiplier, it can land cheaper than the expensive model's own standard price, which is worth watching for.
Glossary
Terms worth knowing.
Ultrafast
A premium inference-speed tier from OpenAI that generates tokens far faster than standard speed, at a steep price multiplier on top of the base model's rate.
GPT-6 Astra
OpenAI's flagship coding-focused model discussed throughout, sold at Standard, Batch, Flex, Fast, and Ultrafast speed tiers, each with its own token price.
Cached input tokens
Previously-processed tokens the API can reuse at a discount off the standard input price; fast-speed tiers can price cached tokens well above their normal-speed cost.
Token throughput
The rate a model generates output, measured in tokens per second; it's the real-world speed difference between a fast tier and standard speed.
Long context
A request that includes a very large amount of input text or conversation history, which can trigger a higher per-token price on some pricing tiers.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphor
Ultra fast may not have been ultra fast to get here, but now it is here. And not only can it build absurdly quickly with a model like Astra, it also can burn through your rate limits and bills even faster. It is still genuinely trippy when you're using it though.
Like I have an app I just whipped together with like that and I'll say make it light mode and funnier. And we'll just watch as it live updates. This is all real time.
So I hope I don't stumble over any words because I don't want to. have to deal with that in post while still maintaining the real time.
I was taking a snapshot of the page. Oh, it's already done. Yeah, again, fully real time.
It's just flying through it. There's a ton of real use cases for this, like prototyping live, working in the loop, getting work done faster, creating experiences for users that feel way more immediate, computer use in a way that you can watch and feels even faster than a human using the computer, and more. but as powerful as it is to use a model that is this fast the price is rough aster's already a little expensive at 50 dollars per million output tokens but on ultra fast it goes up to three hundred dollars per million out and god forbid this is a long context project because now you're at 450 dollars per million tokens out these bills are actually insane when i was first testing out ultra fast i used it for some basic day -to -day work like i had it review two small pull requests that i had open both were under 100 lines of code and they were quite expensive i want you to think about that how much would you guess was the token spend to review two pull requests
And while you think about that, I hope you don't mind me making some of that money back with a quick break for today's sponsor. Gonna ask you to think into your past. This might hurt, so I'm sorry in advance.
I want you to think back to a particularly complex pull request that you had to review in the days before AI. You might have looked through every line of code and thought you saw everything that could possibly go wrong and left a bunch of feedback just to have it go out and realize things were broken. When that happened, you probably realized the best solution is to actually pull down the changes and test it yourself before giving the thumbs up, but that's slow and tedious and annoying.
But if you don't do that, are you actually giving the right realistic feedback for that stuff? Because there are so many times where the code looks totally fine, but then when you go to run it, things work in unexpected ways. Today's sponsor is Graptile, and they asked themselves the same question because they were looking at what their agents were doing and realized that, yeah, they're pretty good at reading code and finding issues, but they do a hell of a lot better when they're given a real computer to run that code on.
And that's why I'm so hyped on T -Rex. t -rex is a new solution that reptiles put together to allow your review agents to run on a real computer that can actually run and test the changes so it's no longer just looking at the code and giving it a thumbs up that it thinks it's fine it's actually testing the code itself when running t -rex on a real code basis we've seen the benefit immediately for example in one of ben's projects for managing all of our youtube stuff this pr would have had a real regression in how we manage our youtube auth keys but since reptile ran it for real it noticed the issue immediately and even provided evidence of these Welcome back.
Do you have a price in mind? I see people in chat guessing like $100 per PR. It was $600 of usage for both.
It's insane. It's crazy how much money you can burn with this. Chat doesn't seem to believe me that it was $600.
two zeros yeah yeah it was but if we go back here and look at the price and realize it went from 50 per mil out to 300 per mil out you realize just how brutal this is a 6x increase is rough especially outside of the output tokens the cash read cost going up to six dollars per mil makes cash reads more expensive than they are to do normal reads on other models.
And the cash right cost at 75 bucks is insane. It's like unbelievably high. I'm going to be so real with you guys.
I just didn't think ultra fast was going to be for me at all. That would make any sense at all. And I still do largely believe that.
I don't think almost anyone should be touching ultra fast right now. It's just... far too expensive, especially over the API rates.
But even the amount you get in the new $500 a month ultra fast plan they offer is just, it's not enough. They are subsidizing it heavily. So in that ultra fast $500 plan, you get about $1 ,000 of usage per week, which takes a whopping 2 .1 hours to burn when you're using ultra fast.
I refreshed my limits right before starting prep for this video. And I had 38 % left in my Codex account, the one that has the $500 plan. i sent two prompts one to do this quick build it took 38 seconds and one to make it light mode and funnier and it took 16 seconds that knocked my usage down from 38 to 37.
that's a one percent hit for under a minute of work let's ask it to do a bit more quick how about turn this into a kanban app with a live chat Make the changes one at a time so I can see them appearing in the dev server as you go. I'm going to knock it down to medium as well.
Not deploy as you go too. Because Lakebed will auto update the live deploy when it changes as well, which I think is cool. But let's just go look at the local version for now.
So again, all real time. We are not doing the usual cuts. I'm sorry, Jeff.
You're the one person who's going to have an easier time. And we see it as it updates while it is going. It's full of useless subtext, which I'll ask it to fix.
But it's going to keep making changes as it goes. And we'll see them come in live. And it's trippy.
You can have a very different working relationship with a setup like this. I found that it actually makes sense to split screen with your app. and your model again.
Because I'm not just asking it to go do the thing and then looking at it when it's done. So if you're struggling a lot with those flows where you ask it to do a thing, the thread disappears, then you come back to it later, this is very good for getting out of that. But it is still just far too expensive for it.
Now I'm going to do... I'm going to move this to the in -app browser.
Oh, that's the deployed version, yeah. I'm going to do the in -app browser so we can see it as it goes a little more easily.
And here's where we're going to start noticing the first thing that you'll want to think about when you use ultra -fast mode. When your token generation took... 10 minutes and your git commands or builds or whatever took under 30 seconds that ratio meant that the token time was so much higher that you didn't care about optimizing those other pieces i almost feel like i want to use a different system prompt when i use ultra fast because i want to steer it away from things that aren't as important that take more time Because when inference goes from 10 minutes to 30 seconds, that 30 seconds of waiting on tool calls goes from a small percentage of your runtime to doubling your runtime.
So I just told it, for example, to not make changes to the preview browser, just change code when I ask. Because I don't want it to call tools that take seconds, because I just want to see the updates while we go. So let's do another here.
There's a ton of useless space being taken up from the top. I want you to delete all unnecessary subtitles and information and really try to clear out the top nav so that way less vertical space is wasted. And now we get to watch it make these changes live.
All gone. 11 seconds. It's very nice.
And if this is the flow you want to be in the loop with the model, I think this makes a ton of sense. And where this gets really crazy is with models like Astra, they handle... getting steered much better so i can do something like say make it dark mode by default again and say i don't want guests to be able to chat it just made the dark mode change in five seconds allegedly working is a bit lame make the copy better and i'm just sending these while it's doing other things forcing it to steer but it will do all of them at the same time fine Can I attach images?
Make that work for attaching images in new tickets and chat. Signed in only. Also, new tasks should be signed in only.
And you can just spam it with things you notice as it is working. And the result feels something that I've never actually felt before coding, where it's the flow state that you used to get when you were hopping between the code and the output. and knowing what was affecting what and just flying between the two back in the day when we would actually write the code ourselves but you can make way bigger changes this way too because you can build real things with this because it is the real astra it's actually running on nvidia i thought it was on cerebrus i just recently learned it's not with this it is magic combine it with voice and the experience feels fundamentally different and i understand why all those open ai employees were so hyped about this it's because of this unique feeling with the live back and forth but it's also because they don't have to pay the bill because opening eye employees get a limited inference that's of course they do they're employees but just from the handful of changes i made here let's see the effect on my usage limits i was at 37 before that like small handful of changes now i'm down to 34.
on the 500 plan i disabled two features i switched to dark mode And I'm now down 3 % of my $500 plan. It's insane.
I'm now going to ask 61Soul on Medium to analyze that thread and see what the dollar value would have been. And while I do that, I will have a sip of a beer because this makes me very depressed. Apparently, OpenAI employees now have to get approval per task for Ultrafast because even they are hitting compute restrictions with it because it's so expensive.
That's a bit of a relief. No longer feels so much as inference for me and not for thee as it did before. Once you get used to ultra fast, everything else feels super slow.
Yeah. If I was paying standard API prices for that Astra usage, it would have been $7 .72. But instead it was $46 .32.
Yeah. Out of curiosity, I want to know what the price would have been using 6 .1 Sol. Instead, it would have been $1 .10.
So if I use 6 .1 Sol at normal speed, it would have been $1 .10. And it's absolutely capable of the same things here. But instead I spent over 40x the same amount to do it with Astra ultra fast.
On one hand, this shows how insane of a value 6 -1 Sol is, but it also shows how absurd the cost is for Ultra Fast. And it really is absurd. So you're probably looking at these numbers and thinking that there's no world in which this makes sense.
And I'm not going to say I disagree. I think it was insane of them to release this at this price. I don't want to shill it.
I don't want to push it. I don't think most people should use it. and it's not surprising that someone like primogen bought the 500 plan saw it was the ultra fast plan and decided to try it out ended up burning over half his limit in an hour they really should not have made this seem like a thing you just click it's the same problem i have with the max reasoning level funny enough when you give users a knob they will turn it all the way and that has resulted in a lot of bad decisions by both OpenAI for giving us the knob and the users for using the knob.
And yeah, just it sucks. So where is the value in this? I'm going to explain this in a strange way with a recently released app I built called Slopalytics.
I don't know if you guys are like me in this way, but when I look at a site like this and I'm like, wow, this actually looks decent. The info is clearly communicated and it's easy to navigate and things work. That must have been made with an anthropic model, right?
Like clearly this was done with Opus or Fable, right? Well, I actually built this entirely with Astra on Ultra Fast. Funny enough, I was working on it with Opus 5 .5 and Ultra Fast at the same time.
And it ended up taking the Opus run 15 minutes or so before it had a first working draft. Meanwhile, it took Ultra Fast well under two minutes. It was close to about a minute and a half of Ultra Fast work on Astra.
to get a first working draft but it was astra so you can guess it was ugly the first draft from opus that took almost 20 minutes looked meaningfully better than the first draft that took a minute and a half using astra ultra fast but by the time i saw the opus version the astra one was better because it only took a minute and a half and i looked at it and i said here are the issues i take with this version you just made And then they updated live in my browser.
I was like, okay, those things are better. These two things still suck. Can you try this?
And then it tried it. It's like, oh, I don't like that. Go back the other way.
Try this other thing instead. And I ended up sending more messages than I've sent in a thread in a long time. After a bit of finagling and learning some things I fucked up in my existing local T3 code, I found the thread.
And I think this will probably be the best way I can showcase why I enjoyed using Ultrafast for this so much. I'd go as far as to say I might not have even finished Slopalytics if it wasn't for ultra fast. It made it way more fun to build this way.
So the first thing that happened is since I was in the thread, I noticed it trying to use my TS dev skill, which is a skill for using tail scale to host dev servers remotely. But I was doing this on my machine, so I didn't need that. And since I was actually paying attention, I corrected it.
I'm on the machine. It can be local host. 30 seconds later, Snapshot now has 688 variants.
i noticed the load was bad though because it started making the page and it had spit out a url at some point so i went there and it didn't work it ended up being roughly two minutes of work total before it had a local host version that had all of the data from artificial analysis and a decent ish ui but i noticed some of the bar charts were bad they were all horizontal which didn't have the comparison i wanted so i said vertical bars make it all fit fine on a normal screen i sent a screenshot 41 seconds later where it had a random error and this is what it looked like at the time better artificial analysis and it was it was sloppy and i told it not acceptable fixed some old css issue had it live i said use more logos and colors consistent with the labs use artificial analysis for inspiration a minute later it was updated i i almost immediately respond you'll notice the time stamps here it responds at 1 14 a .m i respond at 1 14 a .m He responds at 1 .13.
I respond a minute later, not even. I was in the loop for this. I had stopped paying attention to the Opus thread that was working on the same task.
Change my mind on sort. Averages are useless. Sort by the highest option.
Get rid of average lines. this was all ideas that i had to make the sorting make more sense and to change how i was grouping the models one of the things i think i did here that's the coolest i don't exclusively show the max reasoning levels and then sort by tokens on max reasoning and this is one of the problems i had with the artificial analysis dash one of my final artificial analysis flashbang warnings ever because i will be showing slopolytics in the future i just want to compare quickly the numbers on the home page here are for Opus 5 .5 and Fable 5 .1.
These are all max reasoning, as are the costs. None of these numbers make sense. And it gets even worse if you go down to output tokens, where you see Sonnet at 200 ,000 tokens, Opus at 120 ,000, and then Astra all the way here at 27K.
It makes even less sense when you turn on Opus 5 .5 high, and it appeared somewhere. Can you find it? I happen to know where it went.
It's here. Opus 5 .5 is really far on the left, even though 5 .5 max is over there. It is impossible to see the relationships between models in a family in this view.
So instead, I group the numbers by the model family. And then the order is based on which one has the highest number, which has some fun side effects. For example, when I put in Sonnet, it's very much the worst model token efficiency wise.
But if I turn off max, which is one of the things I have here is my little toggles between low, mid, high, and max. I can turn off max reasoning and now Sonnet has been moved a bit to the side because it is unreasonable on max effort, but it's actually quite reasonable on X high. And it's even more reasonable on high, dropping down to be lower than Opus for the amount of tokens it's using.
Not only would I probably not have spent the time building this if I didn't have ultra fast, I don't think I would have come up with all of these UI patterns. Because I had to do back and forth where I asked it to do a thing, it didn't look right, and I would just ask it to do something else. And even though Astra is way, way, way, way worse at design than the Anthropic models are, I made a better design because I was in the loop with it.
And it's been a while since I was in the loop. I even made a change in T3 code where when threads are running, they get auto -hid because I don't care what the thread is doing until it's done. Unless, of course...
The model is working with me and we have this real quick back and forth. When I went back to this thread just now, I noticed that I was on 6 .1 Sol without Ultra Fast because there is no Ultra Fast for 6 .1 Sol. I cannot wait till there is.
But I was on 6 .1 Sol regular speeds and I made the switch not because I was so scared of wasting all my money, but because my food was in the oven and I wanted to be able to get up and do it and not feel like I was missing as much. So I ended up going and getting my food out the oven, coming back, and it was still working.
But then I switched back to Ultra Fast and I was flying again. And what you'll see in this thread is a very different prompting style than I normally do. Picture with not acceptable.
Use more logos and colors consistent with the labs. Change my mind on the sort. Averages are useless.
Sort by highest. Get rid of all the useless titles and subtitles. It's so cringe.
And like all these are being sent within like minutes of each other. 115, 116, 118, 114. That's like five prompts in four minutes.
And all the replies are under a minute. Some of them are 20 seconds. The sort change, 18 seconds.
You don't have time to get distracted. Same with share button. It's just a URL.
I was telling you to like remove things. Default view doesn't need to have a URL hash. Only do that when they make changes.
Sort orders all wrong. Default should be highest last. Can ever draw a line showing where it cuts off other models?
Oh, and if I click it, it should persist. And hovering other models will show the gap in score and cost. This is another one of the crazy UX things I added that I would never have come up with if I didn't get to go back and forth this way.
When I want to compare two models in this chart, let's say I want to compare Sonnet 5 on X high to other things. Like, let's see how it compares to Fable 5 .1. Now I have this awesome comparison view that appears in the top left where it shows you the multiplier difference for the cost, the tokens, and the intelligence scores.
I only came up with this because I was staring at it while it was being worked on and could have the back and forth with the model and try different things. I ended up doing a lot of back and forth here. i forgot what the first version looked like i screenshot it because i hated it yet here it was like covering things up i changed my mind three times as to where this card should go and just went back and forth a whole bunch until i was happy with where it ended up i was still unhappy with how it was organizing the panel so i just told that i want these three columns and then it did it ui feels a bit flickery when hovering around not sure the best solution maybe delay it did it and i hated it i changed my mind this feels awful it took 25 seconds to do it 10 seconds to undo it And I sent another follow -up.
Panels should be top left, not right. Top right's a forbidden area. Need that for my face in my videos.
Also, can we move the sidebar to the right in a way that doesn't suck? 32 seconds later, it was done. Hovering names in 2D charts should make the line more prominent and others more transparent.
This is in these charts. So I can hover Opus 5 or Sonnet and it makes it clear which is which. That was a thing that I randomly came up with when I was scrolling around.
So I just threw it there. Then it was fixed. Good work.
Sadly, it's still hideous and it needs a bunch of cleanup. Get on that. Spin up some separate sub -agents to work on branding.
And it did. It came up with a bunch of names. I hated all of them.
I asked it to have other threads come up with things. It came up with a bunch of garbage. Turns out models are still bad at naming things even if they're fast.
I came up with Slopalytics. And here's where it became Slopalytics. But like, do you see how many messages I did in this thread?
There's like 50 plus. I never send that many messages in a thread. it felt entirely different to build this way also i just ran real tps numbers for my usage i saw regular at 30 tps fast at 60 ish and ultra fast at over 300 between 320 and 340.
crazy but i also found the cost for all of my threads working on this that main thread was 306 dollars and the follow -up where i did a bunch of additional work was 250. If I was not using Ultra Fast and instead I'd used a normal Astro, it would have went from $550 to $90. And if I had used Sol 6 -1, which would have been just as good for this work, by the way, $12.
So now we have to ask the question, is saving all that time worth $540 or so to you? The answer is probably no. and if you're comparing it that way you should never ever ever use ultra fast it's just straight up not worth it it is far too expensive it burns your limits hilariously quickly it made the 500 tier make no sense at all and it's just it's strange it's bad it should not be included in the product of a bunch of warnings so that's why it's going to hurt me a lot to defend it because this is not a real comparison here this is a comparison of how much astro ultra fast costs compared to soul 6 -1 The difference here isn't I could have been more patient and saved 500 something dollars.
The difference is that I just wouldn't have built it. If I had to wait between every step, I would have lost the context of what I'm thinking about. I would have lost track of the things I wanted to change.
And I have to rebuild that context every time a step finishes. So I found myself writing smaller requests and asking for smaller changes instead of doing what I do with Opus, which is look at the output, make a list of the 20 things I want changed, send it, and then expect it to come back in between five minutes and two hours.
And this in the loop work is super useful for certain things like refining a UI. It's silly to put it this way, but I can make better UIs with Astra than I can with Fable and Opus. Not because Astra is better at design, it's objectively worse, but because design is all about how users interact with and perceive things.
And if I could spend more time on that perception, and I can spend more time in the loop here, I can fly. But this is where I have some hope, and also where I think our friend BotCooper's catching on.
Am I tweaking, or if the multiplier stays 6x, then 6 -1 Sol Ultra Fast would come in notably cheaper than 6 Astra, and much cheaper than 6 Astra Fast. Well, considering that 6 -1 standard pricing for Sol is $12 .03, and Astra's would have been $90, yeah. In a hypothetical world where 6 -1 Sol comes out with an ultra fast mode that happens to be the same price, multiplier specifically, not only is it 10 times cheaper than Astra Ultra Fast, it actually comes out around the price, if not cheaper, than Astra Standard.
You see where we're going here. The point of Astra Ultra Fast is not to be used. It is to be experimented with where lunatics like I can waste far too many tokens on it.
just to see what it's capable of and get a feel of these flows because we're hoping for a future where you can get these speeds at a price that doesn't make you feel sick and if we do end up in that future where sol61 is just as fast on ultra fast as astra is at a tenth of the price i could actually see that becoming my default model i already have 61 soul fast as my default for a ton of when i open a new project in fleet which is my repo for managing things Oh, it's currently pinned to Astra.
I'm actually just going to fix that now because that is not what I want it to be on. I changed it on other machines. This is now going to be 6 .1 Sol high fast because I think it's a really good balance of speed, capability, and performance in general when I'm trying to do tasks where I'm managing the machines in my network, not for writing code I'm deploying, but for working on my systems directly.
I got one more thing for y 'all. Tebow recently teased. 6 .1 coming soon.
A lot of people assume that was 6 .1 Astra, but it was actually a reply that was cut off here about Ultra Fast. This is a public confirmation that 6 .1 Sol Ultra Fast is coming soon. I think that this might end up being Cerebrus.
Not sure just yet. Could be Cerebrus. It could just be the crazy NVIDIA hacks they're doing right now.
Either way, 6 .1 Sol Ultra Fast is coming. And I think it might be the differentiator that makes this $500 Ultra Fast plan go from feeling overpriced for what you're getting. to a genuinely incredible value.
So now we are left with one important question. What should we use ultra fast for? The only thing this is good for is super really high priority issues you're trying to resolve, like an incident response for a high severity issue.
Assuming you work at OpenAI, because it is even for high severity incidents, probably not worth the money unless the money isn't real. It's just so expensive. It is basically impossible to justify.
There is... One other time you can justify it though. And for me, that time is right now.
I have nine manual resets on my ultra fast account because I went to an event where they gave a six and then I have all the others that we've all been getting. One of those expires in an hour. That means I have one hour to burn 34 % of this account's limit.
That's a great time to use ultra fast. And since our friend Tebow made a pretty bold announcement today. over the next 28 days each day we'll either ship one thing that is a clear improvement and relevant for most codex and work users or we'll ship a full reset let the improvements begin this means there will be a lot of opportunities for me to burn my ultra fast because whenever a reset is promised they're always late whenever they say like it'll be out at 10 am it's out at noon earliest so just start burning Since I'm using Opus for my serious work anyways, my codex account being at zero is fine.
It's not going to block me a whole lot. So as soon as there is a whiff that a reset is coming, I burn that account to the ground. Sometimes the reset happens before I hit zero.
Sometimes it happens an hour or two after I hit zero. Sometimes it happens the day after. I'm using my ultra fast when my account's about to reset anyways.
So that's when I use it. I don't know when you should use it. Maybe for incident response here and there.
If you're an OpenAI employee, I'm sure you'll use it a lot for those types of things. Use it when you don't care to lose it. Don't upgrade to the $500 tier just for this yet.
It's not worth it. With 6 -1 -Soul Ultra Fast coming soon, it might suddenly become more worth it. And I'll be sure to keep you guys in the loop when I have more info on that.
But for now, $500 tier, I don't think you should reach for it. What about the $200 tier though? You don't get Ultra Fast, but you do get 6 -1 -Soul.
You do get Astra. You get Fast Mode. You get a lot of other things.
You get all these resets. Is the $200 plan still worth it? Well, for that, I'm going to need another beer and probably another video that will be coming up very, very soon.
So make sure you're subscribed if you're not, so I can cover all of these things and make sure you know what's going on in this crazy AI world. Now you know that UltraFast is worth avoiding. Hopefully soon, you'll also know if that $200 plan is worth using or avoiding too.
Until next time.
God, peace nerds.
The Hook
The bait, then the rug-pull.
OpenAI's new Ultrafast speed tier for GPT-6 Astra generates code live, in front of your eyes, for a price that turns a two-PR code review into a $600 bill. What starts as a roast of an absurd price tag turns into a genuine case for a workflow most people still can't afford.
Frameworks
Named ideas worth stealing.
03:12list
Astra's five-tier pricing ladder
Standard: $50/M output
Batch
Flex
Fast
Ultrafast: $300/M output, $450/M on long context
OpenAI's pricing page breaks GPT-6 Astra into five speed tiers, each with its own token price, and Ultrafast carries the most extreme multiplier.
Steal forany pricing-tier comparison explainer
CTA Breakdown
How they asked for the click.
VERBAL ASK
27:53subscribe
“make sure you're subscribed if you're not, so I can cover all of these things”
soft tease for a follow-up video on the $200/month plan rather than a hard sell
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Theo puts Meta's Claude Code clone, Muse Code powered by Muse Spark 1.2, through benchmarks, a codebase audit, a game rewrite, and a live integration test to see if the price is the whole story.
A 23-minute rebuttal of three viral claims about Anthropic's returning Fable model — that it's nerfed, that its subscription pricing is a bait-and-switch, and that it's too expensive to run.
A 2 a.m. field report on GPT-6.1 Sol, the surprise OpenAI model that matches Claude Opus 5.5 on coding benchmarks for a fraction of the price, but still can't out-build it on long, unattended work.
A developer who used to type 160 words a minute lost the use of one hand, and rebuilt his entire coding workflow around a whisper, a fleet of machines, and agents he no longer reviews before they merge.
A four and a half hour Labor Day stream where two entire YouTube videos get filmed live, one-handed, between sub thanks, a ban, and forty agents running in the background.