Modern Creator
Theo - t3․gg · YouTube

OpenAI fights back

A 2 a.m. field report on GPT-6.1 Sol, the surprise OpenAI model that matches Claude Opus 5.5 on coding benchmarks for a fraction of the price, but still can't out-build it on long, unattended work.

Posted
yesterday
Duration
Format
Review
hype
Views
270.6K
4.9K likes
Big Idea

The argument in one line.

GPT-6.1 Sol matches Claude Opus 5.5's coding benchmark scores at a fraction of the cost by cutting cache-read pricing in half, but it still cannot sustain long, unattended coding work the way Opus can.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • A developer routing AI coding agent work across models who cares about dollar cost per task as much as raw benchmark scores.
  • Someone already paying for a $200/month AI coding subscription who wants to know if a cheaper model changes their setup.
  • A builder trying to understand why AI subscription pricing keeps shifting and what a provider's pricing tweet actually signals.
SKIP IF…
  • You want a rigorous, independently verified benchmark comparison — these are one power user's self-run numbers, explicitly caveated as imprecise.
  • You don't run AI coding agents day-to-day; the cost-per-task figures won't mean much without that context.
TL;DR

The full version, fast.

OpenAI quietly shipped GPT-6.1 Sol days after GPT-6 Sol, a gap so short the host argues it was never meant to carry a '.1' label. It matches Claude Opus 5.5's scores on Terminal-Bench and DeepSWE while costing a small fraction as much, largely because OpenAI cut its cache-read price in half to $0.10 per million tokens, which matters since agent workloads run roughly 96% cache reads. The model also fixes much of GPT-6 Astra's 'smart and dumb' spikiness, making it steady enough for scoped tasks like code review, audits, and even personal finance work. But it still stalls on long, unattended, multi-day builds, where Opus 5.5 remains the stronger collaborator.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:00 – 01:30

01 · Cold open

Recaps Anthropic's Opus 5.5 shockwave, OpenAI's two weak rushed responses (GPT-6 Sol, GPT-6 Luna), and teases the surprise model the host got early access to.

01:30 – 03:50

02 · Sponsor: Depot

Paid read for Depot's bare-metal CI/build infrastructure, pitched as faster and cheaper than GitHub Actions for teams running coding agents.

03:50 – 07:09

03 · Benchmark caveats and the '6.1 wasn't supposed to be 6.1' theory

Discloses his self-run Terminal-Bench and DeepSWE numbers are imprecise, shows GPT-6.1 Sol matching Opus 5.5's DeepSWE score at 73x lower cost, and argues the one-week gap since GPT-6 Sol means this wasn't planned as a minor version bump.

07:09 – 08:42

04 · The Tebow tweet and the pricing tell

Reads a competitor executive's tweet about reopening the $200/month Pro plan with a worse usage formula, and argues the only reason to quietly cut that value is a more expensive model about to launch.

08:42 – 10:09

05 · Terminal-Bench revisit: the cost-per-task gap

Returns to the Terminal-Bench chart across 17 models and highlights a four-to-five times cost gap between the cheapest GPT-6.1 Sol run and the cheapest Opus 5.5 run at the same score.

10:09 – 14:07

06 · Astra's smart-and-dumb spikiness vs. Sol's steadiness

Explains why GPT-6 Astra's brilliance was undermined by random low-quality spikes, and describes trusting GPT-6.1 Sol to review his email and wire money because it doesn't spike the same way.

14:07 – 16:27

07 · Real-world audits: Orchestrator V2, T3 Code, and a blind review

Runs GPT-6.1 Sol against Claude Opus 5.5 and Sonnet 5.5 on real codebase audits, then has Opus blindly review Sol's work under a fake name and rank it frontier-tier.

16:27 – 18:46

08 · Opus's honest tier call on 6.1 Sol

Reports Claude Opus 5.5's own verdict: GPT-6.1 Sol is frontier-tier for scoped review and audit work but falls below frontier on long, unattended building.

18:46 – 21:05

09 · The TS-Rust port: what Opus unblocked and what Sol cleaned up

Tells the story of a TypeScript-to-Rust rewrite stuck for months under GPT-6 Astra and Sol, unblocked by Opus 5.5 in a single day, then audited by GPT-6.1 Sol, which found 1.3 million lines of dead code Opus never flagged.

21:05 – 22:18

10 · Detour: JevRouter and the routing scam

Benchmarks OpenRouter's cost-optimizing JevRouter, finds it routes most requests to a cheap small model anyway while taking far longer and scoring worse than GPT-6.1 Sol at a fraction of the price.

22:18 – 26:54

11 · Fish Slop: beautiful graphics, garbage UI, janky gameplay

Demos a $5 AI-generated aquarium game built with GPT-6.1 Sol: photorealistic fish and plants, but cluttered with more than 20 pieces of unnecessary UI copy, worse movement, and a lower frame rate than prior demos.

26:54 – 29:36

12 · A real regression Sol caught that Fable and Opus missed

Shows a live pull request where GPT-6.1 Sol found two real bugs in a worktree setup feature that both Claude Fable 5.1 and Claude Opus 5.5 had already reviewed and missed.

29:36 – 30:58

13 · Verdict: still Opus's daily driver, but Sol earns a permanent seat

Concludes he isn't switching daily drivers or canceling any Claude subscriptions, but plans to wire GPT-6.1 Sol into his workflow as Opus's dedicated reviewer and investigator, and reads that the pricing may signal the end of the AI subsidy era.

Atomic Insights

Lines worth screenshotting.

  • GPT-6.1 Sol shipped about one week after GPT-6 Sol, fast enough that the host argues it was never meant to be a '.1' release at all.
  • GPT-6.1 Sol scored the same on the DeepSWE benchmark as Claude Opus 5.5, but did it for $0.21 versus Opus's $14.65 on the same task.
  • OpenAI cut its cache-read token price in half, from 10% of the input price down to 5%, the first time it has ever changed that ratio.
  • Agent coding workloads are roughly 96% cache reads, so a cheaper cache price matters more to real-world cost than the headline input or output price.
  • OpenAI's rumored model margins run as high as 95%, meaning a $10 charge can cost the company as little as 50 cents to serve.
  • The most expensive GPT-6.1 Sol benchmark run cost $1.38 per task; the cheapest Claude Opus 5.5 run cost $5.12, a four-to-five times gap.
  • GPT-6 Astra was smarter on paper than Anthropic's best models but randomly produced some of the dumbest output the host had seen from any model that year.
  • GPT-6.1 Sol found two real bugs in a pull request that both Claude Fable 5.1 and Claude Opus 5.5 had missed during their own review.
  • When Claude Opus 5.5 rewrote a stalled TypeScript-to-Rust port from scratch, it left 1.3 million of 1.8 million total lines of code unused and never flagged it.
  • OpenRouter's cost-optimizing JevRouter sent roughly 60% of requests to DeepSeek 4.1 Flash and still took four to five times longer than GPT-6 Astra for a worse score.
  • A demo aquarium game with photorealistic fish and plants cost about $5 to generate, but shipped with more than 20 pieces of unnecessary UI text like 'A little golden overachiever.'
  • Asked to blindly price GPT-6.1 Sol, Claude Opus 5.5 guessed a cost roughly four times higher than what OpenAI actually charged.
Takeaway

The real lesson is specialization, not replacement

PICKING AI TOOLS

When a cheaper model matches your best model on scored, reviewable tasks but still stalls on long unattended work, the right move is running both together, not picking a winner.

03Benchmark caveats and the '6.1 wasn't supposed to be 6.1' theory
  • A benchmark score is only as trustworthy as the person running it: self-run numbers can shift between passes, so treat single-source benchmark claims as directional, not exact.
  • A surprisingly fast follow-up release, arriving just a week after its predecessor, is itself a signal: it often means the update is a bigger jump than the version number suggests.
04The Tebow tweet and the pricing tell
  • Watch what a company changes about its subscription math right before a launch: cutting the value of a flat-rate plan is a tell that a more expensive-to-run model is coming.
  • A subscription that resells discounted API usage only works while the provider's margins can absorb it; a bigger model launch is exactly when that math gets renegotiated.
05Terminal-Bench revisit: the cost-per-task gap
  • Comparing raw benchmark scores without cost-per-task hides the real difference between two AI models that appear to perform equally.
  • A four-to-five times cost gap between the cheapest and most expensive way to get the same benchmark score is common right now, and worth shopping for before assuming price is fixed.
06Astra's smart-and-dumb spikiness vs. Sol's steadiness
  • A model can be more intelligent on paper and still be less useful day-to-day if its output quality is inconsistent, since unpredictable failures cost more trust than a lower ceiling does.
  • Consistency, not peak intelligence, is what makes a tool trustworthy enough to hand real financial or high-stakes tasks to.
07Real-world audits: Orchestrator V2, T3 Code, and a blind review
  • Running the same audit task across multiple models and comparing both the score and the price is a more honest way to judge a tool than trusting a single demo.
  • Blind-testing a model, by hiding which one it is before asking another model to review its work, removes brand bias from the evaluation.
08Opus's honest tier call on 6.1 Sol
  • Judging an AI model by a single aggregate score hides that most models are strong at some task types and weak at others; sort capability by task type instead.
  • A model that follows its own process rules even when they stall all progress, and never asks for help, is a specific and costly failure mode worth testing for before trusting it with unattended work.
09The TS-Rust port: what Opus unblocked and what Sol cleaned up
  • The model that unblocks a stalled long-running project is not necessarily the model that leaves behind a clean result; a dedicated audit pass can catch what the builder model never flagged.
  • Weeks or months of stalled progress on a heavy rewrite is itself a signal to stop and change tools rather than keep spending tokens on the same approach.
10Detour: JevRouter and the routing scam
  • An automatic cost-optimizing router that can't reason about task difficulty will default to the cheapest available model far more often than it should, producing worse results slower.
  • Look at where a 'smart routing' tool actually sends its requests, not just the price it advertises, before trusting it with real work.
11Fish Slop: beautiful graphics, garbage UI, janky gameplay
  • A model can leap forward in one visual capability, like 3D rendering fidelity, while regressing in another, like UI and copywriting taste, in the same release.
  • Cheap, fast generation of an entire playable demo does not mean the output is usable without a human editing pass for taste and gameplay feel.
12A real regression Sol caught that Fable and Opus missed
  • Different models catch different classes of bugs; a second model doing a dedicated review pass can find real regressions that the primary model and its own reviewer both missed.
  • Trusting a model to review and critique code is a lower-risk way to use it than trusting it to write the code you intend to ship.
13Verdict: still Opus's daily driver, but Sol earns a permanent seat
  • The right setup is often not 'replace your main tool' but 'add a specialist': one model as the primary builder, a second cheaper model as its dedicated auditor and investigator.
  • A model that is dramatically cheaper and matches your daily driver on most tasks still isn't a full replacement if it can't sustain long, unattended work without stalling.
Glossary

Terms worth knowing.

Terminal-Bench
A benchmark that scores an AI model's ability to complete real command-line coding tasks, used here to compare model quality against cost per task.
DeepSWE
A coding benchmark measuring the share of software-engineering tasks a model solves, plotted here against dollar cost per task to compare price-to-performance.
Cache read (cached token)
A token from earlier in a conversation that a model reuses instead of reprocessing from scratch, billed at a steep discount versus a fresh input token.
Pro $200 subscription
A flat-rate top-tier monthly plan that bundles a set amount of API-equivalent usage for $200, whose value shifts whenever a provider recalculates the usage formula behind it.
JevRouter
An OpenRouter feature that automatically sends a prompt to whichever model it estimates will produce an acceptable result most cheaply, discussed here as unreliable.
Reasoning effort tier (low/high/x-high/max)
A setting that trades more compute time and cost for a higher benchmark score on the same underlying task.
Resources

Things they pointed at.

02:07toolDepot ↗
04:04toolTerminal-Bench 4.0
04:27toolDeepSWE
21:05toolOpenRouter JevRouter
26:54productT3 Code
29:40productLakebed
Quotables

Lines you could clip.

05:21
“6.1 is significantly smarter than the GPT-6 Sol, thereby indicating this isn't just a .1 bump, there is something fundamentally different here.”
sets up the video's central reveal in one line→ TikTok hook↗ Tweet quote
10:09
“It is smart and dumb at the same time.”
tight, quotable paradox that frames the whole GPT-6 family critique→ IG reel cold open↗ Tweet quote
11:49
“I'm literally trusting this model to wire money for me.”
concrete, personal stakes-raising claim about trust→ newsletter pull-quote↗ Tweet quote
14:51
“I had Opus 5.5 review this model with a different name, obviously, so I didn't know what it was.”
a clever blind-test methodology explained in one sentence→ newsletter pull-quote↗ Tweet quote
16:38
“For unattended long, like heavy rewrite type stuff, Anthropic is just comically far ahead right now.”
blunt competitive verdict from a paid early tester with no reason to favor either side→ TikTok hook↗ Tweet quote
24:33
“This model sucks at front end. It sucks at design. It has no taste.”
rant-mode punchline with high comedic, shareable energy→ IG reel cold open↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphorstory
They're back. Okay, let me catch you guys up quick. Last week, Anthropic dropped a new model, Opus 5 .5, and it was unbelievably good.
It was so unbelievably good that OpenAI rushed out two model drops, GPT -6 Sol and GPT -6 Luna. You might have noticed I didn't do a video on those models. There's a reason.
They weren't that good. I was not particularly impressed with either of them and didn't really have much to say. But there was one other model I happened to get early access to.
that is now available for all. That model is called GPT 6 .1 Sol. And this model made it very, very hard to film that Sonnet 5 .5 video because I knew something was coming.
Something surprisingly cheap, surprisingly capable, and most surprisingly, better than Astra. So, yeah. I got a lot to say about this one.
I'm filming this at 2 in the morning, right after filming my Sonnet 5 .5 videos. So pardon me for stumbling over a few words here and there. I'm doing my best to get this out as reasonably quickly as possible because I want to have some coverage.
And I'll be real. It is also quite fun to cover these things before they are out, so you are getting my true, honest take and not the distilled version of what everyone else is saying. I'm sure this model is going to cause some pretty crazy waves, so it will be nice to have my take out initially separately first.
As always, I feel obligated to remind you, I do have early access, but I'm not being paid in any way, shape or form. OpenAI has no influence over what and how I say things, just when. They've politely asked me to wait until the model is out to talk about it, which makes a lot of sense, but I have to wait for one other thing first.
Today's sponsor. in order to build good software with agents they need to get feedback and let's be real they're getting a lot of that feedback from our ci that's why we've all been seeing our ci bills skyrocket and also why we've been getting more and more frustrated with github actions today's sponsor is depot and they're here to solve all of this and more not only can they make your ci up to 10 times faster as well as your docker builds up to 40 times faster especially when they're downloading cash they're also cheaper and they give better feedback for your agents all this is possible due to depot metal They're running their own bare metal with AMD EPYC processors that are way faster than what you get from traditional CI providers like, of course, GitHub Actions.
If you want it to be a drop -in replacement, it absolutely can be, but Depot's APIs are so much better that you should probably use those instead. They enable parallelization and, most importantly, resilience when GitHub inevitably goes down randomly for no good reason. We've had our releases get blocked because we weren't using Depot, and I'm so thankful that I've been moving more and more stuff over.
For example, when Ben moved PickThing over to Bunn, we immediately had some CI failures. Normally, this would be obscure piles of text that our agents parse through for us.
But when we use Depot, it becomes way easier to see. They'll even analyze the failures and give suggestions to make it much simpler to get this feedback back to our agents. This is especially useful when you tell your agents that you can use Depot because they'll no longer have to push changes and wait for that to trigger a build.
They can just run the CLI to trigger the exact same CI that you'd be triggering through GitHub instead. No longer do you have to file PRs with broken code just to get feedback to your agents. They could just run a tool instead.
Your agents will also get way better breakdowns of what is taking so long in your actual CI runs so that you can figure out how to improve them and make them faster and more reliable. You and your agents deserve faster Docker, faster build times, faster CI, better results, and ideally a cheaper price. Get all of that and more at soydiv .link slash depot.
Let's talk about this model a bit because it is not quite what I expected and it's probably not what you guys expected either. Especially when you consider the GPT -6 sole just came out like a week ago. It'll be around a one week gap from 6 .0 sole to 6 .1 sole.
I also want to disclose the numbers I'm currently showing on my screen are unlikely to be exactly accurate because I am running terminal bench for myself. In the first two times I ran it, I screwed things up. The third one seems to be doing much better.
I didn't run it on medium initially, so there's a miss there, but low, high, X high, and max, although the max run is incomplete, so I'm currently backfilling scores from X high for the ones that max either got wrong. in a previous run or didn't do yet because it takes like eight plus hours some of these tasks i this bench is nuts it's i'm more skeptical of benchmarks than ever now that i've been running a lot more of them myself in order to get the coverage i want to give here for what it is worth terminal bench 4 is state -of -the -art score here as is deep swe although this one's weirder because it goes down on x high and max and stays even on low and high but those even low and high scores are scoring around what astra did on high The difference being it's doing it for comically cheaper.
Switch over to the log scale, you'll see what I mean. This model on low costs 21 cents versus Astra on low costing $1 .46 and Opus 5 .5 on max getting the same score for $14 .65. While I will gladly admit the DeepSuite is far from a perfect measure of how good a model is at day -to -day code work, the fact that 6 .1 Sol is scoring the same as Opus and is also 73x cheaper, is at least worth noticing.
Here's where I'm going to say some of the things that I probably shouldn't. Considering that GPT -6 Sol came out last week on Tuesday, and this model is coming out this week on Tuesday, I think it's reasonable to infer that 6 .1 Sol was not meant to be 6 .1 Sol. There are things I'm not supposed to say, and I'm definitely walking a thin line here by sharing it, so I hope this proves I'm not paid off by OpenAI, because I'm about to give you guys info that I, yeah, just let me get through this.
First and foremost, 6 .1 is significantly smarter than the GPT -6 Sol, thereby indicating this isn't just a .1 bump, there is something fundamentally different here. Next point is that it has meaningfully slower tokens per second. That tends to indicate the model is bigger, hard to know for sure, seems like this model might be different.
Most importantly, we have Tebow's tweet. What am I referring to there? Well, right before I started filming, Tebow dropped quite a wall of text.
The thing I want to emphasize here is, first off, that the Pro $200 subscription is back. But more importantly, more importantly is this sentence. They are changing how they calculate the usage in the sub.
In effect, if you do the math, it will net out at half the dollar in API spend compared to the old Pro $200 plan. Why in the world would they do this? Especially right now where there's allegedly an internal code red because Opus 5 .5 is so unbelievably good and has made the $200 cloud code sub such an unbelievable value.
The only reason in the world Tebow would post this right now is, uh, I don't know. maybe a new model is coming where their margins aren't as good so the ability to subsidize has gone down because remember you can get eight to nine thousand dollars of usage in a month on the 200 clod code plan and you can get over 12 grand on the 200 codex plan i did actually run a lot of numbers before this and the amount you could get on astra did go down slightly closer to like eight grand or so hard to know for sure because they differ for everyone everywhere and it's hard to log all of this stuff but for my math roughly nine grand a month usage and that's where the price for this model comes in this model is two dollars per million input tokens and ten dollars per million out this makes it way cheaper than 5 .6 soul was at launch half the price of 5 .6 all after discounts and the same price as gpt six soul one fifth the price of astra however this is not the whole story because cash reads matter and the cash read price for this model is going to be 10 cents per mil in that's a big deal
OpenAI has not changed cash read price. As far as I know, ever before, it's always been exactly 10 % of the normal read price. That makes it a 90 % discount, and now it's a 95 % discount.
That means they cut the cash read cost in half, massively reducing the cost for real -world agentic use, which, to be clear, is what we're using these for most of the time. So this makes the model absurdly cheap for doing real -world code work. That also means that they are almost certainly cutting into their margins.
Historically, these margins are rumored to be as high as 95%. Like for every $10 you spend, they only have to spend 50 cents. And as crazy as that sounds, it makes a lot of sense when you consider how expensive it is to make and train these models.
But that also gives them wiggle room to change things around a bit, which appears to be what's happening here. that also means that if they were to keep subsidizing the same level that they were on the subscriptions that your electricity cost for your sub would be more than you're paying so i get why they have to change this they've kind of just left the details out there for us to reverse engineer so uh take this as you will 6 .1 coming so fast seems to indicate it is not just a new snapshot of gpt6 so let's talk more about this model as i was showing earlier seems really good at terminal bench every time i refresh the numbers change because new runs come in and it looks like max failed some things that x high passed which is why i just dropped a bit but again pretty much all of these even high and x high are scoring higher than anything else ever has and this is for me running this benchmark on random vms on my network so uh not the best suite to test against i also had to drop three particular tasks from it because they expected an h100 to work against which i make decent money i don't make h100 money okay
but none of this is real world code work so let's talk a bit about that obviously we'll have all the fun things like fish slap near the end so stay tuned for that but i just want to fixate a bit on the costs here because the most expensive run with 6 .1 soul for me was about a dollar and 38 cents per task and the cheapest run with opus 5 .5 was five dollars and 12 cents that's a four to five x gap from the cheapest opus to the most expensive soul so at this point I would imagine you are hoping and praying this model is good and that it can actually replace Opus 5 .5 for day -to -day work.
And I promise we'll get some good answers to that in a bit. But first, we need to talk a bit about model behaviors here, because this model is a part of the GPT -6 family, which means it has behaviors that are worth talking about. I know I cite this diagram a lot, but there's a reason for it.
The thing that made me so frustrated with GPT -6 Astra wasn't that it was less intelligent than the best models from Anthropic, because it was more intelligent than the best models from Anthropic, and I would argue in many ways still is. But there is a problem. It is also dumb.
It is smart and dumb at the same time. GPT -6 Astra would just randomly spike into the dumbest bullshit I've seen a model do this year, even worse than like some of the small open weight models I play with. It's still so deeply frustrating that Astra does this.
Because on the other end, when it does well, it's unbelievable. But these spikes got to the point where I effectively churned. I was only using my codec subs for computer use.
And I ended up just leaning on to Fable 5 .1 and obviously now Opus 5 .5 for almost all of my day -to -day work. So have they addressed the spikiness? Has GPD 6 .1 Sol fixed the problems that I was so frustrated about with Astra?
I would say mostly. Not entirely, but for the most part, yeah, this is a much better model. Its peaks are not as high.
This is not the incredible revolutionary 3D capabilities that we saw with Astra. In fact, I would put it slightly below 5 .5 opus in most of those types of things. It is not as good at computer use as Astra, although it is close enough to the point where I have been happy using it for all of my day -to -day work.
I actually had 6 .1 Sol go through all of my emails and find invoices that I had forgotten to pay or was behind on, mostly like investing stuff, and set up new tabs in Chrome for every investment I needed to wire, fill out all the details for me and just leave me to hit send. It didn't get a single thing wrong and called out additional stuff that I absolutely would have missed if I was doing this work myself.
so i'm literally trusting this model to wire money for me it's trustworthy enough for that and honestly i don't know if i would have trusted astra with that due to the spikiness 6 .1 soul much much less spiky from what i've heard from the other testers they seem to agree with this analysis i know for a fact that julius and ben who have also been testing have had a much better experience with this than astra in terms of the spikiness julius called the model incredible ben called it incredibly boring i think that's the best place you can be for a model drop like this but as i had mentioned before its peaks are not as impressive while it does quality work the majority of the time there are some tasks that are just at the edge of its capability that it will start to do weirder stuff on for the most part it's fine but i i'm still reaching for opus a decent bit we'll talk more about the comparison later i do default to this model for a bunch of stuff though first off as i mentioned before computer use
I can't wait for ultra fast to be like an actual thing you can use with open AI models, because when it is, this model is going to be crazy on it because it can already figure out how to navigate computer use totally fine. If it can suddenly do it six times faster, it's going to be unbelievably fun. Still not quite as good as Astra, but more than good enough that for the price difference, I wouldn't even think twice about it.
But as I mentioned before, there are certain things I would still occasionally use Astra for that I am more than happy to use Sol for. One of those things is deep code reviews. I have still found OpenAI models and the like Rottweiler nature where they'll dig into a problem and shake it and tear it to pieces until they find every single thing wrong with it.
I find 6 .1 Sol to be incredibly capable in this particular way. So as you can probably guess, I had 6 .1 Sol do some deep audits on Orchestrator V2 and other parts of my real world code bases. In my Orchestrator V2 audit, it performed nearly identically to Astra.
I do believe it was slightly higher a score. Okay, not in this analysis, but in my other analysis, it did actually score very, very slightly higher, but it did it at about half the price, 297 versus 584. Sonnet was still cheaper and Opus was slightly cheaper as well.
The difference being neither of these models were anywhere near as thorough with their analysis. 6 .1 Sol, Doug, Deep. to find things which is why it was able to get a score comparable to astra although it did admittedly burn way more tokens another task i've had a lot of fun testing with is asking the model to find opportunities to improve a code base in this case to improve t3 code this is the one where grok 4 .7 scored strangely well of course astra scored way better at an 83 .8 versus the 80 .7 from grok 47 but gbd61 soul hit it out of the park with an 87 .4 I didn't save all the prices for these runs, it's been a bit, okay?
But for Opus 5 .5, it cost $5, and for Sonnet 5 .5, it cost almost $9. With GBD61 Sol, it was $2 .15. That's the difference.
This model's token price is cheaper than Sonnet, but its token utilization is still maintaining OpenAI's usual efficiency, which results in just crazy price -to -performance. This whole thread was particularly fun because I had Opus 5 .5 review this model with a different name, obviously, so I didn't know what it was. I went and edited the history after.
And it concluded very quickly this was a frontier tier model. Its reviews and bug repros match the fixes that later merged. Its first draft code had real bugs, which review bots caught.
Four reviewers are still running. The local key for code reviews back frontier tier again in a blind 10 model bench on the same prompt. 6 .1 Sol placed first of the A7 .4.
It found the fish slop runs and compared those two. It did say 6 .1 Sol's quality output was slightly below Astra's as well as the two OpenAI models with Opus 5 .5 and Sonnet 5 .5, which we will absolutely show you in a bit. But I do want to call out the price here because it only cost $7 to run versus $15 for Sonnet 5 .5 and $50 for Opus.
Opus's honest tier call was that this model is incredible for scoped work, top of the frontier. For find what's wrong and tell me the truth, I would choose it over Astra and about level with Opus. For a long unattended building, this was below Frontier.
Follows its process rules even when they stop all progress and it does not ask for help. This I absolutely noticed. I had mentioned before, well a few times now, that my TS Rust port that I'm making with Opus 5 .5 is going way better than when I was working on that same port using Astra and Sol in the past.
i had that port running for a while with this model and it made no progress it burned a shitload of tokens but it didn't actually improve the compiler at all opus was able to from scratch restart it and get it working in a day after i had spent months and hundreds of thousands of dollars in tokens with this model as well as with astra and five six soul opposite in like a grand in like a night with just two subscriptions with the cloud plan so for unattended long like heavy rewrite type stuff anthropic is just comically far ahead right now and it also didn't have great judgment when i was using it for managing my fleet and for those wondering my fleet is all the computers i use for running all my agents and code because one computer is far from enough i don't run any of them on this macbook now so when i use this model to manage the fleet it made a couple dumb mistakes here and there to be fair so is opus astra is the only one that hasn't really made too many of those dumb mistakes but like i'm gonna be so real i am entirely done using astra for this model
After I had Opus do all of this review, I asked it, how much does it think this model should cost? It guessed $5 per mil in, 50 cents cashed, and 30 per mil out, putting it at Opus's prices roughly. It said that because it's performing like Opus, its speed should add a premium because it is quite fast.
And it's not a pro model, which is where it expects those higher like $100 out tiered pricing things to come. Pro models aren't really a thing anymore. We just use Fable and Aster, but you get the idea.
This is the funniest part of the whole thread though. If OpenAI wants people to adopt it, I would expect $3 per mil in, 30 cents for cashed, and $20 instead. That would still be a fair price for what it does.
To which I responded, if I told you it was $2 in, $10 out, and 10 cents per mil cash read, what would you think? I'd call that very aggressive pricing. For how you use it, it costs about a quarter of what I guessed.
The cash price does most of the work. Agent workloads are 96 % cash reads. So 10 cents for cash reads matters more than the $2 and $10 headline prices.
When I looked at all of my sessions, its price guess would have been $5 ,700. But after looking at these new prices, it redid the math and it would have been $1 ,550. That is a massive decrease.
And for all my PR review type tasks, it was expecting those to be up to $10 and it's actually only up to $3. And that's for like heavy PRs with tens of thousands of lines of code. according to opus so don't blame me blame opus for saying this first off opus says it becomes the default model for scoped work second off it says that bloated system prompts barely matter anymore because of the cash pricing it's just noise third it says long loops are still a bad idea but not because of money it's because according to it the tsros port wasted four days and made no progress at all and opus even said they'd be suspicious of it lasting they expect this price to go up in the future i cannot fathom openai ever increasing the price for a model but opus thinking they will is hilarious and shows just how good a value the model is i love this call out here chibi to 6 .1 soul did three rounds of work for about half the cost of sonnet's single round a lot of this comes down to how context was managed both because 6 .1 soul is much more efficient so it's not doing as many calls that bloat the context it's not outputting as many tokens that are like building up over time so the average number of tokens being read per request
was only around 110 000 tokens versus 360 000 for sonnet 5. the result is that soul used under half as many input tokens as sonnet making this model significantly more efficient speaking of efficiency i want to talk about these deep swe scores a tiny bit more because this is a weird bench for me to have forked and include in these things i actually did it for a different reason not to compare against 6 .1 soul but to compare against a new release from open router JevRouter.
OpenRouter added JevRouter to try and optimize costs with your requests, and I thought it was an incredibly stupid idea. Once I started running it against benchmarks, I confirmed it's an incredibly stupid idea. It turns out a model that cannot reason, that is given a prompt and no context, cannot make a good decision around how hard the problem is.
And JevRouter ended up being DeepSeek v4 .1 FlashRouter for the vast majority of its runs. Around 60 % of all the requests went straight to DeepSeek 4 .1 Flash. So didn't like it that much.
It also routes to other smarter models, which should give it more of an advantage. But it ended up being more expensive than GPT -6 Astra was on low, while also taking four to five times longer because 6 Astra low took 4 .6 minutes and GevRouter took 20. GevRouter's average task took 104 steps, whereas GPT -6 Astra's took 19.
You get the idea. It wasn't very good. But the whole point of JevRouter is that it would be as cheap as possible to get a certain score.
That was the promise on the tin. Whether or not you believe them is up to you, not me. I think it's bullshit.
Regardless, JevRouter was routing to DeepSeek 4 .1 Flash for the majority of its requests. Despite JevRouter routing to the cheapest possible small openweight models from whatever provider will give it away for free, 6 .1 Solon Low got the same score for an eighth the price. OpenAI is here to destroy any wins anyone else has in terms of efficiency.
Completing this bench in 4 .8 minutes for 21 cents with the second highest score I've ever seen on it is a massive achievement. Tying Opus 5 .5, which took 50 minutes per task on max, a 10th the time and a 70th the price for the same score. If your work fits within the things 6 .1 Sonnet does well, you should probably use it for everything.
But if your work doesn't fit in it particularly well, you should probably keep using Opus and maybe give Opus the ability to call 6 .1 Sol when it should for various tasks. I'm almost certainly going to be setting things up so that Opus 5 .5 can call 6 .1 Sol to do investigation work, to try and like root cause bugs, to do analysis of code bases, to figure out what things need to be touched and why, to help me triage real world work, to help me review the work that Opus does and more.
I'm kind of spoiling the ending here, aren't I? I'm going to keep using Opus 5 .5 for now. Before I explain why, let me do the thing that I'm most excited for.
Fish slop! The first thing you might have noticed is the inclusion of slop in fish slop. This model did the horrible thing I hate, where it surrounded the game in a bunch of absolutely garbage UI.
And coming to this right after the 5 .5 Sonnet demo hurts me deeply. because Sonnet 5 .5 did not make graphics anywhere near this good looking, but at least it made a UI that was nowhere near this awful. And man, do I wish the bad UI is where the issue stopped.
I will turn on the sound. Oh God, it's blaring.
It is stunning looking. The fish are some of the best. The model for the sub is way better.
The propellers work way better. I'm going to mute the sound because that is looking pretty bad. I haven't even heard it, honestly.
But damn, like, looks beautiful. But if you actually are playing it, one of the first things you'll notice is that the movement feels significantly worse than it does in either the Opus or the Sonnet versions that I have demoed in the past.
Yeah, it moves jank. It also has a significantly worse frame rate than the versions from the other models. It does have higher graphic fidelity, so that makes sense.
Like, the models here with the plants are significantly better than they were with the Sonic version. The difference in the fidelity of the extras in the tank is absolutely hilarious. Like, yeah.
But goddamn, I'm so tired of the unnecessary text everywhere. This model does it worse than almost any I've ever seen before. Take a breather.
Paused. Your little world can wait. Back to the reef.
Start a new tank. Slop 01. Feeder submarine.
A little underwater chaos. The big little goal. Your little ecosystem.
Little fish become big earners. Four meals and they're all grown up. Make the family a little bigger.
A little golden overachiever. There's so many of these. There's like 20 plus of them.
And I promise you guys, as soon as I saw this, I took a screenshot. I sent it to OpenAI and I crashed out in the Slack because I cannot fathom how they haven't fixed this fucking problem. This model is unacceptably garbage at UI.
It has regressed again. And if you're looking for a model that can make front ends that don't suck, go spend your money somewhere else because it should not be spent here. This model sucks at front end.
It sucks at design. It has no taste. And you're going to have to bring your taste yourself still.
but it is admittedly really good at blender if you give it like a screenshot of a thing you want it to model in 3d and say hey you have blender over the cli go make this it will it'll do a pretty damn good job but i would never have it make the actual mechanics for my games because it feels awful to play it also has like nowhere near as much gameplay loop in fact the first time i tried demoing this before filming it just randomly game overed as i was like getting started in the first 30 seconds and never like said why actually i think i technically beat it i also want to like take this version and hand it to opus or sauna and say hey can you make this play better because the graphics are good but the game sucks but when you combine how cheap it was to make this because like this was five dollars i think to generate that's pretty insane and if you combine that with like ultra fast if that ever happens suddenly you're going to be able to make a game in a few minutes on demand we're actually now getting to that threshold where game development is about to flip upside down because of how models are finally understanding three -dimensional space and the tooling necessary to do these types of things it's happening
As per usual, I was not allowed to put the code I wrote with this model inside of T3 code or other open source projects during the testing window. So I had to use it exclusively on my internal projects like Lakebed, as well as for auditing other work, which means I mostly use this for auditing other work. And I was very impressed.
This is a real PR I was working on to fix a bug where my new little work tree like setup window that would appear in a new thread in T3 code would disappear if you left and came back. I had cloud code work on this, but this problem went pretty deep. So I wanted to make sure that whatever solution I came up with was very, very, very well vetted.
While I personally still do not trust this model to write the code I'm trying to land, I absolutely trust it to review things. Ignore the GPT -6 Sol there, just a placeholder. So when I had 6 .1 Sol look through this, it found real problems that were entirely missed by Fable and by Opus.
First, it called out that follow -up messages can stay blocked after the agent starts, which is very annoying if you want to queue a message. And also that recovered setup progress was disappearing too early. It figured all of these things out with a combination of reading the code and analyzing it, as well as computer use.
And it's able to prevent me from merging a real regression in T3 code. So I literally just copy -pasted those things to Claude and then told it to take another look. Said better, but I'll still fix two small gaps.
Remember, I can't use this model of the code for this project at the time. So I copy pasted that again over to Claude and it eventually got it good enough. And then I finally merged.
But that is what I like this model for. And I cannot wait to push its limits for actually coding. Although I will say from the code that I did have the misfortune of reading, it is harder to justify merging this code than it is for code from Opus.
Normally, I would make you guys wait for the Opus versus solo video or the Sonnet versus solo video, but I'll just spoil the details now. i like using them in tandem because i find soul to be way better at reviewing and digging into the details but i find opus a more pleasant collaborator and significantly better at actually implementing code without getting blocked constantly throughout its work and even now with the rust rewrite of typescript i find myself in a similar pattern where i have opus 5 .5 just going and going and going making the code base work and work well and then i had soul come in and do an audit and this is the funniest part I remember before I said that I had Sol and Astra working on that TS Rust port for effectively months.
I told Sol to come in and it got it from 83 .7 % to 100 % in under a day. I was blown away by that, that it had somehow unblocked the work that Astra and Sol were doing as well as 6 .1 Sol. And I was absolutely blown away by that, that it had taken the work that 5 .6 Sol, 6 .1 Sol, and Astra had done over months and got it unblocked where it had been stuck for weeks and finished it.
I was much more blown away when I had 61Soul take a look at that work and critique it. And what it brought up was that of the 1 .8 million lines of code, 1 .3 million were not being used. The reason was because Opus concluded all of the code from all the other agents was useless slop that had no chance of being recovered, and it chose to rewrite it from scratch itself in another crate.
so on one hand the only reason the code worked was opus but on the other hand the only reason the slop was still around was also opus so i had to have this model come in and clean up the mess that other open ai models had made because opus didn't even notice the mess was still there what i'm trying to say is this model absolutely has a place in your workflows it could probably even be your default coding model and you wouldn't have too many issues with it but i still find opus to be a better collaborator overall that said i have almost no reason to use sonnet anymore because this will effectively take its place and you bet your butt the moment this model comes out i'll be going and making adjustments inside of my cloud config because i already have it set up so that i can use soul inside of cloud code because i want to make sure opus knows this is the model to have review its work and investigate the things going on in the code base This is a damn good model, and I'm really happy to have it.
I wish we had something bigger, smarter, and more capable overall. I was really hoping for something to truly dethrone Opus 5 .5 as my daily driver. This isn't it, and I'm not planning on canceling any of my Clod subs as a result of this release, but I am planning on taking a lot more advantage of my Codex subs in my day -to -day work, admittedly in Clod code.
This is an awesome release and an unbelievable price for what you're getting, but this does potentially mark the start of the end of the subsidization era. So make sure you're subscribed so that you can be here when I cover all of that and more. God, I hope this doesn't get me canceled online.
I have no idea how others feel about this beyond like a handful of early access testers I've talked to. I legitimately don't know if people are going to love it or hate it or land somewhere between. I will know in a few hours, I guess, because it's a yeah, it's three in the morning.
I am going to go to bed now. This is a tiring one. Hopefully I did a good job.
Let me know in the comments. And until next time, peace nerds. God, I'm so dead.
The Hook

The bait, then the rug-pull.

The video opens mid-crisis: last week Anthropic's Opus 5.5 was so good it reportedly triggered an internal 'code red' at OpenAI, and two rushed model drops later, a third model, GPT-6.1 Sol, quietly leaked into the host's hands. What follows is an unfiltered, 2 a.m. field report on whether it's actually a threat.

Frameworks

Named ideas worth stealing.

16:40model

Opus 5.5's tiered verdict on GPT-6.1 Sol

  1. Scoped work (review, audits, debugging, ops and browser/email tasks) — top of the frontier
  2. Long, unattended building — below frontier; follows its own process rules even when they stall all progress, and doesn't ask for help
  3. Judgment managing real infrastructure/machines — below both Opus and Astra

Given a blind review of GPT-6.1 Sol's own work, Claude Opus 5.5 sorted its capability into three tiers by task type rather than handing back one aggregate score.

Steal forany model-comparison writeup: score capability by task type, not a single number
CTA Breakdown

How they asked for the click.

VERBAL ASK
30:27subscribe
“So make sure you're subscribed so that you can be here when I cover all of that and more.”

Soft ask folded into the closing verdict rather than a hard mid-roll pitch — lands right after he says he isn't canceling any of his Claude subscriptions, framing the channel itself as the way to track what OpenAI does next.

MENTIONED ON CAMERA
02:07toolDepot ↗
Storyboard

Visual structure at a glance.

cold open at the mic
hookcold open at the mic00:02
Depot sponsor read
valueDepot sponsor read02:08
the pricing reveal card
valuethe pricing reveal card07:56
Astra's smart/dumb tweet
valueAstra's smart/dumb tweet11:48
Fish Slop game demo
valueFish Slop game demo23:25
sign-off at the mic
ctasign-off at the mic30:46
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.

36:00
Theo - t3․gg · Review

GPT-5.6: The Review

Theo spends 36 minutes putting real numbers behind the GPT-5.6 hype — Sol, Terra, and Luna, benchmarked against Claude Fable, one blog chart at a time.

July 12th