Modern Creator
Theo - t3․gg · YouTube

OpenAI should be scared of this one

A Sonnet 5.5 review that says don't use it yourself. Let Opus call it.

Posted
today
Duration
Format
Review
educational
Views
64.1K
1.8K likes
Big Idea

The argument in one line.

Sonnet 5.5 is not worth picking over Opus 5.5 for day-to-day coding, but as a sub-agent that Opus or Fable calls for codebase audits it is the best value on the market.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You pay for Claude Code or the Anthropic API and want to know whether switching a workflow to Sonnet 5.5 will actually save money.
  • You run multi-agent setups and are deciding which model should handle codebase research, audits, and planning sub-tasks.
  • You've been burned by max reasoning effort bills and want the token-budget explanation for why.
  • You compare Anthropic and OpenAI model tiers for real code work rather than benchmark headlines.
SKIP IF…
  • You want a front-end design shootout. The review covers it briefly and the verdict is that Sonnet 5.5 is the weakest of the three Claude models for UI.
  • You're not running agentic coding tools. The whole cost argument hinges on cache reads dominating agent workloads.
TL;DR

The full version, fast.

Sonnet 5.5 looks cheap on the pricing table at $2 in and $10 out, half of Opus 5.5, but cache reads stayed at $0.20 per million while Opus and Fable got deep cache discounts, and cache reads are most of an agent's bill. In like-for-like coding, Sonnet ends up costing about the same as Opus, burns more tokens per task than any model on Cursor Bench, and is only nominally faster. Max reasoning effort makes it worse: it raises the floor on reasoning tokens by up to 15x, so avoid max and low on every Claude model. The one place Sonnet wins is as a tool. On a large codebase audit bench it scored slightly above Opus at half the price and in a quarter of the time. Use it as the research sub-agent Opus or Fable calls, not as the model you prompt.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:00 – 01:32

01 · Cold open: the model nobody expected to care about

Anthropic's run of releases, Sonnet 5 as a low point, and the setup: the benchmarks look bad, the host is impressed anyway, and this release is aimed squarely at GPT-6 Sol.

01:32 – 03:12

02 · Sponsor: CodeRabbit change stack

Sponsor read for CodeRabbit's new change stack view, which groups a PR's changes by intent instead of alphabetical files.

03:12 – 07:29

03 · What the announcement actually claims

Walks the official post: 30% faster and up to 30% cheaper than Sonnet 5, best Terminal-Bench 4 score to date at 70.6%, first Sonnet to beat Pokemon Red from screenshots, and fewer Claude-isms in its writing. Theory: Anthropic learned to distill its own models.

07:29 – 11:05

04 · The pricing catch: cache reads

Input and output are half of Opus, but cache reads are identical at $0.20. Since cache reads dominate agent work, Sonnet's real-world cost lands near Opus, and on the Artificial Analysis index near Fable 5.1.

11:05 – 14:39

05 · Max effort is a floor, not a ceiling

Reasoning levels are budgets. Max raises the minimum tokens the model must burn, up to 15x more than xhigh, and can hurt accuracy by forcing overthinking. The $7.60 max run versus $0.59 on medium.

14:39 – 18:51

06 · Benchmarks, token hunger, and the speed illusion

Frontier Code and Terminal-Bench charts where xhigh beats max, why low effort is also a trap, Cursor Bench showing 271,920 tokens per task, and a 94 TPS model that still finished the game demo 7 minutes slower than Opus.

18:51 – 20:46

07 · Design showcase: not the front-end model

Runs through Dara's Which AI Made This landing-page comparisons. Sonnet's designs are opinionated, sometimes hideous, worse than Opus and far behind Fable 5.1.

20:46 – 24:38

08 · The real use case: a tool for Opus to call

The turn. Sonnet should not be the model you select. On a six-model audit of the Orchestrator V2 PR it scored slightly above Opus at half the price and in about five minutes. That makes it the research sub-agent Opus and Fable should orchestrate.

24:38 – 26:22

09 · Behind the curtain: TypeScript in Rust

A side project rewriting the TypeScript compiler in Rust, a wrong-thread scare that turned out to be Sonnet reading context correctly, and 1.6 million lines of legacy slop deleted mid-filming.

26:22 – 29:20

10 · Fish slop: the demo that impressed

A 3D submarine game built from one prompt in about 40 minutes. 3.5x the input tokens of the Opus run at similar total cost, but one of the best results seen from any model, running at 120 FPS. About $16 at API prices.

29:20 – 31:46

11 · Subscription math and the verdict

The $200 plan works out to about $2,300 a week of usage. Final call: don't prompt Sonnet yourself, let Opus wield it for deep dives and hunch-checking. OpenAI has lost on smartest, cheapest, most efficient, and best orchestrator.

Atomic Insights

Lines worth screenshotting.

  • Sonnet 5.5 cache reads cost $0.20 per million, the same as Opus 5.5, so the headline 50% discount vanishes on agent workloads where cache reads dominate.
  • Cache reads are roughly 5% of Fable 5.1 spend, 20% of Opus 5.5 spend, and over 50% of Sonnet 5.5 spend.
  • Reasoning effort levels are budgets, not instructions: low to xhigh often changes token usage by only 5 to 8 percent on simple tasks.
  • Max effort raises the floor on reasoning tokens instead of the ceiling, producing up to 15x more tokens than xhigh for a single step up.
  • Forcing a model to keep reasoning past the point it is done makes it second-guess correct answers and lowers benchmark scores.
  • Sonnet 5.5 on max cost $7.60 on one bench where xhigh, high, and medium cost far less, with medium landing at $0.59.
  • Sonnet 5.5 used 271,920 tokens per task on Cursor Bench, more than Opus at 218k and almost double Gemini 3.8 Flash at 162k.
  • Sonnet 5.5 streams around 94 tokens per second on OpenRouter versus about 70 for Opus, yet finished the same game build 7 minutes slower because it wrote so many more tokens.
  • Sonnet 5.5 jumped Terminal-Bench 4 from 10.3% to 70.6%, the highest score posted, which says as much about the benchmark as the model.
  • Anthropic dropped the no-reasoning instant mode because these models fall apart without a reasoning budget, so low effort is also a trap.
  • On a six-way audit of a hundred-thousand-line PR, Sonnet 5.5 scored 7.81 for $1.24 while Opus scored 7.71 for $2.40 and Astra scored 8.22 for $5.84.
  • Sonnet finished that audit in about five minutes; Opus took nearly twice as long and Astra almost three times as long.
  • A full 3D submarine game built by Sonnet 5.5 would have cost about $16 at API prices, roughly the price of one Fortnite skin.
  • The $200 Claude plan yields roughly $2,300 per week of API-equivalent usage, close to $10,000 a month.
  • Sonnet 5.5 is the weakest Claude 5.5 model for front-end design, still ahead of OpenAI and xAI, but not the model to reach for on marketing pages.
  • OpenAI now trails on the smartest model, the most efficient model, the cheapest model, and the best orchestrating model at the same time.
Takeaway

Sonnet 5.5 is a tool, not your model.

WHAT TO LEARN

Half-price input tokens mean nothing when cache reads dominate your bill, so Sonnet 5.5 earns its keep only as the fast, cheap research sub-agent that Opus or Fable calls.

01Cold open: the model nobody expected to care about
  • Judge a new model by real-world cost per task, not by the input and output prices on the announcement page.
03What the announcement actually claims
  • Treat Sonnet 5 as a dead baseline: the 30% faster and 30% cheaper claims only look good because the comparison point was weak.
  • A record Terminal-Bench score from a mid-tier model is a reason to question the benchmark, not just to praise the model.
04The pricing catch: cache reads
  • Cache reads are the bulk of agentic spend, so compare cache-read prices across tiers before assuming a cheaper model saves money.
  • Sonnet's cache-read price matches Opus, which pushes its like-for-like coding cost up to Opus level and near Fable on the Artificial Analysis index.
05Max effort is a floor, not a ceiling
  • Reasoning effort settings are budgets that cap thinking; on simple tasks low and xhigh often differ by under 10% in tokens.
  • Max effort raises the floor on reasoning tokens by up to 15x and can lower accuracy by forcing the model to overthink finished work.
06Benchmarks, token hunger, and the speed illusion
  • Avoid low effort too: Anthropic removed instant mode because these models degrade sharply without a reasoning budget.
  • Check tokens per task, not just score per dollar; a fast model that writes almost double the tokens finishes later and costs more.
  • Medium effort is safe on Opus but noticeably weaker on Sonnet, so match the effort setting to the model tier.
07Design showcase: not the front-end model
  • Sonnet 5.5 is the weakest of the three Claude 5.5 models for front-end design; reach for Fable on marketing pages and Opus when you need steerability.
08The real use case: a tool for Opus to call
  • The right question for a mid-tier model is which tasks an orchestrator should delegate to it, not whether you should prompt it yourself.
  • Deep codebase audits and planning are where Sonnet wins: slightly higher score than Opus at half the price in a quarter of the time.
  • Build benchmarks around comprehension and planning tasks, because pure code-generation benches miss where cheaper models excel.
09Behind the curtain: TypeScript in Rust
  • When a model ignores an instruction, check whether it is following the thread's established context before assuming it lost your intent.
10Fish slop: the demo that impressed
  • Budget for token volume, not just price: Sonnet used 3.5x the input tokens of Opus on the same game build and landed at similar cost.
  • A complete 3D game for about $16 in API tokens resets what a prototype should cost.
11Subscription math and the verdict
  • Wire Sonnet in for investigatory sub-tasks such as confirming a hunch about an API so the orchestrating model feels faster and costs slightly less.
  • The competitive gap is now on every axis at once: smartest, cheapest, most efficient, and best at calling other models.
Glossary

Terms worth knowing.

Cache reads
Input tokens the API serves from a stored prompt prefix instead of reprocessing them. Agent loops resend the same context every turn, so cache reads become the largest share of the bill.
Cache writes
The one-time cost of storing a prompt prefix so later requests can read it from cache at a discount.
Reasoning effort
A per-request setting (low, medium, high, xhigh, max) that caps how many hidden thinking tokens the model may spend before answering. It is a budget, not a command to think harder.
Terminal-Bench 4
A benchmark that scores models on completing real tasks inside a command-line terminal, used as a proxy for agentic coding ability.
Cursor Bench
Cursor's internal coding benchmark that reports both score and tokens used per task, which exposes token-hungry models.
Artificial Analysis Intelligence Index
A third-party composite score that plots model intelligence against the cost of running its evaluation suite, so expensive runs show up even when scores are high.
Sub-agent
A secondary model instance that a primary orchestrating agent spins up to handle one scoped task, such as auditing a codebase, and then reports back.
Distillation
Training a smaller model on the outputs of a larger one so it inherits the larger model's behavior at lower cost.
Tokens per second (TPS)
Output streaming speed. A model with higher TPS can still finish a task later if it generates many more tokens to get there.
Resources

Things they pointed at.

05:20toolTerminal-Bench 4.0
17:00toolCursor Bench
18:55linkWhich AI Made This? (Dara's design showcase)
22:09linkT3 Code Orchestrator V2 PR
Quotables

Lines you could clip.

06:42
“Now these just feel like slightly dumber, way faster versions of Fable.”
one-line summary of the whole 5.5 family→ TikTok hook↗ Tweet quote
06:50
“It really feels like they saw everybody else distilling anthropic models, got jealous, and decided to do it themselves.”
spicy industry take with a clear image→ IG reel cold open↗ Tweet quote
12:22
“It's not really increasing the ceiling for how many reasoning tokens are allowed to be used. It's more so increasing the floor for how many have to be used.”
the max-effort explanation in two sentences→ newsletter pull-quote↗ Tweet quote
14:35
“Pretend it doesn't exist. It'll make your life much easier.”
blunt verdict on max mode→ TikTok hook↗ Tweet quote
18:00
“This is not a token efficient model.”
flat contradiction of the speed marketing→ newsletter pull-quote↗ Tweet quote
21:48
“We shouldn't be calling Sonnet directly. It should be used as one of many things that Opus or Fable will orchestrate when it's trying to break up real complex work.”
the thesis of the video→ IG reel cold open↗ Tweet quote
29:08
“For the price of a skin in fortnite you can make your own game.”
vivid cost anchor→ TikTok hook↗ Tweet quote
31:04
“This model is useful as long as you're not the one sending it prompts.”
paradox that begs a click→ TikTok hook↗ Tweet quote
31:10
“They don't have the smartest model. They don't have the most efficient model. They don't have the cheapest model. They don't have the best model for calling other models.”
the OpenAI verdict as a four-beat list→ IG reel cold open↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphoranalogystory
It's hard to ignore just how well Anthropic's been doing recently. From Fable 5 .1 being my favorite code model ever released, to Opus 5 .5 taking over pretty much everyone's workflows, my own included, I haven't chosen Fable for any tasks for over a week now, which is just like mind -blowingly crazy, to now with yet another release.
And this one seems strategically targeted to screw over OpenAI as much as possible. Sonnet is back, and 5 .5 is looking a hell of a lot more promising than Sonnet 5, which, if you don't remember, was one of my least favorite model releases, like... ever.
Sonnet 5 .5 is the model I did not expect to care about. I legitimately thought I would just kind of skip over this one and not talk about it because I've been so impressed with Opus's price and value and just everything I've been doing with it. But I was wrong.
Sonnet 5 .5 is an incredible model. I am actually very, very impressed with it, but not in the ways you might expect. Because when you look at the numbers, it doesn't seem that impressive.
On the Artificial Analysis Intelligence Index, it comes out as more expensive and dumber than Opus 5 .5 at every single level. So why am I so fond of Sonnet 5 .5? It's going to be hard to justify, but I think I can do it well.
The simple way of putting it is there's finally a cheap -ish model from Anthropic that makes sense in the modern era, especially when you compare it to offerings from OpenAI that have historically been closer to this price range. This model almost feels squarely targeted at OpenAI, specifically at GPT -6 Sol. There's a reason I'm filming this video and not a GPT -6 Sol video.
And it's not because Sol is so impressive, trust me on that. I'll explain what I mean and more after a real quick word from today's sponsor. Staying on top of what's happening in your codebase has never been harder to do.
AI has made contributing to most codebases way easier, but it's made keeping track of what's going on in them way harder. I can't tell you how many times I had some weird issue in a codebase. It was like, what the hell happened here?
Who would ever have merged this? And then went to the notes and saw that it was actually changes I made with my agent that I merged because I was blindly trusting the AI reviewers. Today's sponsor is one of those reviewers.
It's CodeRabbit. But I'm not here to talk about how good CodeRabbit's reviews are. Spoiler, they're very good.
I want to talk about a new feature they introduced, which is made actually keeping on top of the changes way easier. It's the CodeRabbit change stack.
Now when you have CodeRabbit set up on a code base, it will give you the simple button that you can click, even signed out, to see what changed in the PR in a way that makes way more sense. Instead of a pile of alphabetically listed files, you get a description of what changed and why. Instead of a jumbled mess of commits that nobody's gonna click through, you get this stack of the different things the PR changes with descriptions of what is changing and why.
And if you look at this, you can see pretty immediately how much easier it is to read through the changes that were made in this pull request. It even summarizes these different parts so you can understand what's going on in each of them.
In this part, it's the detection for the package install, so making changes to how T3 code is persisted on your machine. Here is where we fix the update logic that was breaking for a lot of users. There's even a security blast radius diagram that makes it way easier to see what changes are risky in a given PR.
I also really like the activity view, which is a super quick way to see what's changing in a PR to catch up without having to scroll through the mess that is the thread view in a given pull request on GitHub. Don't let slop take away your understanding of your codebase. Fight back at soydiv .link slash coderabbit.
As I was saying, the benchmarks don't necessarily look great depending on how you frame them, but I want to start with what Anthropic said with their official announcement post. This is the Claude Sonnet 5 .5 announcement. Introducing Sonnet 5 .5, the second model in the 5 .5 family.
The clear upgrade over Sonnet 5, it's 30 % faster and up to 30 % cheaper for most work. I don't love comparing this model to Sonnet 5, because on one hand it's not going to showcase the benefits that well, because like the speed difference and the cost difference isn't that big a deal. Sonnet was relatively cheap.
But also because Sonnet 5 was a garbage tier model and is not a good target to compare against. I said the same thing with the Opus 5 .5 coverage. It made no sense to compare it to Opus 5, which is why I'm thankful they compared it to Fable so much, but here, again, with Sonnet, because Sonnet's the smallest that they're working on right now, then there's Opus, then there's Fable.
Haiku is eventually going to happen, but we'll see. They said it's coming in the next few weeks. I didn't think Sonnet would be here so quick, and honestly...
I'm pretty impressed with it. Sonnet 5 .5 is faster, lower cost, as a complement to Opus 5 .5. Where Opus is built for complex work requiring careful judgment, Sonnet 5 .5 is strongest at well -scoped everyday tasks.
I will say it actually goes quite a bit further. Make sure you stay tuned for the fish slop demo. I'm very impressed with it, that's all I'll say for now.
It's also good for fixing bugs and creating polished documents, slides, and spreadsheets. I don't know how many of y 'all are actually creating polished documents, slides, and spreadsheets with your models, but if you are, let me know in the comments. actually curious and when you're on the way there if you hit that little red button it does help us out a bunch a lot of y 'all aren't subscribed and you'd be amazed at how much it helps the channel and it's also a great way to keep up to date and seeing that a open ai dev day is in not very much time in fact it may have already happened where you're going to want to be to keep up with how open ai is fighting back and trust me they're going to be fighting back anthropic also claims this model has a sharp eye for design this is an interesting call out because uh yeah I'll show you the designs in a bit.
It's not quite what I was hoping for there. And then they call it Haiku 5 .5, which is going to be built for high volume and cost sensitive applications. We'll be joining Claude 5 .5 family in the coming weeks.
So how does Sonnet 5 .5 improve over Sonnet 5? First off, it gets the best score on Terminal Bench 4 to date, which is kind of nuts because it's a Sonnet model, but it's a huge jump. It got a 70 .6 % where previously it got 10 .3.
like this just looks silly seeing sonnet 5 .5 as the highest score ever in terminal bench on one hand it makes me skeptical of terminal bench on the other just kind of skeptical of benches at this point it's really hard to measure the capabilities and the ones i have been doing on my own have on one hand felt even sillier but on the other hand have better reflected my vibes when i use these things so cover my benches in a bit because i was also pretty amused by the scores So it's seven times better a score for terminal bench, which means it's like actually viable for code.
It scored slightly below Opus 5 .5 on GDP Val, which is a bench I don't really care that much about. And it's strong on long horizon work and image understanding. It's the first Sonnet model to beat Pokemon Red reading and working only from screenshots.
I feel like a lot of models can do this now, but it is cool to see these like cheaper tier models being able to do a long running task like playing a game. Seems like Anthropix had a couple like crazy internal unlocks around their RL stuff recently because previously any model that wasn't their biggest and best didn't really get the right vibe is all I can say.
Like it didn't feel like they actually understood the work they were being given. It more felt like they were robots trying to operate in a very specific set of things. Now these just feel like slightly dumber, way faster versions of Fable.
And I did actually get some of that vibe from Sonnet, which is still crazy to me. It really feels like they saw everybody else distilling anthropic models, got jealous, and decided to do it themselves. And turns out they're good at it now that they've been doing it more heavily.
Another important piece, collaboration. Not like, how well do our agents collaborate with each other? It's more about how well can they write and format their outputs to us.
They have been taking on the, like, Claude -isms and all of the horrible Claude -ese stuff that everyone hates and doing everything they can to get rid of it. And as a result, Sonnet 5 .5 writes way more clearly than their previous generation of models. It feels like a better partner for collaboration than Sonnet 5.
The speed helps too, although I think they pushed the speed bit a little too hard here. The numbers are not as impressive as they make it sound. And now we have price.
Sonnet 5 .5 is priced the same as Sonnet 5 at $2 per million in and $10 per million out, as well as $0 .20 per million tokens for cash reads. But it typically needs far fewer tokens to do the same work. It's up to 30 % less per task than its predecessor.
Man, are there some rough edges to this statement. I'm not going to call it outright a lie, but it is intentionally dancing around some important pieces, in particular, max and cash token reads. As I've talked about extensively, cash token reads and cash writes tend to be the majority of costs for our day -to -day agent code work nowadays, which is why the 20 cents per million token cash read number here is a little bit concerning.
Because all of the other models that Anthropix put out recently, the two, Fable 5 .1 and Opus 5 .5, got huge discounts and cash read costs. This model didn't. When you look at the pricing chart at the bottom here, you'll understand why I'm concerned.
If you look at input token costs, Sonnet 5 .5 is $2 and Opus is $4. So it's half the price. Same with output tokens, $10 to $20.
And even the cash write costs is roughly the same factor here with $2 .50 for cash writes for Cloud Sonnet 5 .5 and $5 for Opus 5 .5. So why am I so upset? The first row, cash reads.
Previously, I was saying that cash reads are a very small percentage of my usage. That was the case for Fable 5 .1 because Fable 5 .1 dropped the cash read cost by 90%. Meanwhile, Opus also dropped it, but only around 60%.
Usually cash reads are 90 % off. So if it's $4 per million input tokens, it'll be 40 cents per million cashed input tokens. It's $2 per million input tokens.
It's 20 cents per cash million input tokens. Pardon me for using the Google AI summary here, but every website reporting on this pricing is garbage, in particular Anthropix. I don't know why they don't have a good model API dashboard like OpenAI does.
I hope that they can throw up a prompt quickly to build that for them because this is garbage. Fable 5 .1 costs $10 per million in and $50 per mil out, which is double the cost of Opus 5 .5. However, cash reads are 25 cents per million in.
Remember, all these other numbers for Fable 5 .1 are double, but for cash read, it's only 25 % more expensive. Everything else 2x, cash read 25%. So Sonnet to Opus doubles price for cash writes for input and output tokens.
Opus to Fable doubles again. But the cash read cost doesn't change between Sonnet and Opus, and it only goes up 25 % for Fable. What this means is cash read becomes a more and more prominent part of your cost when you go down the model tiers.
Where it rounds out to under 5 % with Fable, it's closer to 20 % with Opus, and it's closer to like 50 plus with Sonnet 5 .5. Because that number is disproportionately large when you look at the other numbers. It would have been really nice if they could have knocked cash read down to like 15 cents or God forbid 10 cents.
That would have been insane. But they didn't, which means that the costs for using this model don't end up being as much cheaper as you might think when you look at these numbers. because normal input tokens are barely touched with agentic work because we tend to read from cash, and output tokens are a very small portion of the overall costs.
I feel obligated to call this all out because I've looked at the numbers for doing like -for -like work with Opus versus Sonnet for real code tasks, and Sonnet consistently comes out as if not more expensive than Opus does. In fact, with the Artificial Analysis Intelligence Index, Sonnet 5 .5 costs roughly the same as Fable 5 .1 for real -world tasks.
Flashbang warning, since I know a handful of you like those, we're going to artificial analysis. Yeah, this is insane. Sonnet 5 .5 is neck and neck with Fable 5 .1 for the most expensive run they've ever done on artificial analysis.
That should kill any reason to use this model entirely, right? Kind of. It does kill one thing.
Give you a hint. It's one particular word you can see here. Starts with M and ends in Max.
Don't use it. I've been trying to explain why I hate max mode so much, and I don't want to make this whole video about it, but I'll do my best to explain here. Reasoning levels aren't really levels.
You're not saying you should reason this much. They're budgets. They are allowing the model to reason up to a certain amount.
so when you set low you're saying i want you to do this with a small amount of reasoning tokens when you set medium higher x high you're saying i'm okay with you using more reasoning tokens up to a certain point but you're not necessarily increasing the amount of tokens used i've had benchmarks where the gap between low and x high was like five to eight percent tokens like it's not a big difference if the tasks are simple but if the tasks are complex then having the extra budget can absolutely help So what is the issue with Max?
The issue with Max is that it's not really increasing the ceiling for how many reasoning tokens are allowed to be used. It's more so increasing the floor for how many have to be used. It's effectively telling the model it's not done until it does a certain number of reasoning tokens.
And the result is that I've had benchmarks where from low to X high, you get a 5 % increase in token usage. And then from X high to max, you get, and I'm not exaggerating, a 1500 % increase in token usage, literally 15 times more tokens for that one little bump at the end, because you're no longer letting it be done with easy tasks.
and this could often actually hurt performance because if the model is told to keep overthinking the thing it's going to overthink the hell out of it it's going to start second guessing itself and it's going to start getting wrong answers because you're forcing it to think too much it worked more like this where it slightly increased the floor i still wouldn't recommend it but it'd be fine but it's not it's forcing the model to do too much reasoning and i can prove this very easily so on a 5 .5 max reasoning effort was seven dollars and sixty cents If we compare it to something like, I don't know, Astra on Max, $3 .26.
So Sonnet was two times more expensive than Astra for the same bench. It sounds insane until you realize Sonnet 5 .5 can also be run on, I don't know, X High, High, God forbid, Medium. That looks a little less bad, right?
X High is still a bit more expensive than I would like, but it's pretty close to Astra's price. It's cheaper, but still more than I would want. But once you go down to high and medium, you have crazy low prices, $1 .08 and $0 .59 respectively for running the same benches.
Realistically speaking, though, if we go look at the intelligence scores, when you bump down these reasoning efforts, sure, X high now is roughly Astra level, but Sonnet 5 .5 was roughly three points higher than Astra. Do we actually think Sonnet's going to be that much better than Astra? Okay, I feel bad asking that because realistically speaking, Sonnet does have fewer...
dumb spikes that i get so frustrated with so i actually personally would take sonnet over astra for my day -to -day code work call me insane i would just use opus but yeah wanted to call this out because people are looking too closely at the max reasoning efforts and i really don't think anyone should use them like i have not been shown a good enough use case for why max effort makes sense for most users pretend it doesn't exist it'll make your life much easier Back to the benches quick before I start covering the actual intricacies of using the model in particular.
It's speed, which I don't love the way it's been reported on so far. Funny coming back here for the last rant, because as you can see on some benches like Frontier Code, they put two scores in because X High scored better than Max did. Yes, really.
Max put it below Sol. X High put it above Sol. Very interesting.
And GBD6 Sol is kind of just DOA, isn't it? I've... barely had any reason to talk about it and haven't put it in videos for a reason.
It's just not that impressive to me. Hopeful, fingers crossed, we'll get something better in the near future. Everything else here looks pretty good.
Computer use, it's scoring much better than it did before. Still neck and neck with Opus 5 .5. I wish we had the numbers for OS World 2 .1 for GPT -6.
Not just Sol, but Astra, because I still personally use Astra almost exclusively for computer use. Occasionally a review here and there, but even then, iffy. I think it's kind of insane of them to put this chart as the first chart showing performance of the model inside of their reporting.
First off, it kind of shows Terminal Bench isn't the greatest measure because Opus 5 .5 dipped with Max. But also, much worse, shows that Sonnet 5 .5 is more expensive than Opus on Max. It somehow is outperforming Opus.
This is the type of chart that you would see and think something went wrong. not the type of chart that you would publish as the first chart in the blog post but sure kind of sad that the only place it is outperforming opus for the cost is in the max reasoning effort which again you shouldn't use over to frontier code you can see the same pattern where it plummets on max but does decently on x high and high as well medium and low are pretty big drops and similar to how i feel about not using max i don't think you should use any anthropic model on low right now they actually stopped providing the option to do no reasoning on opus and i'm assuming now on sonnet as well previously there was a no reasoning option where it just start responding immediately like an instant mode they don't ship that anymore because these models suck without reasoning And if they're given too little budget to reason, they will suck even harder, as they have proven to in benches like this.
So generally, avoid low, avoid max. And for the most part, medium's fine -ish. Medium's a lot stronger with Opus than it is with Sonnet, so keep that in mind.
Cursor bench was fun. I can look at the official numbers for it here, where Sonnet did outperform with high, kind of. And it does complement the Opus 5 .5 curve relatively well.
But again, the drop from medium Opus to low is just too big. And Opus is already such a surprisingly good value that it's hard for me to justify using much else. I also can't help but notice that all of these are score to cost and none of them are score to number of tokens.
There's a reason for that. We hop back over to Cursor's Bench and click tokens. you'll see that Sonnet does have yet another greatest of all time score, their token usage in CursorBench, where they did 271 ,920 tokens per task, putting it ahead of even Opus at 218k and Gemini 3 .8 Flash at 162k.
It's almost double the number of tokens to Gemini 3 .8 Flash, which is insane. It is more than 5x the number of tokens that GPT -56 Sol and also 6 Astra used. So keep that in mind.
This is not a token efficient model. which should be made up for by the speed right because it's so fast more bad news sadly according to open router the tps for using sonnet 5 .5 through anthropic is around 94 tokens per second which sounds insane when you're used to something like astra going at 30. they also report opus 5 .5 at around 70.
for what it's worth for my numbers using things through the official subscriptions i have seen around 100 tps on average for opus 5 .5 and around 150 for sonnet 5 .5 so it is faster but it's also not very token efficient and result here is that when i used it for something like fish slop it ended up taking quite a bit longer than opus did sonnet took 43 minutes to finish and opus only took 36 and when you measure how much time they spent generating tokens it's an even bigger gap from 39 minutes to 27 minutes because again this model uses a lot of tokens Let's take a quick look at Witch .ai, the design showcase made by Dara that has been very useful to see the capabilities of these new models when they drop.
For reference, I've also opened up Fables and Opus' runs. Let's take a look at the first Sonnet generation. This one is interesting.
It turned the page into like a fake app with the little things on the side to feel more like what the product might feel like. It's a unique style and that's something I've seen the other models do. Not bad, but a bit opinionated.
Let's see what else we got here.
This one's a bit rough, and it has some issues with the scroll, too, where it just rotated from today to last month back to today. These little lines in the background aren't my favorite thing. They hurt the readability.
Don't love it. Next, we have this card design, and it's hideous. We got this highlight -y one.
Pretty boring. Don't love it. The animations are awful.
Then here we have a weird... hierarchy of like a geologic i don't know the term for this type of like cross -section but yeah not great compared to opus definitely a little worse but compared to fable significantly worse i still think fable 5 .1 is the best overall design model in particular for nice looking front -end designs for your marketing site and whatnot But I have found Opus to be relatively steerable towards good designs, and it follows instructions around design much better.
I've not pushed the limits of Sonnet for design personally very much, but from all of the demos I've seen here, I'm not very impressed with its front -end capabilities. Still way ahead of anything OpenAI and especially anything that XAI has, but not the model I'd reach to for front -end. Especially since in like -for -like work, Opus often ends up being around the same price.
If you remember the intro of this video, you might be a bit confused at this point because it doesn't seem like this model is all that great. it's bad at front end it uses too many tokens it's faster but actually slower because of the token differences and it doesn't seem meaningfully smarter or cheaper than opus and day -to -day work so what the hell do i like this model so much for well to be frank i don't think this is a model that you or i should be selecting if you are presented options inside of cloud code between fable opus and sonnet really don't think many people should be picking sonnet i can already tell how the comment section is going to look after i said that Well, Theo, not everyone can afford Opus.
Sonnet's more expensive. Shut the hell up. Seriously, I don't want to have that argument today.
We're talking about how these things operate in our actual real world usage. That's why I'm excited to say for a handful of types of tasks, Sonnet does prove to be meaningfully more cost effective than Opus. Not necessarily the types of tasks I would send a model after, but absolutely the types of tasks that Opus and Fable would.
Its strengths come as a tool for our other models to use. We shouldn't be calling Sonnet directly. It should be used as one of many things that Opus or Fable will orchestrate when it's trying to break up real complex work.
I kind of made a bench for this for my Grok review, where I was trying to find the strengths of Grok 4 .7. And I did it by having all of these models do a really big, deep audit of T3 code, specifically this giant orchestrator V2PR that's been iterated on for far too long. It's hundreds of thousands of lines of code.
I gave all of these models a pretty detailed prompt, asking them to break up all the things being added to Orchestrator V2 to propose a strategy where we can get these things landed into main sooner rather than doing it all through a single giant PR. And I found that at the time, Astra was by far the best at doing these breakdowns.
Fable was not great at it. Grok actually was outperforming Fable. And then Opus, Sonnet, and Soul didn't do particularly well.
I have overhauled this since and updated the way that it is judged, and those changes have brought Fable up to a meaningfully higher score, still below Opus and still far below Astra, which by far had the best plan on how to do these things. I still find that Astra is just like uniquely willing to dig into the details for things, so it performed really well here.
I think that came as a massive surprise to me was Sonnet's performance, where it ended up being around half the price of Opus, you know, what it should be, and also performing slightly better than opus did according again to my automated judging panel that goes through the changes and proposals to make decisions on various axes so remember this isn't traditional code work this is a deep dive type task where the model has to go through a large code base and figure out what can be changed in it and how to explain it to someone else this is the type of task that requires going through a ton of different things to come back with good information is not testing the coding capabilities in traditional sense it's much more analytical i guess where it's like trying to make good architectural decisions and comprehend what's going on in a code base so the cost per point here is insane it is like the best value i've seen in this bench by far more importantly in my opinion is how much time it took because it ended up only taking around five minutes to do all of this work and like yeah sure soul was able to do it even faster but at like half the score
Opus took almost twice as long and Aster took almost three times as long. So Sonnet is a tool that your agents can call on to do this type of research, to help it plan and scope work. If you give Opus or Fable a big task and they want to analyze the code base before starting, they can now call on Sonnet to do that and get results that they're more than happy with.
This is huge and it's one of the biggest strengths and I hope others start to make benchmarks like this. because it is such a good way to see these capabilities and to see the strength of what Sonnet is introducing here. A little peering behind the curtain here.
As you can guess, I have been testing a ton of models over the last few weeks, honestly, and it's been chaos. One of my bigger tests is that I've been working on rewriting all of TypeScript, like the compiler for the language in Rust. There's already a Go rewrite by Microsoft that's really good, but I wanted to see how much further I could push and also potentially get it working in Wasm.
This was meant to be a Hail Mary project that would never happen, but Opus has fully unblocked and has it going really, really far. But one of the things I noticed is that it left over a ton of the slop that Astra and Sol had made, like millions of lines of it, and it wasn't getting rid of it. So I told it to halt and go clean up all of the legacy slop.
And then I had a thread over here in T3 code where I'm consistently monitoring changes as they come in. And this thread is one where I found all these legacy things that need to be deleted. So I asked for an update.
How about now? We should have lots of legacy stuff cleared out. And I scrolled.
None of this is talking about how much legacy stuff was deleted. And I was really confused. I was like, what the hell?
We deleted a bunch of code. Why aren't you telling me about it? Is this a regression in Sonnet 5 .5's behavior?
Maybe it's not as good as Opus at understanding my intent. nope i was in the wrong thread this is one where i was always talking about numbers so this is actually a really good thing is that it was able to make what i would consider a pretty logical decision based on what the contents of the thread are it was continuing to operate the way the thread had instead of trying to figure out what the hell i meant when i said should have lots of legacy stuff cleared out it just did i would argue the right thing here which is nice i was about to crash out about it not understanding my intent but it totally does i asked it how much code has been deleted it's about 1 .6 million lines of code have been deleted since i started filming so i just kicked this job off before filming funny enough so yeah the real reason i came here was to grab my fish slop thread both to see how everything came out in terms of cost but more importantly to show you guys the actual demo ended up costing roughly the same as the opus one did but sadly the opus one did actually use a fable 5 .1 sub agent at some point so it's score
isn't necessarily the most accurate i guess i do have to rerun opus 5 .5 on fish slop at some point to get even better numbers there but the input token difference is insane sonnet55 used way more like more than three and a half times more input tokens than opus did and the resulting cost is actually quite similar although i would guess if i had not accidentally spawned the fable sub agents opus would have even been cheaper how are the results it's one of the best i've ever seen There are some ways where it is worse, like some of the bottles are just not as good as the models that we were getting with Opus.
But it is still, without question, one of the best I've ever seen. I would argue in some ways it is better than the Opus one. And in all ways, it is better than everything I've gotten out of Astra and obviously out of every other lab.
I remember just like two or three months ago being so impressed that KamiK3 could use Blender at all. And here I am now with like a bunch of... actual like 3d models that you can tell what they are properly it's not even like bad like these fish are kind of cute and have their own like little cute unique design to them the coral the seaweed all of this is like decent and i just realized none of the sound is coming through so let me figure out why that is quick one sec a few audio issues later but now at the very least i should actually be able to hear and you can hopefully too one of the things i can't help but notice it's like all the android models have this but sonnet especially does It just feels good to play.
Like, you're seeing a 30 FPS YouTube video of this. I might have to start uploading these for y 'all to try, though, because this is running at a buttery smooth 120 FPS on my laptop. And, like, all the movement and all the mechanics actually feel pretty solid and balanced.
I actually love the animation for the fish, too. It's adorable. A little pelican there.
The craziest thing is the aliens. Where are they? It says that there's an attack.
Where is he? Yeah, there he is. The alien is the best looking I've seen by far in any of these demos.
Ugh. How crazy is it that I threw a prompt at a model and then 40 minutes later for a few dollars it spit this out? If I had paid API prices, it would have been $16.
I've paid more than $16 for games worse than this, shamefully enough. I was a kid at some point. I bought...
crappy playstation games that were made after like the movies that i was watching we've all been there right but like god damn for the price of a skin in fortnite you can make your own game and that's assuming you're paying the api prices obviously this is super super cheap if you're doing it over your subscription which uh by the way because i know a lot of people are always curious what the numbers look like there i did actually run some cost breakdowns for my own usage across my now six clot accounts And from my rough analysis of my real accounts that I emptied, which I killed three accounts in the past few days, you get around $2 ,300 a week of usage on the $200 plan.
That puts you at almost $10 ,000 a month for 30 days with your cloud code sub on the $200 plan. That's going to stretch you pretty far with these models. While it might not look like Sonnet's going to get you much further per dollar than Opus, That is the case for general work.
If you're having it as the model you select to use for things you do, sure. But if you have it as a tool Opus calls to do things like deep dives into your code base to find specific behaviors or characteristics or confirming a hunch it has about something that you found in an API or all of those types of investigatory things, it is really cheap and surprisingly fast too.
I would bet that if you set Opus up properly to use Sonnet in the right times for these sub -agents, that the result will be Opus feeling way faster and getting real work done for slightly cheaper. That sounds like a pretty good deal to me. So while I don't think you should actually use Sonnet 5 .5 yourself, I think it is an incredible addition to this new family of models, and I'm excited to see how Opus chooses to wield it.
I will not be using this model much going forward myself, but I do hope that my Opus is able to because there is clearly value here. While not in front -end building work or day -to -day coding tasks, there is a ton of value in how this model can be utilized to confirm things in real -world codebases, and I would assume for real -world documents, Office, and all that type of stuff too.
That's not what you guys are here for. You're here for the code. And I'm hyped to say this model is useful as long as you're not the one sending it prompts.
But goddamn, OpenAI really needs to respond to this. They have lost in all of the places they are strongest. They don't have the smartest model.
They don't have the most efficient model. They don't have the cheapest model. They don't have the best model for calling other models.
It's... rough for open ai right now and i can't wait to see how they catch up because right now it just doesn't feel like a good value the 200 codex plan gets me nowhere near as much usage and nowhere near as much real world code as i'm getting out of my 200 clawed subs so take that as you will fingers crossed we got some fun announcements coming i have a feeling that things are about to speed up not slow down hopefully this is a useful breakdown you can better know how to wield this model until next time peace
The Hook

The bait, then the rug-pull.

The title promises a threat to OpenAI, and the first ninety seconds set up a puzzle: the host admits the new Sonnet is dumber and more expensive than Opus on every chart, then says it is one of the most impressive releases anyway. The rest of the video is the argument for why both things are true.

Frameworks

Named ideas worth stealing.

10:02model

Cache-read share by model tier

  1. Fable 5.1: cache reads under 5% of spend
  2. Opus 5.5: about 20% of spend
  3. Sonnet 5.5: 50% or more of spend

Input and output prices double at each tier but cache-read price barely moves, so the cheaper the model the larger the slice of your bill that never got discounted.

Steal forany cost model for agent workloads where the same context is resent every turn
11:36concept

Reasoning effort is a budget, not a level

Low through xhigh cap how much the model may think. Max flips it into a minimum, forcing tokens to be spent even when the task is done, which costs 15x and can lower accuracy.

Steal fordeciding effort settings on any reasoning model
21:45concept

Sonnet as a sub-agent tool

Don't select Sonnet as your model. Let Opus or Fable call it for scoped research: codebase audits, confirming a hunch about an API, planning and scoping before real work starts.

Steal formulti-agent Claude Code setups
21:57model

Orchestrator V2 audit bench

  1. Give six models the same huge PR
  2. Ask each to propose how to land it in pieces
  3. Judge the proposals with an automated panel
  4. Plot score against API cost and wall-clock time

A comprehension-and-planning benchmark rather than a coding one. It surfaced Sonnet's strength where traditional code benches missed it.

Steal forevaluating models on analysis tasks instead of only code generation
CTA Breakdown

How they asked for the click.

VERBAL ASK
04:38subscribe
“if you hit that little red button it does help us out a bunch a lot of y'all aren't subscribed”

Buried inside the announcement walkthrough, tied to a promise about covering OpenAI Dev Day. Casual and quick, then straight back to content.

FROM THE DESCRIPTION
Storyboard

Visual structure at a glance.

open
hookopen00:00
the chart that says no
hookthe chart that says no00:59
announcement
promiseannouncement03:23
pricing table
valuepricing table08:32
tokens vs effort
valuetokens vs effort12:54
terminal bench
valueterminal bench15:41
design showcase
valuedesign showcase20:50
audit bench
valueaudit bench22:26
fish slop
valuefish slop27:03
subscription math
valuesubscription math29:28
verdict
ctaverdict31:34
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.

Video of the Day32:54
Theo - t3․gg · Talking Head

Fable is Mythos, and it is really good.

A 33-minute first-take from a developer who spent $3,000 on inference in 24 hours — benchmarks, real demos, session math, and the hidden safety intervention that silently degrades the model without telling you.

June 11th
36:00
Theo - t3․gg · Review

GPT-5.6: The Review

Theo spends 36 minutes putting real numbers behind the GPT-5.6 hype — Sol, Terra, and Luna, benchmarked against Claude Fable, one blog chart at a time.

July 12th
29:32
Theo - t3․gg · Essay

The weird situation with Fable

Theo breaks down how Anthropic silently modified prompts, rewrote its system card, and built invisible safeguards into its most capable model - then got caught.

June 15th