Modern Creator
Nate Herk | AI Automation · YouTube

I Tested Sonnet 5.5 vs Opus 5.5: What You Need to Know

Seven identical prompts, two Claude models, and one real question: when does paying double for Opus actually buy you something?

Posted
today
Duration
Format
Review
educational
Views
73.7K
804 likes
Part of the collectionThe Claude Opus 5 PlaybookEvery Opus 5 breakdown, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

Claude Opus 5.5 costs roughly double Claude Sonnet 5.5, but across seven real work tasks it didn't win twice as often or cost twice as much in practice, so the deciding factor is whether a task has a clear, checkable definition of done rather than which model is 'smarter.'

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You use Claude Code or the Claude API for real client or business work and want a cost-based rule for choosing between Sonnet and Opus.
  • You run repeatable workflows or skills and want to know which tasks tolerate the cheaper model.
  • You're deciding whether Opus's extra cost is worth it for open-ended, creative-direction tasks versus well-specified ones.
SKIP IF…
  • You only use Claude through a consumer app with automatic model switching and never think about per-task cost.
  • You're looking for a formal benchmark comparison rather than a hands-on workflow test.
TL;DR

The full version, fast.

A creator ran seven identical prompts, a landing page, a motion showreel, a 90-day plan, a spreadsheet and pitch deck, a skill-based explainer, a scraped news dashboard, and a resource guide, through both Claude Sonnet 5.5 and Opus 5.5, then compared the outputs, timing, and token cost. Opus is priced at double Sonnet on both input and output tokens, but across the seven rounds the cheaper Sonnet won four and Opus won three, and the real cost gap came out to about $16 total rather than a clean 2x per task. The pattern: prompts with an objective, checkable definition of done favored Sonnet; vague, open-ended prompts that needed judgment favored Opus. When outputs were functionally equivalent, cost and speed decided the round.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:00 – 00:41

01 · Cold open

Sonnet 5.5 launch pitch, then the real premise: seven identical real-work prompts run on both models to find when each one is worth using.

00:41 – 02:07

02 · Pricing and choosing your model

Opus is exactly double Sonnet's price per token. Anthropic's own guidance splits tasks by whether they're well-scoped and checkable (Sonnet) versus needing careful judgment (Opus); the host reframes this as 'does the task have a definition of done.'

02:07 – 07:13

03 · Website design comparison (Perkform landing page)

Same landing-page brief to both models. One invents a stylized 3D product render with flashier scroll effects; the other uses the real product photography and tells a clearer story. Sonnet costs $6.78 and Opus $13.71; despite liking Opus's version better overall, the host gives the round to Sonnet since a cheap re-prompt could likely fix Sonnet's one weak section.

07:13 – 09:57

04 · Motion design comparison (Glaido showreel)

A branded motion/sound-design showreel task. Opus wins clearly on taste, animation, sound design, and even pulls a live quote from the company website. Sonnet: $14.37; Opus: $20.55. Round goes to Opus.

09:57 – 14:15

05 · 90-day action plan (How They AI)

A deliberately vague prompt to plan a podcast's growth to 1,000 subscribers. Opus picks a more realistic launch window, clearer monthly phases, and links to an actual task-level write-up. Sonnet: ~$4.77; Opus: ~$9.25. Round goes to Opus for handling an underspecified goal better.

14:15 – 17:43

06 · Spreadsheets and investor decks (BrightPath Analytics)

Mock company data turned into an Excel model and investor deck. Both models produce formula-driven spreadsheets and comparable 18-slide decks; quality is judged nearly identical. Opus is actually cheaper and faster here ($8.91 vs $9.27), so it wins outright on cost with no quality tradeoff.

17:43 – 20:24

07 · Visual explainers built with a skill (What AI can my PC run)

Same HTML-explainer skill, same prompt on local AI hardware requirements. Opus jumps straight into specs; Sonnet first explains what 'local' means before applying it, and is judged more visually polished but less clear. Sonnet: $1.58; Opus: $3.16. Round goes to Sonnet on the theory that a follow-up prompt would close the explanation gap cheaply.

20:24 – 23:16

08 · Building an AI news dashboard

Open-ended prompt to scrape X for AI news and build an interactive tracking dashboard. Opus's version is judged more visual and informative. Sonnet: $2.32; Opus: $8.23. Despite the quality edge, the round goes to Sonnet because the cost gap outweighs the visible difference.

23:16 – 24:49

09 · Creating a YouTube resource guide (GPT-6 Astra)

A well-defined resource-guide skill applied to the same source video. Outputs are nearly identical in structure and length. Sonnet is about twice as fast and $1.40 cheaper ($1.47 vs $2.87), so it wins on speed and price alone.

24:49 – 25:48

10 · Final results and takeaways

Final tally: Sonnet wins four rounds, Opus wins three. Across all seven runs, Opus ran about 46 minutes longer and cost about $16 more overall, less than the 2x-per-task gap the sticker pricing implies.

Atomic Insights

Lines worth screenshotting.

  • Claude Opus 5.5 is priced at exactly double Claude Sonnet 5.5 on both input tokens ($4 vs $2 per million) and output tokens ($20 vs $10 per million).
  • Across seven head-to-head workflow tests with identical prompts, the cheaper model won four rounds and the pricier model won three.
  • The deciding factor across every round wasn't the task topic, it was whether the prompt had an objective, checkable definition of done.
  • A vague, single-sentence prompt with no defined output format consistently favored the pricier model; a specific, detailed prompt consistently favored the cheaper one.
  • When two models produce functionally similar output, like a spreadsheet with the same formula structure, the tiebreaker becomes cost and speed, not perceived quality.
  • The measured total cost gap across all seven real tasks was about $16, well under the roughly 2x-per-task gap the list pricing implies, because the pricier model didn't always use more tokens.
  • A reusable, well-defined skill narrowed the quality gap between models more than switching models did: with a shared skill in place, both models produced nearly identical documents from the same source video.
  • For a subjective, taste-driven task like motion and sound design, the pricier model won decisively and the gap wasn't close.
  • Given an identical prompt and access to real brand assets, one model still generated a stylized reinvention of the product instead of using the provided real images, trading brand accuracy for visual flash.
  • The cheaper model was judged the winner on a spreadsheet-and-investor-deck task specifically because it was both cheaper and faster while the content quality was rated equivalent.
  • Re-prompting the cheaper model to fix one weak section was judged cheaper than paying double for the pricier model's full run, so a single flaw doesn't automatically hand the round to the expensive model.
Takeaway

Cheaper doesn't mean worse, it means different terrain.

MODEL CHOICE

Across seven identical-prompt workflows, which model won came down to whether the request had a clear definition of done, not which model costs more.

02Pricing and choosing your model
  • Opus 5.5 costs exactly double Sonnet 5.5 on the API: $4 vs $2 per million input tokens, $20 vs $10 per million output tokens.
  • The practical test is whether a task has an objective definition of done: if you can already say what 'done' looks like, the cheaper model usually gets there for less; if you need a thought partner to help define done, the pricier model earns its premium.
03Website design comparison
  • Given an identical prompt, a model can default to a stylized reinvention of a product instead of using real provided brand assets, trading accuracy for visual flash.
  • Picking the cheaper model doesn't require it to be flawless: a follow-up prompt fixing one weak section can cost far less than paying double for the whole run again.
04Motion design comparison
  • For a subjective, taste-driven task like sound and motion design, the pricier model can win outright, and the win isn't close.
  • The actual cost gap for a given task can be smaller than the pricing tiers suggest, since the pricier model doesn't always use more output tokens even when it runs longer.
0590-day action plan
  • A deliberately vague, single-sentence prompt is exactly the case where the pricier model's judgment pays off, more than a detailed prompt would.
  • With a detailed, specific prompt, models tend to converge on a similar result, so the extra spend only pays off when the request itself stays ambiguous.
06Spreadsheets and investor decks
  • For structured deliverables like spreadsheets and slide decks, both models can produce comparable formula-driven, functional output, so quality alone won't separate them.
  • When outputs are functionally equivalent, cost and speed become the deciding factor, and the cheaper model can win without any real quality tradeoff.
07Visual explainers with a skill
  • Even with a shared skill and identical prompt, two models can diverge in structure, one explaining a concept first, the other jumping straight to specifics.
  • A clearer explanation doesn't always come from the more expensive model; visual polish and explanatory clarity can trade off independently of price.
08AI news dashboard
  • Given an open-ended dashboard-building prompt with no defined schema or views, the pricier model's judgment produces a more informative layout by default.
  • A dollar gap between two outputs doesn't automatically track with a large enough quality gap to justify the extra spend, even when one output is visibly nicer.
09YouTube resource guide
  • A well-defined, reusable skill narrows the gap between models more than switching models does, once a shared skill is in place, two different models can produce nearly identical documents.
  • When outputs converge this closely, speed and price decide the round outright rather than any perceived quality difference.
10Final results and takeaways
  • Across all seven runs, the cheaper model won four rounds and the pricier model won three, even though the pricier model is priced at roughly double.
  • The measured total cost gap across seven real tasks was about $16, not the roughly 2x-per-task gap the sticker price implies, because token usage didn't scale with price the way the rate card suggests.
Glossary

Terms worth knowing.

Effort level (high)
A setting that controls how much reasoning a Claude model applies to a task. Both models in this test were run on the same 'high' effort setting to keep the comparison fair.
Definition of done
A clear, checkable description of what a finished task should look like. Tasks with one tend to favor cheaper, more literal models; tasks without one tend to favor models built for open-ended judgment.
Cache write / cache read pricing
Discounted per-token rates a model charges when it reuses previously processed context instead of reprocessing it from scratch, shown alongside standard input and output pricing.
Cloud artifact
A shareable, hosted output a model can generate directly, like an interactive document, instead of only returning a local file.
Skill (Claude skill)
A reusable, saved set of instructions and style rules attached to a prompt so a model repeats a specific output format or design approach consistently across runs.
Resources

Things they pointed at.

Quotables

Lines you could clip.

01:56
“Sonnet 5.5 fits best when the task has a clear spec and a way to check the result.”
Concise thesis statement, reads directly off Anthropic's own guide and anchors the whole video's argument.→ TikTok hook↗ Tweet quote
17:43
“If you know exactly what you want, Sonnet is probably going to be able to do a good job for you. But if you need the creativity and you send an open-ended, very vague goal, Opus is just going to handle it better.”
The clearest plain-English statement of the video's rule, works standalone with zero setup.→ IG reel cold open↗ Tweet quote
24:59
“It didn't cost double the cost in this experiment, which is also interesting.”
Undercuts the assumption from the pricing chapter with the actual measured result, good pull-quote for a cost-focused newsletter.→ newsletter pull-quote↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphoranalogy
Cloud Sonnet 5 .5 is here and it is a clear upgrade over Cloud Sonnet 5, running 30 % faster as well as costing 30 % less for most work. So I've been playing around with it all day on my actual real workflows, things like building websites, taking a bunch of messy data and turning that into spreadsheets, also Excel sheets with different views as well as formulas inside of every single cell, motion design and editing videos, interactive dashboards that have real -time data syncing, project management and planning, and so many other things as well.
But specifically, I've been testing Sonnet 5 .5 against Opus 5 .5 so you know exactly when you should be using which model. Now in all these comparisons, I'll go over how fast they are, the input and output tokens, and how much they actually cost you in dollars. So let's not waste any time and just get straight into this one.
Okay, so before we jump into all of the different demos, I just wanted to lay a foundation here, which is how much does Sonnet 5 .5 actually cost versus Opus 5 .5? The answer is roughly half. So here you can see Sonnet 5 .5 is $2 per million input tokens, whereas Opus is four and Sonnet is $10 for the million output tokens, whereas Opus 5 .5 is 20.
So Opus is roughly double the cost of Sonnet 5 .5. So we're here to see if Opus is double the quality of Sonnet 5 .5 and where to use each model. Because Cloud actually came out themselves and dropped this little guide that said, hey, here's when you should use each model.
If you have something well -scoped for everyday coding, fixing bugs, Sonnet, high volume, everyday development. Sonnet. Polished documents, slides, spreadsheets, one -pagers, Sonnet.
You should basically be using Opus if it's complex work requiring careful judgment, including long horizon, agentic coding, and knowledge work, and for your hardest problems where you need the most intelligence, Opus 5 .5. I really like this section down here. Sonnet 5 .5 fits best when the task has a clear spec and a way to check the result.
And so the way that I interpret this is basically, if you have a task with an objective definition of done, then try that out with Sonnet 5 .5 first. If you need some more creativity and you're looking for a thought partner to help you decide what the definition of done is and it's a bit more subjective, then that's where I would use Opus as my model of choice for that task.
And then what I'd do is I'd run different skills with different models and I'd see, okay, for this skill, Sonnet does better. For this skill, Opus does better. So in the future, when I want to run this skill again, I know which model to use.
Anyways, you guys know the drill here. I gave all of these different sessions the exact same prompt, and now we will be reviewing the results. I did Sonnet 5 .5 and Opus 5 .5 both on high.
when it comes to the actual effort level they both used high so the first thing was building a engaging and immersive but high converting landing page for perk form and i said that they could use key .ai to generate anything they need and you can see i had to correct it key .ai not g .ai so i should actually probably add that as a misspelling inside of my glido to correct this too key .ai.
Okay, cool. All right, so let's take a look at these two outputs. I'm not going to tell you which one was which until the end.
I really like how when this one loads up, we have that text come in really cool. We can see that this can is interactive. It syncs to what my mouse is doing.
Right off the bat, I noticed that this can is kind of a 3D rendering. So it's not actually the exact can, like pixel for pixel the way it should be. So it had to kind of invent that.
I don't love that, but I do like the effect that it... gave for us. I also like how in the back you can see, oh wow, actually if I click and drag, it like scrolls the can.
But you can see in the background, we can see like perk form. We can see text in the back, which I thought was pretty cool. Anyways, let's go ahead and scroll down.
We can see the different flavors here. We can actually switch right from the hero section, which is pretty neat. As we scroll down, we get the can flooding away.
And then we have two drinks, coffee and protein coming into one, which gives us our one can, which is our perk form. We can see we scroll down into a little section where we have a gym bag, phone, keys. And as we scroll, we have the time going up.
And now we grab our perk form. We have a little flavor section. So we have the three different flavors.
I'm able to see that I can click between them here. So we have our salted caramel. We have our vanilla latte.
We have our bold mocha. It's sort of starting to tell a little bit of a story here. And then that is basically the end.
And we have these three dynamic cans coming back into view. Okay, so that was one output. Here is the other one.
So right off the bat, this one isn't as like... Boom, right in your face. It's not as much wow factor, I'd say.
We do have sort of an element where it's syncing to the mouse. You can see as I'm moving, you can see these are on different layers. Also, the key is on a different layer over here.
And I think that honestly, this looks a little bit cheaper when we look at just kind of like the hero section. But as we scroll, we can see that those other two drinks kind of come down here. It's doing a similar style where we're like scrolling through a time.
And I don't love this one as much. It's kind of making all of your eyes and all of your attention go to one spot. So it's definitely telling a different style of story here, even though all of the branding and everything else is consistent.
This is using the actual real generated image inside of the brand guidelines for the product. So that is better that it gave us the actual, I guess, image in a real way rather than back here. If you guys remember, it was like a generated version.
Here's a cool little animation here. We zoom into that image. All of these were different.
independently generated images i think that this looks really nice ultimately this is the better output in my mind the one thing i don't like as much is this section i think that this version tells a little bit of a better story but ultimately this is the better outputs even though i don't like the hero section as much this is the better output so this one was opus 5 .5 and this one was sonnet i think sonic killed the hero section if this would have been a real can and then opus 5 .5 missed that but everything else of opus's version i think was better so if we look at the stats here we can see that Sonnet ran for 28 minutes, whereas Opus ran for 52 minutes.
Sonnet cost 20 million input tokens, 134K output tokens. Opus was 37 million and 148K output. And Opus cost about double here.
So Opus ran for about twice as long and cost about twice as much, which is pretty consistent with what we saw when we actually look at the pricing. The question to me now would be, is Opus's output two times better? I would say no.
So ultimately, what I would probably want to do here is I would re -prompt Sonnet and say, hey, this is really good, but don't use this fake image. Use our actual real images. And I bet that it could swap everything out, and then I bet that it could give us the level of output we were looking for with a real can, and it wouldn't have cost us $13.
So ultimately, that's why I think that we need to give this run, we give the win here. to Sonnet. But either way, this stuff is subjective, so take it as you will.
But that is what we're going to do right now is I'm going to give this first output to Sonnet. And in my latest videos, you guys have been saying like, oh, please, can you use the subscription usage instead of the API billing? No, I think this is a better way to show the actual cost.
These costs are very much correlated to your subscription. So just think about it like that. But the reason I'm not going to do subscription is because I don't know if you're on the 20 bucks a month plan or the 200 bucks a month plan.
I also like to run all of these in parallel and I don't want to like sit and wait 52 minutes and then jot down what percent it was and then do this one for 28 minutes and then jot like, you know, I'm just going to run this in parallel. Anyways, let's move on to number two. Real quick, guys, I've got this completely free website design skill that I'm giving away.
I use this for every single website that I build, including our AI Automation Society site, which I think is pretty slick, and I absolutely love this website. It's called Scrollcraft, and it understands things like scroll -driven animations and layering, but also it has things in there like taste, typography, spacing, depth.
It's just a really good website design skill in general, and it's completely free, so the link for this will be down in the description, but let's get back to the video. All right, for number two, what I did is I said, hey, I need you to make me a motion showreel skill about Glido. And you do the research, make sure that all of this is consistent in branding, typography, you know, all of this kind of stuff.
You are the motion designer here, so just impress me. So let's first look at Sonnet's output.
Refactor this function to use async await instead of promises.
I'll fold those into the next revision. Reply saying it looks good and we can start next week. We try logic to the API client.
Okay, really not bad at all. I thought there was a lot of creativity here where it's kind of showing off what the product actually looks like when you use it, which I thought was really nice. The voiceover that it did, it generated that with maybe 11 labs or something.
I also really liked how these pixels here formed a G and that actually turned into an AI generated video that it made. So I thought that was pretty nice. Overall, I was honestly not expecting a Sonnet model to do this good with this motion show reel.
Now let's go take a look at Opus's output.
Refactor this function to use async await instead of promises.
Okay.
I mean, I don't think you can beat that. Opus 5 .5 is just so good as of now at motion design. It just felt more glido, the background, the different animations it made.
It just felt like even the copy was better. It also went to our website and pulled this quote from me that was on the site, which was very fast, but right here, this was actually the quote from the website. So ultimately this output, much better from Opus.
Now, when we look at the cost here, Opus was faster and it was also... less output tokens, but ultimately it was more expensive. So Sonnet was $14 .37, whereas Opus was $20 .55.
I think though, we're gonna give this one to Opus. If I'm still looking through the lens of if we had Sonnet try again to get up to $20 .5, I still think Opus's would be a better output. It just had a better taste with like motion and the sound design.
And sound design is something that I think is really hard to get right, even as a human sometimes. And the fact that Opus is doing this so consistently. Very good.
So I'm giving this one to Opus. Okay. So for number three, I said very, very vague.
create me a research -backed 90 -day action plan for getting How They AI to 1 ,000 subscribers. So let's go first take a look at Sonnet's output. Okay, so one thing to call out here is both of these models originally, when I shot off that prompt, they just gave me a markdown file.
And then for both of them, I gave them the exact same second prompt, which was turn that into some sort of document that I can view that's easier for me to read. And this is Sonnet's version. It created me an HTML file and it gave it to me locally.
And then Opus did the same thing, except for it put it as a cloud artifact. So that's one little tiny difference. But besides that, this is Sonnet's output.
Let's see. 1 ,000 subscribers in 90 days. It also threw our logo up here, AIS Media, September 28th.
We have day one launch as well as the day 90 goal. It's interesting that it chose to launch this on November 2nd rather than like next month or something, Q4. Anyways, we have a timeline here.
So we have episodes on the bottom. We have subscriber gates and getting to a thousand by day 90. We can see that we have some other things here like promotions, test send, full list, full list.
We have Nate promotion sends out six of them, community posts, five of them, 500 subs, holidays locked. Interesting. We have different phases.
So what to do from October to November and then two weeks on launch and then a month. in the next stage and then compounding. So it's showing us the different sort of milestones that hit in these different phases.
It's showing us the different sources that we can get subscribers from. None of this is interactive. It doesn't look like.
It's just showing us what this could look like. We can also see that we got some scroll animations coming in there, which I thought was a nice touch. As you can see, the text comes in.
It's also looking at shorts and LinkedIn and guest kit and scorecards and other decisions like that. Okay. So this is good.
I wouldn't say that this is like super, super helpful for me at the moment, but not too bad. And here we have the version with also scroll animations from Opus as a cloud artifact. This one shows to start October 1st to December 29th.
So it's actually like a Q4 90 day sprint, which I think is a bit more aligned based on like what's going on in our business right now. So now we have the 90 days. This one gave us a little bit of a different sort of view, which I actually like.
better, I'd say. We can see when we're going to launch. And from now until launch, we'd be kind of setting up auditions and we'd be doing guest applications and we'd be, you know, announcing all of this kind of stuff, recording these episodes.
We'd be looking at getting out shorts and hitting the email list and community posts and all this kind of stuff. I think that this view is a little bit easier to digest than what's going on here. There's just a bit too much going on here and it's not as clear.
Whereas this seems a little bit more clear to me. We have different phases here, but this one's basically each month, which I think is better. Build.
launch, compound. Same thing, where do these actual subscribers come from, the list, YouTube Organic, Nate's channel, things like that. Here are things that we actually need to decide on.
And then we can open the full plan, which actually takes us to, oh wow, this takes us to like an actual very specific write -up with like tasks. So this is basically something that I can throw into my OTA now. And we have these actual tasks that we can check off.
Okay, I think this one's pretty clear. Opus wins this. test.
Now, when we take a look at the stats, Opus ran for 20 minutes, Sonnet ran for 14. Sonnet costs a little under $5, whereas Opus costs a little over $9. So if we had Sonnet take another stab at it, where would Sonnet be?
Now, this one gets really interesting to me because if you think about the actual prompt, I didn't really give it guidelines at all. I said, create me a research -backed 90 -day action plan. That was it.
And if we think about what I said earlier, If you have an objective definition of done, use Sonnet. If you need more creativity and you're looking for a thought partner, use Opus.
And this is definitely an example of me looking for a thought partner and looking for a model to just take an ambitious goal or sort of a vague goal and just run with it. I think that if I would have given both of these models a very, very specific prompt on here's what I want, here's how it should look, here's blah, blah, blah, I think that Opus and Sonnet would have given me a similar result, but Sonnet's would have been cheaper.
So I do really think that that's interesting to think about the way that you talk to these different large language models. But in this case, I can't ignore how much better Opus did with this vague prompt, even with the cost. So this one is going to go to Opus as well.
So right now we're at Opus is winning two to one. And really guys, I'm not trying to come in here and say like which model is better than which, because ultimately Opus is a better model. But I'm just trying to hopefully show you guys like now you might have a better idea of when to use each model and or how to use each model.
So let's hop into number four. All right, so in this one, I gave them a bunch of... company data, mock company data.
And I wanted to get an Excel sheet as well as a investor pitch deck. So we're going to do Opus first. Here is the deck that Opus built.
It generated this image in the back. We've got 18 slides. We've got generated images here as well.
We have the problem. We have the product. We have who we serve.
We have some traction, so some statistics in here based on the mock data. We've got the business model. customers.
Now this doesn't look too bad. We've got visuals. We've got, you know, not too much going on on each slide, but I will say this is the.
infamous clawed design font with the Fs like that and the numbers. It doesn't look super premium in my mind. Like this doesn't feel as premium as I'd like it to, but either way, the slides aren't designed super, super poorly or anything.
So that is the output. I mean, this really isn't bad at all from Opus 5 .5. And then on the Excel sheet, we have a dashboard, inputs, monthly metrics.
We have different views. The dashboard is really nice with all these stats and all of these different visualizations in here. We have the different inputs.
We have metrics, quarterly P &L. What I really like is in a lot of these cells, we see formulas. We see formulas basically everywhere, which means if we had to come in here and change up the data, the whole sheet would update.
I don't like when AI models will give you a Excel sheet with no formulas. And now you have to basically, it's all static. It's all hard -coded data.
And if you needed to make changes, you would have to go ask the agent to make those changes for you. I hate when it does that. So it's really good that Opus is putting I mean, there's tons and tons of slides here, like a ton of slides or sheets, I guess, tabs.
What I like is that they have formulas inside. So it's really not bad at all from Opus. We've even got some color coding and stuff.
Very good Excel sheet. This is interesting because if you guys remember earlier, Opus made an artifact and Sonnet didn't. But in this time, Sonnet made an artifact.
So let's see how this one's looking. This is BrightPath Analytics. We see our intro right there.
We have, I believe this is the exact same font that Opus was using. Anyways, we have company snapshot, the problem, the product, who we serve, business model, traction. I mean, this is very, very similar.
Like there's not really any difference in my mind between the two different decks. They feel and look very, very similar. And Sonnet actually gave us more slides.
Now let's open up the Excel sheet we got from Sonnet. Immediately, it doesn't look as premium, but we go into the dashboard. The dashboard looks just fine.
We have a lot of visualizations here, monthly data. We can see that we have lots of formulas in these calculation cells. Same thing with the P &L.
We have some SaaS metrics. We have assumptions. We have organic forecast.
All of these are different lookups and formulas as well. Wow, lots of different tabs here. Pipeline, customers, marketing, headcount.
So this one honestly might have given us like too much. I don't know. If this was real data and this was my real company, I'd be able to tell you guys better like how useful all this is.
But ultimately... I wouldn't really be able to come in here and say that one of these outputs doing this sort of like document creation was significantly better than one of the other AI models. So I really think in this case, whichever model was cheaper and faster is who I'm going to say won this round.
And wow, in this case, Opus was actually cheaper and faster. So we don't even have to discuss this one. Opus wins this one, no question.
Cheaper, faster, and the outputs were pretty much identical in my mind. Now, what if we had some skills built around it? That would be another interesting thing to test.
But in this case, I didn't utilize a skill. I just said, hey, make these things for me. All right, so here's an example where it actually did use a skill.
I have a skill called HTML Explainer, and I asked it to use that skill, or I didn't ask it to, but it did, where I said, hey, what models could I run on my current device? I don't know anything about this, so make it easy to understand. Lots of pictures, blah, blah, blah.
So let's take a look at these outputs. All right, so this one is Opus. What AI can my PC run?
You can see this is branded. It's just a super simple, you know. html explainer style skill that i use sometimes so we can see we have graphics card memory processor storage we see the graphics card memory versus other things which is like a gaming pc at a data center chip The 32 gigabyte rule, the model has to fit in your graphics card.
So we probably couldn't run this 43 gigabyte model, but we could run a 20 gigabyte model. So that is a nice little simple rule. We can see that we also have shunk copies.
We have speed. So tokens per second on our graphics card. What to install.
Everyday chat. It gives us a suggestion. Coding.
Hard problems. Gemma 4 for fast replies. Super, super cool.
Too big for one PC. Start tonight. It tells us exactly what we would need to do.
Okay, cool. Very simple. Not too bad at all.
Okay, and now we have the exact same skill, exact same prompts, but this is the Sonnet version. So local AI models your PC can run. It actually made us a little visual, which I thought was cool.
32 gigabytes, the number that decides what local means. Okay, so it starts off with what local means, which I don't think the other one did. The other one basically just went straight into our PC.
So that's good that it did that. It broke it down simply with another little visualization. Three places to keep model, desk, shelf, warehouse.
Okay, fast, slow, storage only. What fits? Green fits, amber crawls, red won't run.
So it doesn't exactly tell me why because of the whole gigabytes thing, but it does show a nice little visualization and some other AI models. So that's not too bad. It also shows us speed, pick by the job, similar to the last one, and how to actually try one out right now.
So ultimately, if we really think about how would we want to understand how this stuff works, I do think that Opus's output was better. Like this one is more visual, I'd say, but it doesn't actually explain it to me as well as the Opus one did, unfortunately. So I would say that Opus wins this output.
But if we think about this cost here, similar timing. Sonnet cost half of what Opus cost. So if I prompted Sonnet one more time and said, hey, I am so confused about this, this, and this.
Can you just add some more info here? I bet that Sonnet's output would be better. And I would also argue that I think that Sonnet's output visually is better.
And that's one of the pieces of this skill, which is a very, very... Simple skill. It's not refined at all.
But ultimately, I think that if we prompted Sonnet again, we would get an output better than what Opus gave us for $3. So I'm going to give this one to Sonnet. Okay, so in this next one, number six, I said, scrape X for all the news that's been going on today.
Give me an interactive dashboard so I can view what's going on and see what's coming out and just get prepped for all these different things going on in the AI space. So let's take a look at Sonnet's output. This is a local host that it created for me.
We can see that we can see when this was last scraped. We can hit refresh. We can see how many posts were tracked.
So 386 posts tracked from 676 scraped. We can see some headline news right up top. We can see what is trending.
We can mark fits my channel or hide covered or pinned only. Okay, interesting. We can also see the heat score.
So Claude's on a 5 .5, 100 heat. Metastart's meta enterprise platform, 94 heat. We can maybe click into this stuff.
Nope, I can just hover over. I can open details. Okay, cool.
So I can actually open the actual posts that it's sourcing, which is pretty nice. We can mark it covered. We can add to plan.
So now we get this list with. like a content plan over there. Not too bad.
We can also sort by newest. We can sort by channel fit. We can look at the last 48 hours.
We can see what's coming this week. Dev day by OpenAI. We can see the calendar.
Okay, this is not a bad view at all. And now here we have Opus's version, which off the bat just kind of looks a little bit better to me. You can see that this one's still scraping X.
As you can see, it's showing us a little progress bar. So it is working, but... This view looks a little bit more visual.
It's showing us posts per hour on AI. It's showing us how many posts are rising and cooling off. And this view, in my opinion, is just a little bit better.
I can also switch up here this week. Okay, nice little view here. We can see content plans.
So I can add things to my list. We can see the raw feed as well. So that's pretty cool.
I can see sort of a heat score as well. It's tagging things by rumor or shift or discussion. And it's pulling up real posts on the side here.
And let's see, how do I add something to right there? I can click add to content plan. I can open the top post and it goes to X.
Okay, very cool. So I'm really starting to understand a pattern here, which is what I said at the beginning. If you know exactly what you want, Sonnet is probably gonna be able to do a good job for you.
But if you need the creativity and you send an open -ended, very vague goal, Opus is just going to handle it better. And that's become very, very clear to me. In this case, Opus ran for about seven minutes longer and cost $6 more.
And if we iterated four more times with Sonnet, or as much as it took to get to $8 of output, but we had a clearer goal, then Sonnet would definitely give us that level of output. There's nothing in here that Sonnet would have not been able to code or create for us. It's just that Sonnet doesn't think about that as much.
And it won't go kind of like outside the box as much as Opus will to give you a better experience. It's basically just going to give you what it thinks you want. And it's going to be the baseline.
So if we're sticking with my rule here, because Sonnet was much cheaper, I'm going to give this one to Sonnet. But ultimately, like right now on one prompt, Opus's output was better there. And there's just no denying that.
So let's go on to the last one, number seven. All right. So in this one, I gave Opus and Sonnet a link to one of my YouTube videos.
And I said, create me a YouTube resource guide. So this is using a skill once again. So the skill is defining what.
good looks like essentially. So let's take a look at Sonnet's output first. We have the AIS header up top.
We have this font. We have Nate Hurk linking to my YouTube channel. It starts off with the challenge and the rules.
It goes over the SMP benchmark. It goes over the actual steps of how I built this. We have the problem, the solution, day one and day two results, day three, day four, day five, day six, day seven, final results, lessons learned.
And we have key takeaways down here. They're pretty much the exact same length. This one is like, seven pages, but it's really like, I don't know, there's a big gap here.
Whereas this one, very similar, same link, same title, same header, the overview, the rules, the benchmark, how the system was built, day one, day two, final results, all this kind of stuff. It's very, very similar, which is why this one would be essentially, whichever one was cheaper is the one that I would go with here.
In this case, you can see that Sonnet was cheaper and faster, about half the speed. twice as fast, and $1 .40 cheaper. So this one is gonna go to Sonnet because the outputs were not too different enough to actually be able to say, oh, one's really better than the other.
And it's because in this case, I know we've seen some other examples where it didn't always match, but in this case, I had a skill that I do use a lot. And a model like Sonnet that's less intelligent than Opus was able to just follow my instructions because it knew what I wanted more clearly. So this one goes to Sonnet.
So ultimately in this test, we had... three, four go to Sonnet, and three go to Opus. So we did technically have Sonnet winning this challenge, but I think you guys understand this wasn't really a head -to -head.
It was more like how and when to use each model. And here are the final stats for you guys. You can see across these seven sessions, Opus ran for about 46 minutes longer.
It spent about 38 million input tokens more, as well as actually less output tokens, which is interesting. But it did cost about $16 more than Sonnet 5 .5 and all these runs, which when you think about the fact that it's actually double the cost of Sonnet, It didn't cost double the cost in this experiment, which is also interesting.
I've talked a lot about how I see each of these models fitting into my workflow now. And hopefully you guys have a better idea of how this all should work for you. But ultimately you got to get in there and you got to test it on your own skills and your own prompts and your own processes.
So that's going to do it for this one. That was my experiment today. I hope you guys enjoyed.
I hope you learned something new. And if you did, please give it a like. It helps me out a ton.
And as always, I appreciate you guys making it to end the video. I'll see you on the next one. Thanks everyone.
The Hook

The bait, then the rug-pull.

He opens with the pitch for the new, cheaper model, then flips the real question: not whether Sonnet 5.5 is good, but when it's good enough to skip paying double for Opus.

Frameworks

Named ideas worth stealing.

01:27list

Anthropic's model-choice guide

  1. Everyday coding, fixing bugs, high-volume development -> Sonnet
  2. Polished documents, slides, spreadsheets, one-pagers -> Sonnet
  3. Well-defined agent tasks run repeatedly -> Sonnet
  4. Complex work requiring careful judgment, including long-horizon agentic coding and knowledge work -> Opus
  5. Your hardest problems, where you need the most intelligence -> Opus

Anthropic's own published rule of thumb for when to reach for Sonnet versus Opus, read directly from their guide on screen.

Steal forAny internal team doc setting a default model policy per task type
02:07concept

Definition-of-done test

The host's personal reinterpretation of Anthropic's guidance: if a task already has an objective, checkable definition of done, try Sonnet first. If the goal is vague or subjective and you need a thought partner to help decide what 'done' even looks like, use Opus. Run skills on both once to learn which model a given repeated skill prefers, then default to that model going forward.

Steal forDeciding per-prompt, not just per-project, which model to call
CTA Breakdown

How they asked for the click.

VERBAL ASK
06:46product
“I've got this completely free website design skill that I'm giving away... It's called Scrollcraft... the link for this will be down in the description.”

Soft mid-roll plug placed right after the first comparison verdict, framed as sharing his own tool rather than a hard sell, free opt-in rather than a paid pitch.

Storyboard

Visual structure at a glance.

cold open
hookcold open00:00
pricing table
promisepricing table01:07
round 1 stats
valueround 1 stats05:38
final scorecard: Sonnet 4, Opus 3
valuefinal scorecard: Sonnet 4, Opus 324:59
sign-off
ctasign-off25:25
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.