Modern Creator
Stacked Podcast · YouTube

Claude Opus 5 Just Ended the Benchmark Era

Two operators who spend real money on inference argue that price-per-task, not leaderboard position, is now the only number that matters.

Posted
1 months ago
Duration
Format
Interview
educational
Views
5.5K
142 likes
Part of the collectionThe Claude Opus 5 PlaybookEvery Opus 5 breakdown, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

When a cheaper model reaches the same practical quality as the expensive one, benchmark rankings stop describing anything useful and cost per completed task becomes the real measure of a model.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You pay for model inference out of your own margin and want to know whether to switch defaults.
  • You run agents in parallel where a 2x price difference compounds into a real monthly number.
  • You keep seeing viral claims about a new frontier model and want a worked example of one falling apart.
  • You are deciding between model families and want a practical take on where each one still wins.
  • You are earning well from a low-overhead business and have not settled on what to do with the surplus.
SKIP IF…
  • You want rigorous benchmark methodology, because this is two practitioners trading field notes.
  • You need current pricing, since the numbers quoted here are tied to the launch window.
  • You are looking for implementation tutorials rather than model-selection strategy.
TL;DR

The full version, fast.

Opus 5 lands at roughly Fable 5 quality for about half the price, which the hosts argue removes the reason to run the more expensive model at all outside of narrow cases like penetration testing. Fast mode compounds the advantage, delivering that quality at around 2.5 times the speed. Because cost is mostly reasoning tokens, a cheaper model that finishes in fewer tokens wins twice, which is why they argue for judging models on cost per completed task rather than benchmark position. The Kimi K3 story is their case in point: a zero-day claim that drew 1.3 million views, against a government evaluation where it completed zero of 41 exploit tasks.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Voices

Who's talking.

00:13hostJack
00:13cohostNick
Chapters

Where the time goes.

00:0001:12

01 · Opus 5 lands and changes the math

Sets up the episode around a launch inside the last 24 hours, with both hosts having already tested the model.

01:1202:09

02 · Fast mode: frontier quality at 2.5x the speed

Describes getting Fable-level quality at roughly 2.5 times the speed for about 80 percent of the cost with fast mode enabled.

02:0905:11

03 · The pricing gap and two philosophies

Lays out $5 in and $25 out against double that for the competing model, then frames the choice as intelligence per dollar versus maximum intelligence regardless of cost.

05:1106:25

04 · The Pareto frontier, explained

Walks through a chart of cost against performance and explains that sitting on the frontier means nothing is both smarter and cheaper.

06:2508:47

05 · What the expensive model is still for

Narrows the remaining case to penetration testing and cybersecurity, then covers the refusal cascade that pushes users down to weaker models.

08:4709:58

06 · The model they have not shown

Speculates about internal models more capable than anything released, and what a lab can do with that much unused inference capacity.

09:5813:40

07 · A model asks to be consulted on its successor

Covers a release-materials detail about the model requesting involvement in future versions, and moves into recursive self-improvement.

13:4014:31

08 · Model naming is a mess

Argues naming is driven by internal training lineage rather than marketing, which is why the families confuse the people buying them.

14:3119:20

09 · Zero of 41, and a viral claim falling apart

Traces a zero-day exploit claim that reached 1.3 million views against a government evaluation where the model completed zero of 41 tasks, then explains why distillation may not carry cyber capability.

19:2021:04

10 · Picking a model by taste

Compares writing style across families, argues tone of voice is a legitimate selection criterion, and notes each family has recognizable tells.

21:0424:52

11 · Where models still misbehave, and the golden zone

Admits to a five-hour build that went sideways, then frames the present as a temporary window where industrious builders hold an asymmetric advantage.

24:5230:50

12 · End benchmarks, and what to do with the money

Answers a listener question on cost per token versus cost per task, calls for abandoning benchmarks, then covers investing business surplus and the high-earning-not-rich trap.

Atomic Insights

Lines worth screenshotting.

  • Opus 5 delivers roughly Fable 5 quality at about half the price, which removes most reasons to default to the more expensive model.
  • Cost is mostly reasoning tokens, so a cheaper model that finishes in fewer tokens saves money twice over.
  • Fast mode gets you frontier-level quality at roughly 2.5 times the speed for about 80 percent of the cost.
  • Sitting on the Pareto frontier means no model is both smarter and cheaper, which is a more useful claim than any leaderboard rank.
  • The remaining argument for the pricier model is narrow, mostly penetration testing and cybersecurity work.
  • A model that keeps refusing pushes you down a cascade of progressively weaker models until the work degrades.
  • A Chinese open model that went viral for discovering a zero-day scored zero of 41 tasks on a government exploit benchmark.
  • The viral exploit claim reached 1.3 million views before any robust testing contradicted it.
  • Distillation transfers design quality and general capability but does not appear to carry cybersecurity capability across.
  • Model naming has drifted from marketing into internal training lineage, which is why the families confuse buyers.
  • Different model families have genuinely different strengths, so the useful mental model is a team of specialists rather than one generalist.
  • Tone of voice is a real selection criterion, and the tells of a given model family become obvious with daily use.
  • Models still fail on long builds, so five-hour sessions where you repeat yourself are still normal at the frontier.
  • The current window rewards anyone industrious enough to build with these tools, and that asymmetry will close as the tools get easier.
  • High earning is not the same as wealthy, and the inversion point is when your assets earn more than you do.
  • Set the monthly investment number first, then treat the business job as hitting that number.
Takeaway

Judge models on finished work, not leaderboards

WHAT TO LEARN

Once two models produce comparable output, the only number that still separates them is what it costs to get a job all the way done.

02Fast mode: frontier quality at 2.5x the speed
  • A model at comparable quality for half the price removes most of the argument for defaulting to the expensive one.
  • Speed settings that keep the same underlying model are close to free performance, so check whether one exists before switching families.
04The Pareto frontier, explained
  • Cost is mostly reasoning tokens, which means a model that finishes in fewer tokens beats a cheaper per-token price.
  • Sitting on the Pareto frontier is a more meaningful claim than a benchmark rank, because it says nothing is both smarter and cheaper.
05What the expensive model is still for
  • Keep the expensive model for the narrow cases where it measurably wins, which here means security and penetration testing work.
  • Repeated refusals push you down to weaker models, so quality loss can come from the safety path rather than the model choice.
09Zero of 41, and a viral claim falling apart
  • A capability claim that goes viral is not evidence, and this one drew 1.3 million views against a score of zero out of 41.
  • Distillation appears to carry general capability without carrying cybersecurity capability, so a cheap clone is not automatically a security risk.
10Picking a model by taste
  • Model families have recognizable writing tells, and tone is a legitimate reason to pick one for customer-facing output.
11Where models still misbehave, and the golden zone
  • Frontier models still fail on long multi-hour builds, so budget for supervision rather than expecting one-shot results.
  • The advantage available to people who build with these tools now will close as the tools get easier, which makes timing part of the strategy.
12End benchmarks, and what to do with the money
  • High income is not wealth, and the threshold worth aiming at is the point where your assets earn more than you do.
  • Setting the monthly investment number first, then earning to hit it, keeps lifestyle from absorbing the whole increase.
Glossary

Terms worth knowing.

Pareto frontier
The set of options where nothing else is better on every dimension at once. For models it means no alternative is both smarter and cheaper, so any move off the line trades one for the other.
Fast mode
A setting that returns output faster from the same model family rather than swapping in a smaller model. Speed changes, the underlying quality does not.
Distillation
Training a smaller or cheaper model on the outputs of a larger one so it inherits much of the larger model behavior at lower cost.
Recursive self-improvement
The point at which a model is capable enough to design and train its successor, so progress no longer depends on the pace of human research.
Safety cascade
The pattern where a refusal on one model pushes a user down to progressively less restricted and less capable models until the work quality drops.
Exploit bench
A benchmark that scores whether a model can produce working security exploits that reach code execution, used by governments to assess offensive cyber capability.
Cost per task
Total spend to finish one job rather than the advertised price per token. A model with a higher token price can still be cheaper if it needs far fewer tokens to finish.
HENRY
High earning, not rich yet. Someone whose income is large but whose spending matches it, so their lifestyle collapses if the income stops.
Resources

Things they pointed at.

04:51resourceElon Musk post on the Pareto frontier chart
15:19resourceViral X post claiming a zero-day discovery
17:12resourceExploit bench (government cyber evaluation)
Quotables

Lines you could clip.

02:40
Who is going to be using the expensive model now when the cheaper one is pretty much the same performance at 50 percent of the price?
the entire pricing argument in one lineTikTok hook↗ Tweet quote
05:37
The Pareto frontier basically means that there is no model that is both smarter and cheaper than it.
defines the concept cleanly enough to stand alonecarousel slide↗ Tweet quote
06:00
What is cost, really? Cost is the number of reasoning tokens, for the most part.
reframes pricing in a way most buyers have not considerednewsletter pull-quote↗ Tweet quote
14:40
Two major world governments tested it and found that it scored zero out of 41 on their big cyber test bench.
hard number that punctures a widely shared claimTikTok hook↗ Tweet quote
24:52
Let us just end benchmarks for all time. Who cares about what the numbers say? How is it actually used?
the title thesis said out loud by both hostsReels caption↗ Tweet quote
22:10
Those with the will and the ability to be industrious can build anything they imagine, and this asymmetry will end.
urgency framing without a product attachednewsletter pull-quote↗ Tweet quote
28:30
If you ask very rich people when they feel rich, it is the inversion moment where your wealth makes more than you do.
clean definition of a threshold most people have never namedcarousel slide↗ Tweet quote
27:50
I set a number that I want to invest every single month, and my job is to make enough to hit that investment number.
concrete, copyable financial habitTikTok hook↗ Tweet quote
Topic Map

Where the conversation goes.

00:0006:25denseOpus 5 pricing and performance versus Fable 5
05:1108:47denseModel selection philosophy and the Pareto frontier
08:4713:40steadyUnreleased models and recursive self-improvement
14:3119:20denseKimi K3 cyber claims and distillation limits
19:2023:30steadyModel taste, tone, and real-world failure modes
23:3030:50steadyListener Q&A on cost per task and profit allocation
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphor
So Cloud Opus 5 just dropped. People are calling it the world's most intelligent model. We have breaking news on Opus 5, stuff that's not been covered yet anywhere else I've seen.
If you're new, I'm Jack, and this is Nick, and we make collectively around three quarters of a million dollars every single month. All we do is talk about AI, and we use models all the time. Nick, Cloud Opus 5, where are we at?
Let's start with some of the breaking news. This launched in the last 24 hours. You've tested it, I've tested it.
What's going on right now? Well, Anthropic just shipped. Opus 5.
I made a video on it. You made a video on it. The whole internet has made videos on it, but I don't think people really fully understand the main competitive advantage of Opus 5 versus current frontier LLMs, things like Fable 5, GPT 5 .6 Sol, and even Kimi's new K3.
Opus 5 is not just close to Fable, like 90 % of the intelligence of Fable. Opus 5 essentially is Fable, it's just now offered not only significantly faster, but also way cheaper. So I don't know if you had a chance to play around with fast mode, for instance, but I love using fast mode on the Opus model family.
For the last, I don't know, last month and a half with Fable and stuff like that, I've been unable to do so. But the second that I got back to Opus 5, you know, the second they dropped and then I went slash fast and I just jumped right into it, I was getting Fable quality intelligence at 3x, maybe 2 .5x the speed of Fable for maybe 80 % of the cost of Fable.
with fast mode enabled major unlock i don't think a lot of people are realizing just to what degree you can squeeze out way more tokens per second from these things and then obviously the fact that you know even for people not on fast mode it's approximately the same level of quality is it's just insane it's wild and it's like the quality is one thing so the general consensus is when you look how everybody's used it and you amalgamate it together is that the performance is on par or similar to Fable 5 based on what you're asking it to do.
But the big difference with the Opus 5 model is it costs less to get there. In other words, I think it's exactly the same price as Opus 4 .8. Some people have described it as more robust, and this kind of knits into the heart of the debate between models.
And what is more important? Is it this, and we're going to get into some controversial stuff here as well, right? Is it about what is the...
is multiple on your dollar spend in terms of intelligence, or is it actually what is the most intelligent irrespective of the cost? Two different philosophies, two different schools of thought that you can take when you think about AI adoption. Do I want to be able to get what we call the Pareto curve?
And we're getting on to some of that in a second, but I think that kind of cuts to the heart of the debate. Yeah, I mean, like, realistically speaking, do I even need the most intelligent model anymore? Like, does my work require frontier math skills?
Should I be solving Erdos problems? Like, probably not, right? And so, if we're at the point now where Opus 5 is approximately Fable, and, you know, Fable is basically everything I'd need for intelligence anyway, do I even need, like, a Fable 6 anymore?
I mean, look at these stats, right? It's $5 in, $25 out. Fable's double that, $10 and $50.
Right off the gate, I mean, I'm saving double for reasonably similar token usage. Exactly. Who's going to be using Fable 5 now when Opus 5 is pretty much the same performance, but 50 % of the price?
You've kind of like nerfed Fable. Dude, I just went back to Fable because I wanted to run the same query on both Fable and Opus 5. And I was just shocked at how much longer it took to run Fable.
It was like crazy. I think I was sitting there for like three minutes or something on a query that Opus 5 finished in like a minute. And so I was trying to compare them.
pound for pound uh line for line and opus was just like zoop done and then fable was like still working i was like oh man i don't feels like i'm already in my freaking stone ages do you think opus was like or i don't know like anthropic inherently nerfs the models i'm kind of curious it's a very very interesting one and we have to put our hands up and say like i can see why so many people have model fatigue not the supermodels that walk down the aisles nick the kind of ones that power our lives in the chatbots and it's like every two or three days now we're getting new frontier level models we had 5 .6 salt that feels like that was stone age ago but very recently that actually came out but the philosophy for anthropic right of like We have a Fable lineage, a Fable style model, and then we have the Opus style model.
One question is, why are they different? Like, are they fundamentally different architectures? Why are they splitting them out like that?
It's a very interesting question because I think for the user, it's very confusing when to use what. Because everyone's like, okay, great, new models dropped. How does this compare to other products?
From the consumer's point of view, which is how we think about these things, there's not enough... of a differentiation between the two products. They need people for us to test it and kind of distill what's going on here.
Yeah, agreed, agreed. Check out this tweet that Elon Musk made. Grok 4 .5 and Opus 5 are alone on the Pareto frontier.
This is frontier code with average cost USD per rollout and then the score or performance on a task. And you could see that Grok 4 .5 is like the cheapest, almost cheapest model. Probably the cheapest pound -for -pound for performance.
There's this poor little dinky one here, SW -1 .6. It's like 10 % freaking one cent. I wasn't going to use that.
Anyway, Grog 4 .5 is over here. And then look at Opus 5. What's interesting is that it looks like its performance goes...
Am I right here? It goes down? This Pareto frontier basically means that there is no model that is both smarter and cheaper than it.
Right. Right. That's essentially what that line means.
So that's the ultimate. Yeah, yeah. Exactly.
Because, I mean, you do get good performance on Fable stuff as well. It's just, look at this. It's like two and a half times the cost.
And if you think about it, what is cost, really? Cost is the number of reasoning tokens, for the most part. And so Opus 5, even if you were to divide this by half, okay, not only is it still another half cheap, it's like a quarter of the total token usage.
That's crazy. To me, that's intelligence. That's wild.
I mean, the Opus 5 model is just an immediate replacement. I mean, what's the ball case for Fable 5 now? Probably cybersecurity pen testing stuff, because Fable 5's performance relative to Opus 5 on penetration testing and cybersecurity stuff is higher.
But that's the whole idea behind it. opus like they don't want people to do cyber security stuff right and i mean you can't use it anyway yeah i i've been using opus 5 for god i mean the last 24 hours just burning a hole in my wall i spent more than 400 testing opus 5 outputs with like the first i don't know hour you know i parallelized i ran it across like 50 different agents and i just had to create a bunch of like cool 3d experiences but i also asked a bunch of questions about a variety of topics and i had not once so far gotten like bio blocked where it's like fable five safeguards flag this message has not happened to me with opus five so that's it's an interesting one as well isn't it because i think people are complaining what's called the downward cascade where essentially you're on fable five fable five things it's too controversial you get knocked down to opus five and then if opus five thinks you're too controversial they bump you back down to and this keeps going until you're on haiku three basically
scrambling around under class a hundred percent we're all kind of like in caves basically with haiku three or something imagine the future where these like big large language model companies had just had big like they're like governments and you say something that's like against them and then they just freaking they just nerf you so hard you're back like you do imagine that like what a punishment It would be wild.
It would be absolutely wild. I think we, zooming out as well, Nick, like if we think about why has this been dropped so quickly, Kimi K3 is going to be a huge accelerant for this. The performance to price ratio has come out and there's some stuff on the cyber test.
But I think what's interesting to me is the speed of this stuff. And they just made it more accessible because Fable 5 effectively was a product that no one could really use outside of their subscription. at all and they're almost like different models maybe they're built slightly differently with fairly different comparable performances i mean if we were running an ai company like anthropic we would we would want to split test methodologies for what series of tests and systems and training would release the best model we wouldn't just have one way of doing it right we were probably running 10.
there's probably 10 different ways of doing this running an anthropic right now and they just picked the two best winners they're approaching slightly differently the open and the Fable. And Fable is allegedly a neutered, a handicapped, a bullied and beaten version of Mythos.
What we don't have, Nick, is the stats and the performance of Mythos compared to these guys. And it's interesting they've not really studied them. Maybe because they're trying to keep under wraps or I'm not sure.
Yeah, probably. And it came out like, dude, Mythos came out like two and a half, three months ago now. I mean, they had models that are probably more intelligent than both Opus 5 and Fable now for like three months.
Dude, what are they doing behind the scenes with these things? Is that ever just like, do you ever wake up in the middle of the night sweating and panting being like, what is Anthropic doing with their inference? Like, dude, they could be doing anything.
They have all of the GPUs too. You might be getting, you know, 50 tokens a second with Opus 5 and think, ooh wee, look how fast I'm reasoning now. Dude, they have 5 million tokens a second.
They could run Opus 5 across all of their compute clusters a bajillion million times faster than any of us ever possibly could. They could solve... science for they've probably thrown this thing at every possible like open math problem ever and the implications are kind of staggering 100 there was there's also very spicy detail that came out as well about so anthropics release materials on opus 5 apparently when consulted asked to be involved in decisions about future versions of itself and we've got some more controversial news we're going to have to touch on a few minutes but and essentially it's a model welfare question here.
So we've created a model that said, Nick, Jack, consult me on future development. I want to be involved in the creation of Opus 6. Ethically, mindful potentially, what do you think about this?
Well, I think it's like a mother being involved in the creation of a child. I don't think there's anything inherently weird about that philosophically. And just like children usually exceed their parents, I think Opus 6 would exceed Opus 5.
Just one of those things. It must be crazy being all of the other large language model companies and just seeing Anthropic, RSI, their weight at the top. For those of you guys that don't know RSI, recursive self -improvement, philosophically, it's the point at which models become intelligent enough to build their successors.
And when you're at that point, the intelligence of the model is basically unstoppable. because you don't need human beings to do it. You can parallelize it, shard it across however much infrastructure you have, run a bajillion geniuses in a data center like Dario Amore likes to say.
That must be terrifying.
100%. And it's also like, I think for the guys working for Teams, the fact that you are literally on the cutting edge is super interesting. And for those of you down there, the models already are involved in the creation of further models.
There's no way you're not using Opus 5 to build. opus six but you're going to consult it in capacities which you are comfortable with not ones that opus five dictates although at a certain point if you do concede the fact that the model will supersede the collective intelligence of every human on the planet it will be aware of things that you are not and will make suggestions that you do not understand so there will be judgment calls and decisions that say we don't know why but we're actually going to back it but the beautiful thing about this we we have the rule of science we can containerize.
One is Opus 5 -led development. One is human -led development. So we can split, A -B test these things massively.
So very, very, very exciting time to be alive. Opus 5 is cool. I think as well, Nick, people have got model fatigue as well right now with new models dropping every week from different labs.
And it is almost, I think they are influencing one another like we had recently. Do you remember the... the clod rush on youtube about a month and a half ago where it was daily update daily videos on clod updates and it was like impossible and effectively what happened the market started to influence itself as happens all the time and when models started to get released of mars like shit we're too far behind let's go ahead and nick if you could pull that graph of that pareto graph from from elon musk i i think it's really interesting that grok sits on that line because when you think of grok today you don't think of best model in the on the planet but if you're looking at punching power like pound these are the pound for pound rankings nick who's the best pound for pound model on the planet sits on that line so i think grok's journey is quite interesting it says something about the efficiency of the model and where it's heading and so i wouldn't be sleeping on grok actually i think it's going to be interesting musk actually said in two weeks we're going to have the next grok update
And in four weeks, we're going to have the second Grok update. So let's see where that lands on that line. Jeez, man.
So what, like Grok 6? Or do you think they're going to roll it out incrementally, like 4 .5, 4 .6, 4 .7? Oh, dude, I don't know.
I think the naming conventions, is that driven by the... marketing department or is it a it should be driven by the marketing department i mean if it's not that's a major l i mean opus five fable five mythos like these are all starting to get really confusing and i think part of the reason why is because they're not driven by the marketing department departments i feel like they're actually driven by some weird homage to like i don't know like oh it's an opus model if it was like pre -trained on this date and then it's a fable model if the pre -training is different But then they post -train them, and that's why they go 4 .5, 4 .6, 4 .7.
Do you know what I mean? 100%. I think, like, if you look at GPT, for example, what it needs to do is stick with its Sol Terra Luna pricing.
And the only circumstances under which you want to distance yourself from that is if the models become known as garbage. Like, if they become like, oh, dude, that's like a terrible model, then you can rebrand. But I'd say stick with the model families.
Yeah. But at a certain point, you just want new. You just want to create something new that sounds new, that sounds cool and rock and roll with it.
Agreed. And speaking about new, we have Kimi K3 Cyber Test. Essentially, X said that a Chinese open model was writing a bunch of working exploits for a variety of highly sensitive data.
Well, two major world governments tested Kimmy K3 and found that it scored zero out of 41 on their big cyber test bench, meaning that a lot of people were probably overly exaggerating just how crazy good and scary Kimmy K3 is at cyber. And I wonder why they do that, Jack. What do you think?
Yeah, it's an interesting one, isn't it, actually? Because they're running it through tests, and oftentimes you have these viral claims. I don't know how substantiated some of the claims are, because I think it started with a claim on X, right, Nick, where researchers said that K3 agents found...
I wonder if I can actually pull this one up on my screen so we can have a quick look, actually. That Kimmy K3 exploited the latest Reddit survey with zero data discovered. So let me come and share this, actually, so you guys can have a quick look at this, because...
This is kind of what happened. And this is part of the idea here that people make these climbs and you basically say, hey, the latest crazy model did this thing that you wouldn't even believe. So this was the three that went viral, 1 .3 million views.
The Kimi K3 exploded the latest ready server with a zero day it discovered. All it took was 27 minutes and 32 agents. And it found this particular book and appointed towards a GitHub, basically explaining it in a little bit more detail.
This was the claim. And so essentially from this claim, there was this whole, you know, you know, immediate storm of wow, Kimi K3 is crushing it. I think people want to believe in these super intelligences.
So we're more likely to do it until we get some robust testing. But if I think about my experience with Kimi K3, it is fable five quality, but a third of the price from a design point of view. So it is.
Very, very powerful model. So I haven't personally done it from all the security testing. But what I would probably say, Nick, is that we don't actually have the ability to test Fablefy properly for security testing anyway.
So we couldn't posturally even use it as it is. That's true. Yeah, I mean, like, I'll be honest, I don't actually use these models for any sort of cybersecurity or really security purposes at this point, aside from just hardening my apps.
But, you know, when you actually run them through a really big, robust test and you find that. x is saying one thing literal major world governments people that are actually extremely genuinely concerned with the security implications of these models uh say something different you know it it completed zero out of the 41 exploit bench tasks keeping in mind that like gpt6 is rumored to have absolutely crushed exploit bench mythos is rumored to have absolutely crushed exploit bench and i think the The value here, actually, if you think about it from like anthropics and open as perspective is the whole idea behind Kimmy is that they distilled a big chunk of fable and then, you know, probably GPT for other models and stuff like that.
The main thing that American governments are worried about are. you know, Chinese cybersecurity capabilities. So if the distillation of American models is not producing those cyber capabilities because there are blocks and safeguards around what constitutes an acceptable use of Fable 5 or whatever is working, then that means that this is pretty solid.
You can have Kimmy reproduce the design coolness of Fable, and that's all good, but it doesn't actually become a lethal superweapon.
you know what i mean so like that's that's positive for the american government like now or maybe before They know that like, yeah, sure, people can distill our models and they can get tons of likes and, you know, tons of views on YouTube videos and X posts about how cool the designs and 3D worlds are. But they don't actually get our cybersecurity capabilities, which means we also still have the advantage.
They only scored 32 % on the joint eval below leading US models. Zero out of 41 tasks actually reached code execution. It scored a little bit above GLM's 24%.
So it's not harmless, but 32, 24. I don't know what Fable and Mythos scored. guarantee you they were higher than 32%.
That's interesting. It's almost like the ability to do cybersecurity analysis is like a different skill. It's like in life, we have our mathematicians, we have our, you know, our literary scholars and those who are bodybuilders, right?
Our bodybuilders that sit with the, did they say that, um, stripes are good for broadening the shoulders actually, like usually a great choice, uh, accentuates it. So it's like, I think it's like these models are experts in different things. And I think that kind of makes sense if you think about it.
Do you have a system where you have one Superman that can do everything? Or is it more like the Avengers model where there's a Thor, there's a Hulk, there's a time traveler, there's a wizard? Do you know what I'm saying?
It's kind of like different experts, different areas of expertise. Claude, interestingly, anthropic models seem to really crush the code side of things and also the design. That said, though, guys, a lot of people are really bullish on the GPT models for coding.
and it will routinely catch things that opus misses true i'm also a big fan of gpt 5 .6 souls tone of voice relative to fables like we had fable help us design some of that website that we were just showing you guys earlier and i found there was like a distinctive like fable smell a lot of them are opus 5 smell like it would write you something and it'd be like here are the receipts and you'd be like receipts like i don't really know if dokes meme i don't really know if a real human being would put it that way but anyway when i prompt a gbd 5 .6 soul i just don't get that effect it's so interesting it just cuts straight to the point no bullshit no like great idea nick well i think it's just like yeah here's the answer and i'm like oh my god so i mean i would be game if okay can you imagine if kimmy k3 is distilling anthropic models and then glm is distilling open ai models i would be game to see a model entirely distilled off of gpt that i could then run up on my own
I'm like crazy fast infra. Because then I could just like write everything I've always wanted to write my whole life without it sounding like AI. Yeah, it's definitely the human centipede of kind of like distillation.
It kind of keeps on going down. Yeah, it's funny, dude. I agree.
I agree. I think tone of voice is really, really important and language. But I often think it's style.
A lot of it's taste. And the more you use it, the more you just identify a pickup of these kind of like weird quirks. But even still.
we always keep things 100 % real on the stack podcast. As you know, we'll keep it 100 % real for our 17, the followers and watch this podcast is that they're still not perfect. Like I was, Nick and I were chatting offline about, I was doing a build yesterday and it was going on for five hours.
And like, sometimes your emotions flare up and you're like, dude, like I even said, we've been doing this for five hours. Like just figure it out. Like just.
fix this thing and obviously we know all the strategies and techniques but but even i was like who uses stuff every single day you can still have times where models don't misbehave that don't do the things that they should be doing so they're not yet at the standard where we give these magic prompts and everything's done i believe we're in the golden zone actually um seven years ago it was effectively a load of developers that you had to give you know you drop a pack of red bulls at their desk and leave them in a room for a few hours and your code would appear and probably in seven years time we are at the point where we're at one sentence builds but today we're at the point where those with the will and the ability to be industrious and enterprise can build anything they imagine and that's what i think is the golden zone for sas it's the golden zone for creating applications and for growing your business i guarantee in 20 years time we will look back on this era and say dude we had a
a huge asymmetry advance asymmetrical advantage here and it this will end 100 most people don't know that it exists but in a certain period of time the asymmetry will disappear when the models go uh you know they're basically one talking dumb so take advantage of that uh whilst it exists 100 i agree 100 i find it really funny because you're like seven years ago we would lock devs in a thing and give them red bulls And now, you know, we talk to it and kind of wrangle it.
I bet you in seven years, AI will be locking us in a cage with Red Bulls. If it's diet Red Bull. If it's diet Red Bull.
I've not had an energy drink in some time, actually. Yeah, it's been a hot minute since I've had a Red Bull, dude. Have you had an energy drink at all in like the last little bit?
No, no, I don't think so. I quit a couple of years ago. I remember in the UK, we have this brand called, they have a different like animals on the energy drink.
Like one's called like, rabbit energy or like um gorilla energy and i had a couple of these and i was like this probably is not the best thing in the world for me so i better stop drinking these do you want to feel like a rabbit i want to feel like a gorilla personally no the gorilla energy yeah before you grow like a third arm or something from drinking too much i don't know like all right you want to do some q a let's do that let's see what we're dealing with so now we're back to daily we'll work back to we just started but we're starting daily for anybody that doesn't know we publish these once a damn day And this is quite the transition because we used to publish them once a week.
Our goal is five days a week minimum, Monday to Friday. I think it just makes sense because that's what people's routines and stuff are like. But we'll drop in a Saturday or Sunday episode every now and then just to stoke you guys up.
So what that means is if you guys have any questions that you guys want us to respond to, just ask away. And Jack and I will answer them at the very end. all right so mao mao speaking of says kimmy k3 is cheaper per token yes but is it cheaper per task completed check out the pricing for soul luna opus 5 and so on versus how much it takes to complete a task well didn't we have that we just go back to this lovely graph here that um jack so handsomely put together i mean Yeah, the point he's making is correct.
It's cheaper per token, but it's generally less efficient per token. Wait, hold on. I don't see Kimi K3 here, though, Jack.
You see every model here but Kimi K3. Too dangerous, Nick. What does that mean?
5 .2 is over here. Oh, my God. Are they hiding?
I didn't even realize this. You know, one best way. Like, they just don't put the freaking model on there.
There's so many, you're going to forget. You're not going to see it. That is hilarious.
Dude, can we also talk about, like... Let's just end benchmarks for all time. Yeah.
Let's never use a benchmark ever again. Yeah, we should. We should not use a benchmark again.
Like, who cares about what the numbers say? How's it actually, you know, used? Kasif was watching the pod with his mom, which is a very fun activity Jack and I would recommend you guys do whenever possible.
And it was on the TV. Obviously, he doesn't understand all this deep AI rabbit hole stuff. She asked me the most hilarious thing ever, which was, why are these two guys talking about models?
Out of context. That was so freaking funny. We do love talking about models here, don't we, Jack?
We do. We talk about them all the time.
HTBlind says he can't believe that we don't have more subscribers on here. Neither can I. Where's my tank top?
I don't know. Where is my tank top, Jack? I'm just trying not to look small.
Just take this off. Jesus, man. I thought we were doing tank tops today, Nick.
We had an emergency pre -meeting meeting. When I tank tops, Nick, I'm doing solo over here, so we'll make it like a thing eventually. It's been a lot of heavy lifting.
Yeah, I'm going to have to break out the tanks. Okay, this is a good question that I think people would enjoy. Oh, and the backwards hat too.
Well, cool. Could you share more, Jack, and this is a question to you. Could you share more about what you do with your money?
It looks like your businesses generate huge revenue with low fixed costs, leading with a big cash surplus every month. Entirely correct. What is your approach to this?
Where do you invest and park that money? Are you planning any new, more capital -intensive ventures? Also, what were your most extravagant purchases when she made that kind of money you used to dream about?
Or are all those big splurges still ahead of you? And then we have a little quote here, which, um, I don't know. I think this is a great question, Nick.
I think this is an awesome question. Okay, so what do we do with money? Basically, you want to maintain like high margin business.
So some you want to reinvest. But honestly, the way that I personally think about it is I set a number that I want to invest every single month. And my job is to make enough to hit that investment number, which is from business to like invested.
That's the way that I think about it. And then I have been thinking about this a lot in terms of like strategies, planning, things like that. That's what I tend to do with the dude.
You have to be careful not to let your lifestyle as you continue to grow and grow and grow exceed with your income. So you have to be mindful of that. Like, you know, keep that under control.
But I haven't bought any crazy extravagant things yet. I haven't like gone and bought something crazy. The bits that I have invested in more is like travel.
So traveling well, like flying and staying. And then the other thing is just food. So there's an unlimited budget for anything to do with health.
And the rest, dude, honestly, just keeping it WhatsApp and then just keep investing. The big transition moment that you want to get to, because you can be what's called a Henry, which is high earning, not rich yet. So some people fall into that category.
And the idea with that is that like a... A lot of people you see in business or sometimes, you know, with high -powered jobs have a high income, but they're not wealthy because they make, say, 100K a year, 500K a year, 5 million a year, but the exact same amount of money goes out, which means if they stopped earning, they would kind of like, you know, have to change their lifestyle.
So the idea is you create a series of assets that makes more money than you do. If you ask very rich people when they feel rich, that's the inversion moment where your wealth makes more than you do. So that's kind of like the ultimate.
like level if you're thinking about it that way but personally dude no i've not gone out and bought like lamborghinis but nick i see this a lot especially with people that may start to make 20 50 100 200 400k a month like you start seeing them in gucci you start seeing them with uh lamborghinis um and it just goes in one and straight out the other that's kind of interesting dude i see that people make 20k a month um they'll make 20k a month and then they'll just like shit drip the hell out my old boss actually back when i was doing door -to -door stuff i distinctly remember a moment that i was like all right i probably cannot work here for much longer came into some like sales meeting and i was trying to fire us all up and like get us super excited about hitting the streets and stuff and he was probably making like 20k a month us at the time like you know this was not that crazy big of an agency it was like reasonably big maybe he was making like 30 30 40k who knows but um i didn't you know to me at the time that was like
elon musk wealth to be clear i was poor i was eating rice and beans my food budget was like 125 bucks a month so he came into this meeting and was trying to stoke us up and then he had on like i think there's a gucci scarf or something i don't know if they even sell scarves but it was like one of those super crazy brand names and he was just like so why y 'all doing this or something and i was like uh you know safety and security for my family he's like no bro you're doing it for this you know he was dripped out he had like chains and rings and shit he's like y 'all want that money i want that paper man people need to know who you are and i was just like i feel like i feel like i'm gonna make way more money than you and i will never ever own a 750 dollar qg scarf while doing it so it really depends you know on like who you are as well but i think the good news about jack and i is that we're just not very flashy people we prefer to spend money on health and you know wellness and investing in the future
100%. That's a good question. I think that's really, really timely, very healthy question.
You can't find what you're looking for in physical purchases. Unless it's a tank top. Unless, yeah, or our soon -to -be 17 merch where you'll have cool hats that say, you know, one of the 17.
We'll figure out the best way to do that and then we'll put that up. Dude, and I think, final question, if you were to get, if there was a merch thing, what would it be? Would it be tanks?
Would it be... Models, I don't know. We'll let the people decide.
Let us know down below, guys. Have a beautiful day and we'll catch you inside the next episode. See you guys soon.
The Hook

The bait, then the rug-pull.

Two hosts who between them spend serious money on inference every month open with a claim that sounds like hype and turns out to be an accounting argument. If the cheaper model finishes the same job in fewer tokens, the leaderboard stops meaning anything.

CTA Breakdown

How they asked for the click.

MENTIONED ON CAMERA
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.

36:00
Theo - t3․gg · Review

GPT-5.6: The Review

Theo spends 36 minutes putting real numbers behind the GPT-5.6 hype — Sol, Terra, and Luna, benchmarked against Claude Fable, one blog chart at a time.

July 12th
08:54
Matthew Berman · Review

GPT-5.6 is FINALLY HERE (WOAH)

A 'dot' release plays out like a full generational leap: two five-to-seven-day unsupervised coding runs, a sponsor benchmark, and a live pricing and capability standoff against a rawer, higher-ceiling rival model.

July 9th