Opus 5 is priced at half of Fable 5's per-token rate and often wins on following instructions and verification, but across eighteen real sessions it still spent more in total because it used roughly three times the active runtime, proving that per-token price and total task cost are not the same number.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You're actively choosing between Claude models for coding, content, or research work and want evidence from real tasks, not just benchmark charts.
You run agentic coding sessions and want to understand why a 'cheaper' model can end up costing more per task.
You're building multi-agent or orchestrator workflows and want a concrete argument for splitting delegation from execution across models.
SKIP IF…
You want a definitive 'always use X' answer — the video's own conclusion is that the winner flips by task type.
You need rigorous, controlled benchmarking methodology — this is one person's subjective real-workflow test, run once per task, not a repeated statistical trial.
TL;DR
The full version, fast.
Benchmarks showed Claude Opus 5 beating Claude Fable 5 on several coding indexes at half the per-token price, so the video runs both through nine matched real-world tasks each — bug fixes, a promo video, a landing page, LinkedIn content, audience research, a slide deck, a computer-use snake game, and a structural simulator — grading cost, time, and quality on each. Opus generally followed instructions more precisely and verified its own work more carefully, occasionally producing dramatically better technical results (93/95 vs. 66/95 on one bug hunt); Fable generally won on design taste and speed, and was frequently the cheaper option despite costing double per token. Aggregated across 18 sessions, Opus's total spend ($243) still beat Fable's ($165) in the wrong direction — Opus needed roughly three times the active runtime, so its per-token discount didn't translate into a cheaper total bill. The actionable conclusion: match model intelligence to task difficulty rather than defaulting to the flagship, and consider using the stronger-reasoning model purely as an orchestrator that delegates without writing code itself.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Opus 5 launches priced at half of Fable 5's per-token rate while beating it on several coding benchmarks — Nate sets up the real-workflow test to check if that holds.
01:00 – 03:17
02 · Testing without a harness
Both models generate an Excalidraw diagram of semantic search from an identical prompt with no coding harness and no way to verify their own work; Fable's version reads as more visual, Opus's as more organized and detailed.
03:17 – 05:19
03 · Two codebase bug hunts, opposite results
On the first bug hunt the models tie, with Fable's patch judged cleaner; on the second, identically-structured hunt Opus scores 93/95 to Fable's 66/95 and is still the cheaper run.
05:19 – 06:51
04 · Why verification matters
Anthropic's release notes credit Opus 5 with being much stronger at verifying its own work and iterating until it succeeds; Nate argues how you instruct a model matters more than which model you pick.
06:51 – 09:37
05 · AIS Live announcement video
Same prompt for a 10-second hyped announcement video: Opus produces two variants (vertical + landscape) for $11.19/40 min, Fable produces one for about $7/8 min. Both include some outdated event details.
09:37 – 12:15
06 · Landing page build
Fable builds a comparable-quality landing page for $20.50 in 22 minutes; Opus takes nearly an hour and $35.83 for a similar result.
12:15 – 14:38
07 · LinkedIn post & carousel
Fable's carousel style is preferred and it's cheaper and faster ($6.17/7.5 min vs. Opus's $8.22/longer) — a sign Opus burned more tokens on the same job.
14:38 – 16:58
08 · Audience research report
Both models mine the same YouTube comments and community posts and land on nearly identical pain points and product recommendations; Opus is cheaper despite running longer.
16:58 – 20:18
09 · Video outline & slide deck
Working in a fresh project with no prior style memory, Fable's Excalidraw-style deck matches Nate's usual visual style better than Opus's, though Fable costs more ($40 vs. $33) and runs much faster.
20:18 – 22:39
10 · Computer-use snake game
A prompting mix-up runs Opus against itself twice with the same instructions; one run ignores the '5 games only' instruction and plays roughly 30, the other follows it exactly — same model, wildly different behavior.
22:39 – 27:41
11 · Structural simulator build
Fable finishes a physics-stress-test simulator in 7 minutes for $73; Opus takes over two hours and $112 — the most expensive session in the whole test — and produces the more overwhelming, less usable UI.
27:41 – 29:19
12 · Total cost & token breakdown
Across 18 aggregated sessions, Opus spent $243 total against Fable's $165, driven by roughly 630 minutes of Opus active runtime versus 202 for Fable; cache re-reads dominate spend on both models.
29:19 – 30:53
13 · Final thoughts
Nate's takeaway: match model intelligence to the task, and consider running the stronger-reasoning model purely as a delegating orchestrator instead of an executor.
Atomic Insights
Lines worth screenshotting.
Claude Opus 5 launched priced at half of Claude Fable 5's per-token rate while benchmarking above it on several coding indexes, including a frontier bench and a coding agent index.
On a bug-hunt task run against one codebase, Fable finished in 11 minutes for $5.30 and Opus in 13 minutes for $4.22, and an independent review judged the results roughly tied, with Fable's patch called more immediately reviewable.
On a second, similarly-structured bug hunt, Opus scored 93 out of 95 on a technical review versus Fable's 66 out of 95, passing 4 of 4 tests where Fable passed 2 of 4 — and Opus was still the cheaper run at $6.50 versus Fable's $8.73.
Given an identical creative-video prompt, Opus produced two output variants (vertical and landscape) for $11.19 in 40 minutes, while Fable produced a single variant for about $7 in under 8 minutes.
On a from-scratch landing page build, Fable was both cheaper and about three times faster than Opus — $20.50 in 22 minutes versus $35.83 in nearly an hour — for a comparably-scored result.
On a LinkedIn post and carousel task, Fable came in cheaper and faster than Opus ($6.17/7.5 min vs. $8.22/longer) — since Opus costs half of Fable per token, any run where Opus costs more than Fable means Opus used far more total tokens to do the same job.
Both models independently converged on nearly identical conclusions when mining the same YouTube comments and community posts for audience pain points and product ideas, despite Fable reading fewer total source posts.
Running the identical prompt against the identical model twice produced wildly different behavior in a computer-use test: one run ignored an explicit 'only run 5 games' instruction and ran roughly 30 games chasing a higher average score, while the second identical run followed instructions exactly.
The single most expensive session in the entire test was a structural-simulator build that took Opus over two hours and cost $112, versus Fable finishing a comparable build in 7 minutes for $73.
The simpler, more polished, more usable interface in the structural-simulator comparison was produced by Fable, not Opus — the opposite of what the person running the test expected before the reveal.
Aggregated across 18 sessions (10 run on Opus, 8 on Fable), Opus's total spend was $243 versus Fable's $165, even though Opus is priced at half of Fable's per-token rate.
Opus's total active runtime across its sessions was about 630 minutes versus roughly 202 minutes for Fable — Opus needed about three times the working time, which is why its cheaper per-token rate didn't produce a cheaper total bill.
On both models, re-reading context (cache reads) dominated total spend and raw output generation accounted for only about a fifth of the bill — the cost of an agentic session is mostly the cost of it re-checking its own prior context, not writing the final answer.
Anthropic's stated improvement for Opus 5 is that it is 'much stronger at verifying its work and iterating carefully until it succeeds' — the video treats this as the single biggest practical differentiator between the two models, more than any benchmark score.
Takeaway
The cheaper model per token isn't always the cheaper model per task.
WHAT TO LEARN
Across eighteen matched real-world sessions, the model priced at half the per-token rate still spent more in total because it ran roughly three times longer — proving that per-token price and total task cost are two different numbers, and only the second one should drive a budget.
02Testing without a harness
Without a coding harness or verification loop, the two models produce noticeably different visual reasoning on an identical prompt — style differences show up even before any task-correctness testing begins.
A raw output comparison with no way to verify correctness mainly tests a model's default formatting taste, not its real capability.
03Two codebase bug hunts, opposite results
On an identical bug-hunt prompt run twice against different codebases, the same two models swapped rankings — cost and quality don't reliably move together, and the cheaper run was sometimes also the more correct one.
A single benchmark run tells you almost nothing about which model will win on your actual codebase; run the same task on both before committing to one for a workflow.
04Why verification matters
The single biggest practical differentiator described for the newer model is how carefully it verifies its own work and keeps iterating until an explicit stopping condition is met, more than any raw benchmark score.
The way to get better verification behavior out of any model is to give it an explicit, testable stopping condition, or have it argue across sub-agents until they reach consensus.
How a model is instructed matters more than which model is picked — the same prompt discipline improves either model's output by a similar margin.
05AIS Live announcement video
Given identical creative-generation prompts, one model produced two output variants where the other produced one — models differ in how much they choose to do beyond the literal ask, not just in raw quality.
Neither model caught outdated source facts on its own; better verification, not a better model, would have caught the stale information either way.
06Landing page build
On a from-scratch landing page build, the cheaper, faster model produced comparable design quality to the slower, pricier one — the more expensive model isn't automatically the better choice for design-heavy tasks.
Design-and-creative tasks and code-correctness tasks favored different models in this test set, which argues for matching the model to the task type rather than defaulting to one model for everything.
07LinkedIn post & carousel
Any run where the model priced at double the per-token rate still comes out cheaper is the clearest possible tell that the other model used far more total tokens to do the same job.
For straightforward, on-brand copywriting, an even smaller and older model was judged sufficient — running the flagship model on a task a mid-tier model already handles is often wasted spend.
08Audience research report
Both models independently converging on nearly identical conclusions from the same open-ended research inputs is a decent signal the conclusion is supported by the data, not just model bias.
The amount of source context a model reads in doesn't automatically translate into higher cost or a meaningfully better answer.
09Video outline & slide deck
Working in a fresh project with no prior examples in context produced results that diverged sharply from the creator's established style — a model is only as visually consistent as the context and prior examples it's given access to.
Matching a model's default aesthetic sense to a visually-judged deliverable is worth paying more for, even when the pricier run also takes longer.
10Computer-use snake game
Running the identical prompt against the identical model twice can produce wildly different behavior — one run can ignore an explicit numeric instruction while another follows it exactly, which is hard evidence these systems are non-deterministic even run-to-run.
Don't diagnose a 'bad model' from a single run's unexpected behavior; the identical model can behave completely differently on the very next attempt.
11Structural simulator build
The most expensive single session in a whole test set can come from a model over-verifying and over-testing a task well past the point of diminishing returns — stronger verification isn't free, and left unconstrained it can blow up cost and time on a task that didn't need it.
Judging visual or UX quality without knowing which model produced which output is a useful check against assuming whichever model you 'expect' to be better actually is.
A model capable of much stronger work can still produce a mediocre result on a given run and simply stop short — a single generation's quality isn't a reliable ceiling on what a model is capable of.
12Total cost & token breakdown
Total cost across many aggregated sessions is driven by total active runtime as much as by per-token price; a model that's cheaper per token can still cost more overall if it needs several times longer to finish the same work.
Re-reading prior context dominates the bill on agentic sessions far more than generating the final output does — the expensive part of an agent session is it re-checking its own prior work, not writing the answer.
The single most expensive session in a dataset can outweigh several cheap sessions combined, which makes median cost per session a more honest budgeting number than a simple total-divided-by-sessions average.
13Final thoughts
Running the stronger-reasoning, more expensive model purely as an orchestrator that delegates and reviews — without writing code or executing anything itself — lets a session run for hours without approaching a context-window limit.
Matching the intelligence of the model to the intelligence actually needed for the task beats defaulting to the most capable, most expensive option; both flagship models tested here were judged overkill for some of the simplest tasks.
Glossary
Terms worth knowing.
Coding agent index
A benchmark category that scores how well a model performs when acting as an autonomous coding agent across a suite of tasks, rather than answering single isolated coding questions.
Frontier bench
A benchmark suite intended to measure a model's performance on the hardest, most current tasks relative to other 'frontier' (top-tier) models.
Verification loop
A pattern where an AI agent is given a way to test whether its own output meets a condition, and is instructed to keep iterating until that condition is met rather than stopping after one attempt.
Orchestrator model
In a multi-model workflow, the model responsible for planning, delegating, and reviewing work rather than writing code or producing output directly — kept lightweight so it can run for long stretches without exhausting its context window.
Cache read tokens
Tokens from prior context that a model re-reads from a cache rather than reprocessing from scratch; billed at a lower rate than fresh input but still a real cost that accumulates over a long agentic session.
Context engineering
The practice of deliberately managing what information an AI agent has in its working context — writing, selecting, compressing, and isolating it — rather than relying on a bigger context window alone.
Resources
Things they pointed at.
00:05linkIntroducing Claude Opus 5 (Anthropic blog post)
05:20linkIntroducing Claude Opus 5 (Anthropic blog post, verification claim)
53:50toolCodex (used as independent judge of Opus 5 vs. Fable 5 outputs)
09:37linkNate Herk's free Skool community (full cost/token breakdown document)
Quotables
Lines you could clip.
00:20
“It shows on things like the frontier bench and the cursor bench and this coding agent index that Opus is actually outperforming Fable and it's cheaper, and this really shocked me.”
sets up the entire video's premise in one line→ TikTok hook↗ Tweet quote
13:30
“If Fable and Opus use the exact same number of tokens, both input and output, then Opus would be pretty much exactly half the cost. So when Opus comes in costing more than Fable, it means it was way less token efficient.”
the single clearest explanation of the video's core finding→ IG reel cold open↗ Tweet quote
22:25
“That's just a really good reminder that, like, at the end of the day, these things are completely nondeterministic. You don't know what they're gonna do.”
punchy, standalone caution about agent reliability→ newsletter pull-quote↗ Tweet quote
27:10
“I think that Claude Fable could have easily designed something and built something this level of detail and much better, like far far better... but it just decided to be done.”
counterintuitive claim that output quality isn't a ceiling on capability→ TikTok hook↗ Tweet quote
30:00
“You shouldn't be writing any code or executing anything. You should just be telling Opus what to do, and that saves your session limit with Fable big time.”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphor
00:00Alright. So Claude Opus five is here. And if you start to look at the benchmarks, it's really interesting because it shows us that for a lot of things that I care about, it's actually better than Fable, and it is half the cost of Fable.
00:12Ultimately, Fable five is still Anthropic's most impressive and, you know, strongest model. But for a lot of these things, you know, I've realized when I'm doing knowledge work and when I'm building, you know, my videos or my research or whatever it is, Opus is more than enough power than what I need. And when you look at some of these charts, it's really interesting because it shows on things like the frontier bench and the cursor bench and this coding agent index that Opus is actually outperforming Fable and it's cheaper, and this really shocked me.
00:38So obviously, I like to take all of this stuff with a grain of salt. It's fun to look at, and it's good to look at, but you want to actually get your hands dirty and run these models through your own actual workflows. So today's video, I'm just gonna break down a bunch of different experiments that I ran with Opus five versus Fable five and break down things like the cost, the time, and the tokens so that you can start to understand where you should work in these different models within your workflows.
01:00Alright. So pretty much all of the experiments that I've been running today that I'm gonna show you guys, I did within Claude code, which means we're comparing the models, but also inside of the Claude code harness. And the variable is the same, so it doesn't really change too much.
01:12But I did do a few tests where I was actually in Claude chat, and I was just, you know, seeing how they felt without a harness wrapped around. And let me just show you one quick example. So here I asked Fable five and Opus five to generate me an Excalidraw diagram that accurately and visually explains how semantic search with vectorization works on a large dataset for AI agents.
01:30And it's interesting here because there's no skills that it can use, and it doesn't have any context of me or, you know, any really way to verify. All it did was it spit out a JSON file of Excalidraw for me, and then I pasted it into Excalidraw.
01:44So here's what we got. Fable came back with this version over here where we see we've got, like, our indexing pipeline. We have a large database, and it looks like it actually misspelled this right here, which is interesting.
01:55Large oh, dataset. Okay. It was just, like, not expanded enough.
01:58Same thing over here, vectorize. And this is part of the whole it had no way to verify.
02:03And as you guys know, if you've been kind of building agent loops and stuff, verification is so, so important. So, anyways, large dataset. We chunk it up.
02:11We vectorize it with an embedding model. We then get our embeddings, which is just like the numerical representation of the data. We put it into a vector database here, and then we can actually start to search.
02:21So we've got similarity, and that's on, you know, points being close together. So your question lands here, and it would grab the k nearest neighbors.
02:29And different meaning is farther apart. We've got different clusters here. And then we come down here to the actual query.
02:35So if the user asks how I get my money back, the agent searches the knowledge base, and then it does semantic search. It looks it up with the actual query vectors, and we get the matches back.
02:44So pretty accurate. I will say though, Opus's layout seems a bit more organized.
02:50Right? Like, it's it's got boxes and it's got I mean, this might not be as visual. You could argue You could argue that fables was more visual, which, you know, I think that that's true, but this definitely feels more organized.
03:00It feels a little bit more detailed as well. So that's just a very subjective exam.
03:05A lot of this stuff that I'm gonna be talking about today is just really opinionated and subjective, but I'm gonna still give you my honest thoughts. So in this example, I think that if I wanted to teach someone, I probably would take Opus five's version here.
03:17Okay. So let's start off with the first test I ran, which was basically giving them a huge code base and having them look through any bugs and looking through, like, the expected behavior and some instructions like that. So it had to do some exploration here and help us out.
03:31Right? So I set the goal, and I gave it this prompt. And then what I did is I had Codex review the output that Fable five gave us and that Opus gave us.
03:40So real quick before we look at the results, Fable took about eleven minutes, and it costed $5.30, whereas Opus here took thirteen minutes. So a little bit longer, but it was cheaper at $4.22.
03:50You can also see the breakdown here of input and output tokens and, like, what models they use and stuff like that. But let me switch over to Codex here. This is the actual result.
03:58So head to head, they both pretty much passed everything, which is great. But codex thinks that Fable wins here because Fable's production patch is exactly the one line upstream fix, blah blah blah.
04:08You guys can read through this if you want. But the final ranking here was that Fable and Opus did similar.
04:14Right? But Fables was a little bit cleaner and immediately reviewable. But now let's take a look at the second one.
04:19So I did a very similar example on the second one where I gave them both, you know, the same prompt, the same code base as you can see. Here was the repo. Here was the bug.
04:26Here's example, expect behavior, blah blah blah. So in this case, Opus took twenty minutes ish, and it costed us $6.50, whereas Fable on the exact same prompt, exact same code base, took twelve minutes and costed us $8.73.
04:40So let's go see what Codec said about these results. So here Right? Because Fable versus Opus, we actually had Opus perform better.
04:48Four out of four passed right here, whereas Fable only passed two out of four and left this unresolved. So the technical score for Opus was 93 out of 95, and Fable scored 66 out of 95, which is really, really interesting. And think about the fact that once again, Opus in this case was cheaper for us.
05:04The bottom line was both agents demonstrated strong repository navigation and independently found the core architectural issue, but Fable's patch is functionally close and fixes the user facing update bug while Opus delivered the more accurate, thoroughly tested benchmark passing implementation. And one thing that I was really excited to see in this release blog from Anthropic, if I can keep scrolling down here and find the right spot, they basically talked about how there was a huge improvement in OPUS five.
05:29Okay. Let me just find this real quick. OPUS five is much stronger at verifying its work and iterating carefully until it succeeds, which is huge.
05:35Like I kind of alluded to earlier, verification has become one of the most important things that you can do for your AI agents, essentially saying, hey. Don't stop until you hit this condition, and this is the stopping condition.
05:45Here's how you can test if it's actually done or not, which basically means if you want them to not stop until, you know, a certain metric is hit, like, explicitly 10 out of 10 of this objective criteria. Or if it's something a little bit more subjective, you can have them spin up sub agents that have to argue and debate, and you have to keep going until all five of them come to a consensus or something like that.
06:04Basically, just the ability for the AI to build something, design tests to see if it's done or if it's good, and then keep iterating until the test passes. So that's something I've realized as I've been testing out different models and different harnesses. It's like, yes, that does matter.
06:19But at the end of the day, what matters way more is how you instruct it and how you feed context in. So just keep that kind of stuff in mind. Yes.
06:27It's good to find the best model, but find the best model for your use case and then understand how to talk to the models right. By the way, guys, as I'm editing this video, I just wanted to say that at the end, I go over, like, a snapshot of all of these experiments and see, like, total cost, total tokens, total time to run.
06:43So if you kinda wanna skim through the experiments, feel free. If you wanna jump to the end and see, like, the total consensus, then that's there. So I just wanna let you guys know, but let's get back to the video.
06:51Okay. Let's take a look at our next example here. So in this one, what I did was I said, hey.
06:56Slash goal, build me a ten second hyper edited engaging viral worthy announcement video for AIS Live. So that was our live event that we just did. And I told her that it could look through whatever it wanted in my code base and my entire computer to figure out, you know, the information to use.
07:11So let's take a look at the two examples and how much they costed us. So this one was Opus. Right?
07:15And I'm gonna open up this example right here. This one actually created us a vertical version and a landscape version. So let's take a look.
07:33Okay. That was the vertical. Let me play the landscape real quick.
07:48Okay. So not bad, but not great. It only had ten seconds to work with, and it was pretty fast paced.
07:53And it honestly looked just a little bit, like, computer y, like, not super professional. But, anyways, let's see how much that actually costed. That took forty minutes, and it costed $11.19.
08:13Okay. Very cool. So that one, you know, they they had different sound effects.
08:17They had different music. Also, you'll notice here is that some of that was inaccurate. Don't if you guys went to AIS Live or not, but some of that data was outdated.
08:24Now that's not Fable's fault. That's probably more on me for not keeping that context completely updated inside of my OS. But if it would have been better about verification, I'm pretty confident that it could have found the right stuff.
08:35If it would have dug deeper into my school community and looked for threads, if it would have looked at some other LinkedIn posts or other things that I've done, it could have found out that some of that stuff was inaccurate. Like, some of those speakers didn't end up speaking at this event and stuff like that. So there was a little bit of, like, a context issue there.
08:48But as far as the actual videos, that's what we got. Right?
08:51And they're they're different. You can tell that the models have different tastes. They were given the exact same prompt.
08:56Opus decided to make two. Fable decided to make only one. So let's take a look at the cost of Fable.
09:01This one was obviously, you know, about half the cost, about $7 and seven minutes and forty nine seconds. So Opus took a lot longer here, but it did decide to do basically double the output. And by the way, at the end, I'm gonna show a full breakdown of all of the total costs, time, input output tokens, so just stay tuned for that.
09:17But let's keep flying through some of these other experiments here. And by the way, if you wanna access this entire free breakdown document as well as all the other resources that I ever give away on YouTube for free, just go to my free school community. The link for that's down in the description.
09:28You'll go to classroom. You'll go to all YouTube resources, and then you'll find everything in there for free.
09:34So let's get back to the video. Okay. So I think you guys get the point.
09:38All of these other experiments, we're gonna do the same thing, exact same prompt. I'll show the cost, and we'll look at the outputs. So this time I asked for a one page landing page for a certified AI consultant program.
09:47I told it it could look through whatever it wanted and, you know, make it as impressive as possible. So let me real quick pull up both of these outputs. Okay.
09:55So here is the Opus output, not another AI course, the credential the SMB market hires from. You can see there's a bit of a dynamic three d element in the background. It pulled my logo.
10:03We have an apply button. We can jump down to certain sections. That's pretty cool.
10:07If I go to the nine pillars, you can see what this course is actually designed around. I mean, this does look pretty generic. It has our brand guidelines though.
10:15It uses our colors. It uses our logos. It uses our different buttons and our different design kind of criteria, and it has all of this, which is so far what I can see, all of this is completely accurate information based on my meeting recordings, based on my, you know, internal docs about this cohort.
10:30As you can see right here, we've got this information as well. And, yeah, this isn't too bad.
10:35It's very wordy, but it's also very accurate. So let's switch over now to Fable's version of this and see what it did here. If I can find where it kept this HTML, here it is.
10:45I'm gonna have to open this up in the browser real quick. Okay. So here is Fable's version.
10:49It's very similar. We have the card right here, which is a nice touch because we do actually have, like, a card designed very similar to this. Not another AI course, the credential, the SMB market hires from.
11:00We've got nine pillars, two layers each, 54 job post analyzed, 50 founding seats, claim a founding seat. We can keep scrolling down. So they're very similar style.
11:07Right? This one has a nice little animation here. They both obviously pulled my logo and our brand guidelines that you can see that they're designed very similarly, which is great.
11:15You know? We have other stuff here like the disciplines. The pillars are the same once again.
11:20I don't know. I mean, they're they're obviously designed very similarly. This is a nice touch here.
11:24I think the thing about this is I probably would obviously wanna manually tweak both versions. They're very similar.
11:31I don't think one definitively beats the other. So let's look at the cost and the time.
11:35So Fable was $20.50, and it took twenty two minutes, whereas Opus was $35.83 and took almost an hour.
11:46So this is one of those cases where Fable was actually cheaper and quicker and arguably maybe just like a little bit better, but it was very similar on that side and very subjective.
11:56So maybe one conclusion we can start to draw here is that Fable still kind of wins on the creativity and the design side compared to Opus five. Whereas right now, Opus five is kind of having more of an edge for me on, like, actually following directions, verification.
12:11In most cases, it's gonna be cheaper once again. Okay. Let's keep on moving here.
12:15So experiment number three here was that I wanted a LinkedIn post with a LinkedIn carousel, and you can read the rest of this post here. But, basically, I was trying to raise the stakes.
12:23Right? I was trying to say, hey. You know, if I post this out on my audience, it would be bad if this was, you know, clearly AI generated or whatever.
12:30And I also had to choose the topic. So Opus, let's see what it decided to do. Here is the actual PDF it created for me as the carousel.
12:38I don't love this styling. Right? It's obviously pretty consistent with our brand guidelines, but it just looks a little bit, you know, meh.
12:45It just doesn't look super, super professional. So that's the carousel. It basically chose to write about AI got dramatically better at coding and trust in it went down.
12:53So let's see what the actual post looks like. It probably used my LinkedIn writing skill to do this. So here is the actual post.
12:59We've got some real stats in here. We've got the arrows, a reference to the actual carousel down there. Okay?
13:05And if I go to Fable version now and scroll up here. So Fable used a different style of carousel, which I actually like a lot more. It's kind of like that tweet style.
13:13AI agents went mainstream. Trust didn't. So that's pretty interesting.
13:17It did similar research, and they honestly both came to a similar conclusion on, hey. You know, based on Nate's audience and based on what's going on in space, what should we write about, which is pretty interesting. But, ultimately, I like this deliverable much better.
13:29And, I mean, honestly, I think that when I read through these, the actual content of LinkedIn posts are pretty similar as far as, like, which one do I trust more. You know, I've got a skill built around it. I've got a no AI slop sort of skill as well.
13:41And this is a perfect example of, like, writing a LinkedIn post, generating that content. I think even Opus five is overkill. Like, you could write really good content with Sonnet, you know, Sonnet 4.5.
13:51So those are pretty similar. Let's look at the cost and the time.
13:55Fable costed us $6.17 and took about seven and a half minutes, and Opus here costed us $8.22 and took longer.
14:02So another example where Fable actually came in cheaper and faster, which is quite shocking to me because what that tells us is that Opus is using so many more tokens to actually be more expensive. Because if Fable and Opus use the exact same number of tokens, both input and output, then Opus would be pretty much exactly half the cost, but that's not the case.
14:22So when Opus comes in costing more than Fable, it means that it was way less token efficient as well, which is a little bit concerning. Okay. So let's just keep on moving here, though, because, obviously, all of these experiments I could run is not gonna be the exact same as when you use it.
14:35And, you know, at the end of the day, it's a black box. You're pulling a lever on a slot machine. So let's move on to Fable or sorry, the the fourth experiment.
14:41So here I told it to go to my YouTube channel, pull comments, and then to go to my school communities and look through threads. And I wanna understand what my audience is saying, what the pain points are, the number one product that I could build to help solve their pain points, and the number one best YouTube video that would resonate with them.
14:55So let's open up the HTML here for Opus. What your audience is actually telling you, it pulled a bunch of sources. It pulled things from, like I said, right here, if I can keep scrolling up, 2,200 comments on YouTube, 480 school posts, it looks like.
15:11And here's what we've got. So we got an HTML. And once again, I don't love this font in general.
15:15I mean, this is on our brand guidelines, but I might wanna change that because it just looks very typewriter. It looks very cheap, honestly. Anyways, we've got pain points that are ranked, pricing and scoping, proving the automation worked, n n n versus cloud versus co work, token, blah blah blah.
15:28The number one product to build would be the offer engine. So a cloud code skill pack plus templates that takes a discovery call and then produces a scoped priced offer with a working measurement layer out the other. So a bunch of different skills.
15:41It tells us why. And then the number one video to make would be I sold an AI system to a real business in seven days, real client, real invoice. Okay?
15:48So let's see if Fable came to a similar type of conclusion with what I could build. So once again, this thing looked through YouTube comments as well as it looked through school posts.
15:59It looked through less school posts, though, which is interesting. Let me pull up this HTML. So we have cost and token pricing, error stuck mid build, getting money or sorry, getting clients, making money.
16:11It goes over the YouTube mood, the school mood, pain points once again, which don't seem to be the exact same. They're similar, but, you know, they're not the exact same.
16:20The number one product to build would be the AI consultant kit. So very similar. A package client delivery system that takes a member from I can build automations to a business paid me, not another how to build course, blah blah blah.
16:29Okay? So that's pretty similar. And then I worked as an AI consultant for a real business, real client, real numbers.
16:34Okay. So these are very, very similar results. So this would be a matter of which one do you trust more and maybe which one looked through more data, and that's how you could maybe trust it more.
16:43So let's look at the cost. Fable here spent $10.60 and took ten minutes, almost eleven minutes, whereas Opus spent $8.34 and took twenty minutes.
16:53So a little bit less efficient once again from Opus, but ultimately ended up being cheaper. Okay.
16:58Let's move on to the fifth experiment. I'm sorry if I'm going fast, but I also don't wanna bore you guys just, like, really, really diving into all these because there's a lot of things to go through, but I wanna show you a kind of a wide range of stuff. So this fifth one, your job is to create me a YouTube video outline and a slideshow, an an Excalidraw style presentation for this YouTube video.
17:16I want you to go through past LinkedIn posts, school posts, YouTube videos, and my AIs plus q and a. So there's a lot of things to dig through. And then I basically told it, you are a project manager.
17:24You're in charge of agents. You don't do anything. You just delegate work, and you review stuff.
17:28So that's what I wanted. Okay. So it created the presentation and the outline.
17:32Let's first look at the outline. So context engineering for agents. We have a cold open.
17:36We then move into what changed, why the terms exist, the failure modes of context, and we get into writing, selecting, compressing, isolating.
17:49Okay. So a pretty legit outline as you can see here.
17:53Let's open up the actual Excalidraw slide deck it made for us. Okay. So context engineering for AI agents.
17:59Your AI agent isn't dumb. Your context is. Bigger windows didn't help, so harder to see.
18:05I like that little touch. We've got our Excalidraw style boxes here. Andre Karpathy, Anthropic.
18:11We've got a quote right here. Attention is a budget. And, honestly, this doesn't look very branded the way my other Excalidraw presentations look.
18:20So I'm not sure exactly what happened here, but this doesn't feel exactly right. Poisoning, distraction, confusion, clash, and rot. We've got some other stats here.
18:28So not too bad. Right? I would obviously make some tweaks before I would get ready to start, you know, thinking about how I'm gonna present this, but not too bad, especially for one pass.
18:36Okay. Let's go ahead and see what Fable did here, what kind of topic it shows for us. So if I go to it it created outline, slides, and notes.
18:44So I'm gonna go to the outline, context engineering for AI agents. Wow. Very similar.
18:48Okay. So we've got the hook. We've got the section by section outline, what bad context is costing you, the four moves, right select, compress, isolate.
18:55Okay. So these are finding similar things, which is pretty interesting. I mean, it's looking through them, assuming similar data sources.
19:02So that's kind of good to know. Right? Like, the consistency makes me feel good.
19:06This looks more like what my YouTube video ones typically do look like, though. So that means maybe Fable did a better job navigating into my other project and finding the right skills because I forgot to mention, this directory was a completely fresh one. Both of these are working in completely fresh environments, so it's not inside of my HERC two as all of my normal things are running.
19:25This one had 29 slides, so quite a bit. We've got a big story here, which is something that happened to us. We have these different colors here.
19:34This one looks way more like what I typically am trying to build. We've got this nice visual with context rot. Gets lost in the middle.
19:40You know, we've got these nice visuals here. I would say this one is definitely a better presentation. So once again, Fable is kind of coming in on top when it comes to, like, the actual visual elements.
19:49I'm assuming they both did verification loops of screenshotting and, you know, validating. I like these a lot. These are nice slides.
19:57Yep. I like these slides better. So definitely, I think Fable takes the cake here.
20:00Let's look at cost. So Fable took about eight minutes and costed us $40. Wow.
20:08An hour and fifteen minutes and $33. So I don't know.
20:12I think that Fable wins here even though it was a little bit more expensive because it was still more token efficient and it was faster. Okay. Let's go to the next one, number six.
20:20So this one's interesting. This one's very interesting. I wanted to try to show you guys some computer use stuff.
20:26I compared it a little bit with codex computer use, and, ultimately, I still like codex computer use. I don't know why. They're very similar now, but I think because I just have this bias, you know, already for codex computer use.
20:38I don't know. Anyways, I don't use computers a ton, but I do use it a lot for verification stuff. So here, I told it to use Playwright CLI, and I told it this is something interesting.
20:46Right? I told it to go to Google and play the snake game. So I don't know if you guys you obviously know, like, what the snake game is.
20:52But if I come here, you can just play snake right here, which I used to do in class all the time. So I told these two models to go to Google and to do that.
21:01I wanted it to only play five games and screenshot the score of each game and give us the average. Right?
21:08Okay. You know what's really weird? I actually just noticed something.
21:11So I made a big whoopsies here. I I sent this off as fable, but I only but I actually used Opus.
21:17And then for the Opus run where I thought I was using Opus, I was using Opus. So I did Opus twice here, but I'm not even gonna change that because I wanna show you guys this. Same exact prompt to Opus.
21:28Right? Same exact prompt to the same exact model twice, but we got drastically different results.
21:34In this first version, the one that I thought was Fable, an hour and fifty three minutes, $18.
21:40But look what I had to do here. You were told to only run five games and give me the average score there. I'm not sure why you went rogue.
21:46This thing started running, like, 30 different games, and it just kept running and kept running games. It was trying to, like, maximize itself.
21:53Right here, it got an average of 63. So it was running batches of five games at a time and did that so many times. I was like, why why are you not following my instructions, but Opus is so much better?
22:03But I didn't realize that they were both Opus. So that's just a really good reminder that, like, at the end of the day, these things are completely nondeterministic. You don't know what they're gonna do.
22:12So anyways, this opus run, two hours $18. And then the real opus run that I thought was opus from the beginning, one hour ten minutes and $10.
22:24But this Opus run did much better, which is really weird. 81, 25, 33, 21, 44, which is, like, so, so much better than the other one did.
22:33So I'm not sure exactly what happened there, but, anyways, hopefully, that was kinda interesting to you guys. Okay.
22:39So let's take a look at this last one that I have to show off to you guys today. So this one was a bit longer.
22:45Right? You can read this if you want. But, basically, what I wanted was a simulator where we could see different, like, buildings and vehicles, and we could stress test them with different weather, and we could add weights, and we could even build and design our own structures and then test them.
22:58And this one's pretty interesting. Right? Because I gave the same problems obviously to Fable and to Opus.
23:03But let's take a look at the two differences first and then how long they ran. So I'm not gonna tell you which one's which yet.
23:10Let's just take a look. So here's the first one. As you can see, there's a lot going on.
23:14Like, it's a little bit overwhelming. Right? So if they wanted to build something that was kind of and not intimidating, they failed on that front, but that's not exactly what we asked for.
23:22We can see the different nodes. We can see the different beams. We can add, like, weights to them.
23:27We can see if they're passing or failing. I don't even know. Like, this is one of my first times opening, you know, this.
23:32I just opened them both to look at them, but I didn't use them. We can see I can add, like, snow. Right?
23:36So if I come here and if we want to add snow, we can see how things are changing. And if I add some rain, this is adding it to the whole thing.
23:44Gust factor, gravity, You can see it starts to change colors right there because there was way too much weight on here. And as you move this stuff you know, I'm not an engineer, so I'm not gonna come in here and tell you this is completely structurally accurate, and you could use this for, you know, making sure that your stuff isn't gonna fall and hurt people.
24:05This is a good simulation, right, because you can see where things are being put under pressure and where you need to increase some stability and stuff like that.
24:14I will say this is pretty overwhelming. Like, this UI, I don't really understand what to do. But if someone did understand how to get in here and how to test out this different stuff and, you know, build their own custom things, I could see this being very useful.
24:26And it just goes to show that in less than two hours, I already have this POC that I could go get feedback on and iterate on and stuff like that. We've even got these skyscrapers in here that we can start to add a bunch of different, you know, things to. So, anyways, this was the first version.
24:39We've got a bunch of different machines, and then you could also build your own. So let's take a look at the other version. This one's a bit more user friendly.
24:46Right? And this one honestly does look a little bit more AI made. This looks very like legacy software.
24:51This one clearly looks a little bit more AI, but, you know, the UI I wasn't too concerned with. But this one's a lot simpler. I can see the different things.
24:58I can easily add wind or snow or an earthquake, and I can see how much pressure is being put on these different points. You can see I can change to a skyscraper or a school or a vehicle, and we can do the same thing once again. And this one also makes it way simpler for me if I wanted to build and design my own thing.
25:13So if I wanted to add a few nodes here, I could add one there. I could add one there. One there.
25:17One there. One there. And then I can start to connect these.
25:20Right? So I can connect these here and here and here and there and here and there and, you know, put some triangles in here if we really wanna, you know, start getting fancy and building some nice support. So, anyways, this one just makes a bit more sense.
25:33And I didn't prompt it to say, hey. You know, like, this should be easy to use and people should understand it, the UI should not be intimidating or whatever. But, anyways, like, I can add weight to these different things, and I can start to really put some pressure on this stuff and, you know, see what it's gonna do.
25:47This is saying that this one is unstable. So, anyways, which one do you think was which? Because this one was actually Fable, and this one was actually Opus, which I wouldn't have expected, honestly.
25:58Because I think this one is honestly much more well designed when you think about what you're actually seeing in the data. So let's take a look at how much these cost us.
26:08So Fable, this only took Fable seven minutes, and it costed us $73.
26:13So it was just that that just goes to show how quick Fable can really run your session if you're not being careful. And then when we go to Opus here, this one took us two hours and twenty six minutes and $112. So clearly, Opus got stuck in a different loop or had, for some reason, Opus interpreted the verification criteria a little bit differently, and it ran way more sub agents, it did more stress testing on it.
26:36And maybe that just goes to show this line that we talked about that Claude Opus is much stronger at verifying its own work and iterating carefully until it succeeds. So maybe that just goes to show that in action right there. Because the thing is, I think that Claude Fable could have easily designed something and built something this level of detail and much better, like far far better.
26:55I've seen Fable do things that are way more incredible than this, and it could have, but it just decided for some reason based on the way I prompted it or whatever it was. It just decided to be done. It decided to be done here, which as you guys know, you've played with Fable, it's capable of so much more.
27:09But I think that that's a really good reminder of the fact that once again, these models are not deterministic, but also that Opus five interprets things different than Opus 4.8 and interprets things different than Fable five, which means the first thing that I did when I got OPUS five is I ran my skills.
27:23I ran my regular workflows. I ran my regular things that I do, you know, generating some YouTube stuff and helping me out because I wanted to see how it feels. That's why I always say I take these benchmarks with a grain of salt because it it matters way more about how you talk to it and how it feels and how you prompt for verification and all of that kind of stuff.
27:41Okay. So here are the consolidated results from those sessions that I just showed you guys. Keep in mind, I accidentally ran 10 with Opus and eight with Fable.
27:48It should have been 99, but either way, the numbers are still pretty telling. Look at this. Opus five spent more, which once again means that it's more inefficient, or I could have just said less efficient, with spending tokens because Opus five per token is half the cost of Fable five.
28:07So for it to be more expensive means that it's using way more tokens. We can see the combined output tokens right here, 2,000,000 for Opus and 832,000 for Fable.
28:16We can see the active time for Opus was six hundred and thirty minutes, so an average of an hour for all of these sessions, whereas an average of about twenty five minutes for the Fable sessions. Here's the total. Now we can see on the output speed, we have some stats here with Opus and with Fable, the cost per minutes.
28:30With Opus and Fable, the cost per thousand output tokens, Opus and Fable, API calls per session, and tool calls per session. And just a breakdown of where the money went, so it's a pretty similar split. Right?
28:41A cache reads a lot of it, and then we have cache write, and then we have output and input. And on both of these and on all of these runs, the input was pretty minimal because it was basically just, you know, creating output.
28:53And if you look at all of the sessions from most expensive to least expensive, it's honestly pretty split besides the fact that Opus here owns, like, the most expensive one. But, anyways, I will attach this exact document in my free school community if you guys really wanna check out this, you know, this data session by session.
29:10But what I really wanted to see was this headline up here. This headline is super interesting to me when it comes to the output tokens and when it comes to the time for actually running these models.
29:19So once again, guys, I hope that you enjoyed this comparison here where I showed you some actual things that I have done. Like I said, I rec really recommend that you just get in here with Opus five and you run your skills and just see what feels good and what doesn't. Have it do some verification loops and see if you like the way that Opus feels as an orchestrator.
29:35I still personally like the feeling of having Fable being an orchestrator, and I like to say something like, you shouldn't be writing any code or executing anything. You should just be telling Opus what to do, and that saves your session limit with Fable big time. Because you can be running Fable for multiple hours or even a day, and it won't even hit, like, 300 k or 400 k on the context because all it's doing is delegating.
29:56So try out that tip, but also, you know, start using Opus for that because I think a lot of us don't realize how powerful, like, an older sonnet model really is that you could use that to drive most of your daily knowledge work. And sometimes Opus five and Fable five are both overkill. So really think about what you're doing and matching the intelligence of the model with the intelligence needed for the task.
30:16But, anyways, I hope you guys enjoyed this video. I hope you found it insightful. And if you did, please give it a like.
30:20It helps me out a ton. And as always, I appreciate you guys making it to the of the video, and I'll see you on the next one. Thanks, everyone.
The Hook
The bait, then the rug-pull.
Opus 5 launched cheaper and, on paper, sharper than Fable 5 on several coding benchmarks — but benchmarks aren't the workflow. Nate Herk ran both models through nine matched real jobs, from bug hunts to a structural stress-test simulator, and receipts-checked every one on cost, time, and tokens.
CTA Breakdown
How they asked for the click.
VERBAL ASK
09:37link
“if you wanna access this entire free breakdown document as well as all the other resources that I ever give away on YouTube for free, just go to my free school community”
soft mid-video mention pointing to a free Skool community for the underlying cost/token spreadsheet, not a hard sales pitch
FROM THE DESCRIPTION
PRIMARY CTAWhere the creator wants you to go next.
Nate Herk breaks down the four ways an AI operating system's context quietly goes wrong, then walks through the five habits that keep a growing second brain accurate instead of confidently wrong.