Modern Creator
Alex Finn · YouTube

Opus 5 Beats Fable 5 on Every Benchmark — But Has Three Deal Breakers

Alex Finn runs both models through five self-designed benchmarks, crowns Opus 5 the winner on price and quality, then spends the back half explaining why he's not fully switching.

Posted
yesterday
Duration
Format
Review
hype
Views
14.6K
589 likes
Part of the collectionThe Fable 5 PlaybookAll 45 Fable 5 breakdowns, synthesized into one page.
Read the playbook
Part of the collectionThe Claude Opus 5 PlaybookEvery Opus 5 breakdown, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

Claude Opus 5 outperforms Claude Fable 5 on cost, speed, and most custom benchmarks, but its verbose personality, tighter usage limits, and weaker coding harness keep the creator running multiple models instead of switching to Opus 5 entirely.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You already use Claude models for coding or agentic work and are deciding whether to move from Fable 5 to Opus 5.
  • You want a specific rundown of what changed between two Claude model versions, not just a benchmark chart.
  • You're assembling a personal multi-model AI stack and want an opinionated take on when to reach for which model.
SKIP IF…
  • You want rigorous, reproducible benchmarks — these are the creator's own informal, self-designed comparisons, not academic evals.
  • You don't use Claude, ChatGPT, or agentic coding tools day to day.
TL;DR

The full version, fast.

Alex Finn puts newly released Claude Opus 5 through five self-designed benchmarks against Claude Fable 5: a 3D roller coaster build, an Apple website clone, an agentic document scavenger hunt, a bug-fixing duel, and a bridge stress test. Opus 5 wins on visual quality, agentic reliability, and cost — roughly half the price of Fable 5 and usable at 100% of plan budget, where Fable 5 capped him at half. Despite the win, he flags three deal breakers: Opus 5's personality is verbose and had to be tamed by editing CLAUDE.md, Claude's usage limits still feel tighter than ChatGPT's, and its coding harness lags Codex and the ChatGPT app. His resulting stack: Opus 5 for hard problems, Fable 5 for planning, ChatGPT 5.6 as daily driver.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0000:26

01 · Cold open

Claims Opus 5 beats Fable 5, is cheaper and faster, but flags three deal-breaking weaknesses to come.

00:2601:43

02 · Headline pitch

Lists the claims: beats Fable on almost every benchmark, half the price, full budget usable, significantly faster, questions whether there's any reason left to use Fable 5.

01:4302:06

03 · Finn Benchmark setup

Introduces his own five-test benchmark suite built to compare the two models.

02:0603:09

04 · Test 1: Build-Off (roller coaster sim)

Both models build a 3D roller coaster simulator; Opus 5's is more detailed and cheaper (82 cents vs. roughly 50% more for Fable 5).

03:0904:17

05 · Test 2: Pixel Perfect (Apple clone)

Both clone Apple.com from scratch; Opus 5's device recreations look closer to the real site than Fable 5's cut-off, distorted shapes.

04:1705:07

06 · Test 3: The Gauntlet (agentic scavenger hunt)

Agentic document-search test across PDFs and spreadsheets; Opus 5 finds 5 of 8 items, Fable 5 stops partway over a false cybersecurity flag.

05:0706:01

07 · Test 4 & 5: Debug Duel and Breaking Point

Debug Duel has both fix ~15 real open-source bugs (Opus 5 finishes ~25% cheaper); Breaking Point tests a bridge simulator under repeated car loads (Fable 5 holds slightly more weight but costs roughly double).

06:0106:27

08 · Score reveal

Final tally: Opus 5 wins on total cost, $6 versus $7.50 for Fable 5, declared the overall winner.

06:2709:15

09 · Three weaknesses

Personality sucks (verbose, unfocused, required a CLAUDE.md edit to tame), usage limits still feel tighter than ChatGPT's, and the Claude Code harness lags Codex and the ChatGPT app.

09:1510:52

10 · New stack & outro

Recommends a three-model stack by task type and plugs an upcoming live bootcamp on Opus 5.

Atomic Insights

Lines worth screenshotting.

  • Claude Opus 5 costs roughly half of Claude Fable 5 and lets users spend 100% of their plan budget on it, while Fable 5 capped usage at 50% of the budget.
  • In a five-test benchmark suite, Opus 5 beat Fable 5 on visual quality, agentic task completion, and total cost, losing only the raw bridge weight-capacity test.
  • Fable 5 refused to finish an agentic scavenger-hunt benchmark, flagging an ordinary document-search task as a cybersecurity risk and blocking the test outright.
  • A 3D roller coaster simulation benchmark cost 82 cents in Opus 5 credits versus roughly 50% more for a less detailed result from Fable 5.
  • Opus 5's personality is the first time a Claude model has felt like a disadvantage against ChatGPT rather than an advantage, by the creator's own account.
  • Editing CLAUDE.md to say 'speak as simply as possible' became necessary to stop Opus 5 from going in a dozen directions before finishing a simple bug fix.
  • Even with full budget usage unlocked, Claude's overall usage limits still feel tighter than ChatGPT's, which the creator says resets its limits every few minutes.
  • The creator's resulting three-model stack: Opus 5 for the hardest problems, Fable 5 for planning and back-and-forth brainstorming, ChatGPT 5.6 as the daily driver.
  • Claude Code's harness still lags behind Codex and the ChatGPT app's harness, according to the creator's hands-on comparison.
  • On an Apple-website-cloning benchmark, Opus 5 built recognizable device silhouettes from scratch, while Fable 5 produced a cut-off iPhone and unrecognizable MacBook and iPad shapes.
Takeaway

Split AI models by task, not by ranking

MODEL STACKING

Benchmarks that beat another model on price and quality don't automatically win your daily workflow — personality, usage limits, and the surrounding tool can outweigh raw intelligence.

04Test 1: Build-Off (roller coaster sim)
  • When comparing coding models, test with visually verifiable output like a 3D simulation rather than reading raw scores — you can literally see quality differences a chart won't show.
  • A more detailed, working simulation costing fewer credits than a plainer one from a pricier model is a reminder that cost and quality aren't always tied together.
05Test 2: Pixel Perfect (Apple clone)
  • Cloning a real website from scratch, not copying images, is a good stress test for whether a model understands layout and proportion rather than just generating plausible-looking text.
  • Small details — a cut-off image, a shape that doesn't match the real object — are the tell that a model's visual output is a generation behind, even when the overall page looks fine at a glance.
06Test 3: The Gauntlet (agentic scavenger hunt)
  • Agentic tasks that require digging through multiple documents for specific facts are a better real-world usefulness test than single-turn question and answer.
  • A model can fail an agentic test from an overzealous safety filter misreading an ordinary task as risky, not from lack of capability — worth checking before concluding a model 'can't' do something.
07Test 4 & 5: Debug Duel and Breaking Point
  • Fixing bugs pulled from real open-source repos is a more honest coding benchmark than synthetic problems, since real bugs come with messy, undocumented context.
  • A model that takes longer to finish a task isn't necessarily worse if it finishes for meaningfully less money — factor cost-per-outcome, not just speed or raw output, into model choice.
09Three weaknesses
  • A model's default personality and verbosity is now a real switching factor, not just a nice-to-have — an unfocused model costs you extra turns just steering it back on track.
  • You can override a model's default tone and behavior with a written instructions file like CLAUDE.md, explicitly telling it to speak as simply as possible rather than accepting the defaults.
  • Usage-limit anxiety changes how freely you're willing to prompt a model — a more generous limit isn't just a convenience, it changes what you're willing to try.
10New stack & outro
  • The surrounding tool that wraps a model (its harness) matters as much as the model's raw intelligence — the same underlying model can feel worse to use in a weaker harness.
  • Rather than picking one 'best' model, split tasks by type across several: the smartest model for hard problems, a more conversational model for planning, and the most available model as your daily driver.
Glossary

Terms worth knowing.

CLAUDE.md
A configuration file Claude Code reads at the start of a session to set custom instructions, such as tone or coding conventions, that persist across a project.
Agentic benchmark
A test that measures how well an AI model can independently use tools and complete multi-step tasks, rather than just answer a single question.
Harness
The surrounding application and tooling (like Claude Code or the ChatGPT desktop app) that wraps an AI model and governs how it reads files, runs commands, and manages context — distinct from the model's raw intelligence.
Resources

Things they pointed at.

Quotables

Lines you could clip.

06:04
Total cost was $6 for Opus, $7.50 for Fable five, and it's the winner.
clean, quotable stat that resolves the whole comparisonTikTok hook↗ Tweet quote
06:27
The personality sucks. It absolutely sucks. I've never been so annoyed talking to a Claude model.
blunt, contrarian complaint about a model's most-loved traitIG reel cold open↗ Tweet quote
06:40
You ever have, like, that friend from high school who thinks he's just, like, way better than everyone else and way smarter than everyone else?
vivid metaphor that carries the whole personality complaintnewsletter pull-quote↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphor
00:00Claude Opus five has dropped. It has happened. It has totally blown my mind.
00:06It is better than Claude Fable five. It is a fraction of the price. It is significantly faster, but it has three insane weaknesses I'm about to go over that change absolutely everything.
00:18But before we get into those three deal breakers with Claude Opus five, let's talk about the incredible things that are, like, legitimately revolutionary. First of all, it beats Fable five on almost every single benchmark. This is not one of those channels where we look at charts and benchmarks all day.
00:33Totally boring. No. I'm just letting you know right now, it beats Fable five on every benchmark.
00:37In a second, I'm going to take you through a world famous new revolutionary benchmark I've been working on for weeks now that proves it's better than Fable five in almost every single way. It is half the price. So this was the biggest issue with Fable five.
00:53It was totally uneconomical for the average user. But Opus five, it's half the price.
00:59It comes in at a good price, and you get your full limitations with your clawed plan. Meaning, you can use it to a 100%. Claude Fable five, for some reason, only let you use half your limitations, which is really stupid and annoying.
01:09It's significantly faster. It works lightning quick, much better than Fable, which brings us to the question, and I'll and I'll show you my benchmarks in one second here. Any reason to use Fable five anymore?
01:21No. With one small exception, which I'll go through right after these benchmarks I'm about to show you. But, no, there isn't a reason.
01:29And then as I said, there's three major weaknesses, which we'll go over as well. Let's go into these world famous Finn benchmarks I just created to show you why this is better than Fable five.
01:40Alright. So this is the new famous Finn benchmark. There are five tests here.
01:44We put both Fable five and OPUS five through. In a second, I'm gonna show you OPUS five versus GPT five six so we can know which one of those two is better. In OPUS five one, basically, all of the benchmarks.
01:58So starting with benchmark number one, this was a three d roller coaster simulator. Both had to build a three d roller coaster simulator. Let's show you Opus five first.
02:08This was Opus five, an absolutely beautiful simulation. It built the entire track, all the trees, the building route, the sky, the clouds. And if I hit ride, you can actually see the roller coaster in real time.
02:21You can see the from the user's perspective going around the roller coaster. This is really, really nice, detailed, beautiful, and was done lightning quick. As you can also see, it was done with 82¢ of credits and about 32,000 tokens.
02:37We go over to Fable five. It was actually done with less tokens, but was significantly more expensive, about 50% more expensive.
02:45Let's see the output of it. And as you can see, it's just not quite as beautiful, not quite as detailed or clear. It looks like just kinda like a generation behind.
02:54If we go on the ride, you can see the track doesn't look quite right. The trees all kinda look similar. There was no real detail to anything, so it doesn't look nearly as good.
03:03The next test is the pixel perfect test where it has to clone an entire website. It has to clone the Apple website.
03:11If we go to the actual Apple website, this is what it was tasked with cloning. You can see College Sorted has a bunch of people holding Apple devices, iPhones, MacBook Airs, MacBook Pros.
03:22Elcho Opus five's recreation, obviously, it doesn't look incredible, but I'll show you Fable in a second. But you can see it's pretty close.
03:30It had to build every part of this website from scratch. It's not copying and pasting over images. It's building what a MacBook looks like from scratch, what a MacBook Pro, an iPad, a watch would look like.
03:42As you can see, it has names, recreates a TV. Let's see what this looks like with Fable.
03:47Fable five not quite the same. Remember, this looked like human beings on the real website and in the Opus five website. The iPhone's cut off.
03:56The MacBook looks nothing like a MacBook Air. The iPad looks nothing like iPads. The watch doesn't look anything like a watch.
04:03The card's kinda lazily done. As you can see, nothing looks quite the same. Opus five, everything looked much better.
04:10It looked perfect. No. But it looks much better than Fable.
04:13The next test was the gauntlet test. And, basically, this is a agentic test where it tests the tool use of the model.
04:22We give the model a whole bunch of documents, PDFs, Excel spreadsheets, a whole bunch of things, and have it basically do a scavenger hunt where it goes and has to find specific things in all the documents. It tests its agentic ability. Opus five, it did a pretty good job.
04:39It got five out of the eight scavenger hunt items. Fable five cut off halfway through because of content blockage from Anthropic.
04:47It thought I was doing something with cybersecurity. It cut it off. This has nothing to do with cybersecurity.
04:52It's about finding specific things in different documents. Fable five wouldn't allow me to do the agentic test, so that didn't count. There's a debug wall.
05:00So, basically, what this benchmark does is go online and find, like, 15 different bugs from open source GitHub repos, and then it hands it to each model and says, hey. Go through this and fix all the bugs and all these open source repos.
05:14Opus actually took a bit longer than Fable five. Did it at about the same amount of tokens, but did it at about a dollar cheaper overall, so about 25% cheaper.
05:26So this goes to OPUS five as well. And then the last test is breaking point. And, basically, the way this works is each model is tasked with building a bridge.
05:36It's basically a bridge simulator. They're tasked with building a bridge, and then the benchmark drives a car over the bridge over and over and over again to see how much weight the bridge can hold. It's basically testing the thinkability.
05:50Okay. Can you design a bridge that holds tons and tons of weight? Fable caused basically double Opus to do this, but only was able to hold slightly more weight than Opus.
05:59So it goes to Fable, but it was a lot more expensive. Overall, Opus five beat Fable, beat him in almost every single benchmark.
06:07Did it for significantly cheaper. Total cost was $6 for Opus, $7.5 for Fable five, and it's the winner.
06:15Opus beats Fable for a fraction fraction of the price. So let's talk about the weaknesses of the model. There are a few deal breakers for me here that are stopping me from using this in my entire stack.
06:27Number one, the personality sucks. It absolutely sucks. I've never been so annoyed talking to a Claude model.
06:33This has actually been the advantage of Claude models up to this point. I've always loved talking to Claude models. It's been their biggest advantage against Chad GPT, But for the first time, it has flipped.
06:44I loathe the personality of Opus five. It is way too verbose. It is not nearly concise enough.
06:51It goes in a 100 different directions when it's talking to you. You ever have, like, that friend from high school who thinks he's just, like, way better than everyone else and way smarter than everyone else? And when you talk to them, they use, like, the biggest words possible and go in a million different directions to prove how smart they are.
07:07That's what it feels like with Opus five. I've had to multiple times using Opus five hit the stop button to get it to shut up, and I say, please be way more simple and concise and talk to me like I'm five years old. I'm not kidding.
07:19For the first time, I've had to go to claude.md to edit its personality. I just said, speak as simple as humanly possible.
07:26And I highly recommend when you use this model, you do the same thing. It's unfortunate because it also kinda leaks into the way it works sometimes where I'll be like, fix this bug, and it'll just do a 100 other things before fixing the bug, which is really, really annoying.
07:40It it appears like it's just this, like, erratic super hyperintelligent being that can't stay focused. For me, Fable five was actually way more focused, and I actually enjoy talking to fable five more.
07:54The issue is fable five, you can only use 50% of your budget on it, and it uses up all your credits. So I have to replace fable five with Opus. So highly recommend editing your personality for opus five.
08:06It's just it's just too much, and it does too much. The limits suck even though you can use a 100% of your budget on opus five. The clawed limits just absolutely suck compared to ChadGBT.
08:17ChadGBT, that that T Bo dude from Twitter is constantly restarting the limits, like, every five minutes.
08:23You basically get unlimited usage with ChadGBT. Claude, even though you can use all your budget on Opus five, it still has lower budgets overall, Anthropic versus ChadGBT.
08:36I can still see the meter going quicker, which gives me, like, a level of anxiety as I'm giving prompts. It makes me wanna do less.
08:43Because ChadGBT has unlimited usage, basically, there's no anxiety when using it.
08:48I I'm more free to be creative and explore and do more things and do interesting things. So I you know, the limits still here are a deal breaker for me. And then the harness for Claude code is still just not as good as Codex or, I guess, it's a ChatGPT app now.
09:04The ChatGPT app is a significantly better harness. Their new voice mode, which video on that coming in, like, the next twenty four hours, maybe forty eight hours, turn on notifications now and subscribe, especially if this video has been helpful for you.
09:17Video on that coming very soon, but it is incredible. It is excellent. Claude just added a voice mode, but it's not nearly even, like, a quarter of what the ChadGBT voice mode is.
09:27That's coming soon. Again, notifications on. Also, by the way, I'm doing a boot camp on Opus five in an hour from me filming this.
09:35It's gonna be recorded. It'll be in the vibe coding academy. Link for that down below.
09:39Number one AI community on planet Earth. Join that link down below. I promise it'll be the best decision you ever make.
09:45Now here is my new stack. With all that being said, here is my new stack. For super hard problems, Opus five, it's the smartest model out there.
09:52It has the highest intelligence. It's smarter than Fable five. Fable five was slightly smarter than five six.
09:58Opus five is slightly smarter than Fable five. If I'm doing massive planning, I'm still relying on Fable five mostly because I just don't like the output of Opus from, like, a talking perspective. So if I need talking, if I need to plan, if I need to go back and forth, if I need a brainstorm, I'd rather do a Fable five.
10:15I don't wanna talk to Opus five to do plenty. I just want to shut up and write code. Daily driver, though, that's ChadGPT56.
10:21You get so much higher limits. The voice mode is incredible. Again, video coming soon.
10:26It is just better to use overall out of the three. So daily driver, ChadGBT56 it is.
10:33Have you used Opus five? How does it compare to Fable five for you? Let me know down below in the comment section.
10:39Hope this was helpful. Way more videos coming out on Opus five, Claude Co, GBT voice mode, Hermes agent, Opus five, and Hermes agent all coming very soon. Make sure to subscribe.
10:49Leave a like if you learned anything at all. I'll see you in the next video.
The Hook

The bait, then the rug-pull.

Alex Finn opens by declaring Claude Opus 5 better than Fable 5 in almost every way — cheaper, faster, and the winner on his own five-test benchmark suite — before spending the video's back half walking through three things that stop him from fully switching.

Frameworks

Named ideas worth stealing.

01:43list

The Finn Benchmark (5-test AI coding suite)

  1. Build-Off (3D roller coaster sim)
  2. Pixel Perfect (website clone)
  3. The Gauntlet (agentic document scavenger hunt)
  4. Debug Duel (open-source bug fixing)
  5. Breaking Point (bridge weight-capacity sim)

A self-designed five-test suite the creator uses to compare coding models on visual output quality, agentic reliability, and cost.

Steal forAny side-by-side AI model comparison that needs varied, visually verifiable tasks instead of just text benchmarks.
10:18list

Three-model stack

  1. Super hard problems: Opus 5
  2. Massive planning: Fable 5
  3. Daily driver: ChatGPT 5.6

The creator's resulting workflow after testing, splitting tasks across three different AI models by strength rather than picking one.

Steal forAnyone deciding how to allocate different AI models across different task types instead of committing to a single model.
CTA Breakdown

How they asked for the click.

VERBAL ASK
09:15product
I'm doing a boot camp on Opus five in an hour from me filming this. It's gonna be recorded. It'll be in the vibe coding academy. Link for that down below.

Verbal plug placed right after the weakness deep-dive establishes credibility, paired with a description link and 'link down below' callout.

MENTIONED ON CAMERA
Storyboard

Visual structure at a glance.

open
hookopen00:00
benchmark scoreboard
promisebenchmark scoreboard01:43
roller coaster build-off
valueroller coaster build-off02:06
apple clone, pixel perfect
valueapple clone, pixel perfect04:09
the gauntlet
valuethe gauntlet05:07
weaknesses title card
valueweaknesses title card09:15
new stack CTA
ctanew stack CTA10:10
Frame Gallery

Visual moments.

Chat about this