Modern Creator
Dubibubi · YouTube

I Made Claude Opus 5 and Fable 5 Build the Same App

A blind, three-build test of Anthropic's newest coding models — and the first time this reviewer picked a model other than Fable as the winner.

Posted
yesterday
Duration
Format
Review
hype
Views
8.5K
273 likes
Part of the collectionThe Fable 5 PlaybookAll 45 Fable 5 breakdowns, synthesized into one page.
Read the playbook
Part of the collectionThe Claude Opus 5 PlaybookEvery Opus 5 breakdown, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

Claude Opus 5, launched hours before this video, beat Claude Fable 5 head-to-head on three identical one-shot app builds at roughly half the cost on the first test, overturning the creator's expectation that Fable would win on output quality.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You're deciding between Claude's Opus and Fable model tiers for a coding subscription and want real build-quality evidence, not just benchmark charts.
  • You build with AI coding agents and want to see how two frontier models perform on identical, one-shot, real-world app prompts.
  • You're curious what a trading terminal, kart racer, or office-guessing game built entirely by an AI agent in one shot actually looks and plays like.
SKIP IF…
  • You want a rigorous, controlled benchmark — this is one YouTuber's single-run, three-prompt test, not a statistically significant study.
  • You're looking for prompt-engineering technique — the video shows prompts being copy-pasted, not how they were written.
TL;DR

The full version, fast.

Hours after Claude Opus 5 launched, a creator ran it blind against Claude Fable 5 on three identical one-shot prompts: a professional trading terminal, a 3D kart racer, and an office-guessing game. A third-party AI opened both builds without revealing which model made which, so the review stayed unbiased until the reveal. Opus 5 won two of three blind rounds on output quality, matched or beat Fable 5's build time, and came in at roughly half Fable's cost on the first test — though cost parity varied on the other two. Total spend across all three builds ran about $150 in API-equivalent cost over 122 million tokens. The video also covers Anthropic cutting 80% of Claude Code's system prompt and Opus 5 being its least prompt-injectable model yet.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0000:10

01 · Cold open — the benchmark hook

Opus 5 just launched; the creator claims it beats Fable on nearly every metric at half the price.

00:1001:53

02 · Test rules & first prompt

Same prompt, max effort, single one-shot output, no revisions, deployed live. First build: a professional trading terminal.

01:5302:59

03 · Judges & blind-test setup

Cardboard-cutout judges Boris and Dario introduced as mascots; creator addresses accusations of Fable bias and commits to a blind test via a third-party AI.

02:5904:45

04 · Benchmark deep dive: CursorBench & ARC-AGI-3

Review of Opus 5's published benchmark scores, including a 0.5% CursorBench edge and a jump from under 5% to over 30% on the ARC-AGI-3 novel-problem-solving test.

04:4507:51

05 · Trading terminal reveal & build stats

Both trading terminals reviewed blind via a third-party AI browser session; favicon tell narrows the guess; reveal shows Opus 5 used fewer tokens/tool calls and finished faster at about half Fable's cost.

07:5108:24

06 · Second prompt: 3D kart racing game

Fresh prompt sent to both models: a complete, playable 3D kart racer in a single HTML file using Three.js.

08:2410:07

07 · Tangent: Claude Code's system prompt overhaul

Discussion of Anthropic engineer Thariq's claim that 80% of Claude Code's system prompt was removed for newer models, a leaked GPT-5.6 system prompt for comparison, and Boris Cherny's tweet on Opus 5's prompt-injection resistance.

10:0712:43

08 · Kart racer blind test & gameplay

Both racing games played live; creator guesses which model built which based on drift mechanics, difficulty settings, and visual detail — and gets it backwards.

12:4313:42

09 · Kart racer reveal & build stats

Reveal shows Opus 5 spent far more tokens, time, and tool calls than Fable but produced the build the creator judged more fun and detailed — his first-ever pick against Fable.

13:4214:18

10 · Third prompt: Office GeoGuessr + snack break

Final prompt sent: a first-person 'guess the tech company office' game in a single HTML file.

14:1818:20

11 · Office GeoGuessr blind test (viewer participation)

Creator and viewers guess which tech company's office each build depicts (Microsoft, Amazon, Apple, Google) across two separate builds, without knowing which model made which.

18:2019:19

12 · Final reveal, total cost, and verdict

Opus 5 identified as the more detailed build (double the lines of code); total spend across all three tests lands near $150 and 122 million tokens; creator declares Opus 5 the overall winner.

Atomic Insights

Lines worth screenshotting.

  • On the ARC-AGI-3 novel-problem-solving benchmark, Claude Opus 5 scored over 30%, up from under 5% for Opus 4.8 just two months earlier.
  • Building a professional trading terminal from one prompt, Opus 5 used about 7.9 million tokens at a 98% cache rate and finished in 31 minutes versus Fable 5's 43 minutes.
  • On the trading terminal test, Fable 5 made 56 tool calls and wrote 2,000 lines of code; Opus 5 made 53 tool calls and wrote 1,900, at roughly half Fable's estimated cost.
  • Opus 5 took about 401 seconds to start acting on the trading terminal prompt; Fable 5 took about 668 seconds.
  • For the 3D kart racing game, Opus 5 used 31 million tokens and 130 tool calls over 1.5 hours; Fable 5 used only 10 million tokens and 70 tool calls in 1 hour, at nearly the same dollar cost.
  • Despite spending three times the tokens and nearly double the tool calls on the racing game, Opus 5 produced the build the creator judged more fun and detailed, including a working drift-boost mechanic Fable's build lacked.
  • On the GeoGuessr clone, Opus 5 wrote 3,700 lines of code — roughly double Fable 5's 1,500 — while costing about the same and taking around 30 minutes longer.
  • Across all three one-shot builds, total spend came to about $150 in API-equivalent cost, using 122 million tokens over 7 hours 40 minutes of compute.
  • Anthropic engineer Thariq said the team removed more than 80% of Claude Code's system prompt for the Opus 5 / Fable 5 generation with no measurable drop in coding performance.
  • Boris Cherny, creator of Claude Code, called Opus 5 the least prompt-injectable Claude model yet, saying layered defenses can push prompt-injection attack success rates to near zero.
  • The creator says Fable has won every one of his past head-to-head model comparisons — this was the first time he judged a rival model's output better than Fable's.
Takeaway

Cheaper tokens don't guarantee the weaker build

MODEL COMPARISON

A blind, three-build test found cost, speed, and output quality moving independently across two frontier coding models, so neither price nor speed reliably predicted which build turned out better.

04Benchmark deep dive: CursorBench & ARC-AGI-3
  • A benchmark score means little without its cost — ARC-AGI-3 plots model accuracy against total evaluation dollars, so a cheaper model that scores lower can still be the better deal per point.
  • Jumping from under 5% to over 30% on a novel-problem-solving benchmark in two months signals a model handling situations it has never seen a matching pattern for, not just memorizing more examples.
05Trading terminal reveal & build stats
  • The build that used fewer tokens, made fewer tool calls, and finished twelve minutes faster was still visually indistinguishable from its pricier rival to a blind reviewer.
  • A 98% cache rate on repeated context is what keeps a long AI coding session affordable — reused tokens bill at a fraction of fresh ones, and it can swing total cost more than the model choice itself.
07Tangent: Claude Code's system prompt overhaul
  • AI coding tools are shrinking their own instruction manuals rather than growing them — Claude Code cut over 80% of its system prompt for newer models with no measured drop in coding performance.
  • A model's resistance to prompt injection, meaning hidden instructions buried in a webpage or file that try to hijack it, is becoming as visible a selling point as raw coding benchmarks.
09Kart racer reveal & build stats
  • Spending three times the tokens and roughly double the tool calls didn't just cost more — it bought extra polish, like a working drift-boost mechanic and multiple difficulty tiers, that the cheaper build skipped.
  • A reviewer's live guess about which model built which output can be wrong in both directions — the build he found 'less detailed' turned out to be the one he ultimately preferred.
11Office GeoGuessr blind test (viewer participation)
  • Doubling the lines of code on an identical prompt tends to show up as filled-in detail — office-specific furniture, lighting, and props — rather than as extra features.
  • Recurring defaults, like one model consistently adding a favicon and the other not, become reliable fingerprints for guessing which model made an output even before testing what it does.
12Final reveal, total cost, and verdict
  • Total spend across three full one-shot app builds landed around $150 in API-equivalent cost — a concrete anchor for what testing two frontier coding models properly actually runs.
  • A reviewer who states his own bias upfront, then reports being surprised by the blind result, is a stronger signal of an honest test than one that simply confirms what he expected going in.
Glossary

Terms worth knowing.

ARC-AGI-3
A benchmark that drops an AI into unfamiliar, instruction-free game-like environments and tests whether it can explore, form a plan, and learn from mistakes to solve completely novel problems, rather than answering questions it can pattern-match from training.
System prompt
The hidden instruction set an AI coding tool feeds a model before the user types anything, telling it which tools to use, how to behave, and what not to do.
Prompt injection
An attack where hidden instructions embedded in a webpage, email, or file try to hijack an AI agent into doing something the user never asked for.
One-shot build
Generating a complete, working app from a single prompt with no follow-up corrections, edits, or revisions.
Cache rate
The percentage of an AI model's input tokens reused from earlier in the same session rather than reprocessed from scratch, which bills at a steep discount.
Tool calls
Individual actions an AI coding agent takes — like running a command or editing a file — while completing a task, often used as a rough proxy for how much work happened during a build.
Resources

Things they pointed at.

03:11linkCursorBench
03:42linkARC-AGI-3 benchmark
08:40linkThariq (@trq212) tweet on Claude Code's system prompt reduction
09:07linkPliny the Liberator's leaked GPT-5.6 system prompt
09:45linkBoris Cherny (@bcherny) tweet on Opus 5's prompt-injection resistance
Quotables

Lines you could clip.

00:05
It's beating Fable in nearly every metric, but it's half the price.
cold-open hook statTikTok hook↗ Tweet quote
07:51
Guys, if you have to choose between Fable and Opus and the difference is this, then I'm going Opus all the way.
clear, quotable stance after the first revealIG reel cold open↗ Tweet quote
13:15
This is the first time I'm actually seeing a model produce a better output than Fable's.
reversal moment for a reviewer with a stated Fable biasnewsletter pull-quote↗ Tweet quote
18:49
The winner is going to be Claude Opus five.
climactic verdict lineTikTok hook↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphor
Chord Opus five just launched a few hours ago and the benchmarks are insane. It's beating Fable in nearly every metric, but it's half the price.
So today, we're putting that to the test by giving the exact same prompt to both Opus five and Fable five across three massive builds, a professional grade trading terminal, a fully playable three d kart racing game, and finally, a GeoGuesser clone but for big tech company officers. Same rules for both, a single one shot output, no revisions deployed live on the Internet.
This is Claude Opus five versus Claude Fable five. Let's get straight into it. Up here on the left hand side, we have Opus five on max effort, and then the right hand side, we have Claude Fable five on max effort as well.
So first off, we're building a professional trading terminal, and this one build is actually testing a lot at once. We're testing UI density, real time data engineering, custom chart rendering, and market simulation, because it's gonna be really interesting to see how they can turn fake data and actually turn it into a live market.
I want this app to feel real, because this can't just be a dashboard with some charts on it. The models have to build the from scratch.
We're talking live candlesticks, a streaming order book, real positions with profit and loss updating on every ticker. And every piece has to move together because in a real trading software, one frozen chart, it doesn't just look bad, it's gonna cost you money.
So this could actually be useful if you're looking to trade with your own portfolio using AI or building fintech products or maybe you're just selling dashboards to paying clients. I'm super excited to see what these two models cook up. Let's have a look at the prompt.
Create a professional grade trading terminal called Apex Terminal. Alright. Cool.
I'm just gonna go ahead and copy this. Oh, and if you want access to these prompts, I've left a link in the description so you can test them out yourself. And they are off.
Let's go. By the way, I also made some improvements to the usage tracking dashboard. So we're gonna be seeing the input tokens, the output tokens, how long each build takes, how much it's gonna cost, and I've actually added some additional features here that will tell us a lot more about what's happening in the background.
I'll tell you more about this later. Also, me know if you want me to open source this dashboard so you can actually run your own tests. Leave a comment below, and I'll consider it.
Now for the marking rubric, I've got Boris over here, the creator of Claude Code as a boxer. He's representing Claude Opus five. And on the right hand side, we have Dario, the CEO of Anthropic, representing Fable five.
Now I do need to address the elephant in the room. Do you guys seriously seem to think that I have a bias towards Fable for some reason? Which I don't.
Bruh. To make this super fair, I'm going to be doing a blind testing, so I'm literally gonna have no idea which model built which website. So I'm gonna let these two models cook, and I'll see you in a bit.
Now while they're still running, let's take a closer look at these benchmarks. It's basically beating Fable in everything, and they did actually make a mistake here because they highlighted the wrong one. But, yeah, pretty much everything.
If we look on CursorBench, the gap is actually even more noticeable. A 0.5% difference in score, but for half the price.
It really felt like Anthropic had to come with something big here because OPUS 4.8 feels so outdated now, and g p d 5.6 Soul being as good as it is forced them to leave Fable in their subscription plan. Now we haven't looked at the tests yet, but I would say OPUS five is a pretty big leap and a very competitive model. And I say this because there is one bench in particular that people are going crazy about.
The Arc AGI three bench. A bench designed to test whether an AI can figure out completely new problems on the fly.
Instead of answering normal questions, the model gets dropped into unfamiliar game like environments with no instructions, and it has experiment, work out the rules, learn from its mistakes, and then solve them.
Humans have been able to get a 100% in every environment, but AI models are still terrible at this. OPUS 4.8, for example, was scoring under 5% just two months ago, and OPUS five is now scoring over 30%. So in the space of just two months, Anthropic has gone from barely solving any of the benchmark to completing almost a third of it.
That is a ridiculous jump, and it makes models like g p d 5.6 feel like toddlers. Alright. The results are in.
I am super excited to see what these two models cooked up. Now I can see they've both finished, but I don't wanna look at the terminal too much. Otherwise, I'm gonna spoil myself.
I brought in a third party AI model, and I'm just gonna send her this prompt. Please open both trading terminals on my browser for review. Do not tell me which models did which.
Okay. It's opened up both of them over here. Let's have a look.
Damn. This looks pretty freaking good, man.
On the right hand side, we have, like, thousands of trades happening every single second. Wow. This looks extremely realistic.
Down here, we've got, a buy and sell button. So let's go ahead and buy.
Oh, we got a alright. Alright. Selling selling selling selling.
Oh, alright. Sell sell sell.
We just sold our entire position. We are currently up $600. Let's go.
I can see this being extremely useful if you wanna paper trade and you wanna get good at trading. You can literally just create a exact replica of the market.
Now up here, you can already see one of them has a favicon and one of them doesn't. So if I was to guess, I know that fable does tend to add favicons, so I actually do think this is fable.
Let's have a look at this second option. Okay. They literally look identical.
Like, look. I literally can't even tell the difference. I'm trying to find anything that is different.
We kinda have, like, this heartbeat rate over here. Yeah. Cool.
Sell my position. Are we are we at a profit here? What do we do?
We're at 1,500 profit here. Damn. Everything I buy just going up, mate.
Yeah. This second build, I think I am liking, like, the tiniest little bit more. Everything just feels a little bit less confusing.
Let's have a look at the usage chart. Okay. There you go.
We I have this little thumbnail here, and that looks like that second build.
Yeah. That's the second build. Okay.
Yes. I was right. I was right.
See? That's how good I know Fable, man. So Fable did the first build, this one over here, and Opus did the second build, this one over here.
So let's have a look at the stats. We've got 7,900,000 tokens used by Opus with a 98% cash rate.
Okay. Interesting. Thirty one minutes for Opus and forty three minutes for Fable.
Estimated costs sound exactly like the benchmarks, literally half the cost of Fable. Now we've got a few extra metrics here. Uh, I've done tool calls.
So Opus called 53, while Fable called 56. And then this one also shows lines of code, LOC.
So Opus did 1,900 lines of code. Fable did 2,000.
And then this last one here is first action. How long did it take for the model to do the very first action? From sending the prompt to actually starting the build.
And you can see Opus, it took four hundred and one seconds. Fable took six hundred sixty eight seconds.
We all know how long Fable takes to get things done. So I'm not gonna lie, guys. It's very close.
It's very close. But if you think about it, it shouldn't be this close because we're paying double. We're literally paying double the price.
Guys, if you have to choose between Fable and Opus and the difference is this, then I'm going Opus all the way. I I don't wanna pay an extra $10.20 dollars.
Look how happy Boris looks, man. He's cheering. He's so happy that he won that.
Alright. Let's have a look at the second prompt, guys. Next up, we got a three d kart racer.
Alright? Create a complete fully playable three d kart racing game in single HTML using 3. J s.
Alright. Let's go ahead and copy this prompt, send it into these fresh terminals.
Opus on the left again, sable on the right, and I'll let these two cook, and I'll see you back in a bit.
Now while those models are building out a three d racing game, I think it's only fair I play some Mario Kart for research purposes. Okay.
So there's something Anthropic mentioned that I think is actually really interesting. The way they build ClaudeCode around these newer models has completely changed. And just to be clear, this isn't the model's training, it's the hidden instructions ChordCode gives it before you even type anything.
Tariq, one of the lead engineers behind ChordCode, said they removed more than 80% of its system prompt for models like Opus five and Fable five with no measurable drop in coding performance. A system prompt is basically a giant instruction manual telling the model how to behave, which tools to use, and what not to do.
Like GP five point six's system prompt is over 42,000 words. Now Anthropic actually published a 73 page document as to how they did this, but the main thing is this.
Older models needed loads of strict rules because without them, they would make dumb decisions. But these newer models are apparently good enough that all those instructions were getting in the way. So now they give Claude fewer rules and let it use its own judgment, which kind of fits in line as to why it got such a high score for the AGI ARC three bench.
And apparently, it's also much harder to trick. Boris Churney, the creator of Claude Code says OPUS five is their least prompt injectable model yet. That basically means if Claude reads a website email or file containing hidden instructions designed to hijack what it's doing, it is far less likely to follow them.
So this upgrade isn't just smarter, it needs less hand holding and appears much harder to manipulate. Let's go. Both models have just finished the build for the three d racing game, so let's have a look.
Oh, alright. Damn.
Oh, okay. And this time they both seem to have a favicon as well, so harder to tell whose is Fables and whose is Opus. Uh, we'll go easy mode for now.
Let's go.
Woah. That was fast. Oh my god.
How do I turn this music off? That's so much better. That was way too loud.
Yo. The drift mechanics are nice, man.
Damn. Dude, I have a feeling this might be fable. This is too good not to be fable.
I like that they give you a boost when you're in the air. They also seem to give you a boost when you are and we won.
Boo. First, absolutely smashed second, third, and fourth. I think Fable did this one.
I actually think Fable did this. Let's let's have a look at the next one. Turbo circuit.
There's no difficulty mode here. This feels like an Opus thing. This feels like an Opus thing.
Okay. Because I don't know if I like this one anywhere near as much. It feels way less detailed.
Dude, what the hell was that? I just went through a mountain, bro. When I drift, don't get a boost.
This has to be Opus five. See, how are you meant to anticipate that turn when you can't even see the freaking road? Alright.
I'm actually kinda nervous. I'm it's time to find out which model did what.
Let's look at the benchmarks. Okay. And the results are in.
Damn. Yo. These this is really interesting.
Opus five. So I was way off. Opus five made the funner version.
This is Fable. Opus made this one. Oh my god.
Guys, look at how long it took. So Fable only took an hour and it cost me $25.40. Opus took an hour thirty minutes and it cost nearly the same amount.
Look at how many tokens it used. 31,000,000 tokens. Fable only used 10,000,000.
Opus has done an extra 500 lines of code compared to Fable. It took a 130 actions.
A 130 tool calls compared to Fable's 70. So this is really interesting. Opus not only took longer, it also spent way more tool calls, wrote more lines of code, did three times the amount of token usage, but I'm not gonna lie.
It came out with a better game. We have different levels. We have a boost when we're landing from the sky.
We also have a drift mechanism that works. Oh my god. It's blowing my mind.
If you guys have watched my other testing videos, you'll find that Fable always wins when it comes to the output. Fable always produces a better output. This is the first time I'm actually seeing a model produce a better output than Fables.
Damn. I was not expecting that. This has actually shocked me.
This this is a this has shocked me. I'm gonna have to give that one to Opus again, guys. Oh my god.
Dude, Anthropic has cooked with Opus five. Alright. Now this last test is going to be working a little bit differently because not only will I be doing a blind test, you guys will be doing a blind test as well.
Okay? Here's the prompt. Create Office GeoGuesser, a fully playable first person guessing game in a single HTML file.
I'm gonna go ahead and send these two prompts off, and I'll see you when it's over.
These nuts? Got it. Just crossed over an hour on the prompt.
I'm expecting a good result. Alright. Both models have just finished up.
I have just asked a third party Claude agent to open both builds on the Internet. Remember, you are going to be doing this blind with me this time. So let's see if you can get it right.
First up, we have a pretty clean UI here, Office guesser. You've been dropped inside the office of a famous tech company. No logos anywhere.
Read the furniture, the colors, the culture, then guess who works here. Alright.
Let's go ahead and start. Whose office is this? I wonder whose office this is gonna be.
This is obviously gonna be Microsoft, guys. Come on. That is such a dead giveaway, but this looks really nice.
I'm really impressed. Every single monitor has something going on.
Like, that's kind of cool.
This guy's playing cards. These feel a little bit like Windows.
Yeah. Alright. I'm gonna go ahead and lock in.
Wait. Wait. What?
I had to oh, there was a time limit to guess inside of okay. Okay. Okay.
Let's try again. Let's try again. Clues.
Who's this is Amazon. Yo. We got a roper vacuum.
That's freaking sick. Damn. That's so cool.
I like that the the desk colors changed to wooden now. Work hard, have fun, make history. Is that even their slogan?
I'm going to guess. They got a dog over here. Rufus.
Alright. This is definitely Amazon. Man, that was cool.
This is cool, guys. This is actually really cool. I feel very immersed in this game.
Alright. Next round. Oh, this is Apple.
Without a doubt, this is Apple. And that should be all of them. Hey.
I love that it shows you the clues too. This is really cool. I was actually not expecting this.
Now I can see a favicon here, so that is a little bit of a telltale sign that this is probably Fable. I don't know. I'm just very confident that Fable never forgets to add a favicon, just from my personal experience.
Let's go ahead and check out the next one. Office Guesser. Oh, this looks cool.
This looks this looks cool. Alright. Start shift.
Damn. No. What?
Dude, I can go into different people's offices. Dude, the screens, some of them are on.
What's this? We got user. I gotta guess.
I gotta guess. This I'm thinking this is this is either Apple or Windows.
What do you guys think? What are we what are we locking in here? Alright.
I'm gonna guess Google. I'm gonna guess Google. It was wrong.
Not this one. Is it Apple? Microsoft.
Oh, it's freaking Microsoft, man. I should've gone with my gut.
Damn it. Okay. Oh, this is Google.
Wow. Look how sick that looks. This is giving Google, man.
This is definitely Google. I love this. It's done such a good job.
Can we go down the slide? Is that no. We can't.
This is Google.
Fabric speaker with four colored lights. This one was obvious. This one was obvious.
If if you guys got it wrong, then I'm honestly I'm judging you. So we got them all right kind of this time around. Let's have a look at the dashboard.
Give us a quick refresh. And the thumbnails, if they wanna load any second now, boom.
Okay. So Opus five did this one.
See, I'm telling you, man. The favicons, man. You gotta you gotta be looking at them favicons, bro.
Fable never misses a favicon. Now obviously, Opus oh my god.
51,000,000 tokens. Holy gajesus.
Again, very close to Fable in terms of cost. A hour 58 took an extra thirty minutes again. Now I'm not gonna lie, Fable did a fantastic job, but there was just more detail in Opus' designs.
You can see 3,700 lines of code from Opus and 1,500 from Fable.
That is literally double the amount of code. We can see total it cost us a $150 in API equivalent cost to do everything today.
Seven hours and forty minutes of compute to generate everything and a 122,000,000 tokens. Oh my god.
The winner is going to be Claude Opus five. Let me know in the comments below.
Did you guys know that this was Fable and that this was Opus, or did you have them the other way around? I definitely think Opus absolutely smashed Fable in this test. I'm not gonna lie.
It is the clear winner here. Do you agree with me or not really? If you like the video, be sure to leave a like and subscribe.
And if you wanna see a video where I tested Kimi k three against Fable five, you can check that video out right here. I'll see you over there. Peace.
The Hook

The bait, then the rug-pull.

Hours after Claude Opus 5's launch, this creator pits it blind against Claude Fable 5 across three identical one-shot builds — a trading terminal, a 3D kart racer, and an office-guessing game — to see whether the benchmark hype holds up once real code hits the screen.

Frameworks

Named ideas worth stealing.

03:42concept

ARC-AGI-3 (novel problem-solving benchmark)

Instead of static Q&A, models are dropped into unfamiliar, instruction-free game-like environments and must explore, form a plan, and learn from mistakes to solve them — testing general reasoning rather than memorized patterns.

Steal forEvaluating whether an AI tool can handle genuinely novel problems, not just ones similar to its training data
CTA Breakdown

How they asked for the click.

VERBAL ASK
19:10subscribe
If you like the video, be sure to leave a like and subscribe.

Standard end-of-video like/subscribe ask, plus a plug for a related video comparing Kimi K3 against Fable 5. The description also carries a free-community link and the prompt-list Google Doc used throughout the video.

FROM THE DESCRIPTION
PRIMARY CTAWhere the creator wants you to go next.
OTHER LINKSAlso linked in the description.
Storyboard

Visual structure at a glance.

cold open
hookcold open00:00
blind-test setup
promiseblind-test setup01:53
first reveal
valuefirst reveal07:51
final verdict
ctafinal verdict18:20
Frame Gallery

Visual moments.

Watch next

More from this channel + related breakdowns.

40:03
Pat Simmons · Review

GPT-5.6 Sol: No-Hype Full Review & Testing

A blind, four-way bake-off — GPT-5.6 Sol against Fable, Opus 4.8, and GPT-5.5 — across ten builds and knowledge-work tasks, scored one task at a time without knowing which model made what.

July 10th
Chat about this