Modern Creator
Pat Simmons · YouTube

Fable 5.1: No-Hype Full Review & Testing

Fifteen parallel builds, three models, and a blind reveal system, testing whether Fable 5.1's benchmark gains actually show up in real creative and coding work.

Posted
5 days ago
Duration
Format
Review
educational
Views
1.8K
128 likes
Part of the collectionThe Fable 5 PlaybookAll 45 Fable 5 breakdowns, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

Fable 5.1's official benchmarks promise large gains in agentic coding and automation, but blind side-by-side testing across five open-ended creative build tasks found only marginal quality differences from Fable 5, with cost efficiency, not output quality, standing out as the clearest real-world advantage.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You're deciding whether to switch a coding agent or automation workflow to Fable 5.1 and want more than a benchmark chart to go on.
  • You build with AI agents day to day and care about token cost and run time as much as raw output quality.
  • You want a repeatable method for comparing AI models blind, without brand bias creeping into the judgment.
SKIP IF…
  • You're looking for a benchmarks-only breakdown of Fable 5.1, this video is explicitly the opposite of that.
  • You need results on terminal-bench-style long agentic automation, none of these five tests are that kind of task.
TL;DR

The full version, fast.

Fable 5.1 keeps Fable 5's token pricing but cuts cache-read costs from $1 to $0.25 per million tokens, which Anthropic estimates saves 25% on typical workloads and up to 45% on long agentic runs. To test whether that translates into better output, the reviewer ran the same open-ended prompt (website, 3D world, browser game, motion graphics piece, and a full brand-revival deck) through Fable 5.1, Fable 5, and Opus 5 in parallel, then blind-ranked the results before revealing which model built what. Fable 5.1 didn't win a single one of the five blind rankings, but it came in cheaper than Fable 5 on every task tracked and often finished faster, making cost, not quality, its clearest edge in this round of tests.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0000:54

01 · Cold open

Pat Simmons introduces Fable 5.1's release and previews the five build categories he'll test it on (web design, 3D and simulation, game development, motion graphics, and a brand refresh), kicking off an agent fan-out of 15 parallel builds.

00:5403:51

02 · Pricing and benchmark deck walkthrough

While the builds run, he walks through Anthropic's official launch deck: pricing stays the same but cache reads get cheaper, and benchmark charts show large jumps on agentic terminal coding and automation work but only a marginal jump on everyday coding benchmarks.

03:5104:41

03 · Agent fan-out and SimmonsBench blind-test setup

The 15 builds finish; he explains his SimmonsBench workflow, which hides each model's output behind a lettered column so he can rank blind before revealing which model built what and what it cost.

04:4110:40

04 · Test 1: World-class website (bell foundry vs. cloud atlas)

Three open-ended 'build the ceiling of your web design capability' outputs are compared blind: two of the three independently landed on a bell-foundry concept, and the third built a scrollable cloud atlas. Reveal shows Fable 5.1 at roughly $20, Fable 5 at roughly $40, Opus 5 at roughly $28.

10:4015:10

05 · Test 2: Real-time 3D experience (storms and prairies)

Same open prompt for a real-time 3D world: a storm-tossed sailing ship, an advancing prairie squall line, and a locust-approaching prairie. Fable 5.1 finishes fastest and cheapest at roughly $39 versus roughly $126 for Fable 5.

15:1019:18

06 · Test 3: Playable browser game

Three browser games are compared: THAW (melt frost to grow grass), ACEQUIA (dig irrigation channels), and OSSAME (play as a spider catching flies). None has a real win or lose condition. Opus 5 burns about 8,000 seconds of agent runtime versus about 2,000 for each Fable model and ends up the most expensive at roughly $75.

19:1823:35

07 · Test 4: Motion graphics piece

Three self-contained motion-graphics pieces: a history of the 1854 Soho cholera outbreak and Dr. John Snow's map, an explainer on the Beaufort wind scale, and an espresso extraction profile. The cholera piece is the reviewer's favorite because it teaches something unprompted. Fable 5.1 costs roughly $20 versus roughly $28 for Fable 5.

23:3529:00

08 · Test 5: Revive a dying brand (J.Crew)

A full brand-revitalization deck for a struggling J.Crew: new logo, positioning, copy, and AI-generated lookbook imagery. Opus 5 is ranked first and is the cheapest at roughly $20; Fable 5.1 is ranked last at roughly $32, still cheaper than Fable 5's roughly $46.

29:0029:59

09 · Final verdict and sign-off

Pat Simmons sums up: Fable 5.1 didn't win a single one of the five blind rankings, though the tests never touched the long agentic and automation work Anthropic's benchmarks actually claim the biggest gains on.

Atomic Insights

Lines worth screenshotting.

  • Fable 5.1 keeps Fable 5's exact token pricing ($10 per million input, $50 per million output tokens); the advertised savings come entirely from a cheaper cache-read rate.
  • Anthropic cut cache-read pricing from $1 to $0.25 per million tokens, a 75% reduction on that specific line item, which is why long agentic runs see the biggest overall savings.
  • Anthropic estimates Fable 5.1 costs about 25% less than Fable 5 for typical workloads and up to 45% less on long agentic runs.
  • Fable 5.1's largest official benchmark jump was on Terminal-Bench Science, a new benchmark, followed by a big Terminal-Bench 4.0 gain from about 42% to 55.8%.
  • On everyday coding benchmarks like Cursor Bench, Fable 5.1 moved only from about 70% to 73%, a far smaller gain than the agentic coding and automation benchmarks.
  • Across five blind, side-by-side build tests, Fable 5.1 did not win a single one against Fable 5 or Opus 5, a rare result for a new model release.
  • In the tasks that showed the biggest cost gaps, Fable 5.1 came in noticeably cheaper than Fable 5 every time it was tracked, roughly half the price on the website build and over 50% cheaper on the 3D build.
  • On a playable-game build, Opus 5 ran roughly 8,000 seconds of agent turns versus about 2,000 seconds for each Fable model, and its higher final cost traced back to that longer loop rather than its per-token price.
  • On the one task framed as pure knowledge work, a full brand-revival deck for J.Crew, Opus 5 was both the reviewer's top pick and the cheapest of the three models.
  • Two independently-run models both invented a bell-foundry concept for the same open-ended website prompt, and two separately-run models both invented a storm-tossed boat scene for the same 3D prompt, with nothing in either prompt pointing to those ideas.
  • Multiple builds converged on the same cream-background, serif-font visual style without being asked to, a recurring AI aesthetic tic the reviewer flags as worth explicitly banning in future prompts.
  • None of the three models, when asked to build a playable browser game with no direction beyond 'avoid common genres,' produced anything with a real win or lose condition, only movement and interaction mechanics.
  • The single motion-graphics output the reviewer rated highest wasn't the most stylized one, it was the one that taught a real historical fact (the 1854 Soho cholera outbreak and Dr. John Snow's map) without being asked to be educational.
  • The reviewer's testing method (agent fan-out into parallel builds, blind ranking behind hidden columns, then reveal of model identity and cost) is built specifically to prevent knowing a model's name from biasing the judgment of its output.
Takeaway

Benchmark charts don't predict real-world creative output.

BENCHMARK VS REALITY

Official benchmark jumps for a new AI model release don't reliably translate into visible quality gains on open-ended creative and coding tasks, even when the price drops meaningfully.

02Pricing and benchmark deck walkthrough
  • Fable 5.1 keeps Fable 5's exact token pricing; the advertised savings come entirely from a cheaper cache-read rate, cut from $1 to $0.25 per million tokens.
  • Anthropic estimates roughly 25% lower typical-workload costs and up to 45% lower costs on long agentic runs, since agent workflows lean heavily on cached context.
  • The biggest benchmark gains Anthropic highlighted were agentic ones (Terminal-Bench Science, Terminal-Bench 4.0, automation benchmarks), while an everyday coding benchmark like Cursor Bench moved only a few points.
03Agent fan-out and SimmonsBench blind-test setup
  • A blind testing method that hides model identity behind lettered columns until after ranking removes reviewer bias that a visible brand name would otherwise introduce.
  • Every test used the same deliberately open-ended instruction, demonstrate the absolute ceiling of your capability and pick your own concept, so the comparison measured creative range rather than prompt-following.
04Test 1: World-class website (bell foundry vs. cloud atlas)
  • Two separately-run models both invented a bell-foundry concept for the same website prompt with no shared wording pointing to it, worth remembering before calling any single AI output 'creative.'
  • Multiple builds converged on the same cream-background, serif-font look without being told to, a recurring AI aesthetic tic worth explicitly banning in future prompts if variety matters.
  • In this test set, Fable 5.1 built the website for about $20 against about $40 for Fable 5 and about $28 for Opus 5, roughly half the price of its predecessor for a comparable result.
05Test 2: Real-time 3D experience (storms and prairies)
  • On the 3D-environment test, Fable 5.1 finished faster and cost about $39 versus roughly $126 for Fable 5, a savings of well over 50%.
06Test 3: Playable browser game
  • Opus 5 burned about 8,000 seconds of agent runtime on the game build versus roughly 2,000 for each Fable model, and its higher final cost traced back to that longer loop rather than its per-token price.
  • None of the three models built a game with a real win or lose condition from an open prompt, a reminder that human creative direction still has to supply the challenge and stakes.
07Test 4: Motion graphics piece
  • The favorite motion-graphics output wasn't the most visually stylized one, it was the one that taught a real fact (a historical cholera outbreak and its map) without being asked to be educational.
08Test 5: Revive a dying brand (J.Crew)
  • On the one task framed as pure knowledge work, a full brand-revival deck, Opus 5 was both the top-ranked output and the cheapest of the three models, with Fable 5.1 ranked last.
09Final verdict and sign-off
  • Across all five blind build tests, Fable 5.1 did not win a single one against Fable 5 or Opus 5, a rare outcome worth remembering before assuming a benchmark jump will show up in your own work.
  • The tests that showed the least difference were exactly the tests furthest from where the benchmarks claimed the biggest gains (agentic terminal coding, long automation runs), so a benchmark chart's category matters as much as its headline number.
Glossary

Terms worth knowing.

Terminal-Bench
A benchmark suite that scores how well an AI model completes real command-line and terminal coding tasks, as opposed to single-turn code suggestions.
Cache reads / prompt caching
Reusing input tokens the model has already processed at a steep price discount, instead of reprocessing that same context from scratch on every request.
Agent fan-out
Running the same prompt through several parallel AI agent sessions at once, typically in separate terminal windows, to generate multiple independent builds for comparison.
Agentic coding
A way of using an AI model where it plans, writes, runs, and iterates on multi-step coding tasks with minimal human intervention between steps.
Cursor Bench
A benchmark that measures everyday, single-session coding assistance quality, closer to routine developer work than long autonomous agent runs.
Blind ranking
Judging and ranking outputs before knowing which AI model produced each one, used to avoid brand bias skewing the evaluation.
Resources

Things they pointed at.

Quotables

Lines you could clip.

29:10
Fable 5.1 did not win a single one of these tests, which is pretty rare for a model release.
the whole video's thesis in one sentence, a rare admission for a launch-day reviewTikTok hook↗ Tweet quote
01:41
Cache reads went from a dollar per million tokens down to 25 cents.
the one concrete number behind Anthropic's cost-savings claimnewsletter pull-quote↗ Tweet quote
24:46
Revenue that only holds flat because prices went up isn't a comeback. It's a slower way down.
sharp AI-generated copywriting from the J.Crew brand deck, shows off the model's writing rangeIG reel cold open↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

All right, here we go. The model releases are back. Fable 5 .1 has officially dropped and Anthropic is claiming this is quite a large leap from Fable 5, at least according to the benchmarks.
But of course, this isn't a benchmarks channel. So in this video, we're going to be running Fable 5 .1 across a variety of tests, web design, 3D and simulation, game development, motion graphics. We've even got a full brand refresh in there.
We will compare it to its predecessor, Fable 5, and to its apparently better, but we're still not really sure counterpart, Opus 5. So by the end of this video, you'll see how much of a jump Fable 5 .1 really is and if it's truly worth it when it comes to cost and efficiency so not wasting any more time let's get right into the build and we're going to kick it off like we always do with my agent fan out here we'll hit enter and we'll have five iterm windows pop up across the three different models so we've got 15 builds in total i will have this skill in the link below as well if you ever need to do this and have a bunch of agents working on this stuff at once and while the agents are building let's just quickly gloss over the benchmarks and fable 5 .1 and i put together a deck with some notable callouts from anthropics press release but first i just wanted to a look at this thing.
You know, you've got to hand it to Anthropix. Say what you will about them. The branding for their model releases always just looks fantastic.
You even got little toggles here. You go dark mode. You can go dusk.
Well done, as usual, Anthropic. But let's actually go to the deck that I put together. And I'm going to gloss over everything of note in their press release.
So the first notable callout is pricing. Pricing is exactly the same as Fable 5. $10 per million input tokens, $50 per million output.
But what changed here is cashing. And Anthropic is estimating about 25 % cheaper costs across the board compared to Fable 5 because cashing got quite a bit cheaper. So cash reads went from a dollar per million tokens down to 25 cents, which I know is technically...
75 % cheaper, but of course not all of your tokens are going to be cash reads. So it looks like Anthropic is just averaging it out to like 25 % cheaper costs. But then on long agent runs, which is probably most of what you'll be using Fable with anyway, you're looking at costs up to 45 % cheaper.
So it's quite a reduction in costs, but we'll actually see when we run these tests on these live builds, how true that actually holds. So that is pricing. Now let's quickly look at how 5 .1 compares to Fable 5 and a couple other models with some notable benchmarks that Anthropic called out.
The first terminal bench science which i believe is a brand new benchmark and i'm going to be honest i know nothing about science and none of these tests are going to be related to science maybe if anyone has any ideas on some sort of sciency tests that i can understand as a complete novice we could try that but this is probably the most notable jump among all the different benchmarks you can see fable 5 down here it will 5 .1 quite a bit of a difference so if you're in that field i mean it's a no -brainer to really be stress testing and using fable 5 .1 as much as you can next up we have terminal bench 4 .0 and a pretty big jump from mythos 5 which It's still hilarious to me that a couple of months ago, this was like this incredibly dangerous model.
And now 5 .1 is just straight up outperforming it in every way. But I guess technically 5 .1 probably still has the guardrails, the safety guardrails that Mythos 5 didn't. And this is just estimating pure performance.
That being said, I mean, you can see even it's a much larger gap between Mythos 5 than Fable 5 .1 and Mythos 5 .1. So that is something that is notable when it comes to agentic terminal coding. And then here's that jump from Fable 5 to 5 .1.
So 42 % to 55 .8%. That's actually a pretty big. gain.
And then we have some other models here, like five, six soul, 37%. So agentic terminal coding. I mean, that is a big jump.
I promise we're almost done with these and I'll stop saying big jump for everything, but we have next automation bench. There's like business workflows in general. Nice bump again, nice jump again from Fable 5 to 5 .1.
So this is just automation work across a number of simulated apps. This is actually something that Zapier built, which I didn't know. I was just curious and I was researching this today and random aside, but yeah, pretty impressive jump when it comes to actual automation.
So if that's something you use in your day -to -day notably opus 5 which is cheaper than fable 5 .1 it's pretty close and much better than fable 5. and then five six soul quite a bit lower and then just some quick agentic coding notes as well cursor bench take a look at this not a huge jump when it comes to everyday coding from 70 to 73 it's really just that agentic coding that's a big call out and then we just have humanity's last exam as well so again Not a huge jump.
But all that being said, we'll see how it actually performs when we put it to the test and we can start to see the differences between Opus 5 and Fable 5 with 5 .1. So with that, let's actually get into that now. All right, the moment we've all been waiting for.
We've got all of the builds completed. Done, done, done, done. Across the board.
Let's close these out. I have them connected to Simmons Bench per usual. So we're going to check these out here.
And we're going to do, just like we always do, a blind ranking of each of these. So I have the prompt here. All of this will be in a blog post in the description below, as well as these live links.
But for our first build, the prompt is this. Build one website that demonstrates the absolute ceiling of your web design capability, your taste, your artistic flavor, your technical range. You choose what the site is for.
You pick the concept, composition, typography. neon or purple gradients or going particle fields or anything like that. I gave it access to image generation as well via GPT image gen and just ask it to QA.
So I want to do the same test I did with Opus 5. I just felt like this was a good kind of test of the model's creativity. Instead of giving a brute force prompt, but the outputs were really kind of the same.
I wanted to see what each of these models will come up with. If I just say, build me something that is the absolute peak of your design capability. So that's what we're doing with basically all of these tests across the board.
We're starting with just a website. then we're going to see how different these actually are and we'll guess which model is which and determine if it will 5 .1 is really that noticeable all right so first up column a here i'm already not liking any of these outputs it all kind of looks homogenized in the same but let's just take a look at a okay nice little loading all right what do we have here eight bells in cast bell metal 78 parts copper to 22 parts tin i should also tell these models to not just hallucinate nonsense but okay not one note on this page is a recording every bill is built from its own particles that the ratios its founders tuned into its founder tuned it to okay so it's gonna make some noises yeah i don't know i don't know if you can hear this but it's like different pitches of bells okay yeah kind of creative kind of creative the same kind of like cream background with the serif font just like really drives me nuts i should absolutely be including this as an anti -slop guidance no more of these cream backgrounds with serif fonts but okay all right nice we got some anatomy here so we've got some svg generation we've got uh some more information here way too much text on the page oh nice look at that
nice nice little uh what do you call this half tone yeah half tone two color generation probably with gpt image gen just kind of all over the place like it's just i don't know why it's towing that but i guess it looks cool then five partials okay oh cool so it's making noise you can't hear that but it's actually making noise uh and i like this little graph here too strike all together whoa cool oh wow that's actually really cool it's like going through each of these little pitches yeah all right i mean an out of the box idea that's for sure and then we have some more Okay.
So it's making noises as I click ring the rounds. Oh, cool. Okay.
So it's making, it's kind of making music out of these bells. All right. This is, yeah, this is a creative idea.
And then we have, we've got some more stuff. Blaine Hunt, Blaine Bob Major, ring it. Whoa.
Wow. See, I know anything about music. This is probably pretty cool, but I don't.
Also, how do I stop this? Okay. Foundry, Holloway and Vane is an invention.
And so whether it's eight bells or names or weights on the sound, on the drawing. Okay. Yeah, sure.
Pretty good. Pretty good. All right.
That's the first one. Next up, let's see what creative web design. All right.
What is going on? Another bell. What's going on?
A bell is not in any way in the prompt. I don't know why they're choosing bells for everything. Did this one choose a bell too?
No, I feel that list of clouds. Okay. So that's so weird.
They both chose bells, but all right. Let's check it out. Askenfell.
Every bell we cast sounds five notes at once. It's the same thing. Oh my goodness.
Okay. Octave below. Tears.
all right this doesn't even really play the bells it's only this strike the tenor that does this this is so strange then that it had the exact same idea as the other one okay because it just makes it bigger okay as i make it bigger i'm assuming it's gonna be like deeper yeah it looks like that's the case nice little cool little animation okay we got more bells we've got these exact same rounds i'm telling you there's nothing in this prompt that mentions bells or anything related to this idea so it's very odd that both of them came up with the same idea but okay let's keep going another image gen kind of worse image gen like a weird kind of sepia tone here and then pouring let's not take the exact same idea what is going on do they like cheat on this who knows okay the foundry all right that is so strange anyway third output let's check it out a field atlas of clouds this is similar to generations in the past where you scroll and then it kind of you have this like meter that will keep going down yep yep yep yeah nothing nothing special about this oh but these images are pretty cool can i actually
Now I can't drag this up. Images are like at least kind of unique. Okay, so these are actually going through, I believe it is explaining these different types of clouds.
The Seraphont, a dead giveaway. Seven specimens collected between the mesosphere and sea level. Every plate computed in the browser, none of them photographs.
Oh, really? Put it in the browser. I think it means GPT image.
And I don't think it's doing anything special to actually generate these, but read from the top down. The descent is 85 kilometers. OK, how to read the Atlas in 1802.
A London chemist named Luke Howard gave the clouds their Latin names. I didn't know that. Cool.
So you can actually click through. Nice. It goes down.
OK, we have all the different Latin names of different types of clouds. I'm assuming. do we have any cloud experts watching please if you're a cloud expert drop your commentary in the comments really want to know if this is accurate or if this means anything or if this is just complete nonsense because i'm not a cloud not a cloud guy anyway we're going to keep scrolling here we've got more clouds remember that all of these will be available in the link below so you can actually click through and take a look at this yourself but uh yeah not bad At least it's a different idea.
I'm just going to give it to that model because it's at least a different idea. But it's all the same kind of design. Just very annoying.
This is going to be on the banned list from here on out. But I'm just going to guess this is 5 .1. Oh, look at that.
Wow. All right. And then I don't know.
I don't know which one's going to be. I guess this is 5 .1. It was like, no, I mean, I think I think this one was better.
This is 5 .1. What is going on? So, okay.
All right. Fable 5 .1. Wow.
Did Fable 5 .1 copy Opus 5 and then try to do it better and then it was just worse? All right. Not seeing much of a jump so far in 5 .1, but that is the first test website.
But let's see how well it does on the next test, which is... 3d generation again similar kind of prompt where i'm just saying get creative with it build a one -time real 3d experience that demonstrates the absolute ceiling of your 3d and simulation capability not giving it any sort of concept or anything it's just going to pick that and i said don't build any of the following because last time was building a lot of black holes and solar systems and stuff like that even though those are cool i'm just kind of getting sick of them so i said Avoid all of that.
And I want to see what else it could come up with. I just said, like, build a place with real scale, too. So instead of this, like, floating space environment, I want to see how well it did there.
And then I just keep going on really stress testing, you know, technical ambition, simulation, all that kind of stuff. So let's check it out and see. All right.
I've seen some different generations. That's a good sign. So let's look at A first.
Woke A. All right. All right.
We've got we've got some craziness going on. I can't control this ship. This is about.
This is an intense scene, though. I just can't actually, jeez, do anything. I need to open this in a new tab.
Nope, it's just going crazy still. Didn't do a very good job with QA. The boat looks good.
Got some nice lighting on the floor there. Okay, now we've got some control. Oh, nice.
Okay, so I can actually move around this boat. Wow, these waves are actually pretty realistic. Yeah, this is cool.
So I can go around this boat. This is exactly what I was saying. In the prompt, I wanted to be able to move around the environment, not just be...
controlling a solar system but yeah i mean wow these waves are pretty darn good and then it looks like i have to like it escape i think to actually adjust these controls sea state okay so yeah that gets more aggressive wow these waves are pretty good it doesn't actually look realistic it's like okay i guess it does looks a little odd when you're looking at it from an angle like that but the way the boat is hitting the waves is pretty impressive look at that you've got some water coming up still a little funky but this is pretty darn good and then i can control different views for deck okay cool Click to look around.
Nice. And then we can control break in the cloud. Oh, nice.
Look at that. Sun height. Wow.
Got a lot of controls here. Wind. Whoa.
Wind getting crazy. Whoa. Whoa.
Okay. That's why I was freaking out because you get, if you get aggressive with the wind. All right.
Throttle. Okay. I got to turn this.
Wow. It's freaking. Okay.
That's a little bit of a bug. It changed the wind. It just starts freaking out.
Yeah. This boat looks good. Let's just mess around with the rest of these controls.
Okay. What happens? Rain.
Rain gets heavier. rain gets lighter nice helm that's what helm does yeah pretty impressive i really like that one all right what else we got second one outflow a real -time prairie under an advancing squall line wow big weather theme today for these models click to enter whoa that's pretty cool it's all right kind of weird it just isn't very fluid it's very stiff another way to describe it okay and then it's something i don't know what's happening the grass just is like moving it's just like infinitely moving.
Look at those clouds. Very nice. And then I don't know what's happening here.
There's like thunder, like what's going on. It looks pretty good. I like these little, this little field here.
This is nice. And wind, F, fly. Okay, F doesn't do anything.
Cinematic, cinematic. It's a little smoother, much faster. It looks like it's just like a never ending movement here, but I can go around.
That's cool. Storm. Whoa.
Okay. So I'm pressing these brackets and it's, whoa, does it actually start a storm? Clouds aren't very realistic, but everything else is like, whoa.
that's pretty cool all right creative and looks great well done mystery model third another prairie what what is going on i'm gonna check i promise i'll fix this these models are cheating but this is just too weird to be a coincidence at this point so let's see click to enter okay just like a another prairie looks pretty good i mean all of these are very what is that locus locus approaching bonus points for the locus it also just looks more realistic actual weather and the actual clouds A lot of clouds, a lot of clouds today.
I like these little rolling hills over there. Yeah, all these are pretty darn good. I don't know.
I really don't know. I'm going to say I think I'd like this one the best just because you're actually doing something. And then this one, number two, because it's a little bit more realistic.
And then number three. So let's see. Number one.
There we go. Fable 5 .1. Number three.
Opus five. Okay. Yeah.
Fable five. All right. Interesting.
So Fable, I mean, these two are very close. It was all very close. There wasn't much of a difference in any of these.
Oh, I forgot to mention too. So this is cost. We'll look at costs in the previous build in a second.
But if we look at costs here, Fable 5 .1, $39. Fable five, $126. So wow.
It's over 50 % cheaper. And then Opus five, it's actually, wow. It's actually way cheaper than Opus five and quicker than Opus five.
Yeah. Very interesting. So yeah, this is like probably something, let's take a look at these.
costs here for this one this website but if we show the models yeah fable 5 .1 20 fable 5 40 so 50 cheaper for 5 .1 and then opus 5 28 so yeah again fable 5 .1 the cheapest granted wasn't very good but it was the cheapest and it got this done faster than the other two models all right number three build a playable game similar kind of idea it's a playable browser game give that same prompt i want this to demonstrate the absolute ceiling of your game development capability you choose the game and then i included some obvious answers like asteroid -like game you know games like pong or snake or tetris or anything like that anything with neon or dark glow or gradients anything like that just avoiding all of those generic outputs that i've seen in the past and we're going to see how well they actually build these games so first up thaw howlemure valley the winter never ended here you carried the last warm light i love when these models just spew nonsense sweep it across the frozen valley and watch the green come back up through the snow okay oh all right pretty good big little graphics here hold click to melt the frost okay so i guess i just click here melt oh look at that wow very cool all right not much of a game but uh yeah you can just like aim this orb at the snow and it'll turn into grass okay interesting warmth
It's like, there's no, there's no challenge in this game, but it looks very cool. I will give it that. Okay.
I think that's literally it. Boring. A sequia.
Again, just a ridiculous name. You are the keeper of a mountain spring. Okay.
At least it explains the premise of the game. The fields below are thirsty. Dig channels with your shovel and the water will find its way downhill.
Okay. Click to begin. Okay.
Thirsty. It's ridiculous. right click okay all right i mean it's got some like advanced controls here those are the crops dig beside the field not in it okay i like the little prompts that pop up so okay so i can start digging got it okay i'm right clicking right now and it's i'm digging but it's creating a mound for some reason i'm not sure what's doing that i like this though this is kind of like a minecraft type game It's a little bit more confusing.
OK, left click. I don't know how to shift run. Dig beside the field, not in it.
Oh, right click is piled dirt. OK, so it was right. OK, so that's accurate.
But all right. So if we left click, then we start digging. OK, cool.
Oh, yeah. Nice. So I guess you just keep digging.
Yeah. And then we're going to get to a spring reserve eventually. Oh, this is so ridiculous.
OK, come on. Oh, yeah. Look at a little spring water popping up.
Bedrock. It goes no deeper here. thirsty.
Still trying to figure out what to do here. Okay, I think you get the picture. This game felt a little bit more advanced than the last one.
And then let's look at this one. Oh, some may go summer and orb weavers dawn. She's all right.
Okay. Oh, you're like a spider. All right.
Hey, bonus points for creative idea. Again, not much of a challenge here. Graphics are pretty good.
Okay, and then just flies just keep coming into the net and you fly to the flies.
Like, okay, next iteration of this prompt is going to be make it challenging somehow. This is just ridiculous. Keep doing this nonstop.
Keep catching flies. Movement and stuff's pretty cool, but these models don't seem to understand how like true games work. This is where like, you know, you still need humans to give it some sort of direction, make it feel competitive, make it feel fun.
This is just a spider eating various insects. anyway i actually think i like that one the best so we're gonna go number one for that we're gonna go number two for this one and then number three this one's nothing special so number one fable five i mean okay five one yep five one all right then opus five all right not much of a difference with fable 5 .1 i mean granted i'm making stupid games but there is usually quite a noticeable jump which i'm just not seeing with 5 .1 however let's take a look at costs so okay fable five fifty seven dollars five point one forty seven dollars and then opus five seventy five dollars i'm guessing just way more turns and less efficient yeah so it took 8 ,000 seconds versus 2 ,000 seconds versus 2 ,000 seconds.
So because Opus 5 just kept going in turns, it got more expensive, even though Opus 5 is considerably cheaper than these two models. And again, Babel 5 .1, definitely cheaper. Anthropic wasn't lying, not that they ever would be, but it is interesting to see this in real tests, the difference in cost overall.
Next up, we have motion graphics. So we're getting more into tasks that are relevant to your day -to -day now. I always like to do motion graphics, even if you're not generating motion graphics.
Any sort of video generation is something that these models are still not very good at, at least not in a one shot. Opus 5 was definitely the best I've seen, but I'm curious how it compares to 5 .1.
So I gave it a similar kind of prompt. I said, again. motion graphics that demonstrates the absolute ceiling of your motion design capability i said pick the concept very similar to the other prompts and just gave some terms here kinetic typography choreography variable font animation i don't even know what a lot of these terms are i literally just had another agent search these terms and then just list them out here but we have all these different terms we're going to see how well these models actually generate motion graphics so first up we've got soho london 1854 it was a tough year yeah cholera 500 dead in 10 days everyone playing the air whoa Oh, that's a cool little fade.
A doctor, John Snow from Game of Thrones, really? Blamed the water. Whoa, this is actually really cool.
I like this kind of like history angle. It's going to 600 deaths. The Broad Street pump.
Deaths drew. They took the handle off the pump. No way.
Did they replace it? The outbreak ended. Modern epidemiology began with a map.
Broad Street, Soho. Bravo. Bravo.
I learned something. Those animations were pretty cool. Let's watch it one more time.
Replay. So it landed. 1854.
Cholera. 500 dead in 10 days. Wow, it was tough back then.
Everyone blamed the air. Nice little fade to the air. Nice little text there.
The map isn't exactly correct. It's kind of all over the place. But this nice little zoom in where you focus on the epicenter or whatever you call that.
patient zero, and then boom. It took the handle off the pump. And just like that, the outbreak ended.
Very nice. Very nice. Okay.
Alright, so this one, we're getting into like, so this one's hurricane. Talking about hurricanes. Okay, so this is just like a learning.
Alright. I'll just let it play.
The Buford scale. Calm. Less than one knot.
Less than one mile per hour. Smoke rises vertically. Light air.
Smoke drifts. light breeze wind felt on face gentle breeze moderate breeze okay so it's getting into like what okay wow this is like whoa this is good too geez near gale gale that's showing wow it's getting stronger and stronger and stronger wow it's from hurricane devastation 64 knots and above interesting so again it went like this like historic francis buford buford 1805 I like this one.
This is the first time, just random aside, this is the first time I felt like a model wasn't just spewing nonsense and it actually taught me something when I didn't ask it to. Really like that. This one too is pretty good.
It was a nice visualization of like what Hurricanes are, like how Hurricanes, obviously people know what Hurricane is, but made it immersive. Well done. Okay.
Third one. Okay. Whoa.
All right. We've got something. All right.
I need to restart this, but we've got some nice grain here already. Okay. Espresso extraction profile.
Interesting. pre -infusion three bar water soaks nine bar wow this is cool this is a good example like i don't know this is typically what a model would output i don't know what's going on it's not really teaching anything i'm kind of confused but it looks freaking cool extraction yield 3 .6 grams of coffin salt the other 14 .4 stayed in the puck it's like who cares but also very cool i'm actually i i still think this one just because it's been really well done great svg this one's definitely three all right which one which is opus five geez okay this has to be five one no What?
Wow, 5 .1. Yeah, I mean, 5 .1 just across the board. Not really anything special.
Then we look at cost too. So Opus 5 was actually the most expensive because I'm sure it went through, yeah, quite a bit longer, a number of turns. It looked awesome though.
Definitely worth it. $20 for Fable 5 .1 and then $28 for Fable 5. All right.
So that is motion graphics. Actually pretty impressive. That was, that was actually a good prompt.
I like the thought of just saying, teach me something too, because it gives it, I didn't actually say that, but I will the next time or give some kind of guardrails where it just doesn't have to spew nonsense like it usually does. All right. Final one, which by the way, is reviving a dying brand.
I did this in the Opus 5 video. I liked what the models output there. And I like how it's a test of creativity, writing persuasive copy, putting together decks.
So it's just kind of a pure knowledge word test. So this is the prompt. Again, all the will be in the description below but for this revival we're choosing j crew and we're saying deliver a complete brand revitalization as a single self -contained html presentation we give some context i probably even shouldn't have included this context because it's just going to parrot this context and probably not too much research but we'll see and then we say there's a clothing brand like yeah no duh it's gonna know it's a clothing clothing brand and then we want like a new logo we want type we want everything and then i wanted i wanted a big idea i'm asking things that would actually be presented to a client.
So I'm asking for the strategy, the creative. I even wanted to see the collection. So have it generate images based on this new, whatever this idea is and what a collection of clothes would look like for this kind of new brand rollout.
And then we have a logo, we have color palette, typography, all of that. So let's check it out. Oh boy.
All right. I'm not liking what I'm seeing so far. So J crew.
Oh boy. Like this does not look like clothing brand anyway, but okay. J .Crew, not bad deck design, but this like looks nothing like what J .Crew would look like or any clothing brand, but okay.
Classic became static. Okay, setting the tone. All right, the precedent.
This is the big idea. America recut. Keep every code J .Crew owns.
Change every measurement. Oh boy. The barn coat, the rollback, the rugby, the Chino.
Whoa. all right that looks pretty cool geez good job on the uh on the actual generation so got some clothes here some nice like harsh lighting weird kind of like barn but okay the new crew all right not bad the brand story okay so this is another good test like it introduces like the actual structure of a deck like why would it go strategy showing some of the creative and then get back to the actual brand story doesn't make any sense i wonder if the other models do that strategy the creative move will each of varia i spy a ton of contrast based framing mandate the whole house not a capsule okay barn code all right so looks like they're just coming up with new names for like existing pieces of fashion first boat this is very confusing but love the photography like it generated this with dvd image gen which i guess is more of a compliment to gpt image gen but very interesting cool kind of like fashion shoot here at the very least i mean this is a good job of prompting the image generator first boat and the oxford gown wow interesting looks like something a sumo wrestler would wear the column roll neck wow okay all right oh nice we got some nice good svgs here this is a very detailed deck i will say that oh nice view there the big barn 248 okay crew mark all right so this is like
a little bit more tactical in terms of like the actual designs. I'm assuming that's what the other models will do too. I thought it would be more of a like strategy as a whole, but okay, not going to get nitpicky here.
And then we have this new logo that I came up with and some more stuff really in depth. I mean, it just kept going. We've even got some mock -ups in the actual store.
Social posts doesn't look very real, but hey, it's something. there we go all right that's number one not bad all right next one we've got jay finish it yourself let's check this out in a new tab okay setting up the brief the press verdict more information like too much background information but okay one line positioning j crew stops selling the finished look and starts selling the raw material of one doesn't make any sense but all right the finished brand is a static brand i want the customer j crew lost the one it never went and got okay so it's saying what is it saying saying different demographic secondhand native buys 40 of their wardrobe resold already owns a j crew barn jacket bought for 38 on depop yeah that's probably true okay so it's saying like age 26 to 32 first job with real money okay the last believer okay so at least it's breaking it out into audience segments that seems like something that would go into a deck like this the deck designs it it's pretty good i mean there's text all over the place the idea is
not great okay so they're recommending bringing in martine rose just a bunch of text here retire nothing recut everything interesting okay we've got some svg generation okay again wow this is might even be better looks great i don't know this is something you'd see at fashion week that's just like so ridiculous looks like david byrne for the talking heads in that big coat but okay all right interesting yeah very getting crazy with it getting crazy with it i mean the image generation just is really cool to like bring it to life like this and then we've got some nice svgs here cool looking jacket cut j okay not bad all right more color palette typography and the actual store social posts yeah social posts look pretty good the alteration bar yeah i mean this is that was pretty good too right final mystery models test of knowledge work we've got this deck maybe i need to open another tab yep deck design is pretty good text kind of all over the place okay after hours a catalog company okay all right they've all suggested they bring in some kind of creative director here's some svgs all right not bad okay gpt image gen wow very impressed so they all kind of went with uh you know different style this looks pretty cool these outfits are ridiculous what's going on here i don't know maybe i'd wear something like this
that's crazy all right cool so we just have these svg generations i mean all of these models did really good on the actual image prompting organizing like that coming up with some sort of like creative look too so i'm again not seeing much of a difference so i'd probably say this is number one just because i hated this logo this is third this is second let's see opus five i mean come on fable 5 .1 is last and then second is fable five all right and then costs opus was the cheapest at $20, Fable 5 .1, still cheaper than Fable 5, $32 versus $46.
Yeah, I mean, so there it is, Fable 5 .1 across building and knowledge work. And surprisingly, Fable 5 .1 did not win a single one of these tests, which is pretty rare for a model release. Then again, this is a point release.
Anthropic wasn't making a whole lot of promises in the benchmarks with this, but I still thought we'd see some kind of clear difference and we just didn't, at least in these isolated tests. Now, granted too, we're not running the tests where they said it excels. This isn't terminal.
bench. It isn't long running agentic automations. This is more the day -to -day coding, the cursor bench type benchmarks where there was only a marginal jump.
So take that for what it's worth. That's my opinion. I'll let you be the judge.
Here's what you think of Fable 5 .1 too. Maybe I'm missing some things. I am going to keep putting this to the test.
And as always, if you've got suggestions for future tests, I'm always open to ideas. I love stress testing these new models so you can get a sense of where these models actually perform and when to use one over the other. So thanks for watching.
As always, I appreciate it and I'll see you in the next one.
The Hook

The bait, then the rug-pull.

Anthropic dropped Fable 5.1 promising a 25 to 45 percent cost cut and a stack of benchmark wins over its predecessor. Pat Simmons skips the press release and puts it through five blind, side-by-side build tests against Fable 5 and Opus 5 instead.

Frameworks

Named ideas worth stealing.

03:51model

SimmonsBench blind agent fan-out testing

  1. Send the same open-ended prompt to multiple models via an agent fan-out (parallel terminal sessions build simultaneously)
  2. Load each finished build into SimmonsBench under a hidden, lettered column
  3. Rank the builds 1-2-3 before revealing which model made which
  4. Reveal shows model identity plus token cost and build time per column

Pat Simmons's structured method for reducing brand bias when comparing AI model outputs: build blind, rank blind, reveal last.

Steal forany AI tool comparison video, blog post, or internal model-eval process
CTA Breakdown

How they asked for the click.

VERBAL ASK
00:54link
I will have this skill in the link below as well if you ever need to do this and have a bunch of agents working on this stuff at once.

Soft, mentioned once in passing rather than pushed; the description carries the real CTAs (newsletter subscribe, blog post with prompts and live links, GitHub repo).

FROM THE DESCRIPTION
PRIMARY CTAWhere the creator wants you to go next.
OTHER LINKSAlso linked in the description.
Storyboard

Visual structure at a glance.

cold open
hookcold open00:00
Fable 5.1 announcement + pricing deck
promiseFable 5.1 announcement + pricing deck00:54
SimmonsBench blind compare launches
valueSimmonsBench blind compare launches04:41
final test: revive a dying brand
valuefinal test: revive a dying brand23:35
verdict and sign-off
ctaverdict and sign-off29:00
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.

27:44
Pat Simmons · Review

Opus 5: No-Hype Full Review & Testing

A blind, five-round test pits Opus 5 against Fable 5 and Opus 4.8 across web design, 3D, games, motion graphics, and a SpaceX investment deck — model names stay hidden until the ranking is locked in.

July 25th
36:06
Pat Simmons · Review

Kimi K3 Is Here! (Better Than Opus 4.8?)

Ten identical builds, five models, blind-ranked before the reveal — a real-world stress test of Moonshot AI's new open-source model against GPT-5.6 Sol, Opus 4.8, GLM 5.2, and its own predecessor.

July 17th
40:03
Pat Simmons · Review

GPT-5.6 Sol: No-Hype Full Review & Testing

A blind, four-way bake-off — GPT-5.6 Sol against Fable, Opus 4.8, and GPT-5.5 — across ten builds and knowledge-work tasks, scored one task at a time without knowing which model made what.

July 10th