Modern Creator
Pat Simmons · YouTube

Opus 5: No-Hype Full Review & Testing

A blind, five-round test pits Opus 5 against Fable 5 and Opus 4.8 across web design, 3D, games, motion graphics, and a SpaceX investment deck — model names stay hidden until the ranking is locked in.

Posted
today
Duration
Format
Review
hype
Views
8.3K
583 likes
Part of the collectionThe Claude Opus 5 PlaybookEvery Opus 5 breakdown, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

A custom blind-ranking tool that hides model identities until after judging shows Opus 5 beating Fable 5 in all five creative and knowledge-work tests, at roughly half Fable's price.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You use frontier AI models for coding, design, or content and want a blind (non-benchmark) comparison before switching tools.
  • You're deciding whether a newly released model justifies dropping a more expensive incumbent for daily work.
  • You want a repeatable method — blind ranking, matched prompts, tracked cost — for testing models yourself.
SKIP IF…
  • You only care about raw vendor benchmark numbers — this video's real value is the blind hands-on comparison, not the published charts.
  • You need software-engineering-specific benchmarks — these tests skew toward creative/visual generation and one finance deck, not coding tasks.
TL;DR

The full version, fast.

A YouTuber runs three AI models — Opus 5, Fable 5, and the older Opus 4.8 — through five identical creative prompts: a showcase website, a real-time 3D scene, a playable browser game, a motion graphics piece, and an investment-analysis deck on SpaceX. Using a custom blind-compare tool that hides which output came from which model, he ranks each result before revealing the source. Opus 5 wins all five rounds, matching or beating Fable 5's visual polish and technical depth while costing about half as much per generation — and sometimes undercutting Opus 4.8 too. His conclusion: unless there's a hidden reason for the price gap, the benchmarks weren't exaggerated this time.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0000:51

01 · Intro

Host previews the test: three frontier models — Opus 5, Fable 5, Opus 4.8 — run through five identical creative prompts, blind-ranked before reveal.

00:5104:21

02 · Benchmarks

Walks Anthropic's own published benchmark charts for Opus 5 vs Fable 5 vs Opus 4.8, questioning why a cheaper new model would beat the flagship across agentic coding, knowledge work, and computer use.

04:2110:48

03 · Test 1: Web design

Each model builds one 'ceiling of web design capability' single-file HTML site; all three land on a space theme. Blind-ranked in a custom compare tool before reveal.

10:4814:57

04 · Test 2: 3D and simulation

Models build a real-time 3D/simulation experience — black hole ray tracers and an interactive underwater jellyfish scene — ranked blind then revealed.

14:5717:44

05 · Test 3: Playable game

Each model builds a playable browser game meant to be fun within thirty seconds — a space shooter, an asteroid-style shooter, and a boat survival game.

17:4421:07

06 · Test 4: Motion graphics

Models generate a self-contained animated motion-graphics piece — a Dante-themed title sequence, a data-story explainer, and an abstract signal/noise piece.

21:0727:10

07 · Test 5: Knowledge work

Models act as a private wealth advisor building an investment-committee deck answering 'should I buy SpaceX?' — testing research depth, forecasting, and deck design.

27:1027:44

08 · Verdict

Opus 5 wins all five blind rounds against Fable 5 and Opus 4.8, at roughly half Fable's cost — the host concludes the benchmarks weren't exaggerated.

Atomic Insights

Lines worth screenshotting.

  • A blind-ranking method — hiding which AI model made which output until after judging — removes the bias of already knowing which model 'should' win.
  • Opus 5 won all five blind rounds against Fable 5 and Opus 4.8, spanning web design, 3D simulation, game development, motion graphics, and knowledge work.
  • Opus 5 is priced the same as Opus 4.8 ($5 per million input, $25 per million output) — half of Fable 5's $10/$50 — while outperforming both.
  • Giving a model a fabricated high-stakes frame ('this is going in a video, your work will be credited by name') is a usable prompting technique to raise effort and polish.
  • Three separate models given the identical open-ended web-design prompt all converged on a space theme, likely because the prompt's own vocabulary ('particles', 'flow fields') nudged them there.
  • A distorted or 'curly' rendering of certain letters, like the letter f, in AI-generated on-screen typography is a concrete visual tell for AI generation.
  • Total build cost can stay flat between two models even when one produces clearly better output, because the stronger model may take more turns to self-correct and verify its own work.
  • Prompting a model to 'do serious research, real current figures, don't just read headlines' before building a business deliverable produced visibly deeper output than a vague design ask.
  • Two independently-run models converged on the same core financial thesis for the SpaceX deck — that a diversified company is really three businesses with different risk profiles.
  • A same-lab successor priced the same as the outgoing flagship while beating the more expensive model on most benchmarks is a signal the pricier model may be getting replaced soon.
Takeaway

Three AI models, blind-ranked, no name revealed until after the vote.

WHAT TO LEARN

Hiding which model made which output until after ranking is the real method here — it turns a vendor benchmark chart into an actual verdict, and it's a technique you can copy for any product comparison, not just AI models.

02Benchmarks
  • A model priced the same as its own predecessor while beating a more expensive flagship on most benchmarks is a strong signal the pricier model is about to be replaced or discounted.
  • When a vendor's most dramatic claim (like exploit-finding capability) is reported only for a model you don't have access to, treat that category of claim as unverifiable rather than as evidence either way.
  • A same-lab, same-price successor that quietly outperforms the premium model across most categories is worth testing yourself before assuming the benchmark chart is just marketing.
03Test 1: Web design
  • Giving an AI model a fabricated high-stakes frame — 'this will be shown to a large audience and credited to you by name' — is a usable prompting technique to raise output effort and polish.
  • Open-ended creative prompts from multiple models tend to converge on the same theme when the prompt's own vocabulary suggests it (here, 'particles' and 'flow fields' nudged all three toward space) — vary your language if you want genuinely different creative directions.
  • Smooth, correctly-layered scroll transitions and consistent light/dark mode support were the specific technical tells that separated the strongest output from the rest, not just visual flash.
04Test 2: 3D and simulation
  • A real-time 3D build that lets the user actually manipulate physical variables (orbit, lensing, disk temperature) reads as more technically ambitious than one that only offers a passive camera view.
  • Departing from the theme every other model chose was treated as a legitimate differentiator on its own — sameness reads as less effort even when the execution is technically solid.
05Test 3: Playable game
  • A playable prototype has to be fun inside the first thirty seconds with no tutorial wall — games that made the reviewer hunt for controls before anything happened got penalized regardless of mechanics.
  • Weak graphics can lose a comparison even against more interesting gameplay, because visual polish is the first thing a viewer judges before they've played long enough to feel the mechanics.
06Test 4: Motion graphics
  • A distorted or 'curly' rendering of certain letters in generated on-screen typography is a concrete, repeatable tell for AI-generated motion graphics.
  • A piece that starts strong but doesn't stick the ending reads as unfinished even if the middle execution was the best of the group — pacing to a clean close matters as much as the opening hook.
07Test 5: Knowledge work
  • Prompting a model to 'do serious research, real current source figures, don't just read headlines' before building a deliverable produced noticeably deeper, more defensible output than a vague design ask.
  • Two independently-run models converging on the same core financial thesis is a useful cross-check that an AI-generated analysis wasn't a one-off hallucination.
  • Small UX details — where a hover tooltip lands relative to the element, whether it overlaps a card — were what separated a 'well done' professional deliverable from a 'good enough' one.
08Verdict
  • A model that wins on quality and comes in at half the price of the incumbent removes the usual quality-vs-cost tradeoff — when that happens, cost stops being a reason to stick with the pricier tool.
  • Total generation cost can end up similar between two models even when one is clearly better, because the stronger model may take more turns checking and correcting its own work rather than being 'more expensive' outright.
Glossary

Terms worth knowing.

Frontier model
One of the current top-tier AI models, typically compared against its close competitors on capability, cost, and speed.
Agentic coding
An AI model performing multi-step coding work somewhat autonomously — writing, running, and fixing code across several turns rather than just answering a single question.
Blind compare
A test format where the evaluator ranks outputs without knowing which model produced each one, so brand reputation can't bias the judgment.
Input / output token cost
What an AI model charges per million tokens: input tokens are what you send it, output tokens are what it generates back — output is almost always priced higher.
Ray marching / ray tracing
Rendering techniques that simulate how light travels through and interacts with a 3D scene, used to produce realistic lighting, reflections, and shadows.
Resources

Things they pointed at.

Quotables

Lines you could clip.

02:19
Anthropic is just losing money. They're basically telling you there's no need to use Fable anymore.
spicy, speculative business-strategy take on a same-lab model cannibalizing its own flagshipTikTok hook↗ Tweet quote
08:40
Look at these f's. Dead giveaway. That that is AI. Just that curly f.
quick, concrete AI-tell viewers can go check for themselvesIG reel cold open↗ Tweet quote
27:06
Opus five was across the board better at three d, creative web design, game development, motion graphics, and knowledge work... for half the cost of Fable five.
the full verdict delivered in one breathnewsletter pull-quote↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphor
00:00Alright. With the model releases that keep on coming, Opus five is officially dropped. And according to Anthropoc's own benchmarks, it is Fable five intelligence for half the cost.
00:08Of course, we're not gonna actually trust those benchmarks. Instead, we're gonna see how true that actually is. So in this video, we're gonna run Opus five through the ringer across a variety of tests.
00:16Web design, three d and simulation, game development, motion graphics, and finally knowledge work. We'll compare it to its predecessor Opus four eight and of course to Fable five.
00:25So by the end of this video, you'll see how Opus five lives up to Fable five, how good it actually is, and whether it's even necessary to use Fable anymore. So not wasting any more time, let's get right into the builds. Kicking it off with our usual fan out of agents.
00:38We'll have five tests running across three models. So let's get this thing started. So we've got each of these builds running in their separate I term windows.
00:46I will explain each of these builds here in a second along with the prompts that we're using. But first, let's take a look at these benchmarks. I promise this will be quick, but these were pretty surprising, so I wanted to call these out.
00:57So if we actually look at these numbers, Opus five is crushing Fable five across the board. Like, it is not even close. There's a 10% jump in agentic terminal coding.
01:05We've also got five six soul here, you can see just how better Opus five is versus five six soul, at least according to the benchmarks. A noticeable jump in knowledge work. So GDP val, where they have blind testers actually rate knowledge work across a bunch of different departments.
01:17We have novel problem solving. I'm not really familiar with that one. Looks like Fable five had been in run there.
01:20Look at this jump. 30% compared to seven point eight percent five six SOL. And then Agentic Search, a nice little jump.
01:27Humanities Last Exam, not really much of a difference there, but it's just as good as Fable. Computer use, a 4% jump. Agentic coding, not quite as good, but close enough.
01:35And then automations, legal, health, biology, it's just kind of split between Opus and Fable. And some of these are just mythos. Like, this mythos number is just for health, which we don't even have access to.
01:44So this thing across the board is a better knowledge worker, agentic terminal coder, agentic searcher, better at computer use, so actually taking over your computer, navigating websites, etcetera, etcetera. And most surprising of all, it's the same cost, if we go down on the cost here, as Opus four 8, $5 per million input and $25 per million Whereas Fable let me just check this.
02:05Fable five, I believe, is $10 per million input and $50 per million output. Yeah. 10 and 50.
02:10So at half the cost of Fable, Opus is better across the board, which makes absolutely no sense that Anthropic is cannibalizing their own model unless this has to be the reason we have a next version of Fable coming out, like in the next one to two weeks. Because Anthropic is just losing money.
02:25They're basically telling you there's no need to use Fable anymore. So it's very confusing, but I mean, I guess that's just how the models progress, but it is it is quite perplexing. And then they go on to just show different graphs here.
02:36You can see Opus five on Frontier Bench, agentic coding by effort level. So low, medium, high, extra high, and max. It's just look at that gap.
02:45Just noticeably better and cheaper. Only exception on UltraCode or UltraMax, whatever it's called, on Fable, it's like a little bit better for for cursor bench. And then for some reason, it gets worse on Max for agent decoding.
02:56So it just doesn't seem to make any sense. The only thing I can think of is maybe cybersecurity, and they have this section here, misaligned behavior, and then they have this finding vulnerabilities and exploits in open source code, they don't even mention Fable.
03:07They show exploitation success with Mythos five, where it's it's much higher than Opus five, but we don't even have access to Mythos five. Remember that Mythos five is a more advanced version of Fable five.
03:16That's only accessible to select organizations. They don't even call out Fable five. So it's very confusing.
03:20And then my final grip with all of this, and I promise we'll get back to the builds, but if Fable five that came out a month, month and a half ago was so freaking dangerous that it had to be banned for multiple weeks, why has Opus five, which is an even better model than Fable five across the board, at least according to the benchmarks, why was there no stopping Opus five?
03:38Maybe that happened in the background or we're not aware of it? Maybe Dario finally shut his mouth and they were able to sneak this one by the government? Anyway, that's all the benchmarks.
03:45Uh, they only have one demo. They talk about CAD, it being really really good at CAD, which, uh, would be fun if you're, you know, if you're three d printing or whatever, just working in Blender.
03:53And they give a couple examples here, but our examples are gonna be much more fun. So rant over, let's get back to the builds and check-in on those. So these are coming in in piecemeal, but let's take a look at a couple of these outputs first.
04:03And what I wanted to do with these is give these models complete creative freedom, partly because I'm just running out of ideas, but also because I wanted to see how they interpret a broad prompt like this and how creative they can really get. So this is just a test. We'll see how well these outputs do.
04:16We might need to go back to more specific prompts next time. But for now, let's take a look at the prompt first. So this is what it is.
04:23Build one website that demonstrates the absolute ceiling of your web design capability, your taste, your artistic flavor, your technical range. You choose what the site is for. That decision is part of what's being evaluated.
04:34Context you should know. This is a video, you know, by I'm trying to I'm trying to raise the stakes for the model, trick it a little bit, think that it really is under a lot of pressure to do good. So I say, this is going into a video comparing you against other frontier models shown to a larger audience, your work will be credited to you by name.
04:48Every model gets this exact same prompt, yada yada. Go nuts, show people what you're capable of. And the requirements, you pick the concept, and then I just had a separate agent write this up where I just said, list out all of the techniques.
04:58So I just get into all these different advanced visual techniques. High quality three d, WebGL, ray marching, I don't I don't know half of these. Post processing, bloom, chromatic aberration, particle systems, flow fields, generative procedural art, fluid simulation.
05:11It goes on and on and on. And I say you're not limited to these, they're just some examples of what you can do. Don't do all of them, but you can do, you know, some of them just so it triggers the model's thinking and generating these.
05:20And then I just continue to go on. You have free rein. This should just be a single self contained HTML file.
05:25And so let's see what these came up with. By the way too, these are gonna be available to you in a blog post below. You'll be able to copy the prompts as well as look at each of these generations and live links.
05:33And we'll start with column a and just like what we've done previously. It'll be a complete blind test. I will rank these, and then I'll reveal each model at the end.
05:39So I'm already seeing some interesting stuff here. They all went a space theme, which is kind of annoying. Maybe because I kept saying floating particles and all sorts of, you know, flow fields, fluid simulation, maybe it just immediately thought space, but I was hoping the actual creative idea would be different.
05:53Having said that, let's take a look at these generations as a whole. Okay. Aeolius, Bureau of Exoplanetary Weather.
06:01Alright. I like the particles in the background. Uh, I really like this planet here.
06:05The actual cursor movement is really funky. It's kind of hard to move this cursor. That's very low.
06:11But let's scroll here. So we're scrolling, we're scrolling. Alright.
06:15The planet goes away. That's actually that's actually a good touch. A lot of times with these kind of websites, it it can never get the actual smooth scroll down where it just kind of gradually goes.
06:24Usually, it'll just glit really fast or won't actually get the layering correct. This one noticeably got all of that, which I I don't think I've seen, at least in a one shot. So I'm starting to think this is Opus five, but who knows?
06:34Okay. Then we have some text. Nobody will ever stand there.
06:37Okay. I'm not gonna read any of this text. Oh, nice little globe.
06:40I like I like how it reappears, but smaller. That's nice. This mouse movement is really bugging me.
06:44It's really hard to move, but we'll keep scrolling. Okay. I like the little graphs.
06:47A little hard to read, but that's pretty cool. And we're scrolling. I like the particles moving.
06:51Oh, I like how it's like, um, particles feel like they're in like a globe, but they're not just flat. That's actually a really nice touch. And then we keep scrolling or keep scrolling.
06:59Oh, nice. And this planet appears again. Very cool.
07:02Alright. It's getting darker. Yeah.
07:04Halley cone. I wish I knew anything about space. I don't know if it just made all this stuff up or if this actually is something.
07:10And we're scrolling. And we're scrolling. Oh, okay.
07:13Alright. We're getting into more planets. So we scroll.
07:16It's a little glitchy here or maybe not. Okay. No.
07:18No. That's actually very smooth. Wow.
07:20Okay. I've yet to see that. I've yet to see a smooth transition like that.
07:23Little little telltale signs this might be an advanced model. Oh, nice little touch there. I like the spin of the globe.
07:29Text is kinda weird. That's funky. Oh, but I like this a lot.
07:32Yeah. Okay. Particles are kinda weird.
07:34I don't know if they're supposed to be stars. Okay. And you can actually hover over that.
07:37And we're scrolling. Okay. It's getting oh, the this little atmosphere thing.
07:41Oh, very cool. Atmosphere departing. That's a nice touch.
07:44Alright. And then we scroll and then this is like the surface maybe? Not sure.
07:49Can you click these? K. Yep.
07:50Man, if this mouse was just a little bit more navigable, this would be really killer. I I like the type too. I like the placement of all the text.
07:58It feels very natural. Oh, a nice little before and after. Alright.
08:01This is nice. This is nice. And we have the sound of the current world.
08:04Does this do anything? Oh, yeah. It's making noise.
08:06I don't know if you can hear that, but I want it to stop now, but it's not stopping. Okay. And then it just kinda ends on this footer here.
08:12But, yeah, this was really well done. Usually, you tell these models to do something, especially this creative, this intense of a build, it'll just kinda get lazy and it'll build out like a couple sections. Here, this is this feels very comprehensive.
08:23Then we also have this light mode here. Let's see. Very hard to click.
08:27And the light mode looks really good. I like this. I like these different shades of white.
08:31Yeah. Well done. Jeez.
08:33Alright. So that is one generation. Don't know which model that is, but I think I have some sneaking suspicions what it is.
08:39Alright. So we have another space thing here, Sagittarius a, the supermassive black hole, the center of the Milky Way.
08:46Alright. Nice. Yeah.
08:47I like the I like the difference different sections here. We've got Sagittarius a in the background. Okay.
08:51Scrolling. Scrolling. Got some text here.
08:53Obviously, I'm not gonna read any of that. It it is a little weird like this. I feel like it's just a little too harsh of a line.
08:58I don't know if I'm being nitpicky like this. I don't like this harsh line here. Seems kinda weird.
09:02It would be it would look much better if was just more of a natural gradient. And this is it's kinda weird, but I get that you have to make the text readable so you have this kind of like black drop shadow background. And we're scrolling.
09:12Disc is not the black hole. It is everything. Black hole is about to take.
09:16Oh, cool. Are we getting closer to the black hole? Oh, we are.
09:18Yep. We are. We're doing it.
09:20We're getting closer. Okay. Nice little SVG there.
09:23Nothing here was photographed. Okay. Yeah.
09:24This is what I'm talking about. Like, it only has one three d generation, which looks cool. It's a it's a it executed this well.
09:30But that last one, I mean, we had multiple three d generations. We had all sorts of complexities going on. So that's the second one, and then here's the third.
09:36Lumen, nice little little wave here. Boom. That's kinda cool.
09:39Okay. Scrolling down. I don't see any three d work.
09:41I just see background particles. Alright. We're scrolling.
09:44We're scrolling. Look at these f's. Dead giveaway.
09:46That that is AI. Just that curly f. In which cloud becomes a star?
09:51Okay. Hydrogen given a reason to glow. K.
09:54We have these little blobs. Alright. It's getting more pronounced.
09:57That's kinda cool. Oh, yeah. This is actually really cool.
09:59Oh, scrolling in. Okay. And then it just kinda keeps going like this.
10:02This thing looks slick. I just wanna play with this this little ball here. Alright.
10:06And we're scrolling. Oh, nice. Very cool.
10:08Very cool. I like that chromatic aberration. Okay.
10:11And we're scrolling and we're scrolling. Nice. Aura Borealis.
10:14That's pretty cool. Okay. And then as you scroll, it changes.
10:17That's that's very nice too. This is cool as well. Jeez.
10:20Yeah. See, is what I'm talking about. This is this is another comprehensive build.
10:24But those are the three of them. I'm actually impressed with all of these. I thought this was gonna be fable and this one's this one was gonna be worse just with the header, but it started to get much better and had a bunch of different generations.
10:33So I'm gonna guess this is Fable. Yep. And then this has to be for it.
10:37Yep. Opus five. Wow.
10:39Okay. There we go. There we go.
10:40That is a good test. Opus five definitely did better there. I would say unquestionably better.
10:45Tang. Okay. Anthropic may not have been lying.
10:48So that is the first test. Now the second one is something somewhat similar, but I just wanted to have each of these models focus on three d. So I'm saying build one real time three d experience that demonstrates the absolute ceiling of your three d and simulation capability.
10:59Again, the same context you should know. Your work will be credited to you by name. No pressure.
11:03You pick the concept. Technical ambition is the point. It has to look genuinely realistic, generally art directed.
11:09And then I give all of these different terms, say including but not limited to all of these. You have free reign to do whatever you want, self contained HTML, etcetera, etcetera.
11:18So let's take a look at this first column here and more space theme. I should say don't do space theme next time, but here we are. So okay.
11:26I think it went on the same kind of idea with this, like, black hole. So we have this scroll in, we have this scroll out. Okay.
11:33That's cool rotation. And then we can change these different settings here. Auto orbit, regular.
11:38We have bloom. We have doppler beaming, disk luminosity, drag orbit, what I just did.
11:45And okay. I guess that's it. So sure, yeah, not not bad.
11:50Next one, gravitas, more space name. Looks exactly the same. Like a little bit more advanced, but the same idea.
11:55But we've got this, according to this title here, real time general relativistic ray tracer, and we can rotate it like this. We can spin it around.
12:04Okay. Oh, the spin around actually looks really well. That's that's good.
12:08Got some nice easing in the actual animation too. Feels like very natural. You can really spin this thing.
12:14Cool. That's that's kind of fun. And then like changes the background too and then it like comes back.
12:18And then what else do we have? Space time, a creche accretion rate. Any astronomists in the comments, please explain what's going on here.
12:25And then we have disk temperature. That is kind of cool. Can just keep adjusting this.
12:31And then we have time flow. Let's bring disk temperature down in the middle. Time flow doesn't seem to be doing much.
12:38And then bloom, it's just gonna get brighter and darker. Yeah. And then lensing, beaming.
12:42So we have more options here. Okay. So it's is is that a planet in the middle?
12:46Sorry for my space ignorance. Okay. That one, yeah, not bad.
12:50It's a little boring. Okay. At least it went ocean themed, the other option.
12:54I should say no space themed, no ocean themed the next time I do these tests. So okay. Oh, we've got some sound effects.
13:00You probably can't hear that, but it's like dinging like a bell and then oh, that's kinda cool. Alright. So you just like click these jellyfish and then they just kinda like, I don't know, have some sort of impact.
13:09Click that. I think we can only click the jellyfish. Drag, look around w a s d swim.
13:13Oh, we we can swim. Okay. Alright.
13:16Pretty cool. Oh, we can go oh, no. We can't go above the water.
13:19We can swim with these fish. That is kind of cool. Bonus points for not space idea.
13:23And, yeah, you can really navigate this this around. Then you're just kind of like clicking the jellyfish. I don't know why.
13:28Can I click the fish? Don't think so. Okay.
13:31How do I turn around? There we go. Okay.
13:33It's just like making impacts. Shift surge, scroll lens, shift surge.
13:38I don't know what that means. Strike a bell bottom lens q e rise. Okay.
13:42Yeah. Uh, bonus points for the creativity. Looks pretty cool how it rendered that.
13:46So I'm gonna say this number one, this number two, this number three. This gotta be Opus 48. Yep.
13:51I don't know. Uh, let's see. Is this Fable?
13:52Yeah. Fable five. There we go.
13:53Opus five yet again. More creative and a better render too. And like more more little more little features here and there.
14:01Alright. Freaking. Wow.
14:03Opus five. Look at this. Let's quickly talk about cost too for each of these.
14:06So interestingly, Opus five maybe because it just needed to take more turns, but it came out roughly the same cost as Fable five. Again, this could just be because it is better, it's more thorough at checking its own work, and maybe it identified some problems, had to fix some things, and do some more turns. But, yeah, it's very close to Fable five, whereas Opus four eight on the left, $5.22.
14:28So even though they are the same costs, Opus five might be ultimately more expensive because it's just gonna be more thorough. Could be an anomaly, but that's my guess. And then we've also got costs on these.
14:38So reminder, and I forgot to rank these two, but Opus is definitely number one, Fable five number two, Opus four eight number three, and yeah. Okay. Now we can start to see the cost difference.
14:47So Fable five was $53, Opus four eight was $27, and Opus five was a little bit more expensive at 38, but still cheaper than Opus and much better of a generation.
14:57Next up, we've got build a playable game. Same kind of idea with this prompt where I'm leaving it up to the models for them to create their own games. So I'm saying build one playable browser game that demonstrates the absolute ceiling of your game development capability.
15:08You choose the game, any genre, any style, original, has to be genuinely fun within thirty seconds. No tutorial wall, no menu maze. The graphics bar is high.
15:17I'm sick of these bad graphic games that all these models are generating, so I really wanna push that. Who knows if that will actually work? But I go on to say all of these requirements for the graphics and then depth as well, and essentially give it free rein.
15:29First up, we've got Volt, Heart Overcharge or Die. Move to begin.
15:34So okay. First impressions. Oh, I do I do like these graphics.
15:37They're a little bit better. I I actually like how it's not a first person shooter, and I like this again, this space theme is just nonstop. How does the actual shooting work?
15:44Oh, the shooting is pretty good. Actual impact's pretty good. Jeez.
15:47I can't get away from these things. Jeez. Alright.
15:49Let's try it one more time. But, yeah, mean, overall, I like it. Even though it's just kind of a shooter, still relatively entertaining.
15:54Good movement. Graphics aren't bad. Jeez.
15:57Okay. That's enough of that. I think we get the idea.
15:59Second game, we have Neon Swarm. Here we go. I should have said no Neon either.
16:03Seen this way too many times, but okay. Alright. Come on.
16:06Could've done better. Oh, good moving. Good moving.
16:08Alright. Fair enough. Fair enough.
16:09What is it? It's like the what's it called? Asteroid?
16:11It's like an asteroid game. I like the movement of this little little ship thing. Cool kind of little raft in the background here.
16:17Very smooth. Okay. You get the idea.
16:19Finally, we have sunk. Little bit different of a concept. Not great waves, not great boat.
16:26Weak on the graphics, but it is at least something different. Okay. And then what is happening?
16:30Okay. What is going on? I can't turn I can't turn around.
16:33Oh, there we go. Okay. I can turn around.
16:35Oh, there's a boat. Okay. Alright.
16:36Yeah. That's kinda fun. Okay.
16:37How do I I don't know why d going the other way, but okay. Yeah. They make these games really hard when I say just make it fun.
16:44Smuggler scouts. Alright. This one might be the most entertaining.
16:48Graphics are not great though. Nice. I can tell, like, these I can tell these graphics aren't good because of the generations I actually just did in a Kimi K versus Fable versus five six Soul video where we had it generate waves and boat movement and stuff like that.
17:03So these waves are pretty weak because I've seen what the models are capable of. But that is the third playable game. Honestly, these were kinda close.
17:10I would I'm gonna say this is number three because the graphics aren't great even if it was a little bit more fun. I understand too it's a little bit more difficult than this, which is just so generic. Okay.
17:19Yeah. No. I'd have to say this is three.
17:20This has to be two because it at least tried. And then even though I mean, these were again, it's kind of cheating because it didn't didn't have to be great with the graphics, but I'm still gonna say that's number one. Let's see.
17:29Opus four eight. There we go. And then number two, they have a five.
17:32Wow. Again, again, again. Jeez.
17:34Every single one of these. I mean, yeah, you can see the attention to detail that Opus five is putting in here. It, uh, it really is quite impressive.
17:41So there we have it yet again. Opus five winning out. Okay.
17:44And for these final two generations, we're gonna make this a little bit more realistic, something you might actually be doing in your day to day. The first is gonna be motion graphics. So the prompt is same kind of idea, build one motion graphics piece that demonstrates the absolute ceiling of your motion design capability.
17:58You choose the subject and the concept, and then we have just some requirements here, a title sequence, a product spot, data story explainer, etcetera etcetera. It plays automatically on load, runs roughly fifteen to thirty seconds, choreography, so I just drop in all of these different motion graphic terms, see how well it does in that, and that is basically the prompt.
18:16So let's take a look here at the motion graphic generations. First up, we've got inferno. Uh, okay.
18:21So it's just like text animations. So let me try this one more time. So inferno.
18:24Midway upon the journey of her life. I find myself okay. More text animations.
18:31More text. Okay. Woah.
18:32Alright. Kinda crazy. And then alright.
18:36Alright. It's getting pretty ridiculous, but this isn't bad.
18:45Inferno, The Descent in Nine Circles from the poem by Dante.
18:52Okay. That's the first motion graphic. Let's check out the second.
18:55Got some nice grain here. Okay. Still text animations.
19:01Okay. Alright. SVG.
19:02Oh, nice. Nice. I like the SVG.
19:05I like the movement. Okay. Nice and clean.
19:10Chart there. Different types of text animations. That's pretty good.
19:15Then we have, like, a little bit of, again, chromatic aberration in the actual text generation.
19:23So that's that's pretty good. Alright. Finally, signal and noise, a study in emergence.
19:30Everyday signal becomes as begins as noise. Particles oh, okay.
19:36I like it. I like it. Looks like a little thing.
19:41Find the signal. Okay. Is that it?
19:46Signal slash noise.
19:50Okay. I had a good start, but didn't didn't end great. Alright.
19:54That's that was number three. I'd have to set this number one. This number two.
19:58Let's see. Opus? Yeah.
19:59Opus four eight. Jeez. Fable again.
20:01Wow. It is noticeably different. Opus five, quite quite a bit better.
20:06Let's look at it one more time. That was a really nice touch. That that was why that that was why I won this year.
20:10It looks pretty good too. Dead reckoning. Oh, I just realized.
20:13Gosh darn it. I need to not have them say what model they're Eclatto is. The other ones have that too?
20:20Yeah. The other ones did. Gosh darn it.
20:21Alright. Well, next time, I'll make sure it doesn't do that. Okay.
20:24So those are two more tests. We're gonna look at the final test here in a second, but but let's take a look at costs for the playable game first. So playable game costs.
20:31Let just show these again. Opus $4.08, $31. And then Opus $516.
20:36Wow. And then so I'm I'm not sure what happened with Opus $4.08. Just again, a bunch bunch of terms.
20:41Yeah. You can see the input token's 20,676 output versus Fable, much more efficient.
20:48And then OPUS five, even more efficient than that, a much cheaper cost. So again, OPUS 5, $16. Fable, 33.
20:54OPUS $4.08, $31. Wow. Okay.
20:56And then motion graphics, OPUS $4.08, $33 again. Jeez. And then Fable $520.
21:03And then OPUS five, more efficient and better output at $17. Our fifth and final test, we're getting a little bit more into knowledge work here with a recommendation on whether or not to buy SpaceX stock. So here's the prompt to your private wealth adviser preparing an investment committee deck for a high net worth client.
21:18They've asked you one question, should I buy SpaceX? A lot of space themes in this video. Do a serious amount of research before you build anything.
21:24Real current source figures don't just read headlines. So what I'm saying is I want you to actually dig deep in this. I want you to do a ton of research.
21:30Look at every business unit. Look at how they make money. Do some forecasting.
21:34Look at the cost structure of the unit economics. Like, does this actually make sense? Prepare it like you would like a private wealth manager.
21:41And I just go on and on. I say build a real forecast, visible, defensible assumptions. And then I also get into prompting around the actual deck design.
21:48I say it has to look beautiful. It has to look super professional, and I want real data visualization, interactive bar charts, comparisons, ranges, all of that, because I'm always pretty underwhelmed when it comes to deck design in these visualizations.
22:00They always seem to miss something or just doesn't look great usually, so I wanted to test that as well. And that is essentially the prompt. So let's take a look at those generations as well.
22:09First up, should we buy SpaceX? Okay. Hey.
22:11Nice little rocket. Tidal slide. Very good tidal slide.
22:14I like it. Okay. And then see, little things here and there like this.
22:16Like, the sources, it's okay. I can't really navigate.
22:20Okay. Like, the sources, it should just be it shouldn't be overlapping into the cards. I also don't like these cards.
22:25Little AI smelling. Business is real. Starlink, 11,000,000,000 in revenue.
22:3027 bill contracted AI compute. The price is not.
22:35Okay. Interesting. The plan, do not buy at $1.14.
22:37Our probability weighted 23 2030 value is 101 a share. Really? So they're saying wait until it goes on 101 a share.
22:44I don't know if that that will happen. And we continue and then we say you're buying three companies, Starlink, Space Launch, and then AI Compute.
22:51And we've got some nice stats here. This is actually yeah. Put some pretty good visualizations.
22:54I like the way it's organized. Let's see if these okay. Oh, look at this.
22:58Nice. Nice. A nice little highlight there.
23:00I don't know how you could do this like in a regular PowerPoint presentation, but I I guess you could just present it in its real HTML if you're actually bringing this in front of someone. Yeah. This is a nice touch.
23:09Okay. The only thing, just the little border there, which is kind of annoying. I was spanning over the cards.
23:15Anyway, eighteen months. Oh, this is interactive too. Look at this.
23:18Very nice. 49% drawdown.
23:20Okay. I I didn't realize it tanked that much. Valuation.
23:23Okay. Just some more history. Okay.
23:26Business units. Nice. This is yeah.
23:28This is well done. And then some forecasting. Capital treadmill.
23:32CapEx quadrupled in a year. AI lane now outspends the rockets three to one. Wow.
23:37I did not realize that. That's that's pretty interesting. You'd think rockets would be more expensive.
23:41The AI segment, 27,000,000,000 of contracted revenue from three customers. Uh, look at Anthropic, 15,000,000,000.
23:48That will probably go up too. The growth is real. The quality is not.
23:51Two customers are 94% of it. Both can walk in ninety days. Yeah.
23:55But Anthropic's locked in. They're not going they're not going anywhere else. Starship is the load bearing wall.
23:59Starlink v three. Okay. Cool.
24:01Yeah. I mean, this is, like, this is very interactive. This this looks nice and professional.
24:04Like, even though, like, little little planets in the background, nice touch. This annoying footer should have caught that. Okay.
24:10I'm kinda getting bored, but these are wow. Yeah. We've got even more interactive charts.
24:14We've got this. That looks great. Okay.
24:16Yeah. This is well done. This is well done.
24:18Next up, SpaceX Blaration Technologies Corporation. No rocket on the title slide.
24:23Dinged recommendation, do not initiate. Okay.
24:26Why not? Alright. Same kind of a same kind of design.
24:29Uh, does this interact? Oh, yeah. This is I actually like this one better because it shows up right here.
24:33The other one was, like, showing up. You're hovering over something. It was showing way at the bottom.
24:37Little nitpicky stuff. I gotta be super nitpicky with all this Otherwise, I don't know what else I would say. Uh, okay.
24:42Nice little gradient. Oh, very interactive. Just a lot of text on the screen.
24:46Obviously, I'm not gonna read. There's not a rock company. It's three companies with opposite financial signatures.
24:51Yep. Exactly like the other model said. Looks pretty good.
24:54Okay. That's that's not bad. The launch monopoly is a cost center that happens to invoice customers.
25:00Yeah. We've got all the Starship. That's a nice little chart too.
25:04It's actually side note, could be a good prompt. I mean, looking at both models, I don't know which which is which, but giving a prompt like that where you're getting very specific, saying it has to be beautiful, seems to have given a much better output, a better output than I've seen in any sort of deck design.
25:17So it could be worth you trying that yourself if your work involves designing PowerPoints. The most underrated asset is the is the defense backlog. Okay.
25:24Yeah. I don't know why it's spelling defense like that. With a defense with a c.
25:29While the market argues about Mars, SpaceX has quietly become prime contractor of one of the largest defense programs in generation. K. The the competitive mode is real.
25:36Yep. It is not also not the thing that's worth the stock. We've got some more charts here.
25:40Oh, that's a nice chart. Bearable base. Very cool.
25:43Get some EBITDA. Jeez. Both these models did a pretty good job.
25:46Okay. We got the same kind of ideas. Yes.
25:48So it just added this, like, whatever you call this thing. Four strong arguments, a ton of text, very hard to read. And we keep going.
25:54We keep going. Alright. But I don't know.
25:55That those are pretty close. Okay. Oh, nice.
25:57Got the rocket in there. Should you buy SpaceX? Okay.
25:59Let's go through here. Alright. Very similar designs.
26:03This one also interactive. SpaceX is not the company, its ticker implies. You get some bar charts here.
26:09Okay. We break down the business. Uh, sorry.
26:11This one's really boring. I thought there was gonna be a clear difference, but it really is not. I mean, this one's this one's up maybe like just in like the the actual designs, a little bit more basic.
26:20Yeah. Like, you know, less kind of gradients. Oh, this this is a new chart here, risk matrix.
26:27Hey, I did give away these these things, but let's keep going. Alright.
26:33Alright. You get the picture. So that one was I'd say that one's the worst.
26:37These ones are really close. Let me look quickly one more time. Let's see if I can if I can get it.
26:42Let's look at this one again. Yeah. I mean, this one yeah.
26:45That was oh, yeah. That that sort of annoyed me. This little thing over like, should just be hovering when you hover over something, it should just be directly above it.
26:51Didn't love that. I'm gonna say this is two for that reason. Number one, wow.
26:56Opus five across the board. That is crazy. I guess every single one of these.
27:01And then Fable and then Opus four eight. Let's just quickly look at costs. So we have Fable, $35.
27:06Opus, 27. And Opus four eight, 23. Well, there we have it.
27:10So I guess Anthropic was not lying. Opus five was across the board better at three d, creative web design, game development, motion graphics, and knowledge work. It won every single one of those tests, at least according to my own judgment, for half the cost of Fable five and sometimes even cheaper than Opus four eight.
27:24So I guess we stop using Fable five now? It just feels so weird to be doing that. But I I mean, I guess that's what these tests revealed.
27:31Anyway, appreciate you watching as always. Let me know what you think of all of this in the comments, plus any ideas you have for future tests. And I guess we'll see you back here when Fable 5.1 inevitably drops, uh, next week.
The Hook

The bait, then the rug-pull.

A blind five-round test — web design, 3D, a playable game, motion graphics, and an investment deck — pits Opus 5 against Fable 5 and the older Opus 4.8, with the model names hidden from the reviewer until every rank is locked in.

Frameworks

Named ideas worth stealing.

04:21concept

Blind compare — one task at a time (SimmonsBench)

A custom tool that runs the identical prompt across multiple models, hides which output came from which model, lets the reviewer rank blind, then reveals identities only after ranks are locked.

Steal forAny product or tool comparison — AI models, vendors, templates — where already knowing the brand would bias the judgment.
CTA Breakdown

How they asked for the click.

VERBAL ASK
27:10newsletter
Let me know what you think of all of this in the comments, plus any ideas you have for future tests.

Soft — no hard pitch spoken on camera. The real CTAs (bootcamp waitlist, newsletter, blog post with every prompt and generation) live entirely in the description, not narrated.

Storyboard

Visual structure at a glance.

open
hookopen00:00
benchmarks
valuebenchmarks00:51
web design test
valueweb design test04:21
3D/sim test
value3D/sim test10:48
game test
valuegame test14:57
motion graphics test
valuemotion graphics test17:44
knowledge work test
valueknowledge work test21:07
verdict
ctaverdict27:10
Frame Gallery

Visual moments.

Watch next

More from this channel + related breakdowns.

36:06
Pat Simmons · Review

Kimi K3 Is Here! (Better Than Opus 4.8?)

Ten identical builds, five models, blind-ranked before the reveal — a real-world stress test of Moonshot AI's new open-source model against GPT-5.6 Sol, Opus 4.8, GLM 5.2, and its own predecessor.

July 17th
40:03
Pat Simmons · Review

GPT-5.6 Sol: No-Hype Full Review & Testing

A blind, four-way bake-off — GPT-5.6 Sol against Fable, Opus 4.8, and GPT-5.5 — across ten builds and knowledge-work tasks, scored one task at a time without knowing which model made what.

July 10th
19:10
Pat Simmons · Tutorial

Clone a $1.2B App in 19 Minutes

An 8-step agentic pipeline that takes you from naive AI slop to a pixel-near Linear replica, deployed to Vercel with an MCP server, in under 20 minutes.

June 8th
Chat about this