Modern Creator
WorldofAI · YouTube

Claude Opus 5 Is THE GREATEST AI Model EVER?! Beats Fable & CHEAPER! (Fully Tested)

A hands-on benchmark-and-demo breakdown of Claude Opus 5's launch, stacked against Fable 5, GPT-5.6 Sol, and Kimi K3 across reasoning scores, cost, and a dozen live-generated apps and games.

Posted
today
Duration
Format
Review
hype
Views
1.2K
73 likes
Part of the collectionThe Fable 5 PlaybookAll 45 Fable 5 breakdowns, synthesized into one page.
Read the playbook
Part of the collectionThe Claude Opus 5 PlaybookEvery Opus 5 breakdown, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

Claude Opus 5 narrowly trails or matches Fable 5 on major intelligence benchmarks while costing roughly half as much per task, making it the creator's pick for complex reasoning and coding even though he'd still reach for other models on frontend-heavy or daily-driver work.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You're deciding which frontier model to route Claude Code, an agent, or a coding workflow through and care about cost-per-task, not just raw benchmark rank.
  • You follow AI model launches and want the actual benchmark numbers (Artificial Analysis, ARC-AGI-3, World of AI Bench) rather than just the marketing claim.
  • You're curious what a frontier model can one-shot right now — full games, OS clones, physics sims — as a gut check on where AI-assisted generation has gotten to.
SKIP IF…
  • You want a rigorous, independently-reproduced benchmark study — this is one creator's own benchmark tool plus a handful of anecdotal demo runs, not a controlled evaluation.
  • You need current, verified pricing and availability — model lineups and pricing change fast and this is a single launch-day snapshot.
TL;DR

The full version, fast.

Claude Opus 5 launched positioned as frontier-class but roughly half the price of Claude Fable 5. Across the creator's own World of AI Bench, Artificial Analysis's Intelligence Index (61 vs Fable 5's 60), and ARC-AGI-3 (30.2% vs the prior 7.8% record), Opus 5 lands at or near the top while consistently costing less per task — a 3D city simulation that cost Fable 5 $9.60 cost Opus 5 $4.20 for a comparable result. Hands-on, it one-shot a playable COD Zombies-style game, a macOS clone, two Minecraft clones, a Fall Guys clone, and a browser black hole simulator, with the video highlighting strong reasoning, debugging, and 3D/spatial generation. The counterpoint: Kimi K3 still wins on frontend polish and speed for less money, and the creator personally sticks with GPT-5.6 Sol as a daily driver for its lower latency and higher usage limits — Opus 5's real edge is complex reasoning and engineering work at a lower cost, not being uniformly the best model at everything.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0001:07

01 · Introduction

Cold open teasing the COD Zombies clone, then the creator says yesterday's leak-based take that Opus 5 was 'on Fable 5's level' undersold it after seeing Anthropic's official benchmarks.

01:0701:40

02 · BEST Reasoning Model

On the World of AI Bench leaderboard, Opus 5 ranks above GPT-5.6 Sol and just behind Fable 5 overall, but scores highest of all as the #1 reasoning model at 83.7 on the 'x high' reasoning test.

01:4002:43

03 · Benchmarks

Opus 5 is called the new default on Claude Max and strongest model on Claude Pro, with FrontierBench and knowledge-work scores shown beating Fable 5, Opus 4.8, and GPT-5.6 Sol, and roughly double Opus 4.8's FrontierBench score at lower cost per task.

02:4303:26

04 · Best Model Ever?

Artificial Analysis ranks Opus 5 as the most intelligent model it has ever tested, scoring 61 on its Intelligence Index versus Fable 5's 60 and GPT-5.6 Sol's 59, at about 26% less cost per task than Fable 5.

03:2603:49

05 · ARC-AGI-3

Opus 5 scores 30.2% on ARC-AGI-3, crushing the prior 7.8% record set by GPT-5.6 Sol, with ArcPrize noting genuinely new problem-solving behavior on environments no previous model solved.

03:4905:05

06 · Which Model To Use

The creator's personal model picks: Opus 5 with Max reasoning inside Claude Code for complex reasoning/debugging/engineering, Kimi K3 for front-end-heavy web work at a cheaper price, and GPT-5.6 Sol inside Codex as his own daily driver for lower latency and higher usage limits.

05:0505:37

07 · How To Use

Brief plug for using the World of AI benchmark tool to evaluate models against your own prompts, framed as more useful than relying on any single reviewer's take.

05:3706:20

08 · Apollo Simulation

Opus 5 procedurally generates an Apollo-mission space simulation from scratch — code, animation, and soundtrack, with no premade assets or external models used.

06:2009:02

09 · COD Zombies Clone

The headline demo: a fully playable round-based zombies survival clone with escalating waves, purchasable weapons, a health-doubling 'juggernaut' perk machine, openable doors/areas, and a slide-cancel movement mechanic — built autonomously with no starter codebase.

09:0211:32

10 · MacOS Demo

A polished macOS clone with a working top bar, reactive toolbar, click sound effects, an AI 'Siri' the creator can query, notifications, a photos app, a music player, and a separately-prompted Minecraft clone launched from within it.

11:3212:19

11 · Opus 5 vs Fable 5

Head-to-head 3D city simulation of a UK skyline: Opus 5 completed it for $4.20 versus Fable 5's $9.60 — under half the cost — with the creator calling Opus 5's version more lively and interactive (moving water, cars).

12:1913:13

12 · Fall Guys Clone

Same prompt run against Opus 5 and Kimi K3: Kimi nailed gameplay/physics/movement fast (9 min, $4.40) while Opus 5 spent longer (17 min, ~$13) chasing extra UI polish and visual effects, with the creator judging Opus 5's overall quality slightly ahead.

13:1313:35

13 · Pricing

Opus 5 is available on all paid Claude plans and via API at $5 per million input tokens and $25 per million output tokens, keeping the same 1-million-token context window as its predecessor.

13:3514:14

14 · Voxel Bench

On VoxelBench, Opus 5 at max reasoning generates a detailed voxel recreation of a TIE-fighter-style starfighter, procedurally producing geometry, proportions, wings, and cockpit — the benchmark measures 3D spatial reasoning and procedural generation from code.

14:1414:45

15 · Butterfly SVG

Opus 5 generates an animated SVG butterfly with an ambient background, per-wing color palette, and adjustable controls for animation speed, wing flutter, and aura glow.

14:4515:31

16 · Frontend Demo

The creator concedes Opus 5's front-end output isn't quite on Fable 5's level in every area but is noticeably cheaper, and again points to Kimi K3 as the stronger/cheaper pick specifically for front-end tasks.

15:3116:33

17 · Minecraft Clone

A second, more built-out Minecraft clone in creative mode: moving flowers, water dynamics, varied terrain/mobs, cave systems, placeable torches, and a lava lake, praised for attention to detail.

16:3317:08

18 · Unity Game Development

The creator notes people are using Opus 5 with the Unity CLI to script game assets directly, including improving a scene originally built by Fable 5 with styled water, wind-reactive foliage, better lighting, contact shadows, and depth outlines for a stronger visual atmosphere.

17:0817:58

19 · 3D Demo

Anthropic's own example: a working wind tunnel simulation visualizing airflow over aerodynamic and non-aerodynamic objects, combining interactive controls, object manipulation, and a readable flow-field visualization.

17:5818:34

20 · Blackhole Sim

A fully interactive browser-based black hole simulator (Kerr metric) with real-time adjustable spin, viewing angle, and accretion-disc acceleration, updating the gravitational lensing effect live.

18:3419:13

21 · Conclusion

Wrap-up: the creator calls it a genuine Anthropic comeback and cheaper than Fable 5, recommends trying it via the World of AI benchmark tool, and closes with newsletter/Discord/Twitter/subscribe CTAs.

Atomic Insights

Lines worth screenshotting.

  • Claude Opus 5 scored 83.7 on the creator's own reasoning benchmark, ranking it above GPT-5.6 Sol as the top reasoning model tested, while sitting just behind Fable 5.
  • On Artificial Analysis's Intelligence Index, Opus 5 scored 61 versus Fable 5's 60 and GPT-5.6 Sol's 59 — a narrow lead while costing about 26% less per task than Fable 5.
  • Opus 5 scored 30.2% on ARC-AGI-3, more than quadrupling the previous record of 7.8% set by GPT-5.6 Sol on maximum effort.
  • In a head-to-head 3D city simulation test, Opus 5 finished for $4.20 versus Fable 5's $9.60 — under half the cost for a comparable result.
  • Building a Fall Guys-style clone, Kimi K3 finished in about 9 minutes for $4.40 while Opus 5 took roughly 17 minutes and cost about $13, spending the extra time polishing UI and visual effects.
  • Opus 5's API pricing is $5 per million input tokens and $25 per million output tokens, on the same 1-million-token context window as its predecessor.
  • Opus 5 is now the default model on Claude Max and the strongest model available to Claude Pro.
  • The creator's single biggest complaint about Anthropic's models is overly aggressive cybersecurity safeguards that can refuse or flag legitimate developer prompts.
  • Despite ranking Opus 5 highly for reasoning and engineering, the creator still recommends Kimi K3 for front-end-heavy web development because it's cheaper and does well with UI/physics out of the box.
  • For a daily driver, the creator personally uses GPT-5.6 Sol inside Codex over Opus 5, citing lower latency and more generous usage limits.
  • Opus 5 one-shot a full playable COD Zombies-style clone with round-based survival, escalating waves, a purchasable weapon/perk system, and a health-doubling perk machine, coded without any starter codebase.
  • VoxelBench, used to test the model on a detailed voxel TIE fighter build, measures how well a model generates complex 3D scenes purely from code — spatial reasoning plus procedural generation.
Takeaway

Rank models by cost-per-task, not just leaderboard position.

WHAT TO LEARN

Claude Opus 5 shows that a model narrowly behind the top scorer on every major benchmark can still be the better pick once you weigh cost per task, and no single model wins every category.

02BEST Reasoning Model
  • A model doesn't have to top the leaderboard to be the better buy — Opus 5 trails Fable 5 on most benchmarks but wins on cost-per-task, which is what actually matters for repeated real-world use.
03Benchmarks
  • Frontier Bench and similar coding/knowledge-work benchmarks are increasingly reported alongside cost, because raw capability without a cost figure hides which model is actually efficient to run at scale.
04Best Model Ever?
  • A composite third-party score like the Artificial Analysis Intelligence Index is more useful for comparing models than any single vendor's own benchmark, since vendors have an incentive to cite the numbers that favor them.
05ARC-AGI-3
  • ARC-AGI-3-style novel-puzzle benchmarks exist specifically to catch memorization — a model can ace familiar coding tests while still failing genuinely new reasoning problems, so a 4x jump on ARC-AGI-3 signals a different kind of capability gain than a benchmark score bump.
06Which Model To Use
  • Pick your model per task type, not once for everything: reasoning/debugging, front-end polish, and daily-driver latency/usage-limits are three different axes, and the model that wins one routinely loses another.
  • Aggressive safety/content filters that block or flag legitimate technical prompts are a real cost of using a frontier model in production, not just a minor annoyance — factor false-positive refusal rate into a model choice alongside benchmark scores and price.
09COD Zombies Clone
  • Benchmarks that test procedural 3D/spatial generation from code (VoxelBench-style) are a proxy for a model's real coding and spatial-reasoning ability that's harder to game than text-only coding tests.
11Opus 5 vs Fable 5
  • When two models are pitted on the identical prompt and one 'converges fast' while the other 'keeps iterating for polish,' that's a legible signal of the model's default behavior under ambiguous instructions — useful to know before you build automation around it.
12Fall Guys Clone
  • When comparing two models on the same generation task, track both wall-clock time and dollar cost — a model that takes almost twice as long and costs 3x more (Opus 5's 17 min/$13 vs Kimi K3's 9 min/$4.40 on the same prompt) may still be worth it if the extra output quality matters, but only if you're actually measuring that tradeoff instead of assuming faster/cheaper is worse.
  • The cheapest model on a benchmark leaderboard and the cheapest model for your actual workload can be two different models — Kimi K3 undercut both Opus 5 and Fable 5 specifically on front-end tasks, not across the board.
13Pricing
  • A model's context window size staying flat between versions (Opus 5 kept Opus 4.8's 1-million-token window) is a useful diagnostic — it tells you where the vendor invested its improvement budget (reasoning/efficiency, not raw context).
07How To Use
  • Treat any single reviewer's benchmark tool as a starting point, not a verdict — the video's own advice is to re-run the comparison on your own prompts, because your workload's mix of tasks won't match a generic benchmark suite.
Glossary

Terms worth knowing.

World of AI Bench
The creator's own third-party benchmark platform and leaderboard (woaibench.ai) for scoring AI models on coding, reasoning, and agentic tasks, used as the video's primary comparison tool.
Artificial Analysis Intelligence Index
A composite score from the analytics firm Artificial Analysis that aggregates results across multiple benchmarks into a single ranked intelligence score for comparing AI models.
ARC-AGI-3
A benchmark that tests an AI model's ability to solve novel abstract-reasoning puzzles it hasn't seen before, scored as a percentage of problems solved, used as a proxy for genuine reasoning versus memorization.
VoxelBench
A benchmark that scores how well a model can generate complex, detailed 3D scenes purely from code, testing spatial reasoning, geometry, and procedural-generation ability.
Unity CLI
A command-line interface that lets an AI model script and modify assets inside the Unity game engine directly from text prompts, without using the Unity editor GUI.
Agentic terminal coding
A benchmark category measuring how well a model can autonomously use a command-line/terminal environment to complete multi-step coding tasks without step-by-step human guidance.
Resources

Things they pointed at.

02:43toolArtificial Analysis
03:26toolARC-AGI-3 / ArcPrize
Quotables

Lines you could clip.

01:41
The new king in reasoning and agentic work.
short, punchy claim about Opus 5's ARC/reasoning rankingTikTok hook↗ Tweet quote
04:10
For complex reasoning, debugging, and difficult engineering tasks, I would say the Opus five would probably be my first choice.
clear, quotable recommendation with a concrete use casenewsletter pull-quote↗ Tweet quote
11:32
Opus five completed it in about $4.20, whereas Fable five cost it around $9.60.
concrete head-to-head cost number, easy to caption as a stat cardIG reel cold open↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphoranalogy
00:00Yeah. I cannot believe this. I full on created a COD zombies full on clone with the Opus five, which has multiple levels.
00:08You have multiple weapons, and you have the same survival based gameplay that we see within Call of Zombies being fully cloned with the Opus five. This is just insane.
00:19Yesterday, I was talking about Opus five when we were looking through some of the leaks, and I've basically stated that it's on Fable five's level in many cases. But after looking through Anthropic's official benchmarks which launched today, I can officially tell you that I underestimated this new model.
00:37Across coding, reasoning, agentic workflow, and long horizon tasks, Opus five appears to be better in almost every area than I initially expected.
00:47This is where Anthropic is finally back with the official launch of Claude Opus five. They're delivering a genuinely frontier class level model, and it's thoughtful, proactive, and comes remarkably close to the intelligence of Fable five while costing roughly half the price and in many cases even better.
01:06On the world of AI benchmark, if we are to take a look at the leaderboard, you will notice that Claude Opus five is highly ranked above the GPT 5.6 soul which is incredible and this is where it is slightly behind Claude Fable five. But the wildest part to me is that Claude Opus five is currently ranked number one as the best reasoning model on the world of AI bench.
01:30It scored an 83.7 on x high, which beats g p t 5.6 soul and everything else currently listed. The new king in reasoning and agentic work.
01:40When it comes to benchmarks, this is a model that comes very close to fable five level intelligence at roughly half the price while setting new state of the art results on coding as well as knowledge work. This is with benchmarks like Frontier Bench as well as GDP Evolve.
01:56You can see in most of these benchmarks, it is surpassing Fable five, Opus 4.8 obviously, as well as GPT 5.6 Soul. It is now the default model on Claude Max and the strongest model available to Claude Pro. The biggest improvements is efficiency because the Opus five delivers much stronger performance for the same price as Opus 4.8, more than double its predecessor's FrontierBench score at lower cost per task and comes just within point 5% of Fable five on CursorBench, which is kind of insane to me, while costing about just half of whatever Fable five is doing.
02:35If you want the best AI tools, workflows, and drops before everyone else, join my free newsletter with the link in the description below which is completely free. Artificial analysis also ranked Claude Opus five as the most intelligent model that they have ever tested. On their intelligence index, Opus five scores a 61, narrowly beating Fable five at 60 and GPT 5.6 Soul at 59 while costing around 26% less per task than Fable five.
03:01It also sets state of the art results on AgenTechnology work benchmarks like GDPEvol a a and a a briefcase, making it one of the strongest models available for professional research, coding, and multistep workflows.
03:14And on coding agent index, Opus five is now tied for the first place where it is delivering performance that's remarkably close to Fable five at significantly lower cost. Opus five also delivered a massive result on Arc AGI three where it scored a 30.2 percentage, crushing the previous record of just 7.8 percentage set by g p t 5.6 SOL on maximum effort.
03:39ArcPrize has also observed completely new problem solving behavior, allowing OPUS five to solve environments that no previous model has ever beaten, even outperforming Fable five. So my two cents is that if you're on Claude Max, I would absolutely use the Claude OPUS five with Max reasoning inside Claude code as I still think ClodCode is the best coding harness available today.
04:02My biggest issue with Anthropic is still the overly aggressive cybersecurity safeguards. It can sometimes refuse or flag legitimate prompts which can be frustrating for a lot of developers. Overall though, I genuinely love this model and for complex reasoning, debugging, and difficult engineering tasks, I would say the Opus five would probably be my first choice.
04:24That said, I'd still reach for the Kimi k three. We know that within our benchmark, we have basically displayed that the Kimi k three is currently one of the best models that is available right now.
04:36It was ranked number three until Opus five came, but it is a model that does exceptionally well with web development. It's able to work on front end heavy tasks quite well at a significantly cheaper pricing range than OpenAI's models or even Anthropic. For my daily driver, I would be using the g v t 5.6 sole inside codecs.
04:56The lower latency and much more generous usage limits make it easier to rely on throughout the day. At the end of the day, use whatever works best for you, but that is just my take. Now throughout the next bit of the video, we're gonna be testing out this model across multiple domains to showcase what this model is capable of doing.
05:13And I highly recommend that you use the World of AI benchmark tool to actually run this benchmark to get a better evaluation, obviously, through your own prompts, but through the World of AI benchmark suite so that you can evaluate it across all of the different domains that assess a model and how proficient it is. You also get a lot of other cool things, so if you wanna try this out, you can easily do so with the link in the description below and you can start off for free.
05:37Here, Opus five was requested to simulate the Apollo mission, and this is where everything is procedurally generated with animations and even soundtracks added to this.
05:49This is just incredible. This is where it was able to generate everything from code. No premade assets were used.
05:56No external models. Just procedural generation and the fact that this was fully created is just remarkable. Just take a listen.
06:21Alright. This is gonna blow your mind. This is where I had Opus five fully create a playable COD zombie style clone.
06:29This is just insane, guys. Because it includes classic round based survival. You have escalating zombie waves.
06:36You even have different weapons that you can even purchase after you collect more points. Now I actually played this for twenty minutes before I made video, and I'm just having so much fun with this model right now. The fact that there is a full point system, even a juggernaut style perk machine that doubles your health for a couple of points.
06:57I forgot what it was. But you can see, I can open all of these individual doors that take me to different areas in the map. So let me show you how that would actually look after I get enough points.
07:11So it looks like I have enough points, so I can actually open the courtyard by clicking f, and there's even animations for the tour that goes down. But one thing I realized, I used to play Warzone a lot. If you are to press control, you're able to do a slide cancel, which is just nuts.
07:27That is just remarkable. Now the fact that it coded up all of these different areas within the map is just so cool to me. You even have a storage system.
07:37You have different crates, different perks that you can buy as you're walking through this map. Now this is one of the machines. I forgot which one this is.
07:46Speed cola. You even have different weapons that you can buy from these different, uh, shops.
07:52But let's actually go back and collect some more points.
07:58This right here is the juggernaut machine, which is gonna double your health. But let's actually go ahead and actually purchase one of the weapons. Now the entire map UI game play system all was fully loaded out with a core loop, which was five footed by the Opus five within the World of AI benchmark tool.
08:16There's no massive starter code base or a constant handling. It did this fully autonomously. The fact that it did this is just so remarkable to me.
08:28Not one thing I can say is the scope doesn't actually work in this, which is kind of hard, but it is what it is.
08:55Overall, I am really happy to see the type of quality that this model is.
09:02Here, I had requested it to create a Mac OS clone, which is where it has mimic the latest iOS. You have the starter screen, which it did add, and I really love what it ended up doing with the polished UI.
09:15The toolbar looks to be something that is functional, and it is reactive. You even have the top bar which actually works. Sound effects when you click on certain icons, we can actually use the top bar and that seems to be working quite well.
09:29Each and every component is perfectly attributed to this macOS clone, and I really love everything about this. You even have the ability to open up these individual files. The top right bar, which showcases the battery as well as a couple of other functions, has been thoroughly generated as well, like the top menu bar.
09:48You even have Siri that has been added, AI Siri, so I can ask it stuff. It actually has simulated this. So this is something that I haven't seen with any Mac OS clone, is nice.
09:57You even have a notification tab. And I had also added in something else to this where I had told it to create a Minecraft clone, and I told it to actually use a lot more, uh, expenditure in creating out that Minecraft clone.
10:12But everything seems to be coded out perfectly. You have a photos app, which is really cool. You can open up these individual photos.
10:19You even have a music player.
10:23It seems to be working as well, so that's nice. You have a preview for opening up the photos, but this is the Minecraft clone. Let's actually take a look at it.
10:30This is where it created a pretty basic clone within this, uh, system. You can break different blocks, which is nice, and it gives you that animation as well. The lighting doesn't seem to be thoroughly generated, but it did end up adding mobs, which is nice.
10:45The texture for the mobs is not something that is perfect. I guess it tried to add in a cave system, which is cool, but that is a little buggy. Regardless, I do love what it did and attempted to do with this particular generation.
11:00It is still decent, and I do have a better Minecraft clone that I will be showing a little while later. But moving forward with this demo, you also have the twenty forty game, Minesweeper, a terminal, you have a calculator, a weather app, you even have the clock, activity monitor and the system settings where you can change things like the accent color.
11:20You can reduce transparency. And then there's a lot of functions that we haven't seen with other models attempt to clone a Mac OS, but I just really love the attention to detail with this particular generation. Here is one comparison that really caught my attention and that is the costing because the Opus five was on a head to head test against the Fable five to create a three d simulation of a city.
11:44And for the exact same task, OPUS five completed it in about $4.20, whereas Fable five cost it around $9.60. That's less than half the price while delivering incredible similar results, maybe even better in terms of recreating the full on UK skyline.
12:03Now that is just something that I really didn't expect for the Opus to actually generate. You have various sorts of animations like the water stream, even have different cars moving. I'm not saying Fable five doesn't do that, but it just looks more lively and more interactive from the Opus.
12:19This is just nuts, guys, because Opus five full on created a clone of Fall Guys. Obviously, it's not the exact same thing as Fall Guys, but this looks really similar to the actual game. With the exact same prompt, it also tested Kimi k three.
12:34Kimi definitely nailed the gameplay, the physics, as well as the movements right out of the box, but if you take a look at what Opus five did, it focused heavily on refining the UI and kept on adding unnecessary visual effects that make it more alive. It spent much longer iterating without fully converging, but still, it did a better job with the overall quality.
12:54The difference was pretty significant. Kimi definitely finished at about nine minutes for roughly $4.40, but Opus five ran for around seventeen minutes and costed around $13.
13:05So if your goal is building polished systems, web experiences, you would wanna use a hybrid approach of maybe using both of them. I should've stated this a bit earlier, but Opus five is something that could be accessible with any of the paid cloud plans.
13:20You can easily access it through their chatbot and through the API. It's listed at five dollars per 1,000,000 input tokens and $25 per 1,000,000 output tokens.
13:30And this is a model that still contains the same context window, which is 1,000,000. On the Voxel bench, this is where the Cloud Opus five at max reasoning created an incredibly detailed voxel recreation of the Galactic Empire's iconic twin ion engines starfighter.
13:46And that is just insane because the geometry, the proportions, wings, cockpilot, and overall silhouette are all procedurally generated, and the final result looks insanely polished.
13:58And if you do not know why the voxel bench is important, it essentially measures how well an AI can generate complex three d scenes from code. It tests the spatial reasoning understanding of the model as well as procedural generation and coding ability in general.
14:14Here's where I had requested it to create a butterfly in SVG, and I really love what it did here. It added in the ambient background with each individual wing palette, which is really cool to see.
14:26You even have customization where you can change the animation speed. You can even change different displays that change certain things like the wing flutter as well as aura glow, stuff like that that enhances the overall capability and function of this butterfly display.
14:43Now front end wise, I wouldn't say it is exactly on the same level as Fable five. In different areas, it may be better, but I do think it is a lot cheaper to use over the Fable five. And I like I said before, you would be wanting to use the Kimi k three in most cases for web development tasks like front end, and that is where you can get the best out of the quality as well as pricing from that model.
15:10I'm not saying not to use the Opus five because in many cases, you're gonna be able to use certain functionalities of this model in that particular domain, which would be even better than Kimi. So you would need to flip around and see which works best when you're initially testing out these models. I'm just stating that Opus five definitely delivers some strong front end design taste.
15:31Now this is the Minecraft clone that I've really wanted to showcase to you guys. This is where we have an inventory system. I'm currently in creative, so I'm not gonna be able to break and show the full on, uh, breakage of blocks.
15:43I will be showcasing it in survival in a second, but what I really like is the attention to detail. You have the flowers moving like how it does within the game.
15:52You're able to well, if I have to go into the water, you're able to actually see that it has added in the water dynamic feature from the original sandbox game. You have different mobs. You have different terrains.
16:04The texture of whatever it has added is just remarkable, guys, with this clone. What I really like is that it added in cave systems as well. There's different auras.
16:14You even have the ability to place lights like torches, and then there we go. This is the lava lake that I wanted to showcase.
16:22So it added in a lot of cool features, and it just is remarkable to see that we're in this day and age where an AI model in one shot can create a full on Minecraft clone. Now some people have been using Opus five quite interestingly, like with Unity CLI, and it honestly feels like a game changer for them because you're able to code out different sorts of game assets with this model where some people use it to improve a scene originally built by Fable five and then adding styled water, wind reactive foilage, better lighting, contact shadows, and depth that outlines different components, and it gives the overall atmosphere better visual style than what we previously saw with the Fable five.
17:04And that is a cool way where people have been using this model right now. Here, Anthropic showed a really cool example where it generated a working wind tunnel simulation that visualizes airflow over aerodynamic and non aerodynamic objects.
17:18But what makes this notable is that it's not just a nice three d scene. It combines the interactive controls, object manipulation, and readable flow field visualization in a way that actually communicates with aerodynamic behaviors.
17:31This is the kind of results that starts to push beyond flashy demos and it shows us what you can actually do with this model in different scenarios. If you like this video and would love to support the channel, you can consider donating to my channel through the super thanks option below. Or you can consider joining our private discord where you can access multiple subscriptions to different AI tools for free on a monthly basis, plus daily AI news and exclusive content, plus a lot more.
17:59Honestly, this model is just genuinely leaving me speechless because it built a fully interactive black hole simulator that runs directly within the browser. Just take a look at the quality of output which thoroughly represent what a black hole would be able to do to different forms of light.
18:16You can adjust the black hole spin, viewing angle, as well as acceleration disc in real time while the gravitational lens updates live. It is something that I haven't seen any model actually produce with the same quality as well as depth.
18:31This is just nuts that Fable five is able to do stuff like this. Overall, I am really happy that Anthropic is slowly making a comeback, and I definitely love that Opus five is a lot cheaper than what we have with Fable five. It is something that I would definitely recommend trying out with all the links that I use in today's video, especially the World of AI benchmark tool.
18:50Make sure you go ahead and take a look at that with the links in the description below where you can get started for free. You can also take a look at the universe of AI, which is our second channel. Join the newsletter.
18:59Join the Discord. Follow me on Twitter. And lastly, make sure you guys subscribe, turn on notification bell, like this video, and please take a look at our previous videos so that you can stay up to date with the latest AI news.
19:07But with that thought, guys, have an amazing day. Spare positivity, and I'll see you guys fairly shortly. Peace out, fellas.
The Hook

The bait, then the rug-pull.

Before the benchmarks load, the creator leads with the payoff: Claude Opus 5 one-shot a full playable Call of Duty Zombies clone, and that's presented as proof the model's official launch-day numbers undersell it.

CTA Breakdown

How they asked for the click.

VERBAL ASK
02:29newsletter
If you want the best AI tools, workflows, and drops before everyone else, join my free newsletter with the link in the description below which is completely free.

Soft mid-roll CTA dropped mid-benchmark-recap without breaking pacing, repeated more fully in the outro alongside Discord, Patreon, Twitter, and subscribe asks.

Storyboard

Visual structure at a glance.

cold open
hookcold open00:00
benchmark table
valuebenchmark table02:43
COD Zombies clone
valueCOD Zombies clone06:20
cost comparison
valuecost comparison11:32
black hole sim
valueblack hole sim17:58
outro/CTA
ctaoutro/CTA18:34
Frame Gallery

Visual moments.

Chat about this