The argument in one line.
Claude Opus 5 narrowly trails or matches Fable 5 on major intelligence benchmarks while costing roughly half as much per task, making it the creator's pick for complex reasoning and coding even though he'd still reach for other models on frontend-heavy or daily-driver work.
Read if. Skip if.
- You're deciding which frontier model to route Claude Code, an agent, or a coding workflow through and care about cost-per-task, not just raw benchmark rank.
- You follow AI model launches and want the actual benchmark numbers (Artificial Analysis, ARC-AGI-3, World of AI Bench) rather than just the marketing claim.
- You're curious what a frontier model can one-shot right now — full games, OS clones, physics sims — as a gut check on where AI-assisted generation has gotten to.
- You want a rigorous, independently-reproduced benchmark study — this is one creator's own benchmark tool plus a handful of anecdotal demo runs, not a controlled evaluation.
- You need current, verified pricing and availability — model lineups and pricing change fast and this is a single launch-day snapshot.
The full version, fast.
Claude Opus 5 launched positioned as frontier-class but roughly half the price of Claude Fable 5. Across the creator's own World of AI Bench, Artificial Analysis's Intelligence Index (61 vs Fable 5's 60), and ARC-AGI-3 (30.2% vs the prior 7.8% record), Opus 5 lands at or near the top while consistently costing less per task — a 3D city simulation that cost Fable 5 $9.60 cost Opus 5 $4.20 for a comparable result. Hands-on, it one-shot a playable COD Zombies-style game, a macOS clone, two Minecraft clones, a Fall Guys clone, and a browser black hole simulator, with the video highlighting strong reasoning, debugging, and 3D/spatial generation. The counterpoint: Kimi K3 still wins on frontend polish and speed for less money, and the creator personally sticks with GPT-5.6 Sol as a daily driver for its lower latency and higher usage limits — Opus 5's real edge is complex reasoning and engineering work at a lower cost, not being uniformly the best model at everything.
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →Where the time goes.

01 · Introduction
Cold open teasing the COD Zombies clone, then the creator says yesterday's leak-based take that Opus 5 was 'on Fable 5's level' undersold it after seeing Anthropic's official benchmarks.

02 · BEST Reasoning Model
On the World of AI Bench leaderboard, Opus 5 ranks above GPT-5.6 Sol and just behind Fable 5 overall, but scores highest of all as the #1 reasoning model at 83.7 on the 'x high' reasoning test.

03 · Benchmarks
Opus 5 is called the new default on Claude Max and strongest model on Claude Pro, with FrontierBench and knowledge-work scores shown beating Fable 5, Opus 4.8, and GPT-5.6 Sol, and roughly double Opus 4.8's FrontierBench score at lower cost per task.

04 · Best Model Ever?
Artificial Analysis ranks Opus 5 as the most intelligent model it has ever tested, scoring 61 on its Intelligence Index versus Fable 5's 60 and GPT-5.6 Sol's 59, at about 26% less cost per task than Fable 5.

05 · ARC-AGI-3
Opus 5 scores 30.2% on ARC-AGI-3, crushing the prior 7.8% record set by GPT-5.6 Sol, with ArcPrize noting genuinely new problem-solving behavior on environments no previous model solved.

06 · Which Model To Use
The creator's personal model picks: Opus 5 with Max reasoning inside Claude Code for complex reasoning/debugging/engineering, Kimi K3 for front-end-heavy web work at a cheaper price, and GPT-5.6 Sol inside Codex as his own daily driver for lower latency and higher usage limits.

07 · How To Use
Brief plug for using the World of AI benchmark tool to evaluate models against your own prompts, framed as more useful than relying on any single reviewer's take.

08 · Apollo Simulation
Opus 5 procedurally generates an Apollo-mission space simulation from scratch — code, animation, and soundtrack, with no premade assets or external models used.

09 · COD Zombies Clone
The headline demo: a fully playable round-based zombies survival clone with escalating waves, purchasable weapons, a health-doubling 'juggernaut' perk machine, openable doors/areas, and a slide-cancel movement mechanic — built autonomously with no starter codebase.

10 · MacOS Demo
A polished macOS clone with a working top bar, reactive toolbar, click sound effects, an AI 'Siri' the creator can query, notifications, a photos app, a music player, and a separately-prompted Minecraft clone launched from within it.

11 · Opus 5 vs Fable 5
Head-to-head 3D city simulation of a UK skyline: Opus 5 completed it for $4.20 versus Fable 5's $9.60 — under half the cost — with the creator calling Opus 5's version more lively and interactive (moving water, cars).

12 · Fall Guys Clone
Same prompt run against Opus 5 and Kimi K3: Kimi nailed gameplay/physics/movement fast (9 min, $4.40) while Opus 5 spent longer (17 min, ~$13) chasing extra UI polish and visual effects, with the creator judging Opus 5's overall quality slightly ahead.

13 · Pricing
Opus 5 is available on all paid Claude plans and via API at $5 per million input tokens and $25 per million output tokens, keeping the same 1-million-token context window as its predecessor.

14 · Voxel Bench
On VoxelBench, Opus 5 at max reasoning generates a detailed voxel recreation of a TIE-fighter-style starfighter, procedurally producing geometry, proportions, wings, and cockpit — the benchmark measures 3D spatial reasoning and procedural generation from code.

15 · Butterfly SVG
Opus 5 generates an animated SVG butterfly with an ambient background, per-wing color palette, and adjustable controls for animation speed, wing flutter, and aura glow.

16 · Frontend Demo
The creator concedes Opus 5's front-end output isn't quite on Fable 5's level in every area but is noticeably cheaper, and again points to Kimi K3 as the stronger/cheaper pick specifically for front-end tasks.

17 · Minecraft Clone
A second, more built-out Minecraft clone in creative mode: moving flowers, water dynamics, varied terrain/mobs, cave systems, placeable torches, and a lava lake, praised for attention to detail.

18 · Unity Game Development
The creator notes people are using Opus 5 with the Unity CLI to script game assets directly, including improving a scene originally built by Fable 5 with styled water, wind-reactive foliage, better lighting, contact shadows, and depth outlines for a stronger visual atmosphere.

19 · 3D Demo
Anthropic's own example: a working wind tunnel simulation visualizing airflow over aerodynamic and non-aerodynamic objects, combining interactive controls, object manipulation, and a readable flow-field visualization.

20 · Blackhole Sim
A fully interactive browser-based black hole simulator (Kerr metric) with real-time adjustable spin, viewing angle, and accretion-disc acceleration, updating the gravitational lensing effect live.

21 · Conclusion
Wrap-up: the creator calls it a genuine Anthropic comeback and cheaper than Fable 5, recommends trying it via the World of AI benchmark tool, and closes with newsletter/Discord/Twitter/subscribe CTAs.
Lines worth screenshotting.
- Claude Opus 5 scored 83.7 on the creator's own reasoning benchmark, ranking it above GPT-5.6 Sol as the top reasoning model tested, while sitting just behind Fable 5.
- On Artificial Analysis's Intelligence Index, Opus 5 scored 61 versus Fable 5's 60 and GPT-5.6 Sol's 59 — a narrow lead while costing about 26% less per task than Fable 5.
- Opus 5 scored 30.2% on ARC-AGI-3, more than quadrupling the previous record of 7.8% set by GPT-5.6 Sol on maximum effort.
- In a head-to-head 3D city simulation test, Opus 5 finished for $4.20 versus Fable 5's $9.60 — under half the cost for a comparable result.
- Building a Fall Guys-style clone, Kimi K3 finished in about 9 minutes for $4.40 while Opus 5 took roughly 17 minutes and cost about $13, spending the extra time polishing UI and visual effects.
- Opus 5's API pricing is $5 per million input tokens and $25 per million output tokens, on the same 1-million-token context window as its predecessor.
- Opus 5 is now the default model on Claude Max and the strongest model available to Claude Pro.
- The creator's single biggest complaint about Anthropic's models is overly aggressive cybersecurity safeguards that can refuse or flag legitimate developer prompts.
- Despite ranking Opus 5 highly for reasoning and engineering, the creator still recommends Kimi K3 for front-end-heavy web development because it's cheaper and does well with UI/physics out of the box.
- For a daily driver, the creator personally uses GPT-5.6 Sol inside Codex over Opus 5, citing lower latency and more generous usage limits.
- Opus 5 one-shot a full playable COD Zombies-style clone with round-based survival, escalating waves, a purchasable weapon/perk system, and a health-doubling perk machine, coded without any starter codebase.
- VoxelBench, used to test the model on a detailed voxel TIE fighter build, measures how well a model generates complex 3D scenes purely from code — spatial reasoning plus procedural generation.
Rank models by cost-per-task, not just leaderboard position.
Claude Opus 5 shows that a model narrowly behind the top scorer on every major benchmark can still be the better pick once you weigh cost per task, and no single model wins every category.
- A model doesn't have to top the leaderboard to be the better buy — Opus 5 trails Fable 5 on most benchmarks but wins on cost-per-task, which is what actually matters for repeated real-world use.
- Frontier Bench and similar coding/knowledge-work benchmarks are increasingly reported alongside cost, because raw capability without a cost figure hides which model is actually efficient to run at scale.
- A composite third-party score like the Artificial Analysis Intelligence Index is more useful for comparing models than any single vendor's own benchmark, since vendors have an incentive to cite the numbers that favor them.
- ARC-AGI-3-style novel-puzzle benchmarks exist specifically to catch memorization — a model can ace familiar coding tests while still failing genuinely new reasoning problems, so a 4x jump on ARC-AGI-3 signals a different kind of capability gain than a benchmark score bump.
- Pick your model per task type, not once for everything: reasoning/debugging, front-end polish, and daily-driver latency/usage-limits are three different axes, and the model that wins one routinely loses another.
- Aggressive safety/content filters that block or flag legitimate technical prompts are a real cost of using a frontier model in production, not just a minor annoyance — factor false-positive refusal rate into a model choice alongside benchmark scores and price.
- Benchmarks that test procedural 3D/spatial generation from code (VoxelBench-style) are a proxy for a model's real coding and spatial-reasoning ability that's harder to game than text-only coding tests.
- When two models are pitted on the identical prompt and one 'converges fast' while the other 'keeps iterating for polish,' that's a legible signal of the model's default behavior under ambiguous instructions — useful to know before you build automation around it.
- When comparing two models on the same generation task, track both wall-clock time and dollar cost — a model that takes almost twice as long and costs 3x more (Opus 5's 17 min/$13 vs Kimi K3's 9 min/$4.40 on the same prompt) may still be worth it if the extra output quality matters, but only if you're actually measuring that tradeoff instead of assuming faster/cheaper is worse.
- The cheapest model on a benchmark leaderboard and the cheapest model for your actual workload can be two different models — Kimi K3 undercut both Opus 5 and Fable 5 specifically on front-end tasks, not across the board.
- A model's context window size staying flat between versions (Opus 5 kept Opus 4.8's 1-million-token window) is a useful diagnostic — it tells you where the vendor invested its improvement budget (reasoning/efficiency, not raw context).
- Treat any single reviewer's benchmark tool as a starting point, not a verdict — the video's own advice is to re-run the comparison on your own prompts, because your workload's mix of tasks won't match a generic benchmark suite.
Terms worth knowing.
- World of AI Bench
- The creator's own third-party benchmark platform and leaderboard (woaibench.ai) for scoring AI models on coding, reasoning, and agentic tasks, used as the video's primary comparison tool.
- Artificial Analysis Intelligence Index
- A composite score from the analytics firm Artificial Analysis that aggregates results across multiple benchmarks into a single ranked intelligence score for comparing AI models.
- ARC-AGI-3
- A benchmark that tests an AI model's ability to solve novel abstract-reasoning puzzles it hasn't seen before, scored as a percentage of problems solved, used as a proxy for genuine reasoning versus memorization.
- VoxelBench
- A benchmark that scores how well a model can generate complex, detailed 3D scenes purely from code, testing spatial reasoning, geometry, and procedural-generation ability.
- Unity CLI
- A command-line interface that lets an AI model script and modify assets inside the Unity game engine directly from text prompts, without using the Unity editor GUI.
- Agentic terminal coding
- A benchmark category measuring how well a model can autonomously use a command-line/terminal environment to complete multi-step coding tasks without step-by-step human guidance.
Things they pointed at.
Lines you could clip.
“The new king in reasoning and agentic work.”
“For complex reasoning, debugging, and difficult engineering tasks, I would say the Opus five would probably be my first choice.”
“Opus five completed it in about $4.20, whereas Fable five cost it around $9.60.”
Word for word.
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
The bait, then the rug-pull.
Before the benchmarks load, the creator leads with the payoff: Claude Opus 5 one-shot a full playable Call of Duty Zombies clone, and that's presented as proof the model's official launch-day numbers undersell it.
How they asked for the click.
“If you want the best AI tools, workflows, and drops before everyone else, join my free newsletter with the link in the description below which is completely free.”
Soft mid-roll CTA dropped mid-benchmark-recap without breaking pacing, repeated more fully in the outro alongside Discord, Patreon, Twitter, and subscribe asks.
- 🔥 Become a Patron (Private Discord) ↗
- 🧠 Follow me on Twitter ↗
- 🚨 Subscribe To The SECOND Channel ↗
- 🚨 Subscribe To The FREE AI Newsletter For Regular AI Updates ↗
- 👾 Join the World of AI Discord! ↗
- Claude Code + Ollama = FULLY FREE AI Coding FOREVER! (Tutorial) ↗
- Hermes Agentic OS is The Future ↗
- Hermes Agent The 24/7 Self-Evolving AI Agent! ↗
- AI NEWS ↗
- Blog Post ↗
- x.com ↗
- x.com ↗
- x.com ↗
- x.com ↗
- x.com ↗
- x.com ↗
- x.com ↗
- x.com ↗
- x.com ↗
- x.com ↗



























































