Modern Creator
Nick Saraev · YouTube

I Spent $400 Benching Opus 5 — Here's What It Can Do

A tour of a dozen apps Claude Opus 5 built from scratch, followed by Anthropic's own benchmark numbers on coding, computer use, and cost.

Posted
yesterday
Duration
Format
Review
hype
Views
26.7K
955 likes
Part of the collectionThe Claude Opus 5 PlaybookEvery Opus 5 breakdown, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

Claude Opus 5 rarely posts the single highest raw benchmark score against rival frontier models, but it wins on cost-per-task across coding, computer-use, and business-automation benchmarks, which the creator argues is the number that actually matters for builders.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You build interactive prototypes or demos with AI coding tools and want to see Opus 5's zero-shot creative and technical range before switching models.
  • You run an agency or automate business workflows and care about the new AutomationBench and computer-use numbers specifically.
  • You're deciding between Claude Opus 5, GPT-5.6, and comparable frontier models on cost-per-task, not just raw accuracy.
SKIP IF…
  • You want a step-by-step coding tutorial — this is a demo reel plus a benchmark readout, not a how-to.
  • You only care about consumer chatbot use cases; almost every number here is an agentic or tool-use benchmark.
TL;DR

The full version, fast.

Nick Saraev spends the first half of the video showing a dozen zero-shot apps Claude Opus 5 built — a 3D gallery, a cloth simulator, a predator-prey ecosystem, a sneaker configurator — arguing the model's craft has jumped even where raw benchmark gains look incremental. The second half breaks down Anthropic's release benchmarks against a comparison model he calls 'Fable 5', Opus 4.8, and GPT-5.6 Sol: Opus 5 leads on agentic terminal coding (43.3%), novel problem-solving (30.2% vs 1.5% for Opus 4.8), and business-workflow automation (26.0% vs ~17-18% for rivals), while roughly tying on raw coding benchmarks. His conclusion: Opus 5's real edge is cost-per-task — reaching comparable or better scores for about half the dollar cost of the rival model — plus the lowest misalignment score on Anthropic's behavioral audit.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0001:38

01 · Opus 5 just dropped (3D gallery demo)

Cold open: Opus 5 built a full navigable 3D art gallery with mouse-over sound cues, entirely from a single prompt.

01:3802:30

02 · Orbit launch game

A Kerbal Space Program-style orbital launch mini-game/explainer, letting the viewer launch objects into a simulated planet's gravity.

02:3002:55

03 · Cloth simulation

A Verlet-integration cloth/fabric simulator the creator can grab, tear, and manipulate in real time.

02:5503:28

04 · Fractal generator

A zoomable, pulsing 3D fractal shader with adjustable bloom, height, and dispersion controls.

03:2804:02

05 · Falling-sand cell sim

A falling-sand style particle sandbox (sand, water, oil, fire, steam) the creator interacts with live.

04:0204:32

06 · Predator-prey ecosystem

A cellular-automaton 'Meadow' ecosystem sim where the creator spawns rabbits and foxes and watches population boom-and-bust dynamics.

04:3204:55

07 · Quadcopter flight sim

A procedurally generated flight sim ('SkyHop') with a controllable drone, complete with sound and a landing mechanic.

04:5505:14

08 · Double pendulum

A chaotic double-pendulum simulator with trail visualization; the creator notes a minor rendering glitch here — the only flaw he calls out in the whole demo tour.

05:1405:52

09 · Pixel editor

A working pixel-art editor ('PixelForge') with eraser, undo, onion-skin, multi-frame animation, and resizable canvas.

05:5206:40

10 · Sneaker configurator

A 3D sneaker/product configurator ('Atelier') with material, color, and section customization — the creator notes it doesn't quite read as a shoe, but credits Opus 5 for generating the 3D assets itself rather than pulling stock models.

06:4007:14

11 · Animated population map

An 'Urban Spikes' data visualization animating city population growth (London, New York, Tokyo) from 1900 to 2026 on a rotating globe.

07:1408:20

12 · Wrecking ball simulator

A physics-based 'Wreck-It' demolition sim (wood, brick, glass) with a destruction-percentage tracker, plus a brief flow-field particle-art demo ('Living Ink') shown without narration.

08:2009:05

13 · My take + cost vs Fable 5

The creator's verdict: Opus 5 is 'slightly more intelligent [than] Fable at maybe half of Fable's costs.' A real build (AirLens product page) cost 69 cents in Opus 5 vs 94 cents for the comparison model, using only generated SVG assets.

09:0511:22

14 · Benchmark breakdown

Anthropic's release benchmark table read line by line: agentic terminal coding, knowledge work, novel problem-solving, agentic search, multidisciplinary reasoning, computer use, agentic coding, business workflows, legal, health, and biology.

11:2213:40

15 · Performance per dollar (effort curves)

Cost-vs-score scatter charts across effort levels for computer use, business workflows, multidisciplinary reasoning, and novel problem-solving — Opus 5 clusters at higher score for lower cost than rivals on most charts.

13:4014:34

16 · Alignment scores

Anthropic's automated behavioral audit bar chart: Opus 5 scores lowest (best) at 2.30, ahead of Opus 4.8, 'Mythos 5', and 'Sonnet 5'.

14:3415:53

17 · What this means for AI + outro

Closing take: Opus 5 is an incremental step in a 'slow and steady' release cadence, framed against Kimi K3's recent launch and a prediction that the most capable models will stay unreleased. Ends with a Maker School plug.

Atomic Insights

Lines worth screenshotting.

  • On Frontier-Bench v0.1 (agentic terminal coding), Opus 5 scored 43.3% versus 33.7% for the comparison model 'Fable 5', 21.1% for Opus 4.8, and 34.4% for GPT-5.6 Sol.
  • On ARC-AGI-3 (novel problem-solving), Opus 5 scored 30.2% against Opus 4.8's 1.5% — a 20x jump within the same model family in one generation.
  • On the new AutomationBench (building business workflows), Opus 5 scored 26.0% versus 17.4% / 17.0% / 18.1% for the three comparison models — the widest single gap on the whole benchmark table.
  • A real AirLens product-page build cost 69 cents in Opus 5 versus 94 cents for the comparison model — about 27% cheaper for a result the creator judged qualitatively better and built entirely from SVG, no downloaded assets.
  • On OSWorld 2.0 (agentic computer use), Opus 5's lowest-effort tier already scores about 60% for roughly $9/task, while the comparison model needs nearly $50/task to hit a similar score.
  • On Anthropic's automated behavioral audit (lower = safer), Opus 5 scored 2.30 — the best of four models shown, beating Opus 4.8 (2.85), a model labeled 'Mythos 5' (2.81), and one labeled 'Sonnet 5' (3.35).
  • On DeepSWE v1.1 and FrontierCode v1.1 (straight coding benchmarks), Opus 5 essentially tied the comparison model (68.8% vs 69.7%, 53.4% vs 53.5%) while GPT-5.6 Sol led both outright.
  • The creator's framing of the release: Opus 5 is 'basically slightly more intelligent [than the comparison model] at maybe half of [its] costs.'
Takeaway

Opus 5 wins on cost, not ceiling

COST VS CAPABILITY

Across coding, computer-use, and business-automation benchmarks, Opus 5 rarely posts the single highest raw score — its edge is matching or beating rival frontier models at roughly half the cost per task.

01Opus 5 just dropped (3D gallery demo)
  • A single prompt can now produce a full playable simulation — a 3D gallery, a cloth simulator, a predator-prey ecosystem, a flight sim — each generated by Opus 5 with, in the creator's words, 'virtually zero work.'
  • The strongest demos were self-contained physics or generative-art systems (cloth sim, fractal shader, falling-sand simulator); the weakest were object-imitation tasks like the sneaker configurator, which the creator says didn't quite 'read' as a real shoe.
13My take + cost vs Fable 5
  • On a real build (an Apple-style product page called AirLens), Opus 5 cost 69 cents to generate versus 94 cents for the comparison model — about 27% cheaper for a result the creator judged qualitatively better.
  • The AirLens page Opus 5 built used only generated SVG — no downloaded images or template assets — while the comparison model's version was described as visibly simpler.
  • The creator's overall framing: Opus 5 is 'slightly more intelligent [than] Fable at maybe half of Fable's costs' — the value case is cost-per-quality, not raw capability alone.
14Benchmark breakdown
  • On Frontier-Bench v0.1 (agentic terminal coding), Opus 5 scored 43.3% versus 33.7% for the comparison model, 21.1% for Opus 4.8, and 34.4% for GPT-5.6 Sol.
  • On ARC-AGI-3 (novel problem-solving), Opus 5 scored 30.2% against Opus 4.8's 1.5% and GPT-5.6 Sol's 7.8% — the comparison model had no published score on this benchmark yet.
  • On DeepSWE v1.1 and FrontierCode v1.1 (straight coding benchmarks), Opus 5 essentially tied the comparison model (68.8% vs 69.7%, and 53.4% vs 53.5%) while GPT-5.6 Sol led both — coding parity, not dominance.
  • On the new AutomationBench (building business workflows), Opus 5 scored 26.0% versus 17.4%/17.0%/18.1% for the three comparison models — the largest single gap in the whole benchmark table.
15Performance per dollar (effort curves)
  • On OSWorld 2.0 (agentic computer use), Opus 5's lowest-effort tier already scores about 60% for roughly $9/task, and its highest-effort tier reaches about 70% for about $22/task, while the comparison model needs nearly $50/task for a similar score.
  • Across every effort-level chart shown, Opus 5's cost-vs-score curve clusters with less spread than GPT-5.6 Sol, which the creator says has 'a massive spread of cost per task' between its lowest and highest reasoning settings.
  • On Humanity's Last Exam with tools, Opus 5 reached 64.7% versus 63.9% for the comparison model — a narrow win, but at meaningfully lower cost per task.
16Alignment scores
  • Anthropic's automated behavioral audit scores lower as safer: Opus 5 scored 2.30, versus 2.85 for Opus 4.8, 2.81 for a model labeled 'Mythos 5,' and 3.35 for a model labeled 'Sonnet 5' — Opus 5 was the safest model shown.
  • The creator specifically calls out 'Mythos 5' as the model he ties to a past cybersecurity incident that led the US government to pull back some model access — this is the creator's own characterization, not independently verified in this breakdown.
  • The pitch on alignment: Opus 5 pairs the best cost/capability profile in the chart with the best (lowest) misalignment score too — no capability-vs-safety tradeoff shown here.
17What this means for AI + outro
  • The creator frames Opus 5 as incremental, not a new unlock — 'the slow and steady stream of improvements' continuing from Opus 4 through 4.5, 4.7, and 4.8.
  • He ties the release timing to Kimi K3's launch about a week earlier, framing it as part of an ongoing back-and-forth between open-source/Chinese models and closed-source US frontier labs.
  • His prediction: the most capable ('galaxy brain') models will increasingly stay behind closed doors as AI labs coordinate with the US government, leaving the public with a 'slow and steady drip release' of what labs actually have internally.
Glossary

Terms worth knowing.

AutomationBench
A new benchmark cited in the video that scores models on building business automation workflows rather than general coding tasks.
OSWorld 2.0
A benchmark that measures how well an AI model can operate a computer the way a human would — clicking, typing, navigating software — referred to in the video as 'agentic computer use.'
ARC-AGI-3
A benchmark designed to test novel problem-solving on tasks a model hasn't seen patterns for before, as opposed to memorized knowledge.
Humanity's Last Exam
A very difficult multidisciplinary reasoning benchmark, scored in the video both with and without external tool use.
GDPval-AA v2
A knowledge-work benchmark referenced in the video; scores are shown as raw point totals (Opus 5: 1861) rather than percentages.
BrowseComp
A benchmark for agentic web search — how well a model can find and verify information by browsing, cited in the video as 'Agentic Search.'
Verlet cloth / Verlet integration
A physics simulation technique for approximating motion over time, commonly used to make realistic cloth, rope, or fabric simulations — labeled on-screen during the cloth-simulation demo.
Misaligned behavior / behavioral audit
Anthropic's automated scoring of how often a model does something a human didn't actually want in order to complete a task — lower scores mean fewer of these subversive or unwanted behaviors.
Resources

Things they pointed at.

15:38productMaker School
00:00channelLeftClick (Nick Saraev's AI growth agency)
14:34productKimi K3
Quotables

Lines you could clip.

00:39
Every time I mouse over, you can hear this kind of like a little ding that's, you know, me looking at this work and essentially cataloging it like a Pokedex.
quirky, self-contained visual metaphor for the 3D gallery demoTikTok hook↗ Tweet quote
08:20
Opus five is basically slightly more intelligent Fable at maybe half of Fable's costs.
one-sentence thesis of the entire videoIG reel cold open↗ Tweet quote
13:21
OPUS five is just brutally mocking them all at like 31 or so.
punchy benchmark brag line about the ARC-AGI-3 novel-problem-solving chartTikTok hook↗ Tweet quote
15:07
The actual intelligences, like the really big galaxy brain ones, the ones that are far smarter than anything we saw in this benchmark, I think those are going to repeatedly and more brazenly just stay behind closed doors.
provocative closing claim about AI labs withholding their most capable modelsnewsletter pull-quote↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphoranalogy
00:00Anthropic just dropped Claude Opus five, and it is by far probably the best frontier large language model currently available to us underclassmen. In this video, I'm gonna show you everything that you need to know about how to take the most advantage and use Opus five for the best of its capabilities. This is a website that essentially Opus created itself.
00:19It's a full three d application that allows me to kind of jump in, see the cartography of nowhere. Really, what this is is it's a three d world. And I was able to simulate this and create this entirely in Opus five with virtually zero work.
00:34So what we've done here is we've essentially created a bunch of art. We've then put these art, you know, these these things up on the wall, and I'm just walking through it.
00:42Every time I mouse over, you can hear this kind of like a little ding that's, you know, me looking at this work and and essentially cataloging it like a Pokedex. So this is not easy stuff to do. Right?
00:52I mean, it's not a long time ago that this thing would have been considered a full game and sort of three d experience in its own right. I bet you if I asked Opus five, it could turn this into an actual virtual reality experience in a few seconds, complete with, you know, enemies and hands and lasers and whatever the hell else I want.
01:08Well, I could show you benchmarks and I will in a second. Suffice to say, Opus five kicks ass on virtually all benches and I'm gonna break down exactly what that looks like in a second. But I think the more important thing here at this point is, how does the model feel?
01:21What sorts of outputs can it generate that, you know, I can rate based off taste? Not necessarily little static percentages on a screen that, let's be real, don't really mean anything to us anyway.
01:31So what follows is a massive list of all the impressive things I got Opus five to do, the things that I consider visually stimulating and also pretty interesting. Over here is pretty neat. This is like a three d, essentially Kerbal space program style launch game.
01:45It allows you to build a stable orbit in by changing the trajectory of this. So if I if I launch something, let's say from the launch site in that direction, what I can do is I can actually have it enter essentially the gravity of the planet, which is pretty badass.
02:02Obviously, this is kind of on the simpler design side, I would say. But it's still kind of neat that I could just ask AI to do this and I sort of have my own simulated solar system. It really does feel like simulations are where all this stuff is going.
02:13There's Newton's cannon over here, which actually shows a bunch of these launched. Then we also have Hawkman with a big target orbit over here. We're essentially trying to make this this thing.
02:23So it's both an explainer and it's also kind of a mini game. This is a simulation of fabrics and how they move in the wind.
02:30So it's called sort of a cloth simulation. And you can see I I'm even capable of ripping this simulation apart, which is very neat. You know, I can't say that this uses fabric dynamics, you know, but as accurately as possible.
02:44But, I mean, you know, when I'm looking at this and I'm thinking, damn, this is pretty accurate. This feels like it would respond the same way that, you know, a curtain or something would respond if I were to pull it. How about this fractal generator?
02:55I mean, this thing is insane. You can zoom in and you can go, I mean, however deep you want with this shader. It's honestly quite ridiculous.
03:03You can also see that it's sort of pulsing and growing and, you know, we can increase the height and the size, the bloom, how much it sort of shines, dispersion, and so on and so forth.
03:13King of simulation, I don't if you guys have ever played this game, but this is essentially a collection of cells that are growing in a tube. Every time I press this button down, what I'm doing is I'm dropping some sand down.
03:24So it allows Opus five to design this environment where these cellular automatons are sort of climbing. I can also do, I don't know, cause an oil spill here if I really wanted to, or maybe generate some steam underneath that, I don't know, destroys a lot of that water and and stuff like that.
03:40So this is pretty badass. I mean, it's fun that I can do that. I can also cause fire.
03:44I can regenerate starter scenes. I can stop time. I can do a lot.
03:48This over here is like a living predator prey ecosystem where I can spawn new bunnies. These bunnies can then start consuming, you know, grass in the environment.
03:57But then there's also what looks to be some foxes that are hunting. So, you know, I just caused a bunch of bunnies to appear. Let's get a bunch of foxes to appear, thin the herd a bit, shall we?
04:06And you can see what's happening is now there's a massive preponderance of grass here because the fox have just eaten all of them. However, the fox population is now crashing because there just aren't a lot more bunnies available to eat. How about this quadcopter flight simulator?
04:19I mean, here we are with my little quadcopter drone. I don't know if you guys could hear, but the drone itself is literally making sound. And I'm just proceeding through my totally, you know, procedurally generated environment.
04:31It looks like if I wanted to land, I'd I'd hold x. So I'm just gonna try that and land my helipad right there. Nice.
04:37And I just did. So this is a 100% like a game and yeah, just code that in a few seconds. This one's a double sided pendulum.
04:44I'm sure you guys are probably familiar with what one of these do. Interestingly enough, it does kind of look like it screwed up the design in the top right hand corner. Although that is the first screw up that I've been able to see so far.
04:54You You can add as many pendulums as you want and like redo this. Nothing super special here. Let's move on to something cooler.
05:00Here, designed a pixel editor. So I'm actually capable of drawing just like something in Microsoft Paint. We're very close to everybody being able to design their own sort of custom applications.
05:10But yeah, you can see how it has all the functionality. It has a little eraser functionality. I can, you know, undo whatever I want.
05:16I can, I don't know, do some sort of onion skin or even multiple layers, so I can add an additional frame? And then what I can do is I can actually like make a movie where I have this be my first frame, this be my second frame, all the way up. So that's kind of neat.
05:30Hey, you could you could sort of forge something. You could even make the canvas bigger or smaller, which is kind of neat. And you see, when I made it smaller, it does not look very happy with me.
05:37Here, we're building sort of a sneaker configurator. So I'm not gonna say this is the clearest and cleanest example of a sneaker I've ever seen, but I don't know. It looks like I can click on various things and change a bunch of this.
05:49So let me just shuffle, sort of change the design up. Maybe I want a new color here. You know, that looks nice.
05:54That looks like maybe it's like a Nike style thing. I can change the material and that'll slowly change what's going on right now. Have liquid chrome down.
06:02Looks like I can also save a render to my computer, which is kinda neat, or just stop it entirely. I can then move this thing around, click on different sections if I wanted to change things custom. Not gonna lie, not getting a shoe vibe here, but I think the reason why is because Opus five, it didn't just pull a shoe from the internet.
06:19What this did is it actually created these elements itself, which is a big step up because that sort of compositing is usually not what models do. We have what looks to be here a explainer graph showing the populations of different centers over the course of the last hundred and twenty six years.
06:36So in the top right hand corner, you can see the populations of New York change, London change, Tokyo change, and so on and so forth while it explains it to us. So now we're moving over to Asia and we're seeing that post war Tokyo is now at 16,000,000. And you know, I can also drag this and move this however I want with populations just getting increasingly interesting, I should say, as we go.
06:56I can also zoom in, zoom out, do all this fun stuff. Okay. This one's interesting.
07:00This is a wrecking ball simulator and so I can actually drag this, move this around and then push this towards this brick wall.
07:10And I'm not doing a very good job, mind you. Definitely not coming in like a wrecking ball. But yeah, you can see that that this is falling, which is pretty wild.
07:18You can also change the mass of the ball, make it heavier or lighter, which is kind of fun. Anyway, let's pretend I have a bunch of glass here. There you go.
07:27And there's even a little like slider progress bar that's asking, hey, know, how how close am I to destroying the entire thing? Apparently, 96%.
07:35So this thing's pretty close. Yeah? I mean, this can virtually do anything that you would ever want it to do.
07:40And if it's not clear, uh, the reason why I'm so excited about this is not because it can do all these things. Obviously, are a lot of models out there now that can do things like this. But I think the way that it does the things is a little bit better than those models, a few percentage points better anyway.
07:54And I think more importantly, it's also way cheaper. And I'm gonna show you guys this in a second, but this is far cheaper than any of the other major large language models, especially the ones at the frontier for these sorts of tasks. So what's my take on this whole thing?
08:07Opus five is basically slightly more intelligent Fable at maybe half of Fable's costs. And so we see the cost of this website here, AirLens, which is sort of like a a Mac or Apple style website, apparently trying to sell you some sort of three d panels, was 69¢, whereas Fable five's was 94¢.
08:27And you can see that I actually think Opus five did a better job. I I don't think the websites are really even comparable. If I open both of them up in new tabs, this is the one that Opus five just made, k, where there's this cool kind of, you know, slide thing.
08:40Keeping in mind that this is not, you know, a a resource downloaded from the Internet. It just made these resources with SVG. And then this is Fables over here, which, you know, I think is just way simpler and probably a lot dumber in reality.
08:52So on to benchmarks, you could see Opus five here does pretty darn well on agentic terminal coding. So much so that it's actually considered better than Fable five at the same task and far better than Opus 4.8 and even GPT 5.6 Sol, which scores a little bit above Fable five, at least as of the time of this recording.
09:11On knowledge work, on GDP Val dash a a v two, scored eighteen sixty one, so this puts it squarely above, like, the average human at knowledge work. And novel problem solving using Arc AGI three, it scored 30.2%.
09:25Fable has not actually used this benchmark yet, so there's no stat here. Opus 4.8 scored 1.5%. So as a progression from, you know, Opus 4.8 to Opus five, at least in the Opus series, this is a massive step up.
09:38On AgenTic Search, you see that it scored 90.8% on Browse Comp. A little bit better than Fable five, significantly better than Opus 4.8, but essentially the same as GBT 5.6.
09:49So multidisciplinary reasoning, it scored 56.3% with no tools.
09:54But when tool augmented, it actually outperforms Fable at 64.7 to Fable 63.9.
10:01So that'll be very useful, especially for long steps, sort of human knowledge work tasks. Computer use was 70.6, so we're seeing models get better at just using a computer the way that a human being would.
10:12Nothing super impressive on the agentic coding aspect. The gbt 5.6 actually still outperforms all of them, so they were probably just deep trained on deep s w e v 1.1.
10:22On agentic coding, you can see that it is essentially equal to Fable five, although a big step up from Opus 4.8. And what's cool is there's this new automation bench, which I'm particularly interested in, somebody that does a lot of automations for businesses.
10:36And it scores 26% on that compared to 17.4% for Fable, and this is on building workflows for businesses.
10:44So I mean, obviously, it's a far cry away from humans as of right now, know, basically a quarter of the time it's getting things right. But, you know, it's a big step up from just a couple of generations ago. Legal is at 11.7%.
10:55Health is at 59.8%. And then what's interesting is biology, it scores 49.4% on hard problems, but 90.1% on human solved problems, which is really, really interesting to see.
11:07One of my favorite breakdowns to date is the Agenta computer's performance by effort level. Essentially, they will feed models a bunch of different tasks, and then they will see how many tokens and, you know, conversely, how much money they needed to spend in order to complete those tasks.
11:22And so you can see over the course of, uh, you know, a variety of effort levels from like super low effort all the way up to high, there's there's quite the range. And most of this range is actually caused by GPT 5.6 SOL, um, which has a massive sort of spread of cost per tasks and scores on their lowest reasoning level all the way up to their highest.
11:41Well, anyway, Claude models tend to cluster sort of at the top right hand corner of this graph, which means there's less of a spread and, uh, also less of a a a task cost spread as well. So, you know, if you think about it, the dumbest version of Opus five still scores about 60% on the vast majority of agenda computer use tasks, um, for a cost of, I wanna say, about $9 or so.
12:01That's probably what that looks You know, as you progress up the reasoning ladder, the most intelligent models score something like 70% for about $22 a task, which is quite impressive, especially considering Fable five, you know, can score approximately the same, I wanna say, as the mid level of Opus five, but it does so at nearly $50 as well.
12:22So massive cost saver there. Performance on the new automation bench is just leagues above virtually every other model. As you could see, they tend cluster here around the, I don't know, 15% pass rate or so.
12:33This is almost 25%. So big implications on people that want to use models like this to automate parts of their workflows. And I hate to say it, but we're pretty close to beating humanity's last exam.
12:42I mean, you could see here, even the dumbest version of Opus five scores somewhere around 56 or so. And the smartest one is almost at 65% across the board for a cost per task of $3, which makes it not only better at Fable, a better than Fable at this particular benchmark, but also markedly cheaper.
13:00And it's not even really comparable to OPUS 4.8. There's also the ARC AGI three benchmark and on novel problem solving by cost, it is just massive top right hand corner here. You could see the vast majority of the other models, like OPUS 4.8, GPT 5.6.
13:15So they didn't put Fable on this, I think just because they didn't run it on ArcGa three. But they all scored somewhere around 2%. And then over here, OPUS five is just brutally mocking them all at like 31 or so.
13:25Finally, a mention about their alignment. Misaligned behavior is something that obviously large language model companies like Anthropic really wanna crank down on.
13:33They hate when models do things, you know, that are sort of subversive or things that a human being, you know, didn't really necessarily want the model to do in order to achieve the task. And as you can see here, Opus 4.8 on this 10 part score was a 2.85. Mythos five, which is that massive cybersecurity disaster causer that, uh, led the US government to pulling back a lot of this access was at 2.81.
13:57SONNET five was at 3.35, but OPUS five is way below all of them at 2.3. Which really just tells you not only is this model smarter, significantly cheaper, significantly better at a lot of specific tool calling tasks like the automation bench and variety of terminal bench workflows, but it's also a lot safer, which really makes this a win win across the board.
14:20What does this mean for AI more generally? Well, no real new crazy unlocks here. I mean, Opus five is continuing the slow and steady stream of improvements that we saw with Opus four, Opus 4.5, Opus 4.7, OPUS 4.8, and so on.
14:36This sort of thing comes at a good time because Kimi k three obviously dropped, you know, about think about a week or a week and a bit ago, and that really upset the quote unquote balance of power between, you know, these open source or Chinese models and then the more closed source now American frontier models. I think as we continue to go into the future, what'll probably happen is, you know, these models are getting quite intelligent and they're getting quite ready and capable of disrupting human knowledge work at scale.
15:01Because of the economic implications, I would imagine that large AI model companies like Anthropic, OpenAI, you know, now they're colluding with the US government, essentially. I think what'll probably happen is we're just gonna see a slow and steady drip release of these models.
15:15But the actual intelligences, like the really big galaxy brain ones, the ones that are far smarter than anything we saw in this benchmark, I think those are going to repeatedly and more brazenly just stay behind closed doors. And unfortunately, I don't think us poor little normies are gonna have access to them anytime soon.
15:30So do with that what you will, but, yeah, hopefully now you know everything you need in order to crush it with Opus five integrated into your business workflows and more. If you guys like this video, definitely check out Maker School down below. It's my ninety day automation accountability roadmap where I will guarantee you your first customer selling something built by Opus five or a similar model in ninety days, or I'll give you all your money back.
15:51Have a lovely rest of the day.
The Hook

The bait, then the rug-pull.

The title promises a $400 benchmarking spree, but the figure is never actually stated on camera — instead the video opens with Opus 5 walking the creator through a self-built 3D art gallery, the first of a dozen zero-shot demos before any numbers show up.

Frameworks

Named ideas worth stealing.

09:05list

Opus 5 release benchmark table

  1. Agentic terminal coding (Frontier-Bench v0.1): Opus 5 43.3% | Fable 5 33.7% | Opus 4.8 21.1% | GPT-5.6 Sol 34.4%
  2. Knowledge work (GDPval-AA v2): Opus 5 1861 | Fable 5 1747 | Opus 4.8 1593 | GPT-5.6 Sol 1736
  3. Novel problem-solving (ARC-AGI-3): Opus 5 30.2% | Fable 5 — | Opus 4.8 1.5% | GPT-5.6 Sol 7.8%
  4. Agentic search (BrowseComp): Opus 5 90.8% | Fable 5 87.4% | Opus 4.8 84.3% | GPT-5.6 Sol 90.4%
  5. Multidisciplinary reasoning, no tools (Humanity's Last Exam): Opus 5 56.3% | Fable 5 56.5% | Opus 4.8 49.8%
  6. Multidisciplinary reasoning, with tools (Humanity's Last Exam): Opus 5 64.7% | Fable 5 63.9% | Opus 4.8 57.9%
  7. Computer use (OSWorld 2.0): Opus 5 70.6% | Fable 5 66.1% | Opus 4.8 55.7% | GPT-5.6 Sol 62.6%
  8. Agentic coding (DeepSWE v1.1): Opus 5 68.8% | Fable 5 69.7% | Opus 4.8 59.0% | GPT-5.6 Sol 72.7%
  9. Agentic coding, main (FrontierCode v1.1): Opus 5 53.4% | Fable 5 53.5% | Opus 4.8 46.5% | GPT-5.6 Sol 47.5%
  10. Business workflows (AutomationBench): Opus 5 26.0% | Fable 5 17.4% | Opus 4.8 17.0% | GPT-5.6 Sol 18.1%
  11. Legal (Legal Agent Benchmark, held-out): Opus 5 11.7% | Fable 5 13.3% | Opus 4.8 10.4% | GPT-5.6 Sol 2.5%
  12. Health (HealthBench Professional): Opus 5 59.8% | Mythos 5 66.0% | Opus 4.8 57.4% | GPT-5.6 Sol 60.5%
  13. Biology, hard (BioMysteryBench): Opus 5 49.4% | Mythos 5 46.5% | Opus 4.8 42.4%
  14. Biology, human-solved (BioMysteryBench): Opus 5 90.1% | Mythos 5 89.0% | Opus 4.8 88.5%

Anthropic's own release-day comparison table, read aloud and shown on-screen, comparing Opus 5 against a model the creator calls 'Fable 5', the prior Opus 4.8, GPT-5.6 Sol, and (for health/biology rows only) a model labeled 'Mythos 5'.

Steal forA quick side-by-side reference when choosing which frontier model to route a specific agentic task to (coding vs. computer-use vs. business-automation).
CTA Breakdown

How they asked for the click.

VERBAL ASK
15:38product
If you guys like this video, definitely check out Maker School down below. It's my ninety day automation accountability roadmap where I will guarantee you your first customer selling something built by Opus five or a similar model in ninety days, or I'll give you all your money back.

Soft, single-mention plug placed only in the final seconds after the full benchmark analysis — tied directly to the video's own subject matter (build something with Opus 5) and framed as a money-back guarantee rather than a hard sell.

MENTIONED ON CAMERA
FROM THE DESCRIPTION
Storyboard

Visual structure at a glance.

open — bio slide over webcam
hookopen — bio slide over webcam00:00
demo tour begins
promisedemo tour begins01:38
my take + cost comparison
valuemy take + cost comparison08:20
benchmark table breakdown
valuebenchmark table breakdown09:05
Maker School CTA
ctaMaker School CTA15:38
Frame Gallery

Visual moments.

Chat about this