Modern Creator
Nick Saraev · YouTube

Fable 5.1: The Benchmark Breakdown

A same-day walkthrough of Claude Fable 5.1's benchmark chart, why the real story is cost-per-task rather than raw score, and what the new safeguard numbers mean for how often the model refuses benign questions.

Posted
6 days ago
Duration
Format
Review
educational
Views
26.1K
834 likes
Part of the collectionThe Fable 5 PlaybookAll 45 Fable 5 breakdowns, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

Fable 5.1 posts solid but not revolutionary score gains over its predecessor, and the more consequential change is that it delivers roughly 2.5x more capability per dollar, which the creator argues now matters more than raw intelligence.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You use Claude or a competing frontier model for coding, research, or business automation and want to know what actually changed before deciding whether to switch.
  • You're evaluating AI models on cost-per-task rather than sticker price per token, especially for agentic or automation-heavy workloads.
  • You want a fast, same-day read of a new model's benchmark chart without digging through the official release page yourself.
SKIP IF…
  • You want a hands-on tutorial with prompts or workflows — this is a benchmark readout and commentary, not a how-to.
  • You need independently verified numbers; this video reports Anthropic's own published benchmarks without third-party testing.
TL;DR

The full version, fast.

Fable 5.1 and Mythos 5.1 launched with benchmark gains across agentic scientific research, coding, knowledge work, computer use, reasoning, and especially business-workflow automation, which nearly doubled from 17.1% to 31.4%. But the creator argues the real story is cost: on Anthropic's accuracy-vs-cost frontier chart, Fable 5.1 delivers roughly 2.5x more capability per dollar than Fable 5, and a cut to cache-read pricing lowers real workload cost by 25-45%. He also flags that benchmark scores don't always predict real-world quality (Opus 5 scores well but underdelivers for him in practice), and closes on Anthropic's claim that Fable 5.1 flags benign requests as safeguard violations 60% less often, with an 85% drop in false refusals on biology and medical questions.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0000:14

01 · Cold open

Fable 5.1 and Mythos 5.1 just dropped; the creator promises to make sense of the updates and their practical impact.

00:1401:39

02 · Research, coding & knowledge-work scores

Reads the comparison table: Fable 5.1 scores over 50% on agentic scientific research (vs. 24.7-29% for Fable 5/Opus 5), 55.8% on agentic coding, and 1853 on the GDPVal-AA knowledge-work benchmark.

01:3902:37

03 · Computer use & multidisciplinary reasoning

Fable 5.1 scores 77.9% on OSWorld 2.0 (partial-credit computer use) and 60.9% on Humanity's Last Exam, both frontier gains over Fable 5 and Opus 5.

02:3704:03

04 · Business workflows: the real upset

AutomationBench score nearly doubles from Fable 5's 17.1% to Fable 5.1's 31.4%; agentic coding on Cursor Bench 3.2.0 also improves to 73.4%.

04:0404:58

05 · Benchmarks aren't the whole story

Opus 5 scores well on paper but the creator says he gets lower-quality real results from it than from Fable; flags that informal 'taste benchmarks' will matter over the next day.

04:5906:13

06 · Cost per task: the real headline

Anthropic's log-scale accuracy-vs-cost frontier graph shows Fable 5.1 scoring about 2.5x higher per dollar than Fable 5 at the same mean cost per task.

06:1307:29

07 · Cache-read discount

Cache reads cost about 75% less on Fable 5.1, cutting real-world workload cost by roughly 25% typically and up to 45% for highly agentic workloads.

07:2908:42

08 · Why intelligence keeps getting cheaper

The intelligence behind ChatGPT's 2022 launch now costs roughly 1000x less; the deciding question is shifting from 'can it do this' to 'is this profitable'.

08:4209:29

09 · Safeguards: fewer false refusals

Anthropic's post says benign requests are flagged as safeguard violations 60% less often, with an 85% drop in fallback rate on biology/medical questions.

09:2910:22

10 · Sign-off

Closes on the idea that as models converge in quality, distributing them effectively and pricing them right matters more than incremental intelligence gains.

Atomic Insights

Lines worth screenshotting.

  • Fable 5.1 scored over 50% on Terminal-Bench Science 0.1, roughly double Fable 5 and Opus 5's 24.7-29% range on that agentic scientific research benchmark.
  • On AutomationBench, the benchmark for automating real business workflows, Fable 5.1 nearly doubled Fable 5's score, from 17.1% to 31.4%.
  • On the accuracy-vs-cost frontier graph, Fable 5.1 scores roughly 2.5x better per dollar spent than Fable 5, not just a higher raw accuracy number.
  • Anthropic cut the cost of cache reads by about 75%, which lowers real-world workload cost by around 25% for typical use and up to 45% for highly agentic tasks.
  • The intelligence that powered ChatGPT's original 2022 launch now costs roughly 1000 times less to access than it did then.
  • Despite Opus 5 scoring well on paper across most benchmarks, the creator says he consistently gets lower-quality real answers from it than from Fable, which undercuts how much a benchmark score alone should be trusted.
  • Anthropic says Fable 5.1 now flags benign requests as safeguard violations about 60% less often than before.
  • On basic biology and medical questions specifically, the false-refusal ("fallback") rate dropped by around 85%.
  • Cost per task, not cost per token, is becoming the standard way to compare model pricing, because smarter models can need fewer tokens to finish the same job.
  • The industry's central question is shifting from "can the model do this?" to "is this profitable?" as raw intelligence stops being the bottleneck for most tasks.
Takeaway

Fable 5.1's real story is cost-efficiency, not just higher scores

WHAT TO LEARN

Fable 5.1 posts solid benchmark gains, especially on business-workflow automation, but the bigger shift is that it delivers roughly 2.5x more capability per dollar, which the creator argues matters more now than raw intelligence.

02Research, coding & knowledge-work scores
  • A new model's benchmark jump can be non-linear: Fable 5.1 roughly doubled the prior generation's score on agentic scientific research instead of nudging it up incrementally.
  • When comparing models, look at the specific benchmark task type (research vs. coding vs. knowledge work) rather than a single headline number, since gains aren't uniform across categories.
03Computer use & multidisciplinary reasoning
  • Computer-use benchmarks are often scored on a partial-credit basis, so a model that gets there in a roundabout way still counts as a pass; read the methodology before trusting the percentage.
  • A roughly 5 percentage point gain on a broad reasoning benchmark like Humanity's Last Exam is treated as meaningful progress at the frontier, where marginal gains get harder to produce.
04Business workflows: the real upset
  • The business-automation benchmark nearly doubling in one generation is the number to watch if you build or buy automated workflows, since it's a closer proxy for real deployability than academic reasoning scores.
  • Expect more people to pay for short bursts of frontier-model time to generate a complete automated workflow rather than subscribing to a narrow point-solution tool.
05Benchmarks aren't the whole story
  • A model can out-benchmark another on paper and still feel worse to actually use, so treat published scores as one input, not the final verdict.
  • Watch for informal 'taste benchmarks' in the days after a release; people testing creative or interactive tasks surface real-world quality gaps that academic benchmarks miss.
06Cost per task: the real headline
  • Compare models on a cost-per-task, log-scale frontier chart rather than a flat price list; it shows how much accuracy you actually get for a given budget.
  • A 2.5x improvement in capability-per-dollar can matter more to your bottom line than a modest jump in raw accuracy.
07Cache-read discount
  • If your workload makes repeated queries in quick succession, check whether the provider cut cache-read pricing; it can lower real costs by 25-45% without you changing anything.
08Why intelligence keeps getting cheaper
  • The cost of a fixed level of intelligence has dropped roughly 1000x since ChatGPT's 2022 launch; plan for AI capability to keep getting cheaper, not just smarter.
  • Once a model is smart enough for a task, the deciding factor shifts from can-it-do-this to what-does-it-cost, which should change how you evaluate new releases.
09Safeguards: fewer false refusals
  • A 60% drop in false-positive safeguard flags on benign requests is a usability fix as much as a safety one; fewer legitimate questions get needlessly downgraded to a weaker model.
  • The 85% drop in fallback rate on biology and medical questions specifically suggests the earlier version was over-triggering on entire topic categories, not just risky phrasing.
Glossary

Terms worth knowing.

Terminal-Bench Science 0.1
A benchmark that scores a model's ability to perform agentic scientific research tasks, such as running autonomous research loops without human supervision.
GDPVal-AA v2
A benchmark used to score a model's ability to understand and complete economically valuable 'knowledge work' tasks.
OSWorld 2.0
A computer-use benchmark that tests whether a model can operate a browser or OS to complete a task, often scored with partial credit for indirect but successful paths.
Humanity's Last Exam
A benchmark for multidisciplinary reasoning, designed to test broad knowledge and reasoning across many academic fields.
AutomationBench
A benchmark that measures a model's ability to automate real business workflows using representative company systems and processes.
Cursor Bench
A benchmark that scores agentic coding performance, meant to reflect how well a model handles real coding tasks rather than isolated code snippets.
cache reads
Reusing previously processed context in a follow-up query instead of reprocessing it from scratch, which cuts the cost of making multiple queries to a model in quick succession.
mean cost per task
A pricing metric that measures the average dollar cost to complete a benchmark task end-to-end, rather than the raw per-token price of the model.
Resources

Things they pointed at.

00:22toolTerminal-Bench Science 0.1
01:01toolCursor Bench 3.2.0
01:22toolGDPVal-AA v2
01:58toolOSWorld 2.0
02:08toolHumanity's Last Exam
02:37toolAutomationBench
Quotables

Lines you could clip.

00:43
It's worth noting that it just freaking brutally mogs GPT-5.6 Sol while doing it.
blunt, high-energy comparison line with a built-in punchlineTikTok hook↗ Tweet quote
03:09
This is the real upset here was for business workflows, which is something that is obviously near and dear to my heart as somebody that automates stuff.
personal stake + the video's biggest number (17.1% to 31.4%)IG reel cold open↗ Tweet quote
05:16
So mathematically, we do about 2.5x better per dollar, at least according to this frontier graph, than Fable 5.
the single number that summarizes the whole video's thesisnewsletter pull-quote↗ Tweet quote
08:20
It's now more about how do we distribute these models to the economy in an effective way. And that typically starts with price.
clean closing thesis line, works as a standalone soundbiteTikTok hook↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphoranalogy
So Fable 5 .1 and Mythos 5 .1 just dropped maybe three or four minutes ago, and I wanted to take a second and help you guys make sense of the updates, what they mean for you, and ultimately how you can use technology like this to improve your life, your business, and more. So the way that Fable 5 .1 stands benchmark -wise on probably the four or five most important benchmarks are as follows.
Fable 5 here and Opus 5 both scored around 24 .7 to 29 % on agent scientific research, terminal bench science 0 .1 specifically. Fable 5 .1 dropped and absolutely changed the game here. It scored over 50 % on this task, which is relatively recent as far as benchmarks go.
What we can expect from this, obviously, is significant improvements in an agent's ability to perform scientific research on loops like autoresearch without human supervision. And it's worth noting that it just freaking brutally mogs GPT -5 .6 Sol while doing it. This model is now available to everybody that is watching this video.
Likewise, agentic coding, it scores a new top of 55 .8 % compared to Fable 5's 42 and 52 .3, as well as GPT 5 .6 Sol's 37 .3%. So I would rate this not as a revolutionary upgrade, but as a significant improvement on, you know, yesteryear's or yesterday's generation of models. KnowledgeWork, which is currently scored using GDPVAL -AAV2, it scored 1853 compared to 1723 and 1824 for Fable 5 and Opus 5 respectively, a 1711 for GPT 5 .6 Sol.
What we see here is we see an improvement of about 7 % or so on this model's ability to understand the context of GDP or value creation, as well as its ability to actually go and do it. That's what GDPVAL -AA represents. Some big upsets here on computer use.
I think that the Anthropic series of models up until now probably lag behind mostly in computer use relative to OpenAI's GPT series. And we see here that they scored 77 .9 % using a... partial rating threshold on OS World 2 .0, which just means that it did most of what the initial task was on a browser.
It might have taken a circuitous way or sort of a more multi -step way to get there, but it eventually figured things out, versus a 72 .9 % and 75 .4 % on Opus 5. Multidisciplinary reasoning, we saw 60 .9 % versus a 57 .8 % and 56 .6%, another approximately 5 % improvement. The real upset here was for business workflows, which is something that is obviously near and dear to my heart as somebody that automates stuff.
Business workflows is assessed by automation bench. And essentially, this represents a model's ability to automate real pipelines and systems. that are installed in companies.
They use representative systems and workflows. Obviously, they don't feed it in a bunch of super real ones. But essentially, what this says is this is a benchmark that has effectively doubled from just last generation's Fable 5.
Fable 5 reliably automated business workflows 17 .1 % of the time. Fable 5 .1 reliably automates business workflows 31 .4 % of the time, which means we're likely about to see a lot more Fable 5 .1 -ing of businesses. People will probably be a lot more likely to want to sign up to the Claude Max or API pricing, pay Fable 5 .1 a few tens of dollars, let's say, and then leave with a fully automated workflow, probably generated in Python, or maybe even optimized in Rust or something like that, which is quite exciting.
And then finally, we have Agenda Coding, Cursor Bench 3 .2 .0. It scored 73 .4 % versus Fable 5's 70 .5 % and 70 % for Opus. So what's really interesting to me as somebody that has evaluated a lot of model drops in just the last year is that Opus 5 actually scored pretty damn good on most of these benches.
But you probably wouldn't have known it just talking to Opus 5, because personally, when I use Fable 5 versus Opus 5, I tend to get much higher quality results with Fable, despite the fact that Opus scores higher on benchmarks. And this takes me to, I think, just like a broader point, which is that despite the fact that we see awesome benches across the board for a lot of different models, including, you know, the recent suite of Chinese models that have been launched, like from Z .ai, Moonshot and stuff like that.
These benchmarks aren't actually always illustrative of the real experience of chatting with the models and actually having them do cool things. And so what we're going to see over the course of the next maybe 12 to 24 hours is we're going to see a suite of not like empirical or academic benchmarks, but real world human benchmarks where people, what I call taste benchmarks, where people throw, you know, cool tasks at the model to have it do things like design cool 3D worlds and so on and so forth, maybe design new ways of interaction and so on and so forth.
And I mentioned that just because I want you to keep all that stuff in mind. Benchmarks do not demonstrate everything, although obviously they're a relatively standardized way of doing so. And so in addition to being significantly better than Fable 5 across the board on all of these benchmarks, the thing that I think matters for a lot of people is no longer just, hey, can it do the task?
Because obviously, as we've seen, this is now sort of veering into or perhaps at AGI or artificial general intelligence levels. It's how much money is this going to cost me? And so what Anthropic did here is they publicized scores versus mean cost per task.
Worth noting, this is a log scale, which means it's an order of magnitude per kind of vertical marker here. And then they showed us that Fable 5 .1 scores remarkably higher for less money across the board on an average task at different effort levels. Okay, so Fable 5 down here, for instance, would score, let's say about, I don't know, 15 % or so for a mean cost per task of maybe $17.
Okay, at $17, Fable 5 .1 would score the equivalent of almost 40%. So mathematically, we do about 2 .5x better per dollar, at least according to this frontier graph, than Fable 5. So that tells me that you can expect approximately 2 .5x the, I want to say, cost efficiency per task.
And the reason why we're moving to cost per task as opposed to cost per token or whatever is because as the models get more intelligent, they also require different lengths of tokens in order to complete tasks. A model of yesteryear might technically have significantly lower input or output token pricing. But interestingly, it might also take way more tokens to complete a job.
Whereas new models are highly optimized, such that each token gets them a lot closer to the end result of, you know, in our case, completing a task. And so you're going to see cost per task probably be the predominant way that people measure how cost effective a model is moving forward. And so, you know, they know this.
And they've also changed the way that Fable 5 .1 hits cash tokens. So what are called cache reads, which is where you make multiple queries to Fable 5 .1 in quick succession, has changed significantly. Essentially, the cost of the model has gone down by about a quarter or so across the bar, and they're saying up to 45 % for highly agentic ones.
It's a little too early to know exactly what that means, but suffice to say, they have designed this model with cost in mind. It's not just a model meant to be better in... a sheer intelligence, but then cost way more to run.
This is a model over here that is both significantly more intelligent, and it's also cheaper, which takes me to just the trend of intelligence dropping costs greatly. You know, it wasn't more than a few years ago that we had GPT 3 .5 Turbo as like the smartest model. That was the intelligence that underlied the initial chat GPT drop back in 2022.
The equivalent cost of that intelligence okay, the same intelligence that you would get back then in terms of its ability to do things, you know, complete economically valuable tasks. Although the token pricing doesn't make that really clear, the equivalent cost of getting that intelligence has probably dropped by over 1000 times by now, meaning to get the same level of digital intelligence that you got with chat GPT on drop day, versus now, you can now pay 1000 times less.
I think model companies like this are understanding that they're probably far past the point at which models need to be intelligent in order to be rolled out for economic tasks, such that now the primary thing on most consumers' minds is not, hey, is this possible? It's now, is this profitable? And so Fable 5 .1 is a good move in that direction.
Finally, perhaps most importantly, what they've done is they've improved their safeguards a lot. This was one of the major concerns people had when using Fable. You would ask it a question about biosciences or, I don't know, something, maybe a little cybersecurity -y, maybe how to harden your app or whatever.
And a lot of the time it would say, we've detected that this flags our safeguards protocol, and so we have to roll you back to a dumber intelligence. Here's Opus 4 .8 or something like that. People really didn't like that, right?
What they've done with Fable 5 .1 is they've identified that being one of the most common problems that people have with the model. And they now flag benign requests, requests that sound like they might be a little bit spooky or like they contravene safeguards, but not actually contravene safeguards, 60 % less often. Specifically on biology and then medical questions, which I think a lot of people are starting to use models for, they've reduced the fallback rate by 85%.
So really big news. This brought to you by some random guy on Twitter. No, I'm just kidding.
Obviously take everything I'm saying here with a grain of salt. And I think what you'll find is as models continue to grow more and more intelligent, our ability and really the meaning of discriminating between a few points on a frontier graph tend to become less important. They also tend to just become less sensitive.
I mean, to be honest, if you blindfolded me and then had me fed the audio dictation of Opus 4 .8. Fable 5 and Fable 5 .1, would I really be able to tell the difference? Probably not, right?
So the point that I'm making there is model intelligence has just gone so, so good so recently, and so quickly rather, that eventually at a certain point, how intelligent the models get stops mattering. It's now more about how do we distribute these models to the economy in an effective way. And that typically starts with price, which you see anthropic optimizing for here.
Have a lovely rest of the day. I'll catch all y 'all in the next.
The Hook

The bait, then the rug-pull.

Anthropic's Fable 5.1 and Mythos 5.1 dropped minutes before this video was recorded, and the creator reads the benchmark chart live, line by line, before turning to what he thinks is the actual headline: cost.

Frameworks

Named ideas worth stealing.

06:13concept

Cost-per-task vs. cost-per-token

As models get more capable they often need fewer tokens to finish a task, so a model with a higher per-token price can still be cheaper end-to-end. Anthropic and the creator both frame Fable 5.1's improvement in terms of mean cost per completed task rather than raw token pricing.

Steal forComparing any two AI models or API price lists before switching
CTA Breakdown

How they asked for the click.

FROM THE DESCRIPTION
Storyboard

Visual structure at a glance.

cold open
hookcold open00:00
business workflows upset
valuebusiness workflows upset02:37
cost-per-task frontier graph
valuecost-per-task frontier graph04:59
sign-off
closesign-off10:18
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.

22:52
Nick Saraev · Tutorial

The Viral $1 Website Effect That Looks Like $10K

A five-step AI pipeline — generated painting, AI video, frame extraction, dithering, and free deployment — that turns a couple dollars of image and video credits into an animated, expensive-looking website hero.

July 29th