I Spent $400 Benching Opus 5 — Here's What It Can Do
A tour of a dozen apps Claude Opus 5 built from scratch, followed by Anthropic's own benchmark numbers on coding, computer use, and cost.
July 24thA same-day walkthrough of Claude Fable 5.1's benchmark chart, why the real story is cost-per-task rather than raw score, and what the new safeguard numbers mean for how often the model refuses benign questions.
Fable 5.1 posts solid but not revolutionary score gains over its predecessor, and the more consequential change is that it delivers roughly 2.5x more capability per dollar, which the creator argues now matters more than raw intelligence.
Fable 5.1 and Mythos 5.1 launched with benchmark gains across agentic scientific research, coding, knowledge work, computer use, reasoning, and especially business-workflow automation, which nearly doubled from 17.1% to 31.4%. But the creator argues the real story is cost: on Anthropic's accuracy-vs-cost frontier chart, Fable 5.1 delivers roughly 2.5x more capability per dollar than Fable 5, and a cut to cache-read pricing lowers real workload cost by 25-45%. He also flags that benchmark scores don't always predict real-world quality (Opus 5 scores well but underdelivers for him in practice), and closes on Anthropic's claim that Fable 5.1 flags benign requests as safeguard violations 60% less often, with an 85% drop in false refusals on biology and medical questions.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →
Fable 5.1 and Mythos 5.1 just dropped; the creator promises to make sense of the updates and their practical impact.

Reads the comparison table: Fable 5.1 scores over 50% on agentic scientific research (vs. 24.7-29% for Fable 5/Opus 5), 55.8% on agentic coding, and 1853 on the GDPVal-AA knowledge-work benchmark.

Fable 5.1 scores 77.9% on OSWorld 2.0 (partial-credit computer use) and 60.9% on Humanity's Last Exam, both frontier gains over Fable 5 and Opus 5.

AutomationBench score nearly doubles from Fable 5's 17.1% to Fable 5.1's 31.4%; agentic coding on Cursor Bench 3.2.0 also improves to 73.4%.

Opus 5 scores well on paper but the creator says he gets lower-quality real results from it than from Fable; flags that informal 'taste benchmarks' will matter over the next day.

Anthropic's log-scale accuracy-vs-cost frontier graph shows Fable 5.1 scoring about 2.5x higher per dollar than Fable 5 at the same mean cost per task.

Cache reads cost about 75% less on Fable 5.1, cutting real-world workload cost by roughly 25% typically and up to 45% for highly agentic workloads.

The intelligence behind ChatGPT's 2022 launch now costs roughly 1000x less; the deciding question is shifting from 'can it do this' to 'is this profitable'.

Anthropic's post says benign requests are flagged as safeguard violations 60% less often, with an 85% drop in fallback rate on biology/medical questions.

Closes on the idea that as models converge in quality, distributing them effectively and pricing them right matters more than incremental intelligence gains.
Fable 5.1 posts solid benchmark gains, especially on business-workflow automation, but the bigger shift is that it delivers roughly 2.5x more capability per dollar, which the creator argues matters more now than raw intelligence.
“It's worth noting that it just freaking brutally mogs GPT-5.6 Sol while doing it.”
“This is the real upset here was for business workflows, which is something that is obviously near and dear to my heart as somebody that automates stuff.”
“So mathematically, we do about 2.5x better per dollar, at least according to this frontier graph, than Fable 5.”
“It's now more about how do we distribute these models to the economy in an effective way. And that typically starts with price.”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
Anthropic's Fable 5.1 and Mythos 5.1 dropped minutes before this video was recorded, and the creator reads the benchmark chart live, line by line, before turning to what he thinks is the actual headline: cost.
As models get more capable they often need fewer tokens to finish a task, so a model with a higher per-token price can still be cheaper end-to-end. Anthropic and the creator both frame Fable 5.1's improvement in terms of mean cost per completed task rather than raw token pricing.
00:00
00:11
00:19
00:27
00:34
00:42
00:50
00:58
01:06
01:13
01:21
01:29
01:37
01:44
01:52
02:00
02:08
02:16
02:23
02:31
02:39
02:47
02:54
03:02
03:10
03:18
03:26
03:33
03:41
03:49
03:57
04:04
04:12
04:20
04:28
04:36
04:43
04:51
04:58
05:07
05:14
05:22
05:30
05:38
05:45
05:53
06:01
06:09
06:17
06:24
06:32
06:40
06:50
06:55
07:03
07:11
07:19
07:27
07:34
07:42
07:50
07:58
08:05
08:13
08:21
08:29
08:39
08:44
08:52
09:00
09:08
09:15
09:23
09:29
09:39
09:47
09:54
10:02
10:10
10:18Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
A tour of a dozen apps Claude Opus 5 built from scratch, followed by Anthropic's own benchmark numbers on coding, computer use, and cost.
July 24thFive hours, thirty-six sections, and one argument: clear thinking is an engineering problem, not a discipline problem.
September 2ndA four-and-a-half hour screen share that takes four business systems from a single typed prompt all the way to a scheduled job running on somebody else's server.
August 23rdSix hours of screen-share in which one operator rebuilds an entire marketing department out of prompts, skills, loops and scheduled cloud routines.
August 8thA five-step AI pipeline — generated painting, AI video, frame extraction, dithering, and free deployment — that turns a couple dollars of image and video credits into an animated, expensive-looking website hero.
July 29thA single terminal command turns a text idea into a full scroll-driven cinematic website — no designer, no video editor, one AI coding agent running the whole pipeline.
July 22nd