Claude Sonnet 5 just dropped. I'm changing how I use AI
A 12-minute breakdown of Sonnet 5 benchmarks, a live ChatGPT 5.5 head-to-head, a two-model workflow for Claude Code, and leaked Fable 5 strings.
June 30thA benchmark-by-benchmark walkthrough of Anthropic's Fable 5.1 release, plus a custom test suite proving it beats GPT-5.6 Sol and demolishes Fable 5.
Claude Fable 5.1 costs the same per token as Fable 5 but finishes tasks in fewer tokens, refuses far less, and shows its biggest capability jump in research-style tasks, signaling a push toward models that generate novel ideas rather than just execute known ones.
Alex Finn walks through Claude Fable 5.1's release, arguing it's a clear step up from Fable 5 in every category, with the largest gains in scientific-research benchmarks. Token pricing stays the same, but tasks finish in 25-40% fewer tokens, and the model's content filters trigger far less often, letting it complete tasks Fable 5 used to refuse outright. Head-to-head demos show Fable 5.1 beating GPT-5.6 Sol on a 3D build test, a debugging task, and an Easter-egg hunt, and producing a noticeably closer pixel-perfect clone of apple.com than Fable 5 did. He closes with three ways to use it immediately: audit your AI agent setup, rerun old brainstorms for new angles, and build your own benchmark test suite instead of trusting public leaderboards.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →
Alex Finn opens by calling Fable 5.1 a clear step up from Fable 5 in every way, and warns that using it wrong wastes money.

A quick look at official benchmarks shows normal point-release bumps are 5-10%, but Fable 5.1 doubles some scores, with the biggest jump in scientific research.

Token pricing stays the same as Fable 5, but the model finishes tasks in far fewer tokens, cutting real-world cost by 25-40%.

Alex lists the four changes he cares about: substantially cheaper at the same price, substantially smarter, way better content filtering, and built specifically for research and novel ideas.

Before running his personal tests, Alex frames why he doesn't trust official benchmarks and previews his own benchmark suite.

Fable 5.1 beats GPT-5.6 Sol on a 3D roller coaster build test with noticeably more visual detail, needs fewer tool calls on the debug duel, and edges out on the Easter-egg gauntlet task; it also demolishes the older Fable 5 on every test.

Fable 5 used to refuse the relic-hunt Easter-egg task on content-filter grounds; Fable 5.1 completes it. The pixel-perfect apple.com recreation is dramatically closer to the real site than Fable 5's version, which turned people and logos into abstract shapes.

Export a markdown description of your current AI agent, skills, and sub-agent setup from your main orchestrator, then feed it to Fable 5.1 and ask it to find ways to improve the setup.

Alex reruns prior Fable 5 brainstorming sessions through Fable 5.1 and gets genuinely new angles, arguing the model was tuned to act like a research intern capable of first-principles thinking, the same direction OpenAI is pushing with Astra.

Rather than trust public benchmarks, Alex has Fable 5.1 identify his five most common task categories, design a multi-step test for each, and build a website to track model performance on his own use cases going forward.

Alex predicts a string of new Anthropic releases, then pitches his Vibe Coding Academy community with a 24-hour pricing deadline.
The real signal in a model upgrade is which specific axis moved (cost per task, refusal rate, a named specialty like research) and whether that axis matches your own workload, not a single leaderboard number.
“if you don't use it the right way, you're going to waste tons of money”
“it is doubling the percentages in some ways”
“you might blow up the world with that question”
“Fable 5 absolutely refused to do the relic hunt... that's not the case with Fable 5.1”
“I think what these AI companies are going for is coming up and building their own research interns so they can start the recursive self-improvement loop”
“you shouldn't trust other people's benchmarks... build your own”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
Alex Finn opens by promising Fable 5.1 is a clear step up from Fable 5 in every way, but warns that using it the wrong way burns money fast, then spends twelve minutes proving it with his own benchmark suite, a side-by-side apple.com clone test, and three concrete ways to put it to work today.
Alex's personal breakdown of what actually changed in the release, shown on-screen as a numbered list.
Alex's own five-category custom benchmark website, run against every new model release instead of relying on public leaderboards.
“Join the Vibe Coding Academy. Link down below. Number one community on AI... Special pricing for the next 24 hours. Then it's done. It's going back up permanently.”
Classic urgency close (24-hour pricing deadline) paired with a like/subscribe ask right before it.
00:00
00:14
00:23
00:32
00:42
00:51
01:00
01:10
01:19
01:28
01:40
01:47
01:56
02:06
02:16
02:24
02:34
02:43
02:52
03:02
03:11
03:20
03:30
03:39
03:48
03:59
04:07
04:17
04:26
04:35
04:43
04:54
05:03
05:12
05:22
05:31
05:40
05:50
06:00
06:08
06:22
06:24
06:36
06:42
06:55
07:04
07:14
07:23
07:32
07:42
07:51
08:00
08:10
08:23
08:28
08:38
08:47
08:56
09:06
09:15
09:24
09:34
09:43
09:52
10:02
10:11
10:20
10:30
10:39
10:48
10:58
11:07
11:16
11:26
11:35
11:44
11:54
12:03
12:12
12:22Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
A 12-minute breakdown of Sonnet 5 benchmarks, a live ChatGPT 5.5 head-to-head, a two-model workflow for Claude Code, and leaked Fable 5 strings.
June 30thA leaked OpenAI blog post says GPT-6 Astra is AGI, beats Claude Fable 5.1 on every benchmark, and crosses a cybersecurity threshold that lets it hack on its own.
September 3rdAlex Finn runs both models through five self-designed benchmarks, crowns Opus 5 the winner on price and quality, then spends the back half explaining why he's not fully switching.
July 24thA feature-by-feature walkthrough of Hermes Agent's newest release — mixture of agents, one-command skill learning, and a vibe-coding tool with real git controls.
July 6thA tour of the open source, AI-first Linux desktop that lets you edit the operating system itself just by asking an agent.
August 28thA hands-on tour of Grok Bot's zero-config, cloud-first multi-agent setup, and the six-bot roster running one founder's day-to-day business.
August 17th