Kimi K3 Is Here! (Better Than Opus 4.8?)
Ten identical builds, five models, blind-ranked before the reveal — a real-world stress test of Moonshot AI's new open-source model against GPT-5.6 Sol, Opus 4.8, GLM 5.2, and its own predecessor.
July 17thA blind, five-round test pits Opus 5 against Fable 5 and Opus 4.8 across web design, 3D, games, motion graphics, and a SpaceX investment deck — model names stay hidden until the ranking is locked in.
A custom blind-ranking tool that hides model identities until after judging shows Opus 5 beating Fable 5 in all five creative and knowledge-work tests, at roughly half Fable's price.
A YouTuber runs three AI models — Opus 5, Fable 5, and the older Opus 4.8 — through five identical creative prompts: a showcase website, a real-time 3D scene, a playable browser game, a motion graphics piece, and an investment-analysis deck on SpaceX. Using a custom blind-compare tool that hides which output came from which model, he ranks each result before revealing the source. Opus 5 wins all five rounds, matching or beating Fable 5's visual polish and technical depth while costing about half as much per generation — and sometimes undercutting Opus 4.8 too. His conclusion: unless there's a hidden reason for the price gap, the benchmarks weren't exaggerated this time.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →
Host previews the test: three frontier models — Opus 5, Fable 5, Opus 4.8 — run through five identical creative prompts, blind-ranked before reveal.

Walks Anthropic's own published benchmark charts for Opus 5 vs Fable 5 vs Opus 4.8, questioning why a cheaper new model would beat the flagship across agentic coding, knowledge work, and computer use.

Each model builds one 'ceiling of web design capability' single-file HTML site; all three land on a space theme. Blind-ranked in a custom compare tool before reveal.

Models build a real-time 3D/simulation experience — black hole ray tracers and an interactive underwater jellyfish scene — ranked blind then revealed.

Each model builds a playable browser game meant to be fun within thirty seconds — a space shooter, an asteroid-style shooter, and a boat survival game.

Models generate a self-contained animated motion-graphics piece — a Dante-themed title sequence, a data-story explainer, and an abstract signal/noise piece.

Models act as a private wealth advisor building an investment-committee deck answering 'should I buy SpaceX?' — testing research depth, forecasting, and deck design.

Opus 5 wins all five blind rounds against Fable 5 and Opus 4.8, at roughly half Fable's cost — the host concludes the benchmarks weren't exaggerated.
Hiding which model made which output until after ranking is the real method here — it turns a vendor benchmark chart into an actual verdict, and it's a technique you can copy for any product comparison, not just AI models.
“Anthropic is just losing money. They're basically telling you there's no need to use Fable anymore.”
“Look at these f's. Dead giveaway. That that is AI. Just that curly f.”
“Opus five was across the board better at three d, creative web design, game development, motion graphics, and knowledge work... for half the cost of Fable five.”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
A blind five-round test — web design, 3D, a playable game, motion graphics, and an investment deck — pits Opus 5 against Fable 5 and the older Opus 4.8, with the model names hidden from the reviewer until every rank is locked in.
A custom tool that runs the identical prompt across multiple models, hides which output came from which model, lets the reviewer rank blind, then reveals identities only after ranks are locked.
“Let me know what you think of all of this in the comments, plus any ideas you have for future tests.”
Soft — no hard pitch spoken on camera. The real CTAs (bootcamp waitlist, newsletter, blog post with every prompt and generation) live entirely in the description, not narrated.
00:01
00:21
00:42
01:12
01:33
01:57
02:15
02:28
02:56
03:23
03:32
04:01
04:13
04:40
05:01
05:22
05:41
05:56
06:24
06:45
07:06
07:27
07:48
08:08
08:27
08:50
09:11
09:30
09:45
10:17
10:26
10:45
11:19
11:36
11:49
12:18
12:49
12:50
13:25
13:43
13:58
14:16
14:37
14:57
15:29
15:46
15:58
16:18
16:40
17:04
17:23
17:41
18:16
18:23
18:51
19:05
19:24
19:51
20:09
20:30
20:51
21:15
21:40
22:08
22:11
22:42
23:03
23:24
23:44
24:05
24:18
24:47
25:09
25:28
25:54
26:10
26:36
26:42
27:03
27:43Ten identical builds, five models, blind-ranked before the reveal — a real-world stress test of Moonshot AI's new open-source model against GPT-5.6 Sol, Opus 4.8, GLM 5.2, and its own predecessor.
July 17thA blind, four-way bake-off — GPT-5.6 Sol against Fable, Opus 4.8, and GPT-5.5 — across ten builds and knowledge-work tasks, scored one task at a time without knowing which model made what.
July 10thSame prompts, three one-shot builds, two frontier coding agents — and one surprisingly clear winner.
July 11thA creator builds daemon, a single Mac app that runs every AI coding agent he owns, by giving Claude Fable 5 four rounds of blunt feedback instead of writing a line of the UI himself.
July 7thA creator locks a glassmorphism design first, routes all implementation to Opus sub-agents, and ships a working iPhone app before he runs out of usage credits.
July 3rdAn 8-step agentic pipeline that takes you from naive AI slop to a pixel-near Linear replica, deployed to Vercel with an MCP server, in under 20 minutes.
June 8th