GPT-6 Astra: Full Review, Benchmarks, and Demos
An early-access hands-on with OpenAI's new GPT-6 Astra model — benchmark scorecard, alignment numbers, API pricing, and a run of 3D game and browser-automation demos.
September 3rdA live reaction to Anthropic's Claude Opus 5 launch, walking chart-by-chart through benchmarks that put a mid-tier-priced model ahead of Anthropic's own flagship on almost everything except cyber exploitation.
Claude Opus 5 launched priced the same as its predecessor but performing close to or above Anthropic's larger, pricier Fable 5 model, which the video argues proves cost-per-task — not price-per-token or raw benchmark score — is now the metric that actually matters when picking a model.
Anthropic released Claude Opus 5 at the same price as Opus 4.8 ($5/$25 per million input/output tokens) but with performance that beats Fable 5 — Anthropic's larger, more expensive model — on nearly every benchmark: agentic terminal coding, GDPval, ARC-AGI-3 (30% vs single digits before), BrowseComp, and OSWorld computer use. The one exception is cybersecurity exploit development, where Anthropic appears to have deliberately capped Opus 5's capability while still improving vulnerability-finding accuracy. Anthropic frames the release around cost-per-task rather than price-per-token, arguing that's the metric that actually predicts what a task costs to run, since a cheaper-per-token model can still burn more tokens completing the same task. Sponsor Box AI's own internal benchmark shows a similar jump on real document-grounded enterprise work. The reaction is unusually bullish: Opus 5 is described as outperforming Anthropic's own bigger model at half the price, comparable to GPT-5.6 Sol.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →
Berman opens on the claim that Opus 5 beats Fable 5 on price and performance, recaps Anthropic's model tier history (Haiku/Sonnet/Opus, then Fable), and pulls up the first benchmark table.

Walks the Opus 5 vs Fable 5 vs Opus 4.8 vs GPT-5.6 Sol table: ARC-AGI-3 jumps to 30%, BrowseComp improves slightly, Humanity's Last Exam is flat, OSWorld computer-use gains 4 points, DeepSWE dips slightly.

Frontier Code is flat, AutomationBench improves 9 points, but legal benchmark drops from 13.3% to 11.7% and health benchmark professional also declines — called out directly rather than skipped.

Pivots to Anthropic's efficiency framing — cost-per-task charts for OSWorld computer use and AutomationBench business workflows show Opus 5 sitting above and to the left of Fable 5, Opus 4.8, and GPT-5.6 Sol, meaning better score for lower cost. Explains why price-per-token comparisons (using Kimi K3 as the counterexample) are misleading.

Box AI segment — their own 12-industry document-grounded benchmark shows Opus 5 scoring 68% overall vs Opus 4.8's 63%, with the biggest gains on due diligence and data analysis.

Returns to benchmark charts — HLE cost/performance curve, Frontier-Bench agentic coding (Opus 5 nearly matches its own high-effort score for less), then explains how the ARC-AGI-3 game-based benchmark actually works using the interactive game board UI.

Novel-problem-solving cost chart shows the 30% ARC-AGI-3 jump again, then the OSS-Fuzz cybersecurity results: Opus 5 beats Opus 4.8 at finding vulnerabilities but stays far behind an uncensored model ("Mythos 5") at developing exploits — read as a deliberate Anthropic guardrail.

Switches to Anthropic's own blog post — pricing confirmed at $5/$25 per million tokens (same as Opus 4.8, half of Fable 5), Fast mode at 2.5x speed, mid-conversation tool switching, automatic safety-classifier fallbacks, and a physics demo of Opus 5 building a working wind-tunnel simulation.

Shows a tweet confirming Opus 5 briefly fell back to Opus 4.8 due to safety filters, revisits the ARC-AGI-3 and OSS-Fuzz charts, reads a tweet recommending Opus 5 as daily driver paired with Fable 5 for hard problems, notes three models are now on the cost/quality Pareto frontier, and signs off.
Anthropic's Opus 5 launch shows that a model priced the same as its predecessor can out-benchmark a pricier flagship, and the only way to see that clearly is by comparing total cost to finish a task rather than the sticker price per token.
“Better than Fable, half the price. This is Opus five.”
“It's no longer sufficient to just look at the price per token and think, okay, well, that's the price I pay... The main metric you need to be looking at is cost per task. That is it.”
“Opus 5 is stronger than Opus 4.8 on cybersecurity tasks, but it remains substantially behind Mythos 5 at developing exploits.”
“It's crazy to me that Opus is outperforming Fable. Does that mean the government is going to take it offline soon?”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
Matthew Berman opens mid-sentence, already stunned: Anthropic's new Claude Opus 5 is beating the company's own flagship Fable 5 on almost every benchmark, at half the price — and he spends the next twelve minutes screen-sharing the receipts.
Anthropic's efficiency charts plot benchmark score against total dollar cost per completed task rather than sticker price per million tokens, because a cheaper-per-token model can still cost the same or more if it needs more tokens to finish the same job — illustrated with Kimi K3, which is under half the price of GPT-5.6 Sol per token but uses roughly twice the tokens, netting to a wash.
“Thanks to the sponsor of this video, Box... Opus 5 is coming soon to Box AI.”
Mid-video sponsor read woven into the benchmark narrative — Box AI's own internal eval is presented as corroborating evidence for Opus 5's gains rather than a hard break in the video, then closes with Berman's personal endorsement ("we use it at forward future").
00:00
00:14
00:23
00:29
00:42
00:51
01:01
01:10
01:20
01:29
01:39
01:48
01:57
02:07
02:16
02:28
02:35
02:40
02:54
03:04
03:13
03:22
03:32
03:41
03:51
04:00
04:10
04:19
04:28
04:38
04:47
04:57
05:10
05:16
05:30
05:35
05:44
05:54
06:03
06:12
06:22
06:31
06:41
06:50
06:59
07:09
07:18
07:28
07:41
07:47
07:56
08:06
08:18
08:24
08:36
08:43
08:53
09:02
09:14
09:21
09:27
09:39
09:49
09:59
10:08
10:18
10:27
10:38
10:47
10:55
11:05
11:14
11:20
11:33
11:43
11:52
11:57
12:15
12:20
12:30Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
An early-access hands-on with OpenAI's new GPT-6 Astra model — benchmark scorecard, alignment numbers, API pricing, and a run of 3D game and browser-automation demos.
September 3rdA full walkthrough of Anthropic's twin release, the pricing math, the enterprise data compromise, and the independent benchmark that says the cost story doesn't add up.
September 1stThe mystery model that took over OpenRouter turns out to be a Chinese open-weights release that nearly matches frontier intelligence for a few cents a task.
August 29thEleven power-user habits from someone who has logged over a thousand hours in OpenAI's Codex CLI — model tiers, thread delegation, safety hooks, and remote control from a phone.
July 14thA 'dot' release plays out like a full generational leap: two five-to-seven-day unsupervised coding runs, a sponsor benchmark, and a live pricing and capability standoff against a rawer, higher-ceiling rival model.
July 9thIt's not a hidden mark in the text, it's a loaded die on word choice, and only a key Anthropic won't hand out can read it.
September 2nd