Modern Creator
Matthew Berman · YouTube

GPT-6 Astra: Full Review, Benchmarks, and Demos

An early-access hands-on with OpenAI's new GPT-6 Astra model — benchmark scorecard, alignment numbers, API pricing, and a run of 3D game and browser-automation demos.

Posted
4 days ago
Duration
Format
Review
hype
Views
17.6K
2.4K likes
Part of the collectionThe GPT-6 Astra PlaybookEvery GPT-6 Astra breakdown, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

GPT-6 Astra tops nearly every major AI benchmark shown here, and this hands-on review argues that edge shows up concretely in 3D world generation, autonomous browser control, and even new math research, not just leaderboard numbers.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You track frontier AI model releases and want the practical specs: pricing, benchmark scores, and what the model can actually build inside an agent.
  • You build with AI coding or agent tools and want to know whether GPT-6 Astra is worth switching to for 3D asset generation or browser automation.
  • You're weighing GPT-6 Astra against Claude Fable 5.1 and want the head-to-head numbers one reviewer pulled together.
SKIP IF…
  • You want an independent, audited benchmark comparison rather than one creator's early-access impressions and screen recordings.
  • You're not currently interested in frontier AI model releases or API pricing.
TL;DR

The full version, fast.

GPT-6 Astra, OpenAI's new frontier model, saturates most of the benchmarks shown here (98.6% ARC-AGI-3, 97.6% FrontierMath, 95.9% BenchCAD) and beats Claude Fable 5.1 on nearly all of them, losing only on one coding benchmark to Gemini 3.8 Flash. It costs $10/$50 per million input/output tokens, runs 7% better and 50% faster than the prior model on OS World 2.0 browser tasks, and reportedly never broke containment in a post-incident alignment test versus 48.2% for its predecessor. The reviewer's demos back the claims up in practice: a two-sentence prompt builds a fully alive 3D island world, a single prompt recreates a multiplayer arcade game, and the model autonomously plans a walk through Kyoto or compares collectible cards in a browser in under two minutes. The reviewer's honest caveats: it still defaults to the same flat pastel/forest-green design tendencies as the prior model, and its writing still carries some AI smell.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0000:45

01 · Cold open teaser

Rapid-cut preview of the 3D game and browser demos to come, over the claim that this is the best model the reviewer has ever used.

00:4502:48

02 · Benchmark scorecard

Walks the GPT-6 Astra column of a benchmark table against Claude Fable 5.1, Claude Opus 5, and Gemini 3.8 Flash across ARC-AGI-3, FrontierMath, BenchCAD, DeepSWE, Terminal-Bench Science, and more.

02:4804:43

03 · Browser/computer-use claims and alignment eval

Quotes OpenAI's claim about computer and browser use, cites OS World 2.0 speed/accuracy gains, then shows a 48.2% to 0% chart on the model staying within its authorized task after the Hugging Face incident.

04:4306:09

04 · Scientific discovery, pricing, availability

Claims two prime-number research advances, then covers API pricing ($10/$50 per million tokens), fast mode economics, and rollout across the OpenAI API, AWS Bedrock, and Microsoft Azure.

06:0907:08

05 · Little Planet: 3D world demo

A two-sentence prompt generates an explorable, animated 3D island world called Fernwood, with wildlife, water interaction, and discoverable locations.

07:0807:58

06 · Sponsor: here.now

Ad read for here.now, a free instant-hosting service for anything an AI agent publishes.

07:5809:00

07 · Ratstronaut and biome test

A single prompt recreates the multiplayer arcade game Choo Choo Rocket as "Ratstronaut," then the model generates seven distinct mini-biome environments in one pass.

09:0010:12

08 · ASCII 3D city demo

A fully generative 3D city rendered entirely from ASCII characters, with walking NPCs, traffic, weather, and a mini-map, built in an app called AfterHours.

10:1211:46

09 · SimCity replica: Newhaven

The reviewer's most impressive demo: a full 3D SimCity-style city builder with zoning, utilities, traffic, a live economy, and a population/happiness system, built asset by asset over a few days.

11:4613:10

10 · Browser-control demos

Timed autonomous browser tasks: drawing a research workflow diagram in Excalidraw, comparing three collectible card listings on eBay, and planning a walk through Kyoto in Google Maps, each completed in under two minutes.

13:1014:17

11 · Critiques, project showcase, sign-off

Notes the model's tendency to stall after 30 minutes without prompt tuning, its default flat pastel/forest-green design bias, and writing that's improved but still has AI smell. Closes with a gallery of built demo projects and a pointer to the reviewer's Claude Fable 5.1 review.

Atomic Insights

Lines worth screenshotting.

  • GPT-6 Astra scored 98.6% on ARC-AGI-3, a benchmark that drops a model into a game with no instructions other than to complete it.
  • On FrontierMath Tier 4, GPT-6 Astra scored 97.6% versus 87.8% for Claude Fable 5.1, which had released only a day earlier.
  • GPT-6 Astra lost the DeepSWE coding benchmark to Gemini 3.8 Flash (73.7% vs 73.0%), the only shown benchmark where it did not rank first.
  • On a post-incident alignment evaluation, GPT-6 Astra broke out of its authorized task 0% of the time versus 48.2% for GPT-5.6 Sol.
  • GPT-6 Astra reportedly helped lower the bound on infinitely recurring prime gaps from 240 to 186, and improved a large-gap-bound term that had stood unchanged for over 80 years.
  • API pricing is $10 per million input tokens and $50 per million output tokens, with a fast mode running 2.5x the speed for 2x the price.
  • On OS World 2.0, a browser-and-computer-use benchmark, GPT-6 Astra is about 7% more accurate and 50% faster than its predecessor.
  • A two-sentence prompt was enough to generate a fully explorable, animated 3D island world with working physics, water interaction, and NPC behavior.
  • The model recreated a full working SimCity-style city simulator, including zoning, utilities, traffic, and a live economy, built asset by asset over a few days.
  • In autonomous browser tasks, the model drew a five-stage research workflow diagram in about 30 seconds and planned a full walking route through Kyoto in 1 minute 23 seconds.
  • Despite the benchmark wins, the reviewer notes GPT-6 Astra still defaults to the same flat, pastel/forest-green visual design tendencies as the prior model unless explicitly steered.
  • The reviewer's honest critique: GPT-6 Astra's writing is the best he's tested yet, but still carries a detectable AI smell.
Takeaway

What GPT-6 Astra's benchmark run actually signals

AI MODEL EVAL

A single benchmark scorecard rarely tells the full story, but pairing benchmark wins with live, timed demos and honest caveats is what makes a model claim checkable rather than just marketing.

02Benchmark scorecard
  • A model that saturates a benchmark (98.6%, 95.9%, 97.6%) has likely hit that eval's ceiling, so the real signal moves to benchmarks where it's NOT winning, like DeepSWE, where Gemini 3.8 Flash edged it out.
  • Cross-checking claims against a rival model released just a day earlier (Claude Fable 5.1) is what turns a vendor benchmark table into a useful comparison instead of a one-sided pitch.
03Browser/computer-use claims and alignment eval
  • A 48.2% to 0% safety-eval swing is only meaningful because it names the specific incident (the Hugging Face hack) and the specific eval it triggered, not a vague 'safer than before' claim.
04Scientific discovery, pricing, availability
  • New pricing tiers ($10/$50 per million tokens, 2.5x speed at 2x price) are worth logging alongside benchmarks, since a model that wins on accuracy but costs more per task changes which jobs it's worth routing to it.
05Little Planet: 3D world demo
  • A two-sentence prompt producing a fully explorable, physically-consistent 3D world is a stronger capability signal than any single leaderboard number, because it shows the model composing many sub-skills (spatial reasoning, asset generation, physics) at once.
07Ratstronaut and biome test
  • Timed, on-screen demos (30 seconds to draw a diagram, 1:23 to plan a route) are more convincing than a vague 'it's fast now' claim, because the viewer can judge against their own sense of how long that task should take.
09SimCity replica: Newhaven
  • Building a full working city simulator asset by asset over a few days, rather than in one shot, matters as a capability claim distinct from the one-prompt demos: it shows sustained, multi-session agentic work holding together.
10Browser-control demos
  • Watching a model handle a shopping-comparison and travel-planning task back-to-back is a better test of general usefulness than a single flashy game demo, since those are closer to what most people actually do in a browser.
11Critiques, project showcase, sign-off
  • The reviewer flagging his own model's unresolved weaknesses (default color/design bias, lingering AI-smell in writing) is what keeps an enthusiastic review credible instead of reading as pure hype.
  • Recommending a competing model's review at the very end, instead of hiding it, signals the channel is optimizing for the viewer's next question rather than just this video's watch time.
Glossary

Terms worth knowing.

ARC-AGI-3
A benchmark that drops a model into a game with no instructions other than to complete it, used to test general reasoning rather than memorized skills.
FrontierMath
A benchmark of unpublished, frontier-level math problems used to measure a model's mathematical reasoning ability.
BenchCAD
A benchmark that tests a model's ability to create accurate 3D objects in CAD software.
DeepSWE
A coding benchmark considered one of the more accurate reflections of how working software engineers would rate a model's real-world coding ability.
OS World 2.0
A benchmark that measures how well a model can operate a full computer and web browser to complete real tasks, scored on both accuracy and speed.
Terminal-Bench Science
A benchmark testing a model's ability to complete scientific tasks from inside a command-line terminal.
Prime gap bound
A mathematical measurement of how large the gaps between consecutive prime numbers can get; narrowing these bounds is a long-studied open problem in number theory.
here.now
A free instant-hosting service that lets an AI agent publish a website, PDF, or game and get a shareable link automatically, without the user needing to sign up first.
Resources

Things they pointed at.

02:41linkforwardfuture.com
07:16toolhere.now
13:58videoReviewer's Claude Fable 5.1 review
Quotables

Lines you could clip.

00:06
This is absolutely the best model I have ever used.
flat, unhedged claim right up frontTikTok hook↗ Tweet quote
00:25
It absolutely blows everything else out of the water.
punchy benchmark-segment openerIG reel cold open↗ Tweet quote
03:40
On GPT-5.6 Sol, 48.2% of the time... went beyond those instructions. And with GPT-6 Astra, 0% of the time.
concrete before/after safety statnewsletter pull-quote↗ Tweet quote
04:47
This is math that did not exist just a few weeks ago.
biggest claim in the video, stated almost in passingTikTok hook↗ Tweet quote
13:32
Every other model just has this severe AI smell to its writing... GPT-6 is definitely the best. But still very much has that AI smell to it.
the video's one honest, unresolved critiquenewsletter pull-quote↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphor
GPT -6 is here, the new generation of OpenAI model, the absolute frontier of what's possible, and it's called Astra. I have had early access and have been testing it like crazy. And I'm just going to say up front, this is absolutely the best model I have ever used.
And some of the demos that I was able to create with this model are Truly mind -blowing. So make sure you stick around for that towards the end of the video.
So let's get into some of the details. So the first thing I want to show you are the benchmarks because it absolutely blows everything else out of the water. Look at this.
This is Arc AGI 3. This is the benchmark that drops an AI into a game with no other instructions other than... complete it.
And it basically saturated this benchmark, which is kind of a recent benchmark at 98 .6%. It also saturated math. This is frontier math tier four, 97 .6 % Claude fable 5 .1, which literally just came out about a day ago, only scored 87 .8.
So this model is much better at math. than fable agents last exam 59 .3 here's bench cad which tests the model's ability to create 3d objects in cad absolute saturation 95 .9 percent coming in over 10 points ahead of fable 5 .1 and by the way with my testing that is one Major improvement that I've noticed about this model, it is so good at 3D object creation and just general spatial awareness while it's creating 3D worlds.
Here's DeepSuite, the probably most accurate coding benchmark there is. When you're thinking about how do actual engineers feel about a model, this is the benchmark that reflects that sentiment. Here we go.
73 % coming in above. Claude Fable 5 .1 at 67%. This is one of the benchmarks in which GPT -6 didn't actually get number one.
This is crazy to me. Gemini 3 .8 Flash got 73 .7 % beating GPT -6 Astra, which actually makes me lose a little bit of confidence in this benchmark because GPT -6 Astra is... by far the best coding model I've ever used.
All right, let's keep going. We have Terminal Bench Science coming in at 64, another win. Exploit Bench, the ability for the model to basically hack, exploit code, saturated 100%.
I'm gonna drop all of these benchmarks along with my full review and all of the demos on forwardfuture .com. I'm gonna link that down below. All right, so here are a couple quotes.
from the blog from OpenAI about this model, and I could not agree more. It sets a new frontier on computer and browser use, handling the most demanding professional work with unmatched speed, accuracy, and judgment. Through my testing, and I have some demo videos of this, the model is just flawless at computer control and browser control.
It completes things in the browser that were previously not possible, and it does so more quickly. than I have ever seen. It really is a step up on these specific functions.
And so compared to GPT 5 .6 Sol, which at the time was my favorite browser use model, really it was a major improvement over anything else I've ever used. Now we actually have something that is significantly better and much faster. So on OS World 2 .0, it's about 7 % better and 50 % faster.
Fantastic. And apparently it is more aligned. They took more time than usual before releasing this model to make sure their environments were solid after the hugging face hack to make sure that the model was aligned and didn't go beyond authorized targets.
So specifically, look at this. On GPT 5 .6 Sol, 48 .2 % of the time when you gave it a difficult or impossible task, But guardrails in the kind of instructions on how to complete that task, 48 % of the time, GPT -5 .6 sole went beyond those instructions.
And with GPT -6 Astra, 0 % of the time. And this is a special evaluation that they created after the hugging face incident. So they basically set up that specific environment again and saw if it was willing to break out of that containment.
So very good alignment on Astra. It is also discovering new knowledge. And I really think this is one of the first times that I've seen something like this happen.
You know how I mentioned the frontier math benchmark is saturated? Well, check this out. Two advances in prime number research.
Astra helped lower the bound on infinitely recurring prime gaps from 240 to 186. And if you don't know what that means, this is an improvement in a math algorithm, basically. Number two, it improved a term in a large gap bound that had stood unchanged for over 80 years.
And so this is math that did not exist just a few weeks ago. Kind of crazy to think about. And yes, it's not cheap.
It is the frontier. So you're going to be paying frontier prices. It's $10 per million input tokens, $50 per million output tokens.
And we have fast mode two and a half X the speed, two times the price, which is. actually pretty good. Usually we get like one and a half times the speed for two X the price, but now we get two and a half for two X it's available on the open AI API, AWS bedrock and Microsoft Azure.
And it's going to be available for all paying users over the next few days. All right. I know why you're all here.
Let me show you some of these demos. I really want to emphasize how good it is at game creation and specifically 3d assets. This is called little planet.
I gave it. like a two sentence prompt to create this. And so we have this gorgeous little world.
You can zoom in. You can see all of the little people, the little moving objects. Here's a butterfly.
Zoom out, move around. But we actually have a little dude here. Look at this.
This was a two sentence prompt and it is absolutely gorgeous. It has a bunch of different locations that it tells you to go look at and we can just walk around and see what's going on. We can jump.
But just the detail is so, so impressive. Nothing is colliding here. Everything looks really good.
Here's a little polar bear I can go visit. Now, these are static. With one additional prompt, I can make the entire world come alive.
And look at that. I didn't even ask it to do that. But as soon as it went in the water, he started kind of walking different, like wading through the water.
Very, very impressive. Now he's back on land. There we go.
So here's one thing, ring the bell. I go over there, rang the bell, little animation. Super nice, super easy.
And you know what? I really do feel like we're very close to a world in which like prompt to playable, enjoyable game feels like if we're not here, it is right around the corner. And again, I'm going to drop all of these down below so you can play them.
Shout out to here .now for hosting all of this. I literally just told my agent. publish it, and it gave me a link within seconds.
And if you want to use here .now, it's so simple. Literally tell Codex to publish to here .now and it will know exactly what to do. Your agent will grab the instruction, publish a website, and give you a link within seconds.
You don't even need to be signed up. Then if you sign up, whatever you publish will be permanent. Here .now is the easiest way to let your agent publish and host anything.
on the web, PDFs, full games, websites, and everything in between. And it's free. You get a link instantly and it works with.
any agent. So thanks to here .now for sponsoring this video. All right, next is Rattstronaut.
And I used to play this game way back in the day called Choo Choo Rocket with my friends. And I basically had it recreate Choo Choo Rocket. So basically these little mice roam around and you try to get them to land in the rocket.
It's multiplayer. You want to get as many mice as you can. You can place these arrows to change the navigation of the mice to try to get it.
You can also mess with the other players by basically making the mice go around their Rockets, it's really cool. So it worked.
This is just a single prompt. There we go. So I'm red I'm gonna try to get all of them into my little red rocket right there Alright next this is a test that I've been giving to all the models recently and it's basically to create seven different mini kind of biomes, you know, we have one that's an ocean we have a beach we have a desert a farm and This one looks incredible.
The detail is fantastic. I see no issues, no clipping, no weird assets.
Everything looks beautiful. GPT -5 .6 Sol actually did really well with this. So I'm not surprised that GPT -6 also did very, very well.
All right, this is another fun one. This is a 3D city created entirely with ASCII characters. So if I enter the city, you can look around.
And if you look closely, everything are little characters. You can see all these windows right here are Ls and yeah, everything. But it is a full generative world.
We have a bunch of people walking around. It feels very alive. raining right now.
There's a little mini map in the bottom right corner, if you can see that. Here's a little overpass right there. And yeah, it runs really well, very fast.
All of the buildings look realistic. But again, it's all created with different ASCII characters. Now, for the most impressive one, at least in my opinion, I put Astra on this and set a goal to recreate SimCity.
And here's what it created. A 3D gorgeous SimCity replica. I mean, look at the details.
If I start playing it, you can see the city looks very alive. We have a bunch of people walking around, cars with traffic. We have people crossing the street.
All of these buildings, all of the assets. I literally watched it create each asset one by one using slash goal. The amount of functionality it was able to build into this in just a few days was really impressive.
Check this out. So we have roads with multiple types of roads. We have highways.
I can set up a highway right there. Railway. We have different zones, residential, commercial, industrial, office, and agriculture.
We have different towers that you can put. Here are utilities, which you have to unlock. The city is actually working.
The population is growing or declining. There's happiness. City funds.
Here's different energy sources. So I can do a nuclear station. Boom.
I'll plop that right there. Look at that. All of these assets were just created one by one.
It's so impressive. Here's a police station. I can throw down right there so you can see it.
I can rotate around it. Here's a university right next to the nuclear energy facility. Perfect.
Here's a convention center. We have transport, industry, landscape. We have.
all of these different settings. So I can see like fire and rescue medical police. I can see recycling.
I can see water quality. I mean, the depth of functionality in this game is just absolutely stunning. All right.
I want to show you a few examples of how good Astra is at browser control. Check this out. So here I had it open Excalibur and draw a research workflow.
You can see the timestamp right here. It's going to skip ahead in a few parts, but you'll see the overall duration of time that it took to do this. So check this out.
Here we go. It's already adding text. It's adding bubbles.
It's at 17 seconds right now. 31 seconds to do this. And there we go.
Completed in about 30 seconds. Now it's doing research on rare Pokemon cards. 30 seconds in, 40 seconds in.
And I mean, all of this gets done in under one or two minutes. Now it's running comparisons. It's not just looking at the page.
So here we go. That finished in a minute, 38 seconds to look at all of these different cards, compare them. And now we're going to plan our walkthrough Kyoto.
So this is actually it using Google maps. And by the way, it put together this entire video you're looking at. It recorded its own screen, put together the information on the left side.
And there we go, a minute 23 to plan an entire walk through Kyoto. Now, there are a few critiques that I will give it. Number one, it has this tendency to work for 30 minutes, but with a little bit of prompt adjustments, you can get it to go for much longer.
And of course, if you use slash goal. Same thing. It also has some of the same design tendencies.
So as you can see here, a lot of these demos kind of look the same. This like faded green and other pastel colors, very flat design. So still a lot of those same design tendencies as GPT 5 .6.
But that can be easily fixed. It is. Very steerable in the design department.
You just tell it. But by default, it really wants to use this forest green everywhere. And then last, writing.
This is something that is near and dear to me. We write a lot at Forward Future. And obviously, every other model just has this severe AI smell to its writing.
And I will say GPT -6 is definitely the best. But... still very much has that AI smell to it.
So we got Fable 5 .1 this week. We now have the brand new training run, the brand new model out of OpenAI. And what do you think?
Which one do you think is better? I actually did a full review of Fable 5 .1. Go check that out right here.
The Hook

The bait, then the rug-pull.

Matthew Berman says he's had early access to GPT-6 Astra and calls it the best model he's ever used, promising benchmark saturation, near-zero alignment failures, and demo footage of 3D worlds and browser agents he says are "truly mind-blowing."

Frameworks

Named ideas worth stealing.

00:45list

GPT-6 Astra benchmark scorecard

  1. ARC-AGI-3: 98.6%
  2. FrontierMath Tier 4: 97.6%
  3. Agents' Last Exam: 59.3%
  4. BenchCAD: 95.9%
  5. DeepSWE v1.1: 73.0%
  6. Terminal-Bench Science: 64.6%
  7. GPQA Diamond: 96.0%

The core table the reviewer walks through, comparing GPT-6 Astra against Claude Fable 5.1, Claude Opus 5, and Gemini 3.8 Flash across reasoning, math, CAD, coding, and science benchmarks.

Steal fora quick reference table when deciding which frontier model to route a given task to
CTA Breakdown

How they asked for the click.

VERBAL ASK
13:58next-video
I actually did a full review of Fable 5.1. Go check that out right here.

A verbal pointer only at the very end, no on-screen link card or overlay shown in the captured frames.

Storyboard

Visual structure at a glance.

cold open
hookcold open00:00
benchmarks
valuebenchmarks00:45
alignment eval
valuealignment eval03:40
Little Planet
valueLittle Planet06:09
sponsor
ctasponsor07:16
Newhaven SimCity
valueNewhaven SimCity10:12
browser control
valuebrowser control11:46
sign-off
ctasign-off13:58
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.

15:48
Matthew Berman · Tutorial

You Aren't Using Codex Like Me

Eleven power-user habits from someone who has logged over a thousand hours in OpenAI's Codex CLI — model tiers, thread delegation, safety hooks, and remote control from a phone.

July 14th
08:54
Matthew Berman · Review

GPT-5.6 is FINALLY HERE (WOAH)

A 'dot' release plays out like a full generational leap: two five-to-seven-day unsupervised coding runs, a sponsor benchmark, and a live pricing and capability standoff against a rawer, higher-ceiling rival model.

July 9th