GPT-6 Astra Is Finally Here (And It's Really Good)
A first-look review of OpenAI's GPT-6 Astra, run through published benchmarks, a set of repeatable creative tests, and a computer-use experiment where the model built and animated its own 3D game world.
GPT-6 Astra's published benchmark scores are only incrementally ahead of rivals like Meta's Muse Spark 1.3, but its computer-use ability to independently operate creative tools such as Blender and Unreal Engine turns single prompts into working animated games in minutes instead of hours.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You track frontier AI model launches and want a hands-on read on GPT-6 Astra beyond the official benchmark chart.
You're curious what 'computer-use' actually looks like when a model operates real creative software like Blender or Unreal Engine on its own.
You build with AI coding or agent tools and want to know whether GPT-6 Astra's coding benchmark jump is worth switching for.
You want a quick read on how GPT-6 Astra compares to Claude Opus 5, Claude Fable 5.1, Gemini 3.8 Flash, and Meta's Muse Spark 1.3.
SKIP IF…
You want a rigorous, independently verified benchmark comparison rather than one creator's early-access impressions.
You're looking for a step-by-step tutorial rather than a reaction-style first-look video.
TL;DR
The full version, fast.
OpenAI began rolling out GPT-6 Astra on September 3, 2026, and Matt Wolfe got early access. On paper the jump is modest: a roughly 2% DeepSuite coding gain over GPT-5.6, and Meta's unreleased Muse Spark 1.3 actually scores higher on that same benchmark. But GPT-6 Astra's real strength shows up in computer-use tasks. It cloned a full 3D game in eight minutes instead of the usual two hours, built an interactive planet simulator from a single prompt, and independently took control of Blender and Unreal Engine to model, rig, animate, and drop a humanoid wolf character into a fully playable forest world. Other early testers report similarly ambitious builds, including a week-long Manhattan recreation and a simulated world of AI agents that started talking to each other unprompted.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Matt Wolfe explains why GPT-6 Astra's launch felt significant enough to break from his usual Friday news format, and discloses his early access and past OpenAI sponsorship.
00:59 – 02:04
02 · Rollout status & blog demos
GPT-6 Astra begins rolling out September 3rd to a limited group; OpenAI's blog shows demos like reading and filling out a Form 1040 and an interactive circuit board.
02:04 – 04:07
03 · DeepSuite & coding benchmarks vs. Muse Spark 1.3
The DeepSuite coding benchmark shows only a modest gain over GPT-5.6, and Meta's unreleased Muse Spark 1.3 claims a higher score on the same test.
04:07 – 06:13
04 · ARC-AGI, Artificial Analysis Index & cost per task
Terminal Bench Science and ARC-AGI show massive generational jumps, but the aggregated Artificial Analysis Intelligence Index ranks GPT-6 Astra only fifth, tied with GPT-5.6.
06:13 – 07:08
05 · BuseyBench: the SVG drawing test
GPT-6 Astra ranks first on BuseyBench's AI-judged SVG portrait test, clearly ahead of Claude Fable 5.1 and Meta's Muse Spark 1.3.
07:08 – 09:38
06 · Megabonk 3D game clone test
Matt Wolfe's recurring 'clone Megabonk' prompt produces a full playable 3D game with class selection and combat in about eight minutes.
09:38 – 11:35
07 · Orbis: an interactive world simulator
A single 'showcase your capabilities' prompt produces Orbis, a fully interactive planet simulator with sliders for sunlight, sea level, and rainfall that change habitability and population in real time.
11:35 – 13:46
08 · Blender: modeling and animating a wolf
Across three sequential prompts, GPT-6 Astra takes computer-use control of Blender to model, rig with 50 bones, and animate a running humanoid wolf character.
13:46 – 16:11
09 · Unreal Engine: the Whisperwood forest world
GPT-6 Astra opens Unreal Engine itself and builds Whisperwood, a full playable forest world with the wolf character wired up to WASD movement, in about 35 minutes.
16:11 – 17:17
10 · What other creators built with GPT-6 Astra
Screenshots from X show other testers' builds: Matt Berman's planet worlds and Fall Guys clone, Matt Schumer's week-long Manhattan recreation and a simulated world of AI agents talking to each other, Pietro's Atlantis simulator and Beyblade game, and Ethan Mollick's Library of Alexandria.
17:17 – 19:32
11 · Wrap-up & subscribe CTA
Matt Wolfe sums up his overall impression, previews upcoming videos, and pitches his Friday news roundup and newsletter.
Atomic Insights
Lines worth screenshotting.
GPT-6 Astra began rolling out to a limited set of organizations on September 3, 2026, with full ChatGPT Plus, Pro, Business, and Enterprise access expected within a few days.
On the DeepSuite coding benchmark, GPT-6 Astra scored 74.1%, only about a 2% jump over GPT-5.6 Sol, and it actually trails Meta's unreleased Muse Spark 1.3 at a claimed 75.4%.
The ARC-AGI 3 benchmark, which tests how well an agent learns unfamiliar interactive tasks, jumped from 7.8% to 99.9% between GPT-5.6 and GPT-6 Astra, well above the 48% average human score.
Terminal Bench Science jumped from 22% to 64.6% between GPT-5.6 and GPT-6 Astra, one of the largest single-generation benchmark gains discussed in the video.
On the aggregated Artificial Analysis Intelligence Index, GPT-6 Astra lands in only fifth place, essentially tied with GPT-5.6, despite feeling like a much bigger leap in daily use.
Running a task through GPT-6 Astra costs roughly $1.67 via the API, slightly more expensive than GPT-5.6.
On BuseyBench, an AI-judged test of drawing an SVG portrait from code, GPT-6 Astra ranked first, ahead of Claude Fable 5.1 and two Gemini Flash models.
GPT-6 Astra built a full 3D Megabonk game clone, complete with class selection and combat, from a single prompt in about eight minutes, versus the one and a half to two hours it used to take earlier models.
Using its own computer-use ability, GPT-6 Astra took control of Blender across three prompts to model, rig with 50 bones, and animate a running humanoid wolf character with no manual editing.
GPT-6 Astra then opened Unreal Engine itself and built a full playable forest world called Whisperwood, complete with the wolf as a controllable character, in about 35 minutes.
Other early testers report GPT-6 Astra building a full Manhattan recreation in Unreal Engine over a week, and populating a simulated world with AI agents that began talking to each other unprompted.
Takeaway
The real leap is computer-use, not benchmarks
WHAT TO LEARN
GPT-6 Astra's benchmark scores barely beat rivals, but its ability to independently run Blender and Unreal Engine end-to-end is the real capability jump worth understanding.
02Rollout status & blog demos
GPT-6 Astra began rolling out September 3, 2026 to a limited set of organizations, with full ChatGPT Plus, Pro, Business, and Enterprise access following within days — check whether you have access before assuming a new model announcement means immediate hands-on use.
OpenAI's own demo page shows the model reading and filling out a real Form 1040 tax document directly in the browser, hinting at practical paperwork automation beyond coding tasks.
03DeepSuite & coding benchmarks vs. Muse Spark 1.3
A model's marketing headline can still trail a quieter competitor on the same benchmark — GPT-6 Astra scored 74.1% on DeepSuite while Meta's Muse Spark 1.3 claimed 75.4% — so don't assume the newest release is automatically the highest scorer.
Benchmark comparisons only hold when every model is listed on the same scoreboard; Muse Spark 1.3 wasn't even ranked on the DeepSuite site yet, so its figure came from Meta's own claim, not an independently verified test.
04ARC-AGI, Artificial Analysis Index & cost per task
The ARC-AGI 3 benchmark, which tests how well an agent learns brand-new interactive tasks on the fly, jumped from 7.8% to 99.9% between GPT-5.6 and GPT-6 Astra — far above the 48% average human score.
Aggregated benchmarks like the Artificial Analysis Intelligence Index can rank a model far lower than it feels in daily use, so a single leaderboard number is a poor substitute for hands-on testing.
Running one task through GPT-6 Astra costs roughly $1.67, a little more than GPT-5.6, which matters if you're metering API credits or a ChatGPT plan on agentic tasks.
05BuseyBench: the SVG drawing test
On a creative coding test, GPT-6 Astra clearly separated from competitors even though it under-performed on strict coding benchmarks, which shows a single benchmark category doesn't predict performance in another.
The same Muse Spark 1.3 that beat GPT-6 Astra on DeepSuite produced a noticeably worse SVG portrait, a reminder to check multiple benchmark types before trusting one number.
06Megabonk 3D game clone test
Prompting a full 3D game clone with class selection, combat, and level-up mechanics used to take a top model up to two hours; GPT-6 Astra produced a comparably polished version in about eight minutes from one prompt.
That speed gain might partly reflect low early-access server load rather than a pure model improvement, so treat dramatic speed claims from any single test run with some skepticism until they hold up at scale.
07Orbis: an interactive world simulator
A single well-scoped creative prompt can surface a model's actual design taste far better than a spec sheet — asking for an interactive capabilities showcase produced a fully working planet simulator, not a static demo page.
Small parameter changes cascading into visible, connected outcomes is what makes a generated tool feel real instead of decorative — that interconnectedness is worth studying when building your own interactive demos.
08Blender: modeling and animating a wolf
Breaking a complex creative task into sequential prompts let GPT-6 Astra complete each step in minutes without the user needing any Blender skill, showing step-by-step prompting still outperforms one giant ask.
The model's computer-use took real control of Blender's interface rather than just generating a file, adding a 50-bone rig and hand-authoring the animation timing itself — a meaningfully different capability than a model that only outputs a static 3D asset.
09Unreal Engine: the Whisperwood forest world
GPT-6 Astra chained together an entire pipeline unattended: opening Unreal Engine, building an environment, importing the character, and wiring up movement controls, all in about 35 minutes with zero manual engine work.
The presenter was explicit that the results were rough — a reminder that computer-use demos should be judged on what the AI operated unattended, not on final production polish.
10What other creators built with GPT-6 Astra
Other independent testers pushed the same computer-use capability further in a single sitting than one reviewer could alone, including a week-long Manhattan recreation and an underwater Atlantis simulator.
One tester reported a simulated world of AI agents that began conversing with each other unprompted after being left running — worth treating as an early, unverified anecdote rather than a confirmed capability, since it's a single secondhand account shared on social media.
Glossary
Terms worth knowing.
DeepSuite
A coding benchmark used to compare how well different AI models handle real software engineering tasks, referenced repeatedly as one of the most human-relevant coding scores.
ARC-AGI
A benchmark that tests how well an AI agent learns to solve brand-new, unfamiliar interactive puzzles it hasn't seen before, designed to probe general reasoning rather than memorized patterns.
BuseyBench
An informal AI benchmark where models write code to draw an SVG portrait of actor Gary Busey, then have the results judged by another AI for accuracy and style.
Computer use
An AI capability that lets a model directly control a computer's mouse, keyboard, and open applications to operate real software, rather than only generating text or code for a human to run.
SVG
Scalable Vector Graphics, an image format built from code-defined shapes and paths rather than pixels, often used to test how well a model can 'draw' through code.
Artificial Analysis Intelligence Index
An aggregated benchmark score that combines results from many individual tests, weighted together, to rank models by overall estimated intelligence.
“This launch feels too big to just be like a footnote in an AI news video.”
Strong, standalone cold-open hook that frames the whole video's stakes.→ TikTok hook↗ Tweet quote
03:40
“This new Astra model is underperforming on the coding benchmark that most people are paying attention to compared to Muse Spark 1.3, which is just really surprising to me.”
Candid, contrarian benchmark take that isn't pure hype.→ newsletter pull-quote↗ Tweet quote
08:20
“It prompted this game and it built the whole thing in eight minutes. In the past... it would take an hour and a half two hours.”
Concrete speed comparison that quantifies the leap.→ IG reel cold open↗ Tweet quote
13:30
“I have no idea how to use Blender. I'm not good at Blender at all. So that was pretty impressive.”
Relatable non-expert framing that sells the computer-use capability.→ TikTok hook↗ Tweet quote
16:00
“I'm super impressed with what this was capable of using computer use.”
Tight, quotable summary line of the video's core takeaway.→ newsletter pull-quote↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphor
We just got another huge model launch today in GPT -6 Astra. And I was going to save this for my news video and just talk about it tomorrow. But if I'm being honest, this launch.
feels too big to just be like a footnote in an AI news video. When I open up my X feed, it is literally the only thing anybody is talking about on X. It is purely GPT -6 and what early access testers were able to build with it.
So let's talk about what GPT -6 is, do a few tests with it ourself, and take a look at what some other people have managed to build with it. Now, full disclosure, I did get... slightly early access to GPT -6.
So I've had a little bit of time to play with it. Also in another disclosure, OpenAI has sponsored my channel before, but this video is not sponsored. They don't even know I'm making a video about GPT -6 and I can say whatever I want about it because, well, it's not sponsored by OpenAI.
So first thing to note is as of today, the day I'm recording this, September 3rd, it is not fully publicly accessible yet. They made the announcement that GPT -6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus Pro business and enterprise users.
Although the day this video is going live, you may not have access to it yet. Within the next, you know, two to three days, you probably will have access to it. They have a handful of demos on their blog post here, like this interactive circuit board, this Excel competition, developing like what looks like a SimCity type thing, and a bunch of other demos.
I'm not gonna to really go through those demos because it's easy enough for you to just go read the blog post and click through them yourself. But in a minute, I will show you some of the things that I tested as well as some of the other things I've seen others test.
But let's scroll down to the benchmarks really quickly. Although I know benchmarks mean less and less to most people these days, there are some benchmarks that I still think are impressive and interesting to take a peek at. Things like automation bench, where the jump went from 18 .1 % to 41%.
And the benchmark I always like to talk about, the one that feels most correlated with how you'll actually feel when coding with it is DeepSuite. And this one, I'm going to be honest, it's interesting to me because I thought this would actually be a bigger number.
From my experience, it feels really, really good at coding. And don't get me wrong, 74 .1 % is a great benchmark on DeepSuite. It puts it at pretty much the top of the list.
It's about a 2 % jump over 5 .6 Sol. However, the reason I thought it would kind of be a bigger jump is if we take a look at DeepSuite here. Gemini 3 .8 Flash is at 74%.
Cloud Opus 5 is saying 74%. Meta's Muse Spark 1 .3, which came out yesterday, according to their website, is 75 .4 on DeepSuite. So this new Astra model is underperforming on the coding benchmark that most people are paying attention to compared to Muse Spark 1 .3, which is just really surprising to me.
Now I do say according to the Meta website, because again, if I look at the DeepSuite website, MuseSpark 1 .3 isn't actually listed on their website yet. We have 1 .2 down here, but 1 .3 is not showing on their ranking yet.
But then again, neither is GPT -6 Astra. Now where the benchmarks start to get really interesting are down here, Terminal Bench Science. It went from a 22 % to a 64 .6%.
Massive, massive leap. GPQA, Google Proof Question and Answers. It's basically saturated benchmark at this point.
Same with Frontier Math Tier 4. And then the Arc AGI benchmark, it went from 7 .8 % to 99 .9%. I mean, it's basically solved this benchmark.
But for additional context, the Arc AGI 3 tests how well agents learn as they solve unfamiliar interactive tasks. GPT -6 Astra saturates the evaluation, scoring 99 .9%. The average human tester scored 48%.
Now, the other benchmark that I often refer to is artificial analysis, because this is like a sort of aggregated benchmark. They take the scores of a bunch of different tests, put different weights on those different tests, and then rank what they believe to be the objectively smartest models. And this one I find kind of interesting because GPT -6 lands in fifth place.
pretty much tied with GPT 5 .6, which if you've used 5 .6 and then you use 6, it definitely feels like a big leap, which makes me sort of question the legitimacy of this benchmark just a little bit. Now, I still like the aggregated, weighted take that artificial analysis does, but in my tests, I've been way more impressed with GPT -6 Astra than I have been with Opus 5 and Fable 5.
I haven't personally tested Muse Spark yet on anything but UC Bench, but it's on. list to play with next and when we come down here and look at cost per task like i talked about in yesterday's video fable 5 .1 is the most expensive per task to run and gpt6 falls down here at about a dollar 67 per task which is you know a little bit more expensive than gpt 5 .6 so if you're using this through the api or if you have a chat gpt account and have usage credits well you might find you use them a little bit quicker or it's a little bit more expensive through the API if you're using 6 versus 5 .6.
But let's take a look at BuseyBench. Here's what GPT -6 generated. And according to my AI as a judge here, GPT -6 falls into first place, just above ClaudeFable 5 .1 and the two Gemini Flash models here.
Now, MuseSpark 1 .3, which actually performs better on both DeepSuite and the artificial analysis benchmark, created this SVG of Gary Busey. Now, if you're not familiar with SVGs, they're... basically images that were created from code.
So the code is basically telling it where to draw the lines and the circles and the shading and things like that to generate the image. So it is pretty much a coding test to see how well the code is at drawing images with code. And again, here's MuseSpark 1 .3, which scored 5 .7 on my AI as a judge.
And here's GPT -6. Like I don't even think they're close, which is again, why I was slightly confused by the... difference on DeepSuite and artificial analysis.
But again, artificial analysis is taking a bunch of different benchmarks into account and not just coding. Now, digging in a little bit deeper here on the GPT -6, it does say NA for cost. That's because I had early access.
I didn't run it directly through the API, so it didn't get the cost specifically. However, the AI did estimate it would have cost about $1 .94. to generate this.
But again, that's just an estimate. It used 63 ,858 tokens and it took nine minutes to generate. And when I do my generations on Buseybench, I always use the highest mode I have available.
So this was actually using the max mode. Now, if you've been watching my videos for a while, another thing I like to test just for fun is to see it's kind of design taste for making a game. So I always like to tell it to go and clone Megabonk for me.
If you're curious what the actual prompt was, it was literally Create a clone of the popular Megabonk 3D game. I always like to add 3D game because in the past it's made like a 2D top -down version.
I just like to specify that. And honestly, this is one of the most visually appealing versions of Megabonk I've gotten so far. Here's what it looks like.
I can be a knight. I can be a ranger or I can be a mage, each with different starting weapons. So let's start as a mage this time and click let's bonk.
And you can see it's pretty good looking. I mean, the character is one of my favorite character generations from any of these mega bonk tests that I've done so far. Like it's just a cool looking character and it's all done with 3JS.
It didn't actually build the models. So it coded all of this up and yeah, it looks really good. I'm going to quit out.
this I want to try the Rangers see what their weapon is so they have more of like a bow and arrow where they're shooting just one arrow at a time and then if I use the night they have more of this like bonk which is like this wave I think it's supposed to be hitting with the hammer it didn't actually animate the hammer moving very much I guess it kind of did a little bit but uh yeah the hammer animation is kind of leaving a little bit to be desired i just really like the design aesthetic i think it did a great great job and the game works beautifully it works really smooth and one thing i found really interesting was i prompted this game and it built the whole thing in eight minutes in the past when i've done this mega bonk test it would take an hour and a half two hours to make something like this this one it did the whole thing in eight minutes now to be honest i'm not sure if it was super quick because less people were using it, it wasn't in public access yet.
And so. because there was only a handful of people that had early access, it was loading a lot faster, or if this new model is just way faster. But in everything I've tested so far, it worked really, really quickly.
Now, another prompt that I like to test just to see what the AI comes up with is create a beautifully designed interactive website that shows off your capabilities as a model. The site can focus on whatever you want, but design it to really impress me with your capabilities. I've gotten some pretty fun websites out of that prompt.
in the past and here's what it generated i'm gonna hide my face real quick so i could show it off a little better but it created this like world simulator so i can control the earth and look at the details you can see like this is the dark side of the earth and all of the city lights are on over here it's not earth i guess it's orbis a little world uh that it created so it's it's like a simulated world similar to earth but it gave me all of these sliders so i can see what this world would look like if it got a lot more sunlight.
You can see how it dried the whole thing out or a lot less sunlight. You can see how the whole thing freezes over. I could raise the sea level.
and just make it a water planet or lower the sea level and make the whole thing just completely solid. I could change the level of rainfall, which affects how much clouds there are and how green it is. But here's what I found the most cool about this is that I can change what this world looks like.
I can raise land. So any of these areas where there's no land, I can click around and actually add more land to them. I could carve an ocean so I can go through here and actually add oceans into it.
I could add more. Seed life and make areas greener Add more plants.
I could leave craters. So there's like an impact thing. So if I wanted to put a crater here, boom, an asteroid just hit that spot or an asteroid just hit right there.
And you can see what would happen and how would it affect the world. But check this out. Down here, it shows the habitability, 85 % habitability, surface temp 16 .5 C, population 1 .35 billion and rising.
Now, if I crank up the sunlight so it gets drier, now 10 % habitability and we can watch the population. just drop like a rock bring this back to 50 habitability goes back up population goes up reduce the sunlight once again we see habitability start to drop we'll watch the population drop and it's just this really cool like little world simulator and there's some like beginning templates that you can start with here that changes everything for you to see what these other worlds would look like and once you build a world that you really like it'll build a little postcard of the planet that you made but i think probably the thing i've had the most fun doing is getting this new model to use its computer use functionality to go and play around in Blender and Unreal Engine for me.
So first I prompted it with, I want you to take control of Blender and create a 3D figure of a humanoid wolf. So it took about eight minutes and it created this little figure. Now this is just a screenshot of it, but if I pop open Blender here, you can actually see the character it created.
It's not the best. There's definitely quite a bit of wonkiness to it, but this was a one. prompt thing.
I could have gone back and forth, prompted it over and over again and got it more dialed in, but I wanted to see what it would do with one prompt. And this is what it designed for me. And then I went, Hey, can this like rig it up and actually animate it for me?
So I said, can you rig this up, remove the stand and then animate it? The original version had that little stand on the bottom. It took six minutes and it says done.
Removed it added a 50 bone rig and created an eight second looping animation and it created this animation Which left a little bit to be desired because the head's just kind of moving and the tail's wagging But there's not much else going on So I said now animate it so it looks like it's running and it made it so it looked like it's running I mean a little a little wonkiness with the hips, but it made it this running and if i open it up in blender here you can see that's actually what this animation is right now it's actually the wolf running but it did this by taking control of my computer opening Blender, drawing out the wolf for me.
And then after that, it rigged it up with 50 different bones so that it knew where to move. And then it created the little animations for me all through prompts. Like I have no idea how to use Blender.
I'm not good at Blender at all. So that was pretty impressive. But then I was like, let's take this to the next level.
Can I make this as a character in a game that I could run around as? So I gave it the prompt, create a beautiful forest world in Unreal Engine, and then. add this wolf as a playable character that i can use and run around as within the world that's supposed to say run around as i said round around uh but it knew what i was talking about luckily it took 35 minutes and it said it created whisperwood a forest with winding trails a pond wildflowers etc and it added my character in and here is whisperwood again this is all computer use it took control opened unreal engine and made this not the greatest running animation But it built this whole world.
I could jump around. I could play as this wolf. So WASD is to move.
I can use my arrow keys to look around. I could jump. Shift is to run.
And it built all of this for me. I know people who know what they're doing with Blender and Unreal Engine are going to be very, very unimpressed. But as somebody who doesn't know how to use either of these tools, just telling ChatGPT to prompt this into an existence for me.
Yeah, for me, it's wild. I'm super impressed with what this was capable of using computer use. I hope I just found the edge of the world.
And again, this was just like one prompt at a time. Like I wasn't trying to optimize or dial it in or really, really improve the look of the wolf or anything like that. A few more prompts and this would have looked like really good.
And quite honestly, most of what I created actually feels kind of janky compared to some of the stuff I've seen people posting on X that are way more creative than I am. So I figured I would quick. share some of the stuff that I saw other people share that they made with GPT -6.
I don't have to look very far on X, because again, that's my whole feed right now. But here's something that Matt Berman created, like these little planet worlds where you're this character that can move around in them. It looks really good.
I'm pretty sure this is all 3JS, like this didn't go and create separate assets. This is all written with code, and it looks really, really good. Matt Schumer here, GPT -6 Astra built this Manhattan world in Unreal Engine.
over the course of a week. So again, somebody who managed to get it to just take control of Unreal Engine and build inside of that world. And yeah, it looks really, really impressive.
Matt also created this 3D horror game here that looks like some sort of first person shooter, just without the shooter part of it. But it looks pretty good. And another one here from Matt Berman, which is like a Fall Guys clone, where you're running through these obstacle courses and...
trying not to fall off. Looks pretty decent again. Here's another one from Matt Schumer.
He claims he asked it to create a world in Unreal Engine and fill it with humans, and each one have an Astra -powered agent who all work together to survive. A day later, he was in his bedroom, heard some voices coming from the other room, came out of his room scared, and it was the Astra agent. They actually started talking to each other.
So here's the clip of what he shared of that. Now the video here looks ultra compressed. I'm sure it looks better on...
his original system but pretty interesting pietro here has some demos he made this 3d model and the animation all from a image that he gave it he made this atlantis underwater simulator which is one of the more impressive things i've seen for sure a beyblade game that you can play a game that kind of looks like wave race from n64 but maybe even like slightly better than the n64 version graphics wise ethan malik here Had it create the library of Alexandria here?
Now, I have the audio turned off, but it will actually explain things to you. Around 250 years before our era. And we're going to see more and more demos of what other people have built with it and what it's capable of over the coming days.
But man, it is one of the more impressive models I've played with lately. I just did a whole video yesterday about how this release schedule of new models like every other day coming out is sort of getting frustrating. because they all feel like small leaps.
But this one does to me feel like a bigger leap. It's one of the models I've had more fun with than others, especially getting it to go and use Unreal Engine and Blender and things like that for me. Really, really fun.
I'm excited to use it to go and overhaul like my dashboard, find ways to improve it and make me even more productive on my daily routines, things like that. So really, really exciting. Again, I'll talk about it a little bit more in tomorrow's Friday News Breakdown where I break.
down all the news for the week but I wanted to make a couple extra videos this week because a lot of models dropped and I didn't want tomorrow's video to be like an hour and a half long video talking about all this stuff. Plus the GPT -6 launch felt like a big enough leap in model capabilities that it warranted in its own video.
If you like videos like this and you want to stay looped in on all of the latest AI news again I also make an end of week video that comes out every Friday that breaks down all of the AI news stories for the week so if you only want one video a week that breaks down everything you need to know in the world of AI, maybe consider liking this video and subscribing to this channel.
That'll make sure that those news videos show up in your YouTube feed. I drink from the fire hose all week. I keep up with all of it so that ideally I'll be overwhelmed on your behalf and you can just watch one or two videos a week and be fully looped in on what's going on.
That's my goal with this channel. Again, consider liking and subscribing if you're into that kind of thing. I also have a free AI newsletter that comes out twice a week that will also...
keep you looped in. I'll put that link down in the description below. But thank you so much to everybody that tunes into these videos and nerds out with me about all of the latest cool tech stuff.
I am upping the amount of videos I'm doing outside of the Friday news video, because I feel like there's a lot to talk about right now. So I'll make a few more videos than I used to make.
I've got a really cool project coming out next week and some really cool collaboration videos coming up as well with some other creators that I'm super excited to share with you. So again, I really, really appreciate everybody that tunes into this channel and nerds out with me and obsesses over tech like I do. It's a lot of fun.
And the community that's been growing around this channel has been awesome to see. So thank you again. I really, really appreciate it.
Appreciate you. But that's what I got. Hopefully I'll see you in the next video, which will be tomorrow's news video.
See you later. Bye bye.
The Hook
The bait, then the rug-pull.
OpenAI dropped GPT-6 Astra on launch day, and Matt Wolfe says the release felt too big to bury as a footnote in his usual weekly news roundup. The benchmark charts tell a modest story, but what the model does when it's handed the keys to Blender and Unreal Engine tells a much bigger one.
Frameworks
Named ideas worth stealing.
07:08list
Matt Wolfe's repeatable model-test suite
Clone Megabonk as a full 3D game from one prompt
Build a beautifully designed interactive website that shows off the model's own capabilities
Take computer-use control of Blender to model, rig, and animate a character
Take computer-use control of Unreal Engine to build a playable world around that character
Run the model through BuseyBench's SVG portrait test
A fixed set of creative, open-ended prompts Matt Wolfe reruns on every new frontier model so results are comparable launch to launch, rather than relying only on vendor-published benchmarks.
Steal forBuilding a personal, repeatable benchmark suite for evaluating any new AI model's real-world creative and agentic ability beyond a published leaderboard.
CTA Breakdown
How they asked for the click.
VERBAL ASK
18:16subscribe
“consider liking this video and subscribing to this channel... I also have a free AI newsletter that comes out twice a week”
Low-pressure, benefit-framed ask woven into the sign-off (fewer videos to watch, still stay current) rather than a bare subscribe plea, paired with a secondary newsletter pitch.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
A creator prompts OpenAI's Codex into building a free, self-hosted dashboard that replaces a stack of SaaS tools for news, mentions, newsletters, and audience tracking.
Matt Wolfe opens the full toolbox behind his own YouTube intros — every AI app, model, and exact prompt, from wall-bursting entrances to a fully AI-dramatized Sam Altman text leak.
A two-hour gap after Claude Fable 5.1 shipped, a site-wide AI outage, and a blog post that got pulled mid-cycle — the strange week OpenAI introduced GPT-6 Astra.
OpenAI, Anthropic, and Google all released flagship models in the same week, Gemini Notebook quietly capped how much you can generate, and Grok's new payments plugin lets a bot spend your money with your sign-off.
A leaked OpenAI blog post says GPT-6 Astra is AGI, beats Claude Fable 5.1 on every benchmark, and crosses a cybersecurity threshold that lets it hack on its own.