A leaked OpenAI blog post says GPT-6 Astra is AGI, beats Claude Fable 5.1 on every benchmark, and crosses a cybersecurity threshold that lets it hack on its own.
A leaked OpenAI blog post claims GPT-6 Astra crosses a formal 'critical cybersecurity capability' threshold and beats Claude Fable 5.1 on every major benchmark, and argues price-per-task, not price-per-token, is what will actually decide which AI subscription wins.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You track frontier AI model releases and want the specific claims (benchmarks, pricing, safety threshold) instead of just the headline.
You're deciding which AI subscription to pay for and care about price-per-task, not just the sticker price.
You're curious what 'agentic AI' and 'observability' actually mean in a real release announcement.
SKIP IF…
You want a hands-on tutorial for using GPT-6 Astra — this video only covers a leaked announcement, not a walkthrough.
You're looking for independent verification of the claims — everything here comes from one leaked document the host says he obtained, not from OpenAI directly or from his own testing.
TL;DR
The full version, fast.
Alex Finn reviews a leaked OpenAI blog post for GPT-6 Astra, which CTO Greg Brockman calls AGI. The model reportedly beats GPT-5.6 Sol and Claude Fable 5.1 on every cited benchmark, including a 98.6% ARC-AGI score, and is the first OpenAI model to cross a 'critical cybersecurity capability' threshold, meaning it can find and exploit vulnerabilities without human help. Standard pricing matches Claude Fable 5.1, but price-per-task matters more than price-per-token, he argues. It trained on roughly 100,000 GPUs, OpenAI's largest run yet, and is positioned as an autonomous agent you supervise rather than prompt. Access starts with businesses before the public rollout.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Alex opens with the claim that GPT-6 Astra's creators are calling it AGI and previews a leaked blog post covering how it works, pricing, benchmarks, and safety risk.
00:29 – 01:22
02 · Greg Brockman calls it AGI
The leaked post frames GPT-6 Astra as a generational leap and the first model to cross OpenAI's 'critical cybersecurity capability' threshold: finding and exploiting vulnerabilities without a human in the loop.
01:22 – 02:37
03 · Benchmark blowout
A screen-shared benchmark table shows Astra beating GPT-5.6 Sol and Claude Fable 5.1 across ARC-AGI (98.6% vs roughly 7%), FrontierMath, coding (Deep Suisse, Terminal-Bench), and more by wide margins.
02:37 – 04:16
04 · Pricing and usage limits
Astra's standard-mode token price matches Claude Fable 5.1 exactly, but Alex argues price-per-task is the number that matters, since a model needing fewer tokens per task can be cheaper even at a higher per-token rate; OpenAI also resets usage daily and gives more volume than Anthropic.
04:16 – 06:44
05 · Largest training run ever
OpenAI says Astra trained on roughly 100,000 GPUs at its Stargate, Texas data center, the biggest training run in company history, which Alex uses to argue the scaling hypothesis is proven and that Anthropic's more cautious compute strategy has cost it the lead.
06:44 – 08:53
06 · Bad observability, cybersecurity threshold
The leaked post admits Astra's reasoning is harder to monitor because it uses fewer natural-language reasoning tokens; it also clears the critical cybersecurity capability threshold, and Alex notes it was reviewed by the Trump administration before release.
08:53 – 10:26
07 · "An employee you observe, not prompt"
OpenAI positions Astra as an autonomous agent that navigates websites, spreadsheets, browsers and desktop apps on its own, shifting the user's role from prompting to supervising.
10:26 – 11:23
08 · Rollout: businesses first
Access starts with businesses so they can harden their own systems before the model reaches the general public over the following days.
11:23 – 12:51
09 · Closing hype and CTA
Alex tells viewers to spend the long weekend testing the model and building with it, then pitches his newsletter and asks for a subscribe.
Atomic Insights
Lines worth screenshotting.
GPT-6 Astra is the first OpenAI model described as crossing a 'critical cybersecurity capability' threshold, meaning it can find and exploit software vulnerabilities without human intervention.
On the leaked benchmark table, GPT-6 Astra scores 98.6% on ARC-AGI versus roughly 7% for GPT-5.6 Sol.
GPT-6 Astra's standard-mode token pricing ($30/$50/$60) is identical to Claude Fable 5.1's, while its fast mode runs up to $20/$100/$120.
Comparing AI models by price per token is misleading; price per completed task is the number that actually predicts cost, since a cheaper-per-token model can need far more tokens to finish the same job.
OpenAI trained GPT-6 Astra on roughly 100,000 GPUs at its Stargate, Texas data center, calling it the largest training run in company history.
The model's reasoning is described as harder to observe because it uses fewer natural-language reasoning tokens, which OpenAI frames as both a safety concern and a defense against being distilled by competitors.
OpenAI says it briefed the Trump administration on Astra's cybersecurity capability before release and received approval to ship it.
OpenAI is rolling Astra out to businesses first so they can patch their own systems before the model reaches the general public.
OpenAI's positioning for Astra reframes the user's role from prompting to supervising: 'an employee you observe, not prompt.'
The video frames OpenAI's willingness to spend heavily on compute as the reason it may have overtaken Anthropic for the first time in years, after Anthropic's more conservative compute strategy.
Takeaway
Why 'AGI' claims should be graded on threshold, not vibes
WHAT TO LEARN
The real signal in a leaked-AI-announcement video isn't the AGI label, it's the specific, checkable claims underneath it: a named capability threshold, a price-per-task argument, and a training-compute number worth verifying.
02Greg Brockman calls it AGI
Companies now attach their own 'AGI' label to a release rather than waiting for outside consensus, so treat the term as a marketing claim on release day, not a settled technical result.
A 'critical cybersecurity capability' rating means a model can independently discover and exploit software vulnerabilities, a specific and testable claim distinct from vague 'more capable' language.
03Benchmark blowout
A benchmark table with one model near 99% and its predecessor near 7% on the same test signals either a genuine leap or a benchmark the newer model was specifically tuned for; check whether the test is public and reproducible before trusting the gap.
Coding and agentic benchmarks are increasingly the metric vendors lead with, because they map more directly to paid usage than academic reasoning scores.
04Pricing and usage limits
Matching a competitor's per-token price is as much a marketing signal as a cost decision, since it invites direct comparison shopping.
Price per completed task, not price per token, is the number to model when comparing AI costs, because token efficiency varies by model even at the same nominal rate.
A vendor that resets usage limits daily rather than monthly is signaling it has spare compute capacity, which is itself a competitive claim worth watching.
05Largest training run ever
Publishing the literal GPU count behind a training run is being used as evidence for the scaling hypothesis: more training compute reliably buys more capability.
A conservative compute strategy carries real competitive risk if a rival scales faster and crosses a capability threshold first.
06Bad observability, cybersecurity threshold
As models use fewer natural-language reasoning tokens, it gets structurally harder for anyone, including the lab itself, to audit what the model is 'thinking' before it acts.
The same opacity that raises safety concerns internally is being framed as a defensive moat against competitors reverse-engineering the model through distillation.
A government sign-off on a capability threshold is a checkpoint worth verifying independently, not a guarantee of safety by itself.
07"An employee you observe, not prompt"
The shift from 'prompt it' to 'assign it a goal and supervise the output' changes what skill is actually valuable: knowing what good output looks like matters more than knowing how to phrase a prompt.
Agentic tools that act across websites, spreadsheets, and desktop apps without step-by-step instruction require a different review habit than chat tools: checking outcomes, not intermediate steps.
08Rollout: businesses first
Staging access to businesses before consumers lets a vendor use paying customers to find and patch security gaps before the model reaches a much larger, less controlled audience.
Early access tied to a long weekend or holiday is a deliberate nudge to get power users testing before the wider release news cycle peaks.
Glossary
Terms worth knowing.
AGI (Artificial General Intelligence)
A hypothetical AI system with human-level or broader general capability across tasks, as opposed to narrow AI built for one job. In this video it's a label OpenAI applies to its own release.
ARC-AGI
A benchmark designed to test general abstract reasoning rather than memorized patterns, used here to show a large capability gap between GPT-6 Astra and the prior GPT-5.6 Sol model.
Critical cybersecurity capability threshold
An internal OpenAI safety classification for a model capable of independently discovering and exploiting software vulnerabilities without continuous human guidance.
Observability (AI safety)
The ability of researchers to monitor and understand a model's internal reasoning process well enough to catch dangerous behavior before it acts.
Distillation
A technique where a competitor trains a cheaper model by studying a stronger model's outputs and reasoning traces, used to explain why other labs have kept pace with US AI companies.
Price per task vs. price per token
Two ways to measure AI cost: token price is the rate charged per unit of text processed, while price per task is the total cost to complete a job — the more meaningful comparison when models differ in how many tokens they need per task.
Chain of thought
The step-by-step reasoning text a model produces before its final answer, which researchers use to audit how the model arrived at a decision.
“It's the first model that OpenAI is calling a critical cybersecurity capability threshold, meaning it can independently, without human intervention, find vulnerabilities in software and exploit those vulnerabilities.”
specific, sourced, and scary claim in one breath→ IG reel cold open↗ Tweet quote
03:10
“the price per task is way more important here”
reframes a common cost argument in one line→ newsletter pull-quote↗ Tweet quote
08:54
“It is an employee you observe, not prompt.”
tight, memorable framing of the product shift→ TikTok hook↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphor
AGI is here. ChatGPT6 Astra just released and its creators are literally calling it AGI. It's rolling out to a small group of businesses right now, but I got my hands on a leaked blog post that goes over every single detail about this model and some of these details are actually pretty scary.
In this video, I'm going to cover all the exclusive information. I'm going to talk about how it works, pricing, benchmarks, and why this might be the most dangerous AI model ever. release.
The AGI episode has arrived. This is going to be a banger. Now let's lock in and get into it.
We are here. GPT -6 Astra. It's not GPT -5 -7.
It's GPT -Astra. Let's start off with the big news here. AGI.
This is being claimed as AGI by Greg Brockman, the CTO of OpenAI. It is a generational leap in capability for basically everything. It's the first model that OpenAI is calling a critical cybersecurity capability threshold, meaning it can independently, without human intervention, find vulnerabilities in software and exploit those vulnerabilities.
That is really dangerous and it's why they're slowly rolling it out right now. But the big thing here is Greg Brockman saying, when we fast forward in a couple of years, this is going to be the moment we look back at and say, this is AGI, we have made it. Let's go straight into the benchmarks.
This is actually insane, mind -blowing. It doesn't even look like it's real, to be quite honest with you. It just destroys every benchmark by a country mile.
It blows everything out. It's a giant leap over 5 .6 Sol. And most impressively, it is a gigantic leap over Fable 5 .1 Anthropix Frontier model that released two days ago for the first time.
Ever. I mean, these two companies have been going band for band for years now. One releases a model, the other catches up.
One releases a model, the other catches up. This is the first time one releases a model and the other leapfrogs it, right? Look at this.
ARK AGI, 98 .6%. GPT -560 is at 7%. This is at 98%.
I mean, let's look at what we care about here. Coding, Deep Suisse, 74%, Fable 5 .1 was actually just a 67%. Terminal Bench, 64%, Claude Fable 5 .1, 52%.
I mean, it beats it and it wins in every single benchmark in every single category, not by 0 .1%, not by 1%, but by a very large gap. That is incredible. But what about pricing?
If it's so good, the pricing must be insane, right? Well, no, not really, actually. Let's take a look at pricing here.
GPT -6 Astra standard mode is basically the exact same as Fable 5 .1. But you might be going now, well, Alex, Fable 5 .1 is absurdly expensive. I get about three prompts in my subscription.
So am I going to have the same thing with GPT -6? Well, no, not at all, really. One thing they made sure to be very clear in this leaked blog post is the price per token is high.
It is Fable 5 .1 level and about double of... what GBT 5 .6 Sol is, but the price per task is way more important here. OpenAI has been talking about this a lot, and this is a very important concept you need to know over the last couple of weeks.
They have been talking a lot more in terms of tasks completed rather than tokens burnt. Basically what that means is like you can have two models and one can be way cheaper per token, but if that model requires significantly more tokens to get a task done, then it's actually more. expensive, even though it's cheaper per token.
So a super critical metric we need to pay attention to here is price per task, not price per token. If GPT -6 Astra can get more work done with significantly less tokens, that means it is going to be cheaper than Fable 5 .1. On top of it, let's just talk about how much usage these companies give you.
ChatGPT gives significantly more usage in your subscriptions than Anthropic. They not only gave you more usage, but they're resetting usage basically every single day.
So you are going to get way more out of this model, which is, looks like a leapfrog over Fable 5 .1. And why can they give you so much usage? It's because OpenAI has basically more hardware and GPUs than any other company out there, other than maybe SpaceX AI, because Elon's putting together like the biggest data center in history, but they have significantly more compute than Anthrop...
That's why they can give you so much inference in fact DPT -6 Astra was trained on the largest fleet of GPUs in history a hundred thousand GPUs in Stargate, Texas were used to train this model, which is absolutely mind -blowing and absurd It also proves the scaling theory. It proves that the more compute you have, the better your models are going to be, which is why I personally am investing in like every AI company I can get to on planet earth, Nvidia, Micron, SpaceX, like anyone that is AI infrastructure I'm investing in because it is very clear what's happening here.
The more compute you have, the better your AI models are going to be. And I think it proves it right here, right? Anthropic made a very I think, costly mistake early on trying to be careful with compute, not going all in on compute.
OpenAI and SpaceX did the exact opposite. They went all in. They leveraged everything to buy more GPUs.
And now I think it's finally paying off for them. You're seeing it, right? SpaceX and OpenAI releasing nonstop blast humans, constantly improving more and more products.
Anthropic for a few months now has been stalling. Fable 5 .1 is excellent, right? But it's their first release in like a couple of months and it's really good.
but they're clearly slowing down their pace. OpenAI's pace is exploding, as you can see, and at this point, it looks like they finally, for the first time in years, have taken a lead. I mean, I think since Sonnet 3 .5, OpenAI's never really been like a big leap above Anthropic.
They're always kind of playing catch -up, it felt like. This feels like the changing of the tides, a changing of the guard, the first time OpenAI puts their flag on the ground and goes, hey, we're here to win. You know, I really believe in the world we live in now, the world of $40 trillion of debt, money being printed out, the wazoo.
I really think you need to be, and this is not investment advice, but kind of investment advice. You need to be investing in assets. You need to have your money in growing assets.
And for me, it looks like, I mean, you just pay attention. Everyone's just buying more and more compute. The demand for compute is absolutely insatiable.
I think this model proves it. I'd highly recommend doing some research into AI infrastructure companies and putting your money to work appropriately. Now let's get into the scary stuff.
Let's get into the stuff that kind of makes you nervous, right? And that is the observability. It's the cybersecurity, everything around that.
There are some cybersecurity concerns here. You heard about the Hugging Face incident. They talked about it a ton, how Astra went.
broke out of a sandbox and hacked Hugging Face. Well, there's a lot of things behind this. Number one, the observability for this model is not great.
The smarter these models get, the harder it is to actually monitor their reasoning. It's actually able to manipulate its own reasoning so you can't monitor it, which is insane. They use fewer natural language reasoning tokens so it can influence its own chain of thought.
You basically can't really see the reasoning or how this model is. thinking, which is why there's so much concern about security here. On the positive side, if you can't really see the chain of reasoning, then it's going to be hard for Chinese companies to distill this model.
A large part of how these Chinese companies have been able to keep up with American companies is they use our models, they distill it, which basically means they hit it with a bunch of prompts and look at the reasoning to see how it thinks, and then train models based on that reasoning. You can't really see the reasoning quite as well with this model, so you won't be able to distill it as much.
So that's kind of good news at all of this, but the bad... news is, is like this could get dangerous. We don't know how it thinks really.
We don't know if it's 100 % perfectly aligned. So that is a risk here. It passes the cybersecurity threshold as we talked before, which means it's capable of finding previously unknown vulnerabilities.
and developing exploit chains across well -protected systems without continuous human guidance, which means it can think on its own, find exploits, and exploit those cybersecurity concerns. That's a little scary, but they did say they ran this by the Trump administration to get it approved, and they approved it, which the Trump administration was fickle before, right?
They demanded they took down Fable 5 when that originally came out. So if it passed the security check from them, then it might be good. So what about how it works and how you use it?
Well, they're saying in this leaked blog post, it is an employee you observe, not prompt. This is truly a model built for AI agents. It's built to be an AI agent that independently works, that independently thinks, that you don't prompt.
You just give it goals. It goes, it works independently, and you just observe it. This is where the beauty of AI comes in, right?
This is where the vision and the dream and the positive side of AI comes in is freeing up your time. You get more time to do what you want. You get more time to chase your passions and whatever you want to do with your life rather than sitting there and doing work all day.
This model can go out and do work for you. You just observe it. You just make sure it's doing things well.
You don't need to sit there and prompt it nonstop. This is where the optimism of AI needs to come in. This is where people need to be excited.
People can spend their time doing what they want now. You have more free time to go outside, experience the world, experience the beauty of the universe. You don't need to be sitting around doing work all day because your AI agent can go out and do that work for you.
So this is truly the first AI agent that can go out, navigate websites, spreadsheets, browsers, desktop apps, create documents, presentations, websites, 3D projects. I can't wait. I'm going to launch an AI video game.
company. I promise you right now, I'm going to launch my own AI video. I love video games.
I'm going to launch one. This one seems like it's built for that kind of use case. This is the shift from prompting AI to supervising its work.
More independence for you to pursue what you love. I think that's the beauty of AI. So at the moment, it is rolling out to businesses.
Now, as we speak, businesses are getting access to it. The reason they're doing it this way is so that businesses can use the model to harden their systems, get rid of security vulnerabilities, so that when the larger population gets it, they don't just sit around and start hacking every single website. Everyone else will be getting it over the next few days.
So you, me, everyone else involved, the next few days, we'll be getting it slowly. I will personally be sitting here and refreshing my chat GBT every five seconds. I need access.
I cannot wait to get access to this. I highly recommend you do the same thing. We got a long three -day weekend coming up if you're in America.
I highly recommend carving off time to use this model. When new models come out, you have an opportunity, right? 99 % of the world will have no idea this is happening, right?
But you, because you watch this channel, because you learn from me, thank you very much for learning from me, you know what's going on in the world. You know what the cutting edge is and you know how to use the most cutting edge technology. You need to be using this model this weekend.
Carve out time. Make your own benchmarks, right? Make benchmarks that are personal to you.
I showed you the world famous Alex Finn benchmark in the last video. Make your own world famous benchmarks. Test models out.
Build games with it. Just have fun and create with these models. If you do it right when it comes out, you have an edge over your competition.
You can launch products. You can launch businesses. It's just important you use it.
You use it to build. You use it to be productive. You use it to make amazing things.
This is one of those options. opportunities if you take advantage of it. Make sure to like, subscribe, and turn on notifications.
As this rolls out, I will be rolling out many, many, many tutorials. I also have a newsletter. It's the number one AI newsletter on planet earth.
Check out the link for that down below. I'll be sending out tons of guides and tutorials for free in that newsletter. Sign up for that newsletter down below.
I'll put the link in the comments too. Super critical you do that right now. I hope this was helpful.
I am so, so, so, so grateful you're coming with me through this journey of AGI, the most exciting time to be alive and using technology ever. This is as exciting. Hey, the people that watch say, hey, you're a hype guy.
Hey, too much hype. Hey, shove it. If you're not happy right now, if you're not feeling alive, then you don't have a pulse.
This is the greatest time to be alive using technology ever. I can't wait to use it. If you're optimistic and happy and hyped like me, let's go.
This is the best time to be alive. I will see you in the next video.
The Hook
The bait, then the rug-pull.
Alex Finn says he got an advance look at OpenAI's internal blog post for GPT-6 Astra, and opens by repeating the company's own claim: this is AGI.
Frameworks
Named ideas worth stealing.
03:07concept
Price per task, not price per token
The claim that comparing AI model cost by per-token rate is misleading; the real cost metric is total tokens-to-completion for a given task, since a cheaper-per-token model that needs many more tokens can end up costing more per finished job.
Steal forany AI tool cost comparison or pricing page
CTA Breakdown
How they asked for the click.
VERBAL ASK
11:54newsletter
“I also have a newsletter. It's the number one AI newsletter on planet earth. ... Sign up for that newsletter down below.”
Soft in-video pitch near the end, reinforced by a 'like, subscribe, and turn on notifications' ask right before it; the description also links a paid bootcamp, a second YouTube channel, and his own AI products.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
A 10-minute breakdown of why a better, cheaper AI model being locked behind 20 government-selected companies is a turning point — and what to do before the window closes.