GPT-6 Astra's biggest performance gains come from changing how you operate it, not pushing it harder: default to a lower effort level, let it gather its own outside references through browser and computer-use tools, prune old skills built for weaker models, use voice mode as a multi-chat orchestrator, and tell it explicitly whether to ask before acting or push through to a result.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You use Codex or Claude Code regularly and want to cut AI spend without losing output quality.
You have a pile of old Claude Code or Codex skills installed months ago that you've never cleaned up.
You're building sites or apps with a coding agent and want it to pull its own design references instead of doing that research by hand.
You want to know how to use voice mode for more than dictation.
SKIP IF…
You've never used an AI coding agent like Codex or Claude Code — this assumes an existing workflow to tune, not a first introduction.
You're not working with GPT-6 Astra or a comparable frontier coding model; some of this is model-specific.
TL;DR
The full version, fast.
GPT-6 Astra performs well even on lower effort settings, so cranking every task to max or ultra mostly burns extra cost and weekly usage for a small score bump. The video walks five habits: pick a lower default effort level unless a task is genuinely complex, let the model's browser and computer-use tools gather outside references like design sites instead of doing that research by hand, run a skill audit to prune old Claude Code or Codex skills cluttering the context window, use the improved voice mode as a multi-chat orchestrator rather than a single side channel, and write prompts that explicitly state whether the model should ask before acting or push through to a reviewable result on its own.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Pitches GPT-6 Astra against Anthropic's models and previews five usage mistakes that hamstring results.
00:27 – 04:34
02 · Mistake #1: Wrong effort level
Argues most people default to max or ultra effort needlessly; walks DeepSWE and Artificial Analysis benchmark charts and a side-by-side Dune House homepage build on light vs. max effort to show the score gap is small but the cost gap is large.
04:34 – 07:46
03 · Mistake #2: Underusing browser and computer use
Shows Codex controlling a browser to pull design references from Dribbble and bring screenshots back into the project autonomously, closing with a mid-video sponsor read for Chase AI Plus.
07:46 – 10:22
04 · Mistake #3: Outdated skills clogging the agent
Argues many installed Claude Code/Codex skills built for older models now hold newer models back, and demos a Skill Audit GitHub tool that inventories, evaluates, and reports on a skill library.
10:22 – 15:07
05 · Mistake #4: Not using voice mode as an orchestrator
Compares Codex's voice mode favorably to Claude Code Desktop's, then demos one voice chat directing edits inside an existing project and a second voice chat spinning up and managing a brand-new research chat.
15:07 – 17:12
06 · Mistake #5: Prompting without stating how to handle forks in the road
Cites OpenAI's own guidance that GPT-6 Astra asks for clarification more than older models, then walks OpenAI's prompt templates for either always-ask or execute-to-a-reviewable-result behavior.
17:12 – 17:47
07 · Outro
Recaps the five mistakes, argues Codex is ahead of Claude Code Desktop on browser use, computer use and voice mode, and repeats the Chase AI Plus pitch.
Atomic Insights
Lines worth screenshotting.
Raising an AI coding agent's effort level from low to max can nearly quintuple the cost per task while improving benchmark scores by only a few percentage points.
On the DeepSWE benchmark, dropping from max to low effort cut cost from $12 to $2.19 per task for only a 6-point score drop.
A homepage built with a light effort setting in 13 minutes on 91,000 tokens looked nearly as good as the same brief run at max effort in 30 minutes on 152,000 tokens.
An AI agent with browser and computer-use access can pull design references from sites with no API, like Dribbble, without a human touching the keyboard.
An installed AI skill still occupies context window space through its description even when it's never invoked in a session.
A skill-audit workflow can sort a bloated skill library into buckets like fix now, review for retirement, and preserve, instead of guessing what to delete.
Newer voice modes can act as an orchestrator that spins up and directs several separate coding chats from one running conversation.
OpenAI says GPT-6 Astra asks for clarification more often than older models did, so prompting habits built around the model quietly guessing now cause more interruptions.
A prompt that says to ask for approval only after preparing a concrete, reviewable result stops an agent from stalling on questions it could have already answered by doing the work.
Telling an agent to just figure it out without scaffolding tends to produce a regression to the mean instead of a sharper outcome.
Takeaway
Five habits that quietly waste an AI coding agent's power
WHAT TO LEARN
Most of the friction with a frontier coding model comes from operating habits carried over from weaker models, not from the model itself.
02Mistake #1: Wrong effort level
Higher effort settings cost far more than they earn back in score on standard coding benchmarks, so treat max as a last resort, not a default.
Test a task on a low or medium effort setting first, and only raise it if the output is actually missing something.
Effort level should scale with how complex a task genuinely is, not with how important the task feels to you.
03Mistake #2: Underusing browser and computer use
An agent with browser and computer-use access can gather outside references, like competitor or design sites, on its own instead of you copying screenshots in by hand.
This matters most for services with no API or MCP integration, where manual copy-paste used to be the only option.
Ask the agent to document what references it used and why, so you can audit its sourcing after the fact.
04Mistake #3: Outdated skills clogging the agent
Old skills built for weaker models can actively hold back a newer, more capable model instead of helping it.
A skill you installed months ago and haven't touched still costs context window space just by being loaded.
Running a periodic audit that sorts skills into keep, fix, and retire buckets keeps a library from silently bloating.
05Mistake #4: Not using voice mode as an orchestrator
Treat voice mode as more than a dictation shortcut: a single voice conversation can act as an orchestrator directing several separate coding chats at once.
Match effort level to how you're using voice: light for back-and-forth conversation, higher when the agent is doing real work in that chat.
06Mistake #5: Prompting without stating how to handle forks in the road
A newer, more capable model is more likely to stop and ask for clarification than older models that used to just guess and proceed.
Decide up front whether you want the agent to check in at every fork in the road or push through to a finished result, and say so explicitly in the prompt.
A prompt that asks the model to prepare a concrete, reviewable result before requesting approval prevents it from stalling on questions it could resolve itself.
Glossary
Terms worth knowing.
Codex
OpenAI's command-line AI coding agent, roughly the Codex-side counterpart to Anthropic's Claude Code.
Effort level
A Codex setting (light, medium, high, extra high, max, ultra) that trades compute time and cost for how much reasoning the model applies to a task.
Computer use
A capability that lets an AI agent control a mouse and keyboard inside a real desktop or browser environment, not just call an API.
Orchestrator
In voice mode, a single conversation used to direct and monitor several separate coding chats running in parallel, instead of managing each one by hand.
Context window
The finite amount of text an AI model holds in memory at once; installed skills and their descriptions consume space in it even when unused.
Resources
Things they pointed at.
07:29productChase AI Plus (Claude Code Masterclass + Codex Masterclass)
00:00communityChase AI free community
00:00linkchaseai.io consulting
00:54toolDeepSWE benchmark leaderboard
01:40toolArtificial Analysis Coding Agent Index
05:10toolDribbble
07:46toolSkill Audit (GitHub tool)
15:55linkOpenAI's GPT-6 Astra prompting guide
Quotables
Lines you could clip.
00:44
“Spoiler alert, it is not.”
tight punchline landing right after the setup about maxing out effort level→ TikTok hook↗ Tweet quote
04:10
“But is it significantly better than light? You can make an argument for both.”
an honest, non-hype admission that max effort barely beat light effort→ IG reel cold open↗ Tweet quote
08:28
“You might have 30, 40, 50 skills that have nothing to do with what you do anymore, and they're just clogging up your context window.”
names a relatable pain point for anyone with a bloated Claude Code or Codex skill library→ newsletter pull-quote↗ Tweet quote
10:30
“What we get here inside of Codex destroys its cousin inside of Anthropic Cloud Code Desktop.”
a blunt cross-vendor comparison that invites disagreement→ TikTok hook↗ Tweet quote
16:58
“Don't just come to me with a problem, come to me with a problem and your potential solution.”
a crisp, transferable prompting rule stated as a single line→ newsletter pull-quote↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphor
GPT -6 Astra is giving Anthropic a run for its money, and for good reason. This is the greatest AI model we have ever seen. But if you are using this model the same way you've used older models in the past, you are severely hamstringing your results.
But today in this video, I'm going to be going over the five mistakes you are making when it comes to your Astra usage. And more importantly, I'm going to show you. how to fix them.
Now, the first mistake you're making when it comes to Astra is you are using the wrong effort level. It is unfortunate how many people I have seen open up Codex, throw effort level to max, or even if they're complete freak shows, throw it to ultra and think this is going to automatically get you better outputs. Spoiler alert, it is not.
We need to be very conscientious of what sort of effort level we are taking for our particular project. And unless you're doing some sort of wild, extremely complicated projects, you probably shouldn't ever be going above high.
And in many cases, you should probably be on medium or even light. And the stats kind of prove this. This is illustrated very well here with the deep sweep benchmark, which is a benchmark that's all about long running agentic tasks, the type of tasks that you would expect the much higher effort levels to really thrive in when we compare it to the lower effort levels.
Yet, looking at deep sweep, and I am on the max setting, I get a score of 73 % and my average cost per task is $12. However, on the exact opposite side of the spectrum, if I go too low, I'm at 67%, which is only a 6 % drop -off, yet I've gone from $12 per task to $2 .19 per task.
Now, I know most of you are on a subscription plan. We're talking about usage, but the fact remains, the amount of weekly usage you will burn on something like low will be significantly less than on high, or on max, rather. And as we go down the line from max to extra high, you see we actually got better results on extra high, yet the cost per task was almost half.
We see that again when we go to high. We are at the same exact score as max, yet we go from $12 to $5 .72. And at medium, we're basically at the same exact place.
What's the point here? The point is Astra is extremely efficient even at lower settings. And if we compare that to something like Fable 5, the low setting on GBT6 Astra is basically right in between Fable 5 high and medium.
So that would be about a $7 price point on the Anthropic side. Again, $2 .19 with Astra. And this idea is repeated across multiple benchmarks.
Here's the Artificial Analysis Coding Agent Index. At low setting, we're at $1 .50 at a 62 .6 score. On max, we're at $4 .98 per cost and a 67.
score. So a difference of 4 .4 for the score, yet the price jumps up, you know, almost $3 .50. Now on some of these other benchmarks, we see there is a dip as we go from low to max, like max will give you better outputs on these far ends.
But even here, right on terminal bench, high gives us a better score than max. And again, much cheaper. So what's the takeaway here?
When you are inside of Codex, don't just automatically throw this bar for effort level all the way to the right. In fact, you can probably get away on average with high. Now, I went ahead and tested this out on front -end design, which is a common use case for Astra.
I gave Astra the same prompt with two different effort levels. This was the prompt on light. I said I wanted to create a homepage for Dune House, a fictional boutique desert hotel in Joshua Tree, California.
And this is what it created for us. Honestly, pretty solid. Again, this was using the light effort level in terms of the time it took.
I gave it the prompt at 3 .41 and 13 minutes later it was complete. Total tokens used was 91 ,000. And here's a look at the same result on max setting.
This time it took 30 minutes to complete. We used 152 ,000 tokens and this was the end result. And I mean, this looks pretty good.
But is it significantly better than light? You can make an argument for both. Point being, is there a huge difference in the outcome here for this particular task?
Not really. So when in doubt, when we're working with effort levels, less is more. If you aren't getting the output you like, go ahead and begin incrementally bumping up that effort level.
But there is essentially zero reason why you should start at extra high, max, or ultra. Now, the second mistake you are making when it comes to Astra is you are completely underutilizing its browser. and computer use.
These two things allow us to sort of connect Astra to applications where we don't have a readily available CLI, MCP, or API. For example, let's say we want to improve on this website we created. This was the version we got with the max effort level.
And instead of me manually going out on the web, finding new references, and feeding it to Astra, why don't I have Astra do it on its own? Why don't I send it to a website like Dribbble, which doesn't have an API that I have access to at least, and have it find references for other hotel -type websites, download the images or take screenshots of the images, and bring it into the fold for its next iteration.
It can do all that automatically. I don't have to touch anything. So that prompt will sound something like this.
Hey, so can you, in the browser on the right -hand side, head to dribble, that's D -R -I -B -B -B -L -E .com, and then look up... some websites, some references for hotel websites, because what I want you to do is I want you to go onto Dribbble, I want you to search for hotel websites, and I want you to find references of other hotel websites that look very visually stunning.
Take screenshots of them, do what you need to do, so then you can bring those reference images back into Astra, back into Codex, and then come up with a new version of our website. So we can see here over on the right, it has now pulled up Dribbble. And it has my login because it actually saves those when you log in at all.
It's now searching for hotel website. It's then clicking on individual websites that were listed there. It's capturing screenshots.
And then it's continuing this process with more websites that it thinks sort of fits the bill. It then put all that together to create this website. And I'll turn off my camera so you can see it better.
So it generated the brand new website. And like Codex always does, it also made sure it worked on mobile, ran all the tests. we got something a little bit different and to be honest i kind of like this version better and i like the very prominent image it generated as well and normally we would do this completely manually in terms of finding all these references but you could see how it'd be so easy to scale this where maybe you want to just work it look at dribble it could look at pinterest it could look at twitter or we could apply this to really any scenario where we have some sort of application that we want to grab information from and bring in the codecs but again we don't have an api we don't have an mcp this is where these browser automations and computer use tools are so handy astra even created sort of a reference table so i can see which screenshots it actually used as well sort of its notes and links to the original source on dribble which is also nice now before we go into the third mistake a quick word from today's sponsor me so inside of chase ai plus i have released not only a brand new cloud code master class i also have a codex master class as well so
No matter your technical background or lack thereof, I will teach you how to master these AI tools. Focus on real use cases. I post updates every single week.
So if this is something you really want to dive into, Chase AI Plus is the place for you. There is a link to it in the pinned comment. Now, the third mistake you are making when it comes to Astra is that your skills are all wrong.
Your skills are holding you back because we are in the same place now with Codex that we were at with Cloud Code just a few weeks ago. You remember when Boris Cherny, the maker of Cloud Code, came out and said, you need to delete your cloud .md, you need to delete your skills. Well, the idea there wasn't that we should just delete them for the sake of it.
It was that these models, Astra and Fable, have gotten so good that many of the skills that acted as scaffolding for older models we used to play around with just are irrelevant now. And in fact, in many cases, they are holding you back. Now, this isn't black and white, although I will say a lot of the super heavy scaffolding skills, things like superpowers and GSD, Should probably go by the wayside.
But even if you disagree with that statement, your context window is probably clogged with a ton of skills that you just don't even use. How many skills did you install nine, six, three months ago that you haven't touched since then? Have you gotten rid of them?
If you haven't, they're still there. Like it's just their description, but these add up. You might have 30, 40, 50 skills that have nothing to do with what you do anymore.
And they're just clogging up your context window. So what do we need to do? Well, we need to do an audit.
and i created a skill that does that for you this is the skill audit and this is actually based on anthropix skill creator skill because it has benchmarks and testing in place where if you pointed out a specific skill it will run tests to see does this skill make sense with the current model codex has its own skill creator skill but it wasn't as robust as anthropic so that's why i created this and this is what it does so i'll put a link down below to where you can find it i show you the install and then i also give you a prompt you can run this prompt will then go through all of your skills and your logs and figure out which skills have you not been using at all and that we should probably prune and it's going to take a look at the front matter of your skills see what descriptions make sense and then also have a list of skills for you to take a look at where if you want to you can go one by one through these skills and actually run this skill audit benchmark test against them to see do they make sense with astra
And so group solve your skills into a bunch of different buckets, either fix now, review for retirement, text next, test later, or preserve. Again, this is super simple for you to use. You're just going to install the skill and then run this prompt.
And this will buy you a few things. One, it's going to free up some of our context window that has been bloated through skills we don't even use. And then two, for those skills that are in the gray area of, well, we still want to use them, but we're not sure if they're legit or can we improve them.
It's going to improve them. It's going to benchmark them. And you're no longer going to have a question of, do these actually help me?
Now, the fourth mistake you are making when it comes to Astra is you are not using its voice mode. And its voice mode is best in class. What we get here inside of Codex destroys its cousin inside of Anthropix Cloud Code Desktop.
And it got improved with Astra. Before, when you use voice mode and you use it by just clicking this little thing right here and this bobble is going to pop up, before it was powered by GPT -Terra. And I believe it was Tara on the light effort setting.
Now you can use it with Astra and you can use it with Astra high or Astra low. Furthermore, if I wanted to use voice before, so let's say I want to start a new voice chat, it would be in an entirely different chat panel like you see here. And I'd have to use it as an orchestrator.
So I would say, hey, go do X, Y, and Z. And now it's kind of listening to me right now. I would say, hey, go do X, Y, and Z.
And it would open up a new chat and do that. Now what I can do is I can go into an individual chat, like the one we were just at, and I can use voice mode inside of here. So I have multiple options, orchestrator that controls multiple chats, or I can use it inside an individual chat, and I have Astra.
So, for example, if I wanted to use it inside of here, I'm just going to click this thing, and let's say I wanted to adjust something. Let's say I wanted it to add some sort of form somebody could fill out. on this web page over here on the right and then i also wanted it to test it out so let's try that so i'm going to click on this and i will say when it comes to how you should best use it if you are using it to do things so i'm inside a chat right now i'm having it do something which is add the form i'm probably going to put it on high effort if i'm using it as an orchestrator or i just want to kind of talk to it back and forth i'll probably just put it on light so i'm going to do voice Throw this on high.
And then I'm just going to click it and talk. So what I want you to do right now is I want you to take a look at this website we've built on the right -hand side of the browser. And I kind of would like some place where people can fill out just a form if they want more information, probably at the bottom near the footer and kind of figure out what best practices are for that, what sort of information they should put in there.
Right now... we don't need the functionality to actually work i don't need you to hook it up with resend or anything i kind of just wanted to see what it will look like on the web page can you go ahead and do that for me sure let me take a look and also i want you to notice kind of stop that for a second i also want you to notice how quick and snappy that was when it said sure let me take a look if you haven't played with the voice mode here it is extremely snappy and very responsive and so while that voice mode is working let me sort of demo what else we can do with it so i'm here in a new voice chat i'm going to set this to light because i'm going to have it act as an orchestrator and i can pretty much tell it to like open up a new chat over here on the left use astra and we'll say something like hey can you figure out what the top five gpt voice you know use cases are can you go ahead and spin up a new chat window over on the left hand side doesn't need to be a new project um get astra working on doing some research for us to figure out what are the top five use cases for the new gbt6 astro voice mode specifically looking at the voice mode so open up a new chat with that and i also want you to tell that agent to write up a little document for us like html or something on it setting that up now
So you can see here, it's now created that chat. I can open up that chat. It's over here on the left -hand side, top five voice models.
And you can see right here, it says sent by chat GPT.
And so you can see here the prompt that was sent to this chat. Inside of this chat window, like I said, this can act as an orchestrator. I can have it create a bunch of different chats.
I can control everything from this one voice pane. So this is super useful if you're someone who's like has two, three, four, five, six plus agents going on all at once. And instead of trying to track all those manually, again, have Astra track all of them for you and just talk to this one orchestrator.
And we can see over here with our original command, it is now added and it's working on our little. sign up sheet right here, more information sheet. And you can see sort of the browser control in action with a little cursor.
It's actually testing it out as if it was a person. Done. The HTML brief is ready to open and it separates live voice handling from Astra's role doing the underlying work, which is a handy way to frame it.
So here you can see the five use cases it found. And really, I think the big unlock with voice chat is using it as an orchestrator. Now, the fifth and final mistake you are making with Astra is you are prompting it.
Now, this model is extremely effective, but there are some differences with how GPT -6 Astra handles things compared to GPT -5 .6 Sol. Specifically, when it comes to making assumptions, when it comes to forks in the road, this is coming from OpenAI themselves. And what they tell us is that Astra is more likely to ask for clarification where older, earlier models would make assumptions.
What does this mean for you? Well, this means when you are coming up with your prompts, especially when we're talking about long -running, agentic, complex tasks, you should probably put in some sort of verbiage there about how you want it to handle these forks in the road. Do you want it to always ask you questions, or do you want it to sort of carry the user's intended task to completion?
Now, OpenAI gives us a specific prompt you can use, which is listed right here. Or you can use the template I'm going to put on the screen right now to help guide you for some of these bigger tasks. The big thing here is that you just need to know how you want it to behave.
Are you someone who likes it when AI is constantly checking in, which Astra will tend to do on its own? Or when we're doing longer stuff, do you want it to just do its thing? I'm going to give you a North Star.
I'm going to give you some sort of end state. Go forth and conquer. Don't ask me questions.
Figure out. There's pros and cons to both of these solutions. Just know when you take the route of just, hey, go ahead and figure it out, that's when you tend to get sort of a regression to the mean unless you set up a lot of scaffolding that's sort of pointed in certain directions when it reaches said forks in the road.
Another good option you have here is this prompt that OpenAI gives us where we're going to tell the model to ask for approval only after preparing a concrete reviewable result. This enjoys blocking the task if it could already have done it. right don't just come to me with a problem come to me with a problem and your potential solution so i think some combination of all these and what i showed earlier will get you to a good place if you're having some problems with astros who are just like stuttering its way through tasks so those are the five mistakes that are holding you back when it comes to gpt6 after this model is extremely powerful especially when we look at it inside of codex because codex has so many cool little features that i think anthropic in the cloud code desktop app are sort of falling behind on namely browser use computer use and the voice mode so make sure to try those out let me know how they worked for you as always if you want to learn more about codex and cloud code make sure to check out chase ai plus i'll put a link to that down in the pin comment and i'll see you around
The Hook
The bait, then the rug-pull.
The video opens by pitting GPT-6 Astra against Anthropic's models, then argues most people are getting worse results from it because they're still running it on habits built for older, weaker models.
Frameworks
Named ideas worth stealing.
00:54list
Effort level ladder
Light
Medium
High
Extra High
Max
Ultra
Codex's effort settings trade cost for reasoning depth; benchmark data shown in the video has scores plateauing well before max/ultra while cost keeps climbing.
Steal fortuning cost on any AI agent task that exposes a reasoning-effort dial
07:46concept
Skill Audit workflow
Audit
Usage review
Evaluate
Apply, when requested
A GitHub tool that inventories an agent's installed skills, checks which are actually used, evaluates them against current model behavior, and reports fix-now / retire / preserve verdicts.
Steal forpruning a bloated Claude Code or Codex skill library
A prompt structure from OpenAI's own GPT-6 Astra guide that tells the model when to ask vs. proceed, and asks it to arrive with a problem and a potential solution rather than stalling on a question.
Steal forlong-running agentic tasks where you want fewer unnecessary check-ins
CTA Breakdown
How they asked for the click.
VERBAL ASK
07:29product
“A quick word from today's sponsor, me: inside of Chase AI Plus I have released not only a brand new Claude Code masterclass, I also have a Codex masterclass as well... there is a link to it in the pinned comment.”
Woven in as a first-person aside mid-listicle rather than a hard stop, then repeated briefly in the outro; framed as more free teaching rather than a hard sell.
FROM THE DESCRIPTION
PRIMARY CTAWhere the creator wants you to go next.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
A free GitHub skill turns a one-line prompt into a finished motion-graphics video by pairing GPT-6 Astra, Claude Code, or Codex with Higgsfield's Seedance 2.5.
A hands-on walkthrough of Impeccable, the open-source Claude Code design skill, and its new Live Mode and Worlds features for turning AI-generated web design from generic to genuinely good.
Kimi K3's benchmark charts and rock-bottom per-token price look like a knockout blow to Claude and GPT — until a blind three-way build test and a real cost-per-task tally tell a much closer story.
A breakdown of three automation buckets — sales, research, and content — built on Claude Code and an indexed Obsidian vault, reclaiming five to ten hours a week.