Modern Creator
Mark Kashef · YouTube

I Tested GPT-6.1 Sol vs Astra. Here's What I'd Use

A creator built the same voice-and-research app three times, once with each model and once with both, to see where the price gap in OpenAI's new GPT-6.1 Sol actually shows up.

Posted
today
Duration
Format
Review
educational
Views
12K
135 likes
Part of the collectionThe GPT-6 Astra PlaybookEvery GPT-6 Astra breakdown, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

GPT-6.1 Sol matches most of Astra's capability at roughly a fifth of the price, so the strongest workflow uses Astra to plan a build and Sol to execute most of it, rather than picking one model for everything.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You build with OpenAI's Codex-style agent tooling and want real cost numbers instead of vendor benchmark claims.
  • You're deciding whether to pay Astra's rate on every task or mix in the cheaper Sol model for execution work.
  • You want a repeatable pattern for splitting an AI build between a planning model and an execution model.
SKIP IF…
  • You're not using OpenAI's Codex or agent-thread tooling, and want a general-purpose model comparison instead.
  • You need production-scale numbers rather than one creator's single recorded test run.
TL;DR

The full version, fast.

OpenAI's GPT-6.1 Sol claims near-Astra intelligence for a fraction of the price, so the creator built the identical Mac voice-and-research app three ways: Astra only, Sol only, and a hybrid where Astra plans and Sol executes. Build time barely differed, 42 minutes versus 40. Cost didn't: for nearly identical token counts, Astra's build cost about $203 and Sol's cost about $43. Astra's research tool returned broken citations; Sol needed extra prompts to get voice input working. The conclusion: plan with the expensive model, spend most of the token budget executing with the cheap one.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:00 – 00:15

01 · The claim

OpenAI positions GPT-6.1 Sol as a near-Astra-level model at a fraction of the price; the video sets up a real build test to check that claim.

00:15 – 00:42

02 · The app: a Wispr Flow clone

Three versions of Relay Voice, a Mac voice-dictation app that can also call Gmail and Firecrawl research tools, were each built by Astra, Sol, or a combined workflow.

00:42 – 01:47

03 · Pricing side by side

Astra costs $10 input / $50 output per million tokens with $1 cached input; Sol costs $2 input / $10 output with $0.10 cached input, roughly a fifth of Astra's price.

01:47 – 03:12

04 · Build 1: the combined workflow

The Astra-plans/Sol-executes build demoed live: checking Gmail for recent senders and researching current ChatGPT pricing plans through Firecrawl.

03:12 – 03:43

05 · Build 2: Astra

The Astra-only build works but isn't draggable despite that being requested in the plan; voice capture and Gmail lookup both function, a bit slower.

03:43 – 04:24

06 · Build 3: Sol

The Sol-only build looks the most basic and has some latency, but handles a mid-sentence language switch and basic settings correctly.

04:24 – 05:12

07 · Same prompt, same stack

Astra and Sol were given an identical prompt and identical tool stack, Gemini for transcription, Firecrawl for research, Google CLI for Gmail, so the model was the only variable.

05:12 – 06:32

08 · The prompt and the phase threads

Both models were authorized to spin up their own sub-threads per build phase and self-organize; the version with more upfront planning time self-organized into eight phases.

06:32 – 07:14

09 · Results: time and cost

Astra took 42 minutes and Sol took 40; on nearly identical token counts, 69 million versus 66 million, Astra cost $203 and Sol cost $43.

07:14 – 08:32

10 · Where each model struggled

Astra's research tool returned broken citation links; Sol needed several extra prompts to get voice input, its core feature, working reliably.

08:32 – 09:05

11 · Which one to use

The recommended pattern is hybrid: use Astra for planning and cleanup, and spend most of the token budget executing with the cheaper Sol.

Atomic Insights

Lines worth screenshotting.

  • GPT-6.1 Sol costs about a fifth of GPT-6 Astra on both input ($2 vs $10 per million tokens) and output ($10 vs $50 per million tokens).
  • Sol's cached input tokens cost $0.10 per million versus $1 for Astra, a 10x cache discount on top of the base price gap.
  • Building the same voice-and-research app took Astra 42 minutes and Sol 40 minutes, so raw execution speed barely differs between the two models.
  • For nearly identical token usage, 69 million versus 66 million, the Astra build cost $203 and the Sol build cost $43, a roughly 5x real-world cost gap.
  • Astra's research tool calls returned broken citation links when the task required multiple rounds of web scraping.
  • Sol got most of a voice-dictation app working on the first prompt but needed five to six follow-up prompts to fix voice input, the app's core feature.
  • The build that got the most planning time from Astra before handoff to Sol self-organized into eight phases and needed the fewest fixes out of the box.
  • OpenAI's own evaluations put Sol at parity with Astra on software engineering, close on professional documents, and about 2.1 points behind on computer use.
  • A model given explicit authorization to spawn its own sub-agent threads per build phase can self-organize a complex, multi-part project without a fixed human-written plan.
  • The practical workflow this test points to is using the pricier model to plan and using most of the token budget on the cheaper model to execute.
Takeaway

Plan with the expensive model, execute with the cheap one.

MODEL COST STRATEGY

Testing identical prompts on GPT-6.1 Sol and GPT-6 Astra found near-identical build times but a 5x cost gap, so the right move is splitting the labor between them instead of picking one model for everything.

03Pricing side by side
  • Sol's list pricing is roughly a fifth of Astra's: $2 vs $10 per million input tokens and $10 vs $50 per million output tokens.
  • Cached input tokens are the real lever: Sol charges $0.10 per million against Astra's $1, a 10x cache discount on top of the base 5x gap.
04Build 1: the combined workflow
  • Splitting labor, using Astra to plan a build and Sol to execute it, can produce a working prototype without paying Astra's rate on every token.
  • A lead model that can spin up its own sub-threads is what let one build self-organize into eight execution phases and need the fewest fixes afterward.
07Same prompt, same stack
  • Giving both models the identical prompt and identical tool stack is the only way to isolate the model itself as the variable being tested.
  • Even with identical inputs, the two models produced different UI polish and reliability, so the model choice still shows up in the finished product.
08The prompt and the phase threads
  • Explicitly authorizing a lead AI thread to spawn its own sub-threads per project phase let each model decide its own phase count instead of following a fixed plan.
  • More planning time upfront, letting Astra structure the phases before handing off to Sol, produced the version that needed the fewest post-build tweaks.
09Results: time and cost
  • Build time was nearly identical, 42 minutes for Astra versus 40 for Sol, so speed isn't what separates these two models.
  • Cost was the real gap: comparable token counts, 69M vs 66M, cost $203 with Astra and $43 with Sol, a roughly 5x difference for a similar finished app.
10Where each model struggled
  • Astra's research workflow returned broken citations when it had to decide between single-pass and multi-round web scraping.
  • Sol got the core features working fast but needed five to six extra prompts to fix voice input, the single most fundamental feature of the app.
11Which one to use
  • The practical pattern is hybrid: let the pricier model plan and clean up, and spend the bulk of your token budget executing with the cheaper model.
  • GPT-6.1 Sol is a real upgrade over the original GPT-6, capable enough to use seriously, where the original Sol wasn't worth the API call.
Glossary

Terms worth knowing.

Astra
OpenAI's more expensive, higher-capability GPT-6 model, priced at $10 input / $50 output per million tokens in this comparison.
GPT-6.1 Sol
OpenAI's newer, cheaper model, priced at roughly a fifth of Astra's rate and positioned for day-to-day execution work.
Firecrawl
A web-scraping and research API that lets an AI agent search, scrape, and summarize web pages as a tool call.
Codex threads
Separate AI chat sessions a lead agent can spawn to work on different phases of a project in parallel, each reporting progress back to the lead.
Cached input tokens
Previously seen prompt tokens a model can reuse at a steep discount instead of reprocessing them at full price.
Electron app
A framework for building a desktop application with web technologies; all three test builds were packaged this way for Mac.
Resources

Things they pointed at.

Quotables

Lines you could clip.

06:38
“For very similar token expenditures of 69 and 66 million, we have $203 for Astra and then five times less, $43 for a functioning version of that app.”
the single number that proves the cost thesis→ TikTok hook↗ Tweet quote
08:16
“You might bring in Astra as the plumber or the janitor to clean up some of the mess that Sol might create.”
self-contained, memorable metaphor→ IG reel cold open↗ Tweet quote
06:36
“So on time alone, execution time isn't a big differentiator.”
short, contrarian, needs no setup→ newsletter pull-quote↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphor
So OpenAI just dropped their response to Opus 5 .5 in the form of GPT -6 .1 Sol. According to them, it's almost as smart as Astra for a fraction of the price. So in this video, I put that theory to the test and I had both build a very complex product in a short amount of time.
Let's dive in. Now the app I had both build is a Whisperflow clone that can call different types of tools, whether it's calling a Gmail or using something like Firecrawl to do deep research on a variety of topics. And if you pop over to my apps, you'll see I have three distinct ones, all that do the same thing, but were built in different ways.
Relay Voice was built using Astra as the lead and Sol as the executor. And then for the remaining two, I had Astra and Sol use the exact same prompt to go from zero to the final build to show you what the differences are. In terms of economics and tokenomics, when you put both models side by side, these are the differences.
So on input, Astra is $10, whereas Sol is five times cheaper at $2. And it's five times cheaper on output. So Astra is $50 compared to $10 with GPT -6 .1 Sol.
And one thing to pay attention to is the cashed input. So it's $1 with Astra, where it's 10 times less at $0 .10 with Sol. So there's definitely a huge efficiency gain here.
But we want to see, is there an efficiency gain in actual productivity? According to the OpenAI team, they're almost at benchmark parity, but the biggest lags happen in computer use as well as creating professional documents. So when it comes to executing day -to -day tasks that could use an extra jolt of intelligence, something like 6 .1 is better geared for that, whereas Astra primarily would be used for planning and then using that plan and executing it with Sol.
Now I'll quickly show you all three builds, and then I'll show you what prompt was given to execute them and what the stack looked like before I get to the head -to -head comparison in terms of tokens, time, and overall money spent. This was the final product of using Astra for planning and Sol for execution. So when I click on my option key, not only can I get real -time transcription, but I can also do a series of tool calling, whether it's to my Gmail suite or to something like Firecrawl to go and do some deep research for me.
So on the Google test, I could ask it something like this. Can you tell me how many emails I've gotten in the last hour and the first name of every person that's emailed me? So I'm sending it over.
It's going to quickly realize it needs to use a tool. And there it is. It finds my Gmail.
It quickly runs the Google CLI, finds the emails, shortlists it, then translates it, and then pops it up over here where I can see the name of every single person that sent me a message in the past hour. And you can see right here that I can click into every tool call that it made. And if I want to be able to minimize it, then I can drag it wherever I want.
And let's test it out on the research side. I want you to go and research as of today's date, the brand new latest versions of the OpenAI ChatGPT plans and break down the economics of each. Now, this should realize that it's a research -based request.
It will run Firecrawl, quickly run that search at a blazing fast speed, bring back the sources, create the shortlist, and if we move this over and click down, we should see, there we go, the brand new $500 plan, and it comes with the source. So if we click into it, you'll see it's right here, breaks down the economics, and this worked with around five to seven prompts.
And on top of that, we have history, we have settings, and this would look like all kinds of existing voice dictation apps. Now, this was Astra's app, and I would drag it up for you, but it's actually not made to be draggable, even though I asked for it in the plan. But it seems to work, so if I do Option, Shift, hello, hello, hello, I want to test if this is working and how quick it is.
And ideally, can you check my Gmail? So let's see if it transcribes it. Okay, it realizes it needs to use my Gmail.
It's a bit slower, and it opens it up here, which is fine. There we go. It does break it down.
It does show my email. So this one does work. And last but not least, we have GPT -6 Sol, which looks the most basic of all of them and does have a minimization mode, but it is a bit slower and has some latency.
But let's see if it works. So for this one, it should be Option, Command. Hello, I want to speak.
I want to test whether or not this is working. And if you can pick up if I change languages like the other ones. Bonjour, Monsieur.
Bonjour. So let's see if it can handle that. There we go.
So pretty quick. Not the most beautiful. I also can't drag it.
But if I minimize this, if I go to the settings, you'll see that visibly it looks similar. It has the history. You can have changing of models.
But one thing in terms of setting this up, these buttons aren't working the same way that the Astras are. let alone the one that was built by both. And just as a reminder, for the Astra and Sol version, they had the identical prompt and they used the identical stack, which is Gemini for transcription, Firecrawl for research, and then the Google CLI to access things like Gmail, Google Suite, Google Docs, et cetera.
The overall nuance with this build is that this app has to realize whether or not your transcription should only end up as a transcription or if it first needs to go and call some tools. If so, which ones in what order and then transcribe it to you in the UI. And on top of that, they all had to use computer use to not only create what's called an electron app, which is basically a desktop app for Mac.
And as an extra stretch to the models, I tasked them to use their computer use to go and physically use the app, go on settings, enable things like privacy and security and accessibility. So we could actually insert text that's transcribed through the software. Now this behemoth was the prompt that was given to both models, and they were essentially encouraged to create other threads for different phases of the project.
So you can see right here, we have phase one, two, three, four, five, and six. They decided whether or not they needed four phases, five phases, or two to execute the whole task. So if we scroll over, you'll see that they self -organized their own threads, each with a leader thread.
And the whole goal is that each one of these leader threads would check in. on all of these sub -phases to make sure that they're all online, they're all working, and they're all trending towards the overall goal. Now, interestingly, the one that I gave more time to plan with Astra more cohesively and then spin up all the chats with Sol created eight phases.
And this is the one that, out of the box, worked the best and needed the fewest tweaks. Now I'm going to give you the prompt as well as the entire code base that you can one -click clone into your projects. But this was the overall goal for all the models.
Build a quick local prototype called Relay Voice. I want a Whisperflow -style experience. I asked it to go research what Whisperflow is.
With web research and Gmail assistance, optimize for one reliable personal demo, clear design, and inspectable results. You are the lead chat, and this is a big thing. I explicitly authorize you to create new codex threads for phases below.
So for each model, I'm telling it, you have autonomy over how many different chats you can spawn, what tasks they have, your criteria to judge, and how you use computer use. Now fast forwarding to the results, so we can break those down and then see a postmortem of what flaws they had in that process that had to massage.
Step one is in terms of time, they were very similar. So Astron High took 42 minutes, whereas Sol took 40 minutes. So on time alone, execution time isn't a big differentiator.
But in terms of cost, there was a huge difference. So for very similar token expenditures of 69 and 66 million, we have $203 for Astra and then five times less, $43 for a functioning version of that app. Was it as beautiful?
No, but you could probably apply one Astra prompt to beautify it, but use the majority of Sol for the actual build. Now, both models struggled at different things, which is why I'm telling you the hybrid approach where one does the planning, you take that plan, then you spend all the tokens on execution with the other is probably your best bet, especially if you're on the $20 or $100 plans.
Now, Astra struggled with the research workflow. And this is the part where you send over the prompt, it realizes it needs to do some research, then it uses fire crawl, this is my pathetic attempt at flame, and then it realizes whether or not do I need to do one or multiple rounds of search where it might need to go scrape the webpage or do some deep research on a particular site or resource and then come back with a response in a short amount of time.
And in doing that research, when it brought back the citations, a lot of those citations were broken, whereas with Sol, they actually worked out of the box. And with Soul, interestingly, it got a lot of the main features down, but some of the most fundamental ones, like taking voice as an input, I had to struggle with that.
And beyond the single prompt that I gave it, just to make it work for this demo, I had to give it another five to six prompts of it constantly checking why it's not working. Whereas with Astra, this worked on the first try. So even beyond the planning portion, you might bring in Astra as the plumber or the janitor to clean up some of the mess that Sol might create.
But in terms of comparing 6 .1 to 6, where 6 Sol for me was not even usable, not even the conversation for something like Opus or Sonnet, this one is definitely a lot more capable and a lot more efficient. So hopefully that gives you a quick preview for how both models operate and where it might make sense to use each. I'm definitely going to keep testing both of them on day -to -day client work to see if they have an edge when it comes to Opus.
And I'll report back if there's anything material. If you want to grab that Whisperflow clone, I'll make that available along with the entire prompt and the website that I walked you through all the results down in the second link in the description. And outside of that, if you found this video helpful, please leave a like and a comment on the video really helps the video and the reach.
And I'll be back very soon with yet another video. I'll see you in the next one.
The Hook

The bait, then the rug-pull.

OpenAI says its new GPT-6.1 Sol is almost as smart as its flagship Astra model for a fraction of the price. So the creator built the exact same Mac voice app three times, once with each model and once with both working together, and tracked the tokens, the dollars, and the bugs.

Frameworks

Named ideas worth stealing.

01:47concept

Plan with the strong model, execute with the cheap model

Use the more capable, more expensive model to plan and structure a build into phases, then hand that plan to the cheaper model to execute the bulk of the token-heavy work.

Steal forany multi-agent coding workflow with a cost-sensitive model split
CTA Breakdown

How they asked for the click.

VERBAL ASK
08:32product
“Get the Wisprflow clone plus prompt, second link in the description.”

Delivered as a direct-to-camera close after the analysis, pointing to a free Gumroad download rather than a paid pitch.

Storyboard

Visual structure at a glance.

open
hookopen00:00
pricing table
promisepricing table00:45
results reveal
valueresults reveal06:31
CTA
ctaCTA08:47
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.