Modern Creator
Riley Brown · YouTube

Grok 4.6 Is Actually Good, and Claude Keeps Getting Better

A week where xAI undercuts Anthropic and OpenAI on price, and GrokBot's named-agent structure hints at how every AI platform is about to fix the chat-versus-work problem.

Posted
2 days ago
Duration
Format
Talking Head
educational
Views
17.1K
482 likes
Big Idea

The argument in one line.

Every major AI agent platform is racing to fix the confusing split between chat and always-on work mode, and this week's releases show them converging on the same answer: giving each session a name, a purpose, and its own routines, while Grok 4.6 makes that shift cheap by undercutting Claude and GPT on price at comparable capability.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You use Claude Cowork, ChatGPT Work, or a similar always-on agent platform and want to know how the category is about to change.
  • You're deciding whether to add Grok 4.6 to a coding or agent workflow and want a straight price-to-capability comparison against Opus and Sonnet.
  • You want a ten-minute catch-up on what shipped this week across xAI, Anthropic, OpenAI, DeepSeek, and Google without watching five separate videos.
SKIP IF…
  • You've already used GrokBot and Claude Cowork side by side; the platform comparison here won't tell you anything new.
  • You're looking for a hands-on coding benchmark of Grok 4.6 — this is a news recap with screen-share commentary, not a build-along test.
TL;DR

The full version, fast.

xAI shipped Grok 4.6 and a new agent platform called GrokBot the same week Grok 4.6 leads in economically valuable and long professional tasks, not coding, and it costs roughly an eighth of Claude Fable and a quarter of Opus for combined input and output pricing. GrokBot's real innovation is structural: instead of disposable chat sessions like Claude Cowork or ChatGPT Work, each session becomes a named agent with its own system prompt, cloud computer, and routines living inside it rather than in a global scheduler. That same fix shows up in Anthropic's Claude Tag and OpenAI's incoming Workspace Agents, suggesting the whole industry is converging on personified agents as the answer to the 'chat versus work' confusion users have been complaining about. DeepSeek's V4 Pro launch, by contrast, overpromised and got walked back within a day.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0000:49

01 · Intro

Cold open teasing GrokBot, Grok 4.6, Claude and Codex updates, DeepSeek V4, and Gemini news in one agent-native roundup.

00:4904:16

02 · Grok 4.6

Grok 4.6 leads in economically valuable work, long professional tasks, and legal work, but not coding. Riley walks through a developer's Rust-rewrite anecdote and a combined input-plus-output pricing table showing Grok 4.6 at $8 per million tokens versus $30-60 for Opus, Sonnet, and Fable.

04:1611:40

03 · GrokBot

Full walkthrough of xAI's new desktop and iOS agent platform, built with the Cursor team. Every session becomes a named agent with its own routines, a cloud computer, and a teach-a-task record/replay flow.

11:4014:11

04 · Evolution of Agent Platforms

Riley maps the last year of agent platforms into two shapes: the Cowork/Work-style session-list apps (Codex, Hermes, OpenClaw) versus the newer personified-agent shape (GrokBot, Buzz), with Anthropic answering via Claude Tag inside Slack.

14:1117:08

05 · Claude Updates

Claude Chrome sessions now sync across desktop, web, and mobile. Sonnet 5's discounted introductory price was made permanent rather than rising, a move tied to pricing pressure from DeepSeek, Kimi, Z.ai, and Grok. Riley also flags a community PhoneHarness project for controlling a phone from any coding agent.

17:0818:32

06 · DeepSeek V4

DeepSeek's new Pro model launched claiming near-Fable quality at a fraction of the price, but testers found it barely better than the prior Flash model, and it was pulled after backlash over the claim and a switch to usage-based pricing.

18:3219:42

07 · Gemini 3.7 Flash

Google shipped a fast mid-tier model benchmarked only against older mid-tier competitors like Claude Sonnet and GPT Terra, not against any frontier-tier model.

19:4225:13

08 · The Chat vs Work Problem

Riley unpacks why switching between chat and always-on work modes in ChatGPT and Claude Cowork feels broken, argues GrokBot's real innovation is giving every session a name and system prompt, and previews ChatGPT's incoming Workspace Agents and Codex's new Linux release.

25:1325:21

09 · Outro

Sign-off and subscribe ask.

Atomic Insights

Lines worth screenshotting.

  • Grok 4.6 leads its benchmark categories in economically valuable work, long professional tasks, and legal work, but not in coding, which signals xAI is prioritizing general agent tasks over developer tools.
  • Combining input and output token pricing, Grok 4.6 costs $8 per million tokens versus $30 for Opus 5, $35 for Sonnet 5.6, and $60 for Claude Fable, making Fable 7.5 times more expensive for comparable frontier capability.
  • A developer used Grok 4.6 to repeat a Rust rewrite that Fable had one-shotted the week before, finishing it in about ninety minutes for $55, roughly a tenth of the cost of the original Fable run.
  • GrokBot was built by the same Cursor team that shipped the coding tool, and xAI reportedly considered paying about $7 million for the domain dot.com before settling on the GrokBot name.
  • Instead of creating disposable new chat sessions like Claude Cowork or ChatGPT Work, GrokBot turns each session into a named agent with its own description acting as a mini system prompt.
  • In GrokBot, scheduled automations (called routines) live inside the specific agent that owns them, not in a separate global scheduling tab the way Claude Cowork's scheduled tasks do.
  • Every agent created in GrokBot comes with its own cloud computer and browser, which can be taught new tasks through a record-and-replay flow similar to Codex's record and replay feature.
  • The two agent-native platforms that have gone viral this year, GrokBot and Buzz, both broke from the Cowork/Work session-list pattern used by Claude, ChatGPT, and Codex-style tools.
  • Anthropic's answer to the same structural problem is Claude Tag, which puts agents inside a company's Slack, though it is currently limited to Team plans.
  • Claude quietly made Sonnet 5's discounted introductory pricing permanent instead of letting it rise back to full price, a move tied to pricing pressure from DeepSeek, Kimi, Z.ai, and Grok.
  • DeepSeek's new V4 Pro model was launched claiming near-frontier quality at a fraction of the price, but user testing found it barely better than the prior DeepSeek V4 Flash model, and the model has since been pulled offline.
  • Gemini 3.7 Flash was benchmarked against Claude Sonnet and GPT Terra, which are the third-best and second-best models from their respective labs, not against any frontier-tier model.
  • OpenAI is reportedly bringing Workspace Agents, currently limited to Team plans on web, to the ChatGPT desktop app, according to a public reply from an OpenAI team member.
  • A community project called PhoneHarness lets any coding agent, including Claude Code and Codex, take full control of a physical phone.
Takeaway

Agent platforms are converging on named, personified sessions

WHAT TO LEARN

The confusing split between chat mode and always-on work mode in every major AI platform is getting solved the same way everywhere: by giving each session a name, a purpose, and its own automations instead of a disposable chat history.

02Grok 4.6
  • Grok 4.6 leads in economically valuable, long professional, and legal work rather than coding, which tells you where xAI is actually pointing the model.
  • Combined input and output pricing puts Grok 4.6 at roughly an eighth of Claude Fable's cost, so run the real per-task cost math before assuming the pricier frontier model is worth it.
03GrokBot
  • A named agent with its own system prompt and routines is easier to return to and delegate work through than a growing pile of anonymous chat sessions.
  • When automations live inside the specific agent that owns them instead of a global scheduler, it gets easier to remember what each recurring task is actually for.
04Evolution of Agent Platforms
  • Two structural patterns are competing right now: growing lists of disposable sessions versus small sets of named, purpose-built agents; pick the shape that matches how you actually want to delegate work.
05Claude Updates
  • Watch for pricing pressure to keep pushing mid-tier models cheaper: once a frontier model costs less than a lab's own mid-tier model, there's no reason to use the mid-tier one.
06DeepSeek V4
  • Treat a splashy launch claim with skepticism until independent testing confirms it, DeepSeek's V4 Pro was pulled within a day of users testing the claim.
07Gemini 3.7 Flash
  • When a benchmark chart only compares a new model against older mid-tier competitors, check whether any frontier-tier model was left out of the comparison on purpose.
08The Chat vs Work Problem
  • The chat-versus-work confusion isn't a UI bug, it's a structural problem, and the fix multiple labs are converging on is giving sessions identity instead of forcing users to pick a mode.
Glossary

Terms worth knowing.

GrokBot
xAI's new desktop and iOS agent platform, built with the Cursor team, where every chat session is a named agent with its own system prompt, cloud computer, and automations.
Claude Cowork
Anthropic's always-on agent platform inside Claude where users spin up chat sessions that can browse, use tools, and run scheduled tasks.
GPT Work
OpenAI's equivalent agent platform inside ChatGPT, offering the same kind of tool-using, task-running sessions as Claude Cowork.
Claude Tag
Anthropic's platform for embedding Claude agents directly inside a company's Slack workspace, currently limited to Team plans.
Buzz
An agent-native platform styled like Slack, where users create channels containing a mix of AI agents and human teammates.
Routine
GrokBot's term for a scheduled automation that lives inside a specific named agent, rather than in a separate global scheduling tab.
PhoneHarness
An open-source GitHub project that lets AI coding agents, such as Claude Code or Codex, fully control a physical phone.
Resources

Things they pointed at.

10:23productGenspark Super Agent
17:00toolPhoneHarness (GitHub)
12:54toolBuzz
14:15productClaude Tag
19:08toolDeepSeek V4 Pro
18:32toolGemini 3.7 Flash
24:21productChatGPT Workspace Agents
Quotables

Lines you could clip.

03:03
we get $8 for Grok 4.6, $30 for Opus five, $35 for 5.6 Sol, and $60 for Claude Fable
concrete number stack, easy to turn into a comparison graphicIG reel cold open↗ Tweet quote
05:53
hi, Riley. I'm here. What do you want me around for? Could be email, content code, a specific workflow, or something else entirely.
shows the product doing the thing being described, good demo clipTikTok hook↗ Tweet quote
33:21
It's not clear to me that OpenAI realizes how strange the distinction between chat and work feels in practice
sharp critique quotable on its own, sets up the video's core argumentnewsletter pull-quote↗ Tweet quote
23:24
I think the biggest innovation of Grokbot was kind of making this analogy easier to understand
thesis-level line that summarizes the whole GrokBot segmentIG reel cold open↗ Tweet quote
24:00
the automations should not live at a global level. They should just live within the bot.
quotable design opinion, short and declarativeTikTok hook↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphoranalogy
This was the biggest week of the year for Elon Musk and his AI efforts at SpaceX. They released Grok 4.6, which shows that they are finally catching up to OpenAI and Anthropic.
They also released Grokbot, which is their new super app, which will rival GPT work and Claude Cowork. And I have a lot of thoughts about this platform and what makes it so interesting, which I'll talk about today, but we have way more to cover.
We will also discuss the latest updates inside Claude Code and Codecs, as well as the latest DeepSeek v four model that they launched today. And we even have some news from Gemini. This is an agent native update where we put the latest advancements on the frontier of AI agent platforms and models into context so that we can actually use them to improve our business.
Okay. So we have a lot to cover today. Let's dive straight into the SpaceX update.
So SpaceX said this yesterday, introducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price. And so you'll notice here that the three areas that they led were economically valuable work, long professional tasks, and legal work.
And so if you also notice that it isn't coding. Right? They didn't lead in any of the coding benchmarks.
And I believe that this is because they are focused on general agent tasks because this is simply their priority, which explains their brand new platform that they worked on with Cursor, which is called GrockBot.
GrockBot is their brand new super app, which we'll talk about in just a second, that is focused on non coding work. They are trying to get everyone within a company to interact with AI agents to help them get work done.
They're focused on knowledge work. And I still believe that the best and easiest way to use Grok 4.6, this brand new model, is directly inside Cursor.
As soon as you update Cursor, it will default to Grok 4.6 fast, and you can use it directly inside Cursor.
So I've not yet had enough time to actually do a deep test of Grok 4.6, but I do think there's some interesting use cases to discuss that people have posted on Twitter. So here's DHH.
So Fable, I think this was last week, one shotted a Rust rewrite of the terminal text effects Python library in 11,000,000 tokens. If you don't know what that means, that's perfectly fine.
Fable did a really, really hard thing. And then today, he tweeted that he used SpaceX new model, Grok 4.6.
With just a couple of nudges, it was able to repeat this feat in in about an hour and a half, and the key takeaway is that it was only $55. That was about one tenth of the cost of Fable implementation for the same work.
So take a look at the pricing for these models. If we look at Grok 4.6 compared to Opus, Sol, and Fable, right, if we combine the input and output prices per 1,000,000, we get $8 for Grok 4.6, $30 for Opus five, $35 for 5.6 Sol, and $60 for Claude Fable.
And so that means that Claude Fable five is 7.5 times more expensive than Grock 4.6, 5.6 Sol, 4.4 times, and Opus five, and Grok 4.6 is better than Opus.
It is straight up better, and Opus is still 3.75 times more expensive than Grok 4.6. This is a very good model, and it is a reasonable price, and it is legitimately on the frontier.
And beneath this Cognition post, Elon commented, Grok 4.7 will exceed all current models, which includes Fable. He said, that said, Anthropic is a great company and will probably release improved models soon.
However, the SpaceX training corpus is so awesome and unique that I would be shocked if any model is better at real world engineering than 4.7. And so that was their model release. And if you watch my channel, you know that I actually don't dive too deep into model releases.
I care mostly about practical use cases of AI. Like, how do we take these advancements and actually turn them into, like, real business outcomes? Or how do we actually improve our productivity or make our lives better?
And so now we're going to the next update by SpaceX, is their new GrockBot platform. This is GrockBot, and GrockBot is a desktop app and an iOS app. This right here is the desktop app.
So this app right here was being worked on by the Cursor team. So Cursor was working on this platform for many months, I think like four or five months, and this was going to be their general knowledge worker platform. Cursor is the coding tool, and this platform Grokbot, which internally they were calling SAND, and I believe the name that they were going to use was actually dot.
They bought dot.com for like $7,000,000, and instead they went with Grockbot.
And this platform was going to be Cursor's version of Claude Cowork or GPT Work, but it has some key differences that I think makes it pretty unique. And so one of those things is that instead of creating new sessions all the time like you do in GPT Work or Claude Cowork, where on the left side panel, right, if we were to go to Claude, and if you were using co work inside Claude, you just see all these different, like, chats that get lost while you use it.
What Grockbot did is they just said each session, each one of these sessions is like its own agent.
And so instead of like creating a bunch of new sessions, if we just create a new bot here, what it does, instead of it being like a new session where you just go and type in your request, it'll immediately try and figure out what the purpose of this session is, and then it will actually name the agent. And so it's like, hi, Riley.
I'm here. What do you want me around for? Could be email, content code, a specific workflow, or something else entirely.
Weekly agent updates.
Look at my YouTube and Notion to get context. Your job is to help me with these every week. And so you can honestly think of this as each one of these is like your own little bot.
And you can set up plugins just like any of the other platforms, and you can also set up skills. Like, I have my scrape creator skill right here, and all of these agents share plugins and skills.
It's just that these new bots have their own name, title, and little description. And then when you create automations, they get added right here, and so each session or agent has its own automations or they call them routines.
And you can see here, it just updated help Riley with weekly agent updates by pulling context from YouTube. And I'm gonna say name yourself weekly update.
And then I can change the title, and that will just change this little tag right here. And so I can put title as, like, help with updates.
I don't know. And it just shows up right here. And so this agent has a very specific role for me.
It just helps me with these weekly updates. Helps me do research, and that is gonna be the purpose.
And so whenever I wanna work with this, I would just come to the weekly update agent. And I could say, hey. Every weekday, present me with the AI news for the day at 10AM.
And so I can just ask the weekly update bot or, yeah, Grock bot to create a routine. And so it created this routine and you can see it right here.
If we open up this side panel, can see that we have this weekly agent updates and then we also have weekly AI news. It's 10AM on a weekday, deliver Riley's daily AI news briefing in chat. So it'll give it to me every day at 10AM.
And so this is fundamentally different than Claude Cowork. Right? We could go to Cowork, and we could say every morning at 9AM, do a task.
You can see here it created this morning hello. And notice here that, like, we have all these different chats. Most of these chats, I'll never return to.
It's hard to return to them because, like, these just kind of get lost. And so it's hard to, like, pick up pick back up on previous work that we were working on.
And so you'll notice here that Claude Cowork, when I create this chat, it adds scheduled as, a global setting. So it doesn't really have much to do with this chat session anymore.
The scheduled tasks, and if you have like a ton of scheduled tasks, they live up here in this scheduled section, whereas in Grok bot, they actually live inside of the agent itself or inside the session.
So I created this new session. Now I have a weekly update bot and the routines live within here.
For example, my partnership bot has its own it has its own routines. My content bot has its own routines.
Like, it scrapes from all my favorite creators every morning at 09:16AM. But these bots have different cron jobs or routines that live inside the agent itself.
And I if I press command k, I can get a kind of a zoomed out view of all the different routines that I have. And here, it will actually list the name of the agent and the name of the routine.
So I can see all the weekly update agent routines, the developer routine.
We also have a partnership bot routine, and then I have a to do list bot which prints my reminder to do my most important task every single day. I forgot to mention that every single agent that you create comes with its own little computer.
So your it runs in the cloud. So this is a full computer in the cloud that you can use, and your agent most importantly, your agent can use this browser, and you can sign into your stuff on this browser, and you can even teach a task.
And when you teach a task, you can record yourself. This is very similar to record and replay on Codex. For those of you who watch my content, you could do the same thing, but you use your own computer.
Here, I'm actually teaching the agent to do a task in its computer. Right?
This is a virtual computer running in the cloud. You can actually go in and see all of its files on the computer. It's its own thing that you have full visibility into, which is a very new and interesting thing, especially in a platform like this.
Real quick before the next update, I wanna talk about an update by the sponsor of this video, Genspark. One of my biggest inspirations for getting into agents in the first place was to keep track of everything I do to get things done. I talk for a living.
Podcasts, calls, meetings, random ideas in the card, normally 90% of that just evaporates. So a few months back, I started clipping this to the back of my phone, Genspark's second brain note. I hit record when something's worth keeping.
There's a physical light, so it's never a guessing game whether it's on. It's SOC two and ISO 27,001 certified, and it works in over a 100 languages.
So I use it everywhere, not just at my desk. Here's the part that actually got me. It's not just a recorder.
Twice a day, it reviews what I said and figures out what to do with it. Told someone I'd send them an email, it has the draft ready. Agreed on a time on a call, it's sitting on my calendar ready for approval.
Rift on a video idea out loud, there's a script waiting inside Notion. That's second brain. It remembers.
The part that does the work is Genspark's super agent. Second brain retrieves, super agent executes.
All I do is say yes. They just opened the first release to the public Genspark second brain note, 10% off through the link in the description, speaking of agents that actually get things done for you.
So the last thing that I'll say on this is it's important to note the evolution of these general agent platforms, and I wanna take a quick look into this. You know, in the end of, like, 2025, so if this was, like, end of twenty twenty five and this was kind of the first half of twenty twenty six, we got these platforms that kind of looked very similar.
They were kind of like the evolution of Claude code running in your terminal, and they've evolved into these, like, super apps. So, like, even the Hermes agent looks a lot like Claude Cowork.
Codex has GPT work, which looks similar. The OpenClaw desktop app looks a lot like this, where it's this kind of agent platform where you have a bunch of sessions.
You have skills, plug ins. You have artifacts, automations, yeah, which are, like, scheduled, and they're starting to look relatively similar.
It is very interesting to note that the two previous general agent platforms that have gone viral, which are Buzz and GrockBot, have a new shape to them.
Right? Here, GrockBot, you can kind of like see your team of AI agents. I showed you that you can create new agents really quickly.
Right? You can just create agents. You can give it a purpose, and each one has its own routines.
And then the platform that went viral before this was Buzz. And so this platform, Buzz, allows you to create channels with a bunch of different agents, and you can even add people to your Buzz.
And so Buzz looks exactly like Slack, and it's kind of like this agent native Slack or Slack meant to be used with humans and agents, which I find to be very interesting. And publicly, Anthropic hasn't really talked about Cowork that much.
They talked more about Claude Tag, which is their new platform that allows you to basically create agents inside your company's Slack. To my knowledge, you can only use it with Teams, but this is kind of their focus right now, is creating an agent for groups of people or enterprises or small businesses.
I feel like that's kind of the shift that we're moving into. So maybe the first half of twenty twenty six or the first, you know, first two thirds of the year was about the personal agent, and the rest of this year is kind of about how do you create your own personal team of agents. And then in in regards to Buzz and Claude Tag, how do you add an agent so that your entire team can get access to the same agent so you can collaborate using AI agents.
Okay. So now let's discuss updates coming out of Anthropic. Yesterday, Claude announced that your Claude Chrome sessions now carry over to desktop, web, and mobile.
Conversations are saved, and your skills and connectors work inside the browser. Basically, what Claude and OpenAI are doing is they are inserting GPT work in the case of OpenAI and Claude Cowork in the case of Anthropic directly into your browser.
I actually have both of them set up. I can use Claude here.
And what they announced is I can say, tell me more about this. And whatever I type in here is basically the same as using Cowork.
It has access to the same connectors, the same skills, and everything, and all of the chats, right, I can view the history, all of these chats will actually sync into my Claude app.
I can very easily move this conversation over to the Claude desktop app. And you can see here it named this chat more information request, and I can go back to Claude, and I can see that more information request is right here.
So I can very easily switch from the chat that I had in any browser. Right? It's just a Chrome extension back to Claude.
So it's equivalent to coming here and using Claude except you can do it directly from your Chrome extension in Chrome. The next update to Claude is pretty interesting.
Sonnet five had an introductory price, and it was scheduled to increase back up in price at a certain date.
But Claude has decided that it would actually stay at this cheaper price. OpenAI is doing something similar with their Terra and, uh, Luna models.
So, basically, their non frontier models are getting much cheaper. And I believe, and many believe, that this is due to the pressure from Chinese models out of DeepSeek, Kimi, z dot a I, and then now also Grok. Right?
Because if their middle models are way more expensive than these other alternatives that are actually better, then people have no reason to use them. So this is causing them to lower their prices, and I expect this to continue.
The non frontier models from Anthropic and OpenAI will continue to get cheaper if the competition from China and in The US continues such that their frontier models are cheaper than their middle models. Everyone's just gonna use these models, they'll never use the middle models like Sonnet and Opus and Terra and Luna.
And for the final update regarding Claude, this guy released PhoneHarness. So this isn't actually coming directly from Anthropic, but if you look look up GitHub phone dash harness, you will find a repo which will allow you to fully control your phone from any agent, not just Claude Code, but Claude Code, Codex, etcetera.
You can fully control your phone with an AI agent. I haven't tested this out. I just thought this was really cool.
Thought I'd share it really quickly. Okay. So we were about to talk about the brand new DeepSeek v four model, which was just released this morning officially.
And DeepSeek said that we're, uh, launching DeepSeek v four today, and they claimed that it was nearly as good as Fable. And here's a very quick summary after summarizing everything and everyone's takes on Twitter.
I'm trying to figure out this story, but here's a concise, uh, summary here. So apparently, DeepSeek released a new pro model last night and talked like it was almost as good as the top expensive models like Fable, um, and it was only a tiny fraction of the price.
However, many people started testing this and they came back and they said no. It's only a little bit better than DeepSeek's previous model, which was DeepSeek v four Flash.
And so this was embarrassing and they have basically taken the model down. In fact, if you go to Cursor right now and you switch to DeepSeek v four Pro, I'm using this via OpenRouter.
It's not an official model inside Cursor yet. And if I say hi, it actually will immediately fail.
And so this model just doesn't work right now, and so I can't fully cover it because there is some backlash. Apparently, it's not that much better than DeepSeek v four Flash, and they're also getting some flack for their pricing.
So apparently, the pricing has come up and they went to usage based pricing. But luckily for us, fifty eight minutes ago, Gemini officially released 3.7 flash.
And apparently, this is a very fast model and it's brand new. And I do notice here though, a red flag, is that the models they're comparing against, right, you can see Gemini Flash is being compared to Claude Sonnet and GPT Terra.
So this is the third best model by Claude and the second best model by OpenAI, that's what they're comparing it to. And so they haven't released a new pro model in a while and so this 3.7 flash is a fast mid tier model.
And if you wanna test it out, again, I like to do it inside Cursor. I have it running because I'm using OpenRouter. And if you were to sign up for OpenRouter and just ask Cursor how to set it up, you get it set up in like two minutes inside Cursor.
But I can say, hey. Or you could use it inside the anti gravity.
I'm sure they have it inside anti gravity. You can use the new Gemini 3.7 model. Since it's only a mid model and it's not that good, it's just like pretty fast.
I don't think I'll be testing it too much, but if you wanna test it out, you can. Okay. So for the final big thing that I wanna talk about today is I wanna talk about what I call the chat versus work problem.
Many people are confused about how to use, uh, AI agents, specifically the Claude Cowork and the GBT work features inside these platforms.
Uh, Signal said, it's not clear to me that OpenAI realizes how strange the distinction between chat and work feels in practice and how fragmented the entire ChatGBT experience has become. If you leave it on work, every simple question or search becomes an expedition. It starts thinking, planning, and using tools when you just wanted a quick answer.
What he's talking about is in chat, you now have chat and work. And these are two very different products. Obviously, you know what ChatGPT is.
But gbt work is a bigger thing. Right? Here, you can actually do things on Slack.
You can have it control your Gmail. You can have it fully control your Notion. You can literally get if you set up the right scheduled tasks, you can have it fully reply and send out emails, and you can get it to run like it it is a full agent platform.
And what he's talking about here, it's just hard to understand the distinction between chat and work. I made a full video on g b t work. It's incredibly powerful, but the distinction between chat, work, and then the Codex app is pretty confusing right now.
He goes on to say that there's no good default. Chat is often too limited for real tasks. Work is too slow and cumbersome for normal queries.
Constantly switching between them makes the entire product infinitely more complex. Also, lack of sync between mobile and desktop, plus what's a local chat versus a cloud chat is a mess.
And, again, these are things that I talked about in my, uh, last video. It is relatively confusing, the difference between a local chat and a cloud chat.
Um, ChatGPT went from the most usable, simple consumer experience to confusing AF in such a short period of time. Pretty nuts. I do think it's incredibly powerful, but I do think he's talking about a very difficult problem, uh, which I call the chat versus work problem.
And I tweeted this about the brand new GrockBot yesterday. I showed you the GrockBot platform earlier, and then I tweeted this specifically around how the platform is set up.
I said that I think Grockbot is behind Codex and GPT Work for a lot of reasons. There's a lot of things that GPT Work can do that Grockbot cannot do. But I will say their biggest innovation is personifying the chat sessions.
An agent or a Grock Bot is basically a named session or a named chat session with a mini system prompt. For example, on Grock Bot, if you just click on the name, right, this description is just a mini system prompt. And all of the plugins that you can create, right, I can add Gmail, that is a plugin, I can also use Skills, and all of these skills are global to all of my agents.
They basically just made each chat session have its own system prompt and your its own name so that when I want to create a content script, I'll just go to my content agent. Or if I want to scrape from social media, I'll go to my content agent. If I want to do my weekly update agent, I'll go to my weekly update agent.
If I need something handled in my partnership bot, I will go to my partnership bot. Not only that, but like if something happens inside Slack in this Slack channel, it will automatically ping me, and it stays organized by these chat sessions.
And so I think the biggest innovation of Grokbot was kind of making this analogy easier to understand. And so I think we're gonna see a lot of innovation over the next three months as the Frontier Labs try and figure out what is the best setup for people to use AI agents in their business, and I think there's just gonna be a lot of innovation.
Because this is something I would have never thought I wanted, but once I used it, I was like, okay. This actually makes more sense.
The automations should not live at a global level. They should just live within the bot. I have a feeling that the other labs are gonna copy this design because I do find it's a lot easier to get started and it just intuitively makes sense that you have a bot.
Each bot has its own computer, and it has its own routines. I find it to be an easy interface to pick up and understand. Another thing is that Workspace agents are coming soon to ChatGPT, and I know this because I quote tweeted their Workspace agents.
And so Workspace agents are only on the GPT team plans. It allows you to create these like agents. Right?
You can create agents, and these are as powerful as GPT work, except they do have their own system prompts. You can message these custom agents, and you can even add them to Slack and message them through there. And it's only available to Teams.
And so I tweeted this. I said, wish this wasn't only a Teams plan and only on web. This should be on desktop as well.
And the someone from the OpenAI team, Andrew, said, yes. So this indicates that this is coming very soon to the ChatGPT desktop app, and it won't just be a Teams plan, which I think is really cool.
This is this is the most slept on OpenAI product yet, in my opinion. And then finally, one update is that Codex, you can get and download on Linux.
So Codecs, the Codecs app or the ChatGPT app where I use Codecs, you can get this on Linux now.
So that is a new update on the Codecs side. And, yeah, that's basically everything. So this was an agent native update.
I'm Riley Brown. Thank you guys so much for watching. And, uh, please like.
Please subscribe. It helps me out a ton. I'll see you here for the next
The Hook

The bait, then the rug-pull.

Riley opens by misnaming the company (SpaceX instead of xAI) before diving into the week's real story: Grok 4.6 shipped alongside a new agent platform, GrokBot, that structures itself completely differently from Claude Cowork and ChatGPT Work, and he thinks that structural difference, not the model itself, is the more important release.

Frameworks

Named ideas worth stealing.

02:56list

Combined price per million tokens: Grok 4.6 vs Claude vs GPT

  1. Grok 4.6 — $8
  2. Opus 5 — $30
  3. Sonnet 5.6 — $35
  4. Claude Fable — $60

Riley combines input and output price per million tokens across four frontier-tier models to show Grok 4.6 costs roughly an eighth of Fable and a quarter of Opus for comparable capability.

Steal forAny comparison of AI tool ROI by real cost-per-task instead of headline benchmark scores.
11:40model

Two shapes of agent platforms

  1. Cowork/Work shape (Claude Cowork, GPT Work, Codex, Hermes, OpenClaw): a growing list of disposable sessions, with skills, plugins, artifacts, and automations living at a global level
  2. Personified shape (GrokBot, Buzz): a small set of named agents, each with its own system prompt, routines, and cloud computer

Riley groups the last year of AI agent platforms into two structural patterns, distinguished by whether automations live at the global/session level or inside a named, personified agent.

Steal forDeciding how to structure any multi-agent tool or internal AI ops setup.
CTA Breakdown

How they asked for the click.

VERBAL ASK
10:23product
10% off through the link in the description

Mid-roll sponsor read for Genspark's SecondBrain Note, demoed on-camera clipped to Riley's phone, narratively tied to the episode's recurring theme of agents that execute autonomously.

MENTIONED ON CAMERA
Storyboard

Visual structure at a glance.

cold open
hookcold open00:00
pricing table
valuepricing table02:56
GrokBot demo
valueGrokBot demo05:53
sponsor read
ctasponsor read10:23
platform evolution
valueplatform evolution12:54
DeepSeek V4 walkback
valueDeepSeek V4 walkback17:08
chat vs work thesis
valuechat vs work thesis22:29
Frame Gallery

Visual moments.

Watch next

More from this channel + related breakdowns.

47:51
Riley Brown · Interview

OpenAI Merges ChatGPT and Codex

Riley Brown and Ras Mic dig into GPT-5.6, Codex's background computer-use, and why self-scoring agent loops are turning coding tools into a general operating system.

July 12th