Codex's Browser Agent Automates Literally Anything
A screen-recorded walkthrough of Codex's built-in browser and computer control, from QA-testing a website to drafting an X article on its own.
August 13thOne prompt, nine AI agents, a $318 bill — and a lesson about when "maximum effort" actually pays for itself.
GPT-5.6 Sol on Codex's Ultra tier can autonomously chain research, scriptwriting, voice cloning, avatar generation, editing, and self-QA agents to produce a finished narrated video from a single prompt, but running that many agents in parallel at maximum effort burns roughly double the tokens a lower effort tier would need for a similar result.
A creator handed GPT-5.6 Sol, running at Codex's top "Ultra" effort tier, one prompt and walked away. Sol researched the launch, wrote a script in his voice, cloned his voice and avatar, edited the cut, and ran its own frame-by-frame QA before publishing — no recording, editing, or review from a human. The single task quietly fanned out into nine sub-agents and roughly 458 million tokens, costing about $318 at Codex's official rate card. The real lesson lands in the second half: Ultra's parallel-agent coordination tends to overthink and overdelegate, likely doubling the bill versus what a "High" effort run would have needed for a comparable result — so the creator's day-to-day default stays at High, reserving Ultra for tasks that genuinely need the extra coordination.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →
Sol-narrated, HeyGen-avatar Nate states the premise directly: he never recorded, edited, or reviewed this segment.

Dark dashboard-style motion graphics cite Terminal-Bench 2.1 (91.9% vs 85.6% for GPT-5.5) and BrowseComp (92.2%), plus a 13-task internal test scoring 97% of available points.

Script split into sub-60-second chunks for voice consistency; HeyGen avatar regenerated via browser automation to lock the newest motion engine; HyperFrames edit anchored every visual to the exact transcript phrase that triggered it.

Dedicated agents inspect rendered frames for continuity, avatar presence, and text overflow, and fact-check claims against OpenAI's release notes; any failed frame triggers another fix-and-render cycle.

Real Nate, screen-recording his agent logs: 9 sub-agents, ~458M total tokens, ~86M on the main agent, ~$318 at Codex's official Sol rate — cross-referenced against an OpenRouter pricing table showing Sol at roughly half the per-token price of Claude Fable 5.

Ultra's parallel-agent coordination likely doubled the cost versus a High-effort run for similar output; the lesson is giving capable models vague, outcome-first prompts with delegation and verification built in.
A single prompt can now fan out into nine coordinated AI agents that write, voice, animate, edit, and fact-check a finished video — but running that chain at maximum effort roughly doubles the bill for the same result.
“So I gave GPT 5.6 this prompt, walked away, and when I came back, I got this.”
“That is what Sol is really good at, holding on to the outcome while everything between the prompt and the result keeps changing.”
“I think because it was on Ultra it tended to sort of overthink, overdelegate, and that's where the tokens really started to add up.”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
The first three minutes of this video are not real — no camera, no edit, no review. Nate Herk gave OpenAI's new GPT-5.6 Sol model, running at Codex's highest "Ultra" tier, a single prompt and walked away; what came back was a fully scripted, voiced, animated, and self-verified video wearing his own face and voice. Then the real Nate shows up to break down exactly what that cost.
The six-stage chain Sol's agents ran end-to-end from a single prompt, each stage handed to a different tool or agent and checked before the next.
“if you enjoyed, please leave a like, and I appreciate you guys making it to the end”
brief verbal ask at sign-off, no on-screen graphic
00:00
00:05
00:10
00:15
00:18
00:22
00:26
00:30
00:34
00:38
00:42
00:46
00:50
00:54
00:58
01:02
01:06
01:10
01:14
01:18
01:22
01:26
01:30
01:34
01:38
01:42
01:46
01:50
01:55
01:59
02:03
02:07
02:11
02:15
02:19
02:22
02:27
02:31
02:35
02:39
02:43
02:47
02:51
02:55
02:59
03:03
03:07
03:11
03:15
03:19
03:23
03:27
03:31
03:36
03:40
03:44
03:48
03:52
03:56
04:00
04:04
04:08
04:12
04:16
04:20
04:24
04:28
04:32
04:36
04:40
04:44
04:48
04:52
04:56
05:00
05:04
05:08
05:12
05:16
05:20Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
A screen-recorded walkthrough of Codex's built-in browser and computer control, from QA-testing a website to drafting an X article on its own.
August 13thA hands-on tour of a synced, phone-controllable AI agent team — agent computers, teachable skills, scheduled routines, event triggers, and where it stops making sense versus Claude Code or Codex.
August 12thNate Herk breaks down Boris Cherny's YC interview on cutting 80% of Claude Code's system prompt, then tests deleting his own skills to see what actually changes.
August 12thEighteen real Claude Code sessions later, the model that's half the price per token isn't automatically the cheaper one to actually run.
July 24thA skill file extracted from Fable's own leaked system prompt lets a cheaper model borrow its judgment, without paying for its intelligence.
July 7thA live demo of turning Andrej Karpathy's LLM-Wiki idea into a self-organizing, cross-linked second brain with Obsidian and Claude Code.
July 3rd