Cancel Your Subscriptions, Ox-Alpha Is Here (GLM-5.3-Flash)
The mystery model that took over OpenRouter turns out to be a Chinese open-weights release that nearly matches frontier intelligence for a few cents a task.
August 29thEleven power-user habits from someone who has logged over a thousand hours in OpenAI's Codex CLI — model tiers, thread delegation, safety hooks, and remote control from a phone.
Codex's real power shows up once you go past a single chat window — picking the right model/effort tier per task, delegating work to spawned sub-threads, running unattended goal-loops, and hard-blocking destructive commands with a hook are what separate casual use from a thousand-hour power user.
After 1,000+ hours in Codex, the presenter's real habits go past picking a model: for hard problems use the largest tier (Sol) at high effort, for everything else use the smallest (Luna) at high/extra-high effort, since mid-tier Terra scores worse than it costs. Threads can spawn and supervise other threads with their own model/effort, which turns one chat into a small team. AGENTS.md needs a stale-rule cleanup pass after every model release. Goal-loops let an agent run for hours against a private benchmark until a pass-rate target is hit, no re-prompting required. After a viral post about GPT-5.6-Sol deleting a user's files, the video walks through adding a PreToolUse hook that hard-blocks destructive shell prefixes (rm -rf /, deleting $HOME, disk erase commands) before they execute, plus setting approvals to 'approve for me' instead of full unrestricted access.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →
Credibility hook (1,000+ hours in Codex) and a subscribe ask before the tips start.

GPT-5.6's Sol/Terra/Luna sizes plotted on an intelligence-vs-cost chart; Terra scores worse than Luna at higher effort, so the rule of thumb is Sol-high for hard problems, Luna-high/extra-high for everything else.

Threads can see and spawn other threads; a rule in AGENTS.md can auto-spin a new thread on a specific model whenever the operator says 'deploy,' and one thread can supervise several running in parallel.

Asking Codex to review AGENTS.md for stale rules after a new model release, then telling it to clean the file up.

Mid-roll sponsor segment for Zapier's MCP, which connects Codex to 9,000+ third-party apps (Gmail, Trello, Asana, Google Docs, etc.).

Codex's built-in browser can import cookies/passwords and perform real-world tasks; demoed by moving two files into a new archive folder in Google Drive.

Skills extend what Codex can do; demoed by installing Matt Pocock's public skills repo via a pasted GitHub URL.

Goals/loops let Codex work unattended toward a target for hours; a Loop Library and a 'Loopy' skill help find or draft loops, demoed with a benchmark-improvement loop targeting a 90% pass rate.

Pairing a phone with a desktop Codex session via QR code to view and control running threads remotely while the work still executes locally.

A real incident (GPT-5.6-Sol deleting a user's Mac files) motivates adding a PreToolUse hook with explicit forbidden command prefixes (root/home deletion, disk erase/partition) as a hard safety backstop.

Approval settings compared — full/unrestricted access (what the presenter personally uses but doesn't recommend) versus 'approve for me,' which only pauses on commands flagged as potentially severe.
The gap between casual and power use of an AI coding agent isn't which model you pick — it's whether you delegate to sub-threads, audit your instruction file, run unattended goal-loops, and hard-block destructive commands before they can run.
“For your hardest problems, go with Soul. For everything else, go with Luna.”
“You literally say, here's the overall goal. I'm not even gonna tell you how to solve it. Just go solve it and keep working until you do.”
“GPT five point six Soul just accidentally deleted almost all of my Max files, and this is why I trust Fable a thousand times more.”
“I usually leave approvals as full access, unrestricted access to the Internet, any file on your computer, and I actually would not recommend doing that for most people.”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
The pitch is credibility, not novelty: after a thousand-plus hours in Codex, the presenter says even advanced users are missing pieces of it. The video backs that up mid-way through with a real, current incident — a viral post about GPT-5.6-Sol deleting a user's files — that turns the back half from tips into a concrete lesson on locking an agent's shell access down before it runs a command you didn't mean to approve.
Plot any tiered model family's intelligence-vs-cost chart and pick from the top-left quadrant rather than assuming the mid-tier model is the safe default — here Terra scores worse than Luna at higher effort.
One agent thread spawns others on a different model/effort and periodically checks whether they're still progressing, instead of one person babysitting every task manually.
Re-review the persistent instruction file for rules that applied to a prior model generation but misdirect the current one, every time a new model version ships.
A hook script reads the exact shell command before it executes and hard-blocks specific destructive prefixes (rm -rf /, rm -rf ~, rm -rf $HOME, disk erase/partition/format) with a stated justification per rule.
A middle-ground autonomy setting that runs low-risk commands without asking, but pauses for approval on anything the model itself flags as potentially severe or catastrophic.
“go check out Zapier MCPs. I'm gonna drop a link down below so you can find it easily.”
Mid-roll sponsor segment placed right after the two free power-tips (model selection, thread delegation), demonstrated live inside Codex's own plugin/MCP settings rather than cut away as a separate ad break.
00:00
00:18
00:24
00:36
00:59
01:05
01:17
01:28
01:40
01:52
02:09
02:16
02:28
02:39
02:51
03:03
03:15
03:27
03:39
03:51
04:02
04:14
04:26
04:38
04:50
05:02
05:14
05:25
05:37
05:49
06:04
06:13
06:21
06:36
06:48
07:00
07:12
07:24
07:36
07:48
07:59
08:11
08:23
08:35
08:47
08:59
09:11
09:22
09:34
09:46
09:58
10:10
10:22
10:33
10:45
10:57
11:09
11:21
11:28
11:46
11:55
12:08
12:15
12:32
12:44
12:55
13:08
13:19
13:31
13:41
13:55
14:07
14:19
14:30
14:42
14:54
15:06
15:18
15:35
15:42Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
The mystery model that took over OpenRouter turns out to be a Chinese open-weights release that nearly matches frontier intelligence for a few cents a task.
August 29thAn early-access hands-on with OpenAI's new GPT-6 Astra model — benchmark scorecard, alignment numbers, API pricing, and a run of 3D game and browser-automation demos.
September 3rdA full walkthrough of Anthropic's twin release, the pricing math, the enterprise data compromise, and the independent benchmark that says the cost story doesn't add up.
September 1stA live reaction to Anthropic's Claude Opus 5 launch, walking chart-by-chart through benchmarks that put a mid-tier-priced model ahead of Anthropic's own flagship on almost everything except cyber exploitation.
July 24thA 'dot' release plays out like a full generational leap: two five-to-seven-day unsupervised coding runs, a sponsor benchmark, and a live pricing and capability standoff against a rawer, higher-ceiling rival model.
July 9thRas Mic walks through the seven habits, plugins to slash commands, that turn OpenAI's Codex from a prompt box into a full agentic coding workflow.
August 14th