Trust Me, You're Sleeping on Cloud Agents
A walkthrough of the four ways cloud-hosted coding agents replace manual bug verification, QA, and code review — plus the setup behind all of it.
July 15thA creator who pays for every major AI coding subscription ranks Codex, Claude Code, Cursor, and Devin on models, subsidies, and harness quality — then reveals how to get your employer to cover the bill.
Codex wins on overall subscription value not because its harness is the best of the four — Cursor's is — but because OpenAI subsidizes usage the hardest and its models are the cheapest per unit of coding output.
After viewers complained he never talks about cost, the creator ranks four AI coding subscriptions. Codex, Claude Code, and Cursor each run their own frontier model (GPT, Claude, and Grok respectively), while Devin mixes in its own SWE models; Claude's Fable 5 is the smarter model, but GPT 5.6 Sol is cheaper and nearly as capable for most coding work. All four subsidize usage, but Codex resets limits most aggressively. Cursor's harness scores highest for raw output (9/10) against Codex's 8/10 and Claude Code's 7/10, yet Codex is named the single best subscription once model cost and subsidies are weighed in. The video closes with Codex power-user tips and a script for getting an employer to cover the bill.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →
Comments called him out for never talking about cost; he agrees, tells the Pizza Iola story to prove he knows how to squeeze a deal, then names the four contenders: Codex, Claude Code, Cursor, and Devin.

Breaks down which model powers each harness (GPT for Codex, Claude for Claude Code, Grok/all models for Cursor, SWE/all models for Devin) and calls Fable 5 the smarter model but GPT 5.6 Sol the better all-around pick for cost and general programming.

All four products subsidize usage, but Codex/OpenAI resets and extends usage the most aggressively — he hits Claude Code's limits far faster in daily use.
Ad read for Blacksmith, a CI runner that cut his team's build times from roughly 8-10 minutes to under 2.5 minutes with a one-line config change, plus an agent (CodeSmith) that opened the PR for him.

Uses CursorBench to show GPT 5.6 Sol beating Fable 5 on High while costing less ($5.69 vs $8.77), then reveals harness scores: Codex 8/10, Claude Code 7/10 (saved by its Workflows sub-agent feature), Cursor 9/10 for best raw output — but declares Codex the overall winner once model, subsidy, and cost are combined.

Walks the pricing ladders: Codex $100/$200 (ChatGPT Pro 5x/20x), Claude $20/$100/$200, Cursor $20/$60/$200 (with new 'generous limits' for Grok and Composer), and Devin $0/$20/$200 — all converging near the same $200 ceiling.

Recommends exploring Codex's plugin store, singling out Convex as his favorite for spinning up real-time backends by tagging it directly in chat, then demos the in-app browser running and testing a local app, including a standing '/goal' prompt to hit a Lighthouse score of 100.

Demos Codex's computer-use agent driving a real browser autonomously, shown filling out a RoboForm test form field by field and stopping short of entering a credit card number.

Advises spinning up new threads instead of maxing out context on one long thread, and lays out a model-effort mapping: x-high for coding, light-to-medium for knowledge work, and a warning to avoid Ultra, which he calls token-hungry and low quality.

Closes with the promised trick: pitch your employer to cover the $200/month subscription as a productivity investment and tax write-off, including a scripted pitch aimed straight at the viewer's boss.
Across every metric that isn't raw output quality, Codex's cheaper models, heavier subsidies, and power-user features make it the best AI coding subscription for anyone watching their budget.
“I used to go to Pizza Iola in Toronto for $7. You got three slices... and a dip and a drink for $7.”
“It's not surprising that the winner is Codex.”
“The Claude desktop app is weak. Terrible. I am not a fan at all.”
“Don't use ultra... It's terrible. It's bad. It guzzles tokens.”
“Ask your workplace... that AI will make you more productive and that they should pay or subsidize your subscription.”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
Viewers kept telling him he never talks about cost — so the creator who runs every major AI coding subscription at once breaks down exactly what you're paying for across Codex, Claude Code, Cursor, and Devin, and which one actually deserves your $200 a month.
A 3-row scorecard comparing Codex, Claude Code, Cursor, and Devin across which models they run, whether usage is subsidized, and a 1-10 hands-on harness score.
Maps Codex's reasoning-effort settings to task type: knowledge work sits toward light/medium, coding sits toward x-high, and Ultra is flagged as wasteful for almost everything.
A scripted pitch framing a $200/month AI subscription as a tax-deductible productivity investment the employer should cover rather than a personal expense.
“Thanks to AI, my team and I have been shipping a lot... Blacksmith has solved my problems... The link is in the description down below.”
Mid-roll sponsor read woven into the cost-comparison narrative (ties CI cost savings to the video's cost theme), delivered straight to camera with on-screen GitHub PR proof of the one-line fix.
00:00
00:24
00:40
00:56
01:12
01:33
01:44
02:00
02:20
02:32
02:48
02:57
03:17
03:30
03:52
04:04
04:20
04:38
04:50
05:12
05:30
05:50
06:01
06:23
06:37
06:47
07:05
07:21
07:37
07:53
08:17
08:22
08:41
08:57
09:13
09:29
09:39
09:55
10:19
10:33
10:50
11:08
11:22
11:43
11:47
12:12
12:22
12:42
13:02
13:14
13:24
13:49
14:03
14:11
14:34
14:55
15:06
15:24
15:38
15:54
16:05
16:27
16:43
16:59
17:12
17:31
17:47
18:00
18:19
18:35
18:51
19:04
19:20
19:46
19:55
20:11
20:27
20:43
20:59
21:15A walkthrough of the four ways cloud-hosted coding agents replace manual bug verification, QA, and code review — plus the setup behind all of it.
July 15thHow one developer chained a computer-use verification loop and an automated code-review loop to ship 75+ pull requests with an AI cloud agent, without reading most of the code.
July 9thA builder walks through every infrastructure decision behind his own AI-agent product, arguing that system design - not AI - is what separates a prototype from an app that survives real users.
July 2ndA 22-minute live-demo tutorial showing how design tokens and Claude Design eliminate the inconsistency that makes AI-built apps look cheap.
June 23rdClaude Desktop and the Codex app both quietly shipped native, multi-tab in-app browsers within days of each other — used right, one prompt now opens, drafts, and stages a dozen tabs at once.
July 15thElie Steinbock walks Greg Isenberg through the build-verify-learn loop he's using to run SEO, ads, and product feedback on autopilot.
July 13th