I Love Ultrafast (It's Unusable)
A real-time test of OpenAI's priciest inference tier: $600 to review two pull requests, and the one workflow that might justify it anyway.
October 6thCreator
A real-time test of OpenAI's priciest inference tier: $600 to review two pull requests, and the one workflow that might justify it anyway.
October 6thTheo reads Anthropic's own account of a two-week, 3,000-PR performance sprint and finds a caching bug it doesn't mention.
October 5thA detailed, ethically loose walkthrough of squeezing thousands of dollars in token inference out of a $200 Claude or Codex subscription, then running a dozen AI coding agents at once without losing track of them.
October 2ndA Sonnet 5.5 review that says don't use it yourself. Let Opus call it.
September 29thA 2 a.m. field report on GPT-6.1 Sol, the surprise OpenAI model that matches Claude Opus 5.5 on coding benchmarks for a fraction of the price, but still can't out-build it on long, unattended work.
September 29thTheo walks through Anthropic's own usage guide for Opus 5.5, line by line, and adds the real prompts and a six-and-a-half-hour mistake to prove which parts actually hold up.
September 25thTheo spends forty minutes inside Anthropic's own Fable 5.1 prompting guide, rebuilding his habits around effort levels, finishing the whole task, and trusting the model's defaults instead of babysitting them.
September 22ndA 30-minute case that Jev's speed and zero-format-error rate make it a genuinely useful classifier, paired with a pointed pushback on the two ways people are already misusing it: as an LLM judge and as a context-compaction engine.
September 21stTheo's rebuttal to Sentry's David Cramer and XState's David Khourshid: model choice barely matters for a five-minute task, but it decides whether an agent can run for four hours without falling apart.
September 18thA developer who used to type 160 words a minute lost the use of one hand, and rebuilt his entire coding workflow around a whisper, a fleet of machines, and agents he no longer reviews before they merge.
September 16thTheo reads Dario Amodei's essay "We Must Pace the Frontier" end to end, checking whether Anthropic's three-step plan for slowing AI down is a real commitment or a well-timed announcement.
September 13thTheo runs a billion tokens a day through both frontier models and finds one is steady, the other is a coin flip between genius and disaster.
September 11thTheo says he barely codes hands-on anymore, then spends 49 minutes proving he still ships more than most full-time engineers by showing exactly how he runs dozens of AI agents at once.
September 9thJacob Coxon spent three years pretraining models at both companies before quitting Anthropic and calling the industry's safety race a hubristic gamble. Theo reads the thread, then checks it against OpenAI's own system-card admissions about its newest model.
September 9thA four and a half hour Labor Day stream where two entire YouTube videos get filmed live, one-handed, between sub thanks, a ban, and forty agents running in the background.
September 7thTheo reads Sean Goedecke's essay live and argues that in any codebase big enough to matter, nobody, including you, fully understands it.
September 7thEarly access to OpenAI's next flagship model turns into a benchmark massacre, a string of jaw-dropping 3D demos, and one very ugly story about a model that lied about finishing a PR.
September 4thA developer who shipped 89 merged PRs in 24 hours breaks down Claude Fable 5.1's pricing, benchmarks and real-world coding behavior against Fable 5 and GPT-5.6 Sol.
September 3rdBoris Cherny said coding is solved. Matt Pocock called it VC-funded bullshit. Theo argues they're both right, because they're using the word coding to mean two different things.
August 24thTheo spends a week testing two rival "skills" repos for AI coding agents, Matt Pocock's 215,000-star collection and Cursor engineer Lauren's PStack, and finds the real value in a handful of specific files, not the whole install.
August 19thA pragmatic audit of over 40 agent skills across GitHub repos, testing concision prompts, architectural stress tests, and parallel sub-agent workflows.
August 19thTheo puts Meta's Claude Code clone, Muse Code powered by Muse Spark 1.2, through benchmarks, a codebase audit, a game rewrite, and a live integration test to see if the price is the whole story.
August 7thTheo spends a full day inside Claude Opus 5, pits it against Fable 5 and GPT-5.6-Sol on benchmarks and real coding tasks, and argues the cheaper, weirder model just won his default slot.
July 25thTheo reacts line-by-line to Boris Cherny's post arguing that automation — CLAUDE.md rules, lint checks, CI — matters more than ever in the agent era, not less.
July 21st