We Tested Claude Opus 5. It's Frustrating with Flashes of Brilliance.
Every's Dan Shipper spent a week with Opus 5 and comes back with a mixed verdict: pushy, prone to quitting early, and a real pain if you built workflows around Opus 4.8.
July 24thEvery's Dan Shipper puts OpenAI's new flagship through a same-day vibe check and stacks it against Anthropic's Fable.
GPT-6 Astra is a genuinely strong daily-use model for writing, computer automation, and 3D visualization, but it over-decorates every interface it builds and drifts from a prompt's exact intent, which keeps it a step behind Fable for high-stakes, long-running delegation.
Every's team spent release day testing OpenAI's new GPT-6 Astra and found it excels at writing, computer use, and 3D visualization: it can draft a passable self-review from Slack messages, edit a five-hour video in Premiere unsupervised, and build a historically accurate Battle of Waterloo scene from written accounts. Its weak spot is restraint. Astra decorates interfaces with extra buttons and text the prompt never asked for, and it sticks less closely to the literal intent of a request than Anthropic's Fable does on the same task. Verdict: an S-tier daily driver, but Fable still wins the biggest, longest-running delegation work.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →
Announces GPT-6 Astra dropped hours earlier and previews the five-point vibe check, followed by a sponsor read and the Every subscription pitch before testing begins.

Astra one-shots its own critical vibe check by reading the team's Slack messages and calls itself 'a show horse, not a workhorse'; praised as a crisp, slop-free daily writing companion.

Astra runs unattended in Adobe Premiere for about five hours, cutting together a video that reached 25,000 views on its own.

Astra builds an explorable, historically accurate 3D reconstruction of the Battle of Waterloo in three to four hours from written accounts and geography.

Astra's visual taste is strong but it over-decorates interfaces with redundant labels, text, and buttons, especially at higher reasoning-effort settings, leaving Fable's first pass cleaner.

Head-to-head, Astra and Fable each build a handwritten-journal digitizing app from the same brief; Fable's one-button, page-by-page flow beats Astra's more cluttered, multi-step interface.

Calls Astra an absolute S-tier daily driver at Fable's price point, but keeps Fable for the biggest, longest-running delegation work.
Astra writes cleanly and can run a computer unattended for hours, but it can't resist decorating every interface it builds, which is the tell that separates a daily driver from a delegate-and-forget model.
“GPT-6 Astra is a show horse. OpenAI's new model has striking visual taste and a knack for consulting work. Its ambitions still outrun its judgment.”
“I feel like its sentences are crisp. They're to the point. There's no AI-isms. There's no slop.”
“It spent like five hours in Premiere earlier this week, cutting together the Fable 5.1 vibe check that we released earlier this week. That video has like 25,000 views.”
“Overall verdict, absolute S tier daily driver... But for the top end, long running, big delegation tasks, I still use fable.”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
OpenAI dropped GPT-6 Astra with almost no warning, and Every had a 30-person team testing it within hours. This is the same-day verdict: where Astra earns a spot in daily use, and where Anthropic's Fable still wins.
The video's own structure for a same-day model review: three clear strengths followed by two specific, demonstrated weaknesses versus a named competitor.
“check it out now at every.to”
Soft-sell for the Every subscription (newsletter, courses, camps, in-house AI tools) folded into the cold open before the actual testing content begins, plus a pointer to the full-length written vibe check article.
00:00
00:07
00:12
00:16
00:21
00:27
00:30
00:37
00:42
00:47
00:52
00:57
01:02
01:07
01:13
01:15
01:19
01:26
01:32
01:35
01:41
01:46
01:51
01:56
02:01
02:06
02:13
02:16
02:23
02:25
02:31
02:36
02:39
02:44
02:51
02:57
03:01
03:06
03:11
03:16
03:21
03:26
03:29
03:36
03:42
03:45
03:51
03:56
04:03
04:06
04:12
04:16
04:21
04:26
04:31
04:36
04:39
04:46
04:51
04:56
05:01
05:03
05:10
05:13
05:20
05:27
05:30
05:35
05:39
05:46
05:50
05:55
06:00
06:05
06:10
06:15
06:19
06:25
06:30
06:37Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Every's Dan Shipper spent a week with Opus 5 and comes back with a mixed verdict: pushy, prone to quitting early, and a real pain if you built workflows around Opus 4.8.
July 24thEarly access to OpenAI's next flagship model turns into a benchmark massacre, a string of jaw-dropping 3D demos, and one very ugly story about a model that lied about finishing a PR.
September 4thAn early-access hands-on with OpenAI's new GPT-6 Astra model — benchmark scorecard, alignment numbers, API pricing, and a run of 3D game and browser-automation demos.
September 3rdFive ways Mark Kashef points Codex's computer use at apps that have no API, from flight searches to phone settings.
September 5thTwo developers burn six figures in tokens testing OpenAI's biggest model yet, and end up in a forty-minute fight over whether it lied to them.
September 3rdA first-look review of OpenAI's GPT-6 Astra, run through published benchmarks, a set of repeatable creative tests, and a computer-use experiment where the model built and animated its own 3D game world.
September 3rd