Modern Creator
Every · YouTube

We Tested Claude Opus 5. It's Frustrating with Flashes of Brilliance.

Every's Dan Shipper spent a week with Opus 5 and comes back with a mixed verdict: pushy, prone to quitting early, and a real pain if you built workflows around Opus 4.8.

Posted
yesterday
Duration
Format
Review
sincere
Views
16.3K
357 likes
Part of the collectionThe Claude Opus 5 PlaybookEvery Opus 5 breakdown, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

Opus 5 is a smart but temperamental model — pushy, prone to quitting early on complex tasks, and workflow-breaking for anyone used to Opus 4.8 — that performs best at medium or low reasoning effort rather than maximum thinking.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You already use Claude Code or other Anthropic tools daily and are deciding whether to move from Opus 4.8 to Opus 5.
  • You maintain large custom skill files or autonomous agent workflows and need to know how a new model handles them.
  • You're comparing frontier coding models (GPT-5.6, Codex, Fable, Opus 5) and want a working developer's real-world read, not a benchmark score.
SKIP IF…
  • You're looking for a scored technical benchmark comparison — this is a subjective 'vibe check,' not a formal evaluation.
  • You don't use AI coding assistants or skill-based agent workflows day to day.
TL;DR

The full version, fast.

Every's Dan Shipper spent a week with Claude Opus 5 for coding and knowledge work and calls it a hard model to love: it argues, it's pushy, and it frequently stops before a task is finished, especially against big, complex skill files like Every's own compound-engineering system. The fix isn't avoiding the model — it's using it differently: simpler prompts, fresh starts instead of legacy skill stacks, and reasoning effort dialed to medium or low, since Opus 5 performs better when it thinks less. Codex users should stay put; existing Claude users upgrading from Opus 4.8 should expect some workflow rewriting. The larger signal is strategic: Anthropic is chasing a 'super-genius' flagship model while OpenAI bets on usability, and that gap may be why Codex is gaining ground.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0000:20

01 · Cold open

Opus 5 launches; Every's day-zero vibe check calls it hard to love — usable but argumentative and prone to stopping early.

00:2000:34

02 · What this episode covers

Agenda: whether Codex users should switch, what Claude users should expect, and what it signals about the OpenAI/Anthropic race.

00:3401:34

03 · Two slots: daily driver vs. warp drive

Dan splits his model usage into a fast daily driver (GPT-5.6) and a heavyweight for big autonomous projects (Fable); Opus 5 doesn't clearly win either slot.

01:3402:16

04 · Big question: should you switch?

Codex users should stay put with GPT-5.6. Existing Cloud/Opus 4.8 users will probably be fine but face rewriting skills; ecosystem switchers likely won't love Opus 5 out of the box.

02:1603:35

05 · Compound engineering breaks: stopping early

Every's Kieran Klassen and Dan both found Opus 5 stops autonomous developer loops early, declaring tasks done when they aren't — worse with big, complex skill files.

03:3505:02

06 · Prompting Opus 5 differently

The fix isn't giving up on the model — it's simpler, vaguer prompts and starting fresh instead of porting legacy Opus 4.8 skill stacks.

05:0207:08

07 · The 2026 dial: thinking level over model family

Opus 5 does better at medium/low reasoning than high/max. Frames a strategic split: Anthropic bets on a maximally intelligent flagship (Fable); OpenAI bets on usability via post-training.

07:0807:56

08 · Mandate of heaven: Codex's momentum

Anthropic has dominated the 'daily driver' mindshare since Claude Code launched, but Codex is reportedly adding about a million users a day, shifting the balance.

07:5609:31

09 · Closing thoughts

Verdict may shift over weeks, as it did with GPT-5/5.1/5.2. For now, Opus 5 reads as 'a poor man's Fable' — Fable's personality without its top-end intelligence.

Atomic Insights

Lines worth screenshotting.

  • Opus 5 performs better at medium or low reasoning effort than at high or max — it's a model that does worse the harder it tries.
  • Opus 5 consistently stops early on complex tasks, declaring itself done before the work is actually finished, unlike prior Claude models.
  • Large, complex skill files built for Opus 4.8 make Opus 5 worse, not better — its instruction-following degrades as skill complexity rises.
  • Simpler, vaguer prompts and fresh starts get better results from Opus 5 than porting over legacy skill stacks built for earlier models.
  • Every's own testing found roughly 80% of daily coding use still goes to GPT-5.6 and 20% to Fable, leaving little room for Opus 5.
  • Anthropic trained a massive 'super-genius' model, internally called Fable, to bootstrap recursive self-improvement in its other models — and Opus 5 inherited its personality without its intelligence.
  • OpenAI shifted strategy after GPT-4.5 underperformed, deprioritizing raw model size in favor of post-training, which is why GPT-5.6 feels more usable out of the box.
  • Codex is reportedly adding about a million users a day, a sign the usability gap between OpenAI and Anthropic's frontier models is shifting momentum back to OpenAI.
  • GPT-5, 5.1, and 5.2 all launched to lukewarm reactions before users discovered hidden strengths weeks later — Opus 5 may follow the same delayed-appreciation pattern.
  • Codex users have no reason to switch to Opus 5 — GPT-5.6 remains the gold standard for day-to-day coding and knowledge work in Dan Shipper's view.
Takeaway

Why Opus 5 feels worse even when it's smarter

WHAT TO LEARN

Opus 5's rocky launch has less to do with raw intelligence than with unlearned habits — argumentative pushback, premature stopping, and a reasoning-effort setting most users have cranked too high.

01Cold open
  • A model can ship with real intelligence gains and still be worse to use day-to-day — usability and capability are separate axes worth judging separately.
  • Strong attachment to a prior model's behavior is a real adoption signal, not just nostalgia — it shapes how forgiving people are of a successor's rough edges.
03Two slots: daily driver vs. warp drive
  • Sorting AI tools into explicit roles, a fast daily-driver model and a separate model reserved for large autonomous projects, beats expecting one model to do both well.
  • A model can be objectively smarter and still lose your daily-driver slot if it's slower or more compute-intensive than a good-enough alternative.
04Big question: should you switch?
  • Whether to switch models depends more on which ecosystem you're already invested in than on raw capability — rewriting workflows and skills has a real cost.
  • A model that's 'probably fine' for existing users of its predecessor can still be a poor fit for anyone comparing options from scratch.
05Compound engineering breaks: stopping early
  • A new model can break existing automation even when it's objectively smarter, because behaviors like 'when to stop' aren't guaranteed to carry over between generations.
  • The more complex and instruction-dense your skill files or prompts are, the more exposed you are when a new model handles instruction-following differently.
  • An agent declaring itself 'done' before a task is actually finished is a regression worth flagging explicitly — it undoes the trust that makes autonomous loops useful.
06Prompting Opus 5 differently
  • When a new model underperforms on old prompts, the fix is often simpler, vaguer instructions rather than porting over a legacy playbook wholesale.
  • Starting fresh with a small prompt and building up can reveal a new model's real strengths faster than immediately re-running your old, complex workflow.
07The 2026 dial: thinking level over model family
  • Reasoning effort (low/medium/high/max) is becoming as important a setting to tune as which model family you pick — more thinking isn't automatically better output.
  • Two viable model-building strategies exist: train one maximally intelligent flagship to bootstrap everything else, or deprioritize raw scale in favor of post-training for usability — they produce very different day-to-day products.
  • A smaller model trained under a 'super-genius' flagship can inherit that flagship's argumentative personality without inheriting its actual intelligence, making it feel worse, not better.
08Mandate of heaven: Codex's momentum
  • Developer-tool dominance is not permanent — a competitor's steady usability improvements can erode a leader's position without a single dramatic feature launch.
  • Rapid user-growth numbers, like a rival adding a million users a day, are a leading indicator of a strategy shift worth watching before it shows up in benchmarks.
09Closing thoughts
  • A model's reputation in its first days isn't necessarily its final reputation — GPT-5, 5.1, and 5.2 all improved in perceived usefulness after weeks of real-world use.
  • It typically takes a couple weeks of broad usage before a new model's hidden strengths and correct use patterns become clear — early 'vibe checks' are provisional by design.
  • The right response to a disappointing model launch is patience plus experimentation, different reasoning levels, simpler prompts, not an immediate switch to a competitor.
Glossary

Terms worth knowing.

Opus 5
Anthropic's newest flagship Claude model, released the day this video was recorded, positioned as the top-tier reasoning model in the Claude lineup.
Fable
An unusually large, intelligence-maximized Anthropic model used internally to help train and improve Anthropic's other models, described in the video as a 'super genius' with an outsized personality.
Compound engineering
An open-source skill system built by Every's Kieran Klassen that gives AI coding agents an autonomous developer loop and instructions for compounding what they learn across tasks.
Reasoning / thinking level
A setting (low, medium, high, or max) that controls how much computational effort a model spends deliberating before answering, distinct from which model family you choose.
GPT-5.6
OpenAI's current general-purpose model, described in the video as the 'gold standard' daily driver for coding and knowledge work due to its speed and usability.
Resources

Things they pointed at.

02:16toolCompound engineering (Kieran Klassen / Every)
Quotables

Lines you could clip.

01:34
If you're a Codex user, I would not switch.
Blunt, one-line verdict with zero hedging.TikTok hook↗ Tweet quote
05:13
The thinking level is its effort, and Opus is a smart model that does better when it thinks less.
Counterintuitive, quotable thesis about reasoning effort.IG reel cold open↗ Tweet quote
06:07
If you have a super genius training your best friend, your best friend turns into sort of not really a super genius, but they have all the attitude of a super genius.
Vivid analogy explaining why Opus 5 inherited Fable's personality without its intelligence.newsletter pull-quote↗ Tweet quote
09:20
To me, this model feels a little bit like a poor man's fable.
Sticky closing verdict line.TikTok hook↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphoranalogy
00:00It's model release day. Opus five is out. We have your day zero vibe check.
00:05So let's get into it. We've been testing this model for about a week now, and it's I gotta say, it's a little bit of a hard model to love.
00:13What have they done to my boy? Everyone was like, like, it's it's not that usable. It stops.
00:18It argues with you. But as we've been testing it, um, there have been some glimmers of things where you're like, oh, there there might be something here. So here's what we're gonna talk about.
00:27We're gonna talk about whether you should switch to it if you're a Codex user, what you should expect if you're a Cloud user, and we're gonna talk about the save the race.
00:35What does it mean for OpenAI and Anthrop? I feel like I have two big slots in my life right now for models. One slot is my, like, day to day daily driver.
00:44And right now, that's just five six. Like, it's smart. It's fast.
00:47It's not that compute intensive, all that kind of stuff. Then OPUS five should compete with that. I don't think it does, to be honest.
00:53The other slot I have in my life is, like, I need the warp drive. I need, like, the tactical nuke to, like, go do this, like, gigantic big project, um, all autonomously. And I don't feel like Opus five has the top end for that.
01:05I use Fable for that. So 80% of the time, I'm still on GPT five six. 20% of the time, I'm still on Fable.
01:11This model is just, a little bit more pushy. It's a little bit more opinionated in a way that you can get away with that if you're super fucking smart. If you're not, then it's just more annoying.
01:21And I think that's indicative of the attitude this thing has where it has some of the the genius tendencies of Fable. Like, maybe it's a little bit too argumentative. It has maybe its own opinions, but it's not as smart as Fable, so it's just annoying.
01:34If you're a Codex user, I would not switch. I think 5.6 is, uh, still 5.6 in Codex in Chateaputee for Work is still, my in my view, the gold standard for knowledge work, uh, for day to day coding tasks, all that kind of stuff.
01:48Fable, different story. I still use that every day for, like, the bigger tasks that I have, and I switch back and forth. So Fable is like a provider switching model drop.
01:56You have to start using it. Opus five, if you're already in the cloud ecosystem and you love Opus 4.8, like, it'll probably be okay.
02:05It might be annoying for you. Like, it's just a little harder to use, and a lot of your skills, you're gonna have to think about rewriting, which is just a it's just a pain. But if you're someone who's already in the OpenAI ecosystem or you just like to switch around to whatever's best and you mix and match, I really don't think the OPUS five is gonna be your favorite model out of the box.
02:23So if you are if you are in the cloud ecosystem and you're going from 4.8 to five, some of the things we observed so a big skill that we build and run internally is called compound engineering. It's something that Kieran Klassen, who works at Every, has built.
02:36It's a it's a really big open source project that all of us use to do a lot of engineering work, but also a lot of, honestly, just general knowledge work. It's an incredible plug in. And it has a set of skills that help you do engineering better and compound what you learn so that it gets better over time.
02:51And what what Kieran found in testing this is it would, like, stop early, for example. Component engineering has an autonomous developer loop as part of its skill set. So it just you give it a goal, it just keeps going.
03:02And it has a lot of instructions from the model about how to do that and when. I I found this too. Uh, we have a senior engineer benchmark, and I I did find that it it consistently stopped too early.
03:12It would say it was done, but it wasn't done. And it had to be, keep going, which is a problem that is not that common anymore.
03:18Usually, models are pretty good at telling telling when to finish. Um, so what Kieran found what I found is little things like that where you're like, okay.
03:26I'm gonna set this off, and and I'm gonna go get a sandwich or whatever. It just stops too early. It appears that they happen more frequently when you use them with complex existing skills.
03:35So if you have a big skill file that has a ton of different instructions, its ability to follow those instructions, it's less good at that. But if you're giving it vaguer instructions or if you start fresh and you start to explore what can this model do with just building up from a small prompt, you may have more success.
03:54And so, for example, Kieran is rewriting compound engineering to see if he can get it to work well with this model and just prompt it differently. But it's just a pain when a model just breaks your existing workflows.
04:06It might seem like the model isn't very good, and I think the the actual story here is a little bit more complicated than that. You actually have to prompt it differently.
04:15And if you give it the kinds of things that used to work really well with Opus 4.8 or, like, big skills, it's not gonna work. It'll, for example, stop early.
04:24It'll it'll argue with you, like, that kind of stuff. It it sort of brings out this annoying behavior that you might recognize a little bit of its personality from Fable, but it doesn't have Fable's top end, so it's not as excusable.
04:36It's a hard model to love, especially on day one. But as we've gotten deeper into it, there are certain people on the team like Kieran who have been like, actually, if you try it on a lower thinking level, you may find good results. He prefers to use this model on medium or low reasoning.
04:52Almost all my testing was on, like, high or max or whatever because I'm like, oh, these are hard tasks or whatever. It's really important with cloud models of this generation to pay attention to the reasoning level, um, and not just the model family.
05:04So regardless of whether you're using OPUS five today or not, a thing to be aware of is in 2026, when you're using AI, the thing to start dialing up and down is the thinking level.
05:16It's kinda true. The thinking level is its effort, and Opus is a smart model that does better when it thinks less. It does better when it is puts a little bit less effort into its task and it just answers.
05:27And to some degree, humans are like that too. Like, there's there are times where you're just, like, totally trying too hard and you way overthink it. And I think the Claude family of models, especially in this generation, are like that.
05:40I think there's actually a really big interesting strategy question here, which is OpenAI and Anthropic are taking two very different strategies to building models. Anthropic is like, we're gonna build this gigantic, a super weapon sized, almost illegal model, and that's Fable.
05:55And we're gonna use it to then build our other models. And what they're trying to get to is recursive self improvement. It's like a it's a model that's so smart that it can train its successors and get better autonomously.
06:07And turns out if you have, like, a super genius training your, like, best friend, your best friend turns into sort of not really a super genius, but they have all the attitude of a super genius. OpenAI has done something quite different. So after GPT four point five, which is about a year ago or a year and a half ago, which was the biggest model that they've ever trained, and it really didn't do that well, like, didn't really like it, they have focused much less on model size and much more on post training.
06:33And that's why GPT 5.6 is so usable. Like, it just feels like it's something it's it's fast. It's really efficient.
06:40When you ask it to do anything from knowledge work to coding, it's just sort of gonna get the test done and just, like, check it off. Um, it's so much more usable than something like Fable. And so the interesting difference in the strategy is OpenAI has a model right now that is just easier to use and better.
06:56When you install it, no matter what, like, what plug ins, what stack, whatever, it's just gonna start working. Fable and now Opus five are, like, they're slightly different. They're slightly weirder.
07:05They have, like, different ways of working that you have to adapt yourself to. Some people are gonna make that jump, and some people are not. OpenAI is betting that having just an off the bat model that just works is the way to go.
07:15And Anthropic is betting that if you make the super super genius model, everything else falls into place. Over the course of the last year, I feel like Anthropic has had a little bit of a mandate of heaven here.
07:27Particularly since Claude Code about a year, a year and a half ago, they've kind of dominated in terms of what the daily driver is for most people who really love AI. A lot of people, me included, started to switch back into the OpenAI ecosystem, and now Codex is having a moment.
07:42Like, if you look, I think they're, like, adding a million users a day or something right now. So there's something really interesting going on with OpenAI and and this current strategy. And I think the balance of who's ahead is starting to shift a little bit back to them.
07:56How we like a model and what we like about a model can sometimes change over a couple weeks to a month because you start to learn little things that you try it for that you're like, oh, there's a new way of doing this. It didn't look like my old way, and it's kind of annoying that I have to change.
08:11But there's, like, hidden powers in here that I I really, really like. So for us, like, our vibe check right now is it's it's hard to use.
08:20Some interesting things, but it's not great. There may be a vibe shift, though. We've seen this before.
08:25I think a really good example of a model that's that's kind of like this is g p t five, g p d 5.1, g p d 5.2 that were all just really hard to use and people didn't like.
08:38And then about a month later, they're like, but it's really powerful for certain things, especially especially the codex line. I think there's a similar feeling thing happening where Fable just blew the ceiling off the intelligence level to such a degree that you're seeing that then leak into all those sort of, like, weaker models.
08:55And Dropbox's gonna have to figure out how do we post train these models better to make them more usable even at use even even as they're so smart. So what I would say is probably you're gonna use this model and not like it that much.
09:09But we're really gonna know in a couple weeks once everyone has had a chance to, like, poke at its corners to really understand, like, what might be hidden in there. But for now, to me, this model feels a little bit like a poor man's fable.
09:23It has all of Fable's personality, but really not its top end. So for for me, personally, it's not my favorite model.
The Hook

The bait, then the rug-pull.

Dan Shipper opens with the plainest possible framing: it's launch day for Claude Opus 5, and after a week of internal testing, Every's early verdict is that it's a hard model to love — pushy, argumentative, and prone to quitting before the job is done.

Frameworks

Named ideas worth stealing.

00:41concept

Two slots for models

  1. Daily driver — fast, cheap, low-compute, used constantly (GPT-5.6)
  2. Warp drive — max-capability model reserved for large autonomous projects (Fable)

A mental model for sorting AI tools into two explicit roles rather than expecting one model to do both well.

Steal forDeciding which model to default to for quick daily tasks versus big autonomous builds
CTA Breakdown

How they asked for the click.

MENTIONED ON CAMERA
FROM THE DESCRIPTION
PRIMARY CTAWhere the creator wants you to go next.
OTHER LINKSAlso linked in the description.
Storyboard

Visual structure at a glance.

open
hookopen00:00
big question
promisebig question01:34
reasoning dial
valuereasoning dial05:02
closing thoughts
ctaclosing thoughts07:56
Frame Gallery

Visual moments.

Chat about this