The argument in one line.
Opus 5 is a smart but temperamental model — pushy, prone to quitting early on complex tasks, and workflow-breaking for anyone used to Opus 4.8 — that performs best at medium or low reasoning effort rather than maximum thinking.
Read if. Skip if.
- You already use Claude Code or other Anthropic tools daily and are deciding whether to move from Opus 4.8 to Opus 5.
- You maintain large custom skill files or autonomous agent workflows and need to know how a new model handles them.
- You're comparing frontier coding models (GPT-5.6, Codex, Fable, Opus 5) and want a working developer's real-world read, not a benchmark score.
- You're looking for a scored technical benchmark comparison — this is a subjective 'vibe check,' not a formal evaluation.
- You don't use AI coding assistants or skill-based agent workflows day to day.
The full version, fast.
Every's Dan Shipper spent a week with Claude Opus 5 for coding and knowledge work and calls it a hard model to love: it argues, it's pushy, and it frequently stops before a task is finished, especially against big, complex skill files like Every's own compound-engineering system. The fix isn't avoiding the model — it's using it differently: simpler prompts, fresh starts instead of legacy skill stacks, and reasoning effort dialed to medium or low, since Opus 5 performs better when it thinks less. Codex users should stay put; existing Claude users upgrading from Opus 4.8 should expect some workflow rewriting. The larger signal is strategic: Anthropic is chasing a 'super-genius' flagship model while OpenAI bets on usability, and that gap may be why Codex is gaining ground.
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →Where the time goes.

01 · Cold open
Opus 5 launches; Every's day-zero vibe check calls it hard to love — usable but argumentative and prone to stopping early.

02 · What this episode covers
Agenda: whether Codex users should switch, what Claude users should expect, and what it signals about the OpenAI/Anthropic race.

03 · Two slots: daily driver vs. warp drive
Dan splits his model usage into a fast daily driver (GPT-5.6) and a heavyweight for big autonomous projects (Fable); Opus 5 doesn't clearly win either slot.

04 · Big question: should you switch?
Codex users should stay put with GPT-5.6. Existing Cloud/Opus 4.8 users will probably be fine but face rewriting skills; ecosystem switchers likely won't love Opus 5 out of the box.

05 · Compound engineering breaks: stopping early
Every's Kieran Klassen and Dan both found Opus 5 stops autonomous developer loops early, declaring tasks done when they aren't — worse with big, complex skill files.

06 · Prompting Opus 5 differently
The fix isn't giving up on the model — it's simpler, vaguer prompts and starting fresh instead of porting legacy Opus 4.8 skill stacks.

07 · The 2026 dial: thinking level over model family
Opus 5 does better at medium/low reasoning than high/max. Frames a strategic split: Anthropic bets on a maximally intelligent flagship (Fable); OpenAI bets on usability via post-training.

08 · Mandate of heaven: Codex's momentum
Anthropic has dominated the 'daily driver' mindshare since Claude Code launched, but Codex is reportedly adding about a million users a day, shifting the balance.

09 · Closing thoughts
Verdict may shift over weeks, as it did with GPT-5/5.1/5.2. For now, Opus 5 reads as 'a poor man's Fable' — Fable's personality without its top-end intelligence.
Lines worth screenshotting.
- Opus 5 performs better at medium or low reasoning effort than at high or max — it's a model that does worse the harder it tries.
- Opus 5 consistently stops early on complex tasks, declaring itself done before the work is actually finished, unlike prior Claude models.
- Large, complex skill files built for Opus 4.8 make Opus 5 worse, not better — its instruction-following degrades as skill complexity rises.
- Simpler, vaguer prompts and fresh starts get better results from Opus 5 than porting over legacy skill stacks built for earlier models.
- Every's own testing found roughly 80% of daily coding use still goes to GPT-5.6 and 20% to Fable, leaving little room for Opus 5.
- Anthropic trained a massive 'super-genius' model, internally called Fable, to bootstrap recursive self-improvement in its other models — and Opus 5 inherited its personality without its intelligence.
- OpenAI shifted strategy after GPT-4.5 underperformed, deprioritizing raw model size in favor of post-training, which is why GPT-5.6 feels more usable out of the box.
- Codex is reportedly adding about a million users a day, a sign the usability gap between OpenAI and Anthropic's frontier models is shifting momentum back to OpenAI.
- GPT-5, 5.1, and 5.2 all launched to lukewarm reactions before users discovered hidden strengths weeks later — Opus 5 may follow the same delayed-appreciation pattern.
- Codex users have no reason to switch to Opus 5 — GPT-5.6 remains the gold standard for day-to-day coding and knowledge work in Dan Shipper's view.
Why Opus 5 feels worse even when it's smarter
Opus 5's rocky launch has less to do with raw intelligence than with unlearned habits — argumentative pushback, premature stopping, and a reasoning-effort setting most users have cranked too high.
- A model can ship with real intelligence gains and still be worse to use day-to-day — usability and capability are separate axes worth judging separately.
- Strong attachment to a prior model's behavior is a real adoption signal, not just nostalgia — it shapes how forgiving people are of a successor's rough edges.
- Sorting AI tools into explicit roles, a fast daily-driver model and a separate model reserved for large autonomous projects, beats expecting one model to do both well.
- A model can be objectively smarter and still lose your daily-driver slot if it's slower or more compute-intensive than a good-enough alternative.
- Whether to switch models depends more on which ecosystem you're already invested in than on raw capability — rewriting workflows and skills has a real cost.
- A model that's 'probably fine' for existing users of its predecessor can still be a poor fit for anyone comparing options from scratch.
- A new model can break existing automation even when it's objectively smarter, because behaviors like 'when to stop' aren't guaranteed to carry over between generations.
- The more complex and instruction-dense your skill files or prompts are, the more exposed you are when a new model handles instruction-following differently.
- An agent declaring itself 'done' before a task is actually finished is a regression worth flagging explicitly — it undoes the trust that makes autonomous loops useful.
- When a new model underperforms on old prompts, the fix is often simpler, vaguer instructions rather than porting over a legacy playbook wholesale.
- Starting fresh with a small prompt and building up can reveal a new model's real strengths faster than immediately re-running your old, complex workflow.
- Reasoning effort (low/medium/high/max) is becoming as important a setting to tune as which model family you pick — more thinking isn't automatically better output.
- Two viable model-building strategies exist: train one maximally intelligent flagship to bootstrap everything else, or deprioritize raw scale in favor of post-training for usability — they produce very different day-to-day products.
- A smaller model trained under a 'super-genius' flagship can inherit that flagship's argumentative personality without inheriting its actual intelligence, making it feel worse, not better.
- Developer-tool dominance is not permanent — a competitor's steady usability improvements can erode a leader's position without a single dramatic feature launch.
- Rapid user-growth numbers, like a rival adding a million users a day, are a leading indicator of a strategy shift worth watching before it shows up in benchmarks.
- A model's reputation in its first days isn't necessarily its final reputation — GPT-5, 5.1, and 5.2 all improved in perceived usefulness after weeks of real-world use.
- It typically takes a couple weeks of broad usage before a new model's hidden strengths and correct use patterns become clear — early 'vibe checks' are provisional by design.
- The right response to a disappointing model launch is patience plus experimentation, different reasoning levels, simpler prompts, not an immediate switch to a competitor.
Terms worth knowing.
- Opus 5
- Anthropic's newest flagship Claude model, released the day this video was recorded, positioned as the top-tier reasoning model in the Claude lineup.
- Fable
- An unusually large, intelligence-maximized Anthropic model used internally to help train and improve Anthropic's other models, described in the video as a 'super genius' with an outsized personality.
- Compound engineering
- An open-source skill system built by Every's Kieran Klassen that gives AI coding agents an autonomous developer loop and instructions for compounding what they learn across tasks.
- Reasoning / thinking level
- A setting (low, medium, high, or max) that controls how much computational effort a model spends deliberating before answering, distinct from which model family you choose.
- GPT-5.6
- OpenAI's current general-purpose model, described in the video as the 'gold standard' daily driver for coding and knowledge work due to its speed and usability.
Things they pointed at.
Lines you could clip.
“If you're a Codex user, I would not switch.”
“The thinking level is its effort, and Opus is a smart model that does better when it thinks less.”
“If you have a super genius training your best friend, your best friend turns into sort of not really a super genius, but they have all the attitude of a super genius.”
“To me, this model feels a little bit like a poor man's fable.”
Word for word.
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
The bait, then the rug-pull.
Dan Shipper opens with the plainest possible framing: it's launch day for Claude Opus 5, and after a week of internal testing, Every's early verdict is that it's a hard model to love — pushy, argumentative, and prone to quitting before the job is done.
Named ideas worth stealing.
Two slots for models
- Daily driver — fast, cheap, low-compute, used constantly (GPT-5.6)
- Warp drive — max-capability model reserved for large autonomous projects (Fable)
A mental model for sorting AI tools into two explicit roles rather than expecting one model to do both well.







































































