I Tested Fable 5.1 vs Fable 5 vs Opus 5: Cost, Speed, and Design
Twelve identical builds against one prompt, scored blind, to find out whether the newest model's cheaper caching actually lowers what you pay.
September 3rdThe Opus 5 Playbook · Updated July 28, 2026
Not another summary. We mined 24 creator breakdowns for the parts that matter: how it matches Fable 5 at half the price, which effort to use, what to build, and where the experts disagree — every claim linked to the source.
The 60-second version
People are asking
Straight answers, pulled from across all 24 breakdowns.
Claude Opus 5 is Anthropic’s new default Opus model, launched July 24, 2026 and included in every paid Claude plan (Pro, Max, Team, Enterprise) with no extra usage cap. It’s positioned as near-Fable-5 capability at roughly half the price, and it’s now the default model on Claude Max and the strongest model available to Claude Pro. Anthropic pitches it as state-of-the-art on its own coding evals while costing the same per token as the older Opus 4.8.
It’s close, and cost is the tiebreaker. Anthropic’s own charts put Opus 5 at or near Fable 5 on most benchmarks while costing about half as much per task. Creators split the live tests 1-1 to 2-1 on raw output quality — Naman and Alex Finn preferred Opus 5’s visual builds, while Fable 5 still led on a denser flight-sim city and on multidisciplinary reasoning, legal, and health tasks. The consensus: comparable power, but Opus 5 wins on price.
Same price ($5/$25 per million input/output tokens), much bigger model. Opus 5 more than doubles Opus 4.8 on Frontier-Bench software-engineering tasks (43.3% vs 21.1% in Nick Saraev’s reading), and jumps from 1.5% to 30.2% on ARC-AGI-3 novel problem-solving — a 20x leap in one generation. Chase AI calls the 4.8-to-5 jump a bigger leap than the gap between Opus 5 and Fable 5. The new behaviors are self-verification and error recovery: it checks its work against a goal and iterates instead of one-shotting.
$5 per million input tokens and $25 per million output tokens — unchanged from Opus 4.8 and about half of Fable 5’s rate. But half price on the sheet is not half price in practice: Opus 5 burns more tokens per task (about 37k to Fable’s 33k on Artificial Analysis), so Theo measured the real-world gap at closer to 20-25%, and Nate Herk’s 18 sessions found Opus ran roughly three times longer and sometimes cost MORE in total — $112 versus Fable’s $73 on one structural-simulator build. The clearer win is on subscriptions: Claude plans cap Fable 5 at 50% of your weekly limit but let Opus 5 use 100% of it, so switching roughly doubles how much work your plan covers.
Anthropic published its own prompting guide because pre-Opus-5 habits produce worse results, and the headline rule is a reversal: hand it the entire job in one prompt — goal, assets, and where the finished work should land — instead of breaking it into small sequential steps the way older Claude models rewarded. Then rein it in, because Opus 5 over-delivers: state explicitly what NOT to build or touch, cap how long the response should be, cap the size of the deliverable itself, and don’t tell it to double-check its work (it already verifies as it goes, so asking again just burns tokens). Anthropic also shipped new context-engineering guidance saying you no longer need to over-specify instructions in system prompts and CLAUDE.md the way earlier models required.
It depends on the tier. On the $20 Pro plan, Anthropic’s guidance is to default to Sonnet and switch to Opus 5 only for complex tasks or when Sonnet falls short. On $100-plus Max plans, Opus 5 can be your default for nearly all knowledge work, with Fable 5 demoted to a second opinion when you’re unhappy with what Opus produced — not an automatic upgrade. On effort: low on Pro, medium on Max, as a default that beats fiddling with the setting per task.
Medium is the sweet spot for most work; drop to low for simple tasks and only reach for high on the genuinely hard 20%. Chase AI found dropping from max to low effort cut cost per task from about $22 to $3.76 — over 80% — while low effort still hit a 60% pass rate on a long-horizon agentic benchmark. Every went further: Opus 5 actually performs better at medium or low than at high or max, so cranking the dial can cost you both money and quality.
Yes — it’s state-of-the-art on Anthropic’s agentic coding evals, leading Frontier-Bench at 43.3% and one-shotting playable game clones, a macOS clone, and 3D simulations in creator tests. Its real edge is cost-per-task and self-verification: fixing bugs at the root cause and building its own test harnesses. Caveats: Fable 5 and Kimi K3 still edge it on some front-end polish and cybersecurity work, and Every found it stops early on big, complex skill files. Codex users on GPT-5.6 saw little reason to switch.
Opus 5 vs the field, dimensionalized
Not vibes. Here is what happened in the head-to-head tests creators actually ran.
The effort dial
Opus 5’s low / medium / high effort toggle is the dial that trades cost for capability — and cranking it to the top is usually the wrong move.
⚠️ Note: effort controls how much time and how many tokens Claude spends thinking — not a guaranteed quality boost. Anthropic’s own guide now recommends low effort on Pro and medium on Max for normal knowledge work, which lines up with Every’s finding that Opus 5 can do worse the harder it tries. Match effort to the task; don’t crank it to feel safe. ↗ Paul J Lipsky ↗ Every ↗ Chase AI
Anthropic’s own five rules
Anthropic published a prompting guide specifically because pre-Opus-5 habits produce worse results. Here it is, as broken down by Paul J Lipsky. ↗ Watch
Older Claude models rewarded breaking work into small sequential chunks. Opus 5 performs better given the entire job in one prompt — the goal, the assets it can use, and where the finished work should end up, all in the first message.
Its failure mode is over-delivering. Draw the boundary explicitly: what to build, what to leave alone, and when to come back and ask before proceeding.
Ask for a fixed number of bullet points or a hard word count. Without a cap it sprawls — the same verbosity Alex Finn tames with “speak as simply as possible” in CLAUDE.md.
Response length and work-product size are two different levers. “Three bullets” doesn’t stop it building a twelve-page site — constrain the actual output too.
Opus 5 already verifies and fixes its work as it goes. Asking for another pass burns time and tokens without improving the result — the one old habit that now costs you money.
Use-case playbooks
The real methods creators demoed, with the actual tools and the video to watch.
Steal these
Real instructions from real demos. Copy, paste, tweak.
The map
The honest part you only get by watching all of them. Tap any name to watch.
“Comparable output for half of Fable 5’s price — default to Opus 5 now.”
“A top pick for reasoning and cost, but keep other models for daily driving and cost control.”
“Smart but temperamental — frustrating with flashes of brilliance, and three dealbreakers.”
“Judge it on the numbers and live head-to-heads — cost-per-task is the metric that matters.”
“Hide the model names until after you’ve ranked the outputs — brand reputation is doing more of your picking than you think.”
“The habits that worked on older Claude models actively cost you here — hand it the whole job, then rein it in.”
The 7 moves that pay
Make Opus 5 the default for coding and reasoning to capture the half-price win, but keep Fable 5 for the hardest bugs and GPT-5.6 for latency-sensitive daily work.
The sticker price hides the real bill. Run the same deliverable through both models and compare total dollars and time — Opus 5 and Fable 5 often land within pennies.
Max effort costs more and can lower quality. Start at medium, drop to low for simple work, and only reach for high on the genuinely hard 20%.
Use Opus 5 to produce the plan, then delegate execution and research to cheaper models. It’s how you stay under weekly caps without losing quality.
Opus 5 degrades on big Opus-4.8-era skill files. Simplify prompts, start fresh, and tame the personality in CLAUDE.md before blaming the model.
The per-token discount mostly evaporates once Opus 5’s heavier token use is counted. The durable win is that your Claude plan lets Opus 5 use 100% of the weekly limit where Fable 5 is capped at 50% — that roughly doubles the work one subscription covers.
Every creator who hid the model names until after judging got a cleaner answer than the ones who didn’t — and more than one was surprised by their own result. Run your real task through both, then reveal.
What’s changed
Every new Opus 5 video we break down gets added to this page automatically. Every couple of weeks we go through the new ones and update the guide itself — new answers, new head-to-heads, new things to try. Here’s what changed and when.
Anthropic published its own rules for prompting Opus 5 — and they’re the opposite of how you prompted Claude before. Nine new video breakdowns since launch week, including three that judged it blind.
First edition. This page went live the day Claude Opus 5 launched, built from the first wave of creator breakdowns.
The library
Twelve identical builds against one prompt, scored blind, to find out whether the newest model's cheaper caching actually lowers what you pay.
September 3rdTwo five-minute configuration changes, a custom output style and an on-demand skill, turn Opus 5's dense jargon and wall-of-text replies into plain, scannable answers.
August 14thA five-minute walkthrough of Claude Code's Output Styles feature — the config Anthropic's own team reaches for when responses turn into a jargon-dense wall of text.
August 5thAnthropic published an official prompting guide for Opus 5 — five rules, tested live by building a real website in under three minutes.
July 27thA walkthrough of wiring Claude Opus 5 to live SEO data through DataForSEO to rebuild Semrush and Ahrefs' core reports for the price of API credits.
July 27thA blind, three-build test of Anthropic's newest coding models — and the first time this reviewer picked a model other than Fable as the winner.
July 25thRiley Brown runs Anthropic's freshly-released Opus 5 through a real workload, then puts the new real-time Claude Voice and OpenAI's Codex voice mode side by side to show what talking to an agent can actually do.
July 25thA hands-on benchmark-and-demo breakdown of Claude Opus 5's launch, stacked against Fable 5, GPT-5.6 Sol, and Kimi K3 across reasoning scores, cost, and a dozen live-generated apps and games.
July 25thTheo spends a full day inside Claude Opus 5, pits it against Fable 5 and GPT-5.6-Sol on benchmarks and real coding tasks, and argues the cheaper, weirder model just won his default slot.
July 25thTwo operators who spend real money on inference argue that price-per-task, not leaderboard position, is now the only number that matters.
July 25thA blind, five-round test pits Opus 5 against Fable 5 and Opus 4.8 across web design, 3D, games, motion graphics, and a SpaceX investment deck — model names stay hidden until the ranking is locked in.
July 25thA three-tool workflow — Higgsfield for visuals, a packaged skill for the scroll effect, Opus 5 to assemble it — turns one reference photo into a finished animated site inside a single Claude window.
July 25thAlex Finn runs both models through five self-designed benchmarks, crowns Opus 5 the winner on price and quality, then spends the back half explaining why he's not fully switching.
July 24thA Claude Code YouTuber runs Opus 5 and Fable 5 through the same three-part marketing prompt to see if Opus 5's half-price tag actually pays off.
July 24thEighteen real Claude Code sessions later, the model that's half the price per token isn't automatically the cheaper one to actually run.
July 24thA live, blind seven-model benchmark pits personality against performance — and crowns the model its own host can't stand talking to.
July 24thA screen-by-screen read of Anthropic's Opus 5 announcement, benchmark chart by benchmark chart.
July 24thEvery's Dan Shipper spent a week with Opus 5 and comes back with a mixed verdict: pushy, prone to quitting early, and a real pain if you built workflows around Opus 4.8.
July 24thA launch-day test of Anthropic's new default Opus model — benchmarks, a landing page build, a script draft, and a motion graphic, all run head-to-head against Fable 5.
July 24thA side-by-side test of Anthropic's two flagship Claude models — agentic tool use, cinematic web design, and a from-scratch 3D browser game — judged on identical prompts with no cherry-picking.
July 24thA live reaction to Anthropic's Claude Opus 5 launch, walking chart-by-chart through benchmarks that put a mid-tier-priced model ahead of Anthropic's own flagship on almost everything except cyber exploitation.
July 24thA tour of a dozen apps Claude Opus 5 built from scratch, followed by Anthropic's own benchmark numbers on coding, computer use, and cost.
July 24thTwo AI coding models get the same prompts — a data-driven economy site and an endless flight simulator — and split the win, but only one of them costs half as much.
July 24thA creator walks through five concrete levers — effort level, model delegation, token-saving skills, research offloading, and advisor mode — for keeping Claude Code costs and weekly usage caps under control.
July 3rd