The argument in one line.
Anthropic's newly announced Claude Opus 5 matches or beats the rival Fable 5 model across nearly every benchmark category while costing roughly half as much per task, and the jump from Opus 4.8 to Opus 5 is a bigger leap than the gap between competing labs' flagship models.
Read if. Skip if.
- You're choosing between Claude and a competing frontier model for agentic coding or agent-building work and want current benchmark and pricing specifics.
- You build long-horizon autonomous agent tasks and care about a model's ability to verify its own work across many iterations, not just one-shot accuracy.
- You want a fast read of what actually changed in Anthropic's own announcement without opening the source page yourself.
- You want a hands-on Opus 5 coding tutorial — this is a benchmark read-through of an announcement page, not a build-along.
- You don't care about model-vs-model cost-per-task comparisons or effort-level tuning.
The full version, fast.
Anthropic's Opus 5 announcement shows the model matching rival Fable 5 at high and extra-high effort settings while costing about half as much per task, and only trailing at max effort by roughly 0.5% on CursorBench. The bigger story per the host is the jump from Opus 4.8 to Opus 5 itself — a larger leap than the gap to competing labs — driven by stronger self-verification on long-horizon agentic tasks: the model checks its work against a goal and iterates rather than one-shotting. Anecdotal examples include writing its own computer-vision pipeline to reconstruct a machine part from a drawing it couldn't view, fixing a bug at its root cause rather than the symptom, and building its own test harness to validate a market data feed with no live reference to check against. Pricing lands at $5/$25 per million input/output tokens, unchanged from Opus 4.8 and about half of Fable 5's rate.
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →Where the time goes.

01 · Opus 5 drops — beating Fable 5 on most benchmarks
Comparison table across agentic terminal coding, knowledge work, novel problem-solving, agentic search, multidisciplinary reasoning, computer use, agentic coding, business workflows, legal, health, biology. Opus 5 beats Fable 5 and Opus 4.8 in most categories, trails only on multidisciplinary reasoning. CursorBench: within 0.5% of Fable 5 at max effort, at half the cost; roughly equal at high/extra-high effort.

02 · Token efficiency vs. Sonnet 5
The Opus 4.8-to-Opus 5 jump is framed as the bigger story than the Fable 5 comparison. Unlike Sonnet 5 (described as disproportionately expensive per task despite improvements), Opus 5 is token-efficient relative to both Sonnet 5 and Fable 5.

03 · Visual output: wind tunnel and 3D cell demos
Two embedded Anthropic demo clips: Opus 5 generating a wind-tunnel airflow visualization over a car model, and building an interactive 3D cell artifact. Framed as closing a longstanding gap since Claude models can't natively generate images.

04 · Working with Opus 5: self-verification examples
Three anecdotal examples from Anthropic's page: reconstructing a machine part in FreeCAD from a drawing it couldn't view (wrote its own CV pipeline), fixing an open-source bug at its root cause instead of the symptom, and building its own test harness to validate a market data feed with no live reference. Misaligned-behavior audit and vulnerability-finding also improved.

05 · Safeguards and pricing
Cybersecurity classifiers are less restrictive than Fable 5's (85% less likely to trigger) via a Cyber Verification Program for vetted users. Opus 5 is priced at $5/$25 per million input/output tokens — same as Opus 4.8, about half of Fable 5.
Lines worth screenshotting.
- Opus 5 performs within 0.5% of Fable 5's peak CursorBench score at max effort, but at half the cost per task.
- At high and extra-high effort settings — where most users actually operate — Opus 5 and Fable 5 produce roughly equivalent output, with Opus 5 costing about half as much.
- The jump from Opus 4.8 to Opus 5 is described as a bigger leap than the gap between Opus 5 and its closest competitor, Fable 5.
- Opus 5 is Fable 5's near-equal on nearly every benchmark shown except multidisciplinary reasoning, where Fable 5 still leads unless tool use is allowed.
- Given a machine-part drawing but no way to actually view it, Opus 5 wrote its own computer-vision pipeline to extract geometry from raw pixels and rebuilt the part as a 3D CAD model.
- On a real open-source bug, a competing model's patch fixed only the surface symptom; Opus 5 found and fixed the underlying root cause.
- With no live data feed to validate against, Opus 5 built its own test harness to confirm a market-data-feed integration was parsing correctly.
- Opus 5 scores lowest (best) of Opus 4.8, Mythos 5, Sonnet 5, and itself on an automated misaligned-behavior audit.
- Opus 5's cybersecurity safety classifiers are described as 85% less likely to trigger than Fable 5's equivalent guardrails.
- Opus 5 is priced at $5 per million input tokens and $25 per million output tokens — unchanged from Opus 4.8, and about half of Fable 5's rate.
What Opus 5's benchmarks actually tell you
Benchmark wins depend heavily on effort level and cost framing, and the real story in this announcement is a self-verification jump, not just a leaderboard position.
- Benchmark leaderboards shift with effort level, not just model choice — Opus 5 matches Fable 5 at high effort but only Fable 5 pulls ahead at max effort, by about half a percent.
- A model's true cost advantage shows up in cost-per-task charts, not headline win/loss counts — Opus 5 reaches similar scores at roughly half the spend.
- The jump between two versions of the same model line (Opus 4.8 to Opus 5) can be larger than the gap between competing labs' flagship models.
- Long-horizon agentic tasks reward a model's ability to verify its own work and iterate, not just first-pass accuracy — that's a different skill than one-shot benchmark performance.
- When a model isn't given a way to directly perceive its input, like a drawing it can't view, it can still solve the problem by building its own tool to extract the missing information first.
- Safety classifier strictness and raw capability are separate axes — a model can get more capable and less restrictive at the same time, which changes how usable it is for edge-case-adjacent work like cybersecurity research.
Terms worth knowing.
- CursorBench
- A benchmark (version 3.2 referenced here) that scores agentic coding performance against cost per task at different effort levels.
- Frontier-Bench
- An Anthropic-referenced benchmark suite used to evaluate agentic coding and problem-solving performance, including the machine-part CAD reconstruction example.
- ARC-AGI-3
- A benchmark used here for 'novel problem-solving,' plotting score against total evaluation cost.
- Agentic effort level
- A setting (e.g. high, extra high, max) that trades more compute/cost for higher task performance on a given model; benchmark charts in the video compare models across these levels rather than at a single fixed setting.
- Misaligned behavior audit
- An automated evaluation Anthropic runs to score how often a model's behavior diverges from intended goals; lower scores are better.
- Cyber Verification Program (CVP)
- An Anthropic program mentioned in the safeguards section that gives enterprises and researchers who are already vetted early access to a version of the model with narrower cybersecurity safety classifiers.
Things they pointed at.
Lines you could clip.
“Claude Opus five just dropped in. Anthropic is claiming we now have a model that is getting close to Fable five outputs at half the cost.”
“At max max effort, Fable five does pull ahead, but most people aren't actually operating at max. They're either operating at high or extra high.”
“It went ahead and wrote its own computer vision pipeline to pull the geometry from raw pixels and then reconstructed the whole machine part.”
“If the numbers are true, we kind of almost have a better Fable for half the cost.”
Word for word.
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
The bait, then the rug-pull.
Anthropic just published its Opus 5 announcement, and the host reads it live on screen — a benchmark table and a run of cost-per-task charts that put the new model within half a percent of rival Fable 5's peak score, at roughly half the price per task.











































































