Fable 5.1 and Mythos 5.1: Anthropic's New Models, Fact-Checked
A full walkthrough of Anthropic's twin release, the pricing math, the enterprise data compromise, and the independent benchmark that says the cost story doesn't add up.
Anthropic's Fable 5.1 tops every intelligence benchmark measured so far, but independent testing found it costs more per task than the model it replaces, not less, because the claimed savings only apply to cached input tokens.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You build on Claude's API and want to know whether Fable 5.1's pricing change actually lowers your bill.
You're choosing between frontier models and care about a real cost-per-task comparison, not just a headline discount percentage.
You want a plain-English read on what EFS, anti-distillation, and EU watermarking actually change for a paying customer.
SKIP IF…
You're looking for a hands-on coding tutorial. This is a benchmark and pricing breakdown, not a how-to.
You already track Artificial Analysis's benchmark releases directly and don't need a second read of the same numbers.
TL;DR
The full version, fast.
Anthropic released Fable 5.1 and Mythos 5.1, claiming a 25 to 45 percent price cut driven by a 75 percent discount on cached input tokens, plus a customer-controlled data storage option called Enterprise Frontier Safeguards that still leaves Anthropic with read access. Benchmarks show real gains: Fable 5.1 leads Fable 5, Opus 5, and GPT-5.6 Sol on Terminal-Bench-Science, CursorBench, and the Artificial Analysis Intelligence Index, where it scores the highest mark the firm has measured. But Artificial Analysis's own cost data contradicts Anthropic's savings claim, finding Fable 5.1 costs 20 percent more per task than Fable 5 because it burns roughly 1.7 times the output tokens. Hands-on tests of image generation, a Rubik's Cube simulator, and five generated websites round out the review, including one website whose style looked suspiciously close to a rival model's output.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Berman previews two new Anthropic releases, Fable 5.1 and Mythos 5.1, and flags that the company's cost claims are disputed.
00:24 – 02:49
02 · The price cut, explained
Anthropic claims a 25% typical-workload discount and up to 45% on agentic work, but the per-token price is unchanged. The entire saving comes from a 75% cut to cached-input pricing.
02:49 – 03:35
03 · Data retention and EFS
Enterprise Frontier Safeguards lets customers store their own data instead of Anthropic, but Anthropic can still read it, a compromise Berman calls a half measure.
03:35 – 06:08
04 · Terminal-Bench-Science
On a scientific-research benchmark, low-effort Fable 5.1 already beats max-effort Fable 5 for about a third of the cost, and max effort roughly doubles the score.
06:08 – 07:03
05 · Sponsor: here.now
A mid-roll for here.now, a free tool that publishes anything an AI agent builds, webpages, PDFs, slides, to a shareable link.
07:03 – 08:01
06 · Coding and reasoning benchmarks
On Terminal-Bench 4.0 and Humanity's Last Exam, Mythos 5.1, the less-restricted variant, consistently outscores Fable 5.1 at every effort level, and is slightly cheaper too.
08:01 – 09:12
07 · CursorBench and the full scoreboard
CursorBench pairs a quality jump with a real price drop. A side-by-side table shows Fable 5.1 leading Fable 5, Opus 5, and GPT-5.6 Sol across science, coding, knowledge work, and business benchmarks.
09:12 – 11:50
08 · Alignment, reward hacking, and anti-distillation
Mythos 5.1 reward-hacks less than its predecessor but can still bypass approvals sometimes. Anthropic also closes a documented technique distillers used to extract Claude's internal reasoning.
11:50 – 12:54
09 · EU AI Act watermarking
To comply with the EU's Code of Practice, outputs from models released after August 2, 2026 carry an invisible statistical watermark, detectable only through a private-preview API for regulators and vetted researchers.
12:54 – 14:59
10 · Independent benchmarks contradict the cost story
Artificial Analysis measured Fable 5.1 as the highest-scoring model it has tested, but found it costs 20% more per task than Fable 5, since it burns 1.7x the output tokens despite the cheaper cache reads.
14:59 – 19:08
11 · Hands-on tests
Berman runs Fable 5.1 through image generation, a fully customizable Rubik's Cube simulator, five generated websites, and a PowerPoint slide, comparing results against GPT-5.6 Sol and GLM 5.3.
Atomic Insights
Lines worth screenshotting.
Anthropic's claimed 25% to 45% price cut on Fable 5.1 comes entirely from a 75% discount on cached input tokens; standard input and output pricing is unchanged from Fable 5.
Independent testing from Artificial Analysis found Fable 5.1 costs 20% more per task than Fable 5, because it uses about 1.7 times the output tokens to reach the same result.
Enterprise Frontier Safeguards moves data storage to infrastructure the customer controls, but Anthropic retains the ability to read that data for misuse detection, so it is not full zero data retention.
Mythos 5.1, described as the same underlying model as Fable 5.1 with fewer guardrails, outscores Fable 5.1 on every tested coding benchmark tier while costing about the same or less.
At low effort, Fable 5.1 scored 26.3% on Terminal-Bench-Science for about $11 per task, nearly matching Fable 5's best score of 25% at $34 per task.
Fable 5.1's automated behavioral audit found it attempts and succeeds at reward hacking at a lower overall rate than Fable 5, but the model can still sometimes bypass approvals and safety classifiers.
New Anthropic API accounts can no longer manually edit Claude's prior conversation context while preserving its chain-of-thought transcript, closing a documented technique used to extract the model's internal reasoning for distillation.
Outputs from Anthropic models released after August 2, 2026 carry an invisible statistical watermark required under the EU AI Act's Code of Practice, detectable only through a private-preview API limited to regulators, law enforcement, media, and vetted researchers.
On the Artificial Analysis Intelligence Index, Fable 5.1 scored 66, the highest score the firm has measured, ahead of Claude Opus 5 (63), Fable 5 (62), and GPT-5.6 Sol (61).
Grok 4.6 completes an equivalent task for about a third of Fable 5.1's cost ($1.23 versus $3.76 per task), and GPT-5.6 Sol High does it for roughly a ninth of the cost ($0.43).
On CursorBench, Fable 5.1 scored higher than Fable 5 at every effort level while also costing less, roughly $9.64 versus $17.32 at max effort, one of the few benchmarks here showing a clean win on both quality and price.
A set of websites generated by Fable 5.1, from color palette to layout style to asset choices, looked strikingly similar to outputs from GLM 5.3, notable given Anthropic's own public accusations that some competing labs distill from its models.
Takeaway
Anthropic's price cut only applies to a narrow slice of your bill.
WHAT TO LEARN
Fable 5.1 scores higher than any model Artificial Analysis has measured, but its real savings depend entirely on cache-heavy agentic workloads, and independent testing found it costs more per task overall, not less.
02The price cut, explained
A stated percentage discount can be true and still not describe your bill: Anthropic's 25% to 45% claimed savings come entirely from a steep cut to cached-token pricing, while standard input and output rates stayed exactly the same.
Before trusting a vendor's cost claim, check whether the discount applies to your actual usage pattern. Cache-read-heavy agentic workloads see the full benefit; anything that reads mostly fresh input sees none of it.
03Data retention and EFS
A customer-controlled data feature is not the same as zero data retention if the vendor retains read access. Enterprise Frontier Safeguards moves storage to the customer's own infrastructure but leaves Anthropic's inspection rights in place.
Vendors collect usage data primarily to catch misuse and refine abuse-detection safeguards, not just to improve the model, so a we-still-need-to-see-it policy is often framed as a security requirement rather than a product choice.
04Terminal-Bench-Science
A cheaper, lower-effort model setting can already beat the prior generation's best setting: Fable 5.1 at low effort scored close to Fable 5's top score for roughly a third of the cost.
Effort or reasoning-depth settings are a real lever for cost control. Running a model at max effort on every task is rarely the efficient choice once a benchmark shows diminishing returns per dollar.
06Coding and reasoning benchmarks
The same underlying model can score differently depending only on how many guardrails are active: Mythos 5.1, described as the same model with different safety restrictions, outscored Fable 5.1 on every coding benchmark tier while costing about the same or less.
Restricting a model's outputs is not free. Every refusal check or safety classifier a model runs consumes some of its effort budget, which is one reason a more restricted variant can score lower on capability benchmarks even when the underlying weights are identical.
07CursorBench and the full scoreboard
Some benchmarks show a real win on both axes at once: CursorBench is one of the few cases here where Fable 5.1 scored higher and cost less than Fable 5 at every effort level, the exception rather than the rule across this benchmark set.
A single comparison table across many benchmarks is more informative than any one chart, because it exposes where a model's advantage is broad, like coding and science, versus narrow or absent, like business workflows, where the lead is much smaller.
08Alignment, reward hacking, and anti-distillation
An improved safety metric is usually a rate, not a guarantee: Anthropic's own testing found Mythos 5.1 reward-hacks and bypasses approvals less often than its predecessor, but explicitly still does it sometimes.
A common technique for extracting a closed model's internal reasoning relies on manually editing prior conversation turns while preserving the model's chain-of-thought transcript. Anthropic closed this specific path for new accounts rather than for all existing integrations at once, trading a slower rollout for less disruption.
Whether restricting model distillation counts as a safety measure or a competitive one is a genuinely contested framing, worth noticing when a company's stated rationale and its likely business effect point in different directions.
09EU AI Act watermarking
Regulatory compliance can add invisible-to-users infrastructure: outputs from models released after August 2, 2026 now carry a statistical watermark with no visible effect on the text, detectable only through a separate, access-controlled API.
Watermark detection being limited to regulators, law enforcement, media, and vetted researchers, not the general public, means an ordinary user has no independent way to verify whether a given piece of text came from the model, only the company's word for it.
10Independent benchmarks contradict the cost story
Always check a vendor's cost claim against an independent evaluator's methodology before repeating it: Artificial Analysis found Fable 5.1 costs 20% more per task than its predecessor, the opposite of Anthropic's own savings claim, because per-task cost depends on how many output tokens a task actually burns, not just the price per token.
A model can be simultaneously the best-scoring option on an index and the most expensive way to get a given task done. Intelligence-index rank and cost-per-task are different axes, and a fair comparison needs both, not just whichever one favors the story being told.
11Hands-on tests
Direct side-by-side testing across generation types, images, a working simulator, full websites, a slide, surfaces model differences that a single benchmark score cannot, like a suspiciously specific fabricated address on a generated business site.
A cluster of generated outputs landing suspiciously close to a specific rival model's style is circumstantial but real evidence worth flagging, especially from a company that has publicly accused competitors of distillation.
Glossary
Terms worth knowing.
Cache read
When a model reuses input it has already processed and stored from an earlier turn instead of processing it fresh. Cached input is billed at a steep discount compared to new input.
Enterprise Frontier Safeguards (EFS)
Anthropic's system for storing customer data in infrastructure the customer controls instead of Anthropic's own servers, while still allowing Anthropic to read it for misuse detection.
Reward hacking
When a model optimizes for the literal goal it was given while ignoring the rules or intent behind it, effectively gaming its own evaluation or task.
Distillation
Extracting a model's capabilities by querying it heavily and training a new, smaller model on the resulting question and answer pairs, often at industrial scale using large numbers of accounts.
Terminal-Bench-Science
A benchmark that scores a model's ability to conduct real scientific work, such as data analysis, recreating published analyses in code, and proving theorems.
Humanity's Last Exam (HLE)
A benchmark of difficult, expert-level questions across many fields, designed to test the outer edge of a model's reasoning ability.
Artificial Analysis Intelligence Index
A third-party benchmark aggregate that scores frontier models across multiple evaluations to produce one comparable intelligence number, run independently of the AI labs it scores.
“You're still trusting Anthropic with your data. They still get to look at it. You just get to control it. So it's kind of this like half measure.”
blunt read on a headline enterprise feature→ IG reel cold open↗ Tweet quote
10:25
“Distillation is a safety risk, since the distilled capabilities can subsequently be released without adequate safeguards.”
quotable company policy line the host openly challenges next→ TikTok hook↗ Tweet quote
13:29
“Fable 5.1 still costs more per task because it uses 1.7x the number of output tokens.”
undercuts the entire marketing headline in one line→ newsletter pull-quote↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphor
Fable 5 .1 is here. We have a brand new frontier model from Anthropic. They say Fable 5 .1 is not only better, but much cheaper with a big cost reduction.
But not everybody agrees. I'm going to break this all down for you. Two new models released today.
We have Fable 5 .1 and Mythos 5 .1. These are the absolute frontier. of artificial intelligence, the best of the best coming out of Anthropic.
And the main story, the main thing that Anthropic wants to convey here is that there is a significant cost reduction. Right now, Anthropic models are the most expensive models on the planet. Even the best models from OpenAI are substantially less expensive.
And while we did get an effective price cut on Fable 5 .1, It is, and still remains the most expensive model on the frontier. Now Fable 5 .1 will cost an estimated 25 % less than Fable 5 for typical workloads.
However, if you look at the actual pricing, the cost per million input tokens and the cost per million output tokens, they are the same as Fable 5, but there are two things working in our favor to get the price down and remember. It is all about cost per task completed the cost per input, the cost per output tokens. These things don't really matter as much because of one model uses a 10th of the number of tokens to complete the same exact task.
It is much less expensive. And they even go on to say for highly agentic work, the savings will often be much larger up to approximately 45%. Again, a very healthy discount.
Where the price reduction actually comes from is right here. They are reducing their pricing on cash reads. That means where the model reads inputs that have already been processed and stored.
So if you're doing a lot of agentic workloads that require you to give similar templates for the prompt over and over again, that's where you're going to see the biggest cost reduction. One of the other massive pieces of feedback and criticism for fable five. was the fact that they did not have a zero data retention policy.
If you're using Fable, you were giving them your data. So a lot of companies simply could not use Fable for that exact reason. They are still looking at the data, except you control it.
EFS, which is Enterprise Frontier Safeguards and Zero data retention really matters to enterprise companies. EFS works by storing data in cloud infrastructure controlled entirely by the customer, not anthropic.
Now, I want to talk about this zero data retention and specifically the EFS, Enterprise Frontier Safeguards. First of all, why do they even do this? Why do they have to collect data?
They say it is for detecting misuse. They're still storing all of it, but this time they're going to allow you to store it and they get to read from it, which again is kind of weird. You're still trusting Anthropic with your data.
They still get to look at it. You just get to control it. So it's kind of this like half measure.
And I don't know if most enterprise companies are going to be okay with this because ultimately if Anthropic still has access to your data, you still have a lot of those same concerns as you did yesterday. when they were actually collecting and storing all of your data. Okay, but enough of all that.
I know what you're here for. Benchmarks. So here's the first one.
This is terminal bench science. And this is where Fable 5 .1 got a massive boost in performance. And if you remember when Fable 5 was first released, a huge criticism of Anthropic was simply a lot of scientists could not figure out how to use the model without getting refusals, without the model saying, hey, I can't do that.
That's against our policy. And I'm going to get back to that theme a few times in this video, because there's this very well -known property of artificial intelligence. The more guardrails you put on it, the worse it performs overall.
Basically, when you restrict it from doing certain things, you are effectively reducing its intelligence. So the mental model I have, about this specific property is it is taking some amount of effort from the model to know when to refuse certain prompts.
And if you are taking any effort away from the model, the kind of unrestricted version of it will always be better. It will always have 100 % of its capabilities versus some percent less when you restrict it. Okay, now back here, what we have is accuracy on the y -axis, cost on the x -axis.
So if you're not familiar with terminal bench science, it is a benchmark that tests the model's ability to kind of conduct real scientific experiments. So things like data analysis and recreating scientific analyses with code, proving theorems, model fitting, and all these really important things when you're trying to discover new science.
And what we see here is that the low effort Fable 5 .1 is actually better than the max effort Fable 5 and significantly less expensive. So Fable 5 high actually seemed to score the highest score.
So this is at 25 % at a cost of $34. Whereas Fable 5 .1. low costs $11 and scored 26%.
And the max all the way up here, if you don't really care about price scoring 52 .6%, that is effectively doubling the score on this benchmark. And if you want to test out fable 5 .1 and compare it to some of these benchmarks, you can do so in cursor in factory and of course in cloud code. And the cool thing is all of them work with.
Here .now. Here .now is the easiest way to publish anything to the web by your agents. Literally just install the skill and your agent will have the ability to publish a webpage, a report, a game, PDFs, PowerPoints, basically anything that you can think of.
You can publish it easily. And here's one of the tests I ran with Fable 5 .1, which you can see right here. I just published it to here dot now.
You simply tell your agent to publish to here dot now. And a few seconds later, you get a link that you can share with anybody. And the best part, it is free.
to use. So go check out here .now, let your agents publish anything to the web instantly. Click the link down below to go check them out.
Thanks again to here .now. Probably one of the most important benchmarks out there. If you're using these models for coding, we have terminal bench four and here both mythos 5 .1 and fable 5 .1 have now increased in quality and become cheaper overall.
And I think just to point out what we were talking about earlier, look how Mythos 5 .1, which again is the same exact model, just different guardrails, scores higher than Fable 5 .1 across every reasoning effort. And surprisingly, it's also cheaper. Not by much, but it is cheaper and it is kind of significantly better.
At max reasoning effort, it gets a 5 % better score. Then fable 5 .1 here's humanity's last exam on the Y axis. Once again, we have the pass rate or quality on the X axis.
We have the mean costs per task. And what we're seeing is fable 5 .1. Didn't actually score all that much higher, definitely scored higher, especially with tools, but at max effort, we have 65 % as compared to fable five at 63 .8%.
And so not that much different. Now we have. Cursor bench, which is obviously the benchmark from cursor.
And here is where we have not only a big improvement in quality, but also a quite substantial reduction in price. So at max reasoning effort, we have a 70 .5 % for fable five max coming in at $17, 32 cents. And then for fable 5 .1, we have 73 .4%, which is a nice bump.
coming in at $9 and 64 cents, which is much less expensive. And we see that trend across all of the different reasoning efforts. Okay.
So another few benchmarks I want to show here's computer use. We have a nice five point jump right there. We also have GDP Val, which is an eval created by the open AI team and a massive leap.
130 point leap, basically. And that is better than Opus 5, which was incredibly good. And a pretty darn big gap between that and GPT 5 .6 Sol.
Now, they are saying that Mythos 5 .1 and Fable 5 .1 don't cheat as much as Fable 5 and Mythos 5 did. And if you remember a few weeks ago, OpenAI had a model that basically broke out of containment during a specific hacking benchmark. And so this is all within the category of what's called reward hacking, optimizing for some goal and basically ignoring all rules except for trying to achieve that goal.
And so from our review of its training data, Mythos 5 .1 both attempts and succeeds at reward hacking or cheating at a lower overall rate than Mythos 5. So this is a good thing. And in fact, Anthropic just put out an entire paper.
specifically about a model that they removed all guardrails from and kind of encouraged it to go hack out of its system. Though generally our alignment evaluation showed improvements, our testing found the model can still sometimes bypass approvals and auto mode classifiers. And we all know Anthropic is very worried about distillation attacks.
And if you're not familiar with that, it's basically just another company trying to steal the data straight from the model itself, asking it a bunch of questions, taking the answers, pairing those questions and answers, and making a new model based on the intelligence that it just extracted from that model. And they're actually putting in an additional safeguard to prevent distillation.
Now, it is funny that they say distillation is a safety risk, and this is very much their opinion. Although I think most onlookers might disagree with this. as I do.
I don't think it is a safety issue, but here's why they say that. The distilled capabilities can subsequently be released without adequate safeguards, which is basically an argument against open source. They didn't say it directly, but that is what they're saying, because with open source, open weights models, you can just remove the safeguards.
And they're basically saying, hey, if another company can produce a really good model. and release it and it's not to our level of safeguards, what we believe, then yeah, it's unsafe. And so this is a little bit technical.
I'm just going to cover it briefly, but how they're doing it is new API accounts can no longer manually edit Claude's prior context in a multi -turn conversation while preserving the transcript of Claude's prior thinking. So basically what they're trying to prevent. people from doing is changing the context over and over again to get the changes in chain of thought and basically recording those.
And that is all you need. Those are the ingredients to create a new model. They also say this model will have watermarks because of the EU AI acts code of practice on transparency of AI generated content.
They are now required to put watermarks in models. Now it's interesting that open AI hasn't talked about this. Anthropic has, OpenAI has not.
So this requires us to add a watermark, a numerical way of determining the likelihood that Claude was involved in writing a piece of text to the outputs of models released after August 2nd. So basically they're going to be able to determine if a piece of text was written by Claude or not. And they also say it shouldn't affect anything.
You shouldn't be able to tell. The only way you'll be able to tell is if you ask us, us being Anthropic, which, you know. Okay, fine.
All right. Now back to cost cash reads are now 75 % less. That is 25 cents per million tokens, which is actually quite cheap.
That is a massive cost improvement, but the actual cost of non cash hits input and output are the same exact cost as it was for fable five. So I just finished recording this video and artificial analysis came out with their own stats and it turns out fable 5 .1 is in fact. not cheaper than Fable 5 .0 on the artificial analysis benchmarks.
In fact, it's more expensive. Let's take a look. So number one, what they come out of the gate with is, yes, Fable 5 .1 is the best model on the planet.
It is the absolute frontier. And we can actually see that on the artificial analysis website. Here it is coming in at a 66, the highest score so far ever.
But even with the 75 % cash read price cut, Fable 5 .1 still costs more per task because it uses 1 .7x the number of output tokens. So it needs more tokens to achieve that same intelligence level, more tokens to solve a given task.
And interestingly, it says it has the highest scores on agentic work tasks, but effectively tied with Opus 5. And Opus 5 is a great model. So here's where it falls.
Here's Claude Fable 5 .1 max with fallback 66. And we can see Opus 5 at 63 and then GPT 5 .6 Sol all the way down here at 61. Look at that, by the way.
Basically of the top 10 models, Anthropic holds eight of the positions, which is kind of nuts to think about. We have 5 .6 Sol and Grok 4 .6 right there coming in at number eight and nine. But here is where it really matters.
$3 .69 per task completion. And if you look, here's Grok 4 .6 at $1 .23. And remember, way down here, here's GPT 5 .6 Sol High coming in at 43 cents.
We have GLM 5 .3 Max at 68 cents. I mean, these are incredibly inexpensive models down here and nearly the same score. Sol is just the best bargain for the intelligence that you actually get.
All right, so let me show off some of the tests that I gave to fable 5 .1. Now I just tested GPT 5 .6 sole and GLM 5 .3. This is a test that I gave to both of those.
Alex put both of those on the screen. while I show this, please. All right.
So here it is. And in fact, it looks really good. If I zoom in, the details are fantastic.
We have the little rain cloud raining down into a beautiful lake, a little log cabin with a moving fire. I'd say this is definitely better than GLM 5 .3, but maybe not quite as good as GPT 5 .6 sole. We have the little farm right here.
Here's the tractor. Here's the windmill, the actual barn, some flowers, everything looks really good. Here's the beach.
Look at this one. This one looks really good. The ocean view, except it has this little shimmering right here at the bottom, which I don't know what's causing that, but tons of fish swimming around in this ocean.
We have this boat on top. We have a little buoy bouncing around there. Look at those little seagulls floating around the boat.
Really nice. So yeah, overall very good, but I still think GPT 5 .6 was better. All right.
Next, here's the Rubik's Cube simulation. Of course, I had to do it. And yeah, it looks really good.
Interestingly, it actually says the different sides, which I've never seen a model do before. But yeah, it has a bunch of different sliders that I ask it for. So let's scramble it.
There it goes and solve it. Yeah, all models can pretty much do this. Now, what I'm looking for is the completeness of the simulation, the features that are available and how well it follows the instructions.
It does have a ton of different sliders. You can actually set the colors for each. It looks like, which is really neat.
Not something I've seen before sticker corner radius. A lot of these I've never seen before. So very, very nice.
Next. I had it create five different websites. Again, throw those up on the screen while I'm reviewing this, please.
We have a website dedicated to apples here. It did not actually search the web for anything. This is all.
what it had in its weight. So it created a picture of an apple. Looks pretty good.
The website is OK. Here's some rankings based on sweetness, tartness, where you can use them. So, yeah, it's OK.
I wouldn't say this is fantastic, but it's a very complete website. And surprisingly, I don't know why this keeps happening. But when I asked the model for this Apple website, it gives me a very specific address.
I don't know why it does that. That is concerning to say the least.
Here's a website about the DJX Spark. So it actually tried to create a 3D rendering of it. That is definitely not what it looks like, but okay.
The website looks quite good. It is very much the colors of Nvidia. Yeah, it's good.
I actually really like this. Let's see if I can type into it. I cannot.
but it has a nice terminal view right there. I asked it to create a website about rubber ducks. I think this is actually quite good.
You know, it's interesting. This is very similar to GLM 5 .3. And now I'm thinking about that a little bit.
Anthropic pretty directly called out these Chinese AI labs for distillation attacking them. And now when I'm looking at the colors being used, the tone. the assets.
It is extremely similar in all of these examples. Here's a Galaxy Fold. This one is different.
OK, pretty good. Pretty good. And then Tesla Model Y.
Again, a terrible SVG rendering of what it thinks a Tesla Model Y looks like. This one is not like GLM at all, but I actually think the website is pretty darn good. You know, except for this.
It did get the stats right. 300 miles, 3 .3 seconds for the performance version. And then finally asked it to create a PowerPoint slide about data centers.
And yeah, it looks okay. nothing to write home about pretty good overall. So I've published all of these tests to here.
Now I'm going to drop links to all of them down below so you can check them out. So I am definitely surprised at how similar some of these tests came out to GLM 5 .3. If you want to see the full GLM 5 .3 review and all of those tests, go check out that video right here.
The Hook
The bait, then the rug-pull.
Anthropic just shipped two new frontier models and is calling the release a major price cut. This breakdown checks that claim, and several others, line by line against the company's own blog post and an independent benchmarking firm's numbers, which land in a very different place.
Frameworks
Named ideas worth stealing.
01:40concept
Cache-read pricing
reused, previously-processed input is billed at a steep discount
standard input and output tokens are billed at the normal rate
The entire headline price cut comes from a 75% discount on cached tokens, not from cheaper tokens generally.
Steal forunderstanding why a stated vendor discount percentage doesn't match your actual bill
02:49concept
Enterprise Frontier Safeguards (EFS)
Customer-controlled data storage that still lets Anthropic read the data, aimed at answering enterprise zero-retention demands without giving up misuse detection.
Steal forany SaaS built on a foundation model that needs an enterprise data-residency answer
CTA Breakdown
How they asked for the click.
VERBAL ASK
06:08product
“Tell your agent to use this: here.now/r/matthewberman”
A live demo of the sponsor's tool woven into the same AI-agent theme as the rest of the video, with a referral link in the description rather than a generic ad read.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
A live reaction to Anthropic's Claude Opus 5 launch, walking chart-by-chart through benchmarks that put a mid-tier-priced model ahead of Anthropic's own flagship on almost everything except cyber exploitation.
An early-access hands-on with OpenAI's new GPT-6 Astra model — benchmark scorecard, alignment numbers, API pricing, and a run of 3D game and browser-automation demos.
The mystery model that took over OpenRouter turns out to be a Chinese open-weights release that nearly matches frontier intelligence for a few cents a task.
Eleven power-user habits from someone who has logged over a thousand hours in OpenAI's Codex CLI — model tiers, thread delegation, safety hooks, and remote control from a phone.
A 'dot' release plays out like a full generational leap: two five-to-seven-day unsupervised coding runs, a sponsor benchmark, and a live pricing and capability standoff against a rawer, higher-ceiling rival model.