A 41-minute teardown of Claude Haiku 5.5: the price sheet, the 100,000-token cliff hiding under it, and the one job a cheap model is actually good at.
Posted
yesterday
Duration
Format
Review
sarcastic
Views
132.5K
2K likes
57 · 43
Big Idea
The argument in one line.
A small model earns its place by handling the cheap, high-volume work that surrounds a code change rather than the change itself, and Haiku 5.5 is the first Anthropic small model priced well enough to take that job.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You run coding agents against a paid API and your monthly bill is big enough that cache-read pricing changes your architecture.
You are deciding which model to put in a subagent slot in Claude Code, Codex, or your own orchestration layer.
You build products that call an LLM on high-volume, low-stakes work like classification, summarizing, tagging, or auditing records.
You want a concrete read on where Haiku 5.5 sits against GPT-6 Luna, 6.1 Sol, and Opus 5.5 on cost per task.
You are on a $100 or $200 Claude subscription and want to know what the bundled API credit actually unlocks.
SKIP IF…
You pay a flat subscription, never touch API pricing, and just want to know which model to click in a dropdown.
You want a hands-on build tutorial. Most of the runtime is charts, pricing tables, and other people's demos.
You need vendor-neutral benchmarking. The scoring here comes from one creator's own dashboard layered over third-party data.
TL;DR
The full version, fast.
Haiku 5.5 is the first Anthropic small model worth wiring into a pipeline. It runs roughly 75 percent cheaper than Haiku 4.5, at one cent per million cache reads and fifty cents per million output tokens, but only under 100,000 tokens of context. Past that line every price multiplies by five, so the cheap tier is a design constraint rather than a default. Benchmarks roughly doubled across the board, including a jump on terminal-bench from zero to about 39 percent. The real use case is orchestration. A smart model writes the plan and the code, while dozens of cheap subagents gather context, search the codebase, and verify results in parallel.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Anthropic nails flagship coding models and keeps shipping weak distilled ones. Haiku 4.5 was banned across every agent config here, and a public post said as much: no Anthropic small model was worth using.
02:18 – 03:43
02 · Sponsor break: CodeRabbit
Paid read for CodeRabbit, pitched on security review specifically: inline findings with a one-click fix, plus repo-wide deep scans that take about an hour to return.
03:43 – 05:29
03 · The announcement and the cache-read cut
Reading the Haiku 5.5 post: cheapest, fastest, most capable small model, about 75 percent less to run, positioned as a subagent. Buried in the same post, Sonnet 5.5 cache reads were halved from 20 cents to 10 cents, which lands as the bigger news here.
05:29 – 09:00
04 · Pricing, the 100k cliff, and how to stay under it
The full price sheet: one cent per million cache reads, 12.5 cents per million cache writes, 10 cents in, 50 cents out. All of it only applies under 100,000 tokens of context. Past that every rate multiplies by five. An Anthropic staffer's workaround is to pin Haiku's auto-compact window at 100k so subagents never cross.
09:00 – 11:40
05 · Benchmarks, speed, and the reasoning-off trap
Haiku 5.5 roughly doubles its predecessor almost everywhere, with terminal-bench moving from zero to 39.2 percent and a strong computer-use score. Speed runs 100 to 200 tokens per second. A widely shared report of worse results came from running it with reasoning disabled, which Anthropic models are not trained for.
11:40 – 19:42
06 · Tokens per task and the Pareto line
A long pass through a custom cost-versus-capability dashboard. Maximum reasoning settings distort everything, Sonnet 5.5 struggles to justify its slot, and when GPT-6 Luna and 6.1 Sol are added the honest value frontier narrows to three models with Haiku squeezed between them.
19:42 – 21:55
07 · API credits for subscribers
A release-note detail that resolves a fear about third-party tools losing subscription access: $200 subscribers get $200 of monthly API credit, $100 subscribers get $100. The Python and TypeScript SDKs also pick up computer use and browser use in beta.
21:55 – 24:02
08 · Egg drop: one Opus vs Opus plus ten Haikus
Anthropic's side-by-side demo. The flagship alone runs single-threaded and finishes in 3.5 minutes with 25 designs for 47 cents. The flagship directing ten cheap subagents finishes in under a minute with 86 designs for 14 cents, because parallel attempts became affordable.
24:02 – 27:47
09 · Community demos and the design skill
A programmatically generated video for about 60 cents in under 20 minutes, then a front-end comparison against competitors. The sharpest finding is that switching off the Claude design skill drops the same model straight back into generic model-default layouts.
27:47 – 31:40
10 · Fish slop, and what a Haiku build costs
A recurring game-build test. The result is competent but subtly wrong: no mouse look, broken collision corners, mistimed events. The receipts matter more than the game. About a dollar total, 9.73 million cache read tokens, only 102 fresh input tokens, and 47 of 51 requests over the 100k threshold.
31:40 – 33:49
11 · Coding is not the expensive part
The central argument. Once context is cached, emitting the code file is pennies. The expensive parts are gathering that context, verifying the result, and regenerating after failures. Chasing a cheap model to write the code optimizes the one step that was already cheap.
33:49 – 36:50
12 · Haiku corners itself; Opus routes around it
A live failure. The small model fires parallel GitHub fetches, exhausts 5,000 requests an hour, stores an error body as if it were data, then announces it will wait without actually scheduling a retry. The rerun hands the same job to the flagship with instructions to batch it across many cheap subagents, and it plans around the limit instead.
36:50 – 41:11
13 · The final lineup: Opus, Haiku, and where Sol fits
With the full dataset in, Sonnet adds almost nothing to the frontier while Opus plus Haiku covers the range. A rival lab's model is framed as a second opinion from outside the family rather than a middle rung, used as an auditor before a pull request goes up. Verdict: one small model that is sometimes worth using, which is a real change.
Atomic Insights
Lines worth screenshotting.
Cache reads can be 30 to 50 percent of an agentic bill, so a cut there moves your costs more than any headline token price.
Haiku 5.5 costs one cent per million cache reads and fifty cents per million output tokens, roughly twenty times cheaper than Sonnet on output.
The cheap Haiku 5.5 pricing only holds under 100,000 tokens of context. Cross it and every rate multiplies by five, not two.
You can pin a model's auto-compact window to 100,000 tokens so its subagents are structurally incapable of entering the expensive tier.
Terminal-bench went from a flat zero on Haiku 4.5 to about 39 percent on Haiku 5.5, which is the difference between unusable and usable in a terminal.
Turning reasoning off on an Anthropic model to save tokens produces worse pass rates and can run slower than a competitor trained for that mode.
Maximum reasoning settings distort every cost chart. One model burned about 200,000 tokens per task on max and looked more expensive than a flagship.
A cheaper model that burns more tokens per task can cost more than a smarter one, so compare cost per task and never price per token.
In the vendor egg-drop demo, one flagship model took 3.5 minutes and 25 attempts for 47 cents; the same flagship driving ten cheap subagents took under a minute and 86 attempts for 14 cents.
Parallel cheap attempts only pay when verification is trivial and most attempts are expected to fail.
Generating the code file is the cheap step. Gathering the context to write it and verifying it afterward is where the money goes.
A full game build on Haiku 5.5 cost about a dollar, and 47 of its 51 requests crossed 100,000 input tokens into the five-times price tier.
The real cost of a weak model is not tokens. It is the state it leaves you in, like burning 5,000 GitHub API requests an hour and then stalling without scheduling a retry.
Turning off the Claude design skill collapsed the same model's front-end output back into generic defaults, so the skill was carrying the visual quality, not the weights.
If a mid-tier model never touches the cost-versus-capability frontier, a two-model stack of flagship plus small covers nearly everything.
A rival lab's model is more useful as a second opinion than as a middle rung, because different training catches different mistakes.
Takeaway
The cheap model is for everything but the code.
WHERE SMALL MODELS PAY
A cheap model earns its keep on the searching, reading and verifying that surrounds a change, and loses money the moment you ask it to make the change itself.
01The missing piece in the lineup
A lab that pours everything into its flagship tends to ship badly distilled cheap models, and the cheap tier is where high-volume work actually lives.
A small model you cannot trust gets routed around, so you end up paying mid-tier prices for work that should cost pennies.
03The announcement and the cache-read cut
Cache reads can be 30 to 50 percent of an agentic bill, so a cut there moves your costs more than a headline token price does.
When a mid-tier model's cache reads cost the same as the flagship's, the mid tier stops being a real saving and the lineup collapses to two useful models.
04Pricing, the 100k cliff, and how to stay under it
Haiku 5.5 runs at one cent per million cache reads and fifty cents per million output tokens, roughly twenty times cheaper than Sonnet on output.
Crossing 100,000 tokens of context multiplies every rate by five, so the headline price only holds for short, bounded calls.
You can pin a model's auto-compact window to 100,000 tokens so its subagents are structurally unable to enter the expensive tier.
05Benchmarks, speed, and the reasoning-off trap
Haiku 5.5 roughly doubled its predecessor across most benchmarks and moved terminal-bench from zero percent to about 39 percent.
Anthropic models are trained to reason, so disabling reasoning to save tokens produces worse pass rates and can run slower than a competitor built for it.
Throughput of 100 to 200 tokens per second is what makes a small model usable for browser work and live support, not raw intelligence.
06Tokens per task and the Pareto line
Maximum reasoning settings distort cost comparisons. One model burned about 200,000 tokens per task on max, which made a cheaper model look expensive.
Plot cost against capability and the honest value frontier is usually three models wide, not a full vendor lineup.
A cheaper model that uses more tokens per task can cost more than a smarter one, so compare cost per task and never price per token.
07API credits for subscribers
The $200 subscription tier now includes $200 of monthly API credit, which makes building against the API effectively free up to that ceiling.
SDK support for computer use and browser use turns a fast small model into a practical automation driver rather than a chat endpoint.
08Egg drop: one Opus vs Opus plus ten Haikus
In the vendor demo, one flagship model solved the task in 3.5 minutes across 25 attempts for 47 cents, running single threaded.
The same flagship directing ten cheap subagents finished in under a minute with 86 attempts for 14 cents, because parallel attempts got cheap.
Parallelism only pays when verification is cheap and most attempts are expected to fail.
09Community demos and the design skill
A programmatically generated video demo cost about 60 cents of API usage and took under twenty minutes to produce.
Switching off the design skill collapsed output back into generic model-default layouts, so the skill was carrying the visual quality, not the weights.
10Fish slop, and what a Haiku build costs
A full game build on the small model cost about a dollar, dominated by 9.73 million cache read tokens with only 102 fresh input tokens.
47 of 51 requests crossed 100,000 input tokens, so nearly the whole run billed at the five-times tier despite the cheap headline price.
Cheap models fail differently. Fewer loud mistakes, but subtle wrong choices like broken collision corners and mistimed events.
11Coding is not the expensive part
Generating the code file is cheap once context is cached. Gathering that context and verifying the result is where the money goes.
Regenerating a change eight times because the first seven were wrong costs far more than paying a smarter model to get it right once.
Hunting for a cheap model to write the code optimizes the one step that was already inexpensive.
12Haiku corners itself; Opus routes around it
A weak model burned 5,000 GitHub API requests on parallel fetches, stored an error body as valid data, then stalled without scheduling a retry.
The real cost of a cheap model is not tokens. It is the state it leaves you in when it corners itself.
Instructing a smart orchestrator to batch work across many cheap subagents gets the same parallelism without the self-inflicted rate limits.
13The final lineup: Opus, Haiku, and where Sol fits
If a mid-tier model never touches the value frontier, a two-model stack of flagship plus small covers almost everything.
Treat a rival lab's model as a second opinion rather than a middle rung, because it reasons from different training and catches different mistakes.
A practical loop: the flagship writes the change, a model from another family audits it, and the pull request goes up already scrubbed.
Glossary
Terms worth knowing.
Cache read
A billed request that reuses context the provider already stored from an earlier call. Because the tokens do not need reprocessing, cache reads are priced far below fresh input tokens and dominate the bill in long agent sessions.
Cache write
The charge for storing a chunk of context so later calls can read it cheaply. Writes cost more than normal input tokens, and the stored context expires after a set window, commonly five minutes or one hour.
Context tier pricing
A pricing scheme where crossing a context-length threshold moves the whole request into a more expensive rate card. Holding more tokens in memory costs the provider more RAM, so the higher tier passes that on.
Auto-compact window
The point at which a coding agent summarizes its own conversation to free up context. Setting it low keeps every request under a cheaper pricing tier at the cost of the model forgetting more detail.
Subagent
A separate model instance spawned by a primary model to handle a bounded task such as searching files or checking a hypothesis. Subagents can run in parallel and often use a cheaper model than the one directing them.
Pareto line
On a cost-versus-capability chart, the curve connecting the best score available at each price point. Any model sitting below the line is dominated, meaning something cheaper is just as capable or something equally priced is smarter.
Tokens per task
The average number of tokens a model consumes to finish one benchmark task. It exposes verbosity, which is why a model with a low per-token price can still end up with a high per-task cost.
Terminal-bench
A benchmark measuring how reliably a model completes multi-step professional work inside a command-line interface. Scoring zero means the model effectively cannot operate a terminal unsupervised.
Reasoning effort
A per-request setting controlling how many hidden thinking tokens a model may spend before answering, typically from low to maximum. Higher settings improve accuracy and multiply cost.
Computer use
A model capability for driving a graphical interface directly, reading the screen and issuing clicks and keystrokes. It is latency-sensitive, which makes throughput matter more than raw intelligence.
Drop-in replacement
Swapping one model for another behind an existing prompt and pipeline without changing either. It often fails because prompts are tuned to a specific model's reasoning behavior.
08:16linkLydia at Anthropic on capping the auto-compact window
Quotables
Lines you could clip.
31:52
“coding is not what makes llms expensive”
six-word contrarian claim, no setup needed→ TikTok hook↗ Tweet quote
36:29
“if opus can solve a problem in 100k tokens and haiku burns a million before it figures out it can't haiku ends up more expensive”
self-contained argument that flips the cheap-model assumption→ newsletter pull-quote↗ Tweet quote
23:18
“for the subset of tasks where more attempts is more valuable than smart attempts, Haiku 5.5 is going to slaughter”
names the exact condition for using a cheap model, with a punch at the end→ IG reel cold open↗ Tweet quote
40:37
“Anthropic now has one small model that is sometimes worth using. And that is a very nice change.”
the verdict, landed as an anticlimax→ TikTok hook↗ Tweet quote
13:17
“you also see 4.5 Haiku in here which I am going to turn off because we don't care about that model anymore, can finally be left where it belongs: the grave”
the single most quotable burn in the runtime→ IG reel cold open↗ Tweet quote
30:42
“47 of 51 requests exceeded 100k input tokens, so they used Haiku's higher price tier”
hard number that undercuts the cheap headline price→ newsletter pull-quote↗ Tweet quote
38:38
“It's like you're asking your friend who works at a different company for their thoughts.”
clean analogy for multi-vendor model stacks→ newsletter pull-quote↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphoranalogystory
Anthropic is pretty good at making big models for coding. We've seen that with models like Fable and Opus, and all the way back in the day with models like Sonnet 3 .5 and the idea of tool calls kind of making us software engineers use AI the way that we do today. But there's always been a bit of a weakness with their lineup.
It's the smaller models. And as cool as Haiku 4 .5 was when it came out... It came out in October of last year and it was expensive at the time.
Now it's just laughable. And I have put a lot of effort into adjusting all of my tooling, all of my agent and cloud MDs and all those things to make sure every tool in my system makes it very, very clear to all my agents that Haiku 4 .5 should be avoided at all costs because it's useless. It's just not very good.
That's why I made this post a few weeks ago where I said that Anthropic has no small models that are worth using right now and also dunked an open AI saying they have no large models worth using because at the time they had Astra and 6 .0 Sol. They didn't have 6 .1 Sol yet. And of course Google gets the fun jab at the end saying none of the Google models are worth using.
But the thing I want to focus on here is this idea that Anthropic has no small models worth using. Historically, Anthropic puts almost all their effort into the biggest model and then poorly distills it into the cheaper ones. That's why when they had Mythos and Fable, the quality of Opus plummeted and we ended up with 4 .6, 4 .7, 4 .8, and then eventually Opus 5, which was terrible.
That was all because of their focus on Anthropic's large models. That curse has been broken, though, with both Opus 5 .5 and Sonnet 5 .5, which is why I'm so hyped for today's introduction of Haiku 5 .5. They've been teasing this one for a bit, and I had high hopes but low expectations because this is one of the biggest missing pieces.
The model the other models use to search for files, to triple check changes, to do all the things that a model shouldn't waste tokens on if the model is expensive. I got to the point where I had taught my Claude to call Sol in order to save money because I didn't trust it calling Haiku 4 .5. So how did Haiku 5 .5 turn out?
I don't want to spoil too much. but i think it's fair to say they made something pretty cool the price is right the performance seems solid there's a handful of tricks that make it a little extra special especially considering how small this model is but before we can get started In honor of the nature of this small model, we should take a really small break for today's sponsor.
One of the painful side effects of having such good models for coding is that all of our code bases are getting way bigger and more complex. And while AI review tools are good at helping us avoid bugs, are they actually helping us prevent major security issues? And even if it does find one, is it actually going to let you fix it?
Because let's be real, most of these models that are this good just refuse outright once you start doing security stuff. That's why I've been so reliant on today's sponsor. You've already heard of them.
It's CodeRabbit, but I'm showing something different today. But if you remember a few seconds ago, I said that the models can't really help with the security things because they're not allowed to. Well, CodeRabbit has made all of the deals they need to and found all the models they have to to get real security feedback for you as you're working.
Not only will they give you real security callouts in your pull requests, they'll give you the tools needed to fix them. There's just like a little fix button that will run on their side with the models that are approved to do this. They also have something that's even more important, which is the ability to run deep scans across an entire repo.
You click this little button in the corner, you choose the repo you want to run on, and then half an hour to an hour later, you have a ton of real feedback. I ended up addressing over half the things that were found in both of these reviews. And if I hop to this findings tab, I'm going to need my editor to start censoring things because the things it finds are actually quite useful to know about.
And I'll be real, some of these are bad and I need to go address them. As our software is getting more complex, keeping it secure is only getting more important. Secure your stuff at soydev .link slash coderabbit.
Before we dive into everything about Haiku... I want to show one thing quick on Slopalytics. We currently show all of Anthropic's current models.
Sadly, we don't have Fable 5 .5 yet, but we do have Opus 5 .5, Sonnet 5 .5, and now Haiku. Let's just show the last Haiku model real quick. Yeah, hopefully not much more needs to be said.
Let's go through the announcement. Introducing Haiku 5 .5, the cheapest, fastest, and most capable small model we've ever released. Haiku 5 .5 is designed for high -volume, cost -sensitive tasks.
It reliably handles quick and repetitive workloads, things like summaries, compactions, database queries, and classification requests. It pairs well with Opus and Sonnet as a sub -agent on coding work. And since it's also our fastest model to date, it works especially well for speed -sensitive tasks like live customer support and browser use.
Haiku 5 .5 is available at a much lower price than 4 .5. On average, it's now about 75 % less money to run. Along with this launch, We're making improvements to the value of our model range.
We are having the price of Sonnet 5 .5's cash reads. Woo! I actually missed that detail.
Huge. That was the thing I roasted them on in the Sonnet 5 .5 video. That is a huge change, actually.
That makes the whole lineup make way more sense. That might actually be my favorite thing about this release now. Oh, shit.
I actually didn't know about that. So while I sit here in rage that nobody mentioned this to me, because I'm a huge fan of changes to cash read prices. I have stories I wish I could tell that I can't.
If you know, you know. Let's just say I've been deep in the world of cash pricing as of recent. And the one complaint I had about Sonnet that also made Sonnet 5 .5 much less appealing than 6 .1 Sol from OpenAI is that the cash read price for Sonnet 5 .5 wasn't just too expensive for Sonnet.
It was the same price as Opus. So the thing that made up about 30 to 50 % of your bill was the same for Sonnet. and for opus which meant the price difference between these models was not particularly large if you use them for agentic work cash reads are essential to making these things reasonably cheap so with that in mind let's look at price quick because this is what's exciting to me sonnet 55 went from 20 cents per million cash reads to 10 cents which is the same cash read price as a model like gpt61 soul finally making those a little bit closer in price i hope artificial analysis reprices their charts with this new price so that i could show the different numbers because that's a really big improvement especially for agentic work haiku 4 .5 was that same cash read price and then half the price across the board haiku 5 .5 is a tenth the cash read price it's one cent per mil tokens red 12 .5 cents per million cash token rights 10 cents for normal one mil read, and 50 cents per mil out, where previously it was five bucks per mil out, and Sonnet 5 .5 was 10 bucks per mil out.
This makes it 20x cheaper in output tokens and input tokens than Sonnet, comically cheaper in cash writes as well, about, again, 20x cheaper, and then cash reads are a tenth the price too. There is a catch though. that is up to 100k token context this is a bit of a weird thing because anthropic used to charge when you broke 270 or 250k token context and doubled the price they stopped doing that somewhat recently that makes a lot of sense for tools like clod code because you don't have to like micromanage the context to keep your bill low and cash reads are so cheap it's kind of just fine but 100k token limit seems very small especially for like a long -running agentic work i see this as a couple things and also i'll just share in advance here If you go over 100K tokens, the price doesn't double.
It 5Xs. So you go from $0 .50 per mill out to $2 .50 per mill out. You go from $0 .10 per mill in to $0 .50 per mill in, etc.
The reason for that, infra -wise, is because when you have a certain threshold of tokens in the context, that is now much more RAM you need to hold it all, which does make it more expensive. What I think happened here is they made the price for under 100K way cheaper, possibly like eating some of their own margin there. and they're making up a little bit of it with the over 100k.
I see this a bit differently, though. First off, I see this as a method of making the main use case, which is quick reads and categorization checks like that, as absurdly cheap as possible. Something like Jev is the real competitor to a small model used in these ways, where it's being handed text and categorizing, organizing, all these types of things.
By making under 100k this cheap, they are making it harder and harder to justify spinning up another API when you could just throw it at Haiku. And then the other side is to make sure they don't bankrupt themselves when people use this inside of Cloud Code. This is a way of steering people towards the intended use case by making it really, really cheap if they use it on short, quick response things, but also making it reasonably expensive, still half the price like it was before, but not for free like it is in the under $100K range if you're using it in tools like Cloud Code.
It is worth noting that anthropic employees like Lydia have come out publicly and said, if you want to keep your bill cheap and not worry about this, you can manually set haiku 5 .5's auto compact window to 100k so that you're always working within that cheaper token pricing apparently it's a save per model so this will only apply to haiku including sub agents so if you go do this in your cloud code if you're working on api billing or you're on like bedrock or something this allows you to guarantee the sub agents will never cost a shitload of money so now we talk about everything else all the performance all the use cases and of course how it performs in my silly suite of benches.
As we can see here, it absolutely decimated Haiku 4 .5, getting more than double the score in the vast majority of things. And one that my friends at Anthropic mentioned to me is that they were very happy to see that Haiku was no longer scoring a zero on terminal bench. It made their way up from zero to almost 40%.
also worth noting that as incredible as luna is and i still think gbd6 luna is an unbelievable model gbd6 luna only got a 16 .4 so it more than doubled luna score for code so if you're one of those people that thought luna was decent at code i question your judgment i question a lot about you honestly but at the very least you now have a model that is allegedly more than 2x better at coding for not too much more money one of the most interesting numbers for me here is computer use I've mentioned before that OpenAI has been pretty far ahead in computer use.
And also, I don't think these benches are great because it doesn't show the gap as prominent as it is in my experience. Haiku slaughtered here. It did way better than 6luna did by like almost 40 % or so.
This makes this a very solid small model for doing computer use work. Potentially. I haven't tried it yet.
I've heard good things though. And when you see the speeds, you'll see why I'm excited. because haiku 5 .5 is seeing numbers from 100 to 200 tps depending on the provider i asked a notify by my chat that prime tweeted he was using haiku in some work he was doing with reasoning off for automations he swapped to haiku it was a drop in replacement but he's actually seeing a worse pass rate and a much much higher fail rate as well as it being noticeably slower than what he was experiencing with luna this does not surprise me too much because the big thing that he did here is he turned off reasoning.
He has both Six Luna and Haiku 5 -5 on no reasoning. OpenAI has trained their models to be good at that. Anthropic has trained their models to be bad at that.
With Opus and Sonnet 5 -5, if I recall correctly, I know they did this with Opus. I think they did it with Sonnet. They turned off no reasoning, so you always have to let it reason at least a little bit.
Haiku left no reasoning as an option because he will use it for work that they just want a response in quick, but is not very good at that. i would honestly say that for like categorization work if you're turning off reasoning anyways you should just use jev i have been playing with the open ai decisions api and i have plenty of thoughts to share about that in the near future but for now just use jev unless you need like vision and then haiku or luna makes sense probably hike or probably luna because it's cheaper especially with reasoning off but uh yeah take that as you will for what it is worth when i try out haiku five five on my machines with clod code and my normal clod code sub I was seeing around 180 tokens per second for my just traditional prompting working on things like fish slop, which of course will be showing fish slop soon.
So while Prime's numbers with reasoning off weren't great, Haiku's numbers with reasoning on do seem very good in this tier, which means it's time to hop into slopalytics. Let's look at tokens per task because I am genuinely very curious and for good reason. It seems like this is not the usual wasteful shit show that we see from these smaller models from OpenAI.
For reference, Sonnet 5 .5 on Max did almost 200 ,000 tokens per task average, putting it at, I think, the highest ever on artificial analysis, where Fable did 78k tokens, under half as many. Yeah, it was 2 .5 times the tokens per task with Sonnet 5 .5, which resulted in a cost for Sonnet 5 .5 that was slightly higher than Fable 5 .1.
The reason for this is an unreasonable setting. It's right here. It's Max.
So on Slopaletics, you can click the word max, which makes it go away. And now we have a chart that makes a lot more sense because those max reasoning levels are bad. And now we see tokens per task numbers that make way more sense and also cost per task numbers that make way more sense.
Raikou 5 -5 is now all the way on the left here, cheaper than anything else because I only have X high and max numbers for it. The max number is still pretty cheap, but more than I would want it to be, especially once you add back in models like 6 -1 Sol. because six one soul on low is as cheap as haiku on low largely due to the token efficiency difference but let's go back to the cost versus intelligence because i think this is fun you also see four five haiku in here which i am going to turn off because we don't care about that model anymore can finally be left where it belongs the grave and we'll notice some weird things here all this data is from artificial analysis and it is far from representative of everything these models could or would or should be used for but you'll see that aggregate across all their benches Sonnet 5 .5 performs worse than Opus at the same, if not a slightly higher cost, which meant that Opus and Sonnet aren't as complimentary as one might hope.
But if we look all the way to left here with Haiku 5 .5, things are looking a lot more promising. Because even though these scores are close, Opus was almost five times as expensive as the run was with Haiku. Pretty solid.
I have a fun feature here as well, the Pareto line, which is the best performance at a given price point. And this line moves around based on like, what is the best you can do at that level of intelligence? And you can see some very interesting things here.
Of course, as you expect, Haiku 5 .5 does very well here with X high and max being the best value in that range within Anthropix line of models. Sonnet 5 .5 sneaks in randomly with high because it is smarter than Opus low and cheaper than Opus medium. But after Sonnet high, All the best value, according to this bench, is just Opus 5 -5, medium, high, X -high, max.
But here's where I must admit I was slightly misleading you, because I only am showing anthropic models right now. In fact, I'm only showing anthropic models that are the latest for each of these lines. And also Fable 5 -1 doesn't even light up anymore, because there's no reason to use that when you look at the Pareto line here.
Ready for where things start to get more interesting, though? First off, we can switch to linear pricing where you see the gap isn't quite as big as it seems because the log pricing makes things seem further away on the cheap side and closer on the expensive side than they really are. You can see here that Sonnet X high is a fine value at $2 .75 per task, but then Sonnet 5 .5 max makes literally no sense at $7 .60 per task.
If you're using max reasoning levels, I have questions. But I also will say very confidently, do not use Sonnet 5 .5 on Macs. It makes no sense at all.
But again, still just looking at Anthropic models. What happens if we put on like Gemini 4 Argon? Oh, not much.
How about 3 .8 Flash? Still not much. In fact, it almost looks like we took the Haiku line and moved it to the right and down.
So we don't really need those anymore. We'll just turn those back off. How about, I don't know, Rock 4 .7?
Still not very interesting. Still below all of this. How about 6 Astra?
Oh, they got a point on. Astra on low is around the same intelligence as Sonnet 5 -5 on high. So that adds to the Pareto line.
But here's where things get more interesting. Let's put on GPT -6 Luna. Oh, huh.
That just extended the line a whole bunch to the left. Because Luna is cheaper than Haiku always. It's also dumber always, at least with the numbers they have.
but these are complementary this is a pretty stable line if you turn on the pareto line it has all of luna all of haiku then it briefly dips to astra then sonnet then opus for the top and if you were to turn off astra and sonnet you have a pretty clear simple line here for the best value per dollar but here's where things get messy and turn out the proto line to make this clearer i'm also going to turn off fable and sonic so they don't really make sense here Now what happens when we turn on 6 -1 -6?
Oh.
That's unfortunate. Turn on the Pareto line. And I guess technically Haiku is slightly smarter for meaningfully more money there, but then just barely more.
Going from 21 .3 cents to 21 .4 cents, you get a huge jump in intelligence. And if we were to turn Haiku off here, oh, that's a much smoother line. So if you're actually trying to optimize based on the benchmark scores for the best intelligence per dollar, your stock would look like 6 Luna, 6 .1 Sol, and Opus 5 .5.
There is one more hidden gem diamond in the rough that I'm surprised people aren't talking about as much. Mimo V26 Pro is an underrated open weight model. Pretty damn good for the price.
But other than that, nothing really comes above this line. So if you're trying to get the absolute most value and you're paying API prices, six luna six one soul opus gets you pretty much all of that that said haiku five five is neck and neck now are soul and luna still better per dollar slightly usually however if you are trying to do everything through one provider like you want to just use your anthropic inference because you have a crazy spend target with them and you just use their apis or you're subscribed to cloud code and don't want to also subscribe to codex Or you want models that understand each other better because you're using Opus for everything.
And when Opus prompts Solit, it's worse than when Opus prompts Haiku. I don't know how true that is, but it's real. There's a lot of reasons that you would want a model in that range in this family.
And for once, it's no longer irresponsible to use Haiku like it was before. Haiku 4 .5 was so bad and so unnecessarily expensive that there was no justifiable reason to use it. when i had my tweet that anthropic has no small models worth using opening eyes no large models worth using and google has none worth using somebody replied all three statements are false to which i replied lmao this guy uses haiku 4 .5 and then he got torn to teeny tiny pieces in my replies because there is literally no reason to use haiku 4 .5 at this point in september of this year but also right now especially Haiku 4 .5 was barely usable at the time and quickly became garbage and just wasn't touched because Anthropic was too focused on the big models.
Now they're not. And while I do wish this model pushed the Pareto line a bit more, I am thankful Anthropic is embracing models in this range because together, these all actually look quite reasonable. Also remember, Sonnet 5 .5 got way cheaper with the cash read price change, so that might actually adjust things even more.
Good news, Artificial Analysis updated their data. which means I've updated mine. And now Sonnet is shifted quite a bit to the left, which makes it much more compelling on the Pareto line.
Actually, not that much more. It does make this bottom Sonnet 5 .5 high feel less awkward as a dip. But yeah, that is improved.
Sonnet feels less bad now by quite a bit with that change. Okay, enough of the bench maxing. Let's talk about the model and how it actually works.
Okay, one last fun thing from Anthropix release notes here that is actually quite cool. they want to encourage people using cloud code to build more things with the cloud apis a lot of people myself included were scared with this announcement thinking this might be the end of using your sub in tools like t3 code because yes in t3 code you can use your sub with anthropic openai or whatever else if it's cloud code or codex installed on your computer it will just work with t3 code and a lot of us are scared because they said they would kill that and they haven't yet but they did a cool thing here an actually really cool thing where if you're subscribed as a 20x user so you're on the 200 plan You get $200 in API credit every month.
5x users get $100 in credit. And this is for the API. So you're effectively able to get an API key like you would if you were paying, but they preload it with $100 or $200 every month to throw at whatever.
So if you want to vibe code an app in cloud code that has Haiku or Sonnet built into it and then send it to your friends, you now have $200 of credit for that, which I actually think is really nice of them. Do I say these things? Yeah, fuck it.
I chatted with them a lot about this back in the day because they wanted to make sure they didn't screw up things with how like Cloud Code SDK and Agent SDK integrations were going to end up. They were going to do something a lot stupider with this. And I'm thankful they talked to people like me and that they listened because they turned what could have been an L into a really genuinely cool thing.
If you guys have noticed that I'm being nicer to Anthropicus because they're actually listening now and this is a great example of one of those things. Ooh, actually, one more fun small detail here. They're updating the Cloud Python and TypeScript SDKs to add support for computer use and browser use.
That is potentially very useful for us. T3 Code might level up from this. I just saw Elon reply to them with, congrats on another great model.
I would be surprised if Haiku 5 .5 wasn't nicer to use than Grok 4 .7, if I'm being real. Oh, I missed the artificial analysis jump chart. This is hilarious.
Beautiful. From the furthest right end here to actually competitive. They are a bit more brutal showing it is not on the Pareto line.
But yeah, the data is updated. It's more interesting than they give it credit. Regardless, I'm tired of talking about benchmarks.
I want to showcase the awesome things this model can do. and not just by itself but when it's used alongside models like opus 5 .5 and sonnet 5 .5 the best use cases for haiku 5 .5 aren't clicking it in the ui or selecting it in cloud code and using it it's letting your agents use it and letting your applications use it building it in so when data comes in haiku will audit it and categorize it or having sonnet or opus call haiku for testing different things out exploring the code base to find stuff anthropic made a really cool demo here where they had opus 5 .5 by itself try to build a safe egg drop in this emulation of like egg drop mechanics that they built and on the other side they had opus five five plus ten haiku five five sub agents do it and i think the results here are actually quite interesting the thing that you might notice pretty quickly here is that opus is only trying one at a time because it's running single threaded and if you were to have it run multiple threads it would end up wasting way more money
where on the right we have Opus plus 10 Haiku sub -agents, and it's able to try more things at the same time and save a ton of money and get to a result faster as well. There you go. It took three and a half minutes for Opus 5 .5 by itself to succeed at this task.
It only designed 25 attempts, and it cost 47 cents to do, but with Opus 5 .5 commanding 10 Haiku instances, it was able to do it in under a minute with 86 attempts. way more attempts because it's able to get more renditions out when it's that much cheaper and faster and only cost 14 cents. So for the types of tasks where the majority of answers are wrong, like if you're looking through a code base to find one file out of a bunch and you can't just do like a rip grip call to find it, Haku will be way cheaper and faster because it can do a lot of things at once without causing crazy bills.
so for the subset of tasks where more attempts is more valuable than smart attempts where you'd rather try 10 times and verify each result then spend a little more time trying once and hoping it's right if the gap between a random try from a dumb model and a thoughtful try from a smart model isn't that big haiku 5 -5 is going to slaughter This is a small subset of work, and I am hoping Opus is smart enough to know what that subset is.
I have not pushed it hard enough to know yet, but I am going to go edit my CloudMD to no longer force everything to be Opus or Sol subagents, because this is cheap enough that it would make real sense. But you guys aren't here just to listen to me rant about subagent architecture. You're here to see what the model is capable of.
So let's see some of the fun demos people have been making with Haiku 5 .5. Multiple people from my community have been throwing together video demos because everyone knows Opus 5 .5, the model got way better at programmatically creating video. And this example came from Arsh in the community.
This is insanely impressive for the money. Goddamn.
yeah i am very impressed and this cost 60 cents an api cost to make was built in under 20 minutes that's insane a lot of why haiku could make a video that good is that the anthropic models have pretty good like taste built in is the best i can put it especially around design stuff i think they just have a lot of like really good design data and have obsessively trimmed that data set down to give the model good reference points And you can see that even more so with the front -end design capabilities using Witch AI by our friend Dara.
This is so comically better than Grok. Oh, that's cool, actually. The little animation for how that text came in.
I haven't seen anything like that in one of these demos before, actually.
It's different, but it's cool. It's like a new thing. This is the blueprint -y one.
I don't love this diagram. It's cool. It has hover behaviors and changes things underneath.
Not my favorite.
This is kind of cute. Not my favorite. I like the fonts it shows.
It's not picking cringe fonts as much as models used to. And this one's fine.
If we turn off the design skill, how much worse is it? Yeah, now we're back into classic LM design. It seems like the new Claude design skill helps their models a shitload right now.
yeah the gap to that to this is hilarious i had uninstalled the design skill i definitely need to bring it back because it is it is much better now than it was before these are great and just like for comparison we can look at grok for seven or the design skill off it's even worse my god like this one still kills me i'm still very very amused by this particular one Oh, man.
I'm sorry. I have to laugh a little, okay? Someone in chat said that they legitimately think they might have made better sites in raw HTML in elementary school than what they were seeing from Grok there.
I don't disagree. Regardless, for the people who are picking Haiku as their model for cost -sensitive reasons, or you're using it in the $8 plan with OpenCode and such, this is usable. It's not great.
Like I'm not going to sit here and pretend that this is all you'll ever need doing front end design and building applications. But like a model this cheap being decent is impressive. And it means like the $20 cloud sub might be usable for code.
I'm not going to spin one up to test it. I'm not getting more cloud accounts, even for experiments. I think I have too many already if I'm being real with y 'all.
But regardless of all that, good designs. Out of curiosity, I'm throwing this model at one of my favorite tasks, which is auditing my open pull requests. T3 code has over 1 ,500 PRs right now.
Actually, it's just under 1 ,500 PRs. It's probably 1 ,500 by the time I finish the sentence because we have so many coming in at all times. So I have Haiku 5 .5 auditing all of them.
Specifically in this case, it is addressing all the ones that have performance issues, like PRs that are fixing performance -related stuff. And I have another thread here where it's just going to... We'll go through and categorize all of them.
Essentially told to use six soul for sub -agents and then corrected it later. You get the idea, though. This is going to go through all my PRs and make me a nice HTML page showing them when it's done.
While we wait for that, though, I have my favorite demo. We got a new build, a fish slop. Oof.
This one has some problems. First off, no mouse move. Only moving with Space Shift and WASD.
You gotta turn. No sound either. I don't know if that's because I haven't muted.
Nope, it's just no sound.
Yeah, this is a little rough. This is roughly the same tier as the Grok version, which was offensively bad. I will say for this one, other than the lack of mouse move, there is no...
like a just like a straight up incorrect offensively bad thing so for comparison with um gbd6 astra the demo it made was stunning but it got the like core movement wrong like the mouse movement was way too slow and it felt awful and the game plays like core loop was a little screwy and there were like pretty apparent bugs in it this version doesn't seem to have like like other than the mouse move thing where it just doesn't work at all nothing else is like immediately like this is just fucking wrong which is good it's not it's not making bad decisions it's just not making impressive ones oh yeah the corners are fucked that's a good point actually yeah the way this corner works is just bad okay yeah i'm taking back what i was saying there this does have some fucky stuff it's subtle but it has it also the timing of the second alien attack made no sense it shouldn't have been there that early it is not honoring the original source the way i was hoping it would
It's not as immediately offensively pathetic as a lot of other models do things. Its lows aren't as low as Astra's, but its highs are a tenth as high as Astra's. I think that's the simplest way I can put it.
That build of Fish Slop was about a dollar of usage total. It was a shitload of cash read tokens, 9 .73 million of them, a bunch of one -hour cash writes, because it always writes one hour, which is annoying. New inputs, apparently only 102 tokens because it was caching constantly, and the output tokens were about 30 cents.
47 of 51 requests exceeded 100k input tokens, so they used Haiku's higher price tier. If I had capped the context, the game probably would have come out much worse, but it would have been way cheaper. Look at that.
The first PR review is in. They found 40 performance -related open PRs. Most are server or mobile fixes.
Only a few are clean and high impact. Top of the list below is merge ready because GitHub reports clean. So it thinks that they're ready just because GitHub said so.
It didn't seem to audit them thoroughly itself. Replaying a command no longer freezes the server. Retry went from 18 seconds plus to six milliseconds.
I'll do what I usually do here, which is I open up Opus. I paste the PR and I say something along the lines of, this seems like it is a change worth landing. Can you figure out if that's the case?
Fix anything that you think needs to be fixed and get this merged if you're confident it's ready to go. Full send. And now either that PR or something like it will be merged momentarily.
The full send, full stop combo has changed how I build. I'm so happy with it. While I wait for this PR summary view to generate, I want to answer one of the questions I see happening a lot in my chat.
Would I use this model for code? No. I want to be clear about this because I think you guys are thinking about these things wrong in general.
coding is not what makes llms expensive it's not the process of going from a prompt to a javascript file you can execute that is the majority of your costs your costs come from everything else largely like once the context is in the model and it's cached generating the right code file and putting it in the right place is relatively cheap getting the context needed to do that is expensive Verifying that it worked is expensive.
Writing a shitload of code to test all the edges around it to hook into GitHub so that you can monitor the PR when it's up and all these other things, those are expensive. Regenerating it eight times because you got it wrong the first seven, that's expensive. Those things are expensive.
But the act of going from all of the context and details needed to write code to the code file in your code base, even on the most expensive fable, that is a few pennies. it is not that expensive to generate 500 output tokens once the input tokens are cached what i'm trying to say here is i don't get why people are looking for a dumb model to do the code part the code is cheap the expensive part is all of the prep before you write the code file and all the verification after and haiku is a tool your models can use to do those parts more cheaply is valuable and if that middle part the coding has edges around it like you want to try five versions of the thing and you have a verification system where you know if it worked or not trivially where you can run code to know haiku generating five or eight versions to test could be worth it but for most problems the smarter model writing the code is not that expensive and is preferred in general When the model is unsure about a theory and it wants to vet four things to confirm it or wants to go look at eight apps in the world and like random open source projects to compare against or read through a bunch of documentation to figure out some facts about some idea you have.
That is when haiku is incredible. But the execution of the plan. Sure, if you write a really thorough, detailed plan with a really smart model, a dumb model might be able to implement parts of it.
Cool. Awesome. You already spent the money on the plan, though.
Oof. This is an L for me. I hit a rate limit on GitHub, probably my fault because I was doing a lot of these types of things across a lot of data.
But the script saved the error body as though it was API data and then failed because it would hit errors when it ran JQ on that. The parallel per PR fetches that this run did used most of my 5 ,000 requests per hour. So now I'm locked out of T3 code GitHub integrations for a bit.
the pr dash star .json files parsed find an aggregate so their data is valid now it's waiting for the quota to reset then it will refetch the 181 prs and fix the script so it rejects responses that aren't json arrays then it will recount ci and review cool except for the fact that it didn't trigger anything to spin itself back up it said that it's going to wait until the rate limit's up but it didn't trigger a wait we would show that in the ui if it did so this thread is just done until i tell it to keep going So yeah, this is why I don't like dumb models.
Small models will corner themselves like this, make something harder for not just themselves, but for me. Like this just ruined my T3 code experience for at least the next 50 minutes until my quota gets reset. So I can go back to like doing PRs inside of T3 code.
But since my rate limits are fucked anyways, let's show off how I would actually do this. I took the same prompt from before, but I changed two things. First, I told it to use lots of subagents and workflows to review these PRs and to use Haiku 5 .5 for all subagents throughout the work.
Instead of one more sentence, minimize work that you do yourself outside of actually getting the data and orchestrating the subagents. And now the other change, I'm going to run this on Opus because unlike Haiku, Opus is a smart model and it should not only be able to not hit my rate limits, it should find ways to work around them too.
Here is where it's going to hit those rate limits and it's probably going to find a way to work around it. here it's here's opus being smart 1495 openpr is a huge number to tackle of sub -agents maybe batches of around 15 prs per haiku agent will be roughly 100 agents running beyond the typical small scale workflow guidelines but the user explicitly wants quote lots of sub -agents look at that it's actually being thoughtful this is going to take longer than i plan on sitting here to show you guys but uh you get the idea I think that Haiku is incredibly useful as a tool that you insert into your code and that your agents can use as well.
But I would not select this model as a tool I use in my coding agents because I'm on the $200 tier and I would rather just use Opus. And as big as the price gap is, you have to remember that when a model can't do something and it ends up doing lots of things in the process, it might end up using more money. because if opus can solve a problem in 100k tokens and haiku burns a million before it figures out it can't haiku ends up more expensive so for a lot of work not only is haiku not able to do it haiku will cost more money as it fails to do it and as things have continued to improve across the industry in particular cost performance with both anthropic and open ai we've ended up in a fun place where the price gap between opus 5 low and haiku 5 5 max is not as big as you would think it's 20 cents to 55 cents so it's around 2x more money for opus if you're willing to deal with low also we got the rest of the data for haiku 5 5 because they finally finished running it which makes soul and luna still quite compelling with the pareto line but uh it does fill out for the anthropic lineup quite well here where have all of haiku for the cheap
and most of opus for the less cheap this is a compelling line if the only models you used were opus and haiku you're probably not missing out on a whole lot no sonnet i showed you guys sonnet just doesn't offer much with the pareto line actually it offers nothing now maybe just barely on high there interesting but yeah i i think you can avoid sonnet entirely and just use opus and haiku and have a very very good time with your anthropic models but that leaves us with a question when should you use soul Well, first off, as you see here, Sol is actually quite a good bridge from Haiku over to Opus.
It fits in the middle there. But that's not how I would think of Sol. I wouldn't think of Sol as for things that Haiku is too dumb for and Opus is too expensive for.
I think of Sol as a different family with similar intelligence. Imagine that you have a team of people that work on the same things with you every day, and you have a really hard problem you have to solve. And you ask everyone around you and you all, for the most part, agree on what the path is.
That makes sense because you guys work together every day. So you've started to think about things in a similar way. Sol thinks about things in a different way.
It's like you're asking your friend who works at a different company for their thoughts. And it's a unique perspective. And I found OpenAI models bring a unique perspective in particular with their thoroughness auditing code.
I love 6 .1 Sol as an auditor. I have set up skills for all my agents where they use Opus as the thing that does the coding. And then they consult with Sol for opinions before they put up the PR.
And I'm landing code so much faster simply because Sol catches all the dumb things Opus might miss. And by the time the PR is up, it's already been scrubbed by Opus and by Sol. So I personally am sticking to my Opus and Sol duo.
But when Opus needs to go collect a bunch of data or read through tons of shit or just categorize things, haiku is going to be an incredible addition to my portfolio of models i won't be the one to pick it but i trust opus to do a good enough job picking it at the times where it makes sense and i think the result of all of this is a very compelling set of models this is going to take too long i am curious what the results are but uh think of that all i have to on this model release is a useful tool it's a nice change to see anthropic making small models that are actually usable and useful for real world stuff and it's still honestly a bit trippy seeing an anthropic model so far to the left on the artificial analysis cost versus intelligence right there neck and neck with luna actually i mentioned earlier the like pareto line stuff we only had x high and max for haiku at the time now they have the rest it's so close to luna that this no longer feels like a big missing piece
of the Anthropic portfolio. And with that, I will go back to my opening statement. Anthropic has no small models that are worth using right now.
Anthropic now has one small model that is sometimes worth using. And that is a very nice change. And with that, all I have left to say is that I can't wait for Fable 5 .5.
I understand that they're probably not going to put it out for a while because it kind of is outside of the spirit of pacing the frontier. But I'm so happy with this 5 -5 lineup right now that I'm really hopeful Fable meets the bar they have set with model releases like Opus, Sonnet, and Haiku as part of this line. Anthropic is dominating across the field now.
And if I was OpenAI, I would be very, very scared. And I would be praying that the next training run for Astra comes out unbelievable. Because right now I'm honestly struggling a bit to use my Codex subs because my Anthropic and Claude subs are just so much more useful.
Let me know if you guys agree if I'm over -exaggerating this gap or if it really does feel that way to y 'all. Until next time, peace nerds.
The Hook
The bait, then the rug-pull.
The opening is a setup disguised as a compliment. Anthropic is good at big models, and that is exactly why its cheap ones kept arriving as bad distillations of something better, to the point this reviewer had rewritten every agent config he owns to forbid Haiku 4.5 outright.
Frameworks
Named ideas worth stealing.
21:55model
Orchestrator plus cheap subagents
Flagship model plans and holds the goal
Cheap subagents fan out in parallel on bounded tasks
Each subagent result is verified mechanically, not trusted
Flagship writes the final change itself
Keep the expensive model as the single decision-maker and push search, reading, categorization and hypothesis-checking out to many cheap instances running at once.
Steal forany agent pipeline where most of the runtime is reading and searching rather than writing
23:18concept
More attempts versus smarter attempts
When the gap between a dumb attempt and a smart attempt is small, and you can verify an attempt cheaply, buying ten cheap attempts beats buying one careful one. When verification is expensive or the gap is wide, the cheap model loses on both quality and total cost.
Steal fordeciding which tasks to route to a cheap model instead of guessing by task name
31:52model
Where the money actually goes
Gathering the context: expensive
Generating the code from cached context: cheap
Verifying the result: expensive
Regenerating after a wrong answer: most expensive of all
A cost breakdown of agentic coding that inverts the usual intuition. The generation step everyone focuses on is a rounding error once context is cached.
Steal fordeciding where to spend on model quality inside your own pipeline
06:39concept
The 100,000-token price cliff
Haiku 5.5's cheap rates apply only under 100,000 tokens of context, and crossing that line multiplies every rate by five. The threshold doubles as product design: it makes short classification calls almost free while discouraging long agentic sessions on the cheap tier.
Steal forsetting hard context budgets in any system billed on tiered pricing
13:52concept
The Pareto line read
Plot every model and reasoning level on cost per task against a capability index, then draw the best-available-score curve. Anything beneath the curve is dominated and should be dropped from the stack, which is how a vendor's full lineup collapses to two or three real choices.
Steal forpruning a model roster instead of adding to it
37:55concept
Second opinion from another family
Rather than treating a rival lab's model as a middle price tier, treat it as a colleague from a different company. Similar intelligence, different training, different blind spots, which makes it most valuable as an auditor of the flagship's work rather than a replacement for it.
Steal fora two-vendor review loop before shipping generated code
CTA Breakdown
How they asked for the click.
VERBAL ASK
41:02link
“Let me know if you guys agree if I'm over-exaggerating this gap or if it really does feel that way to y'all.”
No subscribe ask at all. The only hard pitch is the sponsor read at 2:17, closed with a tracked short link, and the ending swaps a growth CTA for a genuine opinion question aimed at the comment section.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Theo breaks down how Anthropic silently modified prompts, rewrote its system card, and built invisible safeguards into its most capable model - then got caught.
A live 14-minute breakdown of the US government export control directive that forced Anthropic to pull Fable 5 and Mythos 5 offline for all non-US citizens — including Anthropic's own employees.
A 33-minute first-take from a developer who spent $3,000 on inference in 24 hours — benchmarks, real demos, session math, and the hidden safety intervention that silently degrades the model without telling you.