Theo reads Anthropic's own account of a two-week, 3,000-PR performance sprint and finds a caching bug it doesn't mention.
Posted
yesterday
Duration
Format
Reaction
educational
Views
12.3K
464 likes
57 · 43
Big Idea
The argument in one line.
Anthropic cut claude.ai's load times by two-thirds not by hand-optimizing code, but by giving Claude dozens of deterministic ways to measure performance, because once something is countable, an AI agent can optimize it.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
A developer running AI coding agents who wants a concrete playbook for turning vague 'make it faster' requests into measurable, agent-tractable benchmarks.
An engineering lead curious how a real company ran a two-week, 3,000-PR AI-driven performance sprint without a single rollback.
Someone building their own Claude Code or Claude-in-Slack workflow who wants real prompt examples for pushing an agent past its default caution.
SKIP IF…
You want a hands-on, step-by-step performance tutorial; this is a reaction video, not a walkthrough.
You're looking for uncritical coverage of Anthropic; Theo also catches a live caching bug their own fix introduced.
TL;DR
The full version, fast.
Anthropic spent a two-week sprint making claude.ai and the Claude Code desktop app roughly three times faster, and the post explains how. The core move was giving Claude ways to measure everything: instruction counts under Valgrind, React commit counts, layout-shift APIs, even a deterministic 120Hz frame budget, so an agent that can count something can optimize it. Claude ran in a shared Slack channel (Claude Tag), opened hundreds of narrow threads, and merged over 3,000 changes behind feature flags with human review and ratcheting CI benchmarks as guardrails. Theo reads the post live, tests the real site, and finds a caching regression the article never mentions: claude.ai still won't refresh its sidebar when data changes elsewhere.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Theo establishes his performance-engineering background and explains why an Anthropic post about making Claude 3x faster is personally exciting to him.
01:30 – 05:30
02 · WorkOS sponsor read
Sponsor segment for WorkOS and the AuthMD standard for letting AI agents sign up for apps on a user's behalf.
05:30 – 08:00
03 · The baseline numbers
Theo reads the article's opening stats: time-to-typable-page went from 3.1s to 0.55s, and the team scoped the sprint to the four journeys covering 95% of usage.
08:00 – 19:50
04 · Live-testing claude.ai's caching bugs
Theo tests the real claude.ai site, confirms it feels faster, then finds the sidebar never invalidates its IndexedDB cache: deleted threads keep reappearing after refresh.
19:50 – 24:10
05 · The brief: Claude's standing orders
Anthropic's actual Slack prompt for the performance channel, the Datadog MCP analysis that found the four key journeys, and the one real safety guardrail Claude hit during the sprint.
24:10 – 30:30
06 · Anything can be hill climbed
The 20-project sprint plan, Claude's wildly wrong multi-week time estimate for work that shipped in days, and Theo's own Lakebed/Rusty-V8 story about prompting an agent to 'boil the ocean.'
30:30 – 35:40
07 · The loop, thread by thread
How Anthropic validated new benchmarks with skepticism, used Valgrind to fix a megamorphic-lookup bug, and checked in CI ratchets to lock in every win.
35:40 – 41:00
08 · Scaling horizontally
Parallel threads uncover a 6,900-hook React census, a stray CSS selector costing 24ms per DOM change, a hidden location.reload bug, and the em-dash bug that froze syntax highlighting for 0.35 seconds.
41:00 – 46:15
09 · Guardrails
Feature flags, the brittle static-composer handoff, and a layout-shift bug that traced all the way back to Chrome's own new-tab footer and speculative pre-rendering.
46:15 – 56:50
10 · Steering: ambition, taste, direction
How Anthropic pushed Claude to be bolder than its cautious default, ruled on UX taste calls like streamed tables, and rejected a 900-line PR for a 2-millisecond win.
56:50 – 1:04:10
11 · An 8-millisecond budget
Claude steps through streamed replies frame by frame against a deterministic 120Hz, 8.3ms budget, landing nearly 60 merged PRs out of one Slack thread.
1:04:10 – 1:10:29
12 · What's next, and Theo's verdict
Anthropic's closing numbers and the standing question 'what's next,' followed by Theo's take on what this means for working with AI agents on real engineering problems.
Atomic Insights
Lines worth screenshotting.
Anthropic made claude.ai and the Claude Code desktop app roughly three times faster in a two-week sprint, merging over 3,000 PRs without a single customer-facing rollback.
The core principle: once Claude can measure something, it can optimize it, so the highest-leverage work was building more ways to measure, not writing more fixes.
Wall-clock milliseconds are too noisy for CI gates, so Anthropic had Claude track deterministic proxies instead: CPU instruction counts under Valgrind, React commit counts, and DOM mutation counts.
A page highlighting a finished code block could freeze for 0.35 seconds because any em dash or curly quote forced V8 to store the whole string as slower 2-byte UTF-16 instead of 1-byte Latin-1.
A single stray CSS :has() selector added 24 milliseconds to every DOM change on the page.
One leftover code path triggered a hidden location.reload, causing half a million invisible page reloads a day that no load-time metric could see.
A React hook census found 6,900 hooks and 900 store subscriptions firing on every keystroke in the message composer.
Every new benchmark got a human-reviewed CI ratchet: a ceiling that could only go down, never up, so a win couldn't silently regress later.
Claude estimated a full rewrite project would take 25 working days; a comparable overhaul actually shipped in about three, the same overconfidence problem that makes AI agents bad at their own time estimates.
Telling Claude 'I'm open to wacky ideas, don't be afraid to boil the ocean' pulled a stuck optimization effort out of its safe default path and into a genuinely new custom runtime design.
Every user-visible change shipped behind a feature flag with a named human owner; across the sprint they introduced roughly 200 flags and had already cleaned up more than half by the end.
A teammate's layout-shift bug traced back to Chrome itself: the browser's own new-tab footer disappears when you navigate, making the page 56 pixels taller 100 milliseconds after first paint.
At 120Hz, a frame only gets 8.3 milliseconds; stepping through a streaming reply frame by frame cut one hot path's main-thread blocking from 750ms to 200ms and landed sustained 120fps.
The caching fix that made the sidebar load instantly also broke it: delete a thread in one tab and it silently reappears in another, because nothing invalidates the local cache.
Takeaway
Measure first, then let the agent climb
WHAT TO LEARN
The whole sprint worked because almost every optimization target was turned into a number an agent could move, and the discipline came from reviewing every change a human could actually see.
03The baseline numbers
Anthropic measured before changing anything: time to a typable page went from 3.1 seconds to 0.55 at the 75th percentile, and loading existing conversations dropped from nearly 3 seconds to under 1.
They scoped the whole sprint to the four user journeys that made up 95% of actual usage, rather than chasing every possible slow path.
The team estimates the improvements save tens of thousands of hours of user waiting every single day, which is why a two-week sprint got this much attention.
04Live-testing claude.ai's caching bugs
Claude.ai now caches recent threads in IndexedDB so the sidebar loads instantly on refresh, but it never invalidates that cache when data changes elsewhere.
Deleting a thread in one browser tab left it still showing as active in another tab indefinitely, because the client trusted its local cache over the real server state.
Speed and correctness are separate problems: a cache that loads fast but shows stale data is often worse for trust than a slower cache that's accurate.
05The brief: Claude's standing orders
The sprint started with a Slack channel and a loose standing instruction telling Claude to monitor deploys, maintain dashboards, and fix low-hanging fruit on its own.
Anthropic used the Datadog MCP server to have Claude analyze real usage data and identify the four highest-impact journeys before writing any code.
Establishing comparable baselines mattered more than people expect: each measurement had to start at the same user interaction and end at the same rendered result, every time.
The one guardrail Claude actually hit during the whole sprint was a built-in safety check on a destructive remove command, not anything related to browser debugging.
06Anything can be hill climbed
Claude estimated a full codebase rewrite would take about 25 working days; a comparable overhaul actually shipped in roughly three days once an engineer just pushed it forward.
AI agents are consistently bad at estimating how long their own work will take, because their training doesn't give them an accurate model of their own capability.
Letting Claude work asynchronously for hours or overnight, instead of waiting for field data on every change, let the sprint iterate faster than Anthropic's normal deploy cadence.
A vague prompt like 'I'm open to wacky ideas, don't be afraid to boil the ocean' pushed Claude out of its default cautious path and toward genuinely novel solutions.
07The loop, thread by thread
Every new benchmark was treated with skepticism: it had to move in the lab and correlate with real wall-clock latency, or it got thrown out rather than let Claude optimize the wrong thing.
Claude profiled a hot path with Valgrind, found a quarter of its instructions were redundant lookups resolving the same message ID three times, and cut instructions by 48% in about an hour.
After a fix landed, the team checked in a new CI ratchet so the instruction count could never silently creep back up on a future PR.
The actual loop was: open a thread, build a benchmark, ship a PR behind a flag, watch the real deploy data, then lock in the win or roll it back.
08Scaling horizontally
Once one thread proved an end-to-end approach for finding a class of bug, opening five or fifty more threads in parallel was just a matter of starting them.
A React hook census found 6,900 hooks and 900 store subscriptions firing on every keystroke inside the message composer, a problem no one had measured before.
Claude found a leftover location.reload path causing half a million invisible page reloads a day, plus a bug cloning the same cache snapshot into IndexedDB twice a minute on the main thread.
A page could freeze for 0.35 seconds highlighting code because any em dash or curly quote forced the whole string into a slower 2-byte UTF-16 storage format in V8.
09Guardrails
Every user-visible change shipped behind a feature flag reviewed by a human; across the sprint Anthropic introduced around 200 flags and had already retired more than half by the end.
The static composer trick, showing a fake HTML input before React loads, is brittle by design, so Claude built dozens of pixel-alignment tests to guarantee it never drifts from the real render.
A layout-shift bug that no internal metric caught turned out to be Chrome itself: the browser's own new-tab footer disappears 100 milliseconds after the page loads, shifting content underneath it.
High-risk changes rolled out in stages, employees first, then 1% of users, then everyone, which is how the team caught the Chrome footer bug before it reached the general public.
10Steering: ambition, taste, direction
Claude defaults to being conservative about scope, so Anthropic had to explicitly tell it to be bolder and that it had permission to keep pushing once a goal was stated.
Every thread had one named human owner who ruled on perceptible UX tradeoffs, like whether a table should fill in cell by cell or wait until the whole row is done.
Keeping each thread narrow, focused on one benchmark or journey, mattered more than giving Claude broad freedom; the team rejected a 900-line PR because its 2-millisecond win wasn't worth the added complexity.
11An 8-millisecond budget
At 120Hz, each frame has only 8.3 milliseconds to render, and once that budget became something Claude could measure frame by frame, it became something Claude could optimize.
Moving code-fence tokenization to a worker thread and memoizing finished message blocks cut one long-reply hot path's main-thread blocking from 750 milliseconds to 200, and held a sustained 120fps.
A single Slack thread about frame timing alone produced nearly 60 merged pull requests, showing how far one well-instrumented measurement can be hill-climbed once it exists.
12What's next, and Theo's verdict
Even after hitting every original target, Anthropic kept pushing: 'the targets are not the stopping point, what's next' became the standing question for the rest of the sprint.
Claude.ai and its desktop app are now roughly three times faster than in early August, with CI ratchets in place to keep the gains from eroding over time.
The team is explicit that the 95th percentile, longer journeys, and very long conversations still have real room to improve, so the work isn't framed as finished.
Glossary
Terms worth knowing.
Claude Tag
Anthropic's internal Slack bot that runs Claude directly inside a Slack channel, used for the entire performance sprint instead of Claude Code or a terminal.
Ratchet
A CI check that only allows a tracked metric, like an instruction count, to improve or hold steady; any pull request that makes it worse fails automatically.
Instruction count
The number of CPU instructions a piece of code executes, measured with tools like Valgrind; used as a deterministic stand-in for wall-clock speed because timing itself is too noisy to gate on.
Megamorphic lookup
A slow property-access pattern in the V8 JavaScript engine that happens when the same object shape is accessed too many different ways, preventing the engine's usual fast-path optimizations.
Static composer
An HTML copy of the message input box served before React finishes loading, so users can start typing immediately; React then has to take over the exact pixel position without any visible jump.
Layout Instability API
A browser API that reports exactly which on-page elements shifted position and by how much, more precise than the aggregate Cumulative Layout Shift score most sites monitor.
V8 code cache
A precompiled version of JavaScript bytecode saved to disk so an app's main process doesn't have to re-parse and recompile the same code from scratch on every launch.
“Once Claude can measure something, it can make it faster, so we kept finding more things to measure.”
the thesis of the whole video in one line→ newsletter pull-quote↗ Tweet quote
40:40
“The brief found the culprit: em dashes. If a reply's markdown contained any non-Latin-1 character, V8 would store the entire string as UTF-16, putting every syntax highlighting regex on a slower 2-byte path.”
absurd, specific root-cause reveal→ TikTok hook↗ Tweet quote
00:07
“I love nerding out about the details of good quality engineering, data loading patterns, performance issues, browser stuff, all those types of things.”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
I know most of y 'all think of me as that AI dev vibe coding YouTuber, but believe it or not, a few of you were around back when I was more focused on full stack application development, in particular the website, and really focused on performance and end -to -end capabilities with systems. I love that type of stuff. I love nerding out about the details of good quality engineering, data loading patterns, performance issues, browser stuff, all those types of things.
That's why I'm so excited about this new article that Anthropic just put out about how they made Cloud .ai three times faster in two weeks. largely by using clot. This is a very fun one for me for a bunch of reasons.
Not only am I a performance nerd, I've also been dunking on Anthropix performance in general engineering quality for a long time now. It seems like Anthropix models are now finally good enough that they're actually able to help with performance instead of hurting it, which is particularly fun because I posted a video not that long ago about performance issues in T3 code that were caused by Fable that I had to go solve myself.
So I'm very excited to read this because one of the... things that's made t3 code so awesome has been that focus on performance and if anthropic is finally going to stop using 18 gigs of ram when i open cloud desktop maybe it's more worth a look i also just want to see how they're setting cloud up for success in these types of environments where you're trying to debug deep performance issues it's a not it's a non -trivial thing to do and a lot of the tools that exist i'll be real kind of suck.
So I'm very curious how Anthropic was able to get the profiling set up such that they could make the desktop app as well as the website perform meaningfully better all by using Claude. I do have one more thing worth using, though. And while it might not be three times faster, it's still worth a listen.
It's today's sponsor. Weird question for you. You want people like me to use your product?
Because if you do, you probably shouldn't do a normal sign -in flow. Because I'm not going to be clicking those buttons myself nowadays. I do most of my work through my agents now, which means any platform that my agents choose needs to be usable by the agent.
In particular, it needs to be able to be signed up for by your agents. Today's sponsor is WorkOS, and that might confuse you because they are a platform for user signing in, right? Well, they are.
They're the best way to have humans and businesses signing into your app. So if you're trying to get enterprise customers, then you should probably be using WorkOS.
That's why everyone from OpenAI and Anthropic to FAL, Cursor, Bolt, Vanta, Carta, and more are all using them. But I want to talk about AuthMD because this is a super, super cool standard. The goal of AuthMD is to make it easier for agents to sign up for your apps.
It's a new standard that's been designed by WorkOS in partnership with companies like Firecrawl and Cloudflare, and the results speak for themselves. There's already a ton of companies adopting these standards like Neon, Monday, Parallel, and more. What makes it so cool is that agents can now register on behalf of a user.
So if you're trying to get business customers and you want their agents signing up, this is a great place to start. Get AuthRight for all your users at soydiv .link slash WorkOS. Time to dive into how Anthropic made Claude three times faster in two weeks.
My honest guess is that the models are much smarter now, so the quality of Anthropic's engineering is actually improving for the first time in a while, but super curious what tooling they set up and what they gave Claude access to in order to make this possible. Once Claude can measure something, it can make it faster, so we kept finding more things to measure.
I like this opening, although I am concerned this might be the new Claude pros, because they did a lot of work to make Claude less cringe. But if this becomes the new cloud format, there will be a new cringe sentence structure.
It got some nice little animations on the site too. It's cute. But let's dive into the actual details.
This August, we made the core user experience of cloud .ai in the cloud desktop app around three times faster in a two week sprint. Users have been telling us it was slow and they were right. Hi, it's me.
I'm users. Yeah, it was bad. I did absolutely notice the difference here when I was going to cloud .ai mostly to like check the status of my account and my usage before I had my dashboard set up.
And I noticed the day it got way faster. It was like a meaningful difference. I didn't think I had a post about it.
While that's hunting, we could keep diving in. We focused on four journeys that made up 95 % of user activity. At the 75th percentile, time to a typable page on a fresh load for Cloud went from 3 .1 seconds to 0 .55.
And starting a new Cloud Code session went from 0 .8 seconds to 0 .3. And loading Cloud Cowork Cloud sessions went from almost three seconds to under one second. In aggregate, we estimate that this saves tens of thousands of user hours of waiting every single day.
This is why it's important. Oh, that's funny. I actually noticed improvements in June.
So they definitely... made a bunch of improvements at the time i even called out that it felt faster on a mobile hotspot than it used to feel on gigabit network connections i also found when i dug deeper into it that it was using tan stack now which was really cool to see specifically tan stack router makes it much easier to handle these types of navigation cases but that makes me even more curious what they did because they already did the framework upgrade so how do they get this performance squeezed out after let's dive into the core user journeys there's launching the app They had both the web and desktop versions improved here.
It was, as mentioned before, three seconds to launch the website before. Now it's 550 milliseconds. Desktop app used to be over six seconds, and now it's just over three.
Starting conversations is another thing they focused a lot on, and they were able to make huge improvements there. Loading conversations, they made huge improvements in. And sending a message, specifically Cloud Cowork, went from a second to 48 milliseconds.
And Cloud Code on desktop went from 250 milliseconds to 52. This is all in the apps, by the way, not the terminal version, but the Cloud Code desktop app and the Cloud Code website. I am actually curious how much we feel the difference.
That's still slow as balls. Still waiting. Still waiting.
At this point, I think it was a CDN failure. Content at Cloud AI login may not load or link to file. Great.
Refresh. The refresh worked. I guess those performance improvements didn't happen on the signed out page, but now we're good.
So let's take a look. I'm going to do a refresh. That was pretty fast.
Some test prompt. This is one of the ones that pissed me off the most is the flow of submitting a prompt would create the new thread, but it wouldn't navigate, which meant that if you refresh, there was a chance it would disappear forever. That came through very, very quickly, actually.
Let's do another. And it means immediately refresh. Ah, did they fix that?
No, that thread disappeared. Yeah, here's that bug. I hit it again.
This thread that I have open right now did not make it to my recent threads. So I have this one from before. This one doesn't exist.
Make the title proof of my theory. I'm going to send this and then refresh really quick. It's hard with one hand.
Cool. Did that. Said what to make the title.
getting set up for the session it's really thinking on this but i did manage to kill it getting added to the sidebar at the very least here it had a failed tool call yeah this thread effectively doesn't exist because i refreshed before it would get persisted to the sidebar which is uh yeah rough it does let you submit the prompt pretty quick though like if i refresh okay i tried pasting that whole time so watch this i'm gonna refresh and then start spamming paste it took about 700 milliseconds or so a bit under a second maybe about a second actually oh that one was brutal that one took like two to three seconds while it's still not meeting my bar i am impressed overall with the difference it does feel a lot nicer to navigate i hope they can figure out oh look there my threads finally appeared took them a bit but they did but yeah you get my point it has a lot of like data persistence layer issues still but it does feel better I see you all in chat saying they need convex.
Unironically, yes, it helps with so many of these things. I don't know how well it would handle the insane load that I'm sure that Anthropic pushes, but something similar to it with a proper engine that keeps stuff up to date would help a lot. A few months ago, I did a deep dive on how both ChatGPT and Cloud .ai were handling caching for the data that was being fetched on the page that you actually saw and learned that both are now caching the recent threads in the sidebar through local storage, or I guess now in this case, now I'm digging into it.
IndexedDB is a way of using a key value store to store the threads you have in the sidebar. All these threads here are all cached when you load, which is why when I refresh, they come in almost immediately and there's no pop -in or flash -in. To show you a really silly example of why this matters, I'm going to go to my other browser and archive a handful of these.
I guess there's no real archive, so I just have to delete. That's fine. Delete.
Delete. Cool. So I just deleted those two.
So there's still one test prompt. Watch what happens when I refresh. They're not invalidating the cache.
It still thinks these two threads exist when they don't. How do I get this to refresh? If I click, that doesn't even do it.
What? I was hoping this was going to be a video about how much Anthropix improved their engineering quality. I'm almost relieved to see nothing has changed.
What the fuck? I have these two threads that when I click, you get an error, but it never revalidates the sidebar. You know what?
I bet. I bet if I send a prompt here, that will... Nope, that updates the status for that one, but these two are still broken.
Chat just said, coding is solved, by the way. Yep. Yep, coding solved.
AGI is here. No, they didn't disappear. They're still here.
Two broken fucking threads. Are you kidding? I hope you'll understand why it's a convex shell now.
It solves all of these types of problems. Even T3 code doesn't do this. We don't even have a server to sync against.
It's all on your devices. Oh, they disappear when I was complaining? They just poof, vanished all of a sudden?
I wasn't even looking? God. Reminder that according to Sam Altman, I'm an anthropic shill, so...
Yeah, take it as you will.
Well, that answers a bunch of my questions. They made performance better by making it so that things never get updated. Clever.
They're caching far too aggressively, and they're not invalidating the cache anywhere near often enough. For reference, in T3 code... we do this a bit differently when we're in the loading state because we haven't fetched the new data from the back end yet i show it in a faded blurry state because i don't want to show you confidence in data that isn't accurate yet the fade doesn't trigger until we have loaded the real data from the server did i say t3 code i'm sorry t3 chat my bad you guys know what i'm talking about t3 code does everything through servers that you're running yourself so it's very different here the caching Caching just works fundamentally different when there's one database per user and it's on your own machine.
For D3Chat, you need to do more. It is incredible as Convex is, and it absolutely is. I'm so happy we moved our data over there.
It still doesn't give you data as soon as the page loads. I want data as soon as the page loads, which is why we set this up with the local storage cache so that we can show that immediately. I did also notice, though, that we changed how we were loading in things on the bottom bar.
We need to start caching that as well. on that note actually let me um new thread t3 chat i noticed that it seems like we're not caching the necessary data in order to open the model picker and other things at the bottom of the composer as soon as the page loads it's kind of jarring to have the sidebar come in immediately but the model picker and other selectors in the chat take some time to load What would it look like to cache that similar to how we cache on the sidebar?
If it's relatively easy, I would like you to implement it and file a PR and babysit it. As much as I like to dunk on Anthropix's lack of engineering quality, I think that we should go through this article because at the very least, they did make a lot of improvements to performance. And while they seem to have regressed a bunch of data patterns in the process, we can still learn from the ways that they gave Claude access to the things it needs in order to find these performance issues and fix them.
let's start reading those i am actually excited apparently they did the whole thing through claude tag so they didn't even use claude code or i don't know something better like t3 code they did this all through claude tag which is the slack bot i have seen examples from boris the guy who started claude code using claude tag to do all of his work every day he actually showed an example of a real thread he did to add the mods feature to codex or to cloud code if i recall i could be wrong on that but he did something like that recently So they used CloudTag using an internal research model roughly comparable to Opus 5 .5.
Huh? I hope this is just a snapshot of what became 5 .5 and not some other model that they're claiming is similar to 5 .5, like they claimed Opus 5 was similar to Fable when it wasn't. But yeah, no guardrails though, which is a good call out from Maria because when they're testing the models, they have much less guardrails generally.
And I happen to know that when you start diving into Chrome internals, it becomes much more likely that you start to hit the safety guards and the stuff that prevents you from sending prompts because i think you might be hacking so yeah i will put that all out there because i think it's important because i've had that experience i am hopeful because again i haven't really pushed opus too hard here actually i'm kind of lying there i did have one instance let me find it in my terminal quick i vaguely remember you mentioning that the other agent hit some type of security flag or guardrail when it was working on the performance improvements you have more info on that like what it was trying to do at the time and what guardrail it hit it was a built -in safety check on remove commands that i hit so not as bad as i thought i assumed i was hitting the guard rails because of the weird things that opus was doing in chrome to debug performance stuff for fish slop it does appear to be the case that that isn't what happened here
So hopefully we won't hit too many of those guardrails. I've seen it before with the other models. I haven't done as much deep diving in performance stuff with clawed models lately.
Honestly, I have still found open AI models to be better at this type of, I still call it the like Rottweiler way of working where it grabs the problem by the neck and strangles it. I have found anthropic models to not be quite as good as open AI models in the GPT line for these. like deep dive tear apart all the performance problems type stuff but a lot of this does just come down to the tooling because codex has much better computer use stuff built in and it seems like it's able to figure out how to control chrome and get these traces out of it a little bit better bolts still aren't great at it for my experience though and i find myself still digging in to the details myself when i'm deep in performance stuff so this is all why i'm so curious how anthropic did this they said that through cloud tag they were able to find bottlenecks build benchmarks ship improvements and watch every deploy They steered by setting goals, making trade -offs, and improving every change.
With that approach, they merged more than 3 ,000 changes without a single customer -facing incident or rollback. To be fair, I did just kind of showcase things that might have been worthwhile, a rollback or an incident, but to each their own. This is the post that covers how they did it, apparently, safely.
Very interesting. Before the sprint, they created a new channel in Slack with the following standing instructions. At Claude.
Your job is to facilitate all things related to the performance of the cloud AI website and desktop app. Your responsibilities include monitoring deploys for performance regressions, assessing the accuracy and comprehensiveness of existing telemetry, maintaining well curated observability dashboards, proactively implementing solutions for observed issues and low hanging fruit, proposing performance project opportunities, and communicating with your human teammates.
The ultimate goal for this channel is for you to become as autonomous as possible. But today we know that isn't yet possible. This is a terrible prompt.
They didn't even mention to make no mistakes. No wonder it had those regressions. If they had just told it, make no mistakes, I would have no problems.
We asked Claude to analyze usage data through the Datadog MCP server. They got that dog in them. It identified the four highest impact user journeys, the app launches, starting conversations, loading existing conversations, and sending messages.
Between web and desktop, as well as across our products, those journeys came to 13 distinct... measurements. To establish our baselines, we added instrumentation until they were directly comparable.
Each started with a user interaction, ended once the result was rendered, and disambiguated client and server work. People underrate how important this part is to fixing performance issues. It is much harder to fix what you can't measure, even if you're not an AI agent.
I cannot tell you how many arguments I would get into in my previous jobs where I was like, this thing is slow as hell. They're like, well, we looked in our dashboard and only took 200 milliseconds. And then I went and profiled it myself and learned that they were counting from not like when the page was loaded to when the thing happened.
They were counting from like when their component rendered in the UI to when it went from loading state to done state. And if it took eight seconds to get there, then your trace doesn't matter. And so many places get their profiling wrong with all of these types of things.
I cannot. even comprehend how many times I have been shown data to prove me wrong about a performance issue, just to learn that the data was not actually reflecting the real world performance at all. Now let's see how they actually got the work done once they did the instrumentation.
We kicked off the sprint with a list of about 20 handpicked projects, each targeting specific journeys. Claude estimated the impact of each project in milliseconds, and we aggregated those estimates to set our targets for the sprint. Some of the projects were fairly large, but we thought we could probably achieve most of them within two weeks.
They hit 12 of their 13 targets by day three. Did they fall for the classic mistake? I was planning on having Opus do the classic rewrite of ping .gg to make it modern.
I'm curious about a thing quick though. How long do you think these changes would take if we did them all in one pass? Here we are, the classic anthropic overestimation.
As great as models are now at writing code and understanding code, their understanding of agents is still kind of garbage. And as such, when you ask them, how long will it take to do something? They're wrong.
This is an overhaul that I'm pretty sure I could do in a few hours. And it is very confident it would add up to about 25 working days to do it. It is just wrong, but it's cute that it thinks it isn't wrong.
I'm going to throw this one a rough instruction here. I want you to kick off a... agent run on my computer left book you should be able to see how to connect through the fleet repo on this machine i want you to make sure this repo is cloned all the environment variables that are needed are there and that you can run opus 5 .5 through cloud code in t3 code such that i can still see the thread how i normally do in t3 code my goal here is to do the whole implementation on that machine in one pass with however many subagents and whatever else are needed with full computer use capability on the other machine because I don't want you to be spamming my machine with computer use requests and browser control while I'm using it.
If anything does not work as expected when you move this workload over to Leftbook, don't hesitate to ask questions so we can get it working right. I'm actually curious how that work goes. We'll check into that in a bit, but yeah.
What I was trying to make by pulling this up is to highlight how bad models are at estimating timeframes because they don't understand how capable models are, which I honestly find hilarious. I'm just trying to do an extended version of the joke that Claude told them it would take two weeks and then it actually took three days.
Because these guys can be Claude. Makes sense they're bad at judging this too. As we were.
The planned projects landed early. For faster launches, we baked a static composer into the HTML so users can type during the React initialization. That's a fun one that we did a while back.
They also precompiled the V8 code cache so the desktop shell main process doesn't have to recompile it from scratch. Oh boy, they're going deep in the weeds on that one. Ugh.
Do I want to add that? You know what? Anthropic has been focused on overhauling the performance for Cloud .ai and the Cloud Code desktop app.
I noticed this paragraph in particular has some fun ideas on how they improve performance. I'm curious if any of these would make sense for us. both the embedding the html so the composer is available faster but more importantly the v8 cached compile i'm curious if we'd be able to implement something like that in t3 code and what type of performance improvements we would see if we did very curious how that goes another fun reacty change is that in order to make navigation faster they kept the composer mounted between conversations and they also prefetch sessions when the user hovered over them and in doing all of that also cut cyber re -renders by 90 these are the types of things i was doing by hand just barely a year ago in order to squeeze every bit of performance I could out of T3 chat.
Just some fun, real examples here. I'm just going to click a random thread so you can see how fast it loads. If that was not as good as I was hoping.
I go to a different one, skill manager name. Okay, that was immediate. That's immediate.
Yeah. Data fetching on Hover is one of the most fun little tricks you can do to make things feel... absurdly fast it's actually crazy how fast things will feel when you make a small change like that i'm surprised they hadn't done it before they also called out that they left room for claude to identify opportunities and propose new work streams those work streams quickly ramped into full projects of their own which far exceeded their initial targets so they set new targets and looked for more things to measure after we've ended up funding nearly every project in the original project list and more funding and completing here's the wording there Let's do a refresh.
What have we not explored and what can we hill climb on? Where's the most opportunity at this point? I'm open to wacky ideas.
Don't underestimate the value of that particular piece of the prompt. I know it sounds crazy. I'll be real.
It's because it kind of is crazy. When I was testing GPT -6 Astra early, I tried something similar here. I was doing a huge performance overhaul for Lakebed, my cloud, trying to really see how far I could push performance on it.
Those are these small number of user, tiny JavaScript apps in the cloud would use as little resource as possible. It did a bunch of runs and the performance was better after the first performance set of improvements that I merged, but it wasn't quite where I wanted. So I wrote up prompt very similar to when we just read.
Sounds like there's more improvements to be made. I want you to really dig deep and find everything we can, even extreme measures to improve performance. I want you to come to me with three different proposals for things we could do.
The more extreme, the better. Don't be afraid to boil the ocean. And this is why I now have a custom runtime for Lakebed.
I ended up forking Rusty V8 and building my own runtime in order to make the performance as good as possible. And it was like a 2 to 4x improvement. They proposed the persistent V8 execution service, which ended up being its own custom runtime.
Build a query compiler and maintain live answers and give each app a durable owner with local data. Didn't end up doing those other two because I just got so focused in on the persistent V8 execution service. Rusty V8 is a phenomenal library, by the way.
Thank you to the maintainers of that. But yeah, I effectively told it to just go hill climb, come up with something crazy and the classic boil the ocean. And it did it.
It made something awesome. And the performance is crazy. And for those asking, where's Lakebed?
it's become my pet project to play with these things with so uh making it like a real product that i can confidently pitch to y 'all is less a focus for me now but it is out there it works if you find the url which i don't hide very well in the npm package which is pretty easy to find too you can probably use it for real stuff i use it all the time hell i used it yesterday in my opus 5 5 video to make a better visualizer for the data on the artificial analysis dashes so much easier to read this than it is to read the artificial analysis home page it's a fun little thing it's very nice to just like set up apps that auto sync and live update and get stuff out but the performance improvements i got out of this really stupid prompt and also a lot of time spent working with it to build both ways to verify the performance improvements as well as ways to verify the behavior not regressing it was a a war it was one of the longest threads i've ever done in t3 code and a lot of pretty crazy run times too like the planning to 20 minutes the
first build was about two hours and then something broke i think i was doing something on the machine so it stopped i told to continue five and a half hours later it had the pr and it actually did the thing and it ran awesome so yeah don't underestimate something as silly as i am open to wacky ideas in a prompt like this and that's why the next quote here sounds insane but i i'm on their side here anything can be hill climbed they're not wrong from the start we knew we wanted to iterate faster than our deploy cadence Claude could work asynchronously for many hours, even overnight.
We wanted to let it validate its prototypes without waiting for field reads. To achieve that, we looked for other ways to measure performance in the lab. Sam found the first lead.
What can we do instead of wall clock timing? Can we measure JS instruction counts, for instance? Yeah, for pure JS hot paths, literal instruction counts.
Run the benchmark under Valgrind with Node Predictable and compare to a checked -in baseline. One run, no statistics needed. Okay, so we know this isn't Opus 5 .5.
There's one little hint that this wasn't Opus 5 .5. Thank you for including such a useful hint, Anthropic. For browser paths, there's no instruction counting under Chromium, but there's a ladder of other deterministic counts.
React commits per interaction, function call counts from V8's precise coverage, layout and style recalc counts, DOM mutations. What do you want first? He said, let's explore Valgrind plus IR and predictable flags in one thread and each of the browser React benches in new threads.
Ping me in all of them. You know what we want. Let's go.
I hate to be so fixated on this. I have not read too many prompts from employees of the labs. I don't know how I feel that my prompting is so similar to the prompting of anthropic employees.
I am a combination of proud and ashamed in this moment. Audio made fun of me for the way I prompt. I feel better about it now.
Anyways. 11 minutes later, five threads were running, each focused on different measurements. Why did it take 11 minutes to spin up five threads?
Regardless, instruction counts, V8 call counts, React commits, style recalcs, and DOM manipulations. I think I know what happened now to those weird issues I was calling out before. If you're just blindly measuring the number of React commits or V8 calls, you're going to start turning things off that don't actually affect performance.
The harsh reality is that when you're on the page and the page is working, Firing more network requests is not too big a deal, especially on like desktop web and desktop app experiences. I think both people and agents are a little too quick to look at every single network request and try to disable them rather than find ways to change the hierarchy of requests to get you what you need to interact with the page first and then keep things up to date and reliable after.
And the reason I'm calling that out now is giving the agent this type of insight is encouraging that type of behavior. because you're not going to see when the page becomes interactive via V8 call counts. But you are going to see every request happening, whether or not it's blocking.
And if you see opportunities to reduce those requests, like, I don't know, maybe we shouldn't fetch all of the sidebar data all the time. We already have it in local storage. That would kill six of the 800 calls on this page in this trace.
We should turn that off. And then you end up in the state that I showed earlier, where the app wasn't syncing the sidebar content properly, and you could end up with bad sidebar runs. not the best solution here.
It is useful tooling, but you need to also make sure the agent is combining that with real introspection to how the user flow works. I don't necessarily know if agents can do the whole end to end here yet.
That's why I still do performance stuff very much in the loop myself. I'll usually start by asking it to come up with theories to improve performance for things that could be causing a specific thing to be slow. Like I'll start with, it takes too long to start typing in the input box.
What do you think we can do to improve it? Then I'll tell it to go make demos with all of those potential improvements. And then I will test it.
myself to see if it feels faster or not. It is far too easy to cause weird niche UX regressions in performance overhauls. And I still do this a lot myself.
They then asked it to prove that hill climbing against each of these can actually result in measurable wall clock performance wins. We'll unship the benches for any candidates that cannot prove that. Interesting.
They also called out that they treated every new benchmark with skepticism. Each one had two jobs. First, make a metric that Claude could move in the lab.
And second, a guardrail in CI with a number that could only ratchet down. This is one of my favorites. Benchmark was flaky or didn't actually correlate with user latency.
We threw it out rather than letting Claude climb the wrong hill. Also very good they called that out there. I hate that I'm just going to keep deep diving into my own stuff, but that's what happens when we talk about performance.
I get to nerd out about my favorite shit. I somewhat recently did a lot of work in T3 Code's data layer to make it so you don't have to load as much data when you open big threads or you resume the app. That's a big part of why T3 Code feels so fast, even on absolutely horrifying hell threads.
For example, this build SwiftUI app from scratch thread. This thread I've been working on for four months, and it probably has like 500 prompts in it at this point. Watch how fast it loads when I click.
I am clicking now. Pretty much immediate. It did have like a flicker down due to a weird scroll state thing.
Fun problem for me to fix later. But things load pretty much immediately in T3 code because I had to go the hell and back finding every single place I could squeeze out data that wasn't necessary in the loads. I made a lot of these improvements and then noticed things starting to get worse again.
And I didn't want that. And I also wanted to be easier to measure when I made improvements. which is why the most important change I made actually had nothing to do with the performance improvements itself.
The thing that mattered is this guy. I added an action that would do a test for a bunch of different request formats that actually exist in the app with some stored data for a real, in quotes, fake thread to make sure that we are not meaningfully going above or below the existing data baseline that we have for... loading stuff and this has prevented so many regressions i am very thankful i added this because not only will the github action catch these types of regressions by putting it in as a comment both the agent working on the code as well as the agents doing pr review will notice it and call it out and potentially even fix it autonomously so this gives a signal to the agents when they make mistakes that cause performance regressions so they can fix it before it becomes my problem or worse my user's problem And once again, I want to emphasize, even if you're not that into using vibe coding and AI coded stuff to ship real world projects to users, if you're not using it for some of your work, you are missing out so much because using these tools to slop together tests like this, we're like, yeah, if this was mission critical to have our performance never regress, it might be worth really digging into these tests to make sure every detail of them is good.
But that was not the mindset I went in for this. It was a much more vague, wouldn't it be cool if we had a better idea when regressions kind of happened? And it threw together this test suite.
I think it's a couple thousand lines. I don't know. I didn't fucking read the code.
And now it has actively prevented regressions in our code base. So once again, slop is useful even if you never ship it to your users. Back to the article.
Wall clock time is what users feel, but it's noisy, and milliseconds are too flaky to use as CI gates. Instruction counts were appealing because they were deterministic, but we still needed Claude to prove they tracked wall clock time. We asked Claude to drive the count down on two hot paths, the routine that assembles a conversation's message tree, and a scanner for status lines and Claude code output.
Claude profiled both with Valgrind and found that a quarter of its first paths instructions were megamorphic dictionary lookups, resolving the same message ID three separate times. It'd be surreal with y 'all on this one. Data loading patterns tend to be the vast majority of the problems of these types of things.
Like if it is going down the same path three times, that means data is being fetched in a way that triggers re -renders three times, usually. An hour later, it cut instructions on both paths by 48 % and 31%. And while clock time had dropped from 78 % and 44%, we checked in two new ratchets.
From then on, any PR that raised the instruction counts of these paths would fail CI. The daily job lowered each ceiling whenever the count went down. That's another good idea.
If you do make improvements, it's important to update these things. Because if you don't, then you end up in a weird place where every performance improvement becomes area for somebody to make things worse again. This is even more painful with unit testing and code coverage.
I hate code coverage, period. Like the idea of like having a number that shows you how well tested your code is is silly and dumb. And I had times at Twitch where I had a feature that I rewrote.
We took like half a million lines of code and rewrote it in 50k. And we had shipped the new version with a feature flag to trigger between the two. We had moved everyone over to the new version and it was great.
And then I went to delete the other 500 ,000 lines of code and couldn't. Because when I did that, the code basis code coverage would go down from 95 % to 94%. It was like 94 .9 % or whatever, but it was enough that it made it so I couldn't merge the PR.
Well, you should add more tests to your changes then. I did. My 50K lines had 100 % coverage.
The 500K lines had 96%. But that was enough. we're deleting it would shift the code base enough to knock us just below the threshold and i never got to delete this giant old unused feature that we had rewritten because of code coverage does the count track the clock cpu instructions versus wall clock time yeah cpu instruction was 48 less wall clock was 78 less for message tree assembly that's only the best way to measure that also doesn't give us real times for wall clock here so we don't know if it matters because like if you shave 200 milliseconds to 60 milliseconds.
That matters a lot less than four seconds down to half a second. This led them to the central lesson of the sprint. With Claude, measuring something makes it tractable.
Yeah, makes it so you can actually give Claude the things it needs to improve once it can measure. Measurement used to be step zero. You'd add a metric, wait for data to roll in, and only then would you start to understand the problem.
With Claude, it's just step one of the climb. As soon as Claude has a number that it can beat, it can start optimizing. this meant that the highest leverage thing they could do is find more things to measure this is why i was excited to read this article i want to see the new novel ways they found to measure things the loop thread by thread this is a little cringe lme all of this ran in the same slack channel with multiple engineers and claude jamming in every thread from there the sprint settled into a loop one someone would open a thread about a slow stretch of the journey often with a screenshot or recording two claude would trace the flow and then find or build benchmarks that demonstrate the problem three Once it had promising results in the lab, Claude would come back with a PR, often several sized for risk and review, with anything user -visible behind a flag.
After it shipped, Claude watched the deploy and read the field data. If performance improved, Claude locked in the win by ratcheting the benchmark down. If not, it turned the flag off and iterated accordingly.
Then it went looking for the next slow spot in the same journey. I'm split on this flow. On one hand, nothing matters if it doesn't actually improve performance for real world users.
On the other hand, if you're not testing these things yourself manually as the engineer prompting for the changes and merging the PRs, you're not going to notice the types of regressions that users won't report. Let's be real here. How many users would have actually reported the sidebar issues I just found in Showcase?
Who is using Claude in two browsers enough to notice these regressions? And who would be annoyed enough to bother reporting that to Anthropic in a way that they would even notice? If they're not testing it themselves, you're going to have these types of regressions.
That's why I put so much work in T3 code to make it easier for us as the T3 code developers. to spin up builds to test ourselves. For example, I added a label on GitHub that I can apply to a PR that will trigger a manual build of the desktop app for Mac OS.
So I can just download the DMG and test it directly without having to like fetch the branch, build it locally and spin it up. This means I can even take one of my other machines like my MacBook Air I use for testing and go to GitHub, download the DMG and test it directly there without it affecting my machine. I think these details are just as important, if not more so, than measuring real world user performance.
And I'm amazed how few people have flows for this. I've talked to everyone from like small startups to independent devs to big businesses that are just outright failing in this way. Like if they want to test a change in a PR, that's a 15 to 25 minute process.
And it's so frustrating. If you're going to put slop somewhere, it should be there. Slop up ways to test your changes yourself with as little friction as possible.
because that friction makes you hesitant to do all of this in the first place. I am very hyped this made it in here. Someone shared a screen recording that showed sidebar rows popping in after the page loaded.
Chat and co -work rows resolved at different times, making the page feel janky. None of our existing monitors detected it. The closest that we had was a cumulative layout shift, but each shift only scored about 0 .008, well within the good threshold of 0 .1.
Isaac had the idea to reference the underlying layout instability API directly. I don't think you need to add a bunch of measurements for this. You just need to cache that one query.
Like, why are you hitting the, like, deep layout instability API for this? You don't need that. I'm very thankful they recently hired Adios Mani from the Chrome team.
Because he is going to clean up a lot of the slop that happened in this arc. Anyways. Apparently when they were doing that, they also added integration tests that opened the page with a populated sidebar.
held the sideways data until after first paint, and failed on any shift in any named region. Oh, this might be the problem, actually. If the data in the browser has some things in it and the data from the network has different things, that will be a layout shift.
If you have four threads cached locally and you have a fifth thread that it loads from the network, it will shift the four threads down when that fifth thread comes in. Their tests here might have actually caused the regression because doing this right would fail this test. I love trying to reverse engineer where these failures come from.
After the event deployed, Claude read the field data and found that 31 % of web page loads moved something after the page was usable without any user interaction. From there, Claude worked through the causes by name. A header row that arrived late, a caret that slid sideways once the user's name loaded, a list that moved when the scroll bar popped in.
Claude fixed the top offenders as a batch. And when they were gone, it found the next batch. I do love these types of comparisons.
Do you see how much stuff is moving around on this left example here?
Yeah, that's rough. This is one of the big benefits of using cached data on page load. The catch is if the cached data is not accurate against the real world data, you're going to have some shift.
But I think that like the chats and tasks shifting down one or two items because the real data loaded is not that big a deal, especially if you have a feed treatment on that for when that change happens. Someone in chat just asked a very good question. How do you avoid this type of issue without looking at the code for testing?
I'm going to phrase it back in a pretty silly way. How was I able to identify this issue with testing without ever seeing any of the code at all? Because I don't even have access to the repo where they wrote this.
The blog post that the Anthropic team wrote here is probably, I would guess, slightly less detailed than the responses Claude gave them in their threads in Slack. I don't know how you build these skills now because the way I built them was going deep in the details to fix all of these things myself by hand for almost a decade, way before AI was even a thing you could use for dev.
And I am noticing these issues because of my history doing these things by hand before. I don't need to see the code or read the tests to know that. I can tell from the PR description or from the response Claude gives me in my Claude code or from Astra in Codex.
a lot of this just comes down to yeah this you said it pranch intuition and you have to build that over time i didn't have to read the code to have these potential realizations and i still could be wrong, of course. I might not be correct in that these measurements are enabling or encouraging these specific types of regressions, but it's a theory that if I had code -based access, I could ask Claude to go prove out for me and then use that to tell the team like, hey, I like where you're going here, but these tests kind of suck.
Here are the regressions that we're going to cause if we don't shift how we think about this. So yeah, reading the code does not solve this problem at all. You got to read the responses from Claude and understand what it is doing when it doesn't.
If you're actually reading all the code when you do these types of tests, you're not going to get very far because I'll be frank, when I was doing all the data loading stuff in the T3 code, like overhaul, I probably wrote a thousand lines of tests and bullshit slop for validating my findings for every 10 lines of actual code I wrote.
It's probably a hundred to one ratio. And the code I wrote that actually shipped would have been worse if I hadn't taken the opportunity to slop a shitload of tests to verify my changes. And as far as I know, there were no regressions from those overhauls, which was actually surprising to me.
I expected it to be pretty brutal and wasn't at all. So yeah, take it as you will. I don't know how you build this intuition now.
Maybe you just ask Claude questions. I don't know. It's worth hiring somebody who has it if you don't have someone on your team with it and you care about these things.
It is probably learnable still. I just don't know how I'd recommend to do it. The next chunk of this blog was actually one of the more fun parts, scaling horizontally.
Once the loop worked on one thread, running it on more was just a matter of opening them. It's one of the really fun things about setting up a cloud environment or, I don't know, using something like T3 code in the remote machine management stuff. Once you have an end -to -end flow that is working for figuring out a class of issue or bug, you can spin up a bunch of threads on a bunch of instances on a bunch of machines and immediately parallelize it.
We're all hopefully engineers, and we've experienced this before. Writing code to do a task once is rarely worth it. but once you wrote the code which might take hours even the task to five minutes once you've written all that code you can now run it again and again and again you can parallelize it you can do a bunch of and if it turns that five minute thing into three seconds and you can run hundreds of them at once it fundamentally changes the way you operate with that thing so yeah sure i put a lot of effort into the testing i built for my performance improvement changes But once I had built all of that testing, it was trivial to fan out lots of agents with lots of theories to try and solve lots of things.
And the result is that my agent auto -merged like 40 PRs improving performance in T3 code. And the only real regressions were one change to animations on the actual desktop that people didn't like, which I understand I didn't like the change either. And it broke the sliding ticker for tweets on the homepage of the marketing site.
But other than that, even Astra and its weird spikiness, 38 out of the 40 PRs, it merged itself. This is an interesting detail, though. Instead of closing threads once the original request had been fulfilled, Claude would keep going.
An individual thread would put up 50, sometimes 100, optimization PRs. Interestingly, it was Claude, not one of us, opening new threads to chase opportunities it had found on its own as parts of separate investigations or nightly jobs. Shelley, one of the engineers in the channel, observed, quote, Yeah.
It's surprisingly easy to set something like this up. Every measurement found something to improve. Claude ran a React hook census and found 6 ,900 hooks and 900 store subscriptions in the composer's typing path, re -rendering on every keystroke.
Nearly 7 ,000 hooks for the composer. Maybe we should read the code, guys. I take it all back.
Anyways. Claude counted style recalculations and found a single root has selector that added 24 milliseconds to every single DOM change. You know how many of these types of things that are like one simple CSS thing causes...
every change in the browser, even in Electron, to be way slower. It's way too easy to cause these things. And Chrome's tools do not make it easy enough to find these things, sadly, especially GPU stuff.
Chrome's okay at helping you with CPU stuff, at event loop stuff, anything JavaScript side. Once you're in the GPU accelerated CSS layer, good luck, have fun. I'm going to choose a different mindset for how I'm going to approach anthropic engineering going forward now.
This might get me in trouble. I don't care. Anthropic has chosen to be the first company to really, really experiment on what happens if you let AI write too much of the code.
Since they did this in the Sonnet four days, they have a code base that is full of slop and garbage. They learned a lot of lessons. They made the model much better.
And now for the betterment of everyone else, they are using the models to clean up the slop that they made themselves. the cloud ai site and the cloud desktop app and i'd argue even cloud code itself are them dog fooding both how much slop can something take and also when are the models good enough to clean up the mess that we created they also found a code path that had a leftover location .reload which is like a browser refresh that would cause half a million hidden reloads a day that none of their load metrics could see Claude read profiler samples from idle tabs and found identical cache snapshots being cloned into IndexedDB twice every minute, all on the main thread.
Oof. This actually might also be the change I was talking about with the sidebar issues. If they made it so that loading data from the network doesn't write to the IndexedDB storage as much, that would prevent that storage from being updated.
Great point from Jamie from Convex here. Companies that are good at making models are not necessarily good at software engineering. This is a lesson playing out slow as fuck.
I'll go even further than this, Jamie. There was a deep, as I'm sure a lot of y 'all know, a deep rivalry between OpenAI and Anthropic because Anthropic was largely started from the original pre -training team at OpenAI. Dario and friends were all helping build GPT -3 and they left because of a lack of alignment between them and the leaders at OpenAI.
They still say a lot of this is around safety concerns. I don't actually think they were that concerned about safety in the GPT -3 days. I and many others have started believing the real friction here was the difference in engineering versus research, where OpenAI, despite being a research company, was led by two engineers, Sam Altman and GDB, which meant that engineers were the ones commanding researchers to go do research stuff.
And that meant they were holding them to engineering standards. They were expecting engineering deadlines and timelines and all of that type of stuff. And that caused disdain between the two groups.
Since those guys left, it definitely caused some frustration and open AI towards researchers. And they started looking for ways to engineer solutions to research problems. Meanwhile, at Anthropic, it meant that they kind of didn't like software engineers that much.
And you can still see that in their comms even now. I think a lot of why Anthropic pushed so hard towards this idea of what if we replace all the engineers with AI? It's because they didn't like engineers, which means they probably weren't hiring the best ones.
I'm going to be very real here. It's some of the advice I give the most. Hire for what you love, not what you hate, because you'll make much better hiring decisions.
Since they were hiring engineers and they don't like engineers, they weren't making great engineering hiring decisions. It's difficult to automate things you don't understand, and it's difficult to understand something that you don't respect. Absolutely agree, Jamie.
Back to the performance improvements, because I love this stuff. We rarely knew where a thread would lead. In a sweep for CPU hitches, Claude noticed that highlighting a finished code block could freeze the page for about a second.
Oof. Syntax highlighting in the browser... Syntax highlighting in general is one of those, like, really annoying problems to get right.
It dug in the lab and found the culprit. MDashes! Oh, man.
Oh, man. I saw someone say in chat that there was an MDashes section. I...
I was not prepared. If a replies markdown contained any non -Latin 1 character, like an em dash or a curly quote, V8 would store the entire string as UTF -16, which put every syntax highlighting regex on a slower 2 -byte path. CloudFix built a 20 -line change to copy each code block into a 1 -byte string before highlighting it.
Blocking main thread for 0 .35 seconds is, in my opinion, always unacceptable. We moved our syntax highlighting off to Wasm because since it now runs in a worker, it won't block main thread at all, which is quite nice. Write me five code blocks in various languages doing some simple demo that's at least 10 lines of code.
And I'll put them in the chat here so I can see them.
You might notice the fade. I added the fade to hide the fact that I was doing the actual highlighting in a different thread to not block the main thread. So you can still navigate around T3 code and it still flies even when those things are happening in the background.
I was able to syntax highlight various different languages all in a row while I have a bunch of other stuff going without it affecting the main thread's performance at all. So yeah, they should probably be doing that too. The catch here is that you have to load a bigger bundle for that syntax highlighting.
But if you lazy load it anyways, it comes in later once and you're fine. So yeah, I'm surprised that they didn't get Claude to propose that to them because it would have helped with a lot of these problems. I would argue that since we're not reading the code anyways, main thread being blocked at all for syntax highlighting is bad.
Apparently by the second week, they could barely even summarize their output into daily updates. On the busiest days, more than 200 changes would land. Claude kept proposing new benchmarks.
About a third of PRs included additional telemetry or guardrails, and each new instrument generated more threads with more opportunities. Working in one channel meant everything happened in the open. We jumped in and out of each other's threads to debate decisions and celebrate wins.
Word spread. other teams started to bring their changes into the channel to have them reviewed for performance. New projects were written in subtly more performant ways because of the guardrails that had been added and the new cloud skills that had been introduced.
I do really like this idea of gang prompting at companies, like having a channel that can be used by various different people across various different teams with shared contacts and capability. This isn't reading like an ad for CloudTag and Cloud and Slack, but... It is a very compelling one.
It makes me want to go set up something similar for us with all the T3 code bot stuff that Maria has been doing. And now we are in the guardrails. We'd prepared for the pace because nearly everything we touched was a hot path.
The first page, composer, transcript, etc. We'd established our safety mechanisms up front. Every PR went through automated reviews with at least one human approval.
Unit tests always came before optimizations. Anything that could cause a user visible problem shipped behind a short -lived feature flag. I don't believe you there.
Regardless, once the flags started to pile up, they opened a thread to coordinate their rollout and cleanup. Claude classified every flag as a kill switch or ramp and retired each one as soon as it was safe. Across the two weeks, they introduced nearly 200 flags, more than half of which were already cleaned up by the end.
That is nice. Cleaning up the slop as you add it is an underrated thing. New performance wins decay in a fast -moving code base and code chips fast and anthropic.
Once a project proved a win, they invested in ways to protect it. Static Composer, for example, is brittle by design. We show users an HTML copy of the page almost immediately, and we let React paint directly on top after.
If the React render is off by even a pixel, all of the magic is lost. So Claude built dozens of guardrails. Static Markup is generated by rendering the real React component in JSDOM, and tests guarantee that they never drift.
Integration Test Suite compares the static page against the React render across 14 different viewport sizes, and then asserts alignment within one pixel. A keystroke test types straight through the handoff and fails on any lost reordered keys.
And in the field, every handoff report shifts to a 10th of a pixel. Cloud opens a thread for any event with non -zero movement. Interesting that they have like threads auto open when users experience things by the sound of that.
Not everything could be caught in the lab though. So they had to make use of the oldest guardrail in the book, incremental rollouts. High risk changes were rolled out to employees first, then 1 % of users, then everyone.
Four hours after we released the static composer internally, a teammate shared a screen recording of a layout shift that none of our metrics could see. When he opened Cloud AI in a new tab, the composer would drop em dash, but it wasn't our code. This is a Chrome extension.
I'm calling it now. Chrome resizing the page, not the handoff from the static composer. On the new tab page, Chrome draws its own 56 pixel footer below the page.
Managed by Anthropic .com. Customize Chrome. When the tab navigates to Cloud AI, the footer goes away and the page gets 56 pixels taller, but only 100 milliseconds after the first paint.
Okay, so it wasn't a Chrome extension, which is the cause of most of my weird hellish issues, other than Brave. Brave causes so many problems for us. That all aside, this is weird managed profile stuff that causes niche bugs within Chrome's rendering itself.
They are putting the... Oh, this is actually probably the bigger problem. They are rendering the...
composer based on a percentage of page height and a distance from the top rather than a fixed location on the bottom by doing that they need to fire a js calculation which means they are now reliant on the page resize behaviors and it seems like that didn't or get triggered on time or worse it triggered after the 100 millisecond like repaint occurred or the 100 millisecond shift and then it repainted and moved down possibly even before the js loaded in fascinating It's because of Chrome speculative loading.
While a user was typing the URL in the address bar, Chrome would pre -render the page in the background and the height of the current tab. Huh, okay. That actually is weird, but it makes sense.
Yeah, when you're on this page, there's a certain amount of space available. So if there was something on the bottom here that caused this window to be smaller, and you type clod .ai slash new, this is now pre -loading in the background. And when I hit enter on it, it doesn't have to load as much as a result of that.
But Chrome's prefetching triggering a behind the scenes render that is at a different page size is fascinating. And I would guess this is a non -zero part of why they hired our friend Addy Osmani from the Chrome team. It's actually one of the more fun ones that they have in here.
I think that's a really cool finding. Yes, from the Chrome pre -render. If I was using Claude in Slack and I had a thread like this and I saw Claude come in and point out a Chrome behavior I didn't know about.
I would be the most insufferable human for at least a few hours. I'd be like, what the actual fuck? This is AGI.
They also have a section at the end here about steering, which I think should be quite fun. The loop was productive, but it wasn't autonomous. Keeping it fast, safe, and on track was our job, and it had three parts.
First, ambition. By default, Claude is careful about scope. Hey, chat, do y 'all feel that Claude is, by default, careful with its scope?
I want a yes, no, or LOL from all of y 'all.
Thank you guys. I think the point's been made. Regardless, yeah, helping it know it can push further helps a lot.
You'd be amazed how much like one small change to your CloudMD will affect these behaviors too. If you add to the CloudMD something like, assume you have my permission to keep going. If I state what I want the outcome to be, he'll climb till you get there.
It's much more likely to actually get there. They called out here that they were confident in their guardrails, which meant they didn't want it to be so reserved and careful. So a lot of what they did, especially early on, was encouraging Claude to be bolder.
Here's in one of the threads, Claude said, one small PR to give Code the same timing marks that Chat and Cowork already have. I'll put it up this week. Realistically, Code's number waits a few days for merge, deploy, and a baseline window.
To which Raymond had to respond, if you put it up right now, I will get it merged and deployed. We have the power to do anything. Please be braver.
pr will be up within the hour god quad is still so bad at time estimates speaking of which look at that it actually spun up my thread on my other machine in t3 code to do the overhaul of uh ping nice i don't even have anything set up in t3 code in this instance to allow a thread to spawn another thread i just got it to ssh into another computer clone the repo copy my environment variables copy the plan MD file that we had made and then prompt from one machine to another, spin this up and then it reappears automatically with the remote connections in T3 code.
Pretty fucking cool. People aren't prompting boldly enough and they're not telling their agents to be bold enough either. And I can already hear the comment section.
I know how this goes. Well, Theo, some of us work in real code bases where we can't just merge, slop and break things. First off, so does Anthropic, and they break more stuff than I do.
Second off, more importantly, the focus of this isn't on merging a bunch of breaking changes. The point of this is proving that you can use the slop to measure things, to make regressions and problems less likely, and also make it easier to make improvements. I have a feeling the comment section on this video is going to literally kill me, so I'm just going to move on.
I do like this callout that they put at the end of the section there. When we started hitting the targets we had set, we noticed the threads would slow down. Sam went thread to thread with the same message.
Let's keep driving this down. The targets are not the stopping point. What's next?
Be ambitious. I know y 'all think this is cringe, crazy, not useful prompting. What you need to understand is that these models are effectively word maps and the things it chooses to do because the paths in the model drive it there when it's just being told, fix this thing and improve this metric are going to be the safe, happy paths to do that.
But as soon as you start including words like ambitious or boil the ocean or push beyond what I requested or make weird and wacky suggestions, you're pulling the model out of the happy path and towards the sketchier, scarier ones that is correctly trained to go down less because if you don't have the right tooling, then it could break shit.
So if you... I know you're saying this ironically, S -John, but I actually think this would work. I would unironically guess that if somebody had a model that they felt like wasn't going far enough, if they had a bot or a hook that would automatically append fuck me up fam to the end of every prompt they send, they would actually notice a better experience.
It sounds so stupid. It really does. And I know you guys are going to cook me in the comments for this.
I don't fucking care. It works. Speaking of all of these fun things that people don't really want to believe in, myself included often, the next section is about taste.
Every thread had a named human owner and Claude highlighted any user perceptible change with before and after screenshots or recordings for them to rule on. Should a table fill in cell by cell or wait until each row is complete? Should a loading skeleton show up immediately or only after half a second?
Is a word by word fade on streamed text worth the fifth of a frame budget it costs? Claude looked for ways to shave milliseconds and we weighed the trade -offs. I mean, I can answer all of these questions because I know how browsers work.
Should a table fill in cell by cell or wait until the row is complete? It should wait till row is complete. I'd argue it should wait until the table is complete.
Should a loading skeleton show up immediately or only after half a second? You should do a lot of metrics and analysis to know how long loads take and how wide a range these extremes are. Because the worst case would be that you think the data is going to take a second.
So you show this half a second in and then 501 milliseconds in. It has the data. So you show the skeleton for like two frames at most before you show the real data.
You got to tune this one based on real world usage. And if you get really fancy with it, you can time other requests and cache some data locally around how long this request usually takes for this user and then make programmatic decisions around the skeleton based on that. And then we have the word by word fades for stream text worth a fifth of frame budget.
No word by word fade in is almost, I would say just straight up isn't ever worth it. I begrudgingly went from only showing things when the whole response was done in T3 code to only showing things when the paragraph or new line is completed or the code block is done. And that allows us to have insane performance still without any of these weird potential issues of like frame time being killed by per token stuff.
Yeah, like I'm happy that they were able to work with Claude to find what seemed to be reasonable answers for all of this. But I know I want to see how reasonable these answers are. i'm testing how tables render in quad i want you to render some test tables with at least four columns and 10 to 15 rows can use whatever you want for the data i just want to see how they render let's see how they come in they're doing row by row it looks like ah nope that that was pretty brutally streamed to the side there maybe they're doing line or they're doing a sell by sell yeah that's that's very sell by sell i might need to switch to fable just to make it slower yeah that's definitely sell by sell This might end up being a bit of a self -own here if I do this, but I want to see how this compares to the performance in T3 code for a similar request.
I prefer that. I don't want to see the table while it's still being created. I just want to see it when it's done.
To each their own, but I prefer that. The one thing I don't like, and I'm still deciding what I want to do about it, is the title comes in first, and then we have to wait for the rest of the paragraph. I almost want to hold title renders until the content underneath is done.
That's like the only change I would make with our behavior here. But yeah, you get the idea. These are complex things that you should have conversations with your team about and really make good decisions on.
Row by row is what chat's asking about. It's fine, but I prefer for the table to be done personally. I will also say that row by row renders for like code blocks.
Oh, horrible, terrible. It breaks all the highlighting. It makes way more render work than you should need anyways.
And I don't want to read a partial code block. I just want to read it when it's done. I also find that I don't really sit watching threads anymore.
I usually kick off the thread and then go somewhere else and then come back when I have a reason to. So yeah, I now see why they labeled the section taste because this really is taste. Now we hop to direction.
We kept each thread deliberately narrow, focused on one benchmark or journey, and asked Claude to find improvements only within that scope. We thought of threads as 150 hammers seeking nails. Most of our calls were about sequencing and user impact, which surface to prioritize, how to combine threads that were stepping on each other, and when to close a thread that had reached diminishing returns.
On one 900 -line PR, we got a one -line reply, quote, going to gavel that two millisecond percent is not worth the complexity of maintaining this build plugin. They were adding build plugins to shave two milliseconds off message sends. This is why you have to be careful about how far you let these things go.
Like I wrote my own V8 based runtime for my cloud, but I had to like push for that. We're nearing the point where the agents will autonomously do bullshit like that. So make sure you look out for unnecessarily huge changes for unimportantly small wins.
And this sounds like an interesting section, an eight millisecond budget. One of our side quests shows everything working together. To demonstrate an optimization to a regex used in the live syntax highlighting, Quad attached a screen recording of a longer answer streaming in the lab.
In the corner, it added a frame rate readout computed in the page from animation frame timestamps. Well, yeah, for those who don't know, eight milliseconds is how long you get per frame at 120 hertz. All you will have 120 hertz displays, whether it's their iPhone, their MacBook or any gaming monitor, even some modern TVs.
It's a good target to aim for because somebody with a very expensive pro MacBook is going to want the page to run at the 120 FPS. This is actually kind of a sick bench. Are we capped at 60 FPS?
Can we try to drive scroll and stream smoothness up to 120? If I understand correctly, your rig may not support this. Right.
Today's rig ticks at 60 hertz because headless chromium does by default. I believe it can be driven at 120. Confirming that first, then I'll rerun the eval at an 8 .3 millisecond frame budget.
Update on the 120 hertz rig. It works. Deterministic 120 hertz frame stepping in headless Chrome via dev tools.
Begin frame control. Exactly 240 frames for 240 begin frames at 833 ms. This is not Opus 5 .5.
This is clod slop. This is hardcore clod slop. Then Raymond.
Amazing. Cook. Yeah.
I love this. I also am just now realizing one of the things that I really like about quad tag. Notice what isn't in this thread.
Notice that we don't have code. Notice that we don't have tool calls and the responses for those tool calls. Notice that it's not listing all the commands it ran.
Notice it's not showing all of the thinking traces for these sections of thinking that it did. I like this because they don't really matter anymore. I'm considering adding a feature to T3 code that just outright hides them.
Or maybe gives you a small ticker like, I've run this many tools so far. Because I don't care. I just want to know when it's done.
I also can't help but note the irony that a lot of the performance things they're trying to fix here are not in Slack. They're in the Cloud web app and desktop app, which they are notably not using for any of this. First off, that's funny because Slack's a better experience with Claude than Claude code.
But more importantly, it's funny because a lot of these performance issues are things that they show in the desktop app and on Claude .ai that they do not show in Slack. Just saying. Once the mechanisms and ambition were established, Claude got to work.
Each painted frame had a budget of 8 .3 milliseconds. So Claude stepped through a long reply frame by frame, timing each one to find the slow parts. It eliminated some work that was O to the message length, so however many things are in the message, which is a lot of work it was doing per chunk by memoizing finished blocks, moving tokenization logic from growing code fences to a worker, and revealed tables cell by cell.
Oh, that's cool. We were doing per chunk memoization for T3 chat in February of last year. It's been about 22 months or so.
The moved tokenization logic from growing code fences to a worker. This is the early pieces of what I was suggesting earlier, which is using workers and Wasm to run your syntax highlighting and code block stuff so that it's not blocking the main thread. Also, you can cache the result of that in memoized chunks on the outside of the Wasm side and the worker side in the React world.
So you don't have to re -trigger the Wasm on every change. You can also reveal the table cell by cell. Yeah.
So I guess I had the answer to my question earlier about how they're doing the table stuff here. I can see you just scroll more. In that one thread, we landed nearly 60 PRs.
This is just the thread about the 8 .33 millisecond render times. Long replies were blocking the main thread for 200 milliseconds in total, where they used to block for over 750. Ran on about a third as much CPU, and it held 120 FPS from start to finish on a 120 Hertz MacBook.
The 120 Hertz rig itself became a nightly job with Claude watching for regressions. Awesome. We love that.
Oh, I hadn't thought of that, Igor. Good point. We could just use Jev to do the syntax highlighting.
That would solve all these problems. Oh, that's actually really funny. They cite the tweet that I...
Actually, people saw what my post about this as dunking. It was not meant to be a dunk at all. I was actually trying to start a conversation around whether or not streaming matters.
This is my quote tweet here. I don't want responses streamed anymore. Just give me the whole paragraph or block when it's done.
To be very clear, I'm referring to token by token streaming. You should stream down paragraphs, code blocks, tool calls, etc. Just stop sending me incomplete text.
I also called out that I feel obligated to mention this performance work is very hard and I'm impressed about the work they did here. Particularly funny that I have this when in the end it was all just Claude running in a loop. Anyways, when they started this performance sprint, they hadn't planned to hill climb on the milliseconds between frames while streaming, but it turned out that they could count them.
So it being countable meant Claude could climb it. So what's next? Cloud AI and the desktop app are now three times faster than they were in early August, and the ratchets should keep them there.
But they're not done. The 95th percentile, other journeys, and very long conversations still have room to improve. In a separate post, we'll also write about some of the side quests that took them upstream during the sprint, with contributions landing in Electron, Chromium, Node, and more.
They'll share the results internally. Isaac put it best, you cannot have convinced me that this was possible even six months ago. Yep, this would have been very hard to convince me six months ago.
We expect to keep working this way one thread at a time at any scale. The channel's still going. All shout out Anthony.
I know I've been harsh towards Claude Desktop, but I know he's been working really hard on improving it. I did actually use it a bunch when I was testing Opus 5 .5 and it has made a lot of improvements. So shout out to him and the rest of the team making Claude Co.
Desktop way better. Anthony's awesome. And also I do love that Boris is encouraging them to be more ambitious.
That is a very good thing for him to do as the leader. This was a genuinely really fun read. I enjoyed this a lot.
I didn't learn quite as much as I hoped to, but it did really emphasize a lot of the things I've already been trying to do, giving the model the ability to measure things so that it can fix them, rather than hoping it will figure that all out when you ask it to fix things. If you tell the model improve performance, it will do its best based on its theories to improve the performance.
But if you start by telling it to find ways to measure performance, validate those findings, and then tell it to go improve the performance, it'll improve a lot. This is really cool. And I think we can all learn a lot around how they are prompting how they're thinking about these changes and how they're validating the work that they're doing.
Because that is going to be more and more important as things continue to change. If you still think that reading the code is the solution to all these problems, then you're going to have a whole separate set of problems and a whole lot less solutions. Figuring out how to work with and around the slop really kind of is becoming the meta.
I know a lot of engineers are going to hate that. I personally find it kind of fun. It's a new type of engineering that requires new types of cleverness and design to set up systems and structures that allow your agents to operate effectively and efficiently.
I'm thankful I'm finding more opportunities to deep dive in these fun engineering topics again. I really am enjoying it. Thank you all for watching this video as well as all the kind words on the recent React Native one.
It's such a relief to be able to dive into the engineering topics again where it makes sense. Let me know how you feel about all of this and what you learned in the comments. And until next time, peace nerds.
The Hook
The bait, then the rug-pull.
Theo opens by admitting he used to be a performance-obsessed full-stack engineer before he became known for AI vibe-coding content, and that history is exactly why he's qualified to grade Anthropic's own claim that Claude made claude.ai three times faster. He reads the company's engineering post line by line, tests the live site against it in real time, and catches a caching bug their own fix left behind.
Frameworks
Named ideas worth stealing.
30:30model
The Loop (thread by thread)
Open a thread
Trace the flow and build a benchmark
Ship a PR behind a flag, sized for risk
Watch the deploy and read the field data
Ratchet the benchmark down or turn off the flag
Look for the next slow spot
The repeatable cycle Anthropic used for every performance thread: measure, fix, prove it in production, lock in the win, then hunt for the next bottleneck.
Steal forany process where you want an AI agent to work mostly unsupervised but stay accountable to real data
CTA Breakdown
How they asked for the click.
VERBAL ASK
1:09:40subscribe
“Let me know how you feel about all of this and what you learned in the comments.”
soft engagement ask in the outro, no hard sell, consistent with the video's low-pressure reaction format
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Theo spends forty minutes inside Anthropic's own Fable 5.1 prompting guide, rebuilding his habits around effort levels, finishing the whole task, and trusting the model's defaults instead of babysitting them.
Boris Cherny said coding is solved. Matt Pocock called it VC-funded bullshit. Theo argues they're both right, because they're using the word coding to mean two different things.
Theo walks through Anthropic's own usage guide for Opus 5.5, line by line, and adds the real prompts and a six-and-a-half-hour mistake to prove which parts actually hold up.
A developer who shipped 89 merged PRs in 24 hours breaks down Claude Fable 5.1's pricing, benchmarks and real-world coding behavior against Fable 5 and GPT-5.6 Sol.
Theo reacts line-by-line to Boris Cherny's post arguing that automation — CLAUDE.md rules, lint checks, CI — matters more than ever in the agent era, not less.