The argument in one line.
The productivity inflection from AI agents is not coming — it is already confirmed by Anthropic's own internal metrics, and the companies that fail to restructure their workflows around it will repeat the mistake companies made when personal computers arrived and productivity did not follow.
Read if. Skip if.
- You are a software engineer trying to understand how far agent-based workflows have actually matured.
- You run a team and want a grounded benchmark: how much code is really being written by AI at top-tier companies?
- You are evaluating whether token usage in your organization is real demand or incentive noise.
- You want a first-person account of what running hundreds of parallel Claude agents actually looks like day-to-day.
- You are watching the SaaS disruption thesis play out and want a framework for which moats survive.
- You want technical deep-dives into model architecture — this is a product-and-business conversation.
- You are already up to speed on Claude Code's capabilities and are looking for new feature announcements.
The full version, fast.
Claude Code grew faster than any product the team had seen across prior careers in tech, with each model release (Opus 4.5, 4.6, 4.7) producing a new exponential inflection. Code-per-engineer at Anthropic is up 250% since launch, and 100% of Claude Code is now written by Claude Code itself. The token-maxing debate misses the real dynamic: companies that restructure their workflows around agents — as some did with PCs in the nineties — capture enormous productivity gains; those that bolt AI onto existing processes see little. The near-term road map centers on longer-running tasks, parallel agent fleets, and the auto-mode safety layer that routes tool-use decisions to a second model rather than a fatigued human.
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →Who's talking.
Where the time goes.

01 · Cold open
Coming-up teaser: growth, tokenmaxxing, sustainability.

02 · Claude Code's explosive growth
Internal release to exponential inflection; each model drop (4.5, 4.6, 4.7) re-inflected; team had never seen growth like this.

03 · What is Claude Code?
Agents vs. chatbots; the 'fancy text editor' bet; tools as the key differentiator.

04 · Using AI agents to book flights
Boris's CoWork story: 8 flights, 5 hotels, wrong hotel corrected; trust ratchet analogy to Waymo.

05 · Token-maxing and real demand
Amazon FT report; HBR PC paradox analogy; 250% code volume at Anthropic; advice: give tokens + psychological safety.

06 · Are AI agents too inefficient?
PDF spiral example; effort controls; intelligence vs. efficiency tradeoff; commenter's 'inherent to LLMs' claim rebutted.

07 · Rate limits and user frustration
Doubled rate limits; weekly limit increase; Colossus capacity; power users running hundreds in parallel.

08 · Beyond coding
QuickBooks, auto-mode safety layer, parallel agents as the next UX frontier.

09 · Claude prompting other Claudes
Boris's own workflow: a Claude that talks to his Claudes. Engineer leverage up; still bottlenecked on good people.

10 · The SaaSPocalypse
Seven Powers framework; network effects gain importance, switching costs collapse; one-app-for-all-software thesis pushed back.

11 · Self-improving AI
100% of Claude Code written by Claude Code since Opus 4.5; Jack Clark's 60%/2028 estimate endorsed; AI safety as the reason Anthropic exists.

12 · Do AI agents need world models?
Jan LeCun vs. Greg Brockman; Boris sidesteps the theory but cites surprising emergent planning (poetry experiment).

13 · Is this the future or a fever dream?
Opus 4.7 hackathon: doctor, electrician, carpenter. People jumped through terminal hoops to use it — the ultimate market test.
Lines worth screenshotting.
- Code-per-engineer at Anthropic grew 250% after Claude Code launched, without any regression in code quality.
- 100% of Claude Code is written by Claude Code — a self-authoring loop that has held since Opus 4.5 in November 2025.
- The HBR computer paradox of the nineties is repeating: companies adopting AI without restructuring their workflows will see no productivity gain.
- The productivity gains from AI tools come from people you never predicted — accountants, marketers, new grads — not your top engineers.
- Token-maxing is a real phenomenon but a small share of demand; the larger signal is that most Claude Code users never hit their rate limits.
- Switching costs as a business moat weakens as AI makes migration easy; network effects strengthen because the code author becomes irrelevant to the network's value.
- Running one agent at a time is already obsolete — Boris Cherny runs hundreds overnight in parallel, sometimes thousands.
- Auto-mode routes tool-use decisions to a second Claude instance, and that second Claude catches unsafe commands the fatigued human would have approved.
- Non-engineers installing Claude Code in a terminal — their first terminal experience — was the early signal that the product had broken out of the developer niche.
- The Opus 4.7 hackathon produced a winner who built and sold a startup; other winners included a doctor, an electrician, and a carpenter with no coding background.
- Claude is starting to generate its own feature ideas for Claude Code, though they are not always good ones yet.
- Effort controls (low/medium/high/extra-high) let you trade token spend against intelligence without switching models.
- The model improves faster than users update their mental model of its capabilities — people who last tested it a year ago still think it is unreliable.
- Jack Clark's 60% estimate that models will start self-improving by 2028 was described as 'seems right' by the person running the product that already writes itself.
The agent inflection is already past tense.
When the team building the product runs hundreds of instances of it overnight, and the product writes 100% of its own code, the question is no longer whether AI agents work — it is whether you have restructured your work around them.
- Each model release (Opus 4.5, 4.6, 4.7) produced a new exponential inflection — a pattern the entire team described as unlike any hypergrowth they had seen before.
- Code-per-engineer at Anthropic grew 250% without quality regression, a benchmark worth measuring in your own organization.
- The jump from chatbot to agent is a single property: the ability to use tools — edit files, run commands, access browsers — rather than just talk.
- The original bet was explicit: the existing model for writing code (a fancy text editor) was so suboptimal that something radically different was worth building.
- The trust ratchet with AI agents mirrors the Waymo experience: white-knuckle approval of every action, then five minutes later you are on your phone while it works.
- CoWork found two missing travel stops and incorrect dates that Boris had provided — catching errors in the human's input, not just executing instructions.
- Give everyone tokens and psychological safety before you optimize — the productivity gains come from people you never would have predicted, not your top engineers.
- The HBR PC productivity paradox is repeating: companies that do not restructure their entire business process around the new technology will see no productivity gain, even with access to the tools.
- Effort controls let you dial token spend vs. reasoning depth without switching models — low/medium for speed, extra-high/max for the hardest tasks.
- The 'loops and spirals' inefficiency is a current-model problem, not an inherent LLM problem — Claude Code itself 18 months ago had the same behavior before the model improved enough to self-author.
- Very few users actually hit rate limits — the most vocal complaints come from a small edge of power users running parallel fleets, not the median user.
- Running hundreds of agents in parallel overnight is now a real use case, and the API offers unlimited tokens for those who outgrow plan limits.
- Auto-mode routes tool-use decisions to a second Claude instance; the second model is safer than a human at the permission prompt because it does not rubber-stamp from fatigue.
- The next frontier is not new categories but longer-running tasks and better UX for orchestrating parallel fleets.
- The leverage stack now runs human prompts a Claude that prompts Claudes — the human is no longer bottlenecked by model throughput, only by the quality of their steering.
- Even with per-person leverage multiplied dramatically, Anthropic is still bottlenecked on good people because the demand growth outpaces the leverage gain.
- Network effects are the moat that strengthens as coding becomes cheap — the value of a messaging app is who is on it, not who wrote the app.
- Switching costs are the moat that collapses — Claude Code can already migrate stacks between vendors, and the model will only improve at this.
- 100% of Claude Code has been written by Claude Code since Opus 4.5 (November 2025) — the self-authoring loop is already live, not a future milestone.
- Claude generates its own feature ideas for Claude Code, but the ideas are not always good yet; the human is still responsible for steering, not just prompting.
- The empirical counterargument to the world-model critique is that models trained only to predict the next token exhibit surprising planning behaviors — including composing the second line of a poem before finishing the first.
- Boris's practical rebuttal is an open invitation: sit down and use Claude Code for one hour, then decide whether world models are missing.
- People without coding backgrounds installing a terminal tool for the first time because they needed what it did is the most reliable product-market fit signal — not press, not benchmarks.
- A hackathon winner built and sold a startup with Claude Code; other winners included a doctor, an electrician, and a carpenter — the non-engineer breakout is already happening.
Terms worth knowing.
- Token-maxing
- An organizational practice where employees are rewarded or measured by the volume of AI tokens consumed, sometimes leading to artificial usage that inflates demand metrics without producing real productivity gains.
- Auto-mode
- A Claude Code feature that routes tool-use permission prompts to a second Claude instance instead of a human, reducing approval fatigue while improving safety by using a model that is not subject to fatigue-driven 'always allow' approvals.
- Effort controls
- A per-session setting in Claude (low/medium/high/extra-high/maximum) that trades token consumption against the model's reasoning depth, independent of model tier.
- Seven Powers
- A strategy framework identifying seven durable business moats: scale economies, network effects, counter-positioning, switching costs, branding, cornered resources, and process power. Boris uses it to reason about which software businesses survive AI-driven commoditization of coding.
- CoWork
- Anthropic's computer-use product — a Claude-based agent that can take over a user's browser and desktop to complete multi-step real-world tasks such as booking flights or configuring software.
- World model
- A proposed AI architecture component that would give a model an explicit internal representation of physical and social causality, allowing it to predict consequences before acting. Yann LeCun argues LLMs lack this; Greg Brockman and Boris Cherny both push back empirically.
Things they pointed at.
Lines you could clip.
“I've just never seen growth this steep, and then it just kept going more and more exponential.”
“I don't write code. I prompt Claude. And actually nowadays, mostly what I'm doing is I have a Claude that prompts other Claudes. So I don't even talk to Claude.”
“The amount of code written per engineer at Anthropic has grown something like 250% since we introduced Claude Code.”
“Most nights I run hundreds of Claudes in parallel. Sometimes thousands.”
“People were jumping through hoops to use it because it was so useful.”
“Claude Code is a hundred percent written by Claude Code. CoWork is a hundred percent written by Claude Code.”
Where the conversation goes.
Word for word.
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
The bait, then the rug-pull.
Boris Cherny has a front-row seat to the fastest-growing product in Anthropic's history — and a disarming willingness to say what he actually thinks about the parts that are not working yet.



































































