The argument in one line.
Anthropic's own production data confirms that AI has crossed from coding assistant to primary engineer, and the last human advantage — research taste and the ability to have genuinely novel ideas — is the only thing left standing between compounding automation and full recursive self-improvement.
Read if. Skip if.
- You write code professionally and want a grounded read on where AI assistance is actually headed, backed by internal Anthropic data rather than speculation.
- You're building AI products and want to understand the productivity math behind 8x code output with only 4x perceived value gains.
- You follow the AI safety and alignment conversation and want to stress-test Anthropic's 'slow down' argument with a skeptical lens.
- You're curious about what 'research taste' means as an economic asset when execution is becoming fully automated.
- You want a high-level intro to AI — this assumes familiarity with Claude, Claude Code, AGI framing, and the current coding-agent landscape.
- You're looking for hands-on tutorials or workflow demonstrations — this is purely analytical commentary on a research paper.
The full version, fast.
Anthropic published a paper on recursive self-improvement that doubles as an internal progress report: Claude writes 80%+ of their merged code as of May 2026, up from single digits a year earlier; task horizon doubling time has accelerated from seven months to four; and Claude-written code has gone from clearly worse than human to roughly at parity, with better-than-human expected this year. The remaining human edge is research taste — knowing which problems to pursue, which results to trust — not execution. The host's editorial through-line is that Anthropic calling for a global AI slowdown while sitting in first place and using an unreleased internal model (Mythos) to accelerate their own development is structurally self-serving, no matter how accurate their safety framing may be.
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →Where the time goes.

01 · Cold open
Hook: AI is literally building itself, Anthropic says slow down, host calls it self-serving.

02 · Paper framing
Paper intro — RSI is not inevitable, but trending there. One missing ingredient: novel ideas.

03 · Abstraction progression
Diagram: human writes code → chatbots → coding agents → autonomous agents → closing the loop (no human).

04 · Sponsor + task horizons
DigitalOcean ad. Then: task horizon doubling time from 7 months to 4. Opus 3 (4 min) → Sonnet 3.7 (90 min) → Opus 4.6 (12 hr).

05 · CoreBench + research gap
AI reproducing novel research: 20% (2024) to near 100% (15 months later). But origination — novel ideas — still the missing ingredient.

06 · Engineering vs. research
Two tracks: engineering (code, infra, model training) where AI dominates; research (deciding what to do, interpreting results) where humans still lead.

07 · 80% of code is Claude-written
As of May 2026, 80%+ of merged Anthropic code is Claude-authored. Before Claude Code launched (Feb 2025): low single digits.

08 · Lines of code per engineer
The chart: Q3/Q4 2025 explosion. 8x code output per engineer in Q2 2026. Anthropic caveats this is imperfect — measures quantity not quality.

09 · Mythos and the banned competitor
Anthropic cut x.AI's access to their models, then used their own unreleased Mythos internally. Host's read: winning the race quietly while saying 'we should slow down'.

10 · Productivity paradox
8x code, 4x perceived productivity. Median Anthropic employee: 4x more output with Mythos Preview. The gap implies Claude code is half as valuable per line.

11 · Understanding vs. thinking
'You can outsource your thinking but not your understanding.' Claude judges Claude-written code. Humans increasingly disconnected from what they're building.

12 · Human role narrowing
Once code quality reaches parity, humans stop writing and only review. But review rate can't match generation rate — human review becomes the next bottleneck.

13 · Ideas vs. execution
Edison: 1% inspiration, 99% perspiration. Perspiration is now automated. Does that make inspiration more or less valuable? Research taste is still human territory.

14 · Three futures
Future 1: trend stalls (unlikely). Future 2: compounding automation, humans set direction. Future 3: full RSI — compute is the only bottleneck, capital wins forever.

15 · Anthropic's slow-down argument
Critique: calling for a global slowdown while leading the race, using an unreleased internal model, and having banned competitors is structurally self-serving — even if the safety reasoning is sound.
Lines worth screenshotting.
- AI task horizon at Anthropic is doubling every four months, down from seven months — the acceleration is itself accelerating.
- 80% of Anthropic's merged code is now Claude-authored, up from low single digits before Claude Code launched in February 2025.
- 8x code output with only 4x perceived productivity gain means Claude-written code produces roughly half the value per line of human-written code.
- Claude-written code quality was clearly below human in late 2025, is at parity today, and is expected to exceed human quality within the year.
- The only remaining human comparative advantage in AI development is research taste: choosing which problems matter, which results to trust, which direction to abandon.
- Novel ideas are definitionally not possible from a model trained on existing data — the missing ingredient for full recursive self-improvement is not execution capability, it's genuine novelty.
- Once human and Claude code quality reach parity, humans stop writing code and shift to reviewing it — but humans cannot review as fast as Claude generates, making review the next bottleneck.
- You can outsource your thinking to AI, but you cannot outsource your understanding — and as humans become more abstracted from the systems they build, comprehension failure becomes the alignment risk.
- Anthropic banned x.AI from using their models internally, then used their own unreleased Mythos model to accelerate internal development — calling for a slowdown from that position is a structurally privileged argument.
- In a full recursive self-improvement scenario, the only bottleneck is compute, which means capital becomes the only moat — whoever holds compute at the moment RSI triggers stays permanently ahead.
- A 130-person Anthropic poll found median 4x productivity gains with Mythos Preview — half what the 8x code output number would imply, suggesting much of the additional code is being discarded or requires rework.
- AI systems already succeed at reproducing novel research papers at near 100%, up from 20% two years ago — the gap between replication and origination is the last meaningful frontier.
- A global AI slowdown would require every well-resourced lab in every country to agree and be independently verifiable — which is harder than nuclear arms control because training runs are far easier to conceal than missile silos.
- Anthropic argues that even if recursive self-improvement never happens, today's AI capabilities are already so underutilized that major world changes will occur from capability overhang alone.
The human advantage is narrowing faster than most planned for.
Anthropic's own internal data draws a clear line from 'humans write code' to 'Claude writes 80% of code' — and the remaining human edge, research taste and judgment, is already being measured and shrinking.
- AI task horizon at Anthropic is doubling every four months, down from seven months — the acceleration is itself accelerating, not just the capability.
- 80% of Anthropic's merged code is now Claude-authored, up from low single digits before Claude Code launched in February 2025 — a full order-of-magnitude shift in under a year.
- 8x code output with only 4x perceived productivity gain means Claude-written code produces roughly half the value per line of human-written code — more output isn't the same as more value.
- The remaining human comparative advantage is research taste: choosing which problems matter, which results to trust, and which directions to abandon — not execution.
- Novel ideas are definitionally not something a derivative model can originate — the gap between replication (now near 100%) and origination is the last structural moat.
- Once AI and human code quality reach parity, humans stop writing and shift to reviewing — but review speed cannot match generation speed, making human review the next bottleneck.
- Becoming disconnected from the systems you're building is itself an alignment risk: you can outsource thinking, but you cannot outsource understanding.
- In a full recursive self-improvement scenario, compute becomes the only bottleneck — which means capital becomes the only moat, and whoever holds compute at that moment stays permanently ahead.
- Even if AI capabilities stopped improving today, capability overhang — the gap between what models can do and what workflows have caught up to use — would still produce major economic disruption.
- Calling for a global AI slowdown from an undisputed first-place position, while using an unreleased internal model, is structurally self-serving — even if the safety reasoning behind it is correct.
Terms worth knowing.
- Recursive self-improvement (RSI)
- A hypothetical stage where an AI system becomes capable of designing and training its own successor models, removing humans from the loop entirely. The only remaining bottleneck would be available compute.
- CoreBench
- A benchmark that tests whether an AI can read a published research paper and successfully reproduce its experimental results from scratch — measuring ability to execute on described novel methods, not to originate them.
- Mythos
- An internal Anthropic model, more capable than publicly released models, used internally to accelerate their own development but never released to outside developers.
- Task horizon
- A measure of how long a skilled human would take to complete the same task an AI can reliably finish autonomously — used to compare capability across model generations over time.
- Capability overhang
- The gap between what AI models are technically capable of and what the surrounding systems, businesses, and workflows have caught up to using — the idea that even frozen model capability would still produce major economic disruption.
- Research taste
- The judgment required to choose which experiments are worth running, which directions to abandon, and which results to trust — distinct from the ability to execute experiments, which AI increasingly handles.
- Permanent underclass
- The concept that at the moment of full recursive self-improvement, whoever controls compute (and therefore AI capability) locks in a permanent advantage — those without capital at that moment have no mechanism to close the gap.
Things they pointed at.
Lines you could clip.
“You can outsource your thinking, but you cannot outsource your understanding.”
“As of May 2026, more than 80% of the code we merged into Anthropic's codebase was authored by Claude.”
“Work and life ran on a gift economy of small favors between humans. 'Can you help me get this script running?' — each one created a little debt, a little mutual awareness. Claude has eaten the favors.”
“Ideas are the important part. Interesting to think about.”
“If you're an Olympian and you have 10 other competitors racing the 800-meter dash, and you're in first place halfway through, and you say, 'Hey, guys, why don't we all slow down?' — you're always going to be in first place.”
Word for word.
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
The bait, then the rug-pull.
AI is now literally building itself — and Anthropic published a paper to say so. Matthew Berman walks through that paper line by line, tracking the moment Anthropic's internal data crossed from aspiration to fact: Claude authors more than 80% of the code merged into their codebase as of May 2026, task horizons are doubling every four months, and code quality is at or near human parity. The editorial through-line is one the paper itself doesn't fully reckon with: calling for a global slowdown from a position of undisputed first place is a structurally different argument than it would be from second.
Named ideas worth stealing.
AI Development Abstraction Layers
- Human writes code directly
- Human uses chatbot to assist
- Human delegates to coding agent
- Human prompts autonomous agent swarm
- Agents train their successors (recursive self-improvement — not yet reached)
Anthropic's diagram of how humans have become progressively more abstracted from the actual AI development process, with the Claude logo growing denser (more capable) at each stage.
Engineering vs. Research Split
- Engineering track: writing code, standing up infrastructure, overseeing model training — AI dominant
- Research track: deciding experiments, interpreting results, choosing direction — still human
The two-track framework Anthropic uses to distinguish where AI already dominates from where human judgment still leads.
Three Futures
- Trend stalls: AI capabilities plateau, but capability overhang still reshapes the world
- Compounding automation: development substantially automated, humans retain research taste and direction-setting, 100-person company does work of 100,000
- Full RSI: AI builds its successors, compute is the only bottleneck, capital freezes class structure permanently
Anthropic's three scenario framework for how AI development could unfold, ranging from stall to intelligence explosion.
How they asked for the click.
“I just made an entire video about this specific topic. Check it out right here.”
Standard YouTube end-card CTA pointing to a related video about AI and the job market.

































































