Every night at 1am, 3am, 4am, and 5am, a set of Claude Code agents build something, grade themselves against a rubric, and write their own rules before the next run.
Posted
2 days ago
Duration
Format
Tutorial
educational
Views
1.2K
45 likes
57 · 43
Big Idea
The argument in one line.
A self-improving agent repeats one loop, make it, grade it against a rubric, write down the lesson, and running that loop nightly across web pages, reels, and essays turns slow manual iteration into compounding overnight progress.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You run Claude Code regularly and want it producing finished output overnight instead of only while you're actively prompting it.
You publish recurring content, web pages, reels, or newsletters, and keep redoing the same back-and-forth feedback with Claude by hand.
You want a concrete rubric for judging whether AI-generated design or writing is actually improving, not just different.
SKIP IF…
You're looking for a no-code tool; this requires setting up scheduled Claude Code runs and writing your own grading criteria.
You expect results from one run; the speaker says it took two mediocre nights before the agent's taste matched his own.
TL;DR
The full version, fast.
Duncan Rogoff builds self-improving Claude Code agents on a nightly schedule: 1am web pages, 3am Reels, 4am Substack writing, 5am long-form video. Each agent makes something, a critic agent grades it on a ten-point rubric (clarity, visual craft, motion, topic fit, layout, typography, restraint, wow factor, narrative), then the agent writes itself a rule from what went wrong so the next night starts smarter. The loop comes from Andrej Karpathy's open-source autoresearch project (change one thing, test it, keep it only if the score improves). By night three, the automated web page beat Rogoff's best hand-built page, 77 to 72.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Defines the core loop: an agent makes something, grades its own work, writes down what it learned, and tries to improve the next time.
00:41 – 01:13
02 · Agents That Work Overnight
The system runs on a fixed schedule, not continuously, staging web pages, shorts, Substack writing, and long-form video across different overnight hours.
01:13 – 02:55
03 · Karpathy's Autoresearch Explained
Andrej Karpathy's open-source autoresearch loop: change one thing in the code, test it, score it, keep it only if it's better. 700 ideas tried overnight, ~20 kept, 11% faster.
02:55 – 04:01
04 · The Learning Loop for Content
The same loop generalizes beyond code: give an agent something you care about, let it score its own output against competitors and its own criteria, and have it improve on its own.
04:01 – 05:28
05 · How to Automate Web Pages
Every night at 1am an agent builds a web page from scratch. Night one was clunky; by night three the design and animation had meaningfully improved.
05:28 – 07:28
06 · Training the Agent's Taste
Manual back-and-forth with Claude works because a human supplies taste turn by turn; training the agent's taste upfront replaces that loop so judgment doesn't require a human present.
07:28 – 10:04
07 · How to Score Design Quality
A critic agent grades every page on ten criteria: clarity, visual craft, motion, topic fit, layout, typography, interaction, restraint, wow factor, and narrative.
10:04 – 10:41
08 · One Focus Area Per Night
Improving every dimension at once wastes tokens and underperforms; cycling through one focus area per night keeps each run specific enough to move the score.
10:41 – 11:04
09 · Agents Writing Their Own Rules
When the agent's own grading catches a mistake, it writes itself a rule so the same error doesn't recur on a later run.
11:04 – 12:27
10 · Self-Improving Instagram Reels
The same pattern applied to Instagram Reels: the first automated attempt was weak, and hand-given feedback corrected it into usable motion graphics.
12:27 – 13:01
11 · Long-Form Video Next
The loop is being extended to AI-generated long-form video next, following the same make-grade-record pattern used for pages and reels.
13:01 – 14:03
12 · How to Add Feedback Docs
A markdown feedback doc, updated by voice dictation, lets the creator steer the agent's direction; the agent reads it first each night before resuming its own self-improvement cycle.
14:03 – 15:08
13 · Self-Improving Substack Writing
For writing, the agent is restricted to draw only from real stories, coaching calls, and stated beliefs from a personal knowledge base, specifically to avoid sounding like generic AI writing.
15:08 – 15:25
14 · Final Thoughts
Closing thesis and call to action: pick one thing you make, give it a judge and a rulebook, let it run tonight.
Atomic Insights
Lines worth screenshotting.
A self-improving agent just repeats one loop: make something, grade its own work, write down the lesson, try again better the next time.
Running Karpathy's autoresearch loop overnight, 700 code changes were tried, about 20 were kept, and the result was 11% faster than making the same changes by hand.
By the third night of automated runs, an AI-built web page outscored the best page its creator had built manually, 77 to 72.
A critic agent scoring AI-generated web pages on ten criteria caught a page built on an unrelated topic that had no fit for the creator's actual channel.
Letting an agent try to improve every area of a design at once wastes tokens and produces worse results than giving it one focus area per night.
Serif fonts read as generic 'AI slop' to an audience trained on AI-generated sites, so swapping typography was one of the highest-leverage fixes made.
When an agent's output goes wrong, having it write itself a rule prevents the same mistake from repeating the next night, compounding its own judgment over time.
A feedback doc, a markdown file updated by voice note, lets a creator steer an overnight agent without sitting through every run.
The same self-improving loop that upgrades a web page also applies to Instagram Reels and newsletter writing, but each domain needs its own hand-trained taste first.
Restricting a writing agent to only pull from real stories, coaching calls, and stated beliefs is a direct hedge against an audience that already hates generic AI writing.
Takeaway
Give the agent a rubric, then let it grade itself overnight.
WHAT TO LEARN
A recurring task gets better on its own once it's wrapped in one loop: produce something, score it against a fixed rubric, write down the lesson, and run that cycle on a schedule instead of only when you're at the keyboard.
01What Is a Self-Improving Agent
A self-improving agent just repeats one loop: make something, grade its own work, write down the lesson, try again better the next time.
The loop is deliberately simple; it's a repeatable pattern that applies to any recurring task you want to improve over time, not a complex system.
02Agents That Work Overnight
The system runs on a set schedule rather than continuously, which is what makes it behave like an agent instead of a standing process.
A single schedule can stage multiple content types overnight without manual triggering, so the operator never has to be present to kick off a run.
03Karpathy's Autoresearch Explained
Karpathy's autoresearch loop is four steps: change one thing, run a test, score it, and keep the change only if the score improved.
Trying roughly 700 small changes in a night and keeping about 20 produced an 11% improvement, purely from compounding small kept changes over time.
04The Learning Loop for Content
The autoresearch loop generalizes beyond code: give an agent something you care about, let it score its own output, and have it improve on its own.
Applying the loop to content means scoring output against competitors in the niche and against stated criteria, not just against a code benchmark.
05How to Automate Web Pages
The first automated run of a new task is usually clunky; real improvement shows up only after a few runs, once the agent has rules to work from.
06Training the Agent's Taste
Manually iterating with an AI tool works because a human supplies taste turn by turn, but it costs that human's time on every single revision.
Training an agent's taste upfront replaces the manual back-and-forth: the agent learns to judge output the way its operator would, without a human present.
Concrete taste notes, like requiring legible fonts, real contrast, and reserved layout space, move a quality score more than vague feedback does.
07How to Score Design Quality
A fixed, multi-point rubric turns subjective quality feedback into a repeatable grading pass instead of a fresh judgment call every time.
Output can be visually strong and still fail a basic fit check, so topic or audience fit deserves its own explicit line in any rubric.
Generic, overused stylistic defaults are worth grading against explicitly, since they're a visible tell that output isn't actually improving.
Restraint deserves its own rubric line because models tend to add more to an output, not less, which usually works against clarity.
08One Focus Area Per Night
Asking an agent to improve everything at once spreads its attention thin and wastes resources without meaningfully moving any single score.
Cycling through one focus area per run, instead of trying to fix everything, keeps each pass specific enough to produce a real improvement.
09Agents Writing Their Own Rules
When a run's grading catches a mistake, having the agent write itself a rule prevents that same error from repeating on a later run.
10Self-Improving Instagram Reels
Porting a working pattern to a new task still requires an initial round of hand correction before the agent's judgment matches its operator's.
12How to Add Feedback Docs
A feedback doc that the agent reads first each run, before resuming its own self-improvement cycle, lets a human steer direction without sitting through every pass.
Capturing feedback by voice note into a simple file lowers the friction of steering the system, so correction happens in the moment it's noticed.
13Self-Improving Substack Writing
Restricting a writing agent to only pull from real stories and stated beliefs is a direct hedge against output that reads as generic AI writing.
Training a writing agent on a personal knowledge base, not just prompts, is what lets it produce output that sounds like a specific person rather than anyone.
Glossary
Terms worth knowing.
Autoresearch
Andrej Karpathy's open-source project where an AI model tries to improve its own training code overnight, testing one change at a time and keeping only the changes that measurably help.
Critic agent
A separate AI agent whose only job is to score another agent's output against a fixed rubric, rather than produce the output itself.
Feedback doc
A markdown file a creator updates with notes or voice dictation; an overnight agent reads it first and implements the requested changes before running its own self-improvement cycle.
Self-improving agent
An AI agent that completes a task, evaluates its own result, records what it learned, and uses that record to perform better the next time it runs.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
analogystory
So recently I've been building these self -improving AI agents that work while I sleep. And just this week, something finally clicked for me and I wanted to share it with all of you guys. I did put together a whole free guide for all of this with a bunch of prompts and stuff that you can use.
If you want to get access to it, it's in the description, but I don't want to take any more time. So let's just get into it. So before I show you all the cool stuff that I've been building, you need a little bit of context and I'm just going to give it to you.
So really like what is a self -improving agent? It's basically an agent that make something.
It grades its own work. It writes down what it learned, and then it tries to get better the next time it does that thing. That's it.
And so now I've recently been setting these up to work at night while I sleep. So what's pretty cool is that this is not on all the time. It actually runs on a schedule that I set.
This is what makes it like an AI agent, right? So at 1 a .m. when I'm asleep, my computer wakes up.
It works on the web pages. At 3 a .m., it's going to work on the shorts. At 4 a .m., It's going to work on my sub stack at 5 a .m.
It can work on my long form video. So literally every single morning I can wake up to improved systems. And that is really like the core idea of this video right here.
Let these agents work while you sleep, not while you're working. So this really all started back in March. And you guys might remember this moment with Andre Karpathy when he published auto research.
Karpathy was a founding member of OpenAI and he worked at Tesla. Now he's at Anthropic and he's kind of a big deal. And this post that he put out.
out about auto research got 11 million views which is absolutely crazy basically his goal was just to have an ai model that could train itself and get better and faster over time without humans needing to do anything so when this happened like because we had all these views and this concept is so crazy like we all knew this was a big deal and i knew it was a big deal but i didn't like quite understand it and i didn't quite understand how to make it actually work for me he actually released this entire github hub repo, open source, totally free.
If you want it like this is the link right here. And his idea is that it just trains this model while he's asleep. And all it does is every single night, it just repeats these four steps.
So one, it changes one thing in the code, then it runs a test. It basically runs a five minute test and it gets a score. If the score is better than it was previously, it keeps the change.
If it was worse, it throws the change away. And then you repeat this process with the next test. That's exactly how it works.
And I promise. to you, there's a reason I'm telling you all this. And so basically he found that the agents were actually able to make improvements on their own.
So he tried 700 ideas in a single night, about 20 of them they kept, and it was 11 % faster than making these changes by hand. And this was just right away, right? So imagine the next night you try 700 ideas and maybe you keep 50 and maybe it gets 15 % faster.
It just improves over time. And so I knew this was super, super powerful in March, but it really only clicked. for me this week.
So really, you don't actually need like auto research or this GitHub repo. It's just this idea, this concept that you give your agent something that you care about and then you have it get better on its own. So this type of learning loop literally works for anything.
So I realized for me, I wanted to apply it to content like a web page, a video, my writing. Right. And so basically I'll make a web page like the one you are looking at right now.
The agent will score it against. other creators in the niche against its own criteria. It's going to figure out how it can get better.
I'm going to talk about how exactly it does that and what things it's looking at. It decides what to fix it. It keeps the best parts and then it grades itself and then it creates new rules.
to pick up where it left off the next night. And so every single night, it starts out better than the night before. And so the first thing I had the agent get to work on was web pages.
Now, why did I do this? I wanted to make more YouTube videos like the one that I am making now. And I wanted to wake up every single morning with a beautiful web page built on a current topic in the AI space so that I could show up to my computer, the website's built, I can learn a little bit about the topic, I can record my video, I can post it, and we're good to go.
And then every single time, I come back to it, the site gets better. The information gets easier to present and clearer for you guys.
And it just makes the whole process that much more faster and that much more seamless. But they're actually hard to create something beautiful and visually interesting and dynamic that tells the right story that gives you guys good information. Like there's a lot of pieces to this.
Right. And I still think it could get 100 percent better. Like this is night three for me.
Right. So basically what happens like every single night at 1 a .m., an agent is going to build a Web page. And the first night that I did it, it looked like this and it was like bad.
Like it had some good concepts. It was all about quad mods. It had some cool animation, but it was a little bit clunky.
It didn't feel very mature or premium and it felt very linear. Like it just had this like sidebar down the side that you were scrolling on. So like it was fine.
But by night three, we're really starting to get somewhere. Like we're improving the design significantly. The animations are getting like really plussed up.
It's starting to look. pretty cool and it's starting to look on brand and feel like me which is arguably one of the most important parts the other thing that is cool about this before we move on is that these concepts actually apply to any future websites that I build myself like not that the agent does like if I'm building something by hand and I say hey I need a website for this thing like these same rules and standards will apply So by night three, it actually started to beat the best pages that I was making by hand.
And what do I mean by hand? And why does that typically yield better results? So by hand, you give Claude like a prompt, right?
Hey, I want to build a website about Claude mods. And Claude comes, it makes you a website. It gives it back to you and you say, hey, this looks good, but the colors are off or the animations could be cooler or I don't really like the text, right?
So it starts to get a little bit better, but you're having this back and forth with Claude, which is great because you as the human like. taste and you know what you want and so that's why you're able to get really good results but then you have to sit there and it takes a lot of your time but what if you could hand all of that thinking over to an agent and the agent is going to base its judgment on your own personal taste so now you have trained the agent to think like you and give feedback like you so the second night that i did this you can see here we dropped off significantly because didn't really have any training to go on and it made like a pretty mid website but then by night three the score started to increase and we started to have like some really nice outputs here right so these are some of the notes that i gave to claw that actually changed it the most like one of the things to look at is the fonts like bold fonts are they legible do we have dynamic colors right is there enough contrast like you can see this first one was like kind of boring like this kind of blended in with the background so i said oh
We give it more contrast. Okay, we're starting to get somewhere. Sizing, how big do you want things to be in the screen?
So big is good, but for me, I like to make these websites like for these YouTube videos, but I need a little more space to put my little head in the corner. So I would argue that this is the most important part. This is the part that's probably the hardest and took me the longest to figure out.
And so paying attention to this may save you the most trouble overall. So the way that I figured these out was actually by building these sites by hand and thinking about which aspects of the site I was personally critiquing. One of the first ones was clarity.
Like I actually have found that quad, especially when it writes, has a tendency to be really, really vague and kind of abstract. Like it uses words like it or she. And you're like, all right, but like, who is she?
Like, what are you talking about? Right. So clarity of information, because if I'm going through these sites and I'm trying to read them to you guys.
It needs to make sense, not just to me, but to everybody who's following along. So is it clear? Does it make sense?
OK, the visual craft, right? Like, what does this actually look like? Can we make this more stylized and more dynamic?
Can we make it look sexier? We have these nice glows now and these frosted glass panels, right? Motion.
OK, so we started out with like things fading on and off. And now we are starting to get to these cool little diagrams, these little glowy orbs traveling along the lines. And this is just kind of cool.
Like, is it better for clarity? Not really, but it's definitely more interesting to look at. Okay, this one was huge.
And I actually only noticed this after the second night because I woke up in the morning to actually a pretty cool website, but it was about. Pi. And you all know that my YouTube channel is all about cloud and cloud code.
And so this site was pretty cool, but it didn't have a fit for my channel or for my audience, and I couldn't really use it. So we had to work on topic fit. So layout and composition, where does the text go?
Where do the graphics go? Where do our buttons go? Do we use images in the background, right?
All of these things. And again, you can see it's starting to get better, like big numbers. People love big numbers.
Trust me, let me tell you that. And so there are still like some issues. with this but again like the idea is that we work from here and we get better And I don't have to do anything.
So typography, what does the text look like? We all know that this is called a serif font with the little lines at the bottoms and tops of letters. And this font just screams AI.
Like we see it all over the place from generic AI slop sites, right? So can we change the typography? Can we change the interaction?
Like you can see how I'm clicking through these buttons. This is cool, right? But at first, like we didn't have any of that.
It had to get built in. Okay, restraint. Like sometimes these models have a tendency to throw the kitchen sink at a page and put everything.
everything on it which looks cool but again doesn't help with this idea of clarity right so what can we strip away to make the message clear and then of course like wow factor like we want some cool moments right like a sexy rendered image or this cool light glow coming down from the ceiling and the last one is narrative Like I said, I'm going to show up and present this on YouTube.
So is it really clear? Like, is there a clear beginning, middle and end? Is there a story and a flow?
Does this read like I am speaking to someone else? Because that's the whole point for me. So again, this is what I'm doing for me.
I think if you're building websites, these are some pretty good things to think about. But I imagine at this point, there are probably some ideas in your head about how you can make this a little bit more specific to you. And so one of the things that I noticed people doing is they actually lack focus in their direction and what they want their agents to work on.
And so what I did is that every single night, the agent will cycle through seven of the areas and it'll just pick one to work on that night. Because if it's trying to improve every single area of the site, like you're gonna end up wasting a lot of tokens and it's not really gonna improve anything all that much at all because it's just gonna get kind of overwhelmed with information.
So like tonight, maybe it's going to focus on typography. Maybe the next night it's gonna focus on clarity of message. The next night it's gonna focus on motion design, right?
So it's just giving it a new task to focus on every day. single night so when something goes wrong it's actually just going to write a rule for itself so it doesn't make that same mistake again because remember the agent is creating a web page and then it's checking its own work and if it notices something it's off it's going to update its rule so that the next time it goes to build a page it doesn't do that again But what's even crazier is that I realized that I don't just have to use this for web pages.
Like I can use it for my short form videos. One of my videos last week, I show you how I create these Instagram reels completely on autopilot. And I realized that I could use the same concept of a self -improving agent to make my videos better, more interesting, more engaging over time.
So I just ran the first test literally today just to see if it worked. And the first test like didn't work great. And I went back and I said, okay, like some things aren't working.
here, but we can do better. And I told it what to fix. And now we have these improved motion graphics.
And I'm really, really happy with this. And so at first, you do have to train these systems by hand so it understands what your personal tastes and rules are. and so we can see that this stack is okay like it's very linear it's straight up but over here in the new staircase when these pieces animate one at a time we have a little bit of bounce and glow and a little pop then in the background you can actually see one two three four and five these nice big numbers so just trying to get 10 better every single night But for these videos, in the same way that we think about the website, like you don't just have to think about the visuals.
Like you can think about, can we do better hooks? Can we write better scripts? Is there something else we can do to show better screenshots, right?
So again, we're just trying to improve all of the areas of what make an engaging Instagram reel. The next thing I'm looking at is doing the same thing for my long form videos. I haven't really even talked about this on the channel yet.
I'm playing around with these AI generated videos, which are coming out pretty good. And I think we can improve them over time. So what's pretty cool is that this is not on all the time.
It actually runs on a schedule that I set. This is what makes it like an AI agent, right? So at 1 a .m.
when I'm asleep, my computer wakes up. It works on the web pages. At 3 a .m.
it's going to work on the shorts. At 4 a .m. it's going to work on my sub stack.
At 5 a .m. can work on my long form video. So literally every single morning I can wake up to improved systems.
And that is really like the core idea of this video right here. Let these agents work while you sleep, not while you're working. The other thing I got an idea for literally while making this website is this idea of a feedback doc.
So if I come to my computer in the morning and I wake up and I look at the brand new web page and it's pretty good, but I see some areas that it could actually be better. Well, now there's actually a feedback doc. It's literally this markdown file right here.
And so what I can do is I can just go in and I can hit my little whisper flow button and I can start talking to the agent and say, yada, yada, yada. We need to make this better and prettier and more dynamic. And I hate your text here, et cetera.
Right. So I can literally just send that off and then save this document. And basically when the agent wakes up in the middle of the night, the first thing it's going to do is it's going to read my feedback and implement the feedback first before it goes about its own self -improving process.
So this is a way where I can have a back and forth with this agent, and I don't really need to be in the loop all the time. And so basically now all you have to do is set your tastes and your preferences and what you like once, and then it's going to get better every time. And then you can also give your input into the system to tell it how to get better every single time.
So the last one that I'm doing, and now I'm just doing this to give you ideas for how this might apply to yourself, is I'm going to be doing this to Substack. And I literally just started writing on Substack. The whole idea is that we know like people hate AI writing and a lot of these companies are like coming down on AI swap.
But the whole goal is to not like trick anyone to think that my AI writing is good or that I'm creating something just on like trending news, right? The whole idea is that I want this agent to like truly be me. Like I've been working on my Obsidian Vault.
or second brain like I know a lot of people have, right? So I'm training this agent to understand like my story, the interactions I have with the community, my personal beliefs and preferences, the way I like to talk and write, the topics I like to talk about, my philosophy, right? all of these things.
And so we train it on that and we train it on what other creators are talking about and what's working for them. Because I can tell my story like as much as I possibly can, but like some things are still true. Like you need a great title, you need a catchy cover image, you need a strong hook, right?
Like those things are kind of like universal truths on the internet. So those are the types of things that the agent can improve over time. So if you enjoyed content like this, make sure you subscribe to the channel.
If you want to learn how to build anything with cloud code and get access to all my resources already done for you, just check the link in the description. If you want to see how I use Opus 5 .5 to create automatic Instagram reels, check out this video right here. I'll see you.
The Hook
The bait, then the rug-pull.
Duncan Rogoff opens by naming the engine behind the video you're watching: every night at 1am, an AI agent writes a new web page, grades its own work against a rubric, and files away what it learned so the next version starts ahead of where the last one left off.
Frameworks
Named ideas worth stealing.
00:20model
The Self-Improving Loop
Make
Grade
Write the lesson
Next night
The core four-step cycle: an agent makes something, grades its own work, writes down what it learned, and tries to get better the next time it does that thing.
Steal forany recurring content or workflow you want to improve without manually re-prompting every time
01:36list
Karpathy's Autoresearch Loop
Change one thing in the code
Run a test
Score the result
Keep it if better, discard if worse
The four-step loop Andrej Karpathy's open-source autoresearch project repeats overnight to improve a model's training code without human intervention.
Steal forany automated A/B-test-and-keep system
07:28list
The 10-Point Critic Rubric
Clarity
Visual craft
Motion
Topic fit
Layout & composition
Typography
Interaction
Restraint
Wow factor
Narrative
The ten criteria a critic agent uses to grade every automated web page, scored 0-10 each.
Steal fora quality gate for any AI-generated design or content output
00:41list
The Nightly Schedule
1am: web pages
3am: Instagram Reels
4am: Substack writing
5am: long-form video
The fixed overnight schedule that stages each content type so agents work while the creator sleeps rather than while he's at the computer.
Steal forany multi-format content operation run by one person
CTA Breakdown
How they asked for the click.
VERBAL ASK
15:08product
“Pick one thing you make. Give it a judge and a rulebook. Let it run tonight. ... Join Claude Code Club, build anything with Claude Code, only $9.”
Reframes the video's own thesis into the call to action, paired with an on-screen 'guessing vs building' before/after graphic pitching a $9 paid community.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
A screen-recorded walkthrough of building a custom Claude Code skill that watches any viral video, breaks it into timestamped beats, and recreates it with Higgsfield's Seedance 2.0.
A three-level walkthrough of prompting, saving a design system, and layering in AI-rendered images to turn one topic into a branded Instagram carousel.
Andrej Karpathy says the best way to brief an AI coding agent is a messy ten-minute voice ramble — so a creator tests it live by building a naturopath website from a single, uninformed viewer request.
A former art director breaks a single sentence into a seven-piece "goal prompt" and gets Fable 5 to autonomously build a playable game, a 3D website, and a cinematic video in one session.