Turn Claude Into a Video Editing GENIUS (in 3 Simple Steps)
A RoboNuggets creator walks through the three-step Claude Sonnet 5.5 workflow he uses to cut, animate, and direct his own videos end to end, with no separate AI video-generation model involved.
Posted
6 days ago
Duration
Format
Tutorial
educational
Views
74.4K
1.1K likes
57 · 43
Big Idea
The argument in one line.
Claude Sonnet 5.5 can replace a dedicated AI video-generation model for motion graphics and editing if you feed it three things in order: a clean transcript-based cut, a codified design system to animate against, and iterative feedback to direct the result to production quality.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You make talking-head or tutorial videos and want to cut editing and motion-graphics costs by doing the work inside Claude instead of a separate AI video tool.
You already use Claude Code or a similar coding agent for other work and want to extend that same tool into video production.
You're comfortable building and refining a reusable prompt or skill file over several rounds of feedback rather than expecting one perfect output.
SKIP IF…
You want a finished, one-click video tool. This is a workflow built from raw Claude prompts and a custom skill file, not a packaged product.
You have no existing footage or transcript workflow. The method assumes you're already recording and need the cut, animate, and direct steps, not sourcing ideas or B-roll.
TL;DR
The full version, fast.
Jay argues Claude Sonnet 5.5, released days before this video at roughly half the token cost of Opus 5.5, is now good enough to handle both the cuts and the motion-graphics animation in his videos without any separate AI video-generation model. His process has three steps. First, transcribe the raw footage with AssemblyAI, free local Whisper, or ElevenLabs, then have Claude identify and keep the best takes. Second, animate the cut footage, but only after codifying a specific design system, borrowed from sites like styles.referral.design and skillery.dev or reverse-engineered from another creator's video with ffmpeg screenshots, because a generic prompt produces generic 'AI slop.' Third, direct the output by reviewing drafts in a timestamped comment-based editor, batching feedback back to Claude, and folding the lessons into a persistent /animate skill that gets sharper with every video.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Showcase of Sonnet 5.5-made cuts and motion graphics, a preview of different design styles (flat vector, isometric, 3D, glassmorphism, cutout), and Jay's quick introduction.
01:39 – 04:44
02 · Context: why Sonnet 5.5 changes this
Sonnet 5.5 as the cheaper sibling of Opus 5.5, independent testing confirming its quality, the fact that every animation is generated as editable JS/HTML code, and credit to Vincent Wei for the original design themes.
04:44 – 08:44
03 · Step 1: cut
Transcribe raw footage with AssemblyAI, free Whisper, or ElevenLabs, have Claude identify and keep the best takes, and review cuts through the Rubric editor with timestamped comments.
08:44 – 12:28
04 · Step 2: animate
Why generic prompts produce 'AI slop,' sourcing design systems from styles.referral.design and skillery.dev, and reverse-engineering an uncredited style from a video using ffmpeg contact sheets.
12:28 – 15:16
05 · Step 3: direct
Why a one-shot animation prompt lands around 80% there, giving timestamped feedback through the Rubric editor, and folding every lesson into an updated /animate skill.
15:16 – 15:46
06 · Wrap up
Recap of the three-step system and a subscribe ask.
Atomic Insights
Lines worth screenshotting.
Claude Sonnet 5.5 runs at roughly half the token cost of Opus 5.5, making AI-generated motion graphics cheap enough to produce for every video instead of just a showcase piece.
Every animation Claude generates for video is built from JavaScript and HTML under the hood, so individual visual elements can be isolated, edited, or turned off independently instead of treated as an opaque rendered clip.
Animating footage with no design system produces generic results. The same default visual style showing up everywhere is becoming the new 'AI slop' that signals an unedited prompt.
A design system only has to be dialed in once. After that it can be codified into a reusable skill and invoked by name on every future video.
Free local Whisper transcribes an hour of audio in about 15 minutes using your own CPU, while a paid service like AssemblyAI transcribes the same hour in about 42 seconds.
Feeding Claude a competitor's video and asking it to pull frames into a contact sheet with ffmpeg is enough for it to reverse-engineer and replicate a design system that was never shared publicly.
A single one-shot prompt to animate footage against a design system typically lands around 80% of the way to a usable result. The remaining 20% requires direct, iterative feedback.
Routing feedback through a timestamped, click-to-comment review tool produces more accurate revisions than describing the same changes in a single chat thread.
The real deliverable of a feedback round isn't the finished video. It's updating the skill file with what was learned, so the next video starts further along.
Reference sites like styles.referral.design and skillery.dev exist specifically as swipeable starting points for a design system, each with a starter prompt already attached.
Takeaway
Three steps turn Claude into your video editor, not just your animator.
WHAT TO LEARN
Cutting, animating, and directing are three separate skills, and Claude needs a transcript, a design system, and your feedback, in that order, before it can actually finish a video for you.
02Context: why Sonnet 5.5 changes this
Claude Sonnet 5.5 shipped at roughly half the token cost of Opus 5.5, which matters because independent testing had already found Opus 5.5 excellent at video editing and motion graphics.
Every animation Claude produces is generated as JavaScript and HTML code, not a rendered black box, so individual layers and elements can be isolated or edited after the fact.
A design system doesn't need to be reinvented every time. Once a style works, it can be codified into a skill and reused on every future video.
03Step 1: cut
Before Claude can cut a video it needs a transcript. AssemblyAI (paid, about 42 seconds per hour of audio) and local Whisper (free, about 15 minutes per hour) are the two go-to options.
Reviewing cuts through a timestamped comment interface is faster than a single back-and-forth chat, and batching all the comments into one message back to Claude gets more consistent fixes.
Feedback on cuts can be used to calibrate the editing skill itself, so the system gets more personalized to how a specific person speaks over time.
04Step 2: animate
Animating footage with Claude's defaults and no design system produces generic results that are becoming recognizable as 'AI slop.' Supplying your own design system is the differentiator.
Reference sites like styles.referral.design and skillery.dev are ready-made libraries of design systems and starter prompts that can be copied straight into a prompt.
When a visual style has no public prompt attached, Claude can still replicate it: feed it a link to the source video, pull frames into a contact sheet with ffmpeg, and ask it to extract the design system.
05Step 3: direct
A one-shot animation prompt against a solid design system typically lands around 80% of the way there. Hitting production quality requires a dedicated feedback pass.
Timestamped, click-to-comment feedback produces more accurate revisions than describing the same changes in prose inside a chat thread.
The real deliverable of the direct step isn't the video. It's updating the /animate skill file with everything learned in that session, so the next video starts further along.
Glossary
Terms worth knowing.
SVG (Scalable Vector Graphics)
A vector image format that scales to any size without going blurry, can be opened and recolored by hand, and can be animated directly by Claude.
Design system
A codified, reusable visual style (colors, shapes, motion language) that gets fed to Claude so its animations match a specific aesthetic instead of a generic default.
Skill (Claude skill)
A saved, reusable instruction set that can be invoked by name in a prompt, letting a workflow like 'animate this footage in my style' be repeated without re-explaining it each time.
Contact sheet
A grid of frame grabs pulled from a video with a tool like ffmpeg, used to let an AI model examine a video's visual style at a glance in order to replicate it.
AI slop
Generic, default-looking AI output that becomes instantly recognizable once enough people use the same unedited prompts with no custom design system.
“All of the cuts, all of the animations, and including the motion graphics that you've seen in this video, all done by Sonnet 5.5.”
states the whole premise in one line with proof already on screen→ TikTok hook↗ Tweet quote
09:07
“These are quickly going to become the new AI vibe coded slop that a lot of people would want to avoid, including yourself.”
names the exact fear viewers have about default AI output→ IG reel cold open↗ Tweet quote
12:09
“If you're just going to prompt Claude generically, then you'll probably be disappointed.”
tight, contrarian, no setup needed→ newsletter pull-quote↗ Tweet quote
14:39
“If you want to get to production level quality, then this step three is the most important, where you give feedback to Claude in order to iterate on those animations.”
the thesis payoff after two steps of setup→ TikTok hook↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
So Sonnet 5 .5 got released early this week and it turns out it's quite incredible when it comes to video editing and motion graphic animations like the ones you're seeing on this page. And it's also half the token cost of Opus 5 .5 so you can potentially get these types of animations for cheaper. And just to show you an example, here at the left is a raw footage of me for one of our videos.
And basically I just asked Sonnet 5 .5 with a single prompt that uses this skill called slash animate in my workspace. And it was able to produce this video on the right. 17 is to ask for SVGs.
So when you ask Claude for new icons, illustrations, or diagrams, one of the best things you can do is to ask for them in the SVG format. That stands for scalable vector graphics so that it scales to any size without going blurry. You can open the file and change the color by hand.
And you can even have Claude animate these icons. So with SVGs, every graphic that you get back from Claude is editable and animatable. So there you go.
All of the cuts, all of the animations, and including the motion graphics that you've seen in this video, all done by Sonnet 5 .5. Now, this obviously uses my own design system, the ones you see in our videos. But you can, of course, use Claude Sonnet for other styles of videos, like this flat vector animation, isometric style graphics, 3D render animations like these, or even ones with liquid motion.
Glassmorphism animations, cutout animations like the ones from Vox, and at this point, pretty much any style that you want. And these are all done using Sonnet 5 .5 and the skill called Slash Animate. So today I'll teach you in three steps how to create a skill like this so that you can use Sonnet 5 .5 for video animations without having to rely on any other AI video generation models.
And if you're new, my name is Jay. I spent over a decade working with brands you probably know, have been in AI since my master's in data science, and now I'm leading our AI business in one of the largest AI communities globally. Let's dive into it.
So to give a bit of context, Sonnet 5 .5 got released a few days ago. And it is basically the cheaper version of Entropic's Opus 5 .5 model. And the reason why this matters is because Opus 5 .5 from a lot of independent accounts, including our own internal testing and in our community, is actually really good.
In fact, in a video I did just before this one, I was talking about how good Opus 5 .5 is when it comes to video editing. And so mainly I was curious, with Sonnet 5 .5, given it is half the cost of Opus, can it actually produce videos that is almost as good as Opus or maybe even better? And it turns out that it absolutely can.
So these videos that you're seeing, they're all created with Sonnet 5 .5 without any extra AI video generation model. And its diversity and taste in terms of just generating these videos is pretty incredible. So you can see here that there's different styles that you can do.
You can have flat animations. You can have 3D renderings. You can have this collage or cutout animation and even this glassmorphism style if you want.
And what's important to remember when it comes to all of these videos is that they're all done with code. So you can literally go into these videos and, for example, isolate each of these elements. And if I just turn off the other layers here, you can see that what Sonnet is basically doing is create all of these animations through JavaScript and HTML, and it handles all of the motion graphics and animations under the hood.
So you can see this particular shape morphing into these different styles. That is all done by Sonnet 5 .5, which is pretty incredible. And so theoretically, unlike AI video generation models, which are basically just black boxes that are pretty hard or costly to replicate, if you have a style that is dialed in like this flat vector style, then you can potentially reuse this just by asking Sonnet to turn this into a design system and make that design system part of a skill and use that skill whenever you want to animate videos like these.
And just to show an example of that, remember for this particular raw footage, what I use here is my own design system. But if let's say I use an isometric design system is to ask for SVGs. So when you ask Claude for new icons, illustrations, or diagrams, one of the best things you can do is to ask for them in the SVG format.
So you can see because Sonnet understands the transcript of what I said in the video, it fit that particular design system and style to create animations specific to what is being said. And if you want to try another design system like this flat vector animation, you can see it adjusted what the visuals look like, but it's still maintaining the core of that message, which is about SVGs in this case.
And then this one is another example that uses more 3D elements. So I probably won't change my own design style to this, but the point being that if you have a core theme, you want and you actually codify that into a skill, then you can actually do pretty cool things with this capability that Sonnet 5 .5 now unlocks that is much cheaper versus Opus.
And just to give credit, all of those design themes that you saw, it is care of Vincent Wei, who is a pretty good follow on X because he does a lot of these animations. And basically what he did here is to create these animations via Opus 5 .5. And so I was curious if Sonnet can do the same.
Later in the video, I'll also show you how I was able to codify these design systems just by feeding Claude this video, which effectively, once you learn that, you can pretty much look and get any video from the internet and then feed it to Claude to create your own design system based on that. All right, so what are now the three steps so that you can fully utilize this capability of Sonnet?
Well, the first one is, especially if you are editing a video that has a transcript or spoken words like the talking head video that I was showing earlier, you of course need a way to transcribe the words being said in that video and for Sonnet to actually cut it. so that any pauses or any false takes are excluded from the final edit.
And for you to do this, it's really simple. You can use a prompt like this to get started, where you basically just give the path to the raw footage, and then you ask Claude to transcribe it with Assembly AI. Now, Assembly AI, we're not sponsored by them at all, but we do use them a lot in order to transcribe our videos.
And once you sign a prompt like this, Claude will anyway guide you on how to sign up to it. I believe they're still offering free credits for initial signups, which would last you a long time in terms of transcription. long footage.
And then Claude will just guide you to where to get your API key, which is in the settings of Assembly AI, which will now let you use their services in order to transcribe videos. And by the way, all of the starter prompts and resources that I'll mention in this video, I just put together this guide, which you can just get below.
And this eight -page PDF, you can just send that to Claude, and you can just grab the prompts and the resources that you would need in order to build out the stuff that I'll mention in this lesson. Now, to be clear, Assembly AI is not the only tool that does this. In fact, if you want a free tool, Whisper is another good one.
However, it does take a while because it just uses your own CPU. It is open source, that's why it's free. But an hour of audio takes around 15 to 15 minutes, depending on the strength of your CPU.
And so with Assembly AI, the good news about it is if you have an hour of audio, that takes around 42 seconds, which is one of the fastest from the ones that we've tried out. If you have an 11 Labs subscription, you can actually use that. And that is...
reasonable cost as well per hour of transcription and it also transcribes things quite fast a bit slower versus assembly ai but if you already have an 11 lab subscription then you can just go and use that And just to give you an example of an output, you can see here this footage, which is around a minute and 50 seconds long, that was all transcribed using Assembly AI.
And it gave me this script of what I said. And then Claude basically looked at this script and identified which parts are worth keeping. So if I play this here somewhere in the beginning, this is at 2x speed.
So you can see there I had a false take. And so Cloud, because it now understands the transcript of what I said, it was able to identify what particular take is the best. Or in this case, I just told it to get the latest take.
Now, when you do this and you send that prompt, most likely Cloud will give you the final edit, which is just an MP4 file in your machine. And you can actually give feedback to Cloud in that same chat session. But I just find when it comes to video editing, it's much more easier to work with agents when you have an interface like this.
And just to give you a background, this... rubric editor it is a tool that we're developing for our community and which we also have started using ourselves where let's say Claude had some errors in terms of the cuts that it makes all you can simply do is to adjust this the same way that you would in Premiere and when you do that that will just add a note which you can see here there's a couple of ones here already that Claude fixed for me and if I copy the comments for all of those clips I can then just send all of those comments to Claude And what I usually do is have it calibrate the skill that we made so that the next time that it cuts a video, it gets more refined in terms of how much to cut so that it becomes much smarter and is much more personalized depending on how I speak.
And again, you don't necessarily need to build yourself an interface like this, but I just find that it is much more helpful if you have one, especially if you do a lot of videos. And by the way, if you want to learn how to build and sell AI systems that businesses actually pay for, then that's pretty much all we do over at the RoboNuggets community, where not only do you get access to the Cloud Living Masterclass, which we update every week and takes you from zero to mastery with the latest on AI, but you also get access to our Agents as a Service course, which walks you through how to actually get paid for all these AI skills that you are learning.
learning. You also get to be part of a genuinely great community of AI builders. In fact you can see just some of the recent wins our members are getting from the program right here.
So if you want to start earning from AI then check that just in the pinned comment below. Now back to the video. So now once you have a video that is cut the next step now is for Sonnet to animate that video.
Now, if you ask Claude to animate the video that you've got using a generic prompt, you probably would be disappointed. So for example, that same raw footage that I have, I just asked Claude to make a version using its defaults. So no extra skills, no particular design systems.
And this is the animation that we've got, which is not necessarily bad per se. But the big problem with this is that because these are Sonnet's defaults, these are quickly going to become the new AI vibe coded slop that a lot of people would want to avoid, including yourself. Because AI slop at the end of the day are just beginner users who are just prompting Cloud without any design system.
And so they're getting these defaults that you would see everywhere. And so clearly the way that you would differentiate yourself is to create your own design system. And then when you ask Cloud to animate this video, you can just invoke or declare that design system so that Cloud would understand the specific aesthetic that you want.
Now, if you don't have a design system, one of the best resources that you can bookmark is styles .referral .design. So let's say if you want this slush design system, they actually have a design markdown file here, which you can just copy, send that to Claude and just ask it to create a design system and a skill, let's say slash slush design.
And whenever you type in that skill in your prompt, you'd be able to copy this particular design. Now, apart from styles .referral, which is mainly for website design systems, what I've actually found recently is the emergence of these websites, which are basically collections of motion graphic videos that are made using, for this one, Opus 5 .5, but basically using Cloud.
So this is from skillery .dev. And again, this is just going to be part of that PDF, which you can just grab below. And the great thing about this is you can just browse any style that you want from here.
Not all of them are great, but there are some that are quite unique that might fit the use case that you have. click on one of them, it provides attribution on where the original post is, including a starter prompt that you can just copy. This is another one called whatchips .com, and it has a lot of launch videos for applications in particular.
And if you click through any of these, that will provide you the link to the X post. But the only limitation with this website is that it doesn't necessarily have the prompt to get you started with replicating this design system. So what can you do?
Well, the good news is you can also use Sonnet 5 .5 to effectively have it watch videos for you and actually to replicate that design system based on what it sees. So for example, this one from Vincent Wei that I saw from X, which has done a really good job to create these design themes. And he did mention that he's going to upload the prompts to GitHub in just a bit.
But since that was not available yet, essentially what I did is to give Claude the link to this X post and have it use a tool called FFmpeg in order to get screenshots of this post and analyze those frames in order. to try and replicate them using sonnet 5 .5 and just to show you what i mean i asked claude here to show me the screenshots that it just took with ffmpeg for that x post and you can see here that it basically made this contact sheet where it sort of watched the video by taking screenshots of it and then it picked out the themes and design systems that i wanted from this video and tried to recreate it using sonnet 5 .5 now just for your awareness if in case you haven't used ffmpeg before that is an open source tool have been around for 20 plus years and basically that is a piece of open source software that lets you automate basic video commands like taking screenshots.
So it's a very good fit for agents is the point. Now coming from these screenshots, I just ask Sonnet to create design systems from them. And that is basically how we arrive at these styles that we can now reuse and add to our library.
And so the big takeaway for this step. is if you're just going to prompt Claude generically, then you'll probably be disappointed. What you first need is a real design system that is attuned to whatever your use case is.
But the thing is, even with a design system, and even if your raw footage is properly cut, if you just send one prompt to animate your footage with a design system, it is very likely the case that you won't get the final result that you want. You'll probably get to 80 % there, but for you to actually polish it, you need to get to step three, which is about directing and actually looking at the beats and all of the editing.
animations that it made, and for you to give feedback to any changes that you want made. To give an example of a clip that I had animated, you can see for this clip that I was talking about cloud design. So I just had cloud add in this UI of cloud design, so very meta.
But basically, this is the version two. So this is not a one -shot prompt to get to this point. Because if you look at the version one of this, you can see that essentially...
In the beginning, it was using these line art animations, which I thought was okay, but it can probably be improved, or at least the video can be improved further if we actually show the actual Cloud design UI. And so when you get these types of results from Cloud, which are suboptimal, let's say, and probably 80 % there in terms of the final look that you want, you can, of course, just give feedback to Cloud via the chat interface, which should still work, but similar to what I mentioned before, it is just much easier to have an interface like this where you can put in comments to these videos.
Like for example, for this rubric editor, which is the toolkit that we're making for our community, you can just click on any part of this video. And when I make a comment here, like for example, actually show the cloud design UI that will not only provide the comment, which I said, it also includes in that comment, the timestamp that it's in.
It includes the specific point where that comment was made. And so once we go through the review process and when we're directing this whole animation and we actually want to give feedback to cloud for its animations to us, we can just watch these animations. give comments to the ones that we want changed.
And then what I usually do is I just copy all of those comments, go back to Cloud, and then send the whole thing in one go. And you can see in this second version of this clip, it was able to accommodate that request of ours. It now used the Cloud design UI.
And in my view, it is a much better experience for the viewer to actually see user interfaces that they are familiar with versus those line art infographic animations. that Claude started with. And so that's really the secret because step one and step two, like these steps, you can pretty much automate and you can pretty much rely on a one shot to give you a result that is 80, 90 % there sometimes.
And even with Sonnet 5 .5, it is actually really capable to give you really good animations already. But if you want to get to production level quality, then this step three is the most important where you give feedback to Claude in order to iterate on those animations. Now, what's important here is in that particular session where you have just communicated to Claude the design systems and every feedback that you have, that is where you can ask Claude to make your slash animate skill so that all the learnings in that particular session, all your feedback that you sent will now be codified into this slash animate shortcut.
And if you already have one, then all that remains to do is to have Claude update it based on skill creation best practices. And if you ask your agent to do that, then this skill will be polished more and more to suit your tastes. And so there you go.
That is essentially how I make those videos and how you can rely on Sonnet 5 .5 to create quality animations for you, even if you're just one person. Hopefully that was useful. And as always, thanks for watching until the end.
And consider tapping subscribe below if you haven't yet, because that just helps me a lot to put out more educational stuff like this. I'll see you all next time.
Cheers.
The Hook
The bait, then the rug-pull.
Jay opens with a showcase reel, cuts and motion graphics all produced by Claude Sonnet 5.5 at half the token cost of Opus 5.5, before admitting the real subject of the video isn't the reel itself. It's the three-step system behind it: cut, animate, direct.
Frameworks
Named ideas worth stealing.
04:44list
The 3-Step Video Automation Framework (Cut, Animate, Direct)
Cut: transcribe raw footage and have Claude keep the best takes
Animate: apply a codified design system to the cut footage
Direct: give iterative, timestamped feedback and fold it into a reusable skill
A sequencing framework for handing video production to Claude: each step has to be solid before the next one works, a sloppy cut makes animation harder to judge, and an undirected animation stays stuck at roughly 80% quality.
Steal forany Claude-driven video pipeline (talking-head edits, promo cuts, motion graphics) where the output currently relies on a separate paid AI video tool
CTA Breakdown
How they asked for the click.
VERBAL ASK
08:03product
“If you want to learn how to build and sell AI systems that businesses actually pay for, then that's pretty much all we do over at the RoboNuggets community...”
Mid-roll pitch for the paid RoboNuggets community (Cloud Living Masterclass plus an Agents-as-a-Service course), framed as a direct continuation of the editing workflow being taught, followed by a second explicit subscribe ask in the final seconds.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
A screen-recorded tour of 25 prompts, sites, and Claude Code skills, sorted easy to advanced, for anyone tired of interfaces that scream default AI design.
Mobbin's new MCP connector lets Claude search 600,000+ real, shipped app screens as design references — trading the generic serif-and-purple-gradient AI look for interfaces grounded in production apps.
How TypeSafe's System 1 model Jev cuts LLM costs by 70% by handling routing, skill picking, and high-volume classification at 4 cents per million input tokens.
Anthropic updated its official skill-authoring guide: keep main files under 500 lines, limit references to one level deep, and replace prompt rules with hooks.
A decision-only model built for picking, not writing, paired with Claude Code across nineteen real automations, from spreadsheet tagging to routing which Claude model handles a prompt.