Modern Creator
Jad M.H | AI Automation · YouTube

How to Make Viral Motion Graphics With AI With 0$ (Turned it into a Skill)

A breakdown of turning 10,000-character prompts, local audio synthesis, and automated critique loops into high-end motion graphics with Claude and zero paid APIs.

Posted
2 days ago
Duration
Format
Technical Walkthrough & Case Study Breakdown
Direct, analytical, and pragmatic
Views
17.9K
431 likes
Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You want to generate complex motion graphics without paid subscriptions to After Effects, Suno, or ElevenLabs.
  • You want to understand the actual technical stack behind viral AI showreels.
  • You are a designer or developer interested in programmatic video rendering using code and FFmpeg.
SKIP IF…
  • You are looking for traditional GUI timeline editing tutorials.
  • You lack local storage or compute to run open-source audio synthesis models.
Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:00 – 00:47

01 · Cold Open & The Zero-Dollar Claim

Introduces the project scope: three motion graphic experiments generated entirely via Claude with zero third-party subscription costs.

00:47 – 01:36

02 · Local Infrastructure & Viewer Retention

Explains that audio was synthesized locally on a laptop to avoid fees, and sets viewer expectations for technical comprehension.

01:36 – 02:36

03 · The Viral X Motion Graphics Wave

Surveys viral posts following Claude Opus 5.5's release, illustrating the sudden explosion of AI-generated motion graphics.

02:36 – 03:33

04 · Case 1: Viral Showreels Analyzed

Examines Stephan Livera's viral showreel tweet and breaks down how basic prompt variations achieved millions of views.

03:33 – 05:12

05 · Case 2: 12-Hour Autonomous Renders

Reviews complex music video generations like Donald's viral animation, addressing claims of 12-hour continuous generation.

05:12 – 06:16

06 · Debunking the 'One Prompt' Myth

Reveals Anthropic internal insights showing that 'one-shot' viral clips actually required 10,000-character prompts and supporting frameworks.

06:16 – 07:13

07 · How Claude Actually Generates Video

Details the technical rendering mechanism: Claude writes web code, a headless browser captures individual frames, and FFmpeg builds an MP4.

07:13 – 08:10

08 · Video 1 Breakdown: The Simple Prompt

Demonstrates a basic 16-second kinetic showreel created with minimal prompt instructions, pointing out aesthetic homogeneity.

08:10 – 09:34

09 · Video 2: Local Audio & Music Integration

Walks through integrating ACE-Step 1.5 to compose an original hyperpop soundtrack locally on a MacBook Air without paid audio tools.

09:34 – 10:56

10 · Video 2 Output & Code-Based Sound Effects

Plays the finished lyric video and reveals how sound effects were synthesized using raw mathematical Python scripts rather than stock audio.

10:56 – 12:38

11 · Video 3: The Multi-Role Crew Prompt

Introduces the narrative prompt architecture where Claude acts as director, writer, animator, sound designer, and render engineer.

12:38 – 14:10

12 · The Autonomous Visual Critique Loop

Explains how Claude reviews midpoint frame contact sheets, assigns ratings, and iterates until every shot scores an 8 out of 10.

14:10 – 15:47

13 · Packaging the Skill & Final Verdict

Packages the pipeline into a downloadable Skool resource and answers whether human motion designers are obsolete.

Atomic Insights

Lines worth screenshotting.

  • Claude does not generate video files directly; it writes code for web pages that render motion, captures screenshots via headless browser, and stitches them with FFmpeg.
  • A 16-second animation rendered at 60 fps translates into exactly 960 individual frame screenshots compiled programmatically.
  • The viral 'one-shot prompt' narrative is deceptive; top-tier outputs rely on prompts approaching 10,000 characters alongside pre-configured frameworks.
  • The prompt constitutes only 10% of final quality; the surrounding harness and pipeline configuration account for the remaining 90%.
  • Without explicit negative constraints banning default styles, Claude defaults to centered typography against standard gradient backgrounds with basic fade transitions.
  • Sound design can be executed without paid asset subscriptions by generating synthetic risers, impacts, and whooshes directly via Python mathematical audio code.
  • Local music generation using open-source architectures like ACE-Step 1.5 allows zero-dollar audio synthesis without recurring API fees.
  • Transitioning Claude from executing a single functional task to inhabiting a multi-role crew (director, animator, composer, editor) dramatically alters structural coherence.
  • Implementing a critique loop where Claude inspects contact sheets of mid-shot frames allows iterative autonomous scoring from 1 to 10 until quality thresholds are cleared.
Takeaway

Architecture, not prompt brevity, determines generative video quality.

THE PRODUCTION PROTOCOL

To transition AI video from generic slides to polished animations, creators must shift from single-line prompts to closed-loop technical pipelines.

  • Treat LLMs as frontend engineers generating canvas code rather than direct media rendering engines.
  • Eliminate repetitive AI visual signatures by explicitly banning default typography and centered layout patterns in your negative prompt instructions.
  • Generate synthetic Foley effects locally through programmatic waveform formulas rather than licensing audio assets.
  • Deploy self-critique loops that parse visual contact sheets to force autonomous quality revisions before export.
Resources

Things they pointed at.

06:45toolFFmpeg ↗
14:30toolKokoro-82M
Quotables

Lines you could clip.

06:00
“The prompt is 10% of the video. The other 90% is the setup around it.”
Directly refutes the misleading 'one simple prompt' trend across AI social media.→ TikTok hook↗ Tweet quote
08:16
“Motion graphics without sound feel like a PowerPoint.”
Concise creative truth highlighting why audio integration transforms visual perception.→ IG reel cold open↗ Tweet quote
11:07
“For video one and two, I gave Claude a task. For video three, I gave it a crew.”
Clear, memorable framing of persona-based agent prompt design.→ newsletter pull-quote↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphoranalogy
everything you're watching right now claude made the story the animation the music every single sound effect and there are 375 of them no after effect no suno no stock music websites and it costed me zero dollars this is the third of three videos i made with claude to test it out the first one from a simple prompt second one with a bit better i didn't think ai could do this yet and today i'll show you all three exactly what changed between them so you can make your own and before anything i turned everything in this video into a skill that you can install and use on your own cloud under two minutes it's free and my school community link is in description as you know me i'm not daykeeping any of this And because the music and the sound all run locally on my laptop, it's literally $0.
No API keys, no credits, nothing on top of Claude. And just so you know, I'm not sponsored. Nobody paid me to make this video.
I don't make a single dollar from the skill buy. I'm doing this because I genuinely enjoy it. And I thought I would share it with you.
So you can go grab it right now. But the only thing I'm asking. I would really appreciate if you stick all the way to the end of this video.
First reason is because if you don't understand how it works, you won't get the good videos out of it. And the second reason it really pushes my videos. I'm not asking for money.
I'm asking for something more important. I'm asking you for your attention and I will try to give you the most value there is. So let's get right into it.
So on September 22nd, Anthropic released Opus 5 .5 and within a few days, my twitter feed was just motion graphics like this one from stefan levera sentence make a dynamic 15 seconds motion graphics video that shows what an incredible motion designer you are and it got 2 .2 million views This one from Tony.
Just one year ago, I paid over $1 ,000 for a video like this. Now I made Opus 5 .5 in less than 30 minutes, which I also did in less than 30 minutes.
And this one is actually crazy. It's from Donald. He posted it.
I spoke to my computer for 5 minutes, Claude worked for 12 hours, and I woke up to this, almost 4 million views.
captions say the same thing one prompt one shot so of course everybody tries it and rex and wong said what everyone was thinking everyone says they created it with one prompt but his video looked like it looked mint so is it really one prong or is it just a trust me bro situation here harik who works on cloud code at anthropic answered on one tweet the post claude one shot this prompt 10 ,000 characters plus skill examples and API keys and Donald's prompts, almost 10 ,000 characters.
That's not one prompt. That's a whole documents. There is this article on Twitter by moves 1 .7 million views and one line stuck with me.
The prompt is 10 % of the video. The other 90 % is the setup around it. So what I did is I took and I fed my AI with all the knowledge and I tested it for my three videos.
One simple prompt, better and best. But first, one thing most people don't get. Cloud can't make video.
It's outputs. It can't output MP4. It only writes text.
So what it does is write code, basically a web page that knows exactly what the screen should look like at every moment. Then a browser opens in the background, takes a screenshot of every single frame and a free tool called FFM and PEG. glues them together into an MP4 at 60 frames a second, a 16 second video is 960 screenshots.
And because it's code, everything is editable. You don't like the color. You can change it, line the render again.
So why did I use the open free source framework for this called hyperframes as a lot of people know, because if you don't give claw the framework, it usually builds everything from scratch. every single time it starts from a setup that already works all right so here's the first video it's a simple prompt i took a viral sentence the show real one and i added a few lines about the look the soft space the colors a nice serif font i'll let you look at it and watch it So 16 seconds, 60 frames, 7 techniques, and labels every one of them in a corner.
Like a real showreel. For one short message, this is actually insane. But here's the thing.
Hundreds of people use the exact same sentence. So all those reels kind of look the same now. And there's no sound, there's no real idea behind it, because the prompt doesn't have one.
So I tried to test that first with a single prompt. now for my second video is a bit better there is a huge difference i would say because motion graphics without sound feel like a powerpoint sound is the start to feel like a real video now the normal way to do this is either suno for the music and 11 labs for the voice a stock site for sound effects that's three subscriptions three api keys and you pay per song per character per download and i didn't want any of that so i found a music model called ace step 1 .5 it's free it's open source and it runs on my laptop and my laptop is not a beast it's a macbook air 16 gigs of ram so for video 2 i told claude you're the songwriter you're the animator and the sound designer with original song about making motion graphics with ai make it on my mac and make a lyric video where every word shows up exactly when it's sung and claude did the whole thing by itself wrote the lyrics the hook the sentence the 60 frames it made 12 versions of the songs then it checked everyone i'll let you watch it and you tell me what you think
There is a fun fact. One line has much effect on the voice. the that the transcription heard something else should insert instead of type it out hit enter same syllables so claude just matched them in order and the sound effects it didn't download a single one it made them with code wishes impact a riser into the drop like sparkles basically a math that sounds like a whoosh and there's one trick i used i told claude exactly what's not allowed what it's not allowed to use every color every font every effect from my computer every effect from my other videos because if you don't ban the defaults it goes back to that center text a grid and everything fading in that's the look that everyone gets and that video too literally costing me zero dollars the only thing that cost me was the clot subscription that i have video three is the best one and honestly this is where it gets crazier you're able to give an emotion and a story in this For video one and two, I gave Claude a task.
For video three, I gave it a crew. The prompt starts with, you're a director, writer, animator, composer, sound designer, and a render engineer. And instead of describing a video, I gave it a story.
It's called One Conversation. It's 2am, a creator with a big idea and no team. They type they have an idea and they don't have a team.
I'll let you watch it.
The good thing is Claude can see images so it can look at what it made after every round. It takes a screenshot from the middle of every shot, puts them all in one page and grades every shot 1 to 10. Story, composition, motion, typography, sync with the music.
Then it writes down the three worst problems and fixes them and goes again. Every shot is an 8 or more. And that's where I saw really a difference.
and the people that are really honest about it say the same thing like maple joseph made a beautiful watercolor short film with opus 515 163 models calls almost seven hours And the post says it's straight. It doesn't really nail it in one shot.
It needs a good sale. I took all these videos from X, the prompts and what to connect it with, with zero APIs, turned it into a skill. So you don't have to write 10 ,000 character prompts.
The structure is already in there. The story, the steps, it can skip the music, the sound effects, the critic loop and all the lessons that I have to learn. i damaged definitely my computer to make this also i want to mention if you want a voiceover you don't need to pay for that either there are free voice models that run on your computer too my three videos don't have a voice the story works with a text on screen but you can add one one thing though the music model takes some space on your desk so check if you have space for that before you install it here's my take are motion designers cooked no everyone If these videos needed someone to decide what the story is, what it should feel like, and what's good versus what's basically meant.
Claude did the work. Still, I had to direct it. A year ago, someone had to pay over $1 ,000 for a video like this.
Now it's as simple as just one conversation. So again, if you want the skill, it's free. It's in my school community.
Link is in description. Install it. Make a video and please show me what you have done.
and thank you so much for staying all the way to the end of this video seriously it really helps me a lot i'm making a carousel motion graphic designer so stay tuned and tell me in the comments which one was your favorite subscribe if you want more skills like these and i'll see you in the next one
The Hook

The bait, then the rug-pull.

The video opens with a cinematic, AI-rendered dark-mode interface before revealing the host in picture-in-picture, declaring that the entire intro was assembled autonomously without After Effects or paid music libraries.

Frameworks

Named ideas worth stealing.

00:28model

The Three Video Progression

  1. Video 1: One simple prompt (task execution)
  2. Video 2: Add local music and sound (sensory expansion)
  3. Video 3: Give it a story and crew (directed iteration)

A staged progression demonstrating how AI output quality scales with prompt specificity, modal integration, and iterative critique.

Steal forStructuring AI software case studies and educational tutorials.
12:40concept

The Automated Critique Loop

  1. Render midpoint contact sheet
  2. Score each shot 1 to 10 across clarity, typography, and sync
  3. Identify 3 worst problems
  4. Iterate until all shots score >= 8

An agentic feedback pattern where multimodal models evaluate rendered image sheets against strict rubrics and apply code fixes autonomously.

Steal forAutomated QA pipelines in content generation systems.
CTA Breakdown

How they asked for the click.

VERBAL ASK
15:15link
“So again, if you want the skill, it's free. It's in my Skool community. Link is in description.”

Pitched as a free community resource eliminating prompt writing, positioned alongside a soft viewer appreciation ask.

MENTIONED ON CAMERA
Storyboard

Visual structure at a glance.

Hook: 375 Sound Effects
hookHook: 375 Sound Effects00:00
Overview of Three Tests
promiseOverview of Three Tests00:29
The Opus 5.5 Wave on X
contextThe Opus 5.5 Wave on X01:38
10% Prompt vs 90% Setup
argument10% Prompt vs 90% Setup05:55
Code-to-Frames Mechanism
technical_detailCode-to-Frames Mechanism06:20
Video 1 Output Analysis
case_studyVideo 1 Output Analysis07:20
Running ACE-Step 1.5 Locally
technical_detailRunning ACE-Step 1.5 Locally08:50
Video 2 Lyric Video Output
case_studyVideo 2 Lyric Video Output09:45
Assigning Crew Roles in Prompt
frameworkAssigning Crew Roles in Prompt11:10
Contact Sheet Self-Critique
frameworkContact Sheet Self-Critique12:45
Skool Community Skill Download
ctaSkool Community Skill Download15:15
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.

11:16
Alex Finn · Review

Claude Opus 5.5 Is the Greatest AI Model Ever Released

An early-access creator says Anthropic quietly shipped a model that's smarter, faster, and cheaper than its predecessor, then reversed a months-long industry trend of AI answers reading like jargon-stuffed smoke-test reports.

September 23rd