Modern Creator
Ahmed Mukhtar AI · YouTube

How to Make AI Video Ads for $6 with Claude Code

An ex-Rolls-Royce ML engineer wires Claude Code straight into BytePlus's Seedream and Seedance APIs, skips the credit-metered ad platforms entirely, and live-builds a JBL speaker ad for $6.64.

Posted
1 weeks ago
Duration
Format
Tutorial
educational
Views
5.9K
149 likes
Big Idea

The argument in one line.

Connecting Claude Code directly to Seedream and Seedance's APIs instead of renting credits from a platform like Higgsfield turns a $17-20 AI video ad into a $6.64 ad with the same finished quality.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • A solo marketer or small agency running paid ads who wants AI-generated UGC or product ads without paying per-credit platform markups.
  • Someone comfortable running Claude Code from a terminal against a cloned repo, not looking for a no-code drag-and-drop tool.
  • A creator who needs a character or product to stay visually consistent across multiple AI-generated clips instead of one-off generations.
SKIP IF…
  • You want a point-and-click tool with no setup — this requires API keys, an IDE, and running Claude Code yourself.
  • You need broad model choice — this pipeline is deliberately built around one provider (BytePlus) to avoid a face-rejection problem that hits cross-provider pipelines.
TL;DR

The full version, fast.

The video argues that platforms like Higgsfield sell access to AI models at a markup: you rent credits, get back raw clips, and still edit, stitch, and caption everything by hand, at $17-20 per finished ad. The creator built a Claude Code pipeline that connects directly to BytePlus's Seedream (images) and Seedance 2.5 (video) APIs, writes a brief, casts a consistent character or product from one written description, generates cheap still keyframes before paying for motion, and renders a finished ad with captions and sound. Staying on one provider for both image and video generation avoids a face-rejection problem common to cross-provider setups. A live build for a JBL speaker ad comes out to $6.64 total, itemized stage by stage.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0001:06

01 · Intro

Cold open showing three finished AI ad clips and the claim that none of it was filmed, building to the $6-vs-$17-20 cost claim.

01:0602:00

02 · Where This Came From

Backstory: a prior AI-ad startup (Lumiere) plus an existing shorts-repurposing pipeline are the two systems combined into this one.

02:0003:12

03 · The Credit Meter

Breaks down how platforms like Higgsfield price access to models, and why buying direct API access removes the markup.

03:1204:42

04 · Why Prompts Fail

Side-by-side of a vague quality-adjective prompt versus a camera-and-lighting-specific prompt, and why reference photos beat descriptions for real products.

04:4205:22

05 · The Four Blocks

States the fixed four-part prompt structure: subject, camera and lens, light, action.

05:2206:24

06 · Keyframes And Casting

Explains generating cheap still keyframes before paying for motion, and writing one reusable character description since the model has no memory.

06:2409:11

07 · The Face Problem

Explains why sending a face generated on one provider to a different provider's video model gets rejected, and why staying on one provider avoids it.

09:1110:48

08 · The Ads Demo

Plays three finished example ads, a Porsche film, a protein-shake UGC review, and a walkthrough of the four pipeline stages that produced them.

10:4821:37

09 · Live Build

A real-time build of a JBL speaker ad from three product photos through brief, casting, keyframes, and clip generation, including a live continuity fix.

21:3722:37

10 · What It Cost

Itemizes the exact API cost of the finished 15-second live-build ad, stage by stage, totaling $6.64.

Atomic Insights

Lines worth screenshotting.

  • AI ad platforms like Higgsfield charge $17-20 per finished ad but only hand back raw generated clips — the buyer still edits, stitches, and captions everything by hand.
  • Connecting directly to a model's API instead of renting platform credits dropped the cost of an equivalent finished ad from $17-20 to $6.64.
  • Vague prompts like 'ultra realistic, 8K, flawless skin, perfect lighting' don't describe anything — they're a wish for quality, and they produce visibly synthetic faces.
  • Describing a shot in real camera terms, an iPhone 24mm f/1.8 lens, 4000K overhead light against window daylight, produces images that read as photographs instead of AI renders.
  • For a real product with an exact shape, like a specific car, a reference photo beats a text description every time, because a photo carries information words can't fully specify.
  • A reliable image or video prompt has four blocks in a fixed order: subject, camera and lens, light, action — reordering them changes what the shot emphasizes.
  • Generating a full video clip on the first try is the expensive way to work; generating a cheap still keyframe first and approving it before spending on motion catches mistakes early.
  • AI video models have no memory between requests, so a consistent character across scenes requires writing one detailed description once and reusing those exact words in every scene, not rewording it.
  • Sending a face generated on one provider to a different provider's video model gets rejected most of the time, because the video model can't confirm it isn't a real person's photo.
  • Keeping image and video generation on the same provider's account avoids the face-rejection problem entirely, since the video model trusts faces traceable to its own image pipeline.
  • A 4-image character casting pass costs about $0.36, cheap enough to audit and regenerate before committing to full-length clips.
  • The pipeline caught its own continuity error automatically, a t-shirt graphic that drifted between keyframes, and offered a cheap targeted fix instead of a full recast.
  • For a 15-second ad, casting and keyframe generation together cost under $0.50; the video clip generation is the dominant cost, at roughly $6.19 of the $6.64 total.
  • Breaking the pipeline into approval stages costs more time up front but catches mistakes, like a mispronounced word in generated dialogue, before paying to regenerate a full ad.
Takeaway

Own The API, Skip The Credit Meter

PROMPT & PIPELINE CRAFT

Treating camera, light, and character description as precise engineering inputs, and paying a model provider directly instead of a credit-metered platform, turns a $17-20 AI ad into a $6.64 one.

01Intro
  • AI ad platforms hand back raw generated clips only; the buyer still has to stitch, caption, and add sound effects manually, on top of a $17-20 per-ad platform fee.
  • The stated cost comparison driving the whole video is $6 versus $17-20 for a functionally identical, fully finished ad including voiceover, sound effects, and captions.
02Where This Came From
  • Building a full video-ad startup end to end forced a deep understanding of how to prompt image and video models for genuinely usable output, not just demo-quality clips.
  • A separate shorts-repurposing pipeline that transcribes long-form video, extracts sections, and auto-posts to five platforms already ran in production for months before being combined with this ad system.
03The Credit Meter
  • Credit-based platforms gate the newest models behind higher-priced tiers, so a $47/month plan can still lack the specific model actually needed for the best output.
  • Buying API access directly from the model provider removes the credit meter entirely: the same clips generate at the provider's raw cost instead of a marked-up credit price.
04Why Prompts Fail
  • Quality-adjective prompts like 'ultra realistic, 8K, flawless skin, perfect lighting' don't describe anything concrete; they read as a wish list and produce faces that look obviously synthetic.
  • Describing the actual camera and lens and a real lighting setup produces images that read as photographs, not renders, using the same model and parameters as the vague version.
  • For products with an exact real-world shape, feeding the model a real reference photo beats describing every detail in words, because the photo carries information text can't.
05The Four Blocks
  • A reliable image or video prompt has four blocks in a fixed order: subject, camera and lens, light, and action, and reordering them changes what the shot emphasizes.
06Keyframes And Casting
  • Generating a full motion clip on the first attempt is the expensive way to work; generating a cheap still keyframe first and approving it catches mistakes before they're paid for.
  • Because the model has no memory between requests, a consistent character across scenes requires writing one detailed description once and reusing those exact words in every scene's prompt.
07The Face Problem
  • Cross-provider pipelines routinely reject realistic human faces at the video-generation step, because the video model can't verify the face didn't come from a real photo of a real person.
  • Keeping image and video generation on the same provider's account resolves the rejection, since the video model trusts faces it can trace back to its own image generation.
  • Automated continuity checks caught a drifting t-shirt graphic between keyframes and offered a cheap, targeted fix instead of defaulting to a full, more expensive recast.
08The Ads Demo
  • A single reference photo can carry a product's identity across six different generated scenes, keeping it visually consistent without describing it fresh in every prompt.
  • A believable UGC-style ad doesn't need a real actor: a plainly-described, non-existent character held the same look and voice across every scene of a product review.
09Live Build
  • Approving each stage before the next one runs limits how much gets spent on a direction that ends up getting rejected or reworked.
  • The system offers ready-made ad angles (problem-led, social-proof, demo-led) generated from a pattern-matched set of high-performing ad structures, not written from scratch each time.
  • Even a well-run live build produced one usable error, a mispronounced word, out of five generated clips, showing regeneration is a normal, budgeted step rather than a failure.
10What It Cost
  • For a 15-second ad, casting and keyframe generation together cost under $0.50; the dominant cost is video clip generation itself, at roughly $6.19 of the $6.64 total.
  • The brief, script, storyboard, caption planning, and final render all run inside a flat-rate coding assistant subscription, so only the image and video API calls are metered.
  • The full measured total for one finished 15-second ad was $6.64, against a stated $17-20 for the same result through a credit-metered platform.
Glossary

Terms worth knowing.

Keyframe
A single still image marking the start of a video clip, generated cheaply before the full motion clip so mistakes get caught before the expensive generation step.
Casting (AI pipeline)
Generating a reference image of a character or product from a written description, then reusing that exact image and description across every scene so the subject looks consistent.
Credit meter
The pay-per-generation credit system platforms like Higgsfield sell, where each render, including failed or discarded ones, consumes paid credits regardless of the model's actual cost.
Character sheet
A set of reference images showing a character or product from multiple angles, generated once and used to lock in appearance before any video clips are made.
Seedance 2.5 / Seedream
BytePlus's video and image generation models, called directly through their API here instead of through a credit-based aggregator platform.
Resources

Things they pointed at.

00:30toolHiggsfield
02:56toolSeedream
02:56toolSeedance 2.5
11:49toolBytePlus ARK API
01:30productLumiere (formerly Lumilime)
11:26toolWindsurf
00:00tooln8n
Quotables

Lines you could clip.

00:00
every ad you're seeing on the screen right now is 100% generated with AI
flat, confident claim with an instant proof-of-concept setupTikTok hook↗ Tweet quote
15:07
afrobeats with no bass is a crime
the punchy ad line generated live, doubles as a quotable one-linerIG reel cold open↗ Tweet quote
22:10
renting the same ad runs $17 to $20; learning what to type cost a lot more than one ad
closing thesis stated as a clean, quotable comparisonnewsletter pull-quote↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphoranalogystory
every ad you're seeing on the screen right now is 100 generated with ai whether it's cinematic ads ugc videos or product motion graphics i have built a system to do this entirely with cloud code and direct integration with the models themselves bypassing any of the model aggregators like hicksfield or founder ai or any other platform that charges you on credit and i think with ai video ads is that anyone can make them now but the problem is they look like ai and it's very easy to tell therefore no one actually runs them and spends money on the ads the systems and platforms that get you to this high quality output similar to what we have here, charge you anywhere between $17 to $20 for a single ad.
And even then, they only just simply hand you over all of the generated clips. You still have to stitch them together, do the transitions, sound effects, captions, and do all of that manually. The ads that you're seeing on the screen right now only cost me $6.
And that includes voiceover, sound effects, transitions, and captions. 100 % ready to be posted. So today, I'm going to show you the entire system end -to -end of how you can get similar results.
And then we will build one live so you can see the complete process and the cost breakdown. and all of the skills will be in my school community so make sure you hit that link in the description down below so all you have to do is download it and point it to your product that being said let's dive into the video so a few months back i actually built a startup in the space so i know what i'm talking about it's called lumilime and it basically takes a single brief to fully produce video adverts from script visuals voiceover music captions in minutes not weeks so by building this i fully got to understand everything in the space of video generation audio generation i basically pulled apart every system out there to really understand how to get the most out of these models and how to prompt them to get the best results.
Taking that and coupling it with my shorts pipeline architecture that basically takes my long form videos and transcribes them, extracts different sections and basically automatically renders the shorts and posts them into my content calendar. And this is actually running 100 % in production. I've been using it for months now.
So all I have to do is just give it my YouTube video and it does everything for me up to the point where it actually posts it in Instagram and TikTok, YouTube X and LinkedIn. taking these two systems and combining them is what produced this result so just to quickly cover the issue with all of the different platforms out there is that you start off signing up to a platform something like hicks field you basically get credits for certain number of generations and you get a bunch of clips back that you're finally happy with but then you have to open up an editor and do it by hand so whether that's being kapka or premiere pro whatever the tool is you have to do a lot of manual work and they also charge quite a lot so anywhere from 17 to 20 just for one advert and that's not including all of the iterations all of the regenerations that you have to do so if you look at something like hicksfield and the price packages you'll see that it doesn't even include the best models and the latest models like c dance 2 .5 which is what i actually use in my system so you would have to go for the plus and that costs around 47 a month again this is just based on credit so you can quickly use all of this up in just a few adverts so basically i wanted to find out what would happen if we owned the whole system so what we're building is direct apis with the models so c dance 2 .5 and c dream
five for the image generation so then there's no credit meter we get the same clips at cost the app does the whole thing the system does basically from start to end like i mentioned with the captions transitions voiceover everything and this only cost us around six dollars for a finished ad so how do we actually get this high quality output and it's all in the prompts right so on the left hand side here we have the typical vague prompt where we just basically specify the quality of the photo and you know we want this person ultra realistic 8k flawless skin perfect lighting and this is the that we get.
And I actually ran this test, the same model, the same parameters, everything is exactly the same. And the only thing I changed is how I did the prompt. And this is the output that you'd get on a left -hand side.
It doesn't look like a perfect woman, which is actually a good thing because it doesn't actually look like AI here. And the idea is you want to actually describe the camera lens. You want to describe the angles in actual cinematic terminology.
And the way that I've actually built it is I've collected all of the different camera settings, all the different cameras, filters, lighting, angles, the exact terminology. technology that actually gets used in film production and this is the key so what we did here is iPhone 24 millimeters and when people use terms like ultra realistic 4K flawless skin they're not really describing anything they're just wishing for high quality images and this is where it gets interesting for products because using images gets a much better result so here for a gt3rs Porsche we spelled out every detail and you can still see it's not 100 accurate and it's not great what I did instead is got a real image from Google to basically have it as input and you know photographs carry the car the words can describe the lighting and atmosphere and a cinematic feel of the image and this is another thing that we've kind of unlocked by building this system is that sometimes it's a lot better to just have a reference image that we can use to generate our product shots and one thing we have to keep in mind is that there is an order to this and i've built the prompts in four different blocks we first want to start with the subject what is in actual frame right if it's a person if it's a product if it's a car this is the first block in the prompt and
what we want to focus on then we go into the camera and the lens you know the real camera body real focal length and third comes the light so the light setup the color the temperature of the light and finally we add the action so what moves and how and by combining all these four blocks this is how we get very good prompts that generate us very good outputs if you want to focus on camera angles instead and movements of the camera you can take the second block of the camera and put that at the top of the prompt so that you can actually have a more dynamic and fluid looking shots and camera angles all right so there's four things that make up the pipeline and one thing i want you to get right from the get -go is we don't actually build by prompting and generating the full clips what we want to do first is generate keyframes these are low cost and they basically give a structure to the clips so there's two ways to do this you can either pin both the start frame and the end frame of the short clips that make up the ad but on the other hand if we just give it a prompt and a start frame the model runs to its own conclusion and this is how i got it to perform really
well the second thing is that we want to use casting we want to create the characters because every scene has a different request the model doesn't really have memory so the way to keep the characters consistent throughout the different clips we want to give it a character sheet and i'll show you what that looks like in just a second but basically we describe the person or the character or the product we give it you know some very detailed descriptions and to improve it we might add some imperfections like a mole or scar you know we want to describe the clothing the body shape the face shape and everything that goes into building that character so we generate the different character sheets and you kind of audit and approve the character that you like and this also costs us pennies to get a very good looking character like you've seen in the ads at the start and this basically fixes the full ad to that character so that we can have consistent characters across the different scenes so one big problem that you'll come across if you try to connect to the apis of the models directly and not using something like Hicksfield is that the models sometimes reject faces and realistic looking people so a way around it is by using the same provider
so for example if we use google nano banana pro to generate the character sheet and then send it to cdance 2 .5 to generate the clip nine times out of ten it will actually refuse to do so because it doesn't want to generate videos of realistic looking humans because it thinks it's actual real person however a way around it is by using the same provider so in the system i've connected to cdream which is under the same umbrella and it actually accepts when we send it a character to generate the images even if the images have faces on them and they look realistic so this is something that might catch you out so keep that in mind and this is a benefit by using hicks field you don't really have this problem but if you're building something like this and you're directly connecting to the models this will be a big problem for you and this system basically does the full thing including that last 20 because we're using cloud code have a very optimized prompt for the brief and the campaigns for example you want to say i want to build a campaign for a lamp so you give it a very simple brief but because of the prompt library that we've built in the system it goes and builds out the full scenarios and the different scenes and the keyframes that should make up the final ad and this is all of course using cloud code so this doesn't actually cost us anything it then generates the cast and the keyframes and once you approve it will go ahead and create the clips and this basically limits the spend because we're only approving
frames that look good this makes us save a lot of money the third thing then is we render this is done locally on my laptop which is going to be also free so we listen back and if we need to change anything the captions or the timings of any of the things we can copy different scenes stitch it together and then we have a finished ad with captions and music and everything that we can go ahead and post so the finish is something like this we have the same clip same words a different finish we also have some very cinematic looking captions as you can see on the right hand side here and this basically takes the ad from good to amazing and professional so on the left you can see this is very subtle it highlights the words that the person speaks and on the right hand side we have very dynamic fluid cinematic looking captions and these finishing touches is what makes or breaks the ads so if you want to grab the ad factory skill and repo i've put everything in my school community and that will be
the first link in the description down below right so let's look at the ads that we've generated with the audio enabled some machines are built this one escaped 9000 rpm 518 horsepower naturally aspirated Porsche 911 GT3 RS so you can see here every sound is generated with the model the tires the voiceover it stitches all of the different scenes and clips together and we basically started with one paragraph and the car stayed the same throughout the whole thing because we're using the product sheet which is very similar to the character sheet but different shots of the product we hear in the scenarios the car at different angles I stopped making shakes a month ago this is why 30 grams, no sugar and I don't have to wash anything.
That's it. That's the whole review.
Here we had a UGC video of a character promoting a protein shake. She doesn't exist, right? The cast was written in plain description and it held the same look throughout the whole scenes.
So sometimes people use different models for lip syncing after generating a video. We didn't have to do this because Seedance 2 .5 is actually really good. So to visualize the four stages, this is what it looks like.
So we start off with the cast. This is the character sheet of the subject from different angles. Once we approve, we go to the key.
frame generation so this is the start key frame of the single scene every scene has a key frame and then we get the finalized clip once we're happy with the clip then we go and do a finished advert stitching together all of the different clips adding the captions and it's ready to go and i want to also iterate again only stages two and three actually cost us anything because we're using claude at the start here this doesn't cost anything and the rendering and the captions also runs on the laptop so this is also free right so let's jump into cloud code and build one from scratch all right so once you go and download the repo for my school community you have something that looks exactly like this what you need to do is open it up with whatever ide that you like here i'm using windsurf but you can use vs code cursor or whatever that suits you you can see i have three different images here because what we're going to be doing is creating an ad for my gbl speakers so i just quickly took some photos of my speaker from the front
from the side and from the top. So the first thing you have to do is get the API keys for the Arc API from Byte Plus. So what you need to do is go to byte plus dot com and inside of there, go to products and then see dance.
You'll come to this page. Then you need to click start now. So this will take you to the platform where you'll see all of the different models, the console.
and what you should focus on is at the bottom left here you'll see api keys and you just need to make sure that you sign up or log in you don't have an account i just signed up with google and once you've signed up you'll get taken into your api key section i've already got one created so all you have to do is just create api key give it a name give access to which models you want under permissions just make sure this is set to all and then click create it'll give you the api key so just copy this and go back into the repo here and paste it under the api key and then what you need to do is convert this file into a normal .n file for the video generation we're using the cdance 2 .5 for the image generation i previously had gemini so this is why we have this here but because we want to keep the same provider and this is how we select the provider you don't actually need the google gemini api key so this can actually be completely removed and make sure the image provider is set to cdream which is the image generation provider that also comes from byteplus so it uses the same api key once you've done these two things make sure to convert this into a normal .n file
and you'll be set to go so let's save this and close this out and then you what you need to do is open up the terminal in this repo and start cloud code so here i've got cloud code running with opus 5 so all we need to do now is basically tell cloud what we want to generate right so i have a bunch of photos that i've added to the root directory of the repo i would like you to use these to generate a high quality ad for my speakers so there's three different photos front of the speaker side of the speaker and the top and we want to use these to create the product sheet following the skills and the process of the entire pipeline to generate high quality and engaging advert for these speakers i want to target young party people that love a bit of afrobeats so make sure you stop at each stage and let me approve before we go to the next stage so i'm just going to let it transcribe what i said and we can just go ahead and send this off so you can see that it starts by invoking the ad factory skill and it will go ahead and look at the photo so you find the photos and i'm just going to let it go through and run this right so we just came back to confirm some of the aspects here so the model this is the party box on core essential 2 the format
I don't want a faceless party. I want a UGC. one how many variants we just want one variant and then for the music we're just gonna let cdance generate it inside the clip and we're gonna go ahead and submit these answers all right so it finished the first stage and this is the brief they came back with so we have the product the price audience exactly kind of what i described it's going to be a high energy street party uh ugc kind of video we just want to do one variant and we've chosen the captions to be the cinematic theme and the industry is e -commerce So we are happy with this.
So I'm gonna say approve It will now go ahead and do the second stage where it's gonna give us three different angles because this is another thing We have a system that basically takes the best performing ads and we built the system using that that gives us the structure for how the ads should be so it's gonna do three different angles uh problem led social proof and demo led so we can see uh which one kind of fits this specific ad and then we can pick one right so we just finished the second stage and it has three different angles that we can take with this app the problem led basically the headline being afrobeats with no bass is a crime i quite like this one most likely the one i'll be choosing the idea here is that everyone asks me what the speaker is so it starts off with a group chat scrolling fast blah blah blah and then someone asks what speaker is that and then there is like a demo led so watch what this button does so extreme micro close up bass boost on the top panel from the actual speaker image that we had uh so what i'm gonna do is tell it let's go with the
problem led angle and then it should have everything to generate the keyframes all right so it just finished the script so before we get to the casting stage it just showed me the actual breakdown of all of the different scenes and i read through this i'm happy so i'm gonna go ahead and approve and then tell it to go ahead with the casting all right so as it's doing the casting image generation you can see in the folder here we have the campaigns and for party box you can start to see everything getting built out from the brief the characters the hooks the script and if you go into the assets this is where we'll get the keyframes so the keyframes get generated and put into the keyframes folder and then characters get put into character and the clips will get into put into clips and then the final output will be under the out folder so let's give it a few seconds and wait until it finishes the character generation so we can see what the different kind of characters look like so we can approve them and you'll see here the cost for the four images from the different angles of the face will cost us 36 cents and that's barely anything all right so we have
the different auditions and the casting so if we go to the character here we can see the different scenes so we have her in the kitchen this is amara so it starts off by creating a character sheet at the different angles then we have her in the kitchen the street and also in the sofa One thing you notice is that the graphic of the t -shirt is slightly different and Claude actually just automatically realized this, which is great.
So these are the kind of things and issues that you'd come across. So basically it's saying that there's a real drift because the t -shirt graphic is not the same in every photo because it didn't actually describe exactly what the image on the t -shirt is. So we can either fix it, right?
And just regenerate costing us 18 cents. So it's given us different options. We can recast entirely, redo the casting, or do we just fix?
the plane T edit so let's go with option one it's the cheapest and let's see if it fixes it right so it looks like it's just finishing up doing the keyframes and it also fixed up the t -shirt so now if you look at the different scenes so these are the starting frames of every single scene so we've got five different scenes starts off with her in a one looks like a kitchen that's very messy in like a party house party kind of vibe we get a close -up of the speakers this looks like the photo that i uploaded we've got side view of the speakers for this one you can see the party is becoming live now with nice colors in her back and everything she looks more happy and you know energetic and excited t -shirt here is kind of consistent throughout so that's definitely fixed and then it finishes off with a product shot of the jbl speakers right so i'm just gonna let it finish verifying everything and then we're gonna go to the next stage of the clip generation
So for this ad, we can see that it's going to generate 21 seconds and it costs around 41 cents per second. So this is actually going to be totaling $8 .60. So I'm going to tell it to go ahead, run the assemble prompt script, and then it will generate the clips and submit it to CDance.
All right, so it's just finished. We got the five clips back for all of the different scenes. So let's go ahead and quickly watch them over.
Afro drinks with no bass is a crime. This is a crime scene. okay and then nice good stuff this is not an actual bass but it does the job right same song same room that's what bass does this looks super realistic i love that yeah finishes off with the product shot so i think in the first scene it mispronounced a word afro janks with no bass is a crime yeah it did says afro janks instead of afro beats so only one clip out of the five we'll need regeneration so that's fine we can go ahead fix that and then generate the final you know completed ad stitched together so i'm gonna let it run this and then i'll be back when it's done right so now it finished doing the first render and put all of the different scenes together and very nice feature that it also does is it verifies where all of the subjects and products are in the frames and then it adds the captions the cinematic captions on top of those so we look at different points of the video just to make sure that the
don't overlap with like the logos so as you can see here it just makes sure the text is not on the faces or on the the main subject that we want to look at which is a very nice feature so it does all of this automatically because of the skill that we have in place so I'm just gonna let it run and finish up the final render and then we're gonna get to go ahead and see what that looks like right then so it just finished rendering the final completed advert so let's go ahead and give it a watch afro beats with no bass is a crime This is a crime scene Same song same room.
That's what base does Lovely I really like that and I think it does a really good job at portraying, you know, the whole process from start to end, as you just saw. I also like how we have a human in the loop at each of the stages so that we can approve the different steps and angle that we want to take with the advert.
This is just a design choice. You can go ahead and skip the all of the human in the loop steps and just give it a brief and then it will produce a fully completed advert for you. They probably have to tweak and regenerate some of the scenes, which is going to cost you more money.
So this is why I like doing it by step by step and break it down so that we can. minimize the regeneration costs we can capture errors and mistakes early in the process and then it's a lot cheaper to fix those instead of having to wait for the final produce ad so let's look at the cost and what this actually cost us so of course it always depends on the length and how many scenes inside of the advert so if you look at an average 15 second ad we have the first step that creates the brief script and the storyboard because it runs inside of cloud code so it's under your monthly plan this is free then we've got the casting so we've got two different images using cdream this is going to cost us 18 cents and then if we use those casting images for the keyframes so three different images for three different scenes this is going to cost us around 27 cents and then to generate the actual clips so 15 seconds using the cdance 2 .5 this will cost us 16 .19 and the caption plan this is also under cloud so there's nothing that we get charged here and then rendering transcribing and captions and it's usually the part where there's a lot of
and forth tweaking of things so it actually makes it really good that both of these steps are free so total cost for a 15 second ad is only 6 .64 so if we compare that to hicksville for example that cost us 17 to 20 dollars to run and with that being said let me know what you think of the system in the comments down below and if you want to grab the full repo make sure you hit the first link in the description down below and as always make sure you hit that like and subscribe button until next time take care
The Hook

The bait, then the rug-pull.

An ex-Rolls-Royce machine learning engineer opens by showing three finished AI ad clips, a Porsche film, a UGC protein-shake review, a skincare close-up, then flatly states none of it was filmed or edited. The rest of the video is him proving it: first the prompt theory, then a live $6.64 build of a JBL speaker ad from three phone photos.

Frameworks

Named ideas worth stealing.

04:42list

The Four Prompt Blocks

  1. Subject (what's in frame)
  2. Camera and lens (real body, real focal length)
  3. Light (setup, color, temperature)
  4. Action (what moves, and how)

A fixed-order prompt structure for image and video generation. Swapping the order changes what the shot emphasizes — putting camera first produces more dynamic, movement-led shots.

Steal forany AI image or video prompt for a person or product, especially when the current output looks generic or overly 'AI'
10:33list

Four Pipeline Stages

  1. Cast (character/product sheet)
  2. Keyframes (approved start-frame stills)
  3. Clips (generated motion)
  4. Finished ad (stitched, captioned, scored)

Only the cast and keyframe/clip stages hit a paid API meter; the brief and final render both run inside a flat-rate coding assistant subscription.

Steal forstructuring any AI content pipeline so approval gates happen before the expensive generation step
CTA Breakdown

How they asked for the click.

VERBAL ASK
00:58product
all of the skills will be in my school community so make sure you hit that link in the description down below

Soft CTA under a minute in, repeated again after the ad demos (~9:00) and once more at the close ('grab the full repo'). No urgency or discount language, framed as 'download it and point it to your product' rather than a hard sell.

MENTIONED ON CAMERA
FROM THE DESCRIPTION
PRIMARY CTAWhere the creator wants you to go next.
Storyboard

Visual structure at a glance.

open
hookopen00:00
origin story
promiseorigin story01:06
four prompt blocks
valuefour prompt blocks04:42
finished ad reel
valuefinished ad reel10:48
cost breakdown
ctacost breakdown21:37
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.

29:59
Danny Why · Tutorial

Claude Code Just Changed CapCut Forever

A screen-recorded build of a full 3D motion graphic inside CapCut, using a Claude Code-generated asset and an image-to-video AI to carry it past the timeline's edge.

August 21st