Automate Your Entire DaVinci Resolve Pipeline with GPT-6 Astra
A DaVinci Resolve editor connects an AI agent straight into the open project, then spends the video testing whether it survives real revisions, not just a first draft.
Connecting an AI agent directly to an open DaVinci Resolve project turns the real test of AI-generated motion design into whether it survives revisions, not whether it produces one impressive first draft.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You edit video for clients and want to know whether an AI agent can actually take specific revision notes, not just generate one polished cut.
You already work inside DaVinci Resolve and Fusion and want to see what stays editable after an AI builds the initial animation.
You produce recurring content, a mascot character, a product line, and need to know if a design can stay consistent across scenes an AI generates separately.
SKIP IF…
You're looking for a step-by-step Fusion compositing tutorial. This is a workflow demo of directing an AI agent, not a lesson in the software itself.
You don't use DaVinci Resolve or Codex. The core trick is specific to that combination of tools.
TL;DR
The full version, fast.
DaVinci Resolve Studio 21.1's native Setup AI Assistants option lets an outside agent, here GPT-6 Astra through OpenAI Codex, read and directly build inside an open project. The video briefs Astra on three jobs (a 15-second Vox-style explainer, a 12-second animated-character gag plus a 5-second sequel scene, and a 15-second product ad built from Higgsfield-generated visuals), then follows each with a specific revision: rebalance the timing, direct a second performance take, swap the opening, recompose for 9:16. Because every layer, text, images, timing, sound, stays editable in Fusion, each revision executes without rebuilding the whole sequence. The real cost comparison isn't the generation price, it's every unused generation plus software plus the reviewer's own time, measured against the versions that actually ship.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Cold open previewing the whole workflow, then Sanji states the plan: three jobs (explainer, animated character, product ad), each followed by a revision request, to see how well an AI agent handles the feedback you'd normally give an editor.
02:00 – 02:45
02 · Connecting Astra (Codex) to DaVinci Resolve
Walks through DaVinci Resolve Studio 21.1's Setup AI Assistants option, confirms Codex is installed and connected, and tests it by asking Astra to name the currently open project before handing it real work.
02:45 – 05:55
03 · Job 1: Vox-style explainer graphics
Astra is briefed to invent a 15-second Vox-style editorial explainer (cutout imagery, bold typography) about a pop-culture idea, built as editable Fusion elements; Sanji then requests a timing revision to give the key beat more room to read and compares both versions side by side.
05:55 – 09:36
04 · Job 2: One character, two scenes
Astra invents a character and a 12-second sight gag (an overconfident office worker attempting push-ups and failing), then reuses the same character in a new 5-second scene; Sanji directs a second take that adds a beat of anticipation before the reaction.
09:36 – 16:53
05 · Job 3: One product photo, a complete ad
Starting from a single phone photo, Astra uses the Higgsfield plugin to generate a matching background and two supporting clips, then assembles a 15-second commercial with typography and sound in DaVinci; Sanji requests an alternate opening that keeps the middle and ending, then a 9:16 vertical version that reads with the sound off.
16:53 – 17:57
06 · Saving the workflow as a reusable skill
Instead of re-explaining project structure for the next product, Sanji has Astra save the whole process, visual style, pacing, project structure, revision pattern, as a reusable skill.
17:57 – 19:10
07 · Where I'd use this
Sanji separates what he'd hand to the AI (build and revisions) from what stays his own call (direction, what's actually ready to publish), and frames the real cost comparison as the complete job, including his own review time, not just the generation cost.
Atomic Insights
Lines worth screenshotting.
DaVinci Resolve Studio 21.1 added a native Setup AI Assistants menu option that lets an outside agent read and directly control an open project, no separate plugin required inside Resolve itself.
Keeping text, images, and timing as separate editable Fusion elements is what lets a single revision request, more time on one beat, execute without rebuilding the whole animation.
A revision that has to fit the same total runtime forces an AI to redistribute time from the rest of the sequence, which tests whether it understands pacing, not just duration.
A character's personality has to read before anything goes wrong: stance, approach, and pacing establish confidence, and the reaction is what actually lands the joke.
Reusing an illustrated character across separate scenes only holds up if proportions, color, and personality are explicitly locked in the brief, otherwise later scenes drift from the original design.
Directing a revision by isolating one variable, add a beat of anticipation before the reaction, keeps the note focused instead of re-litigating the whole scene.
Sound effects have to move with whatever timing change gets made; an impact that shifts without its sound attached breaks the moment it's supposed to land.
Splitting a generative job into two stages, generate the raw visual material first, edit and assemble it second, keeps each instruction to the AI specific enough to actually get followed.
A generated background can look great in isolation and still fail a product ad if it doesn't leave visual room for the product and readable text.
Swapping only the opening while locking the middle and ending lets you hand a client a real creative alternative without re-deriving the whole ad.
A vertical recomposition isn't a crop of the horizontal edit, each scene gets redesigned (product placed above the title instead of beside it) so it still reads with the sound off.
The honest cost of an AI-generated ad isn't the accepted final render alone, it's every unused generation plus the software plus the time spent reviewing and correcting, against what actually ships.
Saving a proven workflow as a reusable skill, visual style, pacing, project structure, revision pattern, means the next similar job starts from a working process instead of a blank chat.
Takeaway
Directing AI Through Revisions, Not Just Drafts
REVISION WORKFLOW
The real test of an AI motion-design agent isn't the first draft, it's whether it can take a specific revision note, keep a design consistent, and still hit a fixed runtime.
02Connecting Astra (Codex) to DaVinci Resolve
DaVinci Resolve Studio 21.1 added a native Setup AI Assistants option that lets an outside agent read and directly control the open project.
Confirming the connection before starting work, asking the agent to name the open project, catches setup problems before they cost you a build.
03Job 1: Vox-style explainer graphics
Briefing an open-ended creative task still needs a hard constraint, here a fixed 15-second runtime, to get a usable first draft.
Keeping text, images, and timing as separate editable elements in Fusion is what makes a single revision request doable without a full rebuild.
A revision that must keep the same total runtime forces the AI to redistribute time from the rest of the sequence, which tests whether it understands pacing, not just duration.
04Job 2: One character, two scenes
A character's personality has to read before anything goes wrong: stance, approach, and pacing establish confidence, and the payoff reaction is what completes the joke.
Reusing a character across scenes only works if proportions, color, and personality are explicitly locked in the brief, otherwise later scenes drift from the original design.
Directing a second take by isolating one variable, add a beat of anticipation before the reaction, keeps the revision focused instead of re-litigating the whole scene.
Sound effects have to move with whatever timing change you make; an impact that shifts without its sound attached breaks the joke's landing.
05Job 3: One product photo, a complete ad
Splitting a generative job into two stages, first generate the raw visual material, then edit and assemble it, keeps each instruction specific enough to actually get followed.
A generated background can look great in isolation and still fail the brief if it doesn't leave room for the product and readable text.
Swapping only the opening while locking the middle and ending lets you offer a client a real creative alternative without re-deriving the whole ad.
A vertical recomposition isn't a crop: each scene has to be redesigned so it still reads with the sound off.
06Saving the workflow as a reusable skill
Once a workflow is proven, saving it as a reusable skill means the next similar job starts from a working process instead of a blank chat.
Character work, explainer graphics, and product ads each need their own saved skill because the consistency requirement is different for each one.
07Where I'd use this
The honest cost comparison isn't the generation price alone, it's every unused generation, the software, and the review time, measured against the versions you actually ship.
Directing an AI agent through revisions still leaves the two decisions that matter with the human: what direction to take, and what's actually ready to publish.
Glossary
Terms worth knowing.
Fusion
DaVinci Resolve's built-in node-based compositing and animation tool, used here to keep motion graphics editable after the AI builds them.
Astra
The GPT-6-based AI agent, accessed through OpenAI's Codex, that connects directly to DaVinci Resolve in this video to build and revise animations.
Codex
OpenAI's coding agent tool that hosts the Astra model and provides the connection point into an open DaVinci Resolve project.
Higgsfield
A third-party AI image and video generation plugin for ChatGPT, used here to create a background and two supporting clips from a single product photo.
Setup AI Assistants
A menu option added in DaVinci Resolve Studio 21.1 that registers and confirms which AI assistants (like Codex) are connected to the open project.
Resources
Things they pointed at.
00:16toolDaVinci Resolve Studio 21.1 (Setup AI Assistants)
“I'm giving ChatGPT the kind of job that you'd normally pay a motion designer for. And I'm doing it all by voice.”
cold-open thesis, works with zero context→ TikTok hook↗ Tweet quote
05:00
“The extra time has to come from somewhere.”
tight one-line insight about timing tradeoffs in a revision→ newsletter pull-quote↗ Tweet quote
18:40
“Astra handles the build and the revisions, whereas I set the direction and decide what's actually ready to publish.”
clean human-vs-AI division of labor line to close on→ IG reel cold open↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphoranalogy
Watch this. I'm giving ChatGPT the kind of job that you'd normally pay a motion designer for. And I'm doing it all by voice.
ChatGPT, connect to DaVinci Resolve. Build a premium motion design intro with bold typography, depth, and punchy transitions. Choose the concept yourself and end on made with AI.
Keep everything editable in Fusion. And just look at what's happening inside DaVinci. It's putting the animation together inside of the actual project.
So let's play the finished version.
The part that matters here is what we can do with that project afterwards. On a real job, you always get feedback. The timing changes, someone wants another scene, or the same animation suddenly needs to work on a phone.
I'm going to give Astra three jobs. First, explain an idea with graphics that I could use in a YouTube video. Then, invent a character and make a short animated scene.
And finally, turn one product photo into a complete ad. After each result, I'll ask for changes. I'll want different timing, another scene with the same character, or a new opening and a vertical version of the ad.
I want to see how well it handles the feedback that you'd normally be giving an editor or an animator. That's where this could save real work. Keeping the design and material we already have, and then directing the next version from the chat.
I'll show you the prompts, the projects, and the revisions. But first, let's get Astra working with Resolve on this computer. Then we'll build the animation.
For this, I'm using Astra and Codex in DaVinci Resolve Studio 21 .1. Resolve now actually has official support for AI assistance, and that's the connection that we're going to use to give Astra access to our project. Make sure that Codex is installed and that you're signed in.
Then, in DaVinci, open the File menu and choose Setup AI Assistance. Wait for the confirmation and Resolve will tell you which assistance is found and configured. Then, we're going to check that Codex is included.
Now, you may need to restart Codex for the connection to appear. Keep DaVinci running with your project open and then start a local task in Codex using Astra. Let's first check the connection.
We'll type, use the DaVinci Resolve integration to tell me which project is currently open. Once it's able to read the project, we're ready to give it the next job. Now, think about the graphics in a documentary video.
A cutout image enters the frame, a label will give it context, and the movement help explain what exactly the narrator is talking about. That's the kind of sequence that we're going to be making here. I'm using Vox as the visual reference, so the brief has a clear direction.
editorial imagery, strong typography, and scenes that always feel connected to one another. I'm going to leave the subject open -ended, and Astra has to choose a simple idea and work out how to communicate it in 15 seconds. We'll say, create an editorial -style animation for a YouTube video.
Choose one simple everyday idea and explain it visually. Give it the feel of a Vox style explainer with cutout imagery, bold typography, depth, and connected transitions. Make the central idea clear without narration.
Check the explanation and build the animation with editable elements in Fusion. I'm giving it room to choose the scenes, but the job is specific. Help somebody understand one idea.
That gives me something concrete to judge when the first version comes back. Let's open Fusion while it builds the sequence. I want the text, images, and timing controls to be easy to find.
If I later put this under a voiceover, I might need a label to appear on a particular word or an image to stay up until the sequence finishes. Keeping those elements separate gives me a way to make those adjustments. Now, here's the first version.
Let's watch it without narration.
Now I'm going to give the main explanatory moment a bit more time. When an image and a label appear together, the viewer needs an extra moment to read the words and connect them to what they're seeing. I want to test whether a longer hold makes that idea easier to follow.
I'm keeping the whole sequence at 15 seconds. Now in a real edit, I might already have the next shot coming in at that point, so the revised graphic needs to fit the same space. Make.
a second version. Give me the main explanatory moment one more time and rebalance the rest to keep the animation at 15 seconds. Keep the visual style and reuse the existing elements.
Adjust the transitions to the new timing and save the original as well. Now, with the same artwork and wording, we can judge what the timing actually does. So let's put both versions side by side.
Follow that main explanatory moment, and then watch how each version moves into the ending. The extra time has to come from somewhere. I'm checking that the longer hold still leaves enough time to read the other labels, and that the ending feels natural.
I choose the version that makes the most complete explanation easiest to follow. Now let's move from explaining an idea to telling a story. For the next job, Astra has to invent a character and give it something to do.
I'm starting with a super simple setup. A character is way too confident about an ordinary task, and then something goes wrong. Now, Astra can choose the character, the setting, and the mistake itself.
I'm keeping those open -ended because the interesting part is how it turns that setup into a visual joke. The complete scene gets 12 seconds and it has to work without dialogue. That puts the attention on the character's movement and reaction.
I'll tell it. Create a 12 -second animated scene about an overconfident character trying a simple task and getting into trouble. Invent the character, setting, and a visual joke that works without dialogue.
Use expressive movement, a consistent illustrated style, and sound effects. Give the final reaction time to land. Build the character and scene as editable fusion elements.
Let's look at the character while the scene is still coming together. For this brief, the personality needs to come through before anything goes wrong. The way it stands, approaches the task, or pauses before acting can establish the confidence that we're looking for.
Then the reaction has to finish the joke. If everything happens at the same exact speed, the important moment can disappear into the rest of the movement. I'm curious what it decides is funny.
Let's play the whole scene.
Now, we'll actually use this character for a second animation. Imagine a brand wants a recurring character for a series of posts. Once the design is approved, the next brief is another situation with that same exact character.
The proportions, color, and personality all need to carry over into the following scene. So I'm keeping the existing design and setting and simply asking for a different action in a five -second scene. I'm splitting this request into what stays, what changes, and the timing.
We keep the existing character and setting, Astra chooses a new action that fits the same personality, and it has only five seconds to show that action, land the reaction, and match the sound. Then we can compare the two scenes and see whether it still feels like the same character. Let's play them together.
The big boss wants him. The herd freezes. Every instinct says, hide.
All but one.
He walks toward it. Pay attention to the character across the cut. The situation changes, but it still needs to feel like the same person is showing up just in another moment.
Now I'm gonna direct another take of that second scene. I'm keeping the action and changing how it builds towards the reaction. The adjustment is super simple.
I want to add a little bit of anticipation before the main action and then make the response even more expressive. That gives the viewer a moment to expect something before it actually happens. I'll say, make another version of the second scene.
Add a brief moment of anticipation before the main action, and then make the reaction more expressive. Keep the character, setting, camera, main action, and story outcomes. Make sure to fit it into the same five seconds, adjust the sound, and save it separately.
Now, I'm being super specific about what stays the same because it keeps the revision focused. We already have the situation, and the pass is about the performance within it. Now, again, let's compare our two takes.
Watch that extra pause before the action and how much time the final reaction gets. The sound needs to follow that new timing as well. If an impact or a reaction moves, its sound effects need to move with it.
Then we can choose which take works better after the first scene and keep both versions in the project. This is how I'd approach a recurring character. Establish the design, build another scene from it, and direct the performance without asking for a new character every time.
Now, for the next job, we have a different starting point. I'm bringing in a real product photo, and we're going to build the visuals and edit around it.
Here's the photo I'm using. This is going to give us an actual product, including its shape, color, and branding.
Now, before creating anything, I separate the job into two parts. First, we need the supporting visuals of background and two short clips. Then, we need an edit that brings those assets together with the product, typography, and sound.
For the visuals, I'm connecting to Higgs field. Astro will use it to generate the material and then assemble the commercial in DaVinci. Open Higgs field and click MCP in the top menu.
Select the ChatGPT tab and then click Add Higgsfield Plugin. On the page that opens, click Add and sign in with your Higgsfield account. Now, you will need an active Higgsfield subscription.
The setup is simply adding the plugin and signing in. There's no API key to configure. Now, Start Creating will open a chat, but for this video, we'll continue in the Codex project that we're already using.
Before switching chats, I ask Astra to save the Resolve connection instructions and the current project name in the project notes. Then, open a new local chat in the same Codex project, keep DaVinci running, and give Astra the following set of instructions. Read the project notes, reconnect to the same Resolve project, and confirm that Higgs field is available in this chat.
Once that's ready, you'll add the product photo. If the plugin opens an upload window for the reference, then put it there. Use Hixiel to create one background image and two short supporting clips for a premium ad based on this product photo.
Choose a visual concept that suits the product and keep the assets consistent. Use those assets and the original photo to build a 15 -second ad in DaVinci. Give it a strong opening, a clear product reveal, changes in scale, and sound design.
End on a clean product shot with a short line that fits the product. Make it 16 by 9 and keep the edit, text, and transitions editable. Let's open Higgsfield and look at the assets that it created for this ad.
Here's the background followed by the two clips. I'm looking at how they all work together. The lighting and the colors need to belong to the same campaign and the composition needs room for the product.
A background can look great on its own and still be too busy behind the label. So I'm checking how it supports the product and whether it leaves enough space for readable text. Now let's go back to DaVinci and look at how those materials come together in an edit.
Astra has to choose which moments to use from the clips, place the product, and build the transitions around the timing of the app. The generated material is simply one part of that work. The opening needs to catch your attention, the reveal needs enough time to register, and the final frame needs to leave you with a clear view of what's being advertised.
Now, here's the first version, and let's let it play through.
Now we have an opening to compare with the alternative. I'm keeping the middle and the ending and asking Astra to introduce the product in a different way. This is a useful request when a client wants another creative option without replicating the entire direction.
we can work with the material already in the project. For this request, I'm changing the opening and keeping the rest of the ad in place. Astra can use the material that we already have to introduce the product differently.
The product itself, the middle, and the ending are all gonna stay the same. We'll keep it within 15 seconds and save a separate version so that we can compare the two options. That gives us another creative option using the work that we've already done.
Let's play both versions with the rest of the commercial attached.
Watch how each opening introduces the product and then follow the transition into the new section we kept. The new beginning still needs to lead naturally into the same ad. Next, we're making a vertical version for mobile.
I'll choose the opening that works best and then use that version for the mobile edit. This needs its own composition. A product placed besides a title in a horizontal frame may need to sit above it in a vertical version.
The text still needs to be readable and the product needs enough space to be recognized on a smaller screen. I'm also asking for the product to appear clearly within the first second. That gives somebody scrolling a reason to understand what the ad is about immediately.
We'll say to take the version that we just selected and make a 9x16 version for mobile viewing. Show the product clearly within the first second. Recompose each scene for the vertical frame, keep the branding and final text readable, and make the sequence understandable without sound.
Preserve the soundtrack for viewers who turn it on and save it alongside the horizontal versions. Let's watch this one without sound first. Follow the product through the sequence and check the final text.
The pictures and typography need to carry the message on their own. Now, bring the sound back in and play the opening again.
The soundtrack should support the movement while the visual sequence stays clear either way. For client work, these are three separate deliverables, the original commercial, an alternative opening, and a version for mobile. Each one starts with the same product and supporting material.
And that's the exact practical reason to keeping your project editable. Once the direction is working, a new request should build on those decisions. Having to replace the product treatment or to rebuild the scene with every revision would quickly eat into your budget.
And the cost comparison needs to cover the complete job. Include the generated assets, the software, and the time spent reviewing and correcting to edit. Every unused generation still belongs in that total.
The useful comparison is what it takes to reach the versions that you're actually going to deliver. That includes the feedback after the first render. Now let's keep the process behind those versions so the next brief has somewhere to start.
For the next product, I don't want to explain the project structure, versioning, format requirements all over again. So I'm asking Astra to save the process as a reusable skill. The instructions should cover how we supply the product, how the supporting assets are created, and how the edit is organized for revisions.
We'll ask it to save the workload for the vertical version of this ad as a reusable skill. Include the visual style, pacing, project structure, and revisions. Explain which source files I provide, assets Higgsfield creates and how to preserve that original look and feel before making changes.
The next brief can begin with another product photo and this workflow. We can then choose a visual direction that fits the new product while keeping the job organized. I'm keeping the actual DaVinci project alongside the instructions.
The project contains the work itself and the skill describes how to approach another job that's just like it. The same principle applies to the character and to the explainer. And you'll want to keep those as separate starting points because each one has distinct requirements.
A recurring character needs consistent design and explainer needs a clear visual sequence. A product ad needs accurate branding and layouts that work across both horizontal and vertical formats. That gives the next request a useful starting point with room to make new creative decisions.
Now, at the start, I gave ChatGPT the kind of job that you'd normally pay a motion designer for. The work we've put it through includes the brief, the first version, and the requests that come afterwards. For my own videos, I'd start with a short graphic that has a clear place in the edit.
Give it the brief, adjust the timing to the narration, and keep the project for the next change. That's a super manageable piece of work to judge from beginning to end. For a character series or a product campaign, consistency becomes the much bigger requirement.
The next scene has to preserve the character. The next format has to preserve the product and the visual direction. Those decisions always stay with me.
And that's exactly how I'm approaching this workflow. Astra handles the build and the revisions, whereas I set the direction and decide what's actually ready to publish. The savings depend on how much of that work carries through to the finished version.
So compare the complete job, including your own time, and keep the projects that are worth using again. Let me know which of these workflows that you want me to break down next. I'll see you guys in the next one.
The Hook
The bait, then the rug-pull.
Sanji Nai-Chien opens by handing ChatGPT a motion designer's job by voice alone, then spends the rest of the video finding out whether it can survive the part that actually matters on a real job: the feedback.
Frameworks
Named ideas worth stealing.
00:00list
The Three Jobs Test
Explain an idea with graphics (editorial explainer)
Invent a character and animate a scene
Turn one product photo into a complete ad
Sanji structures the entire video as three escalating creative briefs for the AI agent, each followed by a real revision request, so the video tests feedback-handling instead of just a single generated result.
Steal forany AI-tool review that wants to test more than a first draft
11:37concept
What stays, what changes, and the timing
When directing a revision, Sanji explicitly separates what must stay identical (character, setting, camera, story outcome) from the one thing that should change, and states the runtime constraint, which keeps the AI's revision focused instead of regenerating the whole scene.
Steal forany prompt asking an AI to revise existing creative work
18:34concept
Complete Job Cost
Generated assets
Software
Review time
The real cost comparison for AI-assisted work includes every unused generation, the subscription cost, and the human time spent reviewing and correcting, measured against the versions actually delivered, not just the accepted final render.
Steal forpricing or evaluating any AI-assisted production workflow
CTA Breakdown
How they asked for the click.
VERBAL ASK
18:56next-video
“Let me know which of these workflows that you want me to break down next.”
Soft, low-pressure engagement ask at the very end, framed as a request for direction rather than a hard subscribe or sales pitch.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
A packaged Claude Code skill turns one sentence into a fully scripted, voiced, and animated Vox-style explainer — no storyboard, no per-scene prompting, no manual editing.
Danny Why runs GPT-6 Astra Ultra through DaVinci Resolve Studio's MCP server, escalating from a 15-second animated loop to a fully illustrated 12-minute talking-head edit.
A creator builds a reusable Codex + Remotion editing template from a reference video's style, then layers AI-generated special effects on top with scripted editor's notes.