How to Create Vox-Style Documentary Animation 100% Free
A single chained chatbot prompt takes a true-crime headline from idea to a finished paper-collage documentary clip, using only free AI tools.
Posted
4 days ago
Duration
Format
Tutorial
educational
Views
56.7K
3.1K likes
57 · 43
Big Idea
The argument in one line.
A single reusable master prompt can drive a chatbot through an entire AI documentary pipeline — niche, research, script, beats, and image prompts — letting free tools alone produce a finished Vox-style short.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You want to run a faceless YouTube channel in history, crime, disaster, or mystery niches without hiring an animator.
You're comfortable directing a chatbot through a multi-step prompt chain and copy-pasting between a few browser tabs.
You want a documentary look built from paper-collage, maps, and stamped headlines using entirely free-tier AI tools.
SKIP IF…
You need an original visual style — this produces a specific 'Vox-style' collage look, not a custom brand aesthetic.
You're not willing to publish AI-narrated, faceless content.
You need finished long-form output on free credits alone — the demo itself only produces a 30-second clip before practical daily-credit limits kick in.
TL;DR
The full version, fast.
This tutorial chains free AI tools into a full documentary pipeline using one long 'master prompt.' A chatbot (ChatGPT, Claude, or DeepSeek) picks a niche, researches a true story, writes a timed voiceover script, and splits it into 14 visual beats with matching image prompts. Google Flow's Agent Mode batch-generates all 14 paper-collage-style images from a single style instruction, then a universal animation prompt animates each one into a short clip — deliberately trimmed to six seconds to conserve free daily credits. A free voice model or ElevenLabs narrates the script with performance direction like 'whisper,' and CapCut assembles the clips, speed-matches them to the narration, layers in AI-generated music, and exports — no paid software, animation skill, or on-camera presence required.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Frames the Vox-style trend as under-explained, then plays a finished MH370 clip as proof the free workflow works before explaining any steps.
01:44 – 03:13
02 · Picking a chatbot and finding the master prompt
Any of ChatGPT, Claude, or DeepSeek works; opens the creator's Google Doc containing the reusable master prompt for this workflow.
03:13 – 04:43
03 · Loading the engine-prompt source document
Pastes the master prompt into the chatbot, then uploads an ~8-page PDF covering visual styles, niches, and generation rules the chatbot needs.
04:43 – 06:11
04 · Choosing a niche and a documentary idea
Chatbot offers 8 niche categories; picks crime & documentary, gets 10 ready-made video ideas, and selects the MH370 disappearance.
06:11 – 07:58
05 · Setting run-time; AI researches and drafts the voiceover
Sets a 30-second demo length; the chatbot searches roughly 17 websites for facts, then writes and word-counts the full narration script.
07:58 – 09:58
06 · Splitting the script into 14 visual beats and image prompts
Chatbot divides the narration into 14 numbered story beats, then outputs a single .txt file containing all 14 image-generation prompts.
09:58 – 12:40
07 · Google Flow Agent Mode batch-generates all 14 images
Configures Flow (Agent Mode on, 16:9, Nano Banana image model, OmniFlash video model), saves a standing 'multi-image prompt' as agent instructions, pastes all 14 beat prompts, and the agent generates the full image set unattended.
12:40 – 14:37
08 · One universal prompt animates every image
Chatbot writes one animation prompt that works across all 14 images, rewrites it for 6-second clips instead of 10 to conserve free credits, then drags each image in one by one to animate and preview.
14:37 – 16:48
09 · Voiceover, CapCut assembly, and export
Downloads all 14 clips, generates an ElevenLabs v3 'whisper'-style narration, assembles everything in CapCut, speed-matches footage to the voiceover, adds Suno-generated music, skips captions since the collage already carries heavy on-screen text, and exports in 4K.
Atomic Insights
Lines worth screenshotting.
A prewritten master prompt can carry a general-purpose chatbot through niche selection, research, scripting, and image-prompt generation without any custom API integration.
The workflow works with ChatGPT, Claude, or DeepSeek interchangeably, since the intelligence lives in the prompt chain, not the specific model.
Breaking a finished script into numbered story beats before generating any images is what keeps an AI-produced documentary structured instead of random.
Batch-generating an entire 14-image set with one agent-mode instruction replaces 14 separate manual generate-and-wait cycles with a single queued job.
One universal animation prompt can be reused across every still image in a set, as long as the images share a consistent visual style.
Shortening AI-generated clips from 10 seconds to 6 seconds is a deliberate credit-conservation move on free-tier video generation plans.
Free-tier voice generators now support performance-style direction like 'whisper' or 'excited' without requiring professional voice-acting knowledge.
Matching video clip speed to a fixed voiceover length, rather than editing the voiceover to match the footage, keeps narration natural while hitting an exact runtime.
Heavily text-forward visual styles — maps, stamped dates, headlines — can make burned-in captions redundant rather than helpful.
A documentary can be built entirely from AI-researched facts, an AI-written script, AI-generated images, AI-animated motion, and an AI voice, with a human only assembling and directing.
Free daily credit allowances on tools like Google Flow (50/day) and Noise AI (2,000/day) are generous enough to complete a short documentary without paying.
Takeaway
One chained prompt can run the whole pipeline
AI WORKFLOW
A single master prompt can carry a chatbot from niche selection through research, scripting, beat-breakdown, and image prompts, letting free tools like Google Flow generate an entire documentary's visuals unattended.
01The hook and the 30-second proof clip
Framing a tutorial around a trend with no existing complete free walkthrough is itself a strong hook — the gap in existing content becomes the pitch.
Opening with the finished output before explaining any steps proves the workflow works before asking for a viewer's time.
02Picking a chatbot and finding the master prompt
The workflow is chatbot-agnostic: ChatGPT, Claude, and DeepSeek all work, so free-tier access to any one of them is enough to start.
A prewritten master prompt turns a general chatbot into a specialized production tool without any custom GPT or API setup.
03Loading the engine-prompt source document
Uploading a single reference document gives the chatbot all its visual style, niche, and formatting rules in one shot, instead of re-explaining them every session.
Separating the instructions (master prompt) from the knowledge (source PDF) lets the same base prompt be reused across unrelated projects.
04Choosing a niche and a documentary idea
Asking the model to first generate a menu of niches, then ten concrete video ideas inside the chosen niche, turns a blank-page problem into a multiple-choice one.
True-event documentaries work well for this format because they already have widely available source material for the chatbot to research.
05Setting run-time; AI researches and drafts the voiceover
Letting the chatbot search dozens of live sources for facts before writing the script anchors the narration in verifiable events rather than the model's own memory.
Committing to an exact runtime up front lets the chatbot word-count the script to match, avoiding a script that's the wrong length for the planned visuals.
06Splitting the script into 14 visual beats and image prompts
Converting a finished script into discrete numbered beats before generating any images is what makes an AI-produced video feel structured instead of random.
Generating all image prompts as a single downloadable file, instead of one at a time in chat, is what makes the next batch step possible.
07Google Flow Agent Mode batch-generates all 14 images
An agent mode that holds a standing style instruction separately from the per-scene prompts means every image shares the same visual language without repeating the style rules 14 times.
Batch-generating a full image set unattended turns what would be 14 manual generate-and-wait cycles into one queued job.
08One universal prompt animates every image
A single reusable animation prompt can drive motion for every still image in a set, as long as the images share a consistent visual style.
Deliberately shortening clip length to fit a free daily credit allowance is a real constraint worth planning for before generating anything, not after running out.
09Voiceover, CapCut assembly, and export
Free-tier voice tools now support performance-style direction that shapes tone without professional voice-direction skills.
Speed-matching visual clips to a fixed-length voiceover, rather than the reverse, keeps narration natural while still hitting an exact runtime target.
Skipping captions is a legitimate choice when the visuals already carry heavy on-screen text — more text on top would be clutter, not clarity.
Glossary
Terms worth knowing.
Vox-style animation
A documentary visual style built from layered paper-collage imagery, maps, stamps, and photographs animated with subtle motion, popularized by explainer-video creators.
Master prompt / Engine prompt
A single long, reusable instruction set that turns a general chatbot into a step-by-step production assistant for a specific creative workflow.
Story beat
A discrete moment or unit of a script assigned its own image and short animation, used to structure a video scene by scene.
Agent Mode (Google Flow)
A mode that lets an AI image/video tool hold a standing style instruction and generate an entire batch of assets from a list of prompts without manual, one-at-a-time generation.
Nano Banana
The image-generation model used inside Google Flow for this workflow's paper-collage-style still images.
OmniFlash
The video-generation model used inside Google Flow to animate each still image into a short motion clip.
Universal animation prompt
A single motion-description prompt written to work across every image in a set, instead of writing a unique animation prompt per scene.
“There's a new trend called Vox style animation going viral across social media right now.”
clean cold-open hook naming the trend→ TikTok hook↗ Tweet quote
00:52
“I searched online to see whether anyone had properly explained how to create these Vox style animations using only free tools, but I couldn't find a complete tutorial.”
establishes the content gap that justifies the whole video→ IG reel cold open↗ Tweet quote
01:43
“Everything you just watched was created entirely with AI using free tools.”
the reveal line right after the proof clip→ newsletter pull-quote↗ Tweet quote
13:56
“The agent handles the entire process automatically saving you a huge amount of time.”
the payoff line for the batch-generation step→ TikTok hook↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
There's a new trend called Vox style animation going viral across social media right now. Basically, creators are using layered paper style visuals and smooth motion graphics to tell documentary, history, crime, and educational stories.
And the number of views these videos are pulling in is actually insane. They're extremely satisfying to watch and at first glance, you might think they were created by professional editors using expensive software and complicated animation tools, but that's not the case. I'm about to show you how to recreate this entire style using AI.
I searched online to see whether anyone had properly explained how to create these Vox style animations using only free tools, but I couldn't find a complete tutorial. The few videos I found relied on paid software and subscriptions and you already know we don't do that here.
I always try to show you workflows that you can test without spending money. So in this video, I'm going to guide you through the complete creation process from generating the story and visuals to animating the scenes, creating the voice over, and editing everything together. I'll also give you the exact master prompt I used to create my own video.
Before the end of this tutorial, I'll explain how you can get that prompt completely free. But first, check out this thirty second clip I created using the exact workflow I'm about to show you. 03/08/2014,
Kuala Lumpur International Airport. Malaysia Airlines flight 370 lifts off for Beijing carrying 239 people.
Forty minutes later, near waypoint IGARI, its transponder signal disappears from civilian radar.
Military radar later tracks the Boeing seven seven seven turning west across Malaysia.
Everything you just watched was created entirely with AI using free tools. Now let's break down the full process step by step. For this workflow, the first thing you're going to need is a chatbot.
As you can see, you can use DeepSeek. DeepSeek is a Chinese chatbot that gives you free access, so you don't need to pay for anything. You can also use Claude.
Claude is another powerful chatbot and one of my personal favorites. Simply create an account, sign in, and use the free version available on the platform. But for this video, we're going to be using ChatGPT which is still one of the best chatbots for almost anything you want to create.
Once you have your chatbot open, the next thing you'll need is access to the Google document I created specifically for this workflow. Inside this Google document, you'll find the complete master prompt I built for creating these Vox style animations. I took my time to build this prompt carefully and every instruction inside it serves a specific purpose.
Once you get access to the document, copy the first part of the master prompt. Make sure you copy everything exactly as I'm doing so you don't leave out any important instructions. After copying the prompt, return to your chatbot and paste it into the chat.
Once you paste it, wait for the chatbot to format and understand the document properly. After that, click send. The chatbot will immediately begin processing the prompt.
This should only take a few seconds. Once it is finished, it will ask you to attach the source material required to begin the workflow. The source material can also be found inside the Google document.
Scroll down until you see the section labeled engine prompt. Click on the engine prompt link and it will take you to a Google Drive document containing approximately eight pages of information covering the entire workflow. Once the document opens, download it to your device.
It doesn't matter whether you're using a phone, laptop, or desktop computer. Once the file has been downloaded, it will be saved directly to your device. Now return to your chatbot and upload the downloaded file.
Since I'm using a laptop, I'll simply drag the document into chat GPT and drop it inside the chat. Once the file has been attached, click send. This document contains all the important information the chatbot needs to execute execute the workflow correctly, including the different visual styles, storytelling formats, niches, and generation instructions.
After the chatbot processes the file, it will ask you which category of Vox style animation you want to create. As you can see, we have options such as crime and documentary, history, money and power, disaster and survival, mysteries and the unexplained, technology, sports, and several other categories.
You can also type in your own custom niche if the one you want is not included in the list. For this demonstration, I'm going to choose crime and documentary. All I need to do is type one into the chat and click send.
Once you make your selection, the chatbot will automatically generate 10 different video ideas within that category. As you can see, it gives us ideas such as how the nineteen seventy eight Lufthansa heist unfolded at JFK Airport, what really happened to Malaysia Airlines flight three seven o in 2014, and several other interesting documentary For this demonstration, I'm going to select idea number two.
What really happened to Malaysia Airlines flight three seven o in 2014. I'll type two into the chat and click send. The chatbot will now ask how long I want the video to be.
You'll have different options starting from approximately thirty seconds and going up to five minutes. However, you can also enter a custom duration. You could create a ten minutes, twenty minutes, forty minutes, or even one hour documentary depending on the type of content you want to produce.
Since this is only a tutorial, I'm going to keep the video at thirty seconds. I'll type thirty seconds into the chat and click send. Once I send that, the next phase of the workflow begins.
The chatbot will search the web and gather the necessary information about the selected story. It can analyze news reports, forums, discussions, Reddit posts, YouTube videos, Google results and other available sources to collect the most relevant details.
This helps make the final story more accurate, structured and suitable for a documentary style video. We simply need to wait patiently while the chatbot completes its research. Once the research is complete, the chatbot will begin counting and arranging the words paragraph by paragraph.
After that, it will generate the complete voice over script for the documentary. As you can see, the script begins with the incident that occurred on 03/08/2014. For now, we're going to leave the voice over script where it is because I want us to generate the complete visual animation before returning to create the voice over.
Later in this tutorial, I'll also show you some free tools you can use to generate professional voice overs. To continue the workflow, scroll to the bottom of the response and type proceed. Once you type it, click send.
The chatbot will now analyze the voice over script and divide the story into individual visual beats. These beats represent the different scenes we need to generate and animate. As you can see, the chatbot has created a total of 14 beats.
That means we're going to generate 14 different images and animate all 14 of them. Once the story beats are ready, scroll to the bottom and type next. Click send and the chatbot will begin creating the image generation prompts for the entire workflow.
Don't worry, the image generation process will be much easier than you might expect. We're not going to copy and generate every prompt individually. Instead, I'll show you a faster workflow that generates all the images automatically.
After processing the story beats, the chatbot will provide a TXT file containing all the image generation prompts. Click on the TXT file to open it. Inside, you'll find the complete image prompt for each of the 14 story beats.
Select and copy all 14 image generation prompts. Once you've copied them, the next tool you'll need is Google Flow. Open Google Flow and create a new project.
Google Flow provides 50 free credits every day and those credits refresh every twenty four hours. We're going to use those free credits to generate the images and animate them. Once your Google flow project opens, you should see an interface similar to the one on my screen.
The first thing you need to do is turn on agent mode. After turning on agent mode, click the menu icon and open the settings tab. For the image generation settings, set the aspect ratio to 16 to nine.
Leave the output at one times and keep the image generation model set to nano banana. For video generation, also keep the aspect ratio at 16 to nine.
Leave the output at one x and select OmniFlash as the video model. Once everything is properly configured, save your settings. After saving the settings, click the section labeled agent instructions.
Select add instruction. Now return to the Google document and scroll to the bottom. Inside the document, you'll find another prompt called the multi image prompt.
Copy the complete multi image prompt exactly as it appears. Return to Google flow and paste it into the agent instructions box. Once the prompt has been pasted, click done.
Now return to chat GPT and copy the complete set of 14 image generation prompts again. We need to recopy them because the multi image prompt is already saved inside Google flow as the main instruction. Return to Google flow and paste all 14 image prompts into the main prompt box.
Check that everything has been pasted correctly and then click send. Agent mode will now begin reading and analyzing the prompts to understand exactly what we want to create. Once it has finished processing the instructions, it will automatically begin generating all the images.
You don't need to copy each image prompt individually or generate the images one by one. The agent handles the entire process automatically saving you a huge amount of time. As you can see, it is generating all the scenes for us without the stress of repeatedly copying prompts between chat GPT and Google flow.
Now we simply need to wait patiently until all the images are ready and just like that our complete set of images has been generated. You can see how clean and detailed everything looks. The images match the exact Vox style animation aesthetic we're trying to recreate with newspaper textures, photographs, maps, documents, bold typography, and layered visual elements.
Now that all the images have been generated, the next step is to animate them individually. Return to Google flow and turn off agent mode. Open the settings tab and switch to video mode.
Keep the aspect ratio at 16 to nine, leave the output at one times and set the video duration to six seconds. Six seconds is the best option for this workflow because we're using free credits and we don't want to waste too many credits generating longer clips. Now return to chat GBT and close the TXT file.
Scroll to the bottom of the response until you see the instruction asking you to type next for the video prompt. Type next and click send. The chatbot will generate a universal animation prompt that can be used for every image.
We don't need to create a different animation prompt for each scene. This one universal prompt will work across the entire set of images. However, you may notice that the original prompt instructs the video generator to produce ten second clips.
Since we want six second clips, return to the Google document and locate the prompt called timing. Copy the timing prompt and paste it into chat GPT. This instruction tells the chatbot to rewrite the universal animation prompt for six second videos instead of ten second videos.
Click send and wait for it to regenerate the prompt. Once the updated universal animation prompt is ready, copy the entire thing. Return to Google flow and drag the first image into the timeline.
Paste the universal animation prompt into the prompt box and click send. Before generating the rest of the clips, we'll wait and preview the first animation to make sure the prompt works correctly. The animation process usually takes around ten to fifteen seconds, so wait patiently while Google Flow processes the image.
And just like that, our first image has been animated. Let's preview the result.
You can see how the visual elements slide into the frame, the pin attaches the photograph to the board and pieces of tape appear to hold the image in position. The movement feels clean, controlled and very similar to the Vox style documentary animations we're trying to recreate. Now that we know the prompt works, we're going to repeat the same process for the remaining images.
Drag the second image into the timeline, paste the universal animation prompt and click send. Then drag in the third image, paste the same prompt and generate it again. Continue repeating this process until every image has been animated.
I'm not going to make you watch the entire generation process so I'll speed this section up slightly. We're now down to the final image. We just need to wait for that last clip to finish processing and just like that the final image has been animated successfully.
When we preview the clips you can see how smoothly everything flows from one image to the next. Every image has now been generated and animated. The next step is to download all the videos individually.
Download each clip one at a time and make sure they are saved in the correct order. As you can see, I'm downloading my clips in ten eighty p because I'm using a pro account. If you're using the free version, you can download your clips in seven twenty p and upscale them later during the editing stage.
Once all the clips have been downloaded, return to chat GPT. Scroll back to the top until you find the voice over script generated earlier. Copy the entire voice over script.
To create the voice over, you can use Noise AI which provides approximately 2,000 free credits every day. You can also use eleven Labs which is one of my personal favorites and the tool I'll be using for this demonstration. Open 11 labs and select the v three voice model.
With the v three model, you can add performance keywords that control how the narration sounds. For example, you can use instructions such as fast, slow, excited, calm, or whispering.
For this video, I'm going to use whisper because I want the narration to sound slightly mysterious and documentary like. After entering the keyword, paste the complete voice over script into the text box and click generate. Once the voice over is ready, listen to the available options and choose the one that best matches your video.
I like the way this version sounds so I'm going to download it. Now we have all 14 animated clips and the complete voice over ready. The next stage is to combine everything inside an editing program.
For this tutorial, I'm going to use CapCut. Open CapCut and create a new project. Click import and select all the downloaded video clips, the voice over and any additional audio files you plan to use.
I saved everything inside one folder so I'll import the complete folder directly into CapCut. As you can see, all 14 video clips have now been imported successfully. The first thing I'm going to do is drag the voiceover into the timeline.
We'll use the narration as a guide to arrange the video footage in the correct order. The first clip begins with the March 8 date, so I'll drag that video into the beginning of the timeline. After that, I'll continue arranging the remaining clips according to the order of the story.
I'll speed this process up slightly while I place all 14 videos into the timeline. And just like that, all the footage has been arranged correctly. However, you may notice that the voice over is shorter than the complete video sequence.
To fix this, select all the video clips at once. Open the speed settings and slightly increase the speed until the footage matches the length of the narration. As you can see, everything now fits together perfectly.
The next thing I like to add is background music. Click import and upload the background soundtrack you want to use. You can generate royalty free background music with Suno AI or another free AI music generator.
Once the music has been imported, drag it underneath the voice over in the timeline. Trim the excess audio so the music matches the full length of the video. Reduce the background music volume so it doesn't overpower the narration.
You can also increase the volume of the voice over slightly to make sure every word remains clear. Normally, I would also add captions or extra text. However, this Vox style animation already contains a lot of written information, maps, headlines, labels, and visual details.
Adding more text could make the screen feel overcrowded, so I'm going to leave it out for this particular video. Play through the project one final time and make sure the visuals, narration, and background music are properly synchronized. Once everything looks good, click export.
You can also give your project a name. For this demonstration, I'll simply name it Vox. If your clips were downloaded in ten eighty p, export the final project in ten eighty p or upscale it to four k.
If you downloaded the free seven twenty p versions, you can also upscale the final project during export for a cleaner result. Click export, wait for CapCut to process the project, and your final VOX style documentary animation is ready.
I'll see you in the next one. 03/08/2014,
Kuala Lumpur International Airport. Malaysia Airlines flight three seven zero lifts off for Beijing, carrying 239 people.
Forty minutes later, near waypoint IGARI, its transponder signal disappears from civilian radar.
Military radar later tracks the Boeing seven seven seven turning west across Malaysia. Satellite signals suggest it flew south for hours toward the Indian Ocean. Debris later washes ashore thousands of kilometers away, but repeated seabed searches find no main wreckage or flight recorders.
The ocean still holds MH370.
The Hook
The bait, then the rug-pull.
Vox-style paper-collage documentaries are all over social media, and — the creator claims — every existing tutorial he could find relies on paid software. He opens with a finished 30-second MH370 clip as proof, then reveals it was built end to end with one chained chatbot prompt and a stack of free tools.
Frameworks
Named ideas worth stealing.
02:10model
The Vox-Style Animation Master Prompt
Engine prompt (style + rules)
Niche selection
10 video ideas
Topic + runtime pick
Web research
Voiceover script
14 story beats
Image prompts
Universal animation prompt
Timing override (10s to 6s)
A single chained chatbot prompt sequence that walks a documentary idea from niche selection all the way to per-scene image and animation prompts, one state at a time inside one chat thread.
Steal forAny AI content pipeline that needs to chain research, scripting, and visual-prompt generation without manually re-prompting at every stage.
A three-line prompt that fans Claude out into paired builder and critic sub-agents until every piece clears a stated quality bar — and the one condition that decides whether it helps or hurts.
A content director runs 336 unsorted vlog clips through a Claude Code + DaVinci Resolve Studio pipeline that classifies A-roll from B-roll, proposes cutaway placements against four editorial rules, and drops the picks onto a real timeline — then shows exactly where it still needs a human.
A five-minute walkthrough of Claude Code's Output Styles feature — the config Anthropic's own team reaches for when responses turn into a jargon-dense wall of text.