I Made a Vox-Style Explainer Video With One Prompt (Claude Code + Higgsfield)
A packaged Claude Code skill turns one sentence into a fully scripted, voiced, and animated Vox-style explainer — no storyboard, no per-scene prompting, no manual editing.
Posted
2 days ago
Duration
Format
Demo
educational
Views
8K
469 likes
Big Idea
The argument in one line.
A single packaged Claude Code skill can replace an entire Vox-style video production pipeline — research, script, voiceover, and visually consistent scene generation — collapsing days of work into one sentence of input.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You make explainer, documentary, or history/finance breakdown videos and want a faster way to produce Vox-style motion graphics.
You already use Claude Code and are curious how MCP connectors let it drive external creative tools like Higgsfield.
You're evaluating AI video pipelines and want to know what a 'skill' actually automates versus a single-prompt image generator.
You publish to international audiences and want a workflow that can localize a finished video into another language and voice.
SKIP IF…
You're looking for a hands-on editing tutorial — this video shows the pipeline running, not how to operate Higgsfield or After Effects manually.
You don't use Claude Code or have no interest in an MCP-based workflow.
TL;DR
The full version, fast.
The creator built a Claude Code skill that automates the full Vox-style explainer pipeline: it researches trending topics, proposes three options, asks a few style questions, then writes the script, generates a Higgsfield voiceover, and produces every scene's visual assets in one consistent style — all inside a single conversation, triggered by one connector set up once via Higgsfield's MCP. The core claim is that consistency across scenes, normally the expensive part of AI motion graphics, comes from the skill enforcing one visual language rather than prompting scene-by-scene. As a bonus, the same skill re-renders the finished video with a new voiceover in Spanish, Portuguese, and Hindi, keeping the visuals — turning localization into a one-line follow-up prompt instead of a separate production.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
A single Claude Code skill handled research, scriptwriting, voiceover, scene planning, and asset generation for a 10-minute explainer video from one input sentence.
Generating AI motion graphics one scene at a time is why most AI-made explainers look visually inconsistent from shot to shot.
The skill enforces one visual language across every generated scene, so shot one and shot twelve look like they came from the same designer.
Connecting Claude Code to Higgsfield for image, video, and audio generation is a one-time setup: copy an MCP connector URL and add it as a custom connector in Claude Code settings.
The only manual input in the entire production was a single sentence specifying video length and giving Claude permission to pick the topic.
The skill interviews the user with a short round of style questions — minimal versus bold typography, maps versus charts versus character cutouts, and any visual reference — before generating anything.
FIFA expects roughly $11 billion in revenue from the 2026 World Cup, over 50% more than Qatar's 2022 tournament and the biggest payday in FIFA history.
New York City budgeted about $70 million in hosting costs against just $55 million in new tax revenue, meaning the city loses money even in its best-case scenario.
The same skill and workflow generalizes beyond Vox-style videos to documentaries, educational content, history channels, and finance explainers.
Re-rendering the finished video in Spanish, Portuguese, and Hindi took one follow-up prompt and reused the same visuals with a new AI voiceover per language.
The creator frames the shift as no longer needing to learn the production pipeline yourself — you describe the result and the skill executes the pipeline.
Takeaway
One packaged skill can replace an entire AI video pipeline.
WHAT TO LEARN
The bottleneck in AI motion graphics isn't generating a single good image — it's holding one visual language across every scene, and that's exactly what a purpose-built workflow can automate away.
Generating AI scenes one prompt at a time is the main reason AI explainer videos look visually inconsistent — colors and style drift shot to shot.
A packaged workflow that plans the whole project up front (script, scene breakdown, asset list) before generating anything is what keeps a multi-scene video visually coherent.
Connecting an AI coding tool to an external creative platform via a connector is typically a one-time setup step, not a per-project chore.
Letting the tool ask a short round of concrete style questions up front (minimal vs. bold, maps vs. charts, reference or not) captures taste without requiring a full creative brief.
Researching what's currently trending before locking a topic is a cheap way to raise the odds a piece of content actually gets watched.
Once a finished video and its visual assets exist, producing a localized version in a new language can be a single follow-up instruction rather than a new production.
A workflow built for one video format (in this case, Vox-style explainers) can generalize to adjacent formats — documentary, educational, finance — without redesigning the process.
Glossary
Terms worth knowing.
MCP (Model Context Protocol)
A connector standard that lets an AI coding tool like Claude Code call out to external services — here, used to let Claude Code trigger image, video, and audio generation inside Higgsfield.
Higgsfield
An AI creative-generation platform used in this workflow to produce the images, voiceover audio, and rendered video clips that make up each scene.
Claude Code skill
A packaged, reusable set of instructions that Claude Code follows to complete a specific multi-step task — in this case, the entire Vox-style video production process, from research through final assembly.
Vox-style explainer
A documentary-style short-form video format popularized by Vox, built from narrated voiceover paired with a consistent set of illustrated or animated scene graphics.
Visual language
The shared set of colors, style, and design choices that make every scene in a video look like it belongs to the same production, rather than a series of disconnected images.
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
00:00I made this entire Vox style explainer video with one prompt using only Claude Code and Higgs Field. This Sunday, the biggest World Cup in football history crowns its champion in New Jersey and someone somewhere is quietly their to able They're And roughly $11,000,000,000 from this single tournament, which is over 50% more than Qatar earned and stands as the biggest payday in FIFA history.
00:36And honestly, I think that this completely changes the way that people create motion graphics with AI. Now just a few years ago, making a video like this would have required an illustrator, a motion designer, and hours inside of After Effects.
00:50Today, it all comes down to one sentence because the entire process from the script to the voice over to every scene you just saw was handled by one Claude skill that I built. I didn't storyboard anything.
01:04I didn't write prompts for every scene. I didn't even write the script. The skill did.
01:08Now in this video, I'm gonna show you exactly how it works from one sentence to the finished video. So let's get started. Now before we open Claude code, there's one thing that people get completely wrong about AI motion graphics.
01:23Most people are generating one scene at a time, one prompt, one image, and then they repeat that very cycle. The result ends up being a project where every scene feels slightly different. You've got different colors, slightly different style, and at the end of the day, that's not how videos like Voxes are made.
01:41A great explainer keeps one visual language from the beginning until the end. The only thing that changes is the story. And that consistency is exactly what's expensive.
01:51Scenes, shots, assets, the whole planning of the project before you even generate anything. Now if that list sounds like a lot, I've got good news for you.
02:01You don't have to do any of it because I already did. I spent an entire week figuring out how videos like these are actually made and I packed everything into one Claude skill. So here's the shift that I want you guys all to catch.
02:16The skill isn't a shortcut for people who already know the workflow. The skill is the workflow. It knows that a video like this needs a script.
02:25It knows that that script needs to break into different scenes and that every single one of those scenes needs its specific visual assets in a single consistent style. It knows that the whole thing needs a voice over and ultimately it understands how all of that gets assembled at the end. And everything that I'm about to show you, the skill, the complete workflow, all of it is in the description.
02:46The whole thing. We'll come back to that in a second. Now before the magic starts, there is exactly one manual step in this entire process.
02:53And the good news is it only takes about a minute. We need to connect Claude code with Higgs field. That way Claude can actually generate the images, the video, and the audio.
03:04So head over to Higgs Field and open MCP and CLI. Next, copy your MCP connector. Then go back to Claude code, open settings, click connectors, and choose add custom connector.
03:17Paste in your connector and you can give it any name you want. I call mine Higgs Field, then click save. That's it.
03:25Now Cloud Code can communicate directly within Higgs Field. Every generation request that you're about to see, images, voice over, the whole thing, It all goes through this connection and you never have to set it up again. Now watch how little I actually do here.
03:39This is my favorite part. I open up Cloud Code, load the skill, and all I do is type, I want a two minute explainer video. You can pick the topic, something people are curious about right now.
03:51That's the entire brief. Think about that for a second. I didn't give it a topic, a script, or upload a storyboard.
03:58I mean, I told it how long the video should be, but ultimately that's it. And here's where it really gets interesting. The skill doesn't just start generating random stuff, it actually starts a conversation with me.
04:10Now the first thing it does is it goes off and researches what's trending right now. What people are actually searching for and actually talking about. Now a minute later, it'll come back to me with three topic options.
04:22Maybe it's history, finance, or geopolitics. I read through them, pick the one I like, or honestly I could even tell it you choose and let it decide.
04:31Then it'll ask me a few questions about taste. Do I want a more minimal style or do I care about bold typography? Do I need maps, charts, or character cutouts?
04:42Do I have a visual reference that I want it to follow? Now in total, it takes about thirty seconds to answer and that's it. That's the complete list of decisions that I make in this entire production process.
04:55Everything that happens here on out goes on without me. So what happens after I answer those questions? Let's just watch because everything is happening live inside Cloud Code right on my screen.
05:06First, it'll write the script. Hook, story, payoff, it'll build everything around the topic that we've picked. Then comes the voiceover.
05:14It generates it automatically through Hakesfield audio. I don't open a single menu. I don't record anything.
05:20Honestly, I don't even hear the final product until it's done. Now the visuals. It breaks the story into scenes, and for each one of those, it'll decide what needs to be shown.
05:29Maybe it's a map, an icon, or a chart. So here it goes. The first generation request straight to Higgs Field.
05:37A few seconds later, the very first assets start coming back to me. And just look at that. Then you've got the next scene.
05:45And notice every asset comes back in the same visual language. Scene one and scene 12 look like they came from the same designer because the skill is enforcing the style on every single request. By the end, every scene has exactly the assets it needs.
06:02And because everything happens inside of one conversation, I never jump between tools or prompts or redo things. One chat, start to finish.
06:11The pipeline I spent months learning is simply running and it's doing so on its own. Once every scene has been generated, the skill asks me one final question. Would you like me to assemble the project?
06:23I simply click yes. From there, Cloud Code starts putting together everything automatically. It organizes every single generated asset and prepares everything for the final animation.
06:35Now at this point, the hardest part is already finished. So let's step back for a second and think of what just happened. It did the research, the script, the voice over, the scene planning, created dozens of consistent visual assets, and developed a project structure.
06:50And all of that came from one sentence. Instead of spending hours organizing files and building my project manually,
06:57I jumped straight into animating the scenes and exporting the final video. Now let's take a look at the final result. This Sunday, the biggest World Cup in football history crowns its champion in New Jersey, and someone somewhere is quietly left holding an enormous bill.
07:1448 teams, 104 matches, three countries, and five and a half million fans.
07:21Nothing this big has ever been staged. FIFA expects roughly $11,000,000,000 from this single tournament, which is over 50% more than Qatar earned and stands as the biggest payday in FIFA history.
07:34Contracts are simple. FIFA keeps the tickets, the sponsors, the broadcast rights, and the merchandise, while host cities pay for security, transit, stadiums, and fan zones.
07:44New York City budgeted about $70,000,000 in costs against just 55,000,000 in new tax revenue,
07:50which means the city loses money even in its best case scenario. If we go through it scene by scene, you'll notice that everything feels consistent from the beginning to the end. Every transition follows the same visual language, and it feels like it's part of a single production.
08:05And that's exactly what we wanted. That's the thing that simply cannot be recreated by generating one scene at a time. But here, it's automatic.
08:15Now the best part here is that it isn't limited just to Vox style videos. You can use the exact same skill for documentaries, educational videos, history channels, finance explainers, really for any kind of style that you can imagine.
08:29The workflow stays exactly the same. And let's remember what we started with. I told Claude that I want a two minute explainer video and I told it that it could pick the topic.
08:39Now let that sink in. That sentence and this video. The distance between those two things is the whole point.
08:47But before we wrap up, there's one last thing that I have to show you guys. Because the first time that I tried this, it genuinely blew my mind. I typed one more line.
08:56Now make this video in Spanish. New voice over and the same visuals. Now a few minutes later, it pumped out a second complete video.
09:37Now think about what that actually means. Every video that you make can be published for every audience on the planet for almost nothing. You're not just translating subtitles, you're shipping a native version of the video into every market.
09:51And that's exactly why I think this changes the way that people create videos. You don't need to learn the pipeline anymore, you simply describe the result.
09:59The pipeline already lives inside of the skill. Now everything you've seen in this video is built into that very same skill. So if you've ever wanted to create videos like these, everything you need is already waiting for you.
10:12I'm gonna leave the skill, the complete workflow, and everything I used in this video in the description below. So go out there and build something awesome. I'll see you guys in the next one.
The Hook
The bait, then the rug-pull.
The creator opens by playing the finished product first — a polished, fully-narrated World Cup explainer — before revealing it came from a single sentence typed into Claude Code.
CTA Breakdown
How they asked for the click.
FROM THE DESCRIPTION
PRIMARY CTAWhere the creator wants you to go next.