Modern Creator
FuturMinds · YouTube

GPT-6 Astra Made This 3D Explainer Video (No Editor)

A walkthrough of Codex and GPT-6 Astra generating four fully narrated 3D explainer videos, from a GPU teardown to space debris, without opening a video editor.

Posted
3 days ago
Duration
Format
Tutorial
educational
Views
2.3K
17 likes
Part of the collectionThe GPT-6 Astra PlaybookEvery GPT-6 Astra breakdown, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

Codex running GPT-6 Astra can direct an entire 3D explainer video end to end, from scene generation through narration, music and sound, by calling a stack of pre-built creative skills instead of a video editor.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You want to produce narrated 3D or cinematic explainer videos without touching a traditional video editor.
  • You already use Codex or a similar coding agent and want to see it directed toward video and creative production instead of software.
  • You're evaluating whether GPT-6 Astra can run a multi-tool creative pipeline, 3D, video generation, voice, music, from natural-language prompts.
  • You want real numbers on how long AI-generated explainer videos take and what they actually cost to produce.
SKIP IF…
  • You're looking for a drag-and-drop editor tutorial. This workflow is entirely prompt-and-code driven, inside a developer terminal.
  • You need production-ready, broadcast-accurate visuals today. The creator flags these as simplified schematics and illustrative reconstructions, not verified technical diagrams.
TL;DR

The full version, fast.

A creator wires OpenAI's Codex agent, running the new GPT-6 Astra model, into a full video pipeline: Three.js and GSAP build and animate 3D scenes, HyperFrames arranges footage, narration, music and sound effects into a timeline, and fal (for Veo 3.1) and Google Gemini generate footage and voice. One setup prompt installs the whole toolchain. From there, four demo videos, a GPU teardown, a cinematic Egypt sequence, a two-ball physics experiment, and a space-debris documentary, are produced from a single prompt plus one round of revisions each. Render times run 20-45 minutes, paid generation costs range from near-zero to $6.50, and the closing advice is to cap spend, request short proofs, and save working styles as reusable skills.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0000:14

01 · Intro

Cold open stating the premise: these videos were made with GPT-6 Astra in Codex plus animation tools and AI services, prompts and resources shared in the creator's free community.

00:1401:40

02 · How Setup Works

Overview of the pipeline: Three.js builds 3D objects, GSAP animates them, HyperFrames arranges the timeline, and generation services like fal and Gemini supply footage, voice and music.

01:4002:17

03 · Setting up the tools in Codex

A one-time setup prompt has Codex install Node.js, the HyperFrames CLI and skills, Three.js, GSAP, and fal/Gemini API clients inside a new Codex project.

02:1704:22

04 · Demo 1: GPU explainer

A 90-second GPU teardown video generated from one prompt, then refined once for smoother movement and cleaner sound, walking from the physical package down to tensor cores.

04:2206:28

05 · Demo 2: Egypt video

A cinematic pyramid-construction sequence using paid generation (Veo 3.1, ElevenLabs, Stable Audio 3), costing $3.80 against a $4 cap, refined once for camera pacing.

06:2807:53

06 · Demo 3: two-ball physics experiment

A physics explainer comparing equal-size, different-mass balls falling through air, visualizing gravity, drag and buoyancy up to terminal velocity.

07:5308:51

07 · Demo 4: space video

A space-debris documentary sequence, validated first with a 10-second proof clip, citing real orbital-debris figures before rendering the full flight.

08:5109:38

08 · Bonus video and actual costs

A bonus cyclotron video is mentioned, weekly Codex usage (about 65% of a 5x plan) is disclosed, and per-video paid generation costs are broken down.

09:3810:08

09 · Tips

Closing practical advice: cap total spend including retries, request short proofs, reuse good assets, save working styles as skills, and storyboard before generating.

Atomic Insights

Lines worth screenshotting.

  • Codex, running OpenAI's GPT-6 Astra model, can direct 3D animation, AI video generation, narration, music and sound design as one agent instead of a team of specialized tools.
  • A 90-second AI-generated GPU teardown video, narrated and scored, took about 20 minutes to produce from a single prompt.
  • One follow-up prompt asking for smoother camera movement and cleaner sound effects added 30-35 minutes without requiring a full rebuild of the video.
  • A 45-second cinematic Egypt pyramid sequence, using paid video, voice and music generation, cost $3.80 against a self-imposed $4 spending cap.
  • Across four AI-generated explainer videos, the creator used about 65% of a weekly Codex usage allowance on a 5x subscription plan.
  • Per-video paid generation costs ranged from $0 for the GPU explainer, which used a free local voice model, to $6.50 for the space-debris documentary.
  • Requesting a 10-second proof clip before committing to a full render let the creator validate camera direction without paying for a wasted full-length generation.
  • The workflow separates narration, music and sound effects onto independent tracks so any one layer can be revised without regenerating the whole video.
  • A reusable skill file lets a working visual style be saved and reapplied to future videos instead of re-explaining it in every new prompt.
  • The two-ball physics demo modeled gravity, drag and buoyancy as separate forces, showing how drag grows until it balances gravity at terminal velocity.
  • The space-debris sequence cited roughly 47,000 tracked orbital objects and 1.5 million untracked fragments between 1 and 10 centimeters.
Takeaway

One Agent Can Direct an Entire Production

AI Video Pipeline

Codex running GPT-6 Astra orchestrated a full explainer pipeline, from 3D scene generation through narration, music and sound, by calling a stack of pre-defined creative skills instead of one giant prompt.

02How Setup Works
  • The pipeline chains four tools: Three.js builds 3D objects, GSAP animates their movement, HyperFrames arranges the timeline, and generation services like fal and Gemini supply footage, voices and music.
  • After the footage lands, the agent adds explanatory graphics, generates narration from the same script, and drops music and sound effects onto their own separate tracks before final render.
03Setting up the tools in Codex
  • A one-time setup prompt has Codex install Node.js, the HyperFrames CLI and skills, Three.js, GSAP, and API clients for fal and Google Gemini inside a single project.
  • Answering the setup wizard's questions about API keys and installs is the only manual step; everything downstream runs from natural-language prompts in the same Codex session.
04Demo 1: GPU explainer
  • A 90-second GPU teardown video, generated fully locally with narration, music and sound effects, took about 20 minutes to render from prompt to first preview.
  • Asking for one narrow revision, more detail and smoother movement through the same object, added another 30-35 minutes rather than requiring a full rebuild.
05Demo 2: Egypt video
  • The cinematic Egypt sequence used paid generation, Veo 3.1 for footage, ElevenLabs for narration and Stable Audio 3 for construction sound effects, all assembled through HyperFrames.
  • With a $4 cost cap set in advance, the full cinematic pyramid video cost $3.80 to generate, including consistent reference images, camera moves and narration.
06Demo 3: two-ball physics experiment
  • The two-ball physics demo specified equal-size, different-mass balls, a paused moment to explain the forces at play, and a vacuum comparison, and rendered in about 20 minutes.
  • The explainer visualized gravity, drag and buoyancy as separate forces, then showed how drag grows until it balances gravity at terminal velocity for each ball.
07Demo 4: space video
  • For the space-debris film, the creator validated direction first by requesting a 10-second proof clip before committing to render the full connected flight sequence.
  • The finished sequence cited real figures: roughly 47,000 tracked objects in orbit, 1.5 million smaller fragments, and a 2009 collision that created over 2,000 tracked pieces of debris.
08Bonus video and actual costs
  • Across all four demos the creator used about 65% of a weekly Codex usage allowance on a 5x plan, on top of separate per-generation charges from fal and Gemini.
  • Recorded external generation costs varied by video: the GPU explainer had none recorded, the pyramid video ran about $3.80, physics about 3 cents, and the space video about $6.50.
09Tips
  • The practical advice: set a hard total cost cap that includes retries, request a short proof clip before committing to a full render, and reuse footage that already works.
  • Once a visual style proves out, save it as a reusable Codex skill, and have the agent research the topic and draft a storyboard before generating any video.
Glossary

Terms worth knowing.

Codex
OpenAI's coding agent CLI that can install tools, write code and execute multi-step workflows from natural-language prompts, used here for video production instead of software.
GPT-6 Astra
OpenAI's newly released flagship model, used in this video as the model powering Codex's decisions.
HyperFrames
A CLI and set of skills (structure, motion, review, sound) that arranges 3D video, graphics and audio into an editable video timeline.
Three.js
A JavaScript library used to build the 3D objects and scenes shown in the explainer videos.
GSAP
An animation library used to control camera and object movement inside the generated 3D scenes.
fal
An AI generation marketplace accessed via API, used here to reach video, voice and music models such as Veo.
Veo 3.1
Google's video generation model, accessed through fal, used to animate a reference image into moving footage from a written camera prompt.
Google Gemini
Used for selected video and voice workflows in the pipeline, including Gemini TTS narration.
ElevenLabs
A paid AI voice-generation service used for narration in the cinematic Egypt demo.
Stable Audio 3
An AI model used to generate music and construction sound effects for the Egypt demo.
Kokoro
A free, local AI voice model used for narration on the GPU explainer so that video avoided paid voice generation entirely.
Resources

Things they pointed at.

00:34toolHyperFrames skills
00:41toolGSAP
00:49toolfal (fal.ai)
00:49toolGoogle Gemini
01:40toolKokoro (local voice model)
04:05toolVeo 3.1 (via fal)
04:05toolElevenLabs v3
04:05toolStable Audio 3
Quotables

Lines you could clip.

04:13
I think this is just amazing. These kind of animated videos required a lot of skill and money in the past. These can truly revolutionize education space.
Direct, enthusiastic verdict that summarizes the whole video's thesis in one breathTikTok hook↗ Tweet quote
09:36
My practical advice would be to set a total cap that includes retries, ask for a short proof and reuse good assets.
Concrete, actionable tip stated cleanly with no setup needednewsletter pull-quote↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

These videos were made using GPT -6 Astra in Codex with a few animation tools and AI services. I described what I wanted, reviewed the results and asked for changes. You can find all the prompts and resources in my free school community.
First, here's the simple setup behind what you're watching. GPT -6 Astra is OpenAI's newly released flagship model and I wanted to see what it could actually build. So here we are using Codex which is our AI agent and it is powered by a GPT -6 Astra model.
I've used HyperFrame skills, including the explainer workflow, animation, and audio guidance. Then we give it access to the tools. 3 .js builds 3D objects, GSAP controls their movement, HyperFrames arranges the video and graphics, and we can connect generation services such as Fall and Google Gemini.
Fall gives us access to different models for footage, voices, and music in one place. We used it for some of our generations. For example, our cinematic pyramid explainer.
So when we ask Codex to create that video, first image generation creates a consistent reference for the location and its look. Then Astra sends that image and a movement prompt to a video model called VO and we are accessing it through fall. The return shots go on to the video track.
Now we have the footage, but we still need to explain what we are seeing. So Astra adds the freeze, the outline and the label as a graphic layer. Next comes the narration.
The picture goes to the voice model. and audio is aligned with the pictures. Then music and the sound effects go on to their own tracks.
Finally, we review the whole thing and render it into mv4. Now that's the process you're watching here. Now let's set up these tools.
Let's start a new session in Codex and paste this setup prompt. We are asking it to set up this Codex project for 3D explainers and cinematic AI videos. It will install and set up Node .js, Hyperframe CLI and all of its related skills.
and 3 .js and gsap we are also asking it to set up fall client for vo 3 .1 11 labs and stable audio and it should also set up google gemini integration which we will use for audio generation just answer all the questions it raises and follow all the instructions to set up the api keys etc and this will complete successfully and now we are ready to make our first video let's open a new session in the same project And let's start with the first video which is GPU explainer.
I'll paste this prompt. I'm asking it to create a 90 second explainer video. The important part is the journey.
Take the module apart, move inside the chip, then put it back together. We are also asking for narration, music, sound effects, all generated locally. Now let's send it.
Okay, so the first preview is ready. And that took around 20 minutes. I'd like to push this further and let's ask for more detail and smooth movement through the same object.
I also want the music to develop more and the effects to be cleaner. I'll ask it to change only those.
Alright, so the final export is ready. From sending the prompt to this file, it took around 30 -35 minutes. Now let's watch the full video with sound on.
It removes heat. Below it, the board connects the central package to the server. Now we separate the package from the board.
The compute chip and high bandwidth memory sit together on this package. H100SXM has 80GB of HBM3 memory. These stacks store data for the processor.
The exploded gaps are only for explanation. The parts return to position. We move inward from the physical package to an architectural schematic of the die.
An on -chip cache helps reuse data and reduce trips to external memory. Beyond it, many streaming multiprocessors, called SMS, work in parallel. Inside 1SM, tensor cores accelerate matrix multiply and accumulate operations, multiplying groups of numbers, then adding the results.
This is a central ingredient in iWorkloads. Shared memory lets cooperating threads exchange data nearby. The diagram shows representative units, not a transistor floor plan.
Cooling, memory, and parallel computation all support the work. I think this is just amazing. These kind of animated videos required a lot of skill and money in the past.
These can truly revolutionize education space. Now let's move on to our second demo. Let's start another session and this time we'll try a different look.
A cinematic journey through ancient Egypt and here's the prompt. There are three important things to notice. Consistent reference images, a camera route to the work site and a pause that explains one stone.
There's also a $4 cap on the generated media. Let's run it.
Alright, so this is finally done and it took around 45 minutes to complete. And by the way, just for your reference, I haven't enabled fast mode in my plan. That's why it's taking a little more time.
It used CDAN's VO3 .1, 11laps V3 for narration and stable audio 3 for the construction sounds and hyperframes for assembling everything together. And it costed around $3 .88 because we had set up a cap of $4. And these are some of the images it created.
It then created videos from these images and put them together. Let's watch the video.
I want the camera movement to feel smoother here.
Let's ask for a gentle descent. and a slower climb reusing the footage we already have.
Alright, so this revision is complete. Let's run the final video.
toward the sky ropes tighten levers turn effort into movement one stone settles a civilization rises the thing to keep in mind is that we just gave very simple prompts and we refined only once and that has led to this output which is really good now let's move on to our third demo and this is going to be a little different let's start a new session for our two ball experiment and paste this prompt basically in this video we want to illustrate why a heavier ball would drop faster than a lighter ball in the presence of air.
So we are specifying equal size balls with different masses, a pause to explain the forces and a vacuum comparison at the end. Let's send it and see how far this prompt gets us.
It took around 20 minutes to complete and this is the final video it has created.
Freeze the fall. Gravity pulls down. Drag pushes up.
Buoyancy adds a smaller lift. At the same speed, both feel the same air forces. But on ten times the mass, their effect is ten times smaller.
Faster motion builds drag until the forces balance. That's terminal velocity. Still moving, no longer accelerating.
And all of this was created locally, except the audio, which is great. Now for the space example, here's the prompt I used. We're asking for a connected flight past satellites and debris, with labels and narration added afterward.
I asked it to give me a 10 -second proof first, so I validated the direction, gave the inputs, and asked it to complete the whole film. Let's watch it.
Above the clouds, an invisible highway circles Earth. 400 kilometers up, the space station spans a football field and completes an orbit every 90 minutes. Higher, satellites share space with broken panels, rocket bodies, and metal shards.
Across Earth's orbits, we track nearly 47 ,000 objects. Another 1 .5 million fragments measure 1 to 10 centimeters. A pea -sized chip can disable a satellite.
Rocks are natural meteoroids. This wreckage is ours. In 2009, one satellite collision created over 2 ,000 tracked fragments.
Above 1 ,000 kilometers, debris can remain a century or longer. Every collision creates more targets. Which one gets hit next?
I created a couple of more videos to test the capabilities, and this cyclotron video was one of them, and I found it really interesting. You can watch the full video in my free school community. Across creating these videos and playing around with the results, I used around 65 % of my weekly codex limit.
on my 5x setup on top of that some services charge for the media they generate for example fall and gemini the final gpu video had no recorded external paid generation api charge whereas pyramid video used around 3 .80 the two ball physics video was about 3 cents across its recorded versions and the space video used about 6 .50 so i hope you got an idea of how much it costed and Where's the reference?
So my practical advice would be to set a total cap that includes retries, ask for a short proof and reuse good assets. Once you find a style that works for your niche, you can save the useful parts as a Codex skill. And before you ask Codex to start working on a video, have it do some research on your topic, create a proper storyboard, and then once you're happy with it, then jump into the video part.
I hope you found this video useful. If you have any suggestions or questions, please let me know in the comments. Thanks for watching and I'll see you in the next one.
The Hook

The bait, then the rug-pull.

The title's claim is the setup: no editor, just Codex directing GPT-6 Astra through a stack of 3D and generation tools. What follows is the actual build log, prompts, tool list, revision requests and dollar costs included, for four separate AI-made explainer videos.

Frameworks

Named ideas worth stealing.

00:34model

HyperFrames Skills Wheel

  1. Structure
  2. Motion
  3. Review
  4. Sound

The four skill categories HyperFrames uses to turn a single creative brief into a finished video: structuring the timeline, animating motion, reviewing output, and handling audio.

Steal forAny prompt-driven creative pipeline that needs to break one request into repeatable sub-skills.
09:36list

Cost-control checklist

  1. Set a total cost cap that includes retries
  2. Request a short proof clip before committing
  3. Reuse footage that already works
  4. Save a working style as a reusable skill
  5. Research the topic and draft a storyboard before generating

The creator's closing checklist for running paid AI generation workflows without runaway costs.

Steal forAny AI video or content workflow that calls paid generation APIs.
CTA Breakdown

How they asked for the click.

VERBAL ASK
00:20link
You can find all the prompts and resources in my free school community.

Referenced in the cold open and again near the end pointing to a bonus video; the description carries the Skool community link plus separate Hostinger VPS and n8n affiliate links.

MENTIONED ON CAMERA
FROM THE DESCRIPTION
PRIMARY CTAWhere the creator wants you to go next.
OTHER LINKSAlso linked in the description.
Storyboard

Visual structure at a glance.

open
hookopen00:00
behind the scenes
promisebehind the scenes00:14
GPU teardown 3D render
valueGPU teardown 3D render03:04
closing CTA
ctaclosing CTA10:08
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.

23:59
Riley Brown · Listicle

GPT-6 Astra Feels Like AGI

One creator spent a week and $1,500 in credits with early access to OpenAI's new model and came away with seven reasons he thinks the agent era just changed.

September 7th