A Free Claude Code Harness That Edits Video Without a Timeline
Kev open-sources the agentic editing stack he used to cut this very video, six reusable skills and a mascot-animation trick included.
Posted
yesterday
Duration
Format
Tutorial
hype
Views
1.7K
29 likes
57 · 43
Big Idea
The argument in one line.
An AI agent can now take a raw voiceover and produce a fully edited video in one shot, which means the timeline-based editor stops being the only way to cut a video.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You're a solo creator who edits your own talking-head videos and want to automate captions, cuts, and animated graphics instead of hiring an editor.
You already use Claude Code and are comfortable running local Python, ffmpeg, and Whisper to set up a new agent workflow.
You want to reverse-engineer another creator's editing style into a repeatable skill instead of eyeballing it shot by shot.
SKIP IF…
You need frame-accurate manual control over every cut. This hands editing decisions to an agent, not a timeline you drag clips on.
You're not willing to install and run a local Python/ffmpeg/Whisper environment. This is a developer-grade tool, not a hosted app.
TL;DR
The full version, fast.
Kev argues the editing timeline is becoming optional: his free, open-source Claude Code harness takes a raw voiceover and produces a fully cut, captioned, and animated video in one shot. Every dropped video splits into two agent pathways, Whisper for word-level timestamps and OpenCV for face-crop framing, which the agent then uses to choreograph captions, graphics, and a custom mascot file (marks.py) that animates any logo or brand asset. The repo ships six skills out of the box, including a /reverse-engineer skill that watches any video and builds a new editing skill that copies its style, which Kev used to clone a Fireship-style edit for this exact video. The takeaway: install the repo, point an agent at your voiceover, and let it build and refine its own editing skill through conversation rather than a manual timeline.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Kev argues AI agents now complete entire video edits in one shot, the same way they vibe code, and declares the CapCut/Premiere timeline era over.
01:16 – 01:40
02 · Why I'm retiring HyperEdit
Kev reveals his previous open-source video editor, HyperEdit, is now obsolete because it still relied on a traditional timeline.
01:40 – 02:06
03 · The new open source repo
Kev introduces the new repo, an agentic harness that lets anyone run the same free AI editor he used to build this video.
02:06 – 02:21
04 · The tools (quick version)
Kev skips a tool-by-tool rundown, saying how the system works matters more than listing every backend tool.
02:21 – 03:38
05 · marks.py: bring any character or logo to life
The most important file in the stack: it lets the agent animate any imported character, mascot, or brand asset, including the Claude Invader logo, and pre-load brand colors and typography for reuse.
03:38 – 04:09
06 · Giphy API: every meme on earth, free
The agent can pull reaction GIFs automatically from the free Giphy API based on the topic or joke being made in the voiceover.
04:09 – 04:55
07 · How it actually works
Kev states the goal plainly: delete the timeline entirely, and let a creator have a full back-and-forth conversation with the agent until the edit matches what they want.
04:55 – 05:21
08 · Pathway 1: Whisper word-level timestamps
Every submitted video first runs through local, free Whisper transcription, which produces word-level timestamps that every caption, beat, and counter syncs to.
05:21 – 05:47
09 · Pathway 2: vision, face detection + the band
The second pathway uses OpenCV to find the speaker's face, crop it, and build the framed band the rest of the edit is composited around.
05:47 – 06:10
10 · Choreography: animations, graphics, effects
With the transcript and vision data combined, the agent choreographs captions, graphics, icons, logos, and effects into a customized pipeline unique to that video.
06:10 – 06:22
11 · render.py + the QA agent
After editing, render.py merges every file into the final video, and a separate QA agent watches the render to catch mistakes before it's done.
06:22 – 06:56
12 · The 6 skills inside the repo
Kev lists the skills shipped with the repo out of the box: short-form and long-form talking-head editors, reverse-engineer, and Fireship-style long-form and short edits.
06:56 – 07:28
13 · /reverse-engineer: learn any editing style
Kev explains how /reverse-engineer studies any video online and builds a new skill that replicates its style, demonstrated by cloning a Fireship-style edit used on this very video.
07:28 – 08:01
14 · Final thoughts
Kev closes saying this is the most fun he's had making videos, ties it to his best month of social growth yet, and signs off.
Atomic Insights
Lines worth screenshotting.
A Claude Code agent can take a raw voiceover and produce a fully edited video in one shot, no timeline required.
Every video dropped into the pipeline splits into two parallel pathways: Whisper for word-level audio timestamps and OpenCV for face detection and framing.
Word-level timestamps from Whisper are what let every caption, beat, and on-screen counter sync exactly to the spoken word.
A single file called marks.py lets the agent animate any imported character or logo, including splitting it into variations and recoloring it on command.
Brand assets like colors, typography, and logos can be pre-loaded into the agent once, then reused automatically across every future edit.
The Giphy API gives the agent a free, searchable library of reaction GIFs it can pull in automatically based on what's being said.
The /reverse-engineer skill watches any video a creator points it at and builds a brand-new, reusable editing skill that mimics that style.
Kev used /reverse-engineer on a Fireship-style video to build a new skill, then used that exact skill to edit the video explaining it.
After editing finishes, a dedicated QA agent watches the rendered output specifically to catch mistakes before the file is considered done.
The repo ships with a built-in short-form and long-form talking-head editor, so the first test run requires no custom skill-building at all.
Because the agent remembers how a video was edited, subsequent edits in the same style get faster and more consistent without re-explaining anything.
Kev is open-sourcing this specifically because he thinks the old editor timeline, and the industry built around it, won't survive this workflow.
Takeaway
An agent can now edit your video the way it writes your code.
WHAT TO LEARN
Dropping a raw voiceover into an agent that splits transcription and vision into two pathways, then choreographs captions and graphics against word-level timestamps, replaces most of what a manual timeline used to require.
02Why I'm retiring HyperEdit
Retiring your own older tools in public, the way Kev retired his prior editor HyperEdit, is a credibility signal worth more than pretending it still works.
A tool built around a manual timeline becomes obsolete the moment a newer tool removes the timeline requirement entirely.
03The new open source repo
An agentic workflow is still something you install and run yourself; it has no hosted app, so it only pays off if you're willing to set up a local environment.
Open-sourcing a working personal tool is a faster way to prove a new workflow than trying to explain it in the abstract.
04The tools (quick version)
How a system works end-to-end is usually more useful to a reader than a tool-by-tool feature list.
Skipping a granular breakdown of every backend component keeps the explanation focused on the decisions that actually matter.
05marks.py: bring any character or logo to life
Treating editing style as a reusable, named skill instead of a one-off habit is what makes an edit repeatable across every future video in that style.
Pre-loading brand colors, typography, and logos into an agent once removes the need to manually re-apply brand identity on every single edit.
06Giphy API: every meme on earth, free
A free API like Giphy can give an agent access to an entire category of content (reaction GIFs) without adding any cost to the pipeline.
Letting an agent select contextual reaction content automatically removes a manual search-and-insert step from the edit.
07How it actually works
Letting an agent iterate through full back-and-forth conversation, rather than demanding a perfect first output, is what makes a complex creative task tractable.
Stating the end goal in one sentence (delete the timeline) keeps a complex multi-step system easy to explain to someone new.
08Pathway 1: Whisper word-level timestamps
Precise word-level timestamps are the hidden dependency behind anything that claims to sync captions, beats, or counters exactly to speech.
Running transcription locally and for free removes a recurring cost that would otherwise scale with every video produced.
09Pathway 2: vision, face detection + the band
Splitting a problem into independent pathways, like audio timing and visual framing, lets an agent solve each half with the right specialized tool before recombining them.
Automated face detection and cropping removes a manual framing step that would otherwise need to be redone for every new piece of footage.
10Choreography: animations, graphics, effects
Combining transcript timing with vision data is what lets an agent place graphics and effects at the exact right visual moment, not just the right time.
A choreography step built from your own transcript and footage produces a pipeline that's harder for anyone else to directly copy.
11render.py + the QA agent
Separating the act of creating from the act of checking, by running a dedicated QA pass after generation, catches errors a single pass would miss.
A single render step that merges every generated asset keeps the final output consistent even when many separate tools contributed to it.
12The 6 skills inside the repo
Shipping a small menu of working example skills lowers the barrier for someone to test a new system before they build anything custom themselves.
Naming each skill after its exact output (short-form, long-form, reverse-engineer) makes the menu self-explanatory without extra documentation.
13/reverse-engineer: learn any editing style
A tool that can study an existing style and generate a new, reusable skill from it turns imitation into infrastructure instead of a one-time copy.
Reverse-engineering a specific creator's visual grammar, as Kev did with Fireship's style, is more valuable when it's captured as a repeatable skill than as a one-off homage.
Explaining a tool by using the tool to produce the very video that explains it is a stronger piece of proof than any amount of description.
Glossary
Terms worth knowing.
Agentic video editing
A workflow where an AI agent makes the actual cutting, captioning, and animation decisions for a video, instead of a human dragging clips on a timeline.
marks.py
A file in the harness that lets the agent bring any imported character, mascot, or logo to life as an animated on-screen element it can reposition, recolor, or duplicate.
OpenCV face-crop band
The vision pathway of the pipeline: it detects a speaker's face in the raw footage, crops to it, and builds the framed band the final edit is composited around.
Word-level timestamps
Timing data from a transcription model that marks the exact start and end of every spoken word, used to sync captions and on-screen graphics to speech precisely.
/reverse-engineer
A skill in the harness that studies an existing video's editing style and generates a brand-new, reusable skill that replicates that style on future footage.
QA agent
A second agent that reviews the fully rendered video after editing to catch errors before the output is treated as finished.
“The traditional editor timeline as we know it is dead.”
blunt, declarative claim that frames the whole video→ TikTok hook↗ Tweet quote
01:04
“A $5 billion industry just got flipped on its head and the leaders of the industry are not positioned to continue leading the race.”
specific dollar figure plus a competitive-threat claim→ IG reel cold open↗ Tweet quote
01:20
“Even HyperEdit is now completely useless because that also uses a video editor timeline.”
self-deprecating admission that raises credibility→ newsletter pull-quote↗ Tweet quote
06:08
“The agent uses something called render.py to merge all the files, and then we have another QA agent spin up to watch the rendered video.”
concrete technical payoff after a long build-up→ TikTok hook↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphorstory
With the recent leap in what AI can do with video editing, the entire content industry is going to be flipped on its head. The traditional editor timeline as we know it is dead. Cloud doesn't just edit videos alongside us now.
it completes the entire video edit in one shot. And what that means is the way that we use AI to vibe code is the exact same way that we are going to use AI to vibe edits and vibe create videos. And after realizing this, that made me really excited and also really sad at the same time.
And the sad news is we got to say bye to an entire era of how the entire world edited videos by using stuff like CapCut and Adobe Premiere Pro to literally sit in front of a screen and click through a timeline. And that... is officially over because not only is what we're seeing through these cloud editing workflows next level, they're better in many ways than what was originally possible.
We're getting faster turnaround times. We don't have to pay editors to create edits. And we are talking about using agents in a way that evolves over time.
And so these skills get even better through our creativity. Things are freaking wild. A $5 billion industry just got flipped on its head and the leaders of the industry are not positioned to continue leading the race.
It's ridiculous. If you guys have been following this channel, I already built and open sourced a video editor called HyperEdit, but even HyperEdit is now completely useless because that also uses a video editor timeline. So I was like, dang, what is the future of video editing and video creation?
And then it hit me. We got to build an agentic harness for these AI models so that everyone can use a free AI editor in the way that I'm using it. And it's easy to set up.
And so that's exactly what I did. So today is to fully break down my new open source repo called. agent edits personally after using this for the past week it's been ridiculous i'm gonna go through how it works all the tools that are associated it's completely free all the tools that are being used but it's good to know and the five best skills that are being shipped inside of this repo that you can use right now just like the one that you're seeing edit this video and also guys if you're new here please like and subscribe but let's get right into the video so first off guys we're going to talk about the tools but i'm not going to bore you breaking down every single tool because you guys can obviously read around what they do and i think breaking down how it works is a little bit more important than all the tools and all the backend stuff you can ask however the most important tool in this entire stack that i can explain to you guys that you need to know is this thing called the marks pi file
What the marks pie file allows you to do is bring in any character into your video editor pipeline and it can animate this type of character. And so it can be like an icon, but it could also be like a mascot. And so in this example, I'm going to pull up the Claude invader logo and have him like walk around on my hand and then maybe have him split into like a hundred different variations and then maybe change colors and then maybe have him wave to the audience and tell you to like this video.
And essentially what this is, is a file that the agent has gotten onto the computer and then turned it into this animated type of code that it can play around with. And then it saved it as a mark Python file inside of its skill templates.
And that one's really important because what this allows you to do is literally like pre -train your agent with any type of personal media and then bring that personal media to life. So if you have logos or colors or typographies or things that are similar to your brand that you want to show up in your videos, This is a very simple hack to do that.
Now, aside from everything that's built into the project that is free, I wanna talk about some additional tools that I've added to my editing system that makes life a lot more fun. So the first one is the Giphy API. And the Giphy API has all of the different memes and GIFs in the entire world.
And the beautiful thing about the Giphy API is that it's completely free as well. So if I'm talking about like the NBA and I'm talking about the Los Angeles Lakers, it'll then go pull in Luka Doncic from the Giphy API. Or if I make a joke and then it wants to like...
have like a reaction to that joke then we have like a gif being like oh my god and it'll pull that from the giphy api so if you stayed around for that that was the most boring part of this entire tutorial because it's ridiculous i want to talk about now guys how it actually works so the end goal of this agent here is to again delete the timeline and so you want to build skills that you literally just drop a video in and it gives you the exact output that you're looking for but the other beautiful thing about the way that this thing is set up that i'm going to show you guys is that you're going to be able to have a full conversation with your agent.
So if this seems confusing at any point in time, just remember that you can literally just have a full conversation with your agent back and forth until the skill that you're trying to build out is built. And so it's really... Quite simple.
And the video can be edited again and again and again. And the agent will remember how it's being edited. And that's because of the tools.
If you read about them, you'll understand how that works. But when a video then gets submitted, it gets split into two pathways. The first pathway is the transcription model, which uses whisper, which is a free local open source audio understanding model that literally gives you like word level timestamps.
And the beautiful thing about the transcription step, and it does more than you think. is that every caption, every beat and every counter that you see on the screen is synced up perfectly because of this step. Then the second step that the agent does is the vision model sets the stage.
And so we need something to find my face, crop my face inside of the video edit. and then build a band of room around me so that real edits can actually take place. From there, the agent can gather all of the relevant information.
So it's going to be able to get the transcript with the keyword timestamps, and it's going to understand the entire flow of your video through OpenCV. It can then start choreographing what happens next. And so it'll use its tools to design entire animations, graphics, video effects, icons, logos, shapes.
And it can all be in accordance to the skill that you're building. And so in this way, you're also creating a customized video editing pipeline that no one else can really replicate. And that's also what's going to set you apart over time.
After all of that editing happens, the agent uses something called render .py to merge all the files. And then we have another QA agent spin up, which is a quality assurance to watch the rendered video, make sure everything's okay. So now guys, I want to talk about the six skills that are going to be inside of this code base the moment that you get it.
And you can use these skills to test whether everything's set up and you have all your dependencies installed. So the first two skills you guys have seen in other tutorials and they're very straightforward. The short form talking head and long form talking head are exactly as they sound.
So you send in a short form video that is unedited and it'll completely create the edited version. You can also send in notes and some additional things that you want added to the edit. So it's not just the video that you send it.
The second one is the long form talking head, exact same thing as skill number one. The most important skill I want to talk about though is called reverse engineer. And essentially this is how you can teach your agent any other video editing workflow that you see on the internet.
And you can even find like a video on YouTube. And that's what this entire video is based on to show you guys that it actually working. I grabbed a YouTuber called fire ships video style, and I've first engineered it to work for my YouTube videos in the exact same way.
And so the reverse engineer skill created a new editing skill that I can now. use and improve upon and so for me guys the reason why this video is so weird for me to film is because i'm literally just sitting here and yapping but i'm also visualizing about like what ai is about to do with the edit because i'm literally i'm literally just sitting here guys so yeah Hope you guys enjoy.
But jokes aside though, guys, I'm really excited to get this in your hands because this has been the most fun I've ever had making videos. And the month of September, it was the most I've ever grown on social. So both of those together is obviously some correlation.
And I've been using all of my video edits on short for my long form using this exact same tool. Yeah, I hope you enjoy it and I'll see you next time.
The Hook
The bait, then the rug-pull.
Kev opens by declaring the editing timeline dead: the same AI coding agents creators already use to vibe code can now vibe edit, taking a raw voiceover straight to a finished video with no CapCut, no Premiere, and no manual cutting.
Pathway 2: OpenCV vision does face detection, crop, and framing band
Every video dropped into the harness splits into an audio-timing pathway and a visual-framing pathway, which the agent recombines to choreograph the final edit.
Steal forany agent-driven media pipeline that needs both precise speech timing and visual framing before it can make editing decisions
The repo ships with six pre-built editing skills so a new user can test the full pipeline before ever building a custom skill.
Steal foranyone building a creator-facing agent toolkit who wants a starter menu instead of forcing users to build skills from scratch
CTA Breakdown
How they asked for the click.
VERBAL ASK
01:44subscribe
“if you guys are new here, please like and subscribe”
A quick aside dropped mid-explanation rather than a dedicated pitch segment. The real monetization (Creator OS, 1:1 calls, Creator University, No Code Academy) lives entirely in the description links and is never mentioned verbally in the video.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Kev wires the tool-dispatch model Jev into his open-source editor HyperEdit, so a chat prompt triggers real cuts, captions, and media pulled straight from an Obsidian vault.
Paul J Lipsky hands his entire YouTube edit, silences to motion graphics, to Claude Opus 5.5 running through an MCP-connected editor, and walks through exactly what still needs a human pass.
A creator builds a reusable Codex + Remotion editing template from a reference video's style, then layers AI-generated special effects on top with scripted editor's notes.
A developer wires Claude Code up to a free CLI tool called Buttercut and has it read raw vlog footage, then assemble a full rough cut in Final Cut Pro without a human touching a timeline first.
A creator hands her raw footage to Claude Code and walks through exactly what it costs, and how to prompt it, to get a fully edited long-form video back.