Modern Creator
Bart Slodyczka · YouTube

How to Separate Any Video's Subject From Its Background With a Free Claude Code Skill

A free Claude Code skill that uses Meta's SAM2 to split a video's subject from its background, then stacks dimming, text, and motion-tracked labels between the two layers.

Posted
3 days ago
Duration
Format
Tutorial
educational
Views
23.4K
274 likes
Big Idea

The argument in one line.

A free Claude Code skill pairs Meta's local SAM2 model with plain-English prompts to separate a video's subject from its background, unlocking layered effects no flat video timeline can do on its own.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You edit your own talking-head or product videos and want background separation without learning Premiere or After Effects masking.
  • You already use Claude Code for other tasks and want to extend it into your video editing pipeline.
  • You want a concrete walkthrough of running a local AI model against your own footage inside a coding agent, not just a feature announcement.
SKIP IF…
  • You already have a rotoscoping or masking workflow in DaVinci Resolve or After Effects that works for you.
  • You don't have a GPU or Apple Silicon machine and don't want slower CPU-only processing.
TL;DR

The full version, fast.

The creator gives away a Claude Code skill that pairs with Meta's local SAM2 model to separate a video's subject from its background automatically, no manual rotoscoping required. The skill has to run inside Claude Code on your own machine, not a cloud sandbox, because it needs direct GPU access to download and run SAM2 locally; it auto-detects your hardware and picks the right model size, then tracks an object across every frame once you point at it. Demonstrated on a Mac mini unboxing clip, the separated subject layer lets Claude add background dimming, text placed behind the subject, and point-and-track motion graphics labeling a port, all from plain-English prompts. A companion technique splits footage into an HTML page of scenes with per-scene comment boxes, so edit notes across many clips get submitted and executed in one batch instead of one prompt at a time.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:00 – 01:21

01 · Intro: separating the subject from the background

Previews the before/after of the separation skill, what it unlocks (background dimming, text behind the subject, motion graphics between layers), and promises the skill free on GitHub.

01:21 – 02:41

02 · How the object separation skill works

The three preconditions for using the skill: run it in Claude Code (not Claude Cowork) for GPU access, let it auto-check your hardware and download the right SAM2 size, then point at an object once to track it across every frame.

02:41 – 06:34

03 · Using the skill in Claude Code

Sets up a dedicated project folder, picks Opus 5.5, uploads the raw Mac mini unboxing footage plus the skill, describes the clip's two visual parts, tells Claude the SAM2 model is already downloaded, and lists the desired edits (dim background, add text, motion-track the HDMI port).

06:34 – 07:24

04 · The result: Mac mini separated from the background

Reviews the finished separation: clean outline around the fingers, a smooth fusion as the box hands off to the Mac mini, processed in about 10 minutes on an M5 Max with 48GB RAM.

07:24 – 12:55

05 · My editing workflow: the scene-by-scene HTML review page

Splits the footage into numbered scenes inside an HTML review page, each with a comment box for edit notes, then submits all scene notes at once so Claude can execute every edit in a single batch.

12:55 – 14:43

06 · The final edits: text behind, motion tracking, step labels

Reviews the three finished edits: white text placed behind the box with a dimmed background, a point-and-track dot and line labeling the USB-C and HDMI ports as the Mac mini rotates, and step 1/step 2 text swapping as the power and HDMI cables go in.

14:43 – 15:11

07 · Wrap-up

Notes that plain text overlays don't always need the separation step, then signs off pointing to the GitHub skill and his community in the first comment.

Atomic Insights

Lines worth screenshotting.

  • A free Claude Code skill pairs with Meta's SAM2 model to separate a video's subject from its background automatically, no manual rotoscoping required.
  • The skill must run inside Claude Code on your own machine, not a cloud sandbox like Claude Cowork, because it needs direct GPU access to download and run SAM2 locally.
  • The skill checks your CPU, GPU, and memory first, then auto-downloads the size of SAM2 that actually fits your hardware.
  • Point at the subject once and SAM2 tracks it across every remaining frame of the video automatically, no manual keyframing.
  • Once a video is split into a subject layer and a background layer, you can darken the background without darkening the subject, add text behind the subject, or insert motion graphics between the two layers.
  • A 27-second Mac mini unboxing clip split into 799 frames took about 10 minutes to process on an M5 Max with 48GB of RAM.
  • Both 16:9 landscape and 9:16 vertical selfie-mode footage work with the separation skill.
  • For batch editing across many scenes, split the footage into an HTML page of numbered scenes, each with its own comment box, then submit all the notes at once so Claude executes every change in one pass.
  • A point-and-track label, a dot and line that follows an object as it rotates on camera, works well for highlighting a specific feature during a product demo or talking-head segment.
  • If the final effect is just text placed over footage with no need to pass behind the subject, the object-separation step may not be necessary at all.
Takeaway

Separate subject from background, then stack effects between the layers

VIDEO EDITING

A local AI model that tracks an object across every frame turns one flat video into two editable layers, unlocking background dimming, text-behind-subject, and point-and-track motion graphics from plain-English prompts.

02How the object separation skill works
  • Run GPU-dependent Claude Code skills in Claude Code itself, not in a cloud sandbox like Cowork, since the sandbox can't reach your machine's graphics chip.
  • Let the skill probe your CPU, GPU, and RAM before it downloads a model, so it pulls the SAM2 size that actually fits your hardware instead of the biggest one by default.
  • A single point-and-click on the subject is enough for SAM2 to track it through every remaining frame, no manual keyframing required.
03Using the skill in Claude Code
  • Keep a dedicated project folder for ongoing video work rather than starting a fresh, unattached session each time, so Claude retains your personalization and process.
  • State the structure of your footage up front, for example that a clip has two visually distinct parts, so Claude knows where object tracking needs to switch targets.
  • If a model is already downloaded locally, say so explicitly and ask Claude to search your device for the existing copy before it downloads a duplicate.
  • Both 16:9 landscape and 9:16 vertical selfie-style footage work for this technique, so a phone recording is a valid source.
04The result: Mac mini separated from the background
  • Expect roughly 5 to 10 minutes of local processing for a short clip split into hundreds of frames, scaling with your hardware.
  • Review the separated output at the exact cut points between objects, like a box handing off to a product, since that's where masking errors are most likely to show.
05My editing workflow: the scene-by-scene HTML review page
  • For a long edit, split the footage into numbered scenes rendered in an HTML page so you can review and comment on each one visually instead of describing timestamps in chat.
  • Attach a comment input box to each scene so you can leave specific edit notes per scene without interrupting your own review flow.
  • Submit all scene notes in one batch and let Claude execute every edit and any needed re-render in a single pass, instead of prompting one change at a time.
06The final edits: text behind, motion tracking, step labels
  • Plain-English instructions, like placing white text behind the subject and dimming only the background, are enough to direct the layered edit once separation is done.
  • A point-and-track label, a dot plus a line plus text that follows a feature like a port as the product rotates, is a strong way to call out a specific detail in a product demo.
  • Skip the object-separation step entirely when an effect is just text or graphics overlaid on footage with no need to pass behind the subject; save it for effects that actually require the extra layer.
Glossary

Terms worth knowing.

SAM2 (Segment Anything 2)
Meta's free, locally-run AI model that finds an object's outline in one frame and follows it through every subsequent frame of a video.
Claude Code
Anthropic's coding agent that runs directly on your own machine, as opposed to a cloud sandbox, giving it access to local hardware like a GPU.
Claude Cowork
A cloud/virtual-machine mode of Claude that cannot access your computer's graphics chip, which is why GPU-dependent skills have to run in Claude Code instead.
Object-separation skill
A custom Claude Code skill, free on the creator's GitHub, that uses SAM2 to split a video into a subject layer and a background layer so they can be edited independently.
Opus 5.5
The Claude model used throughout the video for both running the separation skill and executing the batch scene edits.
Resources

Things they pointed at.

02:26toolSAM2 (Segment Anything 2) by Meta
03:31toolOpus 5.5 (Claude model)
Quotables

Lines you could clip.

00:00
“I just taught Claude a new skill that lets it separate the subject of the video from the background of the video.”
clean cold-open hook stating the whole premise in one line→ TikTok hook↗ Tweet quote
00:29
“That means I can do things like darken the background without darkening my face, I can add text behind me, or I can even add motion graphics in between those two layers.”
concrete payoff list with no setup needed→ IG reel cold open↗ Tweet quote
14:20
“In this case, I realized we didn't need to do the object separation skill just because the text is not overlaid on a Mac mini... you might just need to combine a couple of skills and techniques together to get your desired output.”
honest meta-lesson about when NOT to use the technique→ newsletter pull-quote↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

I just taught Claude a new skill that lets it separate the subject of the video from the background of the video. Let me show you how it works. So here's my video editing folder and here's the object separation skill.
All I need to do is give Claude the skill and my raw footage and then ask it to run the separation. Now look at this. This is what happens.
You get the subject of the video, which in this case, it's me, and it's separated from the background of the video so that now instead of just one flat layer, we actually have two layers to work with. So that means I can do things like darken the background without darkening my face, I can add text behind me, or I can even add motion graphics in between those two layers.
And the separation skill is what makes all that possible. So in this video, I'm going to give you the exact skill for free, it's going to be available on my GitHub. And I'm going to take you through the step by step process so that by the end of this video, you can make these edits for yourself.
And by the way, if you like all the edits for this video so far, including the Windows XP, the visual storytelling, all the motion graphics and transitions, I actually had Claude edit this entire video using my AI editor workflow. So if you want to turn Claude into your 24 -7 AI video editor as well, I've just created a school community where I'm going to be sharing all my skills, techniques, my entire blueprint, including behind the scenes builds that I do for my own YouTube channel.
so that you can also turn Claude into your video editor too. I'm going to have that link in the first comment below. Okay, so let's take a look at that object separation skill.
So let's first start by jumping across to GitHub and then go into the skills folder and open the object separation skill. And the whole skill is just one file. And if you want to use it, you just need to copy it.
But before we try it, there are three key steps in using this skill. So the first step is that we need to run it inside Cloud Code and not inside Cloud Cowork. Cloud Cowork runs inside a separate virtual machine.
So it can't access your computer's graphics chip. And this skill needs your real working machine because it downloads. a small AI model, and it runs it locally.
So cloud code works directly on your computer. So yeah, that's why we need to use cloud code. Now, step two is where the skill checks your hardware.
So it looks at your CPU, at your graphics chip and your memory, and then it picks the right version of that small local model and then it downloads it. And that model is called SAM, which is Segment Anything from Meta. It's free.
And after you download it, you just point it at an object once and it finds the outline and then it follows it through every frame of your video. It comes in different sizes and the bigger the model, the more accurate. the outline but honestly I found that even the small models do a really good job.
If you've got an Apple M chip or an Nvidia graphics card it's going to run super fast but even if you've got an older laptop without a graphics card it's still going to work so just it'll just take a little bit longer. All right, so now that we understand how the skill works and those different components, we are ready to get across to our Cloud Code environment.
So I've got my Cloud Desktop app open. I am inside Cloud Code. I've got a blank project folder just so when I use the skill, you're going to have the exact same experience as me.
I'm not using my actual editing folder because it's got so many things inside it. We get a much different output. Now, whenever you are working on motion design stuff and video editing, if you are going to be doing this a bit more seriously, definitely create yourself a project folder.
just start a new session and leave it not attached to a folder because you are going to lose like that personalization that you do so if you use this skill and you get a video out that you really like you can just ask to remodify the skill that you just imported to adapt to your to your new process so you definitely want to be inside a folder and we're going to be using opus 5 .5 on medium as well and this is literally Opus 5 .5 is the best model that I've been using to do this stuff.
Even above GPT models, I've tried doing similar motion graphics inside the ChatGPT desktop app and right now, no model compares to Opus 5 .5. Let me just start by explaining what I want to do in this first chat with Claude. Hey, I'm going to be giving you a raw video file of me doing an unboxing for a Mac mini.
And I'm also going to attach a skill for you. That skill is going to get you to separate the subject from the background. And you'll see the video that I'm uploading is actually got two main parts.
The first is me unboxing from the main packaging. And then next I'm holding the Mac mini as well. So when you're going through this process, just note that you might have to switch between the box and then the Mac mini.
And now that I've set the frame for what we're doing in this task, I want to mention one more thing to Claude. is that i already have the model downloaded onto my computer so i don't want to re -download another model so i already have the model from the skill downloaded onto my device please scan my device it might be saved to a specific folder so just try and find that model and don't re -download a new one and let's just say skill is below here paste in that skill And now for the video that I'm going to be uploading, this is just taken.
I basically took a video of me unboxing a Mac Mini from like a couple of days ago. And I edited the video down to be this specific 27 second length. And I'm just going to play it to you for you so you can see.
I've just chopped it up to make it a little bit more engaging. And basically, as you see, I open up the box. Once I open the box, I'm now unpackaging the Mac Mini.
And there's a couple of things that I want to do with this clip. So after we actually separate the object, which is the Mac Mini from the background. I not only want to like dim the background and then I want to add some text just to show you how easy it is, but I want to do one more feature where, as you can see that little slot here that says HDMI, when I'm spinning the Mac mini around, I'm just going to use plain English to prompt up Claude.
to do like a motion tracking and have like a little dot that says, you know, HDMI input. It's going to look a little bit silly because it'll be a little bit out of place. But the purpose is I want to show you once you separate the object from the background, there's a bunch of really cool things that you can do.
So I just want to give you a couple of examples in this video. And I'm just going to plunk that video into here as well. And now we're ready to go.
Awesome. So we can see that Claude has found the model on my device already. I've got SAM2 and called split up the video into 799 distinct frames.
So now when you run that local model against each of the frames, it's actually able to track the object across each frame, you don't have to go in, you don't have to click on something and say, this is the thing that I want to track, it's able to do that automatically. Now this process, depending on how big the model is, how powerful your machine is, might take about five or 10 minutes.
Now just in case you don't know what kind of video to upload for this test, you can either go for landscape mode, so 16 by nine format, which is is what I've uploaded here. But you could also just get your phone and do nine by 16 formats.
So just pop your phone, go into like selfie mode and take a step back a couple of meters just so you can see more of your body or at least your upper body against a nice background. And you can basically just upload from your phone as well. So both video formats will work for this process.
Okay, so it looks like Claude has finished running the process. Let's take a look at how it came out. Okay, so yeah, pretty good.
We have the outline against the fingers and then we switch from the box to the Mac Mini. Really good. As soon as we place the fingers, this is really nice as well because when we place the fingers, it knows that it's not part of the actual Mac Mini object.
And I just want to look at this part again. So as we unbox, that is a pretty smooth transition. Like the box is there.
it kind of like fuses from the Mac mini really nicely. And then around here, the box disappears. So this came out really, really good.
And again, it was like 10 minutes and I have an M5 Max with 48 gigs of RAM. So I do have like a nice computer. Maybe it'll take a little bit longer if you have a bit of an older computer, but hey, the process will still work.
And it looks like there was a little bit of like zappiness happening there, but I don't think it's actually gonna impact our end result. So yeah, so far, this is looking really good. Now there's another technique that I do when I'm actually editing.
my intros or like extended extended b -roll so i'm going to show you here as well i'm just going to prompt it up may as well just give you some more of my tricks to my trade but basically i just open up a html page and i split this into each of the separate scenes so like here this is like scene one and scene two is when i have the box flipped so scene one is me ripping it open scene two is me spinning this box scene three might be this part scene four now is the mac mini and scene five is unpacking the mac mini Then scene six should be basically me spinning the Mac mini around.
So you can kind of like tell Claude what scenes you want to split it at. I might just describe it a little bit here, but the idea is we have that like HTML page just because it's easy for us to view in a preview and we go through each of the scenes and then we have like a comment input box and we can drop a comment for that specific scene, hit enter, and then have Claude kind of go through and make all the edits at once and then re -render if it needs to, or basically do whatever it needs to all in one big batch instead of incrementally.
This is awesome. you now take that Mac mini video and then split it into little scenes like at the very start when I'm unpacking the cardboard box that'll be scene one then when I'm taking the Mac mini out of the packaging box scene two scene three would be when I'm unwrapping the the excess like plastics scene four would be when I'm spinning the Mac mini around and basically build out those like five six seven scenes whatever it is plunk them into a HTML page Next to each scene, put a comment input box that I can submit some notes about edits that I want to make.
And then once I've added all my edits, I'll just come back into the chat and let you know to go off and make those changes for me. So if you do go across to my community, since I'm launching it today, I want to be filling out my entire process across like this next week. So instead of just prompting up freehand like what we're doing here, I'll just give you all of my skills.
So that you can run my exact process, which is probably going to be a little bit more optimized than just prompting like this. But hey, but really, as long as the process that you're doing works for you, that's all that matters. And awesome looks like we have that side panel available.
So let's just go through. So we have the first scene, which is just one second here, we can add some notes, our second scene, add some notes. So yeah, I think for this mini unboxing example.
The scenes are really, really short. But typically when I'm doing my videos, you might have a scene that's like five seconds to 10 seconds long. So it's a little bit easier to work with some longer, longer scenes.
So I think some of these might just be a little bit too difficult for us to work with. And we have the separated green. I guess this is going to show us the exact same video, just separated.
So I might just choose a couple of these for this video. and just add some of these effects. So this first one looks interesting.
Maybe we can find a longer one. Let me just mute this. I've got voiceover going over each clip.
Okay, so let's first start with this one. Hey, can you take this, and I want to put a big white text behind the box that just says unboxing. So I'm going to submit this as my first comment, and then we'll be able to do this first technique, which is basically putting some white text behind the box.
And actually, while you're there, can you also just dim the back? so that the white text is on that dimmed background and it just pops a little bit more and has nicer contrast. I just wanted to update this a little bit more because yeah, that white text will pop a little bit nicer.
Now for our next scene. Okay, this might be interesting for us to try something else. We are unpacking it.
Okay, so we actually probably could have done a bit better job over here because as we were unwrapping this, I realized that it was mapping the actual Mac mini. But since we're peeling off the side, I probably messed up a bit of that section there. So we could definitely go back and rework that.
But I'm going to leave that frame for now. I just want to try a couple of other things. Okay.
awesome so this is the next one that i want to work on these are like i think usbc slots like thunderbolt 4 and if we just play all the way through as we spin this we're going to get to yeah okay they're back here so i just want to do this thing which is like object tracking and i want to have like a little dot in the center of this html input and then have like a line coming out and then like the little word kind of saying HDMI just over here.
Now, it's very obvious that this is a HDMI input because the word is right there. But this motion tracking technique is super interesting, especially if you're like doing a talking head and you have something in your hand and you quickly want to point or you want to trace that object and kind of move it around and have some text pointing to it.
It just makes for more engaging videos. Awesome. So for this part, I want to do two things.
When I show the front of the mini, I want to have some motion tracking, which is kind of pinning a point to one of the front USB -C slots. And just, you know, either saying, you know, two times USB -C is like two slots in the front, whatever you decide. And then for the back where the HDMI slot is, I also want to have like a little white.
pin and then a line coming out and then like the word hdmi coming out there as well so basically imagine it's like a tech review and then you have like text just showing all the different inputs and outputs that are there So let that feedback just be saved. And let's see, should we do one more thing here?
I guess we could do like a bit of an explainer and just say connect the power as step one and then connect the HDMI as step two. And for this clip, when I plug in the power, can you just have like step one dot dot plug in power. And then as the HDMI comes in, just replace the previous text and say step two dot dot plug in the HDMI cable.
So that's just a very basic one, but we may as well just do one more edit. Okay, so the notes are ready. I'm going to submit those.
All of my notes are ready for your edit. Thank you very much. And let's send that off.
And Claude has finished with those three edits that we requested. Let me just close this side window here. So the first one had the unboxing text behind and it was basically a darker shade across the entire background, but not including the box itself.
So that actually looks pretty cool. I mean, we can definitely fix the text and probably, you know, lift it up a little bit, make some edits. But if you have that HTML section on our side, like I showed you guys before, you can just come back into Claude and just reprompt to like push the text up or make it bigger or whatever you want to do.
But this looks really nice. It actually looks really neat there. Next, we had the emotion tracking.
So being able to see, okay, two times USB -C slots. And as we rotate to the back and you have it fully tracked at the HDMI slot as well. I just want you to like pay attention to.
As we're moving the Mac Mini around, I just want to put some arrows down. So we first start tracking over here. As we move the Mac Mini a little bit more, we then have the next position over here.
And then as we finish, we have the final position over here. So you can really see that we are tracking the precise location of where that HDMI input is. And then we have our HDMI text as well.
So again, very silly for this example here, but you can definitely see this use case. And then the final thing we want to take a look at is... We just had the text.
Okay, here we are in the bottom, right? We have the text for plugging in the power and then plugging in the HDMI So very nice very neat way to do like explainer videos showing the steps In this case, I realized we didn't need to do the object separation skill just because the text is not overlaid on a Mac mini.
But if we did want to like isolate the cable and the Mac mini, we could have just run the object separation skill onto those two different components and then darken the background and then had the text pop a little bit better as well. So you might just need to combine a couple of skills and techniques together to get your desired output.
And that's it. That's the entire process. I hope you guys enjoyed the other nuggets that I've given you from my own AI.
video editing process the side panel for editing the different scenes as well as just trying to open up your imagination by doing the motion tracking effect as well like i said the object separation skill is going to be in my github i'm going to link that in the first comment below and if you do want to join my school community and get my entire video editing blueprint i'm going to drop that link in the first comment below as well alright guys thank you very much and i'll see you in the next one
The Hook

The bait, then the rug-pull.

A YouTuber who edits his own channel with Claude Code gives away the exact tool he built for it: a free skill that finds a video's subject in one frame with Meta's SAM2 model, then tracks it through every remaining frame, splitting the footage into a subject layer and a background layer that can be edited independently.

Frameworks

Named ideas worth stealing.

01:21list

3 rules for running the object-separation skill

  1. Run it inside Claude Code, not a cloud sandbox like Cowork, since the skill needs direct GPU access
  2. Let the skill check your hardware first; it auto-picks and downloads the right size of SAM2 for your machine
  3. Point at the object once; SAM2 tracks it through every frame on its own

The three preconditions the creator lays out before running the skill: local GPU access, hardware-matched model size, and one-click object tracking.

Steal forany Claude Code skill tutorial that depends on local GPU access
CTA Breakdown

How they asked for the click.

VERBAL ASK
01:21product
“I've just created a school community where I'm going to be sharing all my skills, techniques, my entire blueprint, including behind the scenes builds that I do for my own YouTube channel.”

Soft-sell mid-intro pointing viewers to 'the first comment below' rather than a spoken URL, repeated near the end alongside the free GitHub skill link.

Storyboard

Visual structure at a glance.

open
hookopen00:00
skill explainer
promiseskill explainer01:21
running separation
valuerunning separation05:45
scene review workflow
valuescene review workflow09:56
final tracked result
valuefinal tracked result13:28
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.