Modern Creator
Brad Bonanno | AI Automation · YouTube

How I Watch ANY Video with AI in Seconds

A free agent skill sends any video straight to Gemini's API, so your AI gets back timestamped answers instead of a wall of screenshots.

Posted
4 days ago
Duration
Format
Demo
educational
Views
32.9K
954 likes
Big Idea

The argument in one line.

A free agent skill routes any video link straight to Gemini's API, letting an AI answer specific questions about hours of footage without downloading the file or burning the main model's own token budget.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You use Claude Code, Claude Desktop, or ChatGPT for research and want it to answer questions about a specific video instead of you scrubbing through it yourself.
  • You make videos and want to cross-reference your own retention data against what you were actually saying and showing at the exact moment people left.
  • You keep a personal notes system (a wiki, Obsidian, or similar) and want video explanations filed into it automatically instead of re-watched from memory.
SKIP IF…
  • You don't use an AI coding agent or chat tool with plugin or skill support, so there's nowhere to install this.
  • You're looking for a video editing tool. Nothing here cuts, renders, or exports a clip.
TL;DR

The full version, fast.

AI models can't natively watch video, so Brad Bonanno built a free agent skill that routes a video link to Google's Gemini API instead of feeding the model hundreds of screenshots. Gemini watches the footage and hands back timestamped findings, which Google's own benchmark shows costs 87% fewer tokens than raw frame-dumping while answering more questions correctly. He uses it to find one buried detail in hours of footage, to line up retention-drop timestamps with the actual content, and to pull explanations into a personal notes wiki. Setup is a one-time plugin install plus a free Gemini API key, and a local WhisperX engine still exists for when the model needs to see literal pixels instead of a description.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:00 – 00:30

01 · Cold open: the claim

Teases that AI can now watch video for free without burning usage limits, backed by a Google benchmark showing the agentic approach beats raw frame-dumping while using 87% fewer tokens.

00:30 – 01:38

02 · How agentic video understanding works

Explains that even top models can't parse video natively, so the free watch skill sends the link and the question to Gemini, which scans the footage and returns timestamped findings instead of hundreds of screenshots.

01:38 – 02:31

03 · Use case 1: finding one buried detail

Demonstrates asking Gemini to locate a single demonstration inside a three-and-a-half-hour Andrej Karpathy video (a tokenizer tool) by description alone, then jumping straight to the timestamp.

02:31 – 03:53

04 · Use case 2: diagnosing retention drops

Feeds Claude his own video's audience-retention numbers and has watch analyze the footage around each drop to line up what was on screen and being said at the exact moment viewers left.

03:53 – 04:30

05 · Use case 3: auto-saving notes to an LLM wiki

Has watch pull explanations and visual examples out of a video and hand them to Claude to file into a personal wiki of linked notes, so a half-remembered explanation becomes searchable instead of re-watched.

04:30 – 05:59

06 · Installing the skill and getting a free Gemini key

Walks through adding the watch plugin marketplace in Claude Desktop or ChatGPT, then creating a free Google AI Studio API key on Gemini's free tier (no billing required) and handing it to the agent.

05:59 – 06:14

07 · Updating an existing install

For anyone already running a previous version, updating is a one-step plugin-marketplace refresh that preserves existing engine preferences.

06:14 – 07:04

08 · Local engine as the free fallback

Shows the non-Gemini path: WhisperX transcribes locally for free, and three presets (Efficient, Balanced, Token Burner) trade frame coverage for Claude token spend; he recommends Balanced with WhisperX as the default for local work.

07:04 – 07:57

09 · Choosing Gemini vs. local per task

States his actual routing rule: Gemini handles most lookups and Q&A since it does the watching for free and saves Claude's allowance, but he still switches to local when he needs Claude to see the literal pixels, like recreating a motion graphic frame by frame.

07:57 – 08:07

10 · Outro

Points to the free setup guide link and the next video.

Atomic Insights

Lines worth screenshotting.

  • Even frontier models can't parse video natively, so an agent skill has to translate it into text the model can reason over.
  • Sending a YouTube link straight to an API means no download step, which means no bot-detection wall to trip.
  • Letting the model decide where to look more closely beats dumping every frame: Google's benchmark shows 87% fewer tokens for more correct answers.
  • You can find one buried moment in a three-and-a-half-hour video by describing what you remember seeing, with no timestamp required.
  • Pairing audience-retention drop timestamps with the actual footage turns a vague analytics number into an inspectable moment you can act on.
  • The free Gemini API tier covers up to eight hours of video analysis per day with no billing enabled.
  • The skill works on 30+ platforms, not just YouTube, so TikTok and Instagram footage gets the same treatment.
  • A video can be treated as a source to mine for a personal notes wiki, not just something watched once and half-remembered.
  • A free local engine (WhisperX) stays available for when you don't want to depend on an external API at all.
  • The right engine depends on the task: cloud for answering questions, local when the model needs to inspect actual pixels, like matching a motion graphic frame by frame.
Takeaway

How to make AI actually watch video

AI VIDEO WORKFLOW

A free agent skill hands Gemini the video link instead of screenshots, so your main AI gets timestamped findings back without burning its own usage limit.

02How agentic video understanding works
  • Even frontier models can't parse video natively, so the skill's whole job is turning a video into text the model can reason over.
  • Routing a YouTube link straight to Gemini's API means no download step, so there's no bot-detection wall to get blocked by.
  • The agent decides where to look closer instead of being handed every frame, which is why Google's benchmark shows fewer tokens for better answers.
03Use case 1: finding one buried detail
  • You can locate one specific moment in hours of footage by describing what you remember instead of knowing a timestamp.
  • The same search works on your own raw screen recordings, not just published video, so it doubles as a footage-logging tool.
04Use case 2: diagnosing retention drops
  • Pairing retention-drop timestamps with the actual footage turns a vague metric into a specific, inspectable moment.
  • It can't tell you why someone clicked away, but lining up the drop with what was on screen is enough to spot a pattern across videos.
05Use case 3: auto-saving notes to an LLM wiki
  • Treat a video as a source to mine, not just something to watch once: have the agent file the useful parts into a personal, linked note wiki.
  • That turns "I remember seeing that explained somewhere" into a search instead of a re-watch.
06Installing the skill and getting a free Gemini key
  • The entire pipeline runs on Gemini's free API tier, up to eight hours of video per day, with no billing required to start.
  • Setup happens once per environment; after that, handing the agent a video just works without reconfiguring anything.
07Updating an existing install
  • Updating preserves your existing engine preference, so switching to Gemini is a deliberate choice, not an automatic reset.
08Local engine as the free fallback
  • A free local engine (WhisperX) exists specifically for when you don't want to depend on an external API at all.
  • The three local presets are a direct dial between how much the model sees and how many of your own tokens it costs: Efficient for a quick look, Balanced for normal use, Token Burner for full coverage.
09Choosing Gemini vs. local per task
  • Default to the cloud engine for anything that's just asking questions about a video, since it's free and saves your own model's budget.
  • Switch to the local engine specifically when the model needs to inspect actual pixels rather than a text description, like matching a motion graphic frame by frame.
Glossary

Terms worth knowing.

Agentic video understanding
An AI approach where the model decides which parts of a video to examine more closely, rather than being fed every frame indiscriminately.
LVBench
A benchmark Google cited to show its agentic video approach answering more questions correctly than static frame-dumping, while using far fewer tokens.
WhisperX
A free, local speech-to-text engine used as a fallback when a video is transcribed on-device instead of through a cloud API.
Resources

Things they pointed at.

Quotables

Lines you could clip.

00:07
“AI can now watch videos for free without burning up your usage limit.”
cold-open hook, states the entire value prop in one line→ TikTok hook↗ Tweet quote
01:58
“Say you remember seeing that tool but you've forgotten what it's called. All I need to do is describe what I remember and ask watch to find it.”
concrete demo of the core use case, easy to picture without the video→ IG reel cold open↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

AI can now watch videos for free without burning up your usage limit. Up until now, the only way that AI could watch video meant feeding it hundreds of screenshots. But Google just released something that lets AI understand video perfectly, and you can get it for completely free.
Look at this test from Google. Across 50 video understanding questions, the agentic version gets more answers right while burning 87 % fewer tokens. So I plugged it straight into the watch skill, and now you can use it with any AI tool.
Here's how it works and the best. three use cases for it. Even the best models like Fable and Astra can't understand video out of the box.
So watch is a free agent skill that you can give your AI to give it a way to understand videos. Once it's connected, you just hand it a link and tell it what you want to know. So I'll give it Alex Hormozy's video, how to win with AI in 2026 and ask my Claude to break down the hook and why this video works.
Behind the scenes, the watch skill sends the link and my question to Gemini. Gemini looks through the video, finds the parts that it needs and sends the analysis back to my AI. with the timestamps.
That's what agentic video understanding means here because Gemini can decide where to look more closely. It uses the pictures, the audio, and the transcript together to answer your questions. And the YouTube link goes straight to Google's supported API.
So there's no downloading involved and you never get blocked. My AI gets the findings instead of hundreds of screenshots. And the video analysis runs completely free on Google's free API tier.
And this works exactly the same in ChatGPT as well. And I'll actually show you how to set up both of them at the end of this video. And now that I've got...
the answer back in my main Claude conversation I can ask Claude to use that information that I just got from the video to help me script my next one. Let me show you what else you can do with watch. First up is using AI to find one tiny detail buried in hours of footage.
So let's test it on Andrew of Kapathi's deep dive into LLMs like ChatGPT. It's a three and a half hour long video and somewhere in here is a tool that breaks text into tokens. Say you remember seeing that tool but you've forgotten what it's called.
All I need to do is describe what I remember and ask watch to find it. Slash watch this video, find the demonstration where Kapathi types text into a tool that displays it as colored pieces with numbers beside it.
Tell me the tool's name and website and give me the timestamp. And we're going to open up that timestamp against it to check if that's the actual screen. There it is.
That's the useful part right here. You can ask it about something that only appears for a few seconds and have the AI look for it. And it works on your own footage too.
So if you've got a folder of screen recordings and need a shot where I open a particular menu, I can have my Claude use watch to find it then give that section straight to my editing tools. Number two is using AI to work out what was happening in your video when people stopped watching and then help you actually improve it.
I'm using my book to skill video here as the example. This is the video where I gave Claude five of the best business books of all time and turned them into skills. It's got about a hundred thousand views and because it's my video I have the actual retention data so we can see where people leave then I can look at what I was saying and showing at that specific point.
So I'll give Claude the numbers and then have watch analyze the footage around those drops. Watch this video, use the retention analytics to figure out the three biggest audience drops and tell me what caused it.
Now, Claude can't tell me exactly why somebody clicked away, but it can line up the drop with the actual content and help me look for a pattern. And then once we apply this to multiple videos in my catalog, we can start to incorporate all of those learnings into my scripting for the next videos. And this doesn't just work on YouTube either.
You can actually apply this to any social media because the skill supports over 30. 1 ,300 sites, including TikTok and Instagram. And you don't have to point it just at your content either.
You can point it at other people's content and figure out what made their videos work. So that's using watch to improve your content. But before we get into the next one, if you want to build an AI operating system that runs your entire business, that's what we'll do in my next FounderOS Bootcamp.
I'll hand over my exact setup, all of my AI playbooks, and teach you to adapt them into your business. Join the waitlist below to get first access. And number three is having AI take notes on the...
videos that you watch and save them straight into your LLM Wiki. Because you've probably watched something useful, thought you'd remember it, then have to go and find that entire explanation again a week later. So I want to use AI to keep the useful parts somewhere it can read the next time that we're working on something.
Here, I'm using the same Kapathi video from the first demo, and I'll ask Watch to pull out the explanations and visual examples, then hand them to Claude to save those into my LLM Wiki. These notes are being saved into an LLM Wiki, which is just a collection of linked notes. that your AI can read but you could just as easily have your AI build these out into documents for your notes as well.
So now that you've seen what you can do with the watch skill let me show you how to get it installed in five minutes. The watch skill itself is free and the Gemini model it uses underneath has a free API tier so this entire thing should cost you zero because on the free Google tier you get up to eight hours of YouTube videos per day and if you hit that limit you just wait for it to reset it doesn't automatically start charging you.
Claude still uses his own allowance for the analysis and to put the answer to work but Gemini and I handles all of that video analysis using Google's allowance. So it's essentially free.
And the best part is you only need to do this setup once in the environment where you want to use watch. After that, you can just give it a video as part of whatever you're working on and it'll just work. And if you want the detailed instructions for the setup in ChachiBT and Claude, head to the link below where I've got the full setup guide that steps you through the entire process.
But if you're in Claude desktop, open the customize menu, go to plugins, add the watch repository as a marketplace, then install watch. to the plugin menu as well and then add a marketplace, hit install plugin and you're done. Once you've done that, open up a new chat and type slash watch and go through the setup process.
For the Gemini backend to work, we'll need an API key, which is what lets your AI connect into Gemini. Head to Google AI Studio, go open API keys, click create API key and select your project. Check that the project is on the free tier.
You don't need to enable billing to get started. Give that API key to your agent and it'll save it for you. If you're not comfortable doing that, you can just ask it to open your environment file and save it there manually.
Now give it a video link and ask it to use watch to watch the video. If you've already installed a previous version of the watch skill, all you have to do is jump into your plugin marketplace and hit update. Then ask it to select the Gemini engine or the new local transcription option we'll talk about in a second.
If you install it as a standalone skill, ask it to update it from the same repository, then start a new session. One thing to watch out for, updating keeps your existing preferences. So if you wanna use Gemini, your agent to switch engines and add your key.
And if you don't want to use Gemini at all, the original workflow is still there. Choose local and the skill prepares frames and a transcript for your AI to read. And this version of the skill adds WhisperX, which transcribes your audio on your computer completely free when captions aren't available.
For the local setup, there are settings that control how many frames Claude sees and how many tokens it uses. Efficient adds a smaller selection of frames for a quick look. Balanced picks frames around scene changes like a slide or a screen demonstration and token burner removes the cap entirely for more coverage but uses way more of your court allowance i'd start with balanced and choose a whisper x for the local transcription if you want to go the local route let the model finish its one -time installation and then you're ready to use it i use gemini for most things like asking questions about a youtube video or finding a particular moment it handles the watching entirely so we never have to download the video we just hand the url claude takes the findings so that saves claude having to read all of those frames, but they're actually still jobs where I want to use the original local engine.
Say I'm recreating a motion graphic. I want Claude to see the actual frames directly while it's building so we can inspect the layout, pull another frame of that specific timestamp, and then compare that with its own render. With Gemini, it's working from Gemini's description.
With local, Claude can inspect the pictures itself. That uses more of your Claude allowance, but that direct visual reference can be worth it when you're trying to get an animation looking right or something that needs that particular level of... detail.
But the good news is that you can pick your engine at whatever time and switch up whenever you need. So you don't have to pick one forever. The setup guide for the watch skill is free at the link below there.
You'll find the link to download it directly. And if you want to join my next FounderOS bootcamp, that's going to be down there as well. If this video was useful, hit subscribe.
Thanks for watching.
The Hook

The bait, then the rug-pull.

The video opens on a bold claim: AI can watch an entire video for free without touching your model's usage limit. Brad Bonanno backs it with Google's own benchmark before walking through exactly how his free "watch" skill routes any video link straight to Gemini instead of making your AI read hundreds of screenshots.

Frameworks

Named ideas worth stealing.

06:20list

Local engine presets

  1. Efficient
  2. Balanced
  3. Token Burner

Three settings that trade how many frames the local engine pulls, and how many of the agent's own tokens it burns, for how much visual detail the model sees.

Steal forany local tool where you want a quality/cost dial instead of one fixed setting
CTA Breakdown

How they asked for the click.

VERBAL ASK
03:33product
“if you want to build an AI operating system that runs your entire business, that's what we'll do in my next FounderOS Bootcamp. I'll hand over my exact setup, all of my AI playbooks, and teach you to adapt them into your business. Join the waitlist below to get first access.”

Quick soft mid-roll pitch dropped between use case 2 and 3, then straight back to content.

Storyboard

Visual structure at a glance.

open
hookopen00:00
how it works
promisehow it works00:45
use case 1
valueuse case 101:46
use case 2
valueuse case 202:35
setup
valuesetup05:19
local engine
valuelocal engine06:20
engine choice
valueengine choice07:27
CTA
ctaCTA08:05
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.