How to FINALLY Use Local AI in 45 Minutes
A 44-minute walkthrough of the five-layer stack behind running open-weight AI on your own hardware — and the trick of using Claude Code itself to build the whole thing for you.
July 22ndA Claude Code creator mines his own chat history into a personal test pack, then builds a slash-command benchmark that tells him in one run whether a new model release is actually worth switching to.
Public benchmarks test websites and math problems that barely overlap with real work, so mining your own chat history into a personal, repeatable rubric is the only reliable way to know whether a new model is worth switching to.
Every model launch triggers the same wave of tutorials: throwaway demos and recited benchmark scores from tests — coding arenas, math olympiads — that rarely resemble real work. This video replaces that ritual with a personal benchmark: feed your archive of past Claude or Codex conversations into the AI, let it isolate the tasks you actually run — emails, audits, briefs — and freeze them into a reusable test pack. Add a personal rubric scoring quality, instruction fidelity, token cost, speed, and turns, and one slash command replays your real work against any new model, reporting in 30-40 minutes whether it's actually worth switching or just noise.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →
Opens with the familiar cycle of a new model dropping and a flood of interchangeable tutorials building throwaway demos and reciting benchmark numbers that don't mean anything for the viewer's actual work.

States the video's promise: instead of watching another tutorial, viewers will get a slash command that runs their own benchmark against real work so they can judge new models for themselves.

Public benchmarks like Web Dev Arena, Math Olympiad, and SVG Unicorn tests measure a different population of tasks entirely, with almost no overlap with real work like client emails, audits, and reports.

Explains feeding an archive of past Claude/Codex conversations — over 1,000 threads and several gigabytes — into the AI so it can isolate the recurring tasks a person actually executes day to day.

Warns that letting a model grade itself introduces bias, so the fix is defining a personal rubric first, then turning the mined history into a reusable, weighted test pack.

Once the rubric and test pack exist, comparing any two models becomes a single typed command that kicks off the automated evaluation.

Runs the actual terminal demo, comparing Opus 4.8 on low effort to Opus 5 on low effort for a copywriting task, then confirming the test plan and effort levels before trials begin.

Once trials launch, a browser dashboard spins up and tracks each model's session in real time, confirming both models are executing at the correct effort level.

Walks through the three separate benchmark runs queued up: one isolating a single email-triage task, one comparing low vs. high effort across two models, and one generic run left to the model's own judgment.

Opens a finished benchmark report showing a single task tested to a dead-even result between Fable 5 and Opus 5, illustrating that one tied test isn't enough evidence to justify switching models.

Breaks down the personal rubric into five measurable numbers — quality, instruction fidelity, tokens used, speed, and number of turns — so 'better' is judged on multiple axes instead of one vague impression.

Reviews several completed comparisons, including a four-way Opus 5 vs. Fable 5 test at low and high effort on a debugging task, where instruction-following and a 'soul' score separate the models more than raw quality.

Closes by pointing viewers to the free downloadable /benchmark skill and the paid Early AI-dopters community for going deeper, then asks for a like and comment.
A repeatable, five-number personal rubric run against tasks mined from your own chat history tells you more about whether a new model is worth switching to than any public benchmark score ever will.
“So let me know if this sounds familiar.”
“Their benchmark never sampled your life.”
“You'll be able to compare Opus 4.8 to Opus five to Fable five to whatever brand new model comes out.”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
Every time a new model drops, YouTube floods with the same 100 tutorials — someone building a website you'll never use and reading benchmark numbers as if they mean something for your actual work. This video swaps that ritual for a personal benchmark built from your own chat history.
A scoring rubric a viewer writes for themselves so model comparisons aren't judged on a single vague impression, and so the AI isn't grading its own work.
The four-step loop that replaces watching a tutorial every time a new model launches.
“Grab the skill for free (link in the description)”
Soft CTA delivered as a lower-third over the talking-head close, paired with a pitch for the paid Early AI-dopters community, then a standard like/comment ask on the final frame.
00:00
00:10
00:12
00:20
00:29
00:33
00:38
00:45
00:51
00:57
01:03
01:10
01:14
01:20
01:26
01:32
01:38
01:44
01:51
01:57
02:04
02:10
02:17
02:23
02:30
02:36
02:43
02:49
02:55
03:01
03:07
03:13
03:19
03:25
03:31
03:37
03:43
03:49
03:56
04:02
04:09
04:15
04:21
04:27
04:34
04:40
04:46
04:51
05:00
05:05
05:13
05:19
05:24
05:30
05:36
05:43
05:50
05:56
06:03
06:10
06:16
06:22
06:28
06:36
06:40
06:46
06:53
06:59
07:06
07:12
07:18
07:24
07:31
07:38
07:43
07:51
07:55
08:01
08:05
08:13A 44-minute walkthrough of the five-layer stack behind running open-weight AI on your own hardware — and the trick of using Claude Code itself to build the whole thing for you.
July 22ndA 24-minute Earth-layers framework for building AI operating systems that don't decay.
June 30thA 36-minute blueprint for moving a personal AI agent stack into a locked-down, compliance-ready AWS environment — built over a month and nearly 10 million tokens.
June 25thA 9-minute system for mining your JSONL session logs, measuring the behavioral gap between Fable and any other model, and injecting a distilled playbook at every session start.
June 14thA breakdown of why maxing out effort settings on Claude, GPT, Grok, and Gemini rarely makes the output better — and the framework for picking the right level every time.
July 13thEighteen real Claude Code sessions later, the model that's half the price per token isn't automatically the cheaper one to actually run.
July 24th