A 34-minute in-person conversation about why the thing wrapped around the model matters more than which model you picked.
Posted
3 weeks ago
Duration
Format
Interview
educational
Views
72K
721 likes
57 · 43
Big Idea
The argument in one line.
The model is only the brain; the harness around it decides what actually gets built, so your leverage comes from owning portable skills and assets rather than staying loyal to one provider.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You already use Claude Code or Codex daily and keep second-guessing which one to run for a given task.
You are building a personal AI operating system out of skill files, rules, hooks and context documents and want it to survive the next model release.
Your skills folder has grown past what you actually use and you suspect some of it is now making outputs worse rather than better.
You are evaluating local open-source models and want a clear picture of what they can and cannot do without a harness around them.
You run client work through agents and need to tell a model problem apart from a setup problem when something breaks.
SKIP IF…
You have never used a terminal-based coding agent, because the whole conversation assumes you already know what a skill file and a CLAUDE.md are.
You want a feature-by-feature benchmark comparison of Claude Code versus Codex, because the argument here is deliberately the opposite of benchmark chasing.
You are looking for step-by-step setup instructions, since this is a discussion of principles with one short live demo and no walkthrough.
TL;DR
The full version, fast.
The model is only the brain in a jar. Everything that lets it read files, edit code, run shell commands and loop until a task is finished lives in the harness around it, which is why a local model can write your HTML but cannot spin up your server. That reframe changes what you optimize. Stay unloyal to any provider, build skills and rules that work with any model, and treat them as assets you port between Claude Code, Codex, Hermes or your own open-source harness. It also changes maintenance: every layer decays at a different rate, so audit skills monthly, keep rules scoped to projects, and delete the ones a smarter model no longer needs.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Cold-open montage of the strongest lines, then the setup: two creators sitting down in person in Montenegro after an AI event, framing the question of what stays relevant as tools churn.
02:05 – 07:19
02 · Models vs. harnesses
The brain-in-a-jar diagram. Read, write, edit and bash are limbs the harness supplies, not model capabilities, and the agentic loop is what turns output into finished work.
07:19 – 09:26
03 · Local models in LM Studio
Live demo: a local open-source model is asked to build a landing page and host it. It produces the HTML and cannot serve it, because it has no limbs.
09:26 – 10:29
04 · Clay sponsor
Sponsor read for Clay, a data enrichment and orchestration platform, demoed through its CLI inside Claude Code.
10:29 – 13:37
05 · Choosing and switching AI agents
Zero loyalty to providers, total loyalty to harness and assets. Every skill gets converted to work on both Claude Code and Codex, and the two are set against each other to fight over a plan.
13:37 – 16:47
06 · Hermes, OpenClaw, and Pi
Why Hermes beat OpenClaw on harness quality, how to massage any harness into the verification loops you like, and reverse-engineering your own harness from pi.dev plus your JSONL conversation logs.
16:47 – 19:17
07 · Making your skills portable
Skill files share a structure: YAML header, kebab-case name, a description full of trigger words. A conversion skill reads each provider's documentation and writes one version that fires reliably in all of them.
19:17 – 21:46
08 · Understanding your AI system
You can outsource the thinking but never the understanding. Then the March 2026 Claude Code harness leak, and the discovery that most of it was crutches rather than intelligence.
21:46 – 24:14
09 · Keeping your AI OS updated
The rot.md file. Identity and core context stay stable for months, hooks for about six, rules can go stale in a week, and skills need the most aggressive auditing of all.
24:14 – 30:01
10 · When models outgrow your skills
Why deleting all your skills every six months is defensible, the YouTube-advice analogy for an obsolete skill, and the tradeoff that skills also act as guardrails you may still want.
30:01 – 33:37
11 · Organizing projects and global rules
Twenty-five isolated operating systems instead of one, a promotion rule that keeps global config tiny, and the diagnostic question of model versus harness versus organization.
33:37 – 34:08
12 · Final thoughts
Mastery is knowing how the whole factory works end to end. Before blaming the model, audit your setup, your skills and your bloat.
Atomic Insights
Lines worth screenshotting.
A local model can write the HTML for your landing page but cannot start the server, because serving files is a harness capability, not a model capability.
Read, write, edit and bash are not baked into any model. They are limbs the harness bolts on, which is why the same brain behaves completely differently in two tools.
The agentic loop is just ask, execute, read the result, repeat. How well a model survives that loop depends almost entirely on how sophisticated its harness is.
Be loyal to your harness and your assets, never to a provider, so you can swap the brain out in a day when a better one appears.
Claude Code behaves like a gifted artist that wants to ideate and push back; Codex behaves like a surgeon that follows the exact instruction to completion.
Putting Codex and Claude Code in a loop to critique each other's plan for ten minutes surfaces the gaps neither one catches alone.
Roughly 70 to 80 percent of day-to-day agent work can run on open-source local models, with the expensive models reserved for genius-level planning tasks.
The leaked Claude Code harness was mostly crutches, including regex lists of swear words, which is evidence the magic lives in the scaffolding rather than the model.
Judge a model release by how much the provider improved the harness around it, not by which benchmark it crushed.
Every layer of an AI operating system decays at a different rate: identity lasts months, hooks last about six months, rules can rot in a week.
Skills decay fastest of all, because a skill is a crutch and a smarter model may accomplish the same goal from a vague prompt without it.
Boris Cherny's advice to delete all your skills every six months works because you have to prove each one still earns its place rather than assuming it does.
Most people have hundreds of downloaded skills loading on every run and actually use five, which is pure context bloat with no upside.
Keeping twenty-five separate operating systems instead of one giant one gives you a small blast radius when an experiment goes wrong.
Everything stays project-scoped until it earns promotion to global, so you always know the running total of what applies to every single run.
When output degrades, ask whether it is a model problem, a harness problem or an organization problem, because only isolation makes the answer findable.
You can outsource the thinking, but you can never outsource the understanding, which is why you still need to know where every file lives.
Anthropic telling people to give models more ambitious tasks is a hint that rigid skills may be capping what the model would do on its own.
Takeaway
The harness decides what the model can do
WHAT TO LEARN
Treat the model as a leased brain and the harness around it as the asset you own, then maintain that asset on a schedule because every layer of it rots at a different speed.
02Models vs. harnesses
Read, write, edit and bash are supplied by the harness, not the model, which is why the same brain behaves differently in two different tools.
The agentic loop is ask, execute, read the result, repeat. How far an agent gets depends on the harness quality, not raw intelligence.
A better model does not gain new abilities, it uses the same tools with more judgment about when and which to reach for.
03Local models in LM Studio
A local model asked to build and host a landing page will write the HTML and stall at hosting, because serving files needs a limb it does not have.
Running a task in a tool with no harness is the fastest way to see where the model ends and the scaffolding begins.
Roughly 70 to 80 percent of routine agent work can sit on local open-source models, with expensive models reserved for planning and genius-level tasks.
05Choosing and switching AI agents
Stay loyal to your harness and your assets rather than a provider, so a better brain can be swapped in within about a day.
Claude Code leans toward ideation and pushback while Codex follows exact instructions to completion, so match the tool to the step, not the whole project.
Having two agents critique each other's plan in a loop for ten minutes surfaces the gaps that neither one catches working alone.
06Hermes, OpenClaw, and Pi
What separates one harness from another is the verification loop, and that behaviour can be copied into an open-source harness you control.
Your agents' JSONL conversation logs record which tools fired and in what order, which makes them a reverse-engineering source for harness behaviour.
Point an agent at a harness's documentation and your own logs and have it rebuild the parts you like into something you own.
07Making your skills portable
Skill files share a structure across tools: a YAML header, a kebab-case name, and a description packed with trigger words that decide when it fires.
Different providers weight different parts of a skill file, so one portable version needs a description beefy enough to trigger reliably everywhere.
Python inside a skill is already portable. The part that needs adapting is how the agent knows when and how to invoke it.
Auditing your whole ecosystem monthly catches skills you never use, pairs that should merge, and ones that have drifted out of date.
08Understanding your AI system
You can outsource the thinking but never the understanding, which means still knowing which documentation to feed in and what to do with the result.
The leaked Claude Code harness was full of crutches like regex swear-word lists, evidence that scaffolding is doing more work than the model gets credit for.
Judge a release by how much the provider improved the harness around the new model, not by which benchmark number it beat.
Trust empirical results from pushing a model on your own work over the performance a provider claims its intelligence should produce.
09Keeping your AI OS updated
Track decay per layer: identity lasts one to three months, hooks around six, rules can go stale weekly, and skills rot fastest of all.
Rules deserve the most frequent revision because every new task exposes edge cases the existing rules never anticipated.
Hooks that enforce safety, like stripping client data before a push, stay viable far longer than any instruction about how to do work.
10When models outgrow your skills
A skill is a crutch for the model, so a smarter model may reach the same goal from a vague prompt without the skill injected every run.
Delete skills aggressively on a schedule, because the test is whether the skill still adds value, not whether it worked when you wrote it.
Removing a skill can also remove a guardrail, so the rules layer has to absorb the constraints the skill was quietly enforcing.
Watch for the moment a model outgrows a skill: the instructions you once needed become advice the model would have ignored anyway.
11Organizing projects and global rules
Keeping many small isolated operating systems instead of one large one limits how far a bad experiment can spread.
Make everything project-scoped by default and promote to global only when it earns it, so you always know the full set of global rules.
When quality drops, separate model problem from harness problem from organization problem, because only isolation makes the real cause findable.
12Final thoughts
Before concluding a model is bad, audit your setup, your skills and your accumulated bloat, since the same model performs differently per environment.
Mastery here is knowing how the whole factory works end to end and where each piece sits, not memorising which tool is currently best.
Glossary
Terms worth knowing.
Harness
The software wrapped around a language model that gives it the ability to read and write files, run shell commands, call tools and keep looping until a task is done. Claude Code, Codex and Pi are harnesses; the model is swappable inside them.
Agentic loop
The repeating cycle an agent runs: it calls a tool, the tool executes, the result comes back as either success or an error, and the agent uses that result to decide its next action.
Skill
A markdown instruction file, usually with a YAML header naming the skill and describing when to trigger it, that teaches an agent a specific procedure. Often bundled with scripts or assets the agent can call.
AIOS
An AI operating system: a personal stack of identity documents, core context, skills, rules, hooks and agent definitions that an AI tool loads so it works the way you want across many tasks.
rot.md
A file that records how fast each layer of your AI setup goes stale, so you know to refresh rules weekly, skills monthly and identity documents only every few months.
Hooks
Automated checks that fire on a defined event, such as stripping sensitive client data before anything gets pushed to a public repository.
LM Studio
A desktop app for running open-source language models locally. It gives you a chat interface over the raw model with no tools attached, which makes it a clean way to see what a model cannot do on its own.
Pi
An open-source agent harness, available at pi.dev, that you can install on your own machine and modify, as opposed to using a provider's closed harness.
JSONL conversation logs
The line-by-line transcript files that coding agents write to disk, recording every tool call and verification step. They can be read back to reverse-engineer how a harness actually behaves.
Blast radius
How much of your setup an experiment can break. Keeping configuration scoped to one project rather than global means a bad change affects only that project.
CLAUDE.md
A project or global instruction file that an agent reads at the start of a session to pick up standing rules and context about how to work in that environment.
24:06linkBoris Cherny's tweet on deleting all skills every six months
28:40linkAnthropic model launch blog posts (benchmarks and prompting guidance)
20:18linkThe March 2026 Claude Code harness .map file leak
Quotables
Lines you could clip.
20:48
“This stuff isn't like magic at all. There is a whole factory of workers that are making this model look way smarter than it is.”
lands the thesis in two sentences with a concrete image and no setup→ TikTok hook↗ Tweet quote
11:16
“My number one goal is to never be loyal to a provider, to only be loyal to my harness and my assets. And I will switch the brain interchangeably.”
a clean stance statement that works as a standalone philosophy clip→ IG reel cold open↗ Tweet quote
19:27
“You can outsource the thinking, but you can never outsource the understanding.”
aphorism, quotable verbatim, no context required→ newsletter pull-quote↗ Tweet quote
12:15
“So I see Codex as a surgeon and Claude Code as a gifted artist.”
one-line comparison that settles a tribal debate people are already having→ TikTok hook↗ Tweet quote
24:06
“Skills and agents decay incredibly fast, to the point where Boris Cherny dropped a tweet saying you should delete all of your skills every six months.”
named source plus a counterintuitive instruction, built-in controversy→ IG reel cold open↗ Tweet quote
24:56
“You want to make sure that your skill is actually adding value and not holding back this dragon that gets only bigger with time and smarter with time.”
vivid metaphor that reframes skills as a cap rather than a boost→ newsletter pull-quote↗ Tweet quote
02:57
“It is a brain that can tell you and give you this output of the HTML, but it can't go and spin up a local server on your computer. It doesn't have the limbs for that.”
the single clearest explanation of the model-versus-harness distinction→ TikTok hook↗ Tweet quote
33:25
“Mastery is just understanding how the entire factory works end to end and where each piece lies.”
closing line, works as a standalone definition of competence→ newsletter pull-quote↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphoranalogystory
All right, so Marsh, by the end of today's episode, what will everyone have learned from you? My goal is that by the end of this video, you understand that the harness of a model is much more important than the model itself. It feels to me like Claude Code is like a wise old owl, and then it feels like Codex is like the Rottweiler.
It'll obey your commands, and it will just keep going until it's done. Because whenever we talk about, oh, Claude's better, or Codex is better, you have this brain they're all fighting about. But everything around the brain is actually what gets it to tick.
This stuff isn't like magic at all. There is a whole factory of workers that are making this model look way smarter than it is. My number one goal is to never be loyal to a provider.
To only be loyal to my harness and my assets. And I will switch the brain interchangeably. So I think that is really important to be thinking.
You're building up your own IP. And you need to make sure that you're protecting that. And I always think of the quote.
You can outsource thinking, but you can never outsource the understanding. Skills and agents, though, decay incredibly fast to the point where Boris Cherny dropped a tweet saying you should delete all of your skills every six months. All of them.
Did he? Do you know why? Why did he say that?
All right. So Mark, thank you so much for joining us today in person, which is awesome. We're in the beautiful country of Montenegro, which has been so much fun.
We're here for an AI event and we figured why not sit down together and hopefully drop some sauce today. So I'm super, super pumped to be sitting down. If you guys don't know who Mark is, then hopefully after this video, you start seeing his videos on your YouTube feed because his stuff is absolutely gold.
I've been watching him for a while. I actually was watching him before I started making content. So pretty cool moment for me to be able to sit down with Mr.
Kashaf here today. But yeah, I'm super excited to dig in. Been getting, obviously, tons of questions in the community and discussion around, oh, should we be using Hermes now?
Or we've seen this local thing called Pi. And there's just so many tools going around. And I think we want to make sure that what we're building at the end of the day is still relevant next year and the year after that.
Because we don't know what might happen to Cloud Code next or Codex. So super excited to dig in. And yeah, thanks for kind of throwing together an escaldron.
Let's start getting into it. Absolutely. So the first thing I want to do is just really lock in on this diagram, specifically this little brain in a jar here, because whenever we talk about, oh, Claude's better or Codex is better, Gemini, maybe one day might be better.
You have this brain they're all fighting about, but everything around the brain is actually what gets it to tick. So if we look at the different parts here, so. We have the ability to read.
We have the ability to write and edit files. We have the ability to use what's called Bash, which basically takes control of your computer, makes folders, moves folders. All of these things aren't baked into the model.
And you see this if you ever use a local model. You say, make me a website and spin it up on my local computer. It can do the first part, but it can't do the second part.
It is a brain that can tell you and give you this output of the HTML, but it can't go and spin up. a local server on your computer. It doesn't have the limbs for that.
So the more you start thinking about models in the sense that they are the brain, but everything around them allows them to interact and have hands and legs, then you start to really separate what is the importance of the brain versus everything around it and how can you enrich everything around it so you're less dependent on that brain.
Because we are at a point where many local models, whether it's a Kimmy or insert name of open source model here. they can do like 80 % of the day -to -day work. You might not get the same firepower as you would with Cloud Code or Codex, but for more and more roles, increasingly for next year especially, I see a world where you run 70 % to 80 % on local, assuming you have the hardware, and you bring in the geniuses for genius -level tasks or planning.
I love that because I think when you start to really talk about Cloud Chat in the web versus a Cloud Code, there is... a gap there, which honestly, I wish they didn't call it Cloud Code. Because the code part is all of these tools that you mentioned.
And so just to walk us through a really practical example, what is the difference between you asking Cloud Chat to, let's just say, research something for you and create a PDF versus when you might ask Cloud Code to do that exact same task? Yeah, so Cloud on the web versus Cloud Code. Yeah, so Cloud on the web will have slightly different tools.
Let's say you're using Cloud Chat. And all it can really do is do the researching part, create the PDF part. But some parts in between, maybe calling additional platforms or moving files on your computer, it might not have access to local files on your computer.
It would have to do a lot of work, assuming things in the background that it can't actually touch and feel and see. With the cloud code as a harness, it can not only interact with your local computer, it can interact with the cloud. So you have the best of both worlds.
So you have one that is telling you, hypothetically, here's what we could do. And it can do some stuff. Increasingly, it's getting better.
I see a world where Cloud Chat evaporates completely. And all we have is CodeWork, which is a light version of the harness of Cloud Code. So everything that we will interact with will have a harness.
It's just to what extent is it highly capable to do the tasks you're looking for? 100%. And what I think is so cool about Cloud Code, and when I started learning about it, I remember how intimidated I was to start learning about it back in maybe January of this year.
But when I just... started asking it questions it feels like magic because yes you're interacting with the same model you might be used to but all of the harness stuff really just happens automatically because the harness is essentially built to understand here are the tools that i have as you can see in this diagram which well done on this diagram by the way it knows what's in there same way like you think you need to pick up a glass of water yeah your hands and your shoulder and your it just works together to do it for you so i think that it's really cool to see something like this even though It might at a glance look like you might have to know how to do the bash or the read or whatever.
Yeah. But the model just takes care of it. Yeah.
And what I want to focus on is although these come out of the box, right? That's the whole point of the Cloud Code Harness. That's what made it amazing.
You have the read, you have the edit, all the stuff is done for you. And even with things like Pi, which is kind of a very vanilla open source version where you can build your own harness. That's all cool.
But where you come in and where your channel has been really adding tons of value is... What else can you add to this factory? So now you have things like Lego blocks that are modular.
And these are those skills, these plugins, all of these additional things that you can layer on. So then you have this ecosystem. We have this orchestra.
We have the brain in the middle telling everything else exactly what the goal is. And depending on the intelligence of the model, it might become better at knowing, ah, for this, I need some bash with a write. And I'll need to go through what's called the agentic loop, which is purely you ask thing.
Thing gets executed. Result of thing happens. The result could be an error.
It could be a success. It takes that stimuli and it keeps going in that loop. And its ability to keep going in that loop, comma, well, is fully based on how sophisticated that harness is.
A better model will know how to use tools better. It's kind of like bringing an expert handy person who's worked for a year and studied for many ages versus someone who has... all the battle scars, the really powerful types of tools and has a tacit knowledge of when and where to use them a little bit better.
They'll perform infinitely more powerfully than the first one. The same concept here. Both are going to be smart.
The model itself is great. But if I showed you right now an example, can we pop over to, let's say, an LM Studio? So let's pop over here.
Yeah. So what is LM Studio to anyone that's never used it before? Yeah.
So think of LM Studio as your ability to run chat GPT with... an open source model, it doesn't have a harness by default. So you can just interact with the brain, which is beautiful, because it'll show you exactly why we have an issue here.
So if I go, and I'm just using QN27B here, I have some more powerful models, but I don't want this computer to explode while I'm recording it. So I'm just going to say using the beautiful Guido. Okay, so I want you to make a very basic landing page for my AI consultancy.
I want you to call it prompt advisors, and I want you to spin it up locally on my computer so I can host it and show the entire audience. Now, if I run this over, no matter how much time this takes, it will be able to tell me that it can't spin it up. The reason why is it does not have the limb of being able to interact with my computer to even create that server.
You get this exact same request to Codex or Cloud Code. it's not going to sweat even twice because it comes with that out of the box. So now it's going to create what is the HTML itself because all it can do is take input and get output.
Imagine if your brain was in a jar. All you can get is some form of stimuli and release the signal of what you think the answer would be. Same concept.
So one can theoretically do the thing, but can't take it from A all the way to the touchdown. The whole point of the harness is how do we go from the output of this very intelligent model to some form of tangible output in your hands.
Totally. Yeah. And I actually remember hearing some stories when, you know, Claude first started to come around, ChaiGBT first started to come around before we had the harnesses.
And you would hear these stories of these companies that were building products because they were having Claude write code and then just copying and pasting it into whatever they needed to actually build the code and host it and all that kind of thing. And that's just, it's another one of the examples that goes right along with like, this is still inherently very powerful.
but not as powerful as when you kind of give it the whole agentic loop that you talked about earlier. All right, guys, real quick, huge thanks to Clay for sponsoring this part of the video. Now, one of the most common questions I get is where to actually find leads for cold outreach and how to learn enough about these people to send them something that they'll actually open.
Because a raw list of names doesn't tell you things like if they have decision -making authority and how to reach them or what they even care about. So Clay is a data enrichment and orchestration platform, and it gives you access to over 250 data providers and AI research tools all in one spot. So instead of asking your agent to dig up whatever it can find online, you can research a whole list at once and pull the right person to contact, their work email, and the size of their company.
And if one provider comes back empty, Clay just moves down the list to the next one until it gets you a result. And what's really cool is the logic you build stays attached to your data, so the same steps can be run again and again on every new lead that you add. And you can see each step that it took, like which provider, every value came from, and what it cost you.
You can build all of this from Cloud Code with the Clay CLI, which is what I'm doing right here. So try Clay using the link in the description, and you'll get 2 ,000 free credits. Now let's get back to the video.
Okay, so you talked about how important these pieces are, because this is where you can really start to add in your own subject matter expertise, which is really what helps make the system feel more like it's yours. And what's cool about this stuff is that as you build on maybe these skills, context files, plugins, whatever it may be, You're not locking yourself in to that harness because these can be used across other harnesses and other models as well.
So the question I wanted to ask you is, as you switch through these things and, you know, your Codex can touch your AIOS, your second brain, whatever everyone's calling it these days, Hermes, OpenClaw, whatever comes next. How do you personally, Mark Kashyap, how do you think about the way that you switch between those harnesses?
And like, you know, if you like Codex for certain tasks, Hermes for certain tasks, what does that look like for you? For sure. So my number one goal, is to never be loyal to a provider, to only be loyal to my harness and my assets.
And I will switch the brain interchangeably. I have zero loyalty to that. So although I use Cloud Code a lot, it's more so I'm used to it.
I know the rhythm of it. I know what to expect. But I will create all of my skills so they can work with any language model, open to closed source.
And I'll prioritize really battle testing it with every single thing that I can. So I'll try with Cloud Code, then have the skill I call slash poly skill. It'll convert any skill for Cloud Code and optimize it for Codex.
Okay. So I'll make sure that every skill is eligible for both. Okay.
And that all of the different models know exactly where to find the same assets. So I have one link that has all of my core assets. They're agnostic of working with any of those models.
So I can move around as needed. Yeah. Now, with your question, Codex recently, I've been running 60 % of the time, although I've been very loyal to Cloud Code for the vast amount of time.
Claude Code is amazing at ideation and planning to an extent. It's a visionary. It likes to go back and forth.
It likes to judge you. But the one thing it doesn't like to do sometimes is follow the exact instruction and the exact way you gave it. So I see Codex as a surgeon and Claude Code as a gifted artist.
And many times I have to bring in Codex to look over the plan of Claude Code and I make them fight in a loop for 10 minutes. different rounds until cloud code finally has all the missing parts that codex could see all the things it wasn't anticipating totally yeah yeah i heard this tweet that or i saw this tweet that i thought was awesome and i want to see if you agree and i think you will because i have a very very similar philosophy of the way i think about the two but it basically said like it feels to me that cloud code is like a wise old owl you can talk to it you plan with it it'll you know push back on you a little bit and then it feels like codex is like the rottweiler that will grab onto the task and it will just It'll obey your commands and it will just keep going until it's done, essentially, because I think that the verification loops inside Codex feel really, really sharp to me.
But I think it's important that we kind of have that acknowledgement of, you know, which harness is best, which model is best. And it's more so which one is best for this specific task. It might be, you know, a five -step process, but for step one and two, maybe that's where you go for the Codex and then, you know, or vice versa.
So I think that's good to hear you say as well. Now, where do you see, you know, I think that Hermes and OpenCloud kind of get bucketed in together as well. Where do you see the differentiation there and, you know, with Codex and Cloud Code as well?
Well, with Hermes, what are you doing? You are bringing in the Hermes harness. You're just looping in whatever model you want.
So the reason why people have to really make their Hermes agent tailored to them, a special snowflake, is depending on how they want to use these models. With Hermes, they have to keep hacking Hermes' harness. All of these skill MD files, they're all great.
But the thing that made Hermes better than OpenClaw is its harness. So with many tasks, if you go one -to -one, Hermes agent versus Codex, you'll have a different result, vanilla. But you can eventually massage Hermes agent's harness to do this verification loop that you mentioned earlier that you really like about Codex.
So that performs very similar to it. So you can do a level of monkey see monkey do. So one thing that I did, and you can go, obviously no affiliation here, it's open source.
If you go to pi .dev, you will have this harness that you can bring onto your computer. You can copy with one command, or if you're feeling daunted, what I did is I take this link, I feed it to Codex or Cloud Code and say, go read the documentation, fan out some agents, learn about this whole harness thing. Once I do that, I can then ask it to go through what are called JSON -L files.
Basically, every conversation you have on your Codex and Cloud Code exists in your computer. So I'll have it go look through all the conversations because they have the metadata of what tools were called, what verifications were done, in what order. And then I can have it monkey see, monkey do.
How do I start to make my own version of the harness that would react and do the same things based on similar types of tasks? You can start to reverse engineer all of the things that you love about Cloud Code and Codex. bring it to your own harness and eventually you can have one harness for everything where you bring in all these models as a brain that's leased you swap out as you need so that's where i think we will get to where right now everything's tribal right youtube is tribal x is tribal i am team codex i am team cloud code i am team open source you all are neanderthals right for me i'm anti -tribal i am how do i make a system where any brain that could serve me for The best speed for the best rate at the best time can be swapped in with little to no acclimation needed.
So it's not a pain for me. If Gemini wakes up tomorrow, truly wakes up and now it becomes amazing. Okay.
It would probably take me 24 hours to swap everything to Gemini. And I'd love that nimbleness. Yeah.
Because while Codex is amazing today, Cloud Code might come because they're probably going to IPO sooner, make a big push. And if we finally get a mythos that's not nerfed or neutered to infinity. You might want to move all your stuff to there.
Yeah. Yeah. And who knows what could happen from a price perspective as well for us as consumers.
So I think that is really important to be thinking you're building out your own IP. Essentially, you need to make sure that you're protecting that. And it's nimble.
So I love that point there. Now, you mentioned something earlier about making sure that your skills and your whole ecosystem is model agnostic. And you have a special skill that you use to make sure that they can work and are optimized for different.
models and harnesses as well. What does that actual process look like? Because typically when we see like our skill files or folders, it's usually a markdown file.
And that's sometimes associated or, you know, kind of like also has in there maybe a few Python scripts or whatever the skill does. Maybe there's some assets that go along with it, but ultimately you've kind of got just like a master markdown file. So how do you actually make sure that Codex can pick it up and use it as well or other agents could pick it up?
Absolutely. So the main thing to remember is that, like you said, All these skill files structurally look similar.
They have this thing at the top called YAML, where it's the name of the skill and what's called a kebab case. You have the description. And in the description, you have a series of trigger words where when user does X, I want you to invoke the skill to do Y.
So what I did is I offload this. I asked Codex, go and take a look at all the documentation from Cloud Code. Look at your own documentation.
Look at documentation from, let's say, this specific other provider and go see how to build a versatile, Swiss knife skill that will work for all of them as optimized as possible. So Codex prioritizes some things about skills that Cloud Code doesn't and vice versa.
So how do we make sure that both are included? Now, when it comes to scripts, Python is Python, luckily. So that is already generic on its own.
No need to worry about that. How you invoke that Python, how it knows when to use it and how to use it, that's where you might need to massage it a little bit. So even when I use this slash poly skill, it will always look for these slight differences knowing, ah, Codex might miss this in the way that you're triggering it using cloud code.
So let's make the description that much more beefier. So the likelihood that it picks it up across the board is much higher. Totally.
So I just have AI do the dirty work to go see AI documentation. And I update these monthly on a cron job. So every month, I will auto audit my entire ecosystem, refine all my skills, see where I'm not using skills that I've added.
Because a lot of people, your audience and mine, have bloated skill repositories where they downloaded some awesome skills thing. They have 300 of them.
They load every single time and they use five. So it reduces the number of skills that I have. It combines the ones where there's an opportunity for a compound skill.
And then it makes sure that they're all model agnostic. I love that. Yeah, I think what's really important there is, you know, what you said, you have AI do the dirty work, but you are still very much in control and you still understand.
And I always think of the quote. You can outsource the thinking, but you can never outsource the understanding. And I think it's a great mindset shift to realize that even us as creators, a lot of the things that we don't know or that we need to learn, we have AI help us with it.
But we still understand how to feed in the documentation for it to look through. And we understand now that it has this knowledge. what to do with it.
And I always kind of say this in my videos, even though it might like hurt the views. Sure. Is that ultimately like just use it as your, your thought partner, as long as you're not outsourcing everything.
I think that's a really important way to think about how you as a person continue to learn more too as well. Absolutely. And one thing you can do is again, if we move into this world of you owning your own harness, you can have the same task be executed and then cloud code or codex can watch it and it can run it within the terminal.
on your PI harness, run it on the other model providers, observe exactly what happened and what was the end result and try to continually understand and reverse engineer what happened with the others that's not happening with yours. Back in March of 26, we had the cloud code harness leak. If you remember that.
It was a .map file, it was leaked to the whole world. I spent four to five days, and I'm pretty sure you went and made a couple of videos as well, looking through every single piece of it. And the coolest part was, The majority of the harness was full of all of these crutches they gave to the brain, the model, to not swear at the user.
To detect when you had swear words through a list of regex of all the swear words you would have that would tell the brain how to react. So once I saw there were so many crutches for this supposedly AGI level model, it made me wake up. Ah, this stuff isn't like magic at all.
There is a whole factory. of workers that are making this model look way smarter than it is. And as soon as you understand that one concept, that's when you snap out of it.
You start looking at model benchmarks very differently. You are not as wowed by what bench it crushed. You're more wowed by how well the model provider do in now improving their harness to have a symbiotic relationship with this brand new model's brain.
That's why I care about empirical. How well does this do when I push it versus how well are they telling me? it should perform based on how smart it is.
Absolutely. There's a question I wanted to ask you that I get a ton. And it revolves around this whole harness idea.
It revolves specifically around this idea of building out your own second brain or OS. The question that I get a lot is about as you every month are adding in new things, whether that be skills or an LLM wiki, how do you yourself make sure that it's staying optimized.
And I put that in air quotes because I don't, I truly don't believe that there's only one optimal way to do it. I think it's just a matter of making sure that you can feel when it's maybe searching too long for something that should find right away, or it's hallucinating information because of the bloat in a certain folder.
So I would love to hear just kind of like dive into your brain a little bit. Sure. How do you think about keeping that organized and efficient?
Sure. So the biggest thing that I've done is create what's called a rot .md file. Rot meaning decay.
So if you have different layers of an AIOS and you have entire courses on this, you covered this at length. You have five to six layers. One could be your identity, then your substrate, which is your core context that shouldn't change that much.
Then you have your skills, your rules, your hooks, you have your agents, and then you have additional things you can add on. All of these different parts decay, become obsolete or rot. at different rates.
So who you are, what you do, your goals, your aspirations, unlikely to change very quickly. So you could have one to three months, maybe even maybe longer, if it's a company actually using this, where it's relevant. Skills, like I said, I update and optimize on a monthly basis.
I just made a brand new one. I like the word action. And then you have, let's see, your rules.
Your rules can change daily. I have rules that change daily because as I do new tasks with the same AIOS, I find new limitations and new edge cases. So I find ways to refine and make better skills that are more compound.
My rules probably decay every week. My hooks though, so things that it always checks. So if I push a client project to GitHub, it wants to make sure that there's a hook that fires to remove any PII, any sensitive data of that client that I don't want living on my GitHub.
That may not decay for six months. I might change it a little bit, but it won't really be viable to be decaying at a very high rate. Skills and agents though, decay incredibly fast to the point where Boris Cherny dropped a tweet saying you should delete all of your skills every six months.
All of them. Did he? Do you know why?
Why did he say that? Because a skill is basically an extra crutch. that we're adding to the system to help the brain do things or the brain better understand how to use all these tools to do the thing.
But if the model truly gets way smarter, it might not need the skill to shortcut what tools to use as disposal. It might be able to accomplish the goal of the skill with purely a vague prompt versus the vague prompt plus the skill that's injected every single time. So you want to make sure that your skill is actually adding value and not holding back this dragon that gets only bigger with time and smarter with time and more powerful.
So skill is one thing you want to audit very aggressively because you might not have a skill issue with that thing you built it for three months down the line, six months down the line. With things like, let's say, agents, you hire agents like you hire employees. You might not always need a bookkeeping agent.
Because if the model gets good enough, you might be able to give big prompt. It would know exactly what are the 18 to 20 different subtasks based on all the memory, assuming your memory gets better, that it needs to execute. So we could live in a world where you have a handful of skills, a handful of rules, and those skills are actually very specific tacit knowledge that it would never know no matter how smart it was versus step -by -step instructions.
Wow. Yeah, that is really interesting. make it a point to every time a new model comes out and i have you know i've got the new opus model or whatever plugged into plug code i always make it a point to have that model run through all my skills just to make sure they can use them in the same way and things like that but i have never thought really to you know i hadn't seen that tweet i hadn't thought to go back to some of the like through all of them and say like do we even need this and what happens if you try to run that process without the skill exactly so given your meteoric rise on YouTube.
If I told you, hey, Nate, here's how you make a YouTube video, right? Two years ago, this might've been a helpful skill for you to learn from me before I started before you. But now I give you that same play.
You'd look at me, tilt your head and be like, have you looked at our subscriber counts? Right? You don't need this skill anymore.
You have outgrown this skill. That's really interesting. So as a model becomes very diverse and has training data, you might not need this skill anymore.
It might be obsolete for where the model's at. Have you had an example yourself where you've been able to remove a skill because the model and the artists are now just can crush it out of the water? Basically, now if I tell it, go and read all the conversations that we've had and tell me 15 different things we can optimize.
I used to have a skill that would tell exactly where to look. for the files that are associated with our conversations, how to break it down so it doesn't blow its context window with 50 ,000 tokens at once times 100. I used to have to really explain that bit by bit.
And now I went from skill to now we're moving to skill. And now I have one basic sentence and it knows exactly to optimize. Based on my CloudMD, I'm also going to optimize my context window, my context use on its own.
It can do that without me telling it to. That makes me think of something that's really interesting. because of that ability for it to be more creative and arguably a better problem solver because it knows how to get to that end goal, that might also be kind of a scary thing.
Because in some cases, maybe the skill keeps it on stricter guardrails. Potentially. Because sometimes now if it has to try to find its own way from point A to point B, it might try to create some things that could potentially be harmful to the system.
something like that as well. But that's where the rest of the AIOS comes in. Because like the AIOS is not just a skill.
You have the rules, you have different layers to babysit. So we might live in a world where we need way less skills and much stricter rules. And your CloudMD now just basically says, go and follow these rules for these kinds of situations.
And that could be sufficient. Yeah. Yeah.
So some of these layers, even the concept of assigning an agent. and creating your own agent that's all spun up the same way every time, that could become obsolete. So I'm not saying it is, but I'm saying we should be open -minded that this stuff will evolve.
And as long as you keep your core assets nimble, then you should be good to go. I love that. You know, I did see, which I thought was really interesting.
I don't remember if it was Opus 5 or Fable 5, but one of those drops from Anthropic, you know how they drop the blog with benchmarks and just like how to use it, how to prompt it. And one of the things I remember being at the top.
was give it more ambitious tasks. And you might just skim through that, right? But to me, that made me think, well, why do they feel the need to put that in there?
Maybe they feel like people aren't pushing the models to their true limits and getting out of the model's way. And that's when I started to run these just interesting experiments. I'd set a slash goal and I'd say like, kind of something along the lines of like impress me, like build me this, but impress me and show me what you can really do.
And I just thought that that was interesting that they put that in there. And it just shows like that kind of makes me think that whole skill thing, like the skill that we maybe have baked in a long time ago. And now we just kind of blindly use because it works.
We are almost, yeah, kind of baking in or limiting what the model could truly do because it's kind of going down these guardrails. So I think that's really interesting. You know, as we sort of start to wrap up here today, I'd be interested to hear from you.
What is something that you think that you do? within your cloud code or your harness setup that most people don't do and that you think is important that everyone should be doing? So the main thing I think about is many times I have let's say 25 different operating systems.
So some people like to create one big one. I create very specific ones that live in an isolated world. Okay.
So my tax and finance lives very differently from my consulting OS, which is very living very differently from my school OS for content for there versus everything else. So me segmenting every single part of my business, my education offers, everything we do for enterprise clients, each thing has its own set of operating systems.
It's more to upkeep for sure. But as you build more, you start to build more leverage. So I have one mega system that depending on its audit of every single operating system I have will come up with a series of things that we need to make that specific thing better, that specific operating system work better, and one that generalizes across all of them.
So as you create different folders, you'll have some project level skills, rules, CloudMD, and then global. Some people make everything global, which is awful because as you add new things, you might notice underperformance because you forgot that you have this global specter of rules. applying to every single thing.
So for me, I have a different concept where I think about promotion. Everything is project until it deserves to be promoted to global so that I know at all times what is the running total of everything that's running globally. And I have a very few number of things that are global.
Everything's project specific. But that also gives you this flexibility to have a very small blast radius if I want to be experimental, if I want to audit. and push a certain folder for a certain type of task and not have that bleed over to another project and not know why is this not working?
Is the model worse? Is the harness worse? Or is my setup worse?
So because I'm so meticulous about is this a model problem? Is this a harness problem? Or is this an organization problem?
When I can isolate things, I can find the issue faster. So Fable is amazing with my tax, but all of a sudden, It is horrific with my consulting operating system.
It might not be a model issue. It might not be a skill issue. It could be an organization issue for that model that I might have to adhere to for this specific project.
And that helps me control bloat and find the area of resistance that I need to move to actually get the reforms I'm looking for. I think that's really smart. I think especially because these things are so, so autonomous and agentic.
You have to be able to... find the actual variable that killed the thing. Otherwise, there is no learning there.
There's no improvement there. For the most part, it can help you find out why it went wrong and where, but I think that level of isolation, and I think the key is there that you know how your different folders are drilled down. You're not just blindly trusting that it can find everything.
You still intuitively, I have a feeling that if you are looking for a specific deliverable, you could find it just by clicking through your files and folders because you kind of know where the things live and you know how to drill down. Yes.
And I think that that's just how that's the point again, that you cannot outsource your understanding. You still have to understand where everything is and how it works. Yeah.
The last thing I was going to say is that this becomes an acquired skill. And if you want to get to mastery, mastery is just understanding how the entire factory works end to end and where each piece lies. So if you want more leverage, if you're running into issues where you're kind of being intellectually lazy and saying, oh, this model sucks.
Before you pass fully judgment on it, you want to make sure that maybe it sucks, comma, for the setup that you currently have. So when in doubt, audit your setup, audit your skills, audit the bloat, because different models...
The Hook
The bait, then the rug-pull.
The cold open gives away the whole argument before the handshake: the harness of a model matters more than the model itself. What follows is a 34-minute attempt to prove it, starting with a drawing of a brain in a jar and ending with twenty-five separate operating systems.
Frameworks
Named ideas worth stealing.
02:07model
Brain in a jar vs. the harness
The model (reasoning only)
Read
Write and edit files
Bash and computer control
Modular add-ons: skills, plugins, context files
The model is a brain that can only receive stimuli and emit a response. Everything that lets it touch the world is bolted on by the harness, and modular additions like skills stack on top of that base layer.
Steal forexplaining to a non-technical buyer why two products running the same model behave completely differently
06:14concept
The agentic loop
Ask for a thing
Thing gets executed
Result returns as success or error
Result becomes the next stimulus
A model's usefulness is its ability to keep running this loop well, and that ability is determined by the harness rather than raw intelligence.
Steal fordebugging why an agent stalls partway through a multi-step task
11:16concept
Provider-agnostic asset strategy
Keep every skill, rule and context file in one location and make each one work with any model, open or closed. The brain is leased and swappable; the assets are owned. Switching providers should take about a day.
Steal forany tooling decision where you are worried about vendor lock-in
22:41model
The rot.md decay layers
Identity: one to three months or longer
Substrate / core context: slow
Hooks: around six months
Rules: weekly, sometimes daily
Skills and agents: fastest of all
Each layer of a personal AI operating system goes obsolete at its own rate, so maintenance is scheduled per layer instead of all at once.
Steal forscheduling a review cadence for any living documentation system
31:31concept
Project-first promotion rule
Nothing starts global. Every skill, rule and context file lives at project level until it proves itself and earns promotion, which keeps the global set small and the blast radius of an experiment contained.
Steal forany configuration system where defaults quietly accumulate
31:59list
Model, harness, or organization
Is this a model problem?
Is this a harness problem?
Is this an organization problem?
When the same model performs well on one project and badly on another, the variable is rarely the model. Isolating setups is what makes the real cause findable.
Steal fortriaging any flaky automated system before swapping components at random
CTA Breakdown
How they asked for the click.
VERBAL ASK
10:16link
“So try Clay using the link in the description, and you'll get 2,000 free credits. Now let's get back to the video.”
The only hard ask in the video, and it is the sponsor's. It runs at 9:26, right after the local-model demo lands, and is framed as a tool the host actually drives from Claude Code via the Clay CLI. The host's own offers sit in the description rather than the script, so the body of the video carries no self-pitch at all.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
A 14-minute demystification of agent loops for non-hardcore-coders: what they are, why the done-check matters most, and three live demos that prove loops get you closer — not perfect.
A 14-minute tour of Printing Press — a CLI factory that turns any site (even ones without an API) into a token-efficient command-line tool your agent can call.
A screen-share walkthrough of all 18 building blocks of OpenAI's Codex, from project folders to voice-controlled sub-agents, aimed squarely at people who don't write code.
A five-hour, sixteen-module course that takes a non-coder from installing the Codex desktop app to running a second brain, shipping skills, hosting automations on Trigger.dev, and pricing the work at 10% of the value it creates.
A screen-share walkthrough of using Codex to plan and build automations, then hosting the generated code on Trigger.dev instead of burning your weekly usage limit.