Modern Creator
Mark Kashef · YouTube

These Mods Take Claude Code to Another Level

How reverse engineering Codex via MCP and automated sub-agent routing turns Claude Code into an autonomous desktop operator.

Posted
yesterday
Duration
Format
Tutorial
technical
Views
8.7K
165 likes
Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • You use Claude Code or desktop LLMs and need background GUI automation without losing mouse control.
  • You want to coordinate multi-model agent teams where a lead model delegates subtasks to Haiku or Sonnet sessions.
  • You are building custom MCP tools and plugins to extend developer workflows.
SKIP IF…
  • You only use basic chat web interfaces without local CLI or desktop tool access.
  • You have no need for local macOS system automation or multi-session agent orchestration.
Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:00 – 01:22

01 · Claude uses Codex to control Mac apps

Demonstration of Claude Code executing a slash command to take over macOS Calculator and TextEdit via Codex's background tools.

01:22 – 01:57

02 · What Claude Code mods let you change

Explains how mods allow users to add tools, tailor task execution, and build custom workflow buttons into Claude Code.

01:57 – 02:41

03 · Claude decides, Codex clicks

Details the cognitive-motor split where Claude acts as the decision-making brain while Codex provides the execution limbs.

02:41 – 03:53

04 · Connecting the tools through MCP

How local desktop tools can be surfaced via MCP, and why reverse engineering required logging Codex's background activity.

03:53 – 06:54

05 · The computer-use build prompt

Step-by-step review of the multi-section prompt used to build the launcher, persistent helper, and permission handling.

06:54 – 07:21

06 · What it took to get the mod working

Reflecting on overcoming Claude's initial refusal by using iterative testing and slash-goal directives.

07:21 – 10:07

07 · Threads: one chat runs a team

Introduction to the Threads mod, enabling a single lead chat to launch and coordinate parallel sessions across different models.

10:07 – 10:45

08 · How Threads works behind the scenes

Explains the underlying UI inspection, tmux session spawning, and live transcript monitoring.

10:45 – 13:41

09 · The Threads build prompt

Dissecting the prompt architecture that defines worker roles, model assignment (Opus, Sonnet, Haiku), and UI status panels.

13:41 – 14:40

10 · Testing each piece with a goal

The strategy of incremental integration testing from one worker session to complex swarms.

14:40 – 15:40

11 · Build the features you wish existed

Final call to action, resource availability on GitHub and Gumroad, and community invitation.

Atomic Insights

Lines worth screenshotting.

  • Frontier model desktop clients often run local MCP daemons whose tools can be intercepted and borrowed by competing models.
  • Bypassing repetitive permission prompts in agentic computer use requires a persistent local helper client that maintains session state and auto-approves calls.
  • When an LLM claims a system integration is impossible, forcing it to inspect installed tools and run live diagnostic audits will frequently reveal viable integration hooks.
  • Using a slash-goal command with explicit progressive unit tests forces the model through complex architectural debugging that single-shot prompts fail to solve.
  • Multi-agent delegation works best with tiered model intelligence: Opus for high-level architectural planning, Sonnet for code execution, and Haiku for linting and review.
  • Running sub-agents in background terminal multiplexers like tmux isolates session memory while keeping the primary workspace uncluttered.
  • Reverse engineering desktop software starts by instructing an LLM to monitor background logs and socket traffic while the user triggers the target feature manually.
  • Exposing local computer-use tools over an existing desktop client subscription avoids costly agent API metered token charges.
  • Session isolation in multi-agent systems prevents new tasks from inheriting stale state or competing for foreground focus.
Takeaway

Local AI models can bridge each other's native strengths through background protocols

MOD ARCHITECTURE

You can overcome client feature limitations by forcing LLMs to inspect background daemons and write persistent local helpers.

  • Reverse engineer native app capabilities by running diagnostic prompts while monitoring background client logs and socket ports.
  • Build a persistent helper daemon alongside your MCP launcher to prevent agent workflows from hanging on repetitive permission prompts.
  • Structure multi-tier agent swarms with a coordinating lead model and dedicated lightweight models assigned to sub-tasks.
  • Never accept an LLM's assertion that a system capability is impossible without prompting it to inspect installed tool configurations.
  • Test complex custom mods incrementally by writing unit test harnesses that validate single worker execution before scaling.
Glossary

Terms worth knowing.

Claude Code Mod
A custom plugin or configuration add-on extending the Claude Code CLI and desktop environment with new commands, hooks, and external tools.
Model Context Protocol (MCP)
An open standard enabling LLM applications to securely expose and consume local context, tools, and background server utilities.
Computer Use
An agent capability allowing an AI model to interpret screen pixels, simulate mouse clicks, and input keystrokes into native operating system applications.
Persistent Local Helper
A lightweight background script acting as an MCP client that manages continuous daemon connections and handles authorization handshakes without user intervention.
Lead Thread Architecture
A pattern where an orchestrator model manages the lifecycle, instructions, and aggregated outputs of independent worker sub-agents.
Resources

Things they pointed at.

00:14toolCodex ↗
08:41tooltmux
Quotables

Lines you could clip.

02:21
“Claude essentially became the brain with access to the limbs of codex.”
Vivid conceptual summary of the hybrid multi-model architecture.→ newsletter pull-quote↗ Tweet quote
04:22
“one annoying thing about Claude is many times it will tell you that a lot of these hacky connections aren't possible from the get-go”
Addresses a widespread frustration among developers building agentic workflows.→ TikTok hook↗ Tweet quote
11:47
“Before deciding a feature is impossible, inspect the installed mod authoring API, the current CLI for help, and distinguish a missing tool in this chat from a capability that can be implemented.”
Actionable, rigorous meta-prompting instruction for agent tooling.→ IG reel cold open↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphoranalogy
All right, so watch this. I'm going to show you a very interesting example of using mods to their fullest potential. I'm going to use Cloud Code to use Codex's computer use to execute a task.
All I have to do is I'm going to do slash codex dash cu, and I'm going to say on. Now this should be activated, and then I'll send over this prompt right here. What this prompt will do will make Claude code ignore its own internal browser and instead find and use a bridge to use codex as computer use to execute this task.
So you can see right here, it's starting to use the computer use module and you can see at the very bottom it's activated and we should see it open up the calculator app. There you go. Manipulate it, take notes of exactly what it calculates.
In this case, some hypothetical subscription and come up with a result.
All right, so we sped it up, but you can see here it was able to manipulate the calculator, write its findings in a text document, bold it and format it so we can read it. So in this video, I'm going to walk you through how I put this together and how you can too. And I've got another mod to show you that I think you're really going to like, especially if you've ever used something like Claude and thought, why can't I just do this?
Because now you can build a lot of these missing features yourself. So if you want to take your mod game to the next level, then let's dive in. So in my last video, I walked you through what mods are and how they work.
But as a quick refresh, you can think of a mod as an extra add -on that allows you to tailor and configure how cloud code works for you. You can add a tool, change how it handles certain tasks, or give yourself a button of some sort for something that you always do all the time and you want to create a shortcut for. So I want to take my basic mods and go a step further.
Could I use mods to actually add features that I always wish Claude had, or be able to emulate features that I really like in Codex that I've been patiently waiting for the Anthropic team to put together? Let me show you how I approached the computer use one. So when I sent Claude that prompt, it had to decide which button to press.
Based on the goal it extracted from the prompt, it realizes that it needed to use a mod. And once it realized which mod to use, it then passed on this information to use the set of tools from Codex. And we basically made these tools available at the discretion of Claude.
So instead of having to use codex models to orchestrate codex functions, we could use Claude to orchestrate those same tools. So in this case, when we told it to use the calculator, Claude essentially became the brain with access to the limbs of codex. And every single time it clicked or made an action, all of that information, that visual confirmation, would propagate back to Claude that would then make the next decision as to where it should click next.
Now you might be wondering how it's possible for us to access the tools that should be only designed for the Codex app. The main thing to understand is that a lot of these tools can be exposed as an MCP and pretty much any LLM, not just Claude, can access them if you build a proper bridge to enable that connection. Now I'm going to break down phase by phase exactly what I asked Claude to do, but essentially I asked it to watch Codex open up and use its computer use and basically listen in for all the services it was using.
And I did this to try to reverse engineer how it could possibly access these controls. Once it figured it out, it created a script that basically is called a launcher. And this launcher allows us to launch into the next step.
Now the next obstacle is if we're going to rent a tool from another service, one thing that pops up are permissions. So Codex would constantly ask, can I use this tool? Can I click here?
And we needed a way that Claude wouldn't have to deal with this over and over again. So after pushing it a bit, it created a little program that allows it to maintain a live connection to a codec session. And this bridge, otherwise known as a helper, doesn't only allow the connection to stay open, but handles all the permission layers behind the scenes.
So anytime it asks for permission, it automatically gets an approval from that session. Alright, so with that out of the way, here's the complete build prompt. So the first part of the prompt is obviously breaking down exactly what the North star is.
And in our case, I said very clearly, I want a cloud code mod that makes you use Codex as computer use instead of your own. Whenever I ask you to do something on my Mac, like use a calculator or write a note, Codex can click and type inside my apps in the background without taking over my mouse. Now, one annoying thing about Claude is many times it will tell you that a lot of these hacky connections aren't possible from the get -go until you probe it and you ask it, well, I'm going to open up the Codex app and I want you to watch me run a computer use and I want you to listen in for all the tools that are being run behind the scenes.
Maybe there's a way for us to access them. Usually when you have it run it and actually check, it'll realize that it's wrong and it is possible. And that's the thread you want to pull on to reverse engineer the rest of the connection.
So we go from explaining the goal to over explaining how to connect to Codex. So you'll see here as an insurance policy to make sure every single time we update the Codex app, that this doesn't break. We always want to connect it to the latest and greatest version of that app.
Over here, this is where we're mentioning the launcher. And this is what's connecting to the MCP server. And this is the name of that specific MCP server that gives you access to computer use.
And like we mentioned, we had a bridge between Codex and Claude. And this is the instruction for that bridge. Build a persistent local helper that acts as an MCP client.
to that server, it must initialize the connection, handle tool calls and app approval requests, et cetera. Now, because this is experimental, we wanna make sure that the first time it uses a tool that asks us permission. Once we tell it you're good to go, we don't wanna be asked again.
So in this case, we just say, make it ask me before it uses an app for the first time, give me three choices, just like you can see here, allow for this chat, always allow or don't allow. Now maybe sometimes I still want to use Claude's internal browser because it's an easier way for it to check its own work for a variety of things like going through an app that it put together or a webpage it just needs to scroll through.
So I just want to make a way that I can turn it on and turn it off very easily, but by default have it on for every single chat. And one extra note here, instead of me having to use the agents API and pay extra, this allows us to use an existing subscription to the existing Codex app and not have any extra costs in between.
Now this next part is more technical, so I'll let you read the prompt as I'm obviously gonna make this available to you as well as the prompt for the next thing I'm about to show you down the second link in the description. But the TLDR is I wanted to make sure that if I use this tool in an existing chat and then I created a new chat that I wouldn't be stuck using that tool and having it focused on servicing that prior chat.
Then obviously to make sure it works, we have to test it and using a calculator is probably one of the easiest things you could do to make sure it's working end to end. Part seven essentially asks Claude to create a series of troubleshooting steps in case, for whatever reason, the computer use tool from Codex misfires. In number eight, I had AI help me with quite a bit, which is telling it exactly where these files should go so I wouldn't have a very bloated mod library.
Now, realistically, did I take this one prompt and have it one shot it and work on the first time? Absolutely not. But after a few tries, I gave it a slash goal.
I told it exactly what to troubleshoot based on all the errors it was having. And eventually, maybe after an hour, it finally got it. Now, obviously for you, I'm not going to make you go down the same rabbit hole if you want to emulate this exact same mod.
So I'm going to make available also my GitHub repo with all the underlying code so you can take it, plug and play, retrofit it, and do as you wish. Now, the other mod I wanted to show you and walk you through is called Threads. I know I'm not in Claude, I'm in Codex and I'm here for a reason because I was always jealous and waiting for a point where Claude could do this exact same function that I'm about to show you that I've shown in prior videos before.
Hey, so I want you to spin up a series of different threads with different models, depending on the level of intelligence we need for different parts of this task. But I want to do a full research sweep on what are the best models to combine together. Should it be Opus 5 .5?
and GPT -6 Astra, what effort levels? Should it just be GPT -6 models? Should it just be Opus models or anything like Sonnet 5 .5 or Opus 5 .5?
I want you to spin up a series of threads, have one lead thread and have that lead thread organize that sub threads to make sure they're always reporting back and we have one cohesive answer at the very end. Absolute mouthful there, but if I send this over, the beautiful feature that it has is it can spin up all of these threads, name them, run them in parallel, have a thread that's overseeing the rest of the threads, and this becomes a very helpful functionality at scale.
And boom, just like that, you have five threads on the go, all working together, and this is not something you can do, at least in the Cloud desktop app. You can MacGyver the terminal to get to some level of behavior, and obviously you can use things like Herder if you are more technically inclined. But if you just wanted to use a desktop app, there must be a better way.
If you go spin up a brand new chat on cloud desktop and ask it something like, can you spin up threads in the desktop app with different models? If I ask, and in this case, I'm telling it to not use the mod that I've already put together. It says it can't do that.
And the best it could do is run sub agents on a chosen model. But since there's a mod for this, we could send the exact same prompt and I'll say, make sure to leverage the thread mod and it should be able to start spinning up. the exact same behavior in a different way from the way Codex does it.
And there we go. It starts using the threads creation portion of the mod and they should pop up at the bottom left hand side right here. You can see right here, it's beginning the thread process.
It pops up a pane on the right hand side where we can keep track of exactly what is happening, what model is the lead and what are the other threads being spun up and exactly how much they're spending in the process. And one super interesting thing is you can add your own, like I said, buttons where if I want to steer the conversation, if I want to steer the task for all the different threads, I can do so from one centralized place in this pane.
And there we go. We were able to spin up all these different threads and emulate the exact same experience from Codex. And in many ways, upgrade that experience.
Cause again, we can manage all of it from here. We have full transparency. We have all the transcripts and we could add a lot more richness and complexity to this initial feature.
And when they're done, the unified work ends up in this main chat. Now again, big picture. Trying to explain this feature in plain words, even by having AI metaprompt for me, was a bit tricky.
So what I had to do was very similar to what I did with computer use. So I told Claude, I was about to open the Codex app and create multiple threads from one conversation. And I wanted it to audit all the logs from Codex to see exactly what it was doing, how it was deciding to create those threads and how it created it, most importantly for me, in the Codex app itself, not just some terminal.
Typically, it's a lot easier to create this behavior in a terminal experience, but if you want it on a main UI desktop app, you have to MacGyver it a little. So here's the prompt that I sent over and I'll go over the most pertinent parts. So in layman's terms, I told it I want to build a Cloud Code mod called Threads.
Before you write anything, load the plugin authoring skill and basically remind yourself how mods work. So go and check your documentation so you fully context prime yourself before we take the next step forward. Then I break down the concept.
I want this chat to be the boss of a small team. When I say something like start a Haiku helper to list my files and an Opus helper to review my readme, the mod should start each helper in its own real Cloud Code session on the model I asked for. And ideally, if I tell it to use its judgment, similar to what I showed you in Codex, it should be able to do that too.
Now at this point, I was already burnt by Claude, so I explicitly told it, before deciding a feature is impossible, inspect the installed mod authoring API, the current CLI for help, and distinguish a missing tool in this chat from a capability that can be implemented. So I started pushing on it, and I started having it break down, how does the Codex app work?
And how does the Claude code app work? And then it told me, there happens to be tools that manage the sidebar. There are tools that manage the threads and the chats.
So then by allowing it to recognize what tools it has access to and giving it the goal and actually doing a slash goal, I kept pushing it and pushing it to realize that this was possible. It just has to be a lot more creative. Now, so much of the first mod, we also needed the concept of helper here as well.
So in this case, I said, each helper is a normal cloud code session. And if you're confused on what models you can use, go and literally use Claude dash dash help. Now in the terminal, this just allows you.
to troubleshoot all the core functionality of Claude, which it could use itself to remind itself. And then we spoon fed it what to look for, which is setting the model, giving the session a name, showing it in the Claude desktop sidebar, in which case we tell it, go and look at the different tools you can use to manipulate the sidebar, skip the permission prompts, because initially it would create a brand new session, then ask you to go into that session, click enter to get permission to execute that prompt.
And then once you got that working, we went to the next level of granularity, which is effort level and adding extra instructions. Now the inspiration for this pain here didn't just come out by itself. We actually told it to see the progress of all the threads.
We needed to always have a way to keep track of are any of these threads stuck? And if so, can you keep an eye on it? So then we were able to give it an idea on how to create and how to log the API equivalent cost of running all these extra sessions.
Now these next sections here go into the devil in the details to explain how to guide different types of work, in what order, and exactly what this might look like in practice. So in this case, you might make a plan with Opus. You want to spend your tokens wisely there.
Then it would leave some notes. Those notes would then go to, potentially, a Sonnet helper. And then once Sonnet is done, it might pass it off for a final check or a final result by Haiku to take an extra scan.
Now, one really important thing for very complicated mods is giving it an explicit set of different things to test. So ideally I gave it a what's called slash. And basically this does burn tons of tokens, but sometimes when Claude is refusing or denying that something's possible and you throw it a slash goal, in this case, you make it something specific, which is, okay, you have this feature.
I want you to battle test this. And now I want you to start creating one thread, sending it a message, sending it a step -by -step plan and watching it execute to the end. If that works, let's move on and create two threads and so on.
So giving it that slash goal. gives it an extra push to see what's possible and evaluate it. And one extra nuance here is I wanted to make sure that every single thread it spins up was something that was remote control eligible.
So if I want to walk away from my computer and manage those sub threads that I could, and even this was something that I had to specify after quite a bit of iteration. And obviously there's way more detail there, but I'll let you read the prompt for yourself and see what parts interest you the most. Now, where are those super easy to put together?
No. Is it still possible? And is there pretty much unlimited possibility with what you can do if you could just imagine exactly what the goal is, how you want to accomplish it, and most importantly, how you can get something like Claude to test its own work?
Yes, and all you have to do is really take something like all of my prompts from my last video, this video, the code, and even feeding all that code gives something like Claude and Opus 5 .5 enough inspiration to take on the next big hairy task that you want. And that's pretty much it. So if you wanted to grab all the resources for these two mods that I showed you, again, they'll be down in the second link.
in the description below. And as always, if you want to take your learning to the next level and go much deeper on topics like Moz to fully understand how the system works and how you can apply it, check the first link down below for my early adopters community. For the rest of you, if you found this helpful and novel, I'd super appreciate a like and a comment down below.
Really helps the video and the channel. And I'll see you all in the next one.
The Hook

The bait, then the rug-pull.

The video opens immediately with a live software demonstration showing Claude Code executing actions on Mac desktop applications via Codex's backend without moving the user's cursor.

Frameworks

Named ideas worth stealing.

02:15model

The Brain and Limbs Model

  1. Orchestrator Brain (Claude Code)
  2. Local Bridge Daemon (Persistent Helper)
  3. Execution Limbs (Codex Computer Use)
  4. Visual Confirmation Feedback Loop

A design pattern separating high-level cognitive decision-making from operating system execution tools hosted by separate AI clients.

Steal forarchitecting hybrid AI agent systems without vendor lock-in
13:27framework

Tiered Multi-Model Team Routing

  1. Opus: Architectural planning and high-level reasoning
  2. Sonnet: Main execution, code writing, and tool implementation
  3. Haiku: Rapid output validation, review, and file listing

Allocating different tiers of model intelligence and token cost to corresponding levels of cognitive difficulty in an agent swarm.

Steal fortoken-efficient multi-agent build pipelines
CTA Breakdown

How they asked for the click.

VERBAL ASK
15:13product
“If you wanted to grab all the resources for these two mods that I showed you, again, they'll be down in the second link in the description below.”

Directs viewers to free GitHub/Gumroad prompt templates followed by a pitch for his paid Skool community.

MENTIONED ON CAMERA
FROM THE DESCRIPTION
PRIMARY CTAWhere the creator wants you to go next.
Storyboard

Visual structure at a glance.

Hook: Live Demo
hookHook: Live Demo00:00
Automated Calculator Control
proofAutomated Calculator Control00:43
Brain & Limbs Diagram
conceptBrain & Limbs Diagram02:26
Computer Use Build Prompt
breakdownComputer Use Build Prompt03:58
Internal Browser Toggle
featureInternal Browser Toggle05:58
Threads Multi-Model UI
demoThreads Multi-Model UI08:07
Thread Control & Cost Panel
featureThread Control & Cost Panel09:41
Threads Prompt Architecture
breakdownThreads Prompt Architecture11:27
Step-by-Step Testing Strategy
frameworkStep-by-Step Testing Strategy13:36
Community Call to Action
ctaCommunity Call to Action15:22
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.