How to Actually Build & Sell Software with AI as a Non-Techie
Veteran engineer Dave Ebbelaar breaks down why he no longer writes code by hand, how to climb the four-rung software ladder safely, and why full-stack generalists now outpace hyper-specialized teams.
You are a non-technical builder looking to turn app concepts into production-ready software using AI agents.
You are an experienced developer trying to understand how autonomous coding tools like Claude Code and Codex alter modern architecture.
You want to build and sell automations or SaaS products without risking fatal security breaches or unmanageable code rot.
SKIP IF…
You are looking for beginner command-line installation tutorials for coding tools.
You work in a heavily regulated enterprise where autonomous AI coding agents are strictly forbidden.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Dave Ebbelaar reveals that after ten years of professional Python programming, he no longer writes a single line of code manually, relying entirely on AI agents.
01:28 – 08:19
02 · A Solid 10X & The Rise of the Generalist
Dave analyzes why full-stack generalists now dominate software creation, contrasting historical specializations with modern AI-driven output multipliers.
08:19 – 09:32
03 · Sponsor: CodeRabbit
Nate Herk presents CodeRabbit's Change Stack interface for managing rapid multi-file pull requests generated by AI agents.
09:32 – 12:42
04 · The Four-Rung Software Ladder
Dave outlines his four-rung framework, explaining how user surface area and data liability scale from personal scripts up to public SaaS.
12:42 – 21:04
05 · Building with Agents: Intent Over Specs
The conversation explores why rigid spec-driven planning has been superseded by intent-driven building directly on code branches using frontier models.
21:04 – 21:40
06 · First Client SOP Resource
Nate highlights his free standard operating procedure guide for landing initial paid AI automation clients.
21:40 – 29:40
07 · Automated Evals & LLM as a Judge
Dave explains how to benchmark subjective AI outputs by creating alignment loops between human reviews and judge LLM prompts.
Dave details multi-tier architecture design, avoiding spaghetti code via Matt Pocock's skills, and implementing core security measures like Supabase RLS and firewalls.
42:42 – 1:09:33
09 · The Infinite Genie & Outro
Dave shares his perspective on the unprecedented entrepreneurial freedom unlocked by AI agents and explains where builders can find his tutorials.
Atomic Insights
Lines worth screenshotting.
Coding agents have flipped software engineering from deep language specialization to full-stack generalization because handoffs between siloed developers now create bottlenecks that move slower than AI agents.
A senior developer working with modern coding agents experiences at least a 10x speed multiplier in core languages, but near-infinite leverage when expanding into previously untouched languages like JavaScript.
The meta of AI software construction has shifted away from exhaustive upfront spec-driven documentation toward direct, intent-driven prototyping directly in production git branches.
Risk in software follows a four-rung ladder: personal tools carry zero liability, internal team tools require basic data hygiene, client B2B software introduces contractual obligations, and consumer SaaS exposes founders to catastrophic data breach liabilities.
Vibe coding without reading code is viable for personal tools, but unmanaged feature churn inevitably produces spaghetti codebases with deeply buried dependencies that break unexpectedly at scale.
Objective evals can run against latency benchmarks and unit test datasets, but subjective outputs require an LLM-as-a-judge loop calibrated against human scoring to achieve 90%+ alignment.
Unless a founder possesses an overwhelming technical reason to use an exotic database, cloud-hosted Supabase remains the default choice due to built-in auth, vector capabilities, and relational flexibility.
In AI-generated software architectures, row-level security in PostgreSQL and IP whitelisting firewalls between frontend and backend services prevent bots from exploiting leaked credentials.
Coding agents operate as an infinite genie in a bottle, where the sole operational bottleneck for solo entrepreneurs has transitioned from technical talent to weekly API token budgets.
Takeaway
Architecture and data hygiene determine whether AI-generated software scales or crashes
THE BUILDER BLUEPRINT
Autonomous agents remove technical barriers to code generation, shifting the developer's role from writing syntax to designing robust system boundaries.
Stop obsessing over syntax memorization; your leverage comes from understanding system architecture, databases, and how components communicate.
Match your technical scrutiny to your software's rung on the ladder: prototype freely for yourself, but enforce strict boundaries for external users.
Avoid rigid specification phases when frontier models like Claude Code can generate running prototypes directly from prompt-driven intent.
Address code rot proactively by enforcing modular codebase design so rapid experimentation does not create invisible cascading failures.
Calibrate subjective agent performance by training an LLM judge on human-reviewed test sets before automating feedback loops.
Enforce row-level security on your database and whitelist IP communication between frontend and backend services to neutralize leaked credentials.
Glossary
Terms worth knowing.
Vibe Coding
The practice of building software purely through conversational prompts with AI models without reading, writing, or manually validating underlying source code.
Intent-Driven Development
A development methodology where a human builder articulates the end goal and behavioral constraints in natural language, delegating code synthesis and system scaffolding entirely to autonomous agents.
The Four-Rung Ladder
A framework categorizing software by risk and complexity: Level 1 (personal tool), Level 2 (team internal tool), Level 3 (B2B client project), and Level 4 (public consumer SaaS).
LLM as a Judge
An architectural pattern where a secondary language model evaluates, grades, and validates the outputs of a primary AI agent against an aligned rubric of quality.
Row-Level Security (RLS)
A database security feature that restricts database query results down to specific rows based on the requesting user's identity and security attributes.
AI Slop
Low-quality, bloated, or fragmented code generated iteratively by AI models that lacks modular structure and introduces hidden regression bugs.
Radical, counter-intuitive statement from a ten-year veteran software engineer that immediately hooks viewers.→ TikTok hook↗ Tweet quote
03:15
“coding was really freaking hard.”
Relatable reflection validating beginner frustration while highlighting modern leverage.→ newsletter pull-quote↗ Tweet quote
11:26
“Now your AI agents can move way faster than you can communicate and can keep up with teammates.”
Sharp analytical explanation for why solo generalists now outperform large developer teams.→ IG reel cold open↗ Tweet quote
1:01:24
“like we as developers now have a genie in a bottle.”
Evocative quote encapsulating the transformative shift of prompt-based software creation.→ TikTok hook↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphoranalogy
so you've been coding in python for over 10 years now how different do you view the software engineering space now that we have all these ai tools i don't write a single line of code anymore not a single so everything goes through codex or for cloud code think about any skill you would learn you're not going to start at the elite level but right now we all start with these same really great tools that can help us produce code if you had to put a number to it how much faster do you think you're able to move now well compared to what when I was just starting out, probably 100x.
You can build useful stuff with one prompt. Heck, you can even build entire software products that you could sell even if you have a limited technical background. Do you have any examples of maybe landmines that you've stepped on when trying to scale something?
The stakes get bigger and bigger and bigger as you start to build software that's used by more and more people. There are certain obligations when it comes to data privacy and protection. Security breach, that's one thing that you cannot take back.
Once those emails and that personal data is leaked, you can't go back. If I was starting today with... technical experience, I would be most excited about this.
All right, Dave, so you've been coding in Python for over 10 years now. Your team has delivered 50 plus AI solutions with Glido. You are shipping new features and updates like every week.
I just want to know, how different do you view the software engineering space now that we have all these AI tools? Oh man, what a question to kick this off with. Well, obviously it's completely different.
It's completely different. What happened over the past three years and especially what happened in the last year is just wild. It's just crazy in terms of what we can do right now.
And we're all experiencing this, right? We're all using these tools. We're all producing lots of code.
We're building. And so first of all, it's like really exciting. So if you ask me like what is mainly different from when I started out, back in 2013, so like over 10 years ago already, well, first of all, coding was really freaking hard.
And learning to code as well. And you would just spend so much time working on this tidy little program, looking up Stack Overflow, and then trying to make stuff work, right? So you could build small little things.
And then if you worked on something long enough, you could build something useful, maybe. right now fast forward in the world that we live in today you can build useful stuff with one prompt and we all know that we see it on your channel like every week just one prompt and you you can build stuff so what that then also means is that we can now just think bigger as developers as builders and we can now do stuff and build stuff that just wasn't realistic or possible a couple of years ago so that means we can build automations for yourself for personal productivity for your company inside the company for clients heck you can even build entire software products that you could sell even if you have limited a limited technical background that's now all possible and it's getting faster every week if you had to put a number to it how much faster do you think you're able to move now Well, compared to when I was just starting out, probably 100x, but that's not fair because I was a complete noob.
But let's say, let's do a little bit more of a realistic take. Let's say we go back to 2022. At that time, I was already working as a freelance data scientist, data analyst for like two or three years at that point.
So I was professionally writing code, mostly Python code. almost every day for already three years okay if i look at that point right now and see like how much i could ship and produce and do and all of that compared to what i can do right now i would i would still put it like a solid 10x like a solid 10x of actual useful like production ready code that you can ship and on some projects it's maybe even more because i can now also work on projects and work with languages that before were just not possible so i come from a python background data scientist later ai engineering so i used to make back -end systems but if you were to ask me to build like a shiny like front -end or dashboard i have like i have never written a single line of javascript like actually in my life pretty much i only started to pick that up once the coding ages but i was like yeah like let's just try it right so that is i think interesting like like a solid 10x on on the work that i already did professionally but like probably even more than that on other like possibilities that i just couldn't do before gotcha so before kind of pre -ai engineers would kind of typically specialize in either like more of the back end at a very high level like a back end or a front end and it was a very different way you thought about writing that code oh yeah for sure and yeah of course you have full stack engineers as well but then
Even with full stack, it was probably focused on specific languages. And back then, you also had these 10x engineers, right? The old legends of these.
And you have these crazy developers that could do it all. But generally speaking, mostly people would either do backend, frontend, security, infrastructure. They would have one or two languages that they're comfortable with.
And that is now... That is now completely flipped on its head. So now, instead of really having to specialize, everyone that's building can sort of be a jack of all trades there, more of a generalist and feel confident that they are kind of able to write high quality code in maybe some of the areas that they weren't super proficient in beforehand.
Yes, for sure. And I also would highly recommend... everyone to try and do that because why is that the case if you don't if you stick to specialization where of course like there's always value in specialization right so maybe people watching if you're working in an enterprise let's say you're working for meta or google or whatever and you're part of like a highly specialized team you need to be top one percent of what you're doing right you need to be a hyper specialist but let's be real Most people don't work in the top 1 % at Meta, Google and all of that.
Most of us are like we're building, we like to build stuff. And there you can better nowadays be a generalist because if you specialize, means you focus on one part of the project, let's say, means you also need other people to deliver something end to end, to ship something. And now your AI agents can move way faster than you can communicate and can keep up with teammates.
So if you cannot work across the entire stack, you are quite quickly going to hit a bottleneck where your agent is already ready, right? It's ready to ship, it's ready to work on the next part. But if, for example, I don't know, there is not an integration with the front end because you only do backend or you don't know how to work with the databases and adjust your backend to the database, right?
There is a dependency. And that creates problems. And I've seen really in, first of all, the projects that I've worked on, the teams that I've worked with, but also now what we do in Glido and how we build Glido, that we have very, of course, we have separate roles and responsibilities and we have more people, but we give people complete ownership on a particular aspect of the product where they can really ship something from start to finish.
on their own without needing someone else other than like maybe a review from someone in the process, right? But not like actual work or code that needs to be developed. So I think being a generalist is the best way to go.
It's now possible. So you could just get better at engineering, learn engineering principles, and then which language or framework you apply that to starts to matter less and less. Real quick, guys, I just have to take a second to tell you about the sponsor of today's video, CodeRabbit.
So I know that a lot of you guys are building with coding agents now, and the code definitely piles up faster than you can actually read it because the agent doesn't just touch one file at a time. It goes through your routes, your schemas, your tests, and your config in a single pass. And then you open the pull request and GitHub gives you an alphabetical list of files.
So you're rebuilding the logic in your own head. So Change Stack is Code Rabbit's review interface for that exact problem. What it does is it splits the pull request into cohorts of related work and orders them by what depends on what.
So it reads it in the order that it was actually built. It opens on an overview page with the summary, the walkthrough, and what's blocking merge. Merge conflicts and failing CI checks you can fix from the page in one click.
There's also a timeline of every change, approval, and comment. so you can see why each decision got made. And the semantic diff view is the part that I would probably use the most, because when code gets moved, a normal diff shows it deleted on one side and added back on the other, burying the real change.
The semantic view shows you what actually changed, and you can ask the agent questions about the PR right on that page. So if you're shipping more agent code than you can read, there's a free 14 -day trial, no credit cards required. The link is in the description, and huge thanks to CodeRabbit for sponsoring this part of the video.
Now let's get back to it.
I'm really excited for this conversation because my audience is typically people that are coming from a non -technical background and it seems like your audience is like the exact opposite. And one, one thing that I hear from my audience, like when, you know, I'm building stuff and I always sometimes wonder, is this right?
You know, cause the only really source of truth that I can trust is my agent. And that's not always going to be the source of truth that you want to trust. And I'm curious, you know, I want to zoom out and I want to start from the very beginning, but.
Real quick before we transition into that, I'm just curious, how did you feel about the whole vibe coding era? Like when that term really started to pop up and you as a software engineer, there was probably so many alarms going off in your head when everyone was just picking up these tools and started building apps and stuff like that.
Yeah. So first of all, I found it very exciting because the thing is, I was also very much part of that whole vibe coders journey, even though I have. experience as an engineer like i've said i was now also creating front -end applications fully custom code websites and i wasn't reading the code and i couldn't care so like that would also put me in the category of like five coder and even though that i technically know how to like set up an architecture and how to secure your applications and all of that right that of course helps and it adds a little bit of um just extra experience to it but i'm very much pro pro vibe coding you just need to know what the stakes are depending on what you're working on right because if you're building so if you're building your personal website by all means like vibe away there's like zero risk right but it's always about like risk versus reward trade -off so i think it's awesome that everyone can build right now i can build more uh that's super fun anyone can build and then the question is you just need to start being careful like once you let other users other people use your
software and especially when personal data is involved right so that's that's i think where you need to draw kind of like a fine line in terms of like how far can you really just push things without really understanding what's going on and just going off of the looks or like checking whether something works and diving deeper into let's see if this actually is secure because now it's not your only your own data that you're working with but it's also other people's right so i think that is a Yeah, so that is kind of like my view on the whole vibe coding term or trend, really, that was going on.
I hear you. Yeah, I mean, it's very cool because we see all of these stories of people, you know, from whatever sort of background that just have this idea and then they start talking to an AI agent and they're able to actually run with it. And we've seen some pretty cool exits.
We've seen... People just change their lives because they've been able to take their idea and turn it into something. So let's actually like kind of walk through that process a little bit.
If someone has an idea and they want to sit down this weekend and they want to start building that out, how would you do that? What sort of steps and planning do you go through in order to start building some sort of app or software?
Great question. So let's walk through this. So first of all, you need to get a rough idea.
of how big the thing is that you want to build because when it comes to building products and let's focus on software products uh right now to make it a little bit simpler right um there's different levels different levels to things okay so first of all i would say ask yourself the question is this something is this a software tool that just you want to use like because you want to have some type of tool either something that you're currently paying for right now that you don't want to have a subscription for anymore or just something that doesn't exist or you want to combine multiple tools into one like is it something that you just want to use so so that that's like one question because another level could be uh for example yes it's for me but it's also for example for people within my company so maybe it's like a team thing a company thing or maybe it's something where you say i want to build something and i want to sell that to other people so i want to have this as like an ai agency offer maybe right so i sell it to other companies or i sell it to other people and that could be partly productized meaning you build something you package it and you sell it
partly a surface, right? So there's, you build something, but then there's a surface element to it as well, right? So you help implement it into businesses.
This is all related to, for example, like AI automation, right? As you know, no single automation is ever 100 % the same. So there's also a customization component to it.
And then the other level is, do you want to build like a software product, like fully handoff that people can... just jump on a subscription. They could just go to a website, create an account, download it, get on a subscription and use it.
For example, what we're doing with Glido, right? So there's these four different categories that your ID could fall into. And why is that important?
Because as you progress up that ladder, I would position it as a ladder, it gets more complex and the stakes are bigger. Again, summarize. If you just build something for you, you use it.
stakes are very low just just go and build just open cloud code and go for it if you use it inside your team other people use it it's still considered internal right but you may already want to be a little bit more careful like who's going to use this what are what kind of like data is going to be in this application now then if you sell it to to a company there's even like a legal aspect involved, right?
Because you probably have a contract. There's terms of service. There are certain obligations when it comes to data privacy and protection.
So there's a bigger risk. And then the other level is when you actually build a software product and you let consumers use it, we have the whole idea of data privacy and personal rights and all of that when it comes to data and probably more people using it. So just the risk gets bigger.
So let's pause there for a second because that's what I think it's good to start there. So was that all clear? Yeah.
So it sounds like before you start building, you want to kind of have an idea of... So essentially you said how big, and that makes me think like, how many, how many people will use this? How many hands will this touch?
And would it be fair to say that a good way to get started is you kind of start on that first rung of the ladder where it's like, I'm going to build this first just so I can use it. And that's sort of the way I'm thinking about it. And that doesn't mean that later, if I wanted to have my team use it, I couldn't scale up the back end because, you know, it sounds like as you move up the ladder of.
requirements and people, you have to think about other things like the architecture a little bit more and the database and just making sure that it can handle the stuff, the privacy. But it's totally fine to start off on that first rung. And then when you need to move up, you can start to do that, right?
It's not like you choose, okay, this is going to be a personal tool. And then later, maybe six months later, you realize you want to scale it up. You're not like stuck, right?
You're always able to continue to move up the ladder. Yeah. I think that's a very good clarification because of course you can always like make the ambition.
bigger right of the project and i think One thing to watch out for is, for example, skipping levels and being someone, for example, no technical background and jumping straight into building custom software and then selling that to people and letting companies rely on it. So you skipped a couple of steps in terms of your homework in order to figure out, okay, what does this look like?
What does this mean? What data is flowing through it? So I think that is a very natural way to progress in there because you're totally right.
So when we make that a little bit more... kind of like tactical as to what's different, right? There is not much different nowadays in terms of the process.
It is still, even for my projects, like I don't write a single line of code anymore, not a single. So everything goes through nowadays, either through Codex or through Cloud Code or even through Grog Build, depending on kind of like who has the best models. So everything goes through the coding agents.
So that's also interesting like for the audience. that I think this is the very first time in history where someone with zero coding skills and someone who has been doing this for over a decade use the exact same tools and the exact same process. Which is quite wild to think about.
I'm not just like, now that I'm explaining this, I'm just like making this up, but now that I think about it, it's actually pretty wild because think about any skill that you would learn, right? if you just start with i don't know playing basketball or anything like that you're not going to start at like the elite level like your exercises and what you do is going to be way different right but right now we all start with the same really, really great tools that can help us produce code.
So the only thing that is different is what instructions do you give it and how do you review the output, right? And then that's the skill. That is where the reps come in and that's where also the stakes get bigger and bigger and bigger as you start to build software that's used by more and more people.
But I think that is very motivating. for people to watch. You have these tools.
You can jump on a $20 subscription. You probably want to get to $100, $200. Otherwise, you know, you run out.
You'd be sitting there all week like, man, when is this reset coming? You probably want a little bit more, but like that. For most people, that is within reach.
And from that, you can start building. You can start to build really, really cool things. And yeah, start building with something that...
you would use that you think is cool and then start to make it better. And then over time, and we can get into this deeper if you want to, but for the most part, vibe coding is totally fine. But there is one aspect of it that is security, where you just cannot afford to make the risks.
Because let's say even if you build an application and it's vibe coded and the architecture is bad, What's the worst thing that can happen? People will use it and at some point it will get a little bit slower or it will break.
But you'll notice, right? And you'll start fixing it. So there's nothing really...
And that is something you'll be aware of because people start complaining. It's like, huh, maybe we need to revisit that. But a security breach, that's one thing that you cannot take back.
Oh, sorry. We'll fix it in the next version. You know, like that's out.
Once those emails and that personal data is leaked, you can't go back. So that is one aspect that we could get into that if you want to. But that is, I think, the most important thing as you level up security.
Totally, totally. And I definitely want to get into that. And I think we'll just kind of keep working our way up here.
That's why I love the ladder analogy, because you really do have to earn your way up to the next rung. You can't just jump from the ground to rung four or you're probably going to fall and you just have to start from the bottom again. It sounds like the first piece is kind of you figure out how big and you think about high level, like how big am I building this and how much risk is there involved?
Once you kind of have a good idea of that, what is the next step to start that building process? What does it look like for you to sort of map this out in your head to sort of plan out the requirements? How does that process look?
Or if that's not the next step, what is the next step? Yeah. No, that is definitely the next step.
So when you have this idea and when you want to start just building. What I've observed over the past year, pretty much, one year ago, everyone was really pushing spec -driven development, the importance of plans, and you can probably remember this, right? And I've just found that as the models get better and better and better, and the models that we have to right now, Like the planning phase and the generating the specs becomes like it becomes less important, I found.
So what does that mean, spec planning? So spec development refers to kind of like a way of building software where you first use AI not to produce code, but to produce specs. And what are specs?
Specifications. So you describe how the project or the product should work in natural language. And you would use AI for this.
So you would pretty much say like, I want to build X, Y, Z. I want to build this. And then let's create some specifications for this.
So it would say the AI would start to create a document. Okay, product needs to do this. It needs to do that.
It needs to do that. It needs to do that, et cetera. And you would create this plan and the specifications.
And then once you had all of these documents, only then you would read those and you would approve those. Only then you would like go into development mode and you would actually ask your agents to start writing the code and to start building the theme. That was pretty much the meta one year ago when we just got these like big breakthroughs with Cloud Code, Opus 4 .5, 4 .6, right?
Like the good old glory days of AI agents. And now it's just like getting crazier and crazier. So the thing is, There's still value in that.
And I don't think it's wrong to go through that process, even just to help you think about what it is that you want. But the thing is, what I find myself doing most of the time right now, I just like I go straight into building and like I let the model figure out what it like, how it should structure things. And there, of course.
certain kind of like standards that i kind of like stare the model towards when i'm for example working with a python programming language just because i've so much experience with that but like the summary of this it gets easier and easier and easier to just tell it what you want and it will start to it will start to build it will start to build it so coming back to you you started this question so let's say people have this idea and what is the next step I would say right now, get on Codex, get on Cloud Code, use GPT -6 Astra or Opus 5 .5, which are currently like two top frontier models.
Opus 5 .5, really good so far. Super good. Super good.
And just tell it what you want to build. I love that. Yeah.
And as simple as it sounds for people just starting out, there's not much more to it. And then you just look at the output and... You tell it what you like, what you don't like, and that will already get you really, really far.
Yeah. And the cool thing about that is that I think humans are really good at explaining what they want. Like, you know, it's easy to say, I want this, I want this, but you really just have to sort of go through that grill me process where you have AI interview you about what specifically do you want?
What do you not want? What does this look like? What is your vision?
And I was just talking to a few Google engineers and they were explaining this as saying like the plumbing, the boring stuff on the backend that, that we used to have to think about where humans used to have to set that all up before they could start to be creative and think about the vision. That plumbing is pretty much just taken care of now because you explain what you want and the agents know how to go set up all that stuff so that you can really just, they kept calling it intent driven automation or intent driven auto engineering.
And I thought that was pretty cool. I wanted to see, do you agree with that? Do you agree with that narrative or that mindset?
Yeah, for sure. I really like that. I like that term, intent driven.
And what does it mean? You just tell the model what you want. Like you have a certain intent as a builder, you have an idea and you want to make that a reality.
And the thing is often when you start with something, even though you might have like a rough idea of where this thing should go. But until you actually see it in front of you and use it, you cannot really give the best feedback I found. So you can like, then that's why I've moved away from like big plans and specs and all of that.
I just like, just build me this. And then I see it and I look at it and I test it and it's like, hmm, this feels a little bit off or can we make that faster? this doesn't really look right like whether whether it's visual or actual logic right or the output here is wrong or can we fix that right and um that's also now the pretty much the like the development process with glido the tool that we're building right now it's mostly like ideas and then just using natural language like it's actually fun we're using glider to build glider right so we use glider to talk to it and it's like hey can you try this you know like let's build it out And often I do that straight on the production code base.
I'll do it on a separate branch, right? So it's not like immediately shipped, but there's not even like a playground. I just create a new branch and it's like, build it.
And then I can directly see what it looks like in the product itself. And then I can iterate on it. And then that's, I think, a perfect description of like intent -driven development.
You just tell it what you want and it will build it. By the way, guys, I've got this completely free SOP for you about getting your first AI automation client. It's going to go over the exact steps that has been proven for hundreds of our AIS Plus members to get their first paid gigs.
It goes over the one sentence service pitch that can get you started today, why your first client should cost you money, the five minute video that answers can this person actually deliver before you've actually... received any money. What to do when you have zero case studies.
There's so many good things in here that are going to help you out. Even if you already do have clients, I would recommend grabbing this because like I said, it's yours completely free. So if you want to grab this, there's a link for it down in the description.
Let's get back to the video. That's so cool. It's really interesting how you kind of pointed out how the meta a year ago was doing all this kind of planning, but now you say you're kind of building first more and then really just getting it in your hands and playing around with it.
I remember when I used to use plan mode first for everything, not even if I was building automation, but I would just use plan mode default just to sort of brainstorm. And it sounds like maybe, do you think the models are just getting better and the models and harnesses are getting better? So that's why the meta has shifted or?
It's just like, and it's for sure the combination. So that's also important to understand. And I think even like people that are actively working.
on cloud code on the development i think i even saw this on x like entropic engineers mentioning this as well that plan mode is not even really needed anymore because yes like models get better but um it is essentially also within the harness because what what what is really a plan right it is first essentially getting you give it you always start with a goal or a task i want to build this and then Before, it would help the model to first plan out.
So before just getting into code generation, it would first zoom out a little bit. Okay, let's just first reason over this. Let's see if everything makes sense.
And then once we have the plan, I can now execute on it. And it feels like the engineers working at OpenAI and Tropic is essentially... By making those models and the harnesses better and more capable, they just merge this process as like a trait that the model and the harness combination already have without like turning the plan mode on.
That makes sense. So now let's say we have started to throw in some of the I want statements and now we're getting some stuff back. Yeah.
In automation. You know, we run a lot of evals and we see, okay, cool. We have this, you know, we have these inputs, we have these outputs.
This is, you know, 90 % correct. Let's make some iterations and see if we can get this to like 98, 99. How do you think about from a, you know, an app perspective or a software perspective, essentially the eval process?
What does it look like to verify? What does it look like to see if the things actually work rather than just assuming it does or having you, you know, go in there and click through a few times? Yeah.
So. This, of course, depends very much on what you're building, of course, because output quality, so what does good look like, is always tied to the application that you're building, right? So let's talk about Glido.
What does good look like? Glido works well. If you press a button, people can talk.
That will be processed in under 500 milliseconds, and the text will be pasted back into your application, like exactly the way you want it to see there, ready to paste text. So it will correct some errors and all of that while leaving generally most of your speech intact. Like that's what people want from a dictation app.
So it starts with that, getting clear on what does good look like.
the second question is can we find examples of that can we find examples of cases where it went well and in glido for example we don't use user data because we really designed this application to be private by design. But we use the data of our accounts internally, meaning the developers that work on Glido, right?
So we use our own kind of like data sets that we collect. We have a bunch of people. So that data set then also grows.
And in there, we can run these experiments. So we have data sets in there. Like when we put this into our system, we expect this to come out.
And that essentially then becomes an default, right? That you can run and that you can check against. The cool thing is that as models and harnesses get better and capable, whenever you have some data and when you don't have some data, you can even ask the model to generate some realistic data for your use case.
And again, whatever you're building, whether that's a chatbot, an automation of any sort. And once you give... these models an objective in terms of like here's what i want this to look like here's what i want you to improve on i want it to be faster i want it to look more like this i want it to be better i want to have less errors once you give that and let that run in a loop And that pretty much means you tell it to go and improve it, use the data, generate more if needed, do research, use sub -agents, run this in the loop until it gets better.
It's a very messy prompt, but even just saying, giving that prompt to a model like Opus 5 .5, given the thing that you're building, will kick off this research experiment where it will go on and it will try to find ways to make it better. What we find over and over again is the better models and harnesses get, the easier all the things that we have to do as developers get.
Planning, spec different development, quality assurance, debugging, security, evaluations. All of these things also get easier as the models get better. So I think that's a really interesting property that I found.
And also just something that you need to be aware of because you do need to ask the model, right? So you do need to be aware of like, what are e -files? How can I set up a loop in order to make my application better?
So just that realization and asking it is the right starting point. Yeah. So it sounds like defining what does good look like?
And that's something that you would do with a human as well. You know, you're setting the expectations of what do you want? Because if you don't align with a human or an agent on what you're looking for, then when it comes back with something that isn't aligned, it's maybe not that it failed.
Maybe it's just that you didn't, you weren't clear enough on that. And then from there, giving it a way to essentially verify that, you know, whether that's you verifying or the agent verifying, I think kind of a combination of both, but what is good? Here's how you verify it.
So what you were saying about this whole like loop, it made me think of when Carpathia's like auto research came out. And I think that's when I remember having this big moment of like, wow, like agents can actually see their goal and prove it and then keep researching and keep trying again until they hit it. And I thought that was incredible.
Now, I've got an interesting one for you to kind of stem off that, which is a lot of the best auto research or goal prompts or loops are when you have an objective milestone to make them hit. what about how do you think about it if it's subjective if it is something that's more like taste or you know if the definition of good isn't something that can be objectively proven yeah um that's very true whenever you have something that's very easy like it's either good or bad But because you have a data set with examples, those are the best scenarios or making things faster, right?
So for Glido, we also have a lot of infrastructure that we need to manage, models that we deploy. And there also code and architecture around that. If you just give it an objective, like, hey, here's what we're at right now in terms of latency.
You can test that by running a test. Now go make it faster. You know, it's super clear because every time it runs an iteration, it can see.
am i above or below am i going to in the right direction so what if you have a use case that doesn't have this and uh let's think about an example of this to uh make it a little bit more tangible so my most common is like video editing and i wanted to like i'm saying like yeah make it feel professional make it feel engaging it's like how does it know yeah yeah yeah yeah so Yeah, so that really is a tricky, it's really a tricky example.
I'll give you one more and then I'll also consider this one. Okay. But we have, for example, also done a lot of work in customer support, in customer care.
And the thing is, when a customer asks a question and the AI agent gives it a reply, right? There are hundreds, thousands, probably even like millions of different ways that you could reply to a customer and they could all be correct or they could all be wrong because the final goal in this case is that the customer is happy, right?
So like they got a good reply, but whether you use a different word, yes or no, that is all subjective. meaning there can be differences. So one of the techniques that you can use for that is what you would call an LLM as a judge to run that at scale.
And how does this work, right? So the best way to check the outputs of your AI agents when it's subjective, the best way, which is also the most expensive and non -scalable way, is you review everything by a human or multiple humans. You let the AI do the work, the human reviews it and says, this is not good, this is good, et cetera.
But then what's the point of automating it? Because in the case of like customer care, if a human needs to review every ticket, they could probably just reply themselves, right? Because it's probably easier to reply than to review and give good comments.
So that's unscalable. But if you do that for a little while, you could create a little data set around it. When the customer asks this, And the AI said that.
The human said, this is good. Yes or no? Because.
Now you start to collect data. And then this whole concept of an LLM as a judge is you use another language model. So another like independent language model call or multiple calls or however you want to set it up in order to review the output.
So you let an LLM look. at the, let's say, the reply to the customer. And you let the LLM say, is this good?
Yes or no? And why? Now, the challenge is in the beginning, as you know, like an LLM might say, yes, this is good.
Or no, this is not good. But then how do you know that is good, right? So you like shift the problem from like an LLM that you need to check to another.
So there's a little trick for that. And that is essentially called creating alignment between the human reviewer and the LLM. And you can do it by, let's say you have 100 examples and you let the human review 100 examples and you list all of this.
You have an Excel spreadsheet for this, whatever. This is like manual work. Nobody wants to do this.
It's unscalable, but you need to do this because this is where you can essentially capture the taste. So human will say, I like that. I don't like that.
That should be better. And now what you want to do is you let an LLM. run over all of that and you see where they where do they agree where do both the human and the llm say this is good and this is bad and you'll get a score so you'll have a percentage so let's say the alignment is 80 or maybe maybe even like 50 so 50 of the time the llm agrees with the human yes or no this is baseline and now once you have that data this is where the interesting thing comes in now you can use ai again so now you can say hey here's human data here's llm data Let's optimize the system prompt so we can create more alignment.
You know, now you can do loops. So then the LLM is a judge. It's a system prompt and a model.
That's pretty much it. And then if you let an LLM review that, so like, oh yeah, I can see the LLM thinks a little bit different about this and this and that. And I thought this was good, but actually the human thought this was not correct.
So I'll change that in the system prompt. You run it again. Alignment is at 75%.
Cool. Let's do another round. 80%, 90%, 95%, 99%.
And now if you have a large enough sample size and you do this occasionally to avoid model drift, you'll have an LLM that with two degree of certainty agrees with how a human would review it. That's generally how it is done on real world projects. Coming back to your video editing example, you would follow a same process.
But I would say it's probably even trickier because assessing what good looks like in an edit or an animation or a cut, there are some things that are obvious. Like some, you can make an error. But if like a title has like a certain animation pattern or something like that, and you don't really like that, that's tricky.
But it's mostly because of the LLM's ability to inspect the visuals. right i think that's what makes it more tricky so as multimodal capabilities get better you can apply this exact same process even to optimizing and training your video editing agents yeah i think the interesting thing there too is like let's say you have worked on refining this llm as a judge and it's really really good and then if you switch the model it might not be as good anymore because the model might interpret it differently but what's cool about that is you're you're putting in that time, which is a tedious process because you're essentially creating these standards, but then you can use sort of that loop that, you know, auto research loop on it to have it continuously optimized, which I think is very, very cool.
And it's, it makes me think of, it makes me think of like sports teams. They'll have these scouts that will go out and watch, you know, high school games or college games or whatever it is. And there are stats like height, weight.
you know how many times can they bench press 225 there's stats like that that are objective but then that's also like the team trusts the scout's opinion to say this person has like this guy has potential he's explosive he has great game sense like that stuff that you can't pick up on paper as much and it's like how do you get your llm to have the taste and the judgment of that scout or of you you know what i mean and i think that that is super super cool um now As we transition into now we've been running some verification, we have something that we trust that's working.
What if we want to turn this from a one person app or a team app? What if we really want to scale it? What are the things that you start thinking about?
We don't have to get super, super technical, but what are the high level things you start thinking about around infrastructure, databases, security, that sort of things? And if you have any like examples of maybe landmines that you've stepped on when trying to scale something, that'd be super interesting. Yeah, let's dive into it.
This is my fun zone, pretty much. I like this. So let's indeed make it a little bit more like tactical and therefore also technical for the audience.
Because up until this point, also my recommendations for people may have been quite obvious. Just use cloud code, right? Just ask the model, bro.
And the funny thing is, that's true. That's what you should do. But there is still...
part when you really want to take things to the next level where it is very worthwhile to like train yourself and to get a better understanding of how software works and I keep coming down to describing that as just the high -level architecture of your application and the software that you're building And there are two distinctions between that.
And this is very worthwhile to research a little bit what I'm about to say. So this could even be you with Codex, with Cloud. Spend some time asking the following terms that I'm going to mention.
So architecture. First of all, we have it at what I call the system level. And this means...
How do all of the individual components, like your backend, your frontend, and your database, and potentially even other services, how do they talk together? Because most software products are a combination of those things, right? And it's very important to understand what the role is of each and every one of those, how they communicate to each other.
And what are some of the traits, the unique things, and also the things that you need to care about when it comes to those individual things, right? So this depends heavily on, first of all, what are you building? Are you building a web application, a desktop application, a mobile application, something that combines all of those?
So it depends. It also depends on what language do you use. Do you, for example, use...
languages where you have a very clear separation between your backend, let's say, that is built in Python, for example, or in Rust or in C? And do you have a frontend application? So the visual layer that is maybe in JavaScript using Next .js.
So then you have usually like completely different code bases. Now, you can also have projects that, for example, just work in one stack. So you do everything in TypeScript, for example.
So TypeScript, that whole project, it is your backend and it is also your frontend, your visual component to that. So those are all things that matter based on what you're building. So now that you get a little bit more serious, that is a good first question to ask your AI agents.
Hey, I'm building XYZ. What would be the best? Stack and then stack is the important word here.
What would be the best stack in order to build this? Do I need a separate back end and a front end? What kind of database should I use?
Can I put everything together? So those are then questions that you can go through and it will give you examples. So you can just ask like, hey, what does that mean?
What's the simplest solution? What fits the best for kind of like what I'm building? Because that is a really important starting point.
Because once you log in on a programming language, like within the language itself, you can do a lot, right? You can scale up, you can scale down, you can refactor, you can build. But switching a project to a different language, even though with Coding Agent, it becomes easier and easier and easier, that is usually not a trait that you want to make, right?
We had to do it with Glido at some point. It was a lot of work. Like we started out on a different stack than what we use right now.
Because in the beginning, we didn't do our research properly. It was also a little bit because of team structure. We had someone working on that and the other one working on that.
And then we decided to go in a direction. And then later we figured out we need to do this differently. So that's a landmine that we stepped on.
It was choosing the kind of like things we're familiar with rather than picking. the tool or the stack that was really like most optimized for what we eventually wanted to build and in the past um this would pretty much have not been possible so if you are comfortable with it with a stack you've been developing that for 10 years you most likely want to use that right but i think that now changes you could now very well say i'm confident with that stack but there is there's evidence there's actually evidence that this stack is way better for what we're trying to do Let's go with that stack.
So architecture at the system level, backend, frontend, database, how does everything talk to each other? Then you also have architecture at like the project level. So this is when you go inside, let's say your GitHub repository and how you structure your projects in there.
This is less important than the higher level system architecture, but it's still important to have like a rough understanding of how do you structure projects in terms of files and folders. So what do you put in there? And how do you structure your code in terms of functions, classes and modules that you create?
So when you're creating code and these agents get better at it every month, but you can still have these like. spaghetti code bases. You build something here, then you build something there, then you build something there.
And this can happen quite quickly when you are vibe coding and you go from idea to idea to idea and you experiment a lot, which is a new way of building, right? You just ask it to build something and it might build something. Now you have this artifact in there.
Then you decide, let's leave it there. I don't really like it. Let's come back to it later.
You build something else. You forget about it. So now you have this mess.
You have this junk sitting in there and that might have. some logic in there, which you still really use and which you still really need. But it's got all of this bloat around it that was just a test feature that you wanted to build.
Now you start to build on top of that and you create these new features. It's all great, but they still depend on this kind of like messy little part that's really deep way into your code base. So this is how you get spaghetti code base, right?
So over time, it gets harder for your AI agents to manage and to find where everything is. And this is where you'll run. bugs at scale so all of a sudden things stop working right it works fine for three months and then all of a sudden things don't work anymore and it's because you remove the feature that was not no longer needed but it contained a little part that you actually still need it so understanding how that works at a high level is very valuable to do and I'm gonna give people like one tip to look into there's a really good skill for this and it's called I think it was recently changed, but it's from Matt Pocock.
So this is like, he's famous for his skills. Great engineering skills. Yeah, great engineering skills.
So it used to be called improved code -based architecture, but I think he refactored that a little bit to, I think, code -based design or something like that. But if you - We'll find it, we'll link it in the description. Yeah, if you search for Matt Pocock and then find some of his skills where he talks about - improving architecture about how to create deep modules and locality and how to work with the seams.
That is a skill that I very often use to review my code bases. And he even mentioned that he created that skill to fight AI slop and make code bases that are easier to navigate by AI. And imagine how cool that really is, where I just...
set pretty much there's there's two components that i stay are there are still i think very high leverage to learn more about architecture at like the system level and the code base level for the code base level there's literally just a skill that you can use and if you just read that skill and use it in your coding agents you'll get better at it to me that's just amazing that that's how you totally get how you get better at it and then yeah and then the like the front end database back end Spend the day researching that.
If you're serious about building stuff, just ask questions. Even do it on lunch break, go out, take a walk. Start to learn a little bit about what that means, how they communicate to each other, and also what the different security aspects are of those different levels.
Because security applies to your entire application, but there are specific things that your database handles. your backend handles, your frontend handles, and the communication layers between those. Yeah, I think it's really cool that you brought up the skills of Matt Pocock because one of the things that I was thinking about as you were going off there was what you mentioned at the beginning, which was the fact that really, under the hood, software engineering is still the same, but the human role has changed.
Because of the fact that these things are so smart and can move so fast. But still the principles of building good code are still the principles of building good code. And for someone like me, or for a lot of people probably watching this, maybe when you started saying like Python, C, Rust, they were like, what are those things?
But that's been around for so long. And people like Matt Pocock, people like you, know what it looks like to have a good code base. And that means that you can leverage the subject matter expertise and experience of other people.
if you ask the right questions, if you have it do research, if you leverage skills like that. And I think that that part is really cool and hopefully gives people a lot more comfort in the fact that they're not alone in building this app because there's so much stuff out there from people that understand what does good look like when maybe you can't define good as well as someone else can.
So the last piece I wanted to touch on then is, I know you hit on a little bit, but... What should people be thinking about when it comes to the security stuff? Not only from a data privacy, but also how do you think about what's the potential of your app getting hacked?
That sort of thing too. How do you think about that, especially if you have no background in any sort of cybersecurity sort of stuff? So the cool thing about this is that with...
very little, like with just the basics, you can go really far. So of course, security is this big domain and there's different levels to it. But I also don't come from like a security background.
I wouldn't even consider myself a security expert. We do have yours on our team who knows way more about that. And I've picked up most of the things from him over the years.
Like I've said, the cool thing is there are actually like, if you do a handful of few things, it's almost already impossible for you to get hacked. So let me give you a couple of those things to look into. And there is a little bit of overlap between like what we would call infrastructure.
So that is the whole concept of, okay, we talked about architecture at like the system level. backend, frontend database, right? Infrastructure pretty much means like, where do all of those services live?
Where are they deployed? And how do they talk to each other? And what are, for example, the firewalls around them?
There's a little bit of overlap in there. And again, very high leverage thing that you can look into. Even if it's just like from a research perspective to learn more about this.
So let me walk you through some examples to make this a little bit more tangible. So let's start with like a database. For most people building something, you want a database, right?
Most applications need a database. You need to store user accounts, passwords, any information that you might need in your database. And right now, the most popular option for that is Supabase.
Superbase is a really popular database and Superbase is amazing. I would say to anyone, if you don't have a very strong reason to not use... uh super base use super base meaning that if you're non -technical if you're just starting out just use super base it is flexible enough that you can do anything with it don't be like uh bothered by people saying oh you really need like an unstructured database for this use case or like oh you need a dedicated factor database like super base can do all of that just use that and then like i said if you have a very strong argument as to why you don't like Superbase, then you're probably technical enough to make the decision on your own to get what I mean, right?
So that makes sense. Yeah. So database, use Superbase.
So Superbase, you can also self -host that, by the way, but for most people, I would recommend just use it in the cloud. That means go to superbase .com, create an account, create a database. There you can get on a plan and you have your database.
Now, then what I... recommend to do is again, spend a couple of hours talking about the security aspects and best practices that you can set up around super base. Super base is great out of the box, but there are some things you need to be aware of.
This mostly has to do with row level security. And again, I'm just putting out these terms in here, not to go into them, but for people to like research into that. So row level security is important.
Like look up what that means, how you can work with that. make sure that your database has like a strong password like common sense strong password make sure that your user account that you can log into on super base strong password and also potentially two -factor authentication this is all very simple stuff right but this is just avoiding like people and random bots being able to like randomly guess a password and can get into something now then also So that's your database.
So now your data lives in there. If you use Superbase in the cloud, you put on row -level security, and you have strong passwords, and you make sure that all your services that connect to it use the credentials in environment variables without actually putting it into code, you're already in a good spot. There's more that we can do, but this is already, from a database perspective, a good start.
Now, then, when it comes to your backend, which is the second most important thing, because for your backend, usually you have data flowing, right? So your frontend application may load data directly from your database, or it can also go through your backend, right?
So your backend talks to your database and then your backend like... puts that or flows that data through to your front end. So the backend is another part where you need to be aware of, first of all, what language do you use?
And then second, where do you deploy that? So where does that code live? So when you start building something, it's on your laptop, it's local host.
You're running things, you're testing things, but when you deploy something, you actually put it in the cloud and this is where people can get access to it. So there are, of course, various... A lot of different services that you can use.
You have tools like Railway and Render that make it really easy one -click deployment. You can rent a virtual private server. You can use Azure, AWS, GCP.
So different types of ways that you can set this up. But the most important thing is that when you deploy your backend, that you set up ideally some type of firewall. rule around this, where you restrict certain traffic to it.
And this is, again, something that you just need to do some research in, depending on which platform do you use, which language do you also use, and then use your AI coding agents to talk about, okay, what is like a firewall? How can I set that up? Now, this also applies to your database, and you can do this in the Superbase dashboard.
But why this is important is... Usually there is a part of your application that is user facing. That's typically your front end.
Meaning that's what people actually see. They go into the browser or in a desktop application and they use that. And then your backend and your database, no one should be able to touch that.
Only your user application should be able to pull information from that. And your front end application is that also deployed somewhere.
Again, can be any of the services that you use. Maybe you use it on Vercel. The thing is that deployment there uses an IP address.
So that is a specific address from which it reaches out to either your database or to your backend. And I know it gets a little bit technical, but a firewall pretty much means you set up a rule where you say in your backend and in your database, only open the door. If the request is coming from the IP address that I whitelisted, which is my known front end or desktop application.
And there is very different rules depending on how you set up your infrastructure and how you set up your architecture and your applications. But that is the general best practices around security. Setting up the proper firewalls and then making sure secrets.
API keys, credentials, all of that are handled correctly. So you cannot read them in your application code. People cannot read them, for example, in a web application, inspecting the code in the browser, right?
So if you protect your secrets and your credentials, you have row -level security on, and you put proper firewalls on your database, your backend, and whitelist it for your frontend application, that is security in a nutshell. And again, it's technical, but it's more so for people like to take this snippet, summarize it, put it into like a chat and then reason through this, what that means for your application.
Yeah, I love that. That was a masterclass. And I think one of the things that have become very clear throughout this conversation with different technical concepts that you've brought up is that a lot of this sounds like it's just about awareness.
If you can bring up this conversation to your agent and let it go do the thinking and the understanding of implementing. But you were the one who had to say, hey, think about this and do this, right? Exactly.
I think that's very cool. Yeah, it's those unknowns unknowns that you in the beginning as a new builder, developer need to become aware of so that you can ask the right questions to your AI agents. It's not about the skills or even the intelligence or the years of training.
It's just the awareness. being able to like ask the right questions. And then also with the base level understanding that you have, being able to somewhat reason over those answers, right?
Where you can say, yeah, that is important. And let's not get into this right now because those are not kind of like the basics. This goes way deeper.
That's not important for my application right now at this stage. I love it. Let's wrap up here with one final sort of like big statement from you.
What I want to hear from you is, Basically, I want you to start the sentence like this. If I was starting today with no technical experience, I would be most excited about this.
Like, what is the opportunity and what gets you excited? Like thinking about the future of where AI is headed and what you're able to do with it in this, you know, software engineering or building apps sort of space. Cool.
I really like this. So, yeah, here's a good answer. If I were just starting out and I see what's going on around me with all of this technology, what I would be most excited about is the opportunities it can create for each and every one of us.
So I, out of university, quite randomly happened to land on a freelance gig and worked self -employed pretty much ever since then. Started to build businesses off of that. And that has given me insane amounts of freedom and getting to work on exciting things that I want to build.
And I think that is now within reach for more and more people because of this gigantic transformation that we as the world will be going through in this AI transformation. And you might be thinking, yeah, when the tools get easier, everyone can do this. True, but we still need implementers.
We still need builders. The fact that it is possible and that we can do it doesn't mean that everyone and every business will actually take the initiative to start doing this, right? So there are just going to be so much opportunities.
to for like pretty much um entrepreneurial things to build things on your own to to sell that to make that available to make content to help local businesses or maybe not even local but to help larger businesses to be able to be excited about the technology about the automation about the transformation and then getting better at that craft as you do it and then being able to like monetize off of that And if that is exciting to people watching, which I know it probably is because I know your audience, right?
Then I would say like, go all in on that. And if you, even if you're not just starting out, but you may already have a job, like figure out how you can do this on the side, right? I, for example, have an entire community of freelance developers and data professionals.
And I also share that with people because the opportunities are real. You can do this next to a full -time job. You can...
pick up a project and while you're doing your day job, you could just kick off another agent and let it do the work, right? You can even use your phone to talk to Claude on lunch break, give it a master prompt. And when you're back home, like it has built the automation for your client.
So I would say that's the most exciting thing. I recently watched a podcast. It was on Lex Fritman and we had DHH.
I don't even like it's David blah, blah, blah. DHH is what he goes by online typically. And he described it as having a genie in a bottle.
Like we as developers now have a genie in a bottle. I like that. Meaning that you can just think about it.
And you can prompt it into existence. And you're not even limited to three wishes. You can just keep asking.
You're just limited to your tokens. That's where you got to be careful. To your weekly token budget.
Yeah, yeah. But I think that is a great... like description and you know this as well nate because i see all these videos from you whenever and every time a new model releases like the stuff that it can do just gets crazier and crazier and crazier and it used to be like this yearly timeline where if you looked one year back like the models were like extremely way way better i think now we have it already like on a quarterly basis where you have these like oh yeah really holy fuck this is so much better than what we had right and that will That will continue.
I don't see a world where that is going to stop. So you just need to learn how to like use these tools and then build stuff with it and then use that to ultimately create the life that you want. Because like building stuff just for the sake of it is fun.
But I think we have a great opportunity to like help yourself, help your family and to create a life around these tools that was just not possible. years ago yeah i think the only way we see the the progress slow down is if this whole like pacing the frontier discussion that's going on right now if there's an actual standard in place for regulation and for it's not gonna i guess kind of like a a national and global coordination so we'll see what happens there but it's not gonna happen you don't think so they go full steam ahead there's no stuff oh man we could do we could probably do a whole another hour talking about this stuff i think it's really interesting as well but um this space has been you know so exciting so fun i'm so glad that i've been able to connect with people like you and hang out in person and just you know you're super smart guy so i appreciate you coming on this has been a super informational episode so really appreciate you coming on dave where can people get in touch with you or learn more from you if they're interested cool appreciate it nate uh just go to my youtube channel dave abelard mostly videos for a more technical audience but that said
What does that matter nowadays? Like I make sure that my tutorials, I also create these kind of like handbook guides for them nowadays because it's so easy. I used to just create tutorials and teach them.
Now I create a tutorial and I just say to AI, like create a whole entire documentation page for it so people can walk through it step by step. So the cool thing is. I mostly do kind of like big builds, how to build this.
For example, I just launched a video, how to build an entire company knowledge base and actually do it properly. Like something, not just like the toy project, but actually something that you could like offer to a company. And it's for a technical audience, but you could also just use your AI agent and point it at the GitHub repository in the handbook and it can build it for you.
So yeah, YouTube, Dave Abelard, that's where you find me. sweet i love it well thanks so much for hopping on today dave and maybe we can do it again sometime my pleasure nate talk soon awesome all right see you dave
The Hook
The bait, then the rug-pull.
A dramatic declaration from a ten-year veteran software engineer who reveals he has completely ceased writing code manually, delegating entire production codebases to AI agents.
Frameworks
Named ideas worth stealing.
12:42list
The Four-Rung Ladder
A tool just for you
Used inside your team
Sold to a company
Software anyone can sign up for
A hierarchy mapping software complexity, legal exposure, and technical rigour against user volume and data sensitivity. Moving up each rung requires stricter architectural governance.
Steal forEvaluating the risk profile and security requirements of an AI software project before writing code.
37:24model
LLM as a Judge Alignment Loop
Collect human evaluations on 100 sample outputs
Run an independent LLM judge across the same samples to measure baseline agreement
Iteratively refine the judge system prompt to align with human taste
Deploy the automated judge in an autonomous feedback loop to self-improve the primary agent
A systematic workflow to automate quality assurance on subjective model outputs without ongoing manual human evaluation.
Steal forAutomating quality control for non-deterministic AI workflows like customer service, video editing, or content generation.
CTA Breakdown
How they asked for the click.
VERBAL ASK
28:15link
“FREE First Client SOP: https://app.aiautomationsociety.ai/optin/first-client-SOP/”
Delivered as a brief mid-interview presenter cut explaining a step-by-step PDF SOP designed to help students secure their first paid AI client.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Nate Herk builds ClientPack, an AI tool that turns discovery-call transcripts into branded client proposal decks, from a blank folder to a live paid subscription in about eight hours of prompting three different AI agents.
A five-prompt AI agent build produced a working Calendly clone with live calendar sync and Stripe payments, then the real bill showed up: five days of agent runtime and about $15,000 in inference.
A 26-minute live benchmark that runs three real builds side-by-side and reads the session logs to settle the Claude Code vs Codex debate with actual numbers.
A 25-minute live build that covers why agentic workflows command premium fees, how to structure them with the WAT framework in Claude Code, and how to sell the result on value rather than hours.