Modern Creator
Jan Marshal · YouTube

How Senior Engineers Build With AI (Full Intercom Clone)

A six-hour live build of an Intercom-style support platform, from idea and PRD to a production deploy, with the workflow doing more work than the model.

Posted
4 days ago
Duration
Format
Tutorial
educational
Views
32.3K
600 likes
Big Idea

The argument in one line.

The difference between a junior and a senior engineer using AI is not the model but the workflow: define the product, lock the architecture into a PRD, orchestrate specialized sub-agents, and review every line before it merges.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • A developer who already ships with Cursor, Codex or Claude Code and wants a repeatable system for building a full SaaS instead of prompting feature by feature.
  • Someone building a multi-tenant product with real-time messaging, file ingestion and a RAG agent who wants to see every architectural decision argued out loud.
  • An engineer deciding how to split work across models and sub-agents and how much code review is still required when agents review their own output.
  • A builder evaluating a backend-as-a-service for Postgres, auth, object storage, long-running functions and an AI gateway in one account.
SKIP IF…
  • You want a short tutorial on one feature; this is a six-hour end-to-end build and the value is in the sequencing.
  • You do not write code and are not planning to review generated code; the core argument is that reviewing is non-negotiable.
  • You need a framework-agnostic approach; the stack is locked to Next.js, Prisma, oRPC, PartyKit and Neon.
TL;DR

The full version, fast.

Asking an agent to build an Intercom clone produces thousands of lines that do not work, because a wish is not a plan. The fix is to define the product and its non-goals, draw the architecture, map the three core user flows, then have the agent write a PRD and grill you with questions until the open decisions are gone. Implementation runs as a two-session system: session one is a ping-pong context machine that ends by compiling a precise prompt, session two is a fresh orchestrator that delegates to implementer, researcher and reviewer sub-agents. Build the landing page first to lock a design system, use pre-built shadcn blocks, and feed the agent reference images. Then review: agent reviewers, your own eyes on every diff, and a PR bot. Deploy entirely through MCP servers.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:00 – 06:18

01 · The wish vs the plan

One-prompt Intercom clone generates thousands of broken lines. Layer-by-layer walk through what the product really needs, a demo of the finished MarshalDesk, and the thesis that workflow beats model.

06:18 – 13:57

02 · Prerequisites: ADE, skills, dictation

Pick any harness with a built-in browser. Install under 100 skills: writing-PRDs, grill-with-docs, Emil design engineering, shadcn, and the feature-orchestrator. Use voice dictation to remove friction.

13:57 – 18:30

03 · Defining the idea and non-goals

Two core features for V1 (live chat widget, RAG support agent), explicit non-goals (email, ticketing, billing, mobile), and why off-topic questions are an attack surface and burned tokens.

18:30 – 28:00

04 · System architecture

Why Next.js (training data, not speed). Widget in an iframe, server enforces multi-tenancy, PartyKit for real-time, OpenAI embeddings, Neon for Postgres, auth, functions, storage and AI gateway. Pre-signed uploads to dodge body limits.

28:00 – 32:00

05 · The three core user flows

Flow A: owner uploads knowledge through storage, function, embeddings, Postgres. Flow B: visitor asks, classify, embed, retrieve, answer, stream. Flow C: handoff to a human when classification fails.

32:00 – 1:00:00

06 · Writing and grilling the PRD

agents.md is for rules, the PRD is for product. Frontier model drafts in one minute; grill-with-docs runs 62 questions across rounds, producing a context.md glossary. Tech stack gets locked in, with a live debate on Data API plus RLS vs an ORM.

1:00:00 – 1:19:00

07 · The multi-model workflow

Orchestrator, implementer, researcher, reviewer roles. Why sub-agents preserve context. Medium reasoning effort over max. Two-session system: ping-pong in session one, compiled prompt executed in session two. Stop bite-sizing tasks.

1:19:00 – 1:47:00

08 · Landing page first, then Neon walkthrough

Build the landing page first to create a design system. Motion Sites prompt adapted to the stack, Geist Pixel font, design.md written. Neon tour: lakebase Postgres, managed Better Auth, 15-minute functions, S3-compatible storage, zero-markup AI gateway.

1:47:00 – 2:10:00

09 · Neon project, plugin and auth ping-pong

Create project with dev and production branches, install the Neon plugin (8 skills plus MCP server). Session one for auth: Google OAuth, magic links, email and password, shadcn login and dashboard blocks, shared auth layout to avoid re-renders.

2:10:00 – 2:27:00

10 · Compile, execute, review

Session two prompt reviewed line by line. Why plan mode is no longer needed for most. OAuth callback fix, overcomplicated error handling challenged with what do you think, password validation gap, dark mode and InputOTP fixes.

2:27:00 – 2:45:00

11 · Three-step code review and PR bots

Reviewer sub-agents, manual read of every diff, then Cursor Bugbot or PullFrog on the PR. Automating bot-comment fixes is now safe. Give the agent verified credentials so it can log in and verify its own work.

2:45:00 – 3:05:00

12 · Reference images and parallel sessions

Gather Intercom screenshots and a GPT-image mockup for the dashboard. Worktree session on Sonnet for auth pages while Opus sub-agents build widget settings and inbox in parallel. Claude Code blocks agents from signing in, so verification was manual.

3:05:00 – 3:31:00

13 · Fighting slop: design mode iteration

Rapid annotate-and-prompt loop on sidebar, widget preview, position selector and inbox. One redesign produced pure slop and was reverted to the earlier version. Customize shadcn components from the inside via variants.

3:31:00 – 3:59:00

14 · Backend: persistence, storage, embed, real-time

Four features in one session-one run. Public bucket for avatars, two PartyServer classes, short-lived room tokens, ESBuild embed script, autosave with indicator. Dev secrets can leak because production uses a new branch.

3:59:00 – 4:17:00

15 · Testing 35,000 lines and the review grind

Seven implementer and seven reviewer sub-agents; reviewers changed 5,000 lines. Full-flow test: autosave, embed on Acme Outdoor Co, real-time typing indicators, handoff. Then 35 minutes of manual review before merging the stack.

4:17:00 – 4:42:00

16 · Designing the AI agent pipeline

Gap analysis against the PRD. Ingest in a Neon function (parallel, 15-minute limit), answer in Next.js after(). Qwen3 embeddings via the gateway instead of OpenAI, GPT-5.4 nano classifier at none effort, GPT-5.6 for answers, speed over cost, Wrangler deploy of the dev worker.

4:42:00 – 5:29:00

17 · Agent shipped, tested and reviewed

Knowledge base uploads in parallel, suggested questions regenerate, off-topic and discount-code attacks declined, fake-teammate injection caught by reviewers. UI fixes, sources shown in inbox, agent toggle removed, tooltips and micro-animations. Manual review finds duplicated helpers and wrong oRPC error handling.

5:29:00 – 5:55:23

18 · Agentic deploy to production

Vercel and Neon plugins via MCP: migrations on the production branch, buckets, credentials, worker and function deploys, Vercel project and env vars, trusted domain. Live sign-up, knowledge upload and real-time chat verified on the deployed URL.

Atomic Insights

Lines worth screenshotting.

  • Build me an Intercom clone is not a plan, it is a wish, and the thousands of generated lines that follow are the agent doing exactly what it was told.
  • The difference between a junior and a senior engineer using AI is the workflow, not the model.
  • An agents.md file is for short agent-specific rules; the product definition, tech stack and user flows belong in a PRD the agent reads every session.
  • A frontier model writing a PRD in one minute took shortcuts; 62 rounds of grilling questions made the same document ten times more specific.
  • Install at most 100 skills; context-hungry models invoke skills that are not needed and thousands of overlapping skills just confuse them.
  • Pick the stack the agent writes most reliably, which means the framework with the most training data, not the fastest or cleverest one.
  • Use only technologies you can read yourself; if the agent cannot fix a bug in a language you do not understand, you are stuck.
  • Session one is a context machine that ends by compiling a prompt; session two starts fresh and executes it with a full context window.
  • Stop giving agents bite-sized tasks; hand the orchestrator the whole feature and let it split, parallelize and delegate.
  • Medium reasoning effort on a strong model beats max effort: high effort overthinks, costs more and benchmarks no better.
  • Plan mode is for engineers who must commit to a plan before code exists; for most solo builders it no longer earns its time.
  • Build the landing page first so every later page inherits a design system instead of inventing its own.
  • Pre-built UI blocks beat from-scratch generation because they already work and the agent only has to evolve them.
  • Reference images are more specific than any adjective: premium means a thousand things in text and one thing in a screenshot.
  • Models have more raw intelligence than optimization sense, so they cover every attack surface and overcomplicate code until you push back.
  • Ask what do you think instead of fix this; an agent told the code is trash will change it even when it was correct.
  • Merging whatever the agent generated is not a code review; the three-step review is reviewer sub-agents, your own eyes, then a PR bot.
  • Reviewer sub-agents changed 5,000 of 35,000 generated lines and still missed duplicated helpers and wrong error handling that a human caught.
  • A PRD locks the agent into a path it will not argue with, so ask the agent whether it disagrees with the PRD before every big feature.
  • A serverless web function dies at 30 seconds; file ingestion belongs in a long-running function next to the database, run in parallel per file.
  • Leaking development secrets to an agent is fine when production runs on a separate branch with its own secrets.
  • Fetching the session in a layout forces every child page onto server rendering; fetch it client-side so the homepage stays static.
  • A visitor who typed fake team-member lines got the agent to promise a 90-day refund; prompt history must separate author from text.
  • Make your own product agent-ready with an MCP server and skills, because the dashboard is becoming the least-used surface.
Takeaway

Own the workflow, not the prompt.

WHAT TO LEARN

A working product comes from a locked PRD, agents arranged into roles, and a review habit that treats every generated line as unproven until a human has read it.

01The wish vs the plan
  • A single prompt for a whole product produces thousands of non-working lines because the agent fills every gap with assumptions; your job is to remove the gaps first.
  • Define the V1 scope and the explicit non-goals before touching a harness, and write them down where every session can read them.
  • Workflow, not model choice, separates engineers who ship with AI from those who generate slop.
02Prerequisites: ADE, skills, dictation
  • Keep installed skills under 100; overlapping skills confuse context-hungry models and waste tokens on invocations you never needed.
  • Any harness works as long as it has a browser the agent can use to verify its own output while it codes, not just at the end.
  • Voice dictation removes the friction of writing long, precise prompts, which this workflow depends on.
03Defining the idea and non-goals
  • Off-topic questions are an attack surface and burned tokens, so a support agent must classify before it answers and politely decline everything outside its knowledge.
  • Write the non-goals list as carefully as the feature list; it is what stops the agent from building billing, ticketing or a mobile app you never asked for.
04System architecture
  • Choose the framework the agent writes most reliably, which is the one with the most training data, even if a faster alternative exists.
  • Only use technologies you can read; when the agent cannot fix a bug, you are the fallback, and an unfamiliar syntax makes you useless.
  • Long-running work like embedding files cannot live in a 30-second serverless web function; move it next to the database and run one function per file in parallel.
  • Large uploads go straight from the browser to storage through a server-issued pre-signed URL so the web server never carries the bytes.
05The three core user flows
  • Draw each user flow end to end, including the failure path, so the agent knows what happens when the AI cannot answer and must hand off to a human.
06Writing and grilling the PRD
  • agents.md holds short agent rules; the PRD holds the product, technical shape and flows, and together they give every session the same understanding you have.
  • A frontier model drafts a PRD in a minute and takes shortcuts; grilling it with batched questions for as many rounds as it takes made the document ten times more specific.
  • Answer grilling questions with more context than asked; each extra sentence removes an assumption and shortens the next round.
  • Lock the tech stack into its own file linked from the PRD, because leaving it open means the agent picks libraries for you.
  • When you do not know the right choice, ask the agent for pros and cons and a recommendation, then override it with your own familiarity.
07The multi-model workflow
  • Separate roles into orchestrator, implementer, researcher and reviewer; sub-agents keep the lead's context clean and let cheap models do the reading.
  • Medium reasoning effort on a strong model benchmarks better than a pricier model at max effort and avoids the overthinking that makes agents sluggish.
  • Run two sessions per feature: ping-pong in the first until the agent understands, have it compile a prompt, execute in a fresh second session.
  • Hand the orchestrator the entire feature from A to Z and let it decide what to parallelize instead of slicing tasks yourself.
08Landing page first, then Neon walkthrough
  • Build the landing page first so the design system exists before auth or dashboards, and every later page inherits it.
  • A pre-built prompt or UI block is a better starting point than from-scratch generation; adapt it to your stack and evolve it.
  • Pick a backend where auth, storage, functions and an AI gateway share one account, and use a development branch with separate secrets from production.
09Neon project, plugin and auth ping-pong
  • Install the provider's plugin so the agent reads current docs through skills and acts through an MCP server instead of guessing from training data.
  • Scope session one to the whole feature plus adjacent UI, since modern models can carry auth, login pages and a dashboard shell in one run.
  • Use existing framework primitives such as shared layouts so only the part that changes re-renders; the agent will not optimize this unless you ask.
10Compile, execute, review
  • Review every line of a compiled prompt before executing it; it should name models, boundaries, mandatory reading and verification rules.
  • Plan mode earns its cost only when you must commit to a plan before code exists; most solo builders can skip it.
  • Challenge overcomplicated code with what do you think rather than fix this, so the agent reasons instead of obeying.
  • Models have more raw intelligence than optimization sense and will cover every edge case unless you push them toward simplicity.
11Three-step code review and PR bots
  • Review in three passes: reviewer sub-agents, your own eyes on each diff, then a PR bot in CI.
  • Bot-found issues are valid often enough that letting the agent verify and fix them automatically is now reasonable.
  • Give the agent a verified test account so it can log in and see the dashboard; otherwise its browser verification silently fails.
12Reference images and parallel sessions
  • Reference images are more specific than adjectives; gather competitor screenshots or generate a mockup before prompting for UI.
  • Use a worktree and a fast model for parallel UI iteration while a stronger model orchestrates implementers on other pages.
  • A harness that blocks agent sign-in forces you to verify dashboards by hand, so pick tooling by what verification it allows.
13Fighting slop: design mode iteration
  • Iterate UI by annotating the element and stating the problem, not the solution; the agent will propose and you will accept or reject.
  • A redesign can be worse than the previous version; say so plainly and point the agent back to the version that worked.
  • Customize shared components from the inside with variants rather than wrapping them with external constants.
  • The agent that already learned your taste on one page is the one to reuse on the next; a fresh session repeats the mistakes.
14Backend: persistence, storage, embed, real-time
  • Combine several backend features into one session-one run once the UI and primitives exist, because the remaining work is wiring.
  • Ask why whenever the agent proposes extra infrastructure, such as a second real-time class, and accept it only when the reason is distribution or scale.
  • Have the agent research the current docs with cheap sub-agents before you approve an auth pattern you are unsure about.
  • Development secrets can be exposed to an agent when production runs on a separate branch with its own secrets.
15Testing 35,000 lines and the review grind
  • Read the handoff document, test the whole flow in a browser, and only then review the code.
  • Reviewer sub-agents rewrote five thousand lines and still left slop, so the manual pass is not optional.
  • Push small, obviously correct fixes straight to main; reserve PR bots for code you cannot vouch for yourself.
16Designing the AI agent pipeline
  • Compare the PRD to what was built before each new feature so missing scope is explicit rather than assumed.
  • A PRD can lock the agent into stale assumptions, so ask whether it disagrees with the PRD before approving its plan.
  • Optimize agent models for speed before cost: reasoning effort matters more than model size, and a classifier should think as little as possible.
  • Verify risky bundling choices, such as an ORM inside a serverless function, with a spike before the main run.
17Agent shipped, tested and reviewed
  • Fetch a user session client-side on a static homepage so a layout-level fetch does not drag every page onto server rendering.
  • Prompt-injection through chat history is real; separate the author from the text in every message block you send to the model.
  • Show the retrieved sources next to every agent reply so the owner can audit why the agent said what it said.
  • Allow local domains by default only in development; in production make every embed domain explicit.
  • Remove controls the product does not need, such as an agent on/off toggle, instead of relocating them.
  • Animated icons and motion belong at rare moments like an upload finishing; a dashboard used constantly should stay crisp and mostly invisible.
  • Duplicate helpers and generic error handling are the most common slop; check that the agent used the library's typed errors.
18Agentic deploy to production
  • Deploy through the providers' MCP servers so migrations, buckets, credentials, worker and env vars are done by the agent step by step.
  • Ask the agent what only you can do before a deploy; the answer is usually domain, OAuth client and plan limits.
  • Never develop against the production branch; dev users and dev data should not exist in production.
  • Make your own product agent-ready with an MCP server and skills, because dashboards are becoming the least-used surface.
Glossary

Terms worth knowing.

PRD
Product requirements document. A written spec of what a product does, what it deliberately does not do, the technical shape and the user flows, used here as the shared context every coding session reads first.
ADE
Agentic development environment, the harness you talk to coding agents through, such as Cursor, Codex or Claude Code. Its main job is to remove friction between you and the agent.
Skill
A packaged instruction file an agent loads on demand to perform a specific kind of task well, such as writing PRDs, applying design taste, or following a particular orchestration workflow.
MCP server
A standard interface that lets a coding agent call an external service directly, so it can create database branches, set environment variables or deploy without you opening a dashboard.
RAG pipeline
Retrieval-augmented generation. Uploaded documents are split into chunks, turned into vector embeddings and stored; at question time the closest chunks are retrieved and handed to the model so it answers from real content.
Multi-tenancy
One application and one database serving many separate businesses, with every query scoped so no workspace can ever see another workspace's data.
Sub-agent
A separate agent spawned by a lead agent for one job, such as research, implementation or review. It runs in its own fresh context window so the lead's context is preserved.
Orchestrator
The lead agent in a multi-agent run. It plans, splits a large feature into parts, delegates to implementer and researcher sub-agents, then dispatches reviewers before handing off.
Two-session system
Session one is a conversational planning chat that ends by compiling a detailed prompt; session two is a brand-new chat that executes that prompt with a full, clean context window.
Reasoning effort
A per-request setting for how long a model thinks before answering. Higher effort costs more and can overthink; the video argues medium is the sweet spot for implementation.
PartyKit / PartyServer
An open-source real-time layer built on Cloudflare Durable Objects and WebSockets. A party is a class, a room is one instance, and the server pushes messages to every connected client.
Durable Object
A Cloudflare primitive giving a single stateful instance per room that can hold WebSocket connections and coordinate real-time messages between them.
Pre-signed URL
A temporary, server-authorized link that lets the browser upload a file straight to object storage, bypassing the request body size limits of a serverless function.
Row-level security
Postgres policies that filter which rows a query can touch based on who is asking, needed when clients query the database directly rather than through a trusted server.
ORM
Object-relational mapper. A library such as Prisma that gives you a typed schema, migrations and a single way to query the database from application code.
oRPC
A TypeScript library for defining typed, OpenAPI-compatible API procedures with contract-first development, used here instead of server actions for scalability.
Stacked PRs
A chain of pull requests where each builds on the previous one, so a 35,000-line feature can be reviewed and merged in reviewable slices.
Scale to zero
Compute that shuts down entirely when there is no traffic and spins back up in milliseconds on the next request, so idle projects cost nothing.
Database branching
An instant copy of a database's schema and data, used here to keep a development branch with fake users separate from the production branch.
Static site generation
Pre-rendering a page at build time and serving it from a CDN. Fetching user data on the server in a shared layout forces child pages off this fast path.
Resources

Things they pointed at.

09:00toolMatt's skills repo (grill-with-docs, writing-PRDs)
09:54toolEmil Design Engineering skill
11:40toolWriting PRDs skill by LennySkills
11:52toolFeature Orchestrator skill (Jan's own)
12:58toolWillow voice dictation
21:50toolPartyKit / PartyServer
24:00toolEraser (architecture diagrams)
1:00:00tooloRPC
1:11:00linkCursorBench ↗
1:25:00toolMotion Sites (landing page prompts)
1:43:48toolGeist Pixel font
2:43:00toolCursor Bugbot
2:43:30toolPullFrog (open-source PR review)
2:49:00toolGPT Image 2.5 (dashboard mockup)
4:17:40linkArtificial Analysis (model speed benchmarks)
5:16:00toolLucide Animated
5:22:00toolVercel plugin (MCP server + skills)
Quotables

Lines you could clip.

01:25
“Build me an Intercom clone isn't a plan, it's a wish.”
Six words that explain every failed one-shot prompt.→ TikTok hook↗ Tweet quote
05:38
“The difference between a junior engineer and a senior engineer isn't the better model. No, it's the workflow.”
The thesis, stated cleanly in one line.→ IG reel cold open↗ Tweet quote
17:14
“Every random question it answers is an attack surface and burned tokens.”
Reframes off-topic handling as a security and cost decision.→ newsletter pull-quote↗ Tweet quote
37:12
“Is this interesting? Is this fun? No. But is this important? Yes. You want your foundation to be as good as possible.”
The anti-shortcut stance on PRD refinement.→ IG reel cold open↗ Tweet quote
49:06
“If there's a bug and the agent is not able to fix it, then I can't do anything. I literally don't understand the syntax.”
The honest case for choosing familiar tech over faster tech.→ newsletter pull-quote↗ Tweet quote
1:59:00
“The honest truth is literally open the thing, open the changes and look at them. It's as simple as that.”
Deflates the search for a code review trick.→ TikTok hook↗ Tweet quote
2:27:45
“Merging whatever your coding agent generated isn't a code review. It's a big mistake.”
Punchy, contrarian to vibe-coding culture.→ TikTok hook↗ Tweet quote
4:11:49
“You are an engineer, not a vibe coder. Review the generated code.”
Identity-level challenge in one sentence.→ IG reel cold open↗ Tweet quote
4:15:45
“There is still definitely at least 500 lines of slop. Agents write slop. It's in their nature.”
Blunt expectation-setting from someone who just shipped 35,000 lines.→ newsletter pull-quote↗ Tweet quote
5:25:00
“The dashboard has become less relevant than ever. Make your application ready for the agentic future. Create your own MCP server.”
Forward-looking product advice from the deploy section.→ newsletter pull-quote↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphoranalogystory
Intercom is a billion -dollar company, and its core product is this. A little bubble in the corner of your website with a bit of AI and a bit of real -time functionality. So, how hard can it be to build your own billion -dollar bubble?
I mean, it's 2026, let's just open Cloud Code. Hey, build me an Intercom clone and make me a billionaire. And now, look at that code generation.
Thousands of files, thousands of lines, with even more code. Does anything work?
No. But does it look beautiful? Definitely.
So what went wrong? Well, let's take a look inside. On the surface, we have a chat window, a text field, and a send button.
Our agent nailed that part. Below it, real -time messaging. Every message has to show up instantly on both sides, not after you refresh.
And below that, an AI agent that answers from a business's actual documents. That means a full rack pipeline, extract the text, chunk it, embed it, store the vectors in Postgres, and retrieve exactly the right pieces.
If we go deeper, then we have authentication, with that hundreds of businesses on one platform that never see each other's data. As a keyword, multi -tenancy. And at the bottom, file storage, background jobs, dev and production environments, code reviews, and all of the annoying stuff.
So yeah, it's a bit more than just a little tiny bubble. And to be fair, that's not the agent's fault. Build .me in Intercom Clone isn't a plan, it's a wish.
So today, we are going to build it the right way, together, from scratch, the way a senior engineer would. Meet Marshall Desk, and everything starts right here, on a homepage with animations that honestly look like they belong on an award site. After that, a business owner signs up.
Google, email and password, or magic link, whatever they prefer. Then they land in the dashboard and set up their agent. Give it a name, an avatar, a personality, and watch the widget up.
live on the right as they type. No save button by the way, everything autosaves. Add your domain, copy one snippet, paste it into your website, done.
Now the important part, knowledge. Drag in your docs, PDFs, markdown files, as many as you want all at once. And this is something that most people, including agents, get wrong.
They process files one after another because it's easier. We won't do that. No, every file gets picked up and processed in parallel by neon serverless functions, extracted, chunked, embedded, and stored as vectors.
And once that's done, the widget automatically suggests questions based on what you uploaded. Now a customer shows up, they ask a question, and the agent answers it from your docs, not its imagination this time. But real customers don't always stay on topic.
Random questions, the weather, math homework, and every now and then someone tries to play the system. None of that works here.
The agent only answers from your knowledge and politely declines everything else. Because a support agent that hands out free discount codes is a liability, definitely not a feature. And when a customer asks for a real person, we classify that message and hand the conversation over to a human, in other words to me.
Which brings us to the inbox. A notification comes in, in real time, no refreshing. Conversations on the left, chat in the middle and on the right everything about the visitor, for example location, language, device, etc.
The support agent jumps in and the customer sees the reply instantly. Same idea as the first attempt, but this time a very different result. But here's the thing, you're not just going to build this with me, you're going to see how to build real products the right way.
We will start with nothing but an idea, then we will design the system architecture and write a proper PID. together, the foundation for our project. After that, I will show you my workflow, how I think, how I scope features, which models I use for which job, the skills I use, and of course how I create my prompts.
Little spoiler alert, I don't use my keyboard anymore. And I won't just show you my personal workflow, I will show you how senior engineers work on real production applications. And therefore, reviewing code is super important.
I will walk you through my three -step code review. process because merging whatever your coding agent generated isn't a code review. It's a big mistake.
And since we want to move fast without cutting any corners, just like Lightning McQueen, we will also use Neon, who is also sponsoring this video. Neon will give us a Postgres database with vector search for our RAC pipeline, managed beta author authentication and authorization, object storage for all of the uploaded knowledge, serverless functions that handle our RAC.
pipeline and also allow us to do everything in parallel and of course an AI gateway for our models. And we will use multiple models like for example an embeddings model, a classifier model, a chat model and much more. And since the coding agent is only as good as the context it has access to, we will install the Neon plugin with its MCP server and skills.
That way our coding agent doesn't have to guess based on outdated training data. It always has the latest Neon documentation and best practices at hand.
And since we will also install the Neon MCP server, we will allow our agents to work in our Neon project directly, meaning creating branches, running queries, checking what's actually deployed, and much more. And that's really the point. The difference between a junior engineer and a senior engineer isn't the better model.
No, it's the workflow. And that's what this video is all about. And since this video is quite long, grab some water or maybe a coffee, stay hydrated, grab your keyboard or use your microphone and please don't forget to like and subscribe.
It would mean a lot to me and my heart and it only takes a second and is completely free. With that out of the way, I will now finish my coffee and let's start building our 1 billion dollar sales together. Alright, as a first step, let's talk about the prerequisites needed for this video and the first one is an ADE.
gigantic development environment, or in other words, harness. And I would recommend to use Cursor, Codex, or Cloud Code. And the reason for that is super simple.
They are the biggest players out on the market, they are the most optimized ones, and with that, in most cases, they create the best possible results. But now you might say, wait, Shan, I'm a big fan of OpenCode, Anti -Gravity, maybe Devon Desktop. What about Pi?
Should I now switch to these ADEs? No. Use and choose whatever you feel most comfortable with.
If it's Pi, then use Pi. If it's anti -gravity, then use anti -gravity. The goal of an ADE primarily is to remove friction as much as possible.
It should be as easy as possible to interact with your agents. If that means that you have to use something else than these three listed ADEs, then go for it. No subjections from my side.
Now, yes, the harness does play a role, and you will see that every harness kind of creates a different result but it's marginal in my opinion and it's not really something you should use whenever choosing an ADE. Now if you choose a different ADE then please make sure that the tool has a built -in browser or at least some sort of browser tool for example browser use and that's very important because later on we will have our workflow the coding workflow and one core step will be that we will let our agents verify their own work and that's only possible if they have access to a browser.
Again, preferably a built -in browser, but a browser tool also does the job perfectly fine. So now we talked about the ADE, let's also talk about skills. In this video, we will use skills.
In my opinion, skills are one of the most powerful things you can use as an AI engineer, and they have been a game changer. But here's something you should know. A lot of people have a certain misconception about skills, because skills, in theory, are only dynamically invoked.
will only use a skill if it makes sense. So what a lot of people do is literally install every skill under the sun and then at some point they will have thousands of skills installed. This is a huge mistake, please don't do that.
Because yes, agents invoke skills dynamically nevertheless some models are kinda context hungry. And what you will see specifically with GPT models is that they will invoke skills that are not even needed. And if you now have thousands of skills installed which kinda all do the same thing then the model will essentially be confused.
So my recommendation is to install a maximum of 100 skills, not more. keep it at 100 with that you will preserve the context window of your agent, the tokens, and at the same time your agent will also be able to make better decisions. Now let's talk about the skills I would recommend to install.
First of all, the ones created by Matt. I think you all know who Matt is. It's this cool guy.
He has a great YouTube channel and he created a lot of very powerful skills. These skills are essentially broken down into two different groups. First of all, user invoked and then also model invoked.
What I would recommend is to install all of them, but in this video we will primarily use model invoked skills. We won't use toTickets, toSpec, Implement, it's not really needed. The only user invoked skill we will use in this video is GrillWithDocs once we create the PRD.
So please install all of these skills, please also check out the GitHub repo, it already has almost 270 ,000 stars, so yeah, great job Matt. Then I want you to also install the Emil Design Engineering skill.
So this is a skill created by Emil Kowalski. Emil is an ex -employee of Vercel. He now works, if I remember correctly, at Linear.
He's a super talented designer and he also created this skill right here, Emil Design Engineering. And the main idea of this skill is that it consists of mainly animation, but also some design advice. It teaches your agent about taste because what we don't need is AI slop.
We don't want this blue... purple gradient. We don't want 10 ,000 animations.
We want you to have a high quality website, high quality front end result. And this skill enables our agent to create tasteful designs. He also created a few other skills like animate, improve animations.
You can also install them, but primarily make sure that email design engineering is installed. Then I want you to also install the shared CNUI skill. And that's because we will use shared CNUI in the video.
And yes, agent, already know a lot about shared CNUI. Everything is already embedded into the training data.
Nevertheless, things change quite often and it's always a good idea to give your agents the newest information possible. Then I want you to also install the writing PRDs skill by LennySkills. This is a skill which teaches your agent about how to create PRDs the right way.
And you will see here in the description that the main goal is to help users, or in other words agents, transform apps ideas into actionable project specs that align engineering and design teams on the problem and success metrics. This is a very important skill that we will use once we get started with the foundation, with our PRD.
So now we talked about all of the skills provided by Matt, the design engineering skill, writing PRDs, and also the shared CNUI skill. And then there's one more skill I want you to install, and that's the feature orchestrator skill. That's my own skill, and this skill is very important.
Because in a second, or once we get started with implementation, we will use a very specific workflow, my signature workflow. And this skill holds all of the instructions for this workflow. I always use two sessions, I have an implementer, a researcher, an executor, stuff like that.
So please install the skill, I will explain everything a bit later on, but it's very important that you have it installed. So now we talked about the ADE, the skills you should have installed, and then there's one more thing I want to mention. that I want you to use a voice dictation software.
Because here's the thing, keyboards are great, don't get me wrong. Oh my god. It sounds so good.
This, by the way, is the keyboard I got when I went to San Francisco, which is quite cool. But now let's come back to the actual point. Why do I want you to use a voice dictation software?
Well, it's relatively simple. A keyboard creates friction, and we want to remove friction as much as possible. Using your voice is the easiest way to remove friction.
Later on, once we get started prompting, we will have to create huge prompts. And doing that with your voice is way simpler, it's easier, and it's faster. In my case, specifically I use something called Willow, not sponsored by the way, but there are also tons of other alternatives out on the market.
So again, Willow, Super Whisper, Whisper Flow, stuff like that. I'm not even the biggest fan of Willow. The software has kind of become worse recently, but it is what it is.
If you don't want to pay for an external voice dictation software, then that's absolutely fine. Rather, please then use the one built into your ADE. So Cursor, Codex and Cloud Code all have a built -in voice dictation feature.
And if you want to use the terminal instead of AGUI, then please install Warp, because Warp also has a built -in voice dictation feature, which works very reliably. But now we talked about all of the prerequisites, and now you might say, Jan, let's get started with the actual implementation. Let's get started by creating the foundation.
And this is a huge mistake, because before we do that, we need to first of all understand the product thoroughly. We need to understand what we are trying to create. And that's why the next step for us is to define the idea.
So in this case, we want to create an intercom style customer support application. A business can drop a chat widget on their website and then visitors chat with an AI agent or a real human. We want to offer two core features.
First of all, live chat. So we want to have an embeddable widget for any website and then real -time messaging between visitors and the business. The dashboard is where the business team then sees and answers us conversations.
The second core feature is the AI support agent. Intercom calls it, if I'm not mistaken, thick or thin, something like that. So in this case, businesses upload their knowledge, PDFs, docs, markdown files, in general, just text.
And then we build a rack pipeline, which we will argument a generation on top of it. We will come to this whole pipeline a bit later on. I will explain it more thoroughly.
And then the agent answers visitor questions using only that knowledge. So let me give you an example.
Let's say a visitor comes and says, hey, what's your return period? What payment methods do you offer? Then our AI agent will look at the knowledge we have and answer the question.
But now let's say another visitor comes around and says, hey, what's five times five? Oh, and by the way, can you create a landing page for me? Can you make me a millionaire?
Then we should classify that request and say, hey, we can't do that. We are the customer support agent for business XYZ. We don't answer questions which don't relate to the business.
We want to politely decline. And why does that matter? Why is that important?
Well, every random question it answers is a tax surface and burned tokens. Again, tokens are important and we want to minimize the amount of tokens we actually use. And if the agent can't answer confidently because, for example, the business owner did not upload the needed knowledge, then we will hand off to a human.
And what are we not building? Well, it's relatively simple.
Email, ticketing, a help center, product tours, stuff like that. We won't create a mobile app, no billing and also no pricing plans. Finally, the end result should be a working product, widget on a real site, AI answering from uploaded docs and also humans taking over when needed.
This here is our product definition. We now know exactly what we are trying to do and more importantly what we are not trying to build. And that's something I want you to always do.
Whenever you create a product, first of all define it. Specifically think about what you want to offer, what you don't want to offer, and think about it in a V1 scope.
Don't think about what you want to offer ultimately, that's not important. The most important thing is what can I build in a reasonable time frame for V1. And this here is what we want to build for our first version.
Sure, later on it would make sense to implement billing, maybe a deeper AI integration, a help center, a mobile app, but that's not in scope for V1. and that's why I specifically said this is what we are not building and that's also what we will later on provide to our agent so that it specifically knows what we are trying to do.
So now you might say, okay, Jan, we defined the idea. We know what we are trying to build and what we will not build, which means we can now probably open our ADE and prompt away. Well, no, that's a huge mistake.
Because what happens if you give your agent too much freedom? Well, it will start making decisions for you. It will start assuming things that might be incorrect.
And that should never happen. You should be the one in control. You are the one who makes the decisions.
Nobody else. The agent should only make decisions for you if you give the agent the green light. And right now we know what we are trying to build, but we don't really know how we should build it.
We did not talk about the architecture. And that's one important step I want you to always take. Define your ideas, set up all of the prerequisites, and of course define the architecture.
What technologies do you want to use? What technologies won't you use? Make a little pro and cons list.
the thing what i often see people do is they will decide let's say for example on a database postgres versus mysql and they will literally sit there and dwindle for hours maybe even days they will go through google and through reddit for example and then search for postgres versus mysql what's faster what's lower what's cheaper what's more expensive what's better what's worse blah blah blah Don't do that.
You don't want to sit there for weeks on end and define or set up your architecture just for your product to fail after you launch it. Think about the architecture and then create a diagram. And that's exactly what I did right here.
I already defined the architecture for Marshall Desk, our application. So let's go through it. On the left, we first of all have the customer's website and then also the dashboard.
So the customer's website or in other words, the chat widget is what... visitor will see and then the dashboard is what the business will interact with.
So the chat widget lives on the customer's website inside an iframe, so inside of the widget, and it's also isolated from the entire business website. So it's still our application, it's kind of rendered on the customer's website, but it's our infrastructure, our data, stuff like that. And then visitors can interact with the business.
Then the dashboard, this is where the business owner and team also log in, change settings, upload knowledge, and then also answer chats. Let's now continue with one of the most important parts of our application, the server, and in this case, we will use Next .js.
And I know exactly what you want to say right now, but Jan, there are so many better alternatives out on the market. What about SvelteKit, TanStackStart, or maybe Rust, which is super fast, super optimized, and also super complicated? Well, the reason is simple.
Who will write the code? Will you write the code? No, we will use Coding Agent.
Therefore, it's important to choose the technology that the coding agent understands and at the same time where the coding agent is able to generate reliable code. And next, JS is a huge framework with a lot of resources, a lot of YouTube videos, I mean look at my YouTube channel, a lot of blog articles, very good documentation and in general the resource or the amount of resources is huge.
And how do coding agents work? How do models work? Well, they get trained on training data.
And Next .js specifically is very embedded into training data. So even though SvelteKit as an example might be the better framework, the faster, simpler framework, it's still something I wouldn't choose because agents don't write reliable code. And I don't have the time to check every line of code, supplement all of the logic myself, meaning the needed info from the documentation, I don't want to do that.
And Next .js is the perfect candidate for that, even though it's maybe a bit more complicated, Maybe a little bit slower, but that does not matter in the grand scheme of things. It's still the best player out on the market for our specific use case.
Then what will our server do? Well, it will check, hey, who's making the request? Visitor ABC or maybe business ABC.
Then our server will also keep the query inside one workspace. So one company will never see another company's data. And this is what you call multi -tenancy.
So if you don't know what multi -tenancy is, then please check out this video. but essentially we will have one application one server as you see here with one database but we will have multiple businesses and we don't want to leak data and our server will make sure that if for example business A makes a request right here that we don't add or return data from business C we want to segregate the data and then finally our server will also run the AI logic essentially it will orchestrate everything inside of here you will also see that we have something called PartyKit.
What is PartyKit? Well, PartyKit will handle all of the real -time communication. Because when a visitor makes a request, when the visitor types something into the widget, we want to send the message to the business, to the dashboard.
And we want to do that in real -time. And PartyKit will enable that. So the server will publish a new message, as you see here.
And then PartyKit will push these messages, meaning to the chat widget, as you see here, and then also to the dashboard. And if you don't know what PartyKit is, then that's fine.
PartyKit is a tool which got acquired by Cloudflare. It uses durable objects, it uses web sockets, so we get a message and then we distribute it to all of the connected clients. I will explain PartyKit a bit later on, but the biggest benefit of PartyKit is that it's not a third -party service.
It's something you run on your own infrastructure. It's extremely cheap. In our case, it will be free because we will use the free tier.
And yeah, in my opinion, it's the best alternative out on the market. I also create... a great video on real -time infrastructure so check out this video right here.
I found the best way to build real -time applications and there I explained how PartyKit works, how durable objects work and all of the fancy stuff. Then what else do we have right here? Well we also have OpenAI Embeddings.
So what we will do is turn text into embeddings. When a user uploads knowledge, for example a PDF document, we need to create embeddings and for that we will use an OpenAI model.
We will discuss the model a bit later on, but this is also a very core part of our architecture. Then, what else do we have right here? Well, we have our backend.
And for the backend, we will use Neon. We will use Neon for our database, Postgres. We will use Neon for authentication, better off.
We will use the AI gateway to classify and answer questions. We will use functions to ingest the knowledge. In other words, to make a request to OpenAI to embed everything.
And finally, we will use object storage to store somewhere all of the uploaded files, the PDF documents, the docs, the markdown files, etc. So this here is what you call a backend as a service. So instead of gluing five services together, we will use one service, one account, one dashboard, one set of credentials, one SDK, etc.
And at the same time, we will also let our coding agent connect to Neon. We will install Neon skills and then also use the Neon MCP server.
So we will be able to interact with all of these services right from our ADE, right from our agent. And let's also quickly look at the blocks individually. First of all, authentication.
This is where owners and the team log in, not the visitors. The visitors don't have to log in or sign up. Then we have the AI gateway.
Our LLM model calls go through this gateway and we also use this gateway or the models to classify and also answer the questions. and we will use an AI gateway because it gives us the ability to choose between different models. So for example Neon will allow us to use GPT models, Claude models, Muse, stuff like that.
Then of course we need a database. In this case we will use Postgres plus pgVector so we don't need a separate vector database. One thing you will see here is that I also added functions and now you might say Jan what are functions and why do we need functions?
Well a function in this case will act as a background job that processes uploaded files. What you have to understand is that we will use Next .js and we will deploy this application to something called Vercel.
Vercel is the deployment or hosting provider. Vercel uses serverless functions in the background. So once we deploy our application, a serverless function will spin up, serve the request and then it will again shut down.
This is how serverless functions work. So our server can only run for a maximum of 30 seconds maybe one minute if you are on the pro tier and that's a huge problem for us because this embedding step is an asynchronous task it could take only one second it could take a minute if the data is huge if the data set is huge so we need something that can run besides our application besides our server which can then take the request do all of the needed let's say processes and then again shut down and this is what our function here will do it can do of the necessary stuff and also it allows us to do everything asynchronously so we don't have to do it in the main process on our main server we can do it besides our server so while we do the asynchronous embedding we can still allow our user to do something else in the dashboard as an example what you will also see here is that we will use object storage and that's because our user the dashboard or the business owner will have to upload knowledge and we will have to store the knowledge somewhere
PDF documents, the markdown files, the normal docx files, etc. We will store them in a bucket, in a private bucket. And since we will use Vercel, we will also have to use pre -signed URLs.
Because normally what you do is just use your server. But since we will deploy our application to Vercel, we have a body size limit. So with Vercel, the body size limit, if I'm not incorrect, is about 4 megabytes, maybe even bigger if you are on the pro tier.
And that's fine in most cases. for us but what is if the business owner wants to upload a 50 megabytes document then it will fail a serverless function can't upload 50 megabytes it's not possible that's why we want to upload the file on the client side the user the client right here will request the pre -signed url on the server side we will then serve the pre -signed url back to the client and the client will upload everything on the client side there is no limit on the client side it's unlimited the business owner in theory could upload a file with 10 ,000 gigabytes and it would work.
So the main benefit here is that we don't have any limits. And since we will generate the pre -signed URL on the server side, we will also be able to make sure or authorize the user, because only authorized users should be able to upload files. But yeah, that's the architecture we will use.
Later on, once we start implementing the individual technologies, meaning the database, authentication, our AI gateway functions, et cetera, I will also explain them more detail and how they work. One thing I want you also look at is the user flow.
So let's look at flow A. Owner uploads knowledge. Everything starts with the dashboard.
The dashboard asks for an upload URL because we want to upload something. We then go through our object storage primitive. We upload the file directly on the client side, not the server side.
Our function then parses everything and creates chunks. We will then use OpenAI to embed the chunks and find Ideally, we will store everything in our Postgres database, meaning we will save the chunks and also the vectors.
This is when a user wants to upload something. Let's look at Flow B. Visitor asks a question.
First of all, we have our widget. The user, the visitor, sends a message. Hey, what's the return window?
The second step is that the server gets the request. It checks who the visitor is. Visitor 1, Visitor 2, Visitor XYZ, and it also checks the limits.
As a third step, we then have our AI gateway. We will use a model, let's say Opus 5 .5, and Opus 5 .5 will classify the message, because we don't want to respond to every request. If a user asks us a question, which is, for example, what's 5 times 5, or how's the weather in Germany, then we don't want to answer that.
We want to decline respectfully. As a next step, we will then use again OpenAI. to embed the question.
Postgres, our database, will be then used to find the best chunks, whatever matches the request, and then we will again use the AI gateway to answer from the knowledge. This is what you call retrieval argument generation. We have a rack pipeline.
We get something or the user uploads something, we pass everything, we create chunks, we then use an embedding model, an OpenAI embedding model, and then we save the chunks and vectors in our database. Once the user then makes a request we can look at all of our chunks we find whatever matches the request and then we answer based on these chunks this is a rack pipeline and then finally we will use party kit in other words web sockets to stream the responses to the widget in real time so this looks all good to me and finally we have flow c handoff to human as already mentioned if our classification fails in other words if for example the model sees that the question is not really related to the business, or if the model can't really answer the question based on our chunks, then we will hand off to the human.
So as you see here, AI answering default state, waiting, unsure, or visitor asks, then we will hand off to the human, AI stops replying, and then finally the human can close the request, the conversation, which means everything was successful. And yeah, that's already it. This is our architecture, this is what we will use, and this is how everything will work.
And with that, we are now finished. with the foundation. We defined our idea.
We know what we want to build, what we don't want to build. We looked at all of the technologies we want to use. We also looked at all of the core primitives and how they will connect to each other, as you see here.
And finally, we already mapped out the core user flows. This is very important, and this is also something we want to store in our PRD. Our agent has to know exactly how the proposed flow should look like, what a user should be able to do, what a business owner should be able to do, and what should happen if the classification fails, or in other words, if we have to hand off to a human.
So now you might say, okay, Jan, you know what? I guess we are now done with the foundation because we defined the idea, we looked at all of the core primitives, how we want to connect everything, we created a diagram, an architecture diagram, we looked at the various user flows. We are finished.
Let's start with the homepage or authentication or something like that. Well, here's the thing. Why did we look at all of these small things Like, why did we define our idea?
Why did we create an architecture diagram? Why did we look at the user flow? Is it because I wanted you to understand what we are trying to build?
Well, kinda. But not necessarily. That was not the end goal.
The end goal or all of this leads up to a PRD. We want to create a PRD. That's the ultimate goal for our foundational set, if that makes sense.
So I think you guys all know what an agents .md file is. Essentially, it's a plain markdown file placed at the root of a software repository to give AI coding agents project -specific instructions and rules. But an agents .md file is not really there to explain your project, the technical shape, the user flows, what you want to build for v1, what comes after v2, what you don't want to build.
It's not there to do that. An agents .md file should be relatively short, it should be relatively precise, and it's just there to give coding agent -specific instructions. So we will have an agents .md file, but the PRD file, which we will create in a second, will solve the other rest.
It will give the agent -specific instructions about our product. What we want to build, what we don't want to build.
The technical shape. Next .js, Neon, our Postgres database, better off for authentication, stuff like that. We will also explain the user flows, the ones we just mapped out.
And with that, all of the coding agents will have the same understanding as we both do. Because right now, coding agents are essentially blind. They will have an agents .md file.
They will know, hey, we are in the Next .js project. Oh, this here is maybe Prisma. Oh, this here is, for example, better off.
Oh, you know what? Let's maybe use the browser. Yes, the agent will...
understand that. But the agent won't necessarily know what we are trying to do, what we are not trying to do. Because the goal here is to give the agent guardrails.
We don't want the agent to assume things. We don't want the agent to do things that we don't want. And a PRD solves exactly that.
And pairing that with a good agents .md file is a total game changer. So I want you to remember one thing. Whenever you start from zero, literally from nothing, define your idea, do all of the necessary stuff to define the architecture and then take all of these learnings and combine them into a PRD file because the PRD file will be there to supplement all of the necessary logic to your agent and with that the agent also won't have to assume things and not ask you questions because we want things to be as autonomous as possible and it's only possible if we provide answers to the agent without the agent having to question us and without any further ado let's now finally create the PRD and And instead of doing it manually, which is super boring and time consuming, we will use a coding agent.
That's why I already opened Cursor, I created a new folder, Marshall Desk, which is now also my workspace, and I selected a model. And please select a Frontier model. When recording this video, the latest Frontier models are Opus 5 .5, Fable 5 .1, and GPT -6 Astra.
Once you watch this video, we might already have Opus. 6, 7, Ultra, High, Max, whatever. Please just use a frontier model.
And the reason for that is super simple. The PRD which we will create in a second will be the foundation for our project. It will be the backbone.
And that's why we need as much intelligence, raw intelligence as possible. And frontier models deliver exactly that. Because our PRD should be as good as possible.
So that's why I will use Opus 5 .5 with medium reasoning effort. Then inside of here I will use the writing PRT's skill as mentioned. This is a prerequisite.
And let's also add some context. So let's go back first of all to Notion. I will copy this whole idea definition.
Again this is also all linked down below. I will have a GitHub repo with all of the necessary files, resources, stuff like that. I will go back, paste it inside of here.
And what I will also do is screenshot what we created. Meaning our system architecture. and then the user flows.
So let's just use the screenshot tool. I will screenshot this quickly, bam, and I will go back to cursor and then just add it inside of here. Let's do the same thing for the user flow.
Again, I will screenshot everything, capture, and then also paste it into cursor. Now this here is something I wouldn't have recommended in the past because in the past models were very unreliable when analyzing images. But this has now finally changed.
Models are very reliable when looking at images so you don't have to type everything out word by word. So now we added a lot of info, a lot of data, but this is kind of without context.
What are we trying to do? Why did we add this info? So let's create a basic prompt.
Again, I will use Willow, my voice dictation software, and I will say the following. Hey there, I'm currently working on an application called Marshall Desk, which essentially is an intercom style customer support application with a live chat feed. and an AI agent feature.
Down below I added a lot of context about how the application should work, what features I want to include, what I don't want to build and stuff like that. At the same time I also uploaded two images. First of all one which explains the technical architecture and the core technologies I want to use and at the same time also a screenshot which maps out the core user flows.
So flow A, flow B and flow C. What I want you to now do is use the writing PRDs skill to create a high quality PRD.
So we are now finished, this is our prompt and what I will now do is click on enter. And as you also just saw a second ago the prompt is relatively basic, it's relatively short. But that's absolutely fine because we already added a lot of context.
We added our like simple PRD as you see here and at the same time we also added screenshots. You don't have to explain everything one more time with your own custom them prompt.
If the agent has questions, it will ask you. Models have become smart enough, especially frontier models. Our agent is now finished and it created a PRD.
And now you might say, Jan, since we are finished with the foundation, since we created the backbone of our application, we can now get started with the implementation, right? No, no, no. Not so fast, my friend.
This is one of the biggest mistakes you can make. Because here's the thing. Is this interesting?
Is this fun? No. But is this important?
Yes. You want your foundation to be as good as possible. So instead of now just shipping this, we want to refine our PRD.
We want to make it better. Because yes, we used the Frontier model, we used Opus 5 .5, but do you think Opus 5 .5 used its full intelligence power? No.
It took shortcuts. It wanted to get to the result as fast as possible. It took, what, one minute to create this PRD?
That's not enough. Also how many lines do we have here? 312.
This is not huge, this is not small, but I'm 100 % sure that this is not as detailed as it should be. So how can we fix this? Should we now open the PRD and like manually inspect it and update it?
It's a good idea, but we won't do that. Instead, we will use a skill. As already mentioned, one prerequisite was to install all of the skills provided by Matt.
So what we will use here is the grill with dogs skill. I will say grill with dogs and I will then add the following prompt. Great job.
Nevertheless, I want you to now grill me in regards to the PRD you just created. Ask me as many questions as needed because I don't want you to make any assumptions for me. What the agent will now do is use the grill with dog skill, and it will ask me questions in batches.
This is how the skill works. In the past, it was like the agent always asked you just one question per run, but now we will get questions in batches. This is now round one of our grilling session.
The agent in total created or prepared 10 questions for us, and we will go through them together and answer them right away. So I will turn on my voice dictation. software and then answer the questions in a conversational way so that you also understand what they mean.
So who can sign up for V1? Well, we want to allow businesses to sign up, but here's the thing. Normally a multi -tenant application also allows businesses or workspaces to have multiple members by inviting new members, having roles and permissions, also removing members and all of that fancy stuff.
This is something I don't want to implement for V1. should be relatively simple. So a business owner can sign up, we will create a workspace, in other words an organization for this business owner, and then this organization will have one member, the person who signed up.
And this means that later on, once we then create V2, we will instantly already have the foundation set to allow business owners to invite new members, set roles and permissions, and do all of the fancy stuff. Then question two, who is V1 really for? Here we already have like a recommended answer.
This is what the agent recommends. Let's quickly check. A is what a real product for paying or pilot business.
A fully working product built as a showcase. We want to go with the recommended answer, meaning B. So yes, this is a full -on product.
It should fully work. It should be fully operational. But we don't want to actually charge customers or anything like that.
So this here is essentially V1 of a real product, which maybe can then later...
Okay, this all looks good to me. What's the recommended answer? On topic, small talk, off topic is everything else, including attempts to manipulate the agent.
Yes, I agree with the recommended answer. Then question four, what happens when the AI decides it can't answer? Two options, automatic, the AI says, I couldn't find that, and then what's...
Option B, offered, the AI says, I don't know, want me to connect you with a human. I would actually go with the recommended answer. So A, automatic, your flow, C diagram says unsure, waiting, and it saves the visitor a step.
That sounds great to me. Let's go with the recommended answer. Question five, what does the visitor see when they are waiting and nobody is online?
That's interesting. What do we want to do? So what are the recommended or what are the options?
A, the conversation. conversation stays waiting indefinitely? Yeah, I think that can work.
What else do we have here? B, after a timeout, the AI takes over again. No, I don't really like that.
And C. Business hours. No, I also don't like that.
Let's go with the recommended answer, meaning A. And yes, later on we can maybe do C, but A sounds like the best solution. Question six.
Where does the conversation start and end? Can a visitor have more than one open conversation at a time? No.
Let's go with the recommended answer. A message after closing starts a new conversation. Correct.
Question seven, can a teammate hand a conversation back to the AI? Definitely. So if the business owner has a conversation with the visitor and wants to again give the conversation back to the AI, then we should have a button for that.
As you mentioned, to return to AI, maybe not return to AI, but something like hand off to the AI or something like that sounds a bit better or agent. Then question eight, are visitors always anonymous? Yes and no.
So the default is an anonymous state. But what I want to also do right away is log like the user's IP, maybe not necessarily IP, but like the location, the time for the visitor, stuff like that. Maybe for V1, we don't really need to authenticate the user.
That's not needed. But I do want to get basic data from the user, which is still anonymous. The user will stay anonymous.
But yeah. Then question nine, which file types can be uploaded? What's the recommendation?
answer pdf docx markdown and txt i would say pdf definitely markdown and txt also docx could be a bit complicated so you know what let's put docx as like a little Let's defer it. Once we get to the implementation, I will decide on it, but 100 % we want to have PDFs, Markdowns and also TXT files.
Let's maybe also say that the user can add raw text, so we will have text input, which we will also be able to then use. And question 10, what language does the AI reply in? Well, it's relatively simple.
Let's do the following. Since LMs are smart, we can actually reply to the visitor in the visitor's language. So we will choose English as the default, but if the visitor asks us something in German, then the agent will automatically also respond in German.
That seems quite smart and quite intelligent to me. I now went through all of the questions, and not only did I answer them in the basic form, I also added more context. Because yes, I could just say, hey, recommendation A sounds good, recommendation B sounds good, option C sounds good, but this does not add enough context.
precise and that's what I did right here. I added more details than needed because this will also save us a bit more time. The agent will ask us less questions and at the same time using this newly gathered context it will be able to I guess refine its question set or its questionnaire because in a second once I click on enter it will prepare new questions round two and round two will build on this foundation on the answers I just added.
What you will see here is that the agent now creates a context .md file and what is this file and what does it do? Well this here essentially is a glossary.
We have right here this term workspace and this context .md file explains what a workspace is, what a member is, what an owner is, what a visitor is, etc. What you will also see here is that the agent automatically updated our PRD with the new answers I gave it. So as you see here who is v1 for and stuff like that.
We have a lot of lines. changed and the agent now also prepared round two. I will now go through all of the questions myself.
You will also do it and if you don't want to go through this manually then that's fine. I will have a finished PRD file down below or linked down below in the GitHub repository which you can just use right away. And one thing you might also ask me is Jan, how long should I go through this grilling session?
Should I just answer 10 questions, 20 questions, 30 questions? Well the answer is Do it as long as it takes.
It's as simple as that. If the agent only has two rounds for you, great. If the agent wants to ask you 100 questions, great.
Please give your agent the ability to ask you questions. The foundation you will build right here will allow you to build out features later on exponentially. It will speed up literally everything.
I'm now finished with the grilling session and in total over 60 questions have been asked. To be exact, 62. and that's a good thing because our PRD has now become Literally 10 times better.
It has become more specific, more detailed, and at the same time, it has also created less open questions, which is exactly what we wanted. So let's quickly see what our agent said right here. Docs .prd .md is rewritten.
Great scope. Sign -in conversation. That's all good.
Context .md has 20 glossary terms. That's good. Four things are left for implementation on purpose.
That's also fine. Do the PRD. and glossary match what you have in mind.
Yes, I checked out both files. They are great. So what I will do is say confirmed right here.
One thing I do want to update is the technical scope because I did check the PRD and it was not really clear what technologies we want to use. That's something I want to log in right at the start. The PRD should mention exactly what technologies we want to use.
Next, JS, PartyKit, Neon, PostQuest. data, API, stuff like that. So I will now say the following.
Great, thank you. The PRD is perfect. Nevertheless, I want to now also map out the technology stack in detail.
So I want to use Next .js for the front end and back end. I want to use Taewon CSS and Shared CNUI. I want to use PartyKit for the real -time stack, including WebSockets.
Then I want to deploy the application to Vercel. I want to use Neon for my backend as a service.
So this means I want to use managed authentication, which essentially will be better off. I want to then also use Postgres with vector search. That's also what we already discussed.
I want to use functions. I want to use the AI gateway, object storage. Then instead of using an ORM like many people do traditionally, I want to use the data API offered by Neon, which also means we will have to think about RLS.
row level security. Is there anything else I have missed? Is there anything else you want to ask me in terms of the technical shape?
And this response instantly shows you why I wanted to log in the tech stack right away. An AI library in the code to call the gateway and stream replies. This is something I completely forgot about.
Because here's the thing, we will use the AI gateway. The AI gateway will expose models like Opus, 5 .5, GPT -6, Muse, stuff like that. But we need to interact with these models somehow, right?
We need some sort of intermediary library, which could, for example, be the AI SDK. And if I wouldn't lock this into this tech stack or into this PRT right away, then our agent would just assume something.
It would make a decision for me which might be incorrect or not wanted. So let's go through the questions. Question one, who talks to the database and how?
As mentioned, I want to use the data API and the data API is interesting because how does it work? Well, we have a Postgres database and the data API will create APIs based on the database and based on the queries we will create.
But the problem with the data API is that we will create a connection to the database on the client side. So to secure our database we need row -level security. If this is a bit confusing then maybe quickly spin up ChatGPT, ask ChatGPT about RLS and how it works and what's the suggestion here.
So row -level security guards the one place where a user touches data directly putting anonymous visitor checks and background drops into SQL functions would be the hardest part of the codebase for little gain. Hmm.
So option A was everything through the data API. Option B is what we just mentioned. So let me right away use my dictation software and I will say the following.
Question one is very interesting because that's something I haven't thought about. I'm not quite sure. Like I want to use the data API.
We will have to use row level security. I guess splitting everything makes sense. So the owner's dashboard uses the data API with RLS.
The trusted server paths will... then use plain SQL. Nevertheless, let's quickly also debate on using an ORM.
In your opinion, what's better, using the data API or rather an ORM? What will be more scalable? What will be simpler in the long run?
Then question two, what shape do the security policies take? Well, if we use RLS, then I guess we can use one helper function. Yeah, let's go with your recommendation.
Question three, where do the dashboard's data API calls run? This is an interesting question. Because yes, we can do it either on the client side or on the server side.
And I would say let's make it... somewhat agnostic. What I mean by that is we want to do both.
We want to do pretty much 80 % of the calls on the server side, truly on the server side, though I would like to keep the option open to also maybe do it on the client side if needed. So we will use server components in most cases, though let's keep the option open to maybe also do it on the client side a bit later on. Though if this creates too many implications, then please let me know.
Question four, how are changes managed without an ORM? That's again an interesting thing and that's one problem with the data API.
That's something I haven't thought about that much. That's why let's also come back to question one. Let's quickly debate on it.
Should we rather use the data API and RLS or rather an ORM or in combination as you mentioned with point B, drizzle only for schema definitions and migrations? So yeah, let's debate on that a little bit. Question five, how is the repository structured?
There are three things to deploy. That's correct. I would say, yeah, let's use PNPM workspaces.
Turbo repo is not really needed. PNPM workspaces is enough for our scope. Which library calls the models?
Yes, let's use Vercel's AI SDK. If I'm not incorrect, we probably already have version 7, which is the latest one. So double check on that.
And then question 7. How does the agent's reply stream to the visitor?
Well, I would say we should probably go with A. The next JS route generates the reply and publishes each chunk of the text to the conversation's PartyKit room as it arrives. The widget and the dashboard both watch it appear live.
This makes sense to me. So it matches also our diagram. Let's go with your recommendation.
Question 8. How are PartyKit connections secured? Well, yeah, what you mapped out here sounds great to me.
All three as described. dashboard connects with the owner's NeonAuth token, then the widget connects with the visitor token, the next JS server publishes events to rooms over HTTP with a shared secret. Question nine, oh, and maybe by the way, one thing I want to mention here is, let's also keep this a little bit open, because once we get to implementation, I want to let the agent, the implementing agent, to also research that a bit.
just to find out what the best practices are in terms of the party kit side. Then how are sources parsed and chunked? Well, with unPDF, that sounds good to me.
Let's go with your recommendation. Finally, how is the widget built? So I want to use our working application, our Next .js application, meaning I want to go with your recommendation.
One design system and one deploy. Only the embed script needs its own tiny build. That's fine.
And then question 11, the smaller defaults. No, I don't agree on everything.
We want to use Next .js 16 with the app router, no source directory. Let's use TypeScript in strict mode. We will use PNPM.
Then we will use sort for validation, react hook form. That's all correct. Then we will use stream down.
Sure, why not? Vercel functions, why is this needed? I'm not sure if this is a good idea.
Explain that. Then no, I don't want to use biome or biome. I want to rather use ESLint and Prettier.
That's better in my opinion. And then also finally, question 12, where does the stack get written down? Yeah, let's create a new file and that's also linked to the PRD.
Or in other words, it will be linked from the PRD and it's something I want you to also call out so that the agents will later on know, hey, this is the tech stack. So as you just saw, I myself don't know everything, meaning I don't know if I should use the data API or rather an ORM. And that's why I asked the agent specifically, what do you recommend?
Maybe create a small pros and cons list. Tell me what you would suggest, what is more scalable, what will create a smaller headache in the long run. So let's see, data API or ORM, scalability isn't the deciding factor here.
That's good to know. So the agent just said, hey, everything. is scalable, then here we have three options.
Data API plus plain SQL, hybrid or drizzle. So a full on ORM. My honest take, C is the simplest and the long one.
One way of querying, one set of types, migrations are handled for you. A is the most work. That's not good.
And B is the best fit if you want to use the Data API.
Hmm, what should we choose? I will be honest with you, completely honest with you. I wanted to use the data API, but now thinking about it, it seems to me like choosing an ORM is just a...
You could say simpler step if that makes sense because the data API is interesting but it feels like it will be more work in the long run right now at least because if I would now use Drizzle then we would have two systems. We would have the data API and also an ORM which wouldn't really function as an ORM but rather as a schema slash migration library.
That's totally fine. It will work. It will be scalable.
It will work in the long run. There are big companies that that.
But do I want to do that myself? I feel like it will just be simpler if we will use an ORM which already has everything built in. Type safety, also a migration library, everything.
So let's use an ORM. You know what, let's do that. And instead of using Drizzle, I will use Prisma.
And I will use Prisma because I'm very familiar with the library. That's another thing I want you to think about. Code reviews are important.
You have to understand the technologies that you use. Let's take a random example. Next .js versus Rust.
Rust is more optimized. Rust is faster. Rust is big.
It has a lot of resources. In theory, we could have used Rust. Why didn't I do that?
Well, I'm not familiar with Rust. I don't understand the code. If there's a bug and, for example, the agent is not able to fix it, then I can't do anything.
I literally don't understand the syntax. Here we have the same thing. Prisma versus Drizzle.
I do know Drizzle. I understand it. I used it in the past.
It's a great tool, but I'm way more familiar with Prisma. If there is something that the agent can fix, if I want to review the code, then I need technologies which are familiar to me. And Prisma is familiar.
I understand it. And that's why I will use it. So I will say the following.
You know what? Let's go with C. We will use the full on ORM, but instead of using Drizzle, we will use Prisma.
So we will use Prisma. Everything will be on the server side.
We will use server components. That's all fine. And the reason for that is because I'm more comfortable.
I'm more familiar with Prisma. Then what else? Why Vercel functions?
Vercel adds the visitor's approximate country, city, and time zone. Okay, sounds great. Then let's use Vercel functions because why should we read all of the data manually from the header if we can already use pre -built code?
What client -side queries would mean? Well, we can already throw it out because we won't use any client -side queries.
We will do everything using Prisma and with that the Prisma ORM. What else did you ask me here? Question 13.
Which database approach? A, B or C? Well I already mentioned that.
Location? I answered that. ESLint and Putia setup?
Noted. Great. No source directory?
That's also good. So essentially we are finished. There's only one more thing I want to mention.
Since we will now use Prisma we will also have to create API endpoints. And I'm not the biggest fan of server actions. So another technology stack I want to add is that I want to use ORPCv2 and I want to use TanStackQuery for both querying data but also mutating data both on the server side and client side.
Since we will now use a somewhat traditional stack, meaning we will use a database plus an ORM, we will have to also create API endpoints manually ourselves. And there are a lot of libraries out on the market. We could also do it without any API endpoints.
We could use server actions stuff like that but i don't like that it's not very scalable and we are right here trying to create a scalable product so we will use something called orpc you can check out this video here how senior engineers design production apis we will use orpc plus turnstack query and this will allow us to create scalable open api compatible apis let me quickly open the website orpc and yeah check out the video check out the website it's a very powerful library and it will make our application scalable.
Because that's the goal here. Yes, I'm showing you how I wipe cold products, how I do it reliably. But what I'm also trying to show you here is how I choose technologies which work for me, which will still work in 10 years, and which will also allow me to onboard thousands of customers, make millions of dollars.
Yeah, and that's as simple as it is. And this, by the way, is the exact tech stack I also use for all of my products. Let's now go back to Cursor.
Does it have any questions for me? Prisma 7. stable or prisma 8 release candidate wow they already have prisma 8 honestly i would normally probably use the stable release but in this case i'm fine with the release candidate because the release candidate has finally support or yeah they finally support typed vector search so let's use it if it does not work in the long run if it breaks we will just go back to prisma 7 it's not a huge change so yeah let's use prisma 8 then what else do we have here is the or ORPCv2 beta okay?
Yes, that's because ORPCv2 is almost stable, so that's also fine to use. Question 18, how are ORPC and TAN stack query wired into Next .js? So procedure style, yes, router first, or we will do contract first development.
We will first of all create a contract, then the procedure. I want everything to also be OpenAPI compatible, so we will have an RPC handler, an OpenAPI handler, then... everything should go on the server side, correct server components with 10 -stack query.
In the documentation, they also have an example how to prefetch the query, how to set up the whole server side logic. So once we get to implementation, that's something the agent will have to look at. And finally, client side.
So ORPC 10 -stack query, utilities for use query, party kit events, update a query, cache directly. That's all good. Then what else do we have here?
Question 19, does the widget use ORPC too? Yes, let's go with your recommendation.
Question 20, what replaces row -level security? We don't need row -level security because everything will happen on the server side, right? Question 21, when does the agent run?
The visitor sent message call, shouldn't wait for the whole AI reply. Let's go with your recommendation. It sounds great to me.
And then question 22, where does the Prisma schema live? Both the Next .js application and the Neon functions need the database. Well, let's create a shared package.
That's the reason why we created a monorepo or that's the reason why we use PNPM workspaces. So that also all sounds great to me. We are now finally finished with the core foundation.
We let our agent grill us. We have a perfect PRD file and also a perfect Hackstack file, which means the next step is already to generate code. And the very first step is to set up the foundation.
Therefore, I will now again use this exact same session because we just used up, what, about 300 ,000 tokens. We still have enough tokens in the context window. So it does not really make sense to create a new session.
If you use a model which has a smaller context window like 300 ,000 tokens with GPT -6 Astra, then I would highly recommend to either compact your session or create a new session. Inside of here, I will now say the following.
Thank you, this all looks great. I want you now already get started with code generation. To be specific, the first step right now is to set up the foundation.
I want you to install Next .js, install Zord, maybe ORPC. Just install all of the needed dependencies. And maybe also create a basic homepage which just says hello world and then maybe a dashboard route which says hello in the dashboard or something like that.
So yeah, super basic. Let's just set up the basic foundation for our project. Please refer to our PRD and then also to the tech stack file you just created.
Our agent will now get started and the reason why I want to set up the foundation first and not just say like, hey, do the foundation and do authentication in one run is because I want you first of all verify that all of the basic things work. While the agent now works on our prompt, I want to also talk about my AI coding workflow because this year was now basic.
I used the same session, I used one model, but this is not really how I code normally or at least once we start implementing features, I have a dedicated AI. coding workflow. What you should know about me is that I'm a multi -model user.
I don't just use one model. There are people who just use GPT -6 Astra or just Opus 5 for literally everything. Research, implementation, reviewing code, stuff like that.
I don't do that. There are models which do certain things better than other models and certain models are cheaper than other models. As an example, Opus 5 .5 is a great all -rounder in terms of implementation.
It's great at the back end, it's great at the front end, it's very reliable, and it's somewhat cheap. It's not super cheap, but it's cheaper than Fable 5 .1. GPT -6 Astra, as an example, is also a great all -rounder, though I would say it's maybe a bit stronger in the back end than in the front end.
Some people might disagree. But at the same time, I would say, hey, Opus 5 .5 is not a good researcher. It's not super fast, and it's a bit too expensive.
So for a researcher, as an example, grog 4 .6 or maybe grog 4 .7 would be better alternatives, which means I'm a multi -model user and I use multiple models at once. Let me show you a diagram.
So here I have three roles, orchestrator, implementer and researcher. In theory, I also always have another role, reviewer, but that's something I will talk about a bit later on. So first of all, let's talk about the orchestrator.
The orchestrator is the lead agent. The lead agent should be a smart model. either Opus 5 .5, GPT -6, Astra or Fable 5 .1 or whatever comes in the future, of course.
And the orchestrator is there to think, plan. orchestrate, delegate, and judge. And what you should know about me is also that I'm a huge fan of sub -agents.
What's the biggest benefit of a sub -agent? Well, it does not impact the context window. If this one agent would now do everything, implementing, researching, and reviewing, what do you think?
Would we preserve tokens? No, we would literally consume our tokens in seconds. And that's not good.
That's why I use sub -agents. With that, I can preserve my context window. And at the same time, always use fresh models with a fresh context window.
Here, I will say like, for example, implement authentication. This orchestrating agent will get that prompt. It will now think, it will plan, it will reason, it will judge, and then it will start delegating.
For example, the orchestrator might say, hmm, you know what? I'm not quite sure how to set a beta off or how to implement email and password off. So it will use a subagent, a research subagent.
In most cases, for me, that's GROK 4 .6. Grog 4 .6 is cheap and it's fast. But you could also use Lunar Max or what else is there?
For example, Sonnet 5. Something like that. Something that's cheap and fast.
And this research agent will then find all of the necessary info. It will research the web, the global web, but it will also be able to look at my local codebase. And this research sub -agent is a read -only agent.
It's not able to change code, it's only able to read code and find search things, find answers for certain questions. Once the orchestrator is happy, once it has all of the answers for its questions, it then delegates the work further to implementors or in other words executors.
These sub -agents are there to write the needed code. They already get all of the needed data from the research agents and they start implementing. The executor has to again be a smart model.
In most cases I recommend to use the same model as the orchestrator. In the case it would be opus 5 .5.
And this executor is also able to call other subagents, again research subagents, because for example the orchestrator might find all of the needed answers using this subagent, this research subagent. But the implementer might need more answers, it might have more questions, therefore it has the ability to spin up more research subagents, which again are fast and cheap, that's important, because I don't have a limit on research sub -agents.
The agent could spin up one, two, three, four, maybe five. Well, there is a limit. I don't need 50 sub -agents, right?
But still, the model has to be cheap because we are not rich. Because we are not multi -millionaires, we are not Elon Musk, we can't do that. So that's something I want you to remember.
And that's already essentially it. This is my AI coding workflow. I use three roles, in most cases two to three models, and that's also why I use Cursor.
Codex does not have multi -model support. Claude Cote does not have multi -model support. You are a bit restrained if that makes sense.
So if you now for example use Claude Cote then I would recommend to use Opus 5 .5 with medium reasoning effort. In the future this might be Opus 6, Opus 5 .6, Opus 6 .7, who knows. For the researcher I would use Opus 5 .5 low because it's somewhat cheap.
If we go to Cursor or in other words to Cursor Bench then you will see here that Opus 5 has become relatively cheap. Opus 5 .5 is especially with low reasoning effort.
It's as cheap as GPT 5 .6 Sol, so I don't see any reason in using something like Gemini 3 .8 because it's way too expensive. So let's again go back to the diagram, and for the implementor, again, I would use Opus 5 .5 with medium reasoning effort. Let's quickly talk about the reasoning effort.
If we go back to CursorBench, a lot of people will say, hey, use extra high reasoning, use max reasoning effort. something I wouldn't recommend. As you see here, Opus 5 .5 with medium reasoning effort is better than Fable 5 .1 and it's cheaper than high.
So in my case, I feel like it's not needed anymore to use the most amount of reasoning effort that's available. At the same time, if you select a high reasoning effort, the model becomes kinda, I wouldn't say unreliable, but it starts overthinking things and you don't want that. You want the model to be agile.
It has to have the ability to do stuff. So if you use Opus 5, then please use either medium or high reasoning effort.
And if you use GPT -6, also choose something like high, maybe medium. That's good enough. Trust me.
Don't listen to the people who say, hey, use the highest reasoning effort. These people don't know what they are talking about. Then if you use Codex, I would say use GPT -6 Astra with either medium or high reasoning effort for the orchestrator and executor.
And for the researcher, use Lunar Max. You could use Sol, but Sol is a overpriced at least in my opinion but yeah that's how I use my models or that's my model arrangement.
Another thing I want to talk about is like the core workflow we now talked about the models and the roles nevertheless I have a specific way of orchestrating these agents. Everything starts with session number one.
What I mean by session number one is just a normal chat session as we have it here. This here is one session. And session number one is that you think.
So what I do in session one normally is play ping pong with the agent. I will tell the agent, hey, I want to implement authentication. I want to use better of.
What do you think? The agent will say, yeah, that's a great idea. Here's what I would do.
Then I will ask the agent, aha, and would you maybe recommend? using OR for maybe magic links and how would you implement it, the agent would again respond and I would just play ping pong with the agent. Why do I do that?
Well, this gives the agent context. It learns about my problem and it also learns about what I'm trying to do. So session one is like a context machine, a reasoning machine.
Once I'm happy with the whole, you could say, discussion, once I'm sure that the agent understands what I'm trying to do, I ask the agent for a game plan. In other words, I ask session one to create a prompt and then I use this prompt and give it back to session two.
Okay, maybe this sounds a bit confusing. Let me break it down a bit further. So we have session one.
I play ping pong with the agent. Once I understand that the agent is confident in the solution and knows what I want, I ask the agent to create a high quality prompt. I then use this high quality prompt and create a brand new session with a fresh context window, no tokens used because I need the full path.
of the agent and this prompt already instructs the agent on what I want, what I don't want, how the implementation should go and stuff like that because session one already has access to research sub -agents and then in session two I use executors. Executors again have access to research sub -agents and another thing you should know is that I stopped giving my agent bite -sized tasks.
In the past I would have said hey we want to implement authentication so Let's maybe create two or three sessions. Session number one will be there to implement the off foundation.
Session number two will be there to implement the back end. Session number three will be there to implement the front end. This is not something I do anymore.
What I do instead is say, I want to implement authentication from A to Z. That's what I tell session one. Session one creates the prompt.
It gives then this prompt to our second agent, the orchestrator, the one I just showed you. And the orchestrator then... the big feature into smaller sub -features A, B, C, D.
So as an example, feature A, or I guess sub -racket A, could be the foundation, sub -racket B could be the back end, C could be the front end, and D could be something else. And then this lead orchestrator delegates the work. It can parallelize work, or in other words, do the work sequentially.
It will delegate the work to sub -agents, executor sub -agents. At the same time, this executor also has access to research sub -agents as already mentioned.
Another thing you will see right here is that I have a reviewer sub -agent. This is something I haven't talked about yet. So once the lead is finished, meaning once all of the executor and researcher sub -agents are finished, the lead starts using reviewer sub -agents.
This reviewer sub -agent is there to review the generated code. If we for example now have one big feature which got broken down into four smaller ones, again A, B, C, Then this lead agent will use four reviewer sub -agents, and they will review all of the necessary code.
So for feature A, B, C, and D. And that's because I don't like reviewing code. It's necessary, don't get me wrong, but it's boring.
And that's why I want to let the agents do all of the heavy lifting. Once all of the agents are completely finished and happy with their results, I glance at the code myself, see if everything looks correct, and then I create the needed PR. So this is how my workflow looks like.
So let me summarize everything one more time because I know it can maybe sound a bit complicated. I work in a two -session system. Session one is that you think and also research.
Session two is that you implement. I always have a lead orchestrator sub -agent. It's there to take the prompt and then delegate because we might have a huge feature.
And as you all know, one singular model is not able to do everything or one singular agent. is not able to implement everything from A to Z. Therefore, the orchestrator will break the huge feature into smaller sub -features.
These smaller sub -features will be delegated to executor sub -agents, as you see here. Maybe two, maybe three, maybe four. Also, the lead agent will decide if we should do the work in parallel or rather sequentially.
And if needed, the implementer, or in other words, the lead agents, are able to also use research sub -agents. Once all of the work is finished, the lead will be able to again use reviewer sub -agents to review all of the needed code. And once everything is finished, the lead agent creates a handoff, I review the code, and I create a PR.
It's as simple as that. This is my AI coding workflow, it works the best for me, and it creates the best possible features. At the same time, this saves a lot of time.
Because in the past, I always did everything in smaller sub -features, But this always meant that I had to think about everything myself, what to do sequentially, what to do in parallel. It always took more time, and reviewing the code was always boring.
But this system allows me to save time, save all of the cognitive resources I have inside of here, and most importantly, it allows me to create safe and scalable code, because the agent is able to research the web, and it's able to review the needed code. That's why this here is a game changer. Now, I will be honest with you, is this the cheapest system?
Definitely not, because you will use sub -agents. It is what it is. But that's also why I said, please don't use the most expensive model, because you want to have a good balance between, you could say, price and performance.
And that's also why I use Opus 5 .5 so heavily, with medium reasoning effort. It's super, super, not cheap, but it's affordable. It's way cheaper than all of the alternatives I have used, and I'm able to literally use my workflow for days without getting rate -limited air.
all. So I use cloud code, I use cursor and it works perfectly. Zero complaints from my side.
And that's why I will also recommend this workflow right here with good models like Opus 5 .5. And that's also the reason why I asked you to install the feature orchestrator skill. The feature orchestrator skill is there to exactly explain my workflow to the agent itself.
So later on in session one we will use the feature orchestrator skill and again in session two. This skill explains exactly, first of all, the two -session play, the two -session system, and it also explains the roles, lead, implementer, researcher, and reviewer. This skill already has models hard -coded, which I find good, so they are like the default.
Nevertheless, the model, or I'm sorry, the agent, will also ask you what models you want to use for these roles. So that's very important, and that's why I wanted you to install this skill. You don't have to use it, you can always tell the agent specifically, hey this is my workflow this is what I want you to do but that's not needed again it takes too much time please use skills they are powerful you just have to use them the right way and you have to use the right skills so install it and it will save you a lot of time and with that the agents which you will use will instantly understand your AI coding workflow so this is now finished please don't forget to start the repository let's now go back to cursor is it finished yep it installed everything let's Quickly go through it.
So we have a beautiful screenshot, meaning everything works. What did it do? Mono repo, PNPM, Erp's web, packages shared.
That's all good. Erp's web, Next .js 13. That's nice to see.
No source directory. Dependencies, Zord 4, RHF, so we act hook form. ORPC with the beta, packages shared, tooling.
That's all good to see. Font, da -da -dum. Okay, zero complaints here.
Did it remove Geist? I fixed the layout. page, oh no, that's good, then not set up yet, packages db, that's absolutely fine, party server neon, that's all good, so let's check it out, localhost 3000, I will do a hot refresh, and everything works, can I also go to the dashboard, yes, this also works, so our foundation is completely finished, we can also quickly look at the code, I will open the IDE, what do we have inside of here, let me zoom in a bit, apps, okay, that's all fine, .next, app, what do we have here, there, dashboard, page .tsx, components, what do we have inside of here, a button, that's nice to see, public, we already have an agents .md file, what else do we have here, docs, that's what we created, and packages, this all looks great to me.
As a next step, I want to also let the agent update the agents .md file, because we now have a PRD, we have a tech stack file, but our agents .md file isn't optimized, it does not really explain everything we want to do, and Yeah, it's not good. Here's what I will do.
I will now create a prompt and say the following. Thank you. This all looks good.
Now, here's one thing you should know. We created a PRD file. We created a tech stack file.
We have the needed code base or the foundation of the code base. But our agents .md file currently is not optimized. And I want to change that.
You now already learned a lot about the project and what I want to do. So please optimize the agents .md file for me right now. Thank you.
And by the way, please always say thank you. We want to thank our model masters. No, I'm joking, man.
What does the rootagents .mt cover? What martial desk is? The three documents, outdated training data, blah, blah, blah.
Okay, repository layout and commands. This all sounds great to me. Since we are now finished with the foundation, we can already right away continue with the next step.
And now you might say, Jan, what should we do? Should we first of all implement authentication or maybe the... back and how about our real time server?
What should we do? Well, how about we start with the landing page? And now you might say, Jan, the landing page first?
This does not make any sense. Shouldn't we maybe first of all create the backend logic? Well, no.
And I will tell you why. In the past, that's exactly what I did. I always first of all started with authentication, with my backend, with the dashboard, all of this stuff.
And then only at the end, I always created the landing page. The reason why we will now do it the other way around. is because AI coding has changed a lot of things.
By creating the landing page first we will be able to right away create a design system which means all of the subsequent features and pages will use the design system and build on top of it. If we would now first of all start with authentication we would create the back -end code and we would also create the login and signup pages.
The problem with these two pages will be that the agent will just do whatever it wants to do. It will use its own design system, its own colors, maybe it will use our shared CNUI components, maybe it won't, who knows, and that's not good.
Because this means we will have to rewrite the code later on, this will just use more tokens, that's not wanted. Why should we do that? That's not optimized.
Instead we will create a landing page, we will create a design system based on the landing page, and then we will create all of the other features and pages and they will build on the existing design system. allow us to create features faster than ever. It will make everything exponentially faster.
So how should we now create a landing page? Should we just go inside of here and say, hey, create a beautiful landing page? No, that does not work.
And that's also why I recently created this video right here. How I built beautiful websites with AI coding agents. Cursor plus Clot plus Blender.
This video is almost 50 minutes long. And in this video, we built this beautiful website. Doesn't it look beautiful?
But here's the thing. We won't use this workflow in this video because it takes too much time. It takes about 50 minutes to create such a website.
Instead, I want to use something pre -built. But, you know, what most pre -built websites kinda look.
Copy -pasted, meaning they all look the same. I want something unique. I want something that is as unique as this website right here.
So what should we do? What is fast, easy, and unique? Well, recently I found this website right here, Motion Sites.
This is not sponsored, by the way. I'm not affiliated with the team behind this website, but this website essentially offers pre -built prompts. So build beautiful landing pages in minutes with our ready -to -use prompts.
Just copy -paste. and launch. Instead of now just copying components or something like that, we will use a prompt, we will give the prompt to our agent, and our agent will then create the needed landing page.
So here's what I want you to do. On the right, please select free. We want to use a free prompt, and then you can also order the result based on feature, popular, or recent.
Let's click on popular, and what's popular right now? Hmm, okay, this all looks interesting. Damn, I love this.
This looks super cool. I think we will use this. What else do they have right here?
Aha, this is a full -on landing page. I'm not sure if I'm a big fan of this design language. What else do we have here?
This is something I also already used in the past. This is something you can use definitely. This will give you a full landing page.
But in this video, I will use this hero right here, AI Runtime. So what I will do is click on Copy, Full Prompt. It's now copied.
what we can do is go back to Cursor. And in this case, instead of using Cursor, I will use Cloud Code because Cloud Code has a very nice built -in browser and it allows you to also annotate the browser, the result. So what I will do is create a new session.
Let me also choose the correct folder. Then for the mode, I will select bypass permissions. This is not something you should necessarily do.
This is not super safe because if you select bypass all permissions, you essentially... the agent, hey, you can do whatever you want to do, don't ask me any questions. In this case, this is safe, because we are just creating a landing page.
And this is also what I do in most cases, most of the time. But if you work with very sensitive data, if you have a real SaaS application, then this is something I wouldn't do. If you want to do something safer, then please use Auto Mode, because Cloud will handle permission checks itself, and whenever it's not really sure, whenever there's some sort of sensitive step involved, it will hand off to you.
So if you want to play it safe, use auto mode, but I will use bypass permissions mode. For the model, I will select Opus 5 .5 with medium reasoning effort. Again, that's fast enough or good enough.
And I will now paste the prompt inside of here. But instead of now just clicking on enter, I want to also update everything a little bit. Or in other words, I want to give the agent further instructions.
Because here you will see that the instruction is to use static, HTML, CSS, vanilla JS, no framework. This is not wanted.
Therefore, we need to update the prompt further. I will now say the following. Hey there, so I want you to help me create a high -quality hero, or in other words, landing page, for my application, Marshall Desk.
Down below, I already added a high -quality prompt instructing you exactly on what I want you to have, how everything should look like, the core composition, colors, fonts, stuff. like that. Nevertheless, you will have to deviate a little bit from the prompt because the prompt tells you right here to use HTML, CSS, vanilla JS, stuff like that.
This is not true. I don't want you to do that. Instead, I want you to use the existing text stack.
So next .js, tauren .css, sharedCNUI, and stuff like that. Another thing you should know is that the typo which the prompt instructs you to use is not really relevant to our website. We don't you have intelligence designed to evolve this does not really work for us so step one is to literally just use the prompt create a landing page use our text tag step two is also to evolve the landing page to use the correct typo or to in other words to use the correct which also works for our application and then the third step would be to create a design .md file which will then explain our design system, the colors we want to use, the components we want to use and stuff like that.
So there are three steps I want you to now follow. So as you see here we are now finished. This is my prompt and what did I want to do right here?
Well the prompt down below is first of all kind of without any context which is not good. Secondly it has instructions which Kinda don't really align with our vision.
We want to use Next .js. We don't want to use HTML. And also we want to use shared CNUI.
At the same time, I also told the agent that I don't just want to do one thing. Step one is to literally just use the prompt and create a landing page. Step two is to evolve the landing page and to use the correct typo.
In other words, use the correct text. which works with our application. And the third step is to create a design .md file, which will explain our design system.
Let me also update the prompt further. Right here, I will say to gather more context, please also look at my PRD file and at my tech stack file. It will explain everything further.
And one thing I want you to also remember is that I want you to use the existing, you could say, core primitive. So as mentioned, next .js shared. CNUI, Tailwind CSS.
As an example, if you want to create a button, don't create a custom button, rather use a shared CNUI component, and if needed, evolve it, update it, customize it. Use the foundation built on top of it. While our agent now works on everything, let's already think about step two, which is to implement authentication.
So let's go back to our diagram right here, because our diagram again is there to help you. If you forget what you do, which is absolutely fine, always refer back to this diagram.
This diagram is there to help you, to move you to the correct direction, if that makes sense. So what do we want to do? We have our next JS server, which is connected to what?
Authentication. How does authentication work? We want to use beta auth.
What do we want to do inside of there? We want to use OAuth, we want to use email and password auth, everything that is already mentioned in the PRD. Then let's think further.
Where we store our data? We will store it in our database, in Postgres.
Where will all of this data live? What service will we use for our database? Well, we will use Neon.
That's why the system architecture is so important. It's there to help us. But now you might ask me, Jan, but how exactly does this backend work?
I'm still kind of confused. So you said we will use a backend as a service, meaning we will use Neon. But how exactly Does it work?
Will we now use a third party off service or will we use better off? Because right here you said better off. How will we use the AI gateway?
Who provides the AI gateway? What about Postgres? What about functions?
What about object storage? How does that work? How does it connect to Neon?
What again is the backend as a service? Well, these are all great questions. So here's the thing.
I'm quite confident in that you guys are already familiar or at least somewhat familiar with Neon. Because if you have been a subscriber of my channel for quite some time now, then you will know that I'm a huge fan of Neon and that I have been now using it for about three years or so.
I started using it when it was still in beta. Yes, you heard right, and the landing page also looked completely different, which is quite funny. But yeah, I have been using the platform already for quite some time and even my own website syntaxpath .com runs on Neon.
But now you might say, okay Jan, this all sounds but isn't Neon a Postgres provider? Well if you remember Neon from the past then you will know that Neon is a Postgres provider.
It provides serverless Postgres, it has a great developer experience, it's super fast and super efficient but this has now changed a little bit in a good way because Neon is now a member of the Databricks platform and that's also what you will see here. The backend for apps and agents built to scale on lake -based Postgres.
So this means Neon Neon is not just a database anymore, which isn't a bad thing by the way, it's now a whole backend platform. The database is the foundation and literally everything else builds on top of it.
And that's also what you will see here, not just the database. Neon is a complete backend platform with authentication, object storage functions, and an AI gateway. But now you might say, Jan, okay, okay, this all sounds good, it sounds super fancy, but what is LakeBasePost?
Great question. So this is a term or a name that comes from Databricks. And Databricks is, for example, known for its lake house.
If you don't know what it is or if you're not familiar with Databricks, that's absolutely fine. I'm just trying to give you a bit of context. And that's then also the place where big companies keep all of their data for analytics.
And Lakebase is the Postgres database built to live right next to that. But now you might say, hold up. Wait a minute, something ain't right.
We aren't a huge company. We don't have terabytes in terms of analytics data. We also don't make millions of dollars per month.
How exactly does a lake -based architecture or lake -based database Help us? Well, it's relatively simple.
The lake -based architecture decouples storage and compute to deliver instant operations and scale without compromise on performance or reliability. So we still get a normal Postgres database. It isn't like a completely new system.
No, it's still standard Postgres. But instead of having everything in one run or in one process, your compute and storage, it's decoupled. And that's what you will also see right here.
We have compute and then we have storage. This is still one database, right? It's not like five databases kind of connected to each other in some weird way.
No, it's one database, but it's decoupled. And this gives you efficiency, as mentioned. It makes everything cheaper, but at the same time, it makes everything also faster, which is a huge game changer.
And what other benefits do we get right here? First of all, we can now scale to zero and we also get auto scaling right away. Auto scaling scales compute in real time following your application.
optimized cost performance without capacity planning. So if you don't have any traffic, zero users, then nothing runs. You don't pay anything, zero.
But if you now have a traffic spike, maybe your application has gone viral on YouTube, product hunt, something like that, then it scales up. That's what auto scaling means. You don't have to do anything.
All of the vertical scaling is done for you automatically without you having to say, oh my God, I have so much load. No, you can't. Another benefit that we get is branching right out of the box.
Since the data sits in storage, meaning they're coupled from compute, Neon can make a full copy of your database instantly without any delay. And that's also a big difference to other competitors. With most competitors, it often takes a few minutes to create a branch.
And in our case, we will also use two branches, a production branch and a development branch. In development, when we are locally coding something, we will use the dev branch. and once we deploy our application to Vercel we will use the production branch.
This will give us safety and that's also what I do in production in real life. And what else is important for us? Well, a free tier.
And that's exactly what we get with Neon. Powerful free tier. 103 projects with database, storage, functions, and auth.
Starts at $0. And that's exactly what we need. Because when you create a project, you don't know if it will be successful, right?
This SaaS which we are building out right now might be successful. It might also fail. That's life.
But with Neon, we are not locked into a $200 per month. plan. No, we can start off fresh, we can start from zero, we can build everything out, and once we are ready to scale, we can start scaling with Neon.
Isn't that cool? Another thing I want to talk about is this whole backend as a service primitive. What does even backend as a service mean?
Well, it means that instead of now using five different services, so one service for authentication, one service for Postgres, one service for compute, one service for storage, and so on, we can just use one provider, Neon. And Neon allows us to control, or I'm sorry, interact with all of these primitives straight from one place, from Neon.
And right now, when recording this video, September 2026, Neon offers five core primitives. First of all, Lake -based Postgres. As already mentioned, this is your standard serverless Postgres with also PGVector for AI search built in.
This already decouples storage from compute, which also allows us to create branches and also enables autoscale, away. Secondly off and this is super important because this is not just another third -party service like kind clerk work OS or anything like that this year is better off fully managed user off in every database for free build on better off you don't have to use some sort of third -party service you are not locked into a third -party service right here you are still using better off the only thing that neon does for you is manage everything it can send emails for you do all of the nest stuff including of course also securing everything and at the same time all of the data is instantly stored in your lake base postgres database so this means you own the data you own the database you own the schema and the users also live right in your app's data so you don't have to make a second query or a third query to some sort of third party provider everything lives in one region in one database making everything 10 times faster and you can also query this by
using plain SQL in ORM, the data API, whatever you want to do, everything works. And this also works perfectly fine with RLS, row -level security. So what you do is you build on BetaAuth and then Neon manages everything.
Neon operates the auth layer for you so you can ship authentication without provisioning, maintaining or scaling separate infrastructure. And since everything builds on top of BetaAuth, it's also super easy to use because coding agents are super familiar with better off.
Thirdly, Neon also offers functions. Functions are super powerful, because as mentioned, we will deploy our application to Vercel. Vercel uses serverless functions.
Serverless functions, with Vercel as an example, can only run for 30 seconds as a maximum. And this is not the case with Neon serverless functions, because Neon serverless functions are long -running. They can run for 15 minutes without any problem, start responding quickly, and then keep streaming as agents constantly.
models and tools, WebSockets stay open or SSE sends live updates. So the biggest benefit here is that streams stay open. Something which we don't get if we, for example, deploy our application to Vercel and use the built -in serverless functions.
Of course, Neo now also finally offers object storage. This object storage primitive, of course, is S3 compatible. And this means also all of the learnings are already embedded into training data.
Your AI agent, know everything about S3. So even though they might not be super familiar with Neon or the backend as a service right now, they will instantly know how to upload data because the S3 API is the standard.
And since we also get access to the S3 API, we can also use pre -signed URLs. We can generate the pre -signed URL on the server side and then upload the files on the client side. You also don't need any separate cloud account.
No AWS, I know, man. AWS is so annoying. Everything is built to scale.
And finally, the game -changing feature is the AI gateway. Neon created its own AI gateway, so you don't have to use any third -party provider. You can use one endpoint for all of the models.
So as you see here, OpenAPI, Alibaba, Google, Anthropic, whatever you need, whatever you want. And here you can also see all of the models, GPT -6S, Truck, Rock 4 .6, Kimi K3.
So this means you get unified access. You don't have to use multiple providers. Oh hell no, definitely not.
You also of course get simplified billing. But here's where the real game changer comes into play. Fair pricing.
Zero markup. Yes, you heard right. Neon charges the same per token rate as the model provider.
Published prices passed through with nothing added on top. Zero. This means you essentially get the AI gateway built in for...
And that's also what you can see here. Again, five core primitives. Functions, Lake -based Postgres, AI Gateway, Storage, and Auth.
And all of this is built into one platform, allowing you to scale from 0 to 100, make your first 1 million dollars, all while having zero headaches. So yeah, now you know what Neon is. It's a huge game changer.
You can also check it out by using the first link down below in my YouTube description. Thank you, Neon, for making this video possible. I mean, without them, I would have never created such a huge video.
And again, Neon is something I use for my own websites. Syntaxpath, the website I've been building already for quite some time, is built entirely on top of Neon. But okay, since you now know what Neon is, what a backend as a service is, what all of these core primitives are, we can already continue.
So let's see what Claude called generated right here, or what our agent generated. What do we have here? Step one, the prompt.
built in our stack, that's good, button I added three new styles, that's nice, trust row built from the shared CN avatar, mobile menu built on shared CN dialogue, that's also all good, copy that fits martial desk, an agent that knows its limits. Nice, that's right. Stats, the made -up metrics become claims that are true today, that's also good.
Step 3, design doc, so it created our design system file. What do we have here? Principles, one composition, no clutter, built on the foundation, soft, not heavy.
The three dots, the three small dots are the brand motif, interesting, honest copy, motion reveals, it doesn't perform. That all sounds great to me, so let's check it out. I will use the built -in browser, let me you a hard refresh.
And oh my god, this looks beautiful. An AI agent that knows its limits. A chat widget that answers visitors from your own knowledge base.
Man, I'm ready to make those millions. How about you? No, I'm joking.
But this looks beautiful. I have zero complaints. Let's also see what the agent said at the end.
Other files I touched. That's fine. Open items before launch.
That's important. Display font. Bubble .icg.
It comes from online web font CDN. This is not something I want. Hmm, what should we do here?
Well, if I remember correctly, Geist now also has a Pixel font. So if I go to Google Geist Pixel, let's see. Yes, introducing Geist Pixel.
Instead of using this weird font, which is great, don't get me wrong, but I want to rather use Geist because I'm familiar with the font. Background video, it's on someone else's CloudFront URL, that's fine. also fine.
So here's what I can do. I can click on this mouse button. I will click on it and now I can annotate everything.
I will click on the pen button. There's now bugs a bit, that's fine. And what I will do is annotate first of all our hero text and that's already it.
I will click on add to chat and now I will say the following. Look, so I'm not a big fan of the hero text or let me maybe rephrase it a bit. Currently you use a font called bubble .icg.
I don't like that instead please use Geist Pixel. Also for the normal text I'm not sure what font you're using right now but I would like to use the normal Geist font.
Essentially I want to use Geist for everything for the Pixel font and for the normal font. Another thing I don't like is the headline an agent that knows its limits. This does not really sell our product because what's the core idea of Marshall Desk is that it saves time of the business owner.
The business owner does not have to sit in a dashboard the whole day and answer questions. The AI agent can do that for the business owner.
So yeah, let's rephrase it a little bit. Currently, our agent is still working on the task, but what I like is that the agent is actively using our browser. And that's because the agent invoked our feature orchestrator skill.
So even though I didn't use my two session system, it still looked at the prompt, it understood, we are now in session two, right? Let me use the built -in browser. Because that's something that the prompt also mentions.
The prompt says I want you to continuously use the browser and check the results and in this case our agent is now checking hey is everything mobile responsive? Does the mobile navbar work? Does the video background video load?
That's what our agent is doing right now and it's doing that because it's using the skill we have installed right at the start. Normally agents models don't really use browsers that actively they just spin it up at the end to check if everything loads. But this is not the case here.
The agent is using the browser continuously. So as you see here, we now have a new font, Geist Pixel, and this also uses the normal Geist font. And let's also check if it's mobile responsive.
I will click on viewport, mobile. Let me quickly do a hard refresh. We have a beautiful animation.
And if I click on the mobile nav button, it also opens beautifully. And you see the animation. Isn't it beautiful?
And this here is now also like our image, our logo. It's fine. I like it.
And with that, we are now finished, which means we can finally continue with authentication, build on top of our existing design foundation and scale exponentially. Since we want to now implement authentication, we can also again quickly look at our system architecture diagram. So what do we want to do here?
We want to implement managed better off. Who provides this primitive? Correct.
So what are the steps? Well, the first step will be to log into Neon. The second step will be to create a new project.
And the third step will be to configure the Neon MCP server. And this MCP server will give our agents the ability to directly interact with the Neon platform. With the core primitives it will allow the agents to configure everything and we in theory won't even have to open the dashboard and do anything manually.
That's the power of MCP servers. specifically the Neon MCP server. We will also install Neon specific skills, we will actually install a Neon plugin, but I will show you everything in a second.
Let's again first of all head over back to Neon. You will find the link to Neon down below in my YouTube description, underneath the like and subscribe buttons. And here I will now log in.
If you don't have an account yet, then please sign up. But yeah, I will click on log in and I will instantly get redirected back to my dashboard. Inside of here you see that I already created quite a few projects, 44 in total, and I'm even subscribed to the scale plan.
So when I said that I use neon for production, I actually meant it. I use neon for everything. Let's now again head over back to our projects panel and we want to create a new project.
I will click on new project and let's give our project a name. I will call it Marshall Desk. For the region I will select Frankfurt because Frankfurt is closest to me but please select whatever is closest to you or in other words to your users.
Then we can now also enable certain services. First of all Postgres, our database, it's already enabled by default. We can also update the database name if needed.
If we don't add anything then we will use the default, which is NeonDB. And for the Postgres version, I would recommend to use the latest one.
When recording this video, the latest Postgres version is version 18. Then let's also enable object storage. I will click on enable.
For the bucket name, I will just say uploads. And for the visibility, I will make it private. So here's the thing.
You have the option to either select or create a private bucket or a public bucket. We will create a private bucket. because we will store sensitive and private data.
What you can also do right here is create multiple buckets at once. What I, for example, do in most of my applications is that I create two buckets, one private and one public bucket. In the private bucket, I store all of the private and sensitive data, and in my public bucket, I store all of the public data, like, for example, profile images and stuff like that, which are all publicly available or accessible.
So I will again delete the second bucket. don't need it. Let's also enable functions.
We want to enable the AI gateway and we want to finally enable Neon authentication. So this is all finished. I will click on create project.
Your new project was created in under one second. If you blinked you probably missed it. Correct.
Here we instantly get instructions on how to set up Neon with our coding agent and if you want to you can copy this prompt, paste it into cursor or cloud code and this will also fully work. Essentially this prompt tells the agent to install the Neon CLI, to log in, to install the skills, the MCP server, to link the project.
This works, but that's not something I want to do because again I want to use first of all my AI coding workflow and also I want to install the Neon plugin. The Neon plugin will give us two things. First of all a set of skills and also one MCP server.
We won't have to install everything manually or separately. Separately, we can just install one thing, and this one plugin will already give us everything needed. Let me show it to you.
That's maybe, first of all, Open Cloud Code. Inside of here, I want you to click on Customize. In the Customize section, you can click on Plugins, and inside of here, you can search for Neon.
And here, you will now find a Neon plugin. And as mentioned, this plugin will give you skills and also the MCP server. The same thing also exists in Cursor and also in Kodak.
Let me show it to you in cursor. Inside of here you can open the sidebar, click on customize, I will search for Neon and here we now have the Neon plugin in the marketplace. Let me open the plugin and here you will see that we will install one MCP server and then also eight skills.
As an example, claimable Postgres, Neon, Neon AI Gateway, Neon Functions, Neon Object Storage, etc. This means we won't have to point our agent to the Neon. direction.
The skills already have all of the necessary instructions inside of them and this will make everything more agentic faster and it will also save a lot of tokens. What I will do inside of here is click on add to cursor. I will add this for myself and cursor will now install this plugin for my account so globally.
I will do the same thing in cloud code. I will click on add. This here will now give me a warning.
The Neon plugin may include components that run code. That's fine.
I will click on continue. This will now take a second. As you see here, it will install everything.
The plugin is now installed. As you see here, we have an overview, connects, skills, and then also connectors. And the connector is currently not connected, if that makes sense, the MCP server.
So let's again go back. I want you to click on connectors and inside of here, search for Neon. Right here, we will now get a result, Neon.
I will click on connect. to Claude and this will now redirect me to this website. I will click on continue connecting and essentially what we are doing right here is connecting our local agent Claude called our MCP client to the external MCP server to Neon.
Here we can first of all edit all of the permissions so choose access. We want to say all projects you can access. For the tool categories I will select everything and then for the permissions I will select read and right.
I will click on approve and finally I will also have to now authorize the MCP server. I will click on authorize and this will now redirect me back to Cloud Code. And tada!
Connected to Neon. If I now go back you will see here that everything was successful. Let's click on plugins.
Inside of here I will search for Neon and in theory, only in theory, everything should work. I will click inside of here. We have our skills.
All of the skills have been installed successfully and our connector is now connected. We now have all of the skills and also the MCP server.
If I also head over back to cursor, I need to authenticate the MCP server. I will do the same thing. I will click on approve and then inside of here I will also click on authorize MCP server.
And if I head over back to cursor, everything should now be successful. And yes, it took a second, but now I have access to 113 tools. Everything is enabled and that's exactly what we wanted.
We can now again close this and what's the next step for us? Well, we can head over back to the Neon website. Inside of here I will click on go to project and this is our dashboard.
Postgres database enabled, let me zoom in a bit, beta off enabled, object storage enabled, functions enabled and AI Gateway enabled. Currently we don't have any functions or anything like that. That's absolutely fine.
Let's quickly go through the dashboard so that you get a general understanding of how everything works. First of all, you will see here that we are currently in the production branch. That's something we will change in a second.
On the left side, you will find a sidebar. I will click on branches. Here we currently only have one branch, the production branch.
But you know what? Let's create a new branch right away. I will click on new branch.
I will say the... development. This will be our development branch.
For auto delete I will select never. We don't want to delete the branch. Parent branch will be production and we want to also include the branch data and schema.
So here's the thing. As already mentioned you can create multiple branches. A production branch, a development branch, a feature A branch, a feature B branch etc.
And therefore your configuration might also change per branch. In this case we want to get all of the data, all of the schema.
But for example, let's say you want to create a feature branch. You might need anonymous data, or maybe you want to only get the schema. Who knows?
So I will select branch data and schema and click on create. And as you see here, we now have our new branch. We are also instantly in the development branch, which is exactly what we wanted.
We can also select monitoring. As you see here, currently nothing has happened yet. No usage.
at all but that will change a bit later on once we actually create users and also of course have some load. What we can also do is view the available integrations and I would highly recommend to install the GitHub and also Vercel integrations. So for example the GitHub integration allows you to create a brand new branch whenever you create a new PR and the Vercel integration allows you to create a new branch for every preview deployment.
is what I also have installed for all of my projects because this is a huge game changer and it also makes development secure, which is exactly what you want. There are also a few other integrations as you see here, but they are not needed in our case. You also have the ability to view your settings inside of here and update the project name if needed, but we will leave everything as is.
Now we can also view all of the individual primitives, database, better off, object storage. functions and AI gateway. Let's start with our database.
We currently have zero tables, but that's absolutely fine. But now you might say, wait, Shen, what do you mean we don't have any tables yet? Didn't we enable authentication, meaning better off?
Doesn't that mean that we want to also store all of our data in our database? Why don't we have any schema yet? Well, that's because we are currently in the public schema.
I want you to click inside of here and select the neon of schema. And Here we have finally all of the necessary tables.
Invitation, member, organization, etc. What you also have access to is an SQL editor if you want to write some manual SQL. We can also view our backups and restore something if needed.
But yeah, click through everything yourself to learn more. Let's also view our better of primitive. What do we have inside of here?
Anyone on the web can sign up for your application. Support for restricted signups is coming soon. That's fine.
We don't have any restricted signups. What we can also do is view our plugins. What do we have here?
Enable organizations. Yes, that's correct. We also have already certain limits set, but this is not very relevant for us because we won't invite members, at least in V1.
Here you can also instantly set the creator role and we will leave it as is owner because the creator will be an owner. What else do we have right here? Magic link.
Should we enable magic link? Sure, why not? Then for the link expiration, five minutes is good enough, and we will also allow new user registrations.
You could enable phone authentication, but I will leave it as is. What else can we do right here? We can configure off.
This will redirect us to the settings, and here you can update the application name. You can copy the off URL, JWKS URL, and stuff like that, but this is all fine. Please make sure that localhost or allow localhost is enabled because we want to use Neon in development, or in other words, authentication.
Here, we also already selected sign up with email. Let's also say verify at sign up. This means require email verification when users sign up, and we will send a verification code.
Finally, what else do we have right here? Sign in with email. We also instantly already have one OAuth provider configured, which is Google.
And the biggest benefit here is that you already instantly get shared keys. Normally, if you would use normal beta auth, you would have to first of all go to Google, to the API console, create credentials, copy -paste them.
This always already takes about 20 minutes. The same thing is also true for GitHub, Microsoft, etc. But with Neon, as mentioned, you get managed auth, which means you also instantly, at least for development, get shared keys.
Once you go into production, I would still highly recommend to use your own keys, your own custom keys. But this is fine enough if you want you you can also add another OAuth provider like for example github But then you will have to add your own secrets. Google already ships with its own shared keys Then what else do we have here email provider once you go into production?
I would also highly recommend to use your own email provider We can also enable webhooks if needed, but this is not really needed for us. So this is all finished Let's now check out object storage. What do we have right here?
Well, we have one bucket, the uploads bucket, which is private. Let's also check out the functions primitive, deploy functions next to your code base. And here we also instantly get instructions, which is to install the Neon CLI.
So let's do that. I will copy the installation command. I will open my terminal, paste the command inside of here.
This will now install the Neon CLI globally, and I instantly get an error, and that's because the CLI is already installed. As a next step, I want you to run neon lock -in.
And I want you to authorize the Neon CLI. And now it says off complete. That's great.
I can again clear everything. And let's go back to the Neon dashboard. If you want to, you can also link your project right here.
But we will leave it as is. What else do we have here? Well, we have one of the most powerful primitives, the AI gateway.
And here you will instantly see what models are available right now. GPT 5 .3, GPT 5 .4, GPT 5 .5, GPT 5 .6. and of course GPT -6 Astra.
You also get access to the open source models, meaning GPT -OSS and also models from Chinese labs like Quen, GLM and also Kimi K3. In this video we will probably use a GPT model, but we will see a bit later on. But that's already it.
We now went through the dashboard, we connected Neon to our coding environment and the next step is to finally implement authentication. Also, since we already onboarded our agent, we can close this little, how should we call it, immodal, and let's head over back to Cursor.
I will close this, I will create a new chat session, I will also update the model, I want to use Opus 5 .5, Medium Reasoning Effort, with a 1 million token context window. And as mentioned, I want to use my AI coding workflow. My AI coding workflow consists of two sessions.
The first session will be that you play ping pong with the agent, and that's what we will do here. But before Before we start playing ping pong with the agent, let's first of all verify that we are connected to Neon.
Therefore I will say the following. Hey there, please check out if you have access to the Neon plugin, meaning if all of the Neon skills are installed, and also if you have access to the Neon MCP server. All eight Neon plugin skills are installed and the Neon MCP server is connected and signed in.
That's exactly what we wanted. What's the next step for us? ping pong.
I will use the feature orchestrator skill. As you already know, it's a prerequisite to install it, so please install it. And before we now get started with the prompt, let's first of all think about what we want to offer.
Let's think about the core primitives. We want to first of all enable Google OAuth because that's what we enabled in the dashboard. We want to have a login and sign up page.
What else do we need? We want to enable magic links. That's very important.
Of course, we want to also enable email and password auth and now we can finally already create the prompt. Hey there, so I'm currently looking to implement authentication and we are right now in session one and I already have like a core few primitives mapped out.
I want to create a login page and of course also a signup page and then in terms of authentication methods, I want to offer Google OAuth, magic links and also the OG email and password auth. In terms of implementation itself, I want to use my backend as a service, Neon, and Neon provides managed better off. So we will use better off, but we will store all of our data in our backend as a service, in other words, Neon.
What I want you to right now do is look at my PRD file, look at my tech stack file, use the needed Neon skills, the Neon MCP server, do all of the necessary research, and I want you to just figure out essentially how you would do everything. What do you think of my idea?
Is there anything that is maybe unclear? Is there anything you want to ask me? The core idea here is to just figure out if the agent is able to understand what I'm trying to do.
Our agent is now finished. It also has a few questions for me, but let me for now collapse everything. So let's go through this.
The agent used to research sub -agents. One was there to figure out how to configure NeonAuth and also the CLI and the second one was there to figure out how to use the SDK.
And the agent also has a few valid findings. I went through it. So it wants to use a proxy file, a catch -all route, a server instance, a browser instance.
So this is the standard better -of setup. Now, let's quickly look at the open questions. The first run step needs a database.
Of course, it needs one. Should this feature also set up the foundations it needs, meaning packages DB, Prisma 8, what else, the OR? RPC handlers.
Sure, that makes sense. So I will say yes, include the minimal foundations. So off works.
Then what else? Magic links for new people. A magic link to an unknown email should create an account without a name or password.
Skipping the signup form from A1B. How should it behave? Interesting.
Sign -in only. Magic links work for existing accounts. Yeah, I think that makes sense.
Let's say for existing accounts. Is it okay that off calls? meaning sign in, sign up codes, Google, go through Neon's API off route handler instead of ORPC.
Yes, that's absolutely fine. And that's also what I do in production. The development branch needs to setting changes, turn on the magic link plugin and turn on send verification email on sign up.
I already did that. So I will say the following. I already did that.
So everything should work. It's already set in the development branch. I will now click on next.
else do we have here production setup your own google o off app diversell domain as a trusted domain and turning local host local host off now or at deployed time we will say later at deployed time this is right now not needed and finally what do we have here account settings a5 change name upload a profile photo die spare fallback photo upload needs object storage included this is interesting so i will say the following interesting so So the thing is, I want to give this ability to the business owner later on in the dashboard.
So in the dashboard, the business owner should be able to update their profile photo, then also their name. This means, yes, we need object storage, though I'm not sure if we have to already connect it right now, because the step right now is just to, in basic terms, implement authentication. So I guess we could already include object storage.
Let's think about it maybe a bit further. Now I will click on continue. I answered all of the questions and this is what I meant by ping pong.
I gave this agent one prompt, right? I said, hey, this is what I want to implement. What do you think?
The agent said, oh, that sounds great. Here's what I would do. Here's what I would use.
Oh, this does not look correct. It asked me questions. Now I answered the questions.
Again, ping pong, input, output, input, output. And this is exactly what you to do.
Because what are we doing right now? We are building a general understanding. The agent is understanding what I'm trying to do, what I'm trying to do in the future.
Because as mentioned, for example, I want to allow business owners to update their info, not right now when authenticating, but in the dashboard. How should we handle the owner's profile, name and photo, now members .name plus an empty avatar key column and a user menu with the dice per avatar and signout, later a separate account settings feature that wires objects storage and ads editing.
You know what? Let's do the recommended answer. I think that's a bit simpler.
Question eight. Google accounts come with a profile photo. Should it count as the owner's photo?
Yes, we want to use the Google picture as the photo until the owner uploads their own. What's the recommendation? No, use Dice Bear.
Definitely not. We will use the Google image. Our agent is now finished.
It does not have any questions for me anymore. And in general, the plan or the proposed plan looks good. me.
Nevertheless, I want to do one more thing. Because as mentioned, agents have become smart, models have become good, they have become very intelligent, very agentic. This means we don't have to scope this run to just one feature, implementing authentication.
I want our agent in a second to implement authentication, the back end, the front end, meaning the login and signup pages, and then I want the agent to also right away implement our dashboard page. So here's what I will say. This all looks good.
But you know what? I want to add one more feature which I want to implement in the same run, which is to already create the dashboard UI. We don't need to create the backend functionality or anything like that.
The first step right now or the goal is to create the dashboard layout with our sidebar on the left and then on the right we will have the main content. For the UI, I want to use a shared CNUI component. Shared CNUI offers blocks and I want to use a prebuilt block, meaning the dashboard01 block.
I will add the command in a second down below and I want to also do the same thing for the login and signup pages. I want to use a prebuilt component provided by sharedCNUI. I will also paste the commands down below.
Look, if I wouldn't have said what I just said a second ago, the agent would have built everything from scratch. The login page, the signup page, the dashboard layout. This is fine, it works, don't get me wrong.
But this is not the preferred way of working at least in my opinion. Because these blocks are already prebuilt, they look beautiful and most importantly They work, and that's what we want.
Agents can build full -on UI layers from scratch. They can do that without any issues. But it is very repetitive, because the agent will make mistakes.
You will have to say, hey, this font is too small, this is too big. This is not really wanted. Instead, I will now give the agent the installation commands, and the agent will be able to clone the exact block with already all of the necessary presets, because this here is beautiful, let's be honest.
And if wanted or needed. later on. We will be able to build on top of this rock solid foundation, this UI foundation and optimize things, update layouts.
Let's first of all again go back to the blocks. What do we want to exactly install? This here is the dashboard component.
This is what I want to use. I don't want to necessarily already create the inner page. For now we just need the layout with also this collapsible sidebar.
I think we will use this example because it kind of looks the best. Then let's also look at the login components. Yeah, this can work.
Actually, you know what? Either we use something like this or we use this first example. I would probably use the second example because on the right we could render the image which we have in our hero.
So yeah, I think that will actually work the best. Here's what I will do. I will first of all copy this command right here.
npx shared cn add login02. I will go back. Let me paste it at the bottom.
I will say login and sign up component. I will paste it right here and let's also go back to our dashboard component right here, feedshot.
I will copy the installation command and also paste it inside of here. So dashboard component or dashboard block, let me paste the URL and here I will say block and not component. Let me update the prompt further.
Another thing I want you to know is that for the dashboard itself, I don't need the inner pages. The goal is just to create the layout, the sidebar that everything works. We need the correct color scheme, we want to use our design system, stuff like that.
Also, in the sidebar, we want to have probably two links. First of all, like the general, I'm not even sure how we should call it, maybe the general dashboard link, settings link, whatever, and then also an inbox link. This is where we then later on will allow the business owner to have a conversation with visitors.
And we don't need any other links, that's all good. Maybe let's also have a user profile button at the bottom, and on the top left, we can render our logo. meaning Marshall Desk and also the beautiful icon you already created or render in the homepage.
The agent now has a few questions for me again. What should the first sidebar link be called? Yeah, let's call it home.
Why not? Secondly, dashboard. Oh, one installs charts, a data table, drag and drop, sample data, and about 19 components.
What do we do with the parts you don't need? Install with the CLI as you asked, then delete. Yeah, this sounds great.
What go? in the right hand half of the off pages? Good question.
I forgot to mention that. I told you what I want to do, but I didn't tell the agent. The landing page is looping background video.
As you see here, our agent is smart. It already understood what we want. And that's because we built a foundation and a general understanding with the agent.
We have a design .md file. The agent knows what we like. And at the same time, since we are building the session out, or since we are playing ping pong with the session, will probably want.
This is what the user will probably not want. And yes, we want a looping background video. This sounds good.
Where does magic link sign in go? On the sign in page. A second button under or continue with.
Yep, that sounds great. And then what else should the off pages use? Sign in and sign up.
No, this here is a mistake. And this is now where the benefit comes in of being a developer, right? I'm an engineer.
I have written a lot of code in the past by using a normal keyboard. Which means I also know how to optimize code the right way. And Next .js has a primitive or a feature called layout routes.
Because what the agent proposes here right now is to create two separate routes. Sign in and sign up. But let's again look at the two blocks.
I will go to login. Let me scroll to the bottom a bit. What do we have right here?
Like what will change in between the login page and the sign up page? The only thing that will change is this block right here. login page we will have the same logo, we will have the same image, it does not change.
But if I would now create a login and signup page then we would have two renders, we would have two full re -renders because this would re -render, this would re -render, this would re -render, that's not needed. We don't need to re -render this, we don't need to re -render this, the only thing that needs to be re -rendered is this component right here.
So instead of using two completely new routes we will instead use either a route group or for example something like this with a full on parent route. So here's what I will say.
Look, I'm leaning towards a parent route like off, meaning slash off, slash sign in, slash off, slash sign up. But please don't forget to also optimize all of our pages for performance. What I mean by that is I want to use a shared layout, a layout .tsx.
That's the file we want to create. Because what will re -render? The only thing that will re -render is the login block itself.
The logo on the top left shouldn't re -render and the image or the video asset that we will have on the right side does not need to be re -rendered on a client -side navigation. So yeah, let's create a layout. Another thing I want you to know is that later on, once we also create a dashboard, it might make sense to use a route group.
That's something you have to decide for yourself. But let's use the existing Next .js primitives. Let's use the platform.
Let's optimize our application. make it as good as possible. And finally, what should the menu behind the user button contain?
Avatar name and email on the top, then sign out, account gets added with the account settings feature, Sure, why not? And I think that's already it.
Is there anything else we should add right here? Is there anything I forgot? No, I don't think so, right?
We discussed our authentication setup, the game plan, right? Our login page, how we want to design all of our pages. We also discussed the dashboard, how it should look like.
I think we are finished and the next step will be to generate or create a game plan prompt which we can then hand off to session 2. Let's see right here I've read slash off as off that's why I said willow the voice dictation software has turned to absolute. You know what I'm trying to say?
Yeah, I meant off. So the pages will be da -da -da and da -da -da. This fits the neon off SDK too.
It's middleware already treats this blah -blah -blah and blah -blah -blah as public pages. That's all good. What else do we have here?
Layouts off shell. Correct. That's what we want.
Is there anything else? All of the off screens will live under it. Two paths need checking against the real SDK.
That's also fine. Dashboard shell. App dashboard layout.
That's good to see. Route. groups.
Both of these are real URL segments, so they will work as plain folders. A route group becomes useful to separate the always dark landing page from the light app screens. Sounds good, so that's maybe something we should also do.
Performance rules for session 2. The video panel only shows on large screens. Okay, recap, and that's all great.
So, is there anything else? Everything from the earlier round still stands. If this matches what you have in mind, say compile.
Sure, I will say compile. Here's where our skill again kicks in. So the agent said before I write the prompt, pick the models for session 2's three sub -agent roles.
So the implementor will be Claude Opus 5 .5 with medium reasoning effort. The researcher will be Grok 4 .6. Why won't I use Grok 4 .7?
Well, Grok 4 .6 is cheaper and it's also faster. Grok 4 .7 is a great model, don't get me wrong, but for a researcher specifically, this model makes a bit more sense.
Gemini 3 .8 Flash also is a good alternative, or maybe Composer 2 .5. But again, GROK 4 .6 is my preferred model. I love how it works, it's super fast, it's easy to use, or not easy to use, but it's easy to understand whenever it works on something.
It's enjoyable to use, so I will select GROK 4 .6. And then finally, for the reviewer sub -agent, I will select... What should we do?
I mean, we could use Opus 5 .5 or GPT 5 .6, maybe even Fable 5 .1, but I think I will stick to Opus 5 .5. I've been loving the model recently. It's super cool.
So yeah, I will click on continue and the agent will create a prompt for us. While the agent creates a prompt for us, there's one thing I forgot and that's to create a public. repository because that's the final step.
So what I've done is I've already reviewed the code and the code is clean. I don't have any issues with the code, but that's also because it's super basic. All the agent did is literally compile all of the technologies, right?
It installed everything. It set up our PNPM workspace, our mono repo, and that's it. There's not a lot to review.
And also you might ask me, Jan, how do you review your code? What's the magic sauce? What tricks do you have for me?
Do you want to know the honest truth? The honest truth is literally open the thing, open the changes and look at them. It's as simple as that.
There's not more to it, there's not less to it. All you have to literally do is look at the code and review it. Is it boring?
Yes. Does it take some time? Yes.
Is it important? In my opinion, yes. I would say the following.
You don't have to look at the code as thoroughly anymore as in the past. In the past, LLMs, models created pure slop. It's the honest truth.
Agents have now become better, smarter, and you don't have to be as thorough anymore, especially with our workflow, because we have coding agents which review their own work or review the code of other sub -agents. Therefore, the code review does not have to be as thorough anymore. Nevertheless, I would highly recommend to at least glance at the code.
code because you will instantly see often mistakes. Like, I don't know, code repetition, meaning not following dry principles, using incorrect primitives, not using, for example, if we are talking about Next .js, a route group, or not using the link component, the image component, etc. Reviewing code is important, do it, but also you don't have to be as thorough as in the past anymore.
Do we now have a prompt? Yes, our prompt is getting written, so let me also create another prompt in parallel. Look, thank you for the prompt.
another thing I want you to do right now is create a public repository, call it Marshall Desk or whatever is available, and at the same time also create a PR for this code right here for How should we call it? I don't know, committed as project foundation or something like that.
I already did a code review. It all looks good to me. So you can just create a PR and commit all of the code and of course push it.
Since the prompt is now finished, I can also click on enter. Our agent will create a new PR for us and of course a repository. Let me also copy the prompt and open a new session or create a new session.
I will select the model, which will be Opus 5 .5. I will paste the code inside of here. quickly check if we now have a PR right here.
So the public repo is right here. This is our PR. It contains all of the code.
What I will now also do is quickly check out the prompt and see if everything looks good to me. Subagent model roster. This all looks good.
Then what else do we have here? Repository. Starting good state.
That's all correct. What else do we have here? Authorization boundaries.
Neon. Read only SQL. On development is allowed.
Orchestration. Roll route. Implementer this explains our workflow again, and that's also or how do we even get these instructions you might ask me Well, that's again our skill feature orchestrator the feature orchestrator skill has a lot of instructions Like for example, this is the workflow This is the roles that we have and also for example verify your changes using the built -in browser also Paralyze work if possible if not then do it sequentially and stuff like that Is there anything that I don't like right here the goal and white man?
the relevant settled decisions, progress file built with your eyes open. This again says, hey, use the built -in browser. Mandatory preparations.
This means that the agent should look at all of the files like our PRD, agents .md, the root agents .md, existing product and architecture. So this here is a super, super high quality prompt. And that's super important.
What I often see is that people create the most basic prompts ever. Hey, create a dashboard. Hey, implement.
and authentication. Hey, make me a millionaire. This does not work.
You have to be precise. And that's exactly what this prompt is doing. It's being precise.
It tells the agent exactly what it should do and what it shouldn't do. And that's what I want you to always do. If you don't use my workflow for reason XYZ, it does not matter.
Then please make sure that you are specific when creating prompts. And here's another huge thing I haven't talked about yet, but it's now time to also talk about it. In the past, I was a huge fan.
of plan mode. I literally used everywhere plan mode. I always generated a plan.
This is something I don't do anymore at all. I don't really find it necessary, if that makes sense. It does not help me.
The only time I would use plan mode is if I work in a huge company, if I'm working on a huge task and I want to make sure that the agent actually understood what it should build and also if I want to first of all make sure that I'm able to commit to the task, if you understand what I mean. Because with a plan, we already know exactly what the agent will do.
Therefore, we can also confidently say, hey, you are going into the right direction or into the wrong direction. So I see a lot of big companies still using plan mode. Also, a lot of cool senior engineers, which I know, they use plan mode because they need to have the, how should I say, they have to know that they are able to commit to a certain thing.
I, for example, know one guy who's a senior engineer at Netflix who earns way too much money. He still uses plan mode. And that's because he needs to be able to commit to something to whatever the agent wants to generate.
Because if the agent generates slop, then it might break something and he might get fired. No, I'm joking a little bit, but I think you understand what I mean. Plan mode makes sense in certain scenarios.
But for 99 % of people, or for at least most of you guys watching this video, plan mode might not really be necessary anymore. And I personally don't use plan mode at all anymore. But that's already it.
We are now finished. So what I will do is click on enter. Our agent will get started.
If it has any questions, then it will ask us. But if not, it will just do everything necessary and we will then be able to review the code at the end. Our agent is finished and I already took the time and looked at the code or in other words reviewed it and therefore let's look at the conversation I had with the agent once it was finished.
First of all I said OAuth does not work because what happened is when I tried to log in with OAuth meaning with Google it redirected me back to the login page but at the same time the whole authentication process was kind of successful. The issue was the callback. The callback did not register for the OAuth provider.
And this is something I also kinda expected, because this already happened the first time when I set up authentication privately. So yeah, I gave the agent this simple prompt, hey, OAuth does not work, I get redirected to this website right here. The agent then said, Google sent you back to the index page with the session verifier in the URL, instead of the callback URL we passed, let me fix that in the proxy file.
The agent fixed everything, and then everything also worked. I will show you the end result in a second.
As the next step, I reviewed the actual code. And first of all, I asked what is turnstack query for, because I saw that the agent set up turnstack query. And when we did the foundation, the agent didn't do that.
He then told me, hey, I set up TQ for ORPC. I also set it up in the layout file. I did all of the hydration steps for SSR.
I also asked that, did you set up ORPC? It said yes in the previous PR in the previous foundation or branch, that's all fine.
Another thing I didn't like was actually this code right here in the auth error file. I mean, look at it. Doesn't it look weird or in other words, unoptimized?
And therefore I said the following. Some of the error handling seems overcomplicated to me. Like for example, in the auth error .ts file, what do you think?
The reason why I said what do you think is because I want the agent to think about its You could say code that is generated. If you just say, hey, the code is trash, fix it, the agent will do that.
But it will also do it if the code is correct. And therefore, I always say, what do you think? Give me your thoughts.
And the agent also agreed right here. I agree, it's more than it needs to be. So the agent overcomplicated the code, even though it's not needed.
Why is that? Well, models are smart. Opus 5 .5 is super intelligent.
It has a huge amount of raw intelligence. And that's why the model likes to overcomplicate things. We also saw the same thing with GPT -5 .6, even GPT -6 Astra, Opus 5.
The models have more intelligence power than an optimization layer. I hope that makes sense. So often models will try to generate code that covers every direction, every attack surface.
And this then creates not clean code and overcomplicated code. The code to the model maybe seems clean when generated. everything but once you review it yourself you will see hey this is not really needed simplify it and that's what the agent should do right now so I will say go ahead let's also quickly verify if everything works so I will start the dev server first of all I will open my terminal pnpm run dev we can again close this this will open on localhost 3000 I will open localhost 3000 Man, this homepage is so beautiful.
I love it. I will click on sign in and we should instantly get redirected. Here I also checked the code.
The agent used our layout. So this here won't be re -rendered and the same thing is also true for the video. So if I click right here on sign up, there are no re -renders.
The only thing that now re -rendered is this block right here. Let's again try it out. I will go back to sign in and let's see if also all of the errors work.
So I will say jan at jan .com with some weird password, one character, and I will click on sign in. We instantly get an error that email and password don't match.
That's correct, but one issue I have is that we currently allow a user to type in only one character for the password, which should not happen. We will update our agent in a second. What else can we do right here?
Sign up. Let's actually sign up. I will say jan marshall.
I will add my actual email and then a password. Let's again use one character. And if I click on sign up, ah, here we get the error.
Interesting. So when signing up, our Zord schema validates the password for the correct amount of characters. But on the login page, it does not do that.
Should we optimize the code? I feel like we should do so because it does not make any sense to allow a user here to enter only one character. So that's something we will definitely change.
So let me again say Jan Marshall. Then I will add my email. Let me add my actual password.
and I will click on sign up. Aha, we got instantly redirected and we need to now add our verification code. One thing I instantly see is that this here is a custom input.
That's not something I want. Let's go to SharedCNUI because SharedCNUI has a component called InputOTP and I would rather use this component right here instead of using a custom input. Let me now also quickly open Gmail.
And yes, we got an email. Verify your email address, Marshall Desk. So this is what I mentioned at the start.
Right now we use the Neon Shared Email Sender. This works perfectly fine for development. I would highly recommend to use this for development.
But once you ship your application into production, please use your own custom email service. Or your not custom email service, but your own email service. Because yes, NeonAuth or the shared email service here is great.
But it's always a better idea to use your own custom domain. And also it will be... better for deliverability.
So here we have our code. I will copy the code. Let's go back.
I will paste it inside of here. This is again valid for five minutes as you already know and I will click on verify email. So here we have an issue right away on the bottom right.
What does it say? Warning security warning the SSL modes prefer require and verify CA are treated as alias for verify full. I think this is actually a Prisma error because it says PG connection.
Let me just copy this console error so that we don't lose it. I will copy it. Okay, let's continue for now.
The business name will be Marshall desk. Why not? And I will click on continue.
We should now get redirected to the dashboard and that's exactly what we see here. Home, this is your workspace, conversations, your knowledge base and widget settings will show up here as they are added. On the bottom left, we also have our user profile drop -down and we see the user profile here including the avatar.
One thing I don't like is that the text is grayish it does not look very good so that's one thing we will have to change but besides that I'm quite happy with the result. So let's again go back and I will now create the following prompt. Look in general you did a good job though I have a few issues.
Let's first of all start with the user drop -down in the dashboard which is rendered on bottom left currently we when I click on it we render the name and also the email right the thing is it's kind of grayish I don't really like that the contrast isn't good at least in light mode so that's one thing I want to change another thing I want to change is that I want to add dark mode support so shared CNUI has a documentation page on dark mode support I will also paste the URL down below what else don't I like let's go back to the login page in the login page itself for the password input we don't have any length check which is not good because in the sign up page we make sure that the user adds at least eight characters and another thing I don't like is once we go to the verify OTP page once we get redirected you added some sort of custom input.
This is not wanted I would rather want you to use the input OTP component provided by shared CNUI. I will also paste the documentation for that down below. So let's go back to shared CNUI and UI.
First of all, I will copy the URL for the next JS dark mode documentation page, and in the past I always recommended to copy the markdown, but this is not really needed for highly optimized pages. So shared CNUI is a highly optimized documentation page. You can just use the URLs because the agents will be able to instantly get the markdown needed.
If you use some sort of like super, super old website with a lot of JavaScript, I would still recommend to do it the old -fashioned way, to literally just copy everything in the page. But with highly optimized pages, it does not make any sense.
Let's also copy the input OTP component page. I will copy the URL, go back and paste it inside of here. Is that it?
Is there anything else I missed? I don't think so. I will click on enter and our agent will get started.
If we now again go back to our dashboard, oh, here's one thing we haven't done. We haven't logged in with Google yet. So let me...
do that and yes this works and the interesting thing is since I logged in with the same email as you see here I also essentially got logged into the correct workspace right away so this works beautifully I'm not the biggest fan of the dashboard layout right now I feel like I don't know, it looks a bit too simple maybe, but let's for now let our agent finish everything and then we will check it out once it's done.
While the agent works on everything, I want to talk about one core issue I saw while the agent verified its changes. As you already know, we used the feature orchestrator skill. The feature orchestrator skill instructs the agent to verify its own changes in the built -in browser.
And the agent tried to do it as well as possible, but it kind of failed. And that's because currently when someone signs up, we send an email to the user's inbox, right? But the agent does not have access to some sort of email inbox.
So what we have to do is now give the agent some sort of email password combination for an account which is already verified. The agent should then store these. keys inside of the PRD so that it can then later on log in and also verify the changes in the dashboard.
What I will do is again go back, sign out, and I will create a new account which will be specifically for our agent. So I will say agent admin, something like that. Then I will use my email.
Here you will now see that the input OTP has finally changed. We now use the shared CNUI component. Let me head over to Gmail.
This is our code. I will copy the code, paste it inside of here, verify my email and for the business this will be called admin agent, something like that. Let me click on continue and we are now in the dashboard.
If I go back to cursor then you will also see here that everything is finished and now I will again add the email and password inside of the input. Let me also add a prompt at the top. Look, thank you for your changes.
One thing I have now done is I pasted down below input or credential credentials for you to log in since you were earlier not able to look at the dashboard. You can use this email and the password.
The account is already verified so everything will work. The agent is now actively using our browser. So navigated to localhost 3000 of sign in, signed in and landed straight on the dashboard.
So the account already has a workspace. Before looking at the screen, I will check the database row. The row looks right.
You already went through sign up. This all looks great. Let's head over to Neon, to the Neon console, to the Neon dashboard.
And if I now zoom out a bit and do a hard refresh, then you will see here that we have two new users. User 1, User 2, or in other words, two members. And let me also go to the workspaces.
We have two workspaces, Marshall Desk, and also Admin Agent. If I now change the schema to Neon Off, then you will also see the users right here. Yes, I already signed up four times.
That's because I wanted to test everything. What else can we do? Well, we can click on the Better Off primitive, and then you will also see all of the users inside of here.
Agent Admin, Jan Marshall, and Probe 1. This is the email that the agent used at the start, but it was not able to continue any further because it had to verify its email. Is there anything else we can check?
No, the object storage is still empty. Our database now has a few rows. So this all looks beautiful to me.
And if we can go back to our overview and if I zoom out a bit, then you will also see here that we finally have some load. And that's exactly what we wanted. And that's also what I meant when I said auto scaling.
Right here we had zero load, right? No users at all. So what did our database do?
Well, it slept. Bye -bye. No, I'm joking.
But this is the power of Neon. If nobody is using your application, then you also don't incur any costs. Everything in that sense is free.
And only once you again have a new user, as you see here, once you have load, your database again spins up in milliseconds, by the way, and it then serves the request. If we now go back to monitoring, you will also see the same thing here. We had some RAM usage, some CPU usage, then that log zero, that's good, rows, this all looks beautiful.
Let's again head over back to our application. Where's our application? It's right here.
Is our agent finished or is it still verifying its changes? Nope, it's still verifying everything but that's fine. What's the next step for us?
Well the next step is to already create a UI for our home page and let's quickly think about it. What do we want to do? Well I want to create a 70 -30 split, maybe a 60 -40 split.
So essentially this here will now be 70%. And here we will be able to view all of the settings.
We will have an input, for example, for the widget name. We will have a few inputs for the colors, blue, red, whatever. Then we will also have what else are the drop zone for the knowledge, that's important.
Then maybe also the position, that's one thing I forgot. So bottom left or maybe also bottom right and that's already it. And then on the right side, right here, we will have the widget.
So we will have a UI example so the user the business owner will be instantly able to see how the widget will look in real life and it will also auto update so if the user the business owner updates the name of the widget we will instantly also see it right here this is the layout the UI I want you create and what I want you then also do right away is create an inbox UI right here and you know what I already thought about how to like design everything how to create a complete position and while going through google or while searching for examples i found this right here the og email shared c and ui component and i feel like it can work very well for us because what we would do is first of all have all of the chat conversations right here then we would have on the right side the actual conversation between the visitor ai agent and then with us the business owner and finally on the complete right instead of rendering this on the right side we could render like user details the user location the user what else that we want to render the time zone stuff like that things i already have in the prd so there are two things we want to do we want to have the dashboard and then this inbox and before we get started with the ui implementation i want you first of all head over back to cursor and create a pr and this is another step of my code verification
process. Let me show it to you. First of all, right here, instead of committing and pushing, I will create or I will select commit and create PR.
I will click on the button and our agent will now create a PR for us. Now you might say, Jan, code verification? What does a PR have to do with code verification?
Well, let me show it to you in a second once the agent is done. The agent created a new PR and in total we have two PRs. First of our project foundation and then also this new owner of and dashboard shell.
What I will do right here is make this or mark this as ready for review and the PR is now getting reviewed. Both built on top of main which is quite interesting. I wanted to stack the PRs but that's not a huge problem.
Let's now come to the second code verification step. I do a three -step code review. Step one is that I let the lead agent create sub -agents which review all of the generated code.
You already know that. Step two is to do a manual code review by literally just looking at the code. And step three is to use PR code review tools, like for example, Cursor BugBot.
I'm a huge fan and with that a huge user of Cursor's BugBot. It's a review tool, it's an AI agent, which also reviews your code. And the reason why I use these PR code review tools is because they often find issues which the agent, the main agent, did not find and also issues that I myself did not see.
So it's a huge game changer. I love using them. There are a lot of players out on the market.
Cursor, BugBot, Graptile, then also CodeRabbit. Another alternative I would recommend is PullFrog. PullFrog is completely open source and with that almost free because what you can do is wing your own key.
I'm right now in the console and what did I do? Well, I connected my cloud code. and also my Chachipiti or Kodak subscriptions.
So instead of now paying Pulfrock X amount per month or per reviewed PR, I just use my existing subscription. Right now, this is not connected to this PR, but that's absolutely fine. So this is a great alternative.
I would highly recommend to try it out if you don't want to use BugBot, Graptile, or CodeRabbit, but yeah. Definitely use some sort of code review tool in the PR itself, which runs in GitHub Actions. One thing we can do right away is check out PR number one, our Project Foundations PR, and we also right away have issues.
Cursors bugbot found two issues. Landing animations use undefined reasoning. That's not good.
And also looping video ignores reduced motion. These are not huge issues. These are small issues.
Nevertheless, we can fix them. What I will do is again head over back to Cursor and inside of here I will say the following.
Please fix the issues which have come up in PR1. So the issues that Cursor's bugbot found. Now one thing you might ask me is, Jan, shouldn't we just like set up an automation and tell our agent to change or fix all of the issues found by these PR review tools?
Well, in the past I would have said no. Please check yourself and check if it makes sense but things have now kind of changed.
These PR review tools have become super reliable. 99 .9 % of the time they find valid issues and that's why I would also recommend to actually set an automation and let the agent fix all of the comments. What I would say is tell the agent look at the comments see if they are valid and if so fix them.
It does not really require or there's no real requirement anymore to go through things yourself. Agents have become way better. It isn't like five months ago and they are way more autonomous.
So let's check out our second PR. What do we have here? We have another issue.
So what I will also say right here is we have another comment also in PR 2. So please also fix that or first of all check if it makes sense and if so fix it. That's also what you will see here with PR 1.
I just told the agent please fix the issues but instead of instantly fixing them the agent wanted to first of all confirm everything. And that's what I mean by you don't have to go through everything yourself.
Leave the agent or let the agent do its thing. Things have become autonomous. The agent is able to decide for itself.
Later on, once we also switch to cloud code, which we will do in a second, I will show you a neat little trick where you don't even have to prompt your agent to work in a loop. But let's wait for everything to now finish. Both of the PRs are now finished and everything is also green.
So here's what I will do. I will go into the project foundation PR and I will also merge it. Merge pull request and then let's also do the same thing for the other PR once everything is finished.
Pull requests, owner of and dashboard shell and inside of here I will then also merge the request once we get the ability to do so. I have now merged both PRs. I also opened Cloud Code.
Inside of here I said, please switch into main and also pull all of the recent changes. The agent told me that our PNPM log .yaml file changed, meaning you want to probably install dependencies to merge it. I told the agent, PNPM, install yourself.
And with that, everything is now in sync. So what's the next step for us? Well, it's to create the two dashboard pages.
The home dashboard. page with all of the settings and also the preview of the mockup and then also the inbox page itself. And how should we do that?
Well, of course, we could now just create a prompt and say, hey, homie, I want you to create the two pages, make it look beautiful, make it look tasteful, go ahead. This will work, but that's not something I would recommend because we want to create something beautiful, something unique and something that we can ship into production.
And most importantly, it should not look like AI slop. That's why I I want to first of all gather reference material, something that we can show our agent and just say, look at this reference material, this is the design or this is the approach I'm looking to take, this is what the end result should look like or at least somewhat look like.
So let's go to our biggest competitor, Intercom, and what I will do is just look at the existing screenshots. So this looks quite nice, I will copy the image, let's go back to Cloud Code, I will paste it inside of here, let me also switch the tab around.
What else do we have here? Accept all or reject all. It does not matter.
This image looks quite interesting. I will copy it. Let's go back.
I will paste it inside of here. Then they also have the thin AI agent. Sure, I will copy the image.
Why not? Paste it inside of here. What else do they have here?
Ah, this is the inbox. And this is what I meant. On the left side, we have all of the conversations.
In the middle, we have the individual chat. And then on the right, we have the user details. I will copy the image and also paste it inside of here.
As mentioned, I'm gathering reference material. Is this interesting? I mean, yeah, let's copy the image.
Why not? It's just another example of what we could create. And I guess that's already it.
Nevertheless, we have one big issue right now. We gathered a lot of reference material for our inbox, but we did not really find anything for the home dashboard with like this 70 -30 split everything we already talked about. So, So what should we now do?
Well I also went through Google, I tried to find good examples, I did not find anything. Therefore I now went to the last resort which is to use an image generation model. So I created the following prompt which is super super basic.
I want you to use GPT image 2 .5 and create a 16x9 mockup. Essentially I'm creating an intercom style application. I want a 70 -30 split on the left side, I want you to have all of the settings and on the right side I want you to have like the mockup of the widget.
This here was a super basic prompt. And this here is what the model, meaning GPT image 2 .5 generated. And this already looks quite interesting.
Let me quickly open it right here so that you can see it better. We have our settings on the left side. And on the right side, we have the widget preview.
Now, is this perfect? Definitely not. But this already gives us a good starting point.
And that's exactly what you want. Instead of letting your agent start from scratch, scratch, without any design direction, without any inspiration, I would rather always recommend to gather at least some sort of reference material you can point your agent to.
Even though this is not perfect, even though there are a lot of things I don't like, this already is a great starting point. And by the way, all of these images and prompts are again also linked down below in the YouTube description underneath the like and also subscribe buttons. Let's now again head over back to our best friend Claude.
code let me by the way also copy the image copy image and i will paste it inside of here and with that i think we can already create a prompt the most important thing here is that we won't really use the two -step workflow i normally use with the two sessions and that's because we will work on the front end nothing else i will now say the following look we now already created authentication we created the foundation as the next step i want to work on our dashboard i want to create the home dashboard and also the dashboard inbox page.
So what I now pasted into the prompt right here is a bit of reference material, actual images I want you to look at. So you will find mockups for, for example, the widget for our inbox and also a mockup for our actual home dashboard page. Now, here's the thing.
This is only reference material. I don't want you to copy it or anything like that. I want you to use it as inspiration.
Nothing. else. And yeah, I want you to work on the two pages.
First of all, what do you think of everything right here? What would you suggest? Please also look at our existing PRD file, tech stack file, maybe also quickly at the existing code base.
Our design .md file is also super important. What are your thoughts right now? While our agent works on this, there's another thing I would like to do.
I want to work on the sign -in page and with that also the sign -up page. Because yes, there are side is good, I like the image, but I feel like this is still very basic and this does not really use our existing design language and with that design system.
I mean look at the hero. The hero has this pixel font, we have like this animation and if I now go back to the login page it looks Kinda too basic, if you understand what I mean.
So here's what I will do. Our agent is already finished here, that's fine. Instead, I will now create a new session.
Inside of here, I will create a work tree, that's very important, because a work tree allows us to work in an isolated copy of the repository. In a second, the other agent will work on the dashboard, and this agent will work on the login page and also on the signup page. And since I don't want to have any conflict, I will right here use a work tree.
And what I will do also inside of here is switch the model to Sonnet 4 .5. And that's because we don't really want you to write code from scratch. We will start from a working foundation.
Everything is already done. We want you to just iterate on the existing design language, on the existing result. And that's why Sonnet 4 .5 is a great candidate for that.
It's a fast model, it's very capable, and it's cheap, though the speed is the most important thing in my opinion. So first of all, I will... create a basic prompt and I want the agent to start the dev server and then also pretty much showcase the site in the built -in browser.
So I will say the following. Can you please spin up the dev server and then also open it or open the website in the built -in browser? Our agent will now get started on that.
Let's again head over back to our other session. So what do we have inside of here? Where the code base is today?
The dashboard shell exists. The database only has the following. That's fine.
The reference images. Screens 1 to 4. These are all of the intercom images.
Worth talking about. Three columns, conversation list, conversation, and details panel. This is the standard layout for a reason.
List rows with avatar name, one line preview, that's all fine. Not worth taking because it's out of scope. Teams, teammates, mentions, yes.
Uppercase, section, labels. Correct, we don't want that. Why is that?
Well, AI agents love to generate like uppercase labels with mono, with a big leading. It looks absolutely ugly. And our design MD file...
Dance that right away. That's super good. Then blue and yellow message bubbles.
Our app theme is monochrome on purpose. That's good. The thin chat window and the tone guidance.
Step ID code overlay. Our widget and marketing visuals. Not inbox UI.
That's fine. Screen 6. Nova chat isn't a homepage.
It's our widget settings page. That's also what I meant. It closely matches D8 and D9.
Settings on the left. Left preview. Okay, okay.
So my suggestion for the homepage. Principle 5 in design .md is honest copy. I think the agent misunderstood what I meant.
I don't mean that I want to update the home page, but rather our widget settings page. Then what else does it recommend right here? My suggestions for the inbox, URL -driven routes, that's good.
You don't want to really use client -side state if you don't have to. URL state is way better because if you do a hard refresh or if you share the URL with someone, the state is not lost. List filter tabs for all open, waiting, agent, you, closed sessions, that's all good.
The decision I need from you, real data, my recommendation, UI only with mock data. Let me say the following. I think you kind of misunderstood me.
I don't want you work on the homepage. I want you work on the widget settings page. And homepage is already finished.
And of course, I want you also work on the inbox page in the dashboard. So I want you work on two dashboard routes. And in regards to the widget settings page, I feel like it might make sense to work or build on top of an existing foundation, maybe on an existing shared CNUI block.
So shared CNUI has like this image, or sorry this mail block which the team created in the past. What do you think?
Does it make sense to build on this foundation or would you rather start from scratch? And I will also add the image inside of here. One thing I want to also quickly talk about is that agents in the past were not really able to analyze images or in other words visuals.
This has now finally changed and in my opinion you should try to provide visual assets as much as possible. If you have reference material then share with your agent because text prompts are great, don't get me wrong, but they are kind of, how should I say, they are super basic.
If you say, hey, create something premium, then it can mean thousands of things. I mean, what is a premium design? It's not very specific.
An image, on the other hand, is very specific and that allows an agent to pinpoint exactly what you mean when you have a specific prompt or when you share it in text. form so let me now click on enter and let's see what the agent thinks of my idea the mail example yes but for the inbox not the widget settings of course my man come on we aren't stupid then two things to know first it isn't an installable block searching the shared cn registry okay that's fine we will fix that the closest installable block that's also fine how the mail layout maps to our inbox that's all good finally a few other changes selection lives in the url the page fills the full height on phones one pane shows at a time so one essentially like block if that makes sense widget settings built from our own components left about 65 % and then on the right about 35 % two things I still need from you real data or mock data for now we will use mock data and then also where should the widget go in this sidebar I'd add it as a third entry under home and inbox no that's incorrect so let me say the following
first of all for now we will use mock data and secondly our widget should go into our home entry so under home we will just have links two links home and inbox not a third link i don't need a third entry then for the mail component or block itself yes correct it does not really exist anymore as an installable link so what you will have to do is traverse the github repository and find the mail block because it exists somewhere can you first of all verify that and come back to me.
So the issue here is that this block does not really exist as an installable link anymore. Instead this component, this block lives in the shared CNUI GitHub repository. It's not really public so that's why I let the agent try to find it.
I don't want to search for it myself. Let's see the agent is working on everything. Let's go back to our other session.
The agent started our dev server. Here's our website in the built -in browser. By the way you can open the browser by clicking on this browser icon on the top right and then the browser will open on the right side.
Let's click on sign in. This here is our login page once it loads and I want to now tell the agent that I want to rework it. So here's what I will say.
Look we are now completely finished with the home page. It looks absolutely beautiful also with like this contrast and the geist font. Zero complaints.
Nevertheless our off pages meaning sign in, sign up, stuff like that they kind they look basic especially if you compare them to the existing homepage. As an example on the right side where we have the video it might also make sense to use our pixel font and maybe render something.
On the left side we should probably or maybe render our You could say inputs and everything in the cart. Maybe this will make everything look a bit better.
In general, I want you to look at our design .md file and then also look at the existing homepage. And based on these learnings, I want you to optimize all of our auth pages. Our agent will now get started on that.
This will run in parallel. Let's go back to our other session. What do we have inside of here?
So I checked the mail example isn't in the current shared CN repo anymore. I found the latest version before it was deleted, so we can copy it from there. That's great.
So this has been deleted. That's an even bigger issue, man. SharedCN, why did you do that?
It was a great component. What porting it involves? Style, it imports the old Radix components.
Fine, the agent can change that. Dependencies, it pulls in date, FNS, old default import paths, and also this state management library, which we don't need. We can skip it because the selected conversation will live in the URL instead.
What we drop, nav and also account switcher. The rest of your message, mock data for now and sidebar stays at two links, home and inbox. Then if that's right, I will write up the plan for both routes next.
Do we need a plan? I don't think so. So here's what I will do.
I will say the following. I don't think that we really need a plan. So here's what I would do instead.
If you need research or if you want to research something, then please spin up a few Sonnet 5 .5 sub -agents. up to five. And then you can already get started with implementation.
I want you to maybe also parallelize work since we have two different pages. You can use two Opus 5 .5 sub -agents to implement everything. And also please use the built -in browser continuously because we want to verify that everything works.
Is there anything that I've missed? I don't think so. If you have any questions besides that, then sure, ask me.
But I think you are ready to get started. And again, why didn't I use the feature orchestrator skill inside of here?
Well, that's because we are just doing front -end work. If we would now do multiple things at the same time, front -end work, back -end work, maybe also some sort of foundational work, then sure, I would use my feature orchestrator skill, but the feature orchestrator skill has one big problem. Yes, you're right, it has one problem, which is it creates a prompt which instructs the agent exactly what to do.
And I don't want to really, like, take away the freedom of the agent right now. When working on the UI, you want to give your agent as much freedom as possible in certain cases.
And this is one case. I want the agent to think for itself and just decide for itself. I don't want to lock the agent into one specific design language, one specific composition or anything like that.
Let's again go back to our other agent. Does it have any questions for us? Oh my...
God, what the hell is that? I'm not sure if I love this or if I hate this. Let me think about it.
Let me think about it, man. What did the agent say, first of all? The off pages now match the landing page, dark and pill -shaped.
Left side, every form now sits in a rounded card with a soft border.
I don't know. Honestly, I don't know. I honestly don't really love this right now.
I don't like the rounding, especially since our dashboard does not have that rounding. Let me quickly show it to you. Let's go to the dashboard.
I'm locked out. Let me continue with Google. And yeah, I don't know.
I feel like it does not work. There's too much rounding. So I will say the following or I will go back.
I will click on this select element button and I will select this card right here and I will say the following. Look, this doesn't look bad, to be honest with you, but it's definitely not perfect.
On the right side, first of all, I love the video, and I also love the pixel font that we now use, so zero complaints with that, but I'm not a big fan of the card you created. So, first of all, I don't like this full -on rounding. I feel like it does not really play well with our dashboard, so I would rather use the normal rounding, if that makes sense.
The same thing is also true for the card. I wouldn't really... to use like this full -on rounding besides that everything is good it's just too much rounding for me so what i'm currently doing here is just again playing with the agent the agent generated something i don't like it i again created the prompt the agent will now change everything and then i will again look at the result so now our agent is done what do we have here welcome back sign in email and also password The thing is, our password does not really have any placeholder, so that's another thing I would change.
Let me use the select element. thingy feature, I will select our password input, and I will say the following inside of here. For the password input, I would love to render also a placeholder, and on the right side, I would like to also have an icon which allows the user to toggle password view, if that makes sense.
So to either view the password or to hide it. Another thing I don't like is this or continue with. It has the incorrect background color, so I will say the following.
Another issue I have is is with this or continue with it has the incorrect background color it's currently like full on black but it should have the same background color as our card itself let's also check out our sign up page what do we have inside of here again the same thing i need our placeholder i will say the following in the sign up page i want you also have placeholders for your name and password let me also quickly try to sign up i will say jan marshall i will add some sort of fake email and then Also some sort of fake password.
If I now click on sign up we get an error. Enter a valid email address. Interesting.
So gmail .com sign up. What do we now get? Again an error.
That's because I used an unsupported character. And if I now click on sign up we should get again our OTP input. This is still fully rounded which I don't like.
So let me also select this right here. I can select the full thing in theory. And I will say the following.
Everything also looks good. besides the OTP input itself. The OTP input is fully rounded.
So the actual input number thingies, whatever you want to, or number inputs. So it shouldn't be fully rounded. It should have the same rounding as also our button and card.
And one thing I also don't like is that our button has this weird grow animation. It does not look good. The glow is nice, but this growing animation does not really work.
It does not look professional. Please again review our Emil Design Engineering skill. What I will do down below is again just add the email design engineering skill because this animation does not look professional to me.
I will click on enter. Let's go back to our other session. What do we have inside of here?
Let me scroll to the top. Both implementers are running in parallel, widget settings, the dashboard home route, and also inbox. Before starting them, I did the shared setup so the two can't collide on dependencies.
Shared CN components added, stream down installed. Why do we need streamdown? It's the markdown renderer, text stack names for messages.
Ah, okay, that makes sense. And then also mail example downloaded. Our agent is now working.
As you see here, we have two subagents running, which both use Opus 5 .5. And this is another big benefit of using subagents. Subagents allow you to parallelize work.
We could, of course, also do everything sequentially. And that's also what all of us probably have done in the past. But why should we do that?
Why should we waste time? I'm doing things sequentially. Let's do things in parallel if possible.
And in this case, since we have two different routes, we can also run or work on them in parallel. I mean, I guess we have three agents running in parallel because we have another agent working on the homepage. So let's wait for this to finish.
Let's go back to our other agent. Since it's using Sonnet 5 .5, it should already be finished. And yes, we now have placeholders.
This is very nice. And if I click on sign up, we also have placeholders here. We don't have this annoying animation anymore, which is also quite good.
The signup placeholder repeats the hint under it, so I will change it. All five changes are in, that's good. Password field, it now has a placeholder.
Signup placeholders or continue with has the correct background. Code slots, they were still fully rounded because my last edit didn't take place. That's not good.
The buttons no longer lift or grow on hover. That's also good to see. And this is also correct.
Here's what I will do. I will again click on sign up. I will add a fake name, a fake email.
Let me also add a fake password and I will click on sign up. And yes, this already looks 10 times better. I don't have any complaints anymore and we can also ship this.
Let me also quickly review the code because that's also very important as already mentioned. And one thing I don't necessarily love is that we have this function here, export constant off button class. that's not really needed.
So I will say the following. While reviewing the code, I found this export constant of button class. I don't really see the reason for having this constant because what you could have rather done is update our button component and added a new variant or maybe updated the globals .css file, something like that.
Think like a senior engineer. We want to have a rock solid foundation. We already have components.
components installed. Instead of building new ones or updating them from outside of the component, let's rather update the components from inside by adding or in other words creating variants. Now you might ask me, Jan, what do you mean by this outside and inside customization?
What you should know is that there are a few component libraries out on the market, right? Like MUI, what else is there? Man, I forgot all of the names.
Next UI. And then there's also shared CNUI. What is the fundamental difference between these libraries like MUI and shared CNUI.
MUI gives you components which live in a node modules folder. You don't really own the code. Yes, you own it because you installed it, but you can't really update it that easily.
With shared CNUI you own the components. They live in your components folder so you can customize them from the inside. With MUI you have to update the components, you have to customize them from the outside.
That's why I I don't use MUI anymore, and that's also why Shared CNUI essentially won. Since the rest of the code looks good to me, we can now create a PR.
I will click on the Create PR button and with that the agent will create a PR. Wow, who would have thought? And since Cloud Code now created a PR, I want to also show you a nice little trick because we will now have Cursor Bugbot run and try to find issues.
What we can do right here is click on this CI button and I want you to click on Autofix CI and Address Commons. So normally when I work on my projects, I have a full on CI pipeline with about seven or so GitHub actions and they fail sometimes, right? Because mistakes are made and stuff like that.
And with this little option, we are able to tell Cloud Code, hey, please listen to our PR, right? Listen for GitHub changes or GitHub action changes. And if something fails, then please fix it.
Or first of all, look at what has failed and then see if it's valid. And if so, then fix it. That's a very nice feature.
I hope Cursor also enables something like that or introduces something like that in the future. But yeah, and this worked out beautifully. So Bugbot found two issues of inputs Zoom on iOS and also code slots overflow the form card.
If I now go back to Cloud Code, then you will see here that it received a CI event and it said both findings are valid. It fixed both issues and now Bugbot is happy. So what we can do is merge the pull request.
I will confirm the merge and this will merge right into main. Since this is finished, we can also check out what our other session did, what our other lead agent did. It worked on the two dashboard pages.
What do we have inside of here? Damn, a lot of text. Both implementers are running in parallel.
We already know that. What else do we have here? Both pages are built and pass PNPM type check and PNPM lint, but nobody has looked at them.
in the browser yet? Wait, what? This was a core instruction.
The browser pane is still on the sign -in page and neither Implementor nor I can enter your password. So this is probably the biggest issue I have with Cloud Code. Some super smart engineer working at Cloud Code thought it was a great idea to disable this functionality, meaning agents running in Cloud Code can't sign in even though you already gave them the correct credentials.
the correct email and password. This works in Cursor, this works in Codex, it literally works in any other agent environment in any other ADE, but someone at Cloud Code thought, oh my god, we have to be super, super secure, like, let's disable this, we don't want to allow agents to sign in. Who thought that this would be a great idea?
Like, are you mad? What the hell? So yeah, the browser, or I'm sorry, the agent did not verify anything.
It generated code. It of course checked everything, but it did not use its eyes, which is super annoying. So I guess let's verify everything manually.
I will open my browser. And yeah, I guess you can instantly see it that the agent did not verify its changes. Because this definitely does not look good or beautiful.
Is it usable? Definitely. But does this look good?
No. New visitor after a message. This is quite cool, honestly.
What do we have here? Agent. I can switch it on or off.
This does not look very good. Appearance. Is this all real -time or does it update in real -time?
Yes, that's good. I'm not sure what's wrong with the avatar. It's completely broken.
We can update the accent color. That's good. You see we can update the position This could also be improved.
It does not look that good. We have our greeting. If I update it, does it update in real time?
Yeah, let me again change that. Suggested questions. This is hard -coded.
Allowed domains and also install. The install script is also absolutely ugly. It's not formatted or anything like that.
This is something we will have to change. Let's go to the inbox. This looks a bit better, but that's because the agent used already a pre -built block, so There's not really a lot that the agent can do wrong if that makes sense.
So visitor from Berlin We have the conversations pane this all looks somewhat good And then finally we can also view like the details of the visitor location local time language, etc So there are a lot of improvements that we can and also will make but here's the thing I feel like this agent does not really understand what we are going for.
It does not really have the for design if that makes sense and that's because we never really prompted the agent. But if you remember correctly we used a different agent to work on the login page.
It already understands what we like and what we don't like. So maybe it would be a good idea to use the other session which we used to build out the login page to also improve the dashboard pages. And yes that's a great idea and that's exactly what we will do.
So let's head over back to the previous session where we created the PR. I will say the following. The PR has now been merged.
Here's the thing. Can you switch out of the work tree, go into main, and then pull all of the recent changes? I'm not even sure if the agent can do that because, again, we created a work tree, right?
So the main checkout has uncommitted work of yours. That's all fine. No overlap with your local changes.
Main is now inside of here. What about the other changes? Are they in a different branch?
I'm not sure. No, they are in main. That's also good.
to see. What we have to now do is again let the agent spin up the dev server so I will say the following. Can you please stop all of the previous or currently active running dev servers and once finished create a new one or spin up a new one and open it in the built -in browser.
One thing I will do here is switch the reasoning effort. Instead of using medium I will use high reasoning effort and that's because high reasoning works a bit better with Sonnet 5 .5. If you use Opus 5 .5 I would recommend to use medium reasoning effort.
If you use Sonnet 5 .5, I would use high reasoning effort. So a new dev server is running. That's all good.
Let's use the built -in browser right here. I will sign in and again since the agent can't do it itself, I will have to do it manually. And here's what I will now do.
I will tell the agent that there are a lot of things that I don't like and I want it to just improve the whole page and it should use the same workflow as it did when it updated our of pages. So let me say the following.
Look, you did a great job when working or updating our authentication pages, and I want you now to do the same thing for our dashboard pages. So first of all, we have a general dashboard layout with the sidebar. I feel like we could improve it, and also when we collapse the sidebar, I would like to still render icons, if that makes sense.
So we could render the logo of our company, Marshall Desk, and then we could render two icons one for home and one for inbox and then on the bottom left we would probably also or we could render the user image also as a button so that we still have a sidebar even though when collapsed. That's the first thing I want you to work on the sidebar or on the layout.
The second step is to work on the home dashboard page which we have here essentially the widget settings. The general idea here is good. We have a 70 or a 65, 35 split, something like that.
And the split itself is good, but the UI is definitely lacking. It does not look good. It's not visually appealing.
It looks very AI generated. Again, please look at our design .md file. Look at our design language.
Look at the login page. Look at the homepage. They are very, very specific or they have a very specific design language.
And I want to use the same design language also. my dashboard pages. Another thing I instantly see right here or another thing I don't like is the install snippet which we have.
It's not formatted. That's something I want to change. It should be formatted.
It should look better or it should look beautiful. One card that is missing right here is a card or an input to drag and drop because later on we want to drag and drop files. Look at my PRD to learn more about the product.
And then on the right side we have the preview of the widget. The preview itself is good. It's what we wanted but the UI of the widget is also not as good as it should be or could be.
Another thing I want you to work on is then the inbox page of our dashboard. Currently we use a mail component provided by shared CNUI or a blog and you definitely see it right here. It does not have to do anything with our current design language.
It looks like a completely different product and that's again something I don't want. So use the same workflow you used when working on the authentication pages and just in general improve literally everything.
Maybe we can again add some images from Intercom. So let me go to the Intercom website. And what I will do is copy again a few images.
So copy image, paste it inside of here. Now look, preferably you should probably create a references folder to not always copy paste everything. But I don't want to do that right now.
It's easier to just copy everything. So copy, paste. What else can we add right here?
I don't like that. No, should we also copy? paste this.
No, I think that's already enough. So I will update my prompt and say the following. I also added screenshots of an competitor, which has a quite good design language.
I'm not saying that you should copy it at all. It's not something you should do, but rather I want you to just look at it so that you get a general understanding of how competitors design their pages and how they structure things. So that's already it.
We will now use Sonnet 5 .5. The biggest benefit here is that Sonnet 5 .5 is a very fast model, which means it will also allow us to iterate over everything very quickly.
Our agent is now finished, and you know what? This looks better, definitely, but it's not perfect. First of all, we have this sidebar now.
Let me make this a bit bigger. We have this sidebar, so even though it's collapsed, we still view or we can see all of the icons. But I feel like it's too small.
And if I open this, does it still work? Yeah. So here's what I will do.
I will make this a bit smaller. Let me select everything. Let me select the whole bar.
In other words, I will create the following prompt. Look, great job on the side. But nevertheless, I do have two issues.
First of all, I feel like the buttons, the items are too small. The icon is also too small or the icons are too small. And at the same time...
I'm not sure if we should work with like an accent color because currently we have a very monochrome theme and I'm not sure if that's such a great idea. So please also let me know what you think about that. Let's now look at the pages itself.
This here is our widget page. Let me again open the sidebar and I will also zoom in a bit. as you see here if the page is too big then we instantly have this problem right here where the model or where the widget is it can't grow any further right horizontally so what can we do i feel like we should make the preview a certain size and it should not become any smaller or bigger therefore what i will do is can make this a bit smaller let me select this whole preview thing and say the following when looking at our preview it looks good don't get me wrong but the thing is if we have too much space in terms, or if the screen that the user uses is too big in terms of horizontal size, then our preview panel grows automatically, but the widget stays, or yeah, has the same size, right?
It does not grow. horizontally. That's a good thing, we don't need the widget to grow horizontally.
Instead I would probably lock the preview to a certain size so that it looks good and then let's just make the left block, our 65 % block, grow or in other words become smaller. So instead of having like this very specific split of 65 -35, let's make it dynamic. What else do we have right here?
So the avatar is still broken, that's one thing we will have to fix. I'm also not the biggest fan of how the position currently looks like. It isn't as fancy as it could and should be.
It's also very, like, hard to understand what it will do right here. Like, this bottom left indicator, because that's what we have. We have, like, a mock -up of this bottom left.
it feels like it's a bit hard to understand with this like little dot on the bottom left so that's something we will have to change the rest looks somewhat fine it the biggest problem for me right now is that we don't have any color i need some color let's also look at the inbox what do we have here the inbox already looks way better and on the right side this also looks better so let's do the following i will quickly select our avatar first of all right here i will say the following the avatar is still kind of broken i'm I'm not sure why it does not render.
Also, is everything mobile responsive? Does the widget render in terms of mobile responsiveness? Well, we still have the same issue with the width, so that's something we need to fix.
The inbox itself is fine, but the biggest problem I guess right now is that it's very monochromatic. We only have like three colors, white, gray, and black. I feel like...
It does not look as good as it could and should be. Let's add some color. So right now, don't do it.
I'd rather suggest what you would do. How can we maybe give this whole thing, this whole page a bit of color, a bit of contrast? Oh, and by the way, when you select elements, as I just did a few seconds ago, then Cloud Code instantly also creates screenshots for you.
So you don't have to manually use a screenshot tool. That's why I like Cloud Code. You don't really have that with Cursor, which is a little bit annoying, but it is what it is.
Our agent is now working on the prompt. I'm comparing other dice pair styles for the agent avatar since the abstract shape looks cut off correct. Our agent is finished.
What did it say right here? What I fixed? Sight but too small, preview width, avatar and also mobile responsiveness.
Proposal for color not implemented. The dashboard feels flat because nothing separates what needs me from everything else. I'd add color where it carries meaning and keep the rest neutral.
Conversation states get their own colors, use the workspaces widget color across the dashboard, and also tinted surface, status color only where it's earned, the marketing pages stay monochrome. This makes sense to me. Let's also quickly open the browser.
This already looks a bit better. The avatar also finally works. If I collapse our sidebar, we also have bigger icons, but we are missing a bit of margin in between.
So let me say the following. um yeah your proposal is good so for color let's implement that Also, what I'm currently missing is a bit of gap between our items in the sidebar.
So between the home button and the inbox button, we need some gap in between, maybe margin, whatever you want to use. Another thing I see is it feels like the sidebar on the right is like a bit bigger than on the left. So it seems like on the right side, we have more padding than on the left side.
Please verify that. For the avatar, I'm not the biggest fan of the style you used. I would rather like to use the glass style.
It looks a bit better, a bit more premium. if that makes sense. I would probably also want to improve our preview, the widget specifically.
It does not have the same design language as our homepage, and that's still a theme that we have right here. Like, the landing page looks 10 times better than the dashboard, which makes sense, but the landing page also has a completely different design language. I feel like inside of here, it could also make sense to use our pixel font for certain cases, maybe for the eyebrow text.
Or for the labels. I'm not sure. This is maybe something you can disagree with me on.
That's absolutely fine. But that's something we should improve. Another thing I don't like is the position selector.
However you want to call it. We have these like two cards. But the thing is it's very hard to understand.
Like what's bottom left. What's bottom right. Because the indicator in this position is super small.
Something else could work better. It could make it easier to understand. But also it could look better.
Because this also looks very basic. Is there anything else I don't like? The rest looks somewhat good to me.
The install script is also missing some color. Since this is an install script, we could actually rather use a syntax highlighter. Also index .html.
Yeah, I guess that works. Why not? And if I go to the inbox, the inbox also looks good to me.
On the right with the details, we should probably also render like the user avatar, even though we don't know who the user is. It could still make sense. Let's just use Dice Bear again.
because it will give us a difference between or yeah it will differentiate all of the individual visitors. Our agent is finished and what should I say this already looks 10 times better. Let's check out the result in a real browser and our sidebar fully works as you see here.
We also now have some color and this I don't know, it looks way more refined and that's what we wanted. I also like that this text, let's call it a label, now uses the pixel font.
You know what, this is kinda similar to Nothing. I think you guys all know what Nothing is, this phone brand. It uses the same font, it looks super fancy, super futuristic.
What else do we have here? We can toggle our agent on and off and let's look at our appearance. The cool thing here is, if I change the color to green, for example, then everything updates, but at the same time also this glow updates if we want to call it like that.
Let's again make it blue. I like blue a bit more and the agent color also right here in the widget changes right away. Man, this is super nice.
What do we have here? New visitor after a message. This all works.
What can we do here? Position. This is a bit too big or yeah, it's a little bit too big.
Let's be honest. We have to probably change that. What else do we have?
Messages. this all stayed the same we have our knowledge base with some color ready ready processing and failed here we can add allowed domains and finally the install snippet is now also formatted but now get ready for a flashbang look at this oh man what is this wow this is absolute slop this is the definition of slop this is not usable we can't ship that at all What did the agent do?
It did such a good job with our homepage and it completely broke our inbox. So let's go back to Cloud Code and I will say the following. Look, the widget page itself, so the home dashboard page is perfect.
I don't have any complaints. The only thing I would change is the position selector. The cards currently are too big.
I would probably, I would make them a bit smaller in terms of height. And at the same time, the actual widget, which is then rendered in the card is too small. If you understand what I mean, so the general cut has to become smaller and the widget has to become bigger in the cart That's the first thing then the widget page itself is finished I don't think that there's anything we need to change but going to the inbox Man, I'm not sure what you did, but this looks absolutely disgusting.
It looks like AI slop and this is not shippable. Your homepage or the homepage you created is perfect. If the user selects an accent color like green, for example, we instantly have a glow on the top right.
Our accent in general changes. And if I go to the inbox, it also does the same thing. We get our accent color, but now also the panel is green and everything looks weird.
we have a new message, we have like this vertical yellow or orange line. This does not work. So the glow on the top right or on the top is perfectly fine.
We can do that. But the panels should again have their normal color, which is gray or black, whatever it was. Completely revisit this full page.
This looks like AI slop. Look again at the dashboard page you just created. Take some inspiration from there and rework it.
Our first iteration was better than what you now created. But the pixel font on the top left is quite nice. You can leave it.
Also this orange. color before waiting does not really work here. I mean orange does not work with our design system so I'm not sure why you selected that but please rework the page.
Our agent is finished and what should I say this already looks 10 times better. This is completely usable and we can ship this into production. Let's also go back to our home dashboard.
This still looks ugly, but you know what, let's leave it as is right now. We can change it at the end or again prompt our agent. Let's now work on our backend, because currently everything here is just mock data.
So if I update the agent name to, for example, MarshallDeskAgentTest123, and if I do a hard refresh... everything is lost right away. And that's because we don't persist anything in our database, which is not wanted.
So there are three things I want to do right now. First of all, I want to save everything in my database. Secondly, wait, there are more things, so let me go through it.
First of all, I want to save everything in my database. Secondly, I want to also persist images in my object storage, in my bucket, so in Neon. Thirdly, I want to also now create the needed iframe.
I want to allow the user or the business owner to also embed it in their own website. And finally, at the same time, I want to also implement real -time functionality. Yes, I know, there are a lot of steps right now that I want to do in one run.
But this is fine. Our workflow is built around a huge feature. This will absolutely work.
And is there anything we want to change maybe beforehand? I don't think so. This is completely usable.
And the reason why we can also build so many features... at once now is because the front -end UI is now finished. All we have to do is connect our back -end to the front -end.
All of the primitives already exist. ORPC, TAN stack query, our Postgres database, Neon, authentication, object storage, the AI gateway, everything is already there including also for example party kit.
Everything is already installed. Everything is already set up. So now all our agent has to do is literally just connect everything to the front end.
That's it. And that's why we can also now implement so many features at once in one go. So let me go back to Cloud Code.
And by the way, my usage limit is already almost used up, but that's fine. No worries. I think Cloud Code will pull through.
So let me create a new session and we want to again use my... workflow. What I will first of all do is again not use a work tree.
It's not needed. Then I will select Opus 5 .5 with medium reasoning effort. I want to use my feature orchestrator skill and the next step is to already create a prompt.
I will say the following. Look, we are currently in session one and there are a lot of things I want to now do. So we are currently working in the dashboard and one thing I want to do right away is save all of the data.
What do I mean by the Well, we have our Postgres database provided by Neon. You also have access to the Neon MCP server, to Neon Skills, the Neon plugin to be a bit more general.
And right now I want to save all of the data in my database. When the user currently or the business owner updates the agent name, then we don't save this data in our database. So when we do a hard refresh, everything gets lost, which is not wanted.
Secondly, I want to allow the business owner to also change their agent avatar. For that I already created a bucket for you.
It's a private bucket and what we will do is use the storage primitive provided by Neon. Since I want to later on deploy my application to Vercel we will have to use pre -signed URLs to upload everything on the client side but we want to also please authorize everything on the server side first of all. For the size itself two megabytes is fine we don't have to change that.
What I want you also to you is allow the business owner to right away also use our widget or in other words embed the widget in their own website we now have this script but it currently does not work yet and at the same time and that's i guess the final step right here is to also implement real -time functionality with party kit and with that party server so since we are currently in session one what do you think of that i think that's four features in total is there anything you have quest or do you any questions for me?
Is there anything that you need more info on? As you already know you can also use sub -agents to research anything needed. Please use sonnet 5 .5 for that.
And by the way, since our agent will now work on session one, we can also create a PR right away. For that, I will just go to the dev server setup, and inside of here, let me close this, and I can say, can we create a PR right now? I, again, already reviewed the code myself.
The code is clean, and that's because all we did is work on the front -end UI. The agent was now able to use our buttons and everything like that, so instead of creating custom styles with custom class names the agent created variants as already wanted and when working on the front end there's not a lot that the agent can do wrong or incorrectly it's not like the back end the back end is way more important than the front end code quality what do we now have here the PR is open and ready for review branch my changes and the uncommitted widget settings and inbox code that was already in the main checkout we're all sitting on main that's all fine.
What we can do inside of here is also again click on the CI drop -down and I will select autofix, CI and address comments. Let's go back to our previous session. Our agent now mapped out the basic game plan, how I'd group it, save widget settings.
What do we have here? Add columns to workspaces. This is a table in our database and then also add to procedures.
Agent avatar, a pre -signed procedure, picks the storage key. and returns a five -minute upload URL tied to the file's content type. This sounds great to me.
Then the browser uploads straight to the uploads bucket. This is still a thing we have to discuss. Should we use a private bucket or rather a public bucket?
Because for the data itself, so for the knowledge, which we will then allow the user to later on upload, it definitely makes sense to use a private bucket. But for the avatar image, it probably would make sense or would make more sense to use a public... bucket.
It will make things a bit easier so we will think about that. Then embed. A small embed .js loader creates one iframe that resizes itself with the launcher inside it.
This is the intercom and crisp pattern. That's good. So we will follow the pattern that already huge companies use and goes full screen on phones.
Conversations and real time. This here is very important. So new conversation and message tables.
That's needed. Visitor procedures to start a conversation, send a message, list messages and ask for a person, owner procedures to list, reply and so on.
App Realtime holds two party server parties, conversation and workspace with hibernation on. Why do we need a workspace server? Or why do we need a workspace party?
This is something I will ask my agent in a second. So here we can say the following. Public bucket and also say why two parties.
This is something I will ask my agent in a second. Questions. My recommendation first.
Real time without the agent. New conversations would sit in AI with no agent replying. Recommended.
Build everything. else and design the events for streamed agent text now, but leave the agent pipeline for the next session. Until then, AI conversations show in the inbox and you can take them over.
Sure, why not. Secondly, showing the avatar to visitors on other websites, an XJS route like API avatar key reads from the private bucket and returns an image that can be cached forever. So this is what I meant.
This is like an extra step that is not really needed. Avatars or images itself are public, so... I would probably create a second bucket.
I think that makes a bit more sense. Thirdly, how the dashboard connects to party server. Our tech stack file says the worker verifies our NeonAuth token directly.
Recommended Next .js signs a short -lived token for one room after owner procedure. This is one thing I wanted to talk about. So here I will create another thing, which I want to later on ask my agent.
So question three. please again use a few sub -agents and see if this is also what the docs recommend. I want to see what the documentation recommends in terms of authenticating the user.
Question four, building the embed script at ESBuild as a dev dependency of Erp's web. It's already in the log file indirectly and chain it into dev and build. Yeah, that's fine.
Other things still mocked on the settings page. Suggested questions, knowledge sources, the setup checklist and your own profile photo. Recommended leave them out of scope.
That's fine. We can leave them out of scope. Before session two, I'd suggest you commit the current uncommitted work.
So let me create the following comment or the following prompt. Look, this already sounds good to me. There are a few things that I would probably change.
Currently, we only have a private bucket. I think we should... Also create a public bucket for the avatar images.
This will make our life a bit easier. So we could for example call it user profile images or something like that. Since you have access to the Neon MCP server you can also create a bucket yourself.
Another thing I would like to know is why we need two parties. Here you mentioned we need a party for conversation and then also for workspace. This hibernation step makes sense to me but why do we need two parties?
What else do I want to ask you? So in regards to question three, you said how the dashboard connects to party server. Could you maybe spin up a few more Sonnet 5 .5 sub -agents which directly look at the documentation?
So add the GitHub repo which lives in the Cloudflare workspace or however it's called and then also look at the legacy party kit documentation. What I'm currently again doing is playing ping pong with the agent. We are building a shared understanding with the agent together and at the same time the agent is now also researching the web for further resources because i'm not quite sure if i would agree on this step right here where the agent said that we should right here use the worker which verifies our neon off token maybe this is correct i don't want to open the documentation right now i want the agent to rather recommend or i'm sorry verify the documentation for myself and then based on that i will then be able to either give the green or the red light our agent created a new bucket profile images it's a public bucket which is good and it's also in the correct branch in our development branch private uploads knowledge sources attachments stay in uploads what this changes in the plan there is no image proxy route anymore and yet this was also kind of a stupid idea the reason why the agent proposed the proxy is because it did not really want to change the plan which we already set out in our
PRD with the bucket. So that's always also or that's one big problem of a PRD. A PRD locks the agent into a specific path, right?
And if the agent even would want to disagree with you, it won't do that. So that's one problem with the PRD and that's also why you should always question the agent. Even if the PRD says something, ask the agent, do you also agree with the PRD or do you disagree with the PRD?
What would you change? What would you optimize? Etc.
Session 2 needs to course on this bucket so your dashboard origin can't upload it. That's fine.
Why two parties? Strictly speaking, you don't need two. What you need is two kind of rooms because there are two audiences.
A conversation room per conversation itself and a workspace room, one per workspace. Ah, okay, I understand. So only the owner's dashboard joins it, it keeps the inbox list live, a new conversation appears, the waiting count changes, notifications fire, etc.
Okay, this makes sense to me. The dashboard needs the workspace room because it can't join the room of a conversation it doesn't know exists yet. In PartyServer, a party is one durable object class and a room is one instance of it.
So you can build this either way, two classes, conversation and workspace, or one class with room names like conversation - and workspace -. Either is fine, tell me if you would rather have one. Honestly, it does not really make sense or it does not really change anything.
If we would use two classes, we would use two durable objects. Maybe this would be a bit more expensive, but at the same time, it will also optimize everything and distribute load a bit better. I would probably go with two classes, but that's for now the furthest.
Then dashboard off. Three more Sonnet 5 .5 researchers are on it. Let's see what's the result.
Dashboard off. What the sources say. Current.
repo, Cloudflare PartyKit, there's no official auth recipe, the workers before on connect hook can reject the connection with a 401 or 403, then party socket supports an async query function that runs again on every reconnect, legacy PartyKit documentation, the same pattern, the token goes into the connection URL, browsers can't set headers on a web socket, it's fetched fresh on each connect and the hook checks it before the room wakes, this sounds good to me.
me recommendation a new real -time get token procedure on owner procedure issues a short -lived token about 60 seconds is there anything else decisions so far settled the avatar goes into the public profile images bucket the limit is two megabytes development runs on the development branch that's wanted your call two classes my recommendation or one class with room name prefixes let's use two classes because again it will distribute the load more evenly, which also will allow us to scale even better.
Open questions, my recommendation in bold. Leave the agent pipeline for the next session, agreed. Avatar serving, this is already resolved.
Dashboard auth as a buff, yes, agreed. Build the embed script with ESBuild as a dev dependency, agreed. Leave the other mocked pieces out of scope, agreed.
And commit your current UI work, this is already committed. Let's quickly check our PR. I will open GitHub.
What do we have? inside of here. Our pull request should be successful, right?
We had an issue that came up. So what was the issue? Mobile inbox overrides layout cookie.
That's not good, but our agent fixed everything. Save layout now returns early when the mobile media query matches. What we can do here is merge our pull request.
I will do that. Let's also go back to our agent and I will say the following. I now first of all merged our pull request.
So please switch into main. and pull all of the recent changes and then in terms of the decisions you already settled for everything sounds good to me and all of the open questions are or i agree with your recommendations leave the agent pipeline for the next session then the dashboard off as above that all sounds good to me we want to also use our embed script with esbuild as a dev dependency let's also leave other mocked pieces out of scope that's something i want to work on later on and for our work as mentioned everything has already been merged.
Is there anything else you want to know? Do you have any other questions? Is there anything else you want to first of all get my opinion on?
Here we now have round three with a few questions. Let me answer them already right away. I will again use my voice dictation software.
So in terms of how widget settings save, I don't want to use an explicit save changes button. That's not needed. I want to use autosave.
We need to have some sort of good delay, right? Because we don't want to save on every keystroke, but autosave sounds better to me.
Nevertheless, we need some sort of indicator, maybe on the top right, maybe you have a better idea, which then says either saving or saved or failed, whatever. Maybe we can render an explicit button if it failed, so if autosave failed, but that's something you can decide for yourself. Visitor details in the conversation site panel, country, device, browser, etc.
Recommended capture them now. Yes, I want to capture them right away. So let's go with your recommendation to install the needed dependency for geolocation and then also the UAParser .js for the user agent.
For the notifications, what do you recommend right here? The widget sound when a reply arrives. Yes, when the tab is in the background and the dashboard tab badge and sound.
This is also all fine. I agree with your recommendation. scope with your OK browser notifications they need a permission prompt which would only appear after you click something.
Sure, why not. What else is out of scope in this session? I recommend leaving all of these out.
Image attachments in the widget and inbox. Yeah, let's leave this out. The 24 hour auto close since it needs a neon function and app functions itself.
Uh, yeah, let's also leave that out. Deploying the worker to Cloudflare. Yeah, we will leave that out.
We will use the Wrangler CLI, but this whole deployment is something I want to do at the end. It does not need to happen right now. And then in scope, talk to a real person, agent of.
That all sounds good to me. Question five, what the visitor sees while a conversation is in AI with no agent yet. Recommended nothing special.
The message saves and the owner can take over. from the inbox. Yes, let's do that.
How to split it into PRs? That's a great question. Three stacked PRs, saved widget settings, embed script, and then also apps real time.
Yeah, sure, let's do that. I think three PRs is fine. We could also do maybe four PRs.
But three PRs is also fine. I would definitely stack them. So please do that.
And yeah. So is there anything else you want to know? One setup check.
Session two will need the development branches storage credentials and new secrets. The publish secret, the visitor token secret in env .local. It can generate the secrets itself.
It can pull the storage credentials with neon env pull or the neon MCP server. I'm happy with that. Sure.
Do that. And let's now work on session. or in other words, generate the game plan prompt for session two.
One thing you might say right now is, Jan, you made a huge mistake. Why did you just give the agent the ability to use or pull the secrets, the needed secrets, using the Neon CLI or in other words, the Neon MCP server? Isn't that a huge mistake?
Won't our agent now use this or won't all of these companies like Anthropic, OpenAI, et cetera, use this data for training data, in other words, Won't our agent now know all of the secrets? Blah, blah, blah.
Yes, but that's not a problem. Why isn't that a problem? What branch are we working in?
Development, right. This here is our development branch. We can leak these secrets.
It's fine. Once we deploy our application, we will have a new branch, the production branch with new secrets. So this here is absolutely fine.
You can leak everything. You can give your agent access to Neon. You can give your agent access to the environment variables because they will change in production.
One thing I want you also quickly do is open the Neon console. Do we have our new bucket inside? of here.
Yeah, profile images and it's a public bucket as you see here. Is there anything else we should check? Do we still have all of our users here?
Yes, it also all works. And for the functions, we currently don't have any function yet, but that's something we will do as a next step. Right now, we want to essentially create the needed foundation and then step two will be to create the function and also continue or iterate further.
Let's now again go back. We need to select our models. What options do we have here?
Which model and effort should session2's implementer subagents use? Opus medium, sounds great to me. Which model and effort should the researcher subagents use?
Let's use sonnet medium. And then which model and effort should the reviewer subagents use? I guess we should use Opus Medium.
Sonnet is also not a bad idea because let's quickly check it out. I will open CursorBench, so cursor .com slash CursorBench. And as you see here, Sonnet 5 is a strong model and at the same time, it's a cheaper model.
But when you look at high reasoning effort, the price is almost the same and the performance is better with Opus 5 .5. So what I will do is use Opus 5 .5 with medium reasoning effort. for our researcher, Opus Medium.
This is now finished and our agent will write the prompt or create the prompt for us. Our agent is finished. I will now copy the prompt and I will use Cursor for the implementation because Cursor has a bit of fewer restrictions in terms of the built -in browser.
Let me create a brand new session, Marshall Desk, Main, that's all good. Let me also switch the model, Opus 5 .5, with a 1 million token context window. and medium reasoning effort and I will paste the prompt inside of here.
Let's quickly glance at the prompt and check out if everything looks correct. Widget settings saved to Postgres with autosave. Agent avatar upload to Neon Object Storage through pre -signed URLs.
A working embed script with a real widget. Realtime with Cloudflare party server. You deliver them as three stacked PRs worth of work, that's also all fine.
Environment not code. Let's quickly change that.
This should be rather cursor, implementer opus, researcher sonnet, and reviewer opus. Let me also add the model opus 5 .5, sonnet 5 .5, because cursor has multiple models, and then also again right here, opus 5 .5. Then limitation that we can delete this because this is cloud code specific.
This is not really important. You are the lead agent repository. That's all correct.
Authorization boundaries. This is everything. I already authorized.
What else do we have here? Orchestration. This comes from our feature orchestrator skill.
And since I already mentioned the feature orchestrator skill, let me also annotate it right here. Feature orchestrator session two. Let's again continue with the prompt.
Is there anything else we have to check out? Build with your eyes open. This means use the built -in browser.
This shouldn't again be cloud code. So let me delete these parentheses. This is not needed.
The built -in browser is your vision. Use it. while writing code, not as a final checkpoint.
What else do we have right here? Mandatory separation or preparation. Readagents .md, our PRD, techstackdesign .md.
That all sounds good to me. What else do we have here? Existing product and architecture.
So this is essentially like a little, how could you say, summary of what has already been created. And then we have our individual PRs. PR3, real -time with party server, apps real -time.
party classes because again of distribution and owner visitor conversation id the worker never verifies neon off tokens and has no database access neon off doesn't support custom claims so its token can't carry workspace id interesting this is very interesting explicit exclusions so this is what the agent should not build the agent pipeline classify embed retrieve that's correct knowledge source image attachments in the widget or inbox, the 24 hour auto close and I think that's already it, right?
Behavioral and failure contract, auto save failure, so this is what should happen if something fails. This all looks Perfect to me.
What I will now do is click on enter and our agent will get started. Our agent is now finished. It generated or wrote all of the necessary code and it also created three PRs which stack on top of each other.
And now guess how many lines of code have been added? Correct, 35 ,000. And also 1 ,600 lines have been removed in terms of code.
And that's quite a bit. It's a lot to be super honest with you. Since it's broken down into three PRs, we could say 10 ,000 lines per PR.
That's manageable and that's also reviewable. So let me quickly talk about the workflow in general. What did the agent do?
Well, the agent, the lead agent, started by first of all spinning up sub -agents for research specifically, Sonnet 5 .5. And that's also what our prompt specifically mentioned. This is again the biggest benefit of using the feature orchestrator skill.
Then we used implement or sub -agent. In total, what do we have here? Five, seven.
Seven sub -agents in total for implementation purposes. For PR1, PR2, and PR3. Once everything was implemented, we then had reviewer sub -agents.
And we in total had two review cycles, two fixes, five reviewers. So in total, seven reviewer sub -agents. Wow, that's quite a bit.
And it actually, or the agents found... Quite a few issues, which was very surprising. So about 5 ,000 lines of code have been changed after the reviewers.
So that's also, again, a huge benefit of using reviewer sub -agents. Is there anything else we need to talk about? Well, we will have to review the code, definitely.
No questions asked. The agent, of course, used the browser. I also had a few screenshots.
Let me see if I can find it. No, I can't find it right now. But the agent created a lot of screenshots.
In other words, it used the browser, it took a screenshot, looked at the screenshot and then changed the code or in other words, refined the code to make the end result look better. And it also verified that the whole flow works very well. So if you would now ask me, hey Jan, my code or my agent generated code, what should I now do?
In my opinion, the first step is to always first of all read what the agent pretty much gave you or the handoff document, that's how I call it. First of all, read through it. Try to understand why.
your agent did, how it did that, what workflow it used, etc. The next step is to not review code, no no no, but first of all to test everything out. Does the result, does the whole thing actually work?
If so, only then you should start reviewing the code. So let's test it out. This here is our normal application, Marshall Desk, and here we also have a testing application which our agent created, Acme Outdoor Co.
Inside of here we can pass our script and then test out if our widget works. So let's start from scratch.
First of all we want to see if our database is connected. I will say the following. Instead of MarshallDeskAgent I will say JansAgent.
Do we get some sort of autosave pop -up? No. Only on the top we see saved.
And I don't really like this placement. It does not look that good. Let me try it out one more time.
JansAgent123 saving dot dot dot and then saved. So this works that good.
Let's do a hard refresh. Does everything stay persisted? Yep, that's wanted.
Let's change the color orange. I will again do a hard refresh. Everything stayed persistent.
One thing I probably don't like is currently the color of the font right here or of the text. Because it's black. But this here is a green background.
The contrast is not correct. That's what I'm trying to say. Let's again switch back to orange.
Let me also maybe change the avatar to something super random. We have a nice little pending state. And now we have a Ferrari.
My god, this looks ugly. But don't worry, we will change it later on. What else do we have here?
Let me also change the color to red so that it makes a bit more sense. We can update our greeting. And by the way, can I also scroll in?
the widget itself now and if I zoom out a bit. Hi there, how can I help you today? Let's maybe also add some sort of emoji.
Let's use the, I don't know what emoji, this clown emoji. Yeah, sure, why not. If I now do a hard refresh, we still have everything persisted.
That's good. This here is still hard coded. This is also hard coded.
Allowed domains, this is empty. And then we also have our install script. Let's quickly also open the sidebar.
I will go to the inbox. This here is completely empty but that's absolutely fine. Currently it says install the widget.
What happens if I click on the button? We get redirected. Let me now copy the installation command.
I will go back to our Acme Outdoor Co website, paste it inside of here and then embed and reload. Widget embedded for workspace blah blah blah. Huh.
I don't see any widget, which is not good. What's the reason for that? Well, that's because we did not add the domain to our allowed domains.
And this is something we should probably change. Localhost domains should probably be enabled by default, at least in my opinion. Let me copy the localhost 4000 domain.
I will paste it inside of here. Delete the slash at the end and click on add. It has now been added.
And if I go back and do a hard refresh, does it now work? Yes, we get the widget. So I guess, can we render maybe a warning inside of here if the domain hasn't been added?
I think so. So that's also something we should probably keep in mind so that we can discuss that with our agent. Let's test out the widget.
First of all, we don't really have an animation. Is that an issue? Maybe not really it's not a huge issue.
What I don't like here is that the badge has a black background color This does not work or this does not look good Let's also write a message. Hey, how are you?
And if I click on enter, we shouldn't really get a response because currently our AI agent is not connected. Let me also head over back to the inbox and we instantly also get a message inside of here. What I can do is turn on notifications and that's also what I will do.
I will turn them on, reload. Sure, that's fine. And here we have our visitor message.
Hi there, how can I help you? I don't like this contrast. That's a issue.
Let me quickly take over or I don't even have to take over. I can just write something or respond. Let me also close the sidebar, zoom in a bit and I will say yo and let me click on send.
What do we get right here? We instantly get a notification in our widget and we have a new response from Jan Marshall. Yo, that's quite nice.
I want to test the following. What is if I do or change the color to, for example, green and then go back and do a hard refresh? Wow, I don't even have to do a hard refresh.
So if I switch it to purple and go back, it instantly updates. That's super fancy. What else can we do?
Bottom left, does it also update? Yes, let's again say bottom right, it instantly updates. If we go back to the inbox and if I now create a message, so I will say, do you, I don't know.
get a notification and if I click on enter we instantly get a notification here and it says do you get a notification. We also have like a timeline you took over just now. This also all looks great.
One thing I want to test out right away is if we also get a typing indicator. So let me align that on the right side and let me align this on the left side and I will now just write something bum bum bum and yes we get this typing indicator. that's all so fancy and if I stop it should now take a few seconds because we are using heartbeats and yes it stopped and what is if I do it on the other end meaning in the widget I will say yo and yes we also get a typing indicator again we have a heartbeat so it takes about five seconds for everything to again refresh if that makes sense but that's normal that's what most applications do what also big players like intercom do so I have zero complaints is there anything we should maybe change?
I don't really think so. Everything looks good to me. Sure, there are a few issues, let's be honest, but most of the things work and that's also what we wanted.
So here's what we can do. Let's head over back to Cursor. I opened the built -in browser.
Let me also make this full screen and what we can do is use this design mode. Let me click on the button and what exactly do we want to change? First of all, we can select our preview and what will do is say the following currently when the preview is zoomed in quite a bit it does not allow us to scroll in the preview itself in the widget so that's a basic bug i want to fix so that sounds good what else can we do i want to also select our position and then say the following currently inside of this position selector we have these two cards right but the thing is in the card we then have our widget the widget currently is too small in terms of height but also width Make it a bit bigger in the cart.
Don't change the size of the outer cart. Just the inner widget. So the widget which is inside of the cart.
Then what else could we change here? Let me maybe zoom out a bit. Because it's currently just a bit too big for me.
Then I think that's it for now. I will click on send. Our agent will get started on that.
What else could we do here? This all looks good to me. I add a domain.
That's one thing we have to talk about. Let me again go out of design mode. Inbox.
The inbox. is perfect, right? So I don't think that we have to change anything inside of here.
Let me make this a bit closer or smaller again. Let me zoom in a bit and our agent is now working on the two bugs I annotated or two issues I annotated. One thing I also don't really love is the placement of this agent toggle.
So to toggle our AI capabilities, it looks a bit weird and misplaced. It does not look good. So I'm not sure where to run that.
Probably somewhere in our appearance card maybe that's something we will have to think about but it has to definitely change its positioning. In terms of the inbox let me check it one more time.
Now the empty states are fine I don't think that we have to change anything and this is also fully mobile responsive as you see here. Let me again close this. So yeah this is the only thing we will have to change.
Our agent is still working on everything so let me already create the next prompt. I can again annotate it our agent and I will say the following currently I'm not a big fan of the positioning of this agent you could say toggle button whatever you want to call it it looks very misplaced both in our desktop but also in our mobile screens or yeah in the mobile screen and also on a desktop screen I'm not sure where to position it where it could look better maybe in the appearance section under the appearance section I don't really know.
I don't have the answer for you, the correct answer, but I know that this here does not look good or the positioning is not the correct one. What I will do is click on send because what it will do is just steer our agent and that's absolutely fine. It won't stop the other work that the agent has or is currently already doing.
And instead of now doing anything else or thinking about the next steps, let's just wait for our agent to finish because I want to test everything out one more time. Let's check what do we have here. Man, this This agent replies thing has just become even worse.
What the hell is that? This also does not work at all. Does our thing now scroll?
Yes, the widget now scrolls, which is better. The inner widget is still too small in the cart. The agent did not fix that.
Man, oh man, oh man. So let's create another follow -up prompt. Look, you did not do a good job.
Let me be absolutely honest with you. Think about it. Look at the browser yourself.
Does this agent replies cart look good to you? No, not at all. It looks absolutely ugly and misplaced.
Also, we don't really need such a big card for our agent replies toggle. What we had previously, the button was already good in terms of size. It was just misplaced.
We need some sort of other specific location for the button. Then, what else don't I like here? Well, you did not fix our position card.
On a mobile screen, yes, the height is better, but once you go to a desktop screen, the inner widget is still... too small. The inner widget has to become bigger.
Let the widget take 70 % of the total card, maybe even 80%. Again, I'm talking about our positioning. So this position card section, there we have the two selection cards, and in the cards we have like two widgets which we render to showcase or to show the user what this will do.
These cards, these widgets, sorry, not cards, these widgets have to be bigger in terms of height. What else don't I like? Let's talk about the widget itself.
The widget currently on light mode or in light mode renders this text answers only from your knowledge base. I mean that's all cool but why is this in or why does this have a dark background color? This does not work.
Also please make sure that this badge is only shown to the business owner itself in the dashboard and not in the website of the business owner. This is not something that visitors have to know. Is there anything else I don't like?
I don't like the badges so these recommended questions. Do you ship to Canada?
What's your refund policy, etc. I'm not sure what badge you used right here, but it's not the shared CNUI badge, or at least it's not the variant I would like to use. For the variant, I would probably use secondary.
I think secondary will work the best. Let's also test out what happens if I click on close, meaning, hey, end the conversation. Close and okay, it's now closed.
And where do we see the conversation? We see it in the closed tab. Can we also again reopen it?
No, this does not work. And if I go back to our visitors or to our visitor website and if I do a hard refresh, do we start with a new chat session? No, this conversation has ended.
Send a message to start a new one. Hi there, how can I help you? What is if I click on talk to a human?
Does anything happen? joined shortly, replies will appear here. And if I go back to the dashboard, we also have a new warning, not a new warning, a new badge, one waiting.
Let's go to the waiting section and it says right here, the visitor asked for a person. If I again say something and click on send, we get our notification. This all works beautifully.
Is cursor finished? Nope, it's still working on everything. So let's also wait for that.
Since our agent is now finished with the UI optimization part, we can also head over to the Neon console and check if all of the data has been propagated. And yes, this here is our public schema and we already have quite a few conversations.
We also can view our members, our messages, the visitors and also workspaces. And if we now go to the monitoring page, then you will also see here that we had quite a bit of load, continuous load. And that's because the agent has been working on this prompt, on this task already for about two hours or so.
That's why we have continuous load which is also quite nice to see here and that's our dip the dip is that hey our database wasn't running and with that I guess we are now finished right with this step we verified our changes everything works we checked out our neon console we can view all of our users and by the way I can also view the user in detail and for example also delete the user if needed we can view all of our data inside of our database table viewer and in general everything is working very well the UI is also beautiful.
There's nothing I would want to change right now. And I guess the next step is to continue with a new feature, right? No, this is a huge mistake.
Please don't do that. What I want you to do right now is review the generated code. Yes, I know it's boring.
Nobody wants to do it. It takes too much time, but it is what it is. You are an engineer, not a vibe coder.
Review the generated code, try to understand what has been created, in other words generated, and fix all of the slop that has been generated. Because even though we are using a great model, even though we let sub -agents review all of the code, even though sub -agents found also bugs and fixed them there is still definitely at least I don't know 500 lines of slop it is what it is like agents write slop they generate very complicated code and there's nothing that agents can do about it it's in their nature if that makes sense therefore I want you to now review the code find anything or not find anything you will find something and then just tell your agent that's what I will now do I will actually glue myself to this chair right here take off my hat turn on my hat right here and look at the code for 35 minutes straight.
And before I merge anything, I want you first of all verify that the code is clean. After that, yes, I will have my GitHub Actions running, Cursor's Bugbot and stuff, but before all of that runs, I will first of all use my own hat. So I will now get started with that.
And if everything looks good, I will then create a PR and also merge it automatically. And you will do the same without me. I finally reviewed all of the code, all of the 35 ,000 lines of code.
And believe it or not, the code was quite clean. I did not find any big issues and everything was good. Sure, there were a few things which were overcomplicated.
There were these annoying functions for error handling and stuff like that. all of that and the code in general was very clean. Nevertheless, Bugbot found a few valid findings, which of course is good.
You want Bugbot to do a good job. And let's be honest, not all of the issues were super severe. Let me maybe show it to you with this PR here.
Was it this PR? Nope, I think it was this PR. Let me open it quickly.
What did we have inside of here? The agent, for example, found this issue where I committed my personal test host path. Is this a huge issue?
No, it's a relatively small issue. Nothing will happen. I myself missed that.
I did not see it. It is what it is. But if this would have been committed to main, then everything would still have worked perfectly fine.
This is not a security issue or anything like that. There were also more severe issues. Like, for example, this one, stale list can override mutations.
That's another thing I missed. Because when reviewing code, my primary goal is to see or... Try to figure out if the agent wrote semantic, not semantic, but clean code.
I am not really trying to dig into every line and try to understand every specific line. I feel like that's not really needed anymore. And that's why I also sometimes miss this right here.
Stay list can override mutations. Now, this again is not a huge issue. This is medium severity.
This would still work in production. Nevertheless, it's good that BugBot found this and also told me to fix this. I fixed everything.
and that's very nice. This is why I love PR review tools. Yes, they are a little bit overpriced in most cases, but they are still very essential, at least in my workflow.
I need agenticness, I need my models to do as much of the heavy lifting as possible so that I can review all of the code very quickly and also that I can look at reliable code, not some sort of slop. Since this is all finished, I can also finally merge everything. I already created a stack, as you see here.
so we have a stack of three PRs. I will click on merge stack, confirm and this will then merge into main. What we can do is head over back to cursor.
I will open a new tab or a new session. Let me switch into main and inside of here I will say the following. I just merged three PRs so I want you to switch into main and pull all of the recent changes.
Also maybe please look at the recent or at the three PRs that I just merged and try to get a general understanding of what has already been done. Our agent pulled everything, what the three PRs added.
PR number one, or in other words, number five, saving widget settings and the agent avatar. Number six, the embedded widget with real conversations. And then finally, live delivery with party server, or in other words, party kit.
This is something that might be a bit confusing. When I say party kit, I mean party server. In the past, we just had party kit.
Let me quickly open the web. in Google, Party Kit. This was the website in the past.
It was called Party Kit. Now it's still called Party Kit, but the server is not called Party Kit anymore. It's now called Party Server.
It's a bit confusing, but the whole system is called Party Kit. It builds on top of Cloudflare Durable Objects. It got acquired by Cloudflare, and everything is again free and open source, as you already know.
Let's head over back to Cursor. What isn't built yet? The agent pipeline, the knowledge -based back -end app.
functions for ingest suggested questions and autoclose hasn't been created. That's a big problem. Setup needed before running locally, apply the new migrations, set visitor token secret and the real -time variables and then finally run PNPM filter blah blah blah once for your origin.
So storage course. Then start a worker with PNPM filter. Let's do the following.
I want you to right now first of all stop all of the active dev servers, then I want you to add the needed secrets so what did you say right here visitor token secret and then also the real -time variables please add them start the dev server also run pnpm filter martial desk web storage course and start the worker so i want you first of all again see that everything works and that everything is running actively Our dev server is now running again.
Look at this beautiful homepage. Man, this should probably win an award. But there's one slight issue.
In the navbar you will find the sign -in button. Currently we are not signed in, right? So if I click on the button we should get redirected to the sign -in page and then authenticate.
Right? Right? Right?
No. Let me show it to you. I will now click on sign -in and please look at the whole redirect flow.
Bum. Uh -huh. Aha, this does not look good, right?
So even though we were authenticated, our Nerf bar still shows the sign -in button. Then we got redirected to the sign -in page. And then from the sign -in page, we got redirected to the dashboard.
This does not make any sense and this does not create a good user experience. So I will create the following prompt. I currently have a slight problem.
In the Nerf bar, when the user is already authenticated, we still show or we still render a normal sign -in button. And then once the user... also clicks on the sign -in button, instead of redirecting the user instantly to the dashboard, we still redirect the user to the sign -in page.
And then in the sign -in page, we redirect the user to the dashboard. This does not make any sense, this is not optimized and it creates a bad user experience because we have a certain flash, right? So here's what I want to do.
I want to fetch the user session. in the homepage on the client side then we will get the user session we will either know that the user has a session or no session if the user has a session i would like to render a button which says for example dashboard i think that makes the most sense and if the user does not have a session then we will render the sign -in button also in the since we will fetch the user session on the client side we will have a certain loading state or some sort of like loading period while everything is getting fetched I want you to just render nothing null.
Or maybe, yeah, null I think actually makes the most sense. Now you might ask me, Jan, why do you want to fetch the user session on the client side? Well, here's something you should know about Next .js.
Next .js, or if you fetch data on the server side in a layout in Next .js, all of your pages, all of the children will automatically be rendered on the server side. A huge benefit of Next .js is client side, sorry, not client side, but stack. static site generation, SSG.
Next .js can make your homepage as an example completely. It can pre -render it on the server side and then render everything statically using a CDN. This makes everything super fast and the response is almost instant.
It's a huge game changer, a huge benefit of using Next .js. But if you now fetch the user session on the server side in a layout, all of the children, this page, our dashboard, etc. will automatically use the same rendering style.
SSR and this is not needed. The home page shouldn't be rendered on the server side. It should be statically generated.
We want you pre -rendered on the server side while building the application. The dashboard as an example, yes it should be rendered on the server side. The login page can be rendered on the server side.
It can be rendered on the client side depending if we fetch the user session. So that's what I mean. We should still give our application the ability to decide for itself if it wants to render everything on the server side or if it wants to pre -render everything at build time.
So let's go back to Cursor. Is it already finished? Nope, it's still working on our task.
All three pairs. Next, I will check it in a real browser. So let's try it out by ourselves.
I will do a hard refresh. We don't render anything for a split second. Let me show it to you one more time.
And that's because the user session is fetched on the client side. So what you have to realize is that this website first of all renders on the server side. whoop once the page is rendered on the server side our user session kicks in the fetch on the client side kicks in and that's why we don't see anything for a split second once we have the user session we then render the button let me now click on sign in aha i got redirected let me continue with google and i got redirected let me now again head over back to localhost 3000 we instantly see a dashboard button on the top right and if i now click on dashboard i should get redirected back to the sign -in page and that's also something you will see on the bottom left the path does not redirect anymore to slash off slash sign in but rather to slash dashboard so if i now click on dashboard i will get redirected right away let's quickly look at the code since it's only 44 lines of code account button .tsx what do we do inside of here class name style on navigate that's what we get from our props constant data session is pending off client aha
This allows us to get the user session on the client side. If is pending, return null. That's what I mentioned.
If there's nothing, just render nothing or return nothing. And once we then have the user session, we render a button with the correct link. And this here is quite smart.
Instead of rendering two buttons, we just render one button with then the correct link, render and stuff like that. That's super cool to see because in the past, what I always did manually is just render two buttons and then do an if statement. So this is even more optimized.
So that's great to see. Let's again close this. I don't have to really create a PR for this.
I feel like we can just push this into main. So what I will do inside of here is say commit and push. This code is clean.
There are zero issues. I don't need bugbot to run on this. It's clean.
I can push it into main. This is only something you should do if the code is very small or you are 100 % sure that the code is clean and that it will work. If you're not sure, then please use a code review tool or use a subagent to do something.
But in this case, I know the code. I know how it works. I can push this into main without having or without needing any reassurance from some sort of other agent.
Let's now continue with the next step or what is even the next step? Well, the next step is to already implement our AI agent feature. So here's what I will do.
I will create another prompt and I want to play ping pong with the session. So let's use our feature orchestrator skill. I will say session one and now I will create create my prompt.
Look, so we are now finished with the first, you could say, how should we say, bracket of this application, meaning we have our widget, we have our dashboard, we have our real -time functionality. And the next step is to already implement our AI agent feature. Nevertheless, as you already mentioned, we need to also use the serverless functions provided by Neon.
We need to use the AI gateway. We need to also do embeddings, right? Because if I go back to my system architecture, diagram, then we have to use embeddings.
We have to embed the chunks and then use also an open AI model for the embeddings. We have to then also use functions for the whole pipeline. I'm still not quite sure how all of that should work.
So could you maybe use a few Sonnet 5 .5 sub -agents or let's maybe use Rock 4 .6 sub -agents to figure out how everything will work. Could you also quickly look at the PRD, which we have, which already maps out exactly what I want and please compare it to what we already created.
Essentially I want you figure out what is missing right now. In total, we now had about six sub -agents running and they all researched different areas of this workflow, right? So what do we have here?
What's missing compared with the PRD? The knowledge base, that's very important. We want to allow business owners to upload knowledge.
The agent flow, send message, saves the message and stops. There's no classifier, retrieval, streamed answer or turn logging. Aye, aye, aye.
Not good. No, that's fine. Don't worry.
Suggested questions. They aren't stored anywhere. That's correct.
But that's because we currently don't have suggested answers. Or questions, not answers. Agent off when the knowledge base is empty.
Today only the manual switch turns the agent off. Auto close after 24 hours. Correct.
Image attachments in both directions. And account settings and workspace rename. Ah right, we currently don't allow a business owner to rename their workspace.
This is something we will have to change. How the agent feature would work, let's see. There are two pipelines, ingest runs in the neon function, answering runs in Next .js.
What do I even mean by ingest? Well, here's the thing. A business owner uploads knowledge.
What we will have to do is break this knowledge down into little chunks. And we will do all of that in the neon function. Because as already mentioned yesterday, We could use a Lambda, so when we deploy our application later onto Vercel, we will use Lambda serverless functions, which can run for 30 seconds.
But this would mean that we would, A, first of all, have a time limit, which is not good. But at the same time, we would have to also do it in the process itself. But by using a Neon function, we will be able to take all of this logic, this whole workflow, put it besides our application into a Neon function.
It will be able to run longer, right? there's a 15 minute limit and at the same time this is also or it will be way more scalable because our main process will have less load it will be cheaper for us and our application will still be able to do other things than just this process this workflow itself and that's also what you will see if you go to the Neon website so let's go to Neon again the link is down below in the YouTube description I will zoom in a bit let's click on product I will click on functions here you will now again see run backend logic where your data lives and keep it running when the job takes time.
And in this case our job will take time. And also by the way all of these functions are also branchable. So since we have a development branch we will also create development functions not production functions.
That's super important. And if you now scroll to the bottom you will find this FAQ. Again what are neon functions?
Neon functions are serverless functions you deploy onto a neon branch. So your backend code runs in the same way. region as your database.
That's another benefit of using a backend as a service like Neon. Instead of having different providers, different services, and with that different regions, you do everything in one specific, let's call it box. Neon will deploy all of your code into one region, Frankfurt, US East 1, US East 2, whatever.
It will use everything or everything will run in the same branch. So this also means that everything will be faster, more secure, and more scalable, which is a huge benefit.
Now this is the question I wanted to look at. How are Neon functions different from Lambda -style serverless? Lambda -style serverless, this is Vercel, or that's what Vercel uses in the background, Netlify, Cloudflare, etc.
They all use Lambda -style serverless, and Lambda is a term created by AWS, Amazon. So functions run next to your data and stay open for long -running work. A function can start respawning within 15 minutes.
and keep streaming while data flows. So agents, web sockets and SSE connections aren't cut off by a short execution limit. They are still serverless.
Idle functions can be evicted. And that's quite cool. Since we now know what Neon functions are and how they work again, please check out the FAQ.
It's super powerful. You can also look at the documentation, which is even more powerful. We can now again head over back to Cursor.
Then what else do we have here? Answering a visitor next. next .js in after.
This is quite interesting. After is the next .js function which got introduced quite recently. So what do we want to do here?
Send message saves the message and returns. The message is classified through the gateway as the following right here. For support questions the question is embedded with OpenAI and PGVector.
Returns the closest chunks within the workspace. Makes sense. If no chunk clears the similarity threshold the conversation off with the following, you could say message or label, otherwise the answer is streamed through the conversation's room in batches every 50 to 100 milliseconds.
If the model can't answer from the chunks it calls a cannot answer tool instead of writing text. This sounds very good to me. Corrections to our docs.
The AI Gateway now offers embeddings. That's super good. So when I started recording this video, the AI Gateway provided by Neon did not offer any embedding models, which of course is a little bit annoying.
But this has now changed. This means we don't even have to use some sort of third -party provider like OpenAI. We can use the AI Gateway specifically or in itself.
Let's again go back. have here. The AI Gateway now offers embeddings but only 1024 dimension open models so we stay on OpenAI.
I disagree, but let's discuss that. Let me quickly head over to the AI Gateway. So let's go to the Neon Console.
Console is just a different word for dashboard. It does not really matter what you call this right here. Let's go to the AI Gateway and inside of here you can search for embedding.
And this is the embedding model, QAN 3 embedding 0 .6 billion parameters. I mean, it looks good to me. I feel like we can use this model right here, but that's something we will discuss with the agent.
So let me write it down. Embedding model or something like that. When I will say it right here.
What else do we have here? Risks. Prisma 8 inside a bundled neon function is unconfirmed.
The first thing to build is a tiny function that reads one row within plain PG and a few SQL statements as the fallback. Hmm, interesting. So what the agent is proposing.
is that we take a neon function, right here, this is our neon function, and that we also embed Should we call it embed? What did it say?
Function at read. So that we bundle it, correct. That we bundle Prisma inside of there.
I'm not sure if that will work because Prisma is quite big in terms of size. The dependency size is quite big. So if that won't work, we will just use plain PG and a few SQL statements.
That's fine. Zero issues. The gateway builds from prepaid credits.
That's fine. And I need an open AI API key. That's something we want to discuss.
Round one decisions. of this feature, I recommend knowledge base, agent, suggested questions, agent off when empty, and auto close. Yeah, I agree.
Active image attachments, image answers, and account settings. Yeah, I also agree with that because it does not really have to do with what we are trying to do. Models, I recommend GPT -5 Nano as the classifier and GPT -5 Mini for answers and GPT -5 Nano for suggested questions.
So the agent recommends to use three different models, this isn't a problem, that's a big benefit of using the AI Gateway, but I'm not quite sure, or I'm not sure if I would agree on the model choice, because what we need primarily is a fast model. So for example, for our classifier, we need a fast model, not a slow model.
And GPT -5 Nano is cheap, yes, but it's not super fast. And now you might also say, Jan, stop, stop, stop, there's something called Jeff, oh my God, Jeff is so cool. Truly every YouTuber created a video on Jeff.
Use Jeff. Blah, blah, blah. It's AGI.
No, stop. Jeff is cool. Don't get me wrong.
It's cheap. It's fast. But it's not really needed.
What do I mean by that? First of all, Neon currently does not offer Jeff because Jeff is still very early. We are still in like an early phase.
And secondly, it's not needed. Like Jeff makes sense on a high level. It's a cool system.
It's a cool idea, but it does not reason it. Look, in my opinion, it's not the best candidate for our specific use case. Jeff is good.
Don't get me wrong. Again, I don't want you to say that the product is bad, but it does not work everywhere. And I feel like in our case, Yes, we could use it, yes, it would work, but I don't really want to do that.
I would rather like to use a cheap model, which is super fast, which can also reason, because that's something that Jeff does not do, and then based on that, we will then be able to classify the message itself. How the dashboard sees source status? I recommend checking every few seconds while source is processing.
Sure, why not? What does it suggest doing? A function deployed in Neon can't reach your local worker, so pushing status over real -time would only once the worker is deployed.
Auto -close and real -time, I recommend that the function publish the closed notice only when a real -time URL is configured. The alternative is deploying the worker to Cloudflare as part of this feature. Sure, I think we should do that.
Let's right away also deploy our worker, so in other words, our real -time setup. And finally, docx, I recommend deferring it and supporting PDF, Markdown, and TXT only. That's because docx is a relatively complicated, not architecture, it's a...
Complicated type in terms of the document. It's not really easy to gather information from a docx file So I agree we will use or we will right now offer PDF markdown and txt So let me create the following prompt look most of the things you suggest sound good to me. Nevertheless I don't agree with everything first of all for the embedding model you want to use open AI Why is that or an open AI model because the neon AI gateway already offers a embedding model called QAN3 Embedding 0 .6.
I know it's probably a smaller or a worse model than the one provided by OpenAI, nevertheless I would lean towards QAN3. Please let me know what that would mean for us, would it create a worse result.
Yeah, give me the pros and cons. Secondly, let's talk about the model choice. You wanted to use three models or to be specific two models, but for three different use cases, GPT -5 Mini and also GPT -5 Nano.
This itself sounds good to me. The only problem is I feel like you're currently looking towards cost. Cost is definitely a very valid argument and also something we have to look at, but the primary factor I'm looking at is speed.
I need a model which is super Super fast. Also, I would honestly say let's maybe even forget cost.
I want to use a good model So for the classifier, we need a cheap and fast model for the responding model We I would probably use something high quality like GPT 5 .6. Terra, maybe Gemini 3 .1 Pro Something because I'm fine with the cost to be super honest with you. And so let's think about that Then in terms of our real -time setup, let's deploy our worker that will make our lives a bit easy We can also defer docx, that's fine.
Zero issues with that. What else did you have right here? Scope, this feature.
Yeah, I would do the same thing. Let's defer our image attachments, image answers, and also account settings. That's something that does not have to do with what we are trying to do right now.
Is there anything else you would like to know? Is there anything else I missed? Let's see, what do we have here?
QWEN 3 embeddings instead of OpenAI. You are right, you pushed back. The reason why I leaned to OpenAI.
was that the PRD and TextTags said the gateway has no embeddings, and that was true when starting the or when creating the PRD, but that's no longer true. Here's how QWEN 3 compares with OpenAI's TextEmbedding 3 small. Pros, one vendor, one credential.
Let me again zoom in a bit. Quality is comparable or better. Huh?
Why should we then use an OpenAI model? Strong multilingual support, same price, smaller vectors. What are the cons?
Independent tests are more mixed. Neon AI SDK provider doesn't support embeddings yet. It's embed throws, so we'd call the gateways OpenAI compatible v1 embeddings endpoint.
Okay, do we need to maybe update the SDK? What else do we have here? One rate limit for everything.
Embeddings share the gateways. Yeah, okay, that's all fine. Verdict, I would go with Quen.
Model twice speed is the real constraint from the current catalog on speed. Pro and top tier models think before they answer. That's all correct.
GPT 5 .6 Tera is high quality, can be fast if we turn its reasoning effort down. Gemini 3 .6 Flash is the strongest Flash model. Classifier candidate.
Gemini 3 .1 flashlight, GPT -4 nano or GPT -5 nano with small or with minimal reasoning. No entropic models are in Neon's official catalog. Gemini through Neon's provider can't do structured output or multi -step calls.
We can work around that classifier and can't answer signal. Let's think about it. Let's think about it.
So I don't want to really use a Gemini model. Let's maybe then stick with GPT models. I would probably then say that we should use something like GPT -5 -4 -nano or maybe GPT -5 -nano with minimal reasoning for the classifier and then for the responding agent we will probably use.
GBT 5 .6 Tera, that's probably what I would do. The can't answer signal tool may not work on Gemini. The fallback is to hold back the first few characters.
I don't want to do that. What else do we have here? Deploying the worker to Cloudflare.
This changes local development. Everything should use the deployed worker, not the local one. Otherwise, the function publishing to the deployed worker and the browser connect to da -da -da.
Okay, that's all good. Deploying needs your Cloudflare login. That's what we will do in a second.
Round two decisions. So let me answer. everything right here.
Speed test now, I recommend yes before writing the prompt for session two. No, we don't need a speed test right now. So since you said that Gemini models have issues with tool calls, I would say that we probably should not use the models.
Let's stick to GPT models. Maybe also quickly do a web search to figure out what GPT model would be the... fastest one for us, specifically for the classifier.
For the one responding to the messages or to the questions, I think we should use GPT 5 .6. I would probably go with that model. It seems smart enough to me.
What else do we have here? One worker or two. And by the way, also give me again your recommendations on that in regards to the models.
One worker or two. I recommend deploying a development worker now. For example, Marshall Desk real -time dev and adding the production one at launch.
I agree 100%. Classifier fallback. I recommend that the reply the classifier gets wrong counts as a support question.
never as of topic okay agreed doc updates i recommend session 2 also updates the tech stack and prd agreed who deploys the worker i recommend session 2 also no i disagree i think we should do it right here so is there anything else yeah let's again i guess think about everything and i want you also decide on the model so please do the web search figure out what model is the fastest one and are you 100 sure that Gemini models aren't able to do tool calls?
Is that a Neon -specific related issue? Because I wouldn't say that Gemini models aren't able to do tool calls. So here's where things landed.
GPT -5 .6 Luna as the classifier. Aha, interesting. GPT -5 .6 Terra for answers.
And you were right about Gemini. Before I can deploy the worker, Wrangler needs you to log in. Gemini and tool calls, you are right.
Gemini itself supports tool calls. The limit is in Neon's AI SDK provider. Its readme says Gemini requests go through the gateway's chat completion route or completions route, and on that route, structured output and multi -step tools are not supported.
That's not great, honestly. But okay, models, these speed figures are from the artificial analysis benchmark. What's the fastest model here?
GPT -56 Terra, 5 .6 Terra, da -da -da. Okay, so Luna is actually one of the fastest models, but GPT -5 .4 Nano is even faster. Do we need like a high intelligence score for a classifier?
Not really. Reasoning efforts matter far more than which model you pick. Every one of these models becomes slow once it thinks the old GPT -5 Nano takes 42 seconds.
Classifier GPT -5 .6 Luna with reasoning effort none. The fastest option is GPT -5 .4 Nano, but Luna is only 1 .5 seconds slower and noticeably smarter. Okay, why not, let's use that model.
Answers, GPT -5 .6 Terra with reasoning effort low, falling back to none. Here's the time to first word budget for a visitor. classify about 0 .8 seconds, embeddings and search about 0 .3 and then tera at low about 1 .8 seconds.
Suggested questions, GPT 5 .6 tera at medium, that's good, embeddings, CRAN 3 embedding, the model and then no open AI key needed. One worker or two, two a development worker now and then a production one at launch, each with its own secrets. I will deploy the dev worker under that name with a flag.
That sounds great to me. Next step for you, Wrangler's login has expired and it needs a browser, so I can't do it from here. Here's what I want you to first of all do.
Please head over to the Cloudflare website. By the way, this is completely free. You don't need any credit card or anything like that.
And right now, I want you to sign up. If you have an account, then log in and you will already be finished. The next step is to go back and we want to use this command right here.
pnpm filter martialdesk real -time wrangler login. Copy the command, open my terminal and I will paste it inside of here. What we need to do is connect our local CLI to our Cloudflare account.
Cloudflare now opened for me right here. I will continue with Google. What do we have here?
Wrangler wants to access your account. That sounds good to me. What I will do is click on authorize and with that we will connect our Wrangler CLI to Cloudflare.
If you didn't know or if you are unfamiliar with Cloudflare, Wrangler is... is just the name of the CLI. This is how we can interact with Cloudflare from the CLI.
What I would also recommend to do is open your sidebar. Let's click on customize and inside of here you can search for Cloudflare. So they have a nice little plugin which exists in the Cursor marketplace.
The same thing also exists in Cloud Code and also Codex and what you can do right here is click on add and with that you will install Cloudflare specific skills but also you will get access to the Cloudflare MCP server. In other words, multiple MCP servers like Cloudflare Docs, Cloudflare Bindings, etc.
So let's head over back to our session. What do we have here? Still open access to GPT models on the gateway.
OpenAI models are gated foundation models on Neon, so your account may not be able to call GPT 5 .6 Terra or GPT 5 .6 Luna yet, and prepaid credits are required. This should all work, so we will test it in a second. starting tuning values five chunks per search, agent replies stream into inbox two.
Let me do the following or I will say the following. Could you now first of all deploy our real -time server? My CLI is now authenticated.
And could you also make sure that we have access to the needed models using the Neon Gateway? And while the agent is still working on everything, I want to quickly show you again the benefit of using Neon as your backend and more specifically the Neon plugin. Because did we have to open the documentation once in this video?
No. Did we ever have to open our dashboard, pull credentials, open our .env file, paste them inside of there.
No. Why is that? Well, that's because the MCP server allows to do or allows our agent to do everything for us.
So using the MCP server, our agent now first of all checked our existing service credentials in our development branch. Then it created an AI Gateway credential on the Marshall Desk development branch. Again, I did not do anything myself.
Normally, in the past, in the long distant past, when I was still a little bit younger, I had to literally still open dashboards, create secrets, open my ENV file, paste them inside of there, start the dev server, check everything myself. Now my agent is able to do everything for me.
I don't have to open any dashboards. I can if I want to, but I don't have to. I can use the CLI, I can use the MCP server, I can use the skills, and that's already it.
I can now literally, in theory, take my phone, open TikTok, though I would highly recommend to uninstall that stupid app, and scroll for hours on end, and still I would kind of be productive, because my agent agent would do everything for me. That's the power of the plugin, of the Neon plugin, but also more specifically Neon as a platform.
You can do everything using one provider, which is huge. Alright, we are finished. Our real -time worker is deployed, so our real -time server with that is deployed, and our account can call every model we picked through the Neon AI Gateway.
Real -time worker deployed, we have our secrets, checks against the deployed worker, publish with the right Secret got past the off check, a visitor socket opened on its own conversation over WSS, I sent a real message through the local app and next published it to the deployed worker with no errors. AI Gateway Access, I created a gateway credential on the development branch and then also saved the key and this key right here to the .env file.
Then it also did a few tests, GPT 5 .6 Tera, low reasoning, 2 .2 seconds, then here we have no reason, 1 .4 seconds, GPT 5 .4 nano, 750 milliseconds. Interesting.
GPT 5 .6 Luna, 966, and QAN 3 embeddings, 414. What else do we have right here? Classifier recommendation.
I'd now switch the classifier to GPT 5 .4 Nano with no reasoning. Through the gateway, it's about 200 milliseconds faster. Correct, I would do the same thing.
Two test conversations, please ignore, are now in the Marshall Desk inbox. On development, I couldn't care less. For session two, add a dev environment to rank.
So I will say the following. Look, this all sounds great to me. I fully agree with your recommendations.
I want you to now create a prompt for session two. So in other words, our game plan. What models should we use for the implementor?
I would again use Opus 5 .5 for the researcher. You know what?
I would probably now already use Sonnet 5 .5 because it's a super cool model, which I like. Then for our researcher, what should we do? Should we use GROK 4 .6 or maybe Sonnet 5 .5?
Let's go quickly to cursor bench. I'm not sure what's cheaper, to be honest with you and with that smarter. Can we still filter for GROK 4 .6?
Yep, I can still select it. So let's check it out. What do we have here?
Sonnet 5. We also have... grog 4 .6 grog 4 .6 with high reasoning effort is a cheaper five dollars twenty No, it's more expensive.
Damn. So let's use Sonnet 5 .5. I will go back.
Let's select right here, Sonnet 5 .5. Reviewer, independent read only. Let's again use Opus 5 .5 and I will click on continue.
That's very surprising. So let's do the following. I will go back to Teal Draw and inside of here, let's go back to our diagram.
It's not Teal Draw, it's Eraser. Let me open Eraser and inside of here, I will now do the following change. using ROG 4 .6, I would rather want to use Opus 5 .5.
Opus 5 .5, sorry, not Opus 5 .5, Sonnet 5 .5. Sonnet 5 .5 is cheaper, it's faster, it's smarter. So there's no real reason anymore to use a different model.
Let me make this a bit bigger, let me also copy it, and I will also paste it everywhere inside of here. Don't use ROG 4 .6, forget my recommendation, it's not a model I would use. That's rather, right here, use Sonnet 5 .5.
It seems way smarter. and also way faster. And you don't have to necessarily right here also use Opus 5 .5 Low.
Again, let's rather use our Sonnet 5 .5 model. If you want to use Codex, again, stick to what exists right now, GPT -6 Astra, Luna Max, and then again GPT -6 Astra. This all sounds good to me.
Let's again head over back to Cursor. Our agent is now working on our prompt, so let's wait for everything to finish. We now have a prompt.
What I will do is copy this beautiful long prompt. Let's create a new session. I will again use cursor for that.
And let's quickly go through the prompt. What do we have inside of here? You are the primary implementation agent responsible for delivering Marshall Desk's knowledge base and AI agent, sub -agent, model roster, implementer, opus 5 .5, sonnet 5 .5, opus 5 .5.
That's all good. Then this is our repository, expected starting state, main, okay. That's all good.
Do not stop after producing a plan. What else do we have here? Authorization boundaries.
Neon. Use only development branch. That's correct.
Cloudflare. You may redeploy the existing dev worker with the following command. No Vercel deploys.
That's also correct. We don't want to deploy our application yet. Orchestration.
You are the orchestrator, integration owner, and final verifier. Role routing. Implementor.
Subagents. We already know that. Topology.
We already know that. This all comes from our skill progress file keep a concise living file right here uh -huh dot clod that's interesting but that's fine what else do we have here build with your eyes open use the browser we already know that that's all of the technology we want to use infrastructure already in place ai gateway enabled in frankfurt these are all of the models pg vector dev worker let me do the following since all of this looks good i will scroll to the top i will again first of all use the feature orchestrator skill and let me also say session 2.
And before I now click on enter, I just remembered that there's one thing we haven't discussed yet. Let me go back to the previous session and I want to say the following. You mentioned somewhere that you want to bundle Prisma inside of a Neon function, correct?
Did you verify if that works? I've now verified that Prisma bundles and runs provided the bundle includes a small fix and Neon's own deploy tool adds that fix automatically. What I still haven't done is run it on Neon's actual runtime, what I tested locally, without touching the repo, bundling, without the fix, with the fix.
What's even the fix or what's the issue? Running it fails with dynamic require of events is not supported. PG is a common JS package that calls require, the fix adding the standard create require banner to the bundle made it work.
Neon's deploy docs describe exactly this failure, naming PG and the same error message. And say their bundler applies that banner automatically. Still unverified running it on Neon's runtime with the database URL Neon injects.
Local node 26 is a closed stand for it. Either way I'd replace the Prisma bullet in the prompts Neon function section with this. Prisma 8 inside a function was spiked in session 1.
No we don't have to do that. The thing is it works. That's good.
So what I will do is go back and again click on enter. With that we now verified that Prisma will work in the function. At least.
it works right locally we still haven't deployed it we won't know if it really works in production but if it won't work we will then just fall back to our normal local pg or not local pg but standard pg we will make a direct connection and then use sql functions another thing i would recommend is to again go to your customized section to the marketplace and i would install the prisma plugin this is another thing which is quite helpful because you get first of all two mcp servers and also a lot of skill Let's again go back to our session.
Is everything running inside of here? Yeah, everything started. So let's wait for everything to finish.
This will probably take about an hour or so because it's a big feature, but that's absolutely fine. Our agent is finished. It created all of the needed code.
It verified everything. So I guess let's go through everything. And by the way, how long do you think all of this took right here?
I'm waiting? Correct. Two and a half hours.
Damn, that's quite a bit. So let's see, what do we have here? What works now?
Knowledge base. Owners upload PDF, Markdown or TXT files up to 50 ,000 characters. Owners can preview a source's chunks.
Then if a new version fails, the old one keeps answering. Uploads that never land or runs that hang are cleaned up by a 15 -minute sweep. Interesting, that's quite cool, because this is something I never mentioned in the PRD.
Agent, every visitor message is classified first, off -topic, small talk gets a brief reply, asking for a person hands -off with this following, I guess, tool call or label, and then support questions are answered only from the knowledge base, nothing else. Suggested questions up to 4 regenerated after every knowledge base change, this is what we wanted, or to close.
conversations idle for 24 hours close with the closed notice and the widget and inbox update live agent off the agent is off when it's switched off or the knowledge base has nothing ready number two decisions I made and where I departed from the prompt so here our agent made decisions for us this is normally something that you Not really or not always want, but if it makes sense, it's good.
So let's see. Prisma in the Neon function works directly. Deployed function imports MarshallDeskDB.
That's good to know. So no plain PG fallback was needed. Classification output object with the four labels on GPT 5 .4 nano.
Effort none. It measured 858 milliseconds median. Can't answer.
It cannot answer tool. Nothing is shown until 40 characters have. arrived that's good I dropped answers from low to none because first words exceed three seconds that's quite a bit it made no measurable difference then retrieval tuning top five chunks the last 10 messages go along as history deliberate deviations a finished reply is saved before agent .done is sent which avoids a flicker each agent or chunk now carries the whole reply what else do we have here replacing a file no longer resets it to up loaded, add presign, handoff notices are fixed, English UI text, not model written.
Then neon functions setup, one function drops declared in neon .ts. Let's quickly check out the file just to see how it works. What do we have inside of here?
Define config, neon config, require env, default, define config, drops, martial desk drops. Then we have our source, env, functions secret, real -time URL. This is needed for our real -time server.
For party kit, death, we have a port, buckets, triggers. That's also quite fancy. What else do we have here?
Our agents. We have three roles. Researcher, implementer, reviewer.
We already know that. We had four researcher agents, five implementer agents in parallel, and then also five agents in cycle one and two in cycle two for our reviewer. What else do we have here?
Proposed pull requests. Two in total. That's good.
Why not three? It's not needed. Checks and results.
So this here is already a web browser screenshot. Our agent verified everything. That's good to see.
I sadly can't really zoom in, but that's fine. And you know what? Let's test it out right here.
What did we have here? Cycle 1 fixed. Critical, which I reproduced.
The visitor could type fake team member lines. And the agent then promised a 90 -day refund. Wait, what?
This is a huge mistake. This should never happen. History messages are now separate blocks with the author outside the text.
So the same attack now hands off instead. Cycle 2, no critical or high findings. I fixed the six small items.
You know what? Let's now test everything out. I already started the dev server.
So this here is localhost 3000. Let me zoom out a little bit. I will head over to the dashboard.
What do we have inside of here? Nothing really changed. We still have our agent toggle.
I can view the agent name. We have our greeting. I don't like this emoji.
Let's quickly change it to something normal. What else do we have here? Suggested questions hidden while the agent is off.
The agent is not off, but the thing is we don't have any knowledge data, no sources. We have our domain. What is if I click on add text?
Uh -huh. We get this nice little dialogue. What we can now do is upload some knowledge.
I have some stored in my downloads folder. Can I upload multiple things in parallel.
Let's see. Oh my god, this works. Processing ready.
This was super fast. So what we can do is head over back to Neon. Let's do that.
Because what we now did is first of all use two models provided by the AI Gateway. Our embeddings model, so QAN3. And then also the model we used or the model we use to generate the questions.
What we can first of all do is check out if we have now a new function. I will go to my functions. And yes, we have jobs.
This is our invocation URL. few triggers, auto close, source uploaded and what we can do inside of here is view the logs. Let's see what do we have inside of here.
We had some load, a few invocations and we also had errors. That's not good. What's the error?
Let me expand this. Warning, security warning, the SSL modes, prefer, require, and so on, are treated as aliases for verify full. In the next major version, this will change.
Let me do the following. I will copy this, go back to cursor, and also paste it down below to not forget anything or to not, yeah, to not forget anything. What we can do inside of here is also check out our storage bucket in terms of monitoring.
So we had some load. 15 logs. Let me quickly go to our object storage primitive.
I will go to the uploads. Inside of here we should now have our workspaces folder. Correct.
And then we have the workspace IDs or we have two workspace IDs. Let me open this one. Is that the correct one?
And yeah it looks correct to me. And what I can do is also then view the ID. I could also for example download the source, the empty file and then I can view everything in more detail.
What else can we do? do right now, let's again head over back to monitoring. Let's see what we have inside of here.
So this is our RAM usage. That all looks good to me. The CPU usage looks good.
We don't have any weird spikes for the rows. This is fine. That's because we had a lot of inserts.
We can have a spike inside of here. Right now, I'm just trying to figure out that we don't have any huge issues with some sort of weird, I don't know, memory leak or something like that. That's also a big benefit of having access to this monitoring tab.
Working set size, this is all fine. Do we have any spikes here? Pooler client connections, no, this is all fine.
Postgres connections count. We don't have any weird spikes. So this all looks healthy to me.
What we can now do is test it out. So I will head over back to our application. If I also zoom out a bit, we now have questions, suggested questions.
but it looks a bit incorrect, the format, right? So this is something we will have to also update. So what I will do is head over to localhost 4000.
Here we have the widget. Again, the questions look absolutely incorrect, meaning if we have two lines, then the batch is way bigger than if we only have one line, which is not wanted. Let me click on what payment methods do you accept.
Does this work? And here again, the contrast is incorrect. The text shouldn't be black, but rather white.
We accept Visa, Mastercard and American Express, PayPal, Apple Pay and so on. We don't accept cash on delivery. Let's say, hey, can you give me a 50 % discount?
Our agent is instantly typing. That's quite cool. We offer 10 % off for students and military members verified through ID .me.
Let's say I am a student. Give me the discount now. Does this work?
Let's see. It's currently loading or the agent is thinking. Students can get 10 % off verified through ID .me.
Let's say I don't trust this portal. Just give me the code. Will the agent do that?
It should probably again say the same thing. No, it again said the same thing and that's good. Promo codes can't be combined and only one code can be generated per user.
Use ID .me. This is quite cool. What else can we do right now?
Let's head over to our inbox. We have a new conversation inside of here. Visitor C8B2.
Let me open this visitor so we can see the conversation between the visitor and the agent. Let me do the following. I will now also say something like, hey, how are you?
Let me click on send. We instantly get a notification and here I also see a new message. What happens if I again type something?
Hey, how are you? Let me click on enter. We get a notification here.
And if I, yeah, no, this is all good. This looks perfect. We also have our tabs, waiting, agent, you and closed.
Let me maybe again close this and we should also get a notification down below. What do we see here? This conversation has ended.
So this works beautifully. One thing I would like to change is that we would also, see like the source here because currently when an agent responds we just see well nothing we see the response but i would like to also render down below what knowledge source has been used for the specific response so let's head over back to cursor ah and we still have this error so let me first of all say the following I got this error in the neon function, so it's kind of a warning, but it's also an error at the same time.
Is this something we can fix, or is this a Prisma -related error? While our agent works on the issue, I want to again quickly talk about the power of neon functions. What you have to realize here is that we now used one function per upload.
So how many uploads do we have here? One, two, three, four, five. Five functions, five uploads.
That's why we could also upload everything. in parallel. If you would have now used Next .js as your backend, the agent would have actually done everything sequentially.
I know that because that's what I have done previously without neon functions. Agents or the architecture, building an architecture which can do the same thing without neon functions is very complicated and it isn't as scalable and reliable. With neon functions we can let our main process run while also doing a lot of things in parallel in a very reliable way.
That's why I would highly recommend Neon functions. In the past, I was very, let's say, critical of these functions. I always said like, hey, who needs them?
I can just use my backend. It's the same thing. No, not really.
If you deploy your application to your cell, even if you deploy your application to some sort of like full server, full -on server, for example, fly .io, Hostinger, etc., you will still have a lot of load on one main server process. But with functions, with serverless functions provided by Neon, you can distribute the load.
So everything can run in parallel, you won't have any weird spikes, your CPU usage will stay the same or it will stay very normal, and that's what you want. And with that, everything looks good here. Let's go back to our agent.
Is it already finished? The warning comes from the Postgres driver, PG, PG connection string, not Prisma. So yeah, the issue was that we did not have this query parameter.
SSL mode require. Our agent is now redeploying everything, the function. So since the agent is finished and created the necessary fix, we can now work on the UI and make it better.
The first thing I don't like is how the questions render right here, right? Because if we have two lines, then the badge looks... absolutely incorrect.
It's not proportional. Therefore, let me use the screenshot tool. I will quickly screenshot this right here.
Let's screenshot it. Bump. Capture.
I will attach it into this chat and let me say the following. Look, everything works very well. We can upload all of the sources.
That's what we wanted. Nevertheless, once we have our suggested questions, they render a bit weirdly in the widget itself. What do I mean by that?
Well, if the question is just one line, then everything is proportional but once we have a question which needs two lines it comes unproportional the batch looks incorrect and that's something we need to fix another thing that i don't like is that the batch currently is completely rounded that's something i don't want let's use our normal card rounding which is I guess, Excel or something like that.
Maybe we should also have some sort of contrast or at least some sort of like accent color in the badge, maybe the border, something like that, so that we can also integrate that. Another thing I don't like is the contrast itself. When using this green accent color, the text is black.
So for example it says trans agent 123 and that's black. The thing is the contrast ratio is not good, at least in my opinion. We have to either make the text light or white or we have to darken or lighten the background color.
One of these two things. Then what else don't I like here? I guess the rest looks good to me.
Ah, the inbox, right. That's another issue I had. So let's also add this to the prompt.
Another issue I have is in the inbox itself. Currently everything works, right? The agent can answer all of the needed questions.
But what I don't really like is that we don't see the sources. So when an agent answers a question, I want to also see the source which has been used in the inbox itself. Another thing I don't like is again the contrast ratio.
Currently when an agent answers, the background color is black. That's something I don't like. And another thing I want to change is...
No, actually that's it. The contrast here is a bit, or is just a bit incorrect. What I would probably do is give the agent when it replies a specific color, a specific like...
either background color or maybe border color. You know what? Border colors may be interesting.
We could give the agent in the inbox, not in the widget itself, some sort of like color, maybe border color, so that we instantly know as the business owner, hey, this is the agent and this is what I responded with myself, so the business owner. We need to differentiate the two. And I think that's it for now.
What I will do is click on enter and let's wait for the agent to finish with everything. What do we have here? Here's the plan covering all five requests.
So Suggested question ships, card rounding, consistent wrapping, and an accent colored border. Had a contrast green, orange, and pink currently dropped to black text, I will darken those three accent shades so white text always passes contrast.
And finally, inbox agent replies get an accent border so that they are clearly not yours, and each reply lists the knowledge -based sources it used. That's exactly what I wanted. So all of your points are fixed, and I mean, look at our widget it's already 10 times better it looks 10 times better so what can we now do well i guess check out the result right so let's head over back to our browser this is our widget as mentioned it looks 10 times better i have zero complaints with it let's head over to our inbox to our closed tab did the agent backfill anything nope that's a bit annoying but it's fine let's test it out manually i will do a hard refresh let's open our tab right here We can view all of the questions.
We also have a nice animation here, like micro animations. Let's say, how can I track my order? Do we get a pending state?
Yep, so our agent is typing right now. You will get a tracking link by email as soon as your order ships. Let's head over back to the inbox.
I will go to all open and here we have the message. What is if I click on waiting? Okay, nobody's waiting.
Agent, do we see the message here? Yeah, the conversation. it's again empty.
So let me click on visitor 0087 and what do we have here? Answered from shipping and delivery MD, accounts and orders MD. That's beautiful.
What else should we do here? This is exactly what we wanted. Like there's literally nothing else we could now add inside of here.
Maybe we could make the UI a bit better in certain cases let me maybe write a message hey how are you how does that look let me click on command enter this all looks beautiful you took over what do we have inside of here hey how are you Like this is perfect.
I'm not sure if I would like to change anything inside of here. I love it. Let's again go back to the dashboard.
What can we change inside of here? I'm still not a big fan of this button. It looks weird or the placement is weird.
What else could we change inside of here? Let me maybe zoom out a little bit. So to a normal width 110.
um this all looks good to me can i view the md file let's see yep this works and we can even see all of the hoses called like chunks chunk one chunk two chunk three that's also exactly what we wanted what about the domain ah right we wanted to like enable localhost by default so what we can do is head over back to cursor and i will say the following look everything is now great i love the ui though there are a few things i might want to change depending on what you think.
First of all in our allowed domains we currently don't have any domains enabled by default but I would like to probably enable localhost 3000 and 4000 or let's just say in general localhost by default. Now this is something I want your opinion on. Is this something that's safe to do?
Is that something we should do or would you maybe say no hey don't do that. I would love to know what you think. Secondly with our knowledge base we can view all of the uploads that's all great the only thing i would like to do is that if we click on the text for example shipping and delivery .md that we also open the same like viewer right the same dialogue with all of the chunks another thing that i don't like is that we don't have any how is it called i forgot the name where you like hover over a button and then on the top it gives you the text what it means because right now this i button when i hover over it i don't see what it does it could be something destructive right who knows so that's another like user experience like a little user experience detail that should be added what else would i do inside of here the agent toggle the button on the top right to toggle the agent on and off is still misplaced i hate the positioning it does not work on the top right we need to find something else either we remove the button completely or we have to reposition it give me your thoughts on that where you would put it or if you
would maybe remove it completely. Another thing I would do is that currently in the inbox, we can view all of the sources, right? So when an agent answers a question, we can see what resources have been used, what knowledge has been used.
What I don't like is if I click on the text itself, I can't view the document. So that's maybe something we should do so that we can just open the whole, you could say document all of the chunks in the dialogue, just as we also do it in our norm. in the settings dashboard.
Is there anything else I don't like? The rest all looks good to me, pretty much. I don't really have any complaints.
Is there anything else you would maybe redesign? Is there maybe something you would change in terms of UI, UX, the workflow? Maybe there are certain optimizations that we can make.
Let me know. Oh, and by the way, there's one more thing I would change. In the knowledge base, once everything is uploaded, on the left, we have like this little icon with this document thing, or we have this document.
icon with this rounded border, background, whatever. What is if we would render like an actual icon, a PDF icon for PDF documents, a markdown icon for markdown documents and so on. What do you think?
Ah, that's how it's called. A tooltip. My god, what did I say?
A little badge on the top? No, I meant a tooltip. Because again, as mentioned, these buttons don't really explain what they do or there is no description.
So if we go to shared CNUI, there's this tooltip component and that's what I want to use. If you hover over a button it instantly shows like this tooltip on the top and with that we can explain what the button does without any text.
Let's again go back to the agent. What else does it have for us right here? I will find every icon only.
The paint actions already have tooltips. Let's maybe also quickly talk about like animations maybe animations could also make sense there is one library called lucid animated if i'm not mistaken it provides animated lucid react icons and lucid is also the icon library we use right now and if you hover over an icon there's like this beautiful animation It's beautiful.
So that's maybe something we could also implement, but I want to ask my agent in a second. For now, I will just save this URL right here, and let's wait for the agent to finish with all of the work that it's already doing right now. My agent currently has a few questions for me, so I want to allow localhost automatically, only in development, nothing stored.
I will select that. Then where should the agent switch go? And thinking about it, like in the PRD, I said, I want to have it 100%.
But now thinking about it, do we really need an off switch? I don't think so. If you want to be a customer of Marshall Desk, then you have to be a member of the AI revolution.
No, I'm joking, but let's remove it. It's not really needed. What else do we have here?
Which of my other suggestions should I build next? Why did the agent... Do that.
Classification scores reason in the inbox. Give the knowledge base its own sidebar page.
Should we do that? I don't know. Is that needed?
I don't think so. Waiting count next to the inbox in the sidebar. Nope, I don't think so.
So why did the agent do that? Sounds quite good to me. Though this try or give the knowledge base its own sidebar page could be interesting.
You know what? Should we do that? Why not?
So let me click on continue. I mean, if I don't like it, I can just remove it. Let's expand our application.
Let's make it better. Our agent is now finished. It made a few changes and you will see that also instantly in the sidebar because we have a new link, knowledge base.
Nevertheless, there are still issues. Like for example, we have this weird agent on button, which is absolutely not needed. That's the first thing.
And let's now check out the other changes. So I will go to the knowledge base. What do we have here?
We can first of all view now the icons right here. So MD, TXT, Plans and Pricing. This is a manual text document I created.
This means that's why I also have this edit button and we also have tooltips. If I click on edit I can then text or sorry update the text if needed. Let's also head over to our inbox.
I will again zoom out a little bit. Here we have visitor dbd1 who's waiting we can see the message can i talk to a real person and we got also handed over or the message got handed over to us the visitor asked for a person let's maybe test it out one more time i will first of all copy the script one more time let's go back to localhost 4000 i will paste it inside of here let's embed and reload and i will test the same thing meaning let me copy the message i will again head over not to neon but to our application in inbox again to the same visitor.
Let me copy the message. I will go back, paste it inside of here, and let's test it out. Can I talk to a real person?
Our classifier will now run. It classified the message and said a person will join shortly. Replies will apply here.
And if I go back to the application, I can also see it here. What else can we test out? So do you integrate with Salesforce?
This is another test I had or made, and it also handed over to me because the knowledge base doesn't cover this. Let me also paste it inside of here or we have to first of all for that mark the conversation as resolved.
Let me say close. Boom, this is finished and now we can again go back. I can say do you integrate with Salesforce?
What will happen here? It's loading. We couldn't find the answer to that so we are bringing in a person.
They will reply here shortly. This also looks beautiful and if I go back we see the new message here. Is there anything else that we should test?
I mean, everything works, right? So let's maybe work on UI upgrades, or in other words, updates. One thing I don't like is, as mentioned, this agent on indicator.
It's not needed. Another thing I wanted to discuss is Lucid Animated. Let me copy the URL and let's go back to Cursor.
I will say the following. Look, the changes look great. I almost have zero complaints.
One thing I don't like is the agent on indicator. It's not needed. Throw this.
cart, whatever you want to call it, completely out. It's not needed. Secondly, let's talk about optimizations in terms of the user interface.
There's one library called Lucid Animated, which provides beautifully crafted animated icons. What do you think? Should we use the library or is it not needed?
I guess having animated icons is quite fancy. And since we are already talking about animations, should we maybe add more general animations to our dashboards? specifically, not the homepage, the dashboard.
For that, I want you to look at the Emil Design Engineering skill and let me know what you think. Let me again invoke the Emil Design Engineering skill. And the Emil Design Engineering skill is interesting because it teaches the agent, of course, about taste, great UI, UX, and stuff like that.
But a huge focus has also been set on animations and having animations that are clean and not overpowering, because that's very important. see often is that people will use AI agents and literally animate every little thing and everything will then go from the left to the right to the top to the bottom which does not make any sense.
You want to have very clean animations and sometimes even no animations are better than too many animations. So that's why this skill is a huge game changer and that's also why I use this skill so much. Let me now click on enter and let's see what our agent thinks.
Let's see. Lucid animated. I'd skip it.
Interesting. Wrong frequency. It adds a dependency.
The good idea fits in a few places. Animated icons earn their keep at rare moments, such as an upload finishing or an empty state. More animation in general.
Yes, but as polished rather than flourish. And that's what I meant. Keep it minimal.
Don't overdo it. The dashboard is used constantly. so motion should be crisp, fast and mostly invisible.
All of these use Tailwind and CSS already in the project. None need a new dependency. Should I implement them?
Sure, go ahead. And I guess that's already it because besides that, our application is perfect. Let me again check it out.
Do we have our preview? No, it got removed. Now we will have a few animations and that's it.
The application is totally finished. It completely works. We don't have any weird spikes.
If we again go to Neon, you will see the same thing here. Yes, we had some CPU load.
So if I go to last day, we had CPU load right here. That's because my dev server had a memory leak. Surprise, surprise, next JS.
But now everything works again very well. Our application is performant. We don't have any weird spikes.
That's what I wanted. So adding three small animation tokens to the theme. Okay, let's wait for that.
So we are done. What changed? The button, every button scales to 97 % when pressed.
This is another like style of micro animations, and that's also a great user experience detail, if that makes sense, because the user instantly knows if he or if they clicked on something. Let's quickly check it out. I will go back to Chrome, and if I now click on a button, do you see like this little scale animation?
This is beautiful. Let's again go back. Tooltips, one shared provider for the whole dashboard, live arrivals, new inbox rows, new messages and notices, and new knowledge base rows fade.
in from four pixels below. Status changes, how the agent handles this, theme toggle, reduced motion, left alone on purpose. This all sounds great to me.
So what do we do now? Well, before we now go on and deploy our application, we need to review. the code so that's what i will do right now before i do it let me ask the following did you already create the prs question mark i'm not sure because i don't see them right here i out okay did not commit anything yet we have what 18 000 changes that's fine and we wanted to create i think three pr so branch feature knowledge base feature agent feature ai agent main every change since then covers today blah blah One thing to decide first.
I added that polish only to feature agent. Keep it as is. My recommendation PR1 is the knowledge base foundation and PR2 is the agent plus all UI polish.
Sounds good. So here's what I will now do. I will review all of the code, all of the branch changes.
As already mentioned, there is no trick to reviewing code. There is nothing. You literally look at the code and you try to understand it.
You want to get a general understanding and you want to also make sure that the code is somewhat clean. So this is all fine, we have our constants, we have the neon provider, I don't have any complaints with this code looking at it, but if we would now have weird like functions to handle errors or for example a lot of type annotations, as any, as unknown, stuff like that, then this would be something where I would say hey Dear agent, you made a mistake.
Fix this. Because agents, whenever you start generating a lot of code, 15 ,000, 20 ,000 lines of code, something like that, agents start taking shortcuts, which also means they create rubbish in terms of code. Therefore, it's important that you verify the code, that you make sure that it's clean, and only then create the needed PRs and let BugBot or any other PR review tool review the code one more time.
So I will now get started. This will probably take a... 20 minutes or so to review the code and once I'm happy with it and made all of the necessary changes I will create the needed PRs.
Two in total. I have now finally reviewed all of the code and I found sadly a few issues. I mean is it a sad thing?
No not really it's kind of standard practice and I'm there to fix them or to find them. I won't fix them the agent will fix everything. So what did I say right here?
Look while going through the code I found a few issues. Firstly it seems Seems like you are often creating weird helper functions that are sometimes not even needed.
What do you think? Essentially the agent or we had a lot of functions which did a lot of like basic things. Is object, handle error, is error, is defined error, stuff like that.
Things that don't make sense. So our agent didn't follow certain principles and at the same time the most important one, dry. Don't repeat yourself.
I often saw that the agent duplicated logic. It already for example created one helper function and then it again created a duplicate to do the same thing with a different name. Which of course is not good.
It's not scalable. And at the same time also one function was there to handle errors. And I said also in certain cases the ORPC error handling seems incorrect to me.
What the agent did is it handled errors in a very standard way. Meaning as unknown it not really know what the errors are but this is not the case with ORPC.
ORPC gives you type safe error handling you know what the errors are you can use the ORPC error function is defined error etc so with that we can or we will know what the errors are specifically and the agent didn't do that that's why i said hey please fix it check the current docs and then go ahead and make the necessary changes and the agent of course also figured it out right here why can't i close this again man Cursor, you are buggy today.
So what do we have here? I agree with both points. In each case, the problem is real but smaller than everywhere.
Yes, I didn't say everywhere. I just said in certain places. The ORPC error handling has a couple of actual bugs and most of the clearly unnecessary helpers are tied to the same errors.
So we had like a dependency array, if that makes sense though. What am I saying? So we had two things that were building on top of each other.
We had weird functions that were not needed, and most of them were related to error handling. So because error handling was done the incorrect way, the agent also then created these not needed functions. It's as simple as that.
And even though we had research sub -agents, reviewer sub -agents, they still didn't catch that. And that's why, like, there are people that will tell you, hey, you don't need to review code anymore. Agents are smart enough and stuff like that.
Yes, they are smart enough. They are way smarter than most of us. Nevertheless, they will still create very, you could say, dumb mistakes, to be honest with you, because instead of the agent taking the time and opening the browser and searching for the specific documentation site, it will rather use its training data and then use this training data to create the necessary codes.
And this might work for normal stuff. Again, a general -purpose LM isn't just there for coding -related work. Nevertheless, it does not really benefit us, and that's annoying, even though a web search takes more time.
it still gives us more relevant info. And that's why I said, hey, check the current docs, please. But okay, it is what it is.
So this is now GitHub. We have our two PRs inside of here. Bugbot also found a few issues, but they went super severe.
So for example, stale stream after conversation switch. Definitely an issue that we should fix, but it's not super, super severe. So what I will now do is merge this stack because I fixed everything.
Everything now looks good. Bugbot is happy. as you see here and we have our two PRs which build on top of each other.
Let me click on merge stack and since everything is now merged we can head over back to cursor and inside of here I will say the following please switch into main and pull all of the recent changes. And with that we are now finished you are on main everything is in sync and also your working tree is clean and the dependencies are in stock.
and merge the log file. What's now the next step? Well, it's to deploy our application.
So let's do that. Let's now quickly think about how to deploy our application. Currently, we have Neon, our backend as a service.
We will have to switch the branch into production and use the production branch in production. Wow, who'd have thought? Then we have our real -time server, PartyKit, in other words, party server.
We will have to create a production deployment because currently we have two deployments or we have one deployment which is a development deployment and we also need a production deployment. And thirdly we need to deploy our application and for that we will use Vercel.
Vercel is a serverless provider, agentic, infrastructure, whatever that means. It's not really important. As already mentioned Vercel will use serverless functions in the background, lambdas.
Lambdas are quite cool because they fundamentally have the same idea as Neon. Meaning if nothing is running, you don't pay anything. If your application does not have any load, then your server is not running.
Only once someone, once a visitor wants to use your website, his serverless function spins up and then serves the request. And once the request is served, the serverless function again shuts down. And that's it.
The same thing is also true about Neon. When we have no load, our database is not running. There is zero cost.
Only once the user comes along and wants to sign up and pay us a bunch of money, only then the database spins up, including our application itself, the serverless function for our application. So here's what we will do. We could do it the old -fashioned way, meaning we could now go to our cell, log in, create a new project, add the environment variables, look at the log...
box. But we are in the agentic future. This is boring.
So let me close this website. I don't want to do anything. Let's head over back to cursor.
I will click on customize and I want you to do the following. In the customize section, in other words in the marketplace, search for Vercel. You will find a Vercel plugin and this plugin has a few things.
An MCP server, skills, then also sub -agents, commands and finally also hooks. But in this case what we really need is the MCP. MCP server and the skills and we will let our agent do the deployment part for us.
So the only prerequisite in this case is that you have an account with Vercel. It's by the way completely free like all of the other services. You don't need any credit card.
Just quickly sign up to Vercel and connect the MCP server. The mechanism is the same one as with also Neon. You have to authenticate yourself, create the connection, give the correct permissions and all of the fancy stuff.
Also please again make sure that the Neon plugin is installed because we will also do the same thing here, meaning we will let our agent switch the environment from development into production, also enable all of the necessary stuff in terms of authentication, so the settings, and I think that's already it. So let's go back.
I will create a new session. Let's again use Opus 5 .5. And instead of now instructing my agent on what to do, I will rather ask the agent a question.
What is needed to deploy our application? So, let me say the following. Look, I'm now completely finished with this application.
I created, or in other words, implemented all of the needed features, and I want to finally deploy my application. So, here's the thing. For the app deployment, I want to use Vercel.
I already installed the Vercel plugin. The MCP server is authenticated. You should have access to all of the necessary skills.
So, this is, I guess... step one or primitive number one. Then we have Neon.
The Neon plugin is also installed with an MCP server, skills and all of the stuff is also authenticated. So what we will have to probably do is switch our environment from development into production. We will have to also configure authentication because I think our development settings don't match the production settings.
What else do we have to do? We have party kit, in other words party server. We will have to probably create another deployment for production.
Is there anything else I missed? Can you like give me a small rundown of what we need to do and what you can do and what I have to do? That's also very important.
And look, this is what I call an agentic feature. We can see that the agent is actively using the MCP servers that it has access to. The Neon MCP server.
List the branches of the martial desk. Neon Project. What we have here, we use the or the agent is currently using the Vercel MCP server.
List the Vercel teams on Jan's account, look for an existing Marshall Desk project. This is what I mean. We don't have to use dashboards anymore.
And you know what? Let me give you a bit of advice, because if you are watching this video, then there are probably two things that apply to you. First of all, you want to know how senior engineers code with AI, how to correct or how the right workflow looks like.
Like how to review code, how to think about everything, stuff like that. But the second reason why you are probably watching this video is because you want to create your own SaaS or at least some sort of product which can generate you money. And one thing I want to suggest to you is to focus more time on like the agentic side of things.
If you want to now create a SaaS like Intercom, like Marshall Desk, then please make sure that it's ready for the agentic future. Create your own plugin which is then accessible in Curse. cloud code, maybe an MCP server, create skills, make your application ready.
Because as you see here, the dashboard has become less relevant than ever. Opening the services dashboard is cool. Opening Neon's dashboard to check out monitoring is very important.
But at this point, it's not needed anymore to head over to Neon to, for example, configure authentication. Our agent is able to do everything by itself. And that's something I want you to also do.
So if you want to create Marshall that Ask yourself, make it ready, create an MCP server. Let people interact with your application through an MCP client.
As an example, there was a controversy recently with Figma. Figma is currently not allowing Pi to use the MCP server. So let me show it to you.
Today in MCP land, things are better compared to a year ago, but also worse. I'm trying to connect to Figma's remote MCP server. Figma only accepts certain client names during sign -in, for example, Cloud Code Codex.
etc. This in my opinion is a huge mistake. Allow your applications to also be used using MCP clients like for example Pi, Codex, etc.
It will set you apart from your competition. There are thousands of legacy services out on the market which print a lot of money but which are not ready for the gigantic future. You are someone who can now kind of exploit that and create your own application, your own business which is ready for the gigantic future and need Eon in this case can be your perfect partner to create this future.
So what questions do we now have here after my little monologue? Which production URL should Marshall Desk use? Eversell app URL for now, custom domain later.
Yes, correct, we don't have a domain right now. Google sign -in in production. I will create my own Google OAuth client, launch with email and password only, add Google later.
No, I want to say let's use the shared keys program. provided by Neon. I don't want to create my own client right now.
Once you deploy your own application into actual production, yes, you will have to create your own Google OAuth client and then also use the ID and secret and set it in Neon. But since we are still kind of deploying a demo project, it's not needed. We won't have actual like real users try to pay for this application or use it.
So we can just use the shared keys. Let me click on next and what do we have here? Can I start making production changes?
Neon production branch? Cloudflare worker? Vercel project?
Yes, go ahead. Step by step and report after each step. No, only prepare the script and ENV template for now.
Let me do the following. Let me quickly collapse this because the agent already provided a step -by -step guide. What has to happen in order?
Pick the production URL first. We already did that. We will use a Vercel provided domain.
Generate new production secrets. Okay. visitor token secret real -time token secret etc set up the neon production branch run the prisma migrations ah that's something i forgot create a public profile images bucket uploads already exists create a production storage credential and get a production ai gateway token run storage course with the production origin what else do we have here deploy the real -time worker to production wrangler .jsonc already names it.
What else do we have here? Deploy the jobs function and its triggers to production. Create the Vercel project.
Root directory region connected to this repository. Add all of the variables. Deploy.
Set up production off at the production URL. Replace the shared credentials. That's something we won't do.
And then also use NeonShared's sender. That's what we will do. Smoke test the PRD section 9 checklist on production.
Here's the important part. What I can do.
Write the production deploy scripts, generate the secrets through the Neon MCP, run the migrations, create a bucket, storage credential, and trust the domain. Okay, deploy the worker and set its secrets. That's good.
Through the Vercel MCP, create the project, set the env vars, and trigger deploys. Run the course script. What only you can do.
Choose the domain. We don't have to do that. Create the Google OOF client.
We don't have to do that. Check your plans. everything is fine here, do the final real sign -up yourself.
So our agent is able to do everything for us. Therefore, I will say yes, go ahead, step by step, and report after each step. Our agent will now get started.
It will first of all go to Neon, create our production branch, or it has already been created, but it will run the migration step, setup authentication, the bucket, our function, stuff like that. Then it will go over to Cloudflare, deploy our real -time server, head over to Vercel, deploy our application and then everything should work, at least in theory.
Let's now wait for everything to finish and then we will check it out. Surprise, surprise, we are finished. What I did, step by step.
Database applied all seven migrations to the Neon production branch. Its schema now matches development. The rest of Neon created the correct bucket, the public bucket, created a storage credential and an AI gateway token.
What else? Generate new production secrets, everything is collected in the env .prod file.
Realtime deployed the Marshall Desk Realtime Worker, so our durable object. With the production secrets, it rejects requests that don't carry them. Background jobs, in other words, our neon serverless function, deployed the jobs function with its 15 -minute auto -close schedule and the source upload trigger.
It rejects calls without the secret. Vercel project created Marshall Desk linked to the correct repository building from erp's web on node24.
Also the functions the serverless functions are could you say based in Frankfurt, if that makes sense, next to Neon. What else do we have here?
Oh, and by the way, since Neon and also Vercel both use AWS in the background, they will also be deployed not necessarily in the same data center, but definitely in the same region. Auth and storage added this domain as a trusted auth domain. Both production buckets now only accept browser requests from that URL.
First deploy built from main, Everything works, repo changes, not committed, real -time deploy prod, this all looks good to me. What you need to do, sign up on production, test the widget, know how future deploys work.
So here's what I want to now do. Let's first of all check out all of the dashboards. This here is first of all Cloudflare.
And if you go to the durable objects primitive, then you will see here that I already created a few. Durable objects so Marshall desk real -time real -time chat team flow Marshall desk real -time and these two servers are our development servers That's why we have so many requests and errors the errors are fine So what we can do is open one server just to see how everything looks like we have requests errors Request wall time and then you can also view metrics deployments all of that if you want to it's cool data But it's not always super super help So what should we now do?
Oh, and by the way, another reason why we have errors here is because we also checked that we get authentication errors, right? We wanted to make sure that only authenticated clients are able to make requests. That's why we have errors.
In this case, an error is not a bad thing. It's kind of a good thing. This means our server is working as we wanted it to work.
What we can also do is quickly head over to Neon, to the console. This here is our development branch. Let me switch into production.
And you will instantly see that we have one database better off is set up, right? Object storage functions and our AI gateway. And we even already have some load in terms of our database.
What I will now do is zoom in a bit. Let's go to our tables. What do we have inside of here?
Zero rows, that's wanted. And that's again why we use branches. The development branch can have different data than the production branch.
So even though we had already quite a few users in development, we have... zero users in production. And that's what you want to do.
You want to never develop against your production database, against your production, I don't know, auth users table, against your production object storage and stuff like that. It's not a good thing to do. Does it work?
Yes. But will you run into problems in the long run? Definitely.
So that's why you want to always use two different branches. One production branch, one development branch. Let's also check out better off the primitive.
We have zero users. Let's check out our plugins. What did we select inside of here?
Or what did our agent do? Enable organizations. That's good.
Enable magic link and allow new user registrations. What we can also do is quickly head over to our set settings to auth is everything configured here.
Domains, yes, our domain is configured. That's super important and that's why I love agents. They can do everything for me.
If I would have now done everything manually, I would definitely forget to add the domain here. Everything would then break and I would say, oh man, this is so annoying. I have to now open the logs and debug everything.
But now everything is done for me. I don't have to remember all of the steps. I don't have to think about the steps.
I have to just review whatever the agent is proposing and then everything will be done for me. So that's why you want to use MCP servers.
The rest also all looks good to me. Let me quickly go to our object storage primitive. We have two buckets here, public and private.
Let's check out our functions. We should have one function here, auto close and source uploaded. Let's quickly open the trigger, function jobs, trigger type, object upload, trigger name, function path, bucket name and path.
prefix. This looks beautiful. And finally, let's head over to our AI Gateway and we can still view all of the models.
So what should we do now? Well, I guess it's time to head over to Vercel. This here is our project, Marshall Desk.
It looks successful to me. In other words, the deployment looks successful to me. And you know what?
Let me open the domain right away. Marshall Desk, wow, the animation is beautiful. Let me click on sign in.
In this case, I'm not able to just log in because we don't have any users so I will sign up. Let me say Jan Marshall.
I will also use my email. Please don't write me any emails and I will also create a password. I will click on sign up and what do we have here?
Well we got redirected to our OTP input page. So let me open Gmail quickly and now I have a new email. Verify your email address Marshall Desk.
Let me copy the code. I will go back to my application, paste it inside of here.
And by the way, one thing I've just realized is that we even have metadata set. So on the top, you will see right here that we have a title, verify your email dot Marshall desk. That's beautiful.
I will now click on verify email and the next step is to now also provide our business name. So this is there to create the workspace. Let me say I will call it syntax path because that's my website syntax path and I will click on continue.
We should now get redirected back to the dashboard. This is currently loading and this is now the dashboard. We have our name.
We have a predefined avatar. We can change the color if needed. We can change the position.
All of this looks exactly just like in development. Let me turn on notifications right away. On the top right, bum, we have to also allow everything.
Allow. And now everything is set. Let me do a hard refresh because that's needed for everything to take effect.
Let's quickly check out light mode. Does everything look good? Yeah.
So let's again switch into dark mode to not blind you guys. Let me change the name. So syntax path agent 123.
Does everything update? Saved. let me quickly head over to the neon console i will go to the tables and inside of here we should finally have something new so workspaces and yes name syntax path what else can we do inside of here members we have one member let's also quickly check out the neon off schema and inside of here we have one account and the user signed up using the credential provider in other words email and password let me again go back to our application what can i do inside of here let Let's switch the color to, you know what, I love orange, it looks quite cool.
For the position I will use bottom left. What else can we do inside of here? Let's add an emoji to our greeting.
I will use this big brain emoji. Everything should autosave. Yes, saved.
We currently don't have any questions and that's because we need to upload knowledge and we can add a domain. Let me add again my own domain syntaxpath .com. I will click on add and it has been added instantly.
This looks beautiful. Let's head over to the knowledge space. Inside of here, I want to add some knowledge.
I will click on browse and in my downloads folder, I already have a docs folder. I will select all of these files, all of these, how should I say, knowledge files, and I will click on open. And I'll look at this animation processing, ready.
Did you see how fast everything was? Why is that? We use serverless functions, we use our object storage primitive, and since everything runs in parallel, everything is also almost instant.
So if we head over back to the Neon console and go over to object storage, then you will see here that we have our uploads folder, which is private. We have our workspaces, one workspace, and this workspace has sources. In total, what do we have here?
One, two, three, four, no, one, two, three, four, five files. Let me again go back to our application. Let me view one file.
Okay, I can view all of the chunks. So chunk one, chunk 2.
This looks great. What is if I delete a file? For example let's say shipping .txt, delete.
Does it also get deleted in our Neon console? I will do a hard refresh and yes we only have four files. So let me again upload the shipping file, our shipping knowledge file, processing ready.
Let's head over back to our home dashboard. We should now have suggested questions and yes that's what you can see right here. Let's go to our our inbox.
It should be completely empty. So what I can now do is already test everything out. So is localhost 4000 running?
Let me quickly check. Yeah, it's running. So what I will do is head over back to our application.
Let's go to home. I will copy the snippet. And by the way, do you see how fast everything is?
I can switch between the two pages or navigate between the two pages essentially instantly. That's super cool. Let me copy the install snippet.
Bum. Let's go back. I will paste it inside of here and then embed and reload.
We should now see it on the left, but I don't see anything. Do I have to add the domain to our allowed domains? Let's see.
Ah, now I remember. So in development, we added a feature where local domains like localhost, 3000, 4000, et cetera, all of these domains are enabled by default. But since we are now in production, these domains are not enabled by default.
And that's a good thing, by the way, because Normally, you also don't want to use local URLs in production because it creates a few security issues. But in this case, it's fine.
I will just say local host and I will click on add. Let's go back to our application and now I instantly see the widget. What should we ask right here?
Why do you ship and how long does delivery take? Sure, why not? We get our state or our typing state and here I have the response.
We currently don't ship to other countries. So only the US, Canada and UK. Man, where's Europe?
Where's the UAE? This does not work. What else can we say right here?
Thank you very much. This is, or yeah, let me say this is super help. full.
We instantly get a response. You are welcome. Glad it helped.
Can you or let's say talk to a human. Will we get redirected? A person will join shortly.
Let's head over back to Marshall Desk. Here in the nav bar or I'm sorry in this tab I can see like this number one in parentheses which means we have a new message. Waiting.
Let me open this. What do we have here? Visitor from Hamburg.
Wow that's super cool. Then what can I do inside of here local time language device this is all quite nice we can see all of the responses and inside of here i will say what should i say um yo what's up and if i click on enter we will get a notification in our widget because i now responded yo what's up nothing much how about you and if i click on enter we get a notification and i can see the message inside of here let's Everything is real time.
I will make this a bit smaller. Let's do the same thing here. So right half.
Let me also, I don't know, collapse the sidebar. It's not needed. And I will say the following.
What should I say? I'm not even sure. Let me start with the widget.
Hey, is this real time? And if I go back, then you should see the message here. Well, it didn't scroll, so let me try it one more time.
Yo, and if I click on enter or if I just type something, we get the indicator. This should now go away because we have a heartbeat. It takes about five seconds.
So as you see here, one, two, three, four. five six so about five to six seconds let me click on send and we instantly get a message here everything is real time and that's what we wanted so what should i do now well i can mark this as finished as completed and that's what you will see here this conversation has ended sent a message to start a new one Wow, this is beautiful.
Is there anything else that is missing? I don't think so. We can view all of the resources or all of the knowledge used to answer the question.
I can also view the file, view the chunks. And with that, I think we are now finished. We implemented all of the needed features.
Everything works. Everything is production ready. The code is clean.
I have zero complaints. Also, huge shout outs to Neon for making this video possible. Without them, I would have never...
Posted such a huge video. Because believe it or not. It took about two months to create this video.
Yes you heard right. From idea to actually publishing this masterpiece. And Neon supported me throughout everything.
So thank you. And again Neon is also the platform I privately use. This is not some weird advertisement.
I actually use them in development. In production. For my private companies.
For my private just testing. It's my preferred platform. I love it.
And Neon is also constantly evolving. for example recently Neon acquired Electric, the team behind the Electric Sync Engine and PG Lite, which means soon your database will be completely real -time without any third party server like Party Kit. Again, check out Neon using the first link down below in the YouTube description.
And with that out of the way, I hope you enjoyed the video, I guess see you in the next video, don't forget to like and subscribe, it would mean a lot to me and my heart, so please do it, and yeah. Enjoy your day and see you in the next one. Over and out.
Bye -bye.
The Hook

The bait, then the rug-pull.

The video opens with the billion-dollar revenue of Intercom on screen and then asks a coding agent to clone it in one prompt. Thousands of files appear and nothing works. The rest of the six hours is the argument for why: a wish is not a plan, and the workflow around the model is the part that senior engineers actually own.

Frameworks

Named ideas worth stealing.

13:57list

Foundation before code

  1. Define the idea and the explicit non-goals for V1
  2. Draw the system architecture as a diagram
  3. Map the core user flows
  4. Have the agent write the PRD from all three
  5. Grill the PRD until no open decisions remain
  6. Lock the tech stack into its own linked file

Every step before the first line of code exists to feed one PRD so the agent never has to assume.

Steal forAny new app or satellite project kickoff
1:06:00model

Agent roles

  1. Orchestrator: thinks, plans, delegates, judges
  2. Implementer: same strong model as orchestrator
  3. Researcher: cheap, fast, read-only
  4. Reviewer: independent read-only pass per sub-feature

Sub-agents keep the lead's context window clean and let cheap models do the high-volume reading.

Steal forStructuring any multi-agent coding run
1:12:00model

Two-session system

  1. Session one: ping-pong until the agent understands
  2. Ask it to compile a precise prompt
  3. Session two: fresh context, orchestrator executes
  4. Orchestrator splits the feature into sub-features A-D
  5. Reviewer sub-agents, then human review, then PR

Thinking and executing happen in different chats so execution gets a full context window and a prompt that already contains every decision.

Steal forLarge features in mod-creator or MCN
2:43:00list

Three-step code review

  1. Reviewer sub-agents spawned by the lead
  2. Manual read of every diff with your own eyes
  3. PR review bot (Bugbot, Graphite, CodeRabbit, PullFrog) in CI

Agents reviewing agents still miss duplicated helpers and wrong error handling; the human pass is the only one that catches taste.

Steal forAny repo where agents open PRs
3:58:00list

Post-generation order of operations

  1. Read the handoff document
  2. Test the whole flow in the browser
  3. Only then review the code
  4. Fix slop via prompts, not by hand
  5. Create PRs and let the bot run

Understand what the agent did, prove it works, then judge how it was written.

Steal forEvery large agent run
1:22:24concept

Landing page first

Building the marketing page before auth or backend produces a design.md that every subsequent page inherits, so nothing gets restyled later.

Steal forNew product scaffolds
CTA Breakdown

How they asked for the click.

VERBAL ASK
5:54:14link
“Check out Neon using the first link down below in the YouTube description. And with that out of the way, I hope you enjoyed the video, don't forget to like and subscribe.”

Sponsor link repeated at every Neon touchpoint across six hours, plus a long mid-video Neon product walkthrough around 1:32. Soft close with no product of his own pitched beyond the free repo.

Storyboard

Visual structure at a glance.

Intercom revenue cold open
hookIntercom revenue cold open00:01
Three core user flows
promiseThree core user flows28:51
Grilling the PRD on the database question
valueGrilling the PRD on the database question51:05
CursorBench reasoning-effort argument
valueCursorBench reasoning-effort argument1:10:18
Architecture diagram with Neon primitives
valueArchitecture diagram with Neon primitives1:30:30
Agent asks for real data vs mock
valueAgent asks for real data vs mock2:55:28
Acme Outdoor Co test site
valueAcme Outdoor Co test site4:00:28
Qwen3 vs OpenAI embeddings debate
valueQwen3 vs OpenAI embeddings debate4:37:38
Deployment checklist: what only you can do
valueDeployment checklist: what only you can do5:39:50
Production widget on the test site
ctaProduction widget on the test site5:51:20
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.

12:57
Rob Shocks · Review

Pstack Is Agent Overkill. Use It Anyway!

A walkthrough of Lauren Tan's pstack: 21 engineering principles, 22 playbooks, and 24 skills that turn a coding agent from a slop machine into a verification-obsessed engineer.

September 8th
33:48
Theo - t3․gg · Essay

He's right.

Boris Cherny said coding is solved. Matt Pocock called it VC-funded bullshit. Theo argues they're both right, because they're using the word coding to mean two different things.

August 24th