A stronger model still builds a generic landing page from a text brief alone. Feeding it 600,000 real UI screens through Mobbin's MCP is what closes the taste gap.
A more capable AI model does not automatically produce better visual taste; feeding it real design references through a research phase before building is what turns a generic AI landing page into something that could actually ship.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You use an AI coding agent (Codex, Claude Code, Cursor) to build marketing sites and are frustrated that the output always looks templated.
You're a solo builder or small team shipping SaaS landing pages without a dedicated designer on staff.
You want a repeatable process for getting an AI agent to research real products before it designs anything.
You're weighing whether a paid design-reference tool like Mobbin is worth connecting to your AI workflow.
SKIP IF…
You already have an established brand system or a designer doing this work.
You're looking for a benchmark review of GPT-6 Astra's coding or reasoning performance rather than its visual design output.
TL;DR
The full version, fast.
GPT-6 Astra gets a text-only brief and, in one shot, builds a responsive AI-agent-marketplace landing page. It is fast and functionally complete but reads as unmistakably generic: default SaaS layout, floating banners, no real personality. The creator then connects the Mobbin MCP, a server that gives the model access to over 600,000 real app and website UI screens, and has it research competitors and write a design report before touching the page again. That report names specific problems (a hero that reads as a generic tool, cards missing proof), cites real examples like Clay and Vercel, and proposes a page order. Rebuilding from that report, then hand-annotating a few remaining details, produces a page with hover states, a rotating logo carousel, and real brand personality. The conclusion: don't ask AI to turn a brief straight into a finished design. Give it a research phase against real references first, and still supply your own taste on top of what it finds.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Sets up the test: does GPT-6 Astra's raw capability fix generic AI web design, or does good design require understanding a product category, not just clean code?
01:06 – 02:26
02 · The test
Explains the methodology: build the same AgentHub marketing site brief twice with the same model, once with no design references and once after AI-powered research, then compare.
02:26 – 03:43
03 · The brief
Reads out the exact prompt given to GPT-6 Astra: required sections, a product UI preview, and text-only visual direction ('modern B2B SaaS,' 'polished typography').
03:43 – 05:03
04 · Draft 1
GPT-6 Astra builds the first version from the brief alone. The result is functionally complete but reads as clearly AI-generated, compared to Bootstrap-default styling.
05:03 – 05:20
05 · Mobile check
Confirms the first draft is at least responsive across the required breakpoints before moving on.
05:20 – 06:33
06 · Mobbin MCP setup
Creates a Mobbin account, opens its MCP tab, and connects it to Codex via the plugins panel so GPT-6 Astra can query Mobbin's screen library directly.
06:33 – 07:21
07 · Design research
Prompts GPT-6 Astra to research real examples of the product category through Mobbin and produce a report before making any design changes.
07:21 – 08:28
08 · Research report
The agent's report identifies specific problems (generic hero, cards missing proof), cites real comparables like Clay, Cursor, and Vercel, and proposes a page order.
08:28 – 09:35
09 · Draft 2
Rebuilds the landing page from the research report. The result has hover states, avatar interactions, and a distinct brand personality the first draft lacked.
09:35 – 10:52
10 · Annotating changes
Uses Codex's click-to-annotate tool to make small targeted fixes: turning a logo row into a rotating, pausable carousel and adding real company logos to agent cards.
10:52 – 11:18
11 · Final tweaks
Removes an unnecessary bottom row and confirms the annotated changes rendered correctly.
11:18 – 12:11
12 · Before vs. after
Places the two drafts side by side. The first looks generic apart from its final CTA; the second reads as an established product rather than a one-prompt AI site.
12:11 – 13:33
13 · Recap
States the core conclusion: a stronger model builds fast and functional pages, but without real references it still has to invent the design direction from its own assumptions.
13:33 – 13:52
14 · Try Mobbin
Closing affiliate pitch for Mobbin, with its scale stats (1,428 apps, 621,500+ screens, 323,900 flows) shown on screen.
Atomic Insights
Lines worth screenshotting.
A newer, more capable model (GPT-6 Astra) still defaults to generic SaaS-template visuals when working from a text-only brief with no real design references.
Content requirements (sections, CTAs, a product UI preview) are easy for an AI agent to satisfy from a brief; a distinctive visual identity is not.
Signs the creator used to spot 'AI-generated' design included a floating banner over a UI mockup, a small eyebrow-text label, and default Bootstrap-like layout patterns.
Connecting an AI coding agent to a library of 600,000+ real UI screens (via the Mobbin MCP) measurably changed the page's structure and interactions, not just its colors and fonts.
Having the agent write a research report and recommend changes BEFORE rebuilding keeps a human's judgment in the loop instead of letting the AI redesign unsupervised.
A useful research-report prompt asks the model to name where the current design feels generic, summarize patterns from real comparable products, and propose a visual direction, before any rebuild happens.
Naming real, checkable comparables (Clay, Cursor, Vercel, LangChain) and stating what each does well made the AI's design recommendations concrete instead of vague.
A visual 'annotate' tool, click an element and type the fix, was a faster way to handle small post-redesign touch-ups (adding a rotating logo carousel, removing an unneeded section) than re-prompting in chat.
The clearest tell that the first draft was AI-made: its most distinctive section was the final CTA, essentially just an icon, headline, and button.
The stated conclusion of the whole test is that model strength and design taste are separate axes. A stronger model does not close the taste gap by itself.
Takeaway
A stronger model still can't replace real design references.
WHAT TO LEARN
Model strength and design taste are separate problems: closing the gap between an AI-generated page and a shippable one takes a research phase against real examples, not just a smarter model.
02The test
A more capable model doesn't automatically fix generic AI design because good design is about how a category of product should be presented, not just clean code.
The methodology isolates one variable: same brief, same model, the only difference is whether the AI got real design references before building.
03The brief
A text-only brief, even one that specifies a style like 'modern B2B SaaS' or 'polished typography,' gives a model no measurable way to produce a unique visual direction.
Content requirements (sections, features, CTAs) are easy for AI to satisfy from a brief; visual identity is not.
04Draft 1
Even a top-tier model produces a page that looks clearly AI-generated, comparable to default Bootstrap styling, when working from text instructions alone.
Tells of default AI design include small eyebrow text, a floating banner over a UI mockup, and generic SaaS layout patterns repeated across sections.
06Mobbin MCP setup
Design-reference tools connect to AI coding agents the same way any MCP server does: create an account, generate a connection, and authorize it inside the coding tool.
Desktop AI apps add MCP tools through a plugins panel, not a terminal command, unlike CLI tools.
07Design research
Having the agent write a research report before making any design changes keeps a human's judgment in the loop instead of letting the AI redesign unsupervised.
The same research prompt, identify what feels generic, summarize patterns from real examples, recommend a visual direction, works as a reusable template for any AI-driven redesign.
08Research report
A good research report names a specific problem, like a hero section that reads as a generic tool, rather than a vague verdict that the design needs to look better.
Citing named, real-world comparables and stating what each does well makes AI design recommendations concrete and checkable instead of vague.
Being able to open the actual reference screens, not just a text description of them, lets you verify the AI's interpretation before committing to a rebuild.
09Draft 2
Redesigning from a research report, rather than a fresh prompt, can produce hover states and interaction design a first draft never had, not just different colors and fonts.
10Annotating changes
A visual annotate-and-click tool is a faster way to handle small post-redesign fixes than re-describing the location of an element in a chat message.
Small, concrete instructions, turn this row into a carousel, remove this section, are the right size of ask for a touch-up pass.
12Before vs. after
Comparing two drafts side by side surfaces a difference that describing the improvement in words would not.
The one section with real personality in a generic AI draft, often the final CTA, is a useful diagnostic for how far the rest of the page still needs to go.
Glossary
Terms worth knowing.
MCP (Model Context Protocol)
A standard way for an AI coding agent to connect to an external tool or data source, such as a design-reference library, so the model can pull in outside context beyond its training data.
Mobbin
A searchable library of real app and website UI screens, used here as a design-reference source an AI agent can query before building a page.
Codex
OpenAI's coding agent, used here in its desktop-app form to run GPT-6 Astra against the landing page project and connect it to the Mobbin MCP.
GPT-6 Astra
The OpenAI model being tested in this video for coding, research, and agentic work, evaluated here specifically on visual and UI design quality rather than raw benchmark performance.
Annotate feature
A Codex tool that lets a user click directly on an element in a rendered preview and type an instruction, rather than describing the location and change in a chat message.
“I wanted to test whether a better model alone actually fixed one of the biggest problems with AI-built websites. They still tend to look generic.”
States the whole video's thesis in one line before any screen recording starts.→ TikTok hook↗ Tweet quote
03:33
“Even on light mode, it's actually more powerful than GPT-5.6 Sol on high effort level.”
Concrete model comparison claim, quotable on its own.→ newsletter pull-quote↗ Tweet quote
04:03
“What we have here is, in my opinion, a pretty clearly AI-generated site. It looks extremely default. It almost looks kind of like Bootstrap.”
Blunt, specific criticism of the AI-only draft with a relatable comparison.→ IG reel cold open↗ Tweet quote
12:36
“The biggest takeaway here is that a stronger AI model still doesn't automatically give you better taste.”
The thesis restated as a standalone, tweet-length claim.→ TikTok hook↗ Tweet quote
13:46
“The next time you ask AI to design a website, don't just give it a brief, give it references worth learning from.”
Closing line doubles as a practical rule of thumb.→ newsletter pull-quote↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
analogy
GPT -6 Astra is supposed to be OpenAI's most capable model yet for coding, research, and complex agentic work. But I wanted to test whether a better model alone actually fixed one of the biggest problems with AI -built websites. They still tend to look generic.
You can give AI a solid brief, ask for a polished landing page, and even use the best available model, but it will often fall back on the same layouts, the same visual styles, and the same generic SaaS patterns. And that's because good web design isn't just about generating clean code. or making something look polished.
It's about understanding how a specific category of product should be presented, how the page tells the story, how the product itself gets worked into the design, and what visual patterns make the whole thing feel believable. So for this video, I'm putting GPT -6 Astra through a simple test. First, I'll have it build a marketing site for an AI agent marketplace completely from scratch.
Then I'll give it access to Mobbin, have it research relevant real -world products and design patterns, and ask it to completely rethink the same website based on what it learns. Same brief, same model.
The only difference is the design context it gets before building. So let's see how much better AI gets when it actually studies the category first. So here's what this process is going to look like.
We're going to generate a first draft and we're going to do this solely by giving GPT -6 a web design brief. which is essentially going to act as a prompt for GPT -6 to generate this landing page. Now, this is going to show us how capable GPT -6 really is at web design and how much design skill it has.
But at this point, for the first step, we're not going to be giving it any reference images. And the brief is going to include some direction about what kind of visual design we want. But this is just going to be through text -based instructions.
And this is where we're going to see the second step of today's web design process do some real heavy lifting here. We're going to have GPT -6 research what good UI and web design looks like.
And we're going to do this by feeding it over 600 ,000 design references through a really special MCP server that I'm going to show you how to set up. Then once it's done that research, we're going to design a second draft incorporating all the findings it found about best design practices, design patterns, visual directions, content guidelines, all through that research that we had it do.
And then at the end, we're going to look at a before and after of what the website looked like before research and after research. And just a heads up, the MCP server we're going to be using today is a paid feature from a company called Mobin. And you can use my link in the description for 20 % off.
But Mobin is also free to get started with if you want to do the research manually and then feed it to your AI agent. So here's the first prompt that I'm using to design and build this landing page, which is essentially the brief for our website. Build a polished responsive marketing website for Agent Hub.
marketplace where business teams discover and deploy AI agents for sales, support, and more. And although we're creating a marketing landing page, I also want GPT -6 to include a product UI preview. And this is going to make this web design project a bit more all encompassing where it's a web design project at its heart, but it's also including a touch of UI design.
inside it. Then I've also listed off the rest of the sections that I want GPT -6 to include in this first generation of the landing page. And then I'm ending the prompt with a couple instructions about the visual direction I want.
It should feel like a modern B2B SaaS company, which it is, and I want polished typography, subtle interactions, responsive layouts, and also generate your own images if needed, which Codex can do with the images 2 .0 image generation model. Now you'll see down here that I'm using GPT -6 Astra, and by default I actually only use it on light mode.
And that's because GPT -6 is such a strong and powerful AI model that even on light mode, it's actually more powerful than GPT -5 .6 Sol on high effort level. So even on light mode, you should be very happy with the performance of GPT -6 Astra. So with this first prompt in place, let's have GPT -6 get to work.
All right, so now we have the first draft of this landing page running here as a ChatGPT site. So let's take a scroll through and see how it's looking.
All right, so what we have here is, in my opinion, a pretty clearly AI -generated site. It looks extremely default. It almost looks kind of like Bootstrap.
Even the product UI showing in the hero section is looking very AI -generated. Even just over here on this right side, I'm seeing a couple more signs of AI like this small eyebrow text up here and this banner kind of floating over the UI at the bottom. I'm just seeing countless patterns here that look like AI to me, even when we generate with GPT -6 Astra.
Now, GPT -6 Astra is, of course, incredibly powerful, but just because you look at the benchmarks and see that it's a powerful model does not really give any indication about whether it's good at visual design, UI design, and web design or not. And honestly, I'm not surprised at this result, given that we only really told Codex what content we wanted included, which it got right, to be fair, and that we wanted the design to feel like a modern B2B SaaS company.
And when you're just giving it text -based instructions like that about the design of the site, you can't really expect that to have any measurable impact on how unique or polished the design is. But I did ask Codex to include responsive layouts only. So let's take a scroll through the mobile view and just make sure it's at least responsive.
All right, yeah, every section is looking very responsive, so I'm happy with that. But regardless, this is clearly an AI -generated website, so let's move on to the second step of this web design process. And the second step is to have GPT -6 research Mobin's library of over 600 ,000 UI screens to inform a second draft of our landing page.
This step is basically gonna force GPT -6 to understand what good UI and web design looks like. And the only way you can really do that is by giving it actual visual references and giving it over 600 ,000. of them is pretty much the surest way to do that.
So let's get set up with Mobin's MCP server so that we can connect Codex to the Mobin library. Start off by creating a Mobin account. go into your settings, and go into the MCP tab.
And remember, to get started with Mobin and get 20 % off your subscription, use the link down in the description. The setup is really simple. Just go down to Connect More Tools.
And if you're using Cloud Code, Cursor, or anything else via a CLI or from the terminal, then choose any of these options immediately listed here, depending on what tool you're using. But if you're using a desktop app to use AI, like I am with Codex, for example, go ahead and hit Other. Then all you have to do is copy this one -line URL.
Going back to Codex, I'm going to open my left sidebar, then open plugins, then just search for the Mobbin plugin. Hit the plus button, and then sign in with your Mobbin account to authorize it. Now Codex is connected to Mobbin's library of over 600 ,000 design references.
So to have Codex start researching, I'm saying use the mob and MCP to research design patterns for the current Agent Hub landing page before making any changes. Research real examples of the type of site we're working on, and before you actually build or implement any design changes, generate a report that identifies where the current page feels generic, summarizes recurring patterns of relevant examples, finds a better visual direction, and lists the changes it would make and why.
And the reason I want to report before I just have it immediately redesign our website is that I want to understand myself what competitors are doing, have some input about the design direction we go in before it builds anything, and ultimately just bring more of my taste as a designer on top of the research that I'm having my AI agent do.
So let's have GPT -6 get started on this design research report.
So now, all from inside Codex, it's come up with a dedicated report about its findings from Mobbin. Based on its research, the overall proposed direction that it's landed on from my landing page is make Agent Hub easier to evaluate. But that doesn't say much on its own, so let's scroll through what it actually means by that.
It says the hero section feels too general. Some of the cards on our page omit key details and claims lack proof. And it's also given us a recommended page order.
So let's scroll down further and see what these proposed changes are actually based off of. It's found a handful of relevant existing interfaces from Clay, Cursor, Vercel. link chain.
And for each of these examples, it's saying what actually works about the example and how it should be implemented into agents hub landing page. And the coolest part is for each of these examples, I can click on them and then open up the actual reference inside Mobin. I can scroll through these entire landing pages to see what exactly I want to bring into my site.
And I can also scroll down to see similar references as well. And the best part is it's showing me landing page examples that are relevant to my landing page. So with this research report done and all of this context in place for GPT -6 to work off of, I'm telling it redesign the Agent Hub landing page based on the context included in this report.
Completely overhaul the visual style based on the references you found, making it feel appropriate and on -brand for this type of product landing page. So let's leave it at that and see what it comes back with for its second draft.
All right, check it out. So now we have this redesigned website based on that research report that Codex came up with using the Mobin MCP. And at first glance, what I'm loving about it is it looks so much more personalized and customized.
It feels like the site actually has a personality now, and it's gone as far as redesigning the interactive UI that we have down here in this section. The brand identity before really just felt completely AI generated and artificial, but based on the insights that GBT6 found from Mobin, It now feels much closer to a website we could actually ship.
So this is a drastic improvement from what we saw before. We're even seeing some nice little hover effects like with this card here and with these little agent avatars up in the hero section. But I'm still seeing a couple small things I'd like to touch up.
And the way I prefer to iterate on my designs from Codex is using the annotate feature. So I'll enable it by clicking annotate in the upper right. And then just to create a more dynamic landing page, let's have this trust row be a rotating carousel.
Now, this is, of course, a demo landing page that I'm making for the sake of this video. But just to treat this as a realistic web design project, I'm going to say add some more company logos and turn this row into a dynamic carousel. And then I also want this carousel to take up the entire row.
So I'm going to remove these groups of text on the left and right.
And now I'm going to go through and see what else we could make some small improvements to. I like the theme of using actual company logos like we're currently doing up in that trust row. So I want to make sure we do the same thing down in these company chips right here.
So I'll say include the relevant company logos for all of these company chips. Size properly. And I think this section is looking beautiful so far.
I love the visuals going on in each of these cards. But this row way down here at the bottom is looking kind of awkward and unnecessary. And I want to keep this landing page looking as clean as possible.
So I'm going to remove it. Like I said, I'm a huge fan of this annotate tool because it makes it super easy to go through and call out exactly what you want changed or removed or added. So with all of those annotations called out, I'll hit send.
All right, so now Codex has executed those annotations. So let's check them out. We have this automatically rotating carousel down here with more company logos.
And the carousel pauses upon hover. And we also have this pause and play button in the right, which is a nice touch. It's also gotten rid of that bottom row down in this section, making it look a lot cleaner.
And we now have real company logos in each of these cards' chips. So now let's go and look back at our first generation and see how far we've come using the Mobbin MCP. So here's the first version that we had come up with.
Even using GPT -6 Astra, it looks incredibly AI -generated. And there's pretty much nothing about this version of the landing page that makes it stand out from any other site. It does include some nice visuals of a real product UI, but in my opinion, the most standout thing about this version of the landing page is the final CTA here, which somehow has more personality than the rest of the site, even though it's pretty much just an icon text and a button.
And this is a very stark contrast between our latest version of the landing page, where it looks much more editorial, personable, and like a website you would expect to see from an established company rather than just some small startup that generates their website in one prompt with AI. And this just goes to show you the power of having your AI agent do real research against real visual references to understand what good design looks and feels like.
So just to recap, we started off this design process with having GPT -6 Astra create the landing page on its own without any design references. Then after that first draft, we connected Codex to the Mobbin MCP to give GPT -6 over 600 ,000 real design references. Then we had Codex create a report based off of its findings, and then moved on to a second draft, which was a full redesign of the site.
based on that research and that report the biggest takeaway here is that a stronger ai model still doesn't automatically give you better taste the first website generation was usable astra could build the page make it responsive create the product ui and get to something polished incredibly quickly but without strong references it still had to fill in a lot of the design direction from its own assumptions once we added mob in research it had better context for what this type of product should actually look and feel like the page structure became more intentional the visual direction felt more specific to the category, and even the way the product itself was presented inside the landing page changed.
And I think that's a more solid and reliable workflow for AI web design going forward. Instead of asking AI to immediately turn a brief into a finished website, give it a research phase first. Have it study relevant products, identify the patterns that actually matter, decide which ones make sense for your project, and then build.
You still need to provide the taste and judgment, but the less the model has to invent from scratch, the better the starting point becomes. comes. If you want to try this workflow yourself, I'll leave my Mobin link down below.
You can try it for free and it gives you 20 % off a paid subscription to use the MCP server. So the next time you ask AI to design a website, don't just give it a brief, give it references worth learning from. Thanks for watching and I'll see you in the next video.
The Hook
The bait, then the rug-pull.
Griffin Wooldridge takes OpenAI's newest model, GPT-6 Astra, and asks a narrower question than any benchmark answers: does raw model strength fix the number one complaint about AI-built websites, that they all look the same? He runs the same brief through the model twice, once cold and once after feeding it 600,000 real UI screens, and lines the two results up side by side.
Frameworks
Named ideas worth stealing.
01:08list
Research-before-rebuild workflow
Draft 1: build the page from a text-only brief, no design references
AI-powered research: connect a design-reference MCP (Mobbin), have the agent study real examples, and generate a report before changing anything
Draft 2: rebuild based on the research report, with your own taste and judgment layered on top
The three-step loop used to close the gap between an obviously AI-generated page and one that could plausibly ship.
Steal forany AI-driven landing page, app UI, or marketing site build
CTA Breakdown
How they asked for the click.
VERBAL ASK
13:33link
“Get 20% off your Mobbin subscription, use my link in the description.”
Standard affiliate close, reinforced by on-screen scale stats (1,428 apps, 621,500+ screens) right as the pitch lands.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
A hands-on tour of Brilliant, the AI-native vector design tool built to fix the biggest gap in Claude Design and Google Stitch: you can generate a UI in seconds, but you can't actually edit it.
A working product designer wires Claude Code into Mobbin's 600,000-screen UI library through an MCP server, then builds a profile page and redesigns a financial dashboard entirely from cited, real-world reference patterns.
Mobbin's new MCP connector lets Claude search 600,000+ real, shipped app screens as design references — trading the generic serif-and-purple-gradient AI look for interfaces grounded in production apps.
A creator runs the exact same prompts through Claude's new Fable 5.1 and the outgoing Fable 5 to build a pizza ordering app, a cloned award-winning landing page, and a 3D snowboarding game, then compares the results side by side.