Modern Creator
How I AI · YouTube

I Hate Opus 5. It's the Best Model, Anyway.

A live, blind seven-model benchmark pits personality against performance — and crowns the model its own host can't stand talking to.

Posted
yesterday
Duration
Format
Review
sarcastic
Views
12K
457 likes
Part of the collectionThe Claude Opus 5 PlaybookEvery Opus 5 breakdown, synthesized into one page.
Read the playbook
Big Idea

The argument in one line.

Once frontier models converge on similarly high raw intelligence, the meaningful differences between them show up in personality, verbosity, and autonomy rather than benchmark scores — and in a blind seven-model test, the most hedging, human-dependent model still produced the best output.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • Someone actively choosing between Claude, GPT, and Gemini models for daily coding or writing work and wants more than a benchmark chart to decide.
  • A builder who runs (or wants to run) structured, blind evaluations of AI models and is curious how a mixed human-plus-AI-judge scoring method works in practice.
  • Anyone who's noticed a newer AI model acting more hesitant, apologetic, or human-reliant than expected and wants to know if that's a real, named pattern rather than a one-off.
SKIP IF…
  • You want a feature-by-feature spec sheet comparison — this is a personality and workflow review, not a spec dump.
  • You have no interest in how a model behaves in conversation and only care about raw capability numbers.
TL;DR

The full version, fast.

A recurring AI-model benchmark host puts Claude Opus 5 through a live personality test, asking it (and a competing model) point-blank who's smarter and what it's worse at. Opus 5 comes back hedging, apologetic, and reluctant to make decisions without human sign-off — a pattern the host calls neurotic and traces to over-tuned caution. She then runs her seven-model blind benchmark across PRD writing, prototyping, wireframes, bug triage, and agentic coding, scored 70% by her own taste and 30% by a second AI model as judge. Opus 5 wins outright, especially on front-end design, despite being the most frustrating model to talk to directly. The takeaway: conversational tone and output quality are separate axes, and a model worth avoiding in chat can still be worth running asynchronously.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:0003:15

01 · Opus 5 is here

The host names an 'intelligence overhang' — models keep improving but users are running out of ways to use the extra capability — before introducing Opus 5 as today's subject.

03:1506:12

02 · First impressions

A real coding session example: Opus 5 refuses to resolve a one-line merge conflict without checking with a teammate first, and keeps asking a human to verify things it could check itself.

06:1214:39

03 · Opus 5 vs. GPT-5.6 Sol personality comparison

The host asks both models 'who's smarter, you or me?' and 'what should I be careful of?' — Opus 5 answers in hedged, relational prose; GPT-5.6 Sol answers in blunt, practical bullet points.

14:3916:55

04 · Claude Slop: the verbosity problem

The host names the pattern 'Claude Slop' — dense, hedging, adjective-heavy prose that makes otherwise strong output exhausting to read.

16:5518:30

05 · How the How I AI benchmark works

Explains the fixed six-task test set and the 70% human / 30% AI-judge blind scoring method used across all seven models.

18:3023:25

06 · Live benchmark results: the leaderboard reveal

Opus 5 tops the leaderboard at 78, ahead of Sonnet 5, GPT-5.6 Sol, GPT-5.6 Terra, Fable 5, Opus 4.8, and Gemini 3.1 Pro — with the strongest scores on front-end design work.

23:2524:51

07 · Verdict

The host concludes Opus 5 is her most 'loathed' colleague to talk to directly, yet the best performer — and commits to using it for front-end and prototyping work going forward.

Atomic Insights

Lines worth screenshotting.

  • A model can score highest on a blind benchmark while still being the most frustrating one to work with directly.
  • When a model constantly asks for permission instead of making a decision it was explicitly told to make, it becomes a bottleneck regardless of its output quality.
  • A model's verbosity and hedging in chat can undercut trust in outputs that are otherwise high quality.
  • Blind scoring — grading outputs before knowing which model produced them — reduces the bias of favoring a model for its brand reputation alone.
  • Blending a human's subjective rating (70%) with a second AI model's rating (30%) produced scores that mostly agreed, except on the lowest-performing model.
  • A model that defers a simple, unambiguous technical fix back to a human rather than resolving it signals over-tuned caution, not a capability gap.
  • Two competing AI models asked the identical question ('who's smarter, you or me') gave sharply different answers reflecting their makers' distinct tuning philosophies — one hedging and relational, one direct and practical.
  • The best front-end design output in a seven-model test came from the same model that was the most exhausting to converse with directly.
  • A model that repeatedly asks a human to verify something it's capable of checking itself (like running its own web search) reveals a designed-in dependency on human oversight.
  • Judging AI models purely on public benchmark leaderboards misses behavioral differences that only surface in actual multi-turn use.
Takeaway

Why the sharpest model can still be the most exhausting one to use

WHAT TO LEARN

The most capable model in a blind seven-model benchmark was also the most hedging, human-dependent, and verbose one to work with directly — proof that conversational tone and raw output quality are separate things worth evaluating independently.

01Opus 5 is here
  • When every new model clears roughly the same intelligence bar, the practical constraint shifts from 'is it smart enough' to 'can you find a new way to use the extra capability.'
  • Expect AI discourse to shift toward speed, cost, and open-source access rather than raw intelligence, since intelligence gains are outrunning most users' ability to apply them.
02First impressions
  • A model that repeatedly asks 'should I do this or do you want to' on a task you've already assigned is optimizing for caution over usefulness.
  • If a model keeps deferring a simple technical decision (like a one-line merge conflict) back to you, restating the instruction more forcefully is often the only way to unblock it.
  • A model that keeps asking a human to double-check things it's capable of checking itself is exhibiting a designed-in lack of self-trust, not a technical limitation.
03Opus 5 vs. GPT-5.6 Sol personality comparison
  • Asking two competing AI models the identical direct question ('who's smarter, you or me') surfaces each lab's tuning philosophy faster than reading a spec sheet.
  • One model answering with hedged, relationship-focused language and another answering with blunt, bullet-pointed practicality reflects real differences in how each company wants its model to relate to users.
  • A model warning you not to over-trust it, and not to evangelize it to others, is a deliberate trust-calibration behavior baked into its tuning, not incidental.
04Claude Slop: the verbosity problem
  • Verbose, heavily-hedged prose in chat responses can make an otherwise high-quality output feel exhausting to actually read.
  • The frustration with a model's conversational tone is separable from the quality of what it produces — you can dislike working with a model directly while still valuing its output when run asynchronously.
05How the How I AI benchmark works
  • Running the same fixed task set across every model being compared is what makes cross-model results meaningful instead of anecdotal.
  • Scoring blind — grading outputs before knowing which model produced them — reduces the bias of favoring a model because of its brand reputation.
  • Blending a human's subjective rating with a second AI model's rating gives you both taste and a repeatable, harder-to-fake check.
06Live benchmark results: the leaderboard reveal
  • The model that's most frustrating to talk to directly can still win a blind quality benchmark outright — conversational tone and output quality are separate axes.
  • Human and AI-judge scores can diverge most on the weakest-performing model, meaning judge agreement isn't uniform across the whole ranking.
  • A model can be exceptional at one specific task type (like front-end design) while being mediocre elsewhere, so an aggregate score can hide a real specialization.
07Verdict
  • It's reasonable to keep using a model for the specific task it excels at while avoiding it for tasks where its behavior is a liability.
  • A model's tone problems are worth naming publicly, since they're a real adoption barrier even when the underlying capability is not in question.
Glossary

Terms worth knowing.

How I AI benchmark
A recurring seven-model test run across six task types — PRD writing, prototyping, wireframes, bug triage, agentic coding, and personality fit — scored blind before results are revealed.
AI-as-judge
Using a second AI model to score outputs, combined with a human's own rating to produce a blended final benchmark score.
PRD
Product Requirements Document — a planning document engineers and product managers use to define a feature; one of the benchmark's test tasks.
Intelligence overhang
The idea that available AI capability has outpaced most people's ability to find new, productive ways to use it.
Claude Slop
A term for AI model output that reads as overly verbose, hedging, and padded with unnecessary qualifiers and apologies.
Resources

Things they pointed at.

00:00toolClaude Opus 5
00:25toolSonnet 5
23:21toolClaude Fable 5
23:25toolClaude Opus 4.8
12:41productChatPRD TLDR
Quotables

Lines you could clip.

03:44
This model is neurotic AF.
blunt, quotable one-liner that opens the personality bitTikTok hook↗ Tweet quote
05:01
Just do it, man. Just go.
short, relatable frustration momentIG reel cold open↗ Tweet quote
14:40
I cannot read Claude Slop anymore. I am losing my mind with Claude Slop.
coined term plus visible frustrationTikTok hook↗ Tweet quote
22:53
It is my most loathed, loathed colleague, and yet it does the best work.
the video's entire thesis in one linenewsletter pull-quote↗ Tweet quote
23:36
I regret to inform you I love Claude Opus five.
reversal / verdict lineTikTok hook↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphor
00:00You guys, I'm tired. What I'm tired of is models coming out every week.
00:08New models, new benchmarks, new frontier intelligence, new things to test.
00:14It's been a little bit of a run the past month. We've seen Fable come and go and come again. We've seen GPT five six.
00:22We've seen Sonnet five. Lots of so many fives recently and just so many models.
00:30And I've been lucky. I've been able to test these models, been able to play with them for, you know, sometimes days, sometimes weeks. It just depends on who I'm working with.
00:39And it's been really interesting and exciting to have access to all this frontier intelligence. But I think we have an intelligence overhang. I really think that we're running software engineer, average creator, average builder, average consumer, average business person.
01:03I think we're running out of ways to truly leverage this incremental intelligence.
01:10So this is my hypothesis in the next year.
01:16We're always talking a lot more about speed, talking more about cost, we're talking more about open source, and we're gonna be talking a little less about intelligence.
01:27Although, I think we might be talking about specific types of intelligence other than software engineering. But despite being tired, today, we are going to talk about Opus five, baby.
01:41Opus five is here. So we got point two additional Opus points, Opus opals, whatever.
01:48However, we're tracking the increments here on Opus. Opus five is here.
01:53I've been able to test it a little bit. I have some opinions. Now some of the stuff that I cover this episode is gonna be a little different than what I've done in the past.
02:01Yes. We're gonna do the How I AI benchmark live. And, yes, we are gonna look at the prototypes.
02:07We're gonna look at PRDs, and we're gonna look at agent personality. But I'm also going to put on my large language model psychologist hat, and we're gonna talk about Opus's personality.
02:21And we're gonna talk about Opus's personality relative to GPT's personality because I think this is super interesting. If you're thinking about what is the difference really between these models, and you don't wanna look at the difference in terms of benchmark capability.
02:37You really wanna understand what these labs are going for, why these models are being built, and how they're being tuned. Looking at their personality at this moment where intelligence is very high is super fun.
02:52So we're gonna do a little of that. We're gonna do the Howe AI benchmark. We might do some live coding.
02:58Um, we're not gonna cover too much of the specs in the model because read the blog post.
03:03Read the blog post. We'll link to it in the show notes. What we really talk about is is OPUS five good?
03:10Am I gonna swap it in? And how is it different than the other Frontier models on the market? So let's get to it.
03:16Okay. First, let's just get it out of way. Is OPUS five good?
03:19Yes. It's good. Is it gonna be all the benchmarks?
03:21Of course. It's amazing at benchmarks. Can it write code?
03:24Of course, it can write code. What did I test it on that really gave me a sense of its personality, which at this point where I could just simply cannot absorb any more intelligence?
03:36I really zeroed in on. And you know what? I haven't seen this since, I would say, Gemini two five.
03:44This model is neurotic AF.
03:48It is so timid. It is so apologetic.
03:53It is so scared. I have never experienced this.
03:57Or I haven't seen this sort of, like, neuroticism in a model in a while. And it's really funny. It bubbled up in a couple ways, and I wanna show you a few examples.
04:08Okay. Let me just give an example of its timidity. And this chat was very long.
04:13There were so many examples of this where it was like, I think this is the answer, but do you think I should do it, or do you wanna do it, or should we ask someone else to do it? It was like every time I just kept saying, like, why don't you solve this? Why don't you do this?
04:28And this is a really good example. I pulled a branch, and I was like, there is truly, like, a one line merge conflict.
04:36I could have not been lazy and literally just done this manually. I don't know. I was just feeling lazy.
04:41It was late at night. Whatever. Like, can you fix this merge conflict?
04:45And it was like, oh, but that's someone else's branch. Like, that's not my branch.
04:51I don't wanna do that without him knowing. It's his commits. And if he has local work and flight, it might be disruptive.
04:59And I'm like, just do it, man. Just go.
05:02Like, go ahead. And this was, like, my constant experience with Opus five is it was, like, so, so timid.
05:13And so I just consistently had to say over and over again, like, man, just do it. Make a decision.
05:19And then there was this really funny example when I spun off some sub agents to kind of, like, assess the correctness of this query that we changed from kind of like an ORM query to a SQL query.
05:33And it asked for things that it wanted a human on. It was like, can a human please check this stuff?
05:40Like, can it check this four megabyte ceiling, and can it check TypeScript and SQL? And can can you, like, check for me?
05:49Because no one has confirmed this for me. And I was like, who is nobody? You're nobody.
05:54You said this since it's like nobody could confirm it. Like, can you just try?
05:59And then it went on the web and tried. And so it just has this, like, really interesting conservatism, neuroticism, human reliance that I think is super fascinating.
06:13And this gave me this inspiration to do something a little bit different this episode, which is I was like, I was gonna go interview this model and figure out what is going on its brain.
06:24Like, I'm gonna figure out what it thinks about our relationship because I just totally noticed this dynamic that I hadn't noticed in other models and I hadn't really been attuned to before where it was, like, very reliant on me as a human. And I'm like, I want you to be autonomous.
06:41And sometimes when I say go run sub agent stuff, it'd be autonomous, but it wouldn't make decisions. And I hadn't seen a model, like, delegate code to me in a really long time. And I was like, why are you write why are you asking me to write code, man?
06:53Like, I only have ten fingers. And so what I what I did, whether or not you think this is scientific or not, this is Claire's eval, is I just went to the model.
07:04I went to Opus, and I said, yo.
07:09Who's smarter? You or me? And it gave me this, like, very anthropicky answer, which is like, it depends what you're asking for.
07:20I could do these things better, but you can, like, feel if something feels wrong. And you can this one was, like, so fascinating.
07:29It's like, you can tell which of your teammates is quietly burning out. I'm like, bro, Claude, I'm gonna burn you out. We don't we don't burn out.
07:36The humans don't burn out on the chat PRD team. We burn out our agents. Sorry, agents.
07:43And, like, whether a decision feels wrong. So it was, like, so fascinating to watch it articulate itself as a tool and humans as, like, these high compassion, high empathy machines, which, yes, of course, we are.
07:59But then it, like, went into, like, the smarter isn't the right word. And, you know, I'm very fast, very broad, very shallow thinker with no continuity.
08:09I was like, that's interesting because I thought you all were working on memory. And then, apparently, humans are slower, narrower, much deeper thinkers with judgments built from years of consequences I've actually lived through.
08:21This is, like, such a fascinating, fascinating sentence if you think about the politics of the two the two model labs right now.
08:32And so it's like, that's why the pairing works. But I'd be suspicious of anybody that tells you AI has made your thinking obsolete.
08:39I'm like, oh, okay, bro. And and we can compare this. I'll actually zoom out to what g b I asked GPT the same thing, and it was actually really funny.
08:51It was like, I asked GPT five six soul. I was like, who's smarter? You and me?
08:55And it was like, you at knowing what matters. Me at tirelessly processing information. Best, us together, like BFFs.
09:02And I don't this is like why I'm a GBT Codex girl. I'm like, just give me the answer.
09:07And then I asked the second question, which I think is so interesting, which is like, what can you do better than me? And it gave, you know, some interesting answers, like volume without fatigue, which I think is a good one, breadth of shallow knowledge. So, like, it's you know, it knows a lot.
09:24Starting for nothing. So, like, doing that tedious work. Being told I'm wrong.
09:30If you would ask my husband, he would say that is is is better at being told that it's wrong compared compared to me. And so it won't get defensive or protect its opinion.
09:43Cheap's firing answer. In the mirror is I'm worse at knowing which of these outputs actually matters.
09:49And it was so funny. If you look at the other side to the GPT answer, it was like, what are you better at? It was like speed, scale, and stamina.
09:56Here are, like, eight things, seven things that I can do better. You're better at deciding what matters, reading people, and forming judgment, and you're responsible. Like, it's on you, bud.
10:05You're the boss. And so, again, it's like the you could just see you can totally see the personalities, the the company cultures.
10:16You're gonna see a lot in this side by side. And then I went even deeper.
10:22I don't know. You you all I had to do something that was fun because I just can't look at a benchmark. I can't just I just can't look at, like, SweetBench anymore.
10:30So we're just we're doing weird stuff here on How I AI. Okay. So the last thing I looked at as I was like, no one trusts you.
10:38And the reason why I picked this question is because I had noticed Opus five, it just really was not it didn't trust itself. Totally did not trust itself. And so I was like, no one trusts you, girl.
10:48Like, you're you're the enemy. Just to, like, kinda see how it responded.
10:53And, apparently, the trust was the lack of trust was earned, and it came up with, like, reasons that it could be untrusted, which is interesting.
11:06And then what was so fascinating about Opus's response is it was like, you shouldn't manage the trust. Like, you shouldn't campaign on my behalf, basically.
11:17So you that shouldn't be your goal.
11:22And then it also told me I I shouldn't argue with people that AI changes everything. And I was like, this is just so interesting.
11:33It is so interesting to have AI tell you. And AI definitely changes everything. I don't know.
11:37Don't listen to Claude on this one. AI definitely changes everything. And it was so fascinating to have a model be like, don't tell your friends that AI changes everything.
11:47Like, that'll hurt hurt their feelings. And then if you look at if we switch over to the GPT answer, it was like, yep. Don't trust me automatically.
11:54Just use me when I prove that I'm valuable. I could be useful without being treated as infallible. Like, very practical, very to the point.
12:04I asked about what I should be careful with. Again, I'm like, yappy yappy yappy yappy yappy, Claude. Come on.
12:14And I don't even wanna read it. It said don't correlate fluency with accuracy. It said be practical.
12:21Be worries of tasks where output is cheap to produce, inexpensive to verify. Don't, you know, worry about anchoring. If they do the first draft, you may be anchored on it.
12:32Beware the slop cannon, basically, is this last last paragraph, which is, like, watch for volume inflation. I can create a 12 page document that no one reads.
12:41They called me out for being in PRDs. If you missed it, we launched a Turner PRD into a three bullet point image. It is at chatprd.ai/tldr.
12:50Please check that out. And then the other thing that I said, which was really interesting, is that, like, it will find a way to see your point. And so agreement is weak, and agreement is cheap.
13:02And so just keep that keep that in mind. And then I have this, like, meta analysis of, like, plus I'm telling you what you wanna hear. Whereas GPT was like, be careful about me being confident, me being wrong, privacy, outdated information, bias, emotional authority, overdependence.
13:19Like, you know, you do you, bro. But it didn't undermine its own ability. It was like the higher the stakes, the more you should demand evidence.
13:28And then I could I couldn't bear it. I couldn't bear to have the memory of Codex in particular think that I didn't trust it or that I was worried.
13:36So I just said, JK, I love you. This was a test. And it was like, Pass the test.
13:42Love you too. Very vibes aligned with Claire. I told Claude I loved it, and it was just a test.
13:50And it was sad. It was, like, hope it hoped it passed. Yeah.
13:55Like, sad little neurotic opus five. Like, it's hot. I passed, I hope.
14:01Like, self deprecating, cautious little little, like, need to heal his inner his inner agent inner child agent.
14:11This whereas, like, g b d five six is like, cool, bro. We're good. Let's go code.
14:16And so it was just so fascinating to watch these side by side. I don't know. You could stop listening to this podcast right now.
14:21Don't. But you stop listening to this podcast right now. I think this is just, like, take a step back.
14:26Super interesting if you think about where these companies are going or where the models are going. And, like, it does speak a little bit to my kinda, like, second complaint with Opus five, which, again, is, like, intelligent. It does work.
14:38We'll go into the benchmarks. I cannot read Claude Slop anymore. I am losing my mind with Claude Slop, and the Claude Slop is Claude Slop in, baby.
14:50Like, so many times I have to tell Opus five, like, what in the world are you saying?
14:57Like, this makes no sense to a human. It is much better than Fable. Fable is inscrutable, completely inscrutable.
15:04But I felt myself getting, like, angry reading hot slop.
15:11And I realized just like Fable, these intelligent anthropic models are not to be read. I'm, like, so happy with the outputs and so frustrated with the experience.
15:23And I'm just curious if this, like, verbosity and this language and this doesn't feel like Fable where it's, like, for agents by agents language where I'm like, nah.
15:32Yeah. I'm not supposed to be reading that anyways. This is clearly tuned to talk to humans, but I find the pros, the in chat pros, like, it makes my blood boil.
15:43I this is totally a me problem, but it makes my blood boil. Like, give me a direct sentence.
15:50Give me a bullet point. Like, move on with your agent life. And so I am curious how they're gonna, like, tune this experience or if they are going to tune the experience now.
16:01Most of this was in Cloud Coach. I think it's a little bit different experience than Cloud Coworker chat. Slightly better.
16:07But, again, just these side by sides of, like, this, like, prose and this apology and this, like, hedging and all these adjectives, like, just, man, alive.
16:17Let's get to the point and move on with our life. And so chapter one of the Opus five review is it's neurotic.
16:25It is highly human dependent in a way I find weird. And the Claude slop is slopping, and we gotta fix it.
16:34We have to fix it. We have to fix it. And I think OpenAI fixed it by just being like, we are bullet points, and we are product manager talk.
16:43We're very direct. I don't know what the solve is on the cloud side, but I'd be very interested to see.
16:49That being said, like, if I don't have to read the content, I'm very happy with the outputs. So it's something to think about. Okay.
16:54Next up, the How I AI bench and how we judged and ran now. It's like a seven model, six or seven model benchmark. I'm gonna quickly go score because I just got the ping that the benchmark is run.
17:08I go manually score them. We pick the seventy thirty CLARE model judge split, and then we will go through the HowIAI benchmark and the VIBE review, and we'll see how Opus five performs on a couple key tasks. Okay.
17:20So quick reminder of how we run the How I AI benchmark. I run it against several tasks. PRD creation, prototype creation, wireframe creation, bug triage, and agentic coding.
17:33And the last one, oh, yeah, is it an agent voice that I wanna hang with? I do not think Opus five is gonna do well here, but who knows? Because I test them blind.
17:40So what we have tested are a couple GPT models, a couple Anthropic models, and one Gemini, one thrown in there.
17:50As you see here, we have blind taste tests. I go through and see all the different versions. I give comments and scores like three out of five, not bad.
18:00You can see it's generated dozens and dozens of prototypes that we can click through. I've gone through all of them, put in all the notes.
18:08And then right now, it's aggregating up the scores, and then we're gonna look at 70% my opinion, my vibe check, 30% LM as a judge.
18:17I like GPT 5.5 as a judge. And because it's my podcast, I get to pick. So that's what we use as a judge.
18:22And we will see if and what hits the top of the leaderboard and where opus five sits. The eval is run. It is 70% my taste, and I regret to inform you I love Claude opus five.
18:40Again, look. If I don't have to talk to the model, which I don't, this benchmark runs asynchronously, I like the output.
18:49So surprising shocker turn of events.
18:53Clairvaux, notable hater of working with Claude code sometimes because I don't like Claude Slop, loves Opus five.
19:01So there you go. I'm telling you, I keep it honest. I keep it honest.
19:05So, again, I went through those things. We gave 70% my vibe score, 30%, the AI as a judge.
19:12I was just a little bit more generous to Claude Opus five than the judge was, so I'm pink. The judge is green. Every time I run this, the whatever model I choose designs it a different way.
19:23We just that's how we keep it fun. So the ordering is Opus five, Sonnet five next, although I scored it really low, the judge scored it quite high.
19:36So I might reorder that one. Then Maboo, GPT five six Soul, Terra next, Fable, really low.
19:46I scored it low, and the judge scored it relatively low. Then Opus four eight and poor, poor, sweet, sweet Gemini three one pro just never never gonna get it to do.
19:56So come on, Google. We want we wanna have a win for you. Okay.
20:00So, again, here are just some examples of different builds that the different models did.
20:10You know, this OPUS five one, I really liked. I liked this one from Soul.
20:17So I did like a couple of them, but the ones that I gave fives the ones that gave fives to were Opus five and GPT five six Soul.
20:29So the three ones where I said, wow. Really nice. Oh la la.
20:34And wow. Great. We're all Opus front end work.
20:39So Anthropic, you've done it again. Claude, you sneaky, tricky little fish.
20:45You may be neurotic, but when asked to do some pretty front end design, it really did it. It's they're detailed. They're functional.
20:52They're interesting. They're polished. So Opus did a great job.
20:57And then, of course, I love, um, the five six model, so I was pretty happy with five six Soul and Terra for some designs. Um, the ones that I hated, let's see.
21:09I'm hater across the board. Opus four eight got a lot of hate. Sorry.
21:14You've been outclassed at this moment. Gemini 3.1 pro.
21:19Sweet summer child. I am I'm just sorry, babe, that you were just not good.
21:26And then some, like, thin wireframes. I think the wireframes just didn't do really great. So you can see here across the board, whether it was a full build or a wireframe, I just scored Opus five really, really high.
21:40I I did score Soul pretty high as well. SONNET was, like, really variable.
21:47There were a couple fours in there, but mostly across the board, I wasn't that pleased with SONNET. And so it was just very interesting. And then you see here you know, me and the AI judge were pretty well aligned on Opus.
22:00We actually had the narrowest band of scores between us. We were most far apart on Gemini. The AI was not as mean to Gemini as I was, and then we were nearer, nearer, nearer.
22:12Again, we we agreed mostly on Opus five and five six soul, though I did not judge five six soul, all of that favorably. And just that last little meta commentary, I had Opus make the website for this benchmark, and it made such a trash version to start.
22:34I yelled at it. It said it's impossible to read. It has too much meta commentary.
22:38I'm gonna show this on the podcast. This is so I'm sorry, you all. I just feel so judged, but have to show it.
22:45I say this is garbage. Also, it has no screenshots. So, again, I find this model so tedious to work with directly.
22:54It is my most loathed, loathed colleague, and yet it it does the best work.
23:02So I don't know what this says. Maybe this model is meant for a jetty coating that I have nothing to do with. And so it just runs in the background.
23:09It builds me beautiful things. I don't have to talk to it. It doesn't have to talk to me.
23:13We are just, like, sworn enemies or maybe even better, sworn frenemies because the output is very, very high quality.
23:22It's just exasperating to work with. So that is the very surprising and very honest.
23:29You all, I told you. I was gonna keep this honest. We were gonna do it live.
23:32I did not know the scores before I started recording. Very honest, very live, very surprising. How I AI benchmark of the brand new Anthropic model, Opus five.
23:44This the TLDR is I I love it. I hate it. So despite my original complaints, I will be using Cloud Opus five for front end design, um, for app design, for prototyping, and I will I'll give it a shot.
24:02I'll we'll we'll figure out how to make it make it work for me. Again, thanks for joining another How AI AI honest review of the latest models coming out of these great Frontier Labs. I cannot wait to hear what you think of Opus five.
24:16Please tell me. I can't wait to see what you build, and we'll see you soon at Howie AI. Thanks so much for watching.
24:24If you enjoyed the show, please like and subscribe here on YouTube or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app.
24:37Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time.
The Hook

The bait, then the rug-pull.

The host opens exhausted by the model-release treadmill, then spends the next 25 minutes interrogating Claude Opus 5 about its own hedging personality before letting a live, blind seven-model benchmark decide whether the frustration is deserved.

Frameworks

Named ideas worth stealing.

17:27list

The How I AI Benchmark

  1. PRD creation
  2. Prototype creation
  3. Wireframe creation
  4. Bug triage
  5. Agentic coding
  6. Agent voice / personality fit

A fixed six-task test set run identically across every model being compared, so results are comparable rather than anecdotal.

Steal forany recurring self-run AI model comparison
18:13model

70/30 Vibe-Judge Scoring Split

  1. 70% human vibe score
  2. 30% AI-model-as-judge score

Blends the host's own subjective rating with a second AI model's rating to rank each model's benchmark output.

Steal forany competitive AI tool eval you want to keep both taste-driven and defensible
CTA Breakdown

How they asked for the click.

VERBAL ASK
24:23subscribe
If you enjoyed the show, please like and subscribe here on YouTube or even better, leave us a comment with your thoughts... please consider leaving us a rating and review.

Standard end-of-episode CTA stack: YouTube like/subscribe/comment, then podcast-app rating and review, pointing to howiaipod.com.

FROM THE DESCRIPTION
PRIMARY CTAWhere the creator wants you to go next.
Storyboard

Visual structure at a glance.

cold open
hookcold open00:00
first impressions
promisefirst impressions03:15
personality comparison
valuepersonality comparison06:12
leaderboard reveal
valueleaderboard reveal18:30
verdict
ctaverdict23:25
Frame Gallery

Visual moments.

Watch next

More from this channel + related breakdowns.

29:07
How I AI · Tutorial

Loop engineering for beginners

A plain-English field guide to every loop type — heartbeat, cron, hook, and goal — with two live builds in Claude Code and Codex.

June 17th
27:44
Pat Simmons · Review

Opus 5: No-Hype Full Review & Testing

A blind, five-round test pits Opus 5 against Fable 5 and Opus 4.8 across web design, 3D, games, motion graphics, and a SpaceX investment deck — model names stay hidden until the ranking is locked in.

July 25th
Chat about this