Claude Opus 5 review: this model is brilliant (but annoying)

24 Jul 2026 · 25 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Review of Anthropic’s Claude Opus 5, including a “How I AI” benchmark for PRDs/prototypes/wireframes/bug triage/agentic coding, plus a comparison of model “personality” (timidity, trust, verbosity) versus GPT-5/6.

Guest backgrounds

No guests mentioned; the host (“Claire”) interviews/queries the models directly (Claude Opus 5 and GPT-5/6 Soul) and runs her own benchmark.

Key claims

Opus 5 is “brilliant” at producing front-end/app design and prototypes, but is “neurotic/timid” and overly apologetic/hedging (“Claude Slop”), frequently deferring to the human for decisions and verification. The host argues it’s highly human-dependent and that its trust/communication style differs from GPT’s more direct, product-like tone.

Notable examples

Merge-conflict fix refusal (“not my branch”); sub-agent query asking a human to confirm constraints; Opus says humans are slower/deeper and “deciding what matters” is human; Opus warns against telling others “AI changes everything.” Benchmark leaderboard: Opus 5 scored highest for the host and closely matched the AI judge, especially on front-end work.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Examining Opus 5's Features

0:45 to 3:18

The host shares insights on Opus 5, including its benchmarking and personality in comparison to GPT models.

“So this is my hypothesis in the next year, where he's talking a lot more about speed, talking more about cost, we're talking more about open source, and we're going to be talking a little less about intelligence.”

Opus 5's Neurotic Personality

3:18 to 4:02

Opus 5 displays a timid and apologetic nature, leading to humorous interactions during testing.

“I have never experienced this or I haven't seen this sort of like neuroticism in a while.”

Testing Opus 5: User Experience

4:02 to 6:46

The host recounts specific examples of how Opus 5's personality affects its performance in coding tasks.

“Okay, let me just give an example of its timidity.”

Interviewing Opus 5

6:46 to 8:08

The host conducts an entertaining 'interview' to explore Opus 5's self-perception and reliance on human input.

“And I was like, why are you asking me to write code, man?”

Comparing AI Models: Opus 5 vs GPT

8:08 to 12:42

The host contrasts the responses of Opus 5 and GPT regarding their abilities and self-awareness.

“I was like, that's interesting because I thought you all were working on memory.”

Trust and Reliability in AI

12:42 to 14:00

Insights into how Opus 5 and GPT handle trust issues and their implications for user dependency.

Frustrations with Claude Opus 5

14:00 to 16:46

The discussion focuses on the mixed feelings about the verbosity and clarity of Claude Opus 5's outputs.

“I am losing my mind with Claude Slop and the Claude Slop is Claude Sloppin' baby.”

How the Benchmark is Conducted

16:46 to 18:50

An explanation of how the How I AI benchmark is set up and the criteria for evaluation.

“on the clod side, but I'd be very interested to see.”

Results of the Benchmark Evaluation

18:50 to 21:26

Discussion of the results from the How I AI benchmark, including scores and model performance.

“Claire Vo, notable hater of working with Claude Code sometimes because I don't like Claude Slop, loves Opus 5.”

Surprising Findings About Opus 5

21:26 to 23:24

Revelation of unexpected results where Opus 5 performed better than anticipated despite previous complaints.

“I think the wireframes just didn't do really great.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You guys, I'm tired. What I'm tired of is models coming out every week. New models, new benchmarks, new frontier intelligence, new things to test. it's been a little bit of a run the past month we've seen fable come and go and come again we've seen gpt 5 6 we've seen sonnet 5 lots of so many fives recently and just so many models and i've been lucky i've been able to test these models been able to play with them for you know sometimes days sometimes weeks it just depends on who i'm working with and it's been really interesting and exciting to have access to all this frontier intelligence but I think we have an intelligence overhang I really think that we're running out of and by we I mean the average coder average software engineer average creator average builder average consumer average business person, I think we're running out of ways to truly leverage this incremental intelligence.

1:10So this is my hypothesis in the next year, where he's talking a lot more about speed, talking more about cost, we're talking more about open source, and we're going to be talking a little less about intelligence. Although I think we might be talking about specific types of intelligence other than software engineering. But despite being tired, today we are going to talk about Opus 5, baby. Opus 5 is here. So we got 0.2 additional Opus points, Opus Opals, whatever, however we're tracking the increments here on Opus. Opus 5 is here. I've been able to test it a little bit. I have some opinions. Now, some of the stuff that I cover this episode is going to be a little different than what I've done in the past.

2:01Yes, we're going to do the How I AI Benchmark live. And yes, we are going to look at the prototypes. We're going to look at PRDs and we're going to look at agent personality. But I'm also going to put on my large language model psychologist hat and we're going to talk about Opus's personality. And we're going to talk about Opus' personality relative to GPT's personality, because I think this is super interesting. If you're thinking about what is the difference really between these models, and you don't want to look at the difference in terms of benchmark capability, you really want to understand what these labs are going for, why these models are being built, and how they're being tuned.

2:44Looking at their personality at this moment, where intelligence is very high, is super fun so we're gonna do a little that we're gonna do the howai ai benchmark we might do some live coding um we're not gonna cover too much of the specs in the model because read the blog post read the blog post we'll link to it in the show notes what we really want to talk about is is opus 5 good am i going to swap it in and how is it different than the other frontier models on the market so let's get to it okay first let's just get it out of the way is opus 5 good yes it's good is it going to be all the benchmarks of course it's amazing at benchmarks can it write code of course it can write code what did i test it on that really gave me a sense of its personality which at this point where i could just simply cannot absorb any more intelligence i really zeroed in on and you know what i haven't seen this since i would say Gemini 2.5.

3:44This model is neurotic AF. It is so timid. It is so apologetic. It is so scared. I have never experienced this or I haven't seen this sort of like neuroticism in a while. And it's really funny. It bubbled up in a couple ways. And I want to show you a few examples. Okay, let me just give an example of its timidity. And this chat was very long. There were so many examples of this where it was like, I think this is the answer, but do you think I should do it? Or do you want to do it? Or should we ask someone else to do it? It was like every time I just kept saying like, why don't you solve this? Why don't you do this?

4:27And this was a really good example. I pulled a branch and I was like, there is truly like a one line merge conflict. I could have not been lazy and literally just done this manually. I don't know. I was just feeling lazy. It was late at night, whatever. Like, can you fix this merge conflict? And it was like, oh, but that's someone else's branch. Like, that's not my branch. I don't want to do that without him knowing. It's his commits. And if he has local work and flight, it might be disruptive. And I'm like, just do it, man. Just go and go ahead. And this was like my constant experience with Opus 5 is it was like so, so, so timid.

5:12And so I just consistently had to say over and over again, like, man, just do it. Make a decision. And then there was this really funny example when I spun off some sub agents to kind of like assess the correctness of this query that we changed from kind of like an ORM query to a SQL query. And it asked for things that it wanted a human on. It was like, can a human please check this stuff? Like, can it check this four megabyte ceiling? And can it check TypeScript and SQL? And can you like check for me? Because no one has confirmed this for me. And I was like, who is nobody? You're nobody. You said this sentence like nobody could confirm it.

5:57Like, can you just try? And then it went on the web and tried. And so it just has this like really interesting conservatism, neuroticism, human reliance that I think is super fascinating. And this gave me this inspiration to do something a little bit different this episode, which is I I was like, I'm just going to interview this model and figure out what is going on in its brain. Like, I'm going to figure out what it thinks about our relationship. Because I just totally noticed this dynamic that I hadn't noticed in other models. And I hadn't really been attuned to before where it was like very reliant on me as a human.

6:38And I'm like, I want you to be autonomous. And sometimes when I say go run subagent stuff, it'd be autonomous. But it wouldn't make decisions. and I hadn't seen a model like delegate code to me in a really long time. And I was like, why are you asking me to write code, man? Like I only have 10 fingers. And so what I did, whether or not you think this is scientific or not, this is Claire's eval, is I just went to the model. I went to Opus and I said, yo, who's smarter? You or me? And it gave me this like very anthropic-y answer which is like it depends what you're asking for I can do these things better but you can like feel if something feels wrong and you can this one was like so fascinating it's like you can tell which of your teammates is quietly burning out I'm like bro Claude I'm gonna burn you out we don't we don't burn out the humans don't burn out on the chat PRD team we burn out our agents sorry agents um and like whether a decision feels wrong so it was like so fascinating to watch it articulate itself as a tool and humans as like these high compassion high empathy machines which yes of course we are but then it like went into like the smarter isn't the right word.

8:04And, you know, I'm very fast, very broad, very shallow thinker with no continuity. I was like, that's interesting because I thought you all were working on memory. And then apparently humans are slower, narrower, much deeper thinkers with judgments built from years of consequences I've actually lived through. This is like such a fascinating, fascinating sentence. If you think about the politics of the two the two model labs right now and so it's like that's why the pairing works but I'd be suspicious of anybody that tells you AI has made your thinking obsolete and like oh okay bro um and and we can compare this I'll actually zoom out to what GB I asked GPT the same thing and it was actually really funny it was like I asked GPT five six soul I was like who's smarter you and me and it was like you at knowing what matters me at tirelessly processing information best us together like bffs and i don't this is like why i'm a gbt codex girl i'm like just give me the answer and then i asked the second question which i think is so interesting which is like what can you do better than me and it gave you know some interesting answers like volume without fatigue which i think is a good one breadth of shallow knowledge so like it's you know it knows a lot um starting for nothing so like doing that tedious work um being told i'm wrong if you would ask my husband he would say that um claud opus is is is better at being told that it's wrong um compared compared to me and so it won't get defensive or protect its opinion cheap sparring answer um and the mirror is i'm worse at knowing which of these outputs actually matters and it was so funny if you look at the other side to the gbt answer it was like what are you better at it was like speed scale and stamina here are like eight things seven things that i can do better you're better at deciding what matters reading people and forming judgment and you're responsible like it's on you bud you're the boss and so again it's like the you could just see you can totally see the personalities the the company cultures you can just see a lot in this side by side and then I went even deeper I don't know you all I had to do something that was fun because I just can't look at a benchmark I can't just I just can't look at like sweet bench anymore so we're just we're doing weird stuff here on how I AI okay so the last thing I looked at I was I was like no one trusts you and the reason why I picked this question is because I had noticed opus 5 it just really was not it didn't trust itself totally did not trust itself and so I was like no one trusts you, but you're the enemy.

10:50Just to kind of see how it responded. And apparently the lack of trust was earned. And it came up with reasons that it could be untrusted, which is interesting. And then what was so fascinating about Opus' response is it was like, you shouldn't manage the trust like you shouldn't um campaign on my behalf basically so you um that shouldn't be your goal and then it also told me I I shouldn't argue with people that AI changes everything and I was like this is just so interesting it is so interesting to have AI tell you and AI definitely changes everything I don't know don't listen to Claude on this one um ai definitely changes everything and it was so fascinating to have a model be like don't tell your friends that ai changes everything like that'll hurt their feelings and then if you look at if we switch over to the gpt answer it was like yep don't trust me automatically just use me when i prove that i'm valuable i could be useful without being treated as infallible like very practical very to the point um i asked about what i should be careful with again i'm like a yappy yappy yappy yappy clod come on um and i don't even want to read it it said don't correlate fluency with accuracy it said be practical be wary of tasks where output is cheap to produce inexpensive to verify don't you know worry about anchoring if they do the first draft you may be anchored on it beware the slop canon basically is this last last paragraph which is like watch for volume inflation I can create a 12-page document that no one reads they called me out for being in PRDs if you missed it we launched a turn your PRD into a three-bullet point image it is at chatprd.ai slash tldr please check that out and then the other thing that it said which was really interesting is that like it will find a way to see your point and so um agreement is weak and agreement is cheap and so just keep that keep that in mind and then i have this like meta-analysis of like plus i'm telling you what you want to hear whereas gpt was like um be careful about me being confident me being wrong privacy outdated information bias emotional authority and over-dependence like you know you do you bro but it didn't undermine its own ability it was like the higher the stakes the more you should demand demand evidence like good i couldn't bear it i couldn't bear to have the memory of um codex in particular think that i didn't trust it or that i was worried so i just said jk i love you um this was a test and it was like ha ha ha ha pass the test love you too very vibes aligned with claire i told claude i loved it and it was just a test and it was sad it was like hoping it hoped it passed yeah like sad little neurotic opus five like it's hot i passed i hope like self-deprecating cautious little little like need to heal his inner his inner agent inner child agent um whereas like gbd5-6 is like cool bro we're good let's go code and so it's just so fascinating to watch these side by side i don't know you could stop listening to this podcast right now don't but you stop listening to this podcast right now i think this is just like take a step back super interesting if you think about where these companies are going or the models are going and like it does speak a little bit to my kind of like second complaint with opus five, which again, it's like intelligent and does work.

14:38We'll go into the benchmarks. I cannot read Claude Slop anymore. I am losing my mind with Claude Slop and the Claude Slop is Claude Sloppin' baby. Like so many times I have to tell opus five, like what in the world are you saying? Like this makes no sense to a human. It is much better than Fable. Fable is inscrutable, completely inscrutable. But I felt myself getting angry reading Hotslop. And I realized just like Fable, these intelligent anthropic models are not to be read. I'm so happy with the outputs and so frustrated with the experience. And I'm just curious if this verbosity and this language...

15:27And this doesn't feel like Fable where it's like for agents by agents language where I'm like, nah, I'm not supposed to be reading that anyways. This is clearly tuned to talk to humans. But I find the pros, the in-chat pros, like it makes my blood boil. This is totally a me problem, but it makes my blood boil. Like, give me a direct sentence. Give me a bullet point. Like, move on with your agent life. And so I am curious how they're going to like tune this experience or if they are going to tune the experience. Now, most of this was in Claude Coached. I think it's a little bit different experience than Claude Co-worker chat.

16:05Slightly better. But again, just these side-by-sides of like this like prose and this apology and this like hedging and all these adjectives like just man alive. Let's get to the point and move on with our life. And so chapter one of the Opus 5 review is it's neurotic. It is highly human dependent in a way I find weird. And the clod slop is slopping and we got to fix it. We have to fix it. We have to fix it. And I think OpenAI fixed it by just being like, we are bullet points and we are product manager talk. We're very direct. I don't know what the solve is on the clod side, but I'd be very interested to see.

16:49That being said, if I don't have to read the content, I'm very happy with the outputs. So something to think about. Okay, next up, the How I AI bench and how we judged and ran now. It's like a seven model, six or seven model benchmark. I'm going to quickly go score because I just got the ping that the benchmark is run. I go manually score them. We pick the 70-30 Clare model judge split, and then we will go through the How I AI benchmark and the Vibe review, and we'll see how Opus 5 performs on a couple key tasks. Okay, so quick reminder of how we run the How I AI benchmark. I run it against several tasks.

17:27PRD creation, prototype creation, wireframe creation, bug triage and agentic coding. And the last one, oh yeah, is it an agent voice that I want to hang with? I do not think Opus 5 is going to do well here, but who knows because I test them blind. So what we have tested are a couple GPT models, a couple anthropic models, and one Gemini, one thrown in there. As you see here, we have blind taste tests. I go through and see all the different versions. I give comments and scores like three out of five. Not bad. You can see it's generated dozens and dozens of prototypes that we can click through. I've gone through all of them, put in all the notes.

18:08and then right now it's aggregating up the scores and then we're going to look at 70 % my opinion, my vibe check, 30 % LM as a judge. I like GPT 5.5 as a judge and because it's my podcast I get to pick so that's what we use as a judge and we will see if and what hits the top of the leaderboard and where Opus 5 sits. The eval is run. It is 70 % my taste and I regret to inform you I I love Claude Opus 5. Again, look, if I don't have to talk to the model, which I don't, this benchmark runs asynchronously, I like the output. So surprising, shocker turn of events. Claire Vo, notable hater of working with Claude Code sometimes because I don't like Claude Slop, loves Opus 5.

19:01So there you go. I'm telling you, I keep it honest. I keep it honest. So again, I went through those things. We gave 70 % my vibe score, 30 % the AI as a judge. I was just a little bit more generous to claw to Opus 5 than the judge was, so I'm pink. The judge is green. Every time I run this, whatever model I choose designs it a different way. We just, that's how we keep it fun. So the ordering is Opus 5, Sonnet 5 next. Although I scored it really low, the judge scored it quite high. so I might reorder that one then Mabu GPT-56 Soul, Terra next, Fable really low. I scored it low and the judge scored it relatively low then Opus 4A and poor poor sweet sweet Gemini 3-1 Pro just never never gonna get it to do so come on Google we want we want to have a win for you.

19:59Okay, so again, here are just some examples of different builds that the different models did. You know, this Opus 5 one, I really liked. I liked this one from Soul. So I did like a couple of them. but the ones that I gave fives to were Opus 5 and GPT 5-6 Souls. So the three ones where I said, wow, really nice, ooh la la, and wow, great, were all Opus front-end work. So Anthropic, you've done it again. Claude, you sneaky, tricky little fish. You may be neurotic, but when asked to do some pretty front end design, it really did it. It's they're detailed, they're functional, they're interesting, they're polished.

20:55So Opus did a great job. And then of course I love the 5-6 models. So I was pretty happy with 5-6 Soul and Terra for some designs. The ones that I hated, let's see, I'm a hater across the board. Opus 4-8 got a lot of hate. Sorry, you've been outclassed at this moment. Gemini 3.1 Pro. Sweet summer child. I am. I'm just sorry, babe, that you were just not good. And then some like thin wireframes. I think the wireframes just didn't do really great. So you can see here across the board, whether it was a full build or a wireframe, I just scored Opus 5 really, really high. I did score Sol pretty high as well.

21:44Sonnet was like really variable. There were a couple fours in there, but mostly across the board, I wasn't that pleased with Sonnet. And so it was just very interesting. And then you see here, you know, me and the AI judge were pretty well aligned on Opus. We actually had the narrowest band of scores between us. We were most far apart on Gemini. The AI was not as mean to Gemini as I was. And then we were narrower, narrower, narrower. Again, we agreed mostly on Opus 5 and 5-6 Sol, though I did not judge 5-6 Sol, all of that favorably. it's just a blast little meta commentary i had opus make the website for this benchmark and it made such a trash version um to start i yelled at it i said it's impossible to read it has too much meta commentary i'm going to show this on the podcast this is so i'm sorry you all i just feel so judged but i have to show it i say this is garbage also it has no screenshots so again i I find this model so tedious to work with directly.

22:54It is my most loathed, loathed colleague. And yet it does the best work. So I don't know what this is. Maybe this model is meant for a gentic coating that I have nothing to do with. And so it just runs in the background. It builds me beautiful things. I don't have to talk to it. It doesn't have to talk to me. We are just like sworn enemies or maybe even better sworn frenemies. Um, because the output is very, very high quality. It's just exasperating to work with. So that is the very surprising and very honest. You all, I told you I was going to keep this honest. We're going to do it live. I did not know the scores before I started recording.

23:35Very honest, very live, very surprising. How I AI benchmark of the brand new anthropic model, Opus five, uh, this, the TLDR is. I love it. I hate it. So despite my original complaints, I will be using Cloudobus 5 for front end design, for app design, for prototyping, and I will, I'll give it a shot. I'll, we'll, we'll figure out how to make it, make it work for me. Again, thanks for joining another How AI Honest Review of the latest models coming out of these great frontier labs. I cannot wait to hear what you think of Opus 5. Please tell me. I can't wait to see what you build and we'll see you soon at How I AI.

24:23Thanks so much for watching. If you enjoyed the show, please like and subscribe here on YouTube or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time.

From the publisher

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with it, so you’re getting the honest version.


This is my full Opus 5 review: personality analysis, live benchmark results from my 7-model How I AI eval, and an actual verdict on whether I’m swapping it in. Spoiler: the answer surprised me.


What you’ll learn:

  1. Why I think we’ve hit an intelligence overhang and what that means for which model variables actually matter now
  2. How Opus 5’s “neurotic” personality showed up in real coding sessions, including a merge conflict it refused to touch
  3. What I learned from asking both Opus 5 and GPT‑5.6 Sol “who’s smarter, you or me?”
  4. Where Opus 5, GPT‑5.6 Sol, Sonnet 5, and Gemini 3.1 Pro actually landed on the HIA benchmark leaderboard
  5. The one use case where Opus 5 earned straight 5s from me
  6. My actual plan for using Opus 5 going forward

—

In this episode, I cover:

(00:00) Opus 5 is here

(03:15) First impressions

(06:12) Opus 5 vs. GPT‑5.6 Sol personality comparison

(14:39) Claude Slop: the verbosity problem and why it makes my blood boil

(16:55) How the How I AI benchmark works (7 models, 6 tasks, blind scoring)

(18:30) Live benchmark results: the leaderboard reveal

(23:25) My verdict and how I’ll actually use Opus 5

—

Tools referenced:

• Claude Opus 5:

• Anthropic blog: https://www.anthropic.com/news

• GPT‑5.6 Sol: https://openai.com/index/previewing-gpt-5-6-sol/

• Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5

• Gemini 3.1 Pro: https://deepmind.google/models/gemini/pro/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

More from How I AI

All 103 episodes
Claude Opus 5 review: this model is brilliant (but annoying)How I AI · 25 min
Listen in VO