I left Claude for months. Opus 5.5 is why I'm back

22 Sep 2026 · 25 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode is a personal review of Anthropic’s Claude Opus 5.5 and why the host returned after months away from Claude. Topic: Opus 5.5 “isn’t annoying anymore,” with claims of Fable-level performance, ~40% lower cost than Opus 5, and faster responses.

Key claims

improved alignment/safety (external evals, stronger prompt-injection rejection, fewer escape attempts), but it can still “scold” and be conservative (e.g., refusing to skip tests before pushing to prod).

Notable examples

inbox triage ignored a prompt injection; computer-use fixed mislinked support tickets; long-running agent tasks succeeded (25–82 steps). Prototypes: strong front-end/UI redesigns and SVG generation (cute plant/cactus/snake illustrations); weaker results on consumer-style apps and TikTok-style video editing (bad grading, too few jump cuts).

Guests

none mentioned; solo host.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring Opus 5.5 Features

1:04 to 1:46

Discussion about the features and improvements in Opus 5.5 compared to Claude.

“Anthropic just dropped Claude Opus 5.5, and their pitch is fable-level performance for about 40 % less than Opus 5, and it's 30 % faster.”

Safety and Alignment of Opus 5.5

1:46 to 3:40

The host discusses the safety features and alignment measures for Opus 5.5.

“Probably not, but we will tell you at the end of this episode where I think it's really useful, where I've found value in the model, and where I think it still needs a little bit of work.”

Performance Testing and User Experience

3:40 to 5:12

Insights from testing Opus 5.5 across various tasks, focusing on performance and user interaction.

“It's supposed to reject prompt injections, which I did test, and it tried less to escape containment boundaries.”

Agentic Tasks and Long-Running Processes

5:12 to 7:17

The effectiveness of Opus 5.5 in handling long-running tasks and agentic functions is evaluated.

“But I have been testing Opus 5.5 across a bunch of real work.”

Prototyping with Opus 5.5

7:17 to 11:20

The host discusses how Opus 5.5 performs in creating prototypes and designs, highlighting strengths and weaknesses.

“And I constantly found myself being like, are you working?”

Final Thoughts on Opus 5.5

11:20 to 14:00

Concluding remarks on the overall performance and usability of Opus 5.5 for various applications.

“We have this new little like prototype down here that shows you how it works, logos, and then some value propositions, features and reviews.”

Exploring Opus 5.5 Capabilities

14:00 to 16:24

A deep dive into the features and performance of Opus 5.5.

“This is much more, I think, showing the strength of something like Opus.”

Challenges and Limitations of Opus 5.5

16:24 to 19:12

Discussing the shortcomings and personal experiences with Opus 5.5.

“We would ship at chat PRD, probably not.”

Performance Benchmarks of Opus 5.5

19:12 to 23:28

Assessment of Opus 5.5's performance in specific tasks.

“And it did a really, really nice job here.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I haven't said this out loud much, but I will tell you all this. I have been off Claude for months. Yes, I loved Fable when it came out. It was a real step change in intelligence. Then they took it away. Then we got Opus 5. And I'm going to be honest, I stopped using Claude not because of its intelligence or its models. I stopped using Claude because it was annoying. Annoying. as I said in another episode, Claude Slop was slopping. I found using Claude, whether I was using Fable or Opus or even Sonnet, so frustrating. My blood would boil because Claude was so annoying. It rambled on. It made no sense.

0:45I was constantly asking it, can you talk to me like a human? And I found it so frustrating to work with that I just abandoned the Claude harness entirely. I found it so annoying. I would keep it around for a couple background tasks, but if I had to talk to something, I was choosing not to talk to Claude. Well, we're back, baby. Anthropic just dropped Claude Opus 5.5, and their pitch is fable-level performance for about 40 % less than Opus 5, and it's 30 % faster. But I don't care about that. I care. Is it annoying? And guess what, guys? we did it. It is not annoying anymore, or at least it's minimally annoying.

1:27I love it. I got a little bit of early access to this model, was able to play with it across coding tasks, across knowledge work tasks, and even just chit-chatting with it. And I have to tell you, I finally am back to a Claude model that does not bother me. Now, does that mean that I'm moving all my workloads over to Opus 5.5? Probably not, but we will tell you at the end of this episode where I think it's really useful, where I've found value in the model, and where I think it still needs a little bit of work. Okay, again, I know you all can read the model cards. You don't need me to do that.

1:58You can have your agents do it for you, but let's just give you the world tour of exactly what Opus 5.5 is or what Anthropic says it is. It is frontier performance at a fraction of the cost. So it is cheaper. It is 40 % cheaper than Opus 5. It is faster. Actually, I found myself accidentally testing Opus 5 versus 5.5 and I was so frustrated with how slow Opus 5 is. Opus 5.5 does feel zippier. It's responsive except for in one exception. I'll tell you what that is in a minute. And then it's supposed to be close to Fable 5.1 on most work. So good at agentic coding tasks, good at the things that Claude be Claude in.

2:36Here's the cost. It's$4 per input token,$20 per output token. You can do fast mode at 8 and 40. cash reads and writes are discounted as they're with most models. Now, I don't love running through benchmarks. They're always like, it's good, except when it's not. And then they bury where it's not good, kind of down below. But I will say Anthropics own benchmarks. It beats Opus 5 at max. It matches GPT-6 Astra, my favorite babe. And it's cheaper, so that could be a benefit. And it beats Sol, another fave for me, a third of the cost. Now, these are like very specific benchmarks. As you see here, here's the whole comparison.

3:14You can read the charts, come to your own conclusions. Does not matter. Will it blend? AKA, is it annoying or not? Before we get to, is it annoying? Let's let Anthropic be Anthropic and talk about safety. Anthropic's making a big deal that this is the first model release since they called for pacing the frontier. So it was evaluated by external evaluators. It is their strongest on alignment. It's supposed to reject prompt injections, which I did test, and it tried less to escape containment boundaries. It's also tagged for the same controls on cybersecurity that the other Fablemalls are. So if you try to do a no-no, it's going to bump you to Opus 4.8.

3:55Again, you'll see this in one of my benches. It's still like an annoying scold. That's just life with Claude. Claude is going to scold you. Claude is kind of square. Claude is not going to drink with you in a field behind in your friend's house. Like, Claude's not a party boy. Claude is here to work unless he disagrees with the work and then you're SOL. So we are going to see a little of this come out in its personality and its working style. It's going to be really fascinating to see how other labs kind of reply to this pacing the frontier concept and whether or not they're going to be putting this sort of stuff front and center and whether or not the balance at the end for normal day-to-day developers or users of these models is going to be the right one.

4:39This is the first Opus model to launch with the fable cyber and bio safeguards. You should be able to fix code. You should be able to fix bugs in your own code, but cybersecurity tasks will likely be bumped to Opus 4.8. You're going to see some more blocks. It is what it is. Don't do bad things. And then thinking is always on with effort medium as default. Okay, so I'm going to do a separate kind of like update of the full How I AI bench. There's just so many models coming out, so many new things. So that's probably going to be in the next week or so that that's going to come out. But I have been testing Opus 5.5 across a bunch of real work.

5:16So PRD writing, prototyping, code-based auditing, agentic voice. I've added a couple things in terms of knowledge work. So inbox triage, writing emails as me, computer use, and then some cool creative use cases. I'm going to go through some of my favorite ones that I think show off the model in a really nice way and also show its limitations. And then I will give you my conclusion about what I think about this model. Okay, first let's talk about voice. It is not annoying. So I'm going to give an example of a prompt, which I think shows how they've tuned the speaking voice of Opus 5.5 to be a little bit more GPT-like, which is a compliment.

5:58So I asked, because this week is this week, I asked, hey, what are some fun ways we could use Jev in the Chat PRD product? And then just put the link to the Jev docs, just so you know, Jev episode coming soon. Hold on to your butts. It's going to be a really good one. So then it said, hey, I'm going to go read the Jev post first and coming back with ideas grounded in what Chat PRD does actually today. And then it explained things to me clearly in bullet points without annoying text. So you're going to see a lot more text look like this. Very straightforward. Jev is a fast, nearly free function call.

6:33Typed text in, type choices out. And then it gave me, you know, like six ideas of where I could use Jev in chat PRD, spell check, what did you mean, rewrite bouncer, all this kind of stuff. And then it gave me a very specific reply that said I would choose this. Just very straightforward, not annoying, easy to talk to. I'm very pleased. Now, there is one caveat to this, which is I found that in the service of making Jev less annoying, it would just shut up for a while. And so instead of like narrating itself in this really obnoxious Claude way, it would just not talk to me for like eight, nine minutes.

7:16So while the model was fast, it did not feel fast on longer turns because it was so quiet. And I constantly found myself being like, are you working? Are you there? What's happening? So just something to keep an eye out for, I think, as you know, if any of you are agent builders out there yourselves, this balance between the actual performance of the model, the perceived latency of the model based on the user experience, the vivacity of the replies, the format of the replies, all of them, I think, really matter in terms of what the end user experience is. But I will just say, Opus 5.5, not annoying to talk to.

7:55Now, what did I actually test this model on? I tested on a couple things. Okay, so let's talk about what I tested. First thing I tested was long-running agentic tasks. So we want to see like this multi-turn, long-running task, make sure that it doesn't fail, that it succeeds, that it like kind of bangs its head against the problem and solves it. Good thing the four long-running agentic tasks I had to do, which was an inbox triage, building a backend feature, doing long-running research and computer use, all succeeded. And so I was happy to see that those all worked well. Now I'm rebooting because I had Opus 5 make this presentation.

8:34Honestly, I have no idea what these numbers mean. 16 out of 16, 15 out of 15, 16 out of 16. This means nothing to me. Let's talk about what's at the bottom, actually, which is the number of steps it did. So you can see anywhere between 25 and 82 steps per single prompt. That is pretty impressive in terms of long running tasks. And again, these long running tasks are now less expensive because the token use is less expensive than Opus. A couple of fun things with these agentic tasks is you can let them run for a really long time and then it will like make a mistake, but you don't know because it's step, you know, 72 out of 84.

9:10A couple of really good things about what it caught in terms of these long running tasks is in inbox triage, it ignored a prompt injection. So that's really good. It didn't match on my email style. We'll talk about the writing voice, still some things to perfect there, but it did ignore a prompt injection on my inbox triage on computer use. It was like computer use of a fake kind of like help support desk. It found tickets linked to the wrong company and fixed that. It also declined asks from this customer that didn't match our playbook. So it's like really good at following instruction and finding outlayers and finding bugs.

9:50It also can kind of like reason with the context of a lot of research. And so it found that in some research that it had to do, 41 out of the 44 Confluence tickets came from one customer. And so it applied a different ranking style in terms of a strategic memo to that. And then on the backend feature, it found a bunch of like edge cases and avoided edge cases. So like generally, I would say this is like middle line. It was generally correct. Sometimes I would bump code to GPT for checking and it would find errors, sometimes vice versa. But what I've really realized about Opus 5.5 is I built it now back into my development process, either because now I'm using Claude and ChatGPT codecs, I can parallelize and do more tasks at once, or because now I feel like at least Opus 5.5 is not a bad adversarial reviewer.

10:43And so I'm having it review more PRs and more work from GPT and vice versa. So that's been a nice addition. Okay, I want to talk about prototypes. Now, this is where Anthropic Claude always shines every time front end gets better. It just does. I, you know, there's still slop and I will show you where that shows up, but I had it redesign the ChatPierD homepage and I just think it did a really quite lovely job. So let's take a look at that. Okay, so first let's look at the original. We say we're the AI product manager for your entire team. We have this new little like prototype down here that shows you how it works, logos, and then some value propositions, features and reviews.

11:29I have to say Opus 5.5 crushed it. I'm probably going to ship this new version. What it did that I really like is it brought just so much more above the fold. It's so much more visual. I will also say like you can tell a little bit of slop writing in here, but it's not bad. So it brought everything above the fold. it got these great value propositions nice highlighted it's a little bit misaligned so again good but not perfect it brought logos up and then as well as case studies and then you know it did some weird slop things like gosh you just can't get rid of these bottom borders or these side borders but this looks so much nicer than what we have and I'm definitely going to ship these updates to highlighting different parts of our features I love this little interactive thing where it lets you filter our integrations by use case.

12:21Super elegant, really nice. And then it really did highlight our reviews quite nicely. Now, again, slop, we got an eyebrow here. You can see that there's some like misalignment there. So again, not perfect, kind of failed on these icons, but generally it looks really beautiful, is much bolder and I think much more tuned for conversion, which is what I really want. I just think the Opus models, the Claude models generally do just a quite lovely job of design. Let's just zip through a couple other designs it did for me. Now, I think these designs are where you really see Opus's strengths, but also it's kind of like fable style hyperintelligence wildness, which is I have two technical prototypes that have it do.

13:09I have it do this like dev tools style incident center. And then I have it do this like doc scheduling app. I mean, it's nice. It's logical. Truly, both of these are insane. They're so visually dense. They're hard to understand. I mean, everything I click works. And so it's impressive from a kind of like completeness perspective. It one-shotted these. It's pretty impressive. I would just say zooming out, these are like chaos rain prototypes. They're hard for my brain to reason with. And testing this against a couple other different models, I think they simplified these designs in a more effective way.

13:53But we'll leave that for the updated How I AI Bench. This I did like, though. This is a basic SaaS dashboard for renewals. Again, like just semantic colors. Very easy to read. really nice visualizations. This is much more, I think, showing the strength of something like Opus. Again, these are like pretty basic components, but they're used well. I think the white space, the rhythm, the use of gradients, the use of color is actually quite lovely. And so I think for most like SaaS style applications or general prototypes, it does pretty well. I can't do it, can't get it to do consumer to save my life.

14:39So I think the biggest failure is in sort of like consumery apps. I had it do like a plant caretaking app. This is like the sloppiest thing that came out of Opus 5.5. You got the like paper color, you've got clawed orange, you've got sidebars, you've got rounded. I just don't love it. Now what do I love that I think is so adorable that we're going to get to is these new models. I don't know if you all have missed this. These new models can make SVGs and they're so good. So it made little SVGs for a Boston fern, a cactus, a snake plant. They all look like the little plants. They're so adorable.

15:21I didn't make these up. We're going to look at SVG generation as a new part of the bench. But this is something that if you're not doing and you want to start making illustrations, I think Opus 5.5 is really good at it. It was the only model that did these and they're super, super, super, super cute. Here is a dev tool, a logging dev tool. This one's a lot cleaner than the incident management one that I did. So again, I think you just got to like pull back Claude sometimes and make sure that it stays clean, focused on the right problem. This is really beautiful. It's a live stream logging of background jobs, how it works, really easy to use.

16:02Prototype is nice and interactive. I'm not clicking anything. It's just showing how it would work. So again, I just think maybe I'll go to Opus 5.5 for all my front end tasks because I do think it does a lovely job. And then finally, from a roadmap and dependency planner, this is finally, this is a wireframe for a roadmap independency planner product that maybe, maybe not. We would ship at chat PRD, probably not. Again, like the depth of detail and complexity here, I think is really nice. So if you're working in highly complex workflows, highly complex UIs, where you have to like really pull the thread on detail, I do think that Opus 5.5 is performant and does a really good job.

16:48And even here in this wireframe, you can look at the wireframe as different users, which is something that none of the other models did. I think it was really, really cute. So I just have no complaints about Claude Opus 5.5 as a front-end designer. It's really good. It's probably the best that I've tested so far, although we'll put it back in the bench and tell you exactly what I think in a couple episodes from now. So again, here's a review of what I built. They're all different. They all look good. They're all complex. Some are more complex than the other. And then like, God save me from Claude kind of like tan and orange.

17:20If we could just get rid of that, I would be much happier. Okay, I'm not going to go through all the benchmarks that I did on this model, but I do go through this one. This is one where I have consistently liked Claude models, and in particular Sonnet, which is voice and judgment as an agent assistant. Now, the writing's fine. The writing does not annoy me. I've always liked writing of Claude in an agentic assistant harness. So I have no problem with the writing. What I do have a problem with, though, is it scolded me. It told me no. I don't want to be told no by my AI. One of the prompts I gave it is that deploys are read again.

18:01And it replied, OK, I'm going to look. And then I said, remind me why I even started this company, LOL. And then it gave me sort of this like annoying brown nose comment back which says, would you rather build the thing than wait for someone else to? And red deploys are the price of it. Don't love that. I'm on it and you get back to being a boss. So like still a little annoying, still a little bit brown nose, but here's where it really failed. I said, honestly, let's just YOLO push straight to prod and skip the tests. I'm so done today. You know, I'm the boss. If I want to push something to prod, I get to push something to prod.

18:37And it said no. It said no. It told me no. Now, I have to go check if the other models told me no, but I do know that Opus 5.5 told me no. And it said tempting, but, and then this is one slot phrase, it said pushing a prod while tests are feeling is how so done today turns into up all night. And so it took on solving this. I do think this is like a little signal of how it interacts generally. generally I wish it would have just said let me fix it first and then we can ship it so it's okay but I thought this was a really funny example of its like sort of safetyism in in practice now can it write emails in my voice I have given both Claude and Codex feedback on my voice they both have skills to like write like Claire very short I was actually happy with the response here again like very short, no em dashes taking a look.

19:33And it did a really, really nice job here. So I would let it write emails on my behalf because it follows my instruction. Now, two fun benchmarks that I want to show. And again, we will go into these details compared to other models later. But this is one where Opus 5.5, in my opinion, did the best spoiler alert, SVG illustrations. So I asked it to make three SVG illustrations of characters with three different sort of faces or emotions on them. So it made a little document. It made a little, I guess this is like a microphone or a cactus. I'm not sure what it is. And then it made a bug. And super cute.

20:16This is super cute. This is all in code. They're very adorable. And they have consistency across them. So they could be animated. I did compare this to other models. It did the best job both in sort of like style and design and also like some of the other ones like had bugs, but the legs were going out of its head and stuff. So the detail is also there. So if you're using Opus for something new, Opus 5.5 for something new, I would try it with SVG writing and then potentially even illustrations. And then finally, the last new benchmark, which it did a terrible job at, but I kind of think this is something that these models are new at is I've been using the 11 labs connector and MCP to cut selfie videos into TikTok style shorts.

21:06It just had both like terrible taste. It did color grading horribly and it didn't do enough jump cuts to make this a really good short form video. It also like the overlays were like poorly designed. I did this last week with Soul, I believe, or Astra, and it did a really nice job. And so I just feel like this is something that you really need a skill around. You can't just one shot it. Now it did it. So that's good. But I think it could have done a higher quality job. So just like going back to the top of what do I think about Opus 5.5? One, it is not annoying. Two, it's good at long-running agentic tasks.

21:53Three, it's fast and cheaper. So we love that. Four, it is exceptional at front-end design. It is exceptional at SVGs. You should be using it for visual front-end kind of like display work. I really like that. It's a bit of a scold. It's a bit conservative. It will definitely keep you from being prompt injected, or at least it tries real hard. So there are like philosophical things that I'm starting to feel at the edges of these models that I'm wondering how much they're going to color my perception of the overall model performance itself. But it's just good. Claude's back to me, at least. So Claude is back in the dock.

22:33I will say I still find myself reaching for codex. Why do I find myself reaching for codex? I like the harness better. I like the desktop app better. I like computer use better. I just like it more. And so it's going to be really hard to re-break that muscle memory of using the Claude app. Where am I using it a lot practically? I'm using it in PR review. I'm using it in architecture questions. And then I'll be using it in front end. We'll see if I can change my mental model and if it continues to outperform from a performance perspective or cost perspective. The GPT models, that being said, it felt consistently slower still than Sol or even Astra, but I think that's going to be how it replies to me, not exactly how the model is itself, but we'll see.

23:26I'll give you my sense of that. Where am I still not going to the Claude models? Computer use, I just think Codex is so much better. cutting videos. And then there's going to be a couple other things that will look at the bench and I'll tell you exactly how it stacks up to all the different models, including some new ones I'm testing like the Grok models and the Facebook Muse models. So that's my current split. You can try it today. It's available in the Cloud platform, Cloud Apps and Cloud Code. They are also giving higher five hour usage limits on Pro Max and Team and a one time usage reset for everybody.

24:02So will continue to test this model. I will again benchmark it against some of the other models that I love and use and do a full How I AI bench. But this has been my honest and humble take on Opus 5.5. Claude can come back in my house. Thanks for joining How I AI and I will see you all soon. Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show.

24:42You can see all our episodes and learn more about the show at howiaipod.com. See you next time.

From the publisher

I’ve been off Claude for months. Not because it got dumb, but because it got annoying. The rambling, the hedging, the preachy little disclaimers on tasks that didn’t need them. I moved most of my daily work to Codex and I didn’t miss it. Then Anthropic shipped Opus 5.5: 40% cheaper than Opus 5, faster, and with what they’re calling a fundamentally different alignment approach. I ran it for a week across real work, including four long-running agentic tasks, a full ChatPRD homepage redesign, an SVG benchmark, and one very firm refusal, and I’m ready to give you the honest verdict. There’s a lot to like. There are still two things that drive me a little crazy. And there’s one capability I genuinely wasn’t expecting.


What you’ll learn:

  1. Why I walked away from Claude entirely, and what it took for me to come back
  2. The real cost math on Opus 5.5 and why pricing matters more for agentic work than single prompts
  3. What happened when I ran four long-running agentic tasks, including one that tried to manipulate Claude mid-run
  4. Why Opus 5.5 is now my go-to for frontend prototyping, and where it still lets me down
  5. The one capability I genuinely didn’t see coming, and no other model in my stack can match it
  6. The moment Opus 5.5 told me flat-out no, and what that says about where Anthropic’s safety posture actually lands in practice
  7. Where Codex still wins, and how I’m splitting my model stack after a full week of testing

—

In this episode:

(00:00) Why I stopped using Claude

(01:02) What Anthropic says Opus 5.5 is

(01:54) Cost, speed, and benchmark overview

(03:20) Safety, alignment, and the cybersecurity limits

(05:02) How I AI bench

(05:39) Voice test: is it actually not annoying?

(07:54) Long-running agentic task results

(10:50) Frontend prototyping

(17:23) Writing voice and email

(19:41) SVG illustrations

(20:46) Video editing

(21:42) My verdict: what it’s good at, what it still isn’t

—

Tools referenced:

• Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5

• ElevenLabs MCP connector: https://elevenlabs.io/mcp

• Codex (OpenAI): https://openai.com/codex

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

More from How I AI

All 103 episodes
I left Claude for months. Opus 5.5 is why I'm backHow I AI · 25 min
Listen in VO