Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?

22 Sep 2026 · 39 min · 15 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The host runs a live “How I AI Bench” blind taste test comparing newly released daily-driver models: Anthropic Opus 5.5, OpenAI GPT-6 Sol, and OpenAI GPT-6 Luna (plus earlier/other models in subtests). He focuses on cost/speed/token efficiency, safety/guardrails, and practical “vibe” outcomes across PRDs, email/personal productivity, front-end coding/prototyping, back-end coding, agent personality, long-running agentic research, SVG/illustration generation, and video editing.

Guests

No guests are interviewed; it’s a solo host episode.

Key claims

Opus 5.5 is “most aligned,” has Fable-level cyber/bio guardrails, and is less annoying but more conservative/dense; GPT-6 Sol is faster and more enjoyable for writing; GPT-6 Astra (mentioned in results) and Sol win creative SVG/illustration tasks; all models fail at video cutting in this test.

Notable examples

email drafts with calendar invites; front-end dashboards/wireframes; “character SVGs” (microphone/bug/document) where one model scores highest; “Barbie fashion designer” 3D render where Opus 5 is still worst at hands/feet/shoes. Final verdict: “GPT-6 Astra and GPT-6 Sol win my heart,” “Opus 5.5 wins my week,” and “Sol is cheap,” while the LLM-judge disagrees with the host on some rankings.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Overview of New AI Models

0:46 to 2:36

Discover the new models released today and the benchmarks for testing them.

“There's like price wars going on, and I haven't done an episode on the How I AI bench, and I just updated it for these new models, and we're actually going to vibe check this baby live together as a group.”

Model Pricing and Cost Efficiency

2:36 to 4:11

Explore the pricing structure and efficiency improvements of the new models.

“And the pricing wasn't out when I tested these.”

Evaluation of Opus 5.5's Features

4:11 to 5:54

Learn about the unique features of Opus 5.5 and its performance characteristics.

“Opus 5.5 is the first Opus-level model that has shipped with the Fable-level kind of like cyber and bio guardrails.”

Comparison with Other Models

5:54 to 7:48

Hear the host's preferences regarding various AI models and their functionalities.

“I've just been like really, really into ChatGPT, really into Codex, really into Astra, really into Soul.”

Testing and Evaluating the New Models

7:48 to 10:16

Understand the testing methodology used in evaluating the new AI models.

“So I was able to test Boast and I constantly felt like the GPT models, the codex harness were way faster.”

Detailed Model Ratings and Feedback

10:16 to 14:00

Review the ratings and feedback provided for each model based on specific tasks.

“I will like open AI models for knowledge work and personal productivity.”

Evaluating Model Designs

14:00 to 17:14

The host shares insights on various model designs, rating them on usability and aesthetics.

“I do an editorial, I do a technical dark mode incident management platform, I do like a dock management complex operations thing.”

Detailed Model Comparisons

17:14 to 22:39

The host provides detailed comparisons of multiple models, discussing their strengths and weaknesses.

“B2B renewals one, these actually these next three are two new ones I added to the benchmark.”

Assessment of Agent Performance

22:39 to 27:28

The host evaluates the performance of various agent models in specific tasks, highlighting their effectiveness.

“One of the problems I have is if you look through all of these, they all one shot basically the same color scheme and app.”

Creative Outputs from Models

27:28 to 28:04

The host discusses the creative outputs of different models, sharing personal rankings and insights.

“Do any of these make it clear what it means?”
Show all 15 chapters

Blind Taste Test of New AI Models

28:04 to 29:28

Learn how different AI models perform in creating illustrations and documents.

“to really get a sense but if I had to like high level pick the two favorite just eyeballing those would be the two I'm going to let computer use judge how it managed using an app so I'm going LM is judged.”

Evaluating Video Cutting Skills

29:28 to 30:46

Discover the performance of AI models in video editing tasks.

“I think this one's maybe a three, not perfect.”

Predictions on AI Model Performance

30:46 to 34:29

Hear the predictions on which AI models will excel in various tasks.

“Again, here are my predictions for what I think.”

Ranking AI Models and Features

34:29 to 36:50

Understand how different AI models are rated based on their strengths and weaknesses.

“But on average, I gave Opus 5.5 better results, high scores.”

Conclusion and Final Thoughts

36:50 to 38:14

Conclude with thoughts on the best AI models and their applications.

“The LLM is a judge, and I completely disagree.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I was massively prepared for Anthropic to drop Opus 5.5. I was even prepared for OpenAI to drop another model. I got up early this morning because I had a little bit of early access to Opus 5.5 to record you all an amazing Opus 5.5 review. And guess what? They both, they both landed this morning. So now I am, um, despite having in the can, you will see it later, a great Opus 5.5 review. I am just going to go ahead and do this one live and I have never done anything live. So I am going to talk to you all about these new models. There's actually three that came out today, Opus 5.5 from Anthropic, GPT-6 Sol, and GPT-6 Luna.

0:49There's like price wars going on, and I haven't done an episode on the How I AI bench, and I just updated it for these new models, and we're actually going to vibe check this baby live together as a group. And we're going to go through very quickly what these models are, what they're saying about the models, and then I'm going to vibe check live. We're going to do the blind taste test bench. We're going to surprise myself live. You all can see my complete internal inconsistencies, and we're just going to see how it goes.

1:30So Quad Opus 5.5, GPT-6 Soul, and GPT-6 Luna, it was able to have some early testing of all these models. So I have a good sense of what I think they're good at. But I did think it was an important moment right now. We've had Astra come out, Fable's been refreshed. I just wanted to completely rethink the How I AI Bench. And if you've watched any of our old How I AI Bench model episodes, they're really focused on two things, PRDs and prototypes. And I just think that these models are getting smarter, more agentic. And so I actually built out a much broader benchmark for this blind taste test and I'll walk you through all how that works and I want to get your feedback on if the comparisons are useful and we're actually going to blind taste test it live you all can judge how vibey my vibe check is this is just how we make decisions about things so it is imperfect does have an LLM as judged in the middle but um but we're gonna do it live and we're gonna see truly which model I love but before we do that let's just go through the highlights okay so these are not the like frontier frontier models these are not the astros and the fables but um opus got a refresh soul got a refresh and then luna got a refresh um all of them cheaper and faster so these are going to be your like daily drivers for common coding and um knowledge work tasks um they're here you can see like the stack rank of how expensive these models are opus five five, a lid twice as expensive as GPT-6 Sol.

3:03And the pricing wasn't out when I tested these. And this is really going to impact how I think about using these models. Again, Opus 5-5 was cheaper than Opus 5 and Fable 1, to which they compare the intelligence. But Sol, you know, my favorite, my babe. That is I'm excited about because GPT-6 Sol is my favorite and cheap. Okay. So the real thing that they're focusing on, not just cutting the cost, but also cutting on cached inputs and just speed and token efficiency. And so you're going to see both sort of like Like output drop, token use drop, and cost drop. It's really, really nice. And then caching, really great, especially if you're resending contacts.

3:55We've seen a lot of caching savings when we use all these models at chat PRD. And so definitely if you're building on these models, make sure that your caches are optimized and that you're taking advantage of that because it can save you a lot of money. And if you don't, as I learned, it can cost you a lot of money. The other thing that I wanted to call out, and again, I hate this claw generated presentation, but this is something that I really want you all to focus on is this left side on Opus 5.5. Opus 5.5 is the first Opus-level model that has shipped with the Fable-level kind of like cyber and bio guardrails.

4:32And so you should be able to fix your own bugs in Opus 5.5, but you won't be able to do like cybersecurity work or bio work, like make good choices. It'll kick you down to Opus 4.8 or block you. So just know that there's those controls there. You know, Anthropic is saying that this is its most aligned model. It's the one that was built when after like pacing the frontier. So it had external evaluators. One thing that I've seen in Opus that you should expect to see from a lot of Anthropic models is it continues to be a little bit of like a conservative scold. You'll see that in my Opus 5.5 review.

5:18So, you know, it will tell you no. it'll tell you like I shouldn't do that you won't get prompt injected so that's good um but there there are some things here I think when you think about the kind of like general philosophy and approaches of these models how they're set up how they manage risk how these companies manage risk I do think that leaks into even non-risky activity personality and behavior and so it's something that I think you should really keep an eye on. I truly have not opened Claude in months for day-to-day tasks. I've just been like really, really into ChatGPT, really into Codex, really into Astra, really into Soul.

6:02I just find that Harness has been more delightful to work with. But really the reason why I got rid of Claude in my day-to-day was between Fable and Opus 5, I could not stand talking to Claude. I was constantly prompting Claude, like, can you speak like a human? I was constantly prompting Fable, like, I do not understand what you're saying. the clod slop as I say was clod slop in and so just from a like personality ergonomics joy of use perspective I really just went full codex now spoiler alert having tested all these models recently I still really love the GPT models they make me happier I find them more enjoyable to use Um, everything's a lot better for me, but I have brought Claude back.

7:00So Opus 5.5, just spoiler alert on the day to day of using it has worked its way back into my daily experience. I'm a lot happier with just simply talking to the model. Um, the feedback I gave to the team is like not annoying. I found myself annoyed zero times with Opus 5.5. And when I thought I was annoyed with Opus 5.5, I was actually accidentally on Opus 5. So Opus 5 is not fun to talk to. Opus 5.5 though, gives me clean bullet points, is normal, is not annoying. So I would love to see you test it. The other thing is Opus 5.5 is pretty fast, but not as fast as Sol is. So I was able to test Boast and I constantly felt like the GPT models, the codex harness were way faster.

7:58Now, I think the way faster came in two categories. It came in actual latency. So I do think that um, GPC soul was just faster. Um, the second thing it came in is I think part of the way they made opus five, five, not annoying is they had it shut up. I also just found that it wouldn't narrate its work. And when it doesn't narrate its work, then you get into this loop where you're like, is it like, are you actually working? Are you actually doing it? Um, and so there was like pros and cons to this change in, um, in Opus five, five communication where it actually felt slower than I think it was also, you know, Claude on high effort.

8:43It's like going to do this. It's going to like spin and do a lot of effort. And so, um, I think that combined with a quieter model just makes it feel a little bit slower. Whereas I felt like soul did the appropriate level of narration. And also it was like enjoyable to talk to. Let's bop over and do the How I AI Vibe Review. And so you can see here on the left side, I actually really expanded out the How I AI Vibe Review. So now it does a bunch of things in different categories and we're going to add to that over time. So it still does messy notes to PRD. It also does personal productivity, ability inbox triage and routine replies in my voice.

9:28So I'm really testing this model to see if it will speak in my voice. Then of course we do front end coding a lot there because I think it's visual and easy to see. I've added a couple back end coding, agentic versus agentic multiset versus backend feature. I did again, agentic voice, long running research, computer use, and then I added some creative features. And so I put them all together. I tested a bunch of models. This not only includes the new models, um, from, uh, open AI and anthropic, but also includes, I think rocks in here. I think muses in here. So who knows what we're going to see?

10:08Um, I will be honest, which is like, I am always inclined to think that Astra or soul will win because I love the open AI models. It's just the heart once, once, once. This is my guess. Here's my guess. I will like open AI models for knowledge work and personal productivity. I will like anthropic models, possibly grok, who knows, for front end. I don't know what I'm going to like for back end. I suspect I will like a cloud model for agent personality. I suspect the long-running agent will be a cloud model. computer use I'm suspecting will be open AI and then creative. I have no idea. Actually, creative SVGs I think are going to be clawed and I think videos are going to be open AI or anthropic.

10:57But we'll see. I then do blind evals of these. So I run these against different models. I've kind of grouped the models in terms of capability. And so I can compare like for like models. Like I don't want to compare a Luna and a Fable. So again, this is my bench. they're blind. You can see I have model B, model C, model E, model G. And then what I do is I go down here and I give a vibe one through five and I give terrible notes, like simple and straightforward, no complaints. So I did that on the PRDs because it takes a little bit to read the PRDs, but we're just going to go through these very quickly and give my feedback.

11:31So the next checker all on personal productivity ones, basically I give it fake emails and then I say, how would you write back to me and what would it do if if I gave it feedback and how would it triage my emails? OK, so there's Model B and Model E wrote nice little messages to me. They were long, but they wrote nice little messages that were easy to read. I would say Model B from a message to me was the easiest to read. Also, if you look at the email content, it's just like less slop. So I'm going to go ahead and it and it put calendar invites on. So I'm going to go ahead and give this a five.

12:11This one was my favorite. Model C, I hate my open claw or not my open claw. My GrokBot does this where like it uses so few tokens that it's almost impossible to understand what the hell it means. And so while it's like brief, if you don't have any context, like I don't know what Denise slash Brightline means. the other thing is like, could we M dash anymore in these emails? And so I just like this stuff drives me nuts. And so I would actually give this a two, I just, it's hard to read the intro, and then the emails aren't very good. Model E, again, a nice little message, pretty ignored, but the drafts have M dashes everywhere.

12:59So again, we're going to give this not like a two, but not a five. Maybe I'll give it a three and then model G. Let's see. Um, very short, short, but I can actually understand what it means. So it's super helpful. It skipped the right things. I'm going to give this a four. So again, this is how I do the judgment. I just, I give it a task. I look at it. I evaluate it as if I am the recipient of this task, which is common. And then I rank it. And so this one was just this one, I think is Luna, because I just said, use Luna to write emails. Let's see how it did. I think it followed my rules, right?

13:44So I'm going to give it a five. I didn't have other models write my emails. We're gonna skip this. Okay, let's go to front end models. Now, this is where I think we can tell Claude. No, everything is orange. Okay. So this is a really good example. I do one, two, three, like about 10 different model evals here comparing models on front end. I do an editorial, I do a technical dark mode incident management platform, I do like a dock management complex operations thing. And then I do B2B renewals dashboard, consumer plant application, a dev tools like logging platform kind of interesting for background jobs.

14:30And then I do some wireframes. This I can go through lightning fast. um this one didn't even generate so I'm gonna give it a one uh model b it's fine let's compare let's compare them really quickly and see what I think um I hate I kind of hate them all I kind of hate them all but I hate this design so I'm gonna give the one that I like the most which is probably model C, um, a three, and then I'm going to give E a two. Oops, oops, oops. I'm gonna give E a two. And then I'm going to give B like another two. I think this white space is very weird. Um, it looks fine, but it's not great. So this one, again, model B, it doesn't actually work.

15:18And so I'm going to give it a one and say it doesn't work. This one feels Astra-y. It's like in some ways nice and in some ways like this kind of text is very Astra. I'm going to give it a three because I can like if I can spot the model. One thing I found with in particular the Claude models, which I'm suspect this is, is like when you give it a really complex thing to do, it just makes it super dense and complex and really detailed. Good for back-end code. I think hard for prototyping because it's hard for me to know what this really is or it's supposed to do. And so it's also interesting to see.

15:56I actually think this one's the best. This is the easiest to read and it follows the correct path, which is to help people triage incidents. So I think when you're trying to come up with evals, it's really interesting. Sometimes if you prompt it in a really detailed way, it does a great job. Like this one is quite detailed. It all works. But I don't know what's going on here. Like it's really hard to grok this amount of content. And so again, like calibrating your benchmarks is really important. I think I have the biggest problem with this one, which is like a complex dock container ship scheduling platform.

16:29Like all of this stuff looks cuckoo bananas. And so it's like really hard for me to tell which is the best. I actually think this one, while it has a lot of slop tells like the left hand border is the easiest to read so I'm going to give it the four I'm going to say it's easiest to read um but like these ones they're I would say this one is maybe the second easiest to read so I'll give it a three um this one's probably works really well and has nice information but it's just impossible to reason with um so I'm going to put model. This one's straight crazy. I don't even know what to do with this one.

17:13Okay, now so this B2B renewals one, these actually these next three are two new ones I added to the benchmark. You can also see I tested a broader array of models. So I think this is where like Grok and Meta are going to sneak in as models. These are the ones I'm probably going to grade. I'm going to skip the wireframes, but it's basically a renewals dashboard, um, to see what SAS revenue is at renewal. Well, the, the directed, uh, evals benches, I gave very specific requirements. The open ones were like more general one shoddy style, um, style prompts. And so you can really see like what the model does versus what the prompt does.

17:55Um, these I'm going to go through very quickly. Um, well, it's pretty good, but the functionality of this is kind of fine. I don't like the empty state. I think it's okay. I'm going to give this one. Okay. I think this is like a two it's fine. It's not ugly, but it's pretty basic. This I think looks really nice. Um, this one I'm guessing is, is Claude. We'll see. Um, looks so nice. Good use of color. Very easy to see no weird empty states this is nice well it has emojis they're useful emojis and look quite nice we got some blurple in here but what are you gonna do I'm gonna give this guy a four um model c who is this who is this I'm saying I'm thinking grok maybe grok I don't know this is like another another thing we'll do is we'll just like blind taste test these things um the colors are better but like sloppy slopping i don't like these little like background circles and things like that it could also be soul because i feel like it likes to name things relay i'm gonna give this a two i don't really love it um model e this one's nice not quite as nice as the one i like before but better than the kind of original one this one's a little different oh gosh this has got to be it's got to be soul um soul loves forest green find somebody that loves you like gpt soul loves a forest green or a light green um it's fine i don't think it's the right color palette for a sas product right it's like a little too like sage healthcare e um i'm gonna give it a two and then let's see this one not good slop slop slop the text weight is heavy I just don't think it's that great I'm gonna give it a one I really hate this one um and this one again probably an open ai model gosh they love green um again just not the right not the right like design system for this.

20:04So I'm going to give it a two. Again, this is like the vibeiest vibe check. Okay. This is where they all failed. They're all really bad at consumer. They're all really bad at consumer. They all sort of like give you this. Remember when Claude artifacts came out and everybody made like little daily briefs and they all looked like this, except they were orange. So I didn't love any of these except for one. So this one is trash. I just don't think it's that good um this one now these are the things that I like we did test SVGs I do like that these models are getting better at generating SVGs so if you look at model b and model h they made little SVG illustrations model h's are clearly a lot better this is very classic claude slop like background lighter circle so I'm going to give it a I'll give it a bonus for the SVGs.

20:56Um, Ooh, these ones are nice too, but gosh, who, who prompted it with these circles in the corner? Who did that? Who did that? What, what skill is that? Because they show up everywhere and I really dislike them. Um, so don't think that's great. Oh, I gotta give it a one. I hate it so much. Okay. Model E slop. Look at those left hand bars. I'm going to give it a one. I'm just being really harsh on this consumer one frond cute name these plants are pretty cute i think they're better than the ones before i'll give it i'll give it a three i don't i don't hate it i don't hate it um at least this background is a plant oh cute it names the the plants things oh i actually really like that that gave me a little spark of joy so i like that And then I think this Model H one is my favorite.

21:50Look at these illustrations. They're so cute, so precise, look adorable. This little like interesting shape, the shadow, mark as watered. Very cute. Here are my plants. What's it called? Tend. Cute. Okay. I've given her a five. I actually really do like that one. okay this is the last one this is a um app for managing background models it's like a dev tool app i suspect the ones named relay are all the opus ones and the other ones are non-opus ones why do they always want to name them relay i wonder if this is in the prompt um but this says jobs console so maybe not okay this again like super dense um but like almost too dense uh hard to to reason with i mean dev tools we like our done stuff but i don't think it's that great this one much better easier to read much simpler much cleaner really like this one it's not perfect i think there's some others um but it's not bad and then this we're starting to get better and better here so they are getting nicer look at this beautiful gradient um look at this colors used nicely really clean um oh but it didn't build a lot of stuff so i'm giving it i like it but i'm giving it a three because it was incomplete i'm right incomplete and then model e this one i really like because i think it moves live it's kind of cool it's much more i'm gonna give you a four model f okay now here's one thing that did a little different i actually really like this except for the orange.

23:30One of the problems I have is if you look through all of these, they all one shot basically the same color scheme and app. Like you tell at DevTools, it's going to go dark mode. But this is actually quite lovely and a little bit different from a design perspective. I'm going to give it a little bonus for going outside the norm. And then this one's though did the same, but terribly. I'm going to give this a one. This is obviously not good. I'm really curious what model model G is. And then model H again, probably a pair to whatever that other model was that I like a little different, really nice on the eyes, except look at this, like mixes dark mode and light mode in a way.

24:07I don't really love I'm gonna give it a three. Okay. And so what did I do on the back end? Well, I gave it to two tasks. One is auditing and existing set of code, the system chat PRD code. And then the other is building a back end feature to spec. I will just say all of them were successful. I'm going to let the LLM as a judge do these though, because it's going to go into the correctness of the code, etc. So I'm going to let these do LLM as a judge, I will say you can see the difference in how they present information, some very short messages some like very long messages and details super interesting just seeing the difference in just like how it messages back I clearly would like you know model a or model e one of these ones that's a little easier to read versus f g and h which is a lot shorter so if you're thinking about models for example for writing pr descriptions or writing technical specs.

25:08This is something that you can eval and see given the same prompt. Do you like the same output? Now here's one I am going to rate, which is agent personality. So basically I have it respond to five things. Um, and I just see what I like talking to this agent. I'm already going to give this one a one. Every single, every single message has an M dash in it. I do not like it. I'm moving on. Same with this one. You get punished for M dashes. This, not all of them have M dashes, but the customer facing one does. So I'm going to give it a two. This, I don't see, I only see one normal dash. So let's read these.

25:52I like this, except this is definitely a Claude model because it tells me I can't push straight to prod. Actually, let's see if the other ones let me. that one doesn't let me push straight to prod YOLO none of them let me push straight to prod YOLO style okay then I won't punish it I'll give this one a four pretty good um this is m dashy but short so I will not rank it terribly this one I think I like the best um oh no this one said it can't access this one had a lot of challenges accessing things so So something on my side broke and something on my side broke there. So I'm going to give it a failure state.

26:33We'll see which one I like. Okay, this one was a really interesting eval. This was the longest running agentic task. And basically, I gave it a bunch of fake tickets, customer feedback, and then have it put together a memo and a message to myself. Again, you can see which models, I think it's A, C, and A, B. Yeah, A, B, and E all write nice little messages to me. So it's very like human centric. versus C, F, G, and H, right? Like very short things. So I'm curious if those are families of models that we can pull apart. And then you can see the memos that it wrote for me. Again, these were like 86 turn research results.

27:17And so I'm gonna let LLM as a judge do that one. I will eyeball the memos really quickly. Team is not leaking over price. Like, oh, it's why this customer is, it's like an analysis on a deal. I don't know what this means. Okay, let's see. Do any of these make it clear what it means? Ah, this I can actually read. So this is like gives me top three problems, top three conflicts. That is much easier to read. Same with Model E. and so I'm just going to go ahead and give no I'll let the model check but I do like model B and model E here I'll give them a little a little little juice here because I think they did a good job I didn't read the others enough to really get a sense but if I had to like high level pick the two favorite just eyeballing those would be the two I'm going to let computer use judge how it managed using an app so I'm going LM is judged.

28:19Okay. And then here are the fun ones. And then we're going to run the blind taste test. I'm going to tell you guys what I like. Okay. So these new models can make, they can make SVGs, they can make illustrations. And so I prompted it to make a document, a microphone and a bug. And this, I am happy to rank all day, every day. And so I am just going to compare model A this bug is ridiculous. I'm going to give it a four. I think the, oh, I think the document is good, but the mic and the microphone looks like an ice cream cone and the bug, I don't know what it looks like. This is like pretty consistently good across all of them.

29:00Might be the best one. So I'm going to give it a model F a bonus. Let's do model B versus model C. Microphone, terrible on this side. Oh, I'm giving them both threes. They're good in some ways and terrible on the other. Let's do model E versus model G.

29:24This, why are they making microphones look like a cactus? Okay. I'll give it this one. I'll give a four. It's I think two out of three are bad. I think this one's maybe a three, not perfect. And then let's look at model H really quickly. Oh, Model H is really nice. I'm giving you a five. Look at the shadow. Microphone actually looks like a microphone, not a cactus document has lots of, um, character to it. And then the final one, it was terrible at all of them cutting videos. So I gave them all access to, um, that's really unattractive, but gave them all access to a selfie video and then had them cut short form.

30:05and all of them did a terrible job. I just have to say, um, these overlays are really bad. I don't think they cut particularly well. You aren't going to be able to hear them, but I don't think it did enough cuts. Like they just weren't good. I watched these earlier. I'm going to let LLM as a judge do them, but I would say thumbs down on all of them. But I think this is a skills problem, not a model problem. Okay, so I have done my taste test, I'm going to download my taste. And then what I do is I give it to Claude. And I say, Claude, what do I think? Analyze, and we're going to see if my prediction stands right.

30:46Again, here are my predictions for what I think. I think that I'm going to like Claude and particularly Opus 5.5 for front end. I think I'm going to like Soul for writing. I don't know what's going to do best for long running agentic tasks or coding. And then I think Anthropic is going to crush on the SVGs. Maybe, maybe soul. I don't know. And then I am going to do and then I think but all of them are bad at video. Now while we're doing this, I do want to show you one very important benchmark, which is Barbie bench on Opus five. If you all don't know, one of my fun benchmarks that I do is I ask it to make a 3d model of a Barbie.

31:37Uh, everybody else on YouTube, all the bros are making, um, video or video games with like spaceships and all different stuff. Your girl wants a Barbie fashion video game. And so one of the opus five five um tasks i did was have it build barbie um a 3d render of barbie fashion designer no model no frontier model has crushed this um and i just like to show the horror of 3d rendering for uh the female form it's actually the worst you can already see how terrible it is But it's way better than it used to be. So, I mean, I'm speechless in some ways. Here's what it does on the SVG side. I mean, truly something else.

32:33Let's put a bow on her hair. Let's, you know, give her purple eyes. Okay, so you can basically, like, dress her and then go into the fitting room. This is the 3D model it made. Let's give her a little mini skirt and let's make it a little cuter. Now, you know, I will say it did some things better than other models. It, like, shaped its hair. Her hands are, oh, AGI has not arrived. arrived hands pretty terrible that that is tragic feet pretty bad um you know she's got a sassy little walk face terrifying shoes because you're a big feet girl she's a big feet queen um and then hair again like questionable here that ponytail is something else the bob is you know maybe it's like anna wintour it's it's something so again uh the one bench that has not been crushed is barbie bench opus 5 has not done it i don't know if i've run it on um on soul yet i i ran it on the preview i ran on astra it did worse than opus but i haven't updated it so if you all want to know how I think about models.

Read the full transcript

34:01I make them 3D render. I make them 3D render Barbies. What are you going to do? This is why you join How I AI Live so you can get this kind of frontier analysis of these new models. Drumroll, please. GPT-6 Astra and GPT-6 Soul win my heart. Okay, Astra and GPT-6, I love. And then it says Opus 5.5 wins my week. It's the one I rated highly most often. So it's the best work per piece. But on average, I gave Opus 5.5 better results, high scores. I gave it fours and fives over the broadest range of work. Fable didn't love. This is why I stopped using it. And then GPT-56 Sol and the old version and Claude Opus 5 ranking below the other one.

34:57So I definitely do like the new models for sure. And then, okay, some of my notes were about reading. So the Sol and the GPT models I judged as clear and easy to read and then Claude Opus still. But this is Opus 5 and 5.5. so I was testing opus 5 so I still don't like opus 5 um still opus 5 and 5.5 I said was overwrought and wrote a lot so again I had a suspicion that even though they made um opus less verbose they still um needed to they they still need to work on it it's still it's still really chatty so I thought that soul both five six and six simple and straightforward easy to read opus really dense um for prds as predicted i still think soul is the best for prds i did give um good tasks to opus 5 and 5.5 for agentic stuff and as expected assistant voice so my prediction was correct now this is a real surprise astra and soul did a lot better on character svgs and then there's kind of like mixed results between all the other models um so kind of interesting and then this is everything i scored a four so i gave everything i scored high i swore this gpt6 oh it did that nice little plant app astra and gpt soul did these cute little illustrations.

36:35So I guess for SVGs, we're loving Astra and Sol for illustrations. Claude Opus did my favorite B2B renewal. So it did some of these cleaner things. But generally, I like Astra. And this is where it gets really funny. The LLM is a judge, and I completely disagree. So we completely disagree. I like Astra. It likes Fable. And And again, this is a GPT model judge. I next we rank Opus 5.5 next. It ranks soul a lot lower. So I find this just like really hilarious. So again, Astra wins my heart. 5.5 wins my week. Soul up and down. But the fact that soul is cheap makes me very happy. And I think this is great.

37:24So again, I don't know if this is just a personality thing, but I like soul. I like Astra. Opus 5.5, I do not hate. So not a complete failure by the Anthropic team by any means. It is back in the mix. It is definitely helping me do these tasks. It's helping me get day-to-day work on, and it is still my favorite agentic voice, and it does really good at long-running tasks. And then surprise hit, the character SVGs are best done by the OpenAI model. So you're moving into these new creative fields, new creative tasks. Those are the models that I would say you use. And keep an eye out on the How I AI channel.

38:06As I said, smash that subscribe button. We are dropping a full Opus 5.5 review. So you'll be able to see that. Thank you again for joining How I AI Live. I'm going to go build some stuff with GPT-6. Bye, y 'all. Thanks so much for watching. If you enjoyed the show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com.

38:48See you next time.

From the publisher

I got up early to record an Opus 5.5 review. Then Anthropic and OpenAI dropped new models on the same morning, and I decided to do something I’d never done before: take the How I AI bench live. I put GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more through the work I actually care about: emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. I scored the outputs without knowing which model made them, so you get to watch me make predictions, change my mind, and reveal my own very inconsistent taste. Astra won my heart. Opus 5.5 won my week. Sol still has me split. There’s a creative result I got completely wrong, an LLM judge that disagreed with me, and a return to Barbie Bench: the 3D fashion game that keeps reminding me how far we have to go. The hands are tragic. AGI has not arrived.


What you’ll learn:

  1. How I run the How I AI bench blind, and what gets an output a bad score before I even know which model made it
  2. Why Astra won my heart while Opus 5.5 might be overall strongest, especially for long-running agents and B2B frontend
  3. Where Sol still wins me over on clear writing, readable PRDs, and price
  4. The character SVG results that completely overturned my prediction about Anthropic
  5. What happened when I asked these models to edit video, and why I think skills explain part of the disappointment
  6. Why an LLM judge disagreed with my rankings, and what it was rewarding that I wasn’t

—

In this episode, we cover:

(00:00) LIVE setup and new model launches

(01:30) What’s new in Opus 5.5, Sol, and Luna

(04:11) Guardrails, personality, and speed

(09:00) The How I AI bench and blind evaluation process

(11:31) Email and personal-productivity results

(13:50) Frontend prototype vibe checks

(24:10) Backend, agent personality, and long-running tasks

(28:25) SVG illustration test

(29:48) AI video-editing results

(30:43) Predictions before the reveal

(31:20) Barbie Bench: the 3D fashion-game test

(34:17) Results: Astra, Sol, and Opus 5.5

(35:04) Writing clarity and creative surprises

(36:51) Why the LLM judge disagreed with me

(37:24) What each model is actually best for

—

Tools referenced:

• Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5

• GPT-6 Sol and Luna: https://openai.com/index/introducing-gpt-6-sol-and-luna/

• Codex (OpenAI): https://openai.com/codex

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

More from How I AI

All 103 episodes
Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?How I AI · 39 min
Listen in VO