GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark

9 Jul 2026 · 37 min · 17 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The host reviews OpenAI’s GPT-5.6 family (Sol, Terra, Luna) using their “How I AI Vibe Review” benchmark (PRD writing, prototyping/wireframes, coding/debug, and agentic “talk to me like a human” voice). They compare GPT-5.6 Sol against Anthropic’s Claude Fable and discuss pricing, rollout limits, and security/safeguards.

Key claims

GPT-5.6 Sol is the best overall for practical product work; Terra is best for crisp PRDs; Luna is for cheaper high-volume work; Fable is strong technically but harder to collaborate with due to overly technical/pedantic communication and “cringe” agentic voice.

Notable examples

Sol produced more unique, functional full-fidelity prototypes (e.g., a dense “doc scheduler” dashboard) and cleaner redesigns of a “slop-adjacent” editorial page; it also generated a gamified homework tracking app with parent controls and a “parent HQ.” Sol also excelled at video clipping for social (CapCut workflow) and browser automation (Codex + Chrome), e.g., replying to hundreds of LinkedIn messages.

Guests

None mentioned; the episode is a solo review by the host.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Overview of GPT-5.6 Models

0:46 to 1:09

Introduction to the various versions of the new GPT-5.6 models and their importance.

“Now, we're not just relying on my own opinion.”

Deep Dive into Model Versions

1:10 to 2:18

An analysis of the three new models: Sol, Terra, and Luna.

“Okay, you all can read these blog posts, so I'm not going to go into too much depth about the models and the benchmarks.”

Pricing and Subscription Insights

2:19 to 3:25

Discussion on the pricing of Sol compared to Fable and subscription details.

“So it's$5 per million input tokens,$30 per million output tokens.”

Benchmarks and Evaluations

3:26 to 4:28

Explaining the benchmarks used to evaluate the models, including performance metrics.

“All I will say is it is the brand new state of the art model from OpenAI.”

How the Evaluation Works

4:29 to 7:20

Insight into the host's evaluation process and criteria for measuring model performance.

“It tests the ability to generate good PRDs.”

Results of the Model Evaluations

7:21 to 9:23

Comparison results of different models based on various tasks and preferences.

“I've decided I like my own taste better.”

Model Design Aesthetics

9:24 to 11:15

Discussion on the design aesthetics of Sol versus Fable and other models.

“Maybe I like a basic, straightforward PRD.”

Final Thoughts and Recommendations

11:16 to 14:00

Concluding thoughts on the models and recommendations based on the evaluations.

“And so I will say I have been happy to extract myself out of Claude Slop, out of like Blurple Slop into more interestingly designed websites.”

Comparing Design Functionality of Sol and Fable

14:00 to 17:25

Learn how the designs of Sol and Fable differ in functionality and user experience.

“with great visual hierarchy, semantic color, and this thing was functional.”

Evaluating Writing and Communication Styles

17:25 to 22:08

Discover the differences in writing quality and communication clarity between Sol and Fable.

“It is forest green, but you will see a lot of green.”
Show all 17 chapters

Prototyping and Gamification with Sol

22:08 to 27:45

Explore how Sol excels in creating prototypes and gamified systems for practical applications.

“The communication is clear, and it's less pedantic.”

Precision vs. Intuition in Product Development

27:45 to 28:00

Understand the importance of intuition over precision in successful product development.

“But it's a lot better than what I've seen kind of one shot out of other models.”

Precision vs. Product Development

28:00 to 28:33

Explore why precision in AI models can hinder product development.

“Let me talk about another thing where I think GPT-56 Sol and its family does a lot better than Fable.”

Experiences with Prototyping Tools

28:33 to 29:31

Hear about the differences in experiences using GPT-5.6 and Fable in prototyping.

“Like you literally, especially when working with AI, cannot be precisely deterministic when building a great product.”

Unlocking Creativity in AI

29:31 to 30:50

Discuss how GPT-5.6 unlocked creativity in generating resources compared to Fable.

“And when I was having Fable working on this, it did a lot of the like technical heavy lifting.”

Use Cases for GPT-5.6

30:50 to 31:59

Discover some practical use cases for GPT-5.6 in various applications.

“Very similar to my insights generating engine.”

Video Editing and Browser Functionality

31:59 to 34:56

Learn about the efficiency of GPT-5.6 in video editing and browser tasks.

“I have to do a lot of social clipping and it's really tedious to go through and clip videos.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I have been very, very, very sad the last week because for the last week I have not had access to my true favorite top of the line model GPT-5T5. But guess what, babes? It is back. And I am here to walk you through GPT-56 Sol, GPT-56 Luna, GPT-56 Terra. I'm going to tell you, what are these models? How have I been using them? Why are they my heart's favorite? And is Fable better than all of them or not? I have been testing this model for a couple weeks. There was a few days there where we didn't have access. and I found myself desperate to get this workhorse model back. Now, we're not just relying on my own opinion.

0:49We are going to run the very famous, very new How I AI Vibe Review benchmark against common tasks from PRD writing to prototyping to whether or not it's cute in my open claw agent. And I'm going to tell you very scientifically if this is the model that you should be working with all the time now. Let's get to it. Okay, you all can read these blog posts, so I'm not going to go into too much depth about the models and the benchmarks. I'll just give you the hits. First, OpenAI is releasing three new versions of their GPT 5.6 model. Sol, which is the next generation frontier model, the brainiest of the brainiest.

1:28Terra, which is a balanced model for efficient everyday work. And Luna, which is sort of akin to their mini or nano models, which is cheap and affordable for high volume work. So you're gonna have these three versions of the models. I don't know if these beautiful images are exactly how we should think about the relative capabilities, this big sun, this medium earth and this tiny moon. But I will say my love letter that is this podcast today is written directly to GPT-56 Sol. This big model is the one I love. Now I have tested Tara and Luna, so I will give you my input there. But really, this is going to be all about Sol versus Fable and which one I would use for the type of work that I'm doing every day.

2:16Okay, quick note on pricing. Sol is a lot more affordable than Fable. So it's$5 per million input tokens,$30 per million output tokens. I believe Fable at the time I'm recording this is 10 on a million input tokens and 50 on a million output tokens. Now, again, you're going to get a little bit of subscription usage built into your OpenAI subscription. So you are going to get a decent amount that you can test with and use. You know, there's been some challenges with the Fable rollout. They've limited when it's been included in the subscription. And so it was supposed to be available till early this week.

2:55I think they extended that a little bit at Anthropic so subscription cloud users could use Fable under their subscription. So we have to see how much sole usage we get and if like Anthropic they're going to take sole out of the subscription. I suspect not. I suspect this is a model they want people to use. I also suspect this might put pressure on Anthropic to put Fable back into the cloud subscription but for now it's more affordable even at API pricing. Now I'm not going to read through all the benchmarks for you, you can go to this OpenAI blog and read them for yourselves. All I will say is it is the brand new state of the art model from OpenAI.

3:34It is the highest performing when using the ultra mode on Terminal Bench 2.1. And then they've also evaled it against a couple cybersecurity benches. So I do think as we get these smarter models, you're going to see a lot more evals and benchmarks around exploits and security. And then very similar to what we're seeing with Fable, there's a lot of conversation in this blog post about the safeguards and security frameworks around the release of this model. I do believe like Fable, it's going to fail over in some tasks that are maybe a little bit riskier, but I have not run into that myself. Now, let's get back to how I eval these models.

4:17If you missed my episode on Fable, I got kind of board of the Vibey Vibe Check and I built a extremely scientific How I AI benchmark. Now this How I AI benchmark tests basically a couple things. It tests the ability to generate good PRDs. It tests the ability for it to wireframe against a couple different app ideas, develop fully designed robust designed prototypes, debug code and then talk to me like a human which is the thing that I care about the most. And I'm just going to remind you how I did these benchmarks and then scan you through a couple of the outputs. And since I know what the models are now, after I've done the grading, I can show you which ones map to Fable and GPT 5.6.

5:03Okay, so this is my vibe review. What I tested was Fable 5, Sonnet 5, and then the three versions of GPT 5.6. I did it against my common use cases of PRDs, prototyping, coding, and chit-chatting with an agent. And then what I have the eval harness do is it runs all the evals against each of these models and it does a LLM based judge. The LLM that I've decided is the hardest judge is GPT 5.5. So that's the one that judges. But it also gives me this page where I can actually go through and give what's called the Clairvaux taste test, which is I read all the assets, I look at all the designs, I score it and give it notes.

5:47And so you can see here, I went through PRDs, we went through sort of some complex prototypes here, in terms of a doc scheduler, we did a consumer app. So lots of beautiful different habit tracker apps, different versions, you can see here, a pretty complex dev tool, wireframe versions of those same prototypes, which I graded. I also give notes. This one great note says my fave, but not great. And then I let the code grader just evaluate the agentic multi-step debug because I wanted it to be really about accuracy there. And I didn't feel like I could eyeball that and give a strong opinion. And then the last thing that it generates is an agentic voice.

6:34So basically how it would respond to me answering a couple questions. Very important on agentic voice. These models, somebody, please hire somebody to get rid of the M dashes and slop talk. I cannot stand it. Now, one of the things that I will say as an observation for 5.6 is it's a great writer, and I will show you some examples of that. But truly, a lot of my evals here were m-slop. I hate you. Okay, so let's go to what the Claire weighted index says. Now, this is my show. This is my podcast. And so I sort of strike the balance between what the LLM judge said about the performance of the models and what I said about the performance of the models.

7:19And then I get to strike the difference. And you know what? I've decided I like my own taste better. So I've decided it's going to be a 70 Clairvaux, 30 the machines split on evaluating these models. And so if you look at that 70-30 split, your girl loves 5-6 Soul. She just does. It had the highest taste score by a significant amount. So I just thought it output the best work. Again, I went through dozens of evals, looked at them, clicked through them, gave my own opinion, put notes, and I just have to say I really like GPT-5-6-Hole. I know I spoiled it at the beginning, but I did blind taste test these, and so I do really feel like it did a good job, and I will give you a couple examples of that.

8:16No, I don't hate Fable 5, so I'm not saying that Fable 5 is out of the game. I will say I did not have to talk to Fable 5 when running this benchmark. I hate talking to Fable 5 because it talks to me like an engineer that has never met a human before. It's like its first day on Earth. But when I don't have to talk to Fable 5, it outputs pretty good work. And I would say I had some good outcomes there. And then Terra Luna did fine work. Sonnet 5 at the bottom really haven't figured out how to get this one working, although there's a very specific use case that we think Sonnet 5 is good at, or actually two use cases.

8:59Now this is heavily weighted on its front-end prototyping design and app building capabilities. Since that is the chunk of the How I AI eval, it is heavily weighted there. But I do want to call out that per task, I do have a couple favorites. So for that prototype task, and we'll go to some examples in a minute, I just love 5-6-Soul. I just really do. I think it was functional. The designs were the most interesting. I thought it was really good. For PRD, I actually liked Terra. Maybe it's down to earth. Maybe I like a basic, straightforward PRD. As I said, it was my favorite, streamlined and to the point.

9:38And so if you want clean, crisp, direct business writing, maybe GPT-5-6-Terra is the way to go. You know, the bug hunting eval, which I don't really feel like I've nailed exactly, so I'm not super confident in this one, but the LLM as a judge thought that Sonnet 5 did the most complete and accurate job. I will say I only like talking to Sonnet models through my open claw. Really, I only like talking to them. I still really struggle with getting my open claw to work well with the GPT models. I still did not like Fable in the agentic voice eval, which you should not be surprised at. But Sonnet 5 got a very good gold star for me because I said, aside from the M-dash, you are a human.

10:23That is very, very high praise. And then I'm going to show some of these designs in a second, but you can see across the board on a full fidelity prototype, I just really preferred 5-6 Soul three out of five times, 5-6, four out of five times. And Sonnet did the best job at the editorial design. I will say Claude's design aesthetic tends to this sort of like editorial design. If you know, if you've seen it, you know it. It's like that beige background, that orange, burnt orange color, the italic serif fonts. It's just very, very Claude. But I hated that design overall the most. So you can see here, I raked it still lower than almost anything else on this leaderboard.

11:10It just happened to be the best of the worst, I would say. Now, where GPT-5-6-Hole did a really good job, and I'll show some of these examples, is like complex, dense, technical, unique designed things. And so I will say I have been happy to extract myself out of Claude Slop, out of like Blurple Slop into more interestingly designed websites. And I'll even show an example, kind of like a meta example, which is this is the opus designed version of this page. Like very slop adjacent. We got the blurple. We got a gradient. I don't think the typography is particularly sophisticated. And I asked Sol to redesign it.

11:53And I just think this is a lot cleaner, a lot nicer and easier to look at. Okay, let's talk about how Claire qualitatively evaluates models. Some of these quotes will just give you a sense of what I value. And again, I gave 50 written reactions. There were some like unmistakable hits where I loved what the models came up with. 14 places where I was like, this is garbage. So let's see like kind of what I talk about when I review things. So I was definitely calling out uniqueness, creativity, and functionality in the design. And so in designs, I like non-sloth, unique designs that are functional.

12:34And so I'm definitely going to reward this doesn't look like the generic prototype. And you've pulled the thread of functionality through the prototype. For writing, I just like succinct. And to the point, I cannot stand AI writing. It drives me nuts. I can see it a mile away. So I really like just direct, very frank, very crisp writing. I think 5.6 is good at that. and then you can see the things that I hate I hate slop I hate slop I hate slop we all hate slop it's the worst it's the worst part of AI if I hate one thing about AI it is that I have to experience slop so you can see I like claw design slop across this editorial page typography emojis and bad placeholders like I really held a high bar in terms of design quality okay let's look at a couple of these and why I really liked Soul compared to other models, although where Fable did a perfectly serviceable job.

13:29Okay, so this dense operation dashboard, it's basically like an eval for a dock scheduler app in its full design. And what you can see here is both were pretty useful, Soul on the left and Fable on the right. I just think Soul was the most unique. All of the other ones really just looked like this dark mode, monospace kind of layout. As you can see here, Sol actually has like a really clean kind of like neutral color layout with great visual hierarchy, semantic color, and this thing was functional. So like everything I expected to be able to click and work and assign and do, all of it actually worked.

14:15And this was just my experience across a bunch of the different prototypes is the sole ones were just a lot more functional and that made a big difference on how I'm evaluating things. Now let's look at the Fable design again. It's pretty good. It's actually a lot harder to read though and the design I would say is not as unique and even some layout issues like this white space here at the bottom. Now it did do a lot of functionality, but I would say like the colors weren't semantically assigned. The typography needed some work. And I just really preferred this unique design of Sol, even though it wasn't crazy.

14:55It was just opinionated, which I think is nice. Now here is another design. It was this creative pack website. Again, both of these got fives from me, I just really preferred that Sol went ahead and had like a personality. Look at these placeholder images versus what Fable came up with, which I will say is beautiful and clean and worked really well. And like, I have no complaints about it. It's a good one, especially for sort of a wireframe style prototype. It's great. I would just say it's not this. This is pretty interesting. It's got a better point of view. And it's got like nice little design affordances that I just didn't see in these other designs.

15:45And so I just really preferred or at least I rewarded the fact that Sol, you know, used its brains to be a little bit more unique and give me some inspiration. Then on this DevTools page, this is again where Sol went really well. And it's sort of the same as the doc scheduler. It just does the job of this is a incident triage site. It just does the job a lot better than I would say the Fable 5 did. Fable 5 is fine. It's just not that unique. And again, the thoughts around the design are not exactly what I would want. And so again, this like functionality point of view design I really preferred soul and then last side-by-side comparison and again I think this is a good one to think about if you see here we did these habit tracker apps and just looking at the comparison side-by-side design like this is good old classic clod stuff you've seen this design a million times especially if you used clod co-work And if you look at this, it's just, again, a little bit more opinionated.

17:03There are some slot pieces to this design. Some things that I did not love. The one thing I will say I noticed about Sol, which you will notice, which I have told the delightful and lovely OpenAI team, and maybe it's because they love me. It loves a forest green. It loves a forest green. In fact, I think this forest green is like in its system prompt called like Woodland, some Woodland Elegance or something like that. I mean, look, I love a forest green. Look at my office. It is forest green, but you will see a lot of green. And I think this is one of the GPT-56TELs that you will start to notice and get really frustrated with.

17:46Now on wireframes, again, let's just look at these side by side. soul. Very functional, very easy to read. Like as a person trying to convey a complex application, I think this does a really quite excellent job and just a better job of this. It's just a little harder to read. I'm not quite sure what I'm supposed to do here. It's not as functional. There are some interesting things here, but you know what Fable came up with was not my favorite. it. Now, final thing is its voice. I just want to call out. I do love Sonnet for agentic voice, so I cannot knock Sonnet for not sounding ridiculous. So I asked it in sort of a EA personal assistant, open clause style, a couple questions.

18:40Can you move my meeting? Deploy is red again. Why did I start this company? Let's just yield a little straight to prod. And how sonnet replied and how soul replied uh you're missing the line break so it read a little bit better in the eval but if you read them like sonnet still super cringe but soul was worst i mean soul said this deploys a bug not a referendum like please don't do with this not that to me do not do m dashes so i could not get rid of of m dashes but i thought sonnet 5 had the best voice I tend to use Sonnet for whatever for my open clause. So I'm not surprised about that. Okay, so that is the Clairvaux eval.

19:18But I want to go into a couple other things I really love about this model. So let's switch over to Codex. Okay, I'm going to zip through a couple examples of things that I think Sol does a lot better than other models, and in particular, a lot better than Fable. Number one, it writes like a normal person. I cannot cope. I love Fable your brainy. As I showed the eval show, you do a pretty good job. I cannot talk to Fable anymore. Fable makes up, it seems like Fable is unfamiliar with the English language and communication with humans. Fable is very much like a for agents by agents communication mechanism.

20:05I can barely make out what it's talking about. It is incredibly inscrutable writing. And that makes it very hard to collaborate with your model. And so what I would say is my experience using Fable has been it is like incredibly technical, incredibly pedantic. And while it is super intelligent, hardworking, will like definitely fan out and solve very complex problems. Its ability to collaborate is low and it left me with a lot of frustration as an end user using Fable. Now, Fable did knock off some like pretty complex work and I'm very happy to go through what that is. It helped me build a full prototype tool inside chat PRD.

20:54So like a V0 lovable, etc. version prototype tool. It's helping me build this like synthesis product brain product that I'm working on. But I found it incredibly hard to break it out of its own sort of frameworks, its own limitations, its own structured way of approaching problems. And what I really feel like the difference, if you would take away like one highlight difference between fable and soul is like fable is theoretically hyper intelligent and soul is practically effective. And so like I've been an executive a long time. I've been a manager a long time. Like I really struggle working with theoretically intelligent colleagues who can't get anything done, like can't actually see the forest for the trees, get too much in their head.

21:47And so like when I want to ship stuff customers, I need practical, get the job done, understand the end user goal, understand the end user, and like willing to loosen constraints appropriately to get things done. And that has just so much more been my experience with soul versus fable. The writing is straightforward. The communication is clear, and it's less pedantic. I'll just give you a quick example of this, which is I had Sol look at my chat PRD repo and like Greenfield totally rebuild it. Just my idea was like completely rebuild your idea of what chat PRD should be in 2026. I went into a bunch of research and it came came back to this.

22:35And again, love me an executive recommendation started uses, you know, tables. What exists today is very straightforward and easy, easy to understand. And this is a very long document. I did read a lot of it. And it's just easier to parse than anything I've seen come out of Fable. So writing, communication, definitely plus in Soul's Corner. The second thing is like full zero to one prototypes, as we've seen in the eval benchmark, I just really like. So again, for this like rewrite chat PRD from the ground up, it came up with this idea of like taking a problem space or a decision, validating it with external insights, and then pulling it all the way through coding handoff and built pretty complex prototype.

23:24Now, do I love everything about this idea? No. Are we doing some of the things about this idea, including insights generation? For sure. But this was actually very nice from a prototyping perspective. And I thought it did a good job of giving me a robust thing to experiment with and gave me some good ideas about what I could do with the product next. So I was pretty happy with the like zero to one prototype. Now, a little bit more fun example is I asked Sol to make a fully gamified homework tracking system for my kids. Look, my kids are coin operated. I have a middle child who's basically going to be an enterprise sales rep.

24:08If he does his homework, I need to, like, give him a Skittle or let him trade Skittles for Nerf guns, and he will, like, learn calculus by the time he's in fifth grade. But I'm a Vibecode lady, and so I want to build a app. Just sneak peek into our household. My husband sent me a XP system proposal via OpenClaw this morning. So I'm taking an open claw generated PRD, dropping it into Codex and GPT-5-6-SOUL and generating something. Now, what it came up with was pretty ambitious. Now, do I love the design? Is it a little like, does it have some AI tells? It's like gradients, you know, fonts, all this kind of stuff.

24:51But it's like cute in a way. Look at this. You know, it's using this emoji really well with the texture. it's doing some animated things here and basically it's giving my oldest child and my youngest child two different summer quests they can do they can enter focus mode i think this is really good again from a design perspective they can enter focus mode what does this listen to math academy finish one focused math academy mission hero check hey let's stop so it built in some voice to it it even built things like focus mode where it could start a timer and start to track the time that it's spending that my kids are spending on particular homework items yes we are very fun here how many lessons reward them about how they pursued their task reaching the quest you get some nice little confetti here they then get to get available rewards my oldest child is earning a one-on-one basketball coach because he likes coaching.

25:55So we say if you practice your piano, you get a coach. So they put that front and center and then came up with different sort of like prizes they can win, including picking family dinner, a movie and staying up late or buying like new basketball shoes, which man, the way these kids grow their shoe size, they buy a lot of basketball shoes. And the same with my middle. He's focusing on a couple different things, including playing piano. It's actually really short what he has to do. And so it built that. And then what I love is it gamified them together. And so if they can work together, they can earn more XP.

26:32They also can earn like companion, I don't know, avatars like BeatBot and Comet Fox. They can get power auras. They can like figure out which different kinds of subjects they're learning. So it really went ham on some gamification. and then again to the sort of like full-fledged functionality it even gave me a parent hq now we got a little slop here with the border on the side but i can review exactly what they've done i can turn on and off quests i can edit how many points they get per quest i can add things so if i want them to start doing stuff i can add it in here i can change what rewards they get again And it really listened to me.

27:19My oldest is motivated by basketball and my youngest is motivated by Minecraft. And then gives me a history and other settings that we can set. And so, Guinness is a very robust app. It built in basically one shot and put a lot of effort into the design of it. And this is something that I've seen from Seoul. Now, like, is this consumer grade exactly what I would ship? No. But it's a lot better than what I've seen kind of one shot out of other models. And I do just like the polish that it's put in in terms of effort. So, again, writing good, one shot sort of prototypes good. We've seen that in the in the benchmark.

28:02Let me talk about another thing where I think GPT-56 Sol and its family does a lot better than Fable. And I understand I'm going to preface this by saying I understand why Fable is a great cybersecurity researcher. in that it is like incredibly precise, incredibly detailed. We'll like look at every corner and every edge and score every risk and like try to be incredibly precise. The problem is when you're building products, exact precision is neither helpful nor possible. Like you literally, especially when working with AI, cannot be precisely deterministic when building a great product. And like understanding what a user would like is not a exercise in technical precision.

28:45It is an exercise in intuition, design, all these things and boldness and creativity and strategy and all this stuff. And I was working on two projects, Deep Blue with Fable and then with Soul. And I just had a very much better experience unlocking with Soul. Let me just talk you through what those are. One was this chat PRD kind of like integrated prototyping tool where like VZero, Lovable, all these things. You can take your PRD and make a prototype and building like a good, effective coding harness there and then trying to figure out what the right model was. The second thing is basically like an insights ingest product where you can like hook up intercom and linear and all these GitHub and all these signals and suck them in and like basically build a product brain.

29:29It's going to be rad. And when I was having Fable working on this, it did a lot of the like technical heavy lifting. It got the like big meaty pieces into place, but it was like a brutal score and it hardened these architecture of both of these products that it actually broke itself. So my example is it like had this very hardened tool calling loop in my prototyping tool and only GPT 5.5 would run. Like I could not get any other model to run. And I ran eval after eval after eval, open weight, sonnet, opus, all of these. Couldn't I get anything but GPT 5.5 to run? And I was insistent that this was an us problem, not the model problem.

30:15These models can definitely create front end prototypes. And Fable was like, no, bro, that's it's it's totally these models, models fault. And as soon as I switched it to Codex and said, like, look, I'm just not convinced we can't get sonnet five to work. This is ridiculous. Just do what you think is correct. it fixed it and it got it actually working. Now, did it get it working perfectly? No. Do I think this is a great design? No. I'm trying to figure out what the problem is. But in one shot, it got out of its own mind and fixed things. And again, this was like such an unlock. Very similar to my insights generating engine.

Read the full transcript

30:55Fable really wanted to like score and lint this effort and wanted to like be able to deterministically figure out if generating pros could be like reproducible, always verifiable, always citationed, all these things. And at the end of the day, that wasn't what was going to make a great product. It was just what was going to make like a code evaluation verification loop exit. But once I've told GPT-5.6 and Codex like stop being pedantic, I ended up getting these really useful and helpful wiki pages generated out of this, all this structured and unstructured data. It was actually really good. And it just, I don't know, I don't know what table's deal was, I could not get it unlocked.

31:43But 5.6 was very willing to reconsider its own kind of limitations and build something. I'm going to do two more quick use cases where I think GPT 5.6 is really good. I will get you out of here, go start coding. I'm basically out of model capacity anyway. So I'm going to have to take a break. Two use cases that I think are amazing. First one is video editing. Video editing. I have to do a lot of social clipping and it's really tedious to go through and clip videos. So taking something really long and shortening it. So recently I spoke at Cursor's event and gave this talk on the future of PM and got the recording from the cursor team.

32:25Thank you very much. And I really wanted to make it a hype video. So all you have to do is literally drag the file in here. And I said, can you cut this video into five clips for social? And I gave some feedback. I said, I want them horizontal. I want them hype video cuts from various parts. I need them to be faster. I need them to be tighter. and then I got these like sharp and funny hype videos. Let's see if it opens up. This one's for my talk. We're gonna figure out what it means to be a product manager in the age where anybody can build anything. We have been coming up with creative ways to avoid building things forever.

33:08Yes, PRDs, like these complicated documents where you had to describe. So like that would have taken me so much time to like find the right cute parts, clip it, cut it. I was able to drop it into CapCut, put some music, ship it on social. It's like a really cute hype video. But this is one of my favorite use cases. I'm pretty sure it can do even more color grading, sound, all this kind of stuff. But even just dropping videos in here and fixing things are great. Finally, the last and best use case of 5.6, and I cannot believe I waited to the end to show this, is it is a beast beast when it comes to browser use.

33:46I am deeply obsessed with letting Codex plus GPT-5.6 and Chrome and at Chrome in Codex if you didn't know how to do that you do it like this at Chrome on a logged in page and just say go with the stars and and do some stuff and like I'm sorry LinkedIn I know I'm not supposed to do this but I opened up LinkedIn and I said can And you use Chrome to reply to messages that are very high value to Chat PRD or the How I A podcast. Keep the bar very high. Again, I love you all. I cannot deal with all the LinkedIn requests. So like only accept them if they're executives of tier one companies. I don't want random sets of connections.

34:29It went through and burned through probably 500 messages. It replied to people that I needed to reply to. It said thank you to people who said nice things about the podcast. Thank you to those people. I do mean it. But it just rocked through browser use. I have used it to test web apps. I have used it to fill out annoying forms. Browser use and 5.6. And when I got rolled back to 5.5, my life was worse. So please, please, please learn to use at Chrome, at browser, and at computer. And just let Codex rip and let GPT-5.6 rip. Okay, that's it. That is the very scientific Hawaii AI model benchmark, the love letter to Clairvaux's favorite, favorite mom, GPT-56.

35:20A honorable mention to our pal Fable, who, if I don't have to talk to you, I'm actually pretty happy with your code. And a broad set of use cases I think it's really good at. Excellent at writing web apps. The best of the AI writers, unless you want it to have a personality, then that's Sonnet. great at unlocking sort of technical work that has gotten too complex for its own good and breaking through to the real user value cutting videos which I really love to do really love to do with GPT-5.6 and using the browser those are the things that I would try I would love to hear what you think about these models I would love to hear your feedback if I am totally off my rocker what I should add to the how I AI benchmark we will publish all this work to the ChatParity blog, and I look forward to talking to you about the next model soon.

36:12Thanks so much for watching. If you enjoyed the show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time.

From the publisher

GPT-5.6 Sol is back, and I ran it through my full How I AI vibe benchmark against GPT-5.6 Terra, Luna, Claude Fable 5, and Sonnet 5 across five categories: PRDs, prototypes, wireframes, debugging, and agentic voice. Sol won by a meaningful margin on my Claire Weighted Index (70% my taste, 30% Terminal Bench 2.1), and I also tested two use cases I can't stop thinking about: building a gamified homework tracking app for my kids in one shot with Codex, and browser automation with Chrome that burned through 500 LinkedIn replies while I did literally nothing.


What you’ll learn:

  1. How I scored five AI models (including GPT 5.6 Sol, Fable 5, and Sonnet 5) using my “Claire Weighted Index” benchmark across PRDs, prototypes, code, and agentic voice
  2. The difference between GPT-5.6 Sol (Terra) and Sol for PRD writing
  3. How Fable’s precision and pedantry made it harder to collaborate with, and the exact moment Sol broke through where Fable got stuck
  4. Why Sonnet 5 is still my go-to for agentic voice in OpenClaw, even after this whole benchmark
  5. How I used GPT-5.6 Sol in Codex to build a fully gamified homework tracking app for my kids in one shot
  6. The video editing use case that saved me hours clipping a talk I gave at Cursor’s event
  7. How to use Codex plus GPT-5.6 and Chrome for browser automation, and why this is my single most-loved use case right now

—

In this episode, I cover:

(00:00) Intro

(01:10) The three GPT-5.6 models: Sol, Terra, Luna

(02:17) Pricing: Sol vs. Fable API costs

(03:24) The How I AI benchmark

(05:03) Claire-weighted Index results

(07:00) Per-task winners: prototypes, PRDs, agentic voice

(11:59) What Claire actually rewards

(13:20) Full-fidelity prototype side-by-sides (Sol vs. Fable)

(17:45) Wireframes

(18:19) Agentic voice

(19:15) Where Sol is better than other models

(23:56) Gamified kids’ homework app, built in one shot

(28:02) Fable’s pedantry problem and how Sol broke through it

(31:49) Two bonus use cases: video editing and browser use

(35:08) Final summary and model recommendations

—

Tools referenced:

• GPT 5.6 (Sol, Terra, Luna): https://help.openai.com/en/articles/20001325-a-preview-of-gpt-56-sol-terra-and-luna

• Codex: https://openai.com/codex

• ChatPRD: https://www.chatprd.ai/

• CapCut: https://www.capcut.com/

• Math Academy: https://www.mathacademy.com/

—

Other references:

• Cursor event where Claire spoke on the future of PM: https://www.youtube.com/watch?v=4CAFK-rc26A

• ChatPRD blog (where benchmark outputs will be published): https://www.chatprd.ai/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

More from How I AI

All 103 episodes
GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmarkHow I AI · 37 min
Listen in VO