In short
Claire Vo reviews Anthropic’s Claude Opus 4.8 coding model, asking whether it’s as good as claimed, focusing on one-shot success vs failures in edge cases, hallucinations, and strategy grounding.
Guest backgrounds
No guests. Host is Claire Vo, product leader and AI obsessive.
Key claims
Opus 4.8 is intended as a step-change “agent” model that’s more honest, instruction-following, and supports longer-horizon autonomy. Benchmarks cited: Sweebench Pro 69.2% (about +5 vs Opus 4.7). In her tests, it’s fast and tool-friendly but overconfident without validation, hallucinating and shipping bugs.
Notable examples
In CloudCode, it autonomously built a chat PRD prototyping tool in ~20 minutes and worked on a preview branch, but later bug-fixing produced hallucinations. It struggled rebasing existing branches, requiring repeated cycles. In business strategy, compared to Opus 4.7, Opus 4.8 over-rotated on small data points and produced hand-wavy roadmaps; it also claimed it didn’t need to search GitHub/validate bugs.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOOverview of Opus 4.8
0:45 to 1:15
Discussion on the features and expectations of Opus 4.8.
“to talk about what this model is, what Anthropoc has told us about its benchmarks, performance, and what it's good at.”
Initial Impressions and Performance
1:15 to 2:54
Sharing personal experiences and early thoughts on Opus 4.8's coding ability.
“almost 10 points higher than GPT 5.5 and 15 points higher than Gemini 3.1.”
Strengths and Weaknesses in Coding Tasks
2:54 to 5:04
Detailed experiences with Opus 4.8's successes and failures in coding tasks.
“consistently over time with the same types of trouble.”
Business Strategy Analysis with Opus 4.8
5:04 to 9:03
Evaluating Opus 4.8's performance in analyzing business strategies.
“I thought it did a really good job one shot on a surface area.”
Reflection on User Experience and Output
9:03 to 11:14
Discussing the overall user experience and some positive aspects of the model.
“but I consistently got this experience of the model hallucinating or over-rotating on a hypothesis that it had, as opposed to being anchored in true code truth or in true business truth.”
Final Verdict on Opus 4.8
11:14 to 12:50
Concluding thoughts on the model's capabilities and areas for improvement.
“I did find it was enjoyable to work with.”
Transcript
Automatic transcript. May contain errors.0:03Welcome back to How I AI. I'm Claire Vo, product leader and AI obsessive here on a mission to help you build better with these new tools. Today we have a very special mini episode because Anthropic just dropped Opus 4.8, their latest state of the art coding model. And I got a few hours of early access and I'm here to share my very early thoughts about where this model is intended to perform well, where it did a great job and totally impressive. me and where there's still a little bit further to go. Let's get to it. As you can tell, I am not in my regular How I AI studio, and that's because I am so excited to give you my early thoughts on Opus 4.8 and couldn't wait between meetings to share what I thought.
0:44So to get started, I want to talk about what this model is, what Anthropoc has told us about its benchmarks, performance, and what it's good at. So Anthropoc is shipping Opus 4.8. It is supposed to be their step change model for agents. And there's a couple things they've called out that this model does particularly well. It's supposed to be more honest, a less designed flop, longer horizon autonomy on long running tasks, and enterprise ready. So it means it follows its instructions. And they're saying that Sweebench Pro, they're hitting 69.2%, which is almost five points higher than Opus 4.7, almost 10 points higher than GPT 5.5 and 15 points higher than Gemini 3.1.
1:27Now this model is not cheap it's$5 per input tokens and$25 per million output tokens and then same as 4.7 effort defaults to high and fast mode can be a lot faster. This is what they say this is what you're going to read on the blog post and so on paper this is a very exciting model but I want to tell you my personal experience using this model and where I thought it did a really good job. And again, where it did not do a perfect job. And so when I was giving feedback to the team, I said, surprise, surprise, LOL, it's a good coding model. In that when I opened up CloudCode and asked it to do a fairly complex one-shot brand new surface area task, it did a pretty good job.
2:08So I asked in CloudCode, Opus 4.8 to build a prototyping capability in chat PRD. So we make PRDs. I said, let's just go whole hog. Let's compete with the big boys. Let's make an entire prototyping tool. And I gave it some architecture decisions I wanted to make, what platforms I wanted to use, how I wanted it to function. It went through plan and then it autonomously coded for, I would say, about 20 minutes and shipped it. And when I pushed this live to my preview branch, it worked. And so I would say from a one-shot feature it did quite a good job the code was right and it followed the architecture I want.
2:45Where it failed was this last 10 percent and this is really going to be my theme of this episode. It does really really well until it doesn't do well and I found it did not do well consistently over time with the same types of trouble. So I'm curious as you all get your hands on this model if you have the same experience I did where it does like really really really well and then struggles in the edge cases and the details. So what it nailed here is it did take the spec, it planned the work, it shipped the feature. But then as soon as I got it live and started trying to take it to the next level, the next level, the next level, it really struggled and started to ship bugs.
3:23And even more than its inability to finish that last 10%, when it was bug hunting, it hallucinated. And I am going to tell you, I have not seen a straight up hallucination in a very, very, very long time. But over my experience early testing Opus 4.8, both on business use cases as well as coding use cases, it 100 % made up things based on hypothesis, not data. And this was really interesting to me. This was on high effort. And so I don't think it was effort or reasoning. There's something about this model where it's really not grounding itself as effectively as I've seen in other models. Again, this was a one-shot, but then very specifically propped it up on scoped surface areas for follow-ups, like I saw a bug in the preview branch and got these hallucinations.
4:12And so this is a really interesting reflection of this bug. I'm going to have to run at it a little bit more to see if this holds over time with coding use cases, but it was kind of the theme of my test here. Okay, this headline is a little dramatic. It says in real code bases, the edges destroy it. This is not Opus 4.8. This is just Claude Cowork fail here. It doesn't have the screen, the screenshot. So I'll have to show you the GitHub for this. But basically what I saw is when I pointed it at existing code, it also struggled to sort of insert itself and understand the edges of where it was supposed to work.
4:44So let me give you an example of this. I had a couple of branches in flight that I needed to rebase, that I needed to bring up to base because we had shipped a big underlying PR and it kind of messed up the state the code. And so I asked Opus 4.8 to rebase and check these branches for code. And as you can see here, I had to do cycle after cycle of rebase and fixes because it was continuing to ship really edge case bugs into the code. And again, this was my experience. I thought it did a really good job one shot on a surface area. But then when you got into the specifics, it struggled to understand the elevation at which it should be operating.
5:24The third coding use case that I tried was a fun one, which is I pulled up CloudCode and asked it, what are some fun things we can one shot with CloudCode that my nine-year-old would think is rad? And I really tried to push it to say, make it really interesting. Think about the edges of agentic coding. And aside from the code quality itself, which I struggled with, sort of had highs and lows, the other thing I reflected on when I was coding with Opus 4.8 is it just wasn't ambitious enough. And so it gave me this awesome prompt, which was build a game, then play it yourself by watching the screen and tweaking the difficulty until it's fun for a nine-year-old.
6:01Amazing. This is state-of-the-art coding agent. It's going to cook. Let me show you what it actually shipped. It shipped this, which is like fine. Of course, magic. Like I would have never been able to ship this by myself without a lot of effort, but not pushing the edges of agentic coding. And even when I said, great, let's make it 3D, let's do something even more fun. It ships something like this, which again is super cool. I would have never been able to do this, but it's not 10x agentic coding blow my mind impressive. And so this is where I really struggled with Opus 4.8 is I kept saying more, more, do better, do better.
6:41And it just wasn't as ambitious as I've seen other models be. So in terms of coding, I think it is a totally serviceable job. I wouldn't say it's bad at coding. I would just say my experience has been it struggles with the last 10%. It's not exceptional at orienting itself inside existing code bases. And then it's just not that ambitious. Now let's talk about business work. So I also tested Opus 4.8 in Claude Cowork, and I tested it on strategy. And I gave it this very broad prompt. And I tested 4.7 versus 4.8. And I basically said, based on what you can gather about my last three months, where am I spending my time versus where my priorities should be if I want to 10x my business, I gave it access to all the same business context.
7:27And then once it did that analysis, I said, please write me a strategy prompt. And this is where the performance of Opus 4.7 versus Opus 4.8 really became apparent. Opus 4.7 was very numbers-anchored. You can see this table here. I obfuscated some of the numbers, but it was very numbers-anchored. It was very structured and rooted in real data. While both of these exercises did have access to the same data, Opus 4.8 had a harder time discovering the relevant data, and it over-rotated on small data points and took them as truth as opposed to what I experienced Opens 4.7 doing, which is it zoomed out a lot more and put everything in context.
8:12Now, again, this is mutual one shots. It was basically two shot. It was like analyze my time and then give me a strategy to grow my business. But the difference between these two were very, very high. I then asked it a follow up prompt to build a roadmap. And again, 4.7, very anchored in specifics, very good strategy. and 4.8 was incredibly hand wavy. And in fact, with Opus 4.8, it gave me a roadmap. And then I said, we have all this. Did you search through GitHub? Did you look online? And what's really funny, again, with the hallucination is you see here, no, I didn't. This is a common thing that I had Opus 4.8 say to me.
8:50No, I didn't search GitHub. No, I didn't actually look up that data. No, I didn't actually validate that bug. Now, again, this was early access. So I'm not 100 % sure if this is prompting error, if it's the shape of the model, if it's the harness that needs to be tuned, but I consistently got this experience of the model hallucinating or over-rotating on a hypothesis that it had, as opposed to being anchored in true code truth or in true business truth. And so I honestly would continue to reach for Opus 4x7, which I think did an exceptional job on strategy, versus for 8, which I think was a lot more hand-wavy and just over-rotated on things I didn't think was important.
9:30Now, that being said, what positively impressed me? Voice is great. Claude is not an annoying girlfriend, is what I would say. It was easy to read. It didn't have slop tells. It was token efficient. It felt like it was talking enough, but not too much. And it was fast. Now I got early access. Who knows what the production latency is? But with fast mode, I anticipate you'll have this fast experience. So I think the ergonomics were very nice. Now, if we zoom out and I say the writing was very good and then Opus 4.7 wrote this slide. I don't know if I love this slide that much. So hopefully 4.8 would have done a better job with the voice and ergonomics.
10:09But I do think the experience of using the model was very nice. It had no complaints. It was not annoying. It did not have ticks and tells. Just the outputs were not exactly what I wanted. So here's my theory and this is what I saw. It's just over tuned and has kind of narrow vision. So it's smart, it's fast, it's efficient, but it's overly confident absent true validation. That's what I would want you to walk away from in my review of Opus 4.8. It really latches onto specific data points, specific code points. It draws conclusions for them and then says, this must be truth. And so it sort of misses the forest for trees, both in coding and in strategy.
10:48And this might be part of its efficiency. Like I thought it was super efficient, but does that come at the cost of accuracy and would I rather a long-running sort of relentless coding model really going deep and validating its own opinion before shipping. So I didn't quite experience, I would say, this more honest and long horizon autonomy. I did see it was fast. I did find it was enjoyable to work with. I think it followed instructions well, but it stayed too much in scope, if that makes sense, because it didn't zoom out and contextualize the work that it's doing. So my verdict, I mean, all these models are great.
11:27They're all magic. So like, let's, let's be real. Every model is magic. The fact that I could do any of this in just a couple hours is pretty, pretty genius, but I would use it for greenfield prototypes. It's really impressive on a one shot. I think its design is better. It got rid of the italics emphasis words, which were driving me crazy from quad design. It's good at tool use. It's fast. It's not annoying. Where I would test it and really figure out the right prompting strategy and the right harness strategy is with existing code bases and branches with real edge case with strategy work that requires you to think about numbers.
12:03And again, you can probably prompt this, but I would just think about that prompting and I would just double check where it's really confident because my experience was its confidence was not rooted in fact. Again, I'm really excited to see this model come out. There's a couple more features as well in Cloud Code, as well as Cloud AI and Cowork. In Cloud Code, you now have dynamic workflows, which can let you spin off hundreds of parallel sub-agents. And in Cloud.ai and Cowork, you now can set effort control from low to max, which you were able to do in Cloud Code. So these are all really interesting shifts in both the harness and the model.
12:39I would say it's a good model. It's not the most amazing model. It didn't blow my mind. It has some quirks to it, what I think you need to be aware of. But I'm definitely going to keep testing it because with the benchmarks, with the work that's gone into the product, I think it's a model worth keeping your eye on. So that's it. That's my quick review of Opus 4.8 that just came out today from Anthropic. I'd love to hear your experience with it, especially how it does in coding, how it does in design, and whether or not it gives you strategy anchored in reality. Thanks for joining How I AI. Thanks so much for watching.
13:13If you enjoyed the show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time.
From the publisher
I got a few hours of early-access testing with Anthropic’s newly released model Opus 4.8. I walk through real coding, design, and strategy tasks across Claude Code and Claude Cowork, and give you my unfiltered view on what impressed me and what didn’t.
—
What you’ll learn:
- Where Opus 4.8 excels: greenfield prototypes, one-shot features, and fast execution
- Where it struggles: the last 10%, edge cases in existing codebases, and hallucinations
- How Opus 4.8 compares to Opus 4.7 on business strategy work
- Why I’m still reaching for Opus 4.7 on data-heavy strategy and roadmap work
- The new features shipping alongside the model: dynamic workflows with parallel subagents and effort control in Claude.ai and Cowork
- The prompting and harness strategy I’d use to get the most out of it
—
In this episode, we cover:
(00:00) Introduction to Opus 4.8
(00:44) Benchmark performance and pricing
(01:53) First coding test: Building a prototyping tool
(03:00) Where it failed: The last 10% problem
(03:27) The hallucination problem
(04:23) Testing Opus 4.8 on existing codebases
(05:24) The ambition test: Building games for a 9-year-old
(07:03) Business strategy test: 4.7 vs 4.8
(08:23) The roadmap test
(09:17) Final verdict
—
References:
• System Card: Claude Opus 4.8: https://cdn.sanity.io/files/4zrzovbb/website/c886650a2e96fc0925c805a1a7ca77314ccbf4a6.pdf
• Introducing Claude Opus 4.8 on X: https://x.com/claudeai/status/2060042702150930686?s=20
—
Where to find Claire Vo:
ChatPRD: https://www.chatprd.ai/
Website: https://clairevo.com/
LinkedIn: https://www.linkedin.com/in/clairevo/
—
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.




