BONUS: LIVE: Claude Opus 4.7 Just Dropped. Here's What Actually Changed.

17 Apr 2026 · 1 h 2 min · 34 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Live walkthrough of Anthropic Claude Opus 4.7 changes versus Opus 4.6, plus hands-on testing in Claude Code. They also discuss Anthropic’s “Mythos Preview” (withheld cybersecurity model) and a new open-source release of Quen 3.6.

Key claims

Opus 4.7 shows higher benchmark scores than 4.6 except for some agentic/tool-related areas (agentic search lower; agentic terminal coding slightly higher). They speculate Opus 4.7 may be a “distilled” version of Mythos, with some capabilities “nerfed” for safety. Notable improvements: visual reasoning jumps significantly; instruction following becomes more literal; better multimodal/vision (up to 2500px images); improved finance evaluations; better file-system-based memory. Concern: long-context/“needle test” performance drops slightly versus 4.6.

Notable examples

testing Opus 4.7 in a Final Fantasy Tactics-style Renaissance Italy game; asking it to improve portrait/sprite art using vision; adding Gemini 3.1 Flash TTS endpoints and new audio track types in an Electron app.

Guests

Kyle (developer friend co-host). No other named guests; chat participants provide questions.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Overview of Opus 4.7

0:45 to 1:15

Discussion about what Opus 4.7 is and its significance in AI.

“we're gonna we're gonna try this um okay Okay.”

Benchmarking Opus 4.7

1:15 to 2:30

Comparison of Opus 4.7 with previous versions and other models.

“I'm going to go ahead and grab that trail.”

Introduction to Mythos Preview

2:30 to 4:00

Discussion about the Mythos model and its implications in cybersecurity.

“So we'll let you watch that if you want to dive too deep into it.”

Benchmark Insights

4:00 to 5:30

Insights on the benchmarks and performance of Opus 4.7 compared to 4.6.

“And I think there's some credibility to that.”

Visual Reasoning Improvements

5:30 to 7:20

Discussion on improvements in visual reasoning in Opus 4.7.

“that makes it easier to find those exploits that he was talking about, right?”

Instruction Following Enhancements

7:20 to 9:10

Discussion on how Opus 4.7 follows instructions better than its predecessors.

“And I tell it to fix something with the UI, and then it's like, looks clean, looks good, you know?”

Concerns About Context Handling

9:10 to 11:10

Concerns regarding the handling of long context in Opus 4.7.

“Okay, so there will always be an exploit somewhere in the code, do you think?”

Live Testing of Opus 4.7

11:10 to 13:00

The hosts begin testing Opus 4.7 in a project context.

“So for it to go from, you know, Opus 4.6 level, which here you can see, this is their 64k extended thinking version, um, has 78.3 % on this eight needle test.”

Exploring Claude's Capabilities

14:00 to 14:15

Learn about the potential of Claude's coding abilities.

Game Development with Claude 4.6

14:15 to 14:57

Discover how Claude 4.6 assists in game creation without coding.

Show all 34 chapters

Conceptualizing a Strategy Game

14:57 to 15:46

Understand the design of a Final Fantasy Tactics inspired game.

“So I'm really trying to go for the Final Fantasy Tactics vibe and aesthetic.”

Improving Visual Reasoning in Games

15:46 to 16:47

Learn about the visual reasoning capabilities of Claude in game maps.

“No, I think that's probably the best way to put it.”

Evolution of AI Prompts

16:47 to 17:37

Examine how AI models anticipate prompt needs based on user interaction.

“I would say there's something close to this where they're generally taking the ways that people are getting the most out of the models, and then they're building that into the subsequent models.”

Transitioning to Claude 4.7

17:37 to 18:16

Understand the transition and challenges experienced when moving to the new version.

“So what's interesting about this is I don't have in this window, I don't have 4.7.”

Challenges and Feedback on Claude 4.7

18:16 to 19:06

Explore user feedback and challenges faced while using Claude 4.7.

New Features of Claude 4.7

19:06 to 21:46

Learn about the new capabilities and improvements in Claude 4.7.

“Also, did you notice that they changed the desktop UI recently?”

Maximizing AI Efficiency

21:46 to 23:03

Discover strategies for efficient usage of AI models to save resources.

“4.7 is also state-of-the-art on GDPVal, which is a third-party evaluation of economically valuable knowledge and does doing work across finance, legal, and other domains.”

Personalizing AI Workflows

23:03 to 25:31

Learn about various workflows that optimize AI model outputs.

“What do you guys see using 4.7 for though?”

Humor in AI Communication

25:31 to 27:41

Explore the current state of humor in AI-generated content.

“Yeah, I personally, I'm a little more shifting my workloads.”

Introducing LTX Desktop Tool

27:41 to 28:01

Learn about a tool that enhances AI image and video generation.

Building LTX Desktop: A Generative AI Tool

28:01 to 30:08

Learn about the features and expansions of the LTX Desktop app for generative AI in film.

“Meanwhile, so I have my terminal agent over here.”

Solo Development vs Big Companies

30:08 to 35:42

Explore the dynamics between solo developers and large companies in the tech landscape.

“And then, you know, there's all these sort of add-on things I've added since then, like an outline mode with a lot of different flows.”

Future of Cloud Development

35:42 to 37:02

Discuss the potential divisions in AI cloud service access for developers and consumers.

“They decided not to step on the launch of 4.7 or maybe it's coming later today.”

Audio Production Innovations

37:02 to 39:45

Discover advancements in audio production modes for generative AI applications.

“But yeah, there was supposed to be something about that.”

Exploring Opus 4.7 and Quim 3.6

39:45 to 43:12

Gain insights into the new features of Opus 4.7 and Quim 3.6 in AI models.

“as it's working on it over here which is nice.”

Mixture of Experts in AI Models

43:12 to 44:24

Discover the benefits of using mixture of experts in AI models.

“I've been playing around a lot with the Gemma 4 MOE recently.”

Creating a Renaissance Game

44:24 to 46:06

Understand the process of generating a game using AI and visual reasoning.

“But I'll make it open source so people can play it.”

Using Cloud Code Routines

46:06 to 48:26

Learn how to create agentic workflows with Cloud Code Routines.

“faces to look a little bit more like Renaissance style.”

AGI Development and Performance Concerns

48:26 to 51:18

Discuss the rapid development of AGI and concerns over performance degradation.

“But routines are essentially the cloud code equivalent.”

Audio Enhancements in AI Models

51:18 to 54:08

Explore recent audio enhancements and their implications for AI models.

“Somebody commented on that this morning, where you intentionally downgrade the previous version to make the new version look better.”

Testing New AI Features

54:08 to 56:01

Learn about testing methodologies for new features in AI models.

“Let's see, decision taken, all reversible later, hard-coded narrator voice.”

Technical Setup for the Live Session

56:01 to 57:36

The hosts navigate technical difficulties while preparing for the live discussion.

“yeah you know what I'm going to do I'm going to actually switch my computer up here for the moment so I can plug it in I'm going to go off screen for a sec we've got success with I'm getting 20 on my initial track.”

Engaging with the Audience: Questions on Opus 4.7

57:37 to 59:59

Hosts invite audience questions about the latest updates and features of Opus 4.7.

“Anybody in chat have questions if they'd like to be tested out on our Opus 4.7 or on QN 3.6?”

Final Technical Adjustments and Screen Sharing

1:00:00 to 1:01:31

Hosts make final adjustments and prepare to share screens for the audience.

“I'm gonna rejoin this computer real fast.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:08Okay, sweet, we're live.

0:14All right, what's up everyone? I am here with my friend Kyle, who is a developer, and we are going to test out Opus 4.7, which just came out as well as on Kyle's side we just saw that there was a new open source release of Quen 3.6 so we're also going to be playing around with that so uh shout out in the comments uh if you have anything you'd like us to to test out but uh but we're going we're going live right now and we're gonna we're gonna try this um okay

1:05Okay. So, first and foremost, what I'm going to do is I'm going to pull up the Opus Details. I'm going to go ahead and grab that trail.

1:28Let's look at it while people load in.

1:37Okay. So first of all, for anyone joining us, welcome, welcome. So today you are graciously being hosted by myself and my good friend Kyle. Kyle, you're now live, by the way, if you want to say hi. All right. Hi, all you people. What's up? What's up? And we're going to be talking about Opus 4.7. Now, if you don't know what Opus 4.7 is, it's the latest version of Claude. Claude is made by Anthropic. and Anthropix, you could call it their like frontier models or their smartest model state of the art is their Opus series. And so the latest one before today was 4.6 and now we have 4.7, which just came out.

2:21And so we're looking at it right here and we're going to go through the benchmarks really quickly, but I watched a great video this morning that came out maybe like, I don't know, 20 minutes after they made this release that I'm also going to recommend people check out and I'll include a link to that in the comments where Nick, I don't know if I'm pronouncing his last name correctly, but Sarave goes through all of these benchmarks in more detail and kind of talks about them. So we'll let you watch that if you want to dive too deep into it. But the key thing to keep in mind here, so when you look at my screen, we have Opus 4.7 on the left hand side and right next to it is Opus 4.6.

3:03And then you also have GPT 5.4, which for For people who don't know, that's ChatGPT's best new model, at least as of this writing. And then Gemini 3.1 Pro, which is the same for Google. And then over here on the right-hand side, we have something called Mythos Preview. Now, if you haven't heard about Mythos, this is the model that basically is the one that scared the pants off of everybody who's used it. it can apparently find zero-day exploits in any piece of software on the internet, and Anthropic is withholding it from the general public while they test it and basically give banks and all the major tech companies access to it so they can sure up their defenses before they release it to people they don't trust.

3:51So fun stuff there. Yeah, we'll see. We'll see what happens there. Kyle, do you have any thoughts on Mythos? or just what you heard about it i mean it's you know it's always hard to say when you can't really use the thing the benchmarks look insane and i you know i have reason to believe that anthropic wouldn't really pull our leg on that so right i'm i'm treating it as a potential game changer but I don't hold my breath at this point ever. Yeah, I think that's fair. I think that's fair. And so the thing that's interesting about the video that I watched earlier that I mentioned from Nick is Nick basically thinks that Mythos is a distilled, or sorry, Nick thinks that Opus 4.7 is a distilled version of Mythos.

4:45And I think there's some credibility to that. And basically the reason he gives is that all of these benchmarks on the left-hand side here for opus 4.7 are higher than opus 4.6 um with a couple exceptions one of them is the uh the agentic search was one that stood out to me yeah agentic search is lower that's right yeah agentic search is lower and so potentially that could be something that they artificially nerfed, basically. And one of the reasons that Nick gave for this, which I think makes a lot of sense, is essentially that could be one of the capabilities that makes it easier to find those exploits that he was talking about, right?

5:35And the other one is agentic terminal coding. Now, agentic terminal coding is higher on 4.7 than 4.6, but not by much. it's like a little bit better. But he was saying that that's kind of like where a lot of the, like the agent's ability to use tools and bash commands and things like that help it hack systems. So that's why he thinks that that was also nerfed, which I think is a good call. But pretty much everything else here, I believe, is higher. Actually, cybersecurity vulnerability reproduction is lower. So that's another sign that doesn't... That's an interesting one. That feels too on the nose to be an accident.

6:19Yeah, I think that's why. And then also, there was something that was flagged here. I think it might be in the larger report by Lizanne Al-Gaivon X. And basically, the long context is actually worse on 4.7. It looks like the chart that he shared isn't on here. So let me see if I can pull it up.

6:48I'm on my personal computer right now, so that's what you're seeing. Well, let's see what notifications we get. I will say on their other benchmarks, one of the more impressive improvements is the visual reasoning. Seems like it got a huge jump up. So I have reason to believe they focused a lot on vision, which is something I've sort of seen them get a little dragged about recently. Oh, yeah. Yeah, it was actually kind of terrible. Like, I can't tell you how many times I'm using it, it being Opus 4.6 to code, you know, like a video game or something like that. And I tell it to fix something with the UI, and then it's like, looks clean, looks good, you know?

7:31And I look at it, and it's like pixel slop. Like, it looks terrible. I'll have to show you an example later. But it's just hilarious. Like, its vision capabilities were, in my opinion, very lacking. and I have to tell it multiple times. That's something I'm definitely looking forward to testing out a little bit. Yeah, the thing that Kyle's talking about, let's see, where is that? Vision, vision, vision, visual reasoning right here. So this is a huge jump up. So with no tools, it's a massive jump up and with tools, it's significantly better as well. So yeah, so and then this page goes on to talk about Project Glasswing, which Project Glasswing is what I hinted at earlier with Mythos, where basically they're working with a bunch of industry partners, sharing Mythos with them internally to shore up their systems.

8:22And so they're kind of talking about how they're working with their team there here. And then they've got some pull quotes from people who they've let test it out. And here's some of their highlights that they flag on their blog. So instruction following, They say Opus 4.7 is substantially better at following instructions. This means prompts written for earlier models can sometimes now produce unexpected results. So where previous models interpreted instructions loosely or skipped parts, 4.7 takes instructions literally. I kind of hate when this happens. I don't know if you feel the same way, Kyle.

9:02but I when when they all of a sudden become more literal it's like technically a good thing but then you have to be a lot more precise with what you're saying and especially if it struggles with long context tasks which I will pull that chart up in a minute you know that means you have to be a little more careful with what you write yeah it is an interesting thing that you could either you could interpret it as an improvement in the controllability of the model or you could view it as like a slight regression on its ability to pick up nuance yeah i think that's probably true especially if it if it is struggling with long context like that does feel like it's a bit of a regression to me um let me pull that shirt up because i think that was i flagged it immediately i was like oh this is yeah here it is so basically hold on we've got a we've got brazil mentioned in the chat what up hey brazil what's up brazil welcome to the show i can't let one of those go by yeah no i i appreciate that i appreciate that yeah thanks for um thanks for shouting out lol mao mythos preview destroys basically everything yeah pretty much at some point hackers are going to get aversions or someone else is going to make something like it then all the banks are screwed well hopefully not because anthropic's giving them all access to it now so they better use that thing like uh there's no tomorrow and sure it up um i don't know kyle what What do you think as a developer, is it possible to fix every bug?

10:42I don't think so. Okay, so there will always be an exploit somewhere in the code, do you think? I mean, we might get to some new paradigm someday, but it's not going to be before Opus 4.8 or 9. what what what the heck do we do then in that scenario like that's just wild to me that's it's gonna be an extremely active area of research i think it's it's just everything's gonna accelerate you're gonna have i don't know maybe you have to put an open claw at every network endpoint oh wow we'll see yeah that's that's wild i could i could see it though i could see it um thanks for the overview yeah what's up middleton middletown ohio is it middleton or middletown let me know um what up what up harmonious crow yeah appreciate it you're welcome you're welcome so here we go this is the one thing i wanted to flag because this concerned me and the reason this concerned me is that um i use a lot of long context like kyle was talking to you about this earlier this week.

11:53Like I'm, I'm a little bit, uh, undisciplined in terms of just throwing a bunch of stuff into my agent and being like, okay, now read this, now read this, now use this, um, and try to give it a lot of information for it to parse through. So for it to go from, you know, Opus 4.6 level, which here you can see, this is their 64k extended thinking version, um, has 78.3 % on this eight needle test. And I don't know exactly what the eight needle test is, but basically the needle test in general is how well can you pick out an important detail from a large, long context window. So the classic thing is like loss in the middle problem, right?

12:38Where it just loses details in the middle. It's good at keeping track of things that you say at the beginning and things that you say at the end, but in the middle, it just loses focus. Um, so to go. I still have the habit of, um, prompt at the start context in the middle reprompt at the end, which was totally necessary for the first long context models. The newer ones probably shouldn't need that, but it's just kind of baked into my brain now. No, I think that's a good, that's good policy. And I should probably do that as well because I became very undisciplined with 4.6. As you can see, like the, the best jump between 30, like, like this was like a, this was like a twofold, like this is double as good as like the next.

13:16ones and so to see it drop back down for three points or 4.7 uh is a little concerning so that makes me wonder i don't know should i switch to 4.7 maybe for some things but not for others yeah it could be a thing where you want where you want 4.6 to take your million context and bring it down to half a million so 4.7 can handle it better it'll yeah take some testing to see that honestly makes a lot of sense um what up sf bay area oh hey nice nice to nice to see you john peters let's see what's all who else is here chris taylor oh wow you got another uh person from middletown in the chat that's awesome um san diego land of enchantment don't know where that is but right on um all right so should we should we mess with this stuff should we should we put it to work I think we should yeah I think we should let's see what it can do all right so what I'm gonna do is so I was working on a project on my cloud code over here and I'll give you guys a little preview of what I've been working on um so this was a game that I've totally vibe coded with claude 4.6 opus so i have written no code on this whatsoever and it is basically i i gave it a couple prompts to start off with and then i've just been iterating on it as it goes so this whole thing i've been working with is opus 4.6 right um and so i've got it working on something right now if this is the first time you've ever seen claude code um in the desktop app um when it's working on with its like preview mode here this is what it looks like and this is nice because Claude can actually see the screen but this is what I was talking about earlier when um basically like it will like for example let me see I'm gonna hit new game here I'm gonna write um test and I'm gonna go open world so originally what so okay let me give you the pitch of this game actually So this is basically like a Final Fantasy Tactics style game set in Renaissance Italy, right?

15:33So I'm really trying to go for the Final Fantasy Tactics vibe and aesthetic. That's a turn-based strategy for anyone who hasn't played that. I can't think of a bigger example. Yeah, yeah. No, I think that's probably the best way to put it. And I'll show people what it looks like right now. But I was trying to get it to set up this map, right? and originally it was making a map of Italy it was like the most complete garbage thing I've ever seen and it's like this looks great it looks just like Italy and basically now what I'm gonna do is I'm gonna try and test it out and see how well it does with the visual the visual reasoning now that it knows well how to basically how to see better okay so we got people from New Mexico We've got people from Newcastle, Australia.

16:26What's up? So basically, as the models improve, they anticipate a need for prompts and sort of build them right in. Yeah, I think that's true. So basically what you're saying is as the models improve, they themselves anticipate the need for prompts and build them into the system. Do you think that's what they're doing, Kyle? I would say there's something close to this where they're generally taking the ways that people are getting the most out of the models, and then they're building that into the subsequent models. Like before we got internal reasoning in models, the big trend was that you would just tell the model, think out loud, write a plan, think it through, and then respond.

17:12And so then the new models, that just became an internal functionality. And then they're also, you know, focusing on what fields people are using them the most in, and they are trying to expand their capabilities in those specific fields. As for like specific prompts, that's less baked in. Yeah, yeah, I think that's right. So what's interesting about this is I don't have in this window, I don't have 4.7. available to me. So what I'm going to do is I'm going to start a new session. I wonder if the app needs an update. Yeah, let me relaunch it real fast. Probably, that's my guess.

18:06Used up the limit on a deep research and it session limited you. Yeah, well, tell them what happened to you the second you tried 4.7 this morning. yeah i opened it up i gave it one prompt to make a svg of a of a rising phoenix and it got perfectly through it but then it just died immediately at the end because i had already been cranking 4.6 so much this morning so i i have heard from a lot of feedback on x and maybe i'll pull up some of those posts later that it does burn tokens like a lot more interesting yeah it's the the pricing is the same point as 4.6 that's right um but there's always you know different models will prefer to reason a different amount so you can sometimes get a model that costs the same amount but you ask it a question it thinks twice as long and it costs you sort of twice as much yeah oh cool and we do have 4.7 now i see somebody asking to ask questions just ask the questions yeah yeah you can ask questions go for it um does it use usage quickly i think we kind of just covered that um yeah and uh so if you guys have a prompt or something you want to see let's let's test it out uh right now what i am going to do is i'm going to test it on this renaissance bomb fantasy tactics game so um i'm going to say uh what were we last working on per your last plan.

19:37Also, did you notice that they changed the desktop UI recently? Do you use it or are you using it mostly in the terminal? I'm pretty much either quad code or I'm in DS code using one of the various integrations into that. Yeah, yeah. I actually switched between the two. For this one, I pretty much only worked on it in the desktop version just because it has that nice side view where you can actually preview what you're working on. but for other stuff I'll use a terminal emulator like Ghosty somebody asking if ChadGPT 5.5 is coming I have heard no sign of it but I will confidently say yes so there was rumors that Spud was coming soon Spud is like the internal name for whatever 5.5 will be and in the case of if they're launching it today the only thing i've seen is that tibo or tibu uh from the codex team said i'm feeling codex-y like 40 minutes ago so that could mean that something is launching but um all i know is these guys don't like to let any other company have the spotlight for too long so exactly exactly um great let's keep working on that

20:57All right, so I'm going to let that crank. Let's see. What does it excel at? We just did a quick overview of that, but while I'm waiting for this to kick off, I'll go back to Claude Opus 4.7. This is what the Anthropic team says it excels at. So instruction following, which we covered earlier, improved multimodal support, which is the vision thing that Kyle mentioned earlier. It has better vision for high resolution images. It can accept images up to 2500 pixels, more than three times as many as prior cloud models, and this opens up a wealth of multimodal uses, so that's good to know. For real world work, as well as its state of the art score on the finance agent evaluation, which we did not talk about, but it's good at finance.

21:46Let's see. The internal tests show that Opus 4.7 to be a more effective, more professional, sorry, more effective finance analyst than 4.6, producing rigorous analysis and models, more professional presentations, and tighter integration across tasks. 4.7 is also state-of-the-art on GDPVal, which is a third-party evaluation of economically valuable knowledge and does doing work across finance, legal, and other domains. Also, it's better at using file system-based memories. So it remembers important notes across long multi-session work. So maybe it's worse at long context, but it's better at like pulling in context strategically.

22:27You think there's a thing to that?

22:32That could be the case. Yeah, you might be able to have it work with your file system and then more selectively pull in the context and be able to get the correct information without bloating it. But it does seem like that context fall off is pretty bad. That's one thing I might expect them to change in a later push. Yeah. Yeah. Especially if people blow the whistle on it like I'm doing. Cool. So let's see what's going on in the chat. Sweet, thanks for the overview. What do you guys see using 4.7 for though? Oh, yeah. So I'm going to test it on this coding use case right now. effectively complete, no audio assets, no separate main menu, procedural art only for phase 7.5.

23:22I'm going to say read the whole code base and confirm those are true. Because sometimes what will happen is when it looks at an old plan, it will make conclusions, but then it will go read the code base. It'll be like, oh, you're actually done with that. So it's good to tell it to just like read, like read the whole plan. Yeah, I see some mention of haiku up at the top there too. It sounds like it might have farmed some of that out to a sub-agent. Yeah, that's one thing that this one does a lot of, it seems like.

23:58Which I don't really like that. I'd rather you know if you're going to read something and you're going to make conclusions about it, use the smartest version available. Why are you... I think it's configurable. okay um i do something a bit similar to that where i try to use a local model to summarize short bits wherever i can to kind of save my expensive tokens yeah but yeah there's there are some times where you just it's you would rather have it right than have to correct it yeah no we were talking about this earlier where basically i can uh be a bit of a token waster at the moment because I'm on the max plan whereas you're much more efficient because you are not on the max plan and I think that there's good practices to that that everyone should use.

24:54Somebody's asking what we see ourselves using 4.74. Yeah you want to answer that? Like it sounds like for you the answer is almost everything. Oh yeah well I yes so as I just said with the max plan, I do tend to use the highest intelligence version for as much as possible, because if I'm not doing it myself, then I want the system I'm offloading it to, to be as close to like, like doing it perfectly as possible. Right. Um, but that's, I guess, a privileged position. Yeah, I personally, I'm a little more shifting my workloads. Kind of a common workflow I've been following lately is that I'll start with a state-of-the-art model, maybe 4.7 or Gemini 3.1 Pro.

25:47And I'll talk with that for a bit to kind of form a plan and then ask it to make a really thorough plan. And then I'll pass that to a, like a coding agent model. Usually it'll be like a medium one, maybe a Gemini Flash or a Sonnet, and I'll see how far that can go. And then usually the 3.1 Pro or the 4.7 now will come in to kind of clean up the last bits. But that way you use the smartest model to get the plan. You build a lot of the context, you know, doing the cheap, just look up this file, write the simple functions. And then now that all that context is there, it only takes one quick blow from the from the state-of-the-art model to clean it up yeah i like that i i think that's great um i should be doing that honestly uh but not if you have plenty of max yeah yeah well someone just said you guys should give out max plans on the newsletter totally i'm gonna reach out to anthropic and see if they'll do that the problem is i feel like they don't want people to buy max plans anymore because when you max them out it can cost them a lot of money so i don't know if they're gonna try and uh i don't know if they'd be down to give them out but i like that idea we should try it um when do you think air will start to be funny not like forced funny but integrated into how it communicates um i think i think it's funny like one percent of the time right now like every once in a while it'll squeeze in a little zinger that gets a real like like a heh out of me yeah I think that's accurate I will say I do use Claude when I'm writing the newsletter and sometimes it comes up with pretty good jokes so it's not it's not uh it's not not funny but it has to be in the right context I think

27:41let's see here another YouTube channel right now is about to have six instances open at the same time. That's wild. Oh, look at this! SubAgent was wrong! Audio Assets do exist! Okay. All right, we like self-correction. Let's say we're gonna... This is the last message I had queued. And I'm gonna send that to it. Let's see.

28:23Meanwhile, so I have my terminal agent over here. Let me see if the model actually shows up. Not yet. I might have to restart it.

28:43So this is another thing I was working on. For people who don't know, so a couple of weeks ago, we interviewed a company called LTX. And LTX created this tool that was called LTX Desktop. And it's like a desktop app where you can essentially generate your own images or videos and then send them to a timeline and edit them. And what I've been doing is I've been expanding it because it's open source. so i've been adding like i added a director tab where you can enforce a style across all of the shots that you generate you can give it a mood board you can give it a cast list so anytime this character's name shows up this is the guy that you see a location registry so it's the same for a location and then a shot list card and the main thing i wanted to do was create a version where you could essentially as you're writing the script start generating shots and you could have those shots kind of load in as you're writing the script so you can sort of see what it looks like I think this is the way that generative AI in film should work where you're basically like writing it from the screenplay standpoint and you set up all of this system around it to enforce it so that it essentially like enforces your style across the entire script.

30:09And then, you know, there's all these sort of add-on things I've added since then, like an outline mode with a lot of different flows. I added a shot list where you can see all of your shots. In this case, I didn't generate them yet. Storyboard mode, so you can just see just the storyboard as you go. A whiteboard where you can just do ideas. And what I'm working on now is a production mode where you can kind of see the whole script and watch it over here, but that's still in progress. And then I also added a little agent mode here where you can actually talk to the agent and have it do some of the some of the things you want it to do for you.

30:52So I'm going to try and work on that over here with 4.7 if I can pull that up.

31:05It's got you on 4.6 right now, I think. Yeah, so I'm going to have it. So it kind of laid out a basic plan for me here. So I'm going to make it do a plan in plan mode. Or you know what I could do? What I could do is I could copy this. Like, hey, this is what we're working on. Let's make a plan for it. So let me do that. While that's cooking, I'm going to open a new window.

31:41Do you ever run it with skip permissions? Not full YOLO mode. Do you do that?

31:54Oh, yeah. I made an alias on my bash to where I just type YOLO and it comes up in YOLO mode. That's awesome. I do it mostly on my on my Linux laptop that doesn't have all of my life on it yeah good call good call yeah I have a I have a PC where I can do that let's see what's going on in the chat oh yeah I I actually I've been wanting to um share the link to that um but I just wanted to get it to a point where it was like actually ready to ready to share with people um but I'll do that once I do the blog version of this I'll share a link to it let's see someone says I feel like we all just a bunch of idiots trying to create stuff from every format release that's true cloud is a public company cannot have millions of people writing their own I mean that's an interesting angle basically yeah like when when will they pull the plug on solo developers being able to use cloud models to build competitors to like multi-billion dollar SaaS companies I don't know I don't know if it was ever is ever in their interest to to do that but potentially they will just create products where you just run them and it will do exactly what they designed it to do that's a concern for sure where basically they just are saying okay you don't get access to this anymore we're gonna we're gonna tell you what you get access to do Do you think they would ever do that, Kyle?

Read the full transcript

33:24I don't know. I think the cat might be sort of out of the bag. I mean, they could hold on to their mythos, but as soon as a big model is even somewhat available, people are going to start getting distilling data out of it and trying to figure out what ideas they're using. And it's only a matter of time before you have whatever, a GLM or a Xiaomi or ByteDance trying to do their best to catch up. um yeah yeah i see you know there are these giant companies that are also utilizing these tools and to try to solo dev past them is sometimes an insane notion to think that one guy is going to make photoshop or whatever and it's going to win but i think what's happening sort of is that there's an increasingly long tail of more and more specific products that a solo developer can make and like anthropic could make it too but they just don't it's not worth a large company's time or their effort or their care about such a niche little industry right so i think solo developers will proliferate more into more and more uh specific stuff like software made for one person or one small company yeah i think like we did a live stream i don't think you saw this one a couple I think it's a month or so ago now with Ryan Carson, who he previously had a startup where he taught 1 million people how to code.

34:57And his take was, DHH told him that all you need for a successful business is to make a product that 250 people will pay X amount to get$5 ,000 a month. And as a solo developer, that's like a good salary, right? I mean, maybe not in this economy, but maybe you increase those numbers a little bit. But if you can create a niche product that works for a small group of people, it can still be worth it if it's like the perfect product for them, right? Which I think is an interesting framing where we don't all have to work for a big company in order to have a product that can actually succeed. Do you see any divisions in the future between how Anthropic treats developers using cloud code and day-to-day development versus non-dev consumers using it to expand their building capabilities?

35:51Let me think about that for a minute.

35:57I think they're already different. There are divisions. I don't know how much they're trying to make it like a big tent like hey everyone's welcome like come in and start coding with this I do think that they are gonna have different product paths in fact let me double check there was supposed to be a new release from them that was sort of like their lovable killer if you guys remember that I'll drop it let's see

36:37yeah i don't think that came out today but there was some there was some controversy yesterday because basically mike krieger who was formerly of instagram uh he was on the board of figma and he left the board of figma and i think that had something to do with this new product that they were going to launch and rumor was that it was going to launch today but maybe there's more behind the scenes scuffle there. They decided not to step on the launch of 4.7 or maybe it's coming later today. I don't know. But yeah, there was supposed to be something about that. Let me see if I can pull up that tweet.

37:13Maybe they're waiting for GPT 5.5. Yeah. Do we ever answer that question? So it's imminent, but maybe not today. I think the rumors have subsided that it'll be today. But there might be a codex version. What was that?

37:40Yeah, okay. So this is about the design tool. So a new prompt-based design tool for websites and presentations this week. But I don't know. They haven't announced it yet today, so we'll see.

37:58Didn't Ben Affleck just make big money selling a movie program? Oh, that's funny. Bro, just test it. Everyone can see the benchmarks. Okay. Yep. Yep. We're testing it. Okay. So let's see here. The plan, agent produce concrete file. Here it is. I'll zoom in on this a bit.

38:32okay so here's what's happening so I'm trying to I'm trying to edit this production mode here and basically what I want to do is I want to make it so it will read along as you go through the script and I've got a version of that right now but I'm adding some improvements to it to make it work the way that I want it to work so let's see what happens

39:10All right, so it's got its plan. It's adding narration type, audio production mode, it's refactoring it. Basically what I'm doing is I'm trying to separate the different types of audio that it will generate. So it can generate background music, it can generate a narrator if the person just wants to have like the narrator read the script back to them and listen to it back and it can generate dialogue and music and sound effects on a different on a different audio track basically. And so this is an electron app so it updates automatically as it's working on it over here which is nice.

39:55You know what else we can do which I meant to do earlier is

40:05I also wanted to add a new endpoint for this so I'm going to do that

40:13so for people who don't know this is Gemini 3.1 Flash text-to-speech so this is a new AI speech model from Gemini and I've already got Gemini endpoints on this tool so I'm going go ahead and add this but I need to go to the docs.

40:35Here what I'll do is I will go to

40:45yeah this is what I need. Copy this as Markdown. I'm going to allow all edits for this session so it'll just automatically start editing. Also we want to add this new endpoint from Gemini.

41:06Go ahead and paste that in. So while this is cooking I'm gonna go back to when I was playing with Renaissance here and I'm gonna say

41:25I wonder if this version now, yeah, this version now has Opus 4.7, so I'm going to work on this because this is where the whole history is.

41:37I'm still seeing if I can get this QN 3.6 going. It looks like Llama is picking it up, but I need to do something to make it register in the UI. Nice, nice. But it does seem like it's fitting into the 3090s 24 gigs with that 35BA3B. Sick. Let me pull that up. Let me pull up the details on that for people who don't know about this. So what I'm working on while Kyle is working on something else is Opus 4.7. But Kyle is working on Quim 3.6, which just came out as open source today. So for people who don't know, open source means you can run it on your own servers or your own GPU if it's small enough.

42:23And, you know, the GPU is a chip on your computer, you know, if you have one. Otherwise, you have a CPU. And you said you had the 3090 from NVIDIA and you're running the 27B version. Is that what you said? This one you have up the 35B A3B. Got it. yeah so this refers to the number of parameters and i guess this is just like the the the 35b is the total number of parameters in the model but it's a mixture of experts so the a3d is a active three billion so you get sort of the intelligence of a 35 billion with the speed of a three billion It's not exactly as smart or as fast, but it's kind of a best of both worlds is what I've found.

43:12I've been playing around a lot with the Gemma 4 MOE recently. That's, I think that's 27BA4B. And it's crazy. It goes so much faster than their 31B dense model. And it seems almost exactly as smart. so I'm very uh very bullish on MOE right now oh my gosh there's already another there's already another app update oh my gosh okay let me do that right now although my computer's lagging over here I'm gonna go back to this tab so yeah let nobody say that Anthropic is not busy working that's for sure that is totally for sure um someone in chat saying uh final fantasy yes that is he is making final fantasy tactics but renaissance right yep yeah and it's actually going to be portable so anyone can make anything that they want on it um it's just for fun not gonna monetize it or anything but uh what is happening here Okay, let's see.

44:27Is this the one I was just working on? Yeah, I think it was. But I'll make it open source so people can play it. All right, so let's see here.

44:44It's hard to get the window the way I want it right now.

44:50okay yeah so let me show you actually what this looks like

45:00so all this is written by the ai so if the script sucks it's because i haven't edited it or anything yet

45:09let's see i think i'm gonna do full screen with this right

45:26you could probably go to that url in a browser yeah presumably works yep

45:36so all right so we've got the characters so this is what i'm talking about earlier if y 'all can see this so basically off i i was trying to get claude to you know procedurally generate a renaissance style portrait here and it was like yeah this looks great and as you can see it's this ridiculous looking guy in the corner here um so one of the things i'm looking forward to trying to fix is now that it has much better visual reasoning in this new version to try and actually get these faces to look a little bit more like Renaissance style. Okay, so basically the bad guys show up and then you're going to fight them and so you can pick where you want to position your characters like FF Tactics, then you hit play and then you can move them around the grid and then you have actions.

46:28So in this case I'm going to poison drop this guy and I can choose which direction I want to face. I'm going to face this way because they're going attack me. But anyway, what you really want to see is you really want to see how well the coding agent improves this game, right? So I'm going to say, hey, let's work on improving the art system. Right now, I think the characters look too bulky and not what I envision for a renaissance themed art style. Can you use your vision capabilities to review the portraits and the sprite art and make them better match my original vision in the game plan that we're working on.

47:29Let's see if we can improve this. Kyle, when you have something you want to share on Quinn, let me know. Yeah, still trying to get it registered locally. It's downloaded. It even seems like it loads up correctly, but I feel like it might be just that it's so new that Llama needs some little change or I might need to add a little flag. Right. Right. Yeah. I'll let you know if I get it. Yeah. Now, one of the things I want to talk about that is pretty cool that Anthropic launched recently is Routines. So I don't know if y 'all can see this on the left-hand corner, but Routines is sort of like, because someone was talking about how like, oh, they love games, but they're building a business right now.

48:14Well, Routines is basically where you can like build yourself agentic workflows directly inside Cloud Code. And right now, I don't think it's in co-work. There's something called scheduled tasks, which is sort of similar. But routines are essentially the cloud code equivalent. So just to catch people up, if you're just joining, we've got two sessions going on right now. We've got one in the terminal over here. And that's what this is. This is a terminal. And so I've got it working on my project. Looks like it wired up Gemini 3.1 Flash text-to-speech. So that's cool. And I've also got the desktop app going, and that's where I'm working on the Final Fantasy Tactics project.

48:55So what's cool about this is you can start a new routine. You can do it locally or remote. So I'll do local. You can name it, give it a description. You can determine whether or not it needs to accept edits automatically or work in plan mode. You can set the frequency and daily. And what you can do with this is essentially make like agent workflows like you would in N8N or make.com. But you can do it just in Claude, which is pretty awesome. And we wrote about this this morning in the newsletter for a little bit more about how this works. So you can check that out there. But it's pretty cool. Let's see what people are saying.

49:38They released one month after 4.6. Isn't that too fast? Yeah, dude. It's all too fast. they're like constantly trying to one-up each other and solve these problems and like maintain people's attention and uh like basically control the future of the world from a certain perspective like if you believe that AGI will essentially like become the dominant like controlling force which I don't know if I 100 % buy into that but yeah they just believe that like whoever creates like quote unquote the AGI model is going to win this whole thing and to the victor go the spoils so okay this I think tabs context MCP I don't know what that is what is it so this is something for people who our new urge programming like me, I'll say like what is the tabs context MCP and why should I allow it?

50:47And this is a good trick just so that if you don't know what you're signing up for, usually I'll deny it and then I'll ask the question. It will then tell me what it is and why I should use it.

51:10let's see having scary dangerous mythos and not qa their 4.6 and it downgrade level of thinking with bugs not reading files that the one prompts to fix all strange um i'm not sure i 100 percent understand that comment i think they're talking about the the noted reduction in performance that some of the like the amd employee years or whoever that was said various people have been calling out degradation in their own workflows with 4.6 right before this yeah this always happens and i don't know if it's intentional or because they're moving compute to prioritize the new model but like the week before the old model just gets like way worse and you know if you're conspiracy theorists you You say they do it intentionally, sort of like out of the Apple playbook.

52:02Somebody commented on that this morning, where you intentionally downgrade the previous version to make the new version look better. But it could just be that they're just not prioritizing it or using a quantized version of it or something in the interim. But yeah, I don't know if that's going to continue forever for 4.6. If 4.6 is just always nerfed now, or what? but yeah hi i just got here you making a game yes i was annoyed with opus 4.6 drop we should be refunded for the month it dipped as it's not what we paid for yeah man they uh i think they gave you like a comped uh extra usage bonus for your subscription the price you pay for a month they did that when they introduced the the third party harness extra usage charge right but it did also kind of coincide with the 4.6 change.

52:58Yeah. Yeah. So basically, I mean, they've just been annoying everybody lately. So you're right to be annoyed. Yeah. All right. So, okay. So let's see what it said. So I asked it if I should use tab context. Tab context is part of Cloud and Chrome MCP extension. Okay. So first call you make to use Chrome browser automation tools looks for a Okay, what I was planning to do with it, open the game, navigate to the party screen. Okay, I'm going to say, you have the game open in preview. Can't you see that?

53:38Because basically what it's going to do is going to open it in Chrome, which we can do. If it gives me an answer, I'll do that. Let's see if this works. Okay, so what changed? Audio shot type gained narration, new audio production, audio drama. I'm already thinking of something I should do to improve this. Use audio playback, narration in sequential, narration is sequential. Gain node and volume default added. That's good. Let's see, decision taken, all reversible later, hard-coded narrator voice. Okay, so first of all, we should have a drop down to select the narrator, not hard code it. Second, did we look up the docs for the TTS models we have in our repo, or I'll say in our system, to optimize the prompts for audio?

54:39If not, let's go do that now. All right, so now I'll have it do that. And then I'm also going to say, also, give me a series of steps to test this and make sure it works. Okay, so this is interesting. So it's got sub-agents here. I want to try and test some of the newer things that it just launched. So let me go ahead and dig into this.

55:23Oh yeah, it's got extra high effort. A new level between high and max. So the way that you can control effort when you're using it in the... Sorry, it's going to ask me to approve a bunch of searches right now. Okay, so the way that you can use effort when you're in the terminal is you can do backslash effort. My computer's really slow at the moment.

55:52Okay.

55:58And then you can do... Yeah, I pressed the wrong thing. Sorry. I'm working. yeah

56:13you know what I'm going to do I'm going to actually switch my computer up here for the moment so I can plug it in I'm going to go off screen for a sec

56:32we've got success with

56:41I'm getting 20 on my initial track.

56:54No, it's definitely live. I believe it's some difficult questions now.

57:11Do you want to share your screen while you're doing that? Yeah, I'll see if I can get it set up for that. I've got one sec guys, I'm just reseting my computer here. Thanks for hanging in there. Okay.

57:37Okay, it should be up and rolling.

57:56Thank you.

58:20Anybody in chat have questions if they'd like to be tested out on our Opus 4.7 or on QN 3.6? Drop them in the chat.

58:37We got a question.

58:53Yeah.

59:05Plus that.

59:18Is that about accurate? Say that again, sorry. Someone was asking about the text. And I was saying, my guess is that...

59:41Jeez, I keep losing my connection here.

59:49Hold on one sec.

1:00:03I'm gonna rejoin this computer real fast.

1:00:30Thank you.

1:01:04Okay, I'm going to go back to share my screen.

1:01:14And Kyle, did you share your screen? Did you request it? No, I didn't

1:01:23Okay

1:01:28I'm tearing myself twice

From the publisher

Grant and Kyle dive into a comprehensive review and live test of the newly released Claude Opus 4.7, a cutting-edge large language model. This session explores its capabilities for coding and game dev, specifically referencing the "Renaissance / Plan Final Fantasy Tactics RPG Game" project. Discover how this ai model performs under pressure and its potential impact on game design workflows.


🔴 LIVE at 9:30AM PT / 12:30PM ET


Anthropic just dropped Claude Opus 4.7, and we’re putting it through the gauntlet in real time.


Join Grant Harvey (Lead Writer at The Neuron) for an unscripted, warts-and-all test of Anthropic’s newest flagship model.


What we’re testing

- Advanced coding on tasks Opus 4.6 struggled with

- New higher-resolution vision support for images up to ~3.75 megapixels

- File system-based memory across multi-session work

- The new xhigh effort level, which sits between high and max

- Claude Code’s new /ultrareview slash command

- Auto mode for longer, less-interrupted agent runs


Why this matters

Opus 4.7 is the first model Anthropic is releasing with its new automatic cyber safeguards, following last week’s Project Glasswing announcement.


It’s also the direct upgrade path from Opus 4.6 at the same price:

- $5 per million input tokens

- $25 per million output tokens


If you build on Claude, this is likely the model you’ll be using next.


What’s changing under the hood

- New tokenizer, where the same input can map to more tokens depending on content type, roughly 1.0x to 1.35x

- State-of-the-art score on GDPval-AA, a third-party evaluation of economically valuable knowledge work

- Better instruction following, which means prompts written for earlier models may now behave differently

- Improvements across finance agent evals, document reasoning, and long-context tasks


Bring your hardest prompts. We’ll run them live and show you what breaks, what shines, and whether it’s worth migrating today.


Watch part two, where Grant covers Codex for (almost) anything: https://youtube.com/live/OiRkwm3-og0


📰 Full writeup in tomorrow’s newsletter:

🐱 Subscribe to The Neuron (700K+ readers): https://www.theneuron.ai

More from The Neuron: AI Explained

All 106 episodes
BONUS: LIVE: Claude Opus 4.7 Just Dropped. Here's What Actually Changed.The Neuron: AI Explained · 1 h 2 min
Listen in VO