BONUS: New GPT Memory Feature, GPT-5.6 Rumors, Hermes Desktop Agent, New Codex Plugins, MAI-2.5 Image, Etc.

5 Jun 2026 · 2 h 2 min · 55 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

A rapid-fire “bonus” roundup of major AI/product updates: OpenAI’s new ChatGPT memory system (built on “Dreaming”), rumors of a new OpenAI model (“Mercury Alpha” / “GPT-5.6”), Microsoft’s Codex updates (role-based plugins and “sites” for in-app deployment), Hermes Desktop (a new favorite tool), and new model releases including Microsoft MAI-2.5 image and NVIDIA Nemotron 3 Ultra. They also discuss open-source Gemma 12B running on ~8GB RAM via 2-bit quantization, and a separate Anthropic blog on “recursive self-improvement.”

Guests

No guest bios are provided in the transcript. The episode appears to be hosted by Grant and Corey (and sometimes another participant is referenced, but no names/backgrounds are given).

Key claims

  • ChatGPT memory is now “significantly more capable” and compute-efficient, with memories synthesized by Dreaming and reviewable/correctable in a memory summary page.
  • Codex now supports “sites” (create/share an HTML-like site via URL) plus six role-specific plugins (analytics, creative production, sales, product design, etc.).
  • Rumors suggest “Mercury Alpha” / “GPT-5.6” could arrive next week.
  • MAI-2.5 image quality is strong enough to challenge top image models; Nemotron 3 Ultra is NVIDIA’s best model to date.

Notable examples

  • Live image generation: “oil painting of a dog… rain is meatballs” and a photorealistic retrial; they note improved realism but persistent “uncanny valley”/face artifacts.
  • Codex demo: generating a site inside Codex and using plugins connected to Snowflake/Databricks/Tableau and tools like Figma/Canva/Slack/Salesforce.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Recap of Microsoft Build

0:44 to 2:08

Discussion of highlights from Microsoft Build and new hardware.

“Hope everyone's doing good and having a great week.”

Interview Insights with Mustafa Suleyman

2:08 to 4:20

Insights on super intelligence from an interview with Microsoft's AI CEO.

“And I think it's mostly got to do with the NVIDIA chip.”

Rumored New AI Model and OpenAI Memory Feature

4:20 to 6:22

Exploring rumors about a new AI model and OpenAI's memory feature.

“We'll make sure everybody gets that too.”

Introducing Hermes Desktop and Codex Updates

6:22 to 8:12

Discussion on Hermes Desktop and updates on Codex tools.

“We are going to be talking about the rumored new model, which I have a feeling it's not going to come out today, but we'll see.”

Nematron 3 Ultra Review

8:12 to 9:14

Discussion about the Nematron 3 Ultra and its performance.

“It also just came out today, so I'm unclear.”

NVIDIA's Latest Innovations

9:14 to 14:00

Insight into NVIDIA's new AI models and cloud hosting capabilities.

“So if you're not familiar with Unsloth, this is a company that makes smaller versions of models that you can run on your computer.”

Discussion on Image Quality Testing

14:00 to 14:40

The hosts discuss the challenges of testing image quality and tools.

Live Demo of Image Generation

14:40 to 16:40

A live demonstration of generating images using Microsoft's new model.

“Someone in the comments said our microphone is too sensitive.”

Creating Unique Image Prompts

16:40 to 18:20

The hosts experiment with unique image prompts, including playful elements.

“Cloudy with a chance of meatballs, right?”

Evaluating Image Outputs

18:20 to 20:00

Hosts evaluate the quality of generated images and discuss realism.

“Overall, I've been really impressed with their image models.”
Show all 55 chapters

Exploring Variations in Image Generation

20:00 to 21:40

Discussion on the importance of starting new chats for different image outputs.

“So we're going to go back and do this a little differently now.”

Testing Microsoft's Latest AI Models

21:40 to 23:20

The hosts shift focus to testing Microsoft's new AI models and interfaces.

“I've got to stop sharing and go back in.”

Creating Complex Image Prompts

23:20 to 25:50

Hosts create complex prompts to generate photorealistic images with AI.

“What's interesting, it looks exactly the same, which is very interesting.”

Feedback on Generated Images

25:50 to 27:40

Discussion evaluating the photorealism of generated images and their details.

“Okay, that's way bigger than I should be zooming in on this.”

New AI Developments and Insights

27:40 to 28:00

Discussing new AI developments and insights from industry insiders.

“I mean, I don't know him, but like, I think we follow each other on Twitter.”

Introducing Mercury and Jewel Alpha

28:00 to 29:19

Learn about the upcoming Mercury Alpha and Jewel Alpha models from OpenAI.

“But one of the vague post kings of X.com.”

Understanding Model Parameters and Iteration

29:20 to 31:35

Discover how OpenAI optimizes model performance through parameter adjustments.

“and then go through and can cherry pick like, oh, well, this set of data did better than this trillion parameters.”

Codex Updates and User Growth

31:36 to 34:34

Discuss the latest Codex features and its growing user base across various fields.

“and I'm going to pull that up real fast if this is the right post.”

New Plugins for Codex and Their Impact

34:35 to 37:14

Explore the introduction of new plugins for Codex and how they enhance functionality.

“and I actually use HTML when I write the newsletter so that I can actually have it hyperlink my links.”

Rebranding and User Experience Challenges

37:15 to 39:37

Analyze the implications of branding for Codex and user experience improvements.

“we're really I think we're, I think I'm going to say by 4th of July that this will be the way you use chat GPT for the most part.”

Interface Simplification in AI Applications

39:38 to 42:01

Learn how recent interface changes in AI applications aim to enhance usability.

“They've sort of unified it into one app.”

Exploring ChatGPT's New Modes

42:01 to 43:24

Learn about the new simplified interface in ChatGPT's application.

“And I think, you know, whether they stick with dual alpha or not, you know, probably won't.”

Introducing ChatGPT Memory Features

43:24 to 45:34

Discover the new memory capabilities of ChatGPT that enhance user experience.

“However, it historically was never sufficient as a standalone memory system.”

Windows OS Integrates AI Skills

45:34 to 47:41

Understand the significance of Windows adding AI skills at the OS level.

“I solemnly swear I will not use this in the newsletter.”

AI Personification and User Interaction

47:41 to 49:53

Explore the implications of AI personifying objects in conversation.

“Yeah, it was a stupid brand to begin with.”

Anthropic's Recursive Self-Improvement

49:53 to 53:18

Learn about Anthropic's advancements in AI's self-improvement capabilities.

“Piers Anthony is somewhat problematic, so I don't want to get into him, but these books are brilliant.”

Anthropic's AI Code Contributions

53:40 to 56:00

Examine the growing role of AI in coding and development at Anthropic.

“I want to highlight another thing that came out today.”

Claude Code Session Success Rate Analysis

56:00 to 58:20

Learn about the success metrics used for Claude AI's code sessions and task complexity.

“Each bar is the average over the days in that quarter of lines of code merged per active contributor shown as a multiple of the pre-2025 average.”

Future Scenarios for AI Development

58:20 to 1:00:00

Explore potential future trajectories for AI advancements and their implications.

“By solving the verifiable areas, the others will come into focus after that.”

Energy Infrastructure and AI

1:00:00 to 1:02:50

Discuss the challenges and potential solutions regarding energy supply for AI progress.

“And then it kind of talks a little bit more about that.”

Grid Limitations and Power Distribution

1:02:50 to 1:05:30

Understand the complexities of power distribution and the impact of outdated infrastructure.

“You know, some of it's at least 25 years old.”

The Impact of AI on Hardware Prices

1:05:30 to 1:08:00

Examine how AI demand influences the pricing of graphics cards and RAM.

“And it's a thing that's needed not just for AI.”

Microsoft's Fairwater Data Center

1:08:00 to 1:10:03

Learn about Microsoft's eco-friendly data center and the community backlash it faces.

“has gone from 2 to 2500 to 4500 to 5k i have because i have been for a good number of months eyeballing an RTX 6000, the Blackwell.”

The Fairwater Data Center: Innovation and Community

1:10:03 to 1:13:24

Learn about the environmentally friendly initiatives and community engagement of Microsoft's Fairwater Data Center.

“For anyone who hasn't looked up, Grant, would you pull up Microsoft's Fairwater Data Center?”

Public Perception and Concerns about Data Centers

1:13:24 to 1:16:42

Explore the public backlash against data centers and the legitimacy of their concerns.

“you know, some early ones maybe felt a little to people like they were sort of being snuck in the back door.”

Harnessing Natural Resources: Landfills and Methane

1:16:42 to 1:19:52

Discuss the potential of utilizing methane from landfills to power data centers.

“The landfill is creating a pretty much unlimited around-the-clock supply of methane gas.”

Introducing Hermes Desktop: A New AI Agent Framework

1:19:52 to 1:23:39

Discover the features and user experience of Hermes Desktop, a new AI agent framework.

“They put a lot of their own taste into it.”

Comparing AI Tools: Hermes and OpenClaw

1:23:39 to 1:24:00

Learn about the similarities and differences between Hermes Desktop and OpenClaw.

Exploring Automation and Skills

1:24:00 to 1:25:18

Learn about setting up automations and the importance of consistent skills across AI models.

“which shows you like the types of things it can do.”

Managing Skill Versions Effectively

1:25:18 to 1:27:13

Discover the challenges of updating AI skills and the importance of version control.

“I think probably the hardest part about skills is that you'll be updating them frequently, like at least you should be.”

Codex App and Plugin Integration

1:27:13 to 1:28:48

Dive into the features of the Codex app and the new plugins available for users.

“But do you use a Google Drive system to manage yours?”

Using Codex for Data Visualization

1:28:48 to 1:30:02

Learn how to utilize Codex for building interactive data visualizations.

“And I think they talked about some of the new.”

Interview Insights with Scott Hanselman

1:30:02 to 1:32:07

Explore innovative AI engineering projects discussed in an interview with Scott Hanselman.

“One of these times I need to show how I had Codex visualize stuff.”

Exploring Claude Design

1:32:07 to 1:35:31

Find out about Claude Design and its capabilities for creating front-end interfaces.

“And he built an app that talks to the meter all day and builds charts of his performance and sends notifications to his phone to let him know what this needs to be doing.”

Mocking Up Interfaces with AI

1:35:31 to 1:38:00

Learn about creating mock-ups and integrating AI tools for design implementation.

“You can then connect it to your GitHub repo, so it actually reads your code base and understands how everything works.”

Designing UI with AI Assistance

1:38:00 to 1:39:25

Learn how AI can aid in UI design processes from concept to implementation.

“So here's my old version of my thing and what it looks like.”

Human Verification with Tools for Humanity

1:39:25 to 1:41:04

Explore the concept of human verification and its implications in today's tech landscape.

“Oh, we've got another video we should share.”

The Importance of Unique Identity Verification

1:41:04 to 1:42:33

Understand the significance of verifying unique identities in an increasingly digital world.

“a human to be like yes it's me dad uh is is increasingly important and and i feel like they've done it in a way that is as as uninvasive as is possible if i'm honest yeah um Like they're not keeping your data.”

Verifying Identity in the Digital Age

1:42:33 to 1:43:40

Discuss the innovative solutions presented by Tools for Humanity for identity verification.

“Like a World of Warcraft account where you come in and your character's naked and his bank's empty and all his money's gone.”

Building Applications with Verification Technology

1:43:40 to 1:45:19

Consider ideas for applications that leverage identity verification technology.

“I swear to you that a year ago, real agents weren't a thing.”

Challenges in App Development and Verification

1:45:19 to 1:47:26

Explore the complexities and challenges when developing apps and integrating AI.

“I'm liking the direction of where all this stuff is going which is seemingly more useful, more powerful and easier to use for normal people which is like the right direction, right?”

Microsoft vs Google: A Competitive Landscape

1:52:00 to 1:53:20

Explore how Microsoft focuses on user interface while contrasting it with Google's model output.

“And it's something that's always made me a little bullish on them through all of this because, like, they have an audience that's bigger than anybody as far as, like, when you think of Windows machines.”

Adapting AI Products to User Needs

1:53:20 to 1:55:30

Discuss the importance of adapting AI products based on user feedback and interaction.

“I said Microsoft will be competitive by mid-year or something or have a competitive frontier level model.”

Highlights from Recent AI Developments

1:55:30 to 1:58:24

Review recent advancements in AI tools and their implications for users.

“And this is a thing that's now happened like three times.”

Show Closing and Future Insights

1:58:24 to 2:02:10

Wrap up the discussion with key takeaways and a preview of upcoming episodes.

“Make sure you pop by the neuron.ai, sign up for the daily newsletter, and stay tuned.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Elena R.:Oh, howdy and hi, all of the things.

0:02Julius S. D. M.:How's everybody today? Corey, would you mind giving the intro? I'll be right back.

0:07Elena R.:Yeah, yeah. We just hit the button. Grant's getting the email out to everyone to see who all might want to join. So we're getting started. But in the meantime, welcome humans to the Neuron Live. We're excited to have you here today, as always. And this should be an interesting episode. we're still kind of figuring out a few things about what's going on. So we kind of scheduled this with the idea that there would be some announcements today and, and Grant's gathering some information and all of that. And we'll have some more on that in a minute. But for right now, we're just getting set up. Hope everyone's doing good and having a great week.

0:48Elena R.:And we'll see what's happening. What's happening today? All right. Where's everybody joining from? Really curious to see where everybody's from today. I just got home about 2, 2.30 this morning from Microsoft Build in San Francisco. Really fun trip. It was some pretty major announcements, frankly. Yeah, what was the coolest thing you saw when you were there? Oh, it was the Surface Ultra and that dev box, the RTX Spark dev box. holy smoke like that is top tier premium uh hardware it is it is it is ridiculous i got to i got to play with a couple of them while i was there uh and was was pretty floored like uh and and that that rtx spark chip from nvidia is insane uh they also had not just the uh not just the surface version they had the lenovo the dell the everybody else who is releasing uh HP, who's going to be releasing these systems with that new NVIDIA chip, were all there.

2:03Elena R.:And it looks like they'll all be coming in the fall. Hoping to get hands a little early on the Surface Ultra because it's ridiculous. 128 gigs of unified RAM. Lightning fast. Lightning fast. And I think it's mostly got to do with the NVIDIA chip. uh but they they were showing us things in our demo yesterday like uh i don't know if anybody works with cad and you know what a headache spinning cad is trying to trying to show all of the dimensions of your drawing it was just flawless on a touch screen they could blow it up and go woo and the car spins with no loss of detail or or or graininess or anything it was It was absurd.

2:47Elena R.:Showed us a few video games on it. They showed us 3D world design and the ability to quickly spin your 3D world and not have all of that pixelation and messing out that you have. I also got to hold the chip. The chip is about this big. I've got a picture of it on my phone. I'll see if I can send it to us here in a minute. And it's pretty nifty. But they had one of the surfaces for us blown apart, like hanging from the ceiling in all of its slices there to check out uh to be as hot as it is one of the things that was craziest to me is that the fans were like virtually silent like dead silent i uh i have i have this gigabyte laptop over here that i love but when it gets gets a little load on it it sounds like like a drone took off in the living room it is it is so loud it will like like burn into two grand do you mean to be muted or is that like unintentional

3:42Julius S. D. M.:that is intentional sorry i laugh i laughed heartily at that last call okay i saw you but yeah it's because i'm clicking around and doing things over here so yeah keep doing things

3:52Elena R.:we're good uh i hope uh what else was there from build that was cool i got to spend a little time with mustafa sulyman from microsoft ceo of microsoft ai a lot of really cool stuff coming out of them and probably going to continue to be throughout the year. I expect that the seven models we saw released on Tuesday are only the start. It looks like they've, I've got a full interview I did with him that we'll be releasing next Wednesday. We'll make sure everybody gets that too. It'll be in the, we'll share it out in the Neuron. But it's, it's really interesting. It's not super long, but we talked a lot about humanistic super intelligence, about what his team's been doing about all of the restructuring that they underwent late last year like from october through december and how that has enabled them to really zero in and get their focus back on chasing super intelligence and i think that uh uh the stuff they just released is is good like um like the i want to say i mean i code one flash i'm gonna mess these names up because you know how names are.

5:01Elena R.:I want to say that it debuted on SweetBench Pro at like a 53, which is pretty comparable to Opus 4.6. So out of the gate, that's a big hit for Microsoft. The thinking model is really solid. It's really good. And what's interesting about the thinking model is they trained it on 35 trillion parameters of data. Or 35 trillion parameters of data. it has it's a 1 trillion parameter model with 35 billion active parameters which is super duper low for a model that big and they did it by very very carefully curating the data there's no distilled data there is no we went and synthesized data from all this other stuff they went and and had real data and had humans working it and and creating quality data that was corporate friendly copyright friendly all of those things in order to start from this baseline that is really really clean and and he talks about that in the interview as well so there'll be some good stuff coming up

6:06Julius S. D. M.:there um yeah well uh i think the email may have gone out to everybody at this point or it will be soon so i think it would be good to kind of get us get us on track here so yes What are we going to be talking about? We are going to be talking about the rumored new model, which I have a feeling it's not going to come out today, but we'll see. But there was a really cool announcement that just came out out of OpenAI regarding memory, which we are going to talk about. And then I would love to talk about the new stuff that happened with Codex this week, as well as my new favorite tool, which is Hermes Desktop.

6:52Julius S. D. M.:And besides that, there was also some cool stuff that came out this week regarding Gemma 12B, which is an open source model that a lot of people can run that already has a two-bit quant.

7:06Elena R.:It's going to run on eight gigabytes of RAM.

7:09Julius S. D. M.:Yep, yep, that's right. Which means a lot of computers, if you have a MacBook that's like 16 gigabytes or anything bigger than that of unified memory, you should be able to run it. I don't know.

7:21Elena R.:And 8 gigs going back as far as like 2015. You don't have to have a new computer to run that model, which I think is... Right, if it's a powerful one, yeah.

7:30Julius S. D. M.:Yeah. And then there was also something that you got to play with or you at least knew about ahead of time, which was Nematron 3 Ultra, the big baddie that just came out today.

7:44Elena R.:It is madly impressive, too. I'm anxious to talk about that. Actually, I've been this morning. I got home in the middle of the night, and I got up, and I knew I had a number of things to do, one of which is get my full review out on Nematron 3 Ultra.

7:58Julius S. D. M.:Is that out on the website? I'll link it.

8:00Elena R.:No, it's not. That's because I was still writing it and keep getting pulled into other things.

8:05Julius S. D. M.:we'll publish it after this

8:07Elena R.:when you've been out of town a few days and you come back there's kind of this

8:11Julius S. D. M.:thing going on but

8:15Elena R.:it'll be coming out today tonight long story short it's really good 550 billion parameters I think it's 35 billion active as well it runs light and efficient and if you've watched these live streams before when we had a new model out, and you've seen some of the questions we pushed through it, a few of them that regularly trip up other models, it destroyed. Just absolutely mowed through them.

8:45Julius S. D. M.:Actually, I have a question. You got early access to this. I did not. It also just came out today, so I'm unclear. How are people expected to run this, this NVIDIA and Nemetron? How did you run it when you did it?

8:58Elena R.:I ran it. They sent it to me in a notebook. Okay. So I don't run it locally. I don't have enough machine to run a 500 billion parameter model, nor do most of you watching, most likely nothing personal. It takes a hell of a computer. but uh you know you can use it through open router you can use it through any of those type of places i believe it's on you know bedrock and azure or microsoft foundry uh all of the places you can use models and nvidia has its own cloud hosting platform as well where you can go and get it and attach and use it but it's uh it is undoubtedly nvidia's best model to date and uh it's it's strong let's put some visuals on this shall we They told me it would take four DGX Sparks to run it at like speed.

9:48All right.

9:49Julius S. D. M.:So this is what we're looking at here. So if you're not familiar with Unsloth, this is a company that makes smaller versions of models that you can run on your computer. So this one here is an Unsloth version of Nemetron 3 Ultra, which is what we're talking about. that says it can run 2-bit on 200 gigabytes RAM. Wow, that's wild. Now, how would you actually... Here's the official announcement. I wonder where they're pointing people to go try this.

10:26Elena R.:Yeah, I'm not sure where they're pointing people to go. Yeah, open code, of course.

10:32Julius S. D. M.:So if you did want to run this yourself and spin it up on your own cloud server, The way you do that is you would go to hugging face and you would grab it there. Go ahead and retweet that. And then this is the smaller version. So go ahead and highlight that as well.

10:53Julius S. D. M.:Looks like it is on fireworks potentially. Welcome to the frontier. Yes, I did. I like fireworks. They're pretty great. Okay, so I'm going to look at that blog post This just came out today, so I haven't really looked at this yet

11:11Elena R.:Like they kind of Jensen announced that it was coming on I guess US Sunday night From Computex in Taipei But today was street day

11:28Julius S. D. M.:I like the way they outlined this They bolded the best one

11:32Elena R.:And which models they see it comparable to, I think, is worth looking at up there. They're seeing it when you're looking at various classes on these. One thing to know is not every model comes out being expected to compete with the latest Opus or the latest ChatGPT. They put these models out kind of in classes. For example, this new Microsoft Thinking model is Kimi K26 range. I want to say GLM51 was in there, and it's meant to compete in that weight class, kind of medium weight class right now, but getting bigger. And the thing is –

12:08Julius S. D. M.:People don't know these are open weight models, meaning anyone can spin them up on any cloud server they rent or any servers that they have. I mean, they're all very big. That's what the B and T stands for, which means you can't run it on your local computer, but you could run it on a cloud server that you rent.

12:23Elena R.:Nvidia and Microsoft, though, are very much, right now, the best U.S.-based open source game in town. And I wouldn't have said Microsoft before this week. With that said, their image models are fantastic. The new ones, they dropped higher than Nano Banana 2. What? Or scored higher than Nano Banana 2.

12:46Julius S. D. M.:Yeah, Microsoft. I find that hard to believe, but that's interesting.

12:51Elena R.:You really shouldn't. It's good. I've spent some time with it. Well, not with 2.5. 2 is really good. Is that out now? Yeah. Yeah, it's out.

13:04Julius S. D. M.:I believe all seven of those were released the other day.

13:07Elena R.:MAI image 2.5. Actually, I have a thing. I can probably bring it up on screen and show you. I can run it.

13:14Julius S. D. M.:You can run it? Yeah, let's see it.

13:16Elena R.:Yeah, right there. Read the first paragraph. Read that first paragraph. It said the Nano Banana claim. where where where up ahead of now okay let's see this leaderboard

Read the full transcript

13:28Julius S. D. M.:i don't believe it okay it's right below okay i see it's right below uh gpt so it's not the

13:37Elena R.:frontier otherwise we would have talked about it but but it's still up there yeah i mean i mean if

13:44Julius S. D. M.:we're above grok imagine grok imagine's good yeah i i imagine put intended that this is not 1.5

13:52Elena R.:which just launched but no probably not that may not very well not be on there yet wow this is actually it says preliminary maybe that is let me expand my screen a little bit here what is this go

14:04Julius S. D. M.:away okay it's like getting this sizing right you know okay so image quality i don't know

14:14Elena R.:know this if you look at that little exclamation point on the right over by 1388 it says hover over down straight down that's up this one yeah those are based on pre-release testing is what it says

14:28Julius S. D. M.:so i'm guessing maybe that's 1.5 but i i don't know i'm sure i can see it yeah uh but yeah i mean

14:34Elena R.:2.5 like like they came out wow they made some waves this week yeah that's very cool um yeah

14:41Julius S. D. M.:Would you show us this? Can you fire it up? Let me know when you're ready.

14:46Elena R.:Give me just one. Uno momento. I assume I can show this tool I have.

14:53Julius S. D. M.:Are you saying... Someone in the comments said our microphone is too sensitive. Is that... I wonder if that's me.

15:01Elena R.:Let me check my side too. My mic is always stupid quiet. And I struggle with... oh they changed some tools in here this week let me uh give me one second they moved everything classic there's that there's the mic yeah okay well i don't want to change the mic i just want to boost my volume interesting audio yeah the problem when you launch like seven things at

15:30Julius S. D. M.:once is it's hard to know like what it is and i uh i finally fixed the spring in my mic so i can

15:35Elena R.:move it closer to my mouth as well so hopefully that'll help but thanks for calling that out i'm

15:41Julius S. D. M.:very conscious of the fact that this is a problem we don't 100 know it was you so i did ask a

15:46Elena R.:question to confirm okay just to be safe here i am going to close both are pretty because i since i

15:53Julius S. D. M.:don't know how much of this i can show i'm closing the sidebar because i know there are a couple of

15:58Elena R.:things over there that aren't like necessarily out yet okay good good i'll try to make sure i keep up to evan uh that i that i speak up good because it's it's a thing i it's even even uh like doing those interviews this week at at microsoft the it's always an issue like my mic has to be like way closer to my face wait brian says yeah just don't um over correct

16:28Julius S. D. M.:so don't turn your head when you're talking

16:32Elena R.:that's going to be a little difficult

16:33Julius S. D. M.:when you're talking talk directly into your microphone

16:37Elena R.:so here we go this is where we get access to some of Microsoft's models but what do you want to create give me an image prompt idea folks I'm going to move the camera screen over there let's say we want to see create an oil painting of a dog

17:07Elena R.:in the park dancing in the rain, but the rain is meatballs.

17:20Julius S. D. M.:Cloudy with a chance of meatballs. Cloudy with a chance of meatballs, right? Okay. Sorry, that was weird. No, it's all good. Gotta keep it weird. Okay.

17:34Elena R.:Oh, and this is 2.5 flash.

17:36Julius S. D. M.:Okay.

17:38Elena R.:Okay. So, oil painting, dog. It's definitely oil painting looking. You can see the strokes. Looks like meatballs. Okay, so now let's go back, and now let's say they added an edit feature, too, which is wonderful.

17:51Julius S. D. M.:Let's say, actually. Can you zoom in on that a little bit more? or expand your screen a little bit more?

17:58Elena R.:Let me see. Can I? Not very well. Let's see if this helps.

18:04Julius S. D. M.:Yes, that's perfect.

18:05Elena R.:Did that help? Yes. It takes some creativity to adjust screens because, like, the shape of your screen matters, and you always have to. It takes us a minute sometimes.

18:12Julius S. D. M.:Yeah.

18:13Elena R.:Okay, so actually make it photorealistic.

18:20Elena R.:Overall, I've been really impressed with their image models. It feels like the first place that they kind of hit the mark. They've got some really cool voice stuff that's out, and that's maybe still to come. It'll be really neat. Very cool. Their transcription model is supposed to be really good as well, but I haven't been able to get it to work for me on here, so I've got to check in on that. Okay. I don't know if I would call it photorealistic, but it is nifty. dog definitely looks wet meatballs rain okay yeah it looks like it's stage two similar to

18:59Julius S. D. M.:the original yeah it's very it changed very little actually yeah yeah i can switch through them here too i find that this is kind of a problem though any with any image model that you use is that whatever your first version is, it stays pretty close to that. And if you wanted to actually just get something completely different, just start a new chat. There's no reason to continue with the same one if you want something different. It's like pushing a boulder up a hill.

19:31Elena R.:Yeah.

19:33Julius S. D. M.:Now, there are models that are specifically for image editing. So like Flux, I think Flux Context or something like that was really good at.

19:43Elena R.:Yeah, I think it also matters how much variance you're looking for from the original. Like if what you want is to change the meatballs to, I don't know, dragons, you could do that pretty easily. But if what you want is to change the style of the whole image, often that doesn't go far. I mean, it's definitely an improvement. The water looks pretty real, but you can still tell it's not a photograph. So we're going to go back and do this a little differently now. We're going to ask for it to be photorealistic right up front, and it will do much better. Create a photorealistic image of a golden retriever jitterbug dancing in a park in the fall when it's raining meatballs.

20:18Julius S. D. M.:Looks like it says FOGO. I don't know if that's going to mess you up.

20:22Elena R.:Oh, it does. No, I don't think so. Okay. Surely the magic intelligence in the sky can see that it's a typo. Surely, surely. In a, what, 14-letter word? Yeah. Only one's wrong Yeah

20:37Julius S. D. M.:Oh there you go

20:38Elena R.:Okay

20:38Julius S. D. M.:That looks slightly more real Still looks fake

20:41Elena R.:The dog looks less real Everything else looks very real The dog's lighting

20:46Julius S. D. M.:Is what looks fake to me

20:48Elena R.:Yeah I think it's mostly around the face Like I mean definitely Like the hair and stuff is there I don't know The trees look good The meatballs are definitely meatballs Yeah

21:00Julius S. D. M.:I mean it's not bad But it just got that cartoony kind of look to it That's why I would do the same exact prompt, even with the typo, in a completely new chat and see if it gets closer to what I would at least consider real. By the way, for anyone tuning in right now, we are a little bit sidetracked. We're testing, apparently, this new Microsoft model has surpassed Nano Banana from Google, which I found unlikely. So we're testing it live.

21:30Elena R.:What would you want to test? I'm trying to figure out how to get a new chat.

21:35Julius S. D. M.:Oh, yeah. Good question.

21:37Elena R.:Without having to. Hang on. I've got to stop sharing and go back in.

21:44Julius S. D. M.:So this is the new. What interface are you using for this? This is their new interface?

21:48Elena R.:No. This is the one that they give us to use for such things so we have access to them. Let me. There it is. Okay.

22:05Julius S. D. M.:So basically what Corey said is he's using a secret magic interface. What about Kevavocoyote?

22:11Elena R.:I don't think it's secret, but it's not mine to share, I don't think. But I think doing this is fine. All right. What do you want to see? What would be your test? Are you talking to the chat or me? Or you. Or the chat. Either is fine.

22:31Julius S. D. M.:Okay, well, do exactly what I suggested before, where you take the same prompt, copy and paste it, and see if we can get it to look more photoreal.

22:38Elena R.:Create a photorealistic image of a golden retriever, jitterbug dancing. Ah! Typos.

22:56Elena R.:In a park in the fall while it's raining meatballs.

23:04Elena R.:and let's see what we get this time. All right, let's see it.

23:13Elena R.:I've had overall good luck with this, and I've used it quite a bit for some things here and there.

23:19Julius S. D. M.:Is this available in CoPilot right now?

23:22Elena R.:I believe so.

23:24Julius S. D. M.:What's interesting, it looks exactly the same, which is very interesting.

23:28Elena R.:The dog looks exactly the same, doesn't it?

23:30Julius S. D. M.:Okay, well, let's try this. Yeah, it has like a cartoony glow to it. Create a photorealistic image of a woman.

23:44Elena R.:That's a hell of a sentence.

23:46Julius S. D. M.:Yeah, I know. That's what I was reacting to as well.

23:47Elena R.:With blue hair and red eyes leaning against a building in the city, looking to the left of the camera.

24:14Elena R.:And we shall see what we get here. I won't stay on this too long because I know we've got a lot of other things we want to talk about, but I just thought it was cool to show.

24:26Yeah,

24:31Julius S. D. M.:I've got an interesting tweet to share when you're done with this.

24:34Elena R.:They do have another tool for using this that you'll be able to use where there are things that are already there, like the ability to click. Okay.

24:43Julius S. D. M.:I'd say that's quite photorealistic.

24:46Elena R.:Perhaps it's the last Yeah

24:51Julius S. D. M.:Maybe zoom in on her face a little bit

24:54Elena R.:I can't Stupid screen Hang on though I'll tell you what we can do I'm going to save image

25:05Julius S. D. M.:I would say it looks fairly Photoreal The only part that reads off to me Is the face and maybe it's just because he has like red eyes and blue hair and seeing the face to me is what i thought looked looked much more real let me uh let's see the background looked there is that

25:23Elena R.:thing with ai where sometimes the problem is they look a little too perfect yeah uh which is very much a thing okay so now let me reshare all right good song thank you thank you All right, we're going to zoom way in. Oh, shoot, it didn't move. Okay. Okay. Okay, that's way bigger than I should be zooming in on this. It looks like digital.

25:59Julius S. D. M.:It looks like digital art, which it is, I realize. Well, and to be fair, I was looking at, I'm zoomed in on the head,

26:07Elena R.:and on my monitor, it's this big.

26:09Julius S. D. M.:I was zoomed in.

26:11Elena R.:So like, you know, the artifacts you're seeing are because her head is literally this tall right now.

26:17Julius S. D. M.:Yeah. And who knows if we were to zoom in on a real photo, maybe that would I feel like it has. And correct me if I'm wrong. The hair looks great. It has almost like a painted sheen to it. But I don't know if I'm just overly critical of that because we were just looking at an oil painting of a dog. Yeah, that's fair. Yeah. So I don't want to say that that is is actually.

26:39Elena R.:she appears to have the right number of digits that's good yeah uh people walking around the

26:45Julius S. D. M.:street we got cars that looks very real to me everything looks normal as it should in the

26:50Elena R.:background coat looks good yeah hair is slick like like i like that the hair has like depth and texture to it i don't know if you can tell like when you look around like through here and through here. There's definite

27:05Julius S. D. M.:you know. Yeah. Let's see. Yeah, it looks good.

27:11Elena R.:Red hair, blue eyes. Pretty cool, by the way.

27:14Julius S. D. M.:All right. Well, good job, Microsoft. Uncanny Valley. That's the word. Color me a little bit impressed.

27:23Elena R.:A little bit impressed. That's okay. That's okay.

27:25Julius S. D. M.:I mean, certainly a good showing, right? Yeah, we've definitely had worse. Yeah. So I'm going to share my screen again. I'm going to share this tab. And what we're going to talk about now is we're going to switch gears back to. Okay. So Corey, do you know this guy? His name is Chris GPT. Yeah. Okay. I see him.

27:51Elena R.:I mean, I don't know him, but like, I think we follow each other on Twitter.

27:54Julius S. D. M.:He is allegedly an AI insider. I forget exactly what company he's supposedly an insider for. I imagine OpenAI, but I'm not sure. But one of the vague post kings of X.com. And right here, so this live was about Mercury Alpha. There's also now apparently Jewel Alpha, which is coming out. So have you heard of this at all?

28:20Elena R.:Yeah, I saw yesterday, if you follow any of the OpenAI crowd, Riley Brown, who works for OpenAI, shared a video yesterday. and in that video at the bottom of his screen it showed jewel alpha and i think that's where this cut is from and later on in the day i saw that and i was like oh i want to go look at the video and i went back and i said this video is gone oh my god that means it was real yeah okay right so they they definitely took it out other one we keep hearing right mercury alpha mercury alpha

28:52Julius S. D. M.:yeah and i'm gonna go ahead and i'm gonna try and find the original tweets where i saw this

28:57Elena R.:A thing to know is these could also very much be different checkpoints on the same model. Yes. It's quite common for them to do that. And a thing Mustafa explained to me, Mustafa explained that I didn't know, Grant, and you'll find this fascinating, is that when doing a one trillion parameter model, they do a whole bunch of them. They take the data set, mix it up, use different data, and redo it, redo it, redo it, redo it, and then go through and can cherry pick like, oh, well, this set of data did better than this trillion parameters. And this trillion parameters did worse than that trillion parameters.

29:30Elena R.:So you can go through and kind of adjust your percentages of the types of data in there to tweak the caliber of performance you want. And that's just not a thing I'd consider because they were working with, I think he said, I think he said 33, 33 trillion parameters in the thinking model.

29:47Julius S. D. M.:very cool sorry i just i thought that was no context in talking about it's not a thing anybody's ever told me before so is that kind of equivalent to trying to like stitch two different models performance together because i've heard that's really difficult you know the way to the way to think of a model is like they're building this thing and they're going to keep

30:08Elena R.:building this thing and then when it when a new model comes out essentially what they're doing is and save on that model and doing a save as and like here is the model as it was at this moment right and uh so like they've continued training they've continued tweaking they've continued adding new new data new things and they hit save but they can also do that with entirely different data at that point they can be like we're going to pre-train this model with this set this set this set this set and they're all going to be separate models and then we're going to test them against each other and see what the internal team and their beta testers think of it and be able to then kind of dial in this one.

30:49Elena R.:And it also tells them then that this data works well. This data is hot trash.

30:56Julius S. D. M.:Right. Yeah, they can figure out what to get rid of. Yeah. Well, if you look at this post here, so somebody, Leo, just somebody who started a rumor and I guess they accept tips said Mercury Alpha. And then Andrew Curran said, aka GPT 5.6. And a lot of people think it's coming next week. I think when you see something like this come up, it usually probably means that it's coming in a week, a week now. So, you know, adjusting expectations, it's probably coming next week. But that being said, Codex did have some new stuff that they launched this week. and I'm going to pull that up real fast if this is the right post.

31:43Julius S. D. M.:So codecs for every role, tool, and workflow. So I'm just going to read through this real quick because Corey has not seen this at all. And unless you all - I've been out of the loop. So I'm learning with you guys.

31:55Elena R.:Go Grant.

31:55Julius S. D. M.:Unless you all are extreme link clickers like I am and click every link in everything that you read, you probably didn't click into this to learn more about it. So I'm going to read through it really fast. So more than 5 million people now use Codex for work every week. This feels low to me given the fact that this is basically going to replace ChatGPT as their work surface app. But people have not realized this yet or not accepted this fact. So it is what it is. Codex started as a tool for software development, but it's increasingly useful for more kinds of work. Non-developers, including analysts, marketers, operators, designers, researchers, investors, and bankers, make up about 20 % of overall Codex users and are growing more than 3x as fast as developers.

32:42Julius S. D. M.:So that means that if there's 5 million people, there's 1 million people who use it for this use case, right? And I feel like this number is going to grow a lot. So this week, they introduced new ways to do more work with Codex with new plugins. So if you don't know what a plugin is, a plugin is a series of tools and skills that are put together and connectors that let you use a particular suite of tools to do a particular task. And I can show some of those later. Corey, if you want to bring up your screen, you can show some as well. But basically, they adapt codecs to your role and tools, and they're annotations with annotations that help you refine the result in place and a preview of the ability to create interactive websites and apps you can share with your workspace using a URL.

33:28Julius S. D. M.:So this is where they're talking about how they use it. But one of the really cool things that they launched this week is sites. So sites basically let you create a site directly inside Codex. So this is a demo here that I'm showing. Stop me at any point if you have questions. Yeah. This is, you know, the site thing is really cool.

33:52Elena R.:Because, like, that's the thing that's kept, honestly, it's the thing that's kept me using V0 for a good role. It was the ability to just quickly spin something up and hit deploy. And that is a thing that, well, I guess with Google you can because they have their own cloud. If a company has its own cloud, they have that luxury. But not everybody's had their own cloud. So like for Anthropic, for OpenAI, that's an area where they have not had that ability.

34:16Julius S. D. M.:But apparently they have something because this is really cool.

34:20Elena R.:The idea being you could spin up and share a tool with your team in minutes from scratch.

34:26Julius S. D. M.:Yeah, I think this is kind of like what replaced Canvas. Like they got rid of Canvas this week or last week. I think this is what's replacing it. And I think it kind of builds on a post that Darik from Anthropic shared a couple weeks ago, which was basically like, hey, if you're doing anything that's sharing it for other humans, just make it an HTML page. Like just write an HTML. HTML is the better format to use. and I actually use HTML when I write the newsletter so that I can actually have it hyperlink my links. So when I, you know, like I said, I'm a linkaholic. I love including lots of links to things.

35:04Julius S. D. M.:I try to make, I try to include as much information as possible. The easier way to do that when you're working with AI is to make it an HTML website. So this just builds on that and, you know, you can add all sorts of cool stuff on top of it. And then they add some other stuff here, which the plug, I think they announced six, Yeah, six new role-specific plugins that make it more useful for more kinds of knowledge work.

35:32Elena R.:Like non-developers, non-engineers. Yeah. And I'd say this is a very key piece of the puzzle for super apping.

35:44Julius S. D. M.:Yes, definitely for the super app.

35:46Elena R.:That data analytics plugin is sick. I'm anxious to play with that. I haven't yet.

35:50Julius S. D. M.:Yeah, so this is cool. I just added this, I think, which helps analysts and business teams answer questions with data. They can explore product and business data, explain why key metrics change, and create reports and dashboards. And then it connects to Snowflake, Databricks, all your data sources, Tableau. More coming soon. Then there's a creative production plugin, which helps marketing and creative teams turn a brief into assets they can review. And this uses tools like Figma, Canva, Shutterstock, and Fall, which lets you do image models. So if, for example, you wanted to use that Microsoft model, you could use Fall to do that as one option.

36:25Julius S. D. M.:Sales plugin, which helps sales teams bring customer context into the work that moves the deals forward. Sales teams can find high-priority accounts. They can prepare for customer meetings. And they have tools like Salesforce, HubSpot, Slack, Outreach, Clay. That's a cool one. Rocks Unactively. Then they've got product design. so I think you can tell what product design is it's creating prototypes I built out a product team in mine

36:54Elena R.:say that again? I said I built out a product team worth of skills in mine like I have I have a senior product designer project manager, I have a UX specialist I have a YC Combinator advisor skill awesome that was like look at this like it was going to be something great we're really I think we're, I think I'm going to say by 4th of July that this will be the way you use chat GPT for the most part.

37:29Julius S. D. M.:Maybe not for everyone yet. That's going to take time, but I think it will be the premier way to operate will be in Codex. Yeah, kind of like how people have been slowly switching from Claude workspace or Claude web version to Claude desktop and using Codework. I think the same transition is going to happen with Codex. It's just a little bit harder, honestly, because they called it something different. So I agree that I think the brand needs to evolve beyond ChatGPT. I think ChatGPT limits them, even though it's like Google. But I think for serious work, I think it makes sense for them to make Codex their platform for that.

38:05Elena R.:You know, the flip side is, and this is a decision that both OpenAI and Anthropic made, is that using a name like Codex or Cloud Code implies this is a tool for engineers. And people who are not find that intimidating. And you can't depend on everyone to have gone and read your article here to understand that, listen, this is for anybody who wants to use it. You can do all kinds of things. Like I write in Codex. I research in Codex. I have Codex wings spun up all day long doing various things. Now I work. I mostly still use the chat GPT interface, but I do have a number of small tasks that, that I prefer to be handled locally.

38:50Elena R.:Right. Just because I find it's quicker and, and, and you can trace its steps and what it's done better.

38:59Julius S. D. M.:This whole time I've just been thinking that they should just call it work GPT.

39:03Elena R.:I thought you were doing that. I was like, dang, watch crap. No, no, no,

39:07Julius S. D. M.:no, no, no, no, no, no, no, no, no, no, no, no, I've just been thinking maybe they just call it work GPT. Cause I was like, you make it like chat work. And I was like, no, that's kind of dumb.

39:16Elena R.:And I think they'll hang on to chat GPT. I think the name will be chat GPT. I agree.

39:20Julius S. D. M.:I agree. But I think like codex could become, okay. So like cloud code has this really bad problem, which they solved by creating an alternate version called co-work. Right. Yeah. So codex. You download by downloading cloud code. Yeah. Well, no, No, you'd call it desktop now. They've sort of unified it into one app. I didn't know that. It's very nice. Oh, yeah, yeah. Well, and ChatGBT still has a desktop app as well. That's right, which makes the Codex app even more confusing.

39:52Elena R.:Yeah, Cloud and ChatGBT have both had a desktop app for a year, two years. I forget about that. And then the rest of this essentially became like, it was like that was your chat app and this is your code app. The problem is the code app was the natural place to put all of the agents, to put the plugins, to put skills, but all of the things. So now they've built these monstrosities over there that can do all kinds of crap. And, and like, you know, built in desktop agents. Like, like that's a thing that unless you have a business plan, you're not using it. ChatGPT right now.

40:23Julius S. D. M.:Yeah, you're right. Yeah. And I think, I don't know. So that's why I was thinking they need to solve it. The, but the potential way to solve it is say like, Hey, chatGPT is codex. Now you can still chat with codex. inside the Codex app, but everything's called Codex, right? That's one way to solve it. The other way to solve it is say, hey, we're now launching a work version of Codex that's called WorkGPT or something dumb. Like, I hope they've come up with something better than that, but it keeps like the chat, it keeps like some part of the chat GPT alive. And then - We're gonna call it 09A3.

40:59Julius S. D. M.:Yeah, yeah, yeah. God, somebody in marketing over there helped them come up with something better. I feel like one of the most successful model launches of the last year from a naming convention was Nano Banana. I feel like they should just keep their internal names for stuff.

41:18Elena R.:Yeah, and it's funny because I doubt that was the original intent with Nano Banana.

41:22Julius S. D. M.:No, it wasn't. They went on the record. They've done interviews about this. It was just like, I think it was someone on the team called it that because

41:29Elena R.:It's what it was being tested as on Arena, wasn't it? On LLM Arena or something.

41:33Julius S. D. M.:Which is how people knew about it. Yeah, exactly.

41:36Elena R.:And they just stayed with it, which was perfect. It's absolutely what they should have done, is stayed with Nano Banana because it's fun. Like, can you imagine if we were like, oh, well, I was using Strawberry. No, it was Orion. What was it? I love those names. I think those were a lot of fun.

41:48Julius S. D. M.:They are fun, but I get from a versioning standpoint, it's hard to know which one to use. So it makes sense to have 5.4, 5.5, 5.6. Bigger number, bigger intelligence. Yeah. Yeah. So it makes sense. And I think, you know, whether they stick with dual alpha or not, you know, probably won't. But, uh, if I go to chat bbt.com right now, right. So, um, they've, they've sort of landed on like in the app, you just have the instant version, you have the thinking version and you have the pro version. And if you look at all of the different applications, they've sort of all landed on this interface. Um, and I think this is the cleaner way to do it.

42:27Julius S. D. M.:Um, obviously you can come in here and you can um you can set what version you want to use which also makes sense um but they've still sort of simplified it to these three modes and i think that makes sense yeah

42:40Elena R.:yeah i do too i i like that they made that easier i like that everything on the left is now like very collapsible like you can close your agents you can close your gpt's you can close your projects and and just have them as one-liners right there so you don't have to have them all open all of the time and I really really like that because it makes me feel like a just feels like a cleaner UI

43:01Julius S. D. M.:yeah and I don't know if people know this but now you can just add a project at the bottom here so I could say like I didn't know that yeah so I could say write me a story on this and why don't I pull up something fun just to do an experiment um that's a good idea do something so let's look at that let's look at that new memory thing that we just teased um so I'm going to go back to x here for a second.

43:24Elena R.:And I'm going to go home. I'm going to share this tab. And what's that chatTBT memory? I'm sure it will come up.

43:34Julius S. D. M.:Okay. So OpenAI just launched this new thing. So as we've been researching new ways for chat memory to carry context across conversations and keep useful over time, today that work is rolling out as a more capable memory system in chat dbt great let's click on this um so dreaming better memory for a more helpful chat dbt so memory is what helps chat dbt learn your preferences projects and constraints allowing future conversations to share to start from shared context rather than scratch and so they kind of walk you through you know how it's changed over time and then over the last year dreaming has supplemented save memories to create a step function improvement in chat dbt's ability to personalize responses and offset the staleness of saved memories.

44:19Julius S. D. M.:However, it historically was never sufficient as a standalone memory system. Today, we are launching a significantly more capable and compute-efficient memory architecture built on top of Dreaming. The memories synthesized by Dreaming are reviewable through a summary of them made visible in the memory summary page. Interesting. So you can correct the memories, it looks like. That's cool. And then when we think about what a good memory looks like, a few things come to mind, carry forward useful context. So you tell Chachibichi something once and it remembers that information in subsequent chats, follows your preferences and constraints.

44:54Julius S. D. M.:So if you describe a preference like you're a vegetarian, then Chachibichi should take actions that are consistent.

45:00Elena R.:And remember that anytime you talk about food.

45:02Julius S. D. M.:Yep, exactly. Stay current over time. Memory should account for the passage of time. So imagine the user is planning their birthday party for next Saturday. eventually Sunday arrives. That's interesting. So it kind of tells you a little bit more about how to do that, right? So that's interesting. So I'm going to copy this link. I'm going to go back to Chat2BT and I'm going to say, write me a story on this. So I've read it. I know it's interesting.

45:32Elena R.:Two birds with one stone today. No, I'm not going to write about this for tomorrow.

45:36Julius S. D. M.:I solemnly swear I will not use this in the newsletter. but I just want to show you what's happening here. So I've assigned it to my project. So this is my project folder where I have all of the context about the newsletter. Neuron name and story skill. And then now it's using a skill. So this is how, this is basically I've taught it. Essentially what it, everything it was just saying about your preferences over time, I put that into a skill file here that can do that for me automatically. So I already have a version of what they've just launched here. But what they're saying is they're doing all of that context management inside the app now on your behalf.

46:12Julius S. D. M.:So you don't have to manage skills to manage your context. They're going to start doing that for you with the memories. So anyway, so what it's doing is it's then processing, it's fetching the link, it's using my skill the way that it's outlined here, and then it will eventually come up with the main story. And we'll rank it. We'll say, you know, is this neuron worthy or not? Which I think will be fun.

46:36Elena R.:It's a...

46:36Julius S. D. M.:Any thoughts so far?

46:38Elena R.:You know, something – I'm circling back to build here, but I don't know if you caught this.

46:44Julius S. D. M.:It was very much glossed over by everyone, but Windows got skills at the OS level, Grant. Okay, tell me more.

46:54Elena R.:I spoke with – you'll remember Pavan Davaluri? Mm-hmm. He's the executive vice president of Windows and Devices. and yesterday I was chatting with him about kind of some of the things they add and he said the one thing that I feel like nobody freaked out about as much as they should have was Windows skills he said because adding skills at the OS level was a big deal you know imagine if they just worked wherever you went like instead of it being in any AI you're using or whatever I just I really really feel like that's a cool step toward in a truly agentic OS. I also joked about, I said, are we able to quit using AI PC now?

47:38Elena R.:Are we just going to accept that PCs are expected to be AI capable and go back to calling them PCs? He said, I think so. Yeah, it was a stupid brand to begin with.

47:48Julius S. D. M.:Kill the AI PC. The best way to think about AI is it's like the third iteration of computing, right? It's the third iteration of computer. It's whether it becomes down to the OS level or whether it's a layer on top of the operating system, it will be eventually the uncontroversial, obvious way that we interact with our computers.

48:11Elena R.:Low key, though, the goal is to reach a point where a trillion-parameter model can run locally on an average laptop.

48:18Julius S. D. M.:Yeah, that would be amazing. That's kind of the over-a-few-years goal. Yeah. So let's look at this real quick, and then we'll close this chapter. So OpenAI says ChatGPT can now stream up better memory. So your best AI chat should feel less like a customer service form and more like a co-worker who remembers the project, the constraints, and the weird preferences you mentioned three weeks ago. And then it says OpenAI. So it links OpenAI, which is not what I'd like to do. I like to link to the actual thing. So that's already wrong. But AI is trying to make Chat2BD feel more like it was like that with a new dreaming memory system.

48:54Julius S. D. M.:The name sounds like a Pixar short about a laptop with unresolved childhood issues. You know what's a funny thing I've noticed? AI models like to personify objects. Like this is something when they make jokes, they personify objects. And I try to tell them not to do that. So for example, it says it's about, in this case it's a Pixar short, sorry, it's a Pixar short about a laptop with unresolved childhood issues. In this case, the laptop is personified. It's a person that has unresolved childhood issues. It's like the AI is subtly trying to force you to to empathize with machines. Yeah. And it'll do this a lot.

49:38Julius S. D. M.:It'll make a joke like, it's like a filing cabinet with an anger problem or like all sorts of stuff. And it's so weird. It's a very, you know what though?

49:47Elena R.:As a guy who reads a lot of sci-fi and fantasy, it's very much a thing. There's a series of books, the Piers Anthony Zanth books. Piers Anthony is somewhat problematic, so I don't want to get into him, but these books are brilliant. where things happen like inanimate objects do things you don't expect. Like, you know, if you need a pair of shoes, you go pluck them off a shoe tree. If you need, you know, your statues talk, you know, you might go to a different statue for different problems. You might, you could, like, lots and lots of things are a bit alive. And I wonder if that's not maybe a sci-fi thing, maybe a thing from books.

50:33Maybe it's they just would rather us not unplug them and feel like subliminally messaging to us is the way.

50:40Julius S. D. M.:That's my conspiracy theory. Exactly. That it's subliminal messaging from the AI models to try to get you to empathize with objects. But actually, I think that maybe that it's just very forward thinking to your point. And that in the future, you know, we will be talking to all of these devices. like we talk to other people. So maybe we should just come to accept it. It's, but it's just funny. It's one of those things that you have to iron out of it. It's like, no, when you make a joke, like don't make the joke from the point of view of an object, make it from the point of view of a person.

51:10Elena R.:Like,

51:10Julius S. D. M.:like it's like a weird thing that you have to tell it to do.

51:13Elena R.:Preparing us for a future where in two years we're standing in the kitchen, talking to our refrigerator about what we can make for dinner tonight.

51:19Julius S. D. M.:Exactly. That's what I'm, that's what I'm saying.

51:22Elena R.:Which is probably a thing that's going to be happening. Like, Hey, what's inside you that isn't rotten reveal to me your secrets

51:30Julius S. D. M.:yeah so like as you can see here like this is like formatted like a neuron article obviously this is not what everything I would say but like the factual parts like why would I write that myself like that's directly from their page I just make sure that it's accurate you know maybe I change it to focus on things that I think are more important than this from what I read, but yeah, it basically runs through exactly how to do it.

51:56Elena R.:To be fair, though, all you said was write me a story on this. Yeah, I didn't give it any direction. Here's six or seven different sources I'd like you to consider. My take on this is that it's problematic or it's great. The thing I will say is that Chagibut's first memory launch was, I guess, 2024 now. And it has worked better since then. Like the truth is it changed the way, in my opinion, that your AI works for you. Like if I get different types of responses from my work account than my Corey account. And that is because they know very different Corys. You know, one of them knows, you know, Corey, the nerd who occasionally has a car problem or likes good coffee.

52:48Elena R.:and the other one knows Corey who writes and travels a lot and does a bunch of weird things, records a lot of video with a guy named Grant, who he always wants a quippy line about to introduce a new podcast. Speaking of which, if you're here, I know it's a small group, but we're really grateful you're here today. Please take just a moment to hit the subscribe button. We crossed 20 ,000 on YouTube, which is great, because that puts us in at 31 ,000 overall or so between the podcast platforms as well. we're really grateful for every one of you. If you would subscribe, like the video, it helps us continue doing these and show the value in what we're doing and that people care to watch them.

53:28Elena R.:And we enjoy having you here and all the comments you bring. Also, pop by and check out theneuron.ai and sign up for our daily newsletter if you don't already. It's read by 700 and some odd thousand-ish people, and we'd love for you to be one of them.

53:42Julius S. D. M.:I want to highlight another thing that came out today. This may be right before we jumped on. A blog post from Anthropic about when AI builds itself. So this is a blog post all about their progress towards recursive self-improvement. For people who don't know, that's AI that builds itself or self-improves. And they did a huge blog post about this. Let's see. They say, taken far enough and given enough compute, the trend of AI systems delegating to other AI systems is capable of fully autonomous designing and developing its own successor. This is called recursive self-improvement. We are not there yet, and recursive self-improvement is not inevitable, but it could come sooner than most institutions are prepared for.

54:27Julius S. D. M.:And so they're kind of benchmarking their progress towards that. So they sort of give you this really cool visualization here showing the progress from chatbots to coding agents to where we're at now, which is autonomous agents, and then at some point closing the loop. Very interesting and potentially horrifying.

54:51Elena R.:It is. It's, yeah, it's both equally awesome and gives me pause.

55:01Julius S. D. M.:So this is cool. So this is where we're at now. So this is Claude writes a significant portion of Anthropics Code. As of May 2026, more than 80 % of the code we merged into Anthropics Codebase was authored by Claude. Before Claude Code launched in Research Preview in February 2025, this number was in the low single digits. That shift also shows up in the amount of output per engineer, and then they go into how that works.

55:25Elena R.:Last fall last year, they were saying almost all of the code is written by Claude already, weren't they?

55:30Julius S. D. M.:Yeah, yeah. Well, they were saying that they're getting there. I don't think they have made the claim that it's like 100 % really until this year they were making that, but here they're saying more than 80 % of the code they merge. And then it shows you code contributed per person by quarter, right? And so this is like how much code the average person is contributing,

55:53Elena R.:I think is what this is.

55:56Julius S. D. M.:Let me actually look at this.

55:57Elena R.:So is that, oh, that's a person and their AI is why that's going up.

56:03Julius S. D. M.:Each bar is the average over the days in that quarter of lines of code merged per active contributor shown as a multiple of the pre-2025 average. Wow, way to make this super confusing.

56:13Elena R.:That's assuming them using AI, I'm sure.

56:16Julius S. D. M.:Yeah, yeah, yeah. The hatched final bar is a partial quarter. It averages only the days observed so far, not a full quarter. Dash lines mark public announcement dates. So this sort of shows you like when Claude Mythos preview came out, Claude Opus 4.5, et cetera. And so it's kind of this confusing multiple system. Let's see if they have something easier to follow.

56:39Elena R.:Well, lines of code is a lousy production metric. for this it's a good way to measure it I think I guess it's the only way I don't know how else you would measure it yeah

56:51Julius S. D. M.:this is interesting so this is Claude Code Session Success Rate so internal Claude Code Sessions at Anthropic Weekly 4 million trailing mean LLM judge success task complexity is assigned by an LLM classifier that reads a summary of each session and picks one of four levels of a fixed rubric So how to read this session success is determined by a Claude judge. A session is deemed successful if the Claude code agent clearly succeeded at the user's tasks without requiring corrections. So did it get the job done? And it's very high number right now. So this is for trivial tasks, routine tasks and substantial tasks, which has seen the most growth recently.

57:40Julius S. D. M.:wow open-ended problems has actually kind of fallen a little bit which is interesting too

57:47Elena R.:you know and i wonder if you know the reason chan gbt and jim and i have both done a lot around open problems in math things in science and there's been less of that from claude and i can't help but wonder if that's not due to their more intense focus toward coding specifically maybe less of a diverse approach where i feel like the other labs have leaned a little more into all of that yeah i think because claw um

58:22Julius S. D. M.:claude is sort of um coding agent pilled in the sense that they think the best general purpose agent will be the best coding agent because it can do anything on a computer so um they're not trying to they don't have the resources that google has to make a bunch of bets and they don't have the same or the end and they're not willing to diversify on a bunch of different bets yeah

58:48Elena R.:well i guess that there's there's there's a variety of schools of thought like some believe that like if you solve coding the rest will just fall into place kind of others believe you need to lift it up kind of all at once others believe that you know you should specifically be focusing on any area that is verifiable. By solving the verifiable areas, the others will come into focus after that. The truth is we don't fully know that yet, but so far all of them seem to be doing a pretty darn good job.

59:21Julius S. D. M.:Yeah. Depending on what you're looking for. No, it's really true. This is kind of interesting. They have a possible future section at the bottom here. It says what happens next depends on two things. Whether the trend continues and what we choose to do if it does. So we can imagine at least three future scenarios. So let's look at these. So number one is the trend stalls, but today's AI capabilities are widely diffused. So this article features many exponential trajectories, but these trajectories may actually turn out to be S curves. We may be approaching the bend in that curve where it returns to scale, diminish, and the line straightens, then flattens.

1:00:00Julius S. D. M.:And then it kind of talks a little bit more about that. And it says, alternatively, the binding constraint to AI progress could be in the supply chain, not the model. So advancing and diffusing the frontier may require more energy and compute than presently exists. Wow. They're admitting that they don't have no compute.

1:00:18Elena R.:That Musk has barked about for a couple of years now that I think we tend to glaze over, and he's not necessarily wrong on it, is that... I think energy could be our biggest problem. Like an interesting thing, and he always shows this chart that shows China's growth of its energy infrastructure and ours, where ours is sitting pretty flat. And China's energy infrastructure has exploded over the last like five or 10 years. And they're continuously building more and more. And the truth is the biggest bottleneck we may face is power.

1:00:53Julius S. D. M.:yeah i i think certainly in the u.s that's true i think in china probably the bottleneck is chips and chip production which is why they're focusing on doing that and they may offset another yeah you're right um so like chip fabrication grid expansion or interconnect bandwidth so interconnect bandwidth is the thing that is really complicated and wonky that people don't get but that's like how much power can you actually move from one place to another? Right. And, you know, we have famously, we have a bunch of power that is like waiting to get connected to the grid. And it's, it's a backlogged. Right.

1:01:30Julius S. D. M.:And part of the reason for that is because the interconnect bandwidth is, you know, can you explain that a little bit? Explain interconnect bandwidth. Yeah.

1:01:39Elena R.:Explain what you're talking about when you mentioned that there is power waiting to be connected to the grid.

1:01:44Julius S. D. M.:Yeah.

1:01:44Elena R.:And I apologize if I'm just on the spot for something kind of technical.

1:01:49Julius S. D. M.:No, it's okay. And I don't, let me put it this way. I understand at a very high level, not at a very deep level. So let's just give that caveat at the top here. So there are transmission lines, and transmission lines move power from transformers, the physical transformers, and the source of the energy, the power plant, and they move that to the place where the power is used. And so there's been a big effort over the past six years or, you know, since around 2020 when there was an effort to electrify the grid in a big way. There's been a big effort to update the transmission architecture because it's really old.

1:02:34Julius S. D. M.:From, say, giant wooden poles in the air? Well, there's that, okay? Let alone that. You know, so putting those giant poles underground, so grounding the transmission lines, also, you know, updating the transmission lines just in general, because a lot of the grid infrastructure is from like the 1970s and hasn't been updated since then. You know, some of it's at least 25 years old. So that's an issue. And it's also very expensive. So I think there was some statistic that came out, you know, we can look it up. That was basically per, you know, per mile of transmission line, it costs X billion or something like that.

1:03:09Julius S. D. M.:So think about it. If you're going to dig up or retrench and redistribute all of the transmission lines across the U.S., it's going to add up. It's going to be really, really expensive to do that.

1:03:22Elena R.:You know, people don't realize like every winter across the Midwest and the northern part of the country, thousands and thousands and during storm season, thousands and thousands of power lines get dropped. like every ice storm that rolls through as that ice gets heavy snow gets heavy and it drops power lines and power poles like i i lived in an area where we had a major ice storm a few years ago where we got almost a half an inch of freezing rain which is a lot of freezing rain uh and and there was a 20 mile stretch of highway where every power pole dropped to the ground the first one falls and it hits this yanks the second one yanks the third one before you know it it's this domino

1:04:02Julius S. D. M.:that just goes miles and miles and miles of crashed power lines uh the other benefit to

1:04:08Elena R.:underground is when that happens you don't i mean the benefit to underground is that in things like winter storms you don't usually lose your power because they're not out there snapping against each other in the wind they're underground and not collapsing falling blowing up transformers and

1:04:22Julius S. D. M.:things okay all right yeah no no that's great yeah so basically that the the the level that the grid can accept is, you know, only so much power, right? Because there's a lot of coordination, not, not only all of the, you know, what's the actual physical limitation of how much, you know, transmission lines we have and how to move that power around. There's also the regulatory burden of that. So there's all of the legal like approvals that you need to add new power plants to the grid, which, you know, the current administration, they've been focused on trying to remove as much of that burden as possible.

1:04:58Julius S. D. M.:But a lot of the power that wants to be put online is a bunch of excess renewable power. And so the current administration has certain thoughts about that. So I don't know how much of that power has been put online over the past year or so, but there's like the last time I've read, and I think it might've been 2024, 2025, that there was something like, and don't quote me on this, but something like 30 gigawatts of renewable power that was like just waiting to come online. And I assume some of that has happened in over the past years.

1:05:30Elena R.:And it's a thing that's needed not just for AI. The fact is we are a more electrified society and more connected society than we've ever been. You have tons of things. If you go into a house that was built in like the 1940s, 1950s, 1960s, even the 80s, you'll see a quarter of the plug-ins you see in a new house being built. You'll see a quarter of the switches. There aren't switches everywhere. Not every wall has two plug-ins like they do in new construction.

1:05:57Julius S. D. M.:I mean, now every wall has like a quad box of plug-ins.

1:06:03Elena R.:And back then it was, you know, if you go into an old house, you'll see like maybe two plug-ins in a room is pretty common. Well, the apartment that I'm in is from the – Even the houses aren't equipped.

1:06:12Julius S. D. M.:Yeah, the apartment I'm in is from like 1930s, and I think there's maybe like two plug-ins per room. Do they both require adapters? No, thankfully.

1:06:21Elena R.:Somebody swapped the plug itself out then probably Yeah

1:06:24Julius S. D. M.:So it has been updated thankfully But yeah it's a similar issue So one way to look at the AI boom Is everyone seeing like Oh there's so much waste going on With all these data centers that are being built And all of the power And while you could argue And I've heard some people argue That there's no other use for All these chips that people are buying From NVIDIA other than for AI. So if AI doesn't work out the way people want it to, that's like a ton of money that's wasted and we won't have useful infrastructure like we did during the 90s with the dot-com boom and the dark fiber that was laid in the ground.

1:07:02Julius S. D. M.:The other way of looking at that is saying like, well, we actually need all of this power infrastructure. And in order to get the utilities and private companies to build more power plants, we need to give them an incentive to do that. And AI is a great incentive to actually make all of those commitments come to fruition. so actually if you care about having more power more excess power and your your power your price per power actually going down over time you want them to build as much as possible like like you actually want all of this to happen and even if there's an ai crash of some kind we will now have all this excess power that we didn't have before which would be good for all sorts of other use cases and lowering your power bills you know and all that so you you want there to be a positive incentive for them to build power and i think that's something that we should um you know not forget even as we're concerned about you know the power usage of ai and what it means and all

1:07:58Elena R.:that i want to pull up this message from ryan asking if anyone noticed the price of a 5090 has gone from 2 to 2500 to 4500 to 5k i have because i have been for a good number of months eyeballing an RTX 6000, the Blackwell. The thing Grant has in his laptop, I should add.

1:08:19Julius S. D. M.:Speaking of that, I'm going to join on that stream.

1:08:22Elena R.:But you should know they're now$8 ,000 to$10 ,000 to buy an RTX 10 ,000. Yeah, if it has RAM, it's lost its mind right now. And that also means that even entry-level laptops are going to hit. It's going to hit printers. It's going to hit all of this stuff so hard.

1:08:39Julius S. D. M.:and memory is just nuts.

1:08:40Elena R.:Some of that's AI. Some of that is manufacturing issues. Some of that is... There's a variety of things that have played into it. Definitely AI is hitting home on it. Poor gamers, though. They gave us these great graphics cards, and now the catch with those graphics cards is... First off, the crypto people came in and made their price skyrocket in the 2010s for years and years and years because everybody was mining Bitcoin and wanted 18 GPUs in their empty closet to mine, you know, quarters. And so that sent them through the roof. Then AI comes along, same story, sends them through the roof. And some of that is that, you know, I don't know.

1:09:29Elena R.:But I don't see a short-term scenario where RAM gets better.

1:09:37Julius S. D. M.:no because it's just like really hard to um produce it like we would need we would need the we would need like taiwan uh semiconductors to like you know build build not just the plant in arizona but plants in all 50 states in order to try and like solve it yeah or at least increase their capacity equivalently okay there was a lot of data center talk this week uh there was yeah

1:10:02Elena R.:Tell us about that. For anyone who hasn't looked up, Grant, would you pull up Microsoft's Fairwater Data Center? See what you can pull up on it. Last year, we interviewed Scott Guthrie, who was kind of the man behind this whole thing. And the Fairwater Data Center is probably the most environmentally friendly data center in existence. Like, it was very much built with that in mind. And they made this big commitment about earning the privilege of being able this week to earning. What did he say? Yeah, earning the privilege to build a data center in a community, to create a data center community.

1:10:41Julius S. D. M.:And you know why he phrased it that way, right?

1:10:44Elena R.:Because people are fuming mad?

1:10:46Julius S. D. M.:Yeah. They don't know how to backlash. They don't know how to backlash against AI. So they backlash against data centers, which is perfectly reasonable from the point of view where you don't see the benefits and you only see the downside.

1:10:57Elena R.:You know, and the flip side is a number of the problems that you regularly hear barked about have largely been dealt with. Like there was this disastrous data center in Tahoe where like they built this data center and it took all the power. And suddenly the power company gave the people in these towns like, hey, we're just not going to provide you power anymore. And you got like 96 months or 18 months to figure out how you're going to get electricity moving forward because we're just going to do this. and that's awful but it's also not the norm I would say but Fairwater, the key to Fairwater is that its whole cooling system is built in the way like the radiator in your car is it's meant, it's not

1:11:44Julius S. D. M.:constantly pulling in fresh water to cool these things

1:11:48Elena R.:Satya said it uses about the same amount of water in a year as a single restaurant Wow which is you know I mean nobody would complain about new restaurants in town we get excited when new restaurants show up and they come by the dozens

1:12:05Julius S. D. M.:in some cities you know

1:12:07Elena R.:but the thing to know is that the water they put in them is is about that equivalent and can be good for six to seven years before it's necessary to swap I mean there are probably I understand certain amounts that they replace here and there throughout the year for various reasons. But overall, these are built as environmentally friendly as one can be so far.

1:12:37Julius S. D. M.:And hang on, I'm letting you in, Grant. Okay, it's going to be, it's going to have my camera on to start. So we're going to get some echo.

1:12:46Elena R.:Echo, echo, echo, echo. Echo. Sorry.

1:12:51Julius S. D. M.:No, you're good.

1:12:52Elena R.:but uh you know the other thing is they're trying to be self-sufficient with their energy committing to like not raising the community's energy rates like if rates go up they're shouldering that cost themselves um uh also investing in you know a broad variety of local charities you know doing their best to keep its noise down you know looking for industrial park areas that are already loud, already zoned for that kind of stuff. And, you know, essentially trying to come in and win, win their approval, not so much, you know, some early ones maybe felt a little to people like they were sort of being snuck in the back door.

1:13:33Elena R.:And like we had one in Missouri where the voters voted out the entire city council.

1:13:39Julius S. D. M.:Wow.

1:13:40Elena R.:Because they approved it? They approved a data center. An election was next month. everyone in office was voted out the mayor the entire city council and uh yeah yeah it was just wild you know my dad asked me the other day because dad was interested he's like i don't really follow this he said but i see there are a lot of people fuming mad about data centers he said what's what's the story here are they right are they crazy are they i said you know you know as with anything there's a little everything you know they absolutely have some fair concerns. They also have a certain amount of information that I think is maybe a little dated and could stand to probably sit down for a rational conversation and figure out if we were interested here's what it would take.

1:14:30Elena R.:Because you are talking about new good jobs that they create as well. But there are definitely drawbacks and I can see why someone would want them in a community maybe necessarily. But, you know, if you have an industrial park, I mean, I can't envision a scenario where it's any louder than having a place that manufactures cars near you for sure. I think the number I keep hearing is like 60 decibels, something like that as an average, which, yeah, I don't know.

1:14:59Julius S. D. M.:The main criticism that I've heard is that, you know, when you have all these, you know, quote unquote, unregulated gas turbines that it spews a lot of methane and that can have downstream health consequences. Yeah, and that's an issue wherever there's a natural gas turbine.

1:15:14Elena R.:And wherever there is a landfill. We never talked about this kind of thing.

1:15:18Julius S. D. M.:Yeah, you brought this up. This is a good one.

1:15:20Elena R.:I lived in a town in southern Missouri where about four miles outside of town was a large landfill. I won't mention the company it's owned by because they own a million of them. There's another one not far from here. But landfills by nature create methane. as things rot underground they create gas and the way they build landfills they have these pipes throughout them that take in this methane funnel it to a central thing and it's essentially a giant flame in the center that's ignited and burning most of the methane that comes out with that said you drive around near it and you absolutely smell it like i just thought of they need to use those

1:16:02Julius S. D. M.:as excess power plants for data centers.

1:16:05Elena R.:We always jokingly called it, you know, the eternal flame of stink.

1:16:09Julius S. D. M.:And it absolutely stunk if you lived, if you were within a quarter mile, maybe.

1:16:16Elena R.:Yeah. Not far, not far. Like it didn't like overtake a town as far as stink goes, but I mean, your trash had to go somewhere. So like, I mean, I get it.

1:16:28Julius S. D. M.:And the good news is methane comes from biodegrading. yeah it's a natural source right so it's not like you can get that that's not a turbine you can turn on or off um that's like you have to do something with that gas they should just pipe that into like data centers but it's like you're already using something that exists anyway that was our

1:16:47Elena R.:joke yeah put the data centers next to the landfills pipe that natural gas over to power

1:16:51Julius S. D. M.:the data centers and honestly i say that as a joke but it's a good idea it's seriously if you

1:16:57Elena R.:You put the data center next to a landfill. The landfill is creating a pretty much unlimited around-the-clock supply of methane gas. And then you take that, pipe it in.

1:17:09Julius S. D. M.:I don't know. There you go. And, you know, I am not an electrical engineer or an engineer of any way, shape, or form.

1:17:16Elena R.:Yeah, don't quote me on that. Don't quote it in.

1:17:18Julius S. D. M.:Okay. Do your own research. But that sounds interesting. No, but the other alternative to this, right, is solar power and batteries. You can't have just solar power because, you know, it doesn't, you know, you need these things to be available for, you know, like large spikes in the fluctuation that come with training runs. So, you know, you can't just have, you know, solar power only runs for, you know, whatever, 12 hours, 16 hours a day. Yeah. Effectively. But then you can store that power and you can release it when you have those micro fluctuations. And it makes a lot of sense to do it that way.

1:17:51Julius S. D. M.:So you just need to scale batteries. I feel like scaling power energy storage is one of the most important things we can do over the next 10 years. Like, we should have started it 10 years ago. But, like, at least doing it soon. All right. Before we get too off topic here, I do want to talk about Hermes Desktop, which you have not seen at all. Is that right, Corey?

1:18:14Elena R.:No, I haven't. I haven't played with Hermes yet. I've been pretty open-claw stuck. But it's been on my list. I just haven't. The truth is I'm afraid that having it on a system where I already have open claw, I'm going to get a gajillion conflicts.

1:18:28Julius S. D. M.:Will it make my prop responses smell? Don't quote me on this, but I think they have a part of the onboarding. That's funny. I think part of the onboarding is they'll onboard your open claw stuff into Hermes. It's pretty sophisticated. Yeah. Let me pull it up.

1:18:47Elena R.:Yeah, please do. I'm curious to see it.

1:18:49Julius S. D. M.:I hear good things. I think you took my screen backstage, so you have to bring me up.

1:18:54Elena R.:Hang on. Grant's computer, is that the one we want?

1:18:59Julius S. D. M.:Yes, sir.

1:19:02Elena R.:Okay. No? Okay, there it is. Where is... It's not showing up, though, Grant.

1:19:11Julius S. D. M.:Yeah, it is. Whoa. Whoa, cascading Grant and Corey's. All right, so let me pull up a Chrome. Railway's going to delete your project.

1:19:21Elena R.:Just thought you should know.

1:19:23Julius S. D. M.:Yeah. It was intentional. Okay, so let me show this on the website. Hermes Desktop.

1:19:33Julius S. D. M.:They have two websites. No, I don't know. Okay, so Hermes Desktop. Let's talk about it. So if anyone has seen OpenClaw before, you know where this is going. But this is basically, in my opinion, the clawed to open claws, open AI. So it's the creators of this agent. They put a lot of their own taste into it. They have some interesting design aesthetics. And they try to give you as much power as possible without you having to worry about it. So the desktop version is essentially like open codecs. Looks like codecs. Yeah, exactly. It's like codex, but with any model you want for any task you want to do, basically.

1:20:20Julius S. D. M.:And it's pretty, pretty awesome. So before this desktop came out, which just came out this week, you had to install it in your terminal. It's very scary for people who've never done that before. It's not at all what you like. It's not a fun... They made it very easy, but you still get very scared doing that as a non-technical person. Now with the desktop, it is much, much simpler. And so the way that you would install this is you would pick whatever your platform is. You open the download and let me show you this guy. All right. So right now I'm in the settings. Let me switch out of here. Do that by closing this.

1:21:05Elena R.:Keep a close eye out to make sure you don't like expose your.

1:21:09Julius S. D. M.:It's not really connected to anything right now. Oh, that's good. Yeah.

1:21:13Elena R.:I'm always so scared with things like that because there's like so many ways you could accidentally expose all of your keys

1:21:20Julius S. D. M.:yeah good call but yeah so this is what it looks like so when you have it installed and so I'm going to zoom in just a bit here and right now it's running a local agent step fund so I'm going to say hey can you tell me about Herbie's agent and what it's doing is it's not calling OpenAI, it's not calling Claude or anything like that. It's talking with a local agent right on my computer. Okay. So, hey, I'm actually running inside Hermes agent right now. Hermes is an open source AI agent framework by New Research. Think of it as a terminal-based general purpose assistant, similar in spirit to Claude Code, OpenAI Codex, or OpenClaw.

1:22:05Julius S. D. M.:And then these are some of the things that make it unique, which is fun. Oops. So they have a skill system. So it can learn from mistakes and reusable workflows by saving them as skills, which gets loaded into future sessions. We have a whole stream about that if you want to check it out. Cross session memory remembers your preferences, environment details, and task context. It's provider agnostic. So it works with 20 plus LLM providers, open router, and which you can load your own AI keys there, Anthropic, OpenAI, DeepSeq, local models, etc. and you can swap models without changing anything else and then you can talk to it across telegram discord slack etc but you don't really need to do that with this app you can just talk to it directly in the app i found that was the most complicated part about using these agents was that you had to like sign up for a telegram account and then you had to like message them or use them in your terminal but here you can just start a new session like you would on codex which is really nice yeah and i have only just begun to start playing around with this now is like

1:23:07Elena R.:you're still able to deal with it through your telegram or whatever you want and and have it connect to dropping things in your folders your files and google meets and your google drive and

1:23:17Julius S. D. M.:stuff yeah yeah totally so if you go to skills and tools right um it comes with these are 77 tools that are already built in here i didn't have to go in and connect them or anything like that um it's already set up right so these are all things that are built in um and there's a lot power here shows you how this like stuff works and this is what i mean by them kind of imposing their own taste into the tool they've they've already pre-installed all of these things for you because they're like hey you're gonna want something that does this um yeah which is really cool and then you can go and you can look at okay what are these things actually so it's got agents there's creative tools there's data science tools and then there's this thing called tool sets which shows you like the types of things it can do.

1:24:04Julius S. D. M.:And this is really interesting where you can do cron jobs. You can set up automations, code execution, clarifying question, browser automation, image generation, et cetera, et cetera.

1:24:17Elena R.:Wow.

1:24:21Elena R.:I'm curious. And I need to download the OpenClaw app and look at it too because now that I've seen this, I'm curious to see side by side. What's different? What's the same? Because I know they have an app now as well, but this looks super simple. It's an interface I'm familiar with. I can see a million reasons you'd want to use it this way.

1:24:42Julius S. D. M.:You know, something that I think would be really useful is,

1:24:46Elena R.:and why I think the Windows skills discussion is really interesting to me, is because wouldn't you want your Hermes and your Kodaks and your Claude code all using the same skills? Yeah. Like, you don't want to have to fix it every time in Codex, every time you're in Cloud Code, every time you make a change. You wouldn't want to have to make a change everywhere you might use. And I say that there's probably not a lot of skills you would use in all those places, but to have the flexibility of using them in anywhere you opened would be really valuable.

1:25:23Julius S. D. M.:Yeah, I agree with you. I think probably the hardest part about skills is that you'll be updating them frequently, like at least you should be. And whenever you notice an edge case or something that you want to kind of like train out of it or teach it to do differently, then you want to create a new version. And the way that we've told people to do that is, you know, go use the skill creator skill and say, hey, update this skill. And, you know, the interface makes it really easy on Codex or Claude where you can just like one click button and update it. but then you want to keep track of those versions so yeah actually the right way to do it skill to do it i don't believe say that again i don't think you have to use the skill creator

1:26:06Elena R.:skill to make a change though i think you can just tell it hey you know make this tweak i mean

1:26:12Julius S. D. M.:you could if what you want is like a 2.0 or a 3.0 and maybe that is what the answer is i don't know i think they're using they're using something under the hood to to update it and make it a one

1:26:24Elena R.:click plug and replace button. So yeah, I think

1:26:28Julius S. D. M.:if you tell it to just do, hey,

1:26:30Elena R.:update this, it does it. I've done this on

1:26:34Julius S. D. M.:Codex and even in ChatGPT, the work account that we have, you can do this where one message at a time you can update the skill and do it like a versioning system where you basically say update the skill to do this. Here's this issue

1:26:45Elena R.:and it shouldn't ever do that. Please update it to make sure it won't go.

1:26:49Julius S. D. M.:But the right way to think about it is to do it like versioning. So let's redo this and let's call it 1.1 or 1.2 or something like that so you can keep track of it. So if you were doing skills locally on your computer, if you're doing skills in your Hermes agent, if you're doing skills in Codex and Claude, if you use all four of them, that could be kind of complicated to make sure all of them are up to date. But do you use a Google Drive system to manage yours? I feel like I remember you saying this. Yeah.

1:27:21Elena R.:Oh, I thought somebody said they used Windows skills because I think I'm a big fan of the idea of let's have these all in one folder and give that folder, give access to that folder to each of my various tools that might need it. Be like, this is where skills live. This is where skills live. When I say make a skill, that's where it should go, where then it's instantly available to each of those. Like, I feel like that's a thing that should be easy to do. Right.

1:27:48Julius S. D. M.:Where then you could make the change, whether you're in Claude,

1:27:51Elena R.:whether you're in Codex, it wouldn't matter. I would like to think you could, you know, make a change to that core file either way because they're all using the same file type so they can read them.

1:28:03Julius S. D. M.:Yeah.

1:28:03Elena R.:You would just want them to be somewhat agnostic of, like, mentioning different models and things probably.

1:28:10Julius S. D. M.:Yeah, I think that's right.

1:28:11Elena R.:You could say, like, you know, use your best thinking model instead of use GPT-55 on extended, you know.

1:28:19Julius S. D. M.:By the way, something I just thought of. So this is the Codex app I'm looking at now. As you can see, it's very similar.

1:28:25Elena R.:God, it is identical, isn't it? Yeah.

1:28:27Julius S. D. M.:They basically just wanted to make Open Codex. So just like, respect. I get it. I wanted it. I'm using it. I'm happy with it. But if you go here, you can see all the plugins that we talked about earlier at the beginning of the stream. If you're still hanging in there.

1:28:40Elena R.:Build iOS apps plugin, by the way. How about that?

1:28:43Julius S. D. M.:You notice I have that turned on. Yeah.

1:28:47Elena R.:That's a thing we've been waiting on for a hot minute, isn't it?

1:28:50Julius S. D. M.:Yeah. So there's all these cool ones in here. And I think they talked about some of the new. Maybe I can sort by new.

1:29:00Elena R.:Of course not. That would be too helpful.

1:29:03Julius S. D. M.:No. Yeah. If only I was in charge. Open AI.

1:29:08Elena R.:Hey. Tebow.

1:29:10Julius S. D. M.:If only I was in charge of this. I would fix all these problems I find with it. But okay, so let's see. Creative production, that's one that we talked about earlier, right? Sales, this is another one that we talked about earlier.

1:29:23Elena R.:Investment banking? These are ones that we didn't touch on. They're all right here.

1:29:27Julius S. D. M.:Yeah, they're all right here. So these are the new ones. But when you click in here, you can see what they actually have under the hood. So it shows you, okay, it uses these 17 apps. So, oops, I didn't mean to click that. so you can scroll through and see all the apps that it has connected. And then it has 14 skills built in, so you can analyze data quality, you can build dashboards. Let me know if this is not zoomed in enough.

1:29:54Julius S. D. M.:Jupyter Notebooks, KPI reporting, all sorts of stuff. One of these times I need to show how I had it visualize things. Oh, you're a little quiet. What did you say?

1:30:02Elena R.:One of these times I need to show how I had Codex visualize stuff. remember the dashboard i had it built like like visualize a transformer uh visualize i had another one that was like uh an adjustable universe scale or something that it let's see what happens if i

1:30:20Julius S. D. M.:ask it to visualize yeah because i have all these plugins turned on i've not transparently been able to use all of them yet because i just announced like half this stuff yeah so let's actually see what plugins it decides to use to actually do this.

1:30:35Elena R.:Yeah, because the one I did, it opened up these, it basically built HTML pages that were very interactive.

1:30:40Julius S. D. M.:You could blow things up, make them bigger, change a lot of numbers and change how it feels, how it acts, strengthen black holes and stuff.

1:30:50Elena R.:And I can't remember what it used, but it's on my personal account that I don't have access, my personal codex, which I don't have access to on this machine.

1:31:02Julius S. D. M.:It contains a project of different shape. Yeah, it looks like it's building using...

1:31:07Elena R.:Build web data visualization.

1:31:10Julius S. D. M.:There we go. That's it.

1:31:16Julius S. D. M.:They don't exist where exactly we're listed. Interesting. Well, we'll let that do its thing. Yeah, we'll let that do the... Come back to it.

1:31:26Elena R.:The hoodoo that it does.

1:31:28Julius S. D. M.:Let me see if I make sure that they actually work.

1:31:31Elena R.:What do we have that we should make sure people know about right now this week?

1:31:37Julius S. D. M.:A couple of fresh podcasts and recordings I'll drop.

1:31:40Elena R.:I'm going to drop in yesterday or Tuesday I did a live with Scott Hanselman from Build where we talked about basically AI engineering was more the focus of it is what I would say. But it was really interesting. He walked through, showed some cool projects he's doing. This cool thing he's done with, he has a pancreas pump in one arm and a glucose meter in the other arm that are implants. And he built an app that talks to the meter all day and builds charts of his performance and sends notifications to his phone to let him know what this needs to be doing. And it was a really, really cool thing.

1:32:24Elena R.:And I'm grabbing the link. Yeah, that was awesome. That was totally unplanned. unplanned. It was just like, oh, there's this one thing that we really need. I also liked that he talked about, you know, use my coding to spend something up. And if people suddenly show up and seem interested, then fix it.

1:32:42Julius S. D. M.:Yeah.

1:32:44Elena R.:OK, here's that link. Going to drop that one in here.

1:32:49Julius S. D. M.:Scott Hanselman interview.

1:32:54Julius S. D. M.:You know what's one thing that I've been testing lately as well that's kind of interesting that we haven't really showed off or talked about is Claude Design. Have you used that at all?

1:33:05Elena R.:No. I've never tried Claude Design. And that's really strange because I feel like it's a thing I normally would have played with. I remember it and being like, oh, that sounds cool. Yeah. It must have just fallen on a busy time or something.

1:33:21Julius S. D. M.:They haven't fully integrated it into the app yet, which they should. It makes it kind of difficult to work with it, to be honest. Like if you already have a program that you're building and you want to do the design for it. But it is pretty cool. I wonder if I have it pulled up. Let me check.

1:33:43Julius S. D. M.:Yeah, I do have it pulled up. So check this out. So this is it. It basically just makes front-end interfaces for you. And I'm going to zoom in on this. And I'm going to say on this...

1:34:02Elena R.:Oh, is this your project?

1:34:04Julius S. D. M.:Yeah, so this is something I'm working on just for fun. So basically it's a way to work with these models to work from the screenplay as the source of truth as opposed to some random prompt in some website. So I've kind of like trying to figure out exactly how to design it and make it look. But these are all different pages that it's created for me. This sort of shows like different visualizations. And if I wanted to change something to it, I could demo it here. So like in this case, I'm going to say, can you change it so the windows are resizable? so I can make the left bigger or the right bigger whenever I want.

1:34:56Julius S. D. M.:And then the cool thing is once it creates these edits to this front-end design, you can download the code and give it to Claude Code on your desktop, and it can then go implement that for you. And say, hey, review this code. it's a very lightweight to mock way to mock up stuff without actually having to change all of the changes to your interface sorry there's a massive helicopter overhead i can't hear it if

1:35:24Elena R.:that's any consolation oh that's good maybe maybe a slight droney noise but but it's nowhere near

1:35:31Julius S. D. M.:as loud as you okay good well um yeah so basically it's a lightweight way to like mock up stuff make changes um without having to actually change your entire interface and you can kind of feels like zero in a lot of ways yeah yeah yeah and uh yeah so just you can create a design system so the design system is like where you make all the rules for how it looks and then all the elements will you know respond to that then it has all your files here so you can go in and click and look

1:36:01Elena R.:at all the things it's been working on you what it's created read your code yeah and you can

1:36:06Julius S. D. M.:connect it to, that's funny, and you can connect it to

1:36:14Julius S. D. M.:your what's it called? You can then connect it to your GitHub repo, so it actually reads your code base and understands how everything works.

1:36:24Elena R.:Okay. And designs off of that.

1:36:25Julius S. D. M.:So I actually, I like it a lot. My only criticism is like, put this in the app. Why is this web only? This is annoying. Like, put it in the app. Yeah, I agree. as soon as possible and make them integrate with each other as easily as possible because I would love to just be like, great, hand this to code and go implement it. Yeah.

1:36:48Elena R.:In one fluid thing. That's one of the things that I like about V0 is that it also connects to Vercel's web server service. It just speaks between the two. I can go over and say, tweak that and then hit publish and boom, it's over there and it's up. but this is really nice. I like, and also I really dig what you've built there. Oh,

1:37:09Julius S. D. M.:thank you. Yeah. It's, it's coming along. It doesn't look like this right now, which is the problem. So I made a lot of changes to the front end and then having the code version to go in and like actually figure out how to change it all and do it well. But an interesting workflow pattern is, so people have been using codecs for mocking up front ends. And I was experimenting with this yesterday. Let me zoom out. And I said like, hey, here's the current version of my thing. And, you know, I have, so like, here's the old, I don't know if I can click in. Yeah. Here's the old or the current version of my thing.

1:37:49Julius S. D. M.:This is kind of a crappy screenshot. Here's another one. Nope. That's really zoomed in actually. Oh, cause it's zoomed in. It's zoomed in on my side. So try this again. Okay. So here's my old version of my thing and what it looks like. And I'm trying to get it. And then here's the work in progress version of my thing that I'm working on. And this is the ideal version of my thing. And then I show it that. And then I said, I have this problem. I'm trying to create a designful UI. I want to marry these things. I want to make it feel very connected. But I'm losing out on this element that I liked from this first version.

1:38:31Julius S. D. M.:And so I actually had the image model go in and mock up a new version. And this one's kind of hard to see. So I had to do a light mode version as well. So it kind of came up with a solution for a way to potentially implement it. And then now what I can do is I can take this screenshot and take this to Claude Design and say, hey, try to implement this with our design system. and then I can take that and then I can take that to Cloud Code and say, okay, now implement this because it's kind of been mocked up. Yeah, so it's a really cool way of going from like visual idea or like from problem to visual idea to something that understands the thing that you're actually building to then finally integrating it into the thing that you're actually building.

1:39:19I like that.

1:39:20Elena R.:It's cool.

1:39:20Julius S. D. M.:I think you should try it with one of your applications and see how it can get it. I will, I will.

1:39:24Elena R.:I may do that this weekend.

1:39:26Julius S. D. M.:Yeah.

1:39:26Elena R.:You may do that this weekend. Oh, we've got another video we should share. Yesterday, we released an interview with Tools for Humanity. If you don't know who Tools for Humanity is, you might be familiar with WorldCoin or World, WorldID. This is a startup that involves Sam Altman. I don't know his role, but I know that he's an owner and founder. centers around human verification. And it's a really, really interesting discussion. We get a lot into things around how do you know that the person you're talking to is human. And that's what they've built this solution to do. And it feels a little Terminator maybe.

1:40:16Like there's this wall you hold that basically takes your picture and codes it,

1:40:20Elena R.:It goes on the blockchain and directly into the app on your phone. It's not verifying your identity. It's not verifying who you are. It doesn't have your personal information. It can if you want, but it doesn't have to. Essentially, what it's meant to do is verify that this unique person is indeed a unique person. It verifies that you have never been verified on their chain before. And if you have, it will catch that and you will not be verified. but the idea being that at some point we're going to get in this situation where we're in this situation where you know you could absolutely steal someone's voice and likeness and contact their loved ones and scam them and a variety of things like that and to be able to have a way for a human to be like yes it's me dad uh is is increasingly important and and i feel like they've done it in a way that is as as uninvasive as is possible if i'm honest yeah um Like they're not keeping your data.

1:41:21Elena R.:Everything they have on you is broken up on the blockchain. So it's little snippets all over the world. No one person can pull your stuff back together. They can't retrieve you. It very much lives in your app, your instance of world. And I think it's a really neat thing. We had a great conversation about how, like, you know, we're within months to years of agents being, you know, 99 % of Internet traffic. And with that, there's going to be a need for people to be able to verify, yes, I'm talking with a person. Or for an app to verify a purchase maybe or something else. And they really kind of – they've been around for several years now working on this.

1:42:09Elena R.:And they have hubs in, I think, L.A., San Francisco, some other cities as well, where you can go and get verified. And they also do other things. But you can even buy the Orb to get verified. Everything they've done is open source. You can absolutely take their data and implement it into your app. You can implement it into other things. Yeah, I actually love that.

1:42:30Julius S. D. M.:But I was thinking I think there's some application ideas that need something like that that I would love to build on top of and make something cool. Yes.

1:42:40Elena R.:Like a World of Warcraft account where you come in and your character's naked and his bank's empty and all his money's gone. Whatever happened to you, Grant?

1:42:50Julius S. D. M.:I don't think I was ever hacked on WoW, no.

1:42:54Elena R.:I have been hacked on WoW so many times over the years. Now, mind you, I haven't played in a good long while, but it was very much a thing. You would show up and suddenly there stood your character in all his glory with no armor, empty bags.

1:43:07Julius S. D. M.:I feel like now that can't even happen to you because you can just request, you can just tell Blizzard like, hey, I was robbed. And then they have a checkpoint. Where it would get you is if you had been on hiatus.

1:43:20Elena R.:Like if you were in there yesterday and it happened, you could get fixed. But if you had been gone for quite some time, a lot of times you were stuck. I got stuck once.

1:43:29Julius S. D. M.:By the way, Corey, I don't know if you've noticed, but it's working on the Transformer Explorer.html down there.

1:43:36Elena R.:Yeah, it is. There we go.

1:43:37Julius S. D. M.:It's been taking screenshots, making edits. It's very interesting.

1:43:41Elena R.:It is interesting. Wow, it is. It's so cool to watch these happen. I swear to you that a year ago, real agents weren't a thing. No, not like this. It's so crazy. Oh, the interview. I've got the link here it's really cool I was supposed to get verified while I was in San Francisco last week and even spoke with them a couple times they even volunteered to come to the hotel and do it there so I could get verified while I was in town and record it and I wasn't able to so the next time I'm out there or maybe in Vegas later this month I'm going to make sure we can link up with them and get a video getting verified because I think it's pretty neat I think it's important and I just dropped the link into the chat there so anybody who wants to watch can't.

1:44:25Elena R.:It doesn't have a lot of views on it yet for one reason or another, but it's super, super, super interesting. It's worth a watch because it's a problem that's going to impact everyone at some point.

1:44:35Julius S. D. M.:Why is it not letting me open this? Opening the Codex browser? Weird.

1:44:43Elena R.:If it won't tell it, it won't let you open it. And it'll give you a better link.

1:44:53Yeah.

1:44:56Elena R.:it may even give you like a link to the file. You can go drop in a browser of your choosing. Codex browser won't open.

1:45:05Julius S. D. M.:Yeah. That's kind of annoying. Yeah.

1:45:09Elena R.:It'll sort it out. It seems like I ran into the same issue the first time I did it too. Yeah. It gave me a local host link is what it finally did. So just go pop this in your browser.

1:45:18Julius S. D. M.:Yeah. Yeah. It's pretty cool though. I'm liking the direction of where all this stuff is going which is seemingly more useful, more powerful and easier to use for normal people which is like the right direction, right? All of this needs to be going in that direction so I am excited about that and yeah, there's been a lot of cool stuff I don't know, did GPT 5.6 come out while we were distracted?

1:45:51Elena R.:I don't believe so.

1:45:53Julius S. D. M.:I kind of don't think so either.

1:45:55Elena R.:It did look like LM Studio might have dropped a new mobile app, though.

1:45:58Julius S. D. M.:Oh, yes. I did see that. You can take my personal screen down, or you can hide it if you want. I'll share something on.

1:46:10Elena R.:GLP ones may slow down biological aging. Those are interesting drugs.

1:46:16Julius S. D. M.:Yeah.

1:46:17Elena R.:Makes me want to go get my refill.

1:46:22Elena R.:yeah i mean i transcribed 1.5

1:46:26Julius S. D. M.:logan says we are cooking the world's best vibe coding app on android and ios it's going to be

1:46:31Elena R.:so cool well i assumed eventually google would come to the party well have you tried to make an

1:46:39Julius S. D. M.:ios app with um ai studio yet i have not let's do let's do a stream just on that we could do that

1:46:46Elena R.:Yeah, I've yet to try to make an iOS app. And I know that they added it to Codex as well, so maybe we dual them?

1:46:57Julius S. D. M.:Yeah, we could do an iOS app on Codex. Sorry, AI Studio is Android only, so we could do an Android one on AI Studio. Or you can publish it. Okay, that's fair. That's fair.

1:47:10Elena R.:You will have to set up a cloud account with Google, just FYI.

1:47:15Julius S. D. M.:Ah, yes.

1:47:16Elena R.:And they will want all of the money.

1:47:18Julius S. D. M.:If only they would fix that and make it not so annoying to set up and make it really easy. That would be nice. Yeah.

1:47:26Elena R.:Yeah, it's weird. And you know, of course, there's also the element of the you have to go through the whole app store thing. I say that with Android, you might be able to have it as a direct download from like yourself. It might be like bootload it somehow.

1:47:39Julius S. D. M.:This is like I told her, like I can't open this thing and it's going on like an adventure to try and fix it. It's like, no, just make it work. Just make it work. Yeah. Anyway, I am going to stop this.

1:47:52Elena R.:Okay. Yeah, and I think we're probably good to call it a day. I really appreciate everyone who showed up and joined us. I know this was a much smaller chat than usual, but we just wanted to come on on the off chance. The model came and knew there were other things we could talk about. Wanted to talk a little about Build anyways because there's a lot of interesting stuff. A lot more I could say there, too, that we didn't even touch on, honestly.

1:48:17Julius S. D. M.:Yeah, let's talk about it a bit before we wrap. We have 10 minutes.

1:48:20Elena R.:Okay, okay, that works. Let's see.

1:48:27Elena R.:What was I just thinking of? Scout is really cool. Windows Scout.

1:48:33Julius S. D. M.:Tell me about Scout. Is that available now?

1:48:34Elena R.:Scout is right now in Frontier, but is coming this summer. so you should have access to try it out it's essentially it's the first of their autonomous agents they're launching and it is absolutely simple open claw I've got a full article coming out on it but you know you go in and he's like okay I want to use Microsoft Scout and it's like hey what do you want to name your assistant and you name your assistant it's like here pick an avatar or have it create your own you do that and then it's like let's look at your email and your docs and your teams and see where maybe I could help you out. And it comes back with a bunch of recommendations.

1:49:15Elena R.:Like, Hey, do you want me making sure you're ready for meetings? Do you want me to gather this stuff for you in the mornings? And you're like, yeah, yeah. And it'll do that. And here's the kicker. It talks to you in teams. It's just like another employee. So like it'll message you in the middle of the day and be like, Hey, you got an email looks important. Uh, do you want me to reply to this for you? Or are you going to be able to make that next meeting? it looks like you're still in this meeting. Do you want me to mark you as late and send an email, reschedule? Even the ability to say like, listen, I ate dinner from five to six and I ain't changing that for nobody.

1:49:50Elena R.:And if a meeting request comes, it will automatically go and send an email as itself, not as you. It'll be like, hey, I'm Sebastian, Grant's AI assistant. And Grant can't meet tonight at 5, but he does have some availability at 4.30 or tomorrow at 9 a.m.

1:50:08Julius S. D. M.:Take your pick.

1:50:10Elena R.:And it'll do that for you. I jokingly asked him, I said, okay, but what if the request's from Satya? I said, we're thinking it comes from your team. And you're like, no, you're not scheduling a meeting in my meeting, but what if it comes from your boss? He's like, oh, I could set it to do that as well. Ryan's asking which platform is this, Corey. I don't know if that's from right now. If we're talking about Windows Scout, it's going to be part of Microsoft 365 GoPilot right away, or right soon. Right now it's in their Frontier program, which is like their let's roll it out and make sure it works, get some feedback real quick, make some tweaks.

1:50:49Elena R.:But it's really sick. It's essentially, it's OpenClaw, but it's OpenClaw. I asked, I said, do you have like educational materials on it? And he said, you don't really need them. He's like, it asks you for a name. It walks you through the setup. It's just like, it's easier than, you know, opening a bank account or setting up Facebook.

1:51:08Elena R.:And over time, it gets better. That's the other thing, is it continually learns from its own data and the things you're doing. So it will constantly understand more things about you, ask you questions periodically. It'll have anything you want ready to go. and it's absolutely open claw for dummies, for lack of a better way of putting it. I shouldn't say for dummies. That's not cool. But for non-technical people who, you know, right now this stuff is only available to people who either are engineers or not scared of a command line or, you know, willing to break some things. And the fact is that's a very, very small percentage of the population.

1:51:52Elena R.:And something that Microsoft does well, honestly, better than any of them have done, is taking stuff like that and making it available to the rest of people. And it's something that's always made me a little bullish on them through all of this because, like, they have an audience that's bigger than anybody as far as, like, when you think of Windows machines. And at some point, if what they put out is good, it'll be difficult to stop, much like I would say about Google. But it's funny because Google's put out great models and products that just kind of consistently miss the mark or don't seem to interest people.

1:52:33Elena R.:And it's been kind of the opposite with Microsoft where they focus more on the interface and getting the interface dialed in and its accessibility instead of the models. But the truth is the interface is pretty good. Copilot makes sense for a lot of people who maybe don't want to have multiple AI subscriptions, but would like to be able to use JGBT or Claude, or maybe some of these new models that they will come out with as time goes on, because they're very much just getting started. Like this is at the end of this year, maybe not even this quarter as far as releases go. I would absolutely expect, you know, much more to come over the coming months and for them to continue getting better and better.

1:53:13Elena R.:And I thanked Mustafa for making at least one of my predictions from early this year when we did our prediction episode.

1:53:22Julius S. D. M.:Nice.

1:53:22Elena R.:I said Microsoft will be competitive by mid-year or something or have a competitive frontier level model. And I haven't spent enough time with thinking one yet to say that that is the case. But just based on how it's scored, I mean, if we're within two generations, that's, I mean, you know, and when you think within two generations, if it's scored up there with Opus 4.6, that's February. That's not like, that's not like, it's easy to see that and think, oh, that was three models ago, but there are two models ago. But the truth is that means, yeah, like, yeah, that's, that's four months ago.

1:53:58Julius S. D. M.:Well, 4.5 and 4.6 were such a difference, such a breakthrough, that I think if any model local or open... Yeah, 7 was mid. Yeah. I got used to it eventually, but I think 4.8 is better from what I can tell. But, I mean, 4.6 was all I needed. I think this will happen, where they get really excited to release something new, once they're on a roll and then they blow past something that was pretty good and uh they've both done this or something that's kind of a turd well who who knows because it depends on the use case maybe but it just it all happens so quickly that i think you gotta give these things time you know unless it comes out and it is totally ridiculed by people and unusable sure then get to the next one as quickly as possible but if not let it let it simmer that was a big mistake with

1:54:59Elena R.:their their you know constant shipping a couple months ago too was you know some of those were a-list feature ads some of them were pretty mid but the desire to put them out every day the problem is they buried the good ones with some mediocre stuff and a lot of things fell through the cracks instead of giving them some of those maybe the push they really could have used to get going.

1:55:20Julius S. D. M.:I have, I have, they were, they were, they were cooking, but, but they were like cooking for like too many, too many.

1:55:27Elena R.:I have, I have a criticism I want to share of Dario right now.

1:55:30Julius S. D. M.:And this is a thing that's now happened like three times.

1:55:32Elena R.:I have plenty of criticisms of Dario, which was maybe some are fair, some are not. Uh, but what I want to say is that, that when people talk about hitting rate limits and what Claude costs them, their answer is always to come out on a podcast and be like that's because you're using it wrong the fact of the matter is if when you create a product you have a way you envision people using it and they use it a different way people aren't wrong your guess was wrong you created it with something in mind that is not the case and it's your job to

1:56:04Julius S. D. M.:adapt the product to how the user is interacting with it and um and that is a very very common

1:56:11Elena R.:thing and and i i think that blaming every developer has not been like oh well you just

1:56:17Julius S. D. M.:don't know how to prompt it and it's like well that's a you problem to their credit they they released a blog post shortly after the 4.7 launch where they said here's three things we did wrong um and they went into technical detail on it so i i appreciate that very cool i missed that thank you for yeah i kind of read it i like skimmed over it i was like okay this is interesting but and still is kind of a turd so i'm gonna but i'm so i'm gonna like begrudgingly use it but But I still ended up using 4.6 on a lot of my personal projects where I already had it queued up.

1:56:51Elena R.:As a general rule, if you have a product and people are using it in a way you didn't expect, it's time to reevaluate what you can about your product to make it work for the way the people are using it. If you try to get them to use it the other way, they will get frustrated. They will leave. And I mean, that goes with a cell phone. that goes, you know, it doesn't matter what the product is. You know, you've got to assume that they love this. They want to use this, but it is expensive for them. How can we make that better for them? Yes, yes, there absolutely is, Brian.

1:57:27Julius S. D. M.:The Silicon Valley episode, that's what he said. It's not an us problem. Yeah. The user is wrong. The user is never wrong.

1:57:36Elena R.:The user is never wrong. The user is right. Well, here's the thing.

1:57:40Julius S. D. M.:If the user is wrong, and you put out banger like uh x essays like uh tarik does and basically and basically say like this is how we use it and then everybody adapts to the that style like that that is the way actually to solve that problem yeah um maybe tarik should be ceo of anthrabe i'm bullish on that i would yeah i'd be down with that i have yeah i have i have um I won't go into it. Anyway, I think we're at time, right?

1:58:11Elena R.:Yeah, we're at time. Well, everybody, thank you so much for watching. We really appreciate it. We enjoy... These lives are the funnest thing we do, if I'm honest. We really enjoy interacting with you all and trying out what's new and learning as we go as well. Please make sure you like, subscribe. It's a big help to us. Make sure you pop by the neuron.ai, sign up for the daily newsletter, and stay tuned. We'll have some cool stuff to announce in the near future that we can't get into yet. you're going to want to know. And I realize I've been saying that for a minute, but I promise we're getting close.

1:58:42Julius S. D. M.:And if I could give you three takeaway points, right? It is, Hermes agent is pretty sick. If you want a personal open claw system, like a personal agent system, where you can work with whatever models you want, definitely check that out. Codex is getting pretty good for non-coding. I even noticed there was a setting, which I had never seen before in Codex where if you go to settings, you can choose your work mode. Have you seen this, Corey? No. You can choose your work mode, so it's for coding or for everyday tap. Oh, yes. Yes, I have. I have. I've not seen this before.

1:59:19Elena R.:It's been maybe a couple of weeks, Max. It's new.

1:59:23Julius S. D. M.:Yeah. So definitely check that out. But that should be a front-end feature. That should not be something that's buried in settings.

1:59:30Elena R.:That's kind of their code code answer is you could do it either way you want and one's a little more

1:59:35Julius S. D. M.:use codecs for non-coding load up on all the plugins you want and definitely check that out lots of things you can't do in chat

1:59:43Elena R.:GPT that you can do there

1:59:45Julius S. D. M.:yeah exactly and third thing I don't think I actually had a third thing keep an eye out for oh Gemma we didn't even talk about Gemma 12b which is the new local model but what I would do is I would go to your Hermes agent and say hey I'm going to use Gemma 12B. Why don't you install that for me and run it? And it'll do it for you. Do it. The third thing is check out this cool stuff from NVIDIA and their new Nematron model and Gemma 12B and check that out. And have a fun weekend playing with local AI.

2:00:24Elena R.:They said Microsoft Scout can order you a pizza.

2:00:27Julius S. D. M.:Oh. That has been the Corey benchmark for a long time. Key benchmark. And it's like, you can tell it.

2:00:33Elena R.:Do you want me to? And one of the examples they showed was it saying, hey, do you want me to order your dinner in? You know, and it was like, I can get you a pizza or something. And I said, how does it know what kind of pizza to get you? And it's like, you give it your DoorDash. And it's like, oh.

2:00:46Julius S. D. M.:You know what I have not seen a good demo of? Agents buying things. No. Why are they not showing demos of agents buying things?

2:00:53Elena R.:No. Stripes dropped a good tool for it. I mean, for like a wallet system. lots of them are using ChadGBT is using Plaid now It must be that

2:01:06Julius S. D. M.:nobody's got it working consistently enough or like they don't have to Yeah, there's something there because why haven't we seen a good demo of this? Somebody should be doing that. Somebody go build I'm gonna build a little

2:01:19Elena R.:Build a wallet or we will

2:01:21Julius S. D. M.:Yeah, like don't make me do it I'm the least experienced person here

2:01:26Elena R.:I don't want to do stuff especially not in banking.

2:01:30Julius S. D. M.:Fine. I guess I'll go make the thing that's missing that everyone needs to build. Seems like that is where you should start though.

2:01:37Elena R.:Okay. Once again, we're going to leave this time. Thank you once more for joining us. We have a lot of fun. We got a cool guest going to come on here in a couple of weeks. I think I don't know the date yet, but odds are if you watch AI videos on YouTube, you very much know him and be interested. So something cool in the works. and on that note we'll see you back next time make sure to check out that interview with Tools for Humanity and Mustafa on next Wednesday and on that note that's it for today farewell for Netrun humans

2:02:09Julius S. D. M.:farewell for bye bye

From the publisher

Everyone is talking about Mercury-alpha, the mystery model that many believe could be GPT-5.6.


In this live discussion, we're separating fact from speculation and unpacking what would actually matter if OpenAI releases a new flagship model this week.


We'll cover:

🔹 What Mercury-alpha is (and why people think it's GPT-5.6)

🔹 The biggest rumors and evidence so far

🔹 What a new OpenAI model would need to deliver to move the industry forward

🔹 How Mercury-alpha fits into the broader AI agent race

🔹 Codex, Hermes Desktop, and the rise of coding and desktop agents

🔹 What all of this means for AI users, builders, and businesses


Join us live, bring your questions, and help us figure out whether Mercury-alpha is the next major leap in AI or just another chapter in the internet's favorite pastime: model-name archaeology.


👇 Drop your predictions in the chat:What do you think Mercury-alpha actually is?


📩 Subscribe to The Neuron for daily AI insights: https://www.theneurondaily.com/

More from The Neuron: AI Explained

All 106 episodes
BONUS: New GPT Memory Feature, GPT-5.6 Rumors, Hermes Desktop Agent, New Codex Plugins, MAI-2.5 Image, Etc.The Neuron: AI Explained · 2 h 2 min
Listen in VO