In short
A “bonus” roundup of newly released AI tools and models, focusing on what’s actually worth trying. They compare open-weight models vs hosted models, discuss intelligence-vs-cost tradeoffs, and explain how to test models cheaply. They then deep-dive into Unsloth Studio/Desktop for running local models and connecting them to coding agents, plus brief overviews of DeepSeek Harness and Cursor Origin.
Guests/backgrounds
The episode is hosted by Grant and Cory (co-hosts). No other guests appear on-mic; chat participants ask questions (e.g., Johan, Jamie Sepulveda, Paul Clay Design).
Key claims
- Open-weight models (e.g., “27B” class) can be run locally but require enough VRAM; quantization (e.g., 4-bit) reduces requirements at some quality cost.
- For most people, the only reason to care about model rankings is picking the best model for the task at the lowest cost.
- Local AI is the likely future for personal “super intelligence,” though a hybrid approach will persist.
- Unsloth is positioned as a top tool for compressing/running models locally and integrating them into agent workflows without burning API/subscription tokens.
Notable examples
- Quinn 27B (open-weight) and Meta’s Spark; comparisons using Artificial Analysis and OpenRouter.
- Unsloth Studio features: 100% local run, image/video generation (Minimax H3, Flux), OpenAI-compatible API, agent/tool connections, remote access, and “self-healing” tool-call repair.
- OpenRouter: compare models across providers via one API key; credits-based testing.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOOverview of New AI Tools
0:42 to 4:00
The hosts introduce various new AI tools and their features, sharing personal favorites.
“I've been tinkering with both and I'm excited to talk a little about them, what they are and why I think they're big.”
Discussion of AI Model Advancements
4:00 to 5:00
The hosts delve into advancements in AI models and the competitive landscape.
Discussion of AI Model Advancements
5:08 to 5:34
The hosts delve into advancements in AI models and the competitive landscape.
“AI Search is rewriting the rules of brand discovery.”
Deep Dive into AI Models
5:34 to 14:00
The hosts analyze and compare different AI models, discussing their capabilities and shortcomings.
“all right so we want to dive in grant yeah where do you want to start my friend i i'd let's go in order we've got them in our notes here uh let's go to quinn 3a you gotta start with quinn yeah yeah Yeah.”
Overview of AI Models
14:00 to 14:40
Discussion about various AI models and their usability.
“Like, don't let the numbers on this screen fool you.”
Historical Model Performance
14:40 to 17:20
Review of top AI models from a year ago and their performance changes.
“were the top models a year ago or maybe December?”
Understanding Cost and Intelligence
17:20 to 19:20
Exploration of the cost-per-performance metric for AI models.
“And I just have a feeling they're about to drop them again.”
Mixture of Experts in AI
19:20 to 24:40
Explanation of the Mixture of Experts technology and its implications.
“And whatever Rerank is, I was looking at that last night and going like, what the hell is Rerank?”
Practical Considerations for AI Use
24:40 to 28:00
Advice on selecting AI models based on task requirements and cost.
“Before MOE, when you asked a model a question, it raises every node throughout the model and says, you know, hey, how do you calculate pi?”
Token Usage and Financial Implications
28:00 to 28:36
Explore the financial implications of AI token usage and subscription models.
Show all 39 chapters
Comparing AI Systems: Claude vs. OpenAI
28:36 to 30:28
Discuss the strengths and weaknesses of Claude and OpenAI in AI intelligence.
“And it's, you know, the first thing they're going to say is, well, what makes that not good enough?”
The Shift from Token Maxing to Minimalism
30:28 to 31:38
Understand the industry's shift from high token usage to more efficient models.
“Jensen Wong, the CEO of NVIDIA, put this magic fairy dust on everyone.”
The Future of Local AI Models
31:38 to 35:03
Learn about the emerging trend of running AI models locally on devices.
“Amazon famously had a half-billion-dollar monthly token bill.”
Unsloth Studio: Features and Benefits
35:03 to 38:14
Discover the features of Unsloth Studio and how it enhances local AI usage.
“Like that's what they do is they take big models, compress them.”
Integration of AI Models and APIs
38:14 to 42:01
Investigate how local models can integrate with existing APIs for better functionality.
“Explain what that means for normal people.”
Understanding Open Claw and AI Models
42:01 to 43:39
Learn how Open Claw enhances AI functionality with tools and models.
“I don't know about your clawed subscription.”
Discovering and Running AI Models
43:40 to 46:07
Explore the features for discovering and running various AI models.
“Well, let's hop back over here real quick.”
Remote Access and Tool Calls
46:08 to 47:24
Find out how remote access works with AI tools and address tool call issues.
“I mean, I do it with my codex constantly.”
Challenges with Local Models
47:25 to 50:08
Understand the memory limitations and reasoning issues with local AI models.
“The smaller it is, the more you'll notice this.”
Investing in Local AI Systems
50:09 to 52:22
Discuss the costs and benefits of investing in local AI systems for work.
“compacting, which helps it compact the information so you can continue talking to it even after you hit that limit, Onslaught just kills it.”
Overview of New AI Tools
52:49 to 56:00
Get acquainted with the latest AI tools like Cursor Origin and DeepSeek.
“The next thing on the list is, so we're going to talk about some of these that came out, but we're just going to roughly overview them, because I don't think you've used either of these.”
Exploring Plugin-Based Models in AI
56:00 to 58:25
Learn about the concept of plugin-based models and their implications for AI usage.
“which is the code base, and you could use this to run a model like Gwen.”
Introduction to GitForge: A New Tool for Developers
58:25 to 1:01:00
Discover GitForge, a new tool designed to enhance code management for developers.
“Yeah, I've got it here in the in our doc as well.”
Buzz: The Slack for AI Agents
1:01:00 to 1:03:23
Explore Buzz, an innovative collaboration platform designed for human-agent interactions.
“about that next week, so tune in for that.”
Features and Fun Interactions in Buzz
1:03:23 to 1:08:52
Learn about the unique features and fun interactions available in Buzz for agents.
“So they can go and work in the channels you allow them to work in.”
Introducing Bird: Fun AI Tools
1:08:52 to 1:10:00
Get introduced to Bird, a creative tool that brings fun to agent interaction.
“and typing you a message is what the other one means.”
Exploring Fun AI Tools
1:10:00 to 1:16:39
Learn about creative uses for AI tools and their interfaces.
“It's the reincarnation of Clippy made manifest, but now if anything it could be Clippy.”
Real-World Applications and User Experience
1:16:40 to 1:21:19
Discover real-world use cases for AI and how they improve user experience.
“You know, we could talk briefly about GrokBot and all that stuff.”
Personal Assistant AI and Consumer Challenges
1:21:20 to 1:24:00
Understand the potential of AI as a personal assistant and discuss consumer frustrations.
“and eventually you're just talking like, hey, can you do this for me?”
The Impact of Dark Patterns on Software Success
1:24:00 to 1:25:20
Learn how software companies can succeed by being transparent and user-friendly.
“Yeah, you can't be a jerk in this world with these tools.”
Critique of Credit Systems in Software
1:25:20 to 1:26:40
Explore the issues with opaque credit systems in software pricing models.
“their own version that isn't as dark in terms of dark patterns.”
Evaluating AI Tools: Quality Over Cost
1:26:40 to 1:28:10
Discuss the importance of quality in AI tools versus their pricing structures.
“Well, a lot of times the companies that are offering this are like pass-through, you know, they call them layer companies or wrapper companies.”
GrokBot and AI Integration in Workflows
1:28:10 to 1:30:00
Discover the potential of GrokBot in enhancing AI-driven workflows.
“And I was like, I want to try this, but I don't want to try it$300 bad.”
Data Centers: Economic and Environmental Concerns
1:30:00 to 1:38:00
Examine the challenges and benefits of data centers in the context of renewable energy.
“Bots can sign in to your tools, use them just like you do, and come back with finished work.”
The NIMBY Dilemma in AI Development
1:38:04 to 1:39:08
Explore the conflict between AI development and local opposition to data centers.
“noisy there's a ton of traffic like to me it's it's more become it's less about the reality of data centers and more about this is how we rage at AI.”
Rethinking Data Center Efficiency
1:39:09 to 1:41:38
Discuss the need for efficient data centers and innovative solutions in AI.
“But like at a certain point, you need to scale down in order to scale up.”
Promising AI Applications and Public Perception
1:41:39 to 1:43:52
Examine the public perception of AI and its potential to improve lives.
“here's all of the powerful algorithms that we're running in order to make your lives better instead of replace you at your job.”
The Vision of Decentralized Data Centers
1:43:53 to 1:46:26
Discuss the idea of decentralized data centers and self-sufficiency in technology.
“Yeah, I think both perspectives are valid.”
Innovative Business Models for Home Data Centers
1:46:27 to 1:49:04
Consider potential business models for individual home data centers.
“Well, you know, there have been companies that wanted to do that with blockchain.”
Transcript
Automatic transcript. May contain errors.0:00Good to see you. I don't see anybody. I see you Grant. How are you Grant? I'm doing good. Yeah, let's get started in the chat, y 'all. Where are you tuning in from? Let us know. Yeah. Let us know. Let us know. We got Prairie Mountain Man saying howdy. So I'm assuming that means they're in a prairie and or a mountain. Or are a mountain man, in which case they should come on. Judy Nelson says, well, it's time. It is time, Judy. We agree. It is time. And we were, despite our usual punctualness, we were two minutes late today. but it is to your benefit that's right that's that's right we're in we're in better shape now than we were a minute ago um well it's great to have everyone here we're excited to see you today we're going to be by diving into all of the insane number of tools not maybe not all of them but a bunch of the insane number of tools that have dropped in the last week uh specifically uh quinn 38 unsloss new desktop app um cursor origin uh deep seek harness we're gonna touch on those you probably if you're a normal person you probably don't care about those but we'll tell you what they are yeah uh and a couple of really really fun apps that are my personal favorites right now from uh from block jack dorsey's company so yeah i'm excited to learn they're really cool.
1:25I've been tinkering with both and I'm excited to talk a little about them, what they are and why I think they're big. I love when a tool comes around that does something others aren't really doing yet or at least not in a way that's accessible to the masses. This does that with a couple different things, so that'll be a cool one. Might even get into, there's the meta models, there's GrokBot and Grok 4.6 which now sits firmly between open AI and Anthropic on the artificial intelligence, intelligence chart, artificial analysis, intelligence chart, which is a big deal. The, the, the frontier is getting crowded.
2:08Have you tried? Well, the, let's put it this way. The public frontier is getting the public frontier. Yeah. Yeah. I think, I think at this point we can safely say that the, the two uh leaders are far ahead with unreleased models that they have not released yet yeah i'm sure and and and supposedly musk has another one coming in the next week like four seven should be i think next week based on what he said and supposedly it's going to be better than everything that exists publicly right now yeah yeah god knows what they have behind closed doors i assume they They are probably like, you know, the way this works for so many of them is they, they, you know, they're dealing with, you know, when you think GPT-5, it's just one giant model and they're running through and they keep training and training and training and they bust out checkpoints and see how that checkpoint tests.
3:00And then that becomes your 5.6, but, but they're still training, they're still building, they're still expanding. and uh so yeah i always assume there's there's a lot that we don't see and and you know sometimes it's not like gpt-56 is a model they've been training there might be there might be eight gpt-56s and they're deciding which one is the one yeah or i've even heard officially anointed you know the experience of using chat to be too you're really using like eight different models under the hood they just call it like chat gpt yeah five eight six or whatever um they've they've since like actually made that public with like luna and tara and soul and all these different versions yeah are you seeing this people saying we're not here uh i'm seeing that's probably someone who let me write in the chat refresh the video but it happened to me where it just wasn't loading the video wasn't loading what the must be a scam no we're here just refresh your page okay anyway so i want to just acknowledge a couple people in the chat what up idaho missouri i'm assuming brandon ms is missouri right cory yeah um denver toronto new york what's up a lot of uh a lot of east coast folks today a little bit of west coast representation very cool very cool I just dropped the live link in the chat too so hopefully some folks will see it auto repair software you're up auto repair software okay Judy's got it figured out oh good good deal she's going to figure it out good good good I'm glad I hate to see people disappointed.
5:00Before we get going here, just a quick note to say that today's video is, or today's live stream, is brought to you by Arefs. AI Search is rewriting the rules of brand discovery. And Arefs Brand Radar shows you how often your brand appears in AI answers, what sources influence those recommendations, and where competitors are winning visibility. monitor chat gpt google ai overviews gemini perplexity copilot and more all from a single dashboard go check it out in the link in the channel that grant just dropped there for you all right so we want to dive in grant yeah where do you want to start my friend i i'd let's go in order we've got them in our notes here uh let's go to quinn 3a you gotta start with quinn yeah yeah Yeah.
5:48So for people who don't know, so like, let's say you're a normal person and you have used ChatGPT and you've used Cursor. We were just talking a little bit about models. So, you know, the underlying AI model that is under the hood and running when you use ChatGPT or Cursor. Well, there are also models that are open weights. And what that means is that instead of just being hosted by ChatGPT or sorry, being hosted by OpenAI, being hosted by Anthropic, the actual code that makes it work, what's called weights, is public, meaning anyone could host it if you have a computer of the right size. Which is$100 ,000.
6:29Well, it depends on which version we're talking about. But the big ones is like you need a lot of compute to be able to run something like that. Funny enough, mine will not run 27B successfully. Let's explain what 27B is first. So the one that came out this week, right, is what Corey just said, a 27B, 4 billion, parameter model, which means it's a lot smaller than the$100 ,000 computer that Corey was just talking about. Yeah. Well, that means that potentially you could run it on a computer like Corey has. Although, what's your VRAM gigabyte size? What you should know is I'm running, just for the sake of saying this, a$5 ,000 Dell Pro Max with an RTX 4000 ADA, not the BlackWalt, so it's less VRAM.
7:18Which is a graphics card. Yeah, it's like 20 gigabytes of VRAM. And it absolutely, it'll load it, and it'll start thinking, and it'll start to give me some words back. And then it says, your context is too high. And I run out of, like, KV cache, which we won't get into. But essentially, any context window whatsoever, and I pretty much can't run it. So I've been playing with it through Open Router instead. But it kind of kills me. like i've run models bigger than 27b but um which which is you know not always advisable and and i'm running it at see this is going to get really technical and i didn't want to get into that i'm running it at 4-bit which is like when when an open model releases and you see all those benchmarks a thing to understand is that that is most likely not the local model your computer is going to be running usually they're 16 bit uh and in order to make them run on consumer hardwares they shrink them down shrink them down shrink them down shrink them down it's called quantization or distillation uh yeah and um so they shrink it down shrink it down shrink it down so it can fit on consumer hardware with that though of course you lose quality you know when you squish and remove information you are going to sacrifice some amount of quality for that so like you can think of that roughly as losing a certain amount of intelligence basically if you're talking about that chart that you were just um yeah so the four bit model that i'm running is 16 gigs needs 16 gigs of vram to run which is why i should be able to run it uh i don't have a mac studio somebody just mentioned that um yeah i have my own mac's a little older that is a hot machine if you have one of those yeah um let's see where can we to put some visuals to this so cory mentioned this chart earlier this is the artificial analysis intelligence index this is essentially how all of the intelligence of all the different models compared to one another so as you can see right here overall intelligence looks like claude opus by max is the winner at the moment which is bizarre to me because i've yet to hear anyone say a nice thing about opus 5 uh people love fable but my god opus 5 gets some hate uh the way i the way i think about it is you don't talk to opus you talk to fable and fable's like opus's wingman and fable will tell you what tell opus what to do on your behalf so that's the way that it has to be figured out signals a product problem uh in my eyes uh or or a communication problem that maybe they need to be more clear about because like yeah most of what i'm seeing is it working on its own as an agentic tool where people seem to be really complaining um well actually that's where it seems it does make a lot of mistakes but it then fixes the mistakes it's interesting yeah uh yeah i don't know it's just the whole thing that you were talking about earlier with uh they have all these checkpoints and then they release a model at a checkpoint.
10:35But just because you have a checkpoint that's successful internally doesn't mean that it's going to be well-received by everyone else. If you use models via API ever, you know that when you start using an API model, there's new data. And that date is the checkpoint. When suddenly there's the same model, but it's got a different date, that's a fresh checkpoint of the same model. I am increasingly
11:08Disinterested in benchmarks There are things about this one In particular right now that intrigue me Like number one Is that Opus 5 doesn't match what I keep hearing Like I don't argue with Fable I don't argue with 5.6, I don't argue with Grok But that one kind of surprises me but there's grok's up there yeah yeah grok's tied with gpt 5.6 soul pretty wild actually it is it is and kimmy k3 is one point behind them uh that's another open model for people who don't know yeah that one is way too big to run on your own machine oh yeah that's millions of parameters i've heard some people do it uh i talked about this a little bit on our stream with james but i don't know yeah and kimmy's like two and a half trillion or something like that yeah um but here's here's the thing we were just talking about quinn 27b grant can you hover over 27b there since it's on your screen you still see this here's the thing to know that a 27 billion parameter model sitting in the 50s is bonkers and and and you'll also notice it's sitting right there next to luna which jet gpt released in the last 30 days it's not their core state of the art but it very much just got released and and it's a better model than and honestly some i say that i think people have kind of come around to luna a little bit but that a 27 billion parameter model is sitting right there is unheard of in my opinion uh but there's another newcomer right up there too I don't know if you noticed, but Meta's sitting right up there at a nice 57, just four points behind ChatGPT.
12:58Oh, you're right. You're right. The Spark. Have you tried that? Spark's legit, man. It's good. Their coding harness is good. I've even played with Muse Glimmer some, which also came out last week. That's their current open source. Muse Spark 1.2 for Meta will be coming out open source. And I think what they said was the coming weeks. Yeah. So Glimmer would be equivalent to Quinn. But then if you look at the 27B. Yes. It's a big jump. It's a big jump. What I'll tell you that's funny, though, is like if you go, I was last night, open router, both of them. I was using both of them to do a few little like, I had Quinn troubleshooting why Quinn wouldn't work on my desktop.
13:49and it was doing a great job, by the way. And I was asking other questions in another tab to glimmer on the same subject, and they were both really fast. They were both really good. Like, don't let the numbers on this screen fool you. Like, if you see it on this screen, what you should know is there are hundreds and hundreds of models. If you're looking at it right now on this screen, it's definitely usable for something. It's pretty good. Pretty decent. The only thing I would say is not worth your time is Haiku. Fair. Memetron 3 Super is a surprisingly good model. And Ultra. Ultra. We have a question in the chat.
14:39Johan asks, do you remember what models were the top models a year ago or maybe December? well artificial analysis has that chart for you and let me guess before you bring it up grant yeah well i've got it here just don't read the screen okay i'm not looking i'm gonna say claude 4.5 gpt 5.1 because 5.2 is when gpt really took back off the closest is august 5th uh the top model was clawed 4.1 opus so pretty surprised it's not gemini 25 pro there i really rock 4 had a fleeting moment in the sun in july 10th then we had gpt5 then we had in september 29 it was 4.5 sonnet then we had gemini 3 pro in november then clawed 4.5 and i think this this release right here, November 24th, 2025 was what changed everything.
15:41Because you look at how everything's slowly growing, it takes longer, it takes longer, then we get to November of 24, sorry, 25, and look at how quickly, in such a short time span, they've jumped in terms of intel. By the way, did Grok, I mean, go back to Claude, how big was the gap between Claude 4.5 and the next GPT issue? Right here? Yeah. um pretty close so they were actually tied okay yeah they're generally kind of sitting on each other pretty close yeah but yeah so that's so that to answer your question that's like what the best models were about a year ago yeah yeah that november moment was was claude's really hot heyday like november through march february march i mean like still doing well don't misunderstanding but there was this point where it was absolutely on top of the world uh and uh that was there and then uh yeah it's been interesting so so if you're a girl though god yeah yeah time flies in this industry in particular so um if you're a normal person do you care about this like what why should people care about this what's your take on that like this chart me yeah or anybody for me the there's only one reason i think anyone should care at all and that is just figuring out what's the right model for your task you know if you're pretty model agnostic and don't really like you know like to some people these are sports teams and and uh uh if you're not like that and what you want to know is what's the best model or if you're a business what you may want to know is what model will do my task just fine for the most reasonable cost like that's a question you're going to get used to hearing is people talking about the economics of it yeah and this is the chart for that yes this is a chart for intelligence yeah cost per intelligence it's this chart so you can see here i believe it's intelligence or no cost cost is ranked yeah what what's the x and the y the uh oh it's uh it's got to be cost i think because anthropics is the highest yeah yeah it's it's cost per task like they have these token tasks that that it has to do and they average it and yeah so basically like if you're trying to get clawed to do the task it's going to cost you $3, whereas if you're trying to get any other model below Clen 3.8 Max It's going to be the sales tax on a piece of bubble gum for the others Yeah Lots of options down here Lots of options Yeah, and they fall quick Like, I mean Even Sol Max is Dirt Yeah And that's pretty cheap too $1.23 $0.23.
18:46Yeah, it really is. And I just have a feeling they're about to drop them again. Have you noticed how many open router-type sites in the last two days have started offering it at 50 % off? Well, yeah. I mean, I haven't noticed that, but I think that's a good observation. And not only that, something we didn't talk about is that open router was just bought by Stripe for$7 billion. always and stripe declared the singularity to their investors let me show yeah under their own definition of the singularity uh i think a lot of people fall i think that was just like okay good hypey line but a lot of people feel we're kind of there you know what i mean yeah you know it's not the singularity until stripe confirms it just real quick i want to show my screen this is open router for people who don't know so what you do is you get an api key and then you can compare different models together across text, image, audio, video, and some good examples.
19:50And whatever Rerank is, I was looking at that last night and going like, what the hell is Rerank? Why do I not know what that is? Okay, here we go. Real quick, while that's loading, can we address this follow-up question on that last point real quick? can we say the current models are now equal to that time in the model race they're better uh do you mean current open models yes yeah he said current open models i i just didn't say it out loud um it's close i don't know if i could say it's like exactly that but it's close you know when you look at how those models scored on that same intelligence index we just looked at though uh they scored a whole lot lower like you'll notice like when you look at it that it only shows the most recent stuff from anthropic and chat gpt but if you go back and look like quinn 27b matches gpt 5.4 on intelligence wow uh which is almost unbelievable yeah it's almost unbelievable uh yeah so so that's a really big deal i'm sorry grant i didn't know you're good You're good.
21:01I'm interrupting myself just to show you this. So basically, for example, you could add a model. Like, let's say you wanted to compare, let's say, Quinn, and then let's add another one. And then let's say you said GPT 5.4? They might not even offer that anymore. Yeah, they don't offer that anymore. No, you said 4 ,5. Try again. Oh. Typhos. ROC4. No. No, I don't think they offer it because it's too expensive for them to run it. Okay. I vaguely remember that. No, that was 4.5. I was saying 5.4. Oh. You said 4.5 earlier. Oh. Oh, okay. Well, I was meaning 5.4. Okay, let's see this. Hold on. 5.4. What, Mini Pro?
21:49Okay, I guess regardless. So say you want to see how both of these compare against each other. You can give them the same task. ask you have to buy credits to do it and they'll both answer it so you can see yes they'll both answer it and you can you know hook this up so that you can use this in whatever um tool that you use your ai um to code in uh or you can just use it directly here in the chat and you know you pay 20 and you can use that across all these different models so that's why this is a really good website this um so let me share the let me share the website that i was sharing a minute i I just dropped both of those in the chat.
22:27Artificial analysis and open router. Awesome. Yeah, yeah. Yeah. But yeah, so basically the point is you can test all this stuff on open router. But if you do have a machine that is powerful enough, you could run QN 3.7, or sorry, 3.827B via a tool like the second tool that we're talking about today, which is Onslaught Studio. Yeah. Hey, in the model drop down there, add GPT 5.4. Right, but yeah. Which version, though? I guess extra high, since that's what they use on all the other ones on here. They don't usually show medium and stuff. Oh, sorry, I clicked something. Oops. Okay, 5.4. I don't even see it.
23:21It's a scroll. unless i am blind did it let me let me try it again
23:31yeah i don't see it either oh here you go 53 53. yeah gosh sorry clicking really quick and 20b is 52 27b is 52. wow wow yeah that's that's pretty wild yeah so you go back to gbt3 and it's it's past gpt3 and the uh cost per intelligence for task on that one is 33 cents yeah and you know a thing to know about it is is like i said you're usually going to be using a quantized version on open router you're probably not on open router you're probably using the the whole whole thing i'm not yeah because the way that that works is they have like their own cloud server where they host it and then they serve it to you like yeah the other companies do and that thing to know is that like this is not a traditional dense 27b model or sparse what this is is uh and this is a technology that came around two years ago now i guess it's called mixture of experts if you're not familiar the thing to know about mixture of experts is the way to think of ai before moe and after MOE.
24:44Before MOE, when you asked a model a question, it raises every node throughout the model and says, you know, hey, how do you calculate pi? So it's asking every bit of knowledge inside there. What MOE did was created experts throughout that model. So like instead of, think of it this way, if your boss needs to talk to someone in HR, should he call an entire company meeting to do it and then have the conversation in front of them because that's what it was doing before. With this, he's able to say, I want to talk to this department head and that department head and bring in two people instead of 270 billion or something.
25:27However many it is. But as a general rule, these models will come with a number of experts that can be active at one time. It might be 24, it might be 12, it might be three depending on the model and the size. And those and what that does is it gives it the ability to work faster and for a smaller sized model to perform at a much higher weight class. Like, you know, if you had a welterweight walk into scrap with Mike Tyson and win or, you know, not die. Maybe not die. Yeah, just stay alive. All right. That's the end of that tangent. But I felt like in looking at these, it's important. Understanding MOE is helpful.
26:14So for normal people, what I would say is you care about this for two reasons. One, what is most likely going to help you do the task you're trying to do? And a lot of times you just kind of have to figure that out with trial and error, which is why a tool like OpenRouter is really, really helpful. because you can just try it on a lot of different, you know, you can try the same task on a lot of different models for, you know, one, you know, paying one price. And, you know, you can see which one does it. You can even compare costs and say, okay, this one, they both did it, but this one did it for cheaper.
26:46So I'm going to go with that, which brings me to the second point. The second reason that normal people should care about this is you probably have things that you really need a lot of high intelligence for and things that you don't need high intelligence for, but you need to do a lot. in which case price matters. So you want to know which is the thing that I can do that I know the intelligence will give me the power to do, and which is the thing that has enough intelligence to do the thing I need to do on a recurring basis affordably. Yes. So that's my recommendation, and the tools that we just shared, artificial analysis and OpenRouter, are both very, very helpful to do that.
27:26That's a discussion you're going to be hearing a lot in businesses over the next year, because what's going on right now is that the cost to serve a unit of intelligence, which in the way we use them is a million tokens. Let's just say it's a million tokens for simplicity. The cost to serve a unit of intelligence is going down and continuously going down. However, the demand for intelligence in the amount one person uses has gone through the roof. If it's a charted look like this. and uh and uh as a result what's happening is intelligence is way cheaper but instead of using 800 000 tokens this month you're using 200 million 500 million a billion grant i literally had a conversation with a person this morning who's burning a billion a day right now who you know by the way we'll talk later okay yeah i mean you could that's that's pretty uh pretty out of control billion a day i did about a month but but keep in mind this was on the chat tpd subscription this was not paying per token yes no no no no how subsidized those are um buy the subscription for the love of god do not do not go well that's a good point api so like this is all interesting and you know you can see okay pretty much at the top here you've got you've got Anthropic there's a lot of intelligence that's Claude they make Claude then you've got OpenAI which is a lot of intelligence but much cheaper why would you buy something else?
29:03like why sacrifice intelligence? that's the chart that I feel like is going to hurt Claude in an executive meeting I feel like if you walk to a CEO a CFO, product manager and you sit down with them and you have to say and then you show them the intelligence list, and they're like, oh, it's one better. One what? Nobody knows, but it's one. And it's, you know, the first thing they're going to say is, well, what makes that not good enough? And in a big hurry, you've got an argument that's really hard to make. And I think, you know, the argument that we made in the newsletter over the past six months was that the argument was the interface.
29:49Claude was just they made it really easy to use all of these different tools like skills and plugins and their co-work tool and co-work and Claude code they made it really easy and they basically have innovated the interface over the past six months I would say if not a little bit longer and that was the reason why you would still go with Claude even though they were more expensive it's just easier for normal people to use it but chat gpt has caught up and is caught up on intelligence it's caught up on the interface and also uh you know it's cheaper yeah you know and uh and they're both great that's the truth of the matter you know if their prices are competitive you know you do what you mean you can absolutely do what you want anyway if you just don't like money uh or yeah there was this insane idea that had taken off in Silicon Valley.
30:45I don't know. Jensen Wong, the CEO of NVIDIA, put this magic fairy dust on everyone. And he said, token maxing. You need to be spending$200 ,000 a year. If my half million dollar engineer isn't burning$200 ,000 a month in tokens or something like that. I think it was a year. I don't think it was. I think it was crazy. Well, basically this idea got out there that you need to be token maxing, meaning you need to be spending as much tokens as you humanly can because the value that you get is so great but then all the cfos you know speaking of mitchell d lee we're like whoa whoa wait a minute and they saw these bills that were just astronomical may when april's bills came through shortly after he said it at the end of march exactly and i feel like we've the the last three months has been a reaction to that which has been token minimalizing so it's like okay how do we do things as efficiently as possible Amazon famously had a half-billion-dollar monthly token bill.
31:43Yeah. That's out of control. And that's where the fun goes. That's where the fun goes to die. Okay, let's get back on track. What was next, Grant, on our list? So, Unsloth Studio. And I want to talk about this. There was one point I wanted to make about this before we transitioned. Go ahead. I'll share screen on the next one if that's okay. Yes, please do. So, there's one other reason why you might want to care about having a small model. And, you know, that's because you can run it locally, meaning you don't need a data center to run it. And it's my personal belief. I think a lot of people will come around to this in the future.
32:19Maybe it'll take five years to fully play out. But the future of your personal super intelligence, if you want to call it that, is going to live on your computer. You don't want to be renting out your intelligence over the cloud. like people are going to want their own intelligence on their own device so that they can customize it however they want and share their thoughts only with themselves and not to who knows where over the cloud you're going to want this capability let alone the fact that you know if you care about what's happening with the data centers as a lot of people do right now you might want to be like whoa whoa whoa let's not build so many of those let's focus on making the intelligence work on device so we don't need as much of that now there still will be reasons why you want bigger and better models to do amazing things like what just has been happening with the mRNA cancer vaccine we wrote about this morning.
33:09But... What a deal. You don't need... Look at this. QN27D, you can run it on a computer right now. So we can do it. We know it's possible. And I do think that this is the future. And the tool that you can use to run it right now is what Cory's about to share. Yeah. And I half agree with Grant. That's fair. I think the approach will be more hybrid. I think you will still have, you'll always have individuals who just don't care and just want something to work. We might get to a point where local AI is like that, but I think that is farther down. But I think you do get to where technical people very much would prefer that.
33:58and many would prefer that now as far as technical people goes. I argue that many normal people would benefit from it and want it. They just don't maybe realize that it's possible or they don't understand. Yeah, or no. Well, certainly don't know how to do it. And it's not good today. It's not as good as it can and should be. I agree. And I think there needs to be more emphasis there. Yeah, if you get over, Like if what you're doing requires embedding models, transcribing models, 8 or 9B models, those are great. But if what you need are 20, 30B up, they're good. But it still requires a fair amount of computer.
34:43And I would expect that to just continually get lighter. Like if you'll notice, Quinn is 27B, not 30B. Like is kind of the norm. So I think we're going to watch that just continue to trend down in quality. Let's talk about Unsloth because they are one of the top companies making that happen. Yeah, what you should know about Unsloth, before I show you Unsloth Desktop here, is that Unsloth, we were talking earlier about model quantizations, which are where they shrink the model so it'll work on computer hardware. In my opinion, they own that market. Like that's what they do is they take big models, compress them.
35:24So they'll work on consumer hardware and they are almost without fail, always the best version in my opinion. Um, so a few days ago, was it Monday grant? I don't remember. It was really, really. They had two releases back to back. There was a desktop and studio, but they're essentially the same thing. Yeah. I think it's studio, the online version. Perhaps so, yeah. I am not certain. How's my framing here? Whoa, hang on. It was good until that. Is that good? Yeah, this is good. Okay, cool, cool. So this is Unsluth Desktop. It is a tool where you can run these models on your own computer really easily.
36:10But it's so much more than that. For example, if you've ever tinkered with them, you've probably used a tool like um lm studio is the most user-friendly approach in my opinion to how to run a local model on your computer you know you can literally go in look at the models right there click one it'll download it you just hit run and then you're in a chat interface uh but there are areas where in my opinion it gets a little extra complex is missing some features and could be better. And Unsloth came out with this tool here that answers a lot of that. This is what it looks like on the inside. You know, it looks like just a normal chat window.
36:53You go over here. Let's see what they've got. Let's walk through some features first. It's open source. So, you know, you can fork off of it, build your own, change it how you want it. 100 % local run. um it does image and video generation which is wild that's uh video generation is not a thing that uh that they've done it's traditionally you know comfy kind of owns the local creative market in my opinion comfy ui it's a great tool um but they've added in video and image generation with the latest models minimax h3 which is great flux z image laura adapters I want to say SeedDance is in there as well.
37:39Here's the cool part that really, in my opinion, starts to set this apart, is you can connect your own agents. So you can come in here and connect your Claude code. You can connect your Codex or your GrokBot or whatever you're using as your various tools and connect them directly into Unsloth Studio and run some of those local models right there in Codex if you want or right there in Cloud Code. And they'll do calls, and it's just really cool. It says Unsluth exposes an OpenAI-compatible API. That doesn't really mean OpenAI. Explain what that means for normal people. Most APIs in the AI space are called OpenAI-compatible because since they were first, they got to set the standard, kind of, if that makes sense.
38:31So they all work similar. So existing apps, scripts, and SDKs can connect to your... Sorry, just to put a finer point on that, that basically just means that you can call an OpenAI model when you call their API into any application that you build. So if you say to Codex, which is ChatGPT's coding agent, hey, build me an application where I can chat with ChatGPT like my own ChatGPT, then you would connect it to the API and then you can actually stream the models in. So the fact that it's an endpoint, it works the exact same way. So this works the exact same way as that. Exactly. And yeah, and that's what this will do.
39:14It'll let you connect your local models in an interface you use already. So imagine if you could spin up, I don't know, your codecs and you've got an agent running and you want it to go do web searches and you're like, hey, use Quinn 37B in unsloth to do that and it goes and goes and uses the local model to do it instead of burning up your your uh monthly token quota in your subscription now it has excellent web search tool um i was that was a thing that was really impressive in working with it this week um because that's not a thing i felt like was really strong in in specifically lm studio like You know, you can set up tools, but it's complex.
Read the full transcript
40:00Some of them require coding. Some of them are right there and ready. But this just makes it native. It's right there. It's where you expect it. It looks like a little globe. You just click the button, and off you go. Oh, Corey, before I move on, Jamie Sepulveda asked a relevant question. When you connect something like Unslock to Claude or Chat, does that cost against your subscription when you need to purchase tokens? No, these are free. These are local. It's running on your computer. Unsloth is completely free. These local models we're talking about are completely free. And what you will do is use these so it doesn't burn your tokens.
40:41So not every task requires the same model. So you can use it to, you know, maybe it's, like I said, maybe it's web surfing tasks. Maybe it's information gathering. Maybe you're using it for bigger stuff than that, but you could from Codex call this, and this model will run for free instead of one of the OpenAI models that is pulling from your subscription quota. I hope that makes sense. I can do that again if that's not clear. No, that makes sense. I want to add something to it. So one thing you could do is, say, for example, you use Codex, which is OpenA as coding agent. You could say to Codex in its rules, hey, make sure that whenever you have a task that's pretty simple and well specced out, that you assign that to Quen27B to go do for me.
41:34And then that way it will send a request to Quen. when we'll go off and do it, and then you're not costing your OpenAI credits on that. Exactly. I want to real quick address this question as well, because it's a good question. How does Unsloth compare to OpenClaw? A thing to understand is that they're actually very different things and could be used together. OpenClaw is a harness in the same way that Codex and Claude Code are. Could you explain that? Yes. Essentially what it is, is a place that houses is your model and a bunch of tools so your model can think of it like giving your model arms and legs to go do stuff like a spider and and uh and it sits here you know in this little nest and it's sending calls out it's doing different things like it's going to go check your email it's going to write you a google doc it's going to open up a spreadsheet from your boss that kind of stuff now the thing to know about an open clause you can set that up to where it runs on your codex subscription.
42:33I don't know about your clawed subscription. There was a big fuss there about that at one point, but it may have, it's probably fine now. Uh, yeah, we'll circle back to that. Cause I, I don't know. Uh, but the deal is you can also connect your open claw to local models on your existing computer. So you can use both of them. You can use all of them. That's one of the things that made open claw so cool when it first came out was just the, the diversity of what it allowed like it would that wasn't a thing we had had before where you could just connect every app you own to one harness and uh you know and maybe harness is the wrong word for anyone who's not technical hub think hub like it's like your ai hub yeah yeah yeah i'll allow it thank you i think harness is perhaps a little bit more on point just because it's like it controls it like If you think about it like a harness for a horse, you're steering the horse.
43:32That's fair. You're pulling it this way, pulling it that way. It's a similar thing here. It helps you sit in the driver's seat to make sure that the agent doesn't go off the rails. Exactly. Exactly. Well, let's hop back over here real quick. Models that actually run code, of course. That's a given. The latest models. One of the nice things is it has this discover feature where you can go in. And it's going to immediately recognize your system. It's going to know you have this much VRAM, this much RAM, this processor. It's going to know your machine immediately when it gets there. And you're able to click on device, which you see right here.
44:12And on device, we'll immediately grab any model you've ever downloaded before, no matter what app it's in. Like it shows all of my LM Studio downloads just immediately. But in Discover, you can look at other models and see what else is out there. And you can usually filter these by, can I run it? What you'll notice is these little, there's going to be these little check marks. You see right here, you'll see this little green circle with a little letter I in it. That means you can run it. Sometimes they'll be yellow, and that'll mean, like, you might be able to run it. And sometimes they'll be red, which means, give it up, bucko, there's no chance at hell.
44:52You better pull up OpenRouter if you want to use that Exactly But at least you know And once you come in this section right here Where you see this UDIQ4XS That is one of the quantizations And there might be For example for this model For Quinn 3.6 Or 3.627B that we've been talking about That's not right is it? No 3.827B is that there'll be, when you click this, there'll be a dropdown and there'll be 25 different codes like this that don't mean much of anything to you, but it'll tell you - The number matters. So if it's two, if it's four, if it's eight, that just shows you like how much smaller it is.
45:36Yes. So two is the smallest, four is probably what you want most of the time. And each one will have this. Each one will have this and you'll be able to tell really simply by just a glance. I can run that one. I can run that one. And, you know, you want to pick one of those and go from there. What else do we have here? Access your models anywhere. Yeah. Remote access is a function, which is delightful. I have not tried that yet. It's it's intimidating, but it's awesome. Have you done it? Not in unsloth. I mean, I do it with my codex constantly. Exactly. Same. I was getting tires fixed a few couple of weeks ago.
46:15It's sitting at Walmart waiting on tires. And I had like four agents going at home through my ChatGPT app on my phone. I was like, this is wild. This is the future, man. This is the future. Yeah, for people who don't know, basically remote access lets you say you're using codecs with files on your local computer, which you can do. You might not know this. when you use code work or cloud code or chat to these codecs or chat to put you work it can actually access files on your computer and via sandbox and then you can actually turn on remote control and then that means you can access it on another computer or on your phone and so that's how cory is able to on his phone talk to the agents that are running on his home computer now the agents aren't necessarily running there they're still over the cloud but for example if you're using Unslock, the agents would actually be running on his computer physically on there.
47:09This is one feature they don't call out on the front page. And I want to call it out because if you've ever played with a local model, you know what a problem this can be. And that is self-healing tool call. Please explain it. Yes, I'm going to. When you are using especially smaller local AI models like 4B, 9B, getting into The smaller it is, the more you'll notice this. What happens is those tool calls that run when you're like, I want web search. I want you to go grab a file. I want you to do this. Sometimes those tool calls get corrupted.
47:47What happens is there's a glitch or there's a something, and basically it's coming through as code, and it just goes blah and messes up. And what this does is this detects it, repairs it, and sends it on through like nothing ever happened, which is really, really sick. And a thing that I'm really excited about. Okay. One question from Paul Clay Design. So for people who don't know, there's another tool similar to OpenClaw, which we mentioned earlier, called Hermes. And so Paul Clay Design asks, can you use Hermes to connect directly to local models like Quen 3.827B? Yes, you can. Absolutely. You could either do it this way directly through like an Unsloth or an LM Studio or an Olama, whatever local AI tool you use.
48:39Or you could go to Open Router and just get an API connection to that model and use it through there that way. But absolutely doable. Yeah. Yeah, the one thing to keep in mind about OpenRouter is it's still over the cloud, right? So you're still submitting your data, and it aggregates different cloud providers serving the models. So you don't technically know where you're sending that, like what provider you're sending the data to. So just keep that in mind. And, you know, if you're using it for non-sensitive stuff, I'm sure it's fine. but if you want exactly you're using it for a sensitive topic you might want or sensitive information you might want to use something like unsloth for that for sure and i'm dropping in i did a full written review of unsloth desktop this week and i'm going to drop that into the chat right here for anyone who wants to check it out later on um yeah i've had a lot of fun with this tool it's been really cool um the two things i want to flag the two things i want to flag about this before we move on so number one in my own experience testing it the biggest problem with a local version of quen 3.8 27b is memory and this is something that you're going to come up against so they've solved tool calls they've solved a lot of other things they have not solved memory so we just want to talk about tokens roughly equivalent to words you can fit inside your chat before it loses.
50:06And unlike OpenAI and Anthropic, who do this thing called compacting, which helps it compact the information so you can continue talking to it even after you hit that limit, Onslaught just kills it. You hit the memory limit and you can't talk anymore because it just can't handle it. And part of that is just simply a limit of the way these work right now. I run into it occasionally in LM Studio as well, especially with depends on the model some models run heavier if that makes sense and are just a little more difficult and Quinn 27B kind of does because it's dense because it's good the flip side is I can't run it well but man it's good another thing the second thing is reasoning so everyone says like you need to run this thing at the lowest reasoning effort possible because it thinks a lot and that will then lead to the memory problem which means it will spend a ton of time thinking about the answer and then it'll burn through your tokens and then you'll hit your limit and now you have to start a new chat because it can't handle anymore so those are my buy an rtx 6000 right now which by the way you should know grant remember the other day i told you it was 13 000 oh no don't tell me it's more like 17 now oh my by by the end of the week uh i I was looking at them last night for$16 ,900,$17 ,500,$15 ,000.
51:34It's getting to the point where buying a local machine that you can run is like buying a car. And it's going to help you with business. You can at least data center space for less than what it costs to buy an RTX 6 ,000. Yeah, but in the same way that you buy a car to go to work, so there's a benefit to the cost, right? You know, there is a benefit to buying a home data center or home server, let's say, like with a really nice graphic card. And it will help you with work, especially if you can use a lot of these local models and do tasks that you would need to do. Otherwise, spending money with OpenAI, you know, you could save some of that money.
52:13But dust off your checkbook, buddy, because it ain't cheap. But it's like buying a car. It's like$70 ,000,$15 ,000. Yeah. By the way, just a quick note, if you haven't yet, please take just a moment to like and subscribe. We really appreciate it. We love doing those. We love being able to interact with you all, answer questions as we go. And that helps us continue to do it. And also, just another quick shout out to A-Refs. Go check out their new tool. We got the link here in the chat already for you. We really appreciate them sponsoring today's video. All right. What do we have next, Grant? What do we have next?
52:54Let me pull down my screen. Sorry, I've got it. The next thing on the list is, so we're going to talk about some of these that came out, but we're just going to roughly overview them, because I don't think you've used either of these. I haven't got a chance to use these much yet, but they are Cursor Origin and DeepSeek Harness, which do you want to start with? We've been talking about harnesses, so we can talk about DeepSeek if you want. That's fine. That's fine. I have not touched either of these two tools yet, So I'm anxious to to hear a little more of your take. OK, so let me make sure I'm sharing the right thing here.
53:33So big news.
53:38You're cutting out pretty much, Grant, but it might be me. Oh, let me know. OK, I will. Which one of us is choppy to you all? If one of us is choppy, let us know and we'll know whose internet is messed up. It's me. I don't think I can share right now. I got really slow. I think you can share this link.
54:09Corey, you think you can share that link? Screen share that link? Yeah, the cursor. Oh, DeepSeek. I sure can. One minute here. Okay, it's Grant that's choppy. Good to know. Thank you, Rodwin. I appreciate that. It's hard to tell when it's us because if his screen's chopping up, it could be my internet. I never know.
54:34Okay, I'm getting the screen up here. And welcome to DeepSeek. By the way, the new DeepSeek's really good. I used it last night and was quite impressed. Yeah, what did you think? uh it was good it was reasonably quick uh i say that i threw a couple bigger things and it took its time a little but uh i was really happy with the answers uh honestly it was as good as anything else i touch it seems like like i really don't have any specific i know this is a problem yeah like i wasn't doing anything this is why we're not supposed to have this yeah yeah we're really we're really reaching the point we're like yep it's another good one yeah it's one gooder than the other one was
55:24yeah it's starting to get to the point where you know just what do you like better like what do you like using and that may be a field thing i expect it'll be a field yeah like models attraction was feel yeah like like i think a lot of what drew people to claude initially was was was the vibes. Well, anyway, I think I'm kind of cutting in and out, so I'll keep this brief. Basically, DeepSeek Harness, the thing that you need to know is if you're even remotely technical or interested in it, essentially what you could do is you could take the open repo, which is the code base, and you could use this to run a model like Gwen.
56:06You could customize it, or you could use it with DeepSeek, for example, and you can use this on your, you know, for whatever your coding purposes would be. It would be like replacing Codex or Cloud Code with this model. And this whole idea here that everything is a plugin is like the biggest idea from this. And it basically just makes it super easy to use, to connect this thing to anything else is the simplest way to explain that. Yeah, yeah, that everything is a plug-in thing is really cool. Models, tools, skills, sessions, sandboxes. Like, people are used to working in that way. In terms of, like, connecting things via a quick plug-in, via, you know, OAuth or something to hook something up.
56:59And I think that this is definitely the way those things will go. and like it may be that and maybe that answer is is like some kind of an mcp thing i mean i'm sure the plugins are mcp but maybe the answer is that that's that's how we connect all of these things in the future is just just a quick like you know plugging up your tesla at the charger it's just going back to everything is just code that talks everything is good yeah because the agents know how to do it you just say hey you know and that's how i recommend normal people use this like if you actually want to run something like this you just take the link that i just shared in the chat and you say you know set this up for me and you know take that to codex or cloud code or whatever quentin 27b might not be able to do it but you could use one of those other ones i just mentioned and just say like hey help me set this up so i can run my own version of this like help me clone it and, you know, try to pick it out for a spin.
58:01This is such a great explanation. The model is the soul of an agent. A harness lets an agent understand its environment, use tools, and keep working in the real world. I like that. The model is the soul. I think that's a really good way to put it. Also, just a heads up, this is in developer preview right now. so it's it's not like a final version um but it's it's getting it out there to get people using it so they can uh figure out ways to make it better there's a certain point when you're building where you've got to get people to use it so you can figure out where it's broke uh cory would you be able to share uh origin really fast as well we won't spend much time i sure could i sure cover this for normal people i sure can one minute uh let's see here I think I shared the link in the live chat.
58:57Yeah, I've got it here in the in our doc as well. Perfect, perfect.
59:07Okay. This is not it. Origin. Where is it? Where is it? This is like the... Okay, hang on. We're going to do this the old fashion way, guys. Brace yourself. Google. I never use Google anymore. I hear this. It's on their landing page, but they just auto-redirect you there. Yeah. There we go. GitForge for the agentic area. This is Cursor's GitHub killer is what they'd like it to be. From what I've read, it may or may not be, but go ahead. Go ahead. Yeah. Yeah. No, GitHub is basically like Google Drive for your code. It's a lot more powerful than that, but it's where you host your code on the cloud so that you can work with it with your agent.
1:00:01So, you know, traditionally you can write files on your computer. You can still do that, but then to save a copy of it and then continuously make edits and changes to that code and keep track of all of that and be able to reverse if you break something. you can do all of that on GitHub using this tool called Git. And so this is Cursor's version of that. So if you're familiar with Cursor, if you use Cursor, you can basically use them to now host your code instead of GitHub. And that's what manages like your software versioning and things like that. It's a good off-site place to store your code as well.
1:00:39It allows multiple people to work in and out of it at a time. You can serve it publicly. you can do all of the neat things. I don't have a lot more information on it other than that it came and it got people excited. Yeah, I would just say if you want to learn more about how Git works and GitHub, we're going to have a whole live stream about that next week, so tune in for that. Yeah, and one second. Here, we're going to bring up the one of these that I have been having just an obscene amount of fun with. And that.
1:01:18Is going to be this one buzz can you see the screen is it coming through yeah sorry i've got i've got my video blocked so i wanted to peek over and make sure i'm not on the wrong tab this is introducing buzz buzz has been out well close to a month now i'll be dang it doesn't seem like that long and uh this is a really really cool tool so you're all familiar with probably i assume slack teams google messenger or a variety of things like that, ways that people in the workplace keep in touch with one another and communicate back and forth. Well, Block, which is Jack Dorsey's company, has built one of those for agents, and it is sick.
1:02:04It's literally Slack for agents, so you can go in there and you can build agents. It talks here that it's a free, open-source collaboration platform where humans and agents work together in a shared workspace. You can connect your Claude code. You can connect your Codex, your DeepSea harness. You could connect anything you want. And if it doesn't offer it yet, mark my words, they're going to add it. I've been watching these guys build this in public for the last month. And they're shipping updates like day after day after day after day. As people come back with complaints and concerns and bugs and ideas, they ship them as fast as the ideas come in.
1:02:43And this is really, really neat. I'm looking for some screenshots here because I wanted to show you what it looks like. Actually, let's go over here to the GitHub screen because I think there are... When you get down here... I've got to scroll past all this stuff. Yeah, here's what it looks like. It looks like anything else you've ever used. Can you zoom in on that image a little bit? I'm going to see. Yeah, there we go. Yeah, there we go. There we go. And it's just like anything. Now, in this version, you know, you've got humans working in here, but your agents, which you build in here, operate as first class citizens the same way you do.
1:03:26So they can go and work in the channels you allow them to work in. They can communicate with one another. You can at them, as you'll see right here where it says at Fizz. Can you turn that into a clean three beat capture plan? And it's like, absolutely. And here it is. uh so there is there's just a ton of let me go back because there's more pictures you can it has everything a slack type tool has you can do channels you can do announcements you can do dms you can you have an inbox and you can go into the agents tool itself and uh and set them up like it'll have a few dummies in there you can use and i just picked i i i started with three connected to each of the three GBT 5.6 models.
1:04:13So I think they're named Honey, Fizz, and Buzz are my three. But I then created a fourth that runs on Soul, and its name is John the Delegator. Hat tip if you get that, if you've ever watched Sons of Anarchy or happen to know good old music, because that's an old song. And so he basically bosses them around I message him and I'm like Hey John, I want you to do this I want you to go build a landing page It needs copy, it's going to need all the things Plan it out and then get to work on it And then you see a message pop up right below you Like he's typing, just like in Slack And all of a sudden there's a message come out And he's tagging Honey And it's like, Honey, I need you to go make a plan Fizz, I need you to go do this this is very similar to using Telegram.
1:05:11More organized. It feels more organized. Yeah, so if you use Slack, you can organize into threads and channels. Telegram, the interface is just messages. It's not very neat and tidy. It's not as easy to basically compartmentalize different things that you're working on. Or catch the volume of spam that Telegram has a way of doing. That's my biggest issue with it. I get constant spam. I love it. Messages from random people. Yeah. But this, there's, as with that though, there's obviously a mobile version that's available for every operating system already. And free and open. And the way that you use connector agents is through the API thing we were talking about earlier.
1:05:57Yep. So you can connect to OpenAI, you can You can connect Open Router. You can connect to Quend via Unsloth. You can put it all in one place. And if you do it on your desktop, it's just a process. It's just a click-through process. It just takes a minute. Yeah. Excuse me. I think you could probably even have your Open Claw and Hermes agent chat with you in there. No? Here's the other thing. Check this out. Every message, reaction, workflow step, review approval, and get event is a signed event in one log. Same shape, same identity model, same audit trail, and whether the author is a person or a process.
1:06:33In practice, it feels like a team workspace. Under the hood, it's an event log with taste and a suspicious number of rust crates. This is so AI-generated. Yes, it's another AI-adjacent developer tool. We're sorry. The difference is what agents can actually do once they're inside. They can open repos and patches. I mean, that's all the stuff. I will say they beat me to this. I was trying to build this so now I don't have to now I don't have to it's alright, of course he needed more money he's not making money off of this it's open source no he's not he's giving back to the people but you can go and have conversations with your projects you can go have triage a bug without giving it the keys to the kingdom have their own keys oh yeah, agents have their own keys as well.
1:07:26So not just the model they're using. And that's a really big deal. Scoped by identity, not permission flags. So yeah, you can also give them varying levels of access, just like you can with a human user. Yeah. So let me give you a concrete example there. Let's say you're using GPT and Quinn together in the same workspace. Maybe there are certain things that you want Quinn to have access to and not chat GPT because it's like on your computer local, you're working on your trademark project, you don't want to give that information to OpenAI over the cloud. And then maybe there's some things that you don't want Quinn to access because Quinn's not smart enough and it might F it out.
1:08:08The crassness. And then maybe you want to keep Quinn off that project and you want to put OpenAI give OpenAI access to that. So there's lots of ways you could do the permissioning. I like that. That's something I missed the first time around with this. Yeah, it's really cool. and the truth is they made it fun to use is the truth of the matter like your agents use emojis like you'll come through and you'll notice that hang on, let me go back do you ever catch them talking to each other? yeah, you can watch them I mean without you unprompted I have not caught that but what you'll notice every time you tag one the first thing that happens is the two eyes emoji pops up which means I see you Maya and typing you a message is what the other one means.
1:08:56Like, I'm on it. And it's just, it's a cool layout. It's easy to use. It's a good approach. It's also not the only thing that Block has cooked up that I'm going to talk about today. You've got to explain this other thing to me because I still don't get it. Okay. Now, before I show you this, I want you to know that it's among the most traumatic homepages I've ever visited. And it's a lot to take in. But it's a really cool tool. And I'll show you first. That way you can get past the, oh my God, what is that? And then we'll talk about how cool it is and why it matters. So, here we are. Welcome, my dear friends, to Bird.
1:09:47this is another tool by Block that just dropped this is at the fun level what you're looking at are you're able to give your agents bodies you're able to give them little bodies you can use them as pets on your desktop you can do all kinds of things you can make trading cards with them this is cool yeah it's such a fun tool here's a number of little things they've used them for but we'll talk about what this is in a minute but like here you go prairie mountain man's mind is blown yeah yeah it's insane but it's absolutely you know what we don't have enough fun in these tools man this is this is really cool stuff we're alive at a fascinating time with some abilities that our grandparents would have never dreamed were possible and and dang it if that means i want to talk and dance and swiss army knife to do my legwork, then so be it.
1:10:45I like the typewriter. I'm a big fan of copycat. It's the future. It's the reincarnation of Clippy made manifest, but now if anything it could be Clippy. Yeah, honestly, they need a Clippy, if I'm going to be honest. Look at this. Pushback. It does feel a little bit Black Mirror. That's funny. It very much is. The eyes have a real Black Mirror vibe to them. Just a concept, maybe. never be alone again click me never be but can you can you show the interface though like like what what is actually what is it like to actually use this whoa uh yeah let's go back how did they make this this is so funny it's amazing isn't it it's it's just it's hilarious like you could absolutely do it without all of that but i mean if it's an option i would use it uh here's what you should know here is the interface this is it you got widgets and various things you can stick here what this is is like a home base for your agents your skills your automations your projects on your computer so just like with with the other one you connect your cloud code you connect your codecs you connect whatever the heck you want to connect here and and you can operate all of those agents in here but you can also give them a body now why this is cool is if if you lose a computer if if for some reason you just lose a subscription cancel subscription you have access to this where your agents live they don't live in chat gpt they don't live in claude they live in here and as do your skills.
1:12:25So here's another cool thing. It's like, you know, if you use all of these tools or some of these tools like Grant and I do, like I have skills that exist on three different computers that exist in Claude, that exist in Gemini, that exist in ChatGPT, and they're all some level of different because I've updated them periodically and maybe not gotten all of them. This is where your skills live and you can call them from any of your connected harnesses. so if you come in here and make a change that means when you go to claude code when you go to codex wherever you go it's using this version and that is huge it's still hard for me to wrap my mind around that and perhaps we can try and like take it even one step um like simpler for normal people like what when when would you act what's a real world world use case when you would actually use this okay i'll tell you what skills first yeah here's the deal okay let's say i use i'm going to say codex because that's kind of my main i use i have codex downloaded on three different computers those skills don't talk to each other i mean like if i load it on my mac it doesn't work on my pc i have to go and either download and send that skill over there this is going to be fixed soon.
1:13:45But I have to take it over there and do that, or I have to create another skill similar to it, and they're not matched. They don't match. They're not identical. They're going to perform a little bit different. If I want to make a change, it means I have to go to three different computers to make that change to ensure they're all right. So if I say, I want you to end every sentence with, I have to go do that on my Mac. I have to go do that on my PC. I have to go do that on the web interface but with this i can go into that specific scale right here and make that change and that change will be immediately reflected in all of those other apps it's it's skill versioning and think of this as the hub on a wheel and your agents as and your harnesses as spokes.
1:14:36If that makes sense. They all feed right back to here. This is the hub of your skills, the hub of your agents, the hub of your automations. And then you can from Claude call the same automation you could Claude call from over in Codex. Okay, wait. So with these little guys here, are they technically agents that work in all of these platforms or are they skills that work? Okay. Yeah, these are agents. when you make an a don't click here. Hang on, I've got to click there.
1:15:08Oh, wow.
1:15:13Turn your sound down. Oh, the internet's being stupid. It's like a rave, but the music's not loading. Oh my God. We keep getting a glimpse. Yeah, maybe switch off of this. Yeah, I think you're right. We're not prepared for this.
1:15:37Just hit the wrong button. Oh, well. But it's – did you hear the – it's talking still. Can you hear it? No, I'm not hearing that because you're not sharing your screen. It sounds like Teletubbies. Okay, so if you're a normal person, you probably don't want this tool. But if you're a freak – If you're a freak, that one's for you. It's for me. It's so for me.
1:16:11It's so fun. I love that they made that. Yeah. Yeah, that's kind of what I thought. I was like, I love that. You know, like everything else is like, oh, good. Another dark mode interface. Yay. It looks just like every other coding harness. And this looks like every other coding harness if you just threw some monsters from Digital Circus into it. Yeah. no for real and you know what we need more of this i agree i agree it can absolutely be work and be fun uh it does not have to be miserable as long as it does the job like uh who cares i love the idea of skill versioning yeah skill versioning excites me a lot just because of my specific situation i i just i use a number of different machines and and things and uh that's a thing that's been on my God, I need this list forever.
1:17:04What do we have left? Grant, anything? Or was that our last one? You know, we could talk briefly about GrokBot and all that stuff. But yeah, maybe just briefly touch on GrokBot. And if there's anything else you all are curious about, let us know. I remembered him and Eric's show great job very well. but yeah if there's anything you all are curious about or wondering about go ahead and let us know but while we're finishing up here we'll be another yeah 10-15 minutes probably so 10-15 minutes yeah yeah i guess like one other topic i think i would like to talk about and maybe just like save this for when we ramp is like there's this idea that i've heard a couple times recently that I want to discuss, which is consumers don't care about productivity.
1:17:55And all of the AI apps have been focused on productivity. And so far, there hasn't really been a huge breakout AI consumer app. And, you know, lots of people don't like AI because of the data centers and, you know, copyright and all sorts of stuff, like legitimate reasons to dislike it. And so I wonder to myself What would people want Out of these tools What would you actually want As a consumer But what's a normal thing What's a normal thing that your phone can't do for you So you want a robot I use no I use voice mode as like a Cooking helper regularly Things like you know For instance I make my wife a specific milk.
1:18:44Like it's a blend of half and half and 2 % that she uses for her espresso. And I'm always like, you know, it's doing my math. I'm like, hey, I'm making that milk again. And he's like, all right, man, it's a 60-40 blend. How much half and half do you have? And I'm like, I have five-eighths of a cup of half and half. And he's like, then you need three-eighths of a cup of 2%. That's it. But see, that's handled. That's handled by, you know, Chachi B.T. or Gemini. like what's something that they can't do or don't do that people don't do or can't do yeah like what's something like because i think like what what apple is about to do with siri is going to make your phone so much easier to use in theory that that's going to blow a lot of people's minds like you don't have to click into all these different apps to be able to use them after four weeks of having apple intelligence what you should know is i thought it was really cool for about four days okay and then you got over it and then i just went back to using chat gpt voice mode because i is there a consumer use case for this stuff outside of work i think entertainment i think entertainment like you know video image text that'd be good you ever read a book and no one wants to talk to you about it that you just you just watch something you're just Dying to talk to someone about nobody else will.
1:20:07That's a thing. Ooh, nutrition tracking is a good call because I'm kind of doing a thing like that right now. Right. With ChatGPT Health where I'm, you know, it's kind of helping me through making better dietary decisions. And, you know, kind of helping manage my weight, keeping an eye on my health, keeping tabs on things. And that's it's a really good application. Like there are a lot of applications that I think exist inside the tools that exist right now that people just don't think of repairing your car. Like I've literally diagnosed a sound in an automobile or the mechanic. Yeah, I think all of that makes sense to me as like chat.
1:20:56I think what I'm thinking about is more on like the user experience, like the interface side. and I think actually it's basically just abstracting away as much interface as possible. And I think I'm thinking of it as talking is why I'm thinking. I'm not thinking of sitting in a terminal. I'm thinking of I do all of this with voice mode just walking around. Yeah, I think that's right. I think basically like voice replaces you staring and pressing buttons and eventually you're just talking like, hey, can you do this for me? And then it actually does it and you don't have to go in and babysit it and check on it.
1:21:30I think that's probably that version of devices, which seems like what OpenAI is building towards. That is going to be the thing that most people like. That's very much the thing we've entered into over the first half of 2026. The era where that kind of stuff is doable. Like last night, I had ChatGPT complain and demand a refund from a video model company last night. who I guess I shouldn't put on blast until I see if they're going to refund me. You can check in your email right now. Yeah, literally, I'm pulling up my email. Like, let's see, make sure I haven't gotten it. Tough luck. Yeah, that's hilarious.
1:22:13No, I haven't. Okay. Yeah, I think... Go ahead. No, I think, like, yeah, being able to have it be a true personal assistant for you, I think that is probably a use case that the consumers could get behind. But maybe it's not framing it as productivity. It's framing it as like a helper, like a help buddy. Oh, I'm going to blast him. I just got the email. Hang on. Oh, okay.
1:22:43Cut that.
1:22:47Problem with lives. I'm talking about Hey Jen. I don't know if you're familiar with Hey Jen. But their tool has some behaviors that are absolutely not okay. Like, you can use plan mode, but after you plan a video, if you're not staring at it while it works for 10 minutes to prepare the plan, you've got less than one minute before it just assumes you approve and goes ahead and auto-generates a hunk of garbage and eats 10 % of your tokens every time it hits send. And this has happened. I've had it use the wrong avatar. I've had it do all kinds of things. and last night I had ChatGPT reach out and request a refund.
1:23:28It wrote a very nice but terse email explaining why, offering details, and I just got, hey there, thank you so much for reaching out. Unfortunately, we don't issue refunds on successfully generated content. We understand this may have been generated by mistake, but keep in mind that each generation consumes significant processing resources. As a courtesy, we're going to give you 100 credits of my 600 to support my work. So as far as I'm concerned, I'm going to cut that subscription and never, ever work with or recommend you again. See, this is the problem. Dark patterns, as Grant said yesterday.
1:24:08Yeah, you can't be a jerk in this world with these tools. Like, people will replace you. They will find a way to replace you. Absolutely. Absolutely. There are other tools that don't behave in ways that are shady. It's not about the refund. It's about the fact that it's set up to burn your credits whether you need them or not. I wrote about this in the newsletter a couple of weeks ago, but basically I think the world we're entering into is, if you're going to be successful in the software business, you have to ask yourself, do people like you? because that's the thing that's going to matter. Like when it comes to human customers, agent customers, they're just looking for the most straightforward way to get a task done.
1:24:56What's the easiest, what's the best recommended and they're going to just go do that. And a lot of software companies will just build only for agents. I think that's a lot of what's happening now. But for the consumer business, which still exists, you really have to ask yourself, do people like you? Because otherwise they are going to replace you either with their own custom solution or an open source equivalent that realize that a lot of people hate this company and they want to go and make their own version that isn't as dark in terms of dark patterns. Yeah. And I'm just, I'm done with that.
1:25:29Frankly, credits are a dark pattern in themselves. I went on a Twitter tirade about this at like midnight last night. The fact that credits are an untransparent way of, it gives them the ability to just change what a credit is without me ever knowing and change the value of it without knowing. And I very, very much support a company that's willing to charge me a flat rate and then let me pay API fees. I don't care if it costs more that way. I really don't. I understand that you have to keep your tool alive. I'm happy to spend money to ensure that that is the case, if it is a good tool. But I do absolutely, I'm just tired of credits.
1:26:22I think it's sneaky. Yeah, I agree. I was thinking, what makes credits different than paying per token? And I think it's because per token is very cut and dry. It's like everyone knows, it's like 1 million tokens. credits is much more opaque as you said but if you're charging for an image or a video well you know you can't really charge per token for that so no it's like you charge per generation and you just say x amount of generations is yeah yeah yeah something that gives us an idea of what it is uh but um just credits sounds like you're taking nickel tokens and selling them for 50 cents is what it feels like is happening.
1:27:09Well, a lot of times the companies that are offering this are like pass-through, you know, they call them layer companies or wrapper companies. Wrapper was the popular term where it's like someone else is actually creating the model and providing that. Someone else is providing the data center. You're just providing the interface and then you charge a fee on top of that. That's why for all the tools that I'm building for myself and others, I just do everything as bring your own API key because it's like you can decide what you want to pay. You know, one of the problems is like this is a great tool.
1:27:43The truth of the matter is the layout of it, the setup, it's a great tool. It has issues, though. It's not super consistent. Like its whole thing is you create your avatar, but it uses just random stuff that's not my avatar and calls it a successful generation there. And, yeah, I'm not paying for that. I'll go find a better tool I'll figure out how to make it work with Firefly and be way happy I'm also working on something I want you to test I'll share it with you today I appreciate that we'll maybe do another stream on that well everybody thank you do you have any hot takes on it I know neither of us have tried it yet because it's$200 because it's$300 Well, mine was asking for$300.
1:28:35Yeah. Yeah. And I was like, I want to try this, but I don't want to try it$300 bad. I've heard a lot of people like it. I've used the model. The model's really good. Yeah, I use it every day for research. Because with my Twitter Premium Plus thing, I get a decent amount of usage of Grok. And it's a good model. um but yeah firefly is shiny and should give you some serenity as well so i can stop raging against the machine i think he's talking about the the the sci-fi show and movie oh the movie yes and the show i need to watch that again it's been ages man yeah ages and ages and ages that's a good idea yeah grokbot would you say grokbot is most equivalent to Open Claw or Hermes or Perplexity Computer?
1:29:28Is it? Or does it do something that's even more new than that? Yeah. Yeah. It does? Yeah. It's apparently considerably better. Do you want to bring up the... Yeah, let's look at it. Let's read the notes. I like notes. Okay. Like and subscribe if you haven't yet, please. We appreciate you. It's been a lot of fun. I really appreciate when the chat has questions and thoughts and ideas. It makes these a lot more fun. So check this out. This is like a mix of computer. You're cutting out really, really bad.
1:30:19Yeah. Yeah. Hang on. I'll bring it up.
1:30:25Okay. Yeah, you can talk now. You sound okay, I think.
1:30:35All right. Let me bring up my screen.
1:30:42All right. GrockBot. AI teammates you can give real work to. Bots can sign in to your tools, use them just like you do, and come back with finished work. Very cool.
1:30:59Very cool. Oh, I like this. I like the setup here. Account manager, challenge account. You know, this is kind of like what my agents and scheduled tasks are doing, I guess. In workspace? Chattw2 workspace? Yeah. I mean, even just in codecs with scheduled tasks. a lot of a lot of like you know okay I've put this stuff in there like my whole my whole neuron workflow my email workflow is just scheduled tasks it's not like a workspace agent though I would argue it is an agent I think I think my main problem with this is the price point like both chat2bt and anthropic have$100 tiers and$20 tiers. And yes, of course, Grok has$10 tiers and all this sort of stuff, but this is not available at a price for people to test it.
1:32:02So you're either all in and giving it a shot or...
1:32:08I say that. Oh, I'm on the XAI page, not... Yeah. We need the Grok bot pricing page. We got this wrong in the newsletter the other day because it's really confusing. to try to find this information. Okay, let's just see that in here. Yeah. There it is. 200, 300, 200, 300, or... They have 120 for Workspace. Okay, so that's probably the best way to try it. Yeah. Yeah, that's for Teams, so you'd have to buy more than one. Yeah, but like we said, if you wanted to do Workspace Agents with ChatGBT, you just find a friend to go in on. Get a business account. Get like two or three people together on the same business account.
1:32:50and rock it that way. Hey, Paul, thanks. If you're still here, appreciate you being here and for chatting with us. Get Ultra. GrockBot's own computer signs into your tools, routines on a schedule. Yeah, you're like technically kind of supposed to put this on a computer, aren't you, by itself? Like its own machine? Yeah, I think there's a cloud one. There's a cloud one as well, I think. Cool. They host it for you. It looks neat. I'd love to try it if there was a, you know, a little less out of pocket because I've got so many AI subs already. Yeah. I'm sort of the same where I can't throw off my current, yeah, workflow.
1:33:32For additional hundreds of dollars. Yeah. Not worth it. It's cool, though. And X really stepped up. You know, the truth is, and the thing I credit it to is the same thing I credit OpenAI's success right now to, which is smart compute decisions early. I thought you were going to say... Not having Dario as a CEO? No, but both are valid. Yeah, yeah, no, he built a gigantic data center, and it's paid off for him. Yeah, yeah, and like two years ago. Like, it was very much the same time OpenAI was buying compute, so was Musk. And while that may have seemed crazy to anyone who's not staring at it through a microscope, here they are and uh and and sam and elon were right that demand would go insane and uh it did we're going to build the biggest data centers we've ever seen and it's still not going to be enough i do like their their idea to just put them in space where nobody has a problem with them i'm i'm fully down with that i see i have i have mixed emotions about space i i creep out a little over aliens being able to access your yeah maybe maybe what they are is nice little alien blockers we could attach some like i'm not even going to say those words or we'll get kicked off youtube uh but i don't even know what you're saying thankfully yeah but what i what i was gonna say is um i think like you know obviously people in memphis are really upset with the amount of natural gas plants that he's using, you know, generators that he's using to power this data center.
1:35:15And, you know, he's a solar, he's like the proponent of solar and batteries. And he knows that batteries are the way to actually protect these data centers from what's called micro fluctuations, which is during a training run, there's like this really bad spike and then it throws off the whole thing and it can be really costly. So you need to have extra load that you're able to, you know, basically power on exactly when you need it. So just need to switch all of those generators over to solar panels and battery and i think a lot more panels take time because it means real estate yeah i know there's always a trade-off with everything nuclear takes a lot of time because you know it's very dangerous and very expensive to build yeah yeah there's always a trade-off i i i frankly i'm i'm fully pro solar data centers i think solar data centers make a lot of sense i think there's lots of rural area where you can put massive ones there's one not far from me that is is it's five miles long and like one to two miles off the interstate in each direction and it's just miles and miles of like 12 foot square solar panels back to back to back that follow the sun all day long if the sun's up it just keeps following it to maximize what it can do and like it's it's amazing it's it's amazing it takes up you know a few square miles and I assume that one's going to a data center, but I haven't.
1:36:41I heard rumors it was either going to potentially be for Musk's Memphis data center, which could be very well true because it's 150 miles. So there's definitely some getting it wherever it's going to be that has to be done. But the other assumption was that it might have been Zuckerberg's. Oh. Have you read this story recently? It also has paid attention to compute, by the way. Yeah, that's true. But have you read the story how a lot of politicians are now basically putting the kibosh on data centers and not just liberal ones like Republican ones? All of them. Yeah. Yeah. Because everyone realizes that like outside of our bubble that we live in, people are not into this stuff.
1:37:28You know, and what's funny is the states who do take it, like Texas is going to make a lot of money. Texas will rebuild its entire electrical infrastructure that's been a god-awful disaster for most of our lives. The money is going to be ridiculous. but you know in my opinion if you have areas that are you know removed from people like I don't care it's in a data center I mean in a in an industrial park or something like for god's sakes factories are loud they're pumping all kinds of terrible stuff out into the air around you they're they're noisy there's a ton of traffic like to me it's it's more become it's less about the reality of data centers and more about this is how we rage at AI.
1:38:23Yeah. Well, I mean, I think this is their, their, their, here's how we fight back against it. It's a physical manifestation. I agree. But yeah, I also think that there's a lot of people who like chat GPT and like using it and don't want a data center in their backyard. That's fair too. This is based on NIMBY and Sons reporting. Everybody wants more prisons, but nobody wants it to be near them. This all comes full circle to my whole point where it's like, we need to stop thinking about scale. I think scale is the original sin of the AI industry. And it's why we're burning billions and billions of dollars on this.
1:39:00It's why we're building so many data centers so quickly, getting ahead of ourselves to some degree, but then also being behind because you see what scale brings. But like at a certain point, you need to scale down in order to scale up. Because think about how much more you can do if you fix architecture to be much more efficient, to be much more memory efficient, to be able to work. Like you get so much more capacity out of what you've already built. And not only that, if you can make it so that everyone's computer can run these things, then people will have access to it locally. But that's exactly what's happening already, I would say.
1:39:34Like now, an H100 runs these new, more efficient models that have continued to get better and better and better. An H100 runs GPT, I think this statement goes back to 5.4. I want to say it runs GPT 5.4 faster and better than it ran 3.5. I believe it. Which is getting more out of those. My problem is I don't think the idea that U.S. companies are going to just take a break and hand over the global lead in the AI races is a thing that is realistic. Like I get the idea. I don't disagree with it as a concept. I just can't see a world where it happens. I mean, I will say there's a lot of attention being given to efficiency in these as well.
1:40:39When I was at GTC, I went to Startup Open Mic. And there may have been another name for it, but that's what it was. It's like Startup Karaoke. Were they doing stand-up? It was more like a pitch. It was just like, you got three minutes, come make your pitch. and they just got in line and came up and made their pitch. And almost all of them were data center efficiency, data center cooling, data center distribution, all of these big problems around making data centers do what they do at a much broader scale without getting larger, doing more environmental harm, requiring additional electricity. Just like with the H-100s where it's all about getting more out of what you have, there's a lot of interest in that.
1:41:36A lot of interest. I think people would be more on board with data centers in the abstract if companies were saying, here's all of the powerful algorithms that we're running in order to make your lives better instead of replace you at your job. Yeah. And do the tasks that you like to do. Like we need to be using it, this stuff to be doing more things that we couldn't do before, not just figure out how to do all the things that we can currently do. Because that's not how we roll. Oh, sorry. Go ahead. I didn't mean to. No, no. I'm about to agree 100 percent of what you're saying. The problem with Americans is not Americans, humans, is that we have our ideas and they are what they are.
1:42:21And come hell or high water, that's what we're going to do. We're going to go and we're going to chase ideas that validate what we already think. And people just don't change their minds, not as a rule. Like, you know, it's not a situation where they're going to sit down with the AI companies. The AI companies are going to be like, oh, my God, this is terrible. Why would we have ever suggested this? Or where the AI companies are going to say, like, here's this. Where, in my opinion, what I would hope people would see is like the Moderna announcement yesterday and understand that what's happening right now is what's required for that to happen, not just in melanoma, but in lung cancer, in pancreatic cancer, in stomach cancer, in bone cancer.
1:43:03And there are nine other ones already active right now. And AI has played a big role in that. And I think the world needs to see more of the promise and less of the flock cameras and less of the, you know, ways websites are taking your data and ways that, you know, companies are looking to manipulate you. manipulate you. I think when that's all they see and that's what they think, that's what they assume all this is, and that's what a data center represents. And that's fair. I mean, like, I don't like that stuff either. But at the same time, things like that Moderna announcement give me a lot of hope for the future of our species.
1:43:53Yeah, I think both perspectives are valid. I think like there is unquestionably going to be amazing things that we can do with math and with computation that will unlock so much more possibility and benefits for us than we could have ever imagined but at the same time you know that we got to solve these problems and that's why we talk about both you know we talk about both the upside and the downside yeah yeah and and and both are are legitimate concerns i don't i don't at all mean to shoot down the concerns of the against data centers. And frankly, it's their yard. It's their right to decide I don't want this here.
1:44:30And they don't have to tell me why. They can just say no. But I do feel like at the same time, I see the value in the data center itself. And I believe that I can see what's on the other end of this. Maybe not the other end, but what's in the future ahead. And I think we find somewhere where it is okay. I think... space. Hostile takeover of North Dakota. You know, I don't know. Just kidding. Yeah, what if, well, no, funny theory, what if we just decided like one state, we're like, okay, we're just going to turn this into, you know, this is where they all are. Everyone here gets a brand new home in some other country, other state.
1:45:15You pick where you want to live. But we're going to need, you know, I don't know. I don't know. I do think, though that a lot of what will help with this is decentralized servers so like okay we're trying to build like massive data centers like on the outskirts of towns they're like you know miles long buildings like yes okay that's good for scale but you know that's good for the models of today we don't know what the models of tomorrow are going to need um why not have everyone just has like like they should have a battery on their house solar panels on their house have a data server on their house and then we distribute that as a grid system and we sell the energy yeah every individual gets their cut my personal belief system is that we should be as self-sufficient as possible at every node of the system so that you know if something goes down then you know something else can cover it and you know if you have these big centralized points of failure like sure it helps you for serving lots of customers but then it also So like, you know, the war could break out and somebody bombs the data center.
1:46:21And then it's like, oh, there goes, you know, half of the country's intelligence overnight like that. Well, you know, there have been companies that wanted to do that with blockchain. Like, I mean, I remember dealing with a gaming company a while back where that's what they were doing was like when you go to bed, you've got this little app on your thing. in and like companies doing 3d rendering for movies and stuff we're able to tap into additional compute resources from your laptop from your computer and bring all of that compute together into one big bundle that could help them generate movies faster in terms of like you know rendering 3d and cgi films uh how'd that go for them i mean it was really interesting it absolutely worked but I don't know that anybody cares enough to try it.
1:47:06I'll have to find that. I interviewed, I did, I think, a video interview with him. I'll find that and share it with you. It was, that's cool. It's from when I was with a different magazine, but it was, it was interesting. And I think kind of supports the idea that you're talking about on maybe a more smaller scale in terms of like not solar panels on homes, but, but at least that. Yeah, I would love the idea to be like, hi, I want solar panels on every inch of real estate I have. I want a patio. And I want a hose that goes two ways. One way takes energy. The other one rains dollar bills into my kitchen.
1:47:45And, you know, honestly. I just think if, yeah, I think if people, if you want people to feel involved in this whole thing, right, you have to involve them. You can't just say we're going to build gigantic things out in the middle of nowhere that's going to replace you. that doesn't make me feel very involved in the process how about you say hey we're gonna we're we have this new system you know it's decentralized everyone's gonna get their own server their own data set their own mini data center in their home you know and we'll all be more self-sufficient we'll all have more intelligence like we will be a better civilization if we do this we'll be more who pays for the data centers just curious what's your what's your thought there is this a thing that like who pays for it yeah companies just give you these to run and pump the energy off of for them or is this a people will go buy their own in-home data centers there is a lot of different ways you can slice up the business model isn't there a company that is doing this i think there's a company that's doing this it's sort of yeah there's a company that will uh stick i think it's a gb 200 nvidia gb gb 200 i believe on the side of your house yeah uh but i don't know the specifics of how it works we've reached out to them about a podcast but i don't think we ever heard back uh we're gonna do that and do something about data centers in general soon because i think i think i would like to pick the brains of someone who knows data centers and uh because like like there's a lot i don't know and would like to better understand and uh i'm curious to just see how they work so we're going to try and get i don't know an electrical engineer or somebody who knows data centers i feel like uh i feel like this is kind of classic you and me where the longer we go on a stream when we're not teaching people something we just eventually start talking about data centers just eventually land at data centers data centers is the hub where all things lead all right guys it's been fun i i really appreciate everyone who's here thank you for staying so long and listening.
1:49:46Thank you to AREFs again for sponsoring today's episode. We really appreciate that and glad to have you. On that note, we'll see you back here next week when we have a cool one scheduled. Is it scheduled already? Oh yeah, it's already up. I'll link to it one more time. So this is a beginner's guide to GitHub next Thursday, same time. We have hijacked a human from GitHub to come and tell us about GitHub and help us understand GitHub and teach us how we can use it better. It's a thing I think that a lot of people, when they dabble in Vibe coding, are, like, really intimidated by. And I'm saying that based on the fact that I was really intimidated by it.
1:50:28Oh, sorry. The GB200, it is... Grace Blackwell. It is, yes. It is the earlier-gen NVIDIA Grace Blackwell.
1:50:43server rack gpu server rack thing that that runs in data centers all over the place um and i it's it's the current is the gb300 and it it just started going out like december or something but the gb200 is still a monster and uh i think that's what this company's putting on houses we'll find out who they are report back but yeah thank you robert robot cat always here always a real Rodwin, thank you both. We really appreciate it. We'll see you back here next time. Farewell for you. Let us know next week or in the comments if you use any of these tools and what you think. Yeah, I'd love to hear your thoughts.
1:51:25You guys get bobe. All right. I'm going to call it a day. See you all.
From the publisher
AI tools are shipping faster than most normal humans can figure out what half of them actually do.
So Thursday, August 20 at 10 AM PT / 1 PM ET, we’re going LIVE to translate this week’s biggest AI launches into plain English. 😸
The goal: understand what these tools actually are, who they’re for, what’s useful vs. hype, and which ones are worth trying.
No three developers yelling model benchmarks at each other for an hour. Instead, we’re covering:
🤖 Qwen 3.8: What a powerful open model is, why you might use one instead of ChatGPT or Claude, and when that makes sense.
💻 Unsloth Studio: Run and experiment with AI models on your own computer, even if you’ve never touched a terminal.
⌨️ Cursor Origin: Cursor wants to host your code too. Here’s what that could mean for people building websites, apps, and internal tools with AI.
🛠️ DeepSeek Harness: What an “agent harness” is, why everyone keeps talking about them, and whether they matter outside hardcore coding circles.
💬 Buzz: Imagine Slack or Discord, except AI agents can join the workspace, collaborate with humans, and do work.
🐦 Berd: One desktop home for your AI agents, projects, skills, tools, and models. You can also give your agents little animated bodies. Please emotionally prepare yourself.
Plus, we’ll cover the other notable models and tools that dropped this week.
And if OpenAI drops Astra before we go live? We’ll cover that too. If it’s incredible, great. If it belongs in the “cool, another model” bucket, that’s part of the roundup too.
📅 August 20 @ 10AM PT / 1PM ET
Bring your questions. No PhD required. 😸
TOOLS / LINKS:
Qwen 3.8: https://qwen.ai/blog?id=qwen3.8
Unsloth: https://unsloth.ai/
Cursor: https://cursor.com/changelog/origin-code-hosting
DeepSeek: https://github.com/deepseek-ai/deepseek-harness
Buzz: https://buzz.xyz/
Berd: https://berd.xyz/
OpenAI: https://openai.com/index/pacing-model-development-cyber-capabilities/
