In short
Podcast Episode Summary: The AI Daily Brief - "What AI Developers Are Building Next"
Podcast Overview Podcast Title: The AI Daily Brief (Formerly The AI Breakdown) Episode Title: What AI Developers Are Building Next, with Swyx and Alessio of Latent Space Description: A detailed analysis of current trends in AI development, featuring insights from Alessio and Swyx from Latent Space.
Key Themes and Discussions
- AI Development Trends
- Interest in Developer Perspectives: NLW emphasizes the importance of understanding what developers are actively working on as a way to gauge the future of AI.
- The Current Landscape: There is a noticeable shift from broad, generic AI projects to more specific, verticalized applications. Developers are increasingly focused on solving particular use cases rather than aiming for general AGI.
- AI Agents
- Shift from General to Vertical Applications:
- Developers are now creating specialized AI agents designed for specific domains (e.g., finance, legal, compliance).
- The initial excitement around broad AI agents (like AutoGPT) has plateaued, leading to a more pragmatic approach focusing on narrow tasks.
- Challenges:
- Developing fully functional general-purpose AI agents remains challenging; many still require significant engineering and are not ready for widespread practical use.
- Hardware and Software Integration
- Emerging Chip Technologies:
- Discussions around advancements in chip technology and new architectures are critical as companies like Grok and others challenge traditional giants like NVIDIA.
- The podcast highlights the importance of speed and efficiency in AI computations, suggesting that improved hardware will enable new applications and use cases.
- Wearables and AI Integration
- Interest in Wearables:
- There is growing excitement around wearable technology that integrates AI, but societal readiness for constant data collection and privacy concerns are barriers to widespread adoption.
- The economics of hardware vs. software subscriptions present interesting challenges for developers.
- Future Predictions
- Vertical Agent Utilization:
- The trend towards integrating AI into workplaces is expected to accelerate, with companies like Klarna leading the way by replacing traditional roles with AI agents in customer support.
- CEOs and Corporate Leadership:
- Speculation around potential leadership changes at major tech companies, particularly Google, was raised, indicating possible shifts in corporate strategy and direction.
Key Takeaways
- Focus on Narrow AI: Developers are increasingly prioritizing vertical applications of AI rather than broad, general-purpose solutions.
- Hardware's Role: The advancement of chip technologies will play a crucial role in enhancing AI capabilities and expanding its applications.
- Wearables Market: The integration of AI into wearable technologies is promising but faces societal and privacy challenges.
- AI in the Workplace: The trend towards AI replacing human roles in specific sectors is beginning to manifest, with the potential for both positive and negative outcomes.
Conclusion The episode provides valuable insights into the current trajectory of AI development, emphasizing the need for specificity in AI applications and the interplay between hardware advancements and software capabilities. The discussions highlight a maturing landscape where practical use cases are beginning to dominate over theoretical ambitions. This shift suggests a future where AI is integrated more deeply into various industries, fundamentally changing the workforce and operational paradigms.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the AI Breakdown, part 2 of my conversation with Alessio and Swix from Layden Space. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to Breakdown.network for more information about our Discord, our YouTube, and our newsletter.
0:24Hello friends, back again with part two. If you haven't heard part one of this conversation, I suggest you go check it out. But to be honest, they are kind of actually separable. In this conversation, we get into a topic that I think Alessio and Swix are very well-positioned to discuss, which is what developers care about right now, what people are trying to build around. Hello, friends. Quick note before we get to the rest of the episode. You have probably heard me talk about the AI education beta over the past few months. We've had a ton of you participate, which has been amazing. And now we're almost ready to announce something big and something new.
0:56If you want to be one of the first to hear about our new approach to learning AI that is hyper-practical, hands-on, immediately relevant, continuously upgrading, and anchored by community, go to besuper.ai and sign up to be notified when the project goes live. We're getting there in just a few weeks, and I want all of you along for the journey. Once again, that's besuper.ai. I honestly think that one of the best ways to see the future in an industry like AI is to try to dig deep on what developers and entrepreneurs are attracted to build, even if it hasn't made it to the news pages yet. So consider this your preview of six months from now, and let's dive in.
1:36Let's bring it to the GPT-5 conversation. I mean, so I think that that's a great sort of assessment of just how the stakes have been raised. You know, what is your... I mean, so I guess maybe I'll frame this list as a question, just sort of something that I've been watching. Right now, the only thing that makes sense to me with how fundamentally unbothered and unstressed OpenAI seems about everything is that they're sitting on something that does meet all that criteria, right? Because, I mean, even in the Lex Friedman interview that Altman recently did, you know, he's talking about other things coming out first.
2:13He's talking about, he's just like, listen, he's good and he could play nonchalant, you know, if he wanted to. So I don't want to read too much into it. But, you know, they've had so long to work on this. Like, unless that we are like really meaningfully running up against some constraint, it just feels like, you know, there's going to be some massive increase. But I don't know. What do you guys think? hard to speculate. You know, at this point, they're pretty good at PR and they're not going to tell you anything that they don't want to. And they can tell you one thing and change their minds the next day.
2:44So it's really, you know, I've always said that model version numbers are just marketing exercises. Like they have something and it's always improving. And at some point you just cut it and decide to call it GPT-5. And it's more just about defining an arbitrary level at which they're ready. And it's up to them what ready means. We definitely did see some leaks on GPT 4.5, as I think a lot of people reported, and I'm not sure if you covered it. So it seems like there might be an intermediate release. But I did feel coming out of the Lex Freeman interview that GPT 5 was nowhere near. And, you know, it was kind of a sharp contrast to Sam talking at Davos in February, saying that, you know, it was his top priority.
3:27so I find it hard to square and honestly like there's also no point reading too much tea leaves into what any one person says about something that hasn't happened yet or a decision that hasn't been taken yet so yeah that's my two cents about it like calm down let's just build yeah the February rumor was that they were going to work on AI agents so I don't know maybe they're like whatever yeah they had two agent I think two agent projects right one desktop agent and one sort of more general GPT's agent. And then Andre left. So he was supposed to be the guy on that. What did Andre see? What did he see?
4:03I don't know. What did he see? I don't know. But again, it's just like the rumors are always floating around. I think this is we're not going to get to the end of the year without GPT 4.5 or 5. That's definitely happening. I think the biggest question is are Anthropic and Google increasing the pace. Is the Cloud 4 coming out in 12 months? Like 9 months? What's the deal? Same with Gemini. They went from 1 to 1.5 in like 5 days or something. So when's Gemini 2 coming out? Is that going to be soon? I don't know. There are a lot of speculations, but the good thing is that now you can see a world in which OpenAI doesn't rule everything.
4:49And, you know, so that's the best news that everybody got, I would say. Yeah. And Mr. A. Large also dropped in the last month. And, you know, not quite GPT-4 class, but very good from a new startup. So, yeah, we have now slowly changed in landscape. You know, in my January recap, I was complaining that nothing's changed the landscape for a long time. But now we do exist in a world, sort of a multipolar world, where Claude and Gemini are legitimate challengers to GPC4 and hopefully more will emerge as well, hopefully from meta. Yeah. So let's actually talk about sort of the open source side of this for a minute.
5:27So Mr. Large, notable because it's not available open source in the same way that other things are. although I think my perception is the community has largely given them like the community largely recognizes that they want them to keep building open source stuff and they have to find some way to fund themselves that they're going to do that and so they kind of understand that there's like they got to figure out how to eat but we've got so you know there's Mistral there's I guess Grok now which is you know Grok one is from October is open sourced yeah sorry I thought you meant Grok the chip company no you mean Twitter Grok although Grok the chip company I think is even more interesting in some ways but But and then there's the, you know, obviously Lama 3 is the one that sort of everyone's wondering about, too.
6:06And, you know, my sense of that, the little bit that Zuckerberg was talking about Lama 3 earlier this year suggested that at least from an ambition standpoint, he was not thinking about how do I make sure that, you know, meta, you know, keeps keeps the open source thrown, you know, vis-a-vis Mistral. He was thinking about how you go after, you know, how he, you know, releases a thing that's, every bit as good as whatever OpenAI is on at that point. Yeah, from what I heard in the hallways at GDC, Lama 3, the biggest model will be 260 to 300 billion parameters. So that's quite large. That's not an open source model.
6:43You cannot give people a 300 billion parameters model and ask them to run it. It's very compute intensive. It can be open source. It's just it's going to be difficult to run, but that's a separate question of whether it's open source. It's more like as you think about what they're doing it for, You know, it's not like empowering the person running Lama on their laptop. It's like, oh, you can actually now use this to go after OpenAI, to go after Anthropic, to go after some of these companies at like the middle complexity level, so to speak. Yeah. So obviously, you know, we assume Chantala on the podcast.
7:17They're doing a lot here. They're making PyDorch better. You know, they want to that. That's kind of like maybe a little bit of a shot at NVIDIA in a way, trying to get some of the CUDA dominance out of it. Yeah, no, it's great. I love the duck destroying a lot of monopolies arc. You know, it's been very entertaining. Let's bridge into the sort of big tech side of this, because this is obviously like, so I think actually when I did my episode, this was one of the, I added this as an additional war. That's something that I'm paying attention to. So we've got Microsoft's moves with inflection, which I think potentially are being read as a shift vis-a-vis their relationship with open AI.
7:58which also the sort of mischievous large relationship seems to reinforce as well. We have Apple potentially entering the race finally, you know, giving up Project Titan and kind of trying to spend more effort on this. Although, counterpoint, we also have them talking about it or there being reports of a deal with Google, which, you know, is interesting to sort of see what their strategy there is. And then, you know, Meta has been largely quiet. We kind of just talked about the main piece, but, you know, and then there's spoilers like Elon. I mean, you know, what of those things has sort of been most interesting to you guys as you think about what's going to shake out for the rest of this year?
8:35Let's take a crack. So the reason we don't have a fifth war for the big tech wars is that's one of those things where I just feel like we don't cover differently from other media channels, I guess. Sure. In our anti-interest list, we actually say like we try not to cover the big tech Game of Thrones or it's proxied through, you know, all the other four wars anyway. So there's just a lot of overlap. Yeah, I think absolutely, personally, the most interesting one is Apple entering the race. They actually released, they announced their first large language model that they trained themselves. It's like a 30 billion multimodal model.
9:10People weren't that impressed, but it was like the first time that Apple has kind of showcased that, yeah, we're training large models in-house as well. Of course, like they might be doing this deal with Google. I don't know. It sounds very sort of rumor-y to me. and it's probably, if it's on device, it's going to be a smaller model. So something like a Gemma, it's going to be smarter autocomplete. I don't know what to say. I'm still here dealing with Siri, which probably hasn't been updated since God knows when it was introduced. It's horrible. It makes me so angry. So one, as an Apple customer and user, I'm just hoping for better AI on Apple itself.
9:51but two they are the gold standard when it comes to local devices personal compute and and trust like you you trust them with your data and i think that's what a lot of people are looking for in ai that they have they love the benefits of ai they don't love the downsides which is that you have to send all your data to some cloud somewhere and some of this data that we're going to feed ai is the most personal data there is so apple being like one of the most trusted personal data companies I think it's very important that they enter the AI race. And I hope to see more out of them. To me, the biggest question with the Google deal is like, who is paying who?
10:29Because for the browsers, Google pays Apple like 18, 20 billion every year to be the default browser. Is Google going to pay you to have Gemini or is Apple paying Google to have Gemini? I think that's like what I'm most interested to figure out. Because with the browsers, it's like it's the entry point to the thing. So it's really valuable. to be the default. That's what Google pays. But I wonder if the perception in AI is going to be like, hey, you just have to have a good local model on my phone to be worth me purchasing your device. And that's going to drive Apple to be the one buying the model.
11:04But then, like Sean said, they're doing the MM1 themselves. So are they saying, we do models, but they're not as good as the Google ones? I don't know. The whole thing is really confusing, but it makes for great meme material on Twitter. Yeah, I mean, I think like they are possibly more than OpenAI and Microsoft and Amazon. They are the most full stack company there is in computing. And so like they own the chips, man. Like they manufacture everything. So if there was a company that could seriously challenge the other AI players, it would be Apple. And I don't think it's as hard as self-driving.
11:45so like maybe they've they've just been investing in the wrong thing this whole time we'll see wall street certainly think so wall street loved that move man there's there's a big a big sigh of relief um well let's let's move away from from sort of the big stuff i mean i think to both of your points it's gonna can i can i can i drop one factoid uh about this this wall street thing i went and look at when um meta went from being a vr company to an ai company and I think the stock I'm trying to look up the details now the stock has gone up 187 % since Lama 1 which is$830 billion in market value created in the past year yeah if you haven't seen the chart it's actually remarkable if you draw a little arrow on it it's like no we're an AI company now forget the VR thing
12:41it's uh it is an interesting no it's i i think um unless you called it sort of like zucks disruptor arc or whatever he he really does he is in the midst of a of a total you know i don't know if it's a redemption arc or it's just it's something different where you know he he's sort of the spoiler like people loved him just freestyle talking about why he thought they had a better headset than apple like even if they didn't agree they just loved he was going direct to camera and talking about it for you know five minutes or whatever uh so that's a fascinating shift that i don't think anyone had on their bingo card you know whatever two years ago yeah it's still didn't see him fight elon though so yeah i mean hey don't don't don't write it off you know maybe just these things take a while to happen but uh we need to see him fight in the coliseum uh no i think uh you know in terms of like self management life leadership i think he has there's a lot of lessons to learn from him.
13:34You might kind of quibble with the social impact of Facebook, but just himself in terms of personal growth and perseverance through a lot of change and everyone throwing stuff his way. I think there's a lot to say about to learn from Zuck, which is crazy because he's my age. Yeah, right. Awesome. So one of the big things that I think you guys have, you know, distinct and unique insight into being where you are and what you work on is, you know, what developers are getting really excited about right now. And by that, I mean, on the one hand, certainly, you know, like startups who are actually kind of formalized and formed as startups, but also, you know, just in terms of like what people are spending their nights and weekends on, what they're, you know, coming to hackathons to do.
14:20And, you know, I think it's such a fascinating indicator for where things are headed. Like, if you zoom back a year right now was right when everyone was getting so so excited about ai agent stuff right auto gpt and baby agi and these things were like if you dropped anything on youtube about those like instantly tens of thousands of views i know because i had like a 50 000 view video like the second day that i was doing the show on youtube you know because i was talking about auto gpt uh and so anyways you know obviously that's sort of not totally come to fruition yet but what are some of the trends and what you guys are seeing in terms of people's people's interest and and what people are building um i can start maybe with the agents part and then i know sean is doing a diffusion meetup tonight there's a lot of a lot of different things the the agent wave has been the most interesting kind of like um dream to reality uh arc so out of gbt i think they went from zero to like 125 000 get up stars in six weeks and then one year later they have 150 000 stars uh so there's kind of been a big plateau i mean you might say there are just not that many people that can start it you know everybody already started uh but the promise of hey i'll just give you a goal and you do it i think it's like amazing to um get people's imagination going you know they're like oh wow this is this is awesome everybody everybody can try this to do anything uh but then as technologists you're like well that's that's just like not possible you know we would have like solved everything and i think it takes a little bit to go from uh the promise and the hope uh that people show you to then try and yourself and going back to say okay this is not really working for me and um david one from adapt you know they in our episode he specifically said we don't want to do a bottom-up product you know we don't want something that everybody can just use and try because it's really hard to get it to be reliable.
16:20So we're seeing a lot of companies doing vertical agents that are narrow for a specific domain, and they're very good at something. You know, Mike Conover, who was at Databricks before, is also a friend of Latent Space. He's doing this new company called Brightwave, doing AI agents for financial research, and that's it, you know, and they're doing very well. There are other companies doing it in security, doing it in compliance, doing it in legal, all of these things that like people, nobody just wakes up and say, oh, I cannot wait to go on AutoGPT and ask it to do a compliance review of my thing, you know, just not what inspires people.
17:01So I think the gap on the developer side has been the more bottom sub hacker mentality is trying to build these like very generic agents that can do a lot of open-ended tasks. And then the more business side of things is like, hey, if I want to raise my next round, I cannot just like sit around and mess around with like super generic stuff. I need to find a use case that really works. And I think that has worked for a lot of folks. In parallel, you have a lot of companies doing evals. There are dozens of them that just want to help you measure how good your models are doing. again if you build evals you need to also have a restrained surface area to actually figure out whether or not it's good right because you cannot eval anything on everything under the sun so that's another category where i've seen from the startup pitches that i've seen there's a lot of interest in in the enterprise it's just like really fragmented because the production use cases are just coming like now you know there are not a lot of long established ones to to test against and so that's kind of on the virtual agents and then the robotic side it's probably been the thing that surprised me the most at nvidia gtc the amount of robots that were there there were just like robots everywhere like both in the keynote and then on the show floor you would have boston dynamics uh dogs running around there was like this like uh fox robot that had like a virtual face that like talk to you and like move in real time there were um industrial robots um nvidia did a big push on their own omniverse thing which is like this digital twin of whatever environments you're in that you can use to train the robots agents um so that kind of takes people back to the reinforcement learning days but um yeah agents people want them you know people want them i give a talk about the the rise of the full stack employees and um kind of this future the same way for stack engineers kind of work across the stack in the future every employee is going to interact with every part of the organization through agents and ai enabled tooling um and this is happening it just needs to be a lot more narrow than maybe the first approach that we took which is just put a string in auto gpt and pray but yeah there there's a lot of super interesting stuff going on uh yeah uh well he uh let's have covered a lot of stuff there i'll separate the robotics piece because i feel like that's so different from the software world but yeah we do talk to a lot of engineers and that this is our sort of bread and butter.
19:29And I do agree that vertical agents have worked out a lot better than the horizontal ones. I think the point I'll make here is just the reason AutoGBT and maybe AGI, it's in the name. We're promising AGI. But I think people are discovering that you cannot engineer your way to AGI. It has to be done at the model level and all these engineering, prompt engineering hacks on top of it weren't really going to get us there in a meaningful way without much further improvements in the models. I would say, I'll go so far as to say even Devin, which is, I think the most advanced agents that we've ever seen still requires a lot of engineering and still probably falls apart a lot in terms of practical usage, or it's just way too slow and expensive for what its problem is compared to the video.
20:18So yeah, that's what happened with agents from last year. but I do see like vertical agents being very popular and sometimes I think the word agent might even be overused sometimes. Like people don't really care whether or not you call it an AI agent, right? Like does it replace boring menial tasks that I do that I might hire a human to do or that the human who is hired to do it like actually doesn't really want to do. And I think there's absolutely ways in sort of a vertical context that you can actually go after very routine tasks that can be scaled out to a lot of AI assistants. So yeah, I would basically plus one what I still said there.
20:57I think it's very, very promising, and I think more people should work on it, not less. There's not enough people. This should be the main thrust of the AI engineer is to look for use cases and go to production with them instead of just always working on some HEI promising thing that never arrives. I can only add that. So I've been fiercely making tutorials behind the scenes around basically everything you can imagine with AI. We've probably done, we've done about 300 tutorials over the last couple of months. And the verticalized anything, right? Like this is a solution for your particular job or role.
21:33Even if it's way less interesting or kind of sexy, it's so radically more useful to people in terms of intersecting with how, So those are the ways that people are actually adopting AI in a lot of cases. It's just a thing that I do over and over again. By the way, I think that's the same way that even the generalized models are getting adopted. It's like I use MidJourney for lots of stuff, but the main thing I use it for is YouTube thumbnails every day. Like day in, day out, I will always do a YouTube thumbnail or two with MidJourney, right? And you can start to extrapolate that across a lot of things.
22:06And all of a sudden, you know, AI doesn't, it looks revolutionary because of a million small changes rather than one sort of big dramatic change. And I think that the verticalization of agents is sort of a great example of how that's going to play out too. Yeah. So I'll have one caveat here, which is I think that because multimodal models are now commonplace, like Claw, Gemini, OpenAI, all very, very easily multimodal, Apple's easily multimodal, all this stuff. There is a switch for agents for sort of general desktop browsing that I think people need to keep an eye on. It's not mature yet, but it is absolutely coming on the way.
22:45And so just as we're starting to talk about this verticalization piece, because that is mature, that is ready for people to work on, that a lot of people are making really good money doing that. The thing that's on the rise is this sort of drive-by vision version of the agent where they're not specifically taking in text or anything. They're just watching your screen just like someone else would and piloting it by vision. And in the episode with David that we'll have dropped by the time that this airs. I think that is the promise of Adept. That is the promise of what a lot of these sort of desktop agents are.
23:20And that is the more general purpose system that could be as big as the browser, the operating system. Like people really want to build that foundational piece of software in AI. And I would see like the potential there for desktop agents being that, that you can have sort of self-driving computers. You know, don't write the horizontal piece out. I just think we took a while to get there. What else are you guys seeing that's interesting to you? I'm looking at your notes and seeing a ton of categories. Yeah. So I'll take the next two as one category, which is basically alternative architectures, right?
23:55The two main things that everyone following AI kind of knows now is one, the diffusion architecture, and two, let's just say the decoder-only transformer architecture that is popularized by GPT. You can look on YouTube for thousands and thousands of tutorials on each of those things. What we are talking about here is what's next, what people are researching and what could be on the horizon that takes the place of those other two things. So first of all, we'll talk about transformer architectures and then diffusion. So Transformers, the two leading candidates are effectively RWKV and the state space models, the most recent one of which is Mamba, but there's others like the Striped Hyena and the S4H3 stuff coming out of Hazy Research at Stanford.
24:34And all of those are non-quadratic language models that promise to scale a lot better than the traditional transformer. This might be too theoretical for most people right now, but it's going to come out in weird ways where imagine if right now the talk of the town is that Claude and Gemini have a million tokens of context and like, whoa, you can put in two hours of video now. Okay, but what if you could throw in 200 ,000 hours of video? How does that change your usage of AI? What if you could throw in the entire genetic sequence of a human and synthesize new drugs? How does that change things?
25:19We don't know because we haven't had access to this capability being so cheap before. And that's the ultimate promise of these two models. They're not there yet, but we're seeing very, very good progress. RWKV and Mamba are probably the two leading examples, both of which are open source, that you can try them today. and have a lot of progress there. The main thing I'll highlight for RU-UKV is that at the 7B level, they seem to have beat Llama 2 in all benchmarks that matter at the same size for the same amount of training as an open source model. So that's exciting. But they're at 7B now. They're not at 7TB.
25:59We don't know if it'll scale. And then the other thing is diffusion. Diffusions and transformers are kind of on the collision course. The original stable diffusion already used transformers in parts of its architecture. It seems that transformers are eating more and more of those layers, particularly the VAE layer. So the diffusion transformer is what Sora is built on. The guy who wrote the diffusion transformer paper, Bill Pebbles, is the lead tech guy on Sora. So you'll just see a lot more diffusion transformer stuff going on. But there's more sort of experimentation with diffusion. I'm holding a meetup actually here in San Francisco that's going to be like the state of diffusion, which I'm pretty excited about.
26:40Stability is doing a lot of good work. And if you look at the architecture of how they're creating stable diffusion three, hourglass diffusion, and link consistency models, or SDXL Turbo, all of these are like very, very interesting innovations on like the original idea of what stable diffusion was. So if you think that it is expensive to create or slow to create stable diffusion or an AI-generated art, you are not up to date with the latest models. If you think it is hard to create text and images, you are not up to date with the latest models. And people still are kind of far behind. The last piece of which is the wildcard I always kind of hold out, which is text diffusion.
27:18So instead of using autogenerative or autoregressive transformers, can you use text to diffuse? So you can use diffusion models to diffuse and create entire chunks of text all at once instead of token by token. and that is something that MidJourney confirmed today because it was only rumored the past few months but they confirmed today that they were looking into. So all those things are like very exciting new model architectures that are maybe something that you'll see in production two to three years from now. So the couple other trends that I want to just get your takes on because they're sort of something that seems like they're coming up are one, sort of these wearable, you know, kind of passive AI experiences where they're absorbing a lot of what's going on around you and then kind of bringing things back.
28:04And then the other one that I wanted to see if you guys had thoughts on were sort of this next generation of chip companies. Obviously, there's a huge amount of emphasis on hardware and silicon and different ways of doing things. But, you know, love your take on either or both of those. So wearables, I'm very excited about it. I want wearables on me at all times. I have two right here to quantify my health. And, you know, I'm all for them. But society is not ready for wearables, right? Like no one's comfortable with a device on recording every single conversation we have. Even all three of us here as podcasters, we don't record everything that we say.
28:42And I think there's a social shift that needs to happen. I am an investor in Tab. They are renaming to a broader vision, but they are one of the sort of three or four leading wearables in this space. It's sort of the AI pendants or AI OS or AI personal companion space. I have seen two humanes in the wild in San Francisco. I'm very, very excited to report that there are people walking around with those things on their chest. And it is as goofy as it sounds. It absolutely is going to fail, but God bless them for trying. And I've also bought a rabbit. So I'm very excited for all those things to arrive.
Read the full transcript
29:21But yeah, people are very keen on hardware. I think the idea that you can have physical objects that embody an AI that do specific things for you is as old as the sort of golem in sort of medieval times in terms of how much we want our objects to be smart and do things for us. And I think it's absolutely a great play. The funny thing is people are much more willing to pay you upfront for a hardware device than they are willing to pay like an$8 a month subscription recurring for software. right and so the interesting economics of these wearable companies is they have negative float um in in the sense that people pay deposits up front for like like i paid like i don't know 200 bucks for the rabbit up front and i don't get it for another six months uh i paid 600 bucks for the for the tab and i don't get it for another six months um and and then then they can take that money and and sort of invest it in like their next the next events or their next properties ventures.
30:24And I think that's a very interesting reversal of economics from other types of AI companies that I see. And I think just the tactile feel of an AI I think is very promising. Alessio, I don't know if you have other thoughts on the wearable stuff. The Open Interpreter just announced their product four hours ago. It's not really a wearable, but it's still like a physical device. It's a push to talk mic to a device on your laptop. It's a$99.
30:57But again, going back to your point, it's like people are interested in spending money for things that they can hold. I don't know what that means overall for where things are going, but making more of this AI be a physical part of your life, I think people are interested in that. But I agree with Sean. I mean, I've been, I talked to Avi about this, but obvious point is like most consumers like care about um utility more than they care about privacy you know like you've seen with social media um but i also think there's a big societal reaction to ai that is like much more rooted than the social media one um but we'll see but a lot again a lot of work a lot of developers a lot of money going into it so um there's there's bound to be experiments being run.
31:46On the chips, I... Sorry, I'll just keep shipping one more thing and then we transition to the chips. The thing I'll caution people on is don't overly focus on the form factor. The form factor is a delivery mode. There will be many form factors. It doesn't matter so much as where in the data war does it sit? It actually is context acquisition and maybe a little bit of multimodality. Context is king. If you have access to data that no one else has, then you will be able to create AI that no one else can create. And so what is the most personal context? It is your everyday conversation. It is as close to mapping your mental train of thought as possible without physically you writing down notes.
32:27So that is the promise, the ultimate goal here, which is personal context. It's always available on you. You know, we'll see all that stuff. But that's the frame I want to give people, that the form factors will change and there will be multiple form factors, but it's the software behind that. in the personal context that you cannot get anywhere else, that'll win. Yeah, so that was wearables. On the chip side, yeah, Grok was probably the biggest release. Jonathan, but it's not even a new release because the company, I think, was started in 2016. So it's actually quite old, but now recently captured the people's imagination with their MixRAL 500 tokens a second demo.
33:08yeah I think so far the battle on the GPU side has been either you go kind of like massive chip like the Cerebrus of the world where one chip from Cerebrus is about 2 million dollars you know that's compared obviously you cannot compare one chip versus one chip but H100 is like 40 ,000 something like that the problem with those architectures has been they want to be very general you know but like they wanted to put a lot of the SRAM on the chip. It's much more convenient when you're using large language models, but the models outpace the size of the chips and chips have a much longer turnaround cycle.
33:52Grok today is great for the current architecture. It's a lot more expensive also as far as dollar per flop. But their ideas like, hey, when you have very high concurrency, we actually were much cheaper. You shouldn't just be looking at the compute power. For most people, this doesn't really matter. I think that's the most interesting thing to me. We've now gone back with AI to a world where developers care about what hardware is running, which was not the case in traditional software for maybe 20 years, as the cloud is getting really big. my thinking is that in the next 2-3 years we're going to go back to that people are not going to be sweating what GPU do you have in your cloud what do you have you want to run this model we can run it at the same speed as everybody else and then everybody will make different choices whether they want to have higher front end capital investment and then better utilization some people would rather do lower investment before and then upgrade later there are a lot of parameters and then there's the dark horses that is some of the smaller companies like Lemurian Labs MedX that are working on maybe not a chip alone but also some of the actual math infrastructure and the instructions on it that make them run.
35:17There's a lot going on but yeah I think the episode with Dylan will be interesting for people but I think we also came out of it saying hey everybody has pros and cons uh there's no it's different than the models where you're like oh this one is definitely better for me and i'm gonna use it i think for most people it's like fun twitter memeing you know but it's like 99 of people that tweet about this stuff are never gonna buy any of these chips anyway so it's it's really more for entertainment uh wow i mean like this is serious business here right you're talking about you know like like the potential new nvidia if anyone can take like 1 % of NVIDIA's business they are a serious startup that you should look at.
35:59So that's my take on Madx. I'm more talking about like how should people think about it. I think like the end user is not impacted as much. I disagree. I love disagreements because who likes a podcast where all three people always agree with each other. You will see the impact of this in the tokens per second over time. This year I have very very credible sources all telling me that the average tokens per second, right now we have somewhere between 50 to 100 as like the norm for people. Average tokens per second will go to 500 to 2 ,000 this year from a number of chip suppliers that I cannot name.
36:39So like that will cause a step change in the use cases. Every time you have an order of magnitude improvement in the speed of something, you unlock new use cases that become fun instead of a chore. and so that's what I would caution this audience to think about which is like what can you do in much higher AI speed it's not just things streaming out faster it is things working in the background a lot more seamlessly and therefore being a lot more useful than previously imagined so that would be my two cents on that yeah I mean the new NVIDIA chips are also much faster to me that's true when it comes about startups are the startups pushing the performance on the incumbents or are the incumbents still leading and then the startups are riding the same wave?
37:28I don't have yet a good sense of that. It's next year's NVIDIA release just going to be better than everything that gets released this year. If that's the case, it's like okay, damn Jensen. It's like the meme, it's like I'm going to fight NVIDIA. It's like damn, Jensen got hands. He really does. So, I'll see. well awesome conversation guys i guess just just by way of wrapping up uh call it over the next three months between now and sort of the beginning of summer uh what's one one prediction that each of you has it can be about anything could be big company it can be startup it could be something you have privileged information that you know and you just won't tell us that you actually know what does it have to be something that we think it's gonna be true or like something that we think because for me it's like is sundar gonna be the ceo of google maybe not in three months maybe in like six months nine months you know people are like oh maybe that miss is gonna be the new ceo that was kind of like i i was busy like fishing some deep mind people and google people for like a good guest for the pot and i was like oh what about jepteen and they're like well that miss is really like the person that runs everything anyway and the stuff and it's like interesting and so i don't know sergey sergey sergey could come back i don't know like he's making more appearances these days yeah i don't i bet we can just put it as like you know yeah my thing is like ceo change potential but um again three months is too short to make a prediction i think that's that's fine the time scale might be off yeah i mean uh for me i i think the progression in vertical agent companies will keep going.
39:11We just had the other day Klarna talking about how they replaced 700 of their customer support agents with AI agents. That's just the beginning, guys. Imagine this rolling out across most of the Fortune 500. And I'm not saying this is a utopian scenario. There will be very, very embarrassing and bad outcomes of this where humans would never make this mistake, but AIs did and we'll all laugh at it or we'll be very offended by whatever bad outcome it did. So we have to be responsible and careful in the rollout. But yeah, it's rolling out. Alessio likes to say that this year is the year of AI production.
39:47Let's see it. Let's see all these sort of vertical, full-stack employees come out into the workforce. Love it. All right, guys. Well, thank you so much for sharing your thoughts and insights here. And I can't wait to do it again. Thanks for having us.
40:07Thank you.
From the publisher
NLW is back with Part 2 of his conversation with the hosts of Latent Space. In this segment they hone in on the trends Swyx and Alessio are seeing in what developers are building.
Be the first to learn about our new AI education platform: https://besuper.ai/
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
