In short
AI Daily Brief Notes: Episode - Elon Turns On "Most Powerful AI Training Cluster In the World"
Podcast Overview
- Title: The AI Daily Brief (Formerly The AI Breakdown)
- Description: A daily news analysis show on artificial intelligence, examining creativity, industry disruption, philosophical questions, and ethical concerns regarding AI.
Episode Summary Title: Elon Turns On "Most Powerful AI Training Cluster In the World" Description: Elon Musk activates the Memphis Supercluster, aimed at creating the world's most powerful AI by December. Discussions also explore whether it is too late to start AI startups or vertical LLMs.
---
Key Highlights
- Elon Musk's Memphis Supercluster
- Activation Time: 4:20 AM local time.
- Capabilities:
- Contains 100,000 liquid-cooled H100 chips.
- Claims to be the most powerful AI training cluster globally.
- Musk aims to train the world's most powerful AI by December 2023.
- Context: This supercluster was part of a substantial fundraising effort, initially projected at $3 billion but later increasing to $6 billion.
- Competition in AI Landscape
- Current Landscape:
- Discussions around smaller AI models and the competition between models like GPT-40 and Claude 3.5.
- The emergence of lurking competitors, including Musk's xAI and its product Grok.
- Industry Dynamics:
- The podcast touches on how NVIDIA is facing challenges with U.S. AI chip export restrictions.
- Apple introduces new small models outperforming existing competitors.
- Vertical AI Startups Discussion
- Main Question: Is it too late to launch vertical AI startups?
- Insights from Industry Experts:
- Emily from VC: Skeptical about specific companies (e.g., Harvey, a legal AI startup), suggesting they may struggle.
- Bruno Koba (Stanford MBA): Highlights that every vertical is now filled with AI startups and questions whether specialized models can truly compete with general models like GPT-4.
- Gary Tan (Y Combinator): Asserts that it's not too late for LLM-powered alternatives, emphasizing that small teams can create impactful software.
- Challenges Faced by Vertical Models
- Key Issues Identified:
- Specialization vs. Generalization: General models may outperform specialized models, reducing their market viability.
- Companies face a gap between the intention to use AI and the actual application of these tools.
- Capacity Issues: Organizations struggle to learn and adapt AI technologies effectively.
- Lessons from Historical Contexts
- Comparison with past tech cycles (internet and mobile) suggests that current large AI labs can fulfill needs across all verticals, which differs from the previous landscape where startups could fill specific gaps.
---
Key Takeaways
- Musk's Ambition: The launch of the Memphis Supercluster marks a significant competitive move in the AI industry.
- Vertical AI Startups: While skepticism exists about the viability of new entrants, experts suggest that there are still opportunities for innovative solutions tailored to specific needs.
- Current Era: This is characterized as an experimentation phase, with potential for growth among companies willing to effectively integrate AI into their workflows.
- Market Dynamics: The evolving landscape presents both challenges and opportunities as enterprises seek meaningful AI applications.
---
Conclusion This episode of the AI Daily Brief highlights the current competitive landscape in artificial intelligence, particularly with Elon Musk's ambitious plans for AI development alongside discussions about the viability of vertical AI startups. As the industry continues to evolve, the podcast emphasizes the need for startups to find unique niches and provide real value to enterprises navigating the complexities of AI integration.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the AI Daily Brief, is it too late to start a vertical AI enterprise startup? Before that, in the headlines, Elon turns on what he calls the world's biggest AI training supercluster. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. To join the conversation, follow the Discord link in our show notes.
0:23Hello friends, quick note before we dive in. You might notice that the audio feels a little off in the headline section. I realized only after recording that I was recording with the main computer microphone rather than the normal good podcaster microphone. We've done our best to try to fix it up, but bear with me for the three or four minutes of that. And then in the main section, you'll hear the good, nice audio once again. Welcome back to the AI Daily Brief Headlines Edition, all the daily AI news you need in around five minutes. While most of the LLM war discussion recently has been around either A, the push to competition on the smaller model front, something that we've talked about recently and even have a story on today, or, of course, the question of GPT-40 versus Claude 3.5 Sonnet.
1:05Still, there are a handful of lurking competitors who, it feels clear, haven't really fully entered the fray, one of those, of course, being Elon Musk's XAI. Now, Grok has done some things. We're on version 1.5, version 1 was open-sourced, and simply by virtue of the fact that it is an Elon project, meaning that it gets to be integrated into Twitter slash X as well as access that data, and in the fact that maybe at some point there'll be some Tesla connection, means that it would be highly inadvisable to write them off. What's more, Elon has been talking for a while about the fact that they were going to invest a huge amount in computing resources, and as of this morning, we got a little bit more information about that.
1:42At 5.57 a.m. Eastern Time, Elon tweeted, Nice work by XAI team NVIDIA and supporting companies getting Memphis supercluster training started at 4.20 a.m. local time. With 100 ,000 liquid-cooled H100s on a single RDMA fabric, it's the most powerful AI training cluster in the world. This is a significant advantage in training the world's most powerful AI by every metric by December this year. Elon has indicated previously that Grok 3 would be trained on these 100 ,000 H100 chips. And what's more, there's been a flow of a number of Tesla employees over to xAI, so much that some in the Tesla community have even accused Elon of poaching from himself.
2:19This supercluster was a big part of what Elon was fundraising for, that was originally going to be$3 billion raised on a$21 billion valuation, but ballooned at the last minute to$6 billion on 24. Reports are that between half and two-thirds of that will go to simply buy and compute. I think for us one takeaway is that we now have Elon throwing down a clear gauntlet. He wants to have the world's most powerful AI by every metric by December this year. Next up, speaking of NVIDIA, one of the few headwinds facing that company has been US AI chip export restrictions. Throughout the last couple of years, the Biden administration has continuously tightened those rules, forcing companies that sell chips to China to frequently update their offerings to come in under various power thresholds.
3:01Still, it seems for now the calculus is that rather than abandoning the Chinese market, it makes more sense to try to customize offerings to be as close to the state of the art as possible within those restrictions, with the latest efforts from NVIDIA coming in a chip they're calling the B20. This is part of the Blackwell chip line that was announced in March, which is slated to go into production later this year. The information writes, it's unclear how NVIDIA designs the B20 to align with US rules while keeping competitive in China, basically referring to the fact that as these restrictions take hold, more and more of the Chinese market is moving to local alternatives where they can.
3:35Back to small model competition, Apple has recently shared a new small model that outperforms Mistral. VentureBeat points out that Apple has released a package of two models in the DCLM, or Data Comp for Language Models project, on Hugging Face, one with 7 billion parameters and the other with$1.4 billion. VentureBeat writes they both perform pretty well on the benchmarks, especially the bigger one, which has outperformed Mistral 7B and is closing in on other leading open models including Llama 3 and Gemma. Vaishal Shankar from Apple writes, We have released our DCLM models on Hugging Face. To our knowledge, these are by far the best performing truly open-source models.
4:09Open data, open weight models, open training code. Our$6.9 billion base model is competitive with Mistral Llama 3 Gemma Quen 2 on most benchmarks, but we also release our entire training set and pre-training recipe for the community to build on top of. We also have a strong 1.4b version, which significantly outperforms recently released Soda small LM models. We additionally release an instruction-tuned variant of these models that exhibit strong performance on IT benchmarks, like Alpaca Bench. As we pointed out last week, the competition for small model terrain is just as vicious and intense in many ways as the competition for AGI in the state of the art.
4:41Speaking of Apple and competition, AMD recently has been talking a big game. The Verge writes, AMD says its new laptop chips can beat Apple, but still has to prove it. The Verge writes, 2024 will go down in tech history as the year Microsoft was finally able to make Windows laptops into serious competitors to the MacBook. So far, that's thanks to Qualcomm's new Snapdragon chips, which switched to a homogeneous chip architecture, increased clock speeds, and caught up to Apple's speedy and power-efficient processors. But now, AMD says it has chips that can take on the MacBook too, and keep the company's processors in the mix.
5:11In fact, at a recent event, one of the Verge authors writes, I heard AMD brag about beating the MacBook more than I've ever heard a company directly target a competitor before. The general vibe of the piece, however, is prove it. For now, lots of interesting competition to start the week, but that's going to do it for the AI Daily Brief Headlines edition. Up next, the main episode. Today's episode is brought to you by Super Intelligent, the platform for fun, fast AI learning. Super has a ton of new things going on. We recently announced our partnership with Spotify, through which users of that app can now access Super Intelligent content directly from their mobile apps.
5:45We've also just launched the AI learning feed. In addition to seeing the tutorials that we're dropping, there are polls, news items with related lessons, and a chance for people to show off the projects and use cases that are making AI come alive for them. We've also just kicked off the Super Summer Challenge, where each week we'll share a new challenge that you can use to discover new AI tools and use cases. Go to bsuper.ai and use code SUPERFUN for 50 % off your first two months. That's bsuper.ai. Today's episode is brought to you by Venice. The leading AI companies store your entire conversation history and attach it to your identity forever.
6:21Every question you ask, every answer you receive, every image you generate, every thought you share with the machine, it's all being spied on. If you trust all the companies, hackers, and NSA board members that will ever have access to your AI conversations, then rejoice, for you are well served. For the rest of us, Venice is an alternative. Venice is a powerful AI app for text, image, and cogeneration that respects you as a sovereign individual and believes privacy and free speech are not only human rights, but are necessary for civilizational advancement. Private, permissionless, and uncensored.
6:49You can try it for free without an account at Venice.ai. Welcome back to the AI Daily Brief. Over the weekend, I saw an interesting conversation emerge around effectively whether it's too late to start building some sort of vertical AI startup going after a particular industry. I actually think it's a really useful lens by which to understand where we are in the enterprise cycle as relates to AI. And so I want to talk through this a little bit more in depth, bringing into it some of the conversations that we're having over at Superintelligent. Our main customer base at Superintelligent is individuals and enterprise customers who are thinking about how to use AI in their particular professional verticals.
7:31So we have a lot of reps just with this particular conversation. Now, where this started was with this provocative tweet from at Emily in VC, where she wrote, Calling it now. Harvey, the legal AI company, will end up being roadkill on the side of the highway. Complete smoke and mirrors company. More precisely, it's a zero. Likes aren't public, but y 'all would be floored by the people who are liking this tweet. Some of the top founders and investors in the Valley. This is the company she's referring to. Harvey, which frames itself as the trusted legal AI platform. Augment your workflows using domain-specific models trained by and for professional service providers.
8:05In another article by Edward Buxtell, some of Harvey's history was given. OpenAI invested in the company the same month that it launched ChatGPT, which, as Buskell says, means that Harvey arguably had a three - to six-month head start on virtually all of its competitors in legal technology. The company was literally working with ChatGPT 3.5 before any other generative AI products in legal. There is no question that Harvey could have come out of the gate with a go-to-market strategy that could have easily captured the imagination of legal early adopters across the spectrum of specializations. Harvey, with its massive first-mover advantage, could have set the standard for legal LLMs while creating the workflows for ethical use and manage the hallucination issue.
8:42Now, the rest of the article is a bit skeptical, but I'm more interested in why there is skepticism here than in trying to validate it or not, as I've spent no time with Harvey and have no real sense of how the company is actually doing. Two reference points that I think are interesting. The first is a comparison to Bloomberg GPT. For those of you who don't remember, Bloomberg GPT was a specially trained finance LLM that took advantage of all of Bloomberg's data. And basically, Bloomberg had bet that with their specialized data, they'd be able to outperform a generalist model. However, when GPT-4 came out, which of course did not have any specialized finance training, it still beat it on basically every finance tasks.
9:18Wrote Professor Ethan Mollick back then, it is part of a pattern. The smartest generalist frontier models beat specialized models and specialized topics. Your special proprietary data may be less useful than you think in the world of LLMs. Again, I don't know if Harvey is experiencing any sort of version of this. They claim that their custom trained models are preferred 97 % of the time by clients. But I do think that right now with any verticalized models, this is a really important question to watch. It's one that different people are going to disagree on and will have a huge impact on how vertical industry-oriented LLMs function or not in the future.
9:51The other thing, though, that I want to point out is that Emily followed up her own tweet with a screenshot from a big law partner that uses quote-unquote Harvey from her DMs. That law partner said, Harvey has an insane amount of inactive quote-unquote customers that will churn. Big law firms that want to tell clients they use AI signup and then no one uses it. Now, this is sort of presented in the context of this conversation as a big gotcha, like Harvey is tricking people in some way. And of course, this is just an anonymous source in the DMs. But even if this is accurate, I don't think it necessarily means what it's being presented to mean here, i.e.
10:26that Harvey is just a smoke and mirrors company. There is a huge gap right now. It's the gap that we spend literally all day in every day between the interest and excitement and intention to use AI to get value and the actual capacity to do so. There are a litany of reasons why that gap exists. Yes, one piece of it may be this sort of AI theater that this big law partner is arguing is happening, where big law firms just want to tell clients they use AI, even though they don't. But there are other much less pernicious issues as well. There's a capacity issue in helping people actually learn these tools.
11:00And then there's also the difficult, slow process of industries taking the time to figure out what parts of their process these tools actually help with. Now, Bruno Koba took this conversation and brought it up a level from the specifics of Harvey and the legal industry to vertical AI startups in general. Bruno is an MBA candidate at Stanford and writes, I've been investigating vertical AI startups profoundly for the past few weeks. I think we're in a very strange part of the cycle in AI startup funding and development. One literally every vertical, finance, law, healthcare, etc., is now populated with AI startups building on top of foundation models.
11:32If you start today, it's too late to get first mover advantage. Two, however, foundation models are still not quite there yet to solve problems in those verticals in a tangible bulletproof way. Those who build GPT wrappers, even if they refuse to admit the label, rely on OpenAI's next big model to truly prove sticky long-term retention. And those who decide to train their own specialized models like Harvey are apparently not outperforming smart general models GPT-40 Claude 3.5 Sonnet. Three, startups are facing a massive dependence on big AI labs shipping the next big model. And once that happens, it might just be that we won't need GPT wrappers at all.
12:04ChatGPT will be enough for whatever task we hand them. Four, this is fundamentally so different from the internet and mobile innovation cycles. Back then, incumbents could not touch every single vertical with their product offerings, so startups filled the void and became multi-billion dollar companies. E.g. Apple launched the App Store, but wouldn't build apps for food delivery, dating, ride-sharing, etc. Now, when GPT 5 launches, OpenAI can touch virtually every industry they want. 5. My sentiment is that we're living in a kind of born-too-late-to-explore-the-Earth, born-too-early-to-explore-space-5 in this 23-24 cycle in the realm of AI startups.
12:35And when the rocket ships to space are finally ready, i.e. truly smart general LLMs, vertical AI startups will capture much less value versus incumbents when compared to previous cycles. I think here it's worth narrowing the parameters of the conversation, or at least breaking it apart, between, on the one hand, the question of AI startups in general, and on the other, the specifics of verticalized LLMs. As I mentioned before, there is this big question around whether specially trained or fine-tuned vertical models can ever actually outcompete the generalist state of the art, and Bruno is absolutely right that the answer to that question will have a significant impact on how AI rolls out across industries.
13:14Gary Tan, the president and CEO at Y Combinator, however, isn't so sure. He writes, I don't think it's too late to enter almost any software market with an LLM-powered alternative if you want to. We are seeing tiny teams of a few people build valuable software leapfrogging incumbents with even today's frontier models. The war will be retention. Who can serve your vertical better? I responded to Gary's post, which was of course the inspiration for this full episode, and said, we talk to hundreds of enterprise customers for these startups every week. There's no universe in which the game is won right now.
13:42In fact, if anything, enterprises are getting more sophisticated about wanting a combination of powerful NFAI with real product UI and UX. And that's why I wanted to separate the intrinsic question inside Bruno's tweet around the ultimate role for custom-trained verticalized LLM models from the question more broadly of whether the big labs are just going to eat all of the enterprise business. The era that we are in right now is a use case exploration era. It is an experimentation era that is happening both vertically inside the organization, i.e. these big custom models that take advantage of all their data, but it's also happening horizontally, where, for example, individuals across those law firms are experimenting and figuring out which use cases actually save them time and make things easier or better.
14:24Part of the opportunity, I believe, is for companies to come in and actually design product experiences around LLMs that interact positively with those high-value use cases as they get discovered. In other words, I tend to think that for products to be sticky inside these companies, they're going to have to actually be products. They can't just be ChatGPT, but for legal. Now, one thing that I will also note, where Bruno writes, this is fundamentally so different from the internet and mobile. Back then, incumbents could not touch every single vertical with their product offerings. Sort of. But social really challenges that idea.
14:58When I moved to San Francisco in 2008, it was specifically to help change.org compete in the social network for social goods space. It was a time when four years after Facebook had been launched, pretty much everyone was convinced that there was going to be a Facebook for everything. There was going to be a Facebook for social change. There was going to be a Facebook for you name it. Pick your industry vertical, there was going to be a Facebook for that. There were so many social networks for social change at that time that there was even an aggregator platform called Social Actions, which took inputs from all 40 of those platforms and allowed you to do it from a single spot.
15:31Now, of course, at the end of the day, there wasn't a place for a social network for social change. The social network for social change, just like the social network for everything else, was just meta. The things that came later that were also social networks had fundamentally different type of content experiences at their core. Instagram with mobile and photos, Snapchat with mobile and disappearing messages, et cetera, et cetera, et cetera. And yet still, the ecosystem around all of those social channels is huge. So I'm not sure. What I do know is that for aspiring entrepreneurs out there in the AI space interested in particular verticals, our read from all our conversations is that there is still a ton of space and that the only thing that is clear at this stage is that enterprises want partners who can actually help them take advantage of AI to drive real value.
16:18If you think you have a good answer for that, I tend to think that there's more of a market than many might think. For now, though, that is going to do it for today's AI Daily Brief. Until next time, Peace.
From the publisher
At 4:20am, Elon Musk turned on the Memphis Supercluster to begin training what he claims will be the world's most powerful AI by December. Also NLW explores a question: is it too late to start AI startups (or at least vertical LLMs)?
Concerned about being spied on? Tired of censored responses? AI Daily Brief listeners receive a 20% discount on Venice Pro. Visit https://venice.ai/nlw and enter the discount code NLWDAILYBRIEF.
Learn how to use AI with the world's biggest library of fun and useful tutorials: https://besuper.ai/ Use code 'podcast' for 50% off your first month.
The AI Daily Brief helps you understand the most important news and discussions in AI.
Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614
Subscribe to the newsletter: https://aidailybrief.beehiiv.com/
Join our Discord: https://bit.ly/aibreakdown
