In short
The AI Daily Brief: Episode Summary
Episode Title
The 5 Important AI Models Released This Week
Podcast Overview
- Host: NLW
- Description: This podcast provides daily insights into the landscape of artificial intelligence, covering advancements, ethical concerns, and practical implications.
Key Highlights This week’s episode focuses on five significant AI models that were either released or updated, showcasing the rapid advancements in the AI field. Notably, the episode emphasizes a new music generation tool called Udio and updates to several large language models (LLMs).
---
Major Announcements
- Udio: A New Music Generator
- Overview: Udio has emerged as a leading music generation tool, drawing attention away from its predecessor, Suno.
- Key Features:
- High-quality music generation
- Captured the interest of AI enthusiasts on platforms like Twitter.
- Community Reaction:
- Positive reception for its quality.
- Some skepticism regarding its ability to produce songs with emotional depth compared to Suno.
- Discussion on the "hype cycle" – initial excitement may lead to oversaturation and a return to unique, high-quality creations.
- Mistral 8×22B MOE
- Overview: A new model from Mistral that may outperform existing models like GPT-3.5 and Llama 270B.
- Specifications:
- Potentially has up to 176 billion parameters.
- Context length of up to 65K tokens.
- Community Sentiment: Excitement in the open-source community, especially after Mistral's previous closed-source move.
- Google Gemini 1.5 Pro
- Overview: Publicly launched with significant improvements, including a million-token context window.
- New Features:
- Added capabilities for audio and speech understanding.
- Enhanced file handling through a new file API.
- Community Insights: Praised for its video understanding capabilities, offering new use cases.
- OpenAI's Improved GPT-4 Turbo
- Overview: Marketed as a "majorly improved" version of GPT-4.
- Use Cases Highlighted:
- New applications in software engineering and nutrition analysis.
- Community Skepticism: Some users question the actual improvements compared to previous versions.
- Meta's Llama 3
- Overview: Upcoming models set to be released soon, with a shift in Meta's strategy to be competitive with state-of-the-art models.
- Expectations:
- Multiple versions with varying capabilities planned for rollout.
- Aimed at becoming the most useful AI assistant.
---
Discussion Points
- AI Competition: The rapid pace of releases suggests that the AI field is entering a phase where new models are expected as the norm.
- Quality vs. Quantity: As AI tools become more accessible, the focus may shift to the quality of output rather than just the ability to generate content.
- Community Dynamics: The podcast reflects on the evolving relationship between developers and users in the AI space, highlighting the debates on creativity and authenticity in AI-generated content.
---
Conclusion
- The episode encapsulates a week of significant developments in AI, with a focus on the balance between innovation and community reception.
- NLW emphasizes the normalization of frequent announcements in AI, suggesting that the industry is moving towards a standard cadence of innovation that may lead to a more critical evaluation of what constitutes a breakthrough.
---
Additional Resources
- Superintelligent Platform: New platform with over 300 video tutorials on AI usage.
- Subscribe for more updates and insights on AI through the [AI Breakdown newsletter](https://theaibreakdown.beehiiv.com/subscribe) and YouTube channel.
Ending Note
- The episode concludes with a nod to the ongoing evolution in AI, encouraging listeners to engage with the new tools and participate actively in the ongoing discourse surrounding artificial intelligence.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:01Today on the AI Breakdown, we're looking at five new important models that have launched just this week. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our YouTube, our Discord, and our newsletter.
0:24Hello, AI friends. A quick note before we get into today's show. If you've been listening, of course, you will know that we launched Superintelligent yesterday. And for those of you who don't know, Superintelligent is our new platform for teaching people how to use AI. We've got more than 300 fun, fast, and most importantly, useful, practical video tutorials. Each of them has a companion project, which is a set of step-by-step instructions that allow people to actually use the tools and create the things that are being created in these tutorials. And we're adding between 30 and 50 new video tutorials each week.
0:55That has just gone live to the public. It is$20 a month for unlimited access. If you're interested, check it out at bsuper.ai or email me nlw at bsuper.ai if you have questions. It was directly inspired by conversations that I had with all of you AI Breakdown listeners, and so I'm really excited for you to check it out. Because of that launch, though, this week has been a little bit chaotic, and so today, instead of our normal brief in-demain episode format, we're actually doing a bit of a countdown of the five important models that have been launched this week, including four LLM models as well as one music model that has everyone buzzing.
1:26So without any further ado, let's dive in. Welcome to the AI Breakdown. Everyone anticipated, I think, that 2024 was going to be quite the year when it came to AI competition and technical advances. And this week is a great example, a little snapshot that shows just what passes for normal in 2024. We're going to look at five new models that were released this week or updated in some cases, four of which are LLMs. But the first of which is a new music generator called Udio that has absolutely captured the AI Twittersphere's attention. For the last couple of months, the big hot player in AI music has been Suno.
2:02We've talked about them here on the show, and they really represented a major advance from things that we had seen before. However, a couple of weeks ago, we started seeing little drips and drops that there was something even better coming. Last week, the first contraband examples of this music generator's output started to be published to X, and this week we finally got the service itself, which again is called Udeo. Now, just to give a quick example of what this thing can do, let's listen to a snippet of Dune, the Broadway musical.
2:32Pretty amazing stuff. And consequently, there have been a ton of posts like this one from Min Choi, which reads, This is wild. UDO just dropped and it's like Sora for music. The music are insane quality, 100 % AI. I thought it was interesting though, when Bilual C2 tweeted, Anyone else notice this pattern with new AI features? 1. Cool new tool drops. 2. Early adopters rake in insane views because it's novel. 3. Everyone jumps on the bandwagon. 4. Overuse equals oversaturation equals no one cares anymore equals views drop. 5. It always comes back to the creators making something truly unique. It's the quintessential hype cycle in three months, and it only seems to be compressing further.
3:08It's Suno now, but Sora will be the same when it hits GA. If anyone can do it on demand, it loses value because it's not special anymore. At that point, it's back to creative self-expression. Except now you're creating at a higher level of abstraction with your newfound superpowers to make something unique that cuts through the noise. Now, hold aside any specific examples of whether this is happening to Suno or will happen to Udeo, I think it gets at a profound truth about the world that we're moving into. In a world where so much can be created so easily, the quality of what actually breaks through and captures attention is going to consequently have to go up incredibly.
3:41It is extremely unlikely to me that the best songs that come even from a sophisticated application like Udeo are going to come from a random person on Twitter who just happened to stumble into it versus someone who is probably a musician or a songwriter already, adapting to and mastering a new medium. Now, while by and large, most of the folks that I've seen have ranked Udeo ahead of Suno at this point, developer Nick Dobos disagrees. He writes, Udeo has better audio quality, but I haven't found a single song I've liked. Missing a certain something something. Meanwhile, I have around five Suno songs stuck in my head right now.
4:13Another AI creator, Boris, followed up and said, same, Suno generates complete bangers. Udeo, interesting snippets, no control over the lyrics to experimental styles in general. Getting at the broader point that I was just making, Nick Dobos also tweeted, With the rise of Gen.ai, we are going to see an interesting clash of cultures, as new creatives start and learn with Gen.ai tools first. A big cohort of coders, visual artists, musicians, and video makers will go Midjourney, Dali, Suno, and Udeo, Runway before Photoshop, Canva, Final Cut, Ableton, and Fruity Loops. Expect some wildly different styles as this younger cohort learns creativity at a higher abstraction level, and then learns traditional media and productivity tools backwards.
4:47Also, expect 90 % of the previous cohort to cry and whine about it, this isn't real art. The other 10 % will embrace the new tech and build the most amazing things you've ever seen by combining new and old techniques. I could not agree more that this is likely to be the pattern. And whatever you think about this, if you haven't had a chance to play with UDO yet or go listen to some of the creations, it is highly worth taking a few minutes to do so. From there, let's move into our LLM announcements. The first wasn't really an announcement at all. Mistral, as they have classically come to do, just dropped a link to a release of a new model on Twitter without any further explanation.
5:19The model is called 8X22B MOE, and Gigazine writes, Although details are unknown, 8X22B MOE may have more than three times the number of parameters of the model Mistral 8X7B, which has been shown to outperform GPT 3.5 and Llama 270B in many benchmarks. They also add, The total number of parameters may be up to 176 billion. The context length that can be handled is said to be 65K. Now, the open source community is really excited about this one, not only because it appears that it might have increased capacity, but also because the last Mistral model that was announced was their first closed source model, which was to be distributed exclusively through Microsoft.
5:53At the time, I talked about how most of the people that I saw in the open source community were willing to not just give them the benefit of the doubt, but understand that they had to fund the business somehow. But still, seeing them continue to advance on the open source side has a lot of people breathing a sigh of relief and staying up all night to start hacking. Today's podcast is brought to you by Plum. Is your product team struggling to keep up with the incredible pace of AI development? Are you tired of spending countless engineering hours just to test out small prompt changes in your product?
6:21Thankfully, there's Plum. Build cutting-edge AI experiences for your users in a fraction of the time. Say goodbye to the slow, tedious process of hand coding and hello to the future of AI development. Get ahead of your competition and start moving as fast as AI does. Check out useplum.com and shoot me a message to get early access. Another surprise release a little bit earlier in the year came from Google. You'll remember that back in December, facing down intense pressure to try to keep up or catch up with OpenAI, Google announced its suite of Gemini models. The problem with that announcement was that their biggest and most performant model, the one that they said actually beat GPT-4 on many benchmarks, wasn't going to be available until sometime in 2024.
7:01When that Ultra model did finally come out, Google surprised everyone by just about a week later announcing Gemini 1.5 Pro, an even more advanced model that most notably was said to have a 1 million token context window, completely blowing the doors off of everything we had seen before. At the time, Gemini 1.5 Pro was only available to developers through Google's AI Studio, but this week Gemini 1.5 Pro has moved into a public preview period and can be accessed via the Gemini API. In addition to just being more widely available, they've also added native audio or speech understanding capabilities, as well as a file API to make it easy to handle files.
7:38Their announcement post writes, We're also launching new features like system instructions and JSON mode to give developers more control over the model's outputs. Didi at Menlo Ventures writes, Gemini 1.5 Pro's video understanding is the most underrated thing in AI. In 50 seconds it quote-unquote saw an 11-minute YouTube video, around 175k tokens, of the most iconic moments in sports and was able to perfectly, to my knowledge, list all 18 of the moments. there is no other video AI this good. What Didi is getting at is one of the reasons that people have been so excited about Gemini 1.5 Pro and its longer context window, is that it really does open up totally new use cases that just aren't possible with different types of input lengths.
8:16Not to be totally outdone, however, OpenAI announced what they called a quote majorly improved GPT-4 Turbo model, first available through the API and slowly rolling out into chat GPT as well. Now, of course, when a lab calls something majorly improved, It opens it up for people to ask, is it majorly improved? Professor Ethan Mollick writes,
8:47He goes on,
8:59mark? Still, from a marketing standpoint, OpenAI is going hard on what the new capacities mean. They shared a thread on Twitter slash X where they showed off use cases of the new GPT-4 Turbo that include that very impressive Devon AI Software Engineering Assistant. They pointed to a nutrition app that used GPT-Turbo-4 with Vision to get better insights about what was in food. And they talked about TL Draw, which is an incredibly advanced AI-powered UI designing tool that's gotten a lot of buzz on Twitter recently as well. Now, it appears that one of the things that OpenAI is excited about is the increase in co-generation performance with this new GPT-4 Turbo.
9:33Some are skeptical, however. Benjamin DeKracker, for example, tweets, This is why OpenAI staff have been acting all cheeky. Newest GPT-4 Turbo update is doing very well on coding benchmarks. Very well indeed. Here's the thing. I don't really believe it. For example, these also show the old version of GPT-4 beating Claude 3 Opus at co-generation. In my real-world experience, that's not true. Opus has been significantly better. TLDR, the upgrade is probably good, but I'm not convinced it's actually back at number one for code. Real world is what matters, which is always hard to measure. Now on the flip side is Pietro Chirano who writes, Side-by-side comparison between the latest version of GPT-4 Turbo and the previous one, 0125 Preview.
10:11Not only is the new version less verbose and it goes directly into code, but it also rightfully so decides to add a flag to download the highest quality video. Smart. This of course, in some ways agrees with what Benjamin was saying, that it's not going to be evaluations that matter, but ultimately what people find in practice. Now, one meta note that was really interesting to me comes once again from Bilawal Sidhu, who writes, Is it just me or was today the first time OpenAI was unable to overshadow a Google AI announcement? Gemini 1.5 Pro is pretty wild. Just dropped in an audio file, an hour-long video interview, and now it's helping me package it up for YouTube.
10:44Multimodality plus one million context window is clutch. Threw some thumbnails at it, and it has the prior context to help me decide the best option. Titles, tags, description, promotional posts, all doable with rich context, not just transcripts. So the point here, which was something that I noticed as well, is that OpenAI has done a very good job in general of undercutting everyone else's announcements with more impressive things of their own. In fact, I could be misremembering this, but I'm pretty sure that Sora came out right around the same time that Gemini 1.5 Pro was first announced, and it just totally sucked all of the oxygen out of the room.
11:13I tend to agree that from my observational point of view, Gemini 1.5 Pro got at least as much buzz as GPT-4's turbo upgrades, which I think has a lot to do with how much people are just waiting now for OpenAI to actually make a big jump forward. I don't believe right now it's that OpenAI has distinctly found themselves behind, although I do think that many people believe that Cloud 3 is the most performant LLM right now, but more just that they've been at parity for so long, something just feels strange. For my money though, the thing that had people the most excited so far this week were reports that came out on Monday that Meta was planning to launch some versions of Llama 3 as early as next week.
11:49This was initially a report from the information. Their source was a Meta employee and said that the company was planning to launch two small versions of Llama 3 as a precursor to the launch of the biggest version which was expected this summer. This was confirmed at an event in London on Tuesday. Said Nick Clegg, Meta's president of Global Affairs, within the next month, actually less, hopefully in a very short period of time, we hope to start rolling out our new suite of Next Generation Foundation models, Llama 3. There will be a number of different models with different capabilities, different versatility, released during the course of this year, starting really very soon.
12:18Said Meta Chief Product Officer Chris Cox, the plan is to power multiple products across Meta with Llama 3. One thing that's shifted with Meta's discussion, all the way from Zuckerberg on down, is that they are no longer content to just be the best open source model. They want to be state-of-the-art competing against any model. Reinforcing that was Joel Pino, the Vice President of AI Research, who said, our goal over time is to make a Llama-powered Meta AI be the most useful assistant in the world. And so between this Meta Llama 3 announcement, as well as the new Mistral model, the open source community had a lot to be excited about this week.
12:50Anyways, guys, we'll wrap there. But what I will say is that part of why this week feels so reflective to me of just where we are is that all of these big announcements, while cool, aren't getting people to jump up and down and scream and say, wow, what a crazy week. Although I'm sure some of the other YouTubers will have titles to that effect. Instead, this feels, if certainly not like a slow week, then certainly still something that is to be expected. In other words, not too many standard deviations away from the mean. So that's the world you're living in now. Five big models coming out a week and everyone just excitedly going on with their life.
13:20Anyways, that is going to do it for today's AI Breakdown. If you haven't yet, check out besuper.ai. It's the fast, fun, and most importantly, useful way to learn AI through video tutorials and companion how-tos. Appreciate you listening or watching as always. And until next time, peace.
From the publisher
This week witnessed the introduction of five notable AI models, highlighting the field's swift advancements. Udio emerged as a standout music generator, rivaling the previously leading Suno. Updates and new releases included Mixtral 8×22B, Google's publicly available Gemini 1.5 Pro with enhanced features, OpenAI's improved GPT-4 Turbo, and anticipated versions of Meta's Llama 3.
**
CHECK OUT THE JUST-LAUNCHED SUPERINTELLIGENT PLATFORM - 300+ AI video tutorials https://besuper.ai/
**
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
