In short
```markdown
Podcast Episode Summary
Llama 3 Is Here (And Seemingly Better Than Expected)
Podcast Details
- Title: The AI Daily Brief (Formerly The AI Breakdown)
- Episode Title: Llama 3 Is Here (And Seemingly Better Than Expected)
- Description: Meta has released its Llama 3 models, enhancing the landscape of open-source large language models (LLMs) with advanced capabilities and integration into Meta's consumer products.
---
Key Takeaways
Overview of Llama 3 Release
- Announcement: Meta has officially launched its latest open-source LLMs: Llama 3, with initial models of 8B and 70B parameters.
- Capabilities: The models promise improved reasoning and performance benchmarks, indicating Meta's commitment to leading in AI innovation.
- Open-source Focus: Meta's push to enhance accessibility and performance in the AI community through open-source models.
Competing Technologies
- Microsoft's VASA 1 Model:
- A new AI-generated video technology that creates hyper-realistic talking face videos from a single photo and audio input.
- Concerns raised about the potential misuse of this technology for misinformation.
- OpenAI's Assistance API Update:
- New features allow for access to up to 10,000 documents, improving enterprise solutions and competition in the RAG (Retrieval Augmented Generation) space.
Entertainment Sector Implications
- CAA Vault Program:
- The Creative Artists Agency (CAA) is exploring a program allowing top talent to create digital doubles for profit.
- Discussion on the ethical implications of AI in entertainment and its potential to devalue human involvement versus creating new opportunities.
Llama 3's Technical Specifications
- Model Sizes & Performance:
- Llama 3 features 8B and 70B parameters, trained on 15 trillion tokens.
- Achieves a context length of 8K tokens.
- Initial benchmarks show promise, with Llama 3 outperforming some competitors in specific tasks.
- Community Engagement:
- Meta's significant community outreach, including CEO Mark Zuckerberg's appearances on smaller, creator-focused platforms to discuss Llama 3 and its implications.
Industry Impact
- The launch of Llama 3 is seen as a turning point in the open-source AI landscape, potentially rivaling established models like GPT-4.
- Anticipation of further developments, including larger models that could enhance capabilities beyond current benchmarks.
Future Prospects
- Upcoming Models: A larger 400B model is still in training and may soon provide comparable or superior capabilities to top-tier models.
- Community Expectations: The open-source community is excited about the potentials for innovation and development stemming from Llama 3's release.
---
Conclusion The release of Llama 3 marks a significant advancement in open-source AI technology, with implications for various industries, including entertainment and enterprise solutions. As Meta continues to push the boundaries of AI, the community eagerly anticipates the effects of these developments on the market and the ethical considerations that arise.
---
Additional Notes
- For more details on the episode, consider subscribing to the AI Breakdown newsletter and YouTube channel.
- Consensus 2024 will feature discussions on AI-driven transformation alongside cryptocurrency and blockchain innovations.
```
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the AI Breakdown, the anticipated Llama 3 release has just happened. Before that on the brief, Microsoft's VASA 1 model is turning heads. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our YouTube, our Discord, and our newsletter.
0:24Welcome back to the AI Breakdown Brief, all the AI headline news you need in around five minutes. Yesterday, everyone started talking about a new model from Microsoft Research called Vasa 1. Bindu Reddy summed it up this way. She called it the first AI-generated video that looks super real and said it takes a single portrait photo and speech audio and produces a hyper-realistic talking face video with precise lip sync audio, lifelike facial behavior, and naturalistic head movements generated in real time. This is amazing given that the AI-generated video looks very real. So for those of you who are watching this rather than listening to it, let's check out a quick example.
1:00You ever had, maybe you're in that place right now where you want to turn your life around, and you know somewhere deep in your soul there could be some decision. It really is incredibly lifelike. And indeed, that's the way that they describe it in their research paper. The paper is called VASA-1, Lifelike Audio-Driven Talking Faces Generated in Real Time. TLDR, they write,
1:32In the abstract, they write, The model is capable of not only producing lip movements that are exquisitely synchronized with the audio, but also capturing a large spectrum of facial nuances and natural head motions that contribute to the perception of authenticity and liveness. The core innovations, they say, include a holistic facial dynamics and head movement generation model that works in a face-latent space, and the development of such an expressive and disentangled face latent space using videos. Part of what they're so excited about is that the model not only delivers realistic lifelike videos, but supports, as they write, 512x512 videos at up to 40 frames per second with negligible starting latency.
2:07What that means is that conceivably we're not too far away not only from lifelike AI-generated videos, but from actually being able to interact with these lifelike avatars in real time. Now, in addition to producing just lifelike videos from real photos, they say that their method also has the ability to handle photo and audio inputs that are out of their training distribution. For example, they can handle illustrations or paintings, singing audio or non-English speech, none of which were present in the training set. Again, for those of you watching, here's an example of that lack of latency, which makes it opportune for real time.
2:40But you know what I decided to do? I decided to focus all my attention, all my time on listening. So instead of doing something else, I just listened, listened, and listened. Because I'm a true believer that if you're really bad at something... Now the paper points out that this is not a technology without risk. The closer we get to truly lifelike, and especially real-time generation like this, the greater the chance that it's misused for disinformation or misinformation, for impersonating people without their permission. And at this time, they write, we have no plans to release an online demo, API, product, additional implementation details, or any related offerings until we are certain that the technology will be used responsibly and in accordance with proper regulations.
3:26And ultimately, the biggest caveat or asterisk on any of this is that none of us have actually gotten a chance to play with it. This is just a paper with cherry-picked examples that, while super impressive, theoretically might not represent the standard output that you would get if everyone was allowed to use this. Still, it does seem like a fairly significant jump in this sort of avatar capacity, and so is very much something to be watching for in the future. Next up, we get an update from OpenAI. The company has released a new version of its assistance API. Now, the assistance API is OpenAI's tool that allows people to build agent-like assistants that have some specific purpose that they use ChatGPT for.
4:03One of the big changes is, as Brian Romley writes, The new OpenAI Assistance API can access up to 10 ,000 documents in a vector database rag. It's quite useful, and I have already mentioned it to clients that need larger-scale rags. Sully Omar from Cognosis says, New OpenAI Assistant updates are really good. Hard to bet against them. Tested retrieval, and it's insanely fast and almost instant. Seems like they want a piece of the rag pie as well. At what point does it become cheaper to use them versus building your own? For those of you drowning in acronyms, RAG stands for Retrieval Augmented Generation.
4:33It basically refers to the idea of an LLM pulling from a specific set of information for some specific purpose. This is a popular strategy right now for, for example, enterprises that want some version of an LLM to be able to pull from their proprietary information in order to give people in the company insight that is specific to the company. There are tons and tons of solutions for how an enterprise might go about that. And Sully's pointing out that the better that OpenAI and ChatGPT get at it, the more incentive the companies have just to stay in that ecosystem. Of course, ultimately, it's not just a question of technical capacity, but a question of data trust.
5:06And my perception is, at least, that the biggest reason that enterprises are choosing not OpenAI to spin up their own models, to customize open source models, is because of those concerns around privacy. Still, like Sully says, the fact that this is an easy and fast approach that works really well could shift the balance of that conversation a little bit. Lastly today, an interesting one from the world of entertainment. There's been a lot of discourse in the entertainment sector and in Hollywood about artificial intelligence. Last year, the writer's strike and the SAG strike were some of the first instances in which concerns around AI really started to get into the mainstream.
5:39My question at the time was not so much if the concerns were real or not, but whether, if the choice was on the one hand try to ban or prohibit all this technology, or on the other try to profit from it, how long would it take before Hollywood shifted over to the try to profit from its side? Apparently, the creative artist agency CAA, one of the big Hollywood talent agencies, is thinking in a similar way and testing a new program called CAA Vault that allows the talent they represent, or at least a small handful of A-listers that are testing it, to create a digital double of themselves that they can profit from.
6:11Said Alexandra Shannon, CAA's Head of Corporate Strategic Development, On one hand, there's concerns that technology is being misused to exploit name, image, likeness, voice, body of work without consent. But we also recognize that it's creating opportunities for talent and an explosion of creativity in so many different ways. Shannon said, These technologies should not devalue the human. If somebody's digital likeness is being used in a campaign instead of them in person, it is still the value of that person and what they stand for as a representative of a brand. Of course, the concern that many have is that if the marginal cost of production comes down, and there is a much larger supply of celebrity spokespeople in the form of these digital doubles, what will that actually do to the value of any individual instance of that?
6:49It seems like a pretty price deflationary force. Of course, the Brad Pitt brand, for example, is going to continue to retain value, but if there are digital Brad Pits running around everywhere, how much each individual instance of that is going to be able to charge? Ultimately, those are questions for CAA, not for us, And so that is going to do it for the AI Breakdown Brief. Next up, the main AI breakdown. Today's podcast is brought to you by Plum. Is your product team struggling to keep up with the incredible pace of AI development? Are you tired of spending countless engineering hours just to test out small prompt changes in your product?
7:21Thankfully, there's Plum. Build cutting-edge AI experiences for your users in a fraction of the time. Say goodbye to the slow, tedious process of hand coding and hello to the future of AI development. Get ahead of your competition and start moving as fast as AI does. Check out useplum.com and shoot me a message to get early access. Attention, AI Breakdown listeners. Consensus 2024 marks the 10th gathering for all things crypto, blockchain, and Web3. However, importantly, this year's agenda will also dive deep into AI-driven transformation. And the speaker lineup includes the leading minds and innovators at the forefront of this digital renaissance.
7:57Don't miss the Consensus AI Summit to cut through the hype to find where true transformation and opportunity lie. Listeners to this show can get 15 % off registration with the code AIBreakdown. Visit Consensus2024.coindesk.com to learn more. Some of the folks who will be at Consensus this year include Guillaume Verdun, aka Beth Jezos, founder and CEO of Xtropic, as well as spiritual leader of the accelerationist movement, Neil Stephenson, co-founder of Lamina One, and Brendan Eich, the CEO of Brave Software. Again, go to consensus2024.coindesk.com to learn more and get 15 % off registration with the code AIBreakdown.
8:31Welcome back to the AI Breakdown. Last week, reports came out that Meta was on the verge of releasing its latest open-source LLM models, Llama 3. Now, at first, this was an unconfirmed rumor, but then within a couple days, Meta seemed to indicate that yes, this was coming, or at least a set of small versions of their next model were coming in advance of the largest version, which would be coming out over the summer. All week then, we have been waiting with bated breath to see what meta would put out, with tons and tons of speculation on just how good it would be. A couple hours before I recorded this episode, we started to get hints that Llama 3 was about to be dropped.
9:05First, we saw Llama 3.8b instruct be listed on the Azure Marketplace with little nuggets of information like this line. The fine-tuned versions are optimized for dialogue use cases. And then on Replicate.com, we saw pricing for four different models. Lama 370B, Lama 3 8B, Lama 370B chat, and Lama 3 8B chat. People got in their last polls on how good others anticipated this to be. Jan Peleg wrote, Final moments to guess. Lama 3 will be, and then gave the options, State-of-the-art open source software, Same level as state-of-the-art open source software, Same level as state-of-the-art closed software, Or outperforming state-of-the-art closed software.
9:41Outperforming state-of-the-art closed was the lowest ranked option with 8.8%, Same level as state-of-the-art closed had 12.8%, and then state-of-the-art open source and same level as state-of-the-art open source were almost exactly the same, getting 38.9 % and 39.6 % of the vote, respectively. Accelerate Harder asked a simpler version of the same question. Llama 3 will be either amazing or a letdown, with 47.2 % saying a letdown and 52.8 % saying it would be amazing. Just a few minutes later, we got the actual announcement. The AI at Meta account on X wrote, Introducing Meta Llama 3, the most capable openly available LLM to date.
10:16Today we're releasing 8B and 70B models that deliver on new capabilities such as improved reasoning and a set of new state-of-the-art for models of their size. Today's release includes the first two Llama 3 models. In the coming months, we expect to introduce new capabilities, longer context windows, additional model sizes, and enhanced performance. Plus, Llama 3 research paper for the community to learn from our work. Chief AI scientist at Meta, Jan LeCun, gave a few more bits of information. He said 8B and 70B models available today, 8K context length, trained with 15 trillion tokens on a custom-built 24K GPU cluster, great performance on various benchmarks, with Llama 3.8b doing better than Llama 2.70b in some cases.
10:54We'll come back to that 8K context length because it sticks out kind of like a sore thumb relative to other models we've had recently. But as he has started to do, Zuckerberg took to their own networks to talk about the new release. He wrote on Instagram, Big AI news today. We're releasing the new version of Meta AI, our assistant that you can ask any question across our apps and glasses. Our goal is to build the world's leading AI. We're upgrading Meta AI with our new state-of-the-art Lama 3 AI model, which we're open-sourcing. With this new model, we believe Meta AI is now the most intelligent AI assistant that you can freely use.
11:26We're making Meta AI easier to use by integrating it into the search boxes at the top of WhatsApp, Instagram, Facebook, and Messenger. We also built a website, meta.ai, for you to use on web. We also built some unique creation features, like the ability to animate photos. Meta AI now generates high-quality images so fast that it creates and updates them in real time as you're typing. It'll also generate a playback video of your creation process. Enjoy Meta AI and you can follow our new Meta.ai Instagram for more updates. So a couple things that are notable from this. One, Meta is continuing to reinforce their message of open source.
11:59And that's something we'll see even more in some other parts of the announcement in just a minute. Second, as had been intimated, Zuckerberg is clearly not content with just being the state-of-the-art for open source, they want to go after the state-of-the-art in general. You can see the sort of big language they're using. Our goal is to build the world's leading AI. With this new model, we believe Meta AI is now the most intelligent AI assistant that you can freely use. Even if there are little caveats in there, i.e. you can freely use, the ambition to be the best, full stop, is clearly on display.
12:27Another really interesting thing, however, from this announcement is the extent to which they are integrating this into products right out of the gate. This is clearly not meant to just be a developer release, but something that is immediately impacting consumer products as of today. Alongside the announcement, they put out a lot of new information. Some of it's for developers, but some of it, of course, is benchmarks. Almost immediately, people's eyes started bugging out at those benchmarks. Matt Schumer writes, holy S, Llama 370B cleanly beats Claude 3 Sonnet, small enough to host its scale without breaking the bank.
12:59What he's referring to is both the MMLU, where MetaLlama370b claims an 82 versus Cloud3Sonnet79, and HumanEval, where MetaLlama3 claims an 81.7 versus Cloud3Sonnet73. Bindu Reddy from Abacus writes, historic moment. Lama370b numbers are insane. At 82 MMLU, it's far and away the best open source model. GSM8K, Math, and HumanEval are mind-blowing as well. The open source community is definitely going to beat GPT-4 in a matter of weeks. Schumer also pointed out that, quote, Lama 370B cleanly beats Mixtril 8x22B. We've talked a lot on this show about the extent to which Mixtril had stolen some of the open-source thunder from Meta over the course of the end of last year, and this certainly seems to be Meta's clapback.
13:42Overall, Professor Ethan Mollick writes, Meta released their open-source AI Lama 3 today. As a key leader in LLMs, their models are often the most advanced open-source ones out there. Based on benchmarks, the current model is not quite GPT-4 class, but their larger ones, still training, will reach GPT-4 level. And indeed, this is what some people are the most excited about. Schumer again writes, The craziest Lama 3 reveal. The 400B-plus version of the model is on par with Cloud 3 Opus, and it's still training. Soon we'll have a better-than-Opus fully open-source model. The implications are huge.
14:13Meta 3 400B's reported MMLU score is right on par with Opus 3, with their grade school math score and human eval score right around there as well. Aston Zhang from the Meta team wrote, Lama 3 has been my focus since joining the Lama team last summer. Together, we've been tackling challenges across pre-training in human data, pre-training scaling, long context, post-training, and evaluations. It's been a rigorous yet thrilling journey. He continues, scaling is the recipe, demanding more than better scaling laws and infrastructure, e.g. managing high effective training time across 16k GPUs requires innovative strategies.
14:43Still, I think the big implications are summed up by Dr. Jim Phan from NVIDIA who writes, The upcoming Lama 3 400b will mark the watershed moment that the community gains open-weight access to a GPT-4 class model. It will change the calculus for many research efforts in grassroots startups. LAMA-3400B is still training and will hopefully get even better in the next few months. There is so much research potential that can be unlocked with such a powerful backbone. Expect a surge in builder energy across the ecosystem. Now, I think that's absolutely true. In many ways, the story of 2024 so far has been this standardization of the GPT-4 class models.
15:17And so it's perhaps not surprising that open source seems to be catching up there as well. However, surprising or not, as Jim points out, the implications in terms of what people can build and how and for how much are fairly significant. Now, I mentioned that to the extent that there has been any quick critique, it's that 8K context window. Some pointed out that they thought that the community would extend that pretty quickly, and others inside Meta explained it as well. Aston Zhang again wrote, We've set the pre-training context window to 8K tokens, A comprehensive approach to data modeling parallelism, inference, and evaluations would be interesting.
15:49More updates on longer contexts later. In other words, there were clearly trade-offs that they were willing to make based on their goals with this release. Now, speaking of this release, one other really interesting little detail. Meta seems to have taken it to heart to go directly to the community. Because in addition to all the traditional PR strategies, Zuckerberg popped up on a number of creator shows that are beloved and well-known inside AI circles, but not even close to the size of mainstream outlets. Roberto Nixon, who runs some of the most popular Instagram and TikTok channels on AI and future technology, did an interview with Zuck.
16:21And Dwarkesh, whose podcast has quickly become, I think, the most high-value interview show that exists, released an hour and 20-minute episode with Zuckerberg that gets deep not only on Llama 3, but, as Dwarkesh writes, open sourcing towards AGI, custom silicon, synthetic data and energy constraints on scaling, along with, you know, intelligence explosions, bioweapons,$10 billion models, and much more. Overall, I would say the first impressions are extremely exciting. I think even more exciting than the community thought they were going to be when they heard that a couple of small models were coming last week.
16:51I'm sure in the coming days, I will be able to add more context around what people are finding in terms of actual performance. But for now, it's a cool day with lots to explore. That is going to do it for today's AI Breakdown. Until next time, peace.
17:13Thank you.
From the publisher
Meta has released the first of its Llama 3 models, enhancing the landscape of open-source large language models. These models, including the initial 8B and 70B versions, promise advanced capabilities and integration into Meta's consumer products, solidifying Meta's commitment to leading in AI innovation. The release indicates Meta's strategic push to compete in the open-source domain and extend these high-level AI functionalities through Meta AI across its platforms, setting a new standard for accessibility and performance in the AI community.
**
CHECK OUT THE JUST-LAUNCHED SUPERINTELLIGENT PLATFORM - 300+ AI video tutorials https://besuper.ai/
Consensus 2024 is happening May 29-31 in Austin, Texas. This year marks the tenth annual Consensus, making it the largest and longest-running event dedicated to all sides of crypto, blockchain and Web3. Use code AIBREAKDOWN to get 15% off your pass at https://go.coindesk.com/43SWugo
**
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
