In short
Podcast Summary: Meta’s Joe Spisak on Llama 3.1 405B and the Democratization of Frontier Models
Podcast Overview
- Title: Training Data
- Hosts: Sonya Huang, Pat Grady, and other Sequoia Capital partners
- Description: Focuses on AI conversations with leading builders and researchers to understand evolving technologies and their implications.
Episode Details
- Episode Title: Meta’s Joe Spisak on Llama 3.1 405B and the Democratization of Frontier Models
- Guest: Joe Spisak, Head of Product Management for Generative AI at Meta
- Release Date: Just two days post Llama 3.1 405B launch
- Key Topics:
- Innovations introduced in the Llama 3.1 405B model
- The role of open source in the AI ecosystem
- Future of startups in relation to commoditized AI models
Key Highlights
- Launch of Llama 3.1 405B
- Scale and Training:
- Trained on 15 trillion tokens using 16,000 GPUs.
- Emphasis on zero-shot tool use, distillation, and synthetic data generation.
- Capabilities:
- Multilingual support with improved safety and quality.
- Significant enhancements in tool use, allowing integration with search engines and code interpreters.
- Open Source License
- Transitioned to a more permissive open source license to encourage community adoption.
- Meta's decision to open source was driven by a desire for maximal adoption and community engagement.
- Future of Frontier Models
- Commoditization:
- Joe argues that even frontier models will commoditize, which could benefit the startup ecosystem.
- Emphasizes that innovation and differentiation will shift towards how models are applied rather than just their capabilities.
- Implications for Startups
- New startups are encouraged to leverage existing open-source models rather than building their own from scratch.
- Importance of data control emphasized, as ownership of data and model weights becomes crucial.
- Open Source vs. Monetization
- The open-source decision is seen as both offensive and defensive; it's about fostering an ecosystem rather than direct monetization.
- Meta views its AI contributions as a means to drive innovation and improve its own products, rather than primarily profit.
- Challenges Ahead
- Data Limitations:
- Concerns around a potential wall hit with data availability.
- Importance of synthetic data generation as a potential solution.
- Model Development:
- Discussion on whether model development is becoming akin to software development, focusing on execution and iterative improvements.
- Reasoning and Future Directions
- Importance of reasoning in AI models discussed, with coding and math data identified as significant contributors to enhancing reasoning capabilities.
- Future levers to unlock reasoning include better data, improved benchmarks, and clear applications.
Conclusion
- Joe Spisak emphasizes the collaborative effort behind Llama 3.1 405B and Meta's commitment to open-source innovation. The episode provides insightful perspectives on the future of AI modeling, the evolution of open-source strategies, and the potential for startups in a rapidly changing ecosystem.
Episode Structure
- Introduction: 00:00
- Llama 3.1 405B Launch Discussion: 01:28
- Open Source Licensing: 05:02
- Meta's Perspective and Benefits: 07:01
- Commoditization of Frontier Models: 11:16
- Impact on Startups: 12:41
- Mistral Large 2 Discussion: 19:36
- Comparison of Frontier Strategies: 22:38
- Model Development as Software Development: 26:34
- Agentic Reasoning: 29:09
- Future Levers for Unlocking Reasoning: 31:20
- Small Models Insights: 33:09
- Training Data Scale: 34:08
- Concerns of Hitting a Wall: 37:36
- Lightning Round: 39:49
Additional References
- Llama 3.1 405B Paper: Discusses the technical details and innovations.
- Mark Zuckerberg's Essay: "Open Source AI Is the Way Forward."
- The Bitter Lesson: Insights by Rich Sutton on AI and learning systems.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00If I was a fighter right now, I would absolutely adopt open zores. It forces me, though, to look at the engineering complexion of my org, right? And things like, I'm going to need people doing LOMOPS and things like, you know, fit data fine tuning and how to build Ragn things. And APIs, there's plenty of APIs that allow you to do this, but ultimately you want control. Like your mode is your data, your mode is your interaction with users.
0:42Hi, everyone. Welcome to Training Data. Today we're excited to welcome Joe Spisak, Director of PM for Generous AI and Metta, where he leads Lama and third party ecosystem efforts. Joe spent the last decade in AI, leading product at PyTorch and working on initiatives that span protein folding in AI math, many of which have spun out for meta into their own startups. We're speaking to Joe just two days after the Lama 3 .1 405B launch and we're excited to get his view on questions like, where is the open source ecosystem headed? Will models commoditize even at the frontier? Is model development becoming more like software development?
1:19And what's next in agents and reasoning, small models, data and more? Joe, thank you so much for being here today. We're so excited to have you just two days after the Lama 3 .1 405B launch. It's an incredible gift to the ecosystem. We'd love to learn a little bit more about how you, what specific capabilities do you think the 405B is particularly unique at, especially in comparison to the other state -of -the -art models? Oh, thanks so much for having me. This is so much fun. I haven't done a podcast like this in Saudi pre -ckelby, so it's like fun to be in the room and just like, you know, chatting about this cool stuff.
1:57Yeah, I mean, we're like beyond excited and meta. This is something that I think a lot of us have been working on for such a long time, months and months and months. And, you know, we kind of put out that nice little like appetizer, I'll call it in April of like, about three. And like, I was actually like, we're looking at like, be that excited about these models and like, the response was like, through the roof, like, my God. Like, everyone's excited, but they really don't know what's really And so, like, yeah, kind of hold that, kind of had to hold that back for a while and kind of keep it to ourselves.
2:29And then kind of build up for this launch. And the 4 or 5B is a monster. It's a great model. And I think the biggest thing we've learned about 4 or 5B is just a great, it's like a massive teacher for other models. And we kind of had that plan all along, because when you have a big model, you can use it for improving small models or just distillation. And that's how the eight and seventies became and the great models that they are. I mean, in terms of capabilities, like we listen to the community, we listen obviously to our own product teams, right? Because we got to build products for Meta. And I mean, long context was like one of the biggest things people wanted and we have much longer context internally than what we were released.
3:10But we saw just the use cases start to build up. Multilingual, I mean, we're a global company. So we released more languages, many, many more to come because obviously Meta has billions of people on their platform and hundreds of countries. And so I think that was like to me, those are like table stakes things, but they're like really done well on the models. Like I think we spent a lot of time in post training on our different languages and improving them and safety. Just their really, really high quality. So we don't just like pre -training out like a ton of data and say, look at us from multi -lingual.
3:47You know, we actually did a lot of work in our SFT phase and supervised fine tuning and a lot of safety work. I think one of the coolest things that I'm excited about, well there's just a couple of things I'm excited about, but one is tool use. I think models, oh my god, zero -shot tool use, this is going to be crazy for the community. I'm going to show a few examples, like we can show like calling Wolfram or Brave Search or Google Search and it works really great. But it was that tool use is going to be a game changer. The ability to kind of call code interpreter and actually like run code or kind of build your own kind of plug in for things like Ragn and things like that really be state of the art.
4:28I think it's going to be a really big game changer. And I think just the fact that we released the 4 or 5 itself and we changed our license so you can actually use our data. Like that was a big deal. Like, that was a big discussion. We had many meetings with Mark on that. And we, ultimately, like, landing on a place where, you know, this was like this pain point for the community for so long. Like, these closed models, like I can't use the outputs or maybe I can use them, but maybe I'm using them and slightly inscurpulously or whatever. Like, we actually are encouraging people to do it. I'm sure that was a tough decision to make.
5:04Walk us through the things that you did to consider and actually making that leap to open up license. and life and sing in that way. Yeah, licensing is too permissible. Oh, licensing is like a huge topic in itself. Obviously, you've probably spent a little podcast talking about it. I don't want to, but we could. I think we wanted number one, just to unlock new things. I think we wanted to have the 405 and our Lama 3 .1 models differentiate people and new capabilities. Just like we just looked at what people were really excited about in the community, not only in enterprise and products, but also in the research community because we obviously have a research team and you know we work with academia and we talked to folks I mean you know Percy Langan Stanford checks to be all the time saying you know when you get a release it when you know he said can I use it can I use it and Percy like you know stay state -state -state patient but I think we we heard them and we knew kind of what they wanted and yeah I think ultimately we wanted Lama everywhere.
6:08We wanted just adoption, you know, maximal adoption, really the world using it and building on it and I think Mark even used in his his letter he put out like, you know, the new standard or standardized. So I think like to do that, you kind of have to enable so like that or you kind of have to unblock all these different use cases and really look at what the community wants to do and make sure that that you don't have these kind of artificial barriers. And that's what the discussion really was. And so actually even beyond that, we started working with partners like Nvidia and AWS, and they started building distillation recipes and even synthetic data generation services, which is pretty cool.
6:48I mean, you can start to use those and actually create specialized models from it. And the data that you, I mean, we know how good the data is because we used it in our smaller models. It's really good. And it improves our model significantly. So I want to pull on the open star set a little bit more sure and I've read Zux manifesto it was great, but I'm still I've tried to grab my head around like what what's in it for meta This is a massive investment the open source in some ways you're leaving a lot of money on the table because you now have a state of the art Model that you're offering to everybody for free and so I guess Is this an offensive move is this a defensive move?
7:26What's in it for meta? I mean we've So, well, first of all, our business models doesn't depend on this model to make us money directly. So we're not selling a cloud service. We've never been a cloud company. We've always worked, I would say, with a partner ecosystem. All the way back to the five years, I was helping to lead PyTorch and the ecosystem in the community who built around that. Like, we never built a service we probably could have in some way, but it would have been weird. We saw basically going back to PyTorch, we kind of saw it as this kind of lingua franca, kind of bridge, you know, to this like area of high entropy.
8:06It's kind of weird way to say it, but like there's all this innovation happening, how do we kind of build a bridge to it and actually be able to harness all that innovation. And the way to do that is to be open, and it's to kind of get the world building on your stuff. And I think that's, that ethos is kind of carried over in Talaama. And if you look at PyTorch, that was a huge way for us to pull in. At the time, when we really started working on PyTorch in earnest, computer vision and CNNs and all that, if you remember that, old times now. But we actually would see these architectures come constantly.
8:41The people would, in the right code and they'd publish it in PyTorch. And we'd take it internally. We evaluated people with open source models and put them out on model zoos. and we devaluate them and we'd see just how quickly the community was was improving things and we'd actually leverage that. Especially for like integrity applications, where we're released like heat -full memes this way, these other data sets, we just saw the improvements like week over week, month over month, and it was built on something that is like we were using internally. So it was very easy for us to just take it inside.
9:11So I think like Lava is definitely similar in Evergaard where, you know, when we act academia and when companies start to read TV's models or, you know, try and jailbreak them, we want people to do that to art models. And so we can improve. And I think that's a big reason. And it's like, be careful what you wish for. Of course, but like it's the same with Linux, right? The Linux is open source and the kernels open source and people will, you know, it's much more secure when things are transparent. and bugs can be pushed faster. And so that helps us a lot. I think it's, you know, there's also the angle of, you know, we don't want this to turn into kind of a completely closed environment.
9:57Like I think just like today, if you do, if you look at like Linux and Windows, and like in my opinion, there's room for both, right? There's room for closed, room for open, and people use, depending on what they need and the applications, I think that there's gotta be a world of open models and I think there's gonna be a world of post models and I think that's totally fine. What was the primary argument against open sourcing? Was there one? I mean, there was definitely like competitive concerns we talked through, do you wanna give your technology, put it out there and that. And I think we're less concerned about that because we're moving really fast.
10:37Yeah, like if you look back, I mean, I've been, you know, back, I've been in meta like close to what six or seven years now and like in the last You know, you're you're so we've done you know, we got a connect launch. We release purple llama less December. We released llama three 3 .1 before that we released llama two in July Lama one was like in February so like just if you think about the like the pace incredible. The pace of innovation that's coming out of our team and our our company is like just at a crazy pace right now. So I'm not too worried about it. I don't think we're that worried about it.
11:15I'd like to move into your personal views on the broader ecosystem. I think a lot of the questions that people have center around what happens to the value of all these models, especially as meta open sources, more of them at the state of the art level, with Lama 3 .1 with OpenAI launching GPT -40 mini. What is your view on DeModels commoditized even at the seat of the art frontier? This is a great question. I think if you look at just even the last two weeks, I mean, 40 mini is a really, really good model. Input, I think in -put per million tokens is something like $0 .15, $0 .60 out. So it's incredibly cheap to run, but it's also an excellent model.
12:00Like, it's just like they they've done an incredible job and just distilling and getting to something that's like really really performant yet really really cheap. So I think like, you know, Sam is definitely pushing on that. And then if you look at what we've done in like last week and pushing out, I would say like pretty, pretty compelling to their models across the spectrum. I do think like it's rapidly getting to a place where you know, the model is going to be kind of a commodity. I mean, I think there's this frontier of like data where, you know, I mean, we can certainly gather data from the internet.
12:33We can license data. But at some point there is kind of like some frontier of limitations. I think that we're all going to have. And this goes back to our conversation goes, we can kind of the better lesson of data and scale and, you know, compute is that enough. It's probably not quite enough, but it's like compute and data becomes kind of, if you have enough of both, you know, you can kind of get it. like a first -order approximation of the state of the art without anything else is kind of what we've seen. So I do think the models are commoditizing. I think the value is elsewhere. And I look at Meta and I look at our products and I look at what we're building.
13:08That's honestly where the value is. For us, it's Meta AI, it's our agent. It's all the technology that we're going to put into Instagram and WhatsApp and all of our products where we actually are going to monetize where we're actually going to add real value. the model itself, I think, definitely will keep innovating, new modalities, new languages, new capabilities. That's what research is, right? It's pushing in the frontier, the immersion capabilities, and then we can leverage those in products. But the models are definitely pushing in that direction. If that's the case, and all these existing companies that have massive distribution and wonderful applications that they're already out in the wild, can just adopt the state of the art models, what advice would you give to the whole wave of new startups that are trying to make it out there, either building their own models using other city -yard models and then trying to build applications on top.
14:02Yeah, I mean, there's definitely like some model companies or companies that are building you know, they're training pre -training foundation models and it's expensive. It's like I think we're you know, I can't say how much Lama 3 costs but it was very expensive and you know, Lama 4 is going to be even more expensive. And so I to me given kind of the state of play and things, to me, it doesn't make that much sense if I was a startup to try and go and do a pre -training. I think the Lama models are absolutely incredible as foundations to build on. So I do think if I was a founder right now, I would absolutely adopt open zores.
14:40It forces me, though, to look at the engineering complexion of my org, right? And I'm going to need people doing LOMOPS and things like fit data fine tuning and and how to build Ragn things. And APIs, there's plenty of APIs that allow you to do this, but ultimately you want control. Like your mode is your data, your mode is your interaction with users, and you may want to deploy these things onto a device at some point, and have a mixed interaction or something. You might want to have a similar queries running on your device and have very low latency interactions with your users. You might want to split and have a more cloud -based approach from more complex queries, more complex interactions.
15:25And I think like the open source approach gives you that flexibility. It gives you the ability to modify the models directly on the weights. You can run the weights, you can distill them yourself. There's going to be distillation services that allow you to take your weights, distill them down to something smaller. I think that's pretty awesome. We're just now seeing the beginnings of that. So I think like in my mind, like control matters a lot. And ownership of the weights. There are a lot of API services where you'll do fine tuning your mouth. So you're bringing your own data, you're fine tuning, and then you use something like a low -rate adaptation or Laura.
16:02And unfortunately, you don't actually have access to those lower weights at the end of it. You're forced to use their inference. So you're like, hmm, let's see, I'm kind of like hell hostage here. I've given my data. I don't have access to the actual IP that was generated from that data. And now I have to enforce to use their inference service. That's not a good deal. So I think the open source kind of like brings inherent freedom. I think that approach doesn't so. What do you think of Mr. Large was announced I think maybe a day after Lama 3, if you're following? What do you think of them? And I guess more broadly for everybody at the frontier, is everyone kind of pursuing the same recipes, the same techniques, the same kind of compute scale, data, etc.
16:47run. And so like, you know, everyone's kind of going to be roughly similar at the frontier. Or do you think you guys are doing something very different? So, first of all, I'm going to straw. I mean, an amazing team. It was one of my old teams in Fair. Yeah. They were working on through improving and AI mathematics. So, Guillaume and Tim and the team are very, and they're, they're incredible. Joe was just talking about fun, banter. I mean, I'm less than I can. There, there. So, I mean, this was like one of the scrapiest teams that I've ever worked with. I mean, the team, I don't think ever slept.
17:20So it was like basically by day, they're doing like, pushing the less now. Probably the less now. I mean, they would push the state of art and like AI and Thurham Proving and, you know, during the day, we published some work on that. You know, I think what, a couple of years ago, they now, cheese. And but by night, they were basically scrappily grabbing compute to train Lama 1. And so we were building large language models several years ago in Fair. And that team basically just, they were just really ambitious and they were kind of working by night. And that's really where Lama 1 came from. So the team is great.
17:58I mean, I think they're doing really good work. I think they're definitely challenged in that they're trying to also open source models, but also make money. And models like Foro, Mini are not helping them. Because this is why they changed their license, for example, to have a research -related license, which makes sense. Because they were open -sourcing models, and they immediately, their own ecosystem is competing with them in a lot of ways. because they'll release a model, they'll host it, like use this model, but then they have together and fireworks and lepton and all these companies that provide like sometimes a lower cost, you know, per million token, like offering.
18:40So it's really tough business right now. In terms of like large too, I think it's a really good model. I mean, we just on paper, I haven't evaluated it, we haven't looked it internally yet. I think if you look at like artificial analysis, I think they added up and kind of like though it was a little It was under I think like the 70b model In terms of like quality, but you know, that's like a blended, you know, the blend of bunch of benchmarks to make that Distinction but on paper it looks really good. We're gonna evaluate it I think you know to for me anyway like the more the merrier like the more models are out there the more companies are doing this the better It's not like we're not gonna be the only one.
19:19I think that's good that we're not the only one So, and I think like more generally like the Jenny ice base like you wake up every single day and you kind of expect something like this right you expect to Mollody release or something groundbreaking to happen and that's kind of like the fun of being in it so. Totally, totally. Do you think everyone at the frontier is comparable though like are you all pursuing comparable strategies? Yeah, this is actually a good question because you know if you read the Lama 3 paper which was I think 96 pages you ended up at right lots of citations obviously. sharing.
19:50Lots of sharing, lots of like contributors and core contributors and that. So it was a it was a detailed paper and Lawrence and Angela on the team spearheaded writing that and I think that was like part of the hardest things. Like developing the model was like relatively easy compared to writing the paper. It was it was a lot of work pulling that paper together. I think if you look at a lot of three, you know, there was a lot of I would say innovation that happened, but also we didn't, like we also didn't, I would say take on like a lot of research risk either. So I would say like the the primary things we really did with Lama with the 4 or 5B especially was, was really pushing scale.
20:33I mean it was still, you know, we group career attention for example. So we don't GQA and that improves inference time and you know kind of helps solve the kind of quadratic attention computational challenge. We trained on, or 15 trillion tokens. We did post -training. We used synthetic data, which improved the models. The smaller models quite a bit. We trained on over 16 ,000 GPUs on our training runs, which is something we hadn't done before. It's really, really hard to do that, because GPUs fail. And you know. It's not the table. Yeah. I mean, it's everyone's like, oh, I'm just going to train on 100 ,000 GPUs like, good luck.
21:10Right? You better have a really, really great infrotein, a really great ML system. You better be ready to innovate at that level because it's just non -trivial. Everyone says it's easy or says you can do it, it's non -trivial. So I think like, I almost look at Lambda 3 is very similar to the GBD3 paper. So if you were to talk to like Tom, he was a lead off of Tom Brown, now at Anthropic. And there's a reason why Tom was the first author on that paper is because a lot of the innovation was really scale. It was really like how do I take something that's you know like an architecture and like push it as hard as we can push it.
21:47And that involves like a lot at the MLSIS kind of layer and infer layer and like how do I scale the algorithm. And so I think that was really like the mentality we had with like the Lama 3 and Lama 3 .1. And I mean internally obviously we have great research team, the fair, we have research in our org and we're looking at and lots of different architectures and MOE and other things. And so, you know, so I think we, you know, who knows what Lumberworld will be? We're a lot of candidate architectures and we're looking at it. But it's kind of a trade -off. It's a trade -off between how much risk you take on as like for research and potentially how much reward or, you know, the ceiling of the potential improvements versus just taking something that's relatively known and like pushing scale and getting that to improve even more.
22:34So ultimately this becomes a trade -off. I think this is such an interesting point. I actually also think it makes Salama and Metta quite unique in the strategy it's taking. The words that I like to use yesterday were, is model development becoming more like software development? Yeah. I'm curious to hear if you think, I think, unlike what many of the other labs have been doing on pushing more of the research, you guys have been focused on just executing on strategies that you know work. Do you see that representative of the continuous strategy you think as you extend Lama out 4, 5, 6, 7, 8? And then also how do you think the other research labs and maybe some of the other startups in the ecosystem will react?
23:16Will they kind of switch and veer a little bit more to the strategy that you've been taking? I mean it's a really great question. We don't have all the answers for sure. I think there's definitely like some somewhere in the middle right now is kind of where I see things landing where where we will continue to push, and on execution, we'll continue to push models out. We'll continue, it's one of our products to iteratively improve as well. So we want that AI improving constantly. And so there's definitely a software engineering analog here that's happening where you can imagine something like a Lama train.
23:52And new features and new capabilities get on that train and we have a model release. It's actually it's much easier when you start to componentize the capabilities to like we're doing that with safety right now and You saw in the release we released Pump guard and new long guard and and you can iterate on those components externally. It's great Obviously the core model is much more difficult. I do think You know, we'll start to include or start to kind of push on the research side because as well because I need the architecture like he's going to evolve. I mean, you've seen like, you know, what AI2, for example, is done with their Jamba and their, you know, Mamba and everyone kind of thinks Mamba is like a new architecture that's that could have promise.
24:33I think what's interesting though is like to truly understand like the capabilities of the architecture, you kind of have to push the scale. And I think that's what's missing right now in the ecosystem is, you know, if you look at academia, an academia is like a lot of absolute brilliant people there, but they don't have a lot of access to compute. And that's a problem because they have these great ideas, but they have no way to truly execute them at the level that's needed to really understand. Will this actually scale? Because like the the job of paper and and model was really interesting and the benchmarks are great, but they didn't scale it beyond I think like under 10 billion parameters.
25:08So you're like, okay, what happens when you know we trying this in a hundred so like does it actually do still see those improvements or not and no one really the extra least outside of these labs, knows the answer yet. So I think that's like one challenge. So I think like to me, we're gonna get into this hybrid space of, you know, we are gonna push definitely on architecture with a very, very smart and well -accomplished research team. But we also are gonna be like, you know, we are gonna be executing. And I think that's when we start to get like a recipe, you know, we're gonna push it to the limits.
25:40And, you know, we are gonna start, you know, release, we're gonna continue to release more models on it. But in parallel to that, we have to push on architecture. Yeah. And I think it just makes sense because the next breakthrough, you know, at some point you're going to reach like a kind of a theoretical limit, and you need to evolve the architecture. All right, so I see kind of a little bit of an in -between. And obviously, we're really good at execution. I think we're a pretty good execution. But we're also good at research, and we just need to marry those two, so it makes sense. Because like research and products are very different, right?
Read the full transcript
26:10like one is should be pretty deterministic the product side and one is inherently non -deterministic right it's like is this gonna work I don't know it's a really big bet if it fails it's research like it should have a like a non -zero chance of completely blowing up in our face we just need to go in another direction but that's that's what research is so I'm curious about one branch of where a lot of I think model research is happening right now agentic reasoning And you all have announced really great results in reasoning. I'm curious, maybe at a very basic level, how do you define reasoning?
26:46And then, are you all seeing reasoning fall out of kind of scale during pre -training? Are you, is it post -training? Is there a lot of work left to do on the reasoning side? Yeah, reasoning's a bit of a loaded area. I mean, you could argue it's, you know, things like multi -stop, like, you know, and I think the best, unfortunately, the best examples we have are like the, you know, kind of like the sort of semi -gimmi -key, you know, you know, Bob is driving the bus and like he picks, you know, like those kind of like things, right? And if you troll local Lama, you'll see a billion of those, right?
27:18So, but those actually forced the model to take multiple steps to respond to you and think through and logically kind of respond. I think coding is actually really like, you know, when you look at like pre -training. And so I like to answer your question directly, like, reasoning improvements come in both post -training and pre -training. So, what we've learned, which is now like everyone's like, oh, of course, this is the case. But definitely like the last year or so, everyone's kind of learned that, you know, code, having a lot of code in your kind of pre -training corpus really improves reasoning.
27:51But that's what you'd think about it, like, of course, duh. It's step -by -step, it's very logical, it's, you know, code is very, it's just logical by nature and kind of step -by -stop. and if you incorporate a lot of that in your breach training, your model will reason better. And then we of course look at examples in post training and like super of SFT to improve as well. So we look at the breach train model and it depends on how you balance things as well. Because you can balance how well your model reasons with how well it responds and different languages, like ultimately in post -training, like everything's a little bit of a trade off.
28:31Like you can super optimize things for coding if you want to. And we did that with CodeLama. It was really great, but of course the model will suffer like in other areas. And so it ultimately becomes what, like we kind of like Pareto Frontier, like capabilities we wanna, like bring out, if it's a general model. And I think like, yeah, I mean, ultimately it's a trade off. So anyone can kind of pick a benchmark or some some capability and say, I'm gonna super optimize for it and say, By the way, I'm better than GPD4. Well, great. Anyway, I can do that. But as your model is generally capable of GPD4 or a lot more through point one or whatever, I think is a different story.
29:07What do you think are the future levers to unlock reasoning for anyone going forward? I mean, the obvious answer is data. I mean, the more data, the more code and supervised data that you can get, I think is a natural answer.
29:30I mean, I think we need to find applications as well for how we like define it and that would help us like once you've kind of start finding like those kind of killer applications, then you can like, then you kind of know where to kind of focus in terms of your other your gate at exactly what you're solving for like, and this goes back to like evals and like what is what is your e -mail because we're starting to saturate e -vails. And so we tend to as a community like we define a a benchmark or a metric and we just like optimize it a little bit how lot of it. And it's great, but then you actually look at the model in an actual environment and you're like, oh, well, that model has a better MLU score.
30:09Great, but like, how does it actually respond? Well, it doesn't respond as well, but it has a better MLU score. And so I think we need better evils and better benchmarks that allow us to, you know, I would say like find clear line of site to actual interactions. And I think like, you know, the live, what is it called? The Advocates benchmark, the live bench, I think it's called. I can't really give it. It's pretty good. I was looking at that. And of course, like, Alem says and Chef Adoreena, like these are more natural, even though, you know, it's still not perfect. But it's like moving in the right direction of things are like more human like interactions versus like a static data set.
30:52that is not that helpful. So I think once we start to find these other, what reasoning use cases make sense, we're gonna start to generate more data, and you're gonna start to improve the model there, and hopefully that has again, line of sight to a benchmark range e -vow that actually feels like it improves the end product. And a lot of this actually depends on the end product, of course, what is my application? Yeah.
31:20within large research labs, coding and math, have always been two primary categories and trying to unlock reasoning. In the start of Beacos, as in now we're seeing more folks who really want to go from the math angle. Do you have a perspective on whether or not that has led to interesting unlongs? I mean, the answer is yeah, I mean, I think we, if you look at our data, or at least our models, we've like coding and math have been, I would say the primary lovers. So, I mean, I think that's like having more obviously is better, because obviously math is also very logical and very like stepwise. So obviously you can see the pattern here.
32:00The more data you have, like that kind of follows that sort of pattern, the more your model is going to be able to reason. And you can see that in how actually models respond. Like if they start and you ask them to like respond and like step me through your thinking process, right? And it will actually do that. and some models do better than others. So anything like that, I think scientific papers. Also, we had some projects out of fair that trained on archive papers. And you can see, not only is code and math, pure mathematics, but also scientific papers, science is a very logical and how they write things and how they step wise and how they create images of their charts and stuff.
32:45And that also, I think we've seen just general scientific information helps as well. So Galactica was our project. So Robin Ross from the Peabirds and Co team led that still in my opinion, one of the coolest projects ever. They got a lot of bad press, but wow, they were ahead of their time, in my opinion. I'd love to talk a little bit about small models. given the scale capital and the compute that many startups have, the 8B and 70B models are an incredible gift to the ecosystem. And it's funny that you called them appetizers at the start because I think they're super powerful for that set. But they're also really powerful for a number of different applications where you want smaller models.
33:31And so I'm curious to hear, what do you hope to see developers use the 8B and 70B models for? give that they are best in class for their size of model. So it's interesting though, when we released, we released April, Alama 3, we released an 8 and a 70, the appetizers as we call them. The 8P was actually better than the Alama 270B by Leap's. So we were, I had to look at the chart, and it was like, is this right? Is that really the case? And we're like, yeah, it really is. I mean, was that much better? What's the intuition for how that happens? I mean, it was more data. We had, with sub and x more data.
34:14Obviously, we put a bunch more compute as well. So, you know, going back to like computing data, you know, being, you know, we're pushing on those. So I think like, we just, you know, we saw like just like, it's almost like every generation, which is again, the generations are accelerating. you start to see, you know, the benchmarks for like a large model basically get like, you know, pushed down into the smaller like size regime. And so, you know, 70 becomes an eight and, you know, like we have internally we have models where the eight is, you know, like I'm much even smaller than eight actually we're starting to see like really nice benchmarks on even smaller models.
34:53So you continue to see like in that, you know, that the models improve at smaller scale. And that I think is just we're pushing the architecture, we're pushing, pushing scale, and we're starting, we haven't quite saturated yet, and I think that's really interesting. So for me, one of the biggest reasons that I think it, like a small architecture, is useful is obviously on device. Everyone loves to talk about on device, and Apple is talking about that, and Google has Gemma models and Gemma and I running and Android devices. So I think like on device makes sense. I think safety is kind of interesting because one of the things we have our own internal versions of Lama Guard which we used that are orchestrated for our applications internally at the company in Meta and today they're built on an 8B model which is kind of expensive to run if you think about a safety model that's kind of like the secondary model and so I do think internally we've been experimenting with much smaller models in that in that regard.
35:54And it creates efficiency over latency. So, because really, those models are really just classifiers. You know, they're not really autorescitives like chat, like interfaces. They really just classify like an input up -prom to, you know, does that violate, you know, this whatever category in the taxonomy in the output, the model when it generates does it does violate that kind of stuff. So you can actually push those even further. I think that there's also really interesting cases though for on device where you almost have, when you think about privacy and think about data, you want to have your data stay on device, you can think about a rag architecture on device.
36:36So you have data, even your chat history that's on WhatsApp or other things, you can imagine that model having access to data aggregating it and then running some type of almost like a mini -vactory database where you're reusing Ragn and doing your kind of fuzzy search, a fuzzy, or fuzzy matching, and with your small model, and that becomes its own system in itself. And you can basically do things like local summarization. I don't know, I get so many text messages. Hey, summarize my last 15 messages, please, because I've been in meetings and I haven't looked at my phone. And that's super useful, and that I don't have to send data up to the cloud or anywhere else.
37:16So there's like those kind of use cases, I think they were small models actually are going to be really compelling. And then for like super complex queries and things obviously like you have a big model in the cloud that can always service those. But for like many things, I think like on device or even in the edge and on -prem, these small models actually can do pretty good. You talk to kind of scaling up computes and data as the two fundamental vectors to improve performance. I guess there's been a lot of chatter about how we are going to hit a wall or maybe we're not going to hit a wall on data and maybe synthetic data is the answer, et cetera.
37:49I'm curious your perspective on that. Is there an impending wall that we're going to hit most likely of cheap, accessible data? What do you think, how do we scale beyond that? I mean, I think we've shown with this release that synthetic data does help a lot. I mean, I think we've, you know, in pre -training, we train on 15 -trained tokens or you were take. And in post -training, we generated a ton of millions of annotated synthetic data, a lot of it generated by the 4FIB. We obviously paid for annotations as well. I do think synthetic data is like a potential path forward. I think it's going to, like we know now, and the kind of proof is in the models, right?
38:33It's like great to talk about it in that. I do think like, you know, data is going to be a challenge at some point for us. since this is why I think companies are licensing a lot of data. These days, they get access, made open ads licensing data, were licensing certainly data. I think having access to services that generate data, to improve models is important. So I think that inherently is an advantage for a lot of companies. I mean, Google has YouTube, right? They can, I'm sure, is a value to them. So, which kind of implies that bigger companies have an advantage, which is not something that's anything new, right?
39:12We've been talking about this for a long time. In terms of a data wall, I don't know. I mean, we're not there yet. I would say, let's talk another, just do those, just do those, schedule this for a long year. And let's see where we are next year. You know, I'll see my calendar for one year exactly from now and met AI. But you know, let's talk in a year and see where we are. But we haven't hit it yet and we're still scaling and we're still you know, we're still gathering a lot of data and we're generating data and our models are still like continuing to improve so yeah. Let's close it out with some rapid fire questions.
39:47Sure. Sounds great. And what year do you think we'll surpass the 50 % threshold on sweet bench. I good question. And if I've learned anything, it'll be faster than whatever answer I give you. Because I think Eddie benchmarked, well, zero in on it, people are going to go and print out. So I don't have an answer. But it'll be fast. You know, one of the questions we have been asking people is in what year will an open source model surpass the other companies on the front, the other models on the frontier. And we have to take out that question now. Thanks to you all. This is. I mean. It's true. We're almost there.
40:27I mean, I think four of us is incredible. It's definitely in that class, which is incredible. Well, men always open source. Lama. I mean, I think Mark's pretty committed. I use a newsletter. I mean, we've open sourced for years and years now, back to pie George to fair, to Lama models. I mean, this isn't something that's left in the pan for the company. The company's been committed to open and source for a long time. So I would never say never, but like, I mean, the company in Marker really committed. Amazing, Joe. Thank you so much for being here today. And also for all the work that you're giving to the entire ecosystem, I think the entire AI community has very much grateful for all the work that you've done with pushing out Lama and the advancements to come.
41:11It's a huge team. Check out the paper. Look at all the, all the, all the intelligence. I've just spent all of yesterday reading it. We need like the Star Wars, like scrolling text of all the contributors because it was an incredibly big team. I was thinking about that. Yeah. So my head's off to the team. This was a total, I mean, this is absolutely took a village to get the Lama out there. And so proud and excited to represent the team here. So thank you. Thank you.
From the publisher
As head of Product Management for Generative AI at Meta, Joe Spisak leads the team behind Llama, which just released the new 3.1 405B model. We spoke with Joe just two days after the model’s release to ask what’s new, what it enables, and how Meta sees the role of open source in the AI ecosystem.
Joe shares that where Llama 3.1 405B really focused is on pushing scale (it was trained on 15 trillion tokens using 16,000 GPUs) and he’s excited about the zero-shot tool use it will enable, as well as its role in distillation and generating synthetic data to teach smaller models. He tells us why he thinks even frontier models will ultimately commoditize—and why that’s a good thing for the startup ecosystem.
Hosted by: Stephanie Zhan and Sonya Huang, Sequoia Capital
Mentioned in this episode:
Llama 3.1 405B paper
Open Source AI Is the Way Forward: Mark Zuckerberg essay released with Llama 3.1.
Mistral Large 2
The Bitter Lesson by Rich Sutton
00:00 Introduction
01:28 The Llama 3.1 405B launch
05:02 The open source license
07:01 What's in it for Meta?
10:19 Why not open source?
11:16 Will frontier models commoditize?
12:41 What about startups?
16:29 The Mistral team
19:36 Are all frontier strategies comparable?
22:38 Is model development becoming more like software development?
26:34 Agentic reasoning
29:09 What future levers will unlock reasoning?
31:20 Will coding and math lead to unlocks?
33:09 Small models
34:08 7X more data
37:36 Are we going to hit a wall?
39:49 Lightning round




