Soumith Chintala: Meta’s AI Strategy, PyTorch, and Llama

27 Jun 2024 · 36 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Generative Now Podcast Episode Summary

Episode Title

Soumith Chintala: Meta’s AI Strategy, PyTorch, and Llama

Host

Michael Mignano, Partner at Lightspeed

Guest

Soumith Chintala, Vice President and Fellow at Meta, Co-Creator of PyTorch

---

Episode Overview In this episode of *Generative Now*, host Michael Mignano engages in a comprehensive discussion with Soumith Chintala, exploring significant themes related to artificial intelligence (AI) and its implications for the future. The conversation covers the development of PyTorch, Meta's open-source strategies, and insights into AI models like Llama, along with broader implications of AI on society.

---

Key Chapters and Discussions

  1. Introduction (00:00)
  2. Michael introduces Soumith Chintala and highlights his pivotal role in PyTorch and his experience at Meta.
  1. The Creation of PyTorch (01:05)
  2. Background: Soumith discusses his journey in AI, starting in 2009 and his involvement in open-source projects.
  3. Inception of PyTorch: Initially built to address limitations of existing tools, PyTorch was created out of necessity for a more efficient library.
  1. Building a Community Around PyTorch (04:08)
  2. Feedback Loop: Emphasizes the importance of community and user feedback in the growth of PyTorch.
  3. Amplifying Use Cases: Soumith highlights how showcasing projects developed using PyTorch helped build a supportive community around the library.
  1. PyTorch Explained (05:32)
  2. Definition: PyTorch is described as a machine learning library for building neural networks efficiently.
  3. Impact: It is widely used across various industries, including healthcare and automotive sectors.
  1. Meta's Open Source Strategy and Llama (08:40)
  2. Open Source Culture: Soumith speaks about Meta’s commitment to open-source projects, viewing it as beneficial for both the company and the broader AI community.
  3. Llama Project: Discusses the launch strategy for Llama, emphasizing the natural fit for Meta’s objectives.
  1. Implications of AI (14:16)
  2. Generative AI Models: These models are seen as valuable but require careful consideration of their practical applications.
  3. Cultural Sensitivity and Personalization: The need for AI models to adapt to diverse cultural contexts is highlighted.
  1. Open Source vs. Closed Source AI (19:45)
  2. Market Dynamics: Soumith shares insights on the competition between open and closed source AI solutions, suggesting a potential divide based on user needs and cultural considerations.
  1. Audience Q&A (23:45)
  2. Key Questions: The audience engages with Soumith, asking about potential lock-in mechanisms for foundational models and the future of generative AI.
  1. Closing Thoughts (35:03)
  2. Soumith reflects on the evolving landscape of AI, emphasizing the necessity for adaptability and acknowledgment of both social and legal implications surrounding AI advancements.

---

Key Takeaways

  • PyTorch's Growth: Emerged from a need for better tools and flourished through community engagement and feedback.
  • Meta's Open Source Commitment: Integral to its strategy, fostering innovation and collaboration in AI development.
  • Generative AI's Future: While promising, it necessitates considerations of personalization, cultural sensitivity, and practical utility.
  • Market Landscape: A potential divergence between open-source and closed-source AI, with advantages for democratization and adaptability in the former.

---

Final Thoughts The episode provides invaluable insights into the world of AI from a leading figure at Meta. Soumith Chintala’s experiences and perspectives present a nuanced understanding of the current trends and future possibilities in artificial intelligence, especially regarding the balance between technology and societal implications.

For more information and to stay updated, visit [Lightspeed Venture Partners](http://www.lsvp.com) or follow *Generative Now* on your favorite podcast platform.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Hey everyone and welcome to Generative Now. I am Michael Mignano. I am a partner at Lightspeed And for this episode, I was thrilled to sit down with Sumit Shintala for a conversation about generative AI innovation at one of the world's largest tech companies, Meta. Sumit is a VP at Meta, and he co-created one of the world's most beloved ML libraries, PyTorch. And so he has clearly had a front row seat to the rise of AI over the past couple of years at Meta and previously as a researcher at NYU. We spoke about all things AI on the Generated NYC stage, which is a series of live in-person events for founders, builders, and investors hosted by Lightspeed here in New York City.

0:44So if you weren't able to attend, you're in luck because we have the full conversation right here. So without further ado, please help me in welcoming Sumit Shantala, VP at Meta.

1:00Hello. Hey, how's it going? Good. Thanks for doing this. Yeah, for sure. You are famously known as one of the creators of PyTorch. Why don't you give us your story and how it led to PyTorch and maybe tell us a little about one of the world's most beloved ML libraries. Sure, yeah. First, how many people know what PyTorch is? A lot. More than you expected. That's more than I expected, actually. Because one of the things that's interesting these days is AI, especially Gen.ai has gotten so commoditized that people don't actually need to know a lot of lower level details. They kind of call into various models.

1:41And so I was just like interested to understand how many people actually still know the lower level libraries. So PyTorch is interesting. I'm an AI researcher. I started working in AI in 2009 or so. And mostly I was involved in open source, partly because I was bored and also partly because I liked that spirit of learning things online, contributing to things. As I was building a bunch of AI research, the tools I was using at that time weren't as fun. And there was this tool at the time called Torch that was built by some other people. And I kind of started hanging out with them on their GitHub, started answering people.

2:29Eventually, I landed a job at Meta with offers from other AI companies just because I was doing that work. and the way tools go, I'm sure you can relate to this, is as you use tools to express your ideas, you build a bunch of great ideas and then the tools that you used eventually get outdated and you kind of need to build new tools. It's like the updating of technology, right? So at some point, Torch, which was a library we were building and using, was starting to get outdated. I did one research project in 2015, which was around getting to state-of-the-art object recognition, like take a picture, look at objects, identify objects.

3:19And it was super painful. So around that time, I just decided with a few other people that we'll attempt a refresh. Like, let's rewrite it for ideas that are being explored at the current time. And that's how we built it. Honestly, we didn't have a grand vision. We were like, oh, you know, PyTorch is gonna be this big thing that a bunch of people will use in the world. We kind of built it mostly for ourselves. And it just ended up being pretty useful for other people. And simultaneously, AI has expanded, like the market expanded, the addressable market has expanded quite a lot. And I think that just made PyTorch become like a bigger deal than I've actually ever imagined it.

4:08So you didn't expect it to sort of become such an integral important part of the ecosystem, but how did that actually happen? Like, did it just happen organically on GitHub? Like, did you do things that sort of pushed it in that direction? Actually, it's like a basic startup playbook. You kind of build the feedback loops. You kind of understand everyone who's using your product. You go talk to them. You try to help them. And then anyone posts about it online, If it's like they're showcasing their work built on top of PyTorch, we amplified it. We kind of built a community of people whose identity of why they connected with each other was because they used PyTorch.

4:50Like building that ecosystem and then having all these people who know PyTorch talk to each other about PyTorch and building that network of support, that made PyTorch not just be like a great technology tool, but also a great technology tool that had crazy amount of support that you just Googled any problem you had, you got answers that someone else faced and all those network effects. So I think that is what most people in this room do when they go back. I saw someone handing out cards on like, oh, this is like a scheduling agent. You should just use that. It's like, you know, the same kind of tricks, yeah.

5:32So I realize most of the people in the room, I think, raised their hands. But maybe for the people who didn't, we probably should level set a little bit and explain what exactly PyTorch is. Yeah, sure, yeah. Most generative AI is built on top of this fundamental computer science machine learning concept called neural networks. They have nothing to do with, like, brains. They're just, like, mathematical equations that learn statistical patterns from data. And PyTorch is a library that allows you to build these neural networks really, really efficiently. And if you have ideas and thoughts of how to build different kinds of neural networks, it makes it easy.

6:13And you probably are using PyTorch as a consumer of those neural networks between OpenAI's ChatGPT, almost any of the meta apps you use, Tesla cars, NASA. I mean, you don't use NASA rockets, but like AstraZeneca, the drug maker. Like basically it's now percolated into most of the world that does AI and deploys it uses PyTorch. Like I took my dog to the vet the other day and the vet had this like AI cancer screening thing. and they actually use PyTorch. So you never expected this to happen. You never expected all these different applications to be built on it. But now they have, you know, maybe outside of like that experience.

7:00Are there any that, you know, you sort of pinch yourself and you're like, oh my God, I can't believe this is built using this thing that we created. I get very nervous whenever we find a bug. Because, I mean, Tesla cars and cruise taxis or whatever. Like, you know, all these things. If I'm sitting in a Tesla and you turn on the AI drawing, I'm like, I hope I didn't write any books. What's it been like doing all this from inside Meta? You said you ran sort of the classic startup playbook, but you're doing all this inside of one of the largest companies in the world. What's that like? Do you get the resources and the support, or is it really like you are a startup?

7:37Yeah, I think general rule of thumb working in a large company is if you try to follow everything by the book, like basically you don't get much done. So you try to figure out how to build the best things while navigating bureaucracy and stuff like that. Meta has this culture of open source. React came out of Meta, Open Compute, like a bunch of well-known products like Presto. So culturally, Meta was super positive about open source in general. We got the right amount of support as we needed it, but like the shape of people inside of meta aren't really like people who understand how to grow tag and you know that like it's more like you're a large company engineer and you have a well-defined role and there's not a lot of ambiguity in how how you build and do things.

8:36So there's a little bit of navigating that but it's fine. Yeah so you know maybe sort of talking a little bit more broadly about meta and open source. Obviously, Llama is a huge deal and a huge focus. What role has PyTorch maybe played in the development and the rollout, the strategy of Llama? Yeah, I think Llama came extremely naturally to us at meta. I work on the Llama project indirectly. I help them with all their software and hardware side of things. When it came to Llama and releasing it in open source, honestly we didn't really think twice about it. It wasn't even like hmm should we release it and should we what is like the value for us.

9:23It's pretty obvious that we should do this. And I think we've always played this strategy just just releasing technology and if the technology, like the accelerated deployment of the technology helps the world then it probably helps us like that's how we started FAIR which was the AI research lab. We were just like we need AI for meta to progress like a lot of metas product needs need AI like for content moderation and you know various other like things and we're just like AI is not progressing fast enough so let's start an AI research lab that just publishes everything into the open and then just accelerates AI progress.

10:09And that will indirectly help meta and the applied teams use the open research labs technology back and forth. So LAMA is a pretty similar play. We want the acceleration of gen AI into products in meaningful ways. And we need to understand the safety implications of them. And there's no better way than to get people to try to break them in all kinds of ways. So that's just been super natural for us. It definitely makes sense that it's been natural for Meta to just say, hey, we should make a push with Llama because it's good for the world, therefore it would be good for us. But it very much seems like an intentional strategy now, right?

10:51It's one of the main things Zuck is talking about in the earnings call. The company is dedicating many, many, many H100s to the project. It seems like a very, very deliberate strategy now. What can you tell us about the strategy of Lama and open source AI now for meta? Yeah, I think open sourcing Lama was pretty natural. And as I mentioned, it's almost in our DNA to open source. It is intentional, but also we don't think deliberately about it. But generally, the AI race that is happening, I think you mentioned Zuck talks about it and a bunch of earning call and stuff. I mean, what's happening right now is multiple companies are trying to build foundational AI models which have a certain level of intelligence because they see those models as being valuable.

11:47and the current state of things is a bit unknown and weird in that we're in a in a state where we are just starting to see that you can build more and more intelligent models by throwing more and more data at it and throwing certain kinds of data at it we haven't really reached like this the upper cap of intelligence but I think the other interesting thing we haven't reached is how much intelligence is useful from a product perspective. For example, let's just say you take Lama 405b that Zah talked about that we're training and eventually hoping to release. Let's say Lama 405b, which is our largest model, is let's say 30 % smarter than Lama 70b.

12:39Let's just assume that. but let's just say it costs like 10 times as much to serve. These are all made up numbers. But if you take that into account, you need to start thinking about, oh, maybe actually most of business value and product value will be extracted by 70B or 7B or 3B. So maybe building the most intelligent model might not be cost viable, but it might still give you a technological advantage. I think these are the complexities that haven't been figured out, and people are just starting to figure out that having a single prompt that serves all users across the world does not really work.

13:21It doesn't work culturally, it doesn't work contextually. For example, chat GPT has a single system prompt for everyone in the world. It doesn't have any understanding of local sensitivities and stuff. So I think AI models, Gen.ai models also need to go through a personalization revolution. That, I think, is very early in the overall exploration. We're figuring it out. I think people recognize that there's something fundamentally valuable about Gen.ai models. They're doing useful things. How to use them, what are their limits, how to build products out of them. All of these things are pretty unknown.

14:03Yeah. Speaking of OpenAI, in addition to your work at Meta, which I find super interesting, I also enjoy your tweets. And you've put out some really interesting and thought-provoking tweets lately that I thought we could talk about here a little bit. You know, recently you commented on the whole ScarJo situation, which, for those of you who maybe live under a rock and don't know what I'm talking about, OpenAI came out with a new model recently, GBD-4-0. There was a multimodal voice component to it, which sounded like Scarlett Johansson. And now there is potentially a lawsuit from Scarlett Johansson.

14:38There's a lot to the story. I won't get into the whole story here. But I have a question for you, Sumit, based on one of your tweets you put out. Let's just say there is some sort of legal situation. How might this very, very specific micro incident potentially lead to legislation around AI, around training, around output, around input that may have ripples, that may lead all the way to some of the stuff that you are doing at Meta or other people are doing at OpenAI or other large research labs. I think it's super, super interesting that this has gotten very personal on an individual that many people find beloved.

15:14Yeah. People in current society have economic rights based on various aspects, right? Like they have intellectual So property rights, they have copyrights, they have likeness rights. There's various things that people have baked into how they optimize their lives. And some of those rights, I think, are being questioned by AI models. And I think the Scarlett Johnson incident is a way for people to... People have been getting a little unhappy in various fields. for example, the Screenwriters Association, the various, like, I think various creative artists of all kinds have been getting unhappy. They need moments like these to solidify what they're trying to express, right?

16:09They're like, we don't think this looks right. And we all need to coordinate and come together and express that this is not according to the current social norms. and I think these moments are going to allow people to understand each other, organize with each other and then question whether that's the direction in which you want to take society or not. Like do we preserve some of these rights that will be harder to keep around with AI technology or do we want to not preserve these rights and those are things that are both social and legal. I think just throwing it into the legal system would not resolve the problem.

16:51It has to be socially addressed and legally addressed. I think more moments like these probably will appear. I think Scarlett Johnson's one is just one of many that will appear and will test how society responds to this technology and evolve itself. We'll see where this goes. Because I think it's super positive to have people who feel a certain resentment for AI have these moments and organize themselves. Yeah, it almost sounds like what you're saying is, yes, it'll play out in the courts. That could take a decade. In the meantime, it could play out in the market, like people voting with their dollars or their feet or whatever the phrase is.

17:32Maybe in a similar situation, there was also this OpenAI equity clawback situation around on employees leaving, OpenAI really wanting to protect IP and trade secrets. I think there was a lot of blowback against this. I want to ask you, somebody else who works in a very, obviously a big lab that's investing a lot in AI and probably has a lot of trade secrets, like I know this is not the norm, but maybe given the state, the high stakes that everyone's playing in right now, secrets are that valuable and they're worth this type of situation? Or should we all be saying, hey, what did Ilya see? It's not about the secrets, actually.

18:09Okay. I think a lot of the resentment OpenAI got for their situation is that it was a bait and switch. Yeah. That when you showed up on day one, you signed your employee contracts, they didn't tell you that there will be this agreement that you have to sign at the end. and the language of that agreement is unknown to you until you decide to quit. And that agreement's going to have all kinds of extreme language in the rights that you might have. And if you don't sign it, they might take away a bunch of the money that you felt you already earned. I don't think if this is up front, like when the day you show up or before you sign your offer, they say, you know, after, like you have to go on guard and leave or something.

18:57This is super common in finance. Finance does this all the time where someone's working at a hedge fund and they know all the secrets. They quit. They're forced to just chill for like 6 to 12 months while they're getting paid. And that's the thing they signed, right? I think the problem was people who joined OpenAI didn't know when they joined that when they quit, they're going to have to sign a bunch of this stuff or else their equity is going to be wiped out. And the garden leaf concept could actually work because the market is moving so fast that actually the things you know probably will be irrelevant in six months.

19:37That's a really interesting comp. So another thing I'm really curious about to get your perspective on, and maybe your perspective is gonna be biased because you're at Meta, is the future of this space closed source AI like what OpenAI is working on, or more open source, which is obviously Meta's position? I think it's hard to say. I mean, this is the whole like Apple versus Android and US overwhelmingly uses Apple products, but the world overwhelmingly uses Android products. That's probably because America is a pretty rich country and Europe probably also has a bunch of Apple. And I think that's going to be one of those trade-offs in like AI as well.

20:18Like open source AI is probably not going to be as well integrated and smooth in like all kinds of aspects and ways as a vertically integrated product that uses closed source AI. But the advantages are it is way more democratizable, right? Anyone can build a product around it without having to go through a bunch of agreements with some central provider and doing things only in a certain way and stuff. So things like cultural norms, right? Let's just say a closed source AI provider out of California is going to have certain set of biases that might not work for a large part of the world. And that might be something open source AI can help democratize faster.

21:06I'm sure the closed source companies will try to work to personalize this to various markets, but the open source stuff gets personalized much faster. And that's what happened with Android versus Apple. So I don't know which way this is going to go, But I think there's a strong chance that there will be multiple segments of markets in which one wins versus the other. Can anyone compete with Meta on the open source side? This is a capital game and Meta's got a lot of it. And startups, can they match them? I think it depends on how long it's a capital game. Yeah. I think that's one of the interesting things that hasn't been resolved.

21:46that for context what the capital game is to train each of these foundation models it takes tens of millions of dollars hundreds of millions of dollars and it's predicted that it's going to take billions of dollars and in terms of capital expenditure it already takes like a lot of money and the current hypothesis that the world is running on in ai is if you want to train a smarter model, throw more data and compute at the model, and it's going to get smarter. You know, it's going like, you know, up and to the right, like the more data and flops you throw at training, the model is getting smarter.

22:27But it's not clear when it's going to start to saturate. So as companies train the next version of the model, they might realize that actually like it's starting to taper. And we don't know exactly when that happened and when that's going to happen. It's obviously not going to be like, okay, it just keeps going up into the right first. Like we might run out of data, or we might run out of high quality data. I mean, there's all kinds of things, right? I don't know if it's going to continue being a capital game until we capture all energy in the universe or something. It's not like a paperclip game, right?

23:11So we'll find out. I think to answer your question, if it stops being a capital-intensive game and it becomes about some other thing, maybe about human feedback, then it's very much possible that other open-source companies might also start competing because it's not about capital anymore. Maybe it's about democratization or whatever, right? So we'll see, but I think that's the state of the industry, and there's a lot of unknowns. I have questions I could go on all night, but I know there are a bunch of people in the audience that have questions as well, so we're going to shift over to audience Q &A.

23:51We have one coming over here. When you get the mic, just please stand up, say your name, say what you work on, and then fire away with your question, please. Hi, Nick Kirwani. I'm working on a bunch of different things, so really early days. One question I have for you is I think a lot of us are building on top of these foundational models. And the thing I think a lot about is like if you started building on top of iOS a while ago, you got locked in, right? And you didn't have your destiny. It doesn't feel like there's an obvious way they can lock you in because you can kind of switch them. But at the same time, like all our data is running through them.

24:25So I'm curious how you think the foundational models will attempt to create lock-in for the application layer. That's a good question. I think it's less like building on top of iOS and Android, and it's more like building on top of AWS or GCP or Azure, in that the way you interact with a model is via natural language. But the way the model responds is it has a certain flavor to it. GPT responds slightly differently than Claude, and responds slightly different than Llama. and when you're building a production app and you need it to be reliable, you kind of have a test environment where you're testing multiple of these.

25:09I think the lock-ins are going to be created on value adds on top. It's like, hey, if you use our drag system along with our foundational model, then we can show superior reliability or predictability or determinism versus if you use something else Or if you use it entirely within our stack, then we will give you more volume of APIs because the way we do inference while doing multiple services all at once is better. Like maybe latency guarantees and things like that. I think those are the kinds of things that you're going to slowly start getting locked in on. It is possible there are other ways in which you get logged in, but that's my initial thought.

26:00Hi, Matt Gee, CTO of BrightHive. We have a workspace that gives folks AI data agents, if you don't have a data team. So it's no secret that Jan LeCun has kind of a dim view on the future of generative models and is putting a lot of fair's weight around non-generative models and JEPA architectures. I'm curious, you know, he talks about that as obviously being the path to AGI, but being 10 years away. for folks like us in the room that are builders but also want to be building into the future. Two kind of questions around that. One is, do you see that shift towards non-generative architectures as something that is relevant to builders in this room that we should be thinking about right now?

26:51And two, are tools like PyTorch still relevant in world models or are there new open source tools at the low level that need to be built to make that happen? Okay. I think, like, just a correction, I don't think Yanlikan is actually against generative models. He basically thinks autoregressive models, which are current transformer-based models, which is what are used in current generation LLMs, they predict one token at a time. So they take in all the previous context and they predict one word, and they take all the previous context again and predict the next word. Those are called autoregressive models.

27:31And he thinks those models in their current form, like with the Transformer, are prone to making a bunch of mistakes by design. And he just thinks that you have to significantly improve the structure of Transformers to be able to not make mistakes like that. And I think that's actually a very reasonable take. I think there's a lot of like, it's become like a bit of a cultural war rather than like any real skepticism. For example, as of today, you look at ChatGPT or Meta AI, the assistant, or, you know, several other things. They don't just use a transformer model. They use a retrieval system. They use a search engine.

28:14They use a bunch of other things. So when you have other kinds of things that augment your LLM, you might get over some of the limitations that are there if you just use an LLM, which is just an autoregressive model that's predicting one token after the other. You can also explicitly bake in things like reasoning and memory. And I think what he advocates for is like, look, you want to add explicit reasoning, memory, world models, simulators, things like that, and that would make the model better. And I think the world is already moving in that direction. Tool use is the clearest example of like, oh yeah, we realize tool use is going to make the models better.

Read the full transcript

28:58I think he's in agreement with how the world is going towards. He's just making a point that if there was a thesis when GPT-3 first released saying you just need transformers, you just scale them up and everything's going to be amazing. And he was just pushing back against that. So I think you'd mostly be in agreement with him actually. The tools that are like, you know, you asked like PyTorch and all these other tools, Will they be used to build these next generation models? I think neural networks are going to be one part of it and of the whole system. If you think of AI, you're probably going to be like, well, it's a multi-component system and some of these components are neural network.

29:43And for those, you probably will use Citroën and things like that. But Google Search might just be one component. I think a database would just be one component. In terms of humans, we have a generic thing called a brain. We don't understand very much of it, but we do understand that it has a reasoning engine, it has memory percolated all over, it has various other properties, it has a motor control area and a language specialization area and stuff. I think when we build AI that's smarter, more useful in all ways, we're going to figure out how to build a bunch of this stuff. Some of it might be soft components like neural networks.

30:25Some of it might be hard components like an actual database. Some of it might be like somewhere in between. And yeah, I think for some of this, we'll use existing technology that lets you build a soft components. For some of this, you'll just use existing technology that builds those hard components. Some of this might need net new technology. Well, the next question is getting queued up. I got to ask, what's it like in Meta's Slack or Workplace or whatever when Jan LeCun and Elon Musk go head-to-head on X? Honestly, we don't notice or care. Come on. We're just too busy doing work. Good point.

31:05Like, okay, someone, like, if you're on a break having lunch or something, oh, did you see Jan LeCun and Elon Musk tweeted at each other? It's like, ha, ha, ha, great. I don't know if everyone rallied around Jan, if he walks through the halls and people are cheering him or something. No. Okay. Next question. Hello, my name is Philip. I also work on a bunch of things. My question is about bottlenecks. You mentioned there was one bottleneck, The Capital Game, and then you also mentioned there was another one regarding human feedback. that might be a sooner bottleneck. One sub-narrative that I've heard of a lot is the bottleneck of energy.

31:53A lot of people, I think Mark Zuckerberg mentioned that energy will be a very, in the near future, it'll be the first bottleneck. And then I think Amazon also invested into a nuclear project. Do you think, two questions around that. Do you think that that is a legitimate bottleneck from your perspective? Or is it all a overhyped narrative, one? And then two, if it is a legitimate bottleneck, where do you think it lies? Like, is it going to come first before the data bottleneck, or is it going to come after? That's a good question. We run, largely the world runs on capitalism as a concept. And there is enough energy in the world to do various things, power these lights, run a data center and all that.

32:43and there's more new energy sources being created, both sustainable and unsustainable. And they're largely priced out, right? Like you want to buy a unit of energy, you have to pay something. I think if you keep seeing these models get more and more intelligent and you're like, oh, I need to train like a bigger model because it's going to be more intelligent. At some point, economic pressures are going to catch up, both the energy prices will inflate. They're like, oh, there's more demand than supply. Okay. Energy prices will appropriately change. And there's also going to be a value proposition for the model.

33:29It's like, okay, I train a bigger model. Right now, the current race, I would describe it as a winner-takes-all race, where a lot of companies are assuming that if they have the more intelligent model than like anyone else, they can get more value out of it in terms of extracting value from the world than other companies. So they're like, if I have the most intelligent model, it's disproportionately more important than if I have the second most intelligent model. That's the premise that people are running on, which is why there's this like irrational investment right now where without knowing how you're going to make any money out of these models, you're investing into it.

34:15And based on various economic pressures, cash in hand at various companies, amount of investment people are willing to put in, and general inflation and economic things, I think it will be figured out. If we do get to a place where there's just a lot of demand, and you train a larger model, it's getting more important, and it's going to be disproportionately more important, I think there would be a limit, both physical, in terms of how much net new infrastructure you can build at what pace and how much energy do we have at what price. Those eventually will catch up, in my opinion. So I don't think we're going to go build Dyson Spheres around stars or anything crazy like that.

35:03Sumit, thank you so much for the time. I learned a ton. I'm sure they all did as well. Thank you so much for being here.

35:11Thank you so much for listening to Generative Now. If you liked what you heard, please rate and review the podcast. That does help. And of course, subscribe so you get notified every time we drop a new episode. Generative Now is produced by Lightspeed in partnership with Pod People. I am Michael McDonough, and we will be back next week. See you then. Thanks.

35:34you

From the publisher

This week on Generative Now, Lightspeed Partner and host Michael Mignano talks to Soumith Chintala at Generative NYC, a monthly meetup for AI builders hosted by Lightspeed. Soumith is a Vice President and Fellow at Meta and he co-created PyTorch, one of the world’s most beloved ML libraries. He’s had a front row seat to the rise of AI over the past few years at Meta and previously as a researcher at New York University. Michael and Soumith discuss Meta's open-source strategy, and the role of AI models such as LLAMA, and the broader strategy behind AI advancements. 


Episode Chapters

(00:00) Introduction to Soumith Chintala, VP at Meta and  Co-Creator of PyTorch
(01:05) How PyTorch Was Created
(04:08) Building a Community Around PyTorch
(05:32) Explaining PyTorch and Its Impact
(08:40) Meta's Open Source Strategy and LLAMA
(11:03) AI Models and the Future of Generative AI
(14:16) Implications of AI
(19:45) Open Source vs. Closed Source AI
(23:45) Audience Q&A
(35:03) Closing Thoughts


Stay in touch:


The content here does not constitute tax, legal, business or investment advice or an offer to provide such advice, should not be construed as advocating the purchase or sale of any security or investment or a recommendation of any company, and is not an offer, or solicitation of an offer, for the purchase or sale of any security or investment product. For more details please see ⁠⁠lsvp.com/legal⁠⁠.


More from Generative Now | AI Builders on Creating the Future

All 90 episodes
Soumith Chintala: Meta’s AI Strategy, PyTorch, and LlamaGenerative Now | AI Builders on Creating the Future · 36 min
Listen in VO