In short
Podcast Notes: Generative Now | AI Builders on Creating the Future
Episode Title
Mati Staniszewski (Eleven Labs) and Demi Guo (Pika): The Future of Media and Generative AI
Episode Description In this episode, Michael Mignano speaks with Demi Guo (CEO of Pika Labs) and Mati Staniszewski (CEO of Eleven Labs), discussing the evolving landscape of AI-generated content creation, foundation models, and future perspectives in the industry.
---
Episode Chapters
- 00:00 - Introduction
- 02:52 - Impact of Eleven Labs on AI content creation
- 04:54 - Creative applications of Pika
- 05:43 - High-quality AI video generation tools
- 10:40 - AI models for content creation
- 14:20 - Research behind AI advancements
- 17:17 - Future perspectives on AI-driven video creation
- 23:16 - Overcoming challenges in model development
- 26:11 - Teams behind Eleven Labs and Pika
---
Key Takeaways
Introduction to Companies
- Pika Labs
- Founded by Demi Guo with a mission to simplify video creation.
- Focus on removing barriers like cost and technical skills.
- Combination of artistic and technical backgrounds among team members.
- Eleven Labs
- Founded by Mati Staniszewski to enhance audio AI across languages and voices.
- Inspired by personal experiences with poor dubbing in Poland.
- Aims to provide high-quality, emotionally engaging audio experiences.
Generative AI Landscape
- The rapid growth of AI companies is largely driven by foundation models.
- Both Guo and Staniszewski emphasize the importance of specialized teams (research and creative) to develop effective models.
- Discussion on the current model race and its sustainability, with predictions that the focus will shift from model development to product usability.
Quality of Models
- High-quality models result from a blend of rigorous research and creative input.
- Importance of specializing in either audio or video for better control and quality.
Overcoming Challenges
- The development of models requires significant upfront investment before pivoting to product development.
- Emphasis on iterative improvement: models should continue to evolve alongside user needs and product requirements.
Product Development Insights
- Eleven Labs is exploring both consumer applications (reader and dubbing apps) and enterprise-level solutions (APIs).
- Discussion on the balance of providing consumer-facing products versus foundational building blocks for developers.
Competition and Pricing
- Startups face competition from large incumbents who can afford to undercut prices.
- Emphasis on maintaining quality and specialization to differentiate from larger companies.
- Importance of customer validation and trial access to build trust in product quality.
Future Perspectives in Media
- Potential for generative video to become mainstream, with predictions of AI-created content becoming more common.
- Anticipated features in consumer use cases include personalized content and immersive experiences that merge audio and visual elements.
Venture Capital Insights
- Strategy for fundraising should focus on achieving specific milestones rather than optimizing for the sake of fundraising.
- Building relationships and timing are critical for successful capital raises.
---
Conclusion The conversation highlighted the exciting prospects of generative AI in media, the challenges startups face in a rapidly evolving landscape, and the importance of both quality and specialization in building successful products.
For further insights, listen to Generative Now wherever you get your podcasts, and stay connected with Lightspeed Venture Partners through their social media channels.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Hey everybody and welcome to Generative Now. I am Michael Magnano. I am a partner at Lightspeed And this week, I am thrilled to have not one, but two of the most brilliant minds in AI today. I sat down with Demi Guo, co-founder and CEO of Pika Labs, the idea to video platform. And I also spoke with Matty Staniszewski, the co-founder and CEO of Eleven Labs, the generative voice platform. The three of us spoke on the generative NYC stage, which is a series of live in-person events for founders, builders, and anyone creating an AI today. If you weren't able to attend, you are in luck because we have the full interview here.
0:42Enjoy. I'm so excited to be able to speak with all of you tonight. Please help me in welcoming Demi Guo, co-founder and CEO of Pika Labs.
0:56And Maddie Staniszewski, co-founder and CEO of Eleven Labs.
1:03Hey. Thank you. Thank you. Thank you both. Thank you so much for being here. As I mentioned, both of your companies have become very, very well-known very quickly. Never seen a time in tech before where companies just grow and rise and gain notoriety so fast. And both of your companies are no exception. I'm sure most of the people in this room, if not everyone, have heard of your companies. But if not, maybe each of you could just share a little bit about what your companies do and maybe how they got started. Demi, you want to lead it off? Sure, sure. Me and my co-founder started a company called Pika a little year ago.
1:38We were both really deeply interested in artistic pursuits. So, you know, I grew up in a pretty artistic family. I have two sisters, one studied film, one studied fashion, and my family also have a cute little art gallery. I did a lot of poetry. I wanted to be a writer when I was young. And my co-founder also wanted to be an animator. and when we got older, we got more involved in computer science and were entrenched by the power of AI. So in the past decade, we spent time in the top AI industry labs and academic labs such as Google DeepMind, Facebook AI, and Microsoft Research. And most recently, we were pursuing a PhD at Stanford.
2:26So the reason we dropped out from Stanford and started Pika is video making is still extremely difficult. And we realized that we can actually fundamentally redesign the way people make videos using the research we were doing at Stanford. So we set out to remove the barriers of video creation, such as cost and technical training, to bring it to the masses. Awesome. And what about you, Maddy, with 11 Labs? Thanks for having us here and traveling here as well. So at 11 Labs, we do audio AI research and deployment with the goal of making content. university accessible across voices, sounds, and languages.
3:02I'm lucky to have started 11 Labs with my best friend, who I've known now for 14 years, from Poland. And part of the inspiration for 11 Labs came from Poland. There's this peculiar thing that if you watch a movie, a foreign movie in a Polish language, all the characters are voice-overed by one single voice in the movie. So male and female are narrated by one person. Usually no intonation, no emotion. I can assure you it's a pretty bad experience and something that still happens today. So in 2021, we actually went through that experience where my co-founder's girlfriend doesn't speak English and they did need to watch it together.
3:41And that prompted us to look into the problem overall. Of course, expanded that. It's not only the problem there. The whole dubbing is usually very poor and bad. And then if you extrapolate other verticals, all the texts not being available in audio, all the conversations not being possible with true audio AI, and of course, how we can turn existing audio assets and make it available across languages. And we started 11 laps soon after. So both of your companies are building products that help people with a specific problem. But the magic of these products and many companies in AI right now has to do with the foundation models that are being trained by these companies underneath the products.
4:21And I think these models probably have a lot to do with what I mentioned earlier about the rise of these companies happening so fast. Like so many AI companies, including the two of yours, basically found product market fit overnight because of the magic of these models. It's just so mind blowing what they can do. You can't help but love these things and want to use them. And we're now seeing, you know, this sort of leapfrogging of models like every other week. You know, this company releases a model and meets some benchmark next week. It's another company and another company. Demi, I want to ask you, like, what is driving this model race where every week we see one company just, you know, upping the one from the previous week?
5:02And how long do you expect that to last? Obviously, we see a lot of hype right now in AI space and also on the technology itself. I actually think it will die down over time when businesses start to realize that, you know, to focus on how to apply this technology and implement on a mass scale. I think at that point, we will focus less on the minor differences between different technical specs, but focus on which company provides the most useful and intuitive product interfaces, as well as the most clever and easy-to-use applications. How are these companies achieving such high quality of the models so quickly?
5:47Maybe another way of asking, and both of your companies are building some of the best models in your space. It's like, what is the secret to training the highest quality model? What goes into it? What's the magic? Yeah, for us, I think, first of all, like, we try to build a very, the best technical team, both on research engineering side. So, a lot of our founding, for example, a lot of our founding research team members all came from, like, the top AI lab, like, both from, like, the top AI industry labs, such as Google D.M.I., Facebook AI, and also the top academic schools, like Stanford, MIT.
6:20At the same time, I think for us, art is also as important as science when it comes to building a team, but also building a model. So we have a lot of team members from creative backgrounds who are filmmakers, artists, recording. They provide a creative lens in the model development process. Beyond that, of course, we're a very hardworking and efficient team. Yeah, I 100 % echo the team piece. I mean, I'm lucky to have a co-founder who has done a lot of the research before and has been able to assemble an incredible team of researchers. So 100 % that. And in our case, it's also being able to focus that team on a very specific set of goals.
7:08In our case, it's audio. So just fully going with the audio models. Maybe that will change now with the multimodality, but audio in general. And I can tell that there's always that temptation, even in audio space, to how does the video aspects or text aspects can play into it, but just staying true to that one thing to do it. And to your earlier part, we think that the quality of the models and leapfrogging will effectively, it matters differently for different use cases or different problems that people have. So for example, we do on one side, audiobooks. And we think the quality for audiobooks is getting relatively good, that you can produce an audiobook today, listen to it, and it's great.
7:46And now some of the value is shifting from the model itself into the product aspect of how can you produce that a lot quicker, easier, reuse your previous corrections in the future, audiobooks. But then you have those different use cases that the research hasn't caught up yet, and we need to put all the effort to get there. So in media and entertainment, you have very short segments that you want to control. You want to be able to not only produce it, but you want to modify it and maybe direct it to some extent. And that's where the research is going. So picking the right ones, both like what's the research focus and then what the use cases are where we build the product is the combination.
8:22So it sounds like what you're saying is there's a lot of upfront investment that has to go into the foundation model and getting that to a point where that it's usable. and once that happens, the investment shifts to being all about the product. But if the model hasn't met some quality bar yet, everything has to go into just making the model great. Is that kind of what you're saying? That's right. That's right. Although, of course, you sometimes want to balance those two. In 2023, we've launched some of our platform play, the audiobook example, and we had, of course, a lot of the salespeople try to go to the publishers and sell our solutions.
8:54The common feedback we've heard was we need pronunciation editing, we need more of the features to select the specific lines and correct them. And they all seemed right in practice, but the solution in this case wasn't so much on building the product, but actually just going full and building the research, because exactly as what you said earlier, it would have leapfrogged that. So instead of correcting every word, every number to a word, maybe you can just do it automatically. And that's where a lot of that effort is going. What about like in the sphere of sort of generative video models, Like, when do you think this leapfrogging will sort of asymptote or reach its peak and the shift goes to the product?
9:32Like, when will that happen? I think the interesting thing about model, AI model development is it's like a very continuous process. So, like, you know, there's never too bad and never too good. That's how I think about it. Obviously, right now in the video space, it's a bit more prototype than the voice space. So, both on the product and model side, we're always going to improve the model, right? Maybe initially the model is good enough for more consumer use cases. Maybe later on the model is better for even more professional use cases, but the model will keep developing. And the other reason we want to parallelize the product is also because sometimes the product can actually also inspire the model development.
10:12So for example, maybe you realize it's really important on the product side to do, I don't know, like consistent, having like being able to reference a consistent character or personalized object or personalized character. If you really care about this on the product, you then, this inspire you to develop, to spend more like time and energy on the specific model development for the product needs. Makes sense. Maddie, I've noticed that on the product side of 11 Labs, you seem to be taking two approaches where you're building sort of kind of like the building blocks in the form of like APIs and things like that.
10:52And then you've also got these and then apps like the reader app you just came out with or the dubbing app. I mean, on the product side, not the model side, how do you balance sort of consumer apps and, you know, end user applications versus these building blocks for enterprises? The consumer app is a new piece that we are trying out. But in general, so before the consumer piece, the factors we were trying, like, where do we think research is ready? Meaning where we can actually provide a value where that leapfrogging will not leapfrog us. Two, where do we see the pull from the customers? And then three, where can we provide a value over a longer period of time so it doesn't disappear just after we launch and there's another competition taking it over or it's raised to zero or anything like this?
11:32So these are the factors. And in our case, this was publishing first, then media and entertainment, and now media and entertainment and conversational AI use cases. And it's mostly developers, creators, enterprise working. One thing we've realized, a lot of creators coming through the platform were creating and turning books into audiobooks. They're creating newsletters in audio. They were turning their blog posts in audio, specifically with books. One of the things that they always were going against is that AI narration for some of the platforms, like Audible, is currently banned. So you cannot publish that on the platform.
12:06And we have tens of thousands of people doing those books, and they don't have a place to release it. So we decided to open it up with a reader app that you can download. And now you can upload a book and listen to it. And soon those people that create with the platform will have a distribution channel to reach their audience directly, which aligns with our goal of making it all accessible. So then you're an audiobook platform at that point. You're going head to head with Audible. It's likely just a small portion of indie book authors. But hopefully that quality will be at the level of anywhere.
12:39Speaking of competition, so both of your companies are going head to head with incumbents. When we talk about this leapfrogging of models, oftentimes it's the incumbents that are doing the leapfrogging, right? I think yesterday, Meta came out with the latest version of Llama, state of the art. OpenAI may come out with Sora, or they might come out with a voice model. And again, it seems like the incumbents who have a vast amount of resources are in the game, just like the startups. How do startups like Eleven Labs and Pica Labs stay competitive and stay at the frontier of model training when these incumbents just have war chests of cash and compute and everything else?
13:16Demi, what do you think? Yeah. First of all, I think, you know, the basic answers are like, first of all, like, you know, we're trying to be the best technical teams from both engineering and research side. And we also really value efficiency. and also I think like there's like differences between the product and the model, the kind of model we develop versus a bigger company. So for example, like we really care about, really value like building the type of model that can enable more control for the users. That's something like a pure foundation, media foundation model, what Google and Apple is using might not have the same level of control, for example.
13:58It sounds like what you're saying is specialization matters, right? For Google or OpenAI, they may have 50 priorities and they've got a small team working on video. They're not going to do the level of control that like a PicoLabs that is 100 % dedicated to video might do. Is that what you're saying? Yeah, like the experience of delivering a model, the experience that user interacts with a model also matters a lot. Yeah. Yeah. Maddie, what about what about pricing, though? Right. I mean, like incumbents can afford to just keep slashing prices to bring in as many customers as possible, whereas I imagine a startup like your margin matters.
14:32Like you need that margin to be big enough to keep your business running. And, you know, OpenAI just last week or two weeks ago, they released GPT-4O Mini. It's faster. It's cheaper. It just feels like that trend is going to continue. So how do you think about pricing strategy? Yeah, I think the first thing we need to keep, which of course will be increasingly harder, is that quality advantage over those models. So that's the first piece. And we are betting that we will be the best audio AI tool in the world. So that's the first thing. The second, I mean, to your specialization point, in general, we aren't just providing an API, but we are providing a platform with a plethora of different audio tools.
15:10And increasingly, people are using plethora of those tools together to derive the value. And as long as we can provide a value this way, then hopefully we provide a value and then can give a tenth of that as the price that we charge. The third thing that we bring in, so if you just want the standard voice and just make something speak, then of course those solutions will likely win in those cases. But if you do care about the quality control, if you do care about the variety and diversity of different voices, that's what we are trying to like bring in that most of those platforms aren't. The pricing always like is always correlated to the model itself, like the efficiency of model, meaning efficiency, meaning like the model, like you can actually develop a more efficient model, meaning it has less cost and like, you know, faster, basically.
15:56So there's like specific AI algorithm that I really optimize for the efficiency of the model. and this is something like maybe sometime like it's unclear whether the big company actually really cared that much about initially especially initially because um because they have a lot of resources they actually really care about the training the next huge version of model so they might like you know and then the smaller company may actually care more about efficiency to opt for pricing so it's really unclear about like you know whether you maybe you can actually have a a more efficient model that actually just has less cost.
16:30One more thing to quickly add there, because I think, of course, it's easy to say that quality is better and that's why people would flock to either of the platforms. So that's usually the challenge, I think, at least in our case, of how do we prove to the customer the quality and the platform is worth it for them to pay 5, 10, 2x the price of any of the other models. So both a plug for 11 labs here. We did, and maybe something that other companies can do too, We do have a grants program that you can get, and you have free access to 11 labs for three months where you can actually test whether that quality is good.
17:04But I also think it's a good thing to do in general, potentially in your companies when you want to attract the customers, just make it super easy to try out. And then if it provides value, then the price is worth it. Totally makes sense. Demi, switching to product and use cases, the use cases for generative video, I would say, are still playing out. We see things like avatars where companies are leveraging generative video avatars to do sales and marketing. You know, we've heard rumblings about Hollywood integrating AI into some of the production. But we haven't yet really gotten to the point where this stuff is so accessible and so efficient that it's pervasive in consumer yet.
17:44As the model quality improves and the products improve and they start to get really efficient and fast and super high quality, what do you think are the features and the use cases that will emerge with consumers? and generate a video? I feel fundamentally like, you know, what AI really enable is, you know, obviously like helping people to make like AI media, like content creation process more accessible, right? So it makes it much easier to create a video, but also much easier to create a video that are beautiful, basically. So we're actually incubating like a very new platform and use cases, probably more consumer than professional.
18:23I'm not sure how much we can review it now, but that's something we're working on, yeah. So stay tuned. Stay tuned, yeah. And both of your companies partnered on a feature, if I remember correctly, right? What was that? Maybe you can share a little bit about that, Matty. We are continuously experimenting. Effectively, of course, you know, the video is so much more immersive if you have audio, and audio is so much more appealing if you have the visual element. So we tried a few things with narration on the videos, with effectively combining what you have when you watch Pika Generation with audio, Pika did sound effects as well.
19:00They are merging. Maybe at some point, though, we can be able to create a whole personalized experience. Because, like, in the, of course, you can imagine some of the consumer plays where the video is more cluttered to the user, but as much applies to the voice, like, how incredible would it be where you can listen to the audiobook with, of course, with consent and permission from the person, but have your favorite characters, favorite voices all available on any specific movie, clip, audiobook, I think it'll be amazing. Matty, you talked a little bit earlier about use cases, speaking of use cases and some of the areas you're focusing.
19:33What are the areas that maybe surprised you? Where did you see demand and market pull that maybe you hadn't expected? And it definitely keeps happening in different places. First one is maybe a little bit boring, but it was exciting for us at the time, which was the audiobook example. When we released our app, it was like a text box where you could only input 250 text characters or something like that, literally Twitter length. And we had people, book authors come and copy paste their book all the way through to produce the whole thing. And then they brought people more to us saying like, those people also want to do it.
20:08So that's what inspired building more of the interface. But currently some of the, like one of the recent ones, specific ones that's amazing and I probably already said it to a few people here, but it's in the healthcare space where they don't have enough doctors and nurses to be able to cater to all the patients. So they are figuring out how to automate the call to remind about taking the medicine or take notes of how people are feeling. All disclosed that this is AI and people love it because of course so many people that these people are reaching out don't write emails, don't have texts, easily accessible.
20:40And they just so appreciate people checking up on them, taking notes, and then that feeds the actual people calling it. And I thought it was this incredible use case. But I think that conversational space, we'll see a lot more of like new ones. Maybe building on that quickly, for better or worse, like how far into the future are we away from like call centers, all call centers being replaced by AI? Yeah, I'm thinking in my mind, it was better for me to say longer than earlier. But let's say for the early adopters, I would say over the next year or two, there will be some scale over the next year.
21:16They'll be like hot scale deployment. Maybe similar question for you, Demi. Obviously, like there's a lot of, you know, there's a lot of talk of Hollywood and obviously all UGC moving in the way of generative video. Like, do you expect that, I guess, how long until you expect there is a Pixar like film, feature film, full length feature film that is all AI generated? I mean, is that conceivable in the near future? Yeah, that's a really good question. I think kind of the way I feel like is there are two kind of like features, the two type of way that the AI can create a feature of film. So one is like, you know, AI can create different clips and then you can stitch together different clips and create a feature of film, which is actually like pretty, like probably very normal because if you think about a film, like each scene is usually only like less than 10 seconds.
22:10Right. Right, so I think that's something I feel like is pretty foreseeable in the very short-term future, like, meaning, like, one or two years, maybe one year or something. I think for the AI to directly generate an entire film, like, one hour, 1.5 hour film. Like, one shot generation. Yeah, generate 1.5 hour, that's probably a harder question. maybe at that time you also want to generate like voice and like sound effects and then the model else has been very intelligent enough to understand the characters the the plot right like a interesting plot so that's something probably like more long term so maybe i don't know like but probably just a couple years like four years how long does it take to now generate those 10 second clip like in terms of compute time?
23:00Yeah, so for us, I think if you want to, it depends on different version model, you can actually just use one minute, 10 seconds. One minute to generate 10 seconds? Yeah, yes. Got it. Maybe it's like a continuous stream of movie one day or interactive movie. That's what I was thinking. And what is the limiting factor to this? Is it just more data, more compute? What enables the greater speeds and the efficiency of output? For sure. I think there's a lot of room to improve on the efficiency side. Our team also tried to hire a lot of people really strong on the efficiency side. The reason in general the video model is not that efficient is because people haven't really started to optimize for efficiency.
23:42Because efficiency is like an orthogonal AI, like research direction. so that like right now because there's so, people are still working towards a better quality model because there's still some problem with the current model. It's still not like, for example, 100 % hit rate. It's not like 100 % always very good. So that's why people haven't really started to like spend a lot of energy on the AI efficiency model. Got it. But I would imagine there's a lot, because right now we didn't do anything. So there's a lot of room, a lot of potential on the efficiency side. Got it. But speaking of Hollywood and films, Maddie, Eleven Labs recently released some new voices of like legendary legacy celebrities.
24:27I think Judy Garland, James Dean, Berg Reynolds. I'm guessing these are all like licensed through their estates. Kind of a fascinating release. How did that all come together? We'd love to hear the story of that. We have frequently trying to like bring the talent on board with us on that journey. and we were speaking with some of the estates for a while of like how can we empower their incredible voices to be accessible and in some of the works in our case, this was then a range of use cases across audio books and they themselves wanted to hear it back as well so we actually had some of the people in the family like how can we make it happen?
Read the full transcript
25:05So we tried that and soon not only from the estates and people that passed away but we also have now a number of people that are currently some of the iconic voices and we'll be bringing them up. So also stay tuned for that. That would be great. Could this one day evolve into almost like a marketplace where any talent could - For sure. Yeah. That would be, I think like even when you produce say a dab of the future, like how do you make sure that you compensate all the people and control and make it easier in one platform? We would love to make it easy for whoever creates that content. So if you want that specific voice, you can request it.
25:42the voice can accept it and then produce the line. And it's also not so distant because we already do a version of that. So we currently, you can create your own voice, AI voice, and then share it through 11 labs, and then earn money on a specific time period. And then you can earn money for usage on that AI voice. And now we have thousands of different voices, and soon we'll be releasing the amounts paid out to all those people. But for those that played Mortal Kombat, we have the voice behind Mortal Kombat there as well. Oh, that's pretty cool. Switching to venture and capital, the topic of venture capital.
26:16Both of you have been extremely successful at raising venture capital. And I think, you know, in general, this is one of the fastest moving times in recent memory in terms of capital and new startups. You're in a room full of builders, many of whom are also likely in the process or will be in the process of raising capital. Any tips for them? Demi. I think the way we think about fundraising is sometimes not necessarily just like, oh, fundraising for the sake of fundraising, but more about thinking about what is the next milestone. How much do we raise so that we can hit the next milestone so that we can continue to raise the next round or continue to move the company forward?
26:57So I think that's one thing that we really think about. What's our next milestone and how much you can raise to hit the milestone? Of course, if you show company progress, it's always the best. best thing for fundraising. Yeah. I agree with what Demi said as well. A hundred percent. You've been so good at this, Matty. Like, share your secrets with us. Because I certainly, I would have killed to be able to raise as much capital as the two of you had when I was building my company. So yeah, I would love to learn from the two of you. Maybe like on more practical side, and I think everybody will already know this, but I think the first piece is, of course, try to spend as little time on fundraising.
27:36but when you know you might want to do it, you want to probably queue them all up in a very specific same week or few weeks. So you get them all happening at the same time, which helps you, of course, use the time efficiently. But then if the offers are coming in or if you know the terms, you can use them efficiently against each other. So that would be one.
28:02Second, similar to Demi, and I think how Pika did it, It's like if you are trying to accelerate some bets that you are taking, it's probably good timing and a good prop to raise that capital and do those bets earlier or quicker. And then I do go, I would agree with the usual advice. Don't try to raise more than you will likely need in the next few years. Beyond that, it can make you lazy, which I disagree with. I think you'll still push. I just don't think it's worth to give the equity in the company away. And it's better to save it for employees or other people on the journey. Yeah, I also feel like generally like you should really think about, I think it's really easy for founders to go into like a fund to think about fundraising as a game and really optimize for fundraising.
28:45But I think eventually it's really long term wise, it's really about company building. So like, you know, it's really about like what you need for the company, like what like what you need from funding to help the company to continue to, you know, move forward. So don't need to optimize too much for fundraising, even though sometimes the fundraising process might not be the most optimized. Whatever is, the company building is actually more important. But one of the things that must be so hard for both of you right now is these types of businesses, because of the compute, they're so expensive to build.
29:21Do you think that maybe the value of these chips will come down such that you won't need to raise the type of capital you are now, which will then obviously put so much less pressure on the business because you won't have to raise at these crazy valuations. Is that a thought where you're waiting for the cost of the compute to come down? I think it's really hard to wait for the compute to come down, but hopefully the company can monetize. Yeah, of course. I think, yeah. Of course. So I could ask questions all night. You both have such fascinating learnings and insights and stories, but I know there are people in the audience that also have questions.
29:59And so we're going to turn it over to the audience. If you raise your hand, we have a couple of people running mics around. Please just stand up and speak your name, what you do, and then your question and who you're asking it to. Thanks so much. Hey, how are you doing? My name is Daniel Merja. I'm building an OS for venture capitalists. My question is for Matty. My co-founder and I actually wrote the Golang library for your clients. It's open source. So my question is, what do you think about open source? What are the advantages and how are you basically at the meeting for it? And are you going to release more open source models going in the future?
30:38Yeah, and thanks so much for contributing to community. I think like us, we think about developer tooling or a lot of how to make easier of deploying our model. We try to contribute there and the build the usual is the case and examples that we can. for the actual model, given that in our case, you know, that's at least now is most of the advantage. It's slowly shifting towards the product. But today, research that we've done is most of the advantage and where we spend the majority of time, mostly on the architecture level and not on compute level. We do care about that IP and don't plan to open source.
31:18I think there is a point in the future where we would love researchers to join and be able to share their research. but then we will need to shift a little bit away from the research being such a big part of our work over the product so probably the future and not in the short term given it's our main thing hi my name is Esteban so I'm really curious on what you asked or what was the last question about infrastructure in terms of yeah compute is expensive especially using compute for open source even though the cloud providers offer that it's expensive so it makes me wonder with your response and your answer from both of you is that you don't have any on-prem and you just rely on cloud providers.
31:58So in that regard, do you have plans on doing that yourself to just basically offset the cost, which is completely possible? Or are you completely tied and blocked towards the cost that cloud providers offer today? Yeah. So I can talk about first. So right now we're using cloud providers. We actually don't work on, we actually use a lot of small providers, maybe like small startup cloud provider as well. But we use a company like, for example, Together, some other companies. Actually, the reason we didn't build our own cluster is actually, to me, I feel like building our own cluster might be more expensive in some sense.
32:40And the reason for that is, first of all, like all the smaller cloud providers are actually pretty cheap, or a very good price. And then secondly, it's also very accessible. So let's say if I build my own cluster, I have to set up a team, it might need a couple months, right? So it takes time and energy and I have to hire the right people to build a cluster. It takes time, energy, when you hire people. And then building a cluster is also a very specialized job. So especially for a company like us, we're training a foundation model, we need a very large cluster. So it's actually not easy to build a large cluster.
33:15So it's very technical actually. So a lot of providers actually have problems because if you have a couple of thousand GPUs, then there's always error rate. So that's why if you build your own team, you build your own cluster, there's actually a risk like that. So you have to have a really good team. So that's why we use Cloud Provider. And also the other thing is if you buy a GPU, it's a depreciating asset, right? So if we use Cloud Provider, we can always run the latest GPU every year. So if we buy, let's say we buy A100 last year, and now we want to use H100, right? And then maybe later on, we want to do GH200, whatever.
33:58So then if we build our own cluster, then we just basically have depreciating assets, and then we cannot switch to a new GPU. What do you think, Manny? Do you agree? Slightly different. For inference, of course, we use Cloud Providers today. for training, experimenting with a combination. We are also lucky to have investors. I'm not Friedman who has Andromeda cluster that we can use to make it a little bit easier in the search times. I think our compute needs are smaller than video though. So it's also slightly different. I do expect to experiment more with both. So the quick answer would be that likely we will try some in the near future.
34:37I mean, it is interesting. Matty, you and I were talking earlier. there are some companies that are just vertically integrating across the stack. XAI is an example. They just launched the massive data center. I think OpenAI probably does some of this as well. But obviously a very, very costly endeavor, especially with language. Also, if you take a look, because I used to be very, like, the cluster is only cheaper if you have a very large cluster. Okay. Explain that. Yeah, it's because basically there's a fixed cost for building the team, right? The maintenance team and then I think the space and everything.
35:18You know, I actually, like, while we're deciding where to build GPU, I actually consult a lot of people. I forgot the exact details, but there's some fixed cost of maintenance and, like, renting a place and, you know, all those, the people cost. So it only, like, become cheaper and become marginalized out when you have a very large cluster. But if you have a very small cluster, you have to still run the place to hire people. So it's actually more expensive. To build cluster together, it sounds. There you go.
35:49A couple other quick questions. Hi. Hi, I'm Desmere. I'm a product manager at Cash App that happens to be very interested in generative AI. Both of your products are awesome. So questions for both of you. you're both working on like the frontier of engineering development um leading really big teams and exciting products how do you think about like a roadmap and like what to build because there's so much ambiguity in terms of like what you could build and there's so many opportunities like how do you prioritize how do you figure out what's what the right thing to build is yeah i think there's a i heard there's a saying that uh if you're building a company you should spend one third of time thinking about strategy, one third of time recruiting, one third of time just doing company execution stuff.
36:39I think that's kind of the thing about it is like, it basically I always, probably because we always think about our startup all the time. So basically all the free time you kind of are thinking about strategy anyways. I think it definitely helps if you're like, you should definitely keep thinking about like strategy and then you'll have idea and you'll be convicted in some roadmap. Maddie, Demi, this has been fascinating. I have learned a ton. I'm sure everyone in this room has as well. Thank you both so much for the time. We all really do it. A round of applause for Demi and Maddie. Thank you so much for listening to Generative Now.
37:15If you liked what you heard, please rate and review the show on Spotify and Apple Podcasts. That really does help. And if you want to learn more, follow Lightspeed at LightspeedVP on YouTube, X, or LinkedIn. in. Generative Now is produced by Lightspeed in partnership with Pod People. I am Michael Magnano, and we will be back next week. See you then. Thanks.
From the publisher
This week on Generative Now, Lightspeed Partner and host Michael Mignano talks to Mati Staniszewski (co-founder of Eleven Labs) and Demi Guo (co-founder and CEO of Pika). This conversation was part of a special Summer Edition of Generative NYC, a monthly meetup for AI builders hosted by Lightspeed. Demi and Mati discuss the landscape of AI generated content creation, AI foundation models, and perspectives on the future of the industry.
Episode Chapters
(00:00) Introduction
(02:52) The impact of Eleven Labs on AI content creation
(04:54) Demi Guo on the creative applications of Pika
(05:43) Building high-quality AI video generation tools
(10:40) AI models for content creation
(14:20) The research behind AI advancements
(17:17) Future perspectives on AI-driven video creation
(23:16) Overcoming challenges in model development
(26:11) Teams behind Eleven Labs and Pika
Stay in touch:
LinkedIn: https://www.linkedin.com/company/lightspeed-venture-partners/
Instagram: https://www.instagram.com/lightspeedventurepartners/
Subscribe on your favorite podcast app: generativenow.co
Email: generativenow@lsvp.com
The content here does not constitute tax, legal, business or investment advice or an offer to provide such advice, should not be construed as advocating the purchase or sale of any security or investment or a recommendation of any company, and is not an offer, or solicitation of an offer, for the purchase or sale of any security or investment product. For more details please see lsvp.com/legal.




