Decart’s Dean Leitersdorf on AI-Generated Video Games and Worlds

13 Nov 2024 · 47 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary - Training Data: "Decart’s Dean Leitersdorf on AI-Generated Video Games and Worlds"

Episode Overview In this episode, Dean Leitersdorf, CEO of Decart, discusses the innovative approaches his company is taking towards creating AI-generated video games and immersive worlds. The conversation, hosted by Sonya Huang and Shaun Maguire from Sequoia Capital, covers a variety of topics related to AI, consumer experiences, and the future of entertainment.

Key Themes

  • AI-Generated Experiences: The potential of Generative AI to transform how we interact with digital content and connect our imagination with visual representations.
  • Vertical Integration: Decart's approach to creating a fully integrated technology stack, from model training to real-time video inference, to provide a seamless user experience.
  • Overcoming Limitations: The distinction between solving specific problems and overcoming fundamental limitations in technology.

Content Summary

Introduction (00:00)

  • Overview of Decart and its mission to provide AI-driven consumer experiences.

About Oasis (03:22)

  • Discussion of Oasis, Decart's first product that allows for real-time interaction in a game-like setting without a traditional game engine.

Solving Problems vs Overcoming Limitations (05:25)

  • Dean emphasizes the significance of overcoming fundamental limitations instead of merely solving existing problems, which could lead to groundbreaking advancements in technology.

The Role of Game Engines (08:42)

  • Traditional game engines create environments for users to interact with, but Decart's technology offers a more dynamic, AI-driven alternative.

Real-Time Video Inference (11:15)

  • Explanation of how Decart achieves real-time video inference, which allows users to interact with AI-generated content in a genuine way.

World Model vs Pixel Representation (14:10)

  • Discussion on the difference between creating a world model (3D representations) and pixel-based models, highlighting the flexibility of AI-generated content.

Vertical Integration (17:17)

  • The importance of being vertically integrated to optimize performance and reduce costs, allowing Decart to operate more efficiently than competitors.

Building a Moat (34:20)

  • The conversation touches on how Decart plans to establish a competitive advantage through its unique technology and deep understanding of system architectures.

The Future of Consumer Entertainment (41:35)

  • Dean shares his vision for how AI-generated experiences will revolutionize entertainment and allow for new forms of creativity and interaction.

Rapid Fire Questions (43:17)

  • Quick Q&A session covering Dean's favorite AI applications, thoughts on generating video games vs novels, and his favorite scientist.

Key Takeaways

  • Generative AI's Potential: Generative AI has the capability to transform how we visualize and interact with content, blurring the lines between imagination and reality.
  • Vertical Integration as a Strategy: By controlling all aspects of production, from model training to application, Decart aims to gain a significant competitive edge.
  • Importance of Overcoming Limitations: Focusing on fundamental technological limitations rather than just existing problems can pave the way for substantial innovation.

Final Thoughts Dean Leitersdorf's insights reflect a forward-thinking perspective on the future of AI in gaming and consumer experiences. The potential for AI to reshape entertainment and enable users to interact with their digital worlds in unprecedented ways highlights the exciting developments ahead in this rapidly evolving field.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So we launched the races a few weeks ago. And really, when we launched it, the incredible thing from a tech perspective was, oh, this is the first video model that actually runs real time. And you can interact with it. It responds to user actions. You can move around the world. You can break blocks. You can place blocks. And so we got this nice game without a game engine. OK? That's not interesting. Why is this actually interesting? And so to answer that, forget about Oasis. Oasis 1. Think about, say, Oasis 3. OK? Okay, and imagine this. So imagine for a second, just go, just, just put tech aside for a second.

0:36Imagine you're looking at a mirror. Okay, and you have this magical mirror that you can talk to it. Okay, you can tell to do cool things. You can say, hey, I'm here, and here's my hand, and I want to hold a sword. Okay, you can give me a sword. And then you look at yourself in the mirror and boom, there's a sword in the mirror where your hand is, okay. And you move your hand around and the sword moves. And you can be like, no, no, no, make the sword bigger make it blue and it changes and you can be like, okay, now turn me into Game of Thrones and everything around you become Game of Thrones and then you get a crown and everything and you can be like, I don't like my crown, change it a bit and then you start jumping and you move around and the mirror responds to that.

1:13Okay. And that's interesting. Now the reason that's interesting is because it's a completely different experience than anything we've had before on Earth and it It allows us to kind of channel our imagination through screens that we can see. It connects two things. It connects what we see in our minds and what we can see with our eyes. And so that's what we're going with us. How can we, in a sentence, how can Gen A .I. really allow us to connect our imagination to what we see on our screens? And with that, we can take it into really worlds that we didn't explore before. or you can change everything from applications we can't do today, all the way to how we can even interact with computers or hardware.

2:17Hey everyone, I'm Sean McGuire, I'm a partner at Sequoia Capital. Today my partner, Sonya Huang and I are going to interview Dean Latterstor. Dean is a brilliant young mind. He grew up back and forth between Israel and the United States. He was the youngest person to ever get a PhD from the Technion at Israel at 23 years old. At least until his younger brother beat him and got his PhD when he was 21. DeCarte is trying to deliver the light full AI experiences. Really trying to let people interact with their imagination and other people's imaginations in a way that's never been possible before. To do this, they are fully vertically integrated, optimizing everything from as low -level as CUDA kernels up to designing their own models, training the models, and then at the end of the day delivering experiences.

3:04Over the next few months, we're gonna see some pretty impressive launches from these guys. BEEP BEEP BEEP BEEP BEEP BEEP BEEP BEEP BEEP Dean, thank you for joining us today. I was just playing Oasis this morning. I had so much fun. So let me start by asking Oasis, a fully playable AI game engines like what is it? Why did you launch it? A battle racist. So we launched the races a few weeks ago and really you know when we launched it the incredible thing from a tech perspective was oh this is the first video model that actually runs real time and you can interact with it in response to user actions you can move around the world you can break blocks you can place blocks and so we got this this nice game without a game engine okay that's not interesting why is this actually interesting And so to answer that, forget about Oasis, Oasis 1, think about Oasis 3.

3:53And imagine this. So imagine for a sec, just go, just put tech aside for a sec. Imagine you're looking at a mirror. Okay? And you have this magical mirror that you can talk to it. Okay? You can tell to do cool things. You can say, hey, I'm here and here's my hand and I want to hold a sword. Okay? You can give me a sword. And then you look at yourself in the mirror and boom, there's a sword in the mirror where your hand is, okay? And you move your hand around and the sword moves and you can be like, no, no, no, make the sword bigger or make it blue and it changes and you can be like, okay, now turn me into Game of Thrones and everything around you become Game of Thrones and then you get a crown and everything and you can be like I don't like my crown, change it a bit and then you start jumping and you move around and the mirror responds to that, okay?

4:36there. And that's interesting. Now the reason that's interesting is because it's a completely different experience than anything we've had before on Earth. And it allows us to kind of channel our imagination through screens that we can see. It connects to things. It connects to what we see in our minds and what we can see with our eyes. And so that's what we're going with us. How can we, in an, an, an sentence, how can Genai really allow us to connect our imagination to what we see on our screens? And with that, we can take it into, into really worlds that we didn't explore before. It can change everything from applications we can't, we can't do today all the way to how we even interact with, with computers or with hardware.

5:23I love the mirror. Let's take it further. Where are you going with that? Is this, is Are you building a game? Are you building a world model and interactive world model? Like how should I think about what is Descartes? What is Oasis? So let me ask you this. What problems does ChatGPT solve? Homework. Homework, great. And what else does it solve? It makes it easier to talk to computers. Nice. Shaan knows the answer because it's a lot of time. Because I've done a lot of time with this. It's not a classic Shaan. There's not a lot of time with this. But exactly that. The TLBR is challenging, but he doesn't solve any given problem.

6:02It helps you do your homework better. It helps you write emails. It helps you summarize exactly. Now, it doesn't solve a problem. It overcomes some fundamental limitation, which is exactly what Sean was saying, that it overcomes this communication barrier between humans and computers. Computers speak in structure languages. Humans in unstructured languages are languages with complex structure. That LLAMs just bridge that gap. Elect computers and machines interact with each other. and a language that we can both understand. That itself, the second you have that, you get a hundred different things that are solved on top of that.

6:35So we get with the mirror, or what you get with generative interactive video, is you get that communication barrier, and it overcome not just with text, but also with what we can see. Now computers will be able to see the world the way we see it, and they'll be able to show us the world in ways that we can understand. And you solve that, you build a platform that allows you to build everything on top of that. From next -gen Snapchats or TikTok to simulators for fighter pilots. Okay. And that's the cool thing here. And that's, if now we're in 2024, I think one of the most fun things we had at the cart is that we're founding a company.

7:21When you have an opportunity to build something that doesn't solve a problem, but overcomes limitation. 99 % of companies solve problems. When you look at companies that come to pitch Sequoia or pitch any other VC, they start with, here's the problem, here's how big the problem is, that's our tam and everything. And here's how we're gonna solve the problem. And usually the first to stay the same, otherwise you call it a pivot, right? You say, okay, this is the problem I'm solving. If you change the problem, you're solving, you call that a pivot and you 500 times you change the way you're gonna solve it That's 99 % of the companies and that's where you can do in any regular year There moments in history Recently it's been like once every decade maybe 15 years that you actually have the chance to build something that doesn't solve problem But just overcomes limitation and Let me ask you this in a different way Is it the Mac consumer product or an enterprise product?

8:18And is it a hardware company or software company? It is a hardware company or software company. And what problems is it's solve? Okay. And if you try to give me a list of problems that the personal computer solves, you'd have everything from gaming to Excel. And that's a nice thing about this, that you're building an insane piece of tech that you'll be able to productize in so many different ways. Yeah, I love that. One of the things that was so cool about what you've built is that there's no game engine inside as far as I can tell. Like, what do you think that means? Do you think that game engines are in artifacts of the past?

8:54Or like, what does that mean? Game engines were supposed to make it so that we can, so the one player, one person can create a world and a different person can interact with that world, right? That's the purpose of game engines. You have the game developer and you have the user that uses that. And it might go for also for movies or whatever other people use game engines for. Unreal has been used for movies a lot recently as well. Now, that is a very valuable product. And it has lots of advantages to it. The world is very consistent. You can really make things very accurate. The problem is that it does take a lot of time to interact with it.

9:35People like taking the basic game and like turning into a bunch of different things. And you know, it's, it's, as we got into this and we actually saw what people do with it, do you know there's an actual mod to put Pokemon inside Minecraft? Okay. You can walk around the forest and there's Pokemon running around. That's an actual mod someone built. Okay. And so people inherently have this, oh, we have, we got this platform and we want to change it. And so that's a nice thing about mods. What you get here is that because does what's running your game or your environment is an AI? You can interact with it in the ways where you used to interact with the AI.

10:12You'd be able to say, hey, can you turn this into like Elsa themed and then boom, everything becomes Elsa themed. And can you add a flying elephant? And there's a flying elephant in the game. And it's not just there as a picture. You can actually interact with it. You can punch the elephant, and it'll punch you back or whatever you can do in the elephants. And so I think that if this trend were to replace you game engines, it would have to be at the state that you can program for it so that it's some machine that one person can build worlds on and the other can interact with. And that is definitely coming.

10:48And not only that, it's gonna be much easier to program for this. You can just use words, you don't have to write code. And even if you do know how to write code, you can iterate so much faster on it. So basically to summarize this, I think what this will allow us to do is we'll just get modding much much much much faster and we'll get Interactive modding Do you get a little more technical for a second? Like you're the first video model I've ever seen that has real -time inference What are some of the things that go into having real -time inference? Like you know how hard is it and just like give us give us some of the flavors of what goes into that if we go back like like three, four months, like back to the summer.

11:34I don't remember where this was published, but there were a few headlines about, oh, when black wool chips come out, when Vitya's black wool chips come out, we'll get real -time video, okay? Hoppers just can't do it, the H100s can't do it. We have to wait for Vitya's next generation. And I think I heard this from quite a few different sources. It was, there were like two weeks during the summer where everyone was saying that for some reason, okay? And no, H100s can actually do it. Okay? And to pull that off, you have to do two things at once. You have to change a lot of things around the model itself.

12:13Not every video model can be run real time. You have to train the model differently. The architecture needs to look different. Now, it's not major architectural changes, but you do have to make them. On the other hand, you also have to do lots of the systems level stuff. You actually have to write your own good occurials. You have to write, we threw out like PyTorch's, this is a garbage collector and we're like half of it from scratch, okay? And you really have to write everything on the systems level as well To actually pull this off. So Because if you do only one of the two You'll be waiting for someone else to do the other half for you If you if you're only doing the systems level part then You won't be able to pull this off because you won't have a model that's that's ready to be interacted with this way If you do just the modeling stuff, you won't have the systems level support to be able to make it run real time.

13:01Can you say a word on how the model works? Like, you know, a transformer based. Is it similar to like the sorrows of the world? Like, what have you built on the model side? Yeah, TLDR. It's, it's exactly like the sorrows of the world. Just the prompt is user actions instead of text. Like, that's the easiest way to think about it. Like, think of like, we have text to video models, right? You have Sora that you put in a sentence and you get a video. So same thing here, just you put in your prompt is like your keyboard actions and your past frames and it generates the next frame. Okay, so how do you get the data between actions and video?

13:37So yeah, you do have to do some pre -processing steps here that you don't do with regular video models. For example, you do have to take the raw recordings of, hey, this is the gameplay and to label it at each step with the action that's being taken. And so, you know, we train a small model that does that. It actually doesn't need too much data. Like, you can solve that with a small model that doesn't need too many examples. And so you can just have your team just, you know, play for a bit, recorded that. You get a small model and then you use that to label all your data. Super interesting. And are you building a world model or is this just purely pixel representation?

14:14No, so it is. It's the beautiful thing here is that it's purely pixel representation. Now let's compare that to exactly what you were saying with like world models with 3D stuff and the other things. In AI like for for more than a decade, there's been a general question of do you solve stuff end to end or do you take an existing workflow and make something more efficient. Okay, like there could have been two ways to solve this problem. You can say hey, game engines exist, Unity is amazing, Unreal is amazing. Let's just plug into that workflow. Okay, let's build text to 3D. So I'll describe an elephant and I'll get, you know, the 3D mesh of an elephant and that'll be embedded into Unity and Unreal or whatever game engine you're using.

14:56Okay. So I compare that to the end to end solution of, at the other day, what I have is a screen. The screen needs to show something and that needs to work. Okay, and at the end of the day, what people do is they see the other computer screen and they touch their keyboard and they move their mouse and that's your interface and you solve this end to end from from keystroke to frame. Okay. So obviously these these are competing directions. Now over time, I think that there will be some merging between them, which is the the from a technical perspective, they each have their own advantages. The first is much more consistent over time.

15:34It's much easier to say, Oh, here's this object. Here's how it looks. and when it'll come back in two hours, it'll look exactly the same. And the other one, the N2N pixel, the diffusion version that does pixels in the pixel space, that one is much more easy to work around. It's much more flexible. You can really say, oh, no, no, no, change the elephant's tail. It's too big or you can actually edit it live in a way that's just more dynamic. So I do think that long -term though, So these two things will converge. And just if we roughly map this out. So today we really just have prompt two pixels, keystrokes to pixels.

16:20You could in theory say that the right way to solve this and say the next two or three years is to have two models. To have a model that everything's transformers, right transformers went. You have one model that's in charge of holding some state, state of the game, and that's unrelated to pixels. It's like literally just like a LLM wise transformer, okay? It just gets the current state, it gets the new user's action, and just outputs changes to that state. And you have one model that's doing that. And then the second model takes that state and renders it to pixels. So it makes sense that that's roughly where we'll converge, because that will really take into account both the advantage of world models and the advantages of the fusion models.

17:03Do you want to build both of those models? Of course. Yeah, definitely. But yeah, one of the things for me, I will say that we are a bit off. Like it will take some time to reach that stage. Yeah. One of the things for me that really caught my attention about Dean and DeCarte is they have this ambition to be completely vertically integrated. Like these guys understand, you know, literally down to electrons and how I'm serious. Like they understand how electrons move in logic gates and even like alternate logic gates and how you can represent them, you know, in levels even below assembly, you know, how you can change, you know, then in like in assembly, coup de kernels, like you can go, they that your ICs and they're optimizing every single level.

18:01And there, and I think you by doing that, I think they'll always have a 10X plus advantage over anyone that's just on the application layer. Actually, to talk about this, be a Sean loves to talk about this. I think the counter argument would be specialization. There's 10 ,000 very smart people in video at Choose Your Favorite Company working on this. You should focus on building the best possible user experience and the viral loops and things like that. So talk about your decisions to be vertically integrated. Let me actually say something because Dean can't brag about himself the way we can brag about him.

18:39But I've been studying business models my whole life has been a passion of mine from a young age. And from myself, like Google to me is one of the most amazing companies of all time. One of the most amazing business models. I worked at Google for a few years. I really feel like people have the wrong understanding of what was Google's mode. For us, I also think people have the wrong understanding of what is and videos mode today. But for me with Google, like obviously Sergei and Larry had invented PageRank. PageRank was a very beautiful algorithm. But it's actually, it was like a deep insight, but it's very simple to implement.

19:17It's like a very basic graph theoretical idea. And it was a published paper. So like Once, PIDRINT came out, everyone replicated it very quickly. For me, the real advantage of Google was that these guys were some of the best in the world that distributed systems, and at low -level systems optimization. And they had this very profound insight from early on that basically all the other search engines were buying sun micro systems, like server racks. The way they would get fault tolerance was by buying expensive hardware. where as for Google they realize that they can buy just cheap consumer commodity hardware that fails all the time.

20:00You buy Intel Pentium processors that are in your gaming computer or like Sandisk memory and you need five times as many total flops or five times as many bits to get the same performance because of all the failure rates. But the cost per flop is like 150. is you can have a 10x cost optimization, 10x cost advantage by really leaning into distributed systems and getting the most out of the hardware. And what that led to with Google is, for me, when I first started using it, it was this very, very simple front end. It was literally just a white webpage with a search box. It was, I think, a worse front end than Yahoo at the time.

20:44Yahoo also had chat rooms and other, where these kind of flashier, exciting things. But Google had this magical back end. All the magic to me of Google was on the back end. And I think that back end, the performance came from this cost advantage. And it came from the fact that they'd optimized all the way down to the bare metal. And with Dean and DeCart, the story really rhymes with me. And look, we need to say humble. This company hasn't done jack shit yet. We need to, it's a very long way before they deserve a comparison to Google. But, and for the sort of sequoia, you know, led the series A, co -led the series A in Google.

21:25I'm very proud of that. Also led the seed in Nvidia. So, you know, we have good history. Good track record. Good track record. Also, you know, a series A in Apple. But... Marshall break his own. Commercial break is over. But, but anyways, like, it just, I think, I think to really deliver these delightful, say a delightful mirror experience, which is a very simple front end, I think you need this absolutely insane back end that is optimized to the bare metal. And I think it's kind of all or nothing. Like if you can't deliver real time, I don't think it's very good. And I don't think you can deliver real time in the next year without going all the way to the bottom.

22:10And so I just, I don't know, for me, I think you kind of have to do that and these guys are the only ones I've seen doing that. Wow. I love what Chandra said because two things, two things really caught my attention. What is about the vertical integration? We'll touch about them a second. It goes back to your original question. The second is really about, so I won't name names, but I was speaking to someone who's very, very, very executive at Google recently. Okay. And just reminiscing about the past and trying to hear, because I was few months old when Google was founded. Okay, so I was around back then, but not really paying attention.

22:52And knowing you, Dean, you might have been paying attention. Um, you know, so I was trying to understand exactly what happened there, like why that was interesting. it came from like an unrelated conversation. And the way that person brought it up, we're talking about how GPU cluster is just unreliable. Okay, just in general, today, if you try to train a model like the one we trained, on any cluster, okay, whether it's hyperscalers or GPU clouds, that thing's gonna crash every few hours. Okay, and you're gonna have like the weirdest things, okay? You'll have one node crashing, and it'll be because two other nodes have dust on the cable between them.

23:35Okay, and there won't be any error to really tell you that that's what's happening, okay? So your training room will just crash, and you're like, okay, why did it crash? And you'll try rebooting it, and it won't work, and then you'll try removing random nodes until you understand what happens. And that's the state of the entire industry. Okay, pretty much like the only ones training that don't see this are probably Google and OpenAI, because they really built everything down to like Google built everything down to the hardware as well. OpenAI had a lot of time to really focus a lot of these reliability stuff, but anyone else is training from the big companies to the small startups.

24:08They're all experiencing this. And so I was talking to this person who's very, very high up at Google and they said, hey, we're today with like training today is like back where CPUs were in the 90s. Like forget like Kubernetes. There was no VMware. Yeah. Okay. Okay, nothing was reliable and your servers would just crash all the time and you had the exact same thing that most companies didn't want to deal with that. And so they just either paid for the premium service that was somehow better. A, so they both paid more money, but B, they also paid with time. The broken hardware exists before the stable hardware exists.

24:54Sure, we'll get to stable training runs any year. in two years, whenever that will happen. Nvidia will make their chips more stable, they'll make their code more stable, the GPU clouds will figure out stuff around this, that'll happen. It's not the state today. If you want a trained model today, you're gonna face all of that. And so one of the things that it's really challenged, you have to deal with it at the cart, we just, we could deal with it. Okay. The reason we can, so the model that you saw, Oasis, okay, Oasis, Oasis one, Oasis 1 converges from start to finish in 20 hours. Wow. And compare, you can compare that.

25:34We know what the, we have lots of joint work or communication with other AI labs. They were all shocked by this. Now, in talking about the best labs, training the fusion models, for this model, their convergence would usually take around two weeks. And it's not, it's both because they're not not using optimized systems layer stuff, but also because they crash every few hours, or every few days or whatever. We can actually hold it, we can look at, we can actually hold the training run end to end without crashing, we can also hold a training run for a week or for two weeks without crashing. And that reliability part really, really resonates with what happened back then.

26:17Now the thing that it is, is that it's really not simple to pull off. It's, you see, like we have this internal doc, I think it's around 200 pages now, of everything that can go wrong when you're training a model. And it's everything from, if you see this error on this node, then yeah, tell your hardware operators that these two nodes have a problem between them. These other nodes have a problem between them. And all the way to, and here's a fun one. At a certain point, as we're training oasis, we were doing the training run. And we needed some synthetic data to generate as well. And so we said, okay, well, we have this cluster, it has a shit ton of GPUs as well.

27:02Like great, it has lots of GPUs, but there's lots of CPUs and they're being like, they're utilized by like 3 % or something. Okay, we can just use this and just generate lots of synthetic data on the same cluster as the training is happening, okay? By the way, this is like blue the minds of our GPU cloud. They were like, you guys are using the cluster to like 200%. You're using the CPUs, you're using the GPUs, and you're using, we even use like the, then FentyBant to send data around during training. So like we're getting a lot more out of the cluster than Van is like should be expected, okay?

27:36Now that all makes sense. So on one hand, you know, you have this like the GPUs are utilized, the CPUs are not utilized. So you run like synthetic data in parallel. Well, it's not supposed to utilize, it uses just the CPUs. And so it's not supposed to hurt anything. And then your training run doesn't work. Okay, and you get a random error that literally says, the team will know how to say this better, but the error that you get is something like missing lock file and the data loader. Okay, and it's like, how are these two related? Do you know why they're related? They're related like this. The synthetic data gen was using up more RAM, which is fine, but it caused, sorry, no, it was, as, okay, to move the data around between the different nodes, as the synthetic data was being generated, it was using a more network bandwidth than before, and that caused Python's data loader to take one of its lock files that's usually network mapped and move it to be, swap it out to disk, okay?

Read the full transcript

28:38And that caused the state that different nodes had different lock files and that caused the day loaded to crash. Okay, now I'm probably saying this wrong and the team's probably listening to this and like no, Dean, you're getting all wrong. But that's the TLDR vote happened, okay? You did something that was supposed to make sense and you got a random error. And that's the day to day and we have a 200 page doc of all of these things. And so that's why - And this is a simple example that Dean is happy to share. Like there's, you know, there's - 100x harder, more important things that they've had to figure out.

29:17One that I think is also relatively simple, but it just kind of shows the current state of AI and Dean feel free if you don't want to talk about this. Don't talk about it, but they got access to a new cluster and somehow the cluster had not installed memory yet, but the GPUs have some very small amount of onboard memory. And so like most people would just not even be able to use the GPUs. Can you share anything about this story? Yeah, so this is actually a nice story. So, you know, we call this the best place on earth or train of video model. Training a video model isn't just the cluster. It's everything surrounding the cluster.

29:57Okay, you need to have the storage there. You need to have the networking there. There's so much that needs to go into building the best place on earth or train of video model. And we're actually very far away from them. Okay, like I'm assuming that roughly over the next half year lots of this will stabilize and lots of the GPU clouds are working on this But yeah with one of the clusters that we got to there was any storage and by the way It wasn't even with one it happened with a few clusters and different clouds, okay? That you know the clouds they bring the GPUs and they try to focus on so focus on getting the H100s that you know They forgot the memory or the storage and it's fine And it's okay.

30:34And they were gonna start, they would get there, but they try to release everything as fast as possible, which is great, which makes sense. And so, okay, there was no stable storage, storage optimized nodes that you can use, or an S3 bucket or something that you can use. And so we said, okay, well, every node has a few SSDs connected to it. What if we just build our own mini fake distributed file also still on top of that, okay? And that's what we did. And it worked. And there were so many things to overcome to make that happen, but it works at the end of the day. And that's, I think, and it goes back to your question about vertical integration.

31:20Vertical integration, so I'm, Sean knows business much better than I do and has been around all of these fields way longer than I have. Okay, I did PhDs and like, I think you just called you old. I was used Google when it first came out. And I bought Nvidia shares in the IPO, which is also right around the board. So yeah, I think I've heard before it was born now, 96. 99, I think. 99, okay, okay. As far as I see it, I correct me if I'm wrong. Vertical integration usually gives you two things. It gives you a cost reduction like higher margins or whatever, and it gives you the build to move faster.

32:01Maybe gives you a third thing because usually things give you three things, but who knows? So, I think here in AI, the more important part, sure they're both important, but I think the second one is even more important than the first. Because at the end of the day, if you look at all the problems we're facing, great, they will be solved. It'll take time for them to be solved. And if you, you know, I think that there was a great article, I think at the information about how it was like a few months ago that people who leave Google to start startups suddenly realized that nothing works because everything works inside Google and then you go outside like oh there's no storage or oh the my cloud providers and provide me with this I actually need to take care of this and so okay fine over time these things will stabilize and your clouds will provide you what the cloud needs to provide you and you'll have great companies that provide you with like middle layer for the system stuff or even for the the model training stuff will make lots of easier for you but If you really do everything end to end you can you can get to market a year before everyone else you can get to market two years before everyone else and That's and that's I think what's key here because even if we go to to the Google story Or open your story Techno's don't last Right, I think sure Google Google is a great search engine.

33:21Bing is probably not that bad Sure, maybe Google has more data, so they're able to do that. But Microsoft, a huge company that we've been working on, Bing for so long, it's a good search engine. They have the tech. It still doesn't mean that now Bing and Google are balanced. So at the end of the day, the entire game here is, get your tech mode quickly, and then fast two years before everyone else, like Google and like OpenEI did, and work as fast as possible to convert that to a different note. And that's the game here. That's what you have to play. Because we can all say, okay, you know what?

33:58So, quite invested, all good. Let's put the money in the back for a second. Okay, let's get some interest on that. We'll go beyond the beach for like two years, wait for everything to stabilize. We'll come back in two years and then we'll build the same company. And that'll be great, but someone else would have done it before. And that's, that's, I think why we chose to be vertically integrated. I love it. What's your moat gonna be? long term or short term both both perfect short term tech okay short term tech and that's and that's great and we have the best systems layer stuff and we're also doing the model layer stuff as well so we're fully integrating and that's your that's your mode at the end of the day short term long term long term I think that's that's a great question and let me let me share something that I found really interesting.

34:48Okay? So there is a new weaker version of network effects that exists today that didn't exist before. And that network effect is called what people see on TikTok. Now, why is that interesting? Okay, we were, what are the companies that I, I really, that we learned a lot from and that I think is, is an actually really, really good company. They did end up selling to this character. I they didn't end up selling to Google and wanting to go back to training big models. But character, there's a lot to learn from character. And one of the things that the second they took off, they had lots of competition instantly.

35:32Like fine, they were techno lasted for like half a year until meta released open source models. And then other people started running this, they were still vertically integrated. And so they were able to be 10x cheaper than everyone else, which was great. But one of the things that really stood out to me was their TikTok mode. If you go on TikTok and you look for any character, you're at competitor, fine, you'll find a video of that competitor and then you'll scroll and you'll see 100 videos of character. And if you even, you know, if you go on the videos, which are not character, all the comments are full of character.

36:08And if you talk to a random character, I usually they don't even know the competition. And so we have somehow, literally because of TikTok, there is a new mode of what people say about you on TikTok and do you have a many network effect there? A mini brand, I'm not sure if it's a network effect or brand effect, but why is it different from just brand? So it's very similar to It's similar to brand, but it's in your face. Like brand like 20 years ago was okay, did you hear your friends talking about this, your parents talking about this? Here you're always on, like the younger generation, especially they're always on TikTok, and so they just see this instantly.

36:53And so there's even a big question of whether a moat like that could survive for the two, three years until you need, until you get your long -term votes of insane brand like Google's or a distribution brand or something like that or a distribution mode or something like that. So I think we're really in this new market here. Yeah. That we're not necessarily going to have the same modes we had 10 years ago. Hmm. Super interesting. Hardware is always the best mode though. And for this word, like Google, I think, you know, they elevated what was initially like a software mode and a distributed systems mode to becoming a hardware mode.

37:35I personally think that Google has not leveraged that mode enough on the application layer. They haven't had that many really fantastic breakout, consumer products, sense that really is. But they have an absolutely gigantic cost advantage, really because on the hardware layer. When I was at Google, there was this project that just absolutely blew my mind and gave me a prepared mind for a few investments, which is basically Google built optical interconnects to move data in data centers. This, like, one of the papers, if you Google Jupyter Rising like Google data center, you'll find the papers.

38:20And basically these optical switches, by turning them on, basically about doubled the performance the data centers, like these ones switch is mainly rack to rack in data centers, you know, going, moving from electrons to photons. And one of these switches were insanely hard to build. And basically everyone outside of Google, if you ask them at the time, is it possible to build, you know, they say no way. They say no way. Terrible. Per second switch, whatever. They'd say absolutely no way. But they did it. They didn't even know for years that Google had this. And, And it reduced power consumption of the data center by 30 % or something.

38:55And just like, those things are real fundamental notes. I think it's always hard to know what the notes will be for accompanying the future. But I strongly believe hardware is the ultimate mode. In part because there's always going to be an extreme delay to move atoms. like you to spin up fads to get power, to build a power plant, like even in a world with AGI, the time scale of hardware, even in a world with a billion optimists robots, like the time scale to make new hardware will be much slower, or the time scale will be longer. So anyways, I hope to cart has a hardware mode. So I think I agree with you on that.

39:40Like long term. Okay, you know, this actually goes back to when we were founding the cart. So we said, okay, we called it the golden ticket. We got this ticket that you get once in your life of starting a company and at time, we're going back to what we were discussing before. Starting company at a time, we can solve some fundamental limitation and like there is some huge tech shift going on. And we said, okay, there are three huge companies you can build here. That was our analysis of the field. A, you can build an Nvidia competitor and if you, like the next gen chip that's actually built for AI, and it'll be very tough to do, but Nvidia, Nvidia's not just a chip giant, but they're a supply chain giant.

40:27And it's insanely hard to do, but if you hustle your way around, everyone in the industry wants to help you. And so it's doable if you really, if you really excel in the business line. Two was to build the next AWS. US. Like there is, there is an opportunity because the workloads themselves are changing. There is an opportunity to be able to build a new cloud. Very, very, very tough because in that market there's a default winner. If you all lose, the big three will still win. The big three plus Oracle or the other clouds as well. And the third was create new experiences. That new experiences will happen and these experiences will be drastic enough so that the next trillion dollar company can come out of these in five years and out in 30 years.

41:15And so we had to choose one to start with. We chose the experiences one. But a definitely strong second was let's build an Nvidia competitor one. And so we have that lingering thought of one day we'll get back to this. I see why you two are friends. I will close it out with one last question. If everything goes right. What is what is Dakar in 10, 15, 20 years? And what experiences have you crafted? And like, what is the future of consumer entertainment? I don't know if that's the right market. And I'll say this, and I'll give credit to James from Sequoia here, because he's the one who coined this term, Jettoried experiences, GX, okay?

42:00And we call this UX is dead long load gx. Okay, basically we're going to have new experiences that are generated in ways that match how humans want to interact with computers. And then encapsulates everything from character, eyes, and generate experience to real -time video models or generate experiences. And that's what we're going to see. The card at the end of the day is a generated experiences company. We're implementing this with being fully vertically integrated, with having the systems layer. At the end of the day, you're a generated experiences company. You're creating the new, the new wave of experiences that that's going to touch every single person on the planet.

42:49And that's where the card is. Now, the only question is whether it does take 10 or 15 years? Today's agent might take less. It took a long time for the previous Titans to rule the world. I will take that long. The style definitely take at least five years. You operate on a different time scale than a lot of the best AI researchers that are in our orbit and I really respect that about you. Should we close that with a rapid fire round? Okay. Favorite AI app other than Oasis. It has to be between Chattey Patin character. Has to be between Chattey Patin character. There is character for. Not using character.

43:28Sure. Okay. But on the basic notion of that we'll have these apps that are entities that have, that hold some kind of relationship, whether it's friendship or whether it's your hotelarian with hundreds of millions of people, I think that's an insane platform that's going to be the basis for so many things going forward. Yeah. I love that. Favorite AI company could be the same as last. Same as the last answer. Same as the last answer. Okay. Let's see. Why did you first program a computer? First program on computer. When I was 13, bots for runescape. Okay, great game runescape. I bought it the hell out of it for years.

44:08Until six years in, I used a bot that I downloaded from the internet. 25 hours later got banned. Are we going to have AI generated video games first or AI generated novels? And I mean, at the level where I would actually pay for it. You're going to have, the first thing you're going to have is a platform that lets other people use the creativity to create this content because AI is still far away from creating creative content. Ah, super interesting. Okay. Who's your favorite scientist ever? Favorite scientist. That one I like, that one I like. It's, you know, there's a reason we chose the name the cart.

44:47We chose the name the cart. Because, so, so, okay, first of all, I answered the question. and favorite scientist is the Vinci, because I think he's both an insane scientist and engineer and somehow was able to get people to fund his project, okay? He was like, if you go back to the Vinci, like he literally was a great scientist engineer and somehow knew how to raise money from VCs back then, which were kings, okay? So yeah, definitely the Vinci and the card in Tesla are close seconds. The reason we chose the name the card was, We looked at Tesla, we love both that company and the name. We needed someone who resembles the same thing that Nicole Tesla resembled to the company Tesla.

45:37For that, that was the card because I think they're for IAM resembles almost a lot of what are your eyes today. Perfect notes and then, Dean, congratulations on what you've done. Thank you for joining us today. We love this conversation. I'm not gonna congratulate you haven't done jack shit yet. Right? That's build something insane, but I love the sentiment. We can't celebrate until we really win. Yeah. Okay, there's no celebrating small wins.

From the publisher

Can GenAI allow us to connect our imagination to what we see on our screens? Decart’s Dean Leitersdorf believes it can.

In this episode, Dean Leitersdorf breaks down how Decart is pushing the boundaries of compute in order to create AI-generated consumer experiences, from fully playable video games to immersive worlds. From achieving real-time video inference on existing hardware to building a fully vertically integrated stack, Dean explains why solving fundamental limitations rather than specific problems could lead to the next trillion-dollar company.

Hosted by: Sonya Huang and Shaun Maguire, Sequoia Capital

00:00 Introduction
03:22 About Oasis
05:25 Solving a problem vs overcoming a limitation
08:42 The role of game engines
11:15 How video real-time inference works
14:10 World model vs pixel representation
17:17 Vertical integration
34:20 Building a moat
41:35 The future of consumer entertainment
43:17 Rapid fire questions

More from Training Data

All 110 episodes
Decart’s Dean Leitersdorf on AI-Generated Video Games and WorldsTraining Data · 47 min
Listen in VO