World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI

6 Dec 2025

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Latent Space - Episode on World Models & General Intuition

Episode Overview Title: World Models & General Intuition: Khosla's Largest Bet Since LLMs & OpenAI Host: Latent Space: The AI Engineer Podcast Guest: Pim De Witte, CEO of General Intuition Date: [Insert Date] Listen here: [Latent Space](https://latent.space)

Key Themes

  • Introduction to General Intuition and its foundation from Medal, a gaming data platform.
  • The significance of world models trained on human gameplay as the next frontier in AI after LLMs.
  • The interplay between gaming data, machine learning, and real-world applications in robotics and AI agents.

Key Takeaways

  1. General Intuition Background
  2. General Intuition is a spin-out from Medal, a platform that has amassed a large dataset of gaming highlights (3.8 billion clips).
  3. Khosla Ventures made a significant investment of $134 million, marking their largest seed funding since OpenAI.
  1. Gaming Data as a Goldmine
  2. Medal’s clips serve as “episodic memory for simulation,” enabling world models to learn from peak moments of human gameplay.
  3. The platform emphasizes privacy by mapping actions to visual inputs without capturing individual keystrokes.
  1. Building Advanced AI Agents
  2. Discussion on developing fully vision-based agents that can process frames in real-time and predict actions akin to human players.
  3. Importance of actions, memory, and understanding occlusion (e.g., smoke, camera shake) in creating effective world models.
  1. World Models vs. Generative Models
  2. World models are defined as systems that understand a full range of possibilities and outcomes from current states based on actions taken.
  3. There is a discussion on the integration of world models with LLMs, considering them complementary rather than rivals.
  1. Commercial Applications
  2. General Intuition aims to replace traditional behavior trees in gaming and robotics with a more dynamic "frames in, actions out" API.
  3. The potential for applications in simulation environments, including scientific discovery and robotics.

In-Depth Discussions

  • Data Privacy and Collection
  • Medal’s approach to ensuring privacy while maintaining a comprehensive dataset for training.
  • The role of privacy-first action labels in creating a rich environment for developing AI.
  • Visions for Future AI
  • The ambition to create spatial-temporal foundation models that could drive the majority of AI interactions in both simulated and real environments by 2030.
  • Pim envisions the technology powering a significant percentage of interactions in various applications, from gaming to robotics.
  • Pim's Journey
  • Pim shares insights from his background, including experiences from RuneScape private servers and his transition to AI.
  • The importance of building a strong foundational understanding of AI principles and data utilization.

Practical Applications and Implications

  • Game Development
  • How General Intuition's technology can help game developers enhance player engagement through improved bot interactions.
  • The significance of having realistic AI agents that can seamlessly blend with human players to maintain game dynamics.
  • Robotics and AI Integration
  • The applicability of gaming models and data in the realm of robotics, particularly with training models that require less real-world data.
  • The potential for using this technology in complex environments, such as manufacturing and real-world robotics.
  • Open Research and Collaboration
  • The aspiration to revive the culture of open research and collaborate with academic institutions to foster innovation in the field.

Conclusion and Future Vision

  • General Intuition aims to be at the forefront of AI by leveraging its unique dataset and world model techniques.
  • The long-term vision includes creating a significant impact on both simulation and real-world applications, aspiring to redefine intelligence in AI systems.

Additional Resources

  • Follow Pim De Witte: [Twitter](https://x.com/PimDeWitte) | [LinkedIn](https://www.linkedin.com/in/pimdw/)
  • Follow Latent Space: [Twitter](https://x.com/latentspacepod) | [Substack](https://www.latent.space/)
  • Full show notes available at: [Latent Space](https://latent.space)

---

This concludes the detailed notes and insights from the podcast episode on world models and general intuition. The discussions highlight the innovative approaches being taken in the AI field and the implications for future technologies.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hi listeners. As you may know, I recently wrapped up the AI code conference in New York. And while I'm traveling, I do like to visit top AI startups in person to bring you interviews that you don't find on any other podcast that just does a Zoom call. General Intuition, or GI for short, is a spin-out of a 10-year-old game clipping company called Metal, which has 12 million users. But in comparison, Twitch only has 7 million monthly active streamers. Metal collects this data by building the best retroactive clipping software in the world. In other words, you don't need to be consciously recording.

0:28you actually just have Metal on in the background while you're playing, and you hit a button to clip the last 30 seconds after something interesting happens. It's very similar to how Tesla and self-driving does bug reporting, if you've ever done a self-driving bug report in Teslas. The result is that Metal has accumulated 3.8 billion clips of the best moments and actions in games, resulting in one of the most unique and diverse datasets of peak human behavior, actively mining for the interesting moments. They were also very prescient in navigating privacy and data collection concerns by mapping actions to these visual inputs and game outcomes.

1:04As you saw on our Fei-Fei Li and Justin Johnson episode with World Labs, and with the recent departure of Yan Le Kun from Meta, there's a lot of interest in world models as the next frontier after LLMs to improve on spatial intelligence and to work on embodied robotics use cases. DeepMind has been working on this with Genie 1.2 and 3 and SEMA 1.2, and this year, Okonai and I seem to finally agree because they have been betting on LLMs a lot, and they made the news by offering$500 million for Metal's video game ClipData. Our guest today, Pim, turned down that money and instead chose to build an independent world model lab instead.

1:38Kostler Ventures led the$134 million seed round, which is Vinod Kostler's largest single seed bet since OpenAI. We were able to get an exclusive preview of GIs models, which unfortunately we cannot show you directly, but I can confirm they were incredibly human-like and we chose to include the first 11 minutes of the demo discussion, even though I couldn't show it to you. It may be hard to follow, but I tried to call out what was noteworthy for you to know as your likely reaction if you were watching along with us. Now enjoy the world's first look at my first look at Genuine Intuition. So what I'm about to show you is a completely vision-based agent that's just seeing pixels and predicting actions the exact same way a human would.

2:15And so yeah, what I'll show you here is what this looks like four months ago. So again, this is just an agent that's receiving frames, and it's just predicting action. So you can see it has a decent sense of being able to navigate around. It tabs a scoreboard, just like gamers always tab the scoreboard. So these are pure imitation learning. I see. The Z is slicing the knife. Yeah, exactly. So it's doing everything that humans would. In this case, here was the first interesting part that we saw. It gets stuck, and they have memory as well, so you see it can get unstuck. How long is the memory? Four seconds.

2:52Yeah, four seconds. So this was four months ago. This was maybe a few weeks after that. So you can see there's like, it's still doing the scoreboard thing, but they're still quite like, and these are bots too. So you can see that. It's very human. Let's just say that. Yeah. And then, right. So this was really like the early days of research where you can see, right. It does one thing and then goes for another. And then we've been scaling, right. on data and compute, and also we've just been making the models better. And this is where we are now. So what you're seeing is, like I said, pure imitation learning.

3:30This is just a base model. There's no RL, no fine-tuning. This model sees no game states. It is purely capable, not sequence, acceptance. It's purely predicting the actions from the frames. That's it. And this is playing against real humans, just like a human would play. And it's also, it's running completely in real time. So there's absolutely, everything here plays exactly like human.

4:00Do you give it a goal? No. It just figures out it's a goal because obviously it's trained on the same. Yes. And I picked, right, I picked a sequence where also it doesn't do well initially. So you can see like, this is just like a sequence, a random sequence. But this is the, I mean, it looks like it's very well.

4:18Yeah, this will be good. Maybe too good.

4:30This is my favorite part. So you can see it does something that like, here, like Neumann would never do this, then gets unstuck, then has four, realizes switch, and then in the distance.

4:48so you're saying one it makes a mistake that a human will never make but it unstacks itself and two what we just saw is it is doing superhuman things yeah okay yeah um i mean there are things that that demons said obviously um but because it is trained on on the highlights of things that all the exceptional things it's inherited yeah so it's not like move 37 where we rl their way into or something. Yeah, replicating a superhuman. Yeah, exactly. Or like peak human. The baseline of our data set is peak human performance. Yes. Yeah. Okay, so that's the agent. So now what I'm going to show you is we then are able to take those action predictions and we're able to label any video on the internet using those actions.

5:40And so this is just frames in, actions out. Yellow is the model prediction. Or sorry, yellow is ground truth. Purple is the model prediction. And then bottom left is compound error over the entire sequence. And then this is reset per prediction. Reset meaning every now and then you reset? Yeah, so this just means it resets the baseline. And so this basically, a single error in the entire sequence compounds here, but it doesn't compound here. if that makes sense. So, and again, this is just seeing frames, right? It's not seeing any of the actions. And so, you know, so what we did, right, is we trained it on less realistic games and we transferred it over to a more realistic game.

6:24And then, and this is where it gets really exciting, we transferred it over to a real world video, which means that you can use any video on the internet as free training. What was it predicting? It's predicting it as if you were controlling and using keyboard and mouse. So if you were basically playing this sequence as the human. Is there some sense of error? So that's why you transfer to more realistic games first. And then you transfer to real-world video because you can't get a sense from ground truth from the real-world video yet. Let's see. And then... So I'll show you here. This one is also...

7:02This is the same agent that I just showed you. This is playing against other AIs. This one's playing against bots, yeah. The previous one was against players. But with the sniper, it doesn't really matter that much, let's just say. It's like... So one thing that's really interesting is you notice that it behaves differently as it has different items, right? That makes sense. Yeah. Yeah. I think there's also a question about egocentricity versus the third person. Doesn't matter? The third person I think will be very, very helpful if you're, for instance, trying to control multiple objects in an environment later on.

7:45Right now, I think having fully in perception first person is quite helpful. This one's also, this is the policy itself. What do you mean this is the policy? The agent. Yeah, same for the strengths that I just told you about. Yeah.

8:02Like this, where it hides, that to me was just incredible. Just from knowing, being able to predict. The appearance also hides when you see it. Yeah, exactly. And it needs the spatial intuition to go, well, this is hiding and that's not hiding. Exactly. And right while it was reloading, yeah. Okay, so that's the policy and this is a completely general recipe, meaning we can scale this to any environment. is this work closest okay now let's keep going on demos until um i was gonna go into research yeah yeah sounds good um okay so and then this is this is these so what i'm about to show you are the world models um there's a few really really interesting parts about our world models so the first is uh we actually made a decision to uh transfer um sorry we made the decision to um pre-trained world models from scratch, but also we've actually been able to fine-tune open-source video models to get a better sense of physical transfer.

9:08And so one of the things that you'll notice here is our world models have mouse sensitivity, which is something that gamers absolutely want, right? So you can have these very rapid movements, which you couldn't do in any other world model. And so this is a holdout set. So this clip was never seen before at training time. As you can see, it has a spatial memory. This is about a 22nd-ish generation. And here's what's fascinating. This is an explosion that occurs. And you can see that in the physical world, right, the camera would shake. And in the game, that would never happen. So you see the world model inherits the physical world camera shake, but the actual game never does that.

9:49Which is sort of, that to us was quite fascinating, right? Also, the models that I just showed you that we used to transfer over from video. the two of those combined will allow us to like push way beyond games in terms of training um this is another interesting so this is a world model this is rapid camera motion so like again this is stuff that we're literally just taking one second from here in the context and the actions and replaying it here right um and so you'll you never essentially have um uh like what we're saying is the skill that you see in the clips that like the speed and the movement that also pays off at training time when you're doing world models.

10:25This is my favorite example. So this shows that the world model is capable of performing with partial observability. So what you're going to see is, again, you're replaying the actions from here and here, just using one second of video context. Everything after that is completely generated. So what you're going to see is the model is going to encounter, in this case, smoke. Normally now models break down. what you actually see comes out in the same place. And so it's capable of, even with partial observability still maintaining its position in the world. And then here it is also interesting. So this is swiping.

11:05So this gives you like a... Reaction time? Like the fact that it can do depths and like sequences in completely different views, right? So this is a completely different view than if you were to be outside of that view, right? And so it's able to maintain consistency. While zooming in. Yeah, exactly. And so you can see. So even while this goes out of scope, right? Watch. And then it comes back and you'll see it's still there.

11:39Yeah. And so this is the work that Anthony Hu has been working on. I'm just wondering how much gear footage you have to watch in order to find these things. we can ask Anthony I'm sure he's not going to be too excited to play these games afterwards you're not playing it, you just want to great so those were the models these are interesting so we also were able to distill into really really tiny models so this is for instance a long sequence on a very very tiny one you can see it makes a bit more stupid mistakes. It does things that are not as optimal. I haven't seen anything yet. At the beginning, it was running into a low for free.

12:26Exactly. I mean, I do that too. Yeah. It's doing pretty well. Yeah. And again, all these models are running completely in real time. I was thinking, your main model does real time anyway. What's the goal of distilling? Is it cost? Yeah, parameters. Yeah. Yeah. Yeah. This is the interesting one. It peaks the corner. That's what we mean by like the spatial and poor reasoning aspect is humans actually, they sort of simulate the optical dynamics of their eyes and how to actually spatial and reason. The teachers have all the data, right? You've seen all this. Yep. Exactly. And so like even in like real, this is kind of interesting, even in like the real world with, for instance, YouTube data, right?

13:10You have to first solve for pose estimation. Then once you have pose estimation, maybe you do something like inverse dynamics, right? where you basically are able to like somehow label some of the actions that you're seeing. And then you still have to account for optical dynamics of like where your eyes actually looking before the decision because like there's just three levels of information loss. Or when you're playing video games, you're actually simulating the optical dynamics with your hand, right? And I think that's why I think why games are a better representation of switch support reasoning initially than YouTube videos, for instance.

13:42Okay, we're in the GI offices with the CEO Welcome, the way you're welcome. Thank you. Thanks for having us in your office. Yeah, excited to be here. If I'm in New York and you're one of the hottest races of the year, I have to come and visit and thanks for taking some time on the weekends. Yeah. So you've raised 133 million C. So general inflation. Most people don't care about you. I guess the GI is new, but more gamers would have found a middle. Indeed. And before that, you ran probably Waste Lake Summer. Yes. The largest Waste Lake Summer. um what's your reflection on just that journey of life now you're an AI founder yeah you started off like Rootscape yeah I think um so I grew up with Tourette's uh I uh spent most of my time as a teenager coding and playing video games uh so in that sense it doesn't feel that much difference um but I think for uh so yeah so I started the largest privacy of Rootscape worked at Dr.

14:40Subwriters for three years first in Ebola and then on like satellite satellite based map generation for disaster response, which was already like very AI related adjacent. I built some models back then and then started Metal, which became one of the largest social networks and video games. I've always been kind of like AI, like Jason, you know, I'm a self-taught engineer. So for me, the modeling itself always felt a little foreign. I actually had to take a ton of tons of classes over the summer and early this year to get better at it because it still felt like I was really, really good at the infrastructure side.

15:16And I had written our transcoders for Metal myself. So I was very, very familiar with CUDA and the GPU side and all the video infrastructure that we were using for this stuff. But the modeling side itself was still quite foreign. Luckily, obviously, I have really, really good co-founders. But they essentially put a bunch of coursework together for me to go complete, to get really, really good at understanding the fundamentals better. I think for me, I had seen inside of the labs that I had really, really good leadership with fundamentals on top and also the ones that didn't. And I think the ones that did were just like much better.

15:49And so for me, yeah, I wanted to be more like that. So in that sense, it was a bit, it was first very foreign. And then now I feel pretty comfortable with everything. And but yeah, like, I think for there's a lot to be explored starting in video games and also reverse engine. And I think the interesting thing about reverse engineering is it kind of teaches you to look at problems very differently. It's like the ultimate form of deductive reasoning in a way. And so this is just how I think, how I operate. And so for me, it's been a really, really interesting journey. I don't claim to have any of the credentials or skills that some of the other guests do have add-on, but hopefully it will make for a good time.

16:29Yeah, well, your co-founder is definitely bringing a lot of that definability. And you bring a lot of the, I guess, gaining expertise. mostly when i bring to the table just just a little bit of history of metal and bold let's establish metal for those who don't know uh the lady uh twix yeah the year yeah um that's you have more active users concurrent users in twitch something like that yeah on the creator side i think and the reason is because metal is a lot more like instagram than it is like twitch so people um so the way the way to think about metal is it's it's a native video recorder like unlike something like twitch where you actually have to use other software to record and stream to twitch um it's not a streaming software it's actually a video recording software and a lot of gamers love to put things like overlays on top of their videos um and as a result of that we have sort of the largest data set of ground truth action labeled video footage on the internet by maybe one or two orders of magnitude yeah what was an example of an overlay play the only overlay i usually think of is like the case cad yeah yeah also um controller overlays for instance if you're playing um like let's say you're playing uh console yeah like flight simulator you get like you know the joystick and all the things so you get the actual actions that people take inside the games as well as the frames of the games themselves which is a loop right because it's essentially you perceive then you act and there's a state update and then you perceive again you act state update which is like roughly precisely what you use in order to trace to train these agents yeah it's it's almost perfect training data we were showing you were showing me in the demo and we've shown some b-roll here on uh how you don't love key it's very important to log action yeah when did you figure this out well um maybe starting a year and a half ago yeah and we realized that like figuring out the side of the research for us was we very much never wanted to be in a position where we eroded privacy or something like that so we never wanted to actually log like a w or a or s and a d which for researchers the fact that we don't do that like often it sounds strange like why wouldn't you do that but i think for us the privacy yeah i i think you know a lot a lot of the the um the researchers they did they hadn't quite understood yet that you can actually just get away with just doing the actions um and the reason is like at training time having the actual keys is noise anyways like if there is text on the screen and you would want to in theory uh make that um part of training then like reading text from a frame is like really easy and so for us if we actually can so we convert basically hit you hit the input, we convert it to the actual action.

18:59So we had thousands of humans label every single action you can take in every single video game over the past year and a half, which is an enormous amount of action labels. Yeah. So when you act, we get the actual action itself. And then it being said at training time, you can for like the general set of that game, convert back into computer inputs if you want to, but you can never do it for any individual person. And so that for us from like a design perspective was important. So we figured all that stuff out. Then we actually started pushing, like we already had features as well with this. So for instance, like gamers already love to be able to navigate their clips by like things that happened.

19:38So we have an events capture system. And then we also have the overlays where you actually just want to overlay and render the actions on top of your clip. We developed kind of in tandem with the feature set itself. And then obviously when World Bottles became a thing, and it's very, very clear that all the data for this was precisely like that sequence. Yeah, we were able to sort of be first to market, recruit the best researchers and start a lab. Yeah, that's incredible. One more question on Metal before I remove photos of the DI. It's been 10 years. Yeah. What is the, I don't even know how you brought something like this.

20:10I'm just kind of curious, and the opportunity to ask you, what really worked? Yeah. That you became so huge because you're not the only one. Yeah, but I'm sure it's performance and everything, but... A few things that really worked. I think the first was a lot of our competitors were focused on solving the social network and the recorder at the same time. And that never, like our bet was really that we could get so many people to record with us that we could bootstrap the network on top of that. And that worked. So while everyone was sort of distracted trying to bootstrap a social network, we were just focused on building a really, really good capture tool.

20:43And then we got tens of millions of people to use that, which then we were able to bootstrap a network on top of the share behaviors. we already had like the profile behaviors and the share behaviors obviously but the actual content consumption piece and the sharing piece really only came after we hit critical mass it was actually early days during covid when like the network really accelerated fortnight happened which was really important and i think also the fact that discord existed um made it quite a different time than uh when other types of networks of these types had launched because discord essentially was like the connective tissue already between gamers that like never really existed before And so I think those combinations of things really, really made it.

21:19I think we also built a product that, for instance, with most video recorders, you have to remember to start and stop the recorder. So you have to go into the application, then hit start, then start your game. And then, you know, maybe you'll play games for three hours and you'll close the game. Then you have to close your video application. Then you have to process like a multi-gigabyte file. Then you have to upload those somewhere. And so like this was a pain for people. and so what we did is we just ran this kind of recorder when you hit that button it does a retroactive video record so all the recording initially is in memory and then when you hit that button it exports only that sequence to disk and syncs it to your phone and so that that became super popular it also what was interesting about it also means that you're not sort of behaving or acting differently because it's always there and you can just export whatever happens which is also very very helpful for for trading obviously um the thing you went the first to do that Yeah.

22:10The thing you were explaining just before this was it's similar to how Tesla does the bug reports. Right, you're driving and from the having disengaged autopilot, they're like, well, tell us what happened. Exactly, exactly. See, you're driving, Tesla doesn't want to train on the like 10 hours of you driving through a desert where nothing interesting happens. You have the clip button on the steering wheel, something interesting happens either while FSD is engaged. And I'm not sure if you can use it without FSD as well. But you hit the clip button and it basically uses that precise sequence to mark which is then more helpful for training because it's more unique as a training time yeah yeah i mean so one thing i mean we're going to introduce on the inside one thing that i that does pop up as well a lot of life is boring a lot of life is going from a long black a lot of playing games is doing the boring stuff that is not capable somehow using the generalized fight yeah yeah yeah it makes you think right it makes you think yeah yeah it's also quite interesting like i showed you the models like what happens when you increase the size of the context window and how behaviors actually are largely shaped by the size of the context window.

23:15That to me was one of the most interesting parts about the research. Made me think about our own behaviors in a way. Let's talk about also forming a team. On your website, you have 12. I don't know if that's changed now. I meet four, three co-founders. And let's talk about how this team comes together, because you may not visit yourself. You don't have that at the end of the network while you manage the LMS people. Yeah. I started reading all the research papers. By that time, I was already pretty deep into having a decent understanding of not world models, in particular LMS and transformer-based models.

23:53And so there was Genie, there was Sima. Those two were really, really interesting. And Sima in particular was interesting because what they do is they basically take 10 games and then they have a graphic in Sima where you can see kind of the precise actions that are inside of those games that they mapped and I believe they found something like 100 which are actually actions that also exist in the real world and what they did was they didn't I believe it was specifically for navigation they did a 9-1 holdout set so they trained an agent on the 9 games and then they had to play the 10th game the holdout game But then they also trained a specialized agent just on a 10th game and they compared how good they did.

24:35And if I recall correctly, it did roughly as well playing the 10th game on navigation specifically on the holdout on the nine game agent than it did on the one game agent. And that's what was really interesting because that's precisely the type of data that we had. Right. And so for us, the thinking was, OK, what if we did exactly what LLMs did? What if we use, right, this, right, so LLMs for training on predicting, like, text tokens on words on the internet? What if we predict action tokens on essentially what is the equivalent of the common crawl data set, but for interactivity? Vision input?

25:09Yeah, action output. Correct. That's it. Well, I think, well, actually, I'm going to double back a little bit to, like, a question I had, which is, One of the reasons why I thought you would want to prefer keyboard and mouse over actions is the action series is potentially unbounded, right? You can jump, walk left, walk right, but then also look up, look left, flip bench. It's unbounded. So it's huge, isn't it? Yeah, there's benefits to the action space being small to start with. So I think we're going to start with anything you can control using a game controller. but yeah long term we want to actually predict maybe like action embeddings and have models sit inside a general action space to be able to transfer out to other inputs as well yeah okay and then let's basically going on the research side so uh genie sima yeah and then the co-founders yeah so there was the diamond paper there was genie and then there was sima the diamond paper for me was really interesting because they had actually managed to get this world model called diamond running on a consumer GPU, I believe it was a 4090 at 10 FPS, and you could play it.

26:14And they did that on like 90 hours of data, like 95 hours, I think it was 87 hours. And I think eight and a whole that set or something like that. That was just incredible, right? That they had something playable on that little data. So I actually cold emailed the entire group of students and I told them, hey, I think we have this thing. And then it was pretty interesting. So like right when that happened, a lot of the labs also started understanding what we had. And so we started very aggressively, multiple labs tried to bring us in in various ways. And they were part of that, like they basically were seeing that happen.

26:47And I think for them, that also kind of like solidified how real it was. And then when we chose to do our own thing, you know, initially, we thought that we were going to have to just work on world models, right? So we thought, okay, the main benefit of this data set is like Gini is world models. What we didn't realize at the time is that we have so much of this data so we can essentially do these role models in parallel and take the equivalent of like the LLM bet mostly on imitation learning and then use the role models after that to get into like our off stage, right? And so for us - And eventually get rid of the role models.

Read the full transcript

27:17This is something like - I mean, ideally you get rid of the imitation, yeah, the imitation learning, but yeah. We essentially realized that we could get so far on just imitation learning. The way to look at it is we essentially, like let's take the LLM analogy. We essentially have sort of the internet or like common crawl, if you will. And every single lab is trying to simulate that, right? In order to get similar data, in order to train their agents. And so for us, the reason why we say independent and we just said our own thing was we think we can essentially leap every single company that's forced to either be consumers of world models or build world models and take this foundation model bet for spatial-to-bore agents and be in a place where we have a lot of customers years before any of the labs even get there.

28:00and maybe the most similar um comparison is like when anthropic did with code right anthropic just focused really really hard on nailing the code use case their models are incredible for a lot of their customers use it for it so we just want to become incredible at this spatial temporal agent use case and likely that starts in like game simulation and then using world models we can then start expanding out to to other um areas so would you show me a little bit of how you think does generalize our yeah things um but although games is kind of the common pair yeah games and simulation um i would i would specify it as game engines and verticaler so even if you're for instance uh simulating human behavior in omniverse because they're trying to create better training data for factory floors um you can use it yeah maybe meta has a similar data set because of the quest i never really asked them i never really looked into the meta quest specifically so you need a few things you can't just like there's lots of companies that have like maybe recorders but you also need the public graph otherwise you can't train on the data right you can't train on people's like private videos that they have saved somewhere right and so i think you you you need the social network graph components um because these videos need to be on the internet to rank no to train on them yeah i i mean i think i think generally people don't like people don't want to train on like because these things they live on your device usually right yeah um and you can't train on anything that lives on your device like you actually need to go and upload and do your thing right for meta specifically i think also vr the scale of vr is still pretty small The amount of environments in VR that have consumption at scale is probably in the hundreds, whereas on PC, it's probably in the tens of thousands.

29:38And so you get a lot less diversity. The three-dimensional input space of VR is pretty interesting. We see some of this too, obviously. And so, yeah, I do suspect Meta starts using these types of things, but it's unclear to me whether they can get to a similar scale of data or diversity on the environments as we can. Yeah, there's a lot of challenges there. Okay, I want to take this in a few different ways, but I guess let's fill up the papers. Maybe one more to mention is Tire. Yeah, which I actually interviewed at our office, but that too seems like the particular insight that brought it overseas.

30:15Yeah, so Anthony Tu, who led the research on Gaia Tu, is also one of the engineers that joined our team. So it's all the Diamond, the core contributors for Diamond, and then Anthony. And we just had three more researchers showing this week. It's been a good week. And yes, I think a lot of the approaches in Gaia 2 were heavily inspired by Diamond. And then Vin Sa, who was one of the authors of Diamond, also already was at Wave by the time that I emailed them. Anthony also realized what this was and realized that, you know, you could scale world models to a much larger scale and decided just to make the leap as well.

30:48So I think everybody that sees the dataset makes the leap because it's, but it takes a while to wrap your head around because it's like, oh, it's video games, right? Like intuitively, it doesn't make sense. And then when you actually understand and you see, right, how we've been able to transfer it to physical world video and things like that, then it makes sense. And then everybody tends to jump. Don't call it video games. Follow it on a lot of times. Yeah, if I lived in San Francisco, maybe I would. Just a quick note, because we actually cover all these papers in the latest day Super Bowl Club.

31:19SEMA 2 did not seem to have as much intact than SEMA 1. I don't really know why they did it a lot more work. g3 had a ton of impact and but i also felt like because you couldn't play with the model or people it just seems they're an extension of all those days i guess like any quick takes on sima 2 g3 which were both these years yeah i'll talk about sima 2 the steerability of sima 2 was to me the most impressive part because lighting up the action sequences and the text conditioning is quite hard to do, right? And so that, and the fact that they were, like, it's also quite interesting that it means that they can sort of use Gemini as part of the flywheel, right?

31:59Where you can sort of scale this orchestrator as, like, an independent, almost like a puppet master, if you will. And then, like, in theory, Gemini could orchestrate many instances of SEMA, right? That, to me, is the most interesting part, is where I tend to agree with this, where, like, I think our models will initially be used as, like, you'll have like an orchestrator vlm of source that's kind of like managing instances and instructing them um and i think for sima showing that you can do this was was fascinating also the fact that you could um they didn't just have text conditioning but they also were able to do like drawings and markings uh of where to go they really took an interesting end-to-end approach to me uh that i look forward to seeing a lot more of um but you're talking to them like you said it is that the one collaborative yeah i i think that um yeah we're very friendly with deep mind we like them a lot i just saw the team not too long ago and i think you know big fans of their work the different line that i kind of shade from alice's coverage review yeah is uh you are the biggest bet that and they don't cross-line as many since open ai yeah how did that conversation start okay so for now it's style and maybe i'll get slapped in the fingers for revealing this or whatever but uh forgive me if i'm super bad um is he asked you to like draw a 2030 picture of your company and i think he just picks n plus five years whatever i don't know i did the same to you yeah um he asks you to like walk that back from first principles all the way from today and and and yes he expects you to do that flawlessly where he can challenge any assumption any part of the vision that that and he asks you questions right he has a very technical background he also has a bunch of technical people want to see and he truly backs people that have these like very large visions on that vision and the ability to defend it alone um and that's what he did for us um and i think that's why i made that bad so i think also through this uh through this question he he gets to know a lot of things about how technical you are he gets to know how well you think from first principles because if that i if that vision is not connected to something real it's very easy to suss it out by asking good questions.

34:10And then he just backs fully, I think. Like he really gets in your corner if it's the right fit. And yeah, they've been incredible partners. They've opened so many doors for us. I had to ask the question, I think it's a very notable story. Obviously, a lot of work went into it, but it's also worth it. Yeah, for sure. One of the things I also wanted to, I think I asked this question out of Sequence, but one of the things that's exciting about Telling to You is there are a lot of people like you who are founders of businesses and businesses that along the way have a ton of data. And yours happens to be highly valuable.

34:51You pursue, before deciding to do an independence journey, you can also talk to other companies about potential licensing or acquisition and stuff like that. What is your learnings from those periods? Also, one version of this is very simply, how do you value the air? Yeah, I don't think you can value it unless you actually model it yourself and see what the capabilities are. That's my real outcome. You say model, but chain the model. Yeah, but that's obviously not doable for everyone. And also I think my general advice would be as model capabilities increase, and models are also like, these VLMs, they're very, very good at labeling as well, generally, right?

35:35What I was afraid of when I was having some of these conversations was, okay, as the capabilities increase, you're just going to need less ground-free data and you can do more model-based data generation or synthetic data generation. I would recommend if you're going to do large data deals, just try to get a large chunk of equity in the company that you're doing it with, if you can. Now, a lot of them won't do this, but I think that to me would... Or just go do the research, figure out what's actually possible. In our case, we were quite lucky in the sense that this is actually the foundation data right and i think right like that's not true for for every data set i think you know we just happened to hit a particular gold mine but you also did you read kate brady you did the action thing one for five years ago yeah so you eat it work yeah that's the thing like you you have to be grounded right and i think a lot of the um and and i think that's a hard part and i think a lot of what's interesting is you can also kind of look for if like scaling laws already exist on your data type which like for video there were some but for these like input action labeled uh sets there there really wasn't any the other question is like does it go into lms does it go into uh world models does it go into like what type of model is it going to be used for and i think that's an important thing to know and so i just want to you know if you're having these conversations with labs about data just like make sure that you actually understand like what it's going to be used for because that's a very very good way for you to like make that decision yourself about whatever you want to pursue that now a lot of them won't tell you that and i think you know i think in in in that case you generally just don't want to do it because like i think i think for our case like we really cared that like for instance there weren't going to be competing products with game developers built right because we didn't want to like bite the hand that feeds us and i think we are part of the games industry so those questions i think are normal and then we eventually decided you know you just have the data we're just going to go do it ourselves and that's when the rest happened yeah and he assembled the team i can uh think about that i i feel like that's you've aligned a lot of stars in order to make gi happen yeah that other data founders they're at the beginning of the training yes one data founder founders who happen to have data but they have a main business right i don't know if you're there there's two sides to this right there it's really easy to be super naive about it and like i had a lot of people you're just like making this up and and and so for me like doing the work and actually understanding it myself was a really really big part of building that confidence and go start a company but a lot of times it is true that like model capabilities increase so quickly that like the certain data you just don't need anymore yeah um and so i think it is it's really important to like get people to do the work such that you can make these types of distinctions yeah and and and so my recommendation would be go build models with your data see if you can create any sort of capabilities that aren't clearly already there or on path to being there and then figure out where you go.

38:31Yeah. I did want to ask this earlier, but you're giving the opportunity to, we say do the learning, do coursework and all that. And your co-founders gave you some homework. Yeah. Is this like some books? I mean, Coursera? No, this was Francois Fleurais. So he has a little book of deep learning and then he also has a full course that he's published on his website. I went through the entire course over the summer i believe it's like something like 30 or 40 lectures which also take home projects and things like that um and i would recommend anybody uh uh does this it it goes through right history of deep learning like the the topology it takes you through um the literary algebra the calculus eventually end up with like chain rule and by this time you've done like all the the more important concepts it takes you through how do you create neural networks using uh using these concepts that you've learned wow this is super first principles this guy and i've i've had the the uh opportunity to spend some time with him as well he is one of the most first principles people i've met in my entire life i'm convinced like i actually asked him why did you do this course he's like oh because i thought all the other courses weren't right and because because he is so first principles and he can only explain things from like everything you see and how he explains this saying it's everything is from first principles including like the history of deep learning itself was part of of the course and um yes he goes uh um so all so he goes through everything and then uh and by the end of it i think you're like i now have like a pretty good intuitive understanding of how everything works but obviously still right like i like to describe it as um i'm like the the guy who just got his driver's license i can drive the car and like my co-founders are like the f1 drivers that like have done this for years they know where all the um uh where all the the gaps are and and so i i enjoy getting to learn from them the cool thing is also that world models is just like a very very new space and so you know i i got to bring ideas to the table that like you know one thought of and not because i'm great at this just because it's such a new space that like people just haven't tried it yet um so let's get a hit on definition yeah what are world models to you you know in a video model you might predict the next likely sequence or the next most entertaining frame.

40:40What world models do is they actually have to understand the full range of possibilities and outcomes from the current states. And based on the action that you take, generates the next states, right? So the next frame. And so it is a much more sort of complex problem than traditional video models. So to me, it is a world that is accurately generated based on the actions that you take as a result of what's already been generated just a fact check uh that is it needs to understand physics it needs to understand if i'm building a type of material you need the power interacts with some type of material yeah i think the interactions is the most important part i think the reasons why role models are so fascinating one of the things that i did when i was studying over the summer was i tried to actually built a super rudimentary PyTorch physics engine, which I would not recommend writing a physics engine in PyTorch for obvious reasons, but I wanted to be able to, because it's differential, so you can generate this.

41:39It's very difficult. Yeah, exactly. And then you can train. And so I got so many people asked me about why aren't you just simulating or generating this data? And I really wanted to understand from first principles why. And I think the most important thing that I figured out was the compute complexity of simulation goes up really, really rapidly with three variables. First, the numbers of agents in an environment. Second, their DOF, so their individual. Jewels of freedom. Yeah. And then third, the information that each action reveals. so like for instance if you if you have a if you have a text action or a speech action the environment can change so much based on whether you say right water or fire that the outcomes are going to be completely different of like how a human would behave in that type of situation and so it goes up so quickly with those three variables that at some point you just hit a point where you just want to maximally bet on either video transfer or generation of these environments using world models because that type of stochasticity is just incredibly difficult but it's already very, very present in a lot of the video pre-training that goes into these world models, right?

42:48And so I think for us, it is more so about making a maximal bet on video transfer and interacting with things that are difficult to simulate. And the steerability is also really interesting with text than it is on betting against simulation or something like that. And so I think there's still a large market for traditional simulation engines, specifically in areas where video is really hard to get. is this exactly what the big labs are also saying when they're talking to that i honestly haven't talked about the big to the big labs like since we started working on them ourselves i think people are more reserved with what they share with us yeah of course it makes sense that's probably a question how would you contrast your version of all models with the lead yeah yeah i'm fluent yeah so i don't know exactly what young is doing today my understanding it's based on Le Fee-Jepa, like Le Jepa approach, which is...

43:37So I'll start with Fei-Fei Li. I think what's really interesting about Fei-Fei Li's approach is that you in some way are able to reuse the spots, right, in game engines and in things that let you stay in verifiable domain, which I think is a really interesting approach. However, my understanding is they're currently not interactive, which in my opinion is like the whole point of world models, right? It's environments. They're great environments. And I think from a business perspective, I think they picked a really important part of the tool chain. But to me, that's not really a world model. But my guess is they'll get there, right?

44:10They'll start generating... Yeah, they just have a reuse thing. Yeah, exactly. And I think Feifei is one of the founders of the entire space. So I think it's going to be really interesting to me on what maybe that interactive piece looks like for me to really judge their approach. I think... We reviewed, just before we moved to young uh we interviewed her with justin johnson uh her co-founder he was he was more focused on the physics side of things and the interactivity and they just haven't been instead yet but i i i do think that basically that the splats if you just add more dimensions on i guess the forces acting on them then then you get to attract you to the out of the box because you are basically these are virtual atoms that then has all the normal physics applied to them yeah Yeah, I'm excited to see what that looks like when they actually release it.

45:05It's really hard for me to comment on anything. I really like the frame-based approach because all of our training data is in this format. Yes. We actually asked them about this, and they were like, yeah, it's possible, but we're choosing this data. Yeah, and you can also go from splat to frames, right? I'm sure you can write at some... It wouldn't be easy. You'd have to actually render out the environment, do the edge to the shirt. it's not it's not going to be a simple problem but like in theory it has to be something that you can do if you really wanted to so like because it's almost like having a more sort of grounds for three-dimensional representation of the underlying world yeah right so i think it's an interesting approach um it might be overkill right uh uh you're also dealing with like a much larger like degrees of freedom on the output space right so so who knows how well it scales i like the fact that like i think these video models also use things like auto encoders you can actually have the world models predict like much smaller um maybe like a yeah exactly and then you can use like diffusion upscaling or methods like this to actually enrich and so i think that world models just allow a much more or world models in my sense for a much more like controlled space that that that we know really well um i'm not suggesting their approach is wrong i'm just you know like this is i think what we really like about it honestly, Jan's podcast that he did, I don't remember which one it was, but a long time ago where he, where he basically proclaimed LLMs to be a dead end.

46:32Um, uh, it was one of the things that inspired me to do this. I think this is very consensus around world models. People, basically everyone that has this is like stops with the LLMs and just goes through to world models. I would say that the main perspective, I asked this exact question to Nolan Brown from Obed-Eyye and do us like, well, B, learning visible models. So it's basically the difference in considering our system B. What do you want to put down here, everyone? Yes. So, yeah, I'm not one to proclaim LMs or dead ends, personally. I think they're actually quite useful, and particularly as orchestrators.

47:07The way I think about it is, as humans, we had sort of a three-dimensional world, then we invented text, in a way, a compression method, right? So we invented text in order to communicate with each other in a common way, in a way that actually compresses all this information that we are perceiving in three-dimensional space into just like a single sequence. And I think that allowed science of streamer, it allowed so many literature, like so many parts of the world that we charge. So I think it's a critical part of the whole picture. I also agree that it's very, very clear that they do build sort of the internal implicit world models inside LLM's.

47:49And so I think they'll be very helpful as things like orchestrators. The problem is when it comes to the generalization, I think text as a generalization backbone, when most of the pre-training is text or largely text sequences, then I think you want that backbone to be kind of more split-septorial in nature and then also just have text as part of that. And I think the actual argument of LLMs is also, for instance, the autoregressive nature of the prediction itself. So the fact that it's running the entire output through the transformer, and then in order to predict the next token, which doesn't, like, the environment in the real world is continuous, right?

48:32It's always changing. And LLMs kind of just forget about that, right? I think a lot of the argument is in the first, right? So I think the fact that text doesn't necessarily generalize well to a sufficient apporal context and then the autoregressive nature of the prediction and using text for that, right? So I think those are the two main arguments. I think text prediction is just one of the actions that is going to come out of these policies and world models. I think speech and text generation will just be one of the actions that can be a part of that. I think that there will just be labs coming at this problem from both sides.

49:13And everyone ends up in roughly the same place. And the same place will be whatever people think is cool. Right? Like whatever the consumer is closest to AGI. Yeah. And so I don't think there's like a clear answer. I think it's really interesting to come at it from the world modeling side. But it's also because we have to. Right? Because like text is largely commoditized. We can import all the text. I think it's interesting and tempting. Yeah. A lot of tempting, it makes sense that you can probably recover. It's sort of like you're taking a step back. You're starting your branch of the ML research sheet, but you might actually just end up recovering all the other tech stuff emergingly.

49:52Yeah, yeah. We can import a lot of that research, right? A lot of that is... That's really cool on the research side. Let's talk about the stuff that GIS is producing more, like the sort of research and products output. you mentioned the word customers what are your turn customers yeah so we're already working with some of the largest game developers in the world yeah we're also working with game engines directly and so really what we're doing at the moment is replacing essentially the player controller inside of a game engine so anything that you're currently that maybe like behavior trees or things that you're deterministically coding we hope to replace with a single api which is just you stream us frames and we predict actions and that can be inside an engine or it can be eventually even inside the real world hopefully those are then also steerable so the models that you saw were text steerable yet but i think we want to get to a point where they're fully text steerable but you see steerable means like well i want youtube to share if you're anything else yeah i think it's it's sex conditioning on the generation so yeah the ability to to you're right We want to get to a point where you can generally, and that's why it's called general intuition, where we can sort of can mimic the intuition of all these gamers into human-like behaviors in any situation.

51:08As I mentioned, also, Lab is named after the Demis' office quote from AlphaFold, which is, wouldn't it be amazing if we could mimic the intuition of these gamers who are, by the way, only amateur biologists on his path to, he tried to get an AI to train Foldit to generate a lot of data for AlphaFold. And so for us, really, the North Star, what we hope to get to one day is being able to represent scientific problems in three-dimensional space and then have a space-in-the-poral agent capable of perceiving that space and using, hopefully, also the text reasoning capabilities that LLAMs have today in addition to the space-in-the-poral capabilities to be able to work on the other side of that problem.

51:43So that for us is sort of the North Star. That's why we're sort of trying to be hyper-focused, Spanish and poor workloads, the same way that Anthropic was hyper-focused code, and use that to then get into organizations and expand from there. Just as a side note, since you mentioned Anthropic, any idea what they did on this to solve coding? No, out of any lab, I probably know Anthropic the least. I admired him, though. Yeah, well, the current working theory is that they had a super lucky role of the ducks. But, well, and then he compounds from there. That sounds like a nice story. I'm sure it's not that.

52:21Yeah. Okay. So why do the game developers want this? So if you're a game developer, how well you're actually retaining players, it's like if you have a game that's already at scale, it's decently dependent on how good your bots are. So if you're logging in at an obscure time, let's say 3 a.m. in America, and your player liquidity is low, then you need really, really good bots to keep those players engaged. Is this a thing? Yeah, for sure. For sure, we're like 4D and whatever. A lot of human works with this, yeah. And so if you're like, as a human, do I want to play against bots? Usually it's not just bots.

52:56It's like players mixed in with bots because you don't want to play just against bots. But it's better to have a full game than to have like an empty game. Yeah. And so I think as long as it's part of the environment, I think it's okay. That means you also have to sort of grade that skill level. Yeah, which we can do. because we know exactly how good people are at these games. Yeah, I think for us, bots is kind of like step one, right? So what I was showing you is we're building a general agent that can sort of play any game in real time. But really that extends into all of simulation, right? Like in GTA V, for instance, people are generally role-playing real life, right?

53:30And so they're actually behaving in quite aligned ways with the goals they set for themselves. So you have all these examples represented in video games, right? You have truck simulator, power wash simulator. Power and Wash Simulator? Yes, Power and Wash Simulator, where actually the behaviors that you'd want an agent to be able to perceive, they're all there. Okay. Yeah, it's really like how seriously some gamers take Truck Simulator. If you haven't seen these tips, you should watch it. Yeah. They buy the whole truck driving set and they're doing the job of a truck driver. Yeah. What I mentioned to you, we have more people at any given time on Metal playing with steering wheels and like Truck Simulator and these types of games than Waymo has cars on the road.

54:10it's a ridiculous stat, but it's true. Yeah, I mean, so, you know, I used to think that, well, to solve self-driving, you kind of just interplay along the GTA 5. Yeah, I mean, it's not bad for this. Yeah, our bet is not that we can zero-shot any of these things. It's just that, like, the next self-driving company can maybe collect 1 % of the data. Because, right, also, for instance, Clip's already self-select into negative events and adversity, right? And so, like, a lot of our data set, because it already highlights, is really precisely what a lot of these companies spend their last 20 % doing.

54:44I think that's the main argument. If you're another company that's looking at what we're doing, I think the thing that people won't understand is that anything that you're currently doing in pre-training, as long as your robot can be controlled using a game controller, we hope that we can move that to post-training for you. So our bet is not that we can create the next self-driving car company. It's just that the next self-driving car company hopefully only needs 1 % of the data or maybe 10 % of the data. I don't know, right? To be able to deliver a really good product. Yeah. It's also the term that comes to mind a lot is active learning.

55:14I don't know if you've used to identify with that. It got less cool for a bit. And now it seems like down the uptrend, which obviously you have the best data set for the high intensity. You said negative, but feeling for negatively, it could be negative or part of it. Yeah, for sure. I think negative events is just because it's the most common term that people use for like, if you're if you're tesla you want the crashes you want like right um yeah right right right but but it's only gaming it's both yeah so you know the model that you saw obviously had really really incredible moments and and that was largely yeah yeah that um uh that it had a large representation of people at their best yes yeah and worst yeah yeah yeah amazing okay cool uh and you have anything else on the customer development side that you want to sort of flesh off yeah um we're also already working with robotics companies but again that and manufacturing but the key is that the robot has to have gaming inputs so we're like our bet is not that we can transfer over to like higher doff robots and the keyboard and mouse it's it's really just that we can move the hard work of of pre-training hopefully to post training yeah it's kind of like the foundation model that is a very good basis yeah you're gonna straight you're gonna give us frames and and likely some text or you'll license the model to because they're gonna want to post training yeah our business model is initially going to be an API, like the Anthropic API.

56:33But you also saw, for instance, some of the video labeling models that we've been able to develop. So the goal is for any company to be able to take in their video data as well. And we can create first, obviously, custom versions of the policy for you, the agent. If that doesn't work, then we've already working with a customer that is doing we distill a model and they turn that into a product for themselves. So people can engage with you on the agent level, API level, people can engage me on the sort of model level can you also buy data no all right yeah we don't sell data okay cool so that's the that's the business um and is there a world in which i mean i i think this is on your landing page if you are you know frontier labs for for world models is there a world in which there is a more sort of application layer thing that you that comes out like a chat gpt for whatever yeah you're gonna see us launch a few things on on metal itself that are going to blow your mind uh as a result of this this um this agent so i'll leave the imagination for now if people took a great out you know and yeah on the world modeling side like i think one people underestimate is that metal is already one of the largest you know video consumption platforms as well people watch millions and millions of videos a day um so um world model based entertainment and things like that well it's not like a focus for us right now i think we'll be like on the consumer side we have the ability to move very very quickly here and get it integrated in a way that I don't think anyone else can.

57:59Yeah. You could theoretically do a video gen, like the Sora. What is that? It's a BAM. What's the meta one? Meta and Mules? Not Reyes. Oh, the vibes? Yeah. You could theoretically generate clips that nobody play. But you know it's a device. Yeah, I think for us, the games being so human-centric is like a really big part of what makes it special. like I actually I actually just don't think that would work like one thing that we are really excited about though I'll give you one sneak peek of what we're thinking about is what if you could literally replay any of the clips that you have inside a world model or your friends can play them like I showed you a model that already took part of your clip as a contact since the replay entered at walls but it's also how we go from imitation learning to rl right because like it's part of a research rope app anyways to make every single every single clip on that all playable um so uh yeah who's who is to say that that doesn't apply to just the actual clips that you take yeah yeah just can you see more with the rl potential we describe metal as as the episodic memory of humanity and simulation so when you take a clip really the way to think about it is you get the highlight of what is maybe three hours of playtime right you maybe get like two to three minutes of the things that were the most out of distribution right it is genuinely your episodic memory um of that playtime and simulation the things that you most want to remember and share we want to be able to load uh and this is the work that anthony who is doing the reason why we built world models is every crash that you run into in your truck simulator or American truck simulator or a driving game, we want to be able, right?

59:33And again, these are ground truth labels. So we know precisely the actions that lead up to the negative events. They're also title labeled when people upload it onto their platform, they say, oh, good, it's a crash, right? And so we can select all these events. And if we can put them inside a world model, we can go into, right? We can train reward models to then reward based on how you perform in clips that actually contain negative events, for example. And so for us, it's very much about, right, we can create this LLM moment on being an invitation learning, but actually making every single clip on the platform playable at billions of clips scale is how we go from invitation learning to RL.

1:00:09Cool. We covered a lot of it. Is there anything else that you want to do before we wrap up the Lulz and Vision stuff? Yeah, I think for us, this is a very, very ambitious long-term vet. We need the best researchers in the world that want to work on this stuff. It's really exciting not being extremely data constrained. Like we really get to, like we get so many learnings every week that we didn't think were possible and it makes it for a joy working here. Also, the other thing is because we have such a large data mode, we don't have to be as concerned as the LLM companies about publishing because we don't need the ones to be able to.

1:00:44Exactly, no one can replicate the models, right? And so for us, we really want to bring back the original culture of open research, which is why we did the partnership with Qtai in France. I actually didn't. Yeah, we just announced our partnership with Qtai in France, which is an open science lab in Paris, one of the best research labs in the world. Eric Schmidt, I believe, funded in addition to some French people. They are essentially acting as the partner that's currently doing a lot of open research on the data. We also want to partner with universities who, because like we do believe this is the frontier, but it's so data constrained that really everyone has their hands tied behind their back right now.

1:01:21And so we want to help fix that. So for instance, we want to work with universities to build like negative event prediction models for maybe like trucks in India on all the truck data where all these crashes occur. We have all these things that we know we can do that we just have it at the time to do. And so if you're listening to this and you're maybe an academic institution or something and you want access to some of this data in an educational research fashion, I think we're quite open to doing that because we want to educate people. And yeah, and other than that, we just want to work with the best infrastructure and research engineers on the planet as we're going into scaling, you know, runs that have thousands, thousands of thousands, eventually hundreds of thousands of GPUs.

1:01:58Yeah, yeah, amazing. I primed you this as the closing question. It's a little bit of a no cost of that, I didn't know. So what does GR become by the way? Yeah. In 2030, we want to be the gold standard of intelligence. And any sequence long enough is fundamentally spatial and temporal, right? Which I think is... So by nailing spatial and temporal reasoning, you go after the root killer problem of intelligence itself. What the world looks like is we want to have eight. So I sort of group the sequences of AI in three stages. And I credit Andre Goparty for teaching this. Bits to bits, atoms to bits and bits to atoms, and then atoms to atoms.

1:02:36In the atoms-to-atoms stage, I want GI models to be responsible for 80 % of all the atoms-to-atoms interactions driven by AI models. And the reason for that is because we were able to unblock intelligence so quickly, and robotics like intelligence is the bottleneck, that supply chains actually converged on gaming inputs as their primary input methods. And they converged on essentially simpler systems that let us do a lot more, a lot quicker. So we are essentially the 80 % market approach. And then you have lots of companies that have kind of like specialized, maybe humanoid robot OS stacks that are the other 20.

1:03:11And then so I want to be responsible for 80 % of all the atoms-to-atoms interactions driven by these models and be the goal center for intelligence and maybe 100x more in simulation. Because I think simulation will actually be the larger market initially. So I think in simulation, because you have very little constraints, also from a safety perspective, simulation is much easier. so i think a lot of the takeoff initially sits in simulation so a lot of the simulation use cases like what i mentioned scientific use cases i'm really really excited about and so um yeah 80 of atoms-to-atoms interactions uh coming downstream from these types of spatial liberal foundation models and then 100x more in simulation yeah it reminds me a lot of that uh what mark and from this chas zekhberg is to are doing with virtual biology because you can do a lot of simulation and you can do yeah or you can do it a lot faster with interest um amazing thank you for inviting us to your office yeah and thank you for sharing a little bit while you're turning thank you yeah

From the publisher

From building Medal into a 12M-user game clipping platform with 3.8B highlight moments to turning down a reported $500M offer from OpenAI (https://www.theinformation.com/articles/openai-offered-pay-500-million-startup-videogame-data) and raising a $134M seed from Khosla (https://techcrunch.com/2025/10/16/general-intuition-lands-134m-seed-to-teach-agents-spatial-reasoning-using-video-game-clips/) to spin out General Intuition, Pim is betting that world models trained on peak human gameplay are the next frontier after LLMs.

We sat down with Pim to dig into why game highlights are “episodic memory for simulation” (and how Medal’s privacy-first action labels became a world-model goldmine https://medal.tv/blog/posts/enabling-state-of-the-art-security-and-protections-on-medals-new-apm-and-controller-overlay-features), what it takes to build fully vision-based agents that just see frames and output actions in real time, how General Intuition transfers from games to real-world video and then into robotics, why world models and LLMs are complementary rather than rivals, what founders with proprietary datasets should know before selling or licensing to labs, and his bet that spatial-temporal foundation models will power 80% of future atoms-to-atoms interactions in both simulation and the real world.

We discuss:

How Medal’s 3.8B action-labeled highlight clips became a privacy-preserving goldmine for world models

Building fully vision-based agents that only see frames and output actions yet play like (and sometimes better than) humans

Transferring from arcade-style games to realistic games to real-world video using the same perception–action recipe

Why world models need actions, memory, and partial observability (smoke, occlusion, camera shake) vs. “just” pretty video generation

Distilling giant policies into tiny real-time models that still navigate, hide, and peek corners like real players

Pim’s path from RuneScape private servers, Tourette’s, and reverse engineering to leading a frontier world-model lab

How data-rich founders should think about valuing their datasets, negotiating with big labs, and deciding when to go independent

GI’s first customers: replacing brittle behavior trees in games, engines, and controller-based robots with a “frames in, actions out” API

Using Medal clips as “episodic memory of simulation” to move from imitation learning to RL via world models and negative events

The 2030 vision: spatial–temporal foundation models that power the majority of atoms-to-atoms interactions in simulation and the real world

—

Pim

X: https://x.com/PimDeWitte

LinkedIn: https://www.linkedin.com/in/pimdw/

Where to find Latent Space

X: https://x.com/latentspacepod

Substack: https://www.latent.space/

Chapters

00:00:00 Introduction and Medal's Gaming Data Empire
00:02:08 Live Demo: Vision-Based Gaming Agents
00:05:31 Action Prediction and Real-World Transfer
00:08:41 World Models: Spatial Memory and Physics Understanding
00:13:42 From Runescape to AI: Pim's Founder Journey
00:16:42 Building Medal: The Retroactive Clipping Innovation
00:17:59 The Data Realization: Actions Over Keystrokes
00:23:39 Research Foundations: Diamond, Genie, and SEMA
00:32:59 Vinod Khosla's Largest Seed Bet Since OpenAI
00:34:55 Valuing Data and Turning Down OpenAI's $500M Offer
00:38:38 Self-Teaching AI Fundamentals: The Francois Fleuret Course
00:40:28 Defining World Models vs Video Generation
00:41:21 The Physics Engine Experiment and Simulation Complexity
00:43:23 World Labs, Yann LeCun, and the Splats vs Frames Debate
00:46:59 LLMs as Orchestrators: Text as Compression
00:50:08 Customer Use Cases: From Game Bots to Robotics
00:51:25 The North Star: Foldit for Scientific Discovery
00:54:55 Business Model and Open Research Philosophy
00:57:25 Medal's Secret Weapon: Making Every Clip Playable
01:01:59 2030 Vision: 80% of Atoms-to-Atoms AI Interactions

More from Latent Space: The AI Engineer Podcast

All 247 episodes
World Models & General Intuition: Khosla's largest bet since LLMs & OpenAILatent Space: The AI Engineer Podcast
Listen in VO