AI’s Next Frontier: World Models Explained by Christian Keller

27 Nov 2025 · 45 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Beyond The Prompt - AI’s Next Frontier: World Models Explained by Christian Keller

Episode Overview

In this episode of *Beyond The Prompt*, host Jeremy Utley and entrepreneur Henrik Werdelin engage with AI expert Christian Keller. The discussion revolves around the emerging influence of world models in generative AI, exploring how AI learns from various inputs, the importance of multimodal models, and the evolving workflows in research and development due to AI advancements.

Key Concepts Discussed

  1. Information Compression
  2. Text vs. Other Modalities: Christian emphasizes that text inherently compresses information, losing vital details that images and videos can convey.
  3. Example: Watching a game live vs. reading about it the next day illustrates varying emotional responses due to the richness of experience.
  1. World Models and Their Importance
  2. Understanding Change: World models enable AI to learn cause and effect, providing a sense of temporality and change that text alone lacks.
  3. Video's Role: Video input is critical as it captures changes over time, aiding in building a more robust understanding of events.
  1. AI in Research and Development
  2. AI Building AI: Keller discusses how AI's capabilities are leveraged throughout the research process—from generating synthetic data to evaluating model outputs, enhancing experimentation efficiency.
  3. Post-Training Techniques: AI can now guide itself in refining outputs, thus accelerating the research lifecycle.
  1. Shifting Workflows Due to AI Integration
  2. Role Evolution: The introduction of generative tools changes traditional roles and tasks, such as debugging and prototyping, leading to new opportunities for innovation.
  3. Individual Workflows: Encouragement to ask, “Could AI help with this?” opens up new possibilities for creativity and efficiency in personal and professional contexts.

Notable Takeaways

  • Text Compression: Text alone fails to capture the full context and temporality of information compared to images and video.
  • AI's Role in Creativity: AI is not just a tool but a collaborator that can inspire new workflows and solutions.
  • AI's Development Impact: The launch of ChatGPT transformed the pace and nature of AI development and user engagement.
  • Model Limitations and Bias: Keller discusses the nuanced nature of AI biases and the need for diverse model architectures to capture varying perspectives and applications.

Episode Highlights

  • 00:00 Intro: Information Compression
  • 01:13 Meet Christian Keller: AI Expert
  • 02:11 Impact of ChatGPT on AI Development
  • 26:32 Challenges in AI Integration
  • 31:15 Bias in AI Models

Conclusion

In this insightful episode, Christian Keller provides a nuanced perspective on how world models and multimodal inputs are shaping the future of AI. The conversation highlights the transformative potential of AI in both business and personal workflows, urging listeners to rethink their approaches to integrating AI into their processes.

For further exploration of these topics and more, listeners are encouraged to connect through the provided links to the hosts and guest, Christian Keller’s LinkedIn profile.

Additional Resources

  • [Christian Keller | LinkedIn](https://www.linkedin.com/in/kellerchristian/)
  • [Full Episode Transcript](https://podcast.beyondtheprompt.ai/episodes/ais-next-frontier-world-models-explained-by-christiankeller/transcript)

Note For more tips and AI tools, check out [Beyond The Prompt's website](https://www.beyondtheprompt.ai/).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So text is by definition already a compression of information. Information is lost. And so using other modalities helps get other information. And oftentimes when you get text and images or text and videos, with images you get the context of the text in an image, in a setting with the colors. When you add the video, you get the temporality of it that you would not have had before and after how things can change, how things can enter the screen or not. And so I think the first point is that LLMs alone, text-based LLMs, will never be able to learn like a human because they're missing so much information.

0:34We've removed by the sheer definition of putting it in words. Hi, my name is Christian Keller. I'm an angel investor, startup advisor, and currently work as a product lead at the Meta Superintelligence Lab. I've been building AI products and models for the past 15 years, and I'm really excited today to talk to you about AI research and building AI products. Okay, so one of the things, Christian, that Henrik and I love to do is invite our cool friends to introduce to each other. So Henrik has very little introduction. Why don't you tell Henrik and also listeners who may not know you, why they should be interested in this conversation today?

1:15Well, I've been working in AI for the past 15 years. So I'd like to say that it was like before it was cool. And I've taken a pretty deterministic approach to my experience in it, where I've started with building AI products myself, like thinking through problems that I knew and leveraging AI in order to apply them. And then I had some experience in infrastructure around my work on PyTorch and understanding what it takes to actually deploy AI at scale. And at the same time, how researchers also were using these tools. And lately, I've been doing a lot of work in AI research, working on Llama 3, Llama 4, world models also, which I think gave me kind of a good overview of the three kind of main categories I see of the AI ecosystem and what it takes to build successful products in it.

2:04So obviously your career spans long before kind of the chat GPT moment. Maybe we could start with that moment as an inflection point. What changed for you and for the teams that you lead who are building AI systems, what changed for you when ChatGPT came out? And I'm sure there's one element, which is people's awareness of what you do and understanding somewhat of what you do, but then also you and your team's abilities to do what you do in a different way. I'd love to hear about that. So at the time that ChatGPT launched, I was working on PyTorch. And so PyTorch is already quite known in the space of AI and was like...

2:46Could you just explain shortly what that is for people don't know. Oh, right. Yeah. So PyTorch is a framework that's used to build AI models. And so it's the language by which researchers are able to formulate what an AI model should be, what it should do. And it's a translation layer in a sense between these functions that the researchers had in mind with the hardware. And so in effect, when you code normally, you might use Python, right? So PyTorch is on top of it, an additional kind of layer that allows you to build AI models. So for example, it's an open source tool built at Meta initially, but it's used broadly now.

3:23It's part of a foundation. And ChatGPT was built leveraging PyTorch. So it's used in many, many other areas. So what happened to us while on PyTorch, I think what was interesting is that, first of all, people started talking to me about, you know, oh, is AI a thing? Like, have you guys seen it? And it's interesting because from the PyTorch standpoint, we saw a lot of what researchers were doing ahead of time. And language models were not something new. I think the technology had been there for some time. And there were many, many teams that were leveraging them using a key technology called transformers that had kind of changed the space quite a bit at the time.

4:00So we're seeing a lot of innovation already happening. I think the surprising part for a lot of people was how well packaged this technology was in ChatGPT. and how the team that developed ChatGPT was able to make it accessible to a lot of people at scale. So suddenly there was an awareness that wasn't there before of what this technology could do and how well it could be applied. And so I think there was a shift at that time in terms of how people were using it. And why was packaging not a consideration before? I mean, is it the classic engineers aren't thinking about end user consumption or what's going on there?

4:42I don't know, but I think my bet on this would be that oftentimes people are looking at things like assistants. And so you had the Google assistants, you had the Alexas, you had even assistants we're working on at Meta. But each time I think we're considering them more from the perspective of like one use case with one specific deployment. was chat GPT was a way to say, look, we've got this really cool technology. We don't know what people are going to do with it. Let's sort of create the most generic, the most accessible way to do it. And discussion chatting is literally like, you know, the group of society in a sense.

5:14Which probably is a little bit of like a big bet, right? Because I think as a product developer, you would normally kind of like tell people, you need to figure out like, who's the customer? What's the problem? And then kind of be very specific. Like niche down, niche down, yeah. Yeah. So in many ways, I mean, like, do you think they were very purposeful with that kind of understanding that we just this is such a foundational technology? This is more like electricity than TCP IP. So therefore, we just need to kind of like put it out there and basically throw spaghetti at the wall. There's two things that are going on.

5:45I think I don't know what they were thinking about at the time, but it's one of these problems where you've got a really nice hammer and you're looking for nails here. And that's the opposite approach that you want to usually take when you build products, right? You want to find the problem first and then develop the tool that's going to help solve it. So what do you do in this case? Here, I think what was very interesting in their approach is that the chat interface allowed for people to, in a sense, explore what use cases could be interesting. And you've seen some more or less formalization, I think, of some of these use cases in the last several years that were not quite clear initially.

6:23And so I think what they did is they created this learning engine by having all this data from all these users that they were able to kind of capture. And so today you can see use cases that are, I mean, the way I formalize use cases for Generative AI are like one could be UI interface type, something that helps you better communicate with a machine. So, for example, the voice interface, the chat interface. And it's a different way from like being explicit about clicking and typing things. the second thing I see is more around like text generation and so that's initially your chat that could be your your thought leader assistant that could be the editor that's going to help correct your write-up your newsletter or what you want to say so anything that has to do with like interactions text knowledge creation whether it's like with text and images and then there's the kind of automation which is like let's actually do actions and take tasks and entirely kind of not necessarily remove the person from it, but take some chunks of work away entirely that can be either done better or faster, or even if it's not better or faster, maybe cheaper and like, you know, free people up to do other things that are more valuable for them.

7:33You know, what you describe about what the field had been doing in creating narrow assistance versus what OpenAI did when they released JIGPT in creating a general technology is what I've heard called the bitter lesson. I don't know if you all are familiar with that, But it's what researchers call the bitter lesson. And it's been established time and time again that you can work to train a model very specifically. For example, Deep Blue, the model that beat Kasparov, had been trained by chess masters for the purpose of chess. The models that beat Deep Blue aren't trained on chess. They're just told, play chess a bunch of times.

8:08And the bitter lesson basically is the following. any attempt humans have made to train for a narrow task ultimately gets done better by scaling laws of a general model and that's that's been proven again and again that's called the bitter lesson which is to say in a way the it's bitter because there can be a sense of exasperation or fatalism why bother training it for this narrow thing because if scaling laws catch up eventually a generally trained model is going to be able to do this better than what I can train to do it now. It's a little bit also, as you know, I'm obsessed with Kenneth Stanley's book, Why Gradients Can't Be Planned.

8:50He's an AI researcher, turned, sold his company to Uber, then became head of AI Uber, and then went to OpenAI. And he has the same argument. Basically, this book is about how you get to AGI. And his point is that you can't define what it is. You have to have these systems of open-endedness. And then that becomes basically the stepping stones have leaked air and so i think i'm fascinated by that um i have a bunch of questions but one thing i'll say about the chess part is that chat gpt does not beat the current best players at chess so that's something interesting to see which is there is some level of like specialization in in the functions that some models can do and like the counterpoint i think here is that i think there is definitely your point proves that there is value in like building I think these frontier models are very general, both because I think they unlock new capabilities that we were not expecting, and that they allow, I think, for the discovery piece of use cases that we wouldn't even have known would be possible.

9:46But we know today how to take these models and to make them very specialized and even better in a more narrowly, more efficient kind of way to run, which is kind of the other side of the coin here. Is that because that generalism, And obviously the models are statistic kind of like averages in many ways. And often when you want to be kind of beating a chess master, you need elements of originality or differences. And the general models don't do that. And so when you go very purpose is because either you need to work a very specific workflow and you don't really want like the average making that up.

10:23You want something that you know is the best. Or when you sometimes need to have basically outlier answers, you need to kind of train the model to kind of look a little bit off the rapid regression to the mean? There's two things, actually, and you're making a really important point here. The first thing is that I think different model architectures and, you know, will kind of solve different types of problems in some cases. So like the way LLMs are built right now are very much focused on like language or like omni models, maybe various multimodalities, and they're good at doing that. They're not necessarily good at planning initially, right?

10:55That's not what they've been trained for. They've been trained for finding the next most likely token or the next most likely word. So there's a question around which models are best suited for what type of application. But the second point here is that there is something to be said about these frontier models, even when they're post-trained and where when they've actually gone through fine tuning and any other round or, you know, reinforcement learning or whatnot. We've reduced the entropy for them. We've kind of tailored them towards very specific types of answers and very kind of unique approaches or ways we want them to kind of answer some questions.

11:29And so when you take a pre-trained model, there's still a lot of variability in it. There's more variations of the type of answers that they can give you. And so that's an interesting point because you can take that pre-trained model and then specialize it in different ways based on your post-training that you're going to do to make it useful for, I don't know, Maybe knowledge about chess or knowledge about like math theorems proving or knowledge about... Is there a way that users that are just using the models as they are can use that insight to get better outputs of their generalized models?

12:05So, for example, with that thing, if I'm sitting and I'm being asked to make an agent that does HR stuff or brand stuff or whatever it is, obviously you use a rack, I would imagine, or like whatever context window to kind of do some of that. Are there other kind of tricks to make sure you achieve that non-generalizing-ness? I mean, I think it depends like what stage of your product development you're in. Like right now, I'd say most teams I've talked to, like in startups and whatnot, like they usually start with taking a generic model, like the best they can find. And they'll test some prompts and then they'll kind of generate some like system prompts in a sense that will somewhat like help narrow the model usage.

12:51But you're still starting from a model that's been posturing already, that's been kind of narrowed to following the instructions and to having certain languages. So you can move it, but you've got only so much like leeway to kind of move it around. So once you validate that, maybe with your product market fit and see that it's working and you want your tool to improve or your model to improve, then this is when you take a model that's been more like pre-trained with instruction full, for example, and then you fine tune it. And so if you're willing to invest in it, you have a whole strategy you have to figure out around benchmark definition, like data strategy acquisition in order to create models are going to be even better at doing these things.

13:27through that fine-tuning process. The question that I was keen to ask you before we got on today is, I saw a YouTube video of Ian LeCun, who is from Your World. He was making an argument that I kind of understood, but not really. And so I thought I would use the excuse to ask you. He was talking that these language models had built in basically limitations. And so as we're kind of like trying to raise towards AGI, whatever that is, He felt that we needed to kind of start to train on video models or other kind of multimodal models, I think is how I computed it. Is that, do you sound familiar? And can you explain to me why the language models might not be right, the text-based models, and why suddenly having video would make it better?

14:16So I am very much aligned with Jan, I think, on his message. So let me kind of summarize. There's two main points here. The first one is around the modality. So text is, by definition, already a compression of information, right? Like the example I usually give is like you can go and watch a game at the stadium. There's the crowds, there's the smells, there's the game you see maybe not as well because it's further out, you know. But you've got this whole emotion. You've got like a feeling there's a sense of community with the people around you. And then you can like watch it on TV. You get maybe that if you've got a few friends, but you're missing what's happening in the stadium.

14:53You're missing that experience. Then you can read an article about it like in the newspaper the next day. I can tell you you're going to have a very different emotional reaction between like reading the article and like actually experiencing the game in the stadium. So information got compressed. I even tell people like when you read Tolkien, there's like three pages long description of like one tree, right? Like you still haven't touched that tree. You still haven't smelt it. So information is lost. And so using other modalities helps get other information. And oftentimes when you get text and images or text and videos, with images you get the context of the text in an image in the setting with the colors when you add the video you get the temporality of it that you would have had before and after how things can change how things can enter the screen or not and so i think the first point is that is that llms alone text-based llms like will never be able to learn like a human because they're missing so much information that's already we've removed by the same sheer definition of like putting it in words.

15:52The second point, I think, is around, you didn't mention it, but his main point, I think, is around the fact that LLMs are hallucinate, right? They hallucinate by design. Literally, they're just trying to make predictions of most likely things that are coming next. So they're really good at it. And with post-training, we create an illusion of intelligence, right? Because it seems like that's what I would want to hear. And that it makes sense when we having conversation loud that LLM totally gets me, you know, like gets my problems and the like. And so that's actually fine for a lot of use cases.

16:30But when it comes to using LLM for super intelligence or AGI or however you want to call it these days, like you're going to want to solve like really hard problems. You're going to need robustness. And so there's a second concept of world models that Jan pushes for and that I'm a firm believer in, which is we need to create models that can predict a change of state in the world robustly. Not simply say, look, this is the next likely word that was going to be said, but I understand that if in a video I'm holding an apple and I stop the video and then I say, okay, the hand is going to open, what's going to happen?

17:05Well, the apple is going to fall, right? Because there's gravity. We need to understand that if it's easier in the space station, that's going to act differently also, right? And so you need to be able to build models that are going to help increasing this robustness, increasing this level of predictability of outcomes for the models to solve much bigger and harder problems in the world to get to superintelligence. And video is probably the thing that is most available to trade on. But there might be, remember when kind of like suddenly Google put cameras on cars and suddenly there was like street map and it was like mind blown that somebody would actually drive a car through all the streets in the world.

17:43Do you think almost we need to get a point where kind of like you need the Google street car of like the models to go around and kind of like sucking smells and sounds and everything in order to figure out how to map the world? Well, you know, like if you have an iPhone, your iPhone only takes pictures, like has a mini radar for distance in it, like in some cases, depending on the models and whatnot. So you're already getting some of that additional information beyond simply the image or the video that can start getting recorded. I don't know if anybody's using it. But we'll definitely need more.

18:12The videos, what it adds is the temporality. I think that's a really important one. And so it's better to kind of teach cause and effect in the sense if you can see things that happen before and after, for at least in our scale of the universe. But we will need, I think, more. But we don't know yet. I think we haven't pushed the research far enough yet around these JEPAs or world models and videos and others to understand yet if that might be enough to get to what we consider AGI. Okay, so I'm dying to know the other part of our first question, which you could basically describe our interviews as a series of rabbit trails, right?

18:49But going back to the first question, what's changed for you in developing AI models now that we have LLMs? And part of the reason I'm asking this question is I heard an interview with Zuck recently where he talked about one of the things Meta is doing is really using its LLMA models, etc., and tuning it for the development of AI. and I think there's kind of broad belief that in the developer community that using AI to build AI is kind of you know the way to achieve escape velocity right I'm curious kind of from an abstract existential point of view I'm curious to know for a practical point of view for you personally as a developer of this stuff what started changing and when did you feel like whoa all of a sudden I'm riding an electric bike not a regular bike or you know whatever the right Yeah, no, that's essentially what I think it's been continuous.

19:41I think every time AI improves, we find new ways to use it. But I mean, AI has been building AI for quite some time. The whole post-training process, in a way, has a component to it where we generate synthetic data, we generate data traces that are generated by AI models, right? And so initially, and this is all in the LAMA 3 paper, LAMA 4 papers, it's all very public, But you can see that initially what we use is like what we call like SFT, I think, for Find2Me, which is you have humans making annotations, say, OK, what should the answer be? Let's have a human write it because a human will know like what's the best answer there.

20:19And so you would do that. You train models, try to replicate a little bit like what that answer was going to be and kind of guide them towards that. But then what we realized is that, you know, you can scale generating a bunch of different solutions with AI itself much faster than with a human, right? And so we would generate a bunch of data points of outcomes with models. And then we say, well, the humans is still, you know, the better judge here. So the human would look at all these different answers and say, OK, this one's the favorite. Let's like try to train the model to kind of do that. And then we're like, wait a second.

20:53Wouldn't the AI be a better judge? Can we train it?

21:24So now... everything that they're doing right now, they end up having an epiphany where they go, wait, could AI help with that? Is it basically just kind of working its way through the progression of, you know, activities in the workflow? I'm lacking the vocabulary here, but you're getting my chance. Yeah, no, it's definitely that. I wouldn't say like it's that like for every single aspect of research, there's a lot of like... I didn't mean to be making comprehensive claims. I just meant to be describing kind of the way it works. Yeah. No, and definitely. And that's the thing is like, I think about it in my personal life flow, right?

21:55This is exactly how I use AI. And I'm like, wait a second. I'm not using AI for this. Would that be easier and better? And so I ask that constantly to myself to try to stop them. I say, no, I have to eat my own dog food in a sense, right? It's like, I have to keep using these things and understand what the use cases are and test them and researchers do that. What's the last use case you've done yourself where you went like, hey, wait a minute, that was smart. Wait, I can do that with AI? Yeah, I'm dying to know. What was the last one? I'll credit my wife for that one because she's learning French right now.

22:22And she was taking this like B1 level exam for her leveling. And she was like, wait a second, like I'm trying to do more tests, but I can't find these tests online. And so she just worked with various AI models like, you know, Lama, ChatGPT and Gemini to try to like create tests for herself and like learning programs. And so she would like, she didn't like typing. So she would even write everything by hand, just take a picture of the thing. And then it would grade it. But it would grade it because it had provided the model with the online reference to how these things should be graded. And she did some tests to figure out how known rating for some specific examples would come up based on her thing.

23:05And they came out exactly right every time, the exact same grade that were recommended. And so it was great because that allowed her to, one, scale herself, generate a lot of test data for herself to learn faster that wasn't available online. at least she couldn't find it. And so that's an interesting point, right? Because she found this way to do this. This is something she did with a chat interface. But a lot of innovation and products are happening right now where somebody realizes, look, I've done all this with a chat interface. Let me just package it into a product and make it simpler for people to just go and use it.

23:39I'll give you a fun example of that. I'm building this platform called Autos where we're trying to launch 100 ,000 companies a year. It's kind of like an entrepreneurial platform. And one of the things I saw the other day, it blew my mind. And that is the woman that is, she's American Chinese. And she was learning, I think, French too. And one of the things that she found out was that people who grew up in different cultures, they hear sounds differently. And so the tonality. And so she created this little kind of like, almost like guitar auto-tune kind of visualization. So if you want to say, I don't know, bonjour, she could actually go, And she could kind of find the tonality that she had to do to say it correctly in the accent that was correct.

24:22Because it just said to like, you know, phonetically or whatever, what she would kind of see and what she would say would seem right to her. But for a French person to be like, this is completely wrong. Fascinating. So Christian, what you're saying in both your case and your wife's case is there's this moment where you go, could it do that? You know, kind of a epiphany moment, which I love. And I often tell folks, I think sparking your own imagination is kind of the highest priority. And being in an environment where your imagination can be sparked is very important. Do you have practices or routines or rituals either for seeing new use cases and applications and or for auditing your own kind of existing workflows on some kind of a regularity to identify what's the next area of your life you should be inviting AI into?

25:16Or is it more kind of serendipitous and random? No, no, I do. I feel like two or three different things. The first one, I think, is any job that I'm doing that's taking me some time, I always will try to apply AI to it. And what I notice is initially it'll take me longer with AI. But if I commit to it, it'll usually make my life easier after a while once I've kind of narrowed down the kinks. The second one is, being an entrepreneur before, I think I always look for problem spaces. Because all these aha moments come from, this is part of my workflow. This is a problem that I'm solving right now.

25:51And I need to really make it better or I need to scale it. Like in the case of research, like, I mean, I didn't invent these things. Like this is what researchers do. But when they are these moments, it's like, look, I need to be able to scale more. Right. I need to be able to get more data. I need to be able to leverage better the humans in the process. I need to be able to do these things. And so it's having a keen eye on the problems because now we do have this armor, but we are looking for the nails everywhere. So looking for these problems constantly, I think, is an important one. But I think the pitfall long-term with this is that that helps you to optimize existing workflows and existing processes.

Read the full transcript

26:30Not necessarily doing new things. Exactly. And so I think where you see real disruption is when somebody tackles a bigger problem, you know, white page, like from scratch, and doesn't try to replicate any specific kind of workflow that was solving the problem before, that says, okay, I'm starting with AI, I'm starting fresh, like what should the workflow be entirely? I think one of the issues right now, I think is obviously it has to be integrated with the real world, right? You know, I've talked to people that have integrated AI in their organizations and they feel a little bit kind of depleted because they don't really see either their top line growth or basically the ability to be more efficient, right?

27:09And one of the things that I noticed is that you take, for example, a designer that makes dog toy design, because that's my work, and then suddenly they can also... Just randomly, just type it. Just randomly, just pick it up. And then you say to them, hey, now you can also write basically the marketing description of the toy itself, right? Because it's easy, they're doing that. Now that obviously doesn't reduce the workload for the designer, it reduces the workload for the marketing team. Now the marketing team is doing all the things, so you can't just, you know, you're not reducing an FTE just because the designer can do that.

27:39And so there's this kind of interesting patchwork that needs to happen as we're putting AI Ironman suits around people that need to kind of like get overlaid and you probably need to have enough of things happening so suddenly you can think of a new organizational design. Yeah. And so it's interesting and I think you're totally right in what you're saying, but obviously the issue is that all this has to be then deployed to the slowness and the conservatism of humans. And you're making an important point here because it's not just the workflows that change, it's the roles themselves don't make sense anymore.

28:13like the like i think about it like a developer like today you know might be spending a lot of time debugging looking for things like i mean finding where you missed a comma like you know you shouldn't be doing that on your own now they should definitely leverage it uh but that means that i'll give you an example like for the entrepreneur right when you're thinking of and i always have these ideas that i want to test out like when i think of an idea now i'm not going to create like a slide deck or a wireframe right i'm just going to actually build the mvp right Like I can vibe code it or I can do it.

28:42It's going to do the design for me. I'm going to even deploy it. And it'll take me like less than 24 hours from the ideation to the thing. And so suddenly like what's the role of my CTO when I start? What's the role of the, if you're a technical enough person that you can like spin it off, like I wouldn't be able to build like an architecture design for like large scale deployment of some of these products, but building the initial one I can do on my own now. So do these roles make sense? Like is the role of the product manager the same as the engineers, the same as the designer? And are these the same based on the different verticals you are?

29:13And where is it going to go? Is anybody's guess here? I wonder if more, you know, this, you're reminding me of kind of three conversations we had. The first point with Kasser Yunus, who's the CEO of Applied Intuition. He mentioned in our last episode, the kind of the movement towards kind of call it tiger teams or small team models where folks almost guilds, you can imagine an enterprise as a guild of a bunch of small teams doing different things, right? That's kind of one thing. The other thing that you reminded me of is our conversation with John Waldman, who's the CEO of Homebase, Stanford guy, who said that the product development process has completely changed there.

29:49They used to be reviewing 20-page PRD docs. Now he's reviewing clickable, lovable prototypes. So they're spending far more time in the field than they spend in the conference room reading documents, right? And then the third conversation you reminded me of, getting back to your question of are we just solving an existing thing versus doing something from a blank slate? we talked to Martin Reeves, who's the head of BCG's Henderson Institute, which is like their think tank. And one of the things that he found, which I kind of found fascinating was there's very little overlap between the kind of person who is kind of execution oriented or kind of solving the known problem and the kind of person who's doing something who is barely says per his research, only 3 % of people are capable of doing both.

30:33And granted the prior probability of kind of starting from a blank page kind of person is relatively small. So maybe, you know, but it's probably more than 3%, I would assume. But the point is most people who are kind of finding incremental opportunities, I kind of gave Martin this hypothesis. You actually need to do the incremental stuff in order to imagine the new stuff. And he said to me, that presumes that the person who's doing incremental is capable of imagining the new, and that's a faulty assumption, which was just, that was, that was really interesting to me. So anyway, I wanted to, But if you haven't gone into the back catalog, you've now got three killer next.

31:09This is basically just pitched from our episode. And if you enjoyed this. Yeah, exactly. I have. Let me ask you just a slightly changing gears because, you know, we've been a little bit concrete. Now we can go back to the philosophical stuff that we always end up with. So one of my colleagues went to Bhutan a few weeks back and then he did a lesson about entrepreneurship and AI. for like a high school or college, I can't remember. And one of the questions that came out was basically, hey, do you guys use OpenAI or do you use DeepSeek? And he being an American, he was a little bit surprised by because obviously in the US, it's not used as much, the Chinese models over there.

31:49And so he came back with kind of like this interesting thought, which was basically that if we're going to put AI on top of everything we do, the principles and the values and the thinking of the models is going to influence, obviously, how we see the world. Now, meta, obviously, is very big on open source. Your model is kind of like unique because of that everything is very transparent in the way it's and stuff like that. So I was curious on your thoughts on that observation of as we are now using models, do we as humans need to be aware that we're looking at the world through a prism, you know, like through a specific lens, because these models will have different kind of like, I wouldn't call them personalities, but they will have different approaches.

32:37Or do you think that is a moot point because the questions in are going to define the answers out? And so like your biases is the one that's kind of going to take the most weight. Yeah, I'll emulate Deanna a little bit on this one also. I think the point is, the main point I want to make is that all LLMs are biased. Right. It's a question of how you define bias at the end of the day. It's biased because the information we put in, the information we put in is biased because the Internet is biased. We, you know, we can work, filter, try to make it as unbiased as we can, but that's based on what we can measure.

33:13Right. You can only change what you can measure. So at the end, one story short with that is I think this is great that there are so many models out there that allow for so many different viewpoints. because one, we might discover that certain models work better with certain types of application than others. But I think we got to go back and look at what are these models going to be used for. And so 90 % of the time, I think, they could be wrong on this one, but I don't think the bias is going to matter because it's going to be around like, you know, automatically ordering my coffee, you know, or making sure I get the right directions to go somewhere or understanding how to have a more active voice and a passive voice that I'm writing.

33:57Like these things are very like objective in nature. They're really about like language or about actions or about things that are like that. Now, where that becomes tricky is then where for the rest of the use cases, when you use it as like maybe somebody to help you think about a problem, like you want to make sure that the information that's provided back to you is either as complete or as objective or, you know, based on what you see. So the answer point on this is that we're going to want a lot of different models out there that can help users pick and choose based on what they think they get the most utility out of, ultimately.

34:31So I'm not very surprised that maybe in Bhutan, they would use more DeepSea, let's say, than the other open source models that exist from Western companies, let's say. But it's great that we have those coming up now. I think it's such an interesting point. I think the whole bias conversation obviously has a lot of in-person interpretations. I met the woman who at the time ran Google image search. And she told us an interesting story where she says, if you search for a business person on, and she's an Indian descent woman, when you search for a business person on Google image, what it will do is it'll just statistically say how many times did an image appear with that kind of word appeared next to it.

35:13Happens to be a lot of white dudes, right? And obviously, you know, she being like a, you know, high power professional, but she's like, that's maybe the world that I kind of like necessarily indoors, but like it is the statistical model. Now, I shouldn't be the one who suddenly should change the model to a worldview that I kind of like subscribe more. And so I do think it's an interesting example because it just, it is just, it is what it is, right? But then sometimes we do. If people then go like, how do I kill myself? We do kind of like, we don't just add the best answer to that, right? We kind of like put a disclaimer up there because that seemed to be one.

35:49So it's hugely complicated making these decisions. So I agree with you. It's probably great that there just is a lot of models. So you can just kind of go towards the model of choice. But depending on the type of model you use, one of the images is a good example where you have even like language models that helps you like, you can test your model if it's, you know, So for example, the male-female issue, which is like translate this sentence in French from English. Let's say, for example, the doctor is like crossing the street. Like the doctor in French is le docteur. You could say that's most of the case is going to say le docteur, like a masculine word.

36:21But there's other types of professions where there's a different word between like female and male. And every time, like the model would translate it as the male. And so you can kind of test like that also for LLMs, whether or not there is some bias. Now, there's like three ways you can address bias, right? The first one is like data in, so you can filter the data initially and look at what you have. That's very difficult, especially at the scale of the data. But there's still some work being done there. I think that's helping. The second step is depending on the types of models. Some models, not cellular names, but others for images, for example, have what's called like an embedding space, which is that call it like a vector representation of concepts like in this space.

37:00And so there's the concept of male, there's a concept of female. and you can test out whether for certain like professions or certain things like your vector in a sense for that profession is too close to the male one versus the female one. And you can correct for that after the training in some cases. But that's very hard and it depends on the types of model architecture you've got. And the third one is like at post-training, which is if you want to teach the model, for example, the example you said around like, you know, you want to kill myself, like you don't want the model to answer that. So you can teach the model to answer in very specific way for that.

37:32But that's like, let's call it that's at the end, kind of narrowing down the scope of end source. It's not changing what's like underneath in the nuclear model initially. Thanks for giving that introduction. It's super fascinating. We do have like an interesting other episode where we talk about biases in data collection, which... Phoebe Yao, the founder of Pareto AI, is all about eliminating bias in the training side. So yeah, that's a really cool conversation. We, unfortunately, is running out because Jeremy has to go to a keynote. Got it up. Got it up. Professor, I have to profess. Professor Vertolin, what stood out to you out of that conversation?

38:11You know what? I'm always so impressed when you meet people with kind energy who seem just wicked smart and who's just in the middle of all this. And I don't know. I get highly energized by having this conversation. And I felt I literally could ask him like a thousand more questions. So first, I'm just kind of like high on the being able to talk to people like Christian. So that's good. You're buzzing. I love it. You're glowing, Henrik. Are you glowing? Do you have the man crush? Yeah, a little bit, but also just warm and sweaty.

38:43I was intrigued about this thing of, you know, I thought it was very fascinating to hear about this idea of the quality of models we can get out of the type of data we train on. And I think his explanation of text doesn't have the same fidelity as video was intriguing because, yeah, thinking about it and didn't quite understand it when the last time I heard about it. And now I did. And obviously, his other point about hallucination. I found that to be very fascinating. I do think like the whole bias stuff was nice to get just an explanation on because I do think that as we talked about in the pod before, we will increasingly have to be good of understanding what are some of the limitations or some of the areas where using AI will not just in different direction and then be a little bit more active and kind of like understanding that so that we don't necessarily kind of have the same thing that happened with social media happened to all of us.

39:40And so I thought this was a good conversation to have on tape and understand. I think that was my two main ones. You know, for me, the thing that it's a big theme, I think, on the show, but you use AI to use AI. And the more you work with AI, the more you discover ways you aren't. And the more you get, you know, kind of triggers of, oh, I never thought to do that. And what I find is, or AI hypothesis, I haven't tested that, but I bet those realizations start to increase over time. I don't know about you, but I feel like in the last two weeks, I've had more. I can't believe I never thought Tascadia than ever before.

40:26And, you know, we've been in the game for a while, right? I had this this morning with something where I, there's something called Google Scripts, which I don't understand, but I use. and is basically have access to my Gmail, right? And which is my email platform. And I have now all my newsletters I get in and then every morning it looks through all the newsletters and it creates like basically a summary. And I've been using that for a while. It's been very useful. It's one of the only newsletters I read every day. It's basically this kind of like, ah, now - The meta news. It was fine, but it wasn't kind of like, it was a better idea than I found utility out of it.

40:59And so back to the thing about like finding problems, I then took the one from yesterday and went to chat to materials like this is good but it doesn't really do it for me there seem to and so I actually used the model to help me figure out what is the problem I'm trying to solve and then it went like here's the different things you know it writes in a summary language you like to have a more active language and so you should change the tone you know you have these frameworks for how you see the world this a plus one framework you should basically look at not just the expression of what was in it but you should see them in the context the sum of the way that you compute it, whatever it was, right?

41:34And so it has these like five different things, which was like description of the problem. It's like, okay, now here's the code, Google script code, redo it and make it so I can just copy paste it back in. And then I put it in a like, it's completely night and day kind of outcome, right? And so, I don't know, I just thought it was like when he was telling us like, yeah, I could see how you need the problem. Sometimes you can even use AI to help you find the problem. No, it's so important. And it just, it makes me think, you know, many times, And by the way, Henrik, you've been doing this a long time and you're a world leader and you're an expert.

42:04And it was just yesterday that you thought. And the moment you say it, it's like, oh, that's pretty obvious. And yet you just thought of it yesterday. And the point there is not to belittle, but say there are tons of things like this in people's lives. And if they aren't actively realizing, wait, could it help? And I think, you know, we can have this moment where maybe we're self-conscious. Oh, my goodness. You know, forehead slapping. I can't believe I didn't think of that. But I think another step is to be willing to say, have you asked Chachi PT? And that not be an insult. Because the honest truth is for whoever I'm talking to, if I have the idea, chances are they haven't thought of it.

42:41And maybe the kindest thing I could do is be like, hey, you've got an amazing collaborator who's tireless and creative and willing to spitball with you. Have you invited them in? And I think that that does a few things. I mean, one, if they haven't thought of it, it gives them that, oh, you're so right. And two, especially if you're a manager, if they have thought of it and they're kind of wondering whether they should, it gives them permission. And then three, as a manager, it gives you the ability to follow up and then incorporate that into your kind of story kit of cool things that you know your people are doing.

43:18All of that happens if you ask the simple question, did you try AI? right and just like we are having these kind of personal realizations ourselves i think we can actually facilitate these epiphany moments for others if we're bold enough to ask that simple question amen i appreciate you for waking up very early in the morning to do this podcast with us in europe so thank you so much for that and i know you have a uh a keynote can you tell what the keynote is about. I am, yeah, I'm talking with a group of private equity CXOs about shifting their mindset in how they collaborate with AI from treating it like a tool to treating it like a teammate and onboarding AI as a teammate in their companies.

44:07That's awesome. Well, with that, I think it's time to say au revoir. Salut. Au revoir. Au revoir. And as always, we very much appreciate when people share this with a friend because that's how we grow our audience and that's how we get better people on it. And so if you enjoy this and you made all the way here to the end. Share it. We will share it with somebody else. And the, I was about to say, and the keyword or the code word for this episode is baguette. But then I realized I was so stereotyping that I kind of like made a complete mess. can we say croissant is that better anyway let's make let's not make it that let's make the uh path the the pie torch could be could be pie torch pie torch and with that bye thank you bye

From the publisher

In this episode, Christian Keller joins Henrik and Jeremy to explain how world models are shaping the next stage of generative AI. He talks through how AI learns using different types of inputs, and why video adds a sense of continuity, change, and cause and effect that text alone does not provide. Christian shares vivid analogies and clear examples to show what multimodal models make possible.

The conversation moves into how AI is now used throughout the research process, from generating synthetic data to evaluating model outputs. Christian shares how this loop is already in motion and how AI is helping scale and accelerate experimentation. He also reflects on the shift after ChatGPT launched, and how that changed the pace and structure of research work.

Later in the episode, Christian describes how individual workflows are evolving, and how asking simple questions like “Could AI help with this?” often opens new possibilities. He shares examples from his own work and home life, including how his wife built and graded her own French exercises using generative tools.

Key Takeaways:

  • Text removes essential information
    Christian explains that text compresses reality and loses detail, context and temporality. Images and video help restore what text leaves out.
  • World models give AI a sense of change
    Video introduces the before and after and how things move or enter a scene. This helps models learn cause and effect and builds more robust understanding.
  • AI helps build AI
    Models can generate data, evaluate results and support researchers during development. Christian shows how this creates new ways of scaling experimentation and training.
  • Workflows shift when AI handles early steps
    Christian shows how tasks like debugging and prototyping change with generative tools, which reshapes roles and opens new opportunities for innovation.

LinkedIn: Christian Keller | LinkedIn

00:00 Intro: Information Compression
00:37 Meet Christian Keller: AI Expert
01:13 The Evolution of AI Products
02:11 Impact of ChatGPT on AI Development
02:38 Understanding PyTorch and Its Role
07:41 The Bitter Lesson in AI
09:12 Challenges and Future of AI Models
18:57 Using AI to Build AI
23:25 Innovative Chat Interfaces
23:41 Building the Autos Platform
24:35 Epiphanies in AI Integration
25:18 AI in Entrepreneurial Workflows
26:32 Challenges in AI Integration
31:15 Bias in AI Models
38:06 Debrief 

📜 Read the transcript for this episode: Transcript of AIs Next Frontier: World Models Explained by Christian Keller |

 

For more prompts, tips, and AI tools. Check out our website: https://www.beyondtheprompt.ai/ or follow Jeremy or Henrik on Linkedin:

Henrik: https://www.linkedin.com/in/werdelin
Jeremy: https://www.linkedin.com/in/jeremyutley

 

Show edited by Emma Cecilie Jensen. 

More from Beyond The Prompt - How to use AI in your company

All 48 episodes
AI’s Next Frontier: World Models Explained by Christian KellerBeyond The Prompt - How to use AI in your company · 45 min
Listen in VO