In short
Podcast Episode Summary: OpenAI Sora 2 Team
Episode Overview
- Podcast Title: Training Data
- Episode Title: OpenAI Sora 2 Team: How Generative Video Will Unlock Creativity and World Models
- Hosts: Konstantine Buhler and Sonya Huang, Sequoia Capital
- Guests: Bill Peebles, Thomas Dimson, Rohan Sahai from OpenAI
- Description: The Sora 2 team discusses innovations in video generation technology, focusing on how they are compressing filmmaking timelines and enhancing creative capabilities through generative video.
Key Concepts
Sora 2 and Generative Video
- Diffusion Transformers: A novel approach to video generation, introduced by Bill Peebles, that allows for generating video by gradually removing noise, improving quality and maintaining object permanence.
- Space-Time Tokens: Fundamental units used in Sora's video generation that encapsulate both spatial and temporal data, enabling the generation models to have a comprehensive understanding of physical interactions.
- Emerging Properties: Sora 2 improves on Sora 1 by introducing models that respect physical laws in their output, offering a more realistic interaction of objects within videos (e.g., a basketball bouncing off a hoop).
Intentional Design Against Mindless Consumption
- The Sora team emphasizes the importance of designing the Sora product to inspire creativity rather than consumption, aiming to counteract mindless scrolling common in social media platforms.
- They highlight the importance of user engagement and creativity through features such as remixes and cameo functionalities.
World Simulators and Future Applications
- The team envisions a future where generative video could be used for scientific experimentation and simulations, potentially leading to new discoveries in physics and other fields.
- They discuss the potential of digital simulations to create new forms of knowledge work and collaboration in alternate realities.
Product Development Insights
- Team Backgrounds: Each member shares their journey to OpenAI and their roles in the Sora project, highlighting their diverse experiences in video generation and machine learning.
- Product Evolution: The conversation outlines the iterative process of developing Sora 2, emphasizing the importance of user feedback and the balance between technical capability and user experience.
Ethical Considerations and IP Rights
- The team discusses how they are approaching intellectual property (IP) rights, ensuring that creators and IP holders benefit from the new creator economy facilitated by Sora.
- They highlight the challenge of establishing a fair monetization model for both creators and IP rights holders.
Future Vision
- Personal Digital Clones: The potential for users to create digital representations of themselves that can interact and perform tasks within a simulated environment.
- Impact on Filmmaking: The team projects a future where anyone can create compelling video content, democratizing filmmaking and potentially leading to a new era in media creativity.
- Comments on how the platform could evolve into a hub for creativity, social interaction, and even scientific discovery, reflecting a shift in how content is created and consumed.
Conclusion
- The discussion concludes with reflections on the role of technology in shaping creativity, the societal implications of AI, and the responsibility of developers in guiding the ethical use of emerging technologies.
Key Takeaways
- Sora 2 represents a significant leap in generative video capabilities, emphasizing creativity and user engagement.
- Future developments could lead to revolutionary changes in various domains, including entertainment, science, and social interaction.
- The importance of ethical considerations and collaboration with IP holders will be crucial as the platform evolves.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00For OpenAI across the board, it's really important that we iteratively deploy technology in a way where we're not just like dropping bombshells on the world when there's like some big research breakthrough. We want to co-evolve society with the technology. And so that's why we really thought it was important to like do this now and like do in a way where, you know, we've hit this again, this kind of like GPT 3.5 moment for video. Let's make sure the world is kind of aware of what's possible now. And also, you know, start to get society comfortable in like figuring out the rules of the road for this kind of like longer term vision for where there are just copies of yourself running around in Sora and the ether, like just doing tasks and like reporting back in the physical world, because that is where we are headed long term.
0:54Today on Training Data, we sit down with the team behind OpenAI's Sora, Bill Peebles, Thomas Dimson, and Rohan Sahai. You'll hear about space-time tokens, building internal world simulators, and how optimizing for creation instead of consumption is just better for social platforms. This conversation goes way beyond video generation and into questions about how society will co-evolve with powerful simulation technologies. We promised that this was an actual real-world conversation and not a video generation, but we don't know how to prove that to you. Let's jump in. Hey guys, thank you for being here at Sequoia.
1:29Congratulations on Sora. Thank you. Maybe you could tell us a little bit about yourselves and how you got to OpenAI and Sora. Yeah, I'm Bill. I'm the head of the Sora team at OpenAI. I had a pretty traditional path, came through undergrad doing research on video generation, then continued that work at Berkeley, and then started at OpenAI working on Sora from the first day I joined. And I'm Thomas. I work as an engineering lead inside of Sora. I have a bit of a longer story, but I worked at Instagram for about seven years doing some of the early kind of machine learning systems and recommender systems there.
2:05But it was a very tiny company. It was about 40 people. Then I quit, did my own startup for a while, which was Minecraft and the browser, which we've talked about a couple of times. And I think that OpenAI noticed that we had a very cracked product team there. And so they acquired our company. and I've been bouncing around different products inside of OpenAI and on the research side as well on post-training, but super happy we landed kind of together on Sora to bring this thing to life. It was a really cool product in between too, like the global illumination product. Oh yeah, I still believe in it.
2:35Yeah, me too. Awesome, I'm Rohan. I've been at OpenAI for about two and a half years. Started as an IC on ChatGPT, but then as soon as I saw the video gen research, I got quickly Sora pilled and made my way over there. and so currently we have the Sora product team. Before that, just startups, big companies within kind of the valley, a bunch of random stuff. Cool. Well, Bill, you are the inventor of the Diffusion Transformer. Can you tell us what that is? Yeah, so most people are pretty familiar with autoregressive transformers, which is the core tech that powers a lot of language models that are out there.
3:10So there you generate tokens one at a time and you condition on all the previous ones to generate the future. Diffusion transformers are a little bit different. So instead of using autoregressive modeling as kind of the core objective, you're using this technique called diffusion, which at a very high level basically involves taking some signal, for example, video, adding a ton of noise to it, and then training neural networks to predict the noise that you applied. And this is kind of a different kind of iterative generative modeling. So instead of generating token by token, as you do in autoregressive models, diffusion models generate by gradually removing noise one step at a time.
3:45and in Sora 1 we really kind of popularize this technique for video generation models so if you look at all the other competitor models that are out there both in the States and in China most of them are based on DITS diffusion transformers and a big part of that is because DITS are a really powerful inductive bias for video so because you're generating the whole video simultaneously you really solve issues where quality can like degrade or change over time which was kind of like a big problem for prior video generation systems which DITS ended up fixing so that's kind of why you're seeing them proliferate within video generation stacks.
4:16When I try to visualize it, I mean, for each diffusion, you have a matrix of pixels, and then you do the entire video at the same time, which you can basically see as different frames, I imagine. Can you visualize that as, you know, matrix of matrices that basically transforms over time? Yeah, it's a good question. So we really kind of consider things at the granularity of like space-time tokens, which is sort of like an insane phrase. But, you know, whereas, you know, for example, characters are very fundamental building block for language. For vision, it's really this notion of a space-time patch, right?
4:47You can just imagine this little cuboid that composes both X and Y, like spatial dimensions, as well as a temporal locale. And that really is kind of like the minimal building block that you can like build visual generative models out of. And so diffusion transformers sort of consider these almost, you can think of it like a voxel by voxel. And, you know, in the traditional versions of these these diffusion transformer models um you have all of these little space-time patches talking with all the other ones and that's how you actually are able to get properties like object permanence to fall out because uh basically you have full global context of everything going on in the video at every position in space-time which is like a very powerful uh property for a neural network to have yeah and is that the equivalent of the attention mechanism is the objects movement throughout the video yeah that's right so in our like soro one blog post um video generation models as world simulators, we kind of laid out some visuals, which sort of go into exactly your point here, which is really attention is like a very powerful mechanism for sharing communication, like sharing information across space time.
5:52And if you represent data in this way, where you patchify it into a bunch of these space time tokens, as long as you're properly using the attention mechanism that allows you to transfer information throughout the entire video all at once. What are the biggest differences between Sora 1 and 2? And I remember with the original Sora 1, you're already seeing kind of emergent properties where the more you scale, the more it's able to do things like understand physics. Is Sora 2 purely a function of scaling or what are the biggest differences? Yeah, that's a great question. You know, we've spent a long time really just doing like core generative modeling research since the Sora 1 launch to really figure out how we get the next step function improvement and video generation capabilities.
6:32We really kind of operated from first principles, right? So we really want these models to be extremely good at physics. we want them to kind of feel intelligent in a way that I'd say like most prior video generation models don't so by that I really mean you know if you look at kind of any of the previous set of models that were out there you'll notice a lot of this kind of like effects that happen like if you try to do any sort of complicated sequence of like physical interactions right for example like spiking gymnastics classic uh riding a dragon like you did riding a dragon that was fun that that happened for real actually constantine um uh you know they're like very clear problems with the past generation of models that we really like set out to solve with sora too and i think one thing that's really cool about this model compared to prior ones is that um when the model makes a mistake it actually fails in a very unique way that we haven't seen before and so concretely uh for example if uh let's say like the text input to sora is a basketball star wants to like shoot a hoop, right?
7:30Shoot a three throw. If he misses in the model, Sora will not just like magically guide the basketball to go into the hoop, right? To be over optimistic about respecting what the user asked for. It will actually defer to the laws of physics most of the time. And the basketball will actually like rebound off the backboard. And so this is a very interesting distinction, right? Between like model failure and like agent failure. Agent as in the agent that Sora is like implicitly simulating as it's generating video. And we haven't really seen this very unique kind of like semantic failure case in like prior video models.
7:59This is really new with Sora 2. It's kind of a result of, you know, just the investment we put in like really doing like the core generative modeling research to like get this massive improvement in capability. Okay, so not purely a function of scale. You're actually, you know, there's some concept of agents implicit in this. There's things you're doing beyond just scaling up a model. Well, the notion of agents I'd say is actually mostly like implicit from scale. Like, you know, So in the same way where we kind of showed that object permanence begins to emerge in SOAR 1 pre-training once you hit some critical flops threshold, we see similar kinds of things happen as we push the next frontier.
8:36So you begin to see these agents act more intelligently. You begin to see the laws of physics be respected in a way that they aren't at lower compute scales. How does the concept of a space-time latent patch relate to a space-time token, relate to object permanence and how things move through the physical world? Yeah, that's a great question. So I'd say space-time patch and space-time token are more or less synonymous with one another. I'll use them interchangeably. interchangeably, you know, what's really beautiful, right, is when people started scaling up language models from like GPT-1 to GPT-2 to GPT-3, we really began to see the emergence of like world models internally in these systems.
9:14And what's kind of beautiful about this, right, is there's incredibly simple tokenizers that actually go into like creating the data that we train these systems on. But despite this very simple representation, right, you know, like BPE characters, what have you, when you put enough compute and data into these systems, like in order to actually solve this task of predicting the next token, you need to develop an internal representation of how the world functions, right? You need to like simulate things. And like, you know, the models will make lots of mistakes right now at like low compute scales.
9:43But as you continue pushing it from three to four to five, you just see these internal world models get more and more robust. And it's really analogous for video, right? And in many ways, more explicit. So I think it's easier to picture what like a world model or a world simulator looks like with video data, right? because it is literally representing like the raw observational bits of like all of reality. But what's really remarkable, right, is because these space-time patches are just this like very simple and like highly reusable representation that can apply to like any type of data, right?
10:11Whether it's just like video footage of like this set, whether it's like anime, cartoons, like whatever it is, you're just able to build like one neural network that can operate on this vast, extremely diverse set of data and really build these like incredibly powerful representations that model like very generalizable properties of the world, right? It's useful to have a world simulator to predict like how a cartoon will unfold. And likewise, it's useful for predicting how, you know, this conversation might unfold. And so that really puts a lot of optimization pressure on Sora to like grok these like core fundamental concepts in a very like data efficient way.
10:44Did you have to put effort into selecting the data such that it reflected the physical world? For example, I'd imagine if you have data from the physical world, it all abides the laws of physics. but you mentioned anime that might not always abide in the laws of physics did you have to be selective or did it naturally find patterns that separated that out that's a really great question um we did spend a lot of time you know really thinking about you know what does like the optimal data mix for like a world simulator kind of look like and to your point you know i think in some cases we'll make decisions that you know maybe are for like making the model really fun like for example people love generating anime but you know do not necessarily like perfectly represent uh like the laws of physics that are like directly useful for like real world applications so like to put it another way right i think in anime there are certain primitives that are simplified that are actually probably useful for understanding the real world you know people still locomote through scenes for example but like if there's like some crazy dragon that's like flying around that's probably like not so useful for like rocking aerodynamics or something Dragon Ball Z is more or less how I learned athletics.
11:47You know, there you go. The motion and Super Saiyan. I think it is an interesting question like that I do not know the answer to whether somehow like pre-training on simplified representations of like the visual world, whether that's like sketches or like some other modality, like, you know, makes you more efficient, like rocking these concepts. I think it's actually a very interesting scientific question that we need to understand better. Do you think we're close to exhausting the number of pre-training tokens there are out there? Or do you think video data is just so massive and it's actually one of the more untapped vats of data?
12:19Yeah, the way I kind of think about this is the intelligence per bit of video is much lower than something like text data. But if you integrate over all of the data that really exists out there, the total is much higher. So to directly answer your question, you know, I think it's hard to imagine ever fully running out of video data. There's just like so many ways that it exists in the world that like, you know, you will be in a regime where you can continue to just like add more and more data to these pre-training runs and continue to see games for like a very long time, I suspect. You think we'll ever discover new physics?
12:50There's the LLM world of, you know, Einstein thinking the whiteboard. It's equivalent to these LLMs thinking. There's also just the, if you develop a perfect simulator and you just simulate physics better and better, you might learn things about the world that we haven't learned yet. I totally think that this is like bound to happen one day. And like, you know, I think we probably need even like we probably need one more step function change, I'd say, in like model quality to like really get to a point where, for example, you can think about doing like scientific experiments in the models. Like you could imagine, right, one day you have a world simulator that is like generalized so well to the laws of physics that like you don't even need like a wet lab in the real world anymore.
13:24Right. You can just like run biological experiments within Sora itself. And like, again, this needs like a lot of work to like really get to the point where you have a system that's robust enough to do this reliably. but you know internally like again we've used Sora 1 as kind of being like the GPT 1 moment for video it was like really the first time things started working for that modality. Sora 2 we really view as like GPT 3.5 in terms of like it really being able to like kickstart you know the world's creative juices and like really like break through this kind of usability barrier where we're seeing like mass adoption of these models and we're going to need a GPT 4 breakthrough to really get this to the point where this is useful for like sciences as we're seeing now with GPT-5, right?
14:00Like I feel like every day on Twitter, I see another like convex optimization lower bound get like improved by GPT-5 Pro. And I think eventually we're going to see the same thing happening for the sciences with Sora. Do you think you need physical world embodiment to get there? Or do you think a lot of it can be done effectively in Sim? You know, I am like always amazed, you know, every time we like push another like 10x compute into these models, like what just like magically falls out of it with like very limited changes and kind of like what we're training on and like the fundamental like approach to what we're doing.
14:33You know, I suspect some amount of like physical agency will certainly help. I have a hard time believing it will like make you worse at like, you know, modeling like collisions or like something else. Video only is like quite remarkable though. And I wouldn't be surprised if it's actually kind of like AGI complete for like building like a general purpose world simulator. So for this concept of a general purpose world simulator, a world model, where you can do science experiments in that world. Do you think that video is the soul or some combination of video and text are the combined data inputs and you train it on this type of model?
15:12Or does it have to be based on more structured laws of physics that are understood and laws of biology that are understood? I think it probably depends a lot on the specific use case you're kind of envisioning for the world simulator. Like, for example, if you just really want to build like an accurate model of how like a basketball game is played, I actually think like only video data and like maybe audio as well, like kind of sufficient to build that system. Not of me playing basketball. That would be an inaccurate, very bad player of basketball. You know, you actually like Sora's current understanding of how people play basketball, Constantine, may be at your level.
15:49Wow. Okay, that makes sense. You know, it's possible. It's possible. I think he just dissed you. I liked it. It's accurate. But it's better than mine, Constantine. That was like a Sora 1 situation. You're at Sora 2. We'll toss some hoops. Is that what they'll say? You know, I'm down. Okay, perfect. Shoot some hoops. Thanks. Toss some hoops. Thomas' first statement in the podcast. But I'm also at your level, so I'm sorry. I think it is an interesting question. What are all of the modalities that should be present in this kind of general purpose system? system like certainly you know if you add more modalities i have a hard time believing it will like decrease the intelligence i also think there's an argument to be made that um just you know adding more and more does not provide like significant marginal value compared to like you know full mastery of like video and audio for example i think it's an interesting open question i'm not i'm not actually sure right now um and it's something we need to understand more yeah so cool sonia a minute ago mentioned einstein at a whiteboard and obviously that makes me think of you, Thomas, and your hair.
16:49Me too. It had to come. If any hair gives the feeling of space-time tokens, it's definitely yours. At some point, Bill, you're the creator of this revolutionary technology that has changed the way that AI video is created. At some point, you from Sora 1 to Sora 2 said, hey, all together, you said, there needs to be an application around this. There's some benefit to an application. You brought together some of the best product people in the world. How did that crew come together at OpenAI? Yeah, I mean, the story is never as linear as you might think it is. So I think that, I mean, we've had a product team on Sora since the get-go.
17:32Rohan was like spearheading that effort in the Sora 1 days. But I think Bill's right when he says it was really like a GPT-1 kind of moment. We're seeing pockets of very interesting things there. But the models were not like, you know, models without sound, videos without sound. It's like a very different kind of environment. And so we're working on that surface, mostly target on kind of like a prosumer demographic. And separately, I mean, wrong, probably go into more details of all that. Separately, we're also just kind of exploring different social applications of AI inside of OpenAI and like what that could look like.
18:05We had a lot of prototypes, most of which were quite bad. And when we started to see some of the magic was actually with ImageGen before it had been released. we were playing with it internally in a social context and the social context was really interesting to see that what people were doing is you'd sort of like take an image and then you'd have like a chain of remixes of that image where like i don't know there was a it's a duck and then now the duck's on somebody's head and now everything's upside down and they're smoking a cigarette like uh just a lot of weird things yeah it's like and um we were seeing this we're like oh this is kind of like a very interesting thing that's like you know nobody can really do that with like social media because it's so hard to create something or riff on something it like is such a high barrier to entry action um maybe you have to get to go get a camera set up and it's not just like thinking of the ideas there's actually a lot of things involved and so we were like okay this is a very magical behavior how can we kind of productize that behavior um and we're mostly thinking about away from sora um some of the sora research was still ongoing and i mean there are signs of life but it wasn't like quite there yet in productized form bill probably had it in his head somewhere.
19:13I can see the future, but that's fine. I'm a little bit more. Can't quite see the future yet. So we were just exploring that. I think we tried a few things, and then at some point, the research was really just showing very clear value of even iterative deployment style value of like, oh, this is something that people will really want. And so we went into this project like two or three months ago. It wasn't very long. It was like July 4th. Wow. July 4th, yeah. Wow. That's when you disappear, Thomas. That's when I disappeared, yeah. Makes sense. And we just kind of locked in like, OK, we're finally doing it.
19:49That's always a moment. And we started without any magical features, just like, OK, let's just try to get a native video environment where you can hear the audio full screen. And we did some quick generations. Things were showing very cool, very fun, very interesting. and because of that image gen experience we sort of had thought like okay what's the magical here magical thing here is that like barrier to entry is very very low for creation coming from Instagram that's like it's impossible to get people to create on Instagram and that's the most valuable thing that people do so what does that unlock and it's like okay well that remix thing from image gen that kind of could still apply here and so we brainstormed all these things about how could remixes work and like what does a remix mean here um one of those was this like cameo thing which i think also bill had in his head but this uh it was in the ether it was in the ether for sure um but we just were like hacking together things on the product let's see if this works um i i didn't think it would work at all um but it was on the list and there were a few other things on the list some of them were pretty crazy it was like uh why didn't you think it would work i am bad at predicting technology like uh it wasn't super clear to me that you could like you know take a likeness of a person and have that kind of imagined into a video form um and whether it would work or not and so we had early prototypes of different things of like people reacting in the in the video corner or uh stuff like that but when when we saw cameos just start to work and even playing internally like ron do you remember that day where we're like yeah feed is entirely cameo yeah it's entirely it just went from you know we didn't have that feature once we had that feature product market fit on the team all everything we were generating was all of each other you must have seen the meme potential i mean yeah that's i think at first we were just like this is hilarious this is amazing and then like a week later we were like this is still all we do um so there's something here yeah i mean at first we were actually a little bit like, is this good?
21:55Like, Hey, the cameos, it's just all cameos now. Does anyone else care about this? People care about other people doing stuff. And, um, we kind of got to the point where we're like, no, no, this is actually good. Like it's actually, it feels like I'm coming back to see. And it really humanized it a lot where like a lot of AI video is just, um, kind of static scenes that are quite beautiful, quite interesting, might have extremely complicated things going on, but they lose that human touch and it really felt like it was coming back into it so another learning from image gen too like image gen took off and had viral moments because i think you could put yourselves in these scenes in accessible ways that weren't possible before obviously this massive like put me in a ghibli scene um people taking selfies with their idols and stuff like that and so the once you know once you actually kind of thought about it it's like yeah cameo feature makes a lot of sense you put yourself in all these scenes that's way more exciting you and your friends it's novel it's like that's something you could do before yeah and then that combined with remixes i mean came it was kind of a remix to begin with but then you start to think about okay well now i can riff on rohan doing something or whatever it is like with bill had you wrapped in an action figure package and it was it's been remixed like yeah an insane number of times of times yeah so like uh just very very crazy things that kind of go on and very emergent a lot of stuff that i would have never thought of actually how many generations of you guys have been like publicly posted at this point.
23:20I have no idea. I know I'm 11 ,000 or so. I was like a little less than that. Wow. Yeah. What has surprised you about the types of users that are really sticking with Sora? Who is it really a hit with? If you just go to the latest feed, which is just like the fire hose. Astronaut mode. Of everything. Yeah, it's space time Thomas mode. It's wild out there. But that gives you a pretty good snapshot into like just everything happening. I mean, I think we have like almost 7 million generations happening a day. So you can imagine there's just a ton of information there. It's one of my favorite ways to just get product feedback.
23:56It is so diverse, the type of stuff people are doing, the type of people. There'll be like a complete variety of age. Some people just envisioning themselves in scenes that seem like motivation-oriented. People just memeing with their friends. People cameoing some of like the public figures on the platform that have done cameos. So I think the diversity has surprised me. I was kind of expecting this sort of like, you know, the Twitter AI crowd to like heavily dominate the feed. They definitely dominate like the press cycles, at least the ones that, you know, we're most exposed to. But in terms of people actually using this, it's quite a wide variety.
24:32And last thing I'll say is a bigger departure from like the sort of niche AI film crowd that existed before, which is great early adopters. but now you kind of get these. I thought it would start there, but it felt like it started with just a way wider range of people. I think getting to the top of the app store helps with that and just get people who are like browsing and see this thing. My mother keeps cameoing Thomas. Is that right? Just to cameo Thomas. So weird. We got a lot of strange cameos. There are real stories like that. We said 11 ,000. She's done 10 ,000 of them. Thomas, you wrote the original algorithm.
25:11if I'm right, for the Instagram ranking, ranking algo. There was a lot in the Sora 2 blog post about how you guys are clearly being very intentional about how you want to do ranking in the algo. Can you talk a little bit about lessons learned from Instagram and how you're approaching it over at Sora? Yeah, I mean, there's a lot to cover in that. I think that the first thing to think about when we think about these platforms or think about Sora specifically is the thing I was mentioning before about creation. um so uh sore enables basically everybody to be a creator on this platform and that is a very very different environment than something like instagram where you have this like extreme power law of the people that are creating um and the power power law just naturally gets more uh narrow what's the right word there but more uh head heavy yes um so sometimes i feel like i have to defend myself on the instagram uh algorithm side we actually did it for i mean we did it for reason.
26:07It was to solve a problem. It wasn't just kind of like a random decision to optimize for ads or something like that. And the reason we did that was that we noticed that what was happening on Instagram over time was, because it was chronologically ordered, every single person that posted was guaranteed to have the top slot of all their followers. And so if you think about that for a second, the incentive for somebody in that environment is actually to create constantly because they are guaranteed distribution when they create. And over time, because of this power law becoming heavier and heavier, or more head heavy, those type of people, which are great, they provide a lot of value to the ecosystem, but they start to crowd out people you really care about.
26:51And so maybe you follow National Geographic or something, not the Duncan National Geographic, I love them, but if they're posting 20 times a day, your friend's not. They don't have the same optimization objective. They're probably just a picture of their coffee or something. And so you'd have 20 Nat Geo posts and then one picture that you actually really cared about that you never really scrolled to. And there's not too many solutions to that problem if you have a guaranteed ordering. One of them is that you have to unfollow all these accounts that you maybe care about but care about not as much as the person that posts once a day.
27:26And the other is that you have to permute the feed. And so we went with that path. We tried it. We tested it out internally. It was very kind of controversial to do. But I think that you can actually kind of like math this out. It's like a proof that basically over time you're going to have to take control over distribution on the platform in order to prevent these kind of issues and show people what they actually care about. So that's why we did it. And it actually showed a lot of value. I remember the early tests, I won't get into the numbers on them, but they were pretty unambiguous, actually, about this was showing more people that you cared about.
27:59It was improving your experience with the platform. It actually moved creation, which is unusual. It made people create more because they were seeing more content that was accessible to them. But I also think that these things can go astray over time. And I won't say like the Instagram algorithm is unequivocally bad or unequivocally good. But when we started to open up to more unconnected content and ad pressure was very strong, there's also a natural company incentive to optimize for just blind consumption because that's how you make money. So maybe cheaper content or maybe just like get people to scroll more and more and more and more.
Read the full transcript
28:35And that also can encourage people to create less because it's just like a more mindless scrolling mode. um you guys have very concretely committed to doing things to prevent that kind of behavior we have yeah um we have we have a lot of mitigations that are in place but i i think uh what it really comes down to me is just like what are we trying to do as a platform and i think the magic of this technology is that everybody is a creator and so we want this feed to be optimized for you to create to inspire you to create and that can be like sometimes when you think of inspiration you think of like oh it's this beautiful crazy scene that's so elegant when I think about that I think about like a meme culture or something really funny or like well that's cool I've got a riff on that and I think that's a very different brain mode when you're browsing the feed and of course we have lots of other things in place so like I think it starts with incentives that's our incentive right here is to encourage more creation in the ecosystem but there are certainly use cases we want to prevent we're not going to get them right all the time.
29:36It's very challenging. It's a very living system. It's also very hard to write a recommender system when you have no data and you don't know what to recommend or you don't know how the platform is going to evolve. But that's like basically how I kind of think about the incentives of feed. And then Rohan, we have a lot of mitigations in place that I think you've been kind of like thinking about and maybe even more deeply than I have about like preventing maybe the extreme cases. And so I don't know if you want to talk a little bit. Yeah, happy to. But one thing before you, I mean, just one thing to add is that the stated intent of like optimizing for creation is working really well.
30:10Yep. It's almost 100 % of people who like get past the invite code and all that on the app end up creating on day one. When they come back, it's like 70 % of the time they come back, they're creating. And 30 % of people are actually even posting to the feed. So not just like generating for themselves, they're actually like posting into the ecosystem, which is an incredible testament to the model, how fun it is, and to like how what we're optimizing for is actually working pretty well right now. But yeah, beyond that, I mean, like one of the top of mind things is I think we don't want this just to be like a mindless scroll.
30:40And beyond just optimizing for creation in the ranking algorithm, there are things we can do, like trying to just get you out of this sort of flow state of just like consumption and push you into like creative mode. I think there's a great article on this called like the curvilinear nature of casinos where they design it so you never have to make any decisions. It's just like you walk in a circle. There's no windows, all that kind of stuff. we can be very intentional about not doing that. And like, you know, whether it's an in-feed unit that's like, hey, you just kind of viewed a couple of videos in this domain.
31:12Why don't you try creating something? Or other ways to just kind of like push you out of that. We actually have things like that in the product. Yeah, those are some of the things that come to mind. I really commend you guys for what you've done to, you know, make sure that there's a version of the world where video model as world simulator could have just ended up with us, you know, each retreating into our own computer screens and just becoming addicted and just retreating into ourselves. And I think the amount to which you're, you know, prioritizing the human element and the social element, I think that the care you've put into that really shows.
31:44I don't think we would have launched like a feed of just like AI content that wasn't, that didn't have a human feel. Like just, I don't think that excited us. And as soon as we, we like had the product, we had Cameo and we had that feeling internally, um we were like okay this is actually a little different than yeah i don't think it was totally obvious again it was like a pretty crazy sprint to go through this uh and it wasn't like super obvious to us what what would emerge but i think that the idea it makes sense in retrospect but it was a completely not obvious product decision the cameos would be the thing yeah um where it's like of course you just want to see your friends doing cool things so it's like that makes sense um but i I was never actually that afraid of competitive pressure in that, that crazy product phase.
32:29Cause I was like, we, we sort of had these all these non-trivial decisions that are obvious in retrospect, but we're not obvious at the time that we're sort of building on top of each other. It's like, okay, cameos. Well, there's also a version of cameo where you have a crazy flow that's just for you. And it's a one player mode cameo. And you like go through this onboarding flow and do your stuff. But we were already seeing these interesting dynamics where it's like, Oh, I could tag Rohan into my video. That's crazy. Like, and then we can have like an argument or like, I'm going to have a, Anime fight, doesn't matter.
32:56And I was like, OK, so that's actually the human element. That's the magic of this. It's actually strangely more social than a lot of social networks, even though it's all AI-generated content. It's very unintuitive. Totally. Is it a fine-tuned version of Sora 2? Or is it a separate model from what's available over the API, or is it the same? Between the app and product. So we are currently exposing the models in the same state across API and the app. OK, really interesting. What are you seeing people do on the API side? And is it different from the types of things people are doing on the consumer app?
33:30The motivation behind even launching an API is just the support of these long tail use cases. We have this vision of enabling chat GPT scale level consumer audience with this tech. But there's tons of very niche things out there. You can imagine people who are much, with Sora1, we went out and talked in a lot of these studios. What we heard from them is they want to integrate this in this specific part of their stack in this specific way. And we'd love to support all these long-tail use cases, but we don't want to build a thousand different kind of interfaces for this stuff. So that's the kind of stuff we're excited to see with the API.
34:03So far, it's been, you know, it's been those kind of like a little bit more of a niche company, not trying to build like a first-party social app, but maybe, you know, has some either filmmaking kind of audience or kind of people they're supporting, or even just like we've definitely seen some like people trying to, you know, I think there was like some company making, they were doing something with CAD where they were like using Soro. Oh, Mattel. Yeah, yeah, yeah. Oh, that's cool. So there's cool use cases out there. I think we're still getting a sense of what they are. Yeah, I think there's a lot that can be done with these things.
34:37I think about gaming all the time just based on my background. You know, AI and gaming is always a very controversial subject, but it's very clear that there's a place and there's a role. maybe doesn't have to interrupt the creative process. It can enhance it. And I'm pretty excited to see some of those use cases emerge. Do you think the video models are good enough now for people to be able to build video games on top of the API? Or do you think we're still another rev or two away? I have my own take on this. I was going to say, never bet against the ways people can be creative with technology to build.
35:07Like someone will be able to build a game and maybe has built a game already. Will it look and feel like a, you know, obviously there's latency with this model. So you'd have to do all sorts of crazy stuff to get around that. I think that your mind immediately goes to the obvious sort of things that you would do in gaming. And we've seen some of that sort of stuff, certainly in research blogs and that kind of thing. My mind often goes to like, okay, this is like a creative tool that's a little bit different. And the types of games that really excite me there, I'll just go off on one, which is this, there's a game called Infinite Craft, which is the world's simplest game.
35:39It's a web game where you just take elements. It's like fire, water, earth. You have like four elements to start. and you just drag them and it combines into something new and the thing it combines with is like a it's lm based so it's like fire and and and earth might be a volcano um and then volcano plus water might be an underwater volcano or godzilla or something like that you always end up in godzilla for some reason but uh that's a game that like it's like oh it kind of makes sense where it's like yeah you don't really need a crafting tree the lm can derive this crafting and it's a process of discovery.
36:16And so I think there's a lot of untapped stuff in that space where, again, I like the idea of a process of discovery. In fact, my philosophical view on LLMs and video models to some extent is that it is a process of discovery. These are all in the weights. You're just unlocking it with like a secret code, which is your prompt. And I love that. That is very magical. That was always in gaming. That was the thing that like excited me the most. It was discovering something new, especially if it was a true discovery. It wasn't put there by somebody else. Maybe they just enabled the mechanics around it.
36:49So I think there's a huge opportunity in that space of gaming when you think about games in just a different thing and embrace this technology in a very different way. It reminds me of how some of the earliest use cases for GPT-3 were kind of these text games. So it's different from how you think of a playable video game, but actually a lot of these mechanics are very game-like. Exactly, yeah. I think there's still constraints, and I think that's going to be the mechanism design. That's still very human. Like a lot of the early games with GP3, they're kind of like, yeah, it was fun for a minute, and then it kind of went off the rails, and you're like, I don't really know what I'm doing anymore.
37:25But again, this is sort of, in some ways, Sora feels like a little bit of that, where it's got a little bit of gaming DNA inside of it, where it feels very fun and different and exploratory. So I like things like that, and I think there's going to be more use cases that we can't even think of as too creative. What are you guys seeing on the creative filmmaking side? Like, is that an important target market? Do you want to do an empower the long tail? Or do you want to empower the head, so to speak of the creative market? I think it's a really good question. You know, we've benefited a lot from creatives who are really willing to like go all in on, you know, even like the early technology, like Dolly 1, Dolly 2, and really like help steer us along the path.
38:05And like, I think it's important that we continue to, you know, build things for those folks. And we are working on some things that are like more targeted towards like creative power users long term. At the same time, I do think AI is a very democratizing tool at its best. And so what's kind of beautiful about the Sora platform in general is whenever someone kind of strikes gold, you see one of these beautiful anime prompts that goes to the very top of the feed for everyone. Anybody can go and remix that. Everyone has the power to build on top of that and learn from all of these people who come in with this incredible knowledge about how to really get the most out of these tools.
38:43And so I am really excited just to see the net creativity of humanity just increase as a result of this. But I think a big part of that is continuing to empower people who are always at the frontier, which are these more pro-oriented creator type folks. And so we want to keep investing in them as well. We've nerded out for a while, like almost a couple of years now, about that vision of feature film length content. like yes you have these amazing cameos and shorter content but at some point the individual creator it's been something that you've been excited about for a very long time yeah when do we get there you know is there a point where we have a feature film that is created on Sora 2 yeah and how do we consume it is it in the Sora app is it posted somewhere else online do you go to a movie theater and watch it yeah it's a great question I mean I think this will happen in stages to some extent So like if you guys watch the right, the launch video, I mean, that was maybe like Daniel Fraden, who's on the Sora team.
39:41And he already with these tools, right, is able to pump out these like incredibly compelling short stories within like days at most. I mean, he literally made that like all by himself in almost no time. And he's been like continuing to like put new ones out there on like the open AI Twitter since. Clearly, this is like massively compressing the latency that's associated with like filmmaking. I think to get to the point where like really anybody can do this, right? Like any kid in their home can just like fire up the app or sore.com or something and go and make this. It's really like an economics problem of like the video models.
40:16Video is the most intensive compute intensive modality to work with. It's extremely expensive. And you know, we're making good progress on the research team, like really continuing to figure out ways to make this affordable for everyone long term. Like right now, for example, you know, the store app is like totally free. in the future there will probably be ways where people can pay money to get more access to the models just because that's the only way we can really scale this further but you know I think we are not far off from this world where anybody can really like have the tools to make amazing content you know I think there's gonna be like a lot of bad movies that get created by this but like likewise you know there's probably the next great film director who is just kind of like sitting you know in their parents house like still in high school or something and just like has not had the investment or the tools to be able to like really see their vision come to life and we're going to find like absolutely like amazing things from like giving this technology to like the whole world i'm looking forward to the feature film length constantine's greek odyssey coming to theaters near you we're all in it together actually different characters i play the cyclops it's a good one i think um just to touch on that one more thing that something i've learned from recommender systems over and over again is that like oftentimes so the tools getting more people more creative is going to be a huge unlock for just you know making people more more creative in general and because you don't need this access to this like filmmaking equipment all that sort of stuff um but we do consistently see that things content is like also a social phenomenon in a way and like uh movies and all that all all everything you see out there is kind of a bit of a social phenomenon in addition to the actual content itself and so i think we're going to enter a very interesting world where there's so many people creating and so much content out there that even the idea that people aren't paying attention to and watching it is going to become more and more important.
42:04And I think that's actually going to make the quality of content just to kind of elevate because there's just anybody can create and actually it's going to be the consumption that's going to be quite limited, which is very different than the world we live in today. You guys are very thoughtful and intentional about how you treated IP holders. Can you say a word on that? You know, we've been in close partnership with like a bunch of folks across the industry and like really trying to like both show them kind of this like new technology. Right. That is actually like a huge value proposition for rights holders across the board.
42:36Right. And like we're hearing so much excitement from the folks we're talking with. Like they really see this as being like, you know, a new frontier for, again, like, you know, every kid in the world having the ability to like go and like use like some of this beloved IP. and really bring it into their lives in a way that feels much more personal and custom than what's been possible before. At the same time, we really want to make sure that we're doing this in the right way. So we've been really trying to take feedback and really steer our roadmap in a way where we know that both users are going to have an awesome experience getting to use this IP, but also the rights holders are going to get properly monetized and rewarded in a way that everyone wins, basically.
43:15So we're right now actively working on trying to scope out the exact details about how we're going to you know for example make it so if you want to cameo your favorite character from some like beloved film or something um you can do that in a way where uh you have access to it but like monetization will flow back to the rights holder right so really trying to figure out this kind of like new economy for creators uh we kind of just have to create this from scratch right now there's a lot of deep questions about how to do this the right way and you know as with everything with this app we come into it with like an open mind and we hear feedback and we iterate quickly you We're not sure where this is going to totally converge, but we're working closely with people to figure it out.
43:52Really cool. What's ahead? Pets. Yeah. I think, I mean, one... Sorry, what? Pet cameos. Cameo your pets. Is that one of the most demanded features? Great breaker. For me it is. Bill's demanding. I will remind us we were just talking about curing diseases and world models, and now we're to the future. This is something... No, it's actually... So that's definitely true. we've committed to that is coming um but we have we i promise the uh we actually had bill's dog as like when we were playing around with this rocket the goodest boy yeah uh and actually was very very cool um to actually feature a pet you can imagine where that goes it doesn't have to necessarily be a pet um could be anything a clock or whatever you have um well yeah yeah um you have a special clock actually it's really compelling i didn't think it could be so compelling until thomas showed me this clock yeah it's like a sentient clock it's well it's like based on like a real yeah i had a clock my father as a my father was a technology person for a while his company veritas gave him a clock for his like whatever anniversary anyway so i have it on my my uh my uh like table somewhere and uh there's this old simpsons episode where they talk about a walking clock and for some reason that's just been an earworm in my head for the last 30 years And so I always, it's like, you know, they're telling some joke and it's like, is it a walking clock?
45:18It's a walking clock. It's like walking clock. And then it's like, no, man, it's my dog. And so it connected in my brain where I was like, okay, rocket, walking clock. And then so I tried it. Thomas is great. This is what AI enable. Yeah. So it connected to my brain and we've been playing around with this just to see if we can get it to work and whether there's something special there, which is part of the fun of being on the sword team is you get to play with this emergent crazy technology and like maybe it does something you wouldn't even have expected. So I recorded a two second video of my clock and then I gave it some cameo instructions and I said, you're just a walking clock.
45:54You're a walking clock. You talk like you talk, you're a character. And then I generated my first video and it was insane. It was crazy. It was a walking clock. And then I had one where it was talking to Bill and Bill was like, I didn't think it would ever land, the pet cameo feature. And then walking clock's like, here I am. I just landed. So it's coming. It's all internal memes. Talk about emergent IP. Who needs Pokemon when you can have a walking clock? What's the greatest IP? One thing to add in terms of the future, I think on the feature film question, something I think about all the time is like, what will that actually look like?
46:31I think my, I mean, caveat, Bill's the only one who's good at predicting the future here. but my sense is that the you know as we get to longer forms what our equivalent of a feature film will look and feel very very different from what a feature film is today you know i don't know exactly what that looks like but i think on the subject of creators and what what's coming in the world i think a new medium and a new class of creators new class could include a lot of existing creators and um and support existing sort of mediums and stuff like that but i think we're just in the early innings of of what i imagine will be the next film industry rather than thinking about this being a feature film but i think there'll be something new there's some anecdote i hope this is true because i say it all the time but apparently when the recording camera like you know hit the world the first thing people did was record plays this is like the least interesting thing you could do with a recording camera it's like what's the big idea oh we people don't have to travel around acting we can just film them and distribute it and then someone was like wait a minute we can make a film and film in all these different areas and i feel like we haven't we're in like the first inning of so many different sort of things that people will do with this technology especially as the constraints change with latency and length and all that kind of stuff so cool and fun film history uh nerd fact is one of the original videos and we should check this as well but i think the original video was made just down the peninsula uh to settle a bet on if a horse, when it galloped all four legs, it left the ground.
48:04And I could see a world where you have new, that is an example of new scientific discovery. People didn't actually have an answer to that. Now that you have a new simulation format, what are we going to be able to discover in that? It will be crazy. And I think one, one broader point here is, you know, this app right now feels very familiar in a lot of ways, right? It's like a social media network at its core, but fundamentally like the way that like we really view it internally right is with cameo we've kind of introduced the lowest bandwidth way to give information to sora about yourself right aspects about your appearance about your voice etc you can imagine over time that like that bandwidth will greatly increase right so the model deeply understands your relationships with other people it understands you know more than just how you look on any given day um it's you know seen your full, like how you've grown up, all of these details about yourself and will really be able to almost function as like a digital clone, right?
49:03It's like, there's really a world where the Soar app almost becomes this like mini alternate reality that's running on your phone. You have versions of yourself that can go off and interact with other people's digital clones. You can do knowledge work. It's not just for entertainment, right? And it really involves more into a platform, which is really aligned with kind of where these like world simulation capabilities are headed long-term. And I think when that happens, the kind of immersion things we will see are crazy and you know for opening across the board it's really important that we kind of like iteratively deploy technology in a way where we're not just like dropping bombshells on the world when there's like some big research breakthrough we want to co-evolve society with the technology and so that's why we really thought it was important to like do this now and like do in a way where you know we've hit this again this kind of gpt 3.5 moment for video let's make sure the world is kind of aware of what's possible now and also you know start to get society comfortable and like figuring out the rules of the road for this kind of like longer term vision for where again there are just copies of yourself running around in Sora in the ether like just doing tasks and like reporting back in the physical world because that is where we are headed long term so cool so you're building the multiverse actually kind of yeah okay well can can timid me go and find my soulmate somewhere in there I mean anything is possible in the multiverse That's a call for action, everyone.
50:20It is kind of crazy, though, because now I'm going to sound totally cuckoo, but if we're in a computed environment, you're building the perfect simulator. That kind of is the way you ultimately understand and break out of the computed environment, right? Are we getting closer to the heart of the matrix? We're in the matrix. We're in the matrix? Some very deep existential questions. Yeah, yeah. What's your guys' P of we're simulated? Like, this is all... Rising. yeah me too oh what's your i'm low um yeah oh man but yeah it's okay really uh okay i'm just a believer i'm just like you know what this has got to be real yeah i feel like i'm not like solid 60 i don't know like more likely than not at this point i'm i'm there too whoa yeah zero so we might make a calc on it trivially small what's the oracle that that's where 10 will answer your sorority yeah yeah what do you think are the theoretical limits to sora yeah it's actually a great question um i thought a little bit about this like i think there's like a question can you eventually like simulate like a gpu cluster right in sora or something and i i assume there are some very well-defined limits on like the amount of computation you can run within one of these systems, like given the amount of compute you're actually running it on.
51:44Um, I've not like thought deeply enough about this, but I think there are some, like, there's some like existential questions there that need to get resolved. Yeah. Yeah. See, that's why his PCM is so high. Fascinating. Wow. You got a few lightning round questions for the team that we just kind of generated on the fly here. Um, and take your time, jump in whenever you have an answer your favorite cameo on Sora to date and what happened that is so tough I have a hot one
52:17okay so the there was this tiktok trend of I and I got obsessed with them I don't know why but these Chinese factory tours where they're like hello I'm the chili this is the chili factory they get like one like and it's me uh and it's like they're showing their chili factory and they're like it's the chili factory like this is amazing uh and or like there's an industrial uh chemical one there's uh yeah i've lost the name but there's an industrial uh uh chemical factory and uh the first day um i had my cameo options open just because i was like i just want to see what happens and the first day late at night, I opened my cameos and I was starting to get tagged in factory tour cameos that were all in Chinese.
53:06And I was like, I'm in the chili factory. And I was so excited. I get zero likes. I liked it. It was just me. But I was like, I'm the chili factory guy now. I'm like doing the ribbon cutting at the chili factory. Amazing. That's too deep of a cut though. so congratulations fun fact i actually have done chinese factory tours they're amazing in real life and they are truly epic yeah uh there's this one just i saw of mark cuban and jorts dancing around that was pretty good that got me good um but i mean my more back to the like just i scrolling the latest feed and just seeing like the wholesome content of people like doing things with their friends actually i think what brings me the most joy of they're not like super liked but it's like people just like getting a lot of you know value obviously from just like making videos with their friends so sam has so many bangers i like the one of him doing like this k-pop dance routine about like gpus or something it's very good actually i would put it on my spotify if like we had the full song wow it's very good it's like generated by sora it's like like very compelling yeah all right well that leads to the next one because you mentioned spotify what does an ai fully generated ai win first oscar grammy emmy i think the like logical answer is like a short winning an oscar yeah i think that's probably right yeah what would we win it for like for like the jorts yeah the jorts the jorts trilogy yeah yeah we need new content yeah i do think if people stitch things together in an interesting way yeah yeah i think there's You can actually start to make some very compelling storytelling in that.
54:46And I don't think it's like, it doesn't really feel like AI anymore. The content I'm seeing, like that was actually something I noticed with Sora as well. Just like, it wasn't even noticing it was AI. It was just kind of interesting content. That's a more interesting question. Will we know? Oh, yeah. Maybe it's already happened. Maybe it's already happened. I feel like for Oscars, one of the cool things that'll be unlocked is this long tale of epic stories in history, stories of heroism and struggle and all of these things that have been locked up because of the cost of creating. As a history enthusiast, I cannot wait for AI to unlock all of those stories.
55:30Have you seen the Bible video app? No, I haven't. Oh, it's really good. but I'll show it to you after. Like, perfect example. Yeah. Or there's this movie, The Last Duel, a few years ago about this really terrible crime that was committed in medieval France that was historically relevant and, you know, basically says a lot about humanity. And it just got picked up because eventually Hollywood picked up this important story about humanity. But how many more are there in human history? That's going to be really cool. Favorite character from any film or TV show? i have a really random one go for it uh you guys see madagascar oh yeah king julian oh it's played by sasha sasha baron cohen is a lemur he's a lemur yeah absolutely it's just that's a banger it's his humor meets kid-friendly storytelling it's just it's perfect uh i play a lot of video games so i mean your classic answer is gonna be like mario or something like that although i'll do the deeper cut of we were always joking about the rapper yeah yeah yeah for rap of the Rapa, the old PlayStation game, one of the original rhythm games.
56:33And it's got a great artistic style, and it's got a great IP of just this little dog. What is he? A dog. He's a dog, yeah. That's a good pick. When I was a kid, I played the Pokemon trading card game competitively for a while. So I was really in the Pokemon rabbit hole. So like, I don't know. Pikachu or Snowline. Pikachu. Mudkips. Super non-consensus. A fringe, deep cut. Okay, first world model scientific discovery. Most specific possible. Obviously, you're not going to say the discovery. I suspect it will be something related to classical physics. Like a better theory of turbulence or something.
57:18That would be my guess. I was guessing that it was going to be something like that. I was like, yeah. Navier-Stokes, I don't know. Yeah, some fluid dynamics thing that's maybe hard to understand now. There's a lot of unsolved kind of problems there. I think sometimes they call it like continuum mechanics where it's like in between and we don't have good models of them. Something that lends itself to simulation, just like the amount of iterations you can do of a simulation, unlocking something, which I don't, yeah. Something in that realm. The last thing we'll be able to accurately simulate. I do think there's like a set of physical phenomenon for which video data is like a poor choice of representation.
57:55Right. Right. So like, for example, is it really efficient to learn about, you know, like high speed particle collisions or something from like video footage? Maybe. I really think video is at its best when, you know, the phenomenon that you're trying to learn about is just natively represented in the physical world. And so when you need to do like, you know, like quantum mechanics or some other discipline where, you know, it's more theoretical. We don't have like video footage beyond. Can't see it. Yeah. Things that we've like manually rendered for like educational purposes. It feels like a weaker medium for understanding those things.
58:35So I suspect those would come last. I guess it's the things we don't have sensors for. Right. Right. Yeah. Maybe the last things we care to simulate is another way of thinking about the answer. I don't know. I mean, people aren't doing much with smell right now. True. Greenfields. I've been meaning to tell you about that. It's kind of awkward. We're still trying to figure out how to simulate Thomas with bad hair. Oh, yeah. It remains an unsolved problem. Not even Sora can do it. Thomas's hair flow. Just general. Guzzling, ketchup, yes. There was a good round of people being bald. We were all doing bald.
59:07Oh, yeah. Bald gems were good. Actually, kind of cool. That's our use case that doesn't really talk about very much, but it's like. Visualization. when you're bald yeah everybody wants to be bald no it's just like you just see yourself in some different context i think that can be quite powerful even like therapeutic in some ways where you just like see yourself in some context that you either want or don't want yourself to kind of be in and just see see yourself it's a real use case yeah yeah true guys thank you so much for coming from space-time tokens to object permanence world models that will enable scientific discovery, the democratization of creation, all the way to walking clocks.
59:46You guys have covered it all. Thank you so much. And the future is being created by you. Thanks, Konstantin. Thank you, Sonia. Thank you.
1:00:10Thank you.
From the publisher
The OpenAI Sora 2 team (Bill Peebles, Thomas Dimson, Rohan Sahai) discuss how they compressed filmmaking from months to days, enabling anyone to create compelling video. Bill, who invented the diffusion transformer that powers Sora and most video generation models, explains how space-time tokens enable object permanence and physics understanding in AI-generated video, and why Sora 2 represents a leap for video. Thomas and Rohan share how they're intentionally designing the Sora product against mindless scrolling, optimizing for creative inspiration, and building the infrastructure for IP holders to participate in a new creator economy. The conversation goes beyond video generation into the team’s vision for world simulators that could one day run scientific experiments, their perspective on co-evolving society alongside technology, and how digital simulations in alternate realities may become the future of knowledge work.
Hosted by: Konstantine Buhler and Sonya Huang, Sequoia Capital




