V-JEPA, AI Reasoning from a Non-Generative Architecture with Mido Assran - #677

25 Mar 2024 · 48 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

TWIML AI Podcast Episode #677: V-JEPA, AI Reasoning from a Non-Generative Architecture

Episode Overview In this episode of The TWIML AI Podcast, host Sam Charrington interviews Mido Assran, a research scientist at Meta’s Fundamental AI Research (FAIR). The discussion centers around V-JEPA (Video Joint Embedding Predictive Architecture), a novel model aimed at advancing the understanding of artificial reasoning in AI systems, as envisioned by Yann LeCun.

Key Concepts and Discussions

Introduction to V-JEPA

  • V-JEPA is derived from Meta's Joint Embedding Predictive Architecture (JEPA) and focuses on efficient learning by using self-supervised methods to interpret video data.
  • The model aims to bridge the gap between human intelligence and machine intelligence by enabling AI to learn abstract concepts with fewer examples, mirroring human cognitive abilities.

Motivation Behind V-JEPA

  • Traditional machine learning models require extensive training data and compute resources to understand concepts.
  • V-JEPA seeks to be more efficient, similar to how humans learn from minimal examples and with lower energy consumption.
  • The focus is on feature prediction rather than direct pixel-level predictions, emphasizing the need to encode higher-level representations for effective learning.

Core Mechanism of JEPA

  • The JEPA model operates on two signals, X and Y, where the aim is to predict encodings of Y from encodings of X rather than predicting the raw pixels.
  • This approach reduces computational load and enhances efficiency, especially relevant for complex data types like video.

Self-Supervised Learning

  • V-JEPA employs a self-supervised learning framework, allowing the model to learn from unlabeled video data.
  • The model uses masking techniques to predict missing parts of data, enabling it to focus on learning meaningful abstractions rather than irrelevant details.

Joint Training of Encoder and Predictor

  • Both the encoder and predictor in V-JEPA are trained jointly to ensure the embeddings produced are semantically rich and useful for the prediction task.
  • This joint training mitigates the risk of representation collapse, where the model may fail to learn meaningful distinctions.

Challenges and Future Directions

  • Transitioning from image-based to video-based architectures presents challenges, particularly in data acquisition and processing efficiency.
  • The podcast discusses the need for better curated video datasets that parallel the extensive resources available for image datasets.
  • Future research will explore extending the framework to other modalities (like audio) and refining the model to handle various prediction tasks with diverse temporal scopes.

Hierarchical Learning and Multimodal Integration

  • Mido highlights the potential for hierarchical JEPAs, suggesting that different levels of abstraction could improve the model's performance across various tasks.
  • The integration of multimodal data is seen as a key avenue for enhancing the world model's capability, allowing for more comprehensive understanding and reasoning.

Conclusion The conversation emphasizes the innovative nature of V-JEPA and its potential to advance machine learning by focusing on predictive learning strategies over generative ones. The model represents a significant step toward achieving more human-like reasoning capabilities in AI.

Key Takeaways

  • V-JEPA is a promising development in AI aimed at more efficient learning from video data.
  • The shift from pixel-level prediction to abstract feature prediction can significantly reduce computational burdens.
  • Self-supervised learning techniques and joint training of models are critical to achieving meaningful semantic representations.
  • Future directions involve exploring hierarchical models and multimodal data integration to enhance predictive capabilities.

For further details, check out the complete show notes at [twimlai.com/go/677](https://twimlai.com/go/677).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Hey, what's up, everyone, and welcome to today's episode of the podcast. I am Sam Charrington, founder of TwiML and host of the TwiML AI podcast. And today I'm joined by Mido Asran. Mido is a research scientist and Metasphere or Fundamental AI Research Group. If you're watching us on YouTube, be sure to take a moment to hit that subscribe button. And if you're listening on Apple Podcasts or Spotify, it is the follow button that you're looking for to make sure you don't miss any other great content that we've got in store for you. Mido, welcome to the pod. Thanks so much for having me, Sam. So it's a pleasure to be here.

0:40I'm really excited about this conversation. We're going to be discussing vJEPA, which is being billed as the next step toward Jan Lacoon's vision of advanced machine intelligence, referring, of course, to the video version of FAIR's JEPA or Joint Embedding Predictive Architecture. Let's jump right in and maybe have you riff a little bit on the broader goals of this project and unpack the idea of feature prediction that is core to the JEPA approach? A big motivation for the work is really efficient learning in general. So before this kind of project, me and my collaborators were working on self-supervised learning more generally.

1:24And this idea that really noticing there's this big gap between humans and machines in terms of how efficiently we can learn, right? Humans' brain runs maybe on the energy equivalent of an electric razor, right? And we learn from very few examples. You know, I could show you a single example of something and you can acquire and understand those concepts. Machines, on the other hand, require thousands of hours of training, tons of examples, very compute intensive to understand and acquire a concept. This was a problem that we've been thinking about, me and my collaborators, for a few years. And we proposed several approaches in literature, even our own techniques, focused, for instance, on learning from images.

2:03They were relatively effective, published in kind of the typical machine learning conferences or computer vision conferences that you would expect. But what we noticed was that kind of our approaches and our methods were kind of very specific, a lot of very specific inductive biases tailored, for instance, to learning from images. And of course, we want to learn from much more than just images. We want to learn from perceptual data more generally, like video, audio, and so on. And so this is where we felt like, maybe there's something here, we need to step back. And Jan Luka, our, our chief AI scientist, kind of proposed his vision for for AMI.

2:40And actually, he'd been talking about a lot of these ideas for a few years, even before posting his kind of his his document online. And so really, we kind of stepped back and said, Okay, what if we just try more generally, the idea behind the JEPA is, you know, you have two signals from your environment, X and Y, and they're related somehow. And the idea is you want to predict encodings of Y from encodings of X. So again, a very kind of abstract and general framework. The contrast is relative to what we would often do today is, you know, the case of images or video predicting the pixels themselves as opposed to some kind of representation of the pixels or what's happening in the images or video.

3:26That's exactly right. So, I mean, what's common is predicting Y from X directly rather than encodings of Y from X. The idea was really like you want to predict encodings because predicting kind of low-level pixels is A, a much harder problem, B, much more compute-intensive, and C, I mean, you have to ask yourself what the goal is. If your goal is to build a generative model that can produce kind of visually appealing content, that makes sense. But if your goal is to really build a world model capable of, you know, understanding the world, making predictions about it, reasoning and planning, how much do you really need to spend capacity modeling every single kind of blade of grass, every single part of this plain blue sky or cloud?

4:08It's not efficient. That was kind of the idea. And so there were kind of some generative approaches. And what JEPA prescribed was doing this kind of prediction task in this encoded or embedding space. so um and it's you know the idea is you have an encoder so something that encodes your signals x and y and you have a predictor something that can predict encodings of y from encodings of x and that's that's but then you know what does that actually look like how do you implement it and so that's where we started thinking and really identifying some of the kind of research questions to start out with and how to start building up this approach our first work along that direction was the image JEPA.

4:48So it was trying to get this just to work with images. So the first question you ask yourself is, well, what are Y and X here? You know, if you're trying to learn from images. So here you could, for instance, say, well, Y and X are just masked versions of the same image. So I see part of an image and I find predicting codings of the other part of the image. Okay, so that's great. You know, we have Y and X now. We're specifying them using masking, which is convenient because we can leverage kind of a lot of modern tools in the machine learning literature, especially from the NLP domain, leveraging kind of masking.

5:20But we're doing it in the context of computer vision. It unlocks the whole, basically all of the video or images on the internet and the whole idea of self-supervision. Pretty much, exactly. So the idea is now you can learn through kind of passive observation, large quantities of data. Of course, you want to do this responsibly, but it really does unlock a large, vast quantity of data, similar to what kind of humans are exposed to through kind of passive observation. Because humans acquire a lot of understanding of the world, even just throughout the first few months of development, before any kind of capacity for language.

5:53There are these really kind of fun tests that cognitive scientists do with children to try and understand, you know, have they acquired object permanence, this idea that, you know, if an object is occluded, it still exists. Have they acquired, for instance, this understanding of, you know, motion consistency or shape consistency. so you know there's been a lot of studies um kind of child development and you do see these concepts kind of emerge within the first maybe 10 11 months of development a lot of them do and that's before any capacity for language so if we want to get closer it's where it's kind of bridging this gap between human level intelligence or kind of how humans learn in human intelligence and machine intelligence we're trying to step back and really take it from first principles and that's what the JEPA kind of proposes and that's where this research kind of starts to make some progress and propose specific instantiations of this framework.

6:42So it's not that this is the only instantiation or this is the way to do it, but it's kind of first instantiation showing remarkably, actually, that it does work. I'm glad that you agree that it's remarkable. When I think about the polls here, predicting from pixels, it is clear that you end up with a lot of extraneous information that's not germane to the key points that you're trying to learn. On the other hand, embedding seemed very blunt for trying to learn the kinds of things that we're trying to learn. You know, for example, actions and principles about the world, about the environment. Do you have any intuitions about why it works?

7:25Yeah. So actually, I think you kind of hit the nail on the head is like, what are these embeddings exactly? We're saying encodings of some signal embeddings and these embeddings could really. contain anything, right? So in the extreme case, these embeddings have done no kind of semantic abstraction whatsoever, and they just contain the full signal X and Y. And on the other extreme, which is something you actually want to avoid, is that these embeddings could just contain no information about the signal X and Y, it just maybe outputs, it's a constant regardless of whatever you feed the network.

7:57And in that case, the prediction task is very easy, right? But it's trivial. You haven't actually learned anything. So that's really where kind of the research part came in is how do we bootstrap this network? Like we're not really giving it any information. We're just giving it kind of images or video showing it mass parts. The networks, the encoders themselves are randomly initialized. We're not starting from pre-trained networks. The predictor itself is also randomly initialized. So at the start, I mean, you know, it's all kind of random. So how does the network kind of bootstrap itself to actually learn encodings that are semantic and meaningful so the prediction task actually makes sense, but at the same time, eliminate kind of this irrelevant or unpredictable information.

8:38And so you need kind of two kind of forces here. So one is you'll notice actually the network by itself, kind of the way the JEPA is set up, is incentivized to make the prediction task as easy as possible. So if you encode everything about X and Y, that's a very difficult prediction task. On the other hand, if you make it super easy, you just output a constant vector, regardless of whatever you give it, that's a very easy prediction task. So the network is kind of already incentivized towards making things simpler. Now, the challenges don't make it too simple. And so that's where we kind of leverage existing techniques from the self-supervised learning literature.

9:16So not unique to learning from images or video, but kind of more general techniques people have proposed for preventing what's called representation collapse. So that's, you know, So representation collapse comes up in a lot of different settings and scenarios in the context of machine learning and has been previously studied in self-supervised learning as well. And so we kind of set up this architecture and then tried to figure out how we can integrate previous or existing approaches for representation collapse to prevent these kinds of trivial solutions and then left everything else kind of up to the network.

9:47And so are the encoder and predictor jointly trained or are they separately trained in kind of an adversarial kind of way? I find that idea quite interesting, but no, they are jointly trained. There's no adversarial component. So right now, and that's actually one of the reasons you can have representation claps as well is because if you were to, for instance, train the encoder beforehand, so you know it's coding something interesting and then you freeze it. now the predictor is kind of grounded. So you don't have to worry about representation collapse. But because we're training them jointly, it makes the problem much more difficult.

10:27But that's actually one of the hypotheses about the project is kind of the importance of training them jointly. Because when you train your encoder, you have to ask yourself, what objective am I trying to train my encoder with? Like, what is it trying to capture? And if you train them jointly, what it's trying to capture are features that are predictive. Right. And so that's where you eliminate kind of unpredictable or irrelevant information, because it's only going to try and encode whatever is predictable, given kind of even I mean, and that depends on a lot of things, right? Inherently in the world, there is kind of some irreducible amount of uncertainty.

11:04Right. So that is probably going to be eliminated. But also, depending on kind of how much capacity you have in your predictor. So is this a big network? Is it a small network? What is its architecture? that also controls to some extent what the network is able to predict. And so that's going to influence kind of what you end up encoding. And so there's a lot of dynamic factors that kind of come together. And so that's why we believed it was important to train them jointly. And that's why it's important to figure out a lot of these little design details, the predictor architecture, what are you trying to predict from what exactly, because even the masking strategy, you know, if you take an image or a video, and you just mask out little patches here and there, The same way people used to do in natural language processing, they would kind of take a sentence and mask out little spans here and there.

11:48If you do that with perceptual data like images or video, you actually don't really need to do a lot of semantic abstraction. You can pretty much predict these kind of little missing regions just by looking at the neighboring visible regions that you see. And so you get really terrible representations if you do that that way. So that was another thing to figure out is, what are you actually trying to predict from what? And so it turns out we want to predict these big chunks from the video or the image that are masked maybe for the entire temporal duration of the video. Because now the only way the network can predict what's there is if it kind of, A, does some kind of semantic abstraction and kind of discards some irrelevant information, whatever is not predictable.

12:27And so it forces the network to really produce these kinds of semantic embedding. So there's a lot of really interesting dynamics at play. And is the masking approach and strategy, is that learned as well? Or is that based on rules or heuristics? Yeah, I think ideally that should be learned. Right now, it's not. Right now, we kind of figured out some hyperparameters. And we just say, you know, given a video, you mask out, you randomly sample blocks taking up roughly the size of the video. So you end up with actually roughly 90 % of the video masked, but in a very kind of specific way where we've sampled a few different blocks and stitched them together.

13:04So it's heuristic in that sense. We kind of figured out this was a kind of strategy that worked well. But I think ideally also the network itself would figure out kind of what it should mask and attend to. It does seem like in the same way that you want the encoder and predictor to kind of balance one another and find this happy medium of optimal knowledge transfer, the degree of masking plays into that as well. Definitely, definitely. And it's crucial for learning semantic representations. So it's the degree, but also kind of how spatially or temporally connected are these regions? Because if it's kind of sparsely just dispersed throughout the video, again, it won't work.

13:44So I think what you actually want to do, but this, I would say, is still very much an open research question, like what you want to do here. But my opinion about what you want to do is you want to mask out semantic regions in the video. Essentially, like you want to take the real semantic content. So we're not talking about like blue sky or just fields of grass. You know, we're talking about, you know, the objects and their interactions, even temporally, right? Like you want to mask out things where interesting things happen and ask the network to predict that. And that's what seems to lead to kind of more semantic representations and levels of abstraction.

14:21And I think there's lots of interesting connections to, I mean, humans, right, like we, we visually attend to a lot of different things. And actually, if you look at visual attention, again, in children, as they develop, it significantly changes from like, the first month to the first 10 months to the first two years, kind of what we choose to visually attend to and the things we see. So it's clear that that's some important component of the learning mechanism in humans. And probably that's something we should be thinking about also when trying to leverage JEPAs or general mask modeling methods for learning from videos and images.

14:54You mentioned that the mask regions are randomly determined and that at some point in the future, you know, that masking determination should be part of the network. Is it fully random or is there some maybe traditional object detection that tries to identify the semantically interesting parts of the image and then masks those in a random way? Right now it is fully random. It is fully random. So, but again, I do think kind of using more semantic knowledge, maybe using an object detector makes sense. The question, though, is just how do you do it efficiently? Because it's all about kind of efficient learning.

15:39And if you have to run a very expensive object detector to figure out the mass regions, that's not going to be ideal. So the question is, how can you maybe leverage some of that knowledge or some of those insights, trying to take inspiration for maybe implicitly understanding and masking out where the objects in the video or the image, but doing it without having to run this big kind of pre-trained object detector network. And again, ideally, it's something that you would do kind of in a fully unsupervised way. So we don't want to be bottlenecked by the quality of the object detector either, which requires maybe, I mean, current object detectors, right, require the state of the art ones, require lots of labeled examples.

16:19And so one of the key motivating factors is doing kind of large scale unsupervised learning. So learning from streams of videos or images. Can you talk a little bit about some of the enabling work that you're building on here, the paper references, things like video transformers, masked autoencoding frameworks, query-based feature pooling? What are all these components and how do you kind of pull them into the pipeline here? Like I mentioned, JEPA has kind of prescribed a more general kind of learning objective. But really, there's a lot of tools that have been developed by the machine learning literature across various communities that we can integrate.

16:59And then, in fact, we do try to integrate in VGEPA to really explore how far we can push this approach with modern tools. And so one of the architectures that we leverage is a vision transformer. So here, when I mentioned we have this encoder that computes encodings of X and Y, and these networks are randomly initialized, these networks are actually vision transformer architectures. And the predictor as well, which kind of predicts encodings of Y from encodings of X, it's not a vision transformer, but it is a transformer architecture based on the transformer architecture. So it has self-attention, it has kind of multilayer perceptrons, residual connections, and it processes sequences of tokens.

17:41And the reason we choose to use, for instance, transformers for the encoder is because of the fact that we're doing kind of masking. I mean, that's how we chose to design the signals Y and X. In theory, though, right, Y and X, I mean, suppose you're trying to learn from video. Y and X don't have to be mass parts of a video. They could just be the same video but corrupted. You could take that full video, imagine just adding some noise to a part of it, or imagine applying some kind of transformation to that video, you know, maybe like some temporal shift or, or slow-mo, right, or like some spatial cropping or something like that, right?

18:17Geometric or photometric augmentations. So there's lots of ways to design Y and X. The way we've instantiated it in VGEPA is through masking, but I don't think that's the only way to do it. And once you're doing masking, one great thing you can do is leverage transformers to get efficient learning. Because what we do is, for instance, I mentioned we mask 90 % of the video. What that means is now the transformer only needs to process 10 % of the video. It's an architecture where, because we flattened a video into a sequence of tokens, because transformers don't really have any implicit, I mean, just transformers on their own don't have any implicit kind of understanding of, you know, spatial, you know, spatial distance in an image or temporal distance in a video.

19:00That comes into how you take the video and kind of convert it to a set of tokens. And so what we do is given a video, we take 16 by 16 pixels across two frames, and we call that a token, essentially. And then we feed those tokens to the transformer just as a sequence or a set of tokens. And now if you want to mask 90 % of the video, well, just drop 90 % of the tokens. And so you get efficient learning. And that was something, a technique that was kind of explored in literature with masked auto encoders and vision transformers, except they were, of course, trying to do their reconstruction in pixel space.

19:40But so again, it was really trying to integrate lots of these tools. We don't necessarily want to rebuild the wheel from scratch or be building in a bubble. There's a lot of great tools we can leverage. The other one you mentioned was query-based feature pooling. So that's something interesting. Previously, the video community had mostly focused on, you know, after training a video model, you fine tune it end to end, and suddenly it's just very task specialized. You have one model per task, you know, there have been like, there definitely have been works exploring kind of how do you leverage the same backbone across many tasks.

20:12But I would say the dominant focus in the community has been on just fine tuning producing task specific video models. And what we wanted to do was really ask, well, how can we push kind of frozen evaluation? So can we do unsupervised pre-training and get kind of this video encoder that really understands the world to some level and then freeze it? And it's never seen any kind of human supervision or labeled data. And now we freeze that encoder and then train task-specific models on top. So rather than having one huge task-specialized video encoder for every single task, you have a shared video encoder and a bunch of very lightweight task specific models they could be like a single layer neural network for instance or a two-layer neural network tiny tiny and so but the question is how do you take the output of your video encoder and feed it to these task specialized models because kind of the output or the structure of the video embedding you know might require might not be linearly might be hard to linearly combine kind of that embedding to solve a task So what that means is you essentially just need a smart way to kind of just take that output of your video encoder, represent it concisely as a single vector, and that can then be fed to a downstream model.

21:32And that's what query-based feature pooling is. It takes kind of the output of our video encoder, which is a transformer. And so a transformer will output a token, an output token for every single input token, at least the transformer architecture that we use. So we have one vector, one embedding vector for every single input token in the video. So you have something of roughly the same size at the output as it is in the input. And now we want to take that big feature map and convert it to a single vector that we can then feed to just a single layer, single layer neural network or something like that.

22:05And that's what query based feature pooling gives us. Previously, what you would do is you would just average. So that's average pooling. You just take that big feature lab, average it to get a vector. but if the structure of your embedding space is not kind of such that concepts or semantic concepts are linearly linearly separable this average pooling is going to potentially destroy a lot of structure and not work very well so what query-based feature pooling is it's just a single layer of attention you just have this learnable token and it cross attends to the output of your feature map. And so essentially you're just learning, it's a single layer network that just learns to pull this output feature map into a single vector.

22:46And so when combining that with VideoJepa, but also previous video methods, we saw a huge boost in performance in frozen evaluation. So that's something that that unlocked as well. This idea of being able to just take video encoders and use them off the shelf without supervised fine tuning. Yeah. And those smaller networks that you train for specific tasks are what is referred to in the paper as attentive probes. Is that right? Exactly. And so those smaller networks, we've actually, I've mixed kind of the description of the attentive pooling with those networks. So you have a single layer of kind of cross attention and then kind of a very small classifier network on top.

23:29And that's what we call the attentive probe is that pooler and that network together. Got it. Got it. And then in the paper and on the GitHub you're providing, so those attentive probes are trained on a task-by-task basis. So each one is trained separately on a given task. So we look at tasks like action recognition, so someone kicking a ball, someone ironing, someone folding laundry. And the video encoder itself has never seen any of these concepts. But we train these lightweight attentive probes on top or these lightweight task specific models to kind of take the output of the video encoder and then map them to one of those classes.

24:15So someone kicking a ball, someone ironing, and so on. So we released the weights for both the video encoder and the task specific attentive probes on top, which are these very lightweight networks that can just from the video encoder predict an action class. we also looked at other tasks like fine-grained temporal recognition so those ones are tasks like pretending to put something into something but not actually putting it into something or picking something up moving something down moving something from left to right or right to left up and down pretending to throw something but not actually throwing it so those kinds of tasks require a fine-grained understanding of uh of kind of temporality and you can't predict those action classes just by looking at one or two frames in the video.

24:57So there are a lot of kind of video classification tasks where really just you have visual cues and you can kind of understand what's happening just by looking or infer what's happening just by looking at one or two frames. So we looked at those two kinds of tasks that require fine-grained temporal recognition, but also more kind of coarse-grained action recognition. We looked at a few other tasks in the paper as well and released the probes for those kinds of task-specific models on the GitHub. up. There's a note in the paper that kind of suggests that there's a sweet spot temporally between, if I interpreted it correctly, three seconds and 10 seconds in which the network does a really good job, but longer temporal tasks are still a problem.

25:43Did I pick that out right? I think what we do during pre-training is we sample three-second clips of a video. So every time we give a model prediction task, it corresponds to roughly three, three and a half seconds of video. We mask out a part of it and ask it to predict the other part of it. So that's pre-training. For inference, when we come to solve kind of downstream tasks, we do go up to 30 seconds. We can give the model 30 seconds of video and it can make a prediction. But really, I mean, we haven't explored longer range temporal prediction yet during pre-training. So it's not necessarily that we haven't, maybe there's some inherent limitation of going beyond 30 seconds or something like that.

26:25But it's that the pre-training task itself was only focused on kind of this shorter range prediction. And most of the video tasks that we look at, the action recognition, the kind of fine grained temporal recognition, those ones are really, they go up to maybe 10 seconds. you know 10 seconds long so that's kind of what we've evaluated but i think future work is definitely to try and push this especially that learning framework is already pretty general and should be capable of doing it but then there are new considerations right so we mentioned the way the japa is kind of designed it's going to eliminate whatever information is unpredictable and so if you maybe just try taking the same approach and giving it 10 minutes of video and And say you have an efficient architecture that can handle process 10 minutes of video efficiently, for instance.

Read the full transcript

27:12But, you know, suppose we didn't have those kinds of constraints and we gave the model 10 minutes of video and said, predict what's going to happen over the next 10 minutes. You might get some different kind of learning dynamics, right? Because maybe what's predictable in 10 minutes is not the same as what's predictable in 3, 4, 10 seconds. What's predictable in 3, 4, 10 seconds is fine grained motion, right? Because probably it's going to continue along its trajectory. What's more predictable in maybe 10 minutes is something a bit more abstract or semantic. And so probably you want to do a kind of multitask or a hierarchical setup where you're predicting at multiple levels of abstraction.

27:48So we still do want that three-second prediction range. And the question is now, how do you combine that with longer-range prediction without kind of destroying what you're learning at this scale? And so that's where this idea, so Yann Nekun has talked about kind of hierarchical JEPAs in the past. it's not something we've explored in vchepa but i think that's where this kind of architecture idea would come into play you want to predict the multiple timescales and what you encoded those different timescales is going to be different as well there's a sense in which this approach is um you know maybe turning back the clock a little bit like a part of the the way i think about what's happened in machine learning is that like we were really focused on prediction for a long time.

28:30And then we kind of got the generation bug and everyone kind of switched gears to generation. And part of this paper is kind of holding up this flag that's saying, hey, prediction is really important and valuable. And there's a lot that can come out of a predictive approach. Can you talk a little bit about that contrast? I kind of like this idea of kind of going back to old ideas and saying, let's not forget, you know, these kinds of principles that we had come up with. For what it's worth, I won't comment too much other than just saying, I do think generative approaches have value. My personal opinion is the value is different, right?

29:05Like if you're trying to learn a good generative model, there's clearly a lot of social and intrinsic and maybe economic value, even scientific value in doing that. That's very diplomatic of you. Well, I think it's true to some extent, but I think if you want to learn like a, I do think it has a place. So I'm definitely not saying we should not be doing that. But I think if you want to build up build a model that can understand the world kind of and builds up a world model can understand it and use that understanding of the world for not just like you know content understanding and visual perception and more generally perception but to use it to build goal-driven agents so agents that you know identify a goal and then use their world model to kind of reason and plan and really identify decisions like you know make specific sequential sequence of decisions to arrive at their goal there i think it makes sense to focus on prediction because you know if you try and do it in this generative space i i think it's just it's too complex i'm not that it can't be done either but i just think it's very inefficient you need a lot of capacity a lot of compute and the type of uncertainty you have to model now has grown significantly right like you know if you want your encoder to capture something predictable and you're trying to use your model of the world to plan and reason, well, there's a lot of uncertainty now that you have to capture.

30:27That's going to, you're going to have to leverage and capture to come up with a plan to achieve your goals. So the amount of uncertainty you have to capture there is just, it grows exponentially in my opinion. And so that's why I think kind of latent prediction models are just focusing on prediction. Those are the approaches that I think are going to be more efficient, not the only approaches. I mean, generative could work, but I think those are the approaches that are going to be significantly more efficient, efficiently applied by a learning goal-driven agent. So if that's what you want to build, I think that's what you need.

30:59The idea that you are able to apply these different attentive probes to a fixed model suggests that the model or the embedding space that the model creates rather, you know, captures information about all of the tasks that you're ultimately trying to perform. You've talked a little bit about hierarchical JEPAs. Do you think that that hierarchy creates more narrowly constrained representations, or is it, you know, hierarchy that's imposed in some other way? This is, and the question also of kind of how narrow versus how broad your representations are is something the community has been grappling with for, I mean, representation learning community has been grappling with for decades, right?

31:49Because learning, in some sense, learning is all about semantic abstraction, right? So you have some input. What do you keep and what do you discard? Like that's really the question when we're talking about kind of learning. And so now how do you answer that? Well, what do you keep, what do you discard? Depends on what you want to do. And so maybe kind of discarding a lot of information and just retaining kind of shape is going to be really good for some tasks. Maybe if I only keep shape, maybe we're talking here about visual inputs. So maybe if I only kind of retain object shape in my representation, then I can solve specific tasks very efficiently with one, two examples, boom, I got the concept.

32:31But maybe now if I want to not just understand shape, maybe now I have a task that really relies on color. I want to distinguish one color car from another color car. Or maybe actually it's quite important to distinguish color, you know, depending on a task. If you don't have color in your representation, well, that's a problem, right? So even though it was very good for one task, it was not great for the other task. And so you'll find that tasks, depending on the task you want to solve, they often have conflicting requirements for what you want to contain in your representation. And that's the big question is now, like, what is the ultimate representation?

33:06What is this representation that's perfect for all the tasks you care about? What I've kind of been arriving at is maybe there really isn't just one good representation. That's maybe where the hierarchical part comes in. You need several representations at different levels of abstraction. So depending on the task that you care about, you're going to use representations of different levels of abstraction. And I think one thing we also want to move towards is, so it's very practical right now that we can train us as humans, we can train like task specific models on top of these representations. So maybe we choose the ones that are more suited for our task, and we apply them.

33:42But I think ideally, this predictor or the world model, like the predictor in VGEPA can be seen as a kind of first step towards a world model, it's able to kind of model uncertainty and predict missing regions in a video or an image. And I think ideally, what we want to do is not just throw that away after training, like right now in VGEPA, we visually probe it to kind of see what it's actually learned. And we were surprised to see that it's actually capturing spatial and temporal uncertainty because all of our evaluations were focusing on the representation of the encoder, right? We weren't, how do you evaluate a predictor or a world model?

34:14It's still a bit of a growing kind of research area that a lot of people are thinking about, but it's not as standardized as representation learning, for instance. You know, what I think you want to do in the future is have these representations not necessarily only applied or leveraged by kind of human engineers designing task-specific models. But I think ideally, you would feed those into a predictor that can then do the kind of reasoning and planning on its own, given all these encodings or representations. So depending on what the agent is trying to do, this predictor or world model is going to be able to configure kind of the output of the perception module on its own to figure out what is the right kind of level of representations and how do I leverage them and use them for this task.

34:56so it almost automates that a bit more do you see that as an area where LLMs and their reasoning abilities might come into play in some broader you know JEPA based world model architecture or are there other ways that you think that you will get there that's kind of more native to the direction you're going or what we've seen thus far So I think it's important to learn from perceptual signals directly. So it's important that you're trying to predict, you're trying to learn directly from video, from images, from audio, from whatever kind of perceptual modalities or signals that you have. Because I think that builds a very grounded understanding of the world.

35:48It's almost like, imagine trying to understand what a cup is. What is a cup? I can show you a few examples and you might understand it just from a few examples. Now imagine trying to describe a cup. Well, it's round, it has this kind of shape. Usually there's something inside, but that description is a bit too vague. Suddenly it becomes really cumbersome to describe these really intuitive concepts in language. And that's just the example of recognizing a cup. What about interactions that really have a temporal component to them or some kind of understanding of intuitive physics? I think it becomes really complicated to explain some of these perceptual concepts and signals in language.

36:30So I do think it's important to be learning in a kind of perceptually grounded way. But there are also lots of tasks we care about where we want kind of a language interface. And transformers in general, I think they have good capability for kind of symbolic reasoning. So that's once you have a discrete set of concepts, how can you kind of reason over those kinds of primitives that you've already defined? But, you know, a lot of the visual perceptual world is continuous, right? We have continuous pixels, space is kind of open and continuous. And so I do see kind of intersection of these two components coming together.

37:10but I expect the reasoning part, or at least kind of the, maybe there's various levels of reasoning here to talk about. And the kind of lower level, visually grounded, perceptually grounded reasoning about what's happening in the interactions, I think that would be happening from a model that doesn't necessarily have to have language. But when we talk about JEPAs, for instance, we mentioned, for instance, trying to predict in the future, again, the encoder is going and discard whatever information is not really predictable. So one way to kind of make things predictable is to give the predictor, the world model itself, some additional information.

37:46So imagine, for instance, a robot. A robot is kind of interacting with the world and it's trying to predict what's going to happen. And you do have information maybe about like proprioception. So how its joints were moving. So that's information that the predictor can use to try and predict what's going to happen in the future. It's almost the actions that the robot was taking. You can also imagine maybe actions being encoded in text or language, right? So the predictor is modeling uncertainty about the future, you know, given kind of text or information about what's going to happen in the future.

38:18So I don't think they're necessarily orthogonal. I do see these pieces coming together, but I do think for kind of building a world model and understanding of the world, a model that's capable of visually grounded reasoning and so on, you do need to be learning directly from kind of perceptually grounded inputs. The video transformer in JEPA is used in the encoder. Is the predictor also transformer based? Yes. So the predictor is transformer based. It's not necessarily a classical like vision transformer because it operates in latent space. So vision transformers, as they're prescribed, usually part of their prescription is like how you convert an image or a video into a sequence of tokens.

38:59That's called patrification, and that's an important part of it. The predictor operates directly on embedding tokens. So I don't know that I would necessarily call it a vision transformer, but all these architectures are pretty similar, and it is transformer-based. How do you get from a predictor that works well on an understanding of the world that's been learned up through something like VGEPA to, you know, our reasoner and, you know, Jan's vision of advanced machine intelligence. Yeah, I think it's definitely the first step is just showing that you got like the perception module and an initial version of a world model.

39:35Like you said, how do you get there? That's ongoing research, but I think it comes down to now focusing a bit more on the predictor or the world model. So because we mentioned early on in the talk that the predictor and the encoder are trained jointly, and a really important part of this framework is the embedding space? How do we even know that we can get an interesting or semantic embedding space? And so that was kind of this first proof of concept is that we can leverage this approach, instantiate in a specific way, and we get very semantic embedding spaces that are really just bootstrapped.

40:07And what we found kind of, and we just qualitatively probed towards the end of the paper, was this predictor world model that was making these predictions in embedding space. And we were surprised to see, as I mentioned earlier, that it did capture spatial and temporal uncertainty to some extent, because as you also discussed earlier, like it's predicting an embedding space. How, what, you know, what is this embedding space? What is the network actually learning? So the focus, even though it was a bit more on the encoder, I think to get there, you know, using this predictor and this encoder for now building goal-driven agents, doing reasoning and planning, I think that more focus needs to be placed on the predictor.

40:44So that corresponds to some of the things that we had mentioned earlier. So doing longer range prediction of multiple timescales. So the predictor now doesn't just capture three to 10 seconds of uncertainty or information. It can capture information on multiple levels of abstraction at various timescales. Also kind of identifying how do you condition it? So what kind of information do you give in, do you give the predictor to make predictions about the future, especially in a goal driven agent. So you can imagine maybe conditioning the predictor on your goal, or somehow propagating information from your goal, back all the way to the predictor so that it can make kind of intelligent predictions that lead you to your goal.

41:21So I think in short, there's a lot of really interesting research questions and ways to go. We do have kind of a proof of concept learning framework now, a semantic embedding space. And now the question is, how do you beef up this predictor, enhance prediction on multiple timescales? How do you condition on additional kind of information? And how do you leverage kind of and build goals into this learning framework? Can you speak a little bit to some of the challenges that you ran into in essentially getting from iJEPA to VJEPA? And I'm particularly interested in or particularly asking about things related to the challenges of dealing with video files being the computational intensity and the volume of them.

42:10One of the big challenges was going from image to video. I mean, even putting kind of algorithmic components aside, which, you know, there were algorithmic components to figure out. But essentially, you know, even look at the current state of image data sets versus video data sets. So part of kind of reproducibility and wanting to do open science, we wanted to be clear about which data sets we were using, right? Data sets that are being used and explored and well understood by the academic community. and if you look at the state of all those kinds of video data sets and you compare them to the current state of image data sets there's a huge difference both in terms of size and quality and so first thing was you know just finding kind of um the right data to train on good data to train on and um and the fact that it's so much smaller as well these data sets are so much smaller than current large-scale image data sets when you get to efficiency now luckily the masking provides a big boost there, as I mentioned, because we're leveraging transformer architectures and masking.

43:10If you mask 90 % of the video, you only have to process roughly 10 % of it. So that makes a big difference in terms of efficiency. And what we find is actually VJEPA can train much more efficiently than a lot of state-of-the-art image models because of the masking, but also because of the fact that our video data sets are shorter and we're using much shorter training schedules. So the fact that VJEPA is predicting in a kind of abstract or semantic representation space means we need much less training than, for instance, a pixel reconstruction or generative model. So that also helps kind of make video more accessible for training because learning is just much more efficient.

43:48But there are certain other issues as well, like efficient video data loading and whatnot. And so, you know, there are kind of popular open source academic libraries that people are using for doing machine learning with video data. So we leverage those essentially in our data loaders. So they're standardized. Again, academic communities are already familiar with them. So when they want to try and run the code or the models, they're already familiar with some of these basic video decoding libraries. I think though, you know, to make future progress, we really need to address this gap between video data sets and image data sets.

44:27In image data sets, we have large scale curated, well understood. And by well understood, I mean, like people have really inspected and probed these, these data sets to understand what kind of content is there. And, and I think we need kind of the equivalent for video data sets, where you have kind of transparency about what content is there, but we just need to improve the scale and the quality of this data as well. That's where I see a big gap in going from image to video. Efficiency, of course, just to reiterate, it is a factor. But we're already maybe a bit more efficient. Oh, I mean, we're already a bit more efficient than some of the state-of-the-art image models.

45:01So it's not like, it's not a huge jump in terms of compute compared to existing approaches people are already using for learning from images. You've addressed some of the what's next question in talking about, you know, focusing on the encoder, the embedding space, and the predictor. You know, there's this other aspect to what's next, like you went from image to video? Is there some other like domain? Is there an X JEPA that, you know, what's the X in the next X JEPA? Is that something that you're thinking about? Multimodality, definitely. But I think it's not multimodality for the end goal of being multimodal because there's a lot of what's trying to produce kind of joint embedding spaces of multimodal data.

45:50I think it's multimodal to the extent that it improves and builds up a more semantic and capable world model. Um, so if, you know, beefing up the predictor, conditioning it with other modalities, such as proprioception, audio, um, you know, uh, depth, maybe if you have, uh, if you kind of have a depth maps, um, so whatever kind of, um, information that also ideally you want it to be, uh, broadly available in a kind of unsupervised fashion, right? So you don't kill that aspect because one of the reasons the method is so scalable right now is it's fully unsupervised. It can just be applied to streams of video or images.

46:27So that's something you want to retain. The multimodal part, I think, is going to help build up a more semantic and capable predictor, capable of also much longer range prediction as well. But that's where I think that exploration should come in. It's not necessarily building just a better video encoder because it's learned from audio. It's trying to build up a more capable and powerful world model that understands more signals. Well, Amito, thanks so much. This is an awesome kind of explanation of what you're doing with VGEPA and really appreciate you taking the time to share it with us. Thanks so much for having me on the podcast, Sam.

47:06It was a real pleasure to be here.

From the publisher

Today we’re joined by Mido Assran, a research scientist at Meta’s Fundamental AI Research (FAIR). In this conversation, we discuss V-JEPA, a new model being billed as “the next step in Yann LeCun's vision” for true artificial reasoning. V-JEPA, the video version of Meta’s Joint Embedding Predictive Architecture, aims to bridge the gap between human and machine intelligence by training models to learn abstract concepts in a more efficient predictive manner than generative models. V-JEPA uses a novel self-supervised training approach that allows it to learn from unlabeled video data without being distracted by pixel-level detail. Mido walks us through the process of developing the architecture and explains why it has the potential to revolutionize AI.

The complete show notes for this episode can be found at twimlai.com/go/677.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
V-JEPA, AI Reasoning from a Non-Generative Architecture with Mido Assran - #677The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 48 min
Listen in VO