In short
Podcast Episode Summary: Decoding Animal Behavior to Train Robots with EgoPet with Amir Bar - #692
Overview This episode of *The TWIML AI Podcast* features Amir Bar, a PhD candidate from Tel Aviv University and UC Berkeley, who discusses his research on visual-based learning and the creation of the EgoPet dataset. The conversation dives into the challenges of model training in robotics, the limitations of current datasets, and how understanding animal behavior can enhance robotic learning and performance.
Key Topics Discussed
Amir Bar's Background
- Academic Journey: Started as a history major before shifting to computer science and eventually deep learning, spurred by the rise of technologies like AlexNet.
- Professional Experience: Worked at Zebra Medical automating medical imaging and later transitioned to academia for open-ended research.
Research Goals
- Focus on Visual Learning: Amir's research emphasizes learning from large unlabeled datasets, particularly in vision-focused tasks, avoiding the reliance on language-based supervision.
- EgoPet Dataset: A significant contribution of Amir's work, EgoPet consists of egocentric motion videos from animals, aimed at training models for robotic planning and proprioception.
Challenges in Robotics
- Learning Problem: While hardware capabilities have advanced, the primary issue lies in how robots learn and apply policies.
- Limitations of Caption-Based Datasets: Current models often rely on captions that fail to capture the nuances present in visual data, leading to poor performance in complex tasks.
EgoPet Dataset and Its Implications
- Data Collection: The dataset comprises around 80 hours of video footage featuring various animals (dogs, cats, and some exotic species) sourced from platforms like TikTok and YouTube.
- Significance: EgoPet is designed to facilitate better learning of motion planning and sensory processing in robots by mimicking animal behavior.
Key Findings
- Performance on Downstream Tasks: Models trained on EgoPet outperformed those trained on existing datasets (e.g., Kinetics, IN1K) for specific tasks related to animal interactions and locomotion.
- Vision to Proprioception Prediction (VPP): A task aimed at predicting the physical properties of terrain based on video input, showcasing the practical applications of EgoPet.
Future Directions
- Application to Robotics: Aiming for the development of robotic policies that mimic animal behaviors for enhanced navigation and interaction with the environment.
- Integration of Learning Approaches: Proposing a combination of simulation and learning from video data to overcome challenges in robotic learning.
Key Takeaways
- Importance of Visual Learning: Emphasizes the value of understanding visual data over language inputs in machine learning, particularly in robotics.
- EgoPet Dataset as a Benchmark: Provides a new avenue for training robots to understand motion and interaction, potentially leading to more advanced robotic capabilities.
- Broader Implications for AI: Insights from animal behavior can significantly influence the development of AI systems, bridging the gap between biological intelligence and artificial systems.
Conclusion The episode underscores the potential of leveraging insights from animal behavior to advance machine learning and robotics, particularly through innovative datasets like EgoPet. Amir Bar's research represents a promising step towards developing robots with greater autonomy and adaptability in real-world environments.
---
For complete show notes, visit
[The TWIML AI Podcast](https://twimlai.com/go/692).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00One of the problems we have with robotics right now is that we got the hardware right. The problem is mostly a learning problem. how they actually learn these policies and apply them to these robots. If you can watch these videos and learn how animals, dogs, and cats plan and be able to transfer this to, say, a quadruped robot dog, I think it would be very exciting.
0:31All right, everyone. Welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. And today I'm joined by Amir Barr. Amir is a PhD candidate at Tel Aviv University and UC Berkeley. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Amir, welcome to the podcast. Hi, Sam. Thank you so much for having me. I'm looking forward to digging into our conversation. We'll be talking about your research into visual prompting for large visual models. Before we dive in, I'd love to have you share a little bit about your background and how you came to study in the field.
1:09Yeah, and it's actually a funny story. When I started my undergrad, I actually wanted to become a historian. So I did this double major in Middle Eastern history and in computer science. And I thought to myself, you know, I'm going to become a historian, but if that doesn't work, at least I'm going to have an alternative and maybe at least work as a programmer or something. But after the first year of my undergrad, I figured out, well, I love history. History is nice. I love hearing stories and memorize dates and events. But then I didn't see myself sitting in an archive and doing the actual research work, actually becoming a historian.
1:46And so I felt like something is missing. And there was this small hype about deep learning started to grow after the AlexNet moment. And I felt like this is the right moment for me to go back to school and start a master's. And that's what I did. I started a master's program in Tel Aviv. And I was recruited by a startup called Zebra Medical that hired me back then. And my goal was to automate the reading of x-ray scans, CT scans. And I started there as a research scientist. And after a while, I moved to the US and I started a new AI team that I led here in the Bay Area in Berkeley. And after a while, the company was acquired.
2:24And I thought to myself, OK, what's next for me? and I've done some air research in the industry and maybe it's time to do some more open-ended research and start a PhD. So you talk a little bit about this overarching interest that guides your research as being focused on visual prompting for large visual models. Kind of break that down. What is the goal of your research and what are some of the projects that have helped you get there? And of course, we'll ultimately dig into one of your most recent projects, which is EgoPet, which is a data set and a set of tasks for egocentric motion from an animal's perspective.
3:10But, you know, tell us about kind of how you got there through your research. When I started my PhD, I was mostly interested in how can we learn from all these large amounts of unlabeled data, unlabeled images and videos that we have. There's some other modalities as well, like audio and language. I've been trying to focus my research on vision first. I know right now vision and language is very trendy. And a lot of people, for good reasons... You think? For obvious reasons, people have been building on the research on vision and language. And I think that's why the approach. I've been trying to focus on vision only without using language.
3:54And the reason I've been trying to focus mostly on vision, and I think there's like maybe two different arguments for kind of like going vision first before introducing language. And one argument is more maybe philosophical. And it's about the fact that if we look at human evolution, we know that language is something that was developed relatively late in evolution. whereas other visual and sensory motor capabilities evolved much earlier. In fact, if you compare a human to other species, for example, my dog, right? My dog has very basic language. But if I don't look for one second and I leave some food on the table, he would jump on the sofa, jump on the table, grab the food, just in the right second that I'm just looking away.
4:48All of this was just basic language. So I think there's a lot to explore in terms of learning good visual representations before we introduce language. And we know that there's one model that actually works, which is the human brain that has done this process. First, you develop these visual capabilities, visual sensory motor capabilities, and then you introduce language. Now, there's another maybe more practical argument for why go vision first. So let's say that you want to learn from vision language. You want to use language to supervise your vision systems. So most approaches that we have so far, the way they work is as follows.
5:30You collect the data set of paired images and associated captions. Let's say you have an image and let's say the caption of this image is a dog and a cat sitting on a sofa. Now, this caption is very short and precise. There's a dog and a cat sitting on a sofa, but it's also missing a lot of information that we have in the image. What is the relation between the dog and the cat? Are they enjoying themselves or are they grumpy? What is the material of the sofa? Is this a leather sofa? What type of the floor is? The wooden floor? Who's sitting next to who? Is the cat sitting left to the dog? Or vice versa.
6:02The problem with this is that this text actually misses a lot of information. And once you start asking the model to task, let's say, count the amount of objects in the image or specify the relations between the objects in the image, the caption might not be a good enough signal to supervise the training of these models now there are some companies, I don't want to mention a specific company, but there are companies like a company start with O and ends with AI that what they will do, they will try to improve the quality of the captions, right, so they will say okay, so we're missing, for example the number of objects in the image, okay, so let's try to introduce it to the caption, let's maybe ask our annotators to specifically caption these images with higher quality captions and then kind of like recaption our data with this model or introduce an object detector that will first detect the object in the image and then use some GPT-4 capabilities to recaption this image and you can do this and you can kind of like reverse engineer and try to say okay I'm missing this part so they try to reintroduce the caption, return our systems and kind of like...
7:06If you've got an open ended question domain it is very hard to sufficiently capture enough detail out of these images that you can handle any question. You know, what's the pattern on the pillow in the frame of the video? Exactly. For example, you can't anticipate every single question. Exactly. And I mean, recently I tried a very simple experiment with a friend in UC Berkeley, and we took images of dogs from a dataset called Stanford Dogs that have every image of a dog with an associated species. So every image is a specific dog, like let's say Gemma Shepherd or Jack Russell Terrier, etc. And then what we asked GPT-4 to do is to caption these images.
7:50When I say caption, we asked specifically GPT-4 to specify and numerate all the visual features of that particular dog without actually mentioning this exact species or exact breed of that particular dog. And then we asked GPT-4 to classify the image. and classify the text, try to classify the exact breed from either the text or the image. And unsurprisingly, we found out that you can actually get very good performance when you look at the image. But if you look at this caption, just in the visual features, it's actually very hard. And it doesn't work that well compared to the visual input. And it's not surprising, right?
8:29If you ask, for example, let's say, you know, me and my brother, we have very similar facial features, same hair, same smile even, right? But if I ask you, Sam, can you try to nail the differences in appearance between me and my brother? I think it'll actually be very hard for you to nail down. I mean, this is the fundamental reason why deep learning has become so powerful is because, you know, for dog and cat recognition in photos, like how do you describe a cat or a dog? You know, there are things that are captured well in images that are hard to reduce down to language. Exactly. And you can try to go about a top-down approach with introducing language to supervise it.
9:13But then you need to start reverse engineering every possible use case and enrich your captions. And sure, you can make some progress this way, and maybe you can solve 90 % of the problems. A different way, which I think is also interesting to explore, especially now in academia when everybody is putting their bets on, And the first approach is try to do it bottom up. So try to build visual representations that are useful. And a lot of my research is about trying to build these representations. How'd you get started with that? One of our first projects in my PhD was about trying to understand what is an object in a self-supervised way.
9:51Or trying to train neural networks to basically detect objects in images in a self-supervised way. And this was kind of like a very specialized project. It was just about, can we try to learn this concept of objectness from unlabeled data in an unsupervised manner? And maybe if we can do it, let's say we got good representation of what an object is, maybe we can actually improve some performance on downstream tasks like object detection, for example. Yeah, now we've had decent object detection for quite a while. Has that all been supervised? So all the object detection works that we have are basically supervised.
10:25you have data set that's containing a list of bounding boxes. For every bounding box, that covers an object, is mapped to an object, you have an associated object category as well. Usually the way it happens is that you train your detector on this supervised data set, on this labeled data set. The larger data set you have, you can maybe achieve better performance. And what we tried to do back then, we tried to do some pre-training for object detection. and the idea is how can you how can you how can you do it and back then what we did is we relied on this notion of regions in image so we tried to kind of like induce this prior of what is a region in image so you can think about it like super pixels in image like regions that have similar colors in our shape and we had this vast literature of region proposal methods back then that But the idea was basically, you know, you get an image, let's try to extract a large list of region proposal or potential regions in this image.
11:28And they're like very classic algorithms like selective search, MCG and others, Edgebox and others that based on these ideas of hierarchical clustering. So you have some regions in image now. Now, okay, if you have two regions that are neighboring, if they share some characteristics, let's say color or shape, now you can decide, okay, let's merge them into a larger object. And then you can kind of like iteratively go and obtain a list of proposals. And what we did back then is just try to train neural network to fit on this kind of algorithms. And the nice thing about neural networks is that they're very tolerant to noise.
12:05So some of the proposals are very noisy. Some of them have a good signal. but if you try to fit a neural network on this kind of set of proposals it's going to actually be able to perform slightly better than these actual algorithms and that's what we did back then now what we showed is that if you pre-train a detector this way and then try to adapt it to let's say some object detection data set this region prior actually helps you and you can perform better on these downstream detection tasks but that was I think the nice thing about this work was that okay we tried to build up this prior of what an object is on a pre-training stage, but this was a very specific approach, right?
12:45It only adapts to object detection, but what if you want to do something else? What if you care about, let's say, colorization? And that kind of led us to trying to think, is there any way to train one large model that will be able to train on a large amount of data and readily adapt to any downstream tasks that can be specified? And if yes, how do you build such a model? And one of the ideas that we had for this was this concept of analogies, right? So, you know, this concept of analogies, right? Like you have some pattern like A to A prime and you want to apply the same pattern to B such that you get B prime, right?
13:26And our idea was if we have a model that can kind of like reason through analogies, I give it one example of let's say an image and associated segmentation mask. Now I provide a new image. If it can infer the rule mapping the image to segmentation mask, and I can give it a new image and it can apply the same rule over the new image, this model can be a very general model, right? I can fit it with like an input, you know, RGB image segmentation mask. I can fit it with a grayscale image and, you know, a colorized image. I can fit it with an image of some scene in the spring and then a scene in the winter.
14:02And then I want to apply the same transformation on the new image, right? So as long as my model is flexible enough to reason through analogies, maybe I can build a very general vision model that can perform many computer vision tasks. So that was the idea we had back when, I think, 2022. And the way we formed it was just as an in-painting task. And Alyosha Effors, who's one of the professors that we worked with, liked to say they tried the most stupidest thing you can try, which is just do it as an in-painting task. You just take these images, like an input image and an output image. This is the example.
14:33and then a new image, and you stitch them into a single image grid. And now you ask an in-painting algorithm to just fill in one missing part, which is the answer, the output that should be the answer for this query that you have there. And so what are the two images there? The input image is, for example, a black and white image. And what's the other input? Right. So let's say you have one input and output example. So the example could be an RGB image. and its associated segmentation mask. And now you have a new query image, which is a new RGB image. It can be totally different than the example that you've seen earlier.
15:16But what you want is this algorithm or neural network to reason across this example, deduce that there's a rule. And come up with a segmentation mask for the second image. Exactly, and just apply it on the second image, right? And what we did, we just said, okay, let's just stitch these three images in a grid, in a 2x2 grid in a single image and have one missing part, which is the answer to this analogy, which is this new segmentation mask that has to be consistent with the other input image. And let's just try to apply a painting algorithm to solve this problem. And apparently this could work.
15:50I mean, neural networks can reason across these analogies with vision-only examples and solve this problem. puzzles or analogies. But the question is, how do you train them? And these kind of analogies do not appear in natural images like ImageNet data, for example. Images of cats and dogs are not in this kind of grid structure. So it's going to be very hard to utilize this kind of data. But if you think about it, internet data does have similar analogies. Let's say there's a lot of these images of surgeries, like that's a retinography. You have like before and after. You have a grid image and there's some kind of correlation, something happened before, something happens after.
16:34And another source of data that has this kind of analogy is computer vision figures. Because computer vision figures, let's say the figures of the results, they would have like Model A baseline, Model B baseline, hours, and then ground truth or something like this. So if you look at these kind of images, the nice thing about them is that there's like this grid structure and there's correlations between different grid cells in these images. So what we did is we scraped the archive, scraped the computer vision papers and downloaded all the possible figures we could put our hands on back then. And then we trained an algorithm, a mouse autoencoder, by just taking random patches from these images, hiding 75 % of the patches and just trying to reconstruct them from the context of the visible patches in the image.
17:23Were you able to automate identifying these before-after grid types of pictures in the scrape dataset? Or did you have to do that manually? That in and of itself sounds like a hard problem. oh so i think it's a great point so we didn't do any parsing at all over these figures we just downed them as is we then tried to uh you know parse them into small sub figures or anything like this for training we just took random crops out of these images were you downloading pdfs or were you using like html or latex representations or something of these images so there's like the the LaTeX representations for some of the papers, and you can just download the images in the LaTeX.
18:07Got it. And did you apply some classifier to determine whether it was one of these grid side-by-side things, or did you manually... So we applied a classifier just to remove some of the noisy images, like, for example, images of graphs or charts. So we applied a very simple classifier to clean this out. But other than this, and I think it would have probably worked with these images as well, but it would slow the training. And so what we did, we just removed some of these images that did not provide very good signal. But we had like images of architecture, you know, just like deep learning architectures and images of, you know, grid structures.
18:48Some images, it's just like... So you had a very noisy data set, but it was fine, it sounds like. Very noisy data set. and what we did is we just took random crops out of this data set because parsing is just not really something that is very easy to do with this kind of data set and we didn't want to do something like this. I mean, the entire point is that this algorithm is trained from data and can hopefully be improved with more data. So we grabbed these images and now we just take random crops out of them and just hide 75 % of the patches in the image and then try to reconstruct these missing parts.
19:22And what we get is this in-painting algorithm that can reason across different images over grids. And now what we do, we try to apply it to solve analogies. Like we described earlier, we can, in test time, just construct these grid images of, let's say, input image, output segmentation mask, and another input image, and ask the model to complete the missing part. And surprisingly, it worked quite well. So it wasn't state of the art or anything like this, but it could adapt to different tasks like style transfer, colorization, segmentation, in painting, some more stuff that we tried. This was our first work on visual prompting.
19:59And we're very excited about this. I want to make sure we have time to dig into EgoPet. So let's maybe jump over to that. Talk about the broad context for that paper and relate it back to this research thread of yours. Why did you go after this problem of ego motion? I've been thinking a lot about the state of self-supervised learning and what do we do in self-supervised learning? Where do we want to go? And if you think about what we have in self-supervised learning, So a lot of the things we have is learning representations and checking how do they transfer over ImageNet. How good representation we learn for classifying the set of categories we have or the set of classes we have on ImageNet.
20:51What we don't have in self-supervised learning as much is this capability to plan. So how can you learn to do planning from video? And if you end up playing from video, I mean, one way and maybe something that some labs have been focused on is to learn from egocentric data. So you have egocentric data, you have, let's say, a person with a camera mounted on their back, and they're very focused on usually on these human skills, like, for example, cooking or changing a wheel or doing all sorts of manipulation tasks, usually involving hands and objects and hands and object interactions. What I was thinking is that there's something that is missing.
21:39is just basic planning that does not involve any human hand manipulation skills and object manipulation skills just plain planning of locomotion of getting from point a to point b getting from point a to point b and without any language in the loop just you know plain action and different agents like animals are seemed like a very good candidate or very good action and i think another thing that really inspired us is that I was interning at Jan LeCun's lab a couple of years ago at Meta AI and we talked a lot about the cat level intelligence and the dog level intelligence and how cats and dogs can do pretty crazy and amazing things right you know like following objects or I mean some of the videos we have in Eagle Pet for example are pretty crazy we have videos of a cat catching a rat for example or we have videos of a dog running following a human or listening to the sound of a human following its owner.
22:41So I think these are really good examples for how animals plan. And it doesn't seem like we have these capabilities yet in AI. And is the idea specifically that you can apply some type of model that you developed from this data set to something like a spot type of a robot, a quadruped robot? or is the thinking broader than that, that you can learn kind of proprioception primitives or something that are broadly applicable to lots of different scenarios? Right, I think this would be the holy grail. If you can watch these videos and learn how animals, dogs, and cats plan and be able to transfer this to, say, a quadruped robot dog.
23:30and one of the problems we have with robotics right now and the way I see it is that it seems like we got the hardware right in the most part so we already have some papers showing, there was recently this paper called Aloha from Stanford showing how they can do this night teller operation of controlling I think a robot with two arms doing different manipulation tasks And via teleoperation, it seems like they can almost do any task a human can do, right? So this basically tells you that the problem is mostly a learning problem. How do you actually learn these policies and apply them to these robots?
24:15Because the hardware seems like it's almost there already, or at least performed pretty well already. So my question is, how do you learn? And how do you do this learning problem? And the problem about the learning problem in robotics is that you have this chicken and egg problem. In order to collect very useful data, you need to have robots that are very capable and that are already kind of like exploring the real world. And then you can learn better policies maybe. But in order to have these robots deployed, you already need to have these policies in the first place. So the question is, how are we going to go about it?
24:52One way that I think we're thinking about as well is maybe we can learn this from video. If we can try to transfer, like you're saying, this representation that you learned, for example, for locomotion from, you know, these robot videos to these agents, I think, or to these robots, I think it will be very exciting. Now, there is a problem, right? So the problem is, from these videos, how are you going to get something like the amount of torque you need to apply to the motors of the robot in order to induce the right motion, for example? So that is something that is going to be hard to acquire from these videos or the proprioception signal that you get from the joints, from the relative joints of the animal, right?
25:36Do you think that that also needs to be learned or are you kicking that can down to the control theory folks and letting them try to figure it out? I think it's an open question. I don't know the answer. I mean, a wide way to go about it is to say, let's do the low-level part on simulation. So learn this mapping of, let's say I give you this kind of vector of velocities, go in this direction, in this speed, put this rotation, and you train a system to that simulation that gets this vector and knows how much torque you need to apply to move in this direction for that time. And leave this other kind of like high planning or high level planning part to this learning from video.
26:21So you can do the low level part, you can do it in simulation, but then maybe the high level part. you can do from video. So I guess this is one approach you can go about it. And we're maybe, I'm jumping ahead a little bit here, but if I remember correctly, your VPP tasks, so vision to proprioception prediction task, you were as part of that task looking for some fairly low level predictions like surface friction and things like it. And I don't remember if Torx specifically was one, but I thought there was something about motor dynamics. Right. So what we did is that we said, okay, let's actually remember to come back to this and just spend a little bit more time on the data set.
27:05But the broad idea where we're leaving it off here is that you had this idea to collect this egocentric, i.e. first person, if you will, perspective data from animals. And then that would allow you to better train models for locomotion of quadruped types of robots. And so there were a couple of prior works that were mentioned in the paper and data sets, IN1K and K400. Can you just give a high level overview of those? Like you said, what we did is we collected this EgoPed data set. It's around 80 hours of egocentric data of cats, dogs, which comprise the majority part of the data set. There's some other exotic animals.
27:55Sea turtles? Sea turtles, eagles that we found on the internet. And this was mostly all collected from TikTok videos? TikTok videos and YouTube videos. Yeah. So we collected this data and the idea was to see what can we do with this. let's say that we have this data of pets how can we use it? So I mean my idea is try to do some pre-training on this data set to learn good video features and that's one thing that we tried we tried to see okay let's take some off the shelf self-supervised for video algorithm some masking and reconstruction algorithm called MVD and that's what we did we took this algorithm we pre-trained on EgoPet and now let's try to evaluate representations that we learned on some downstream tasks.
Read the full transcript
28:42And then we kind of proposed three downstream tasks. One downstream task was mostly about perception. So we called it the visual interaction prediction. And the idea is in this task, you get a short video, and I need to say, does the agent or the animal perform any interaction with an object or another agent in this video? and if yes, what is the category of this object? For example, if the cat is interacting with another cat, then the category will be a cat. If the cat interacts with a dog, the object interaction will be a dog. So how did that VIP task differ from object detection or object recognition?
29:31Did you, was the task trying to identify a degree of interaction, meaning it wasn't sufficient that there was another cat in the frame from the perspective of this one cat, that the two had to actually be interacting in some way? Exactly. So only if we actually observe an active interaction of the agent with that particular object, then we consider it an interaction. So interaction can be, for example, there's some kind of, I don't know, the cat touches the other cat, or the cat does some meow, or you see that the cat visually attends that specific object. So that's what we consider. And this is supervised, so it ultimately left to the judgment of some human rater.
30:21Yes. So we have a couple of human annotators who looked at these videos and determined there is an interaction, there's no interaction. And what we did is that we evaluated this representation that we learned on this downstream task. So you can take the features and now you can apply only a linear layer classifier on top of it and try to see how well can you classify this interaction. And I think that the interesting thing that we saw in this task is that compared to other data sets like kinetics is, for example, a very common video data set or Ego4D, which is the largest egocentric data set I think we have right now, which is mostly around human activities.
31:03So what we saw that the EgoPet features perform better than these other video datasets. That brings up a question that I had broadly about the paper. And in particular, one of the key claims of the paper is that models trained on EgoPet outperform those trains on other datasets. But is that claim in reference to outperformance on the tasks that you defined against your dataset or were you able to demonstrate that models trained on EgoPet outperformed other models on other datasets? Right. So only on the tasks that we proposed on the EgoPet paper, basically, which are mostly around interaction in this particular setting of, you know, an agent exploring either indoor environments or other environments or in the setting of that or in the other tasks that I talked about that we haven't talked about yet.
32:03One of them is this visual pro-perception, vision pro-perception task and the other one is locomotion task. So on these tasks we found that pre-training when you go prep is beneficial compared to other video data sets. And is that how strong do you think that claim is given that the tasks are kind of defined in the context of the data set that outperformed? It kind of makes sense, right? So the tasks that we have in the Ego paper are very related to this kind of data that we have. One task is this visual interaction prediction. These are essentially interactions of dogs and cats. So if you train on interaction of dogs and cats, you're going to be obviously better on this kind of interaction.
32:46And the other task that we have, which is separate from this data set, is this vision to prop perception task. So in that particular task, what we did is that we took this quadruped Unity RoboDog. Okay, so this RoboDog has a camera mounted on its top part, which is facing the front, taking this egocentric view. And we collected the data set of this video and a paired proprioception. So for every frame, we basically have the proprioception signal that we have, which is basically you can think about as the location of the joints. And what we try to do is given a video of this forward-facing camera, we can try to predict now the future proper perception, which is the future relative location of the joints, or the past proper perception, or the current proper perception.
33:31Now, what does it mean to predict proper perception? So proper perception has been shown to be able to be a really good predictor of the geometry of the terrain. So if you know the proprioception, you can estimate stuff like the friction and different physical properties of the terrain. So if you can predict the proprioception precisely, then it means that the representation that you have captures very well the geometry of the scene that you act in. And to restate, specifically by proprioception, we're saying if you can predict, if your model can predict joint angles over time, then it can ultimately predict both terrain features as well as control dynamics required to navigate.
34:20Exactly. And what we showed in this paper is that if you pre-generate an EgoPet and you transfer the representation to this visual to perception task, then you get better performance compared to using other presentations. if you train on data sets like ego 4d or kinetics and again i think i think it's very intuitive right because when you look at videos of people in egocentric view a lot of the times they're not actually doing locomotion they actually just you know hang out in a very indoor environment the kitchen performing some very specific human skill like cooking or something like this you've reference throughout talking about EgoPet as like pre-training, is the idea that you're pre-training a VIT model on the EgoPet data?
35:11Right, right. The way to think about it, you just do some self-supervised tasks like master encoding on this particular dataset EgoPet, and now you learn a good feature extractor from videos, and now you try to apply these learned features into some downstream tasks. So you train an additional classifier on top of these features. And now you can use the features pre-trained on EgoPath or you can use features from a different pre-trained model like a model that was trained on kinetics for example or on Ego4D. And those other examples of those other models or pre-training processes is that also MAE and MVP and Dyno, iBot that's where those come in as well?
35:53Right so some of the models like single image models are like iBot, MVP they receive one image and they output some representation some of the other models are pre-trained on video data sets like Kinetics or Eco4D and yeah, what we found out by the way was that the models that were trained on ImageNet pre-trained on ImageNet talking for example about approaches like iBot or Dyno their advantage is very good in being able to classify the category in the interaction prediction task. So if you just try to predict that the agent performs or does not perform an interaction, then EaglePet is very good.
36:40But if you compare and if you try to classify the object in the interaction, then the models that were trained on ImageNet are actually very good. And the reason is that if you train on ImageNet, you're exposed to such... You know, cats and balls and other types of objects. of objects. So that kind of makes sense. And so the IN1K, that's ImageNet, and the K400 is what? It's like the kinetic data. So it's an action and cognition data set of different human activities. It's a very general data set that is very popular for pre-training. It's not an egocentric data, but there's some camera motion, but it's not egocentric videos like we have in EgoPet.
37:25So whereas the performance of EgoPet on the first couple of tasks, the visual interaction and the locomotion prediction is less strong because they're tasked to find on that data set, the VPP is stronger in the sense that you are. It's out of domain. You're making predictions. It's out of domain. You're making predictions against the performance of this robot. Exactly. And so where do you see kind of this, you know, your research thread continuing? I mean, exactly what you said in the beginning of the conversation about this. I think the killer application would be to be able to take this data set and directly learn a policy, a cat or dog policy from these videos and deploy it over a robot, a quadriple robot.
38:14I think that will be the killer application. And I think that I'll be very happy if we can get something like this to work. I've been trying very hard since the release of the paper to do it. When you think about a cat or a dog policy, what all does that mean and define for you? Is it that a quadruped robot has the grace that we associate with cats or some characteristics of a dog? Is it specifically with regards to locomotion or is it more nuanced in the way that a cat or a dog might respond to environmental stimuli? How broad are you trying to go? I mean, the very basic thing that you want is just to have capabilities of, say, visual navigation, right?
39:10So you have a robot that explores some environment in a social setting, maybe around humans and maybe around some obstacles. And you want the robot to be able to navigate through these obstacles without hitting a glass wall or something, right? So I think that's the most basic thing that I think we'll be happy to achieve at the first step. Of course, I mean, other people can think about some other crazy scenarios, like trying to actually learn some very dog-style behavior, like following your owner or reacting to treat or you know there's clear constraints on the the the physical aspect of this as well like i'm envisioning one of these quadruped robots like coming up to a wall or a tree and looking up and like you know having the confidence of a cat to you know right right six exits body height and not having the physical capability due to its uh right its implementation right i mean it's possible to try some of this stuff with a drone, right?
40:17You can learn this policy, deploy it on a drone instead, and see if the drone can follow maybe somehow the CAD policy. But yeah, I mean, for some things, it seems like the hardware is already there. For some other things, you know, it might not be there. I agree. Yeah, the broader idea is an interesting one. Like we have over the years, I think, have seen these quadruped robots get more and more capable and move more and more smoothly. But, well, I guess a couple of things. One, it's not always been clear how much of that is autonomous versus pre-planned motion trajectory. But certainly there's something to be learned from the ease with which these animals move through the world right i mean i think you know if you look at some basic actions they can explore the environment a bit they can may perform some particular basic actions like maybe a bit of navigations but how do they perform in say a social interaction like let's say that you that one robo dog faces another robo dog do they are they gonna do some is there gonna be some you know interesting interaction between them no right it's so i think it's still very i think what they're capable of right now is still very basic.
41:40It's still kind of like showing, okay, we can do this basic control. We can... They could move. They can maybe explore the environment. They can do walking. But I don't know if they can do much more complex things than that. If you ask a robodog to deliver a package, I don't know if we're still there. You can actually rely on them to deliver a package and stop at a red light. cross the street on a green light play fetch play fetch yeah so i think that the planning part of being able to navigate the environment in a social setting around humans and being able to interact with humans is still definitely not there awesome awesome well amir thanks for joining us to share a little bit about your research and what you've been working on thank you so much for having me thank you my pleasure
From the publisher
Today, we're joined by Amir Bar, a PhD candidate at Tel Aviv University and UC Berkeley to discuss his research on visual-based learning, including his recent paper, “EgoPet: Egomotion and Interaction Data from an Animal’s Perspective.” Amir shares his research projects focused on self-supervised object detection and analogy reasoning for general computer vision tasks. We also discuss the current limitations of caption-based datasets in model training, the ‘learning problem’ in robotics, and the gap between the capabilities of animals and AI systems. Amir introduces ‘EgoPet,’ a dataset and benchmark tasks which allow motion and interaction data from an animal's perspective to be incorporated into machine learning models for robotic planning and proprioception. We explore the dataset collection process, comparisons with existing datasets and benchmark tasks, the findings on the model performance trained on EgoPet, and the potential of directly training robot policies that mimic animal behavior.
The complete show notes for this episode can be found at https://twimlai.com/go/692.




