In short
Podcast Episode Notes: The TWIML AI Podcast - Episode #722
Episode Overview
- Title: Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
- Guest: Chengzu Li, PhD student at the University of Cambridge
- Host: Sam Charrington
- Focus: Discussing Chengzu's paper on Multimodal Visualization-of-Thought (MVoT).
Key Concepts Multimodal Visualization-of-Thought (MVoT)
- Definition: An approach combining visual and verbal reasoning to enhance the performance of AI models in spatial reasoning tasks.
- Motivation: Inspired by cognitive science principles like dual coding theory, which posits that humans process information better when using both verbal and visual channels.
Core Components
- Dynamic Spatial Reasoning: The ability to understand and represent movement and navigation within a spatial context.
- Token Discrepancy Loss: A technique to align language and visual embeddings to improve the accuracy of generated visual representations.
Background of Guest
- Chengzu Li's Research Interests:
- Focus on multimodal reasoning and spatial reasoning.
- Previous projects included evaluations of models' capabilities on spatial layouts using TopViewRS.
Framework and Tasks in MVoT Task Environments
- Maze: The model predicts the final location after executing a series of movements in a maze.
- Mini-Behavior: Simulates a robot executing tasks, such as moving objects, with a more complex action space.
- Frozen Lake: A classic game scenario where the model must determine if a series of actions are safe to execute on an ice surface with hidden holes.
Input/Output Format
- Input: An image and a corresponding textual question.
- Output: Both an image (visualization of thought) and verbal reasoning.
Methodology
- Model Architecture:
- Utilizes a visual language model (VLM) capable of generating multimodal output.
- Employs recursive reasoning where the model generates outputs iteratively based on previous steps.
Training and Data Collection
- Data collected for the defined tasks, emphasizing the need for high-quality visualizations to ensure accurate reasoning.
- Used LoRA for fine-tuning the model.
Insights and Challenges
- Ambiguity in Textual Information: Complex visuals can lead to misleading outputs if not properly aligned with verbal descriptions.
- Importance of Visualization Quality: High-quality visualizations are crucial for ensuring the model's reasoning process is accurate and effective.
Applications and Future Directions
- Potential applications in robotics (e.g., navigation tasks) and architectural design (e.g., spatial layout optimization).
- Exploration of alternative approaches, such as reasoning in latent spaces, to overcome limitations in context length and computational overhead.
Conclusion
- Chengzu emphasizes the exciting potential for MVoT-like models to improve reasoning and decision-making in various real-world applications, highlighting the need for future exploration and refinement of these methodologies.
Additional Resources
- Complete show notes available at: [TWIML AI Podcast Episode #722](https://twimlai.com/go/722)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Let's say if you have a navigation robot and you ask it like, I want to grab a drink from the refrigerator. So it has to think internally like, where is the refrigerator? Oh, it's in the kitchen. And then it has to find a way to the kitchen, right? If I go forward, can I see the door of the kitchen? And if I can, then this is the right direction. If I can't, then I will try different directions. So, and then thinking this in my mind is a similar process of MVOT.
0:45All right, everyone, welcome to another episode of the Twinkle AI Podcast. I am your host, Sam Charrington. Today, I'm joined by Chung-Zoo Lee. Chung-Zoo is a PhD student at the University of Cambridge. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Chung-Zoo, welcome to the podcast. Hi, Sam. Thanks for inviting me. Looking forward to our conversation. We're going to be digging into your recent paper, Imagine While Reasoning in Space, Multimodal Visualization of Thought. To get us started, I'd love to have you share a little bit about your background.
1:24Yeah, thank you. Thank you for a brief introduction. So I'm Chung Zhu, and I'm currently a second-year PhD student focusing on multimodality reasoning, especially with spatial reasoning. So I'm very fortunate to be supervised by Ivan Voodis and Serge Boulangy at Language Talker Knowledge Lab, University of Cambridge, with a group of fantastic lab mates focusing on different areas. And yeah, I'm very looking forward to discussing my recent work, MVOT, with you for this discussion. Yeah, before we dive into MVOT, talk a little bit broadly about your research interests and the kind of things you've been exploring.
2:06I would say my research basic is driven by the curiosity. So I'm very interested in exploring things that current models cannot do or is not good at, for example, spatial reasoning, and try to improve the performance in this direction. As I've previously mentioned, my research interest basically is multimodality and spatial reasoning. As you can see that, Although I'm in a language technology lab, but my work at VOT is actually emphasizing the potential in visual information during reasoning. So I think, yeah, a big shout out. Thank you for my supervisor for being so supportive during my research.
2:49When I first get in touch with multimodality, it's not really about image. It's about structural knowledge like databases and tables. And it's during my visit at the University of Hong Kong, working with Lin Peng and Tao, who are the professors at HKU. And then I got my master's at Cambridge, working with Simone Teufel. And this is actually my very start, digging into the area of spatial reasoning. So my MPV project topic is to generate the instructions of navigation given a top view map. And yeah, and that's the very start. And later, it seems like current models really are messing around with spatial descriptions when trying to describe navigation path on the map.
3:43And then that brings out my next project with a top view RS. And this evaluates the model's capabilities or performance on reasoning over relative spatial relations between objects in the top view map. And it seems that they're really not good at it, you know, although it seems very straightforward to human. And yeah, and that brings out this work and VOT, which we hope that could change a little bit or improve a little bit of the model's performance in this. So let's talk a little bit about the broad motivation for the paper. You've already alluded to a couple of the key ideas. One is the idea of spatial reasoning.
4:25The other is an idea of multimodality. this is all happening in the context of an explosion in ideas and research around reasoning models and techniques like chain of thought, which is alluded to in the title of the paper. You know, talk about how all these ideas fit together and, you know, where the motivation for this particular work comes from. So I just mentioned my previous work, Top URS, and this MVOT paper actually arises from one of the tasks in top view RIS. So the top view RIS contains one task, which is dynamic spatial reasoning. So imagine there's a navigation path drew on the top view map, and you have to answer certain questions related to the path.
5:15For example, if you look at the path, at which object did the agent take the very last right turn? And so it requires the model to understand the whole path in order to answer this question. And so that's the very first idea of why we're choosing this kind of task in MVOT and this actually origins from my top URS paper. And when it comes to the idea of the methodology and for the MVOT, as you can see from the name, it actually originates or is inspired by the work VOT, which is a work conducted by Wenshan, who is my mentor at Microsoft Research and also a co-host author of MVOT. So basically, VOT tries to have a visualization of thought, but this visualization is described in an ASCII art format, which is basically like a combination of the symbols in text.
6:18And they try to mimic the visualization of thought at each reasoning steps and improve the performance. Was VOT also applied to spatial reasoning and problems? Yes, it is applied to spatial reasoning problems, but within the context of pure textual verbal reasoning. And when me and my mentor actually discussed with each other, we think that it is a very exciting area because both of us are really interested in the cognitive intuition behind our work. So, and I guess this has something to do with some high level intuition about the MVOT paper in this project. It's a little bit related to cognitive science, the dual coding theory.
7:09Dual coding? Yes, the dual coding theory. Okay. So to simplify that, basically the dual coding theory illustrates that the human actually process information through two different channels, verbal channel and nonverbal channel. So for nonverbal channel, it's similar like an imaginary system where the human uses imaginations to process the information. For example, like human naturally form and manipulate the mental imagination to understand the physical world and kind of to conceptualize the visual outcomes once you, for example, you conduct an action. And it is quite straightforward, for example, in education scenarios, like when being educated using verbal and visual materials actually helps to help the student to understand the content more effectively than only using a long textual description.
8:17Right. So, yeah. And I guess, yeah. And this further strengthens us. Like MVOT, I think it's a really good idea. And we should definitely dig into it. Here comes the project. So you talked about the task in your prior work that you carried into the MVOT paper. Dig into that in a little bit more detail and kind of explain the setting and like how you're kind of bounding the problem. So my previous work, Top View RIS, which is a benchmark that evaluates the performance of models in understanding the spatial layouts in Top View Map. And it has four different tasks of increasing complexities. The first task is about top view recognition, which is basically to determine whether this object is presented in the map or whether this room, for example, the bedroom, is presented in the map.
9:21It doesn't really require the model to locate where it is. It only requires the model to answer yes or no. right? And the second task is top view localization, which takes a step forward. Instead of just answering yes or no, you have to specifically tell me where it is. For example, where is the bedroom in the top view map? Whether it's at the top left corner or it is in the middle center part of the map, you have to tell me where it is. And once you have already figured out where it is. Then the third task is to reason over relative spatial relations between different entities. The entities could be an object or a room.
10:06For example, in the top view map, what's the relative spatial relation between the toilet and the bedroom, for example?
10:17But the toilet didn't change its position in the top view map, right? So it cannot move around. So the last task is about dynamic spatial reasoning, which I think is quite interesting because for us humans, when we see a map and we want to reach a destination, we tend to draw the navigation path on the map from your starting point to the destination. And it's very intuitive for us to understand what's going on here. But for the model, sometimes it's a little bit tricky because although you're using a static map, it actually contains a sequential information from the starting point to the destination, which poses a little bit of challenge for the model to understand this.
11:04So that's, and in the dynamic spatial reasoning task, we will ask the model, for example, which direction are you heading for? Or like at which room or at which object did you make the very last left turn or very last right turn? Or at which object you make the first left turn, et cetera. So that requires the model to really follow the path in order to answer this question. When I looked at some of the pictures of the scenarios in the paper, in some ways it reminded me of like maze solving, but without the maze. Did you draw on, and that's something that has been explored in machine learning and reinforcement learning in lots of different ways for a very long time.
11:53Are there parallels and does that work help you at all with this work? Yeah, I think for the maze one, it's obviously a question that has been researched for a long time in RL and other views. But I guess it definitely helps. And I guess the main difference between our work and previous work is that I think we're among the very first ones who tries to tackle the problem with a pure neural model. like, for example, language models and MLMs, I think everything we do with, for example, generating the images or generating the next steps, I think we try to tackle or research these questions in a pure neural setting.
12:44And I guess that's the main difference between our work and previous work. And so talk a little bit about, you know, you've got a language model that understands things verbally and we've got multimodal models now that can incorporate images. How are you incorporating these two components into the visual reasoning or spatiovisual reasoning in MVOT? So previously, the VLMs can only take multimodal input, which is text and image, for example, and then output pure text. So this is what most of the previous VLMs or MLMs can do. But recently, we do see a lot more models that could generate the multimodal output as well.
13:36so whether it is using different decoders or whether it's using the same language model backbone but the output is multimodal so it gives us so it kind of enables us to really have the MVOT working with these models because otherwise we cannot really generate the imagination or visualizations of the model in our tasks. So I guess that's the whole paradigm shift from previously only text generation to multimodal generation. And the current way, for example, we choose ANOE as our model backbone. So ANOE is one of the models that could generate the multimodal output. And the way it is handling this is that it uses discrete tokens to represent the image.
14:43And as you may know, the text is also represented with discrete tokens. So they have a uniformed representation of both text and image. And if you want to generate text, you will just output the textual tokens. And if you want to generate image, you will output the image tokens. And then the image tokens would be fed into an image decoder to decode the image. And I guess that's the whole thing that enables MVOT. Maybe talk a little bit about, I'm trying to make sure that we understand, we talked about the setting. I'm trying to make sure we understand the task and what you're trying to do. What's the input data?
15:34What are you asking the model to do? And talk a little bit more about what it looks like from the perspective of the user or problem. So in MVOT work, there are actually three different tasks. The first one is maze. The second one is mini behavior. and the third one is frozen lake. So before we really dive into the specific tasks, and I'd like to first give you a really brief overview of the whole input and output format so that you can have a better understanding of the MVOT framework. So the MVOT framework, so the model takes an image and the corresponding textual question as the input. And then for the output, it can output image, which is a visualization of thought, and the verbal thought as well.
16:28So the input would be multimodal. The output would be multimodal as well. And in our case, the maze, as you may know, that we require the maze to, for example, we are given the model maze, its starting point, and we are asking the model, if you follow this instruction, where would you locate after executing the actions? So the model is kind of supposed to anticipate the final destination by following the actions. And for mini behavior, it is to simulate a robot carrying out a specific task, which is to pick up a printer from the ground and then carry it to a table and then turn it down. And compared to the maze one, the mini behavior extends the action space to a more diverse selection.
17:27So previously, the maze only contains like go up, go down, go right, go left. But in mini behavior, it asks like, for example, pick up or drop or toggle. And this makes the action space more diverse. And for Frozen Lake, it is a very classic video game. And the environment is an ice lake with some hidden holes. And the model is given a sequence of actions. And it should determine if the action sequence is safe to execute without falling into the holes. And so it compares to the maze. The image is more complex. and because it contains a lot of details in its patterns, like eyes holes or a little elf as the agent on the image.
18:20Yes. And that's the overview of the task and the data, what the image looks like and what's the action space. And when MVOT is reasoning over these different tasks, different tasks. So it first takes a very initial image and the task instructions as the input. And the MVOT at the first step, it will see after executing the very first action, where am I or what the stace currently looks like. And the stace is represented with an image. And then after having this kind of generated image, the MVOT would determine what's the next action. And after figuring out what's in the next action, the MVOT would think after conducting the second action, what's the state would look like, and then generated a visualization to illustrate the reasoning.
19:20And after having the visualizations, then the model should determine the next action and then to anticipate what the state looks like, and et cetera, et cetera. So it's kind of doing it in a recursive way with interleaved text and image, and then at last, conclude the answer. And that's the parallel with chain of thought. It's like, as opposed to asking the model to, in one step, try to figure out where it might end up in the maze or what objects it might come across in the mini behavior, you're asking it to think sequentially from one step to the next and follow that path along? Yes. So I think the main difference between chain of thought is that sometimes because chain of thought is purely verbal, so it tries to describe everything with language.
20:13And sometimes the language is not that effective or accurate in referring certain patterns in the image or trying to describe the spatial layout. And I guess that's it. I guess MVOT, because MVOT uses the image to do so. So at least it's more straightforward and more effective in conveying what's going on and describing the spatial layout in our case. Are there examples that come to mind in your experience, either training this or exploring models or problems like this in prior work of where the model gets confused with just textual information or why that happens? Or, you know, examples that like illustrate the ambiguity you referred to?
21:08I think some of the experiment results in MVOT, actually we do see when the complex, when the image becomes super complex For example, in Frozen Lake, if there are a lot of ice holes in the map or in the ice lake. And it's very hard for the model to really know, really tell the exact coordinates of each hole very clearly and very accurately. But to be honest, we didn't really care about the exact coordinations of each hole, right? We only care about whether the agent would fall into the hole or not. So it's kind of introducing some unnecessary errors in trying to describing the whole image. And sometimes the description, describing itself, it's very hard for the model to do so.
22:11So yeah, and that's the reason why we do see that on Frozen Lake compared to chain of thoughts with coordinates, the FVOT actually achieves a great improvement. One of the elements of your approach is the introduction of this idea of a token discrepancy loss. Can you talk a little bit about that and how it comes into play? So as I previously just mentioned, like for the current models, for example, they use different tokens to represent text and image, right? And the token discriminatory laws is kind of trying to make both parts understand each other. It's similar like, for example, I speak Chinese and you speak English.
22:56And the token discriminatory laws is that try to enable us to understand each other with different languages like language and image. So, yeah, and I think that's the intuition of the token discriminatory loss, because previously, the language model or the MLMs is only optimized in the language embedding. I would say token embedding, because the embedding is intended to be used for language modeling. but when processing the image, because you have to first translate the image into the image tokens and then feed the image tokens for language modeling, right? So when you try to translate the image to image tokens, you're using another set of embeddings, which we call visual embeddings.
23:50But visual embeddings is not involved in training the language model backbone for reasoning. So I think it serves either as an auxiliary information, and it could equip the model to have more information about what's going on in the visual space, or at least make the model be more aware of the visual features, I would say. Yes. And the experiment result actually turns out that it's actually been helpful. And without token discriminatory loss, the model may generate the images that are technically correct. UK is seemingly correct, but the generation doesn't really match the intended meaning or modifications in the corresponding area of the image.
24:46And a lot of irrelevant details would pop up and the important elements would be missing. And we did design several metrics. And we see that after introducing the token discrepancy loss, the performance on all of the metrics gets better. So at least it shows that the visual information is playing an important role or at least it's helpful for the model when generating the visualizations. When you talk about creating extraneous detail or missing important detail, is that specifically on the vision side or do you find the same effect on the tech side as well if the two aren't kept synchronized through the token discrepancy loss?
25:37Yeah, when I mentioned this, it's specifically on the vision side because we see that if we didn't introduce the token discrepancy laws, sometimes the visualization would be totally run. For example, we ask the agent to go up and the visualization turns out the agent actually go right for stuff. And sometimes the agent would just go vanished. We cannot see any agent on the image. So yeah, and that's purely on the vision side. Would you say, you know, is there some intuition why the, it seems like there's better alignment on the tech side, independent of the vision side, but the vision side is maybe less aligned and requires this token discrepancy loss as an alignment tool.
26:27am I hearing that correctly and if so is that you know because would you say that's because of the complexity of you know visual images or just the amount of data that you know is pre-trained on the tech side or do you have any thoughts on that yeah that's a very good question I think it has perhaps it may be related to how image is incorporated into the current language model backbone. Because as you may know, that the current MLMs is most, I would say it's mostly dominated by the language model backbone compared with, especially when it comes to model sizes. Once you converted the image to image tokens, then the image side has already played its role.
27:20And it's all about language modeling joint training together, right? So I think that's the reason why for the visual side, it's not that well aligned or it still has some problem compared to textual reasoning and generation. With the three task types, the maze, mini behavior and frozen lake, were these, are these existing games that were out there and like had data sets or did you create these data sets? So all of these tasks are already well-defined, but they didn't have the data. So we collected the data ourselves and then trained the model. Got it. And when you trained the model, was the model trained with only the input and the desired output as supervision?
28:17Or did you have some supervision signal around the intermediate steps? I think I'm not quite sure what you mean by intermediate step. That's the reasoning process. The reasoning process. Yeah. Yeah. So it's definitely, it is trained with reasoning process. So in MVOT, we actually calculate, we actually optimize the model in a recursive way. So because the current MLMs has a context length, for example, the annoy, I think, has a context length of 4 ,096 tokens, which means that it at most can handle this length of input or output. but an image would cost like 1 ,024 tokens already. So which means that at most they can intake like three images at the same time if there are other texts as well.
29:20So in this case, we wanted to, we do it in a recursive way. So for example, when we feed the model the very first image and the task instructions, then we ask the model to generate the first action it should conduct, right? And then after the model doesn't know what's the first action to conduct, then another trained data would be like, once the input would be the very first image, and the task instructions, and the first action to conduct. And the output would be the generated visualizations. and then once we have the generated visualizations, then another type of training data would be like the very first image, task instructions and the first action to conduct and the first visualizations.
30:18But this image is golden. So it's not generated by the model during training. So this is golden. And then the model should generate the second action. So we managed to keep there are only two images in the input, and then the output would be either, would be the text and image. So at the maximum, there are two images in the input, and at the maximum, there would be only one image in the generation so that we can stick to the context length of the current MLMs. Yeah, I was thinking about my previous question. I'm realizing it's hard for me to not think about this as a RL type of a problem, like where you assign some value to achieving the ultimate destination and let the model figure out all the intermediate steps on its own.
31:21Yeah, that's a very good idea, especially when we see the potential of RL. It's very curious for us as well to see whether it could play a role in this multimodal reasoning. Talk a little bit about as you trained the model, were there any lessons learned or observations that stuck with you and potentially impacted the way you evolved the training process itself? Yeah, I think the lesson that I learned from training MVOT and testing with MVOT is that the quality of the visualization really matters here. Because previously, before we introduced a token discrepancy loss to improve the quality of visualizations, the pipeline of MVOT doesn't work because the visualization sometimes is just incorrect.
32:23And then because it is incorrect, then the final answer would be wrong as well. And I guess the very important thing in an VOT-like thinking system is that how we can assure very accurate and precise visualization during the reasoning process. And this is not really related to the previous image generation task because we really don't care about whether this pixel is correct or not. we're more interested in whether the overall spatial layout is correct and you're not trying to mess around with the image, with external redundant details or make some of the elements disappear during the reasoning process.
33:22Yeah. And I guess we really try hard to assure that the model is behaving as we kind of expected it should be. There was a fine-tuning element of the training process. You started with this Anul base model, and then any particular comments around the fine-tuning approach? I don't know if there's anything interesting there to explore, I guess is the question. Yeah, for MVOT, we basically just use LoRa to tune the model. And it is, it's not, I would say there's nothing special what comes with a LoRa setting. Yeah, but I guess it will be really interesting to see if the MVOT-like reason-share strategy can be extended to more diverse scenarios or to make it really as a foundation model that is more general, you know, because currently MVOT in our paper is just focuses on three very, I would say very toyish tasks.
34:41And it's very interesting to see whether this, how can we assure high quality visualizations when it comes to real world scenarios with more complex images and also more diverse set of actions, how would it behave? And I guess that's a very interesting and exciting direction for us to explore in the future. How toyish were the problems? Like, how did you measure the complexity of the input and what, you know, bounds did you push there? How big were the mazes? How big were the ice patches, that kind of thing. Yeah, so in our work, the size of the mazes and frozen lakes, we try to, in our training data, the size ranges from three to six.
Read the full transcript
35:33And on many behavior, I think it ranges from seven to ten. Meaning three by three to six by six? Yes, yes. Okay. Yeah, and I guess by toyage, I'm not saying that our task is really simple because as you see that GPT-4.0 cannot really solve our task very... But by toyish, I mean that there are lots of details and different scenarios that can be explored in real life scenarios. like for example, robotics or something like, I would say, like architectural engineering. I don't know. When we try to draw a blueprint, it's definitely more complex than a maze, right? And yeah, it contains more details and it puts higher requirements for the model to really when it comes to modifying the visualizations it requires more accurate modifications and manipulations over its imagination yeah and that's the toy i'm referring to yeah yeah can you elaborate on you know in the architecture or some other set of real world scenarios like what it might look like um you know if you had a foundation model that's um you know scaled to those kind of scenarios, what would be the kinds of questions you'd be able to ask it?
37:04And, you know, what would you expect to see there? So, for example, like, I think the very first intuition, I do see a lot of people discussing these days about MVOT applications is robotics. So, for example, if you have, let's say, if you have a navigation robot and you ask it, like, I wanted to grab a drink from the refrigerator. so it has a thing internally like where is the refrigerator oh it's in the kitchen and then it has to find a way to the kitchen right but finding a way to a kitchen I don't think it really happens in a symbolic way like I there I don't think there is an equation that could describe um which action should I take in order to reach the uh the kitchen it's more like if I were the robot it's more like if I go forward, can I see the door of the kitchen?
37:58And if I can, then this is the right direction. If I can't, then I would try different directions. And then thinking this in my mind is a similar process of MVOT. And that's the very first intuition about how these MVOT-like thinking systems can be applied in some specific scenarios. And then, for example, like I just mentioned about architectural engineering and urban engineering, I think the same case, for example, when you try to design your house and how to decorate your home. and you, for example, if I want my balcony to have more sunshine, if I have an AI system that can help me design this, then I'll tell him, like, I want my balcony to be on the side with more sunshine.
38:55Then the AI model should be able to understand, like, which side has more sunshine and then kind of visualize if the balcony is being put there and how would the overall layout of your home looks like in a blueprint and et cetera, et cetera. So I think it's very exciting to kind of imagine what's going on with a foundation model that could reasoning or thinking with visual and verbal thoughts. in this case, yeah. Are there alternative approaches that you are planning to try to tackle the same problem? I guess, yeah, I think definitely there would be alternative ways to tackle this. For example, recently I see a lot of papers that conduct reasoning in latent space, right?
39:56And MVOT explicitly shows the reasoning process with image as the visualizations. And I wonder if latent space could be a more powerful representation than image. I'm not quite sure about that. But I guess that's definitely one of the interesting directions to look into. because after all, I don't think there is a well-acknowledged conclusion of how we human thinks in our brain, whether we think with images or the thinking process is latent presentations. We still don't know. So yeah, and I guess there will be a lot of interesting work going on and I really hope to see what's going next based on our work and recent works on later reasoning.
40:52And that might help with some of the context length limitations you mentioned as well, right? Yeah, yeah. And also about some computational overheads being introduced by generating the images. Well, Cheng Zhu, thanks so much for jumping on and sharing a bit about this paper. It's super interesting. Yeah, thank you for having me. It's a really great discussion with you about our paper. I hope you enjoyed it. Thank you. Thank you.
From the publisher
Today, we're joined by Chengzu Li, PhD student at the University of Cambridge to discuss his recent paper, “Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.” We explore the motivations behind MVoT, its connection to prior work like TopViewRS, and its relation to cognitive science principles such as dual coding theory. We dig into the MVoT framework along with its various task environments—maze, mini-behavior, and frozen lake. We explore token discrepancy loss, a technique designed to align language and visual embeddings, ensuring accurate and meaningful visual representations. Additionally, we cover the data collection and training process, reasoning over relative spatial relations between different entities, and dynamic spatial reasoning. Lastly, Chengzu shares insights from experiments with MVoT, focusing on the lessons learned and the potential for applying these models in real-world scenarios like robotics and architectural design.
The complete show notes for this episode can be found at https://twimlai.com/go/722.




