In short
The TWIML AI Podcast - Episode #646: What’s Next in LLM Reasoning? with Roland Memisevic
Episode Summary In this episode of The TWIML AI Podcast, host Sam Charrington interviews Roland Memisevic, a Senior Director at Qualcomm AI Research. They discuss the evolution of language models (LLMs), the significance of language in humanlike AI systems, and the emerging focus on LLM reasoning. The conversation covers the advantages and limitations of autoregressive models, the role of recurrence in reasoning, and the importance of grounding AI in reality.
Key Points
Introduction to Roland Memisevic
- Roland has transitioned from his previous work at 20 Billion Neurons to Qualcomm AI Research.
- He continues his research focus on combining perception and language for developing situational awareness in AI systems.
The Importance of Language in AI
- Language is seen as a critical element for developing human-like intelligence in AI systems.
- The traditional view of language processing has often focused on nouns; however, understanding verbs and relationships is equally essential for creating intelligent systems.
- Roland emphasizes the need for a rich representation of concepts to enable better feature learning in AI models.
Evolution of AI Models
- Early models were based on LSTM and RNN architectures, which have been largely supplanted by transformer-based models.
- Transformers allow for faster training due to parallel processing, contrasting with the sequential nature of RNNs.
- Despite the advantages of transformers, Roland notes they may lack the deep reasoning capabilities inherent to recurrent structures.
The Role of Fitness Ally
- Discussed as a product of Roland’s earlier company, Fitness Ally integrates these concepts of grounded reasoning and agentic interaction.
- The system acts as a virtual fitness coach, employing language and visual input to engage users effectively.
Reasoning and Autoregressive Models
- There is a distinction made between reasoning processes and the outputs generated by LLMs.
- Roland argues that reasoning is inherently linked to language, and models can perform reasoning-like tasks by generating responses based on learned patterns rather than actual reasoning processes.
Addressing Current Limitations
- Roland points out that LLMs can encounter challenges with tasks that require true understanding versus mere pattern matching.
- The concept of "length generalization" is introduced, highlighting difficulties that LLMs face when extrapolating beyond training data.
Grounding in Reality
- Grounding AI in a perceptual reality is crucial for developing reasoning capabilities.
- Research is increasingly focusing on visual grounding, where models gain contextual understanding through interactions with their environment.
Future Directions
- Roland predicts a resurgence of recurrence in future AI models to enhance their reasoning capabilities.
- He stresses the importance of combining language models with sensory inputs to create agents that can develop a sense of self and understand actions in a broader context.
- The conversation hints at the potential for models to understand concepts like causality and object permanence through grounded interactions.
Conclusion
- The episode wraps up with the recognition that significant progress is anticipated in the field of AI over the next few years.
- Roland expresses optimism about the integration of various components—language, visual inputs, and agentic behavior—to advance AI capabilities.
Key Takeaways
- Language is foundational in developing human-like AI.
- Autoregressive models excel in tasks requiring language but have limitations in reasoning and understanding.
- Future advancements in AI will likely involve integrating recurrence and visual grounding to foster deeper reasoning capabilities and situational awareness.
- The development of AI systems that can engage in continuous learning and adapt based on interactions is crucial for their evolution.
Further Exploration
- For more insights and detailed discussions, refer to the complete show notes on [TWIML AI Podcast](https://twimlai.com/go/646).
---
This markdown summary captures the essence of the podcast episode, providing a structured overview of the key themes and discussions that took place during the conversation with Roland Memisevic.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:09All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington, and today I'm joined by Roland Memesevich. Roland is a Senior Director at Qualcomm AI Research. Before we get into today's conversation, be sure to take a moment to head over to Apple Podcasts, Spotify, or your listening platform of choice. And if you enjoy the show, please subscribe, rate, and review. Roland, welcome back to the podcast. This is three for you, right? Yeah, thanks. Great to be here. And I think it's the third time. Yeah, it's great to be here again. I'm really looking forward to digging into our conversation.
0:45We're going to be talking about LLMs and reasoning and how your work over the years, since that initial conversation has evolved to really take on these topics. Since it's been a while since we've spoken and the first time since you've joined Qualcomm AI Research, I'd love to have you kind of talk through your background and how you got from 20 billion neurons to where you are today. Cool. Yeah. So last time we spoke, I think it was around two years, two or three years ago, something like that. And the first time about five, I think. Yeah, something like that. And we talked about the company that I co-founded and that was 20 billion neurons.
1:26The things we've been doing there, which essentially has been the stuff we've been doing before and then during the company and are still doing now at Qualcomm. We all joined Qualcomm as part of an acquisition around two years ago. So not too long after the last time we spoke. And we are now with the team over here and pursuing essentially the same research direction around combining perception and language so that you can build agents that can drive an end-to-end vision for AI further than we believe other approaches to AI have been able to do that. And so that has been a very long, long story over the years.
2:03I was actually already interested in this particular kind of thing, like building a situated agent rather than modules that solve certain subproblems. Previously, when I was on faculty at the University of Montreal, and then I realized it would be best to pursue this at a company, compute resources, data, all of these things. And now we moved all over to Qualcomm and continuing along the same lines. Awesome. Yeah, I remember what led to our first interview was I think I ran into you and a couple of your colleagues. I think it was like a rework conference in Montreal. Yes, that's right. And you were demoing this data set that you'd created that was trying to capture actions.
2:44Like, I think the way you described it at the time, which I thought was really interesting, was, you know, all of this work we've been doing in computer vision was noun-centric and you were trying to do verb-centric data sets. So how did that work lead you to what came later with the Fitness Ally project and your interest in reasoning? So you are talking about nouns and verbs, which already kind of alludes to the whole reason for all of this, which is essentially language. There has been a really long standing tradition, I guess, over the last 50 years, probably 50 plus years. A tradition or let's say assumption that language is a key ingredient to any kind of human like intelligence.
3:25And there have been many, many efforts like Eliza and all of this stuff over the decades. And I'm definitely a deep believer in language as being a must-have ingredient to build intelligence systems and to make them human-like in some way. So that it's the ingredient to human-like AI. And interestingly, the community at large, I would say, discovered around 2012 after the ImageNet incident happened and people were training networks on ImageNet, which A, worked, okay, which is cool and it kind of revolutionized vision and vision in many ways and so on. But more importantly, that the features, penultimate layer features or whatever, that the model learns in response to being trained on those thousand classes.
4:08Highly universal visual features that can solve many other vision problems besides categorizing cats and dogs and houses and cars and such things. So that was from a language, blunt language perspective. It's true. It was thought of about nouns, not about verbs. but I think it missed a more fundamental point that is behind all of this which is the moment you think about labels not in terms of one of thousand like you say label 37 but you actually give them names like fish but then later you do it for verbs as well and maybe for adjectives and adverbs and etc etc so once you have a rich representation of the kind of concepts you want the model to acquire then you have a much much more fertile ground for the model to develop features that are then useful for downstream tasks, right?
4:55And so the name of the game has been also at the company, how can we drive this end-to-end vision where you train a network supervised on one task, a bunch of tasks, however you want to call that, on something, you train it supervised, not to solve the task, but in order to instill features in the network, representations that are going to make it in some way universally intelligent. And so once you want to do that, want to instill capabilities in the model, you wind up using language. There's no other way than using language to represent the concepts that you want to have as a supervision signal to the network, right?
5:31So in that data set that we discussed some five years ago or something in one of the earlier discussions that we had, it was called Something Something, and it was about adding verbs, but also adding adjectives, adverbs, et cetera, et cetera, et cetera. In other words, adding captions to videos and training a network to make predictions, so predict those captions, detailed predictions like there's an object that fell down over here and it hit the other object whatever crazy weird prediction stuff again not so that the network is able to do this kind of weird stuff but so that the network learns about things like occlusion temporal assistance of objects learns about objects in the first place right when an object moves it doesn't usually teleport itself and just the idea of object comes about partly because there is something persistent about it across time, even though you have motion and these kinds of things.
6:20You can learn about material structure, know something about rigidity and so on and so forth. And so it's a way to instill cognitive skills, base skills, perceptual skills in that case, in the model, hoping that it would make the model smarter. And that whole vision then got pushed further and further to cases where the system was supposed to be chatting with the user in real time, sharing a joint space. So you're like in front of a table with some objects on it or the user is doing things or something like that. And so you have a face-to-face conversation just like we have right now, and you can show things and discuss things and basically immerse the model in a real-time, real-world environment so that it learns common sense capabilities.
7:04So it's common sense that language can easily be used. Maybe some people would say misused, but maybe that's one of the main reasons why language is so successful in evolution, can be used to instill in models, right? So having a shared name for things, we call things fish, and the one individual doesn't call it label 37, and the other one calls it label 501 or something like that, but there is a joint notion, and we both call it fish, and so that way we can teach each other about what the objects are, and we have learning signal in human societies and whatnot. And the same thing carries over, I deeply believe, to AI systems, where once we use language, we can align the representations with the ones that we have fairly well, and that way get some kind of common sense and human-like AI going.
7:50And how did this lead you to Fitness Ally, which was ultimately the product that you were building at 20 Billion Neurons and continues today as a kind of a platform for research? it's a bit broader than that. It was one product. It was one instantiation of getting value out of this weird system that we were basically up to, right? So the idea being, okay, we want to build end-to-end systems. We want to drive this end-to-end learning to the end. The system should be able to have auditory input, visual input. We didn't go as far as sensory input and robotics because we don't want to leave us mechanics and whatnot and didn't feel robotics is quite up for it.
8:28But it's all kinds of sensory and typical human-like sensory input, and specifically vision and audio. And it has to have a body, so it has to be able to convey things by moving around and so on. And it has to give you auditory output, speaking to you, and also visual output in the sense that there's a camera looking at that creature that we built so that it can basically have a sense of a conversation with you. So there's this end-to-end learning pushed all the way to the end, save for robotic, like true embodiment. Sometimes we call this virtual embodiment. So we wanted to build a virtually embodied AI system, a virtual robot, and exposing it in real-world conversations with real people in real time so that it learns something about the world that it would just never learn otherwise if we train it end-to -end, supervised on image-style labels and action labels for videos, etc.
9:16etc. So that was a vision and we were convinced that this is the only way to build human-like AI systems, train them end-to-end on the right setup of tasks. And those tasks have to be of this nature, sensory input, language, back and forth, and learning that way. And so we tried a whole bunch of stuff, turned this into a vision that is also commercially viable. And so we built things like retail assistants that would teach you things about objects at an end-of-aisele display. And And then they tell you, do you want to try this on? Maybe sunglasses. There was one demo that we showcased at NERVES at some point.
9:50And then we had various other use cases where this avatar would just provide value. And then fitness was one of those explorations that we did that stuck very well. And this is probably not surprising. Fitness is a case where you can have real value if you have a companion of sorts that knows you through the days, weeks, months, years, knows how you've been performing your exercises and so on. then guides you by giving you feedback and so on. There's an obvious superficial utility in that system can tell you what to do to avoid injury maybe or to be more efficient. There's a much more subtle and much more powerful one, which we were also after, which is companionship.
10:30So that system, by knowing you and knowing how you've performed your pushups in the past and how many you will be able to do, it's much better at holding you accountable than some bean counting app that basically just has a dial that goes up, a number that increases every time you do one push-up or something, but you feel a little bit more, you know, you don't want to disappoint that trainer in front of you that remembers yesterday you managed to do this many, and today, come on, you're going to do maybe one more, and so on and so forth. So some degree of companionship. And so that got stuck and worked very well, and users liked it, and so we pushed really hard to push that further.
11:07It took a while to land on this particular type of use case and vision as a use case that straddles between these two, which is crazy, right? If you think about it from a business point of view, maybe from a research point of view, it's a system that is really useful and provides real value. And at the same time, it's intimately linked to this end-to-end vision of having an AI system that has to have a real-time language-driven conversation with you about stuff that's happening in front of the camera, which I still believe is the only way to build systems that are human-like, sufficiently human-like.
11:40And so now FitNagala is a system that we had in the app stores, various app stores, and people were subscribing and using. And now after the acquisition, we're still pushing this further as a demo. And for example, combining the language capabilities of typical LLMs, we train with the visual capabilities, also trained in part through language of that perception stream that comes in. It's still one end-to-end system that has to learn and be fine-tuned end-to-end in order to do what it's doing. I was going to mention that a lot of this research predated LLMs. How have LLMs kind of impacted Fitness Ally, but really more broadly, your research interests and research agenda?
12:27Yeah, it's interesting. So we did give us predecessors of LLMs, which used to be LSTM-based RNNs, right? So RNNs had to be a part of this whole stack because you want to have language in order to instill capabilities and so on. So you need to have an agent also that uses language in order to do stuff. We had early experiments of various nature, maybe, for example, showed that if you train a model that has an RNN-based language decoder, autoregressive language decoder linked to the visual input, and you train it across captions, so like, for example, in the something-something scenario of increasing complexity, so instead of doing classification, you do various levels of increasingly sophisticated, detailed description of what's going on in the scene, that then you learn better features for downstream tasks that are better generalizing and that are better infused with a degree of common sense understanding of what's going on in the world.
13:21So we had these kind of studies. They used to be RNN-based. The big difference, the big transition, which I think is going to be a transitory transition, by the way, historically in AI, but the current transition to attention and the lack of true recurrence has some benefits in training. And so they are much better at considering language capabilities that are so deprived of On the other hand, deprived of visual input, perceptual input, but they are better at just handling language and so on. So that's much better right now. And so it's easy, though, since those models are also autoregressive.
13:55So they are recurrent in quotes. They generate one token at a time and they can still be agentic because of that. So they can still have a back and forth conversations and leverage all of those aspects of language. But the good thing about that way of getting at language is the current way of getting at language is that you can train them on parallel hardware, which you cannot as easily do with RNNs. When you have an RNN that generates one token at a time, you generate a linear sequence of events that you're going to have to backpropagate through. So you have a factor T, where T is the length of the sentence or whatever output that you generate as a multiplier in the computational time complexity of training.
14:35There's just no way around it. You're going to have to generate one token at a time for a certain amount of time until you have your attempt to, say, replicate a training case or something. When you look at a transformer, that disappears because now you can just have an input. Every column that processes the inputs can just try to predict the next token. You get losses across and you can just do one. In one step, so to speak, you can process this whole sequence. And so this factor T magically disappears as a constant that influences the training complexity, time complexity. And so you can train them much faster.
15:10So I think a reasonable assumption is that if you would take an LSTM-based old school LLM, whatever, small SLLM, small large language model, it has to be large. So and you would train it for like 100 years, then we would probably get something as good as or probably even better than a top shelf GPT pre-trained language model. It's just that you would have to wait 100 years to train it. And so that's why it's currently the field is dominant. Any language stuff is dominant. large is dominated by parallel stuff. But it's a coincidental technical. I think the point, just to kind of linger on that last point you made, I think you're saying that LSTMs, RNNs are not fundamentally inferior in their ability to capture language, but rather from a practical perspective of training them, they are not as efficient by a significant degree than transformers.
16:10Exactly. In spite of like the difference in, you know, transformers have this ability to capture longer context. Are you saying that if you can, you know, if you can make T arbitrarily long and you had sufficient compute to, to train with an arbitrarily large T, is that part of what you're saying? It's part of it. So you couldn't, you can add other things to that, right? I think the complexity is a key factor. There are benefits, obviously, and the verdict is out. Nobody knows, right? Like what if you had recurrence, yet you have the ability for some long-range stuff. Essentially, what humans are doing is some kind of reverberating shorter memory, right?
16:51So things that you just said, there's an auditory loop. You remember that for a short amount of time, but it's quite accessible to you at any moment. So the last few seconds, you still have them somewhere reverberating around. And so the ability to just grab that and take it and use it and stuff is probably another aspect in which, say, GPT, whatnot, is superior over some LSDM. But there wasn't LSDM or RNN or something like that. Or even back then, people were exploring many, many different incarnations of RNNs with many different clockwork type models where you look back a certain amount of time.
17:26And then RNNs were the initial type of architecture where even attention in an encoder, first in an encoder, decoder kind of scenario got explored. And so it's very natural to have this, to deviate from this linear state-by-state transitions or something like that, even in RNN. By the way, I think it's probably a bit of a trend also right now with papers that are pointing this out. But there's an interesting fallacy for taking that shortcut, pretending to be recurrent, but actually being fully parallel and momentary like GPT style model is. And that is the fact that they can be pseudo recurrent in the sense that they can generate sequences and seem to rely on past actions to do current actions and this kind of stuff.
18:12Where in reality, they can also fall prey to being tempted by pattern matching, essentially, right? So they can basically find similar patterns or they can be doing certain actions just because the context up to now just has a certain signature that makes it look similar to certain training cases, after which that action is also the right action to take or the model learned to take. And so some of the capabilities that seem very striking can maybe be explained away by a certain degree of pattern matching and then something much simpler. And one problem in which this is actually turning into a real problem is length generalization.
18:49So there are various studies lately popping up that study what happens when you take an LNM, train it on a certain length sequence of a trivial task, like multiplying things and parity and whatever, like certain symbolic tasks. And then you try to have it generalized across beyond that window that it was trained on. And so I'm not talking about context window, right? The context window might be much larger than any of those either training or test situations, but it is just that they were trained to solve instances of a certain tiny length and then try to generalize into a slightly larger length and it completely falls apart because it doesn't understand the gist of the problem.
19:27And where does that come from? Well, it comes from the fact that the model isn't actually recurrent. So it doesn't have the same notion that you have about in order to solve this multiplicative what not problem. I'm going to have to take this token here. Then I have to combine it with that one here. And then I have to do this. So this whole sequential nature of arriving at the answer is not there. The model learns this on the training data because it, again, it gets tempted to just do subgraph matching on instances that look similar to training data. And then it just comes up with the answer. And so it seems to being serial in sequential processing those multiplicative problems but in reality it's not and that is a big problem that is actually probably going to require recurrence to come back but now people can approach it from a different angle you can start with some model like that and then sprinkle in recurrent connections and fine-tune them and see what yeah but i think recurrence will be back right it's not i was going to i was going to ask when you say that you think llms are a bit of a transient state it sounds like the next state or n plus n you know n plus one state is some degree of recurrence coming back into the model is that kind of what you're thinking yeah exactly that that's at least one one way in which those models are going to have to evolve and will evolve and it's it's okay right so they provided a lot of value a lot of insight now i want to shift gears a little bit and jump into this idea of reasoning particularly in the context of language.
21:05It's been a topic that I've been super interested in as have many of us in this field. And, you know, one way I've kind of come to, you know, snip or soundbite that I've come to is that LLMs do not reason or cannot reason, but they can produce results that seem like reasoning. And the idea is to reinforce that, you know, they're next token generators. They're not following a process that we might think of as reasoning, but somehow the result of running this next token generation, you know, some number of times looks like or produces the results of reasoning. And I just wanted to throw that out there for you to pick apart.
21:46I don't agree with that. You don't agree with that. Yeah, let's get into it. Tell me why. Tell me a human being that is not producing one micro action at a time, one after the other, right? It's not possible. So at the end of the day, atomically, whatever you want to call this, the robot that we are, the robots that we are, whatever, or when we talk, right, the language that we output, it's always a string of, let's call it tokens, you might call it whatever you want. You can also say phonemes, whatever. You could have done it on the audio level, even if we had more compute and stuff, whatever.
22:21Let's call everything tokens that are model outputs, that are agent outputs. You never do anything but generating one token at a time. And that's just a fact, right? There's space and time, right? The constraints of our physical world and time just advances and we can just do one thing at a time. And that's just what it is. It's just a fact of nature. And so if we build an AI system... I don't know that I necessarily agree with that. Like what we know is that, you know, because we've got, you know, one vocal apparati, the output is one, the output is serial from a word perspective. you know of course but from a biochemical perspective there's lots of things happening at the same time they're just serialized by this output device and i guess that's absolutely fair i agree with that and i agree with partly with the idea that some of the the chains of thought that you see are pretend reasoning and the recurrent the lack of recurrence is just one example of that that's for sure on the other hand want to build an intelligent system i think we We have no other, that is human-like.
23:26So we really want to advance this agenda, like building more human-like intelligence. So we have no other choice than building an agent that has a perceptual input and has one output at a time, an actuator, and it goes through, it is an autoregressive model. It has to generate one thing at a time, influenced by whatever it sees right now and the history of whatever it saw before and whatever it has generated before. So that's it, right? That's a universal definition of an agent, of whatever reinforcement learning agent. And there's a lot of benefits to formulating a system like that versus going for feed-forward networks and solving vision and then solving this and then solving this.
24:03The benefit is that the driving force in neural nets in general, starting with image-net, creature-networks, speech and whatnot, it's end-to-end learning. So you can do supervised learning and you get domain transfer. So that benefit is real and that's what I believe is the 99 % of what drives AI forward. That benefit can carry over to agentic situations where you want to build an agent just as well if you take as the artifact that you're building an agent that has to do one thing. And this is a real agent. It's embedded in a real world that has time marching on and it has to do one thing at a time.
24:39And the environment is unforgiving. And so sometimes it has to do things fast in order to catch up and whatnot. But it is an agent at the end of the day. When we train agents like that, so when we define the AI model not as a function, but as a procedure, right, as an agent that does things over time, then transfer learning, I believe, can play out in the realm of action sequences and decisions and all of this kind of stuff. And that, I think, was a big problem with a lot of reinforcement learning in the past. So in lots of value-based approaches and the whole RL history of many, many years, decades, actually, the engineer was in the driver's seat.
25:16And you have a value function and you design the system based on a time slice that captures the history and the assumptions about the future and so on. And then you use that in order to define an agent yourself. What this is transitioning to now is a situation where the system is a neural net and you leave that damn neural net alone and just define it as an agent and then expose it to various tasks that are agentic in nature. So that are sequentially, you have to do this and this and this and so on and so forth. So that the cognitive skills in referring back to things you have done before and understanding words like first, second, third and so on and so forth.
25:53understanding that there are actions you can choose from and so on, that this emerges in response to the tasks you train it on. And so now transfer learning can play out. And that's this domain and the domain of making a better agent. So from that point of view, I think there's just nothing other than an autoregressive model that will lead us to better AI because you're just missing out on the agent aspect. Humans are agents and our language is deeply informed by humans being agentic. We know about past and future and first, second, third, and this and then that, and so on and so forth. So I think you could claim it's hard to find a significant number of words in the English or in other languages that are not in some way infused with semantics that has to do with gentleness of that human that uses those words.
26:43None of those words are represented in the same way in a GPT-free term language model yet fully because they are not trained in agentic scenarios fully yet, etc., etc. But since it's an autoregressive model, it has the bucket, the potential to absorb this meaning that goes beyond what it sucks up from Wikipedia and web language and so on. And so that's what's exciting about this. So when you think about kind of crafting a research program to explore these kinds of issues, to understand what reasoning means, because that's one of your main focuses that Qualcomm on the research team there, like, you know, talk about how that plays out some of the individual works that you're exploring to kind of get at these questions around what is reasoning and how do you create reasoning capable agents?
27:38systems, agentic systems with, you know, these building blocks that we're left with, LLM for currents, whatever, these autoregressive models? So first of all, I think it's reasonable to proclaim that reasoning exists because of language. That's one distinguishing factor that distinguishes human intelligence from other animals for large part. Not everybody would agree with that, but I think language has to play a key role in that. So when we talk about reasoning, we foremost talk about language. And that incidentally is also the way in which reasoning has made a huge advance in the last few months, actually.
28:15It's because language has entered the realm, right? So it's like, arguably, it's been language had something to do with it. There is the longstanding theory in psychology and cognitive science of dual process scenario or architecture, whatever, that the human brain adopts, which is, there's a system one and a system two. Dan Kahneman, Nobel laureate, has been a very, very outspoken proponent of that view, has written a very famous book, Thinking Fast and Slow, about this and so on. And so that viewpoint says that there is a fairly clearly, a somewhat clearly distinct set of cognitive capabilities, which can be called system one and system two, where system two is the stuff that is serial and slow and deliberate, very human-like, actually, because of that.
Read the full transcript
29:01And then there's the system one stuff, which is fast, reactive, often pattern recognition, often very perceptual, like recognizing in a matter of milliseconds, recognizing as a predator, whatever it is. And that there is this interplay between these kind of, it's not definitely not modules. That's the way in which this can mislead a lot of people, but it's sets of, let's say, clusters of capabilities that humans have. And it's that interplay between the system two and system one that is both fascinating and probably holds the key, at least many believe, to really advancing AI and making it more human-like.
29:37And so system two to many, and so to go back to that original point, is somehow intimately linked to language capabilities. And when we study language, then we are automatically drawn to problems to do with solving long-range serial abstract type problems that only humans can solve. And so that has made a lot of inroads recently because language models, just normal pre-trained language models work very well in tasks that involve mathematical reasoning, abstract reasoning of various sorts and so on. Now, what's interesting to note is that a lot of that reasoning capability is, well, first of all, it's interesting to see that it comes kind of for free.
30:20So surprising to everybody, once you master language really, really well, you're really, really good at doing various kind of mathematical reasoning problems and so on. But it's also interesting to note that a lot of stuff is still missing in these kind of systems. So one playground that is popular now is code generation and code execution. It's mostly code generation, right? The models are asked to generate Python programs that do various things and so on. But it's interesting to note that they arguably fundamentally must be missing some conceptual aspects of the code that they write that completely deprives them of getting creative and good programmers at the end of the day.
30:58And I think a lot is still needed to sort this out. So here's a very simple example. When you are aspiring early computer programmer and you learn about variables and this kind of stuff, you adopt various kind of metaphors about what the heck you're doing in your code. Like, for example, one of the metaphors you adopt is variables are containers. They're like boxes that you can put stuff into. So you can say A equals 2, and then the container A holds the value 2. and you can also put, say, A equals 3 and then it holds another value and so on. And you learn various things that are building upon this metaphor.
31:33Like, for example, if you want to swap the values into variables, like if you want to set A to what B was and B to what A was, you cannot really do this except Python has some shortcuts for that. But normally what you have to do is introduce a third variable C and then set C to A and then set A to B and then set B to C. So you need to basically take this thing, move it here, So you can exchange these two values, which is something you also have to do in space if you want to move two objects around and exchange their position. One has to move out of the way, the other one can go there. So there's some very deep common sense grounding in this stupid, simple concept that I'm just bringing up randomly here of swapping the values of two variables.
32:15Now, arguably, if you look at a GPT pre-trained language model that is good at code generation, it's lacking, I would at least argue, some of that metaphorical underpinning, that idea of a bucket or the idea of some spatial rearrangement of sorts and these kind of things. And it doesn't hurt in a lot of these test problems and so on, because it can be a pretty dead metaphor. You just know that here you have to switch variables because in this type of algorithm, that's what you do in this point in this kind of stuff. But when you then want to build larger systems that rely on similar ideas where you want to shuffle things around and this kind of stuff, my assumption at least is that you're just never going to get there unless the system actually has some kind of imagination of what's going on as it is doing these kind of operations.
33:05So that is actually a very, very difficult and long-range problem. How do you instill this metaphoric analogy type pieces of knowledge in a model that is supposed to generate code and this kind of stuff? And so one path to it, I believe it's the only path, is to link this language model with visual or other perceptual input. Visual is just very convenient. That's why we've all been on visual input for many, many years. So that the stuff that it is talking about and reasoning about is actually grounded in some kind of environment, some kind of reality. It's similar to if it's an agent, it can learn about procedure.
33:47It can learn what the word first means. If it's an agent that is dealing with spatial configurations, say of objects, just to begin with, it can learn about why swapping of variables requires another area that you can move something temporarily to. Don't get me wrong. You could ask a language and LLM to explain the whole bucket aspect of variables. And it would do a very good job at explaining it to you. But that doesn't mean that it would actually understand the notion of some spatial proximity and whatever you have in your head as you're solving these kind of problems. But do you think we can get there with a visual grounding approach?
34:30I think we will get there with a visual and agent type grounding approach, right? So if it's an agent that has to do things actively in some kind of environment, I think we can get a lot of additional signal that informs embeddings. So when people say grounding, they often think of like when the language model says dog, it doesn't have a visual understanding of what it's talking about. So it says dog, but it has no clue what it means. And then that word gets informed about what a dog looks like if we have language models that can also division and this kind of stuff. But this only goes so far.
35:04I think grounding goes significantly further than that. And so some of the things that we are studying in that realm are things like visual reasoning problems where you use language in order to solve them. So you have arrangements of objects in a scene and many data sets. Something, something incidentally had a similar flavor back then, the data set that we discussed before. It's a difficult, finicky task where you have to pay a lot of attention to what's going on. You have to sometimes track objects. You have to sometimes know that an object is hidden. It's still there, even though it's hidden somewhere.
35:35And if it's maybe moving, you know, there's some kind of inertia. It's likely to continue and so on and so forth. So you have to do various finicky things, but they all have to be linked to language. and then you ask questions about what would happen in the scene if this object wasn't there, for example. And then you have to say, well, then it would continue and hit this other object or whatever the domain problem is. And so that's one of the things. Yeah, I just wanted to kind of, well, I guess, ground the conversation in a specific work of research. It sounds like what you're describing is the look, remember, and reason paper or the setting for that paper.
36:10That's right. Though it spills over into many other problems, but these are all very related. They're broad problems, right? Everything is about grounding and understanding what language, the role language can play in building systems that are smart. Yeah, but that's right. The scenario I just talked about was very seriously and fundamentally studied in that paper, LRR. Like what's the setting? How is that problem modeled in that paper? So in that paper, we are using a frozen language model to have the ability to have chains of thought, period. Then we have an adapter that can take a visual input coming from an off-the-shelf, run-of-the-mill vision model, possibly also pre-trained, to project onto the embeddings of that language model so that those embeddings, specifically, they can go across layers, but specifically in higher layers, can have some information about something that is visual.
37:06and that is outside of the model itself and so on. So when the model tests things like left or right or so on and so forth, it's going to have a sense of what that means in this environment that it's training in. So it's basically architecturally super simple. And I'm actually a very big fan of very, very simple architecture. I don't think AI will be solved through architectures, but through smart ways of using data and so on. So it's a very simple architecture. Nothing special about it. Take a language model, take a visual model, adapt them and let them be connected. but with some constraints that have to be present.
37:39So you have to be careful in setting this up. For example, the model has to be able to attend to the visual input in a top-down fashion. That means that if you ask a question about, say, the video that it's going to watch and then let it watch the video, then that model needs to be able to pay close attention to what's going on in the video as a function of the question that you posted. So in many cases, when you talk about vision language models and so on, And it's like there's no such top-down effect that takes place where the very perceptual processing, visual processing in the case of the input, has to be modulated by what you're planning to do with the visual input.
38:20So if you're counting events, then you have to pay attention to one thing. If you're counting objects, to another thing. And if you're saying what would happen if, then it's to a third scene. So there's top-down. There has to be top-down influences between this, so it cannot go just one way. So the embeddings have to be influenced by the visual input, but they also have to be able to determine how you're going to process that visual input. And so a bunch of these things have to be carefully designed so that that very simple system at the end of the day then is able to misuse, if you want to call it that, or use its pre-trained language ability to do reasoning in those visual inference scenarios.
38:57And so a specific aspect there that we explored in that paper is to let the model generate rationales. There is a concept called chain of thought prompting in abstract reasoning where people discovered it's much better to let the model describe in great detail the chain of thought leading to a solution rather than letting it output the solution. So we can do the equivalent thing now in this visual domain where the model can ask to infer certain aspects about that, say, video that we believe is going to be useful down the road to be able to solve this particular task. So we are able to instill cognitive capabilities by phrasing them as language that we think are constituents in solving difficult reasoning problems, like, tell me what would happen if this thing was over there, etc., etc.
39:44And so it goes back to the all the way, right? It's about using language as some kind of tool so we can broaden the conceptual richness that the model is exposed to and then just use supervised learning on a surrogate task, essentially, on some kind of auxiliary task. So we instill the skills that you need to solve a problem. But what's different to 2013 is it's agentic. To some degree, it's video. It's processing in real time. And it solves very hard problems, arguably hard problems. although they are mostly synthetic type problems that are studied in this kind of community and this kind of domain.
40:23In the context of that paper, how do you benchmark the work? Like what are you? Yeah, it's usual. There's accuracy measures of various types of various tasks. And they usually go 92 % versus 95%. And we do very, very well on a lot of those tasks, despite having this very, very simple architecture, which is basically what it is. Are you benchmarking? It sounds like you're using common benchmarks. Are they benchmarks that specifically try to get at examples that confound non-visually trained models? Or is the idea that this visual grounding, by using conventional benchmarks, this visual grounding, if it produces an uplift on the conventional benchmark, it kind of validates the assumption that the visual grounding is generally useful.
41:17They're essentially impossible to solve without combining vision and some kind of reasoning machinery over the visual input, right? What's different now, I think, but this is in which way this is aligned with current trends as well as that it's end-to-end. So we are trying to just train the system end-to-end to be able to do these kind of things rather than building a system on a design board from components and figuring out, do you need to do object detection? Once you know where the objects are, you need to do this and this. Like, no, no, it's none of that, right? There's an end-to-end neural net just gets pixels.
41:49And then there's a language model that can answer questions about and have a conversation, essentially, about what's going on in the video stream. They are probing various reasoning capabilities. One of the benchmarks that we're studying there is called ACA, which is actually a task derived from a task that is administered to children sometimes to check various cognitive capabilities at various ages. And that tests things like causality and confounding factors when you draw inferences about properties of objects and this kind of stuff. and some others are much simpler. For example, there's one, which is something that incidentally we also studied in a real world scenario before, which is essentially the shell game.
42:29So you have a bunch of cups, let's say, transparent cups. There's a marble below one of them and then you move them around and you have to in the end say, where did the marble go? Under which cup is the marble? That benchmark that has this setup in a synthetic way is called CATR, C-A-T-E-R. and what it comes down to on the AI, the cognitive capabilities level is you have to be able to do tracking anonymously because in that particular task, it's even harder. It's like you have to, first of all, you need to remember where that marble is. You need to make sure that you follow the one of the multiple cups that has it, or in this case, it's some kind of cylinders and stuff.
43:08Then the cylinders can be occluded by other objects and then the whole thing can move around. And so you have to basically make sure you pay attention to what's going on in the scene and know about object permanence, basically, and these sort of things to be able to answer it. And that end-to-end model also does very well on that. I mentioned we studied this before. This was actually some exploration that we did in the company at 20BN a long time ago, where you would do this shell game thing. So playing with this marble under the cups, this is an avatar that is looking at you as one of the things that the avatar is supposed to do.
43:43So it's supposed to play this game with you. And then you are the one who moves around those objects. And then the avatar has to prove what it can do by basically saying, okay, I think the marble is over here. These kind of things. Those problems go beyond the types of abilities that classification tasks and so on, captioning tasks in many cases get at. They get more at long-range capabilities, having a certain degree of memory of what happened before what and these kind of things. that way they get much much better at questions around common sense so those problems beyond abstract reasoning that you have in math math word problems and so on are much more closely linked to what humans typically call common sense and that is those things like object prominence and gravity you know there's gravity you know if i let loose then it's going to fall that way and not that way.
44:38So there's this amazingly rich and vast set of properties and facts about the world that you have to carry along with yourself that helps you in many scenarios be really, really good at solving these kind of problems. And AI systems don't have it. And once language and perception grow together appropriately, that ability will be there. And that is a big step up for AI. So another recent paper that you worked on is in the domain of situated chats. This was a paper at the CVPR Embodied AI Workshop. Can you talk a little bit about that one and what it's trying to do? That paper is essentially a logical continuation of the work that we did at 20BN towards the fitness coaching AI system.
45:28And what's different in this system is that we are now using large language models rather than LSDMs and whatnot that are pre-trained also to phrase what kind of improvements you can make, how what you're doing relates to what you've been doing in the past and so on. So now it's much less rigid and much more open-ended and much better simply at generating language. And another change is that in the past, we were always on this mission that we want to have an end-to-end system where everything is happening end-to-end, but at that time, text-to-speech, for example, wasn't quite there yet. And we knew at some point text-to-speech will just be there because neural nets are advancing fast and so on.
46:14But until then, we're just going to have to use whatever we can in order to let that avatar say things. And so in that case, we actually wrote lines and had them acted out and so on for that avatar so that it can say certain things to the user in the right moment. What's going on in this new work now is that the language model decides on the phrasing of certain things that it wants to tell to the user. And a 2023 text-to-speech model is used, turned that into audio to actually issue that, emit that speech back to the user. And so that makes the whole development cycle much, much easier for us and faster.
46:51And it makes the whole system behave much more natural than before. There is still a ways to go, by the way. So right now there's still a state machine, unfortunately in the model, in the whole system that decides based upon the visual input, what has to happen when and then what type of things have to be said to the user and in what kind of scenario. Now we are working towards making that end-to-end by syncing that back into the language model so that the language model becomes the system that is really in the driver's seat and decides everything and behaves the way it should be. So it's not just the communication part, but it's also the thinking part in this whole system.
47:34Okay. You've also got a paper that explores the way that drawing on a canvas type of environment can act as grounding signal for LLMs. How does that fit into, how does that complement the previous couple of papers that we talked about? So, well, there hasn't been a lot of buzz lately around image generation, thanks to... Stable diffusion and like. ...that kicked off with diffusion models, previously DALI one, and so on. Yeah. So here we're taking a completely different route towards generating images in a very, very constrained, simple scenario. It's not so much at the moment about generating beautiful images, but studying how we can generate images by having an AI model that is agentic.
48:23again, right? So we want to have an end-to-end system that is autoregressive and hence it is an agent and that can take action on a pen if you want, or maybe some other types of drawing utensils and then use those and learn to use those in order to create some kind of imagery. In this case it's drawing, so we're just studying this in the context of line drawings and the model has the ability to draw lines and do this for long enough until the object is appearing that you want to appear. The model, incidentally, is very similar, almost identical to the Look, Remember, Reason model where a pre-trained language model has an adapter that allows it to ingest visual information and there can be this back and forth between the visual information and the language capability.
49:16So it's architecturally the same, also by design, because, again, I don't think we need to develop a fancy architecture, but rather use the right overall framework and then fill it with live through training data. What's interesting about this, though, is that as the model draws, since it has visual input, it can learn to make use of what it sees at the moment, right? So the model doesn't have to be ballistic and learn to generate sequences, which a principal language model could, sequences of lines on paper in order to replicate an image, but it actually sees the kind of stuff that it produces.
49:49and as a result of that you can ask it to complete drawings that are started or to erase parts and to basically now have a playground for studying agentic behavior right it can do revisions to what it has done before and all of this is happening in the context of a model that's exposed to an actual environment which is in this case just the drawing environment is it strictly self-supervised or is there a supervision loop? There's supervision in everything. So in this case, we were able to even embark on this because there are various data sets out there, like Quickdraw, for example, that have ground truth strokes associated with images.
50:34So people were drawing things and then you can use those ground truth line strokes as a training signal to just say, okay, if I show you this image, generate this sequence of strokes. If I show you that image and I ask you to replicate it and generate that. If I show you this partial image and I ask you to complete it, then again, we have the ground truth to train the model. And so that was the starting point so that we can fine tune the pre-trained language model on the ability to misuse the outputs it generates to talk about things that it does, drawings. In that case, it's not an extra head.
51:08It's basically just using tokens. And then that is now the starting point to fine tune it this reinforcement learning where the model can get asked to draw something that looks like a fish and then it can make some attempts maybe with high temperature to do it and then a classifier can say how good is what you just tried to do how much does it resemble a fish or whatever and then that can be a learning signal that then allows us to find you that model this rl interesting interesting we talked a little bit in the context of your transitory comment about kind of where you think this is all going, maybe that's a long-term view plot.
51:48I guess I can ask you, what's your viewpoint of the term or the cycle time here? I'm just curious how you play out in time the evolution of AI reasoning. In terms of predicting when what will happen? Yeah, what's next? What are the timeframes? What's your assessment of complexity? like you know i don't want to ask you to pull out a crystal ball but also i do i know i uh got a bloody nose there not too long ago when i was made a bet a few years ago that by i think 2021 or something we're going to have a an avatar video chat like you and i have right now where based upon what the person is saying with realistic rendering through gans and whatever was the rage back then, maybe diffusion models, now whatever it is, you wouldn't be able to tell whether you're talking to a person or an AI generator.
52:47That hasn't happened yet. I still think it's just two years out, but it has been two years out for a while. I think it's not completely unreasonable to assume that we're going to get to that point, but we're not quite there yet. Other than that... So take your predictions to follow with an appropriate size grain of salt. That's right. Yeah. Predictions never really work, though. As some people pointed out, Jeff Hinton, amongst others, either we drastically underestimate or we drastically overestimate. for some reason. I don't know why that is, but it seems to be a constant. Yeah, but I think there's a fairly obvious progression going on right now where language is getting significantly good if we can get recurrence back, which is to be seen.
53:40But if we can get language, we'll keep language, but get recurrence back and combine in the best possible way the ability to process language with perceptual input and taking actions, so building proper agents. To phrase it very carefully, we're going to make significant progress in AI over the next few years. And so one aspect of this is recurrence. You've spoken in length in this conversation about that. Are there other missing building blocks that you think are needed? So, yeah, recurrence goes hand in hand with memory. So that is missing in the same way. There is a total mystery when it comes to memory that may need to be solved, which is something that some people call fast weights, where learning can happen at many different timescales.
54:33I can tell you a fact, like I was born in Germany, and then you're going to keep that. And if I meet you in three weeks, you might still have that piece of information. Like, hey, Sam, where was I born? You're just going to say in Germany. Somehow, right? So you can do fact learning instantaneously just like that. and then there is another kind of learning maybe another kind we don't know what the relationship is but there's another kind of learning which is what we typically do also with neural nets which is synaptic slow synaptic adaptation being exposed to many examples of a task and then slowly adapting the weights in the network so you get better at solving the task and what is the relationship in these two right so does that why can you in three weeks presumably still know that I was born in Germany, if not through some kind of synaptic adaptation?
55:22Or can you maybe? If not, then what has to happen for that synaptic adaptation to take its effect, given that I just told you that thing, and you probably don't even think about this anymore for the next 20 days until I come back and ask you or something like that. So what's going on between these two? At least these two points on maybe a spectrum or maybe qualitatively different types of learning or so. So that's still a bit of a mystery, and then maybe it has to be resolved. Maybe without resolving that, some significant progress in AI is not going to happen. I don't know. Nobody really knows for sure.
55:56But that memory, and it has to do with recurrence in some way, I would assume recurrence in general has to do with memory and so on. So that's clearly one missing piece. Then other than that, I think there is a fairly straightforward linear path. It's going to require a lot of work, but there's a linear paths and improving grounding seriously. So not the industry term grounding that you know something about database entries and stuff, but really infusing the embeddings or let's say brain states or feature vectors as you use language with sensory information and information about yourself. And that goes very, very far.
56:33I think, and I think it's the key missing piece. So it goes as far as potentially having an agent have to develop a sense of the word I, which is also a beautiful and longstanding mystery in AI, which is how can the word I emerge for humans? And what exactly does it mean to an agent to say I, letter I, not the I? So I have a sense of self. And so that may be fundamentally needed in order for the system to solve a variety of problems, because it needs to know that it did that action. It's on itself. How does it interact with other agents and all of these kind of things? And it's a total mystery, a mystery dating at least back to the Buddhist practices, the emergence of Buddhist practices several thousand years ago.
57:25And it's going to be really fascinating to see what we're going to learn about the emergence of the concept AI as we're trying to instill that in AI models going forward. I feel like you're opening the door to the second hour of our conversation, but I'm not going to take the bait. Although I really want to. We're going to have to save that for our fourth interview. That's right. I look forward to that. But Roland, it's been really wonderful to reconnect and to learn about some of the ways that you're exploring neural reasoning. Thank you so much. It was great to be here again. Thank you. All right, everyone, that's our show for today.
58:07To learn more about today's guest or the topics mentioned in this interview, visit twimbleai.com. Of course, if you like what you hear on the podcast, please subscribe, rate and review the show on your favorite podcatcher. thanks so much for listening and catch you next time
From the publisher
Today we’re joined by Roland Memisevic, a senior director at Qualcomm AI Research. In our conversation with Roland, we discuss the significance of language in humanlike AI systems and the advantages and limitations of autoregressive models like Transformers in building them. We cover the current and future role of recurrence in LLM reasoning and the significance of improving grounding in AI—including the potential of developing a sense of self in agents. Along the way, we discuss Fitness Ally, a fitness coach trained on a visually grounded large language model, which has served as a platform for Roland’s research into neural reasoning, as well as recent research that explores topics like visual grounding for large language models and state-augmented architectures for AI agents.
The complete show notes for this episode can be found at twimlai.com/go/646.




