In short
TWIML AI Podcast Episode #686: Language Understanding and LLMs with Christopher Manning
Podcast Overview
- Title: The TWIML AI Podcast
- Host: Sam Charrington
- Guest: Christopher Manning, Professor of Machine Learning at Stanford University
- Focus: The impact of machine learning and artificial intelligence on business and society, exploring significant advancements in natural language processing (NLP) and large language models (LLMs).
Episode Summary In this episode, Christopher Manning shares his extensive insights into natural language processing and large language models, discussing his foundational contributions to the field, the intersection of linguistics and AI, and the future of NLP technologies.
Key Themes Discussed
- Foundational Work in NLP
- Manning's role in developing word embeddings (GloVe) and attention mechanisms.
- Contributions to multimodal machine learning and the evolution of LLMs.
- Surprise in Advancements
- The rapid advancements in LLMs over the last decade surprised even long-term field researchers.
- Although foundational work set the stage, the emergence of powerful tools was unexpected.
- Linguistics and LLMs
- The tension between traditional linguistic theories (Chomsky) and the capabilities of LLMs to learn language structures from data.
- LLMs as evidence that language structure can be learned, contrasting with Chomsky's perspective on innate language acquisition.
- The Nature of Intelligence in LLMs
- Discussion on the distinction between narrow AI and general AI.
- Manning's view that LLMs represent a significant step towards general intelligence but don't yet exhibit true reasoning capabilities.
- Challenges and Opportunities Ahead
- Importance of addressing current limitations in reasoning and knowledge retention in LLMs.
- Future research directions include exploring alternative architectures and better world models to enhance understanding and interaction in AI systems.
Key Takeaways
- Word Embeddings & Attention:
- Word embeddings effectively capture semantic meanings in high-dimensional spaces, aiding the development of NLP models.
- Attention mechanisms revolutionized how neural networks handle sequential data, leading to the development of transformers.
- LLM Capabilities:
- While LLMs can generate human-like text and perform various tasks, their reasoning abilities are still limited and rely heavily on pattern recognition.
- The challenge lies in developing models that can understand and organize knowledge cohesively.
- Future of NLP:
- There's a need for models that can adaptively learn and reason in real-world contexts, moving beyond what current LLMs can achieve.
- Exploring different learning paradigms, such as situated AI and embodied AI, could enhance how machines understand and process language.
- Research Directions:
- Manning is focusing on improving the interaction between language models and world knowledge, aiming for systems that can update and reason about facts seamlessly.
- The exploration of new architectural frameworks to enhance learning efficiency and contextual understanding in language models is a promising area of research.
Conclusion This episode highlights Christopher Manning's influential work in NLP, presenting a thoughtful exploration of the relationship between linguistics, machine learning, and the future of AI. It underscores the ongoing challenges in achieving true intelligence in machines and the potential pathways forward in the field.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Looking behind the veil of language and these questions of reasoning and intelligence and how humans store their knowledge of the world. We have such good language understanding and generation that a lot of what we're missing is then the stuff behind that. And that's going to make all of the difference in giving us intelligent machines, but also intelligent language users.
0:37All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm excited to be joined by Christopher Manning. Chris is Professor of Machine Learning in the Departments of Linguistics and Computer Science at Stanford University, Director of the Stanford AI Laboratory, Founder of the Stanford NLP Group, and an Associate Director of the Stanford Institute for Human-Centered Artificial Intelligence, or HAI. Chris is also the 2024 IEEE John von Neumann Medal recipient, which recognized him for advances in computational representation and analysis of natural language.
1:16Chris, congrats on the recent award and welcome to the podcast. Thanks a lot, Sam. And it's great to be on the show with you. I am looking forward to digging into our conversation. We will have no shortage of things to talk about. You have been involved in a lot of the foundational research that has laid the stage for LLMs and Gen AI. And in fact, I think I want to start us by zooming out and talking about just that. You know, we can circle back to some of these things individually, but you worked on the glove paper, which kind of laid the foundation for us to better understand and use word embeddings.
1:57You worked on some of the early application of attention to NLP. Your work on joint text and image contrastive learning has kind of set the stage for multimodal ML models. I'm really curious to hear, has it surprised you the way all of this has come together to create tools as powerful as LLMs and multimodal models? Or at some point, you know, was it clear what would become possible and it was just a matter of putting the right pieces in place? I think it's absolutely surprising. I think for anyone who's been working in the field for more than a decade, I mean, I guess for me, it's about 30 years.
2:40I mean, it's just astounding, this upward trajectory that we've had. Things started to take off about a decade ago. but really in the last five years, things have just zoomed ahead. And well, you know, you can trace back the predecessors. There's sort of a path that once things started happening with large language models, you could see that there are possibilities from scaling further. But, you know, nevertheless, just the way things have emerged with these abilities of large language models, it's just really very surprising and came in an unexpected direction. And to what degree did your background as a linguist impact positively or negatively your ability to see and contribute in so many ways?
3:25Well, in general, I think me being a linguist in a field where, I don't know, 97 % of people aren't linguists, that they're machine learning people or sometimes physicists or other more general backgrounds, that that's given me a really distinctive viewpoint. And it has allowed me to contribute in complementary ways and showing the role of language. But on the other hand, linguistic knowledge is a good, broad background. I mean, in terms of what's led to these extraordinary advances, there's no doubt at all that the main engine has been in the machine learning and the math. It hasn't really been coming straight from linguistics.
4:10I'm imagining that as a linguist, there is a degree or has been a degree of tension. Maybe, you know, we're long past that, but at some point there was a degree of tension and kind of not wanting to let go of traditional linguistic approaches and needing to embrace statistics. Like, I'm curious, like how that, you know, appeared for you and to what degree you had to grapple with that, what that really meant. I mean, it's a great question and it's absolutely a live issue, very much so to this day. So academic linguistics, especially in the United States, but in general worldwide, in the second half of the 20th century leading into the 21st century, the dominant figure was Noam Chomsky.
4:58And Noam Chomsky is really getting on at this point, but, you know, he's still active. And in his recent pieces about language, Noam Chomsky is loudly proclaiming that large language models are of no interest in linguistics and tell us nothing about human language. Many of Chomskyan linguists who actually dominate linguistics departments certainly in the United States, you know, that has been the picture. That was never my picture. So although I definitely started off as a linguist, you know, how my path shaped was I thought there had to be a lot of learning involved in human language acquisition.
5:43And I was interested in how that could happen. And that led me to start then looking at machine learning and things built from there. And so, you know, there's other more empirically minded linguists who see really good points of connection between large language models and what happens in linguistics. I shouldn't go too deep a riff on the history of linguistics, but there's actually an interesting history here because, you know, way back in the sort of 30s, 40s and 50s, that dominant strand of linguistics was the American structuralists. And really what they believed in was that if you had a body of text in a language, they were thinking of something like a Native American language where they'd collected stories and so on.
6:31What you should be trying to do is define an inductive procedure so you could learn the structure of the language from the texts. Now, you know, in the 1940s and 1950s, basically nothing was known about machine learning or other related fields like information theory, right? Shannon invented information theory in 1951. So, you know, they didn't actually make much progress in this goal, but their goal was to be able to come up with a linguistic structure discovery procedure where you could look through the text of some language and start to learn the structure of sentences. And, you know, really the first thing that made Chomsky famous was trying to argue that no such linguistic structure discovery procedure was possible.
7:17So, you know, despite the fact that they didn't have much in the way of methods, I really think of it as, no, these people before Chomsky were right, that that is the right kind of goal. And we're now actually seeing it happen in the age of large language models. Hmm. And so you mentioned that the Chomskyian viewpoint is that LLMs don't have anything to teach us about language. You know, in what ways do you see LLMs teaching us about language? Yeah. So, I mean, it is important to point out that, you know, human language acquisition, babies starting to learn language is obviously nothing like training a large language model.
7:58And so, you know, there's a lot of difference there, and I don't want to in any way deny it. But the part where it's interesting is that Chomskyian linguistics has been founded on the notion that you can't possibly learn the structure of a human language from the observed evidence. And that leads into his theories of the innateness of human language. And so there's very little learning involved where precisely what large language models show is that actually you can learn the structure of a human language. And so this is something that I've worked on once large language models start coming to the fore around 2018, 2020.
8:40You know, with my linguist hat on, I was not interested just in the fact that these large language models could generate sentences. are as interested in, well, what do they know about the structure of English or the structure of French or the structure of Chinese? And actually, you can poke around the representations of these models and find out, oh, yeah, they know about subjects and objects and predicates and relative clauses, and that they actually have learned to decode the structure of English sentences. And that's part of why they're so good at generating fluent English or whatever language, as we sort of see in the output of ChatGPT or similar models.
9:21So do you see LLMs as, you know, somewhat of an existence proof that this learning does happen? And it sounds like the Chumpkins still resist this. Like, what do they say? Right. Yeah. So it's an existence proof that you can learn the structure of human languages from a ton of data. Not necessarily that we do. Yeah, from a ton of data. So the kind of arguments in the reverse direction that would be made by Chomsky or similar people, well, firstly, the amount of data that large language models use is completely unrealistic for human language acquisition. So humans become good language learners on, you know, somewhere around 50 or 100 million words of data, whereas large language models being trained on a minimum of billions of words of data and the very largest language models now are being trained on a trillion or more words of data.
10:15So there's a huge order of magnitude difference there. But perhaps more profoundly, what Chomsky would like to argue is that these models are much too general, that large language models or similar neural network models kind of suck the structure out of anything. And so that they can as easily learn patterns in DNA or patterns in chemical formulas as they can learn human languages. And so Chomsky thinks that's profoundly wrong because there's a lot of commonality in the structure of human languages. And so he's always wanted to seek a very sort of restrictive model which can describe only what's found in human languages.
11:03Interesting. Interesting. I've always found it super interesting, the kind of the two way street between the biological side or the human side and the computer science side in particular, like the neuroscience side. Neural networks are, you know, quote unquote, inspired by neuroscience. And that kind of pushes the CS side forward. But then the neuroscience folks take what's done on the computer science side and run with it and push our understanding of the neuroscience forward. And it sounds like there's an opportunity, you know, here on a language from a language perspective as well. We've talked a little bit about, you know, Chomsky having his opinion or Chomskyans having their opinion on, you know, the place of an LLM, like what research has to happen from a linguistics perspective going forward to build on, you know, this existence of LLMs and learn more or teach us more about language from a human perspective?
12:05So to start learning more from a human perspective, we need to start exploring models that are much closer to the human context of language acquisition. But, I mean, yeah, I think there is this really useful two sides that can beat off each other that you were just talking about. Because, you know, if you want something closer to human language acquisition, well, human language acquisition is situated that, you know, you're in an environment with things in the environment they're being talked about. And so, well, you need a multimodal foundation model. And well, multimodal foundation models are just the kind of thing that people are starting to work on now.
12:49But even then, most of those models you're saying, yeah, here's a big pile of images and text about them, where it's very clear that it's central to human language acquisition, that it's actually interactional, that you're in an environment and there are people talking to each other, pointing at things, talking about things, looking at each other, and that interaction is highly key to human language acquisition, right? There have been some sort of unfortunate natural experiments where people who can't afford childcare think that maybe if they just leave the TV on all day at home, their kid will learn language the same way as a person with a caregiver.
13:33And that's just not the case, right? The interaction is all important. Which brings to mind other areas of contemporary research around situated AI, embodied AI, themes along those lines. It sounds like you see those as being critical. Yeah, I think they're super important areas to start investigating to kind of get more connection and relevance from the human learning context. Yeah. Did you come at your research and your career broadly trying to create intelligence or were your, you know, steps and interests, you know, different or more discreet or something else? Like, I'm curious, like this thread of intelligence, you know, through your research, is it, you know, very directed or organic or, you know, how do you think about intelligence as a goal?
14:23Yeah. I mean, I think it emerged as definitely not where I began. Where I began was human language seems fascinating. Humans can do amazing things and learning and understanding each other's speaking language. How could we use computers as a way to understand and model that? So I was sort of very focused on language and really that's been the bulk of my career. So it's really as time went on And especially in the neural networks era, when there started to be these very general methods, neural networks, where the same kind of models and methods can be applied to vision, robotics, language, that I slipped into being an artificial intelligence researcher.
15:10and then it's with the enormous success of large language models leading to things like chat gpt and other models obviously claw gemini etc that then everyone is much more concerned about what is intelligence and are these things intelligence so so i sort of have sort of slipped into being concerned with intelligence but it wasn't really what my long-term goal was originally Yeah, that's what I was really curious about, the degree to which that is a primary concern and motivator for you, or is it ancillary to other things? And given that you've kind of slipped into it, do you see where we are?
15:56What do you see as the relationship between LLMs and intelligence? Do you see that we've created some degree of scaled-down intelligence? Do you see it as a small stepping stone? Do you see it as a dead end to what people really want in terms of intelligence? Like, how do you make sense of LLMs and current AI in the context of intelligence? Right. I think it's definitely an important stepping stone and a quite dramatic development, which has given us something very different. So let's say the positive part first. Right. So in general, in AI, there's this longstanding distinction between narrow intelligence and general intelligence.
16:41And everything that was done in machine learning and AI prior to large language models could only be described as building artificial narrow intelligence. because typically what you're doing was you wanted to build a system, whether it was something to recommend movies or to recognize birds and photos or to say whether a piece of text was toxic. Whatever you're doing, you got collected data for that task. You trained your model. You had this thing that could make a decision in some domain, but it had no intelligence whatsoever beyond that. Whereas the goal that people had dreamt of for the entire history of AI was to have an artificial general intelligence, like a human being who can do all sorts of things.
17:31And so in precisely that sense of the definition, I think we've got one, right? That in the age of large language models like ChatGPT, we now have this fairly general intelligence. You can ask it to write a poem. You can ask it to translate this piece of text into Chinese. You can ask it to summarize this long, boring report, you know, get recipe hints or hints on where to go visit in Vienna. Right. You know, it's a very general intelligence. At this point, I have to go on to the caution side, which is, you know, people use the term artificial general intelligence these days in two senses. One is that sort of historical technical sense, but in common usage, it nowadays more often means this is something that's so intelligent, it's as good as humans or better than humans in most respects.
18:27And at that point, I think we need to be really cautious. And I think a lot of people are fooled as to how much intelligence there is in these large language models. You know, they can do absolutely incredible things and producing beautiful text. And I realize that. But, you know, I sometimes use the analogy that a large language model isn't really much more intelligent. than a talking encyclopedia. Because, you know, most of what we're impressed by of large language models is how much stuff they know. And, you know, I think historically, we've tended to value that in human beings as well. You know, this person is really knowledgeable.
19:14But, you know, I don't actually think that's the heart of human intelligence. The heart of human intelligence is being able to adapt to new situations, to very quickly learn new things. That's the real intelligence. And that's not what large language models are doing. Large language models look amazing because they've slurped up all of this knowledge that was being hard won by humans over centuries. And it's all been shoved into the machine that they're not sort of quickly picking up and learning new things going about the world in the same way that human beings do. At the same time, we've seen the quote unquote emergence of properties like the ability to reason in sufficiently large LLMs.
20:03Um, how do you think about, you know, and then like, there's a desire to then extend those capabilities into kind of agentic systems and workflows that combine, you know, the knowledge and the ability to reason, you know, with, um, to create something that kind of can go off and do things in a more, you know, intelligent way. Like how do you parse all of that and, and the limitations that we'll end up finding there? Yeah, I actually don't think we should say at the moment that large language models can reason. They can behave in ways that it looks like they can reason. So, I mean, and again, this comes from them having just huge knowledge from all of the billions or trillions of words that they've read.
20:50So to the extent that there is a pattern of reasoning that's commonly available and they can mimic it, they can produce answers that look like they're reasoning through something. But, you know, it's also obvious in many cases when they make sort of glaring mistakes that they're not actually reasoning. They're sort of just pattern matching. And if there's a pattern of a sequence of steps and they've seen examples of it, they'll apply it to a situation and put in the terms you've used. And sometimes it works and sometimes it's glaringly wrong and they don't really know the difference. and certainly if you give it give a large language model clearer problems like planning problems where you have to sort of plan with some constraints as to you know what days someone's available and how far they have to travel and how you're going to get things to move around various people have looked at planning problems and you know the current large language models just can't do that kind of planning they just can't sort of reason with those kind of constraints so So, yeah, I mean, although there are lots of claims of large language models reasoning, I think at the moment, be cautious.
22:08But on the other hand, you know, there are other places such in playing chess and go where people have hooked up neural nets with search procedures where they really do sort of plan and evaluate positions and so on. I do think in the coming decade, it's quite likely that some of that kind of search and planning technology will be hooked up to large language models. And we really might start to see machines that can reason. And as you mentioned, for tools and agents, once you're sort of connecting those in, that'll give language models new powers to calculate and perhaps plan things using these external tools.
22:50So I think we're on the cusp of some of those things becoming possible, but there's still more to do. I laughed a bit as you started responding because I think that that view of reasoning is the one that I tend to gravitate to. And I've made very similar arguments, talked a little bit about needing to distinguish reasoning as a mechanism versus reasoning as kind of an outcome or an observable behavior. And I've talked about this in lots of interviews. I think what I tend to hear back most is that we really have very little understanding of reasoning as a mechanism in humans. So how can that be the bar?
23:31And, you know, all we can do is observe this thing and see if it looks like looks like reasoning. And from that perspective, LLMs seem to be doing some of it to some degree. Any reaction to that? you know, they seem to be doing it to some degree, and you can show examples where it absolutely looks like the language model can reason, because you'll show it some pattern, you'll ask it some question about, you know, how many kilograms of stones can 10 people carry if one person can carry 20 kilograms, and it'll say if one person can carry 20 kilograms, and then there are 10 people, we should multiply 20 kilograms by 10 people, and therefore it can carry 200 kilograms.
24:27And you think to yourself, oh yes, this totally looks like reasoning, it's understood the problem, it's done the calculation, it's explained it, it is wonderful, then you ask some different question where it should be, you know, obvious what the answer is, and you'll come up with some variant problem, which might be saying if groups of three people can carry a 60 kilogram box, and then you have nine people, how many kilograms can they carry? And it'll go through the same formula and say, okay, you've got nine people and the boxes weigh 60 kilograms, therefore the nine people can carry 540 kilograms.
25:28And it's just, oh no, that's not right at all. It's pattern matched the same kind of framework, but it just did not pay attention at all to what was there. And I think that is the current state of things. So you can show people who want to claim they can reason, can show examples that look just like it can reason. People who want to say can rightly show bloopers like that. And I think the bloopers do show that this is more kind of pattern matching of reasoning patterns. And there isn't sort of good understanding of situations as an awake human does successfully have. You're saying something like sufficiently advanced pattern matching is indistinguishable from reason and yet not reason.
26:15yeah that's sort of it right i mean on the one hand pattern recognition which was sort of once sort of dist as a trivial form of doing things that sort of has really become the centerpiece of ai that we've shown that by taking what's essentially pattern recognition technologies who are pioneered for things like speech recognition and object recognition vision, that you can really push what you can do with pattern recognition technologies. And they've become our most powerful tools in artificial intelligence. You know, a large language model is much more using the approaches of pattern recognition.
27:02But, you know, nevertheless, I think a lot of people, the wise people like myself, I think that to actually get to some of this sort of higher level cognition of planning and reasoning that we do still need to have new breakthroughs and approaches that we don't yet have. And we don't have them. Is it clear? Clear is probably a stretch. But like, do you have a sense for what stones that we need to turn over to hope to find them? Not a clear one. Yeah. So, I mean, I think the way research tends to work, right, is things are sort of stuck somewhere and then people find some really promising way that things can be pushed forward and you climb up those stairs and then it tends to flatten out until someone comes up with a new good idea.
27:57You know, so I think one of the things that people feel like we really need is much more in the way of a world model so that there's sort of a deeper understanding of the world. And to some extent, large language models seem to create a world model. But it's this really weird one since it's sort of this representation above the tokens in a sentence or a paragraph or whatever it is in the language model. It doesn't seem like it's a very actionable world model. language models just sort of generate forward and well people have done some tricks with chain of thought reasoning and things like that but it seems like really for thinking you have to be able to explore outwards and go forwards and backwards and search and so to sort of work out ways to do that kind of exploration and search and thinking we need something like that and some people are working on that now so you know there are ideas and directions that think promising but there's a difference between having good ideas as to roughly what you need and finding a really productive way to push things forward as large language models have been.
29:06Very early on in the conversation, I mentioned some of the key research innovations that you've contributed. One of those is GloVe and kind of our understanding of embeddings. And I'm wondering if you can riff a little bit on the way that you think about embeddings and the relationship between words in a vector space and both in and of itself and the way that that element has contributed to some of the pieces that we've built on top of it. Yeah. So word embeddings are these vector representations of words, so real numbers in a big vector. at least hundreds of dimensions, maybe thousands. On the surface, that seems a really weird way to represent the meaning of a word.
30:00It's certainly not what people have thought about for the rest of human history. But it proved to be a very successful way of capturing the meaning of a word and its relationships to other words, far more successful than methods that had preceded it. And it was also one of the first highly successful cases of doing this unsupervised or self-supervised learning where you're just getting a huge amount of text and saying, OK, relatively simple neural network, cogitate on the relationships of words with their words in the context. and by then doing this optimization process, you've got a word vector for the meanings of words, which just do a really good job at capturing the semantics of words.
30:48You know, early on in my NLP class, I demonstrate, you know, word vectors in the space. And I mean, it sort of actually still seems to me kind of incredible how well something that in retrospect is relatively simple manages to capture word meaning. So word vectors were really the key tool that led to the takeoff of neural network methods in natural language processing in the 2010s generation. But they also have limitations. And although there's still lots of use of word vectors in all sorts of places, for the sort of heart of NLP, we've kind of moved beyond them now. because with the modern transformer large language models, the key difference is that for word vectors, we were learning one vector for a particular word and regardless of the context in which it was used.
31:46And in reality, most words have lots of different meanings depending on the context in which they're used. So a star means something very different in a discussion of Hollywood than it does in a discussion of astronomy. me. Now, the surprising thing is word vectors can sort of cope with that. They kind of make these sort of disjunctive vectors that put all the different senses together. But when we're in a particular context, we want to know more about how to interpret the word in that context. And we can do that with the kind of contextual representations that we now compute for words in one of these large language models.
Read the full transcript
32:26In terms of those contextual representations, do you see, I'm trying to bring together the important role that embeddings still have in practical use of generative AI systems, like retrieval augmented systems, like they're still front-ended by vector search and embeddings. And yet you're saying, we're past embeddings and NLP. Like, reconcile that. for me. Okay. So it's a mix. I mean, so if you're wanting to do things with particular words and their meanings, looking for, you know, gender bias and language or doing vector retrieval, lots of good uses for word embeddings, and they're still very widely used.
33:16But for the center of, where the big advances are happening in NLP, which is with large language models. Well, you know, sometimes people starting off do initialize their transformer with word vector representations for the words. But in principle, you don't have to. You can just start training on enough text and build a huge transformer LLM. And it does. So it learns word. It does learn word vectors. So at the very bottom of a transformer, every token does have a word vector. But the meanings that you're using in your application isn't what you have at the bottom of the transformer. It's what you have at the top of the transformer.
34:04And then that's a context-specific representation of a word. And so then if you're wanting to do things like matching pieces of text for their similarity, you're generally using those top of the transformer representations to say whether the meaning of a piece of text is similar to each other. So, you know, it's a mixture. They're still widely used. But if you're thinking about, gee, what makes ChatGPT great, a lot of the games move somewhere else. Sure, sure. But I'm also wondering, are you hinting at a future where the still widely used applications of embeddings give way to pure transformer architectures like for retrieval, for example?
34:51Is embedding just kind of natural in the way that we'll probably always do these kinds of things? Or is there some analogous pure transformer architecture for retrieval that you're foreseeing? I mean, a lot of these technical questions of sort of how many resources you want to put into things. So I think there's always going to be a place for word vectors. It sounds like you're saying it's kind of an engineering problem and like how much do you want to throw at it? Right. Well, yeah. And like a kind of marginal utility kind of thing. Yeah. Cost, size models, et cetera. Yeah. Got it. Got it. Got it.
35:34We talked a little bit about GloVe and the embeddings. You can talk a little bit about attention and the way you think back on network and the future of attention as a mechanism. Sure. Yeah. So attention is the idea that in the neural networks, you're doing this calculation to find other places in your neural network where you have somewhat similar stuff. So it's a kind of content-based addressing and that you're using information from those other places to help you to decide what to do, what new representations to produce. So attention is really the big exciting new idea that appeared in modern neural networks.
36:22So a lot of the things that people started doing in the 2000s and 2010s with neural networks are actually things that have been done long before. It's just people hadn't had the computing and data scale and some technical ideas to get them to work well. So, you know, around 2013 to 16, you know, the NLP, the big model that everyone was using was LSTMs. But actually LSTMs were invented in the 1990s. And similarly, in vision, everyone was using convolutional neural nets, but convolutional neural nets go back to the 70s. So, you know, it's really sort of all old stuff that was being reinvented. Whereas attention was actually genuinely something new, this idea of doing this content-based lookup.
37:12And it really was a huge breakthrough that improved all of our NLP systems in the mid-2010s. So it was first invented for use in neural machine translation, but then it was quickly extended to be used for question answering system summarization systems. Yeah, so the first version that was invented came from the University of Montreal. But then a second simpler way of doing things, which we call bilinear attention, but much of the rest of the world called multiplicative attention, was then developed by me and students at Stanford. and it's this simpler form of attention that was then the form of attention that came to be used in transformers and well the original paper about the transformer architecture was called attention is all you need and well it's not quite true that the only thing in a transformer is attention because they also have fully connected layers that are important and they also have residual connections which have been you know developed in vision and other places and they're important.
38:18But nevertheless, the distinctive main item in a transformer was making use of this new concept of attention even more extensively than being used previously. And in what ways have attention mechanisms and research around the application of attention evolved since the initial transformer paper? So one answer is kind of not as much as you might think. I mean, because in some sense, you know, the authors of the Transformer paper, people at Google, you know, they're either really, really smart or really, really lucky. Maybe some combination of both. But, you know, they really just seem to nail a good architecture.
39:06And, you know, it's just, I think to many people, including myself, it's just been surprising the longevity of Transformers. That is, it just seemed like this put all the right pieces together roughly right. You know, there have been tiny variants since of moving the layer norm down, but very small changes. And it's still what people are using. I mean, I think there are starting to be some new ideas. I mean, as people are wanting to deal with much longer contexts, I mean, there's interest in hierarchical versions of attention. So you can have sort of tree-structured attentions heading down. And so I think we're starting to see some new ideas.
39:45There's this funny thing that for the last five years that there hasn't been the same kind of innovation in models and architectures that there was for the decade before that. People have been trying to sell the story that just keep with what we have and make it bigger and bigger and bigger and we'll solve AI, which I don't actually believe myself. So what are you most excited about? A, is it solving AI, quote unquote, but then, you know, B, you know, what is what is your your goal or, you know, kind of guiding star? And then what are the research efforts that you're finding most exciting to kind of help you get there?
40:29Sure. Yeah. So there are lots of different things that one can do. But the two things I've been most interested in lately is one, well, actually, I am still interested in neural architectures because I can't believe it'll be Transformers forever. And so I've been working with students on, you know, what are some architectural ideas to do things differently. But then in the other direction, we're in this new world of the large language model or its generalization to other forms of data, the foundation model world. And I think when people look back on things, rather than the age of neural networks being the big breakpoint, the big breakpoint will be when we enter the foundation model era, because that just gave this fundamentally new way of doing machine learning and AI systems, right?
41:20Before, it didn't matter if you were doing neural networks or you're doing support vector machines or you're doing regression trees or something else. It's the same recipe. You collected your data. You labeled your data. You trained your narrow AI system. Whereas really in this modern era that we have this profoundly different thing of you've got this large foundation model and you can just ask it what to do or you can give it three examples for in-context learning and suddenly it'll do things. right so in this very different world but because that world is so different very little of it has been explored so far so just thinking about the pieces of that world almost all of them people came up with some way that roughly worked but there are ways to think of better ways of doing the different pieces of it and for that ecosystem there are sort of lots of things that people haven't built yet and going around building them so if i stick with that bit first you know a couple of the kind of things that I've been doing with students there is, you know, well, if you have one of these humongous models, they know a lot of facts of the world, but facts of the world keep changing.
42:29Someone new is the prime minister or president, and you'd like to be able to sort of cheaply change the knowledge in the model. So how can you do edits for a few facts in the model in a way that's relatively cheap. So that's one area. And then a second area is that, you know, the difference between GPT-2 and GPT-3, which were very good at generating fluent, coherent text, and chat GPT that rocked the entire world and was appearing on national news broadcasts, came from the instruction tuning that allowed you just to tell it what to do or ask a question and get a response. And so that was initially done by reinforcement learning from human feedback, RLHF.
43:18It sort of seemed like OpenAI has this clever reinforcement learning guy, John Shulman, and he did it with PPO, proximal policy optimization. And so everyone else said, well, that's obviously the way to do that. We'll use PPO as well. And a bunch of students and And as faculty, we're thinking about this. And really, there was no reason you had to use PPO. And given the way that people are actually aligning these language models from paired, better, worse data that had been collected offline, it seemed like it was unnecessary and overly complicated and expensive in resources to do traditional reinforcement learning, which was really built for sort of an agent learning online as it wandered around the world kind of things.
44:09And so we came up with this alternative direct preference optimization, DPO, which has been an enormous hit because training it is much more stable. It has way less in the way of choices and hyper parameters that you need to fiddle to get it to work well. It's much less resource intensive, so you can run it without having huge computing resources. So at the moment, Nearly all of the sort of medium-sized open source language models are now being aligned or instruction fine-tuned using the DPO algorithm. So that was a great research success. But I think in retrospect, the biggest success isn't sort of the specifics of the DPO algorithm, but is pointing out, hey, you can do this in other ways.
44:56Because since we published DPO, you know, a whole bunch of people, different colleagues of mine at Stanford, people at DeepMind, people at Berkeley, right? There's sort of now at least five other alternatives to DPO and PPO that people have now used. And to a first approximation, there are lots of other ways that work great as well. We sort of open this jam door of showing that, you know, you can think about this problem in different ways. Yeah. And so that's the kind of sense in which we're in this sort of new world. And a lot of it hasn't been thought through very much. But if I just quickly revert back to other kind of architectural ideas.
45:41Yeah. I hope sort of prominent new alternative architectures will emerge. So one idea that I've been thinking about with students is trying to sort of have models that prefer a soft form of locality and hierarchy that sort of transformers just sort of look out everywhere for any associative signal that they can find. If there's any association that seems predictive, they'll just glom onto their attention. Whereas I think part of how humans learn quickly with little data is that they have a bias that normally things that affect each other are close together. Not always, but sort of 95 % of the time.
46:22And so we've been sort of exploring a couple of architectural ideas about how you might do that. And when you say close together, do you mean spatially in particular and, you know, drawing that analogy to humans? Yeah. Yeah. So does that imply some type of embodied setting? So that would be one way to explore that idea. To be honest, I haven't been doing that. We've been looking at it in the language case still. So close together is meaning close together in terms of what is said or what is in the text. But I mean, I do think the same idea extends out and could be done in those other contexts.
47:02But yeah, I haven't actually done it. Okay. Okay. And so continue, can you talk a little bit about how you have set up that research direction and some of the experiments or promising directions that you're seeing there in terms of leading to new architectures? Yeah. So the main thing that we've done there, which we called push down layers, and this was sort of with Gradj and Shaka-Murthy, that it's still related to the transformer architecture, But we're adding this extra thin layer between the layers of transformers. And its role is to learn something about local structure. And so it's sort of then giving a preference to sort of have things stay close by.
47:49And we were able to show that this could give advantages in faster learning and then better transfer to other different domains of data sets because that was a useful inductive bias to work from. I guess I thought that the that proximity was a part of transformers and why they work already and so I'm looking for greater distinction between what you're describing and yeah so it's actually not so earlier models like an LSTM model was very biased to things nearby affect each other because there's sort of a sequence model, like even older models like HMMs. But in a transformer, it really does attention anywhere inside its context.
48:44And there's no architectural reason to prefer to be influenced by the things nearby to you rather than things that are 173 positions to the right. You know, it's completely... But there's some degree, I guess, of proximity that's enforced by variables like batch sizes and things like that. Right. So to the extent that you've got a certain context window, it has to be within that context window. But, you know, as time has gone by, transformers now have very large context windows, right? It might be a 16 ,000-word context window or some of the recent models are saying there's 100 ,000-word context window.
49:22and within those huge context windows, the model doesn't prefer local versus far away. You know, the model can learn that local is more useful and, you know, transformers do learn that local is more useful. But, you know, that's sort of part of why they need trillion words of training data is to learn these things, whereas maybe learning could be made faster by giving some more biases into the model. mm-hmm and what's the current state of that do you i imagine that there's kind of validating it uh on a small scale and then um part of what's made transformer so successful is demonstrating that you can scale it broadly and you know that scaling laws take over uh how far have you gotten with that yeah i'm we're still at the validate of the small scale stage i'll be honest um yeah Yeah.
50:16Super, super interesting. You know, so before we wrap up, you know, we've talked a little bit about some of your current areas of interest. I'd love to hear you riff kind of more broadly on, you know, just once again, like the way you think about the field and how it's evolved and, you know, where you see everything going. Sure. Yeah. So with the great success of large language models, there's a real sense in which actually there's so much we can do now with language understanding generation, right? Like in earlier decades, almost nothing worked in NLP, right? That you try and get facts out of a piece of text and you were lucky if you could get 50 % of them, right?
50:58Whereas things aren't perfect, but we've just got way better capabilities. And so it almost feels like, oh, we can do language understanding and generation pretty well now. And so the question is then, where do things head from there? And there are many possible directions. You know, there are other ones like multimodal models of doing things between language and vision is certainly a good area. But if I mention the two that I'm most interested at the moment, you know, one is that our current technologies of large language models really only work for major languages. They work great for English or Chinese or Spanish or German.
51:40And you can do reasonably well when you're then going to sort of something like, I don't know, Dutch or Czech. But, you know, there are thousands of languages in the world. and for the vast majority of them, there just isn't much training data available. Some of them actually have millions of speakers, like some Indian languages, but the amount of written text available isn't in the billions of words, let alone the trillions of words. And then there are lots of smaller languages where there are only sort of small speech communities and you'd be lucky to get millions of words of text available.
52:19So there's questions about how to extend some of our breakthrough methods to the rest of the world's languages. And there are interesting ideas there about how you can do transfer learning or exploit the commonalities of human language or human language families. And so that's one interesting area. But perhaps the most profound one is looking behind the veil of language and these questions of reasoning and intelligence and how humans store their knowledge of the world, that these questions are much more coming into focus because we have such good language understanding and generation that a lot of what we're missing is then the stuff behind that.
53:04And that's going to make all of the difference in giving us intelligent machines, but also intelligent language users, because a lot of being an intelligent language user is having that knowledge behind being able to say in fluent sentences. You're actually saying the right things in the fluent sentences. And so what are examples of the things, you know, the stuff behind that? Well, so better understandings of knowledge. So we sort of touched on that earlier, that although to some extent these models know a lot, they don't actually reason well of understanding how facts fit together. So they'll three times say something that is right, but then they'll just say something completely wrong, right?
53:51You get this all the time. You sort of say, where did Christopher Manning get his PhD from? Stanford University. And then you'll just ask it differently or ask something differently. You'll say, was Chris Manning's PhD from Berkeley? And I'll say yes, right? The models just don't have this coherent knowledge behind them. And to start working out better ways to get sort of consistent knowledge and world models behind our language facade, I think is what's going to take us on some of the next steps towards artificial intelligence. And it sounds like you see that as more fundamental or foundational than quote unquote fixing hallucination.
54:36It's a fundamental shift in architecture or approach or the way we think about building these models. Yeah, I think we still need major new developments in neural networks before we'll get there. Yeah, yeah, yeah. Awesome. Well, Chris, there's a lot worse that a researcher can do than being the one that kind of pushes the stuck doors open. So congrats on your success and all the impactful research that you've done and looking forward to kind of keeping in touch and seeing more. Thanks a lot, Sam. It's been great talking to you.
From the publisher
Today, we're joined by Christopher Manning, the Thomas M. Siebel professor in Machine Learning at Stanford University and a recent recipient of the 2024 IEEE John von Neumann medal. In our conversation with Chris, we discuss his contributions to foundational research areas in NLP, including word embeddings and attention. We explore his perspectives on the intersection of linguistics and large language models, their ability to learn human language structures, and their potential to teach us about human language acquisition. We also dig into the concept of “intelligence” in language models, as well as the reasoning capabilities of LLMs. Finally, Chris shares his current research interests, alternative architectures he anticipates emerging beyond the LLM, and opportunities ahead in AI research.
The complete show notes for this episode can be found at https://twimlai.com/go/686.




