In short
Stanford Psychology Podcast Episode 163: Roger Levy - The Science of Language in the Era of AI
Episode Overview In this episode, host Su engages with Dr. Roger Levy, a professor at MIT and director of the Computational Psycholinguistics Laboratory. They delve into Levy's research on language processing, acquisition, and the implications of generative AI on our understanding of language.
Key Themes
- Language as a Cognitive Phenomenon: Exploration of how humans understand, produce, and acquire language.
- The Intersection of Language Science and AI: Insights into how AI models, particularly large language models (LLMs), relate to human language understanding.
- Competence vs. Performance: Discussion of the distinction between linguistic knowledge (competence) and the use of that knowledge in real-world contexts (performance).
Detailed Summary
Introduction to Dr. Roger Levy
- Dr. Levy's background combines mathematics, linguistics, and computational modeling.
- His research focuses on the architecture of language comprehension, production, and acquisition.
Language Processing and Acquisition
- Language allows us to convey and understand abstract thoughts effortlessly.
- Unique aspects of language acquisition in infants, who learn languages without explicit teaching.
- Human language is characterized by complex cognitive capabilities that are still not fully understood.
Key Research Questions
- How do we process and produce language using a finite set of vocabulary and grammatical rules?
- How do we learn language through environmental exposure?
The Role of Computational Psycholinguistics
- Levy emphasizes the need for models that can predict language processing and comprehension.
- Large language models (LLMs) have become significant tools in understanding language structures.
Language Models and Cognitive Science
- Historical approaches to natural language processing relied on symbolic models, but recent advances have favored neural network models.
- LLMs can provide insights into learnability and language processing by analyzing vast amounts of linguistic data.
Generative AI and Language Science
- The success of LLMs does not depend on a scientific understanding of language in the traditional sense.
- Practical applications focus on predictive accuracy rather than theoretical understanding.
Competence vs. Performance Distinction
- Competence: Knowledge of language rules and structures.
- Performance: Actual use of language, which may be affected by cognitive constraints.
- The distinction is crucial for building an accurate model of linguistic behavior.
Cognitive Effort in Language Comprehension
- Cognitive effort refers to the mental resources required to process language, which varies based on sentence complexity.
- LLMs can sometimes mirror human processing but may not fully capture the nuances of cognitive effort.
Garden Path Sentences
- Example of a garden path sentence demonstrates unexpected cognitive effort due to misinterpretation.
- LLMs can replicate the surprisal associated with ambiguous sentences, indicating their ability to model structural relationships.
Future Directions in Research
- Focus on quantitative approaches to language science, leveraging LLMs for theoretical insights.
- Emphasis on understanding the representations of language in the human mind and its implications for cognitive science.
Key Takeaways
- LLMs represent a transformative advancement in processing human language but should not be mistaken for exact models of human cognition.
- Traditional tools in language science remain valuable, and researchers should integrate these with modern advancements.
- There is a promising future for research at the intersection of language science and AI, particularly in understanding how we learn and use language.
Conclusion Dr. Levy's work highlights the complex interplay between language, cognition, and artificial intelligence, advocating for a holistic approach that incorporates both traditional linguistic theories and modern computational methods.
Additional Resources
- [Dr. Roger Levy's Lab Website](http://cpl.mit.edu/)
- [Dr. Levy's Personal Website](https://www.mit.edu/~rplevy/)
- [Research Article: The Science of Language in the Era of Generative AI](https://mit-genai.pubpub.org/pub/ak3evnmm/release/1)
Podcast Information
- Host: Su
- Podcast Website: [Stanford Psychology Podcast](https://stanfordpsychologypodcast.com)
- Feedback: stanfordpsychpodcast@gmail.com
- Twitter: [@StanfordPsyPod](https://twitter.com/StanfordPsyPod)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:09Welcome back to the Stanford Psychology Podcast. His research focuses on theoretical and applied questions in the processing and acquisition of natural language. His work furthers our understanding of cognitive underpinnings of language processing and acquisition, combining computational modeling, psycholinguistic experimentation, and analysis of large naturalistic language datasets to help design models and algorithms that will allow machines to process human language. In today's episode, we discuss his research background together with his recent work the science of language in the era of generative AI.
0:44So without further ado, here is our conversation.
1:09Thank you so much for joining me. I'm really excited to host you today. and I'd like to begin with a question that I hope also serves as a gentle introduction to your work. Language has always struck me as a fascinating concept. It's a tool we use to externalize abstract thoughts, to construct meaning, to bridge minds through shared understanding. The more I think about language, the more I find myself thought intellectually captivated. But with that in mind, what aspect of language do you study and how did your research focus take shape? Did your journey begin with a specific question or a problem or did your interest evolve more organically over time shaped by discoveries along the way?
1:56I'm just curious whether there has been a central thread, a theme or motivation that's guided you in this inquiry. Thank you, Sue. And thank you for having me on this podcast. I appreciate the opportunity to talk with you, and I appreciate this question. Briefly, I describe my research program as focusing on the fundamental architecture of human language comprehension, production, and acquisition. How do we understand what we hear and read, most of the time effortlessly and very quickly? How do we convert thoughts into words and sentences that put understanding of new meanings into other people's heads?
2:38And how do we learn without really being explicitly taught from infancy and in our first years of life from the environment? How do we learn the resources that allow us to do this? Languages are different around the world, yet young children learn this effortlessly regardless of the environment that they're brought up in. So I see it as one of the great scientific challenges and mysteries of the present day. The human mind is arguably the most complex object that we have discovered in the universe. And language is a distinctive property of it, that we don't know any other species that has the flexibility that we have to represent and to convey thoughts and meanings in such a contextually specific and flexible way.
3:25For me personally, how did I arrive in this? I'll actually give just a bit of personal history. I was always very mathematically inclined. And so as an undergraduate, I majored in mathematics, and I got very interested in population genetics. I'll say that language science, I don't think historically has had great PR. It's not as well known in the public at large as I think it should be. It's an amazing area. So I knew very little about it all through undergraduate. But when I was an undergraduate in the early 1990s, I did get very interested in population genetics, which is a very interesting area, evolutionary biology, studying the transmission of information from organism to organism through genes, primarily.
4:07That was an emerging interest from mine. And also in the 1990s, it had become very popular for American undergraduates to spend some time in another country studying and learning from living in another environment. It was after the end of the Cold War, and there was a renewed interest in sort of having the new generation of Americans be exposed to cultures around the world, languages around the world. And I wound up spending several years in East Asia, in Singapore, Taiwan, and Japan. And I became completely captivated with the differences in how people live, even more than anything else, actually the differences in the languages.
4:47From anything I had experienced, I had studied European languages a bit when I was pre-college, but I never thought that that would be something that I would spend a career's worth of time focusing on. But I just became completely fascinated. And that sort of led me ultimately to change my focus from evolutionary biology to the study of language, which is also a means of transmission of information from organism to organism. It's just a different method. It's through cultural means, through conventions that have formed in the context of a human society. And that is endlessly captivating and exciting.
5:23And that led me ultimately, I did my graduate work at Stanford University, my PhD in linguistics there. I arrived right about the same time as Chris Manning started his faculty position there. And I was very interested. I've always been very drawn to words and up. So language science sort of makes a broad division between things at the level of the word and below. So like signs and sign language and sounds in spoken language, morphology. And then there's the words and up, how words come together to form phrases, larger units, syntax, semantics, and pragmatics. I've always been particularly drawn to that area.
5:59And then Chris showed up and taught statistical natural language processing. I had always really loved probability from my mathematics background. And all the pieces eventually fit together and led to what I do now. So it really was an intellectual journey for me, and it was not a straight path. I discovered variously and incrementally the different pieces of the puzzle that I now put together in my research program, which I sort of now characterizes at the interface of linguistics, psychology, artificial intelligence. We see that interdisciplinary research is becoming more and more popular, but it's always really excites me to see really well-known scientists coming from different backgrounds and showing us that given if your background is something else, it can be a superstar somewhere else.
6:44So this was really nice to hear. going back to language sciences and also to connect your background to the work we are going to discuss in this episode a little bit better you mentioned that your research focuses on understanding human language abilities broadly could you briefly explain some of the central goals of language science as you see them for our audience? Absolutely. So I like to start explaining this by doing some back of the envelope calculations or giving, showing you some results of back of the envelope calculations. To first approximation, every day, each of us, adult, fluent or native speakers of English or whatever language you're in the environment of, you hear and read thousands of sentences that you've never encountered before and you produce you speak or write at least hundreds more so your experience of language is novelty all the time but most of the time you are very good at understanding what you hear and read and you're very good at conveying your meanings even if it's a meaning you've never conveyed before a sentence that you've never heard before how do we do that we have a finite set of vocabulary items a very constrained set of grammatical generalizations and rules that we use to structure language, both in comprehension and production.
8:10You know, von Humboldt, back in the early 19th century, characterized language as the infinite use of finite means. And I think that it really is the classic framing of the problem. So we have some knowledge. We have finite knowledge of our language or of multiple languages for multilingual individuals. And we deploy it very rapidly to understand what we encounter to express our meanings. And we do it accurately in a contextually specific way. So how do we do that? And that is a very computationally intensive discipline to describe that. And then the other question is, how do we learn these tools?
8:47How do we learn without explicit supervision from infancy? And that is another set of computational problems. And it intersects with disciplines like machine learning, artificial intelligence, but also sort of human science disciplines, psychology, developmental science, obviously linguistics. So those are the central goals. How do we represent, learn, and use this constrained, finite set of information that allows us unboundedly expressive comprehension and production? And how do we do that? What are the cognitive resources in our mind that actually undergird that capability? How do we actually deploy that knowledge for language use?
9:25Those are the central questions. Mm-hmm. Following that, we are slowly transitioning into the work we are going to talk. You're at MIT, and you have a lab that's focusing on computational psycholinguistics. These days, we are used to seeing the words computation and language side by side a lot. Large language models and the ways algorithms interact with linguistic data have recently captured significant public attention. How do you see these technologies fitting into or influencing the scientific study of language, your study of language, as you describe? You know, it's a really interesting question because clearly one aspect of this is that language science is sort of a broad, inclusive term.
10:12It includes things that we might think of as sort of more basic science enterprises, but also a lot more engineering-oriented areas. So natural language processing is certainly a part of the field. And historically, the stated goals of the field, the shared goals, are in large part applied. We want to actually have machines that do something that seems homologous to human-type understanding and production. And for decades and decades, the field of natural language processing or computational linguistics, which is more or less a synonym, arguably, had bets on how this would be done. And it was using models that used the symbolic hierarchical characterization of language structure that come from the mid 20th century, the work of Chomsky and dependency grammarians in Europe and logicians.
11:01These symbolic tools were really where the bets were. And it turns out that actually the thing that works, at least that we've figured out how to make work so far, is something that seems very different in some ways, these neural network models. And that's a big surprise in some ways. And there's no question that they're better than anything that's before it. And that has immense consequences for the engineering side of language science. And it opens up many opportunities for the scientific parts of language science. You know, things like, for example, learnability. What could be learned from, say, a childhood's quantity of linguistic data?
11:35In the United States, in an English-language-speaking household, the average child maybe gets 10 to 11 million words of exposure a year, give or take a factor of two or three, depending on the household. So it's in the tens of millions, maybe approaching 100 million words during childhood. What can you get out of 100 million words in principle? And we never really had a way of answering that in a strongly affirmative, oh, here you really can learn these things in this way. The modern LLMs, or let's say the LLM technology, allows us to answer those things. So it gives us insights into learnability questions that I think have been long understood to be at the heart of cognitive science of language, and actually cognitive science more broadly, the nativism versus empiricism debate.
12:21That's not to say that they give direct insight into actually how humans are learning it, but the question of what could be learned from such data is a classic part of the question. There are also numerous more instrumental uses of large language models. So for example, in the kinds of theories that I build, a core quantity is very often like the probability of some linguistic unit in context. For example, a word in its context. And language models are the best tools that we have. Like modern neural network language models, LLMs are the best tools that we have right now to estimate those probabilities at scale in a way that correlate reasonably well to human expectations.
12:55And that's only scratching the surface. There's sort of automatic processing of data through LLMs. There's thinking about the kind of theories that you could build, like how could you embed, how could you represent symbolic structures inside distributed vector-based representations, which obviously the brain, at some level of, once you dig down enough, ultimately has to be doing that. How could that be done? These models also give us insights into those questions. So the consequences and the applications are numerous. and also your recent work including the exploratory analysis the science of language in the era of generative ai you co-authored that we are discussing today explores the relationship between the behavior of these language models and human language processing and i assume like what we just discussed now motivates you and your co-authors to take this investigation The analysis mentions that the practical success of language models doesn't necessarily rely on a scientific understanding of language, at least not in the way generative grammar-based theories approaches it.
14:00Could you elaborate on this difference in research strategy or goals between traditional language science and contemporary language model development? Yeah, absolutely. So this actually connects to some of the things that I said earlier. So just to give a little bit more context, the article that you mentioned, The Science of Language in the Era of Generative AI, it's one that my co-authors and I very recently finished. It's sort of an overview survey for a non-specialist crowd of language science, linguistics, and related disciplines and large language models and their relationship. In addition to me, my two co-authors are Yoon Kim, who is a rising star in natural language processing and machine learning here at MIT in the electrical engineering and computer science department.
14:42And Danny Fox, who is a widely known first rate figure in semantics and syntax and generative grammar in the linguistics department. So we really are coming from very different angles, the three of us. Danny and I had taught together, Yoon and I had collaborated together, and we asked, well, what will we find and common ground will we find through our very different points of view if we try to write an overview article together? So that's the context for that. And as you say, one of the many things that we mentioned in this report is that practical success of LLMs doesn't necessarily rely on a scientific understanding of language.
15:17It's certainly not a generative grammar type theory, or at least it doesn't seem like it does. I'd be happy to elaborate on this. And as I said before, in natural language processing, the sort of stated goal is to develop a system that actually can process language in a way that sort of is as effective as humans and understands text and maybe spoken language in ways that are similar to what humans do. And that's characteristic of a broader approach to AI and machine learning-based disciplines of where you actually want to implement a system that actually deals with inputs in an effective way that corresponds well to what those inputs reflect about the world.
15:56The unique thing about language compared with computer vision, say, or maybe aspects of robots navigating the physical environment, other machine learning AI type disciplines, the unique thing about human language is that human language comes from humans. It's a thing about the world that comes from humans, whereas facts about vision, facts about images and about videos, facts about how you navigate a physical body in the world, those are not things that come from humans, right? But language does. So the ground truth is inevitably like a human-based ground truth. Although interestingly, you can actually think about what would a system be like that processes language in a way that's even more effective than a human, for example.
16:38And that actually gets at some interesting questions about the way that the mind works. At any rate, I think the broad characterization is that But it's just not a goal. Like part of the goal that I just described is not develop a scientific understanding of the domain. That's just not the way that modern machine learning works. It is develop a predictive model that is accurate and generalized as well. And that's been the driving force. And it turns out that it wasn't a foregone conclusion that models that are actually hard to understand what they're doing are the ones that work best. But it turns out that that is what is working best, at least right now.
17:13And so that's a very different approach and strategy. It's sort of a prioritization of a system that works well from a predictive and from a generalizing perspective without a prioritization on scientific interpretability. Whereas in traditional language science, including generative grammar, say, there's a much higher premium on scientific interpretability. In the report, you discuss a fundamental concept in language science, the competence performance distinction. Could you explain what this distinction means and what is linguistic competence and how does it differ from linguistic performance?
17:50And why is this distinction considered so important for providing an accurate account of linguistic behavior and nature of human language knowledge? Sure. There are many different ways of describing the competence performance distinction. Different theorists have described it differently. I'm fond of thinking about it this way. I described before that what the human mind does in producing and understanding language is make unboundedly expressive means of use of finite means. So we have this finite knowledge and we use it. The knowledge is the competence. Our ability to use it is the performance.
18:27And there are lots of cases within language and outside of language where our ability to use knowledge or to use information, And that might not be explicitly characterizable knowledge, not necessarily knowledge that we can write down. It might be other kinds of knowledge, like the ability to execute a skilled pattern with your body, for example, procedural knowledge. There's linguistic knowledge, for example, you know, at a very basic level, what are the words that I know? And then there's my ability to use it. And I may not be able to use everything that I know equally, easily or rapidly or in all contexts.
18:58If you take very straightforwardly a rare word, it may take me a little longer to remember what it means than if it's a frequent word. That's at the word level. There are phrasal correspondence of this. But the competence-performance distinction says that, well, you want to build an analytic model, a scientific model of language as a cognitive domain by first characterizing the knowledge and then characterizing how the knowledge is used. and the how the knowledge is used is going to run into limits based on the computational constraints of our mind. So for example, language often taxes your memory.
19:33And there are very famous examples of sentences that if you arrange them in one way, they tax your memory more than others. So if I say, you know, the rat that the cat that the dog chased, killed, ate the cheese, that's very hard to understand but if i say the rat that was killed by the cat that was chased by the dog ate the cheese starts to become easier that was the difference in the arrangement of the words and it turns out for reasons that we can understand theoretically that difference affects our knowledge of language both of those are perfectly fine but one of them taxes our memory in a way that the other doesn't our use of linguistic knowledge interacts with our cognitive resources in that particular case, memory, in ways that make not all applications of the knowledge equally successful.
20:20That's the competence performance distinction. And in the context of LLMs, it's important to remember that even if, in a sense, you could give a human and an LLM the same training data, but at the end, no matter what the internal representations are, the LLM has different computational constraints in using a linguistic input and processing it than a human does. So even if we could, and we don't know how to say, how to sort of perfectly control, like, let's make the knowledge, quote unquote, of the LLM exactly the same, if that's a useful, meaningful thing to talk about, which is not clear that it is, but if it is, then the performance situation would be rather different.
20:59So certainly for studying human language, the Combinance Performance Distinction has a long history in generative grammar research. I continue to believe that it's a very scientifically productive framework for thinking about how to build theories and understand data. It's a more open question whether it's a useful paradigm in the context of modern LLM-type artificial systems. I want to follow up with what you just mentioned. The LLMs, they differ in the way we process these information. And also, they do not have this built-in competence performance distinction as well. So could you elaborate more on that and what that means in practice?
21:42There are different examples that one could use, but a simple way of describing it is that the LLMs, their currency, what they're learning to do is to predict in context, right? And now their training data consists of a lot of well-formed language together with some language that has problems in it. It also has a bunch of non-language stuff like code and weird HTML and so forth. But setting that aside, in terms of the language part, there's a lot of well-formed language and also a reasonable helping of ill-formed language. So there's nothing built in about how an LLM is trained to dissuade it from predicting the ill-formed language in contexts in which it's likely to appear.
22:26So examples of this, and we could get into the theoretical details of how to think about these cases. But for the sake of the argument, let's take a case like, you know, the key to the cabinets are on the table. In most grammatical theory, we would say that that is that is an error because the verb is are. It's a plural form. But the subject word of the sentence is key. That's a singular. So it should be an is not an are. Right. So that's an error. However, it's a common kind of error. It's so common that actually, like for a long time, the New York Times would have like a sort of grammar mavens kind of, you know, oh, look at the mistakes that we found that slipped by the copy editors.
23:07These are such common mistakes that they slip by the copy editors. They happen often and they're hard to catch. So they're ungrammatical, but they happen. OK, so in LLM, that information, that distinction between ungrammatical and grammatical is not built in. So if you have something that happens which is sort of wrong in some sense, the LLM will just learn that pattern. Whereas, you know, the competence performance distinction says, you know, we want to draw a distinction between our competence says that the subject should agree with the verb. We have a body of independent evidence for that. But there's a pattern in performance that has to do, arguably, in this case, also with the way that memory works and the way that language and context of language is represented in memory that makes that kind of error particularly likely.
23:52The frequency of that pattern is the combination of a grammatically correct pattern with a common error. Whereas the LLM is, there's nothing that's built into the LLM that rewards it for learning or using that distinction. Now, it may be an emergent property that you can distinguish within an LLM that one is an error and one is not. But it's not built into the way the LLM is trained. Yes. Switching gears to the language processing, the report discusses how humans process language incrementally, meaning we don't wait until the end of a sentence to start understanding it. This process involves cognitive efforts.
24:31Could you describe what cognitive effort in language comprehension refers to and why it's a key area of study in psycholinguistics? And I think this This is going to be especially interesting because last week I had a chance to have a conversation with Professor Jennifer Hu, who was from your lab. And we talked about cognitive effort. And hopefully once this episode airs, people will have access to that episode as well. Oh, absolutely. I'm glad to. And I'm delighted that you had Jen on the podcast. She's amazing and does wonderful research. And I've really been lucky to work with her. So let's put it this way.
25:09Like many constructs in cognitive science and psychology, we don't know at the end of the day really exactly what it is, but it seems to be the amount of computation of some sort that the mind and indeed the brain has to do. So why do I distinguish between the mind and the brain here? Well, you might measure them differently. So cognitive effort, when you're taking a sort of a mentalistic point of view, you often think about it in terms of time or attention. So something that is cognitively effortful will require you to divert resources from other things that your mind is actively doing, like other things you're thinking about and focus.
25:47It also might make you take more time to do those things. At the level of the brain, you actually may see this in terms of like neural activity. And this is actually measurable in a very coarse-grained way, say by fMRI, which actually measures how much activation there was in a region of the brain recently. That's the oxygenation. Saturation. That's what it's picking up on. The notion of cognitive effort is how to specify those computations that are being done. But there are some things, some tasks that the brain and mind are faced with in particular moments that are more effortful than others.
26:22Now, operationally, like how does that manifest in language comprehension? okay so it's hard to actually pick that up when you're studying spoken or signed language comprehension because the input is just coming in and uh and the comprehender is doing what they need to do they may be understanding it completely they may not be understanding it completely but there are still ways of getting at it so most directly like i said fmri cognitively effortful linguistic inputs will cause greater activation and by the way i should be clear cognitive effort for language understanding is the amount of work that needs to be done by the mind or the brain in order to quote unquote successfully process the input.
27:04That is, get all of the information out of it that it contains in the context in which it appears. So for example, if I say the boys went out to chat, then that may actually be a little surprising to you because you expected the word play. And indeed, you had to get the meaning of the word chat and you had to connect it both grammatically and meaningfully with the context of the previous parts of the sentence. So that amount of work that has to be done is the cognitive effort. You can study it to some extent in spoken or signed language comprehension, but you can really study it in reading very easily because in reading, the control over the pacing of the input is in the eyes or the hands of the comprehender.
Read the full transcript
27:49And so it turns out empirically that things that are cognitively effortful, that is that when we use the theoretical construct of cognitive effort, define it in a principled way, more effortful things take longer for people. They spend longer on them. So that's what cognitive effort is. And it's key because it turns out the cognitive effort varies in exquisitely detailed patterns in very, very temporarily fine-grained manner in language understanding. And indeed, in language production too. That's really interesting. In the report, you use garden pet sentences as an example to show how confusion or just spraisal that comes from those sentences actually produce this unexpected effort or just lingering around the words or specific parts of the sentences.
28:39Could you walk us through why a sentence like this is confusing? And also following up on that, you also cite your work or work of other scientists to show that now you are starting to use LLMs to pinpoint that spraisal or at least estimate that confusion. What advantages do language models offer for testing these? Absolutely. Sure. Garden path sentences are, they're a workhorse, a pun intended for those of you who are in the know, of psycholinguistic research. And despite having been studied for decades, I think they continue to offer fresh insights. So first of all, for those of the listeners who are not familiar with the garden path sentence idea already, I'm going to give an example garden path sentence.
29:30And I just want you to sort of notice as you're listening to it, where you find it's sort of taking you aback and what your subjective experience is of listening to the sentence. So here we go. And I'm going to do this in an intentionally neutral intonation. So it's going to sound a little bit unusually neutral, which is meant to sort of more closely approximate what it would be like to read the sentence where intonation is not available. So here we go. As the dog scratched, the vet removed the muzzle. i probably didn't do that so well but some of you may have had the experience that at the word removed it was surprising and it actually it was confusing even and you may even conclude that there's a problem with the sentence so there's something grammatically wrong with the sentence so let me explain i'll just say it again as the dog scratched the vet removed the muzzle so what's going on with that sentence so let's just take the beginning of the sentence, as the dog scratched the vet.
30:27Okay. So a natural way of understanding what's going on in that sentence, this involves syntax and semantics. So the structure of sentences and their meanings from a grammatical point of view, it's natural to interpret. So scratched is obviously the first verb that you encounter and it needs a subject and maybe an object and the dog is before it. That's where subjects occur in English before verbs. And right after it is this, what's called and noun phrase, the vet, and what a great object for the scratch. So it sounds like the dog is scratching the vet. Dog is the subject of scratch, that is the object of scratch.
30:59But if that's true, that's the case, then how would that word removed fit in? And it actually grammatically, like it doesn't work because removed is clearly not part of the same what's called clause as the dog scratched the vet. But if removed is part of a different clause, then you should have seen the subject of removed first. So you might have expected something as the dog scratched the vet, his assistant removed the muzzle or something like that, but that you didn't get that other phrase. And that's what's surprising about it. Okay. Now, if we go back to the whole sentence, as the dog scratched, the vet removed the muzzle, actually what's going on in that sentence is scratched doesn't have an object in that sentence.
31:37Um, so scratch is the end of that clause that, uh, of its own clause. And the vet removed the muzzle is the second clause. So if I were to give it intonation, it would be, as the dog scratched, the vet removed the muzzle. And in reading, if I don't have a comma there, it's very clear that people misinterpreted that way. So why is that interesting? So what historically, and I think this idea continues to hold, like in a compelling way of thinking about what's going on there, is that, well, it's confusing because originally those aspects of sentence interpretation, the dog being the subject of removed, the VAT being the object of removed, those things have cognitive reality.
32:15And that cognitive reality is directly implicated in your experience of understanding the sentence. Okay. So those are structural symbolic elements of what's in your mind at the moment. And that has a direct causal effect on how you understand the sentence and your cognitive effort. And removed is so cognitively effortful that you may not even be able to successfully recover what it was supposed to mean. That is the idea. It's a way of illustrating extreme cognitive effort, and it's a way of illustrating how the structural symbolic elements of sentence structure and meaning might be directly implicated in cognitive effort.
32:54Now, connecting this to LLMs, so LLMs too are surprised at the word removed. We put the sentences through GBT2 or something, And we get these very high, what are called surprisal values, their negative log probabilities at this word. Okay. And so that's neat. It actually is a really neat way of showing that LLMs actually implicitly do something like represent these symbolic structural relationships. Why would removed be such a surprising improbable word in that context? It's because, I mean, it's right after a noun, it's a perfectly good place to have a word, but not for this noun, not if that noun is the object of this other verb before it.
33:33That's sort of the general idea. Now, one thing that we also do is we build more quantitative models of the relationship between that degree of surprise. It turns out the negative log probability or the surprisal of a word in its context, actually, in the large scale, in ordinary language processing, there's a linear relationship, a very systematic relationship between the surprisal value of a word and how hard it is to understand in its context. What's really interesting is that in these garden path sentences, so that word removed it does what we call disambiguate the garden path there was one interpretation that you started off on and that word says nope that's the wrong interpretation if you recover from the garden path you find that means finding the other interpretation where the vet is not the object of scratch the subject of removed so in these these garden path situations llms do not accurately predict how long words take to read and so they systematically in a sense, even though they sort of qualitatively capture the pattern of these garden path disambiguation words being surprising, they seem to underpredict the magnitude of difficulty that people have.
34:39And that there are many things that you can do with that. Theoretically, I think it's still, there's some really interesting open questions, but one possible interpretation, which I'm quite partial to right now, is it's a kind of evidence for the causal implication of symbolic representations in language processing for humans that might not be obvious otherwise. That's neat because if that's the right interpretation, it suggests a really important difference between how humans process language and how neural network models process language. In neural network models, no matter what the input sequence is, you've got a vector representation of the output at the end, and it's just different vector representations.
35:17For a human, the idea is that there's something qualitatively different going on in these kind of garden path disambiguation situations where you actually had a structural parse that got rejected. So that's at a very high level, and then we can get more technical about it. And that's a lot of what we do is we try to operationalize those ideas technically. Based on all these work and the insights we just talked about, and also the insights you talked about with your colleagues in this report, what do you see as the most promising areas for future research at the intersection of these generated AI models and the science of language, science of communication or producing language, what direction has your work started taking, or do you think it is aligned with what you have been studying?
36:06Yeah, that's a great question. So I think there are many promising areas. One thing I'll say is that in terms of my career, I've always sort of envisioned a quantitative science of language. And a quantitative science of language that, I should be clear, it doesn't just mean counting things. This doesn't mean like counting frequencies of occurrence. It means a science in which the units are very often discrete entities, but there are like quantities that are associated with them, quantities that arise from the configurations of those elements together, and that those quantities are broad-rangingly predictive of how we learn and comprehend and produce.
36:47I think that LLMs are an extraordinarily powerful tool for helping advance that mission of quantitative science of language, both in terms of the instrumental tools for estimating quantities like probabilities and context or like similarities of sentences, but also as a kind of inspiration for theoretical proposals about the structure of linguistic knowledge and the use of linguistic knowledge. So language is definitely discreetly structured in many different ways, but I think it's productive theoretically to think, for example, of, you know, if we have tens of thousands of different words in our vocabulary, that doesn't mean that our meaning spaces, its dimensionality is not tens of thousands.
37:28It's going to have to be something lower than that. The neural network type dimensionality, other dimensionality reduction methods, they give us insights, theoretical ideas as well. So both from technical instrumental use and also for theoretical inspiration. How do we represent the symbolic and vector embedding representations? Those are going to be important questions. But once again, I think we're in a better position than ever to ask, like, what's the representations of language in the human mind and even the brain? Thank you so much for giving all these insights. And I really encourage everyone to visit your report.
38:03And before we wrap up, as a final note, maybe what message would you like our audience of psychology researchers, psychology enthusiasts to take away regarding the potential of using these models as tools or models for understanding human language and cognition? Great. Make sure I have an opportunity to give you a final thanks as well. But let me first start with my last message for the audience. So it's a two-part message. So one is that is, there's no question that LLMs are a better engineering result for artificial systems for human language than anything else that we have on the planet. No question.
38:48They're a transformative technology. However, taking the classic example of don't assume that because airplanes are the only way that man has figured out how to create an artificial flying vessel that airplanes are the right model of birds. Don't assume that LLMs give us the right model of how humans process and learn language. And in particular, I think some of the classic tools that we have, the toolkit of symbolic structured representations, the tools that we have for mathematical linguistics, but also the tools of probability theory, Bayes' nets, Bayesian probability, inference under uncertainty, all those things still are entirely relevant.
39:32Don't forget about them. And don't conclude that because they're not the latest and greatest things in all aspects of machine learning, as it appears to you, that you don't need to learn them. In fact, they're very, very valuable. But the last part of the message is, well, in the modern moment, we are in a better position than ever to continue advancing the core goals of language science and understanding how language is learned and used in the human mind. So it's a great time to be in the field. And just as a final note, thank you very much for having me on the program. And thank you for the kind words as well on the article.
40:06I'm going to convey that to my co-authors, Dani and Dion as well. Thanks so much for being here. Thank you so much for listening. If you enjoyed this podcast, help us make even more people excited about science by leaving us a review on Spotify, Apple Podcasts, or elsewhere, and subscribing to our non-spam all-fans sub stack at Stanford SciPod to connect with other listeners. You can also shoot us an email with your thoughts or suggestions at stanfordpsipodcast at gmail.com. Thank you, and have a wonderful day.
40:43Thank you.
From the publisher
Su chats with Dr. Roger Levy. Dr. Levy is a Professor in the Department of Brain and Cognitive Sciences at MIT, where he directs the Computational Psycholinguistics Laboratory. His research focuses on theoretical and applied questions in the processing and acquisition of natural language. His work furthers our understanding of the cognitive underpinning of language processing and acquisition, combining computational modeling, psycholinguistic experimentation, and analysis of large, naturalistic language datasets, to help design models and algorithms that will allow machines to process human language. In today's episode, we discuss his research background together with his recent work "The Science of Language in the Era of Generative AI".
Roger’s review: https://mit-genai.pubpub.org/pub/ak3evnmm/release/1
Roger’s lab website: http://cpl.mit.edu/
Roger’s personal website: https://www.mit.edu/~rplevy/
Su’s Twitter: https://x.com/sudkrc
Podcast Twitter @StanfordPsyPod
Podcast Substack https://stanfordpsypod.substack.com/
Let us know what you thought of this episode, or of the podcast! :) stanfordpsychpodcast@gmail.com




