In short
Podcast Summary: The TWIML AI Podcast - Episode #626: Are LLMs Overhyped or Underappreciated? with Marti Hearst
Host
- Sam Charrington
Guest
- Marti Hearst, Professor at UC Berkeley
Episode Overview In this episode, Marti Hearst discusses the current landscape of AI language models, particularly focusing on large language models (LLMs) like ChatGPT. The conversation explores their potential to enhance efficiency and their risks, particularly in the realm of misinformation. Hearst expresses skepticism regarding the cognition capabilities of these models compared to human brain nuance and emphasizes the importance of ongoing research in ensuring the safe and appropriate application of AI technologies.
Key Themes
- Skepticism Towards AI Hype
- Hearst has observed past overhyped claims in AI that did not materialize (e.g., early NLP advancements, IBM Watson).
- Expresses caution against contributing to the hype cycle, emphasizing realistic expectations from LLMs.
- The Evolution of Language Processing
- Hearst has a lengthy history in natural language processing (NLP) and observes that significant progress has been made recently with LLMs.
- The current moment is described as transformational, specifically with tools like Copilot and ChatGPT enhancing programming and information retrieval.
- Intersection of Language and Visualization
- Hearst discusses the interplay between text and visual data representation, emphasizing the need for research in this field.
- There’s a growing recognition that visuals and text serve different purposes and can complement each other in information dissemination.
- Concerns About Misinformation
- The ability of LLMs to generate and spread false information is a significant concern.
- Hearst emphasizes the necessity for specialized research to combat misinformation and ensure the reliability of AI-generated content.
- Human-Computer Interaction (HCI)
- The importance of understanding how people interact with AI and the need for user-centric design are highlighted.
- Hearst critiques that many advancements in NLP are driven by engineers who may overlook the human aspect of technology use.
- Future Directions in AI Research
- Hearst anticipates ongoing challenges and opportunities in AI research, particularly regarding AI safety, bias, and the interplay between language and visualization.
- Encourages open-mindedness and passion-driven research from new students in the field.
Key Insights
- Cognition vs. Behavior: Hearst remains skeptical about attributing cognitive abilities to LLMs, arguing that they exhibit behaviors rather than true understanding.
- Emergent Capabilities: The podcast stresses the unexpected emergent capabilities of LLMs and the potential they hold for practical applications.
- Collaborative Potential: Future AI tools may enhance human collaboration through improved interfaces that blend language and visual representations.
Notable Quotes
- "I think it's a sea change."
- "Understanding how people understand things so that we know what to tell the computers to do."
- "The one person band just doesn't exist in this space."
Conclusion This episode provides a nuanced examination of the current state of AI language models and their implications for the future. Marti Hearst's insights underscore the importance of maintaining a critical perspective on AI advancements while recognizing their transformative potential in various fields.
For more detailed information, visit the complete show notes for this episode at [TWIML AI Podcast Episode 626](https://twimlai.com/go/626).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Hey, what's up, everyone? This is Sam. In today's interview, part of our guest host series, you'll hear a conversation led by longtime friend of the show, John Bohannon, director of science at Primer AI and former journalist for publications like Science Magazine, Wired, and others. I'm sure you're going to enjoy this conversation, so let's jump in. Peace.
0:32Good morning, Marty. Good morning, John. So we are just a few miles away from each other across a body of water. I'm in San Francisco. You're across the bay in Berkeley at your office at University of California, Berkeley, where you are the head of the School of Information and a computer scientist who's been in the game for many decades. That's right. So I'm excited to finally do this interview because it's been almost a year in the making. We had a great interview here on the show with Oren Etzioni and former head of AI2 in Seattle. And after that interview, I asked Oren, who would be a really good guest, who would have something worth sharing with the audience?
1:10And you were the first person he said. Well, I listened to that interview and it was an excellent interview. Oren is so articulate and you're a great interviewer. And I'm really honored to be here. And I'm really honored that Oren thought it was worthwhile to recommend me. And you don't do that many interviews, from what I gather. No, I don't. I'm a little bit camera shy, even though I do have to be on camera a lot. Also, I like to be pretty careful about what I say, kind of more from the scientist's perspective. I think Oren is really great at linking science to business and where technology is going.
1:46But yeah, I guess I'm just a little bit shy that way. And that brings me to kind of the big story here. For those listening at home, Marty and I had a chat recently in advance of this interview, just to talk about what we might talk about. And something really striking was that there's this moment, let's call it the chat GPT moment. It's really the large language model moment where artificial intelligence seems to be at some kind of inflection point. And you told me about your long career and how you have kind of seen these moments before. And you're more cautious about speaking publicly to add to the hype cycle because it's often disappointing and often regrettable.
2:25You know, it's easy to say things that you later think were overhyped. Why is this different? Yeah, so I have seen a lot over the years. I've been, I'd say, primarily in natural language processing, that part of AI. And, you know, it's just been a slog in terms of making progress in having machines be able to process language the way people do. I entered it really because I was interested in the brain and interested in language. And I thought it'd be kind of neat if we could have a computer do something with language, which maybe, you know, make cartoons speak, something like that. Animation interested me as well.
2:58We're so far from accomplishing anything that would be realistic, that was more of a scientific endeavor. I think I'm more of a scientist at heart than, you know, I'm not an entrepreneur, for example. So I've seen claims, for example, I remember, I guess, in the early 90s, there was this claim that Oracle bought some NLP company, and it was going to transform everything. And it was just so obviously ludicrous. But you also see, I remember when WebFountain came out with IBM and that was going to transform everything. And it's very much... How about IBM Watson? Well, Watson was a special case in that it was amazing what they did with Jeopardy.
3:36And we can talk about that a bit more. But then there was the claim that it was going to transform healthcare. And again, there was no path from that to directly to healthcare being transformed in the immediate future. So I've learned that if you read the New York Times regularly, as I do, technology is in the business section as opposed to the science section. And that's kind of how technology is talked about in, at least in the US. And of course, there is scientific reporting on it. And your podcast, I think is wonderful in that it goes into a lot of technical details, which is really exciting.
4:08But there's always the business angle when it comes to technology, even when I started out in the late 80s. And, you know, at most there were PCs. Well, and in the late 80s and into the 90s, you became one of the main researchers in search. And search really defined the era that I think is probably coming to a close, the Google era, the era of search, search driving everything. And so you did really see the business side of your research explode and change the world. Are we in a moment like that now? Well, I wouldn't mind talking about search for a few minutes since it is close to my heart. I mean, I was interested in search because I wanted to be able to find things.
4:46I didn't like the library catalog when I was a little kid. And in fact, when I was an intern, I tried to be an intern in my public library in high school, and I was rejected because I wasn't fast enough with filing alphabetically in the card catalog. But I never thought that makes sense. So actually, I always wanted to do a dynamic, smart version of the card catalog, which is what I did in search user interfaces. There's only one spot in the bookshelf for the book representation. What I focused on was search user interfaces. I wouldn't say I was the leading person in search, but I was a leader in search user interfaces, which was kind of a hybrid topic at the time because most of the search field was more on algorithms and not so much on the user interface.
5:25So I brought those two together and that was super exciting because the technology or the kind of framework that I advocated for and showed empirically worked It did become the standard for, it's still a standard, faceted interaction, what you see on a website when you're shopping or library catalogs where you can slice and dice and filter in different ways to find the items that you want. Getting that interface to work well was a big challenge, and that was sort of the breakthrough. What was the big problem with search interfaces before you got into the game? Well, when I got into the game, most software did not have search full stop.
6:01I mean, if you had an application, you couldn't search for material within it. It was just rare. It just didn't happen much. When I got into the game, library catalogs were searched by saying, you know, P-N-Bonneman, comma, J, to find the personal name of the author. I mean, it was command line. And then there were Westlaw and these very expensive tools that you could subscribe to, say, if you were a lawyer. It was all keyword based. But the interface, there was no thought to the interface. So it's just a listing of the output that you got, usually in chronological order. And so there just was no there there.
6:36The web changed things. But even with the web, the initial search was, you know, the 10 blue links, which has actually been really hard to improve on. And I would say until now, which we could get to the new moment, you know, Google's inched towards showing answers to questions. But I remember talking with someone there saying that they were conservative initially because they didn't want to show incorrect information. And I thought that was the right way to go. Isn't that one of the big shifts? It's like once upon a time, the purpose of search was to find a document or find a resource. But nowadays, you want the answer to a question.
7:06It's almost a shift in intention. Actually, I speak to that. I think people always wanted to ask questions, but it wasn't possible to get an answer. So we were just adapting to a bad system. Well, I always like to use the example of this old website called Ask Jeeves, which was an attempt to allow people to ask questions and get answers. And it didn't work because the technology didn't work. But people kept using it and always said they liked it because they liked the idea of being able to get, ask a question and get an answer. And I have some old screenshots of it. It just didn't work. It's like someone saying, oh, people like the mouse, but now we have touch screens and their tastes have changed.
7:44And I'm like, no, no, it's that we didn't know how to do touch screens. We didn't know how to do gestures. technologically. In the early days, it was a bridge to that. So often the interface we see now is the interface people always wanted, but we didn't have the technology to support it. And I'd say that's true for question answering. Now there's an exception for scholars and people doing research who want to see the documents and primary resources, but that's always been a minority. Yeah. So that brings us to the current moment where the machine behind your screen, that's going to try and answer your questions is suddenly, and I really mean suddenly, able to answer it almost like a human, it feels like at times.
8:22I agree. I think it's a sea change. So I gave a keynote talk in October to the Information Visualization Society, the IEEE Society. And in that talk, part of what I did was talked about, you know, this is coming. We are going to see, instead of people developing visualizations manually, it's probably going to be done with text interface. And that's a pretty radical thing to say. And it was a month later that ChatGPT came out. And again, I told the audience at that time that I've been in the NLP field for more than 25 years, maybe 30 years, and I've never said this is a major change. And I say it now, I was saying it right before ChatGPT.
9:01And it is transformational in terms of what we can do with processing language and producing language. It's not transformational in everything, as some of the hype says. Just like we had a mouse and then we had a touchscreen, we had keyword query or statistical ranking, or we had these very complex pipelines for making natural language processing systems. And now it's kind of one relatively simple architecture that does everything as opposed to specific algorithms. And it's kind of head spinning, really. Well, simple schematically, but very complicated in terms of what structure might be hidden in all those billions of neurons.
9:42Yeah, it's simple in terms of what the people have to do and complex in terms of what the program is doing. I actually have an example that I was just trying last night because in the same talk, I gave an example of comparatives being very difficult to process automatically. What's a comparative? So if you have, say, a review of a camera and someone in their regular casual language is saying, oh, the DLSR has a wider angle, but the pixels are not as crisply retained. What are they saying is better than what? There's a lot implied there. And there's an implicit comparison between kind of the overall merits of some camera and then the specific components, the pixels and so on.
10:23And I use that as an example of something that it would be very hard to write an algorithm to process automatically. And one of the reviewers of the paper that I wrote said, yeah, that was true, but I just put this in ChatGPT and it worked really well. So last night I put all these super complex descriptions of reviews of cameras in ChatGPT and it did an amazing job of saying what was being compared to what. But I still say that it would be very hard to write an algorithm to process the language to do that. Absolutely. It's a general purpose tool that does that as a side effect of what else it does.
10:56Yeah, it's sort of an all-purpose reasoning machine. It's something. I don't know what it is. So I watched your keynote and found it really, really remarkable. And something that was gestating in my mind as I watched you walk through all the latest research that you could dig up on the human computer interface and also language and the visual component of people trying to understand complex topics was that we're probably soon heading into a world where you can essentially go to a whiteboard with a model like chat GPT. So at work, when I need to understand something really complicated or communicate something really complicated or collaborate with someone on a really complicated problem, we go to the whiteboard.
11:41It's sort of the best environment to do this. And what that means is you have all the affordances of language, just speaking one-on-one. And you also have this whiteboard next to you that you can diagram things, correct things, point things out visually. And so it's sort of maximum bandwidth, and it feels like the most comfortable way to navigate really complicated things. I think that we've clearly gone way down the road of the chat side of this, the language side of this. You can interact with ChatGPT and talk about really complicated things, maybe even solve problems together. But there isn't yet that whiteboard, but I think it's coming.
12:16And we saw a hint of it with the demonstration video of GPT-4. So it seems safe to say that we're headed towards AI whiteboards and you have been grappling with the nuts and bolts of how you communicate both visually and with language and how the two play off each other, sometimes synergistically. I'd love to pick your brain just on what it's going to mean heading into a world of AI whiteboards. And of course it goes way beyond whiteboards that could show you arbitrary images, videos it generates, things it finds from the internet and actually, you know, points things out and illustrates it. Yeah.
12:53I think that there's a lot of potential for these tools, these large language model-based tools to be collaborators in thinking. I think that's what you mean by the whiteboard. Yeah. But after I'd done my keynote, I did actually ask Chachi Petit to make an outline of a talk on the subject that I had selected. And it was not very creative. It said things that made sense, but it would have been, I guess, somebody who kind of knew the field, but was not innovating, was not seeing the future. And so I don't know that it's capable of doing that yet. I listened to the interview with Sergey, and he here at Berkeley, Sergey Levin on reinforcement learning.
13:32And he kind of pointed out that it's not using technology to kind of do future sequencing, but they're working on it, I guess, or they might work on it. Yeah, no doubt the human is going to have to do most of the intellectual heavy lifting in the beginning. I mean, what is it going to mean for information sharing and explaining when we can use something as powerful as ChatGPT in the language regime, also in the visual regime? Yeah, so, and referring back to that keynote a bit, the topic is the intersection of language and visualization. Because the information visualization community focuses reasonably on how to visualize data, how to visualize information.
14:07And there's been less of a focus of how does language or text overlay on that or interact with that. And I mentioned this in our earlier conversation with John, that for many semesters or many years, I was teaching natural language processing in the fall and information visualization in the spring and thinking about what sort of information can be represented in each modality. And can you convert one to the other directly? And I think the answer is no. They show or they explain different things. Visuals explain different things than text. And if you think about the movie versus the book, that's like the best example.
14:40There are some books written to be made into movies. You think about the Harry Potter series, for example, and they're very true to the original, I think, but there's a lot that don't transfer so well. And a lot of it is about interiority and mood and things like that, that mood is expressed differently with words than with images. And they complement each other, of course, which is why the soundtrack is so important for the film. When you become a grownup, you don't have pictures in your novels anymore. Except you pointed out in your keynote that really lovely classic book by Scott McCloud on how comic books work.
15:11You pointed out that there's a method to it. There's a kind of balance between the visual and the language. And sometimes one can do most of the work and sometimes the other. Couldn't a model learn to do that? Oh, well, I could have model learned to do that. I mean, I think you could give it instructions to learn to do that. I think right now what I'm interested in is how do people understand these things and then how best to express information so that you promote understanding and you don't promote misinformation or you try to combat misinformation. I think it's really important that we understand how, and this is the human computer interaction, the HCI side of the AI-HCI coin, as I think about them, understanding how people understand things so that we know what to tell the computers to do.
15:53Right now, we have people designing visualizations and they don't necessarily know how to put the language on the design and neither will probably the computer. Or if the computer does know, we at least need to know how to assess if it did a good job or not, which I think we need to do more work on. Yeah. But just to go out one step out onto the limb, I know you're very wary of speculation, but this one feels like a safe speculation. I think that there are going to be emergent capabilities with multimodal models that can deal both with the visual and the language side. And we don't know exactly what they'll be.
16:25But if we follow the trend with GPT-3 solely on the language side, I wonder what kind of capabilities even are there to acquire on the visual side. Something that comes to my mind is simplifying something visually. Sometimes as simple as underlining something can make something salient that helps explain the whole. You have a project called Scholarfy with Andrew Head at Berkeley. Is he a student of yours? He was a student and a postdoc, and now he's a professor at UPenn. And it was also in collaboration with people at AI2, hence the Oren reference, too. Yeah. I saw a breakdown of the project. It's so neat.
17:01and one of the really neat insights is when you read something that's got a lot of complicated mathematics in it, your brain is doing a ton of work behind the scenes. And if you had a better interface, for example, click on a variable in a formula and just have it automatically pop out and say, this is what that represents. So you offload some of that cognitive work you have to do. I wonder if those kinds of skills could be learned. I hope so. And you mean the skills of visually showing the information. Yeah. All those tricks that a good visual explainer just knows how to do. Well, I am optimistic that these new models will make that automation tasks that we had more effective.
17:40We worked on algorithms to do it automatically, but PDFs are really tricky to process if you're looking at the image level. And it's very hard to find definitions within a scientific paper because not everything is defined in a crisp way. And so So really what you want to do is generate your own text, but you want it to be accurate. It's based on the text of the paper. So we are actually looking to see if the latest models can help with the automation of that task. But going back to a point you made earlier about creativity or new synthesis with these models, I think someone who was hosted earlier on this podcast pointed out just even the avocado sofa is a synergy of image.
18:17A human had to ask the query, but then the system was able to blend these images together into something new. Although it doesn't blend well if they don't go well together. In case anyone listening doesn't know what the avocado chair is, this was the sort of amazing DALI moment. So the DALI model came with a paper, and in that paper they had some images as examples of what it could do. And one of them was make a chair made of an avocado, something like that. And it was sort of amazingly convincingly good. It really was. Although they are a little cherry picked because if you try to combine two things that don't often go well together or don't appear together, it doesn't work, or at least it didn't work when I was playing.
18:56It'll flub it. Yeah. But still, it's a great example of synergy with these tools. And what I noted in the keynote was that co-pilot these systems that aid in programming rather than – there's been a long debate in HCI about when you're developing, say, a user interface or doing data analysis. Should it be a command line or should it be a graphical user interface, a GUI? And of course, the answer is neither works perfectly. And people who are practitioners use a blend of both. But it seems now, as I said about Ask Jeeves, what people really want to do is just use language to say, do this, do that, and have the program get written.
19:31And then point and use gestures and the interface to tweak it a bit, this multimodality. And again, there are tools to do that, but they're just not perfect. and the more that the algorithms improve, like with ChatGPT, the more effectively we'll be able to help people design visualizations where they don't have to do a lot of coding. It's again, because it works, this general purpose tool that we were talking about as a side effect of being, you know, produce the next word, it's able to do all these other things and I think we don't understand why, but that includes writing code or, you know, being smart about adding things into code and so on.
20:11It wasn't designed for that, but it seems like it will be very effective at making it easier to design visualizations. The problem is, will it design good visualizations? And that's where we still have the human component. Well, I think the safe way to use these things is to generate first drafts and iterate, but that you have to be the human editor who makes the final call and do the driving. Yeah, I agree. The work that we all do in the viz field can help determine what makes a good design, help give guidelines. We do research, empirical research, and then we produce guidelines for practitioners to follow.
20:45So bringing this all back to this moment, you're a natural language processing practitioner. You've spent years trying to teach machines to do useful things with language. And here we are suddenly in a moment where, I don't know about you, but I feel like, wow, a lot of the things we solved, you don't have to worry about anymore. Just sort of more and more and more of all that hard algorithmic hand-rolled feature engineering world is getting eaten up by large language models that can simply speak. And they seem to have cognitive abilities that we would never have dreamed would be in a machine.
21:23Well, I'm not going there with you on the cognitive abilities that I'm very cautious and skeptical about. What should we call them? Behaviors? I guess I don't have my favorite word for it yet. capabilities. It works a lot better than it used to work. There's a lot of people looking into, you know, why does it work? But, you know, each time people start to make some progress on that, then a new model comes out that's even harder to understand because the scale is so much larger and we're not good at thinking at a very large scale. So I think it's going to take years before we understand what's going on.
21:52I don't think it's cognition. I'm very skeptical about that. It's really, I mean, you know, we get into philosophy and the Chinese room, that's an old John Searle thought experiment. I mean, it's unfortunate, I guess, the use of Chinese in that particular example. But the idea being if you replace each piece of your brain with a little like component, electronic component, and you eventually replace every piece, is it still a brain? You know, are you still thinking? There's philosophy thought experiments. You might want to say, oh, this model that's basically just a bunch of numbers, a bunch of weights that have been trained is thinking because you can say that about the brain.
22:29But, you know, I'm not convinced. I think there's a lot more going on in the brain than is going on in these models. I agree with you. They're very good at mimicking, you know, at producing language. And because language is distinctly human, it feels, you know, to a lot of people like it's human. When people are driving in their cars, they name their car, even old cars that had no electronic components. They would name their cars. They would anthropomorphize their cars. They feel a part of their cars. This is what we do with technology. People are going to get used to it and then it's going to become old news.
22:57And I think it's great that we don't have to write all these tokenizers. The LP pipeline didn't work. It was a mess. And there's always new problems and new questions to investigate from a research perspective. Researchers will not be out of business. Of course, it does raise even more societal issues and dangers because of the ability to fake information, to spread misinformation, and for people to not know what's real and what's true. So we're living through a very chaotic moment right now. I think we're going to look back 10 years from now and we're going to go, wow, that was a chaotic time in technology.
23:31That was a chaotic time politically. And hopefully we'll be able to look back and say, thank goodness we made it through. Okay, I'm optimistic we will. Well, the curve you're describing is pretty smooth. It implies that there's going to be another side to this. But if things keep exponentially changing, there won't necessarily be that moment because it'll always feel like it does right now. Well, the technology that this breakthrough with these models and really with training on huge amounts of compute, huge amounts of data, I think it can only go so far, right? We don't know the limits of it, but it's not going to be everything.
Read the full transcript
24:07If you look at people that are trying to study the brain, you know, it just, there's other things going on there, different kinds of structure and so on. You think we're running out of data? No, no, I don't think that's it. I think that the technique, it's a very specific technique. That alone, I don't think it's going to be sufficient for being the same as humans. I'm not saying we could ever do it. Some people do argue that sequence prediction, which is essentially what is driving this whole craze, might be all you need. What do you think about that? Well, it's certainly all you need for certain tests.
24:38We're seeing that now. It's really quite amazing. There is sometimes fine tuning on the other side. But again, who knows? I personally have been wrong about this particular technology. I think like a lot of people, I just didn't know how to think in terms of billions of parameters and we're just not good at that. There were some very ambitious people that just sort of went for it and surprised all of us. I admit it, I did not see this coming and I was surprised by it. And we have certainly in the research community, it's been developing gradually. So Word2Vet came along. So going back even farther, again, when I was doing early in the statistical NLP time, people were looking at SVDs, singular value decomposition and LSA latent semantic analysis, which is similar in a lot of ways.
25:21It was putting words in a matrix, well, words by document matrices and trying to find similarities. Even before that, I was trying to solve the thesaurus or the synonym problem to help with search. So in search, you look for cat and it's really feline and you don't find anything. Going back to the beginning of our conversation. And you didn't want users to have to put in every synonym imaginable for a cat just to find text about cats? Well, library catalogs had synonyms in the early days. They weren't that good. It weren't dynamic. They didn't handle new technology and they were hard to use. So WordNet came along and developed as a linguistic tool.
25:55And I was the first person to download it actually when they had an FTP available. Did work on that. And it was like, oh, we can have a thesaurus. But it never worked. Whenever you had automatically recommended terms for a term, some of them were right and some were wrong. And that was true of SVD and LSA as well. They worked in some cases, they didn't work in other cases. And it wasn't until Word2Vet came along and then people actually then refined it to have different senses, that it actually started to work. And so I was saying, wow, this actually works. And I've seen 20 years of this not working.
26:28And of course, that kept, being refined and being made more sophisticated with the transformers came along and now the really large things. So in the research community, it's been happening gradually. There were a lot of debates about counting versus probabilities and all this. So it's not out of the blue, but I do, again, I admit that in this last year between, you know, the combination of the image plus text generation and these language models where the input could be text. We never thought the input could be text and then the output would be all of these things, right? We thought we had to program things.
27:02And I don't think the people who developed these models expected that either. I believe it was a surprise to them. So it is different now. I don't think everything's solved. I don't think it's AGI, but the tools are much more effective than they used to be. Well, like you said, it's all about capabilities. And it turns out if you teach a very big neural network how to predict the next word on a huge amount of internet text, all these really neat emergent capabilities come into your hands. Couldn't have been predicted. In fact, no one really thought it would work as well as it does, I'm sure, but it does.
27:36I wonder what happens when you teach a model to predict the next image in every YouTube video. What capabilities emerge? Yeah, it's going to be interesting. It should be much better at generating video. I mean, I guess there's already work on generating videos. I feel like generating videos is kind of like that unicorn story moment. So you remember in the early days of GPT-3, when they were trying to show how great it was, they said, look, you can start the first sentence of a story about something that it definitely has never seen. It was something about unicorns and it could just write a story and it's coherent and it's a story.
28:10Well, I think there will be that video moment where you start with an image and you just say, hey, finish this, make this a one minute video from this scene. And that'll happen. But just like with GPT-3, the thing that's going to blow us away are the things we can't predict it'll be able to do. It's going to have capabilities that just emerge. Yeah, well, I have to admit that I was not at all impressed by the unicorn story. And in fact, that's why I was skeptical. I was like, this is clearly cherry picked. And it's like, you know, from a fairy tale and you put anything else in. It's just not useful.
28:42Well, when you do NLP, right, there's different kinds of tasks. Some tasks are easier to evaluate than others, like information extraction. Did you identify the right that accompanies an organization or is it a rock band or whatever? But if you are doing search, it's very hard to know if you have the best ranking in a lot of cases. Or if you're doing summarization, there are many legitimate ways to summarize a paper. And so it's really hard to evaluate summarization. And if you're generating a story, you can generate almost anything and it's a story. So this is why I was not at all impressed by the unicorn example, but it turned out that actually there was more behind it than the cherry picked example.
29:18Although GPT-3... I wasn't impressed with GPT-3 myself until the Instruct GPT version came out. and the thing actually did your bidding. Yeah, well, they improved on it. Yeah, they really improved on it. The problem is initially they were hyping it in ways that weren't helpful. And I know that now they're being more careful. I mean, OpenAI. So, or maybe they did have more behind the scene. They kind of said, oh, well, we know stuff that you don't know and we can't share it. So you want to see, you want everyone to be able to test things. You know, that's what happened with the fake blood testing company and all that.
29:48It was clear from the beginning it was fraud. So you have to, if you're going to make big claims, you need to be able to show your cards. Yep. All right. So zooming out a bit, what do you think is going to be the most exciting things to pay attention to on the research side of your fields? You really have more than one field, but I'd love to just hear your thought. What's in your mind these days, given this kind of big sea change as you describe it, which I agree. Well, there's a lot of people doing a lot of stuff. So there's a lot of people really interested in AI safety and AI anti-bias, all very important.
30:23I think there's also a lot of people looking at the AI human interface, which is something I've been interested for a long time. And that's super important. People doing driving car, self-driving cars have a bit of a headstart, mainly on seeing how hard the problem is. Actually, I had a PhD student, Cecilia Aragon, who looked at projecting LIDAR visualizations for helicopter pilots on the screen. And how could we make that work and have them not crash? Because this could show them, but say a squall was ahead and they might potentially crash if they went into it. And we found that the simplest, most bare bones interface was the very best so that they weren't distracted.
30:57So I've always had questions about self-driving cars and that problem of the attention of the driver and that's really not solved. And the studies I have seen on automatically generated language and interfaces, even some we've done in the ScholarFi, Semantic Scholar project, semantic reader project, we don't have good answers for that. People just start to rely on the automatically generated output. It's natural. And so that's a huge problem that needs to be solved. What are some of the ways that we could help people? If everyone comes to rely on ChatGPT for day-to-day work, what are some of the levers we can pull to help them?
31:35I haven't solved this problem. I mean, it's certainly good user interface design, understanding people. So HCI shows us how to study people and how they work and how they work with technology. So using HCI methods to deeply study that and in different contexts. It's different in a medical setting. There's a colleague that, Ilifar Salehi, here at the iSchool at Berkeley, who is looking at machine translation in a medical setting and when information is not translated correctly and how that can adversely impact marginalized communities when the translation isn't really the right thing and how to get the context right.
32:10So I think each setting is probably going to need some specialized research. And furthermore, one of your podcasts is about how do people inject poison, the training data, and so on. And so you're going to have to be very careful about attacks like that. It's not a field that I'm in. Perhaps it'll be important to have diversification in the different models so that there's ways to check them, make sure that they are safe and appropriate for a particular use. I think the techniques of HCI and ethnography really work independent of the context. It's not AI. People in AI don't necessarily want to sit down with people, humans, and see their details and what they do and so on, but that's the only way to have really working systems that are good for society.
32:52Yeah, I've noticed that there's this kind of mindset of people who build AI, generally people who build AI systems. It's the engineering mindset. The more you can take the person out of the equation, the better because, you know, I want my development environment to be nice and clean and straightforward and I want to build something that I understand. And as soon as you get people involved, oof, people are complicated. But what you're saying is you have to include the person. Yeah. And that's why I'm heartened by the new interest in NLP plus HCI. I've given, you know, there's been workshops I've asked to talk at because I've been thinking about it for a long time.
33:25But it's true that a lot of people who are making the biggest advances in the NLP AI field are mathematicians or physicists, and it's just not what they think about. They like to think abstract way and they're brilliant and they're really improving these systems. But we need teams to work on technology. The one person band just doesn't exist in this space. So what would you say to students coming into these fields just now? Has the advice changed at all? It's hard to know what to say right now. It's moving so fast. What I say to all students is, What are you passionate about? What really interests you?
33:59Do that. Don't do the trendiest thing for its own sake. It's definitely a big question mark right now for research universities and AI labs, other research groups. If there's a big microscope that some people have and you don't, how do you compete? And there's open source effort, things that a hugging face and others are doing. I think the US government's interested as well in giving everybody a microscope, meaning these large language models and the ability to run them. And I think that people are very aware of that issue. But then if you want to do advanced research on this, you know, you get brilliant people like Agent Choi at University of Washington and AI2 that are showing that you can do a lot with much less.
34:38You don't need to have all these parameters and so on. And that's where the university can help. The university people are also going to be looking at how do you save energy? Well, I mean, so is industry, but how do you save energy when you use these? It's very wasteful right now. These are all great areas for research. And of course, understanding what these models are doing, understanding the mind better. Can they help us understand the mind in some way? I'm sure psychologists are thinking about that. There's work already. I've seen work in linguistics on, say, using GANs and adversarial methods to model linguistics in other species.
35:10So there's always more research questions. And if you're interested in being in industry and business, then go do that. And if you're interested in research, then find a problem that just really interests you because that's how you can finish your PhD. And just to bring it to a close, what's coming up in your life that listeners might be interested to know about? Is there a project or an event on the horizon? Well, I think the project that I'm most excited about is working with my PhD student, Chase Stokes, on understanding this interaction between language and visualization. And we just finished a paper that we submitted on if you place text on a chart and the goal of the chart is to make a prediction, how does text impact that prediction?
35:51Like who's going to win an election by looking at this chart? And we actually found surprisingly that the text did not influence the prediction all that much. In this case, where people relied more on the visual input. But in another study we did where it was more, what are you taking away information-wise, then the way the text was used did have an influence. So what we really need to do is understand this interplay more. And I'm just very excited about that topic. I know it's kind of a niche topic, but it's what interests me. Oh, far from niche. We're going into a very momentous political year, and these little interactions between a person and a piece of information can have massive, massive effects.
36:29You're right. We don't really understand how they work, do we? No. I like to work in topics where there isn't a lot of work at that time, like search interfaces and so on. And then when it becomes popular, I tend to move on. I think I can't compete or something. So I have to find a new thing that nobody's thinking about. Uh-oh. I might've just ruined your picnic. Now everyone's going to get interested in language and visualization. No, no, I want that. In fact, the keynote I gave, I think, you know, helped with that. So I want people to be working on this. Marty, thanks so much for talking with us.
36:55It was a pleasure. Really fun. All right, everyone, that's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimbleai.com. Of course, if you like what you hear on the podcast, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening and catch you next time.
From the publisher
Today we’re joined by Marti Hearst, Professor at UC Berkeley. In our conversation with Marti, we explore the intricacies of AI language models and their usefulness in improving efficiency but also their potential for spreading misinformation. Marti expresses skepticism about whether these models truly have cognition compared to the nuance of the human brain. We discuss the intersection of language and visualization and the need for specialized research to ensure safety and appropriateness for specific uses. We also delve into the latest tools and algorithms such as Copilot and Chat GPT, which enhance programming and help in identifying comparisons, respectively. Finally, we discuss Marti’s long research history in search and her breakthrough in developing a standard interaction that allows for finding items on websites and library catalogs.
The complete show notes for this episode can be found at https://twimlai.com/go/626.




