Pushing Back on AI Hype with Alex Hanna - #649

2 Oct 2023 · 49 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: The TWIML AI Podcast - Episode #649: Pushing Back on AI Hype with Alex Hanna

Host: Sam Charrington Guest: Alex Hanna, Director of Research at the Distributed AI Research Institute (DAIR) Release Date: [Insert Date]

Episode Overview

In this episode, Alex Hanna discusses the pervasive hype surrounding artificial intelligence (AI) and the implications of this hype on society. As the director of research at DAIR, Alex highlights the need for responsible AI practices, robust evaluation tools, and frameworks to mitigate the risks associated with AI technologies while promoting community-centric research.

Key Topics Discussed

  1. Background of Alex Hanna
  2. Alex's journey to AI began as a sociologist using machine learning for research on social movements.
  3. He transitioned from academia to Google, focusing on machine learning fairness and ethical AI before joining DAIR.
  1. DAIR and Its Research Agenda
  2. DAIR was founded to focus on AI's social implications, emphasizing community engagement and ethical practices in tech development.
  3. Key projects include:
  4. Machine Translation for Low-Resource Languages: Work on Amharic and Tigrinya languages in collaboration with Lesan.AI.
  5. Spatial Apartheid Project: Utilizing computer vision to study segregation in South Africa.
  1. Do Datasets Have Politics?
  2. Alex discusses a research paper that analyzes the politics embedded in computer vision datasets.
  3. Key findings include:
  4. Developers often prioritize speed and universality over careful consideration of data collection ethics and representation of marginalized groups.
  1. AI Hype and Its Implications
  2. The current AI hype cycle has been exacerbated by the release of tools like ChatGPT.
  3. Alex argues that hype can lead to dangerous applications in sensitive domains like healthcare, where AI tools may provide harmful advice.
  4. The historical context of AI hype, as discussed by figures like Joseph Weizenbaum, reveals a pattern of inflated expectations leading to societal harm.
  1. Critical Applications of AI
  2. Alex emphasizes the need for a critical approach to AI applications, particularly in healthcare, where unregulated tools can cause significant harm.
  3. He advocates for a nuanced perspective, recognizing that while some applications of AI can be beneficial, many are poorly evaluated and can lead to harmful outcomes.
  1. Frameworks for Evaluating AI Use
  2. There’s a lack of comprehensive frameworks for responsibly implementing AI tools across different sectors.
  3. Alex suggests that evaluation should be context-specific, drawing from professional and academic communities to assess risks effectively.

Key Takeaways

  • Community-Centric Technology Development: AI should be developed with input from affected communities to ensure it meets their needs.
  • Importance of Data Ethics: Data sourcing and representation are crucial to the legitimacy and fairness of AI systems.
  • Caution Against Unchecked AI Deployment: The hype surrounding AI necessitates a sober evaluation of its real-world implications, particularly in critical areas like healthcare.

Resources Mentioned

  • DAIR Institute: [dair-institute.org](http://dair-institute.org)
  • Teheku Media: An organization focusing on indigenous language technology.
  • Lesan.AI: A startup working on low-resource language technologies.

Conclusion

The episode provides a critical examination of the current AI landscape, advocating for responsible practices and highlighting the social responsibilities of AI researchers and developers. Listeners are encouraged to approach AI with a critical mindset, understanding its potential impacts on society.

For complete show notes, visit [TWIML AI Podcast Episode #649 Show Notes](https://twimlai.com/go/649).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:09All right, everyone, welcome to another episode of the TwiML AI podcast. I am, of course, your host, Sam Charrington. And today I'm joined by Alex Hanna. Alex is Director of Research at DARE, the Distributed AI Research Institute. Before we get into today's conversation, be sure to take a moment to head over to Apple Podcasts, Spotify, or your listening platform of choice. And if you enjoy the show, please leave us a five-star rating and review. Alex, welcome to the podcast. Thanks for having me, Sam. I'm looking forward to digging into our conversation. It's been, I think, maybe a year and change since I spoke to Timnit and kind of the, I think she had just started DARE or you all had just started DARE.

0:54And so I'm really looking forward to kind of learning more about what that year and change is, you know, what you've been doing in that year and change. But before we dig into that, why don't you start us off with a little bit of introduction? How'd you come to work in AI and AI ethics in particular? Thanks for the introduction, Sam. And I'm looking forward to talking about there's almost two year history at this point. Is it almost two years? Wow. It'll be two years in December, December 2nd. Okay. And yeah, so I can tell you a little bit about myself. So I come to AI in a very roundabout way.

1:31My training is as a sociologist and there's not a lot of social scientists within AI, although there should be more for many reasons. I feel like we'll be discussing that a little bit. We will be discussing that for sure. I got into AI kind of through the back way. I was using a lot of machine learning methods actually in my dissertation. And so I was using a supervised learning technique to identify news articles which mentioned protest. And this is because in some of the work that I currently still do and work that I had been doing earlier. An interest of sociologists that study social movements is identifying kind of the who, what, when, where, why protest for the instance of identifying, you know, what motivates protest or what makes it happen?

2:21What are the demands of folks and how do they win? And so that got me interested in using some automated methods to look into that. After that, I was a professor briefly at the University of Toronto. decided that the academic track wasn't for me. I went to work at Google, initially working as a curriculum designer within machine learning, but was still very much in the conversation around machine learning fairness and algorithmic discrimination. So I connected with some friends of mine working on that. I got more into that space, learned much more about it. And one of the things that I was focusing on was how much of the time within the conversation machine learning, there was very little attention paid to data.

3:05Coming from a sociology background, much of the focus is how data is collected, how data is constructed, how that data may or may not have validity, what kind of errors happening in measurement and operationalization. And so that led me very much into focusing on these things. So I was getting involved with many of the academic communities around fairness. FACT, for instance, is a large conference, had been going to the conference since 2018 and started going every year since then. Then at Google, I eventually transferred to work with doctors Taneet Debrou and Meg Mitchell on the ethical AI team that they had constructed at Google.

3:51I was very excited to do so. At Google, I was the first research scientist that was a social scientist ever hired on that ladder. And that opened the door because many social scientists now work at Google Research focusing on these issues, which is great and I'm happy that it opened the door for many, many other folks. And so everything happened at Google. We know that story. We'll go into it. Taneet was fired. We'll refer back to that previous podcast for more on that. Yeah. Go back to the podcast on how everything happened there. So everything happened, Taneet went and started DARE shortly after DARE's announcement, December 2021.

4:31I joined three months later in February 2022 as director of research, employee number three. So yeah, so that was my background. That's what led me to where I am now and where I'm at at DARE. Awesome. Awesome. Did you start to craft a research agenda from that blank slate? I would say that in terms of doing it, there are two things. I mean, again, reiterating, it wasn't quite a blank slate just because a lot of the work that we were already working on, especially around data, data documentation, was a bit prescient, right? I mean, people are still, if you saw the news today, Tanit and Meg and Emily Bender all were represented in Times AI 100, and notably for the prescience of the stochastic parents.

5:24Not stochastic parents, stochastic parents. Stochastic parents. Some parents are stochastic. It feels like that sometimes. It does, right? If you're a teenager, parents are very stochastic. And so because of that paper, much of that was pre-send. But I will say in another register, a lot of the ways we had set the research agenda is by bringing on researchers and especially research fellows that had a research agenda already. And so I say employee number three because someone who is already at DARE is our fellow Rasecha Safala. Rasecha is a grad student now at Mila in Montreal, and her work at DARE was on the spatial apartheid project.

6:14So that project, and I imagine Janine may have mentioned it when she was on this show, is that that project was using computer vision technology to detect the persistence of segregation and the persistence of spatial apartheid in South Africa. The history of South African apartheid is that Black people and non-white people, so-called colored people that were not classified as Black, but typically Indian and Asian, were separated out into township areas. And in those township areas, those are separate from the neighborhoods in which the wealthy white population lived. And so even though South Africa formerly abolished apartheid in the mid-90s, there's been persistence of that.

7:04even though the census in South Africa does not maintain divisions between townships and neighborhoods anymore, Rassetcha's work revealed many of the persistence, even though, I mean, this is well known that this is persisting, but now we actually have a view on this, where this is. And so those data can then be used to identify things like how long it takes social services to get to certain people, the amount of schools, the amount of hospitals, the kind of time it takes for ambulances to get to a certain place, et cetera, et cetera. And so that itself is a research agenda that we're still pursuing and multiple people are pursuing.

7:43One of our full-time people and also an author on that paper, Nyalin Morosi, who's based in Los Soto, has been working with Resecha on that. And so fellows have come in, and I think in the areas that we want to focus on and have set research agendas by being able to say, you work at bringing on because this is an important dimension that we're focusing on and we want to empower you to focus on that and do work on that. So that includes people like Asma Lashteyka has been doing work with Lisan.ai and developing language technology that works for the Horn of Africa. Adrian Williams, who is a former charter school teacher and Amazon delivery driver who's focused on wage theft via surveillance of Amazon workers, especially drivers and flex drivers.

8:34Crystal Kaufman, who is an organizer with Tricopticon, has also written and done organizing around the rights of data workers, the people who are fueling all the data that goes into AI. And so we bring in folks because we know they have those expertise and we let them do what they need to do. So in terms of coming in with kind of a green field, the sort of research agenda, is that people already have these knowledges, whether they're academics, whether people with lived experience, and we bring them in and help them build those skills and publish original research and work on that. And I certainly spoke with Timnit about this and talking about DARE, But, you know, how do you articulate kind of the common thread that runs through these various research efforts that you described?

9:25What does DARE care most about? I think the common thread is that we are focused on the notion that AI is not inevitable, that it could be a tool that would be useful in some contexts. And those contexts tend to be rather narrow in some guises. So for instance, things like machine translation or automated speech recognition. Those are actually pretty useful technologies. They could be useful assistive technologies. They could be useful in expanding the scope of people that could use computing. It could provide different interfaces for people who maybe have a very hard time typing. Asma Lash and Timmy and I have this paper, Work in Progress, where we talk about an internet for our grandmothers, where Asma and Tanit are talking about their grandmothers who couldn't read or write.

10:24I'm thinking about my grandmother who is speaking in English, spoke Egyptian Arabic, even Coptic in the home. These are languages that aren't really quite well supported in in different kinds of technologies. Even if Google or Meta says that they have these automated speech technologies that work well for these languages, they work quite poorly. And if we were able to provide interfaces for different modalities, then that could be a great use for AI. But right now the use for AI is kind of going all into these different kinds of very extractive uses for large language models. We're trying to automate it at work, trying to put people out of jobs, trying to do it in such a way to threaten the labor of many people, the things that many workers and many people aren't asking for.

11:13So the common thread really is finding technology that works for people based in our communities. And also, I would say the second thing is acknowledging that knowledge that comes from communities is a form of knowledge. It's a way of knowing. We use this big term in the philosophy of science sometimes, epistemology. That means how we get to know certain things. And I think one of the things that we really thrive with and dare is knowing that there are multiple different ways of knowing. And that could be lived experience. That could be a PhD. That could be both. But acknowledging that is where we start from.

11:56And how does that particular point kind of play out in the research? I think it plays out in the research of seeing how people set agendas here. Again, where we came into this project with people into the organization, bringing people, bringing fellows in and saying, determine here what is the most important thing. Okay, we want you to write this out. How is this the most important thing? Write this out. Okay, let's talk about what it means to develop research on this. There's a quote from General Gordon Baker from the Revolutionary Black Workers in which he says, our focus is to turn thinkers into fighters and fighters into thinkers.

12:36And I absolutely love that because the kind of thing that I'm thinking about, he's talking a lot about turning organizers and having them go through political education, but also people who are learned, we need to bring them into advocacy. And I think about that a lot at DARE because I think about a little twist on that. How do we turn researchers into fighters and fighters into researchers, right? We have these people that are being brought in that have these huge wealths of knowledge from how they are, from everything that they've experienced through their labor, through their activism. How do we turn them into researchers and really bring in that evidence that has both the kind of legitimacy within academic domains, but also is going to be stuff that is useful for this kind of goal of ours, this north of ours, to build technology that works for people.

13:33When you think about building technology that works for people, can you give us some examples of projects that kind of squarely focus on that particular goal and some of the outcomes of those? Yeah, so I want to revisit the work I mentioned earlier, Asma Lash's work at Lassan, in which he's been focusing on building machine translation and automated speech recognition tools for two languages that are on the Horn of Africa, Tigrinya and Amharic. These are languages spoken by, I think, Amharic is spoken, I want to say, by 20 or 30 million people in Ethiopia. Tigrinya is spoken by about 3 million people in the Tigray region of Ethiopia.

14:21And these technologies work very poorly in, when you look at Meta's work or Google's work on this, and you use those tools, they actually work very poorly in doing those things. even if they advertise that they've done so. I think Meta even had a video in which they advertise their ability to translate or do automated speech recognition of Amharic. And Asma Lash did an analysis of that and found that it worked very poorly. And so developing those tools with the cooperation of people in those communities has been critical. So he's been working on the development of those tools with the cooperation of those speakers, sourcing those data in ethical ways, checking it with people from the community for community use.

15:12In some ways, we're also very inspired by other efforts in these directions. Teheku Media, for instance, is a organization that's based in Eritrea or New Zealand in which the people involved are all from the Tia Raiari, majority indigenous community in New Zealand. And they've done a bit of work. They're not an academic group per se, but they're both kind of an indigenous and traditional cultural knowledge preservation group and an engineering group. And so what they've been doing is focusing on collecting data from indigenous elders, from speakers, Aftey Ryo Miori, and being able to develop machine translation and automated speech recognition tools that work for that community.

15:57They've compared this with other tools that have been released by, for instance, OpenAI and found how poorly it does in that language. Something else that they've done is they've safeguarded the data that they use to train those tools because those themselves are considered under a certain kind of data sovereignty that they want to keep and maintain. And so we take a lot of inspiration from that project. And it's something that we kind of bring in and thinking about as something that should be exemplified as a way of building tech for people that works. You mentioned in that approaches to sourcing data in a way that is, I guess, more fair to the folks that are contributing to these data sets.

16:41This is something that you've written about quite a bit and is clearly core to what DARE is trying to, well, A, raise awareness around. And maybe that's kind of the way to start the question. Like when you engage around kind of this conversation of data collection, is it primarily an awareness raising thing? Is there research that goes into that? Is there a way to study that as a phenomenon to drive change around it? Or is it primarily kind of this building of awareness? Well, there's definitely ways of building research around it. I mean, much of my research is focused on where these data come from.

17:17This paper that I wrote with Morgan Klaus Schaerman and Romy Denton focused the name of the paper is Do Datasets Have Politics? That paper focused a lot about the sourcing of data and what computer scientists and other related researchers focused on important value-related in datasets. There's been other studies, including the work from Pang and the author's first name and Arvind Naranian, that focus on the afterlives of data that have been considered unethical. They focus on three data sets that have been retracted, many different versions of this data still out in the wild. And a lot of it is research into what exists.

18:04The effort is to change practices. and changing practices is very difficult, but the first part of it is understanding how pervasive the problem is. And so some of the changes have been proffered. So for instance, NeurIPS now has a data set and benchmark track. That data set and benchmark track is an effort to bring more data and have data be a valued contribution in its own right, especially data that meets a certain bar of ethical commitments. Every paper in that track needs to have a data sheet, needs to be explained, needs to be open. If there's any restrictions, it needs to be explained. But this is still not the case.

18:48I mean, that is any organization that wants to go to market with something like ChatTPT or any of these other tools released by large entities don't have this reasoning to release those data. They don't have the compunction or anything compelling them to be transparent or release those data. They hide behind promises or excuses of trade secrecy or if there's someone that would build a nefarious version of the language model or whatnot. And I find those arguments to be disingenuous. I mean, there needs to be transparency into what those data sets are and how existing kinds of problems, bias, quote unquote, hallucinations, although I hate that word hallucinations, or like misinformation and falsehoods, they're perpetuated from those training data.

19:45We just have no way to do any auditing of that. So awareness is one aspect of it, but it's also changing scientific practice and developing regulation and legislation that is going to protect different subjects in the data set development and the model development process. Can you talk a little bit in more detail about the do data sets have politics paper? Kind of what's the, provide an overview of that paper and the kind of the setup and the issues that you investigated there. Yeah, absolutely. So this is a paper that we wrote two, two years ago, started three years ago. And the focus on this paper was looking at a particular sector of machine learning, most specific, more specifically computer vision.

20:32so what we attempted to do is construct almost a population level overview of all computer vision data sets that we had come up with so we looked into many different methods of this we looked at citation patterns i don't think we looked at papers with code for this one but we did search every we searched ieee archives we tried to find basically every type of image data set that we could within the past, I want to say a decade or two years, the past 20 years, we found around 500 to 700 datasets. I think we started, we had 500 initially in the population. We took a 100 dataset sample of that in addition to the 14 most highly cited datasets, and we coded them for 100 different variables, and then also did a qualitative analysis of things that computer scientists valued in those constructions of data sets.

21:36So let me explain each of those things. In terms of the data sets, the different kinds of data that we focused on and coded for, we coded for where did the data come from? Is there any licensing around the data instances? Are there any people in the data set? Can you identify people from their faces? Do they have any consent around it? Was there any licensing around it? As for the data itself, was it held on a repository that had restricted use? Is there any privacy considerations mentioned in this? Is there any ethical considerations? Is the data even still available? Can we access it? Can we audit it?

22:13Are people being good data stewards? And then as for the qualitative variables in terms of the values of the data set, the kinds of things that we focused on is, is there any value-related language around this? So for instance, if the author of a dataset writes something like, we use this dataset because we wanted a larger dataset because larger data means that it serves us a better benchmark, or we wanted to have people in multiple different poses so we could have better out-of-assemble fit, or yada, yada, yada. These are the types of things in which people are making a value judgment on that.

22:48And we did an analysis of that. We used a method that's commonly used in social science called grounded theory, in which you look at lots of different texts, you see what themes emerge, and then you bend them into different categories. And we found four common themes. We found that data set developers typically focused on universality compared to particularity. They want to try to cover every single instance. That, however, has the problem where you may have people at the margins fall out of the data set. You might have people who are quote unquote edge cases, not actually be included in the data set.

23:29And that was at the cost of having particularity, of having things, of having a narrowly scoped problem in which people are, problems are well defined. We found a intent towards speed rather than care. So we need to get this thing, we need to collect this as much as possible. We need to quickly label these. We relied on Amazon Mechanical Turk workers rather than having people or experts judge these things and take the care needed to treat these data with a certain amount of care. That was a common theme. We found the last two things are a focus on impartiality rather than positionality. So focusing on trying to have a data set that would say unbiased or this kind of mythical unbiased biases of a data set.

24:14We know that all data sets, as the title suggests, have politics. They have a particular view of the world. And actually acknowledging that view is what comes at the cost of claiming impartiality. And then lastly, a focus on doing the work of building the model versus the work of building the data set. So, so much of the work focused on the building of the model. This reflects a lot of the qualitative work that my former colleague at Google, Nithyan Sabasian, has also shown in doing interviews. People want to do the model work. They want to build these models, these top-line metrics. You beat state-of-the-art rather than a slow plotting work of doing data, ensuring that this stuff meets a criteria of quality, that people have been paid sufficiently, that you have consent.

25:03Consent is needed and obtainable. Nobody wants to do that data work, which is much slower. And you can even see that by the volume of papers given to describing data sets. A new data set may be released and it gets two paragraphs and an eight page paper. Most of the paper is spent describing the math and the methods and how this paper beats your state of the art. Got it. So maybe shift gears a little bit. One of the topics that you've been kind of outspoken about recently is all the hype surrounding AI. I think listeners of this podcast will be familiar with that hype. You know, a lot of it's come about since the release of ChatGPT.

25:44So, you know, we're nine months into it, into this latest iteration of the hype cycle. This latest iteration of the hype cycle, you know, AI hype has been an issue for a while, but it's, I think we're at new levels here. Why is the hype cycle kind of an interesting and important thing to, or the level of hype, an interesting and important thing to talk about and highlight for you? Well, it's really interesting how we came about this, because I think the papers that we were writing around data, a lot of it came out in 2021, where Emily and I were writing and thinking together with the larger people.

26:23And so this was really launched when Blake Lamone, who I did work with at Google, was fired by Google for claiming that this model was sentient Lambda. But shortly after a VP at Google Blaze, he wrote this very long, literally, I think, near 10 ,000 words, maybe 15 ,000 words on this kind of idea of AI sentience. and didn't refute any of Blake's claims, but was effectively giving some credence the idea that these large language models were sentient. Same thing with Ilya Switzkever, who said something of the nature that large language models are slightly sentient, and then decontextualized sort of tweet.

27:11Sam Altman giving some credence to this and saying that I am a stochastic parrot, and so are you. And so I'm like, wow, let's dig into this. And I want to pick up on something you said, because I've been reading a lot of history of AI lately. And AI hype is not only new, it's actually quite very, very old. It's probably as old as AI itself, right? And I want to give two shout-outs here, one to Abebe Berhani and a second to Ben Tarnoff. Abebe Berhani had a piece in Real Life Magazine called Fair Warning. And a lot of it was about, it was a reading of Joseph Eisenbaum's Computer Power and Human Reason, very much dealing with this kind of idea of AI hype and the kind of risks that we have at this.

27:59Ben Tarnoff has written a longer piece for The Guardian going into Weisabomb's life, going into the way that he had been a person that was digging into, you know, this is the person who wrote the Eliza chatbot, right? And he was struck at how many people were fooled by this thing, that people were really taken by a few simple rules given to this chatbot written in the 1950s and how it did a few different things. One of the things that it did is make people panic or hype, depending on which side of the coin you are, what this would do to your jobs. Eliza was purported to be a Rogerian psychologist.

28:44Many psychologists, Eisenbaum writes in computer power and human reason, or even saying, well, this thing is going to take jobs. We're actually going to be able to have a psychologist in every hospital, and this can take on any number of patients needed. And he was very struck by that. And it did two things that he argues. One of the things it did is it produced this amount of hype. And the second thing it did, and kind of unreasonably so, the second thing he felt that it did is it devalued what it means to be human and devalued what it means to be a particular species at this point in time. And so Weizenbaum was very critical of AI boosters.

29:25He was at MIT. He'd done arguments with Marvin Minsky, the head of the AI lab at MIT. And Minsky was just taking oodles and oodles of defense funding to develop these different tools without being very critical or reflexive of these kinds of operations. And so why is it important to tackle AI hype now. One, it's just out of fever pitch. It seems that you can't turn anywhere the same way that two years ago or three years ago, everywhere you looked was blockchain or crypto or NFTs. AI is being deployed in every which way. Every kind of thing since ChatGPT has become something available to mass market users, you effectively are seeing new and horrible ways in which someone thinks, let's slap a chatbot on it and use it in some business or social service use case.

30:21And so there needs to be someone out here countering those breathless claims. And that's where Emily and I see our role is really taking these with a really sober mind and addressing these. You mentioned new and horrible use cases. Are there some that come to mind for you? So many, Sam, so many. The things that horrify me the most are really the medical use cases. Those cases, I mean, take this as a page right out of Eisenbaum again, but it's those cases in which things are being used for talk therapy or being used for people who are in mental health crises. Recently, there was actually an article published today in The American Prospect about the National Eating Disorders Association and how in the face of their unionization efforts, the whole staff was cut for a chatbot named Tessa.

Read the full transcript

31:18Tessa was doing things like providing, was quickly taken out of commission after they found out that it was giving advice to people like weight loss strategies, things that people with eating disorders don't need to hear in marked contradiction of the kinds of things people in crisis need to hear. The same thing has happened with doctor's services and diagnostics. Martin Schreckelli, the guy who got arrested for jacking the price of insulin. Oh, what's the drug? Yeah, he posted on Twitter some AI tool called DrGupta.ai that was supposed to be helpful as a diagnostic. And this has been done for other more reputable firms as well.

32:01Google said that there, MedPalm 2 was being tested at the Mayo Clinic, Glass.ai. Oh, is that the drug? Yeah. These things have been put in medicinal and clinical settings. And that's, I think, the thing that's one of the most horrifying cases for me. One of the very curious cases that's been pretty alarming is mushroom identification. People have been using LLMs to generate mushroom identification books for amateur mushroom hunters. And if you, I know it's wild. 404 Media had an article on this. I think Samantha Cole wrote it. And it was about how these things are flooding Amazon. And if you have some made up mushroom and says it's safe to eat, and then someone eats it and dies from it, that's literally a death on the hands of this chatbot.

32:52And these things are just flooding Amazon. So there's a lot of horrible use cases. And these are the ones when I'm thinking about kind of direct bodily harm are the ones that I have top of mind, but there's a lot of other stuff out there too. How do you parse through the guns don't kill people, people kill people argument? Like it's not the technology, it's the misuse of the technology. It's where it's helpful to be a sociologist, right? Because you don't focus, and this is why Emily and I work so well together. She's a linguist and I'm a sociologist. As a sociologist, what I pay attention to are the organizations and collective incentives and which drive people to certain kinds of behavior and how certain organizations are incentivized to do so, right?

33:39So, okay, guns might not kill people, but certain people... But you're putting this tool out there and creating... There's an existing incentive structure for them to kill people with. So therefore it's a systemic issue and not an individual choice per se. Yeah. And I mean, that's the, that's the situation in which we're in a funding environment in which funders are, VCs are fighting hand over fist and giving out money like it's water to try to get some ROI on some AI tool. Then yeah, then it's going to be, people are incentivized to use these things and to use them quickly. The last time I checked PitchBook data, this industry had$44 billion in investment with trillion dollars in valuation.

34:31I'm sure if I go back to PitchBook, that's probably gone up$10 billion since the last time I looked in the last quarter. And so if you see just the sheer volume of money that's going out, then it doesn't matter if an individual LLM isn't going to kill people or not. If an LLM is sitting in a closet and it's being used for a scientific purpose only, that's not what's happening. There's a whole infrastructure around trying to turn the vestment off these things. Do you decry all medical uses of LLMs or AI broadly, or is it more nuanced than that? Is it just the irresponsible uses, some of which you just mentioned?

35:14I just mentioned the most egregious versions of these things. I don't decry all of these usages. I mean, I think there can be usages in which there are certain situations in which healthcare providers or people in social services could use these to some degree. However, there's been very little evaluation of these things in clinical settings. There's been very little public evaluation of these things through peer review. If they've been done through peer review, those benchmarks have their own problems. This is kind of the issue. And we recently did a show on our podcast with Dr. Roxana Naneshju, who is an incoming professor at Stanford on the uses of LLMs in medical evaluation and diagnostic.

36:03And, you know, much of the cases, for instance, Google did an evaluation of their MedPalm models and found something like initially a 68 % accuracy on the U.S. medical licensing exam and then an increased accuracy, I think, up in the 80s on that exam. But the problem is that that's not even a good evaluation for clinicians. That's the first step that allows entry into a medical program. There's much more that has to do with diagnostic and treatment plans. Yeah, it's like the LLM can pass the bar, so therefore it should be allowed to be a lawyer. Right. And we have an episode on that too with Kendra Albert, who works in the Harvard Cyber Law Clinic.

36:50We've talked to experts about these things and they're very critical as well. So if there's a place in which evaluation is robustly defined, where it is outlined in a way that has both construct and face validity, where the use case, if it goes wrong, has some type of recourse where there is a close human supervision, where you have a robust process, then yeah. I wouldn't be opposed to it. But that's not what's happening. These things are being put out. The kind of scientific papers that are written about it don't look very different from press releases. It's not kind of slow, thoughtful, read-upon evaluation work.

37:36And that's just not what's happening. Are there frameworks that you can point to or would suggest for folks that are, you know, I've got this, you know, shiny LLM tool. I want to use it for thing X. How do I know if that's a good idea? You know, you know it when you see a thing or there are 10 frameworks that are already published. Just pick any one of them. What tools do folks have for seriously evaluating the applicability and not just LLMs, any AI driven tool to a given problem? I know one of you mentioned, Beba Burhana, we spoke too long ago, years ago. One of the things she really focused on at the time was kind of being human centric or, you know, having a view that is centered on the people that are impacted by whatever the tool is, as opposed to, you know, a tool centric view.

38:28And I know that's the theme that kind of is carried through a lot of DARE's work. Are there frameworks that you would point people to, you know, for thinking this through? Or is this an area that we need to continue to develop? Yeah, I think there's some frameworks that are emerging. So one of them, the NIST has a risk management framework that they've been working through and trying to assess on if you're thinking about a tool, what would it mean to assess risk in this particular view? So I think that could be a helpful thing. In terms of different evaluation frameworks, I think that's a bit harder.

39:00I think it needs to be pretty particular to a use case. I don't really believe in this kind of idea of kind of like a general purpose technology. I mean, that is a thing that OpenAI likes to say that these things are, but that itself is problematic in many guises. And so I think identifying things that are more commonly accepted by a particular scoped academic community would be helpful to look at. So I would say, you know, are there things within the health or health evaluation, particular types of goals that would be well-scoped? Are there ways that communicating with people who are providers or professionals, would that be a process?

39:44Does that exist in a particular view? Or could that be a thing that you can engage certain kinds of professional associations with? I mean, I think those are all places to start looking for these things. But I just think these things are so new. None of those have been developed in cooperation with particular professional communities and societies. And regarding LLMs as a general purpose tool, is the objection there that it leads people to believe that you can take them off the shelf and tell them to do anything and their output is valid for doing that thing? Yeah, absolutely. I mean, this is being very critical of one particular work that OpenAI put out, which was GPTs or GPTs, a paper which got a lot of traction, which suggested that certain kinds of technologies would replace something like 10 % of jobs and affect 20 % of them.

40:41And first off, that paper has many issues, one of them being that the people actually rating those were OpenAI employees. So that also presents a face validity issue just from their own internal metrics as a ranking system. But the fact that many of these things also foreclose the possibility of other technologies. I go back to Weizenbaum here because surprisingly prescient, he's actually very critical of that notion of even a computer as a general purpose technology. We use computers for everything now, though, but that also forecloses a certain kind of notion of how people want to be recognized and computed in certain kinds of systems.

41:20And you also have to think about where Weizenbaum is writing. He writes this in 1976. He flees Germany in light of the Nazi occupation, the rise of the Nazism. And he effectively says, yeah, if Nazis had computers, they would have used them and it would have exterminated people faster. And we, I mean, and IBM, for instance, still hasn't apologized for the use of their counting machines for the kind of tally of people in camps. And so the kind of notion of computing as a kind of device, I mean, can be seen as a certain kind of project which forecloses other possibilities. And I think any kind of technology that claims to be generalizable can have that view, especially if it tries to take over kind of traditional knowledges and traditional ways of doing things.

42:09I think it's a longer conversation and I didn't mean to open that box, but it's also like, I also already mentioned Weizenbaum. So I think he did have some prescience in determining and talking about the way that certain technologies become generalizable and what they do to our imagination of what technology can be. In that last response, you just at the very end kind of grounded on traditional ways of doing things as like the touchstone. And the implication that I thought I heard was that, you know, having technology as a tool that replaces traditional ways of doing things is you didn't necessarily say that it was bad.

42:49But the implication was that you start from a perspective of, you know, it's bad and it needs to prove itself in some way. trying to necessarily formulate a question around this, but I'm mostly trying to get your take on that because that seems overly pessimistic or something. I'm not saying that we should start from the perspective that all technology is bad. I love indoor plumbing. I love pens. I love... Just not computers? I love computers. I can't lie. I've loved computers since I was four. I can't pretend like I don't like computers, right? Computers don't fascinate me. I have a degree in computer science.

43:28That was a dream of mine since I was five. And I'm glad I had that dream. At the same time, what I'm saying is that what are the ways in which these technologies will serve us that don't have externalities that are going to harm us? right? What are the ways in which these things could be viewed in certain kinds of ways? What are the ways in which, you know, we're going to develop machine translation that would be a helpful way of helping, you know, our grandmothers access the internet, while also acknowledging that machine translation has a history of having a colonizing force, or as a force of war making and Cold War spying and intelligence.

44:10Amanda Pallotta, Amanda Lynn Pallotta, one of my co-authors on other words, has a blog post she wrote for The Gradient, which talks about machine translation and the way machine translation shifts power. And she talks about the kind of development of machine translation in the Cold War era, basically used to translate from Russian into ways that would be more legible by intelligence officials. So these things are not, they're not value neutral. They're really value laden. If there's a way we can twist those to our own ends that work for communities, then that's great. But we also need to know, recognize that these things have certain kind of politics and histories that lead them to act in the way that they are now.

44:53Mm-hmm. You're just continuing with translation as an example. I think easy to see that it has its benefits as well as the drawbacks. How do you approach balancing benefits and drawbacks as a sociologist? That's a curious question. I mean, the as a sociologist part, I think is the thing. I think one of the aspects about this is seeing how these things are being used in context, Seeing what the political economy of these things are, who is making money, who's gaining power, status and capital through these things. And if it seems to be the case that these technologies tend to accrue to people who already have a lot of power, that is resulting in more harm than that seems like an issue.

45:39If it is instead something that is maybe a technology that is helpful in some limited sort of context and would benefit people disproportionately that are not already accruing many benefits, then that would be a benefit. But it's a trade-off in every case. I mean, it's hard to kind of talk about this in a kind of general case. I mean, in the translation case, we're at the point in which machine translation, I think, has gone to a certain place. To have a certain kind of access to the digital world, you need to have some, there are elements of the internet that are just completely inaccessible unless you have some translation into English or German or Spanish or a Western language.

46:23I mean, I guess Chinese, translating to Chinese, to and from Chinese is also in Mandarin more specifically. And so given that so much of the internet is, and the web, and therefore commerce and industry, is so inaccessible, then it seems like that one's a cat that's out of the bag. And in that way, it's making it accessible to people so they're able to access that world and exist and live within that world. Is translation the cat that's out of the bag or English and Western languages being dominant on the internet is the cat that's out of the bag? I don't know. What's the cat and what's the bag here, right?

47:06Yeah, yeah, no, I mean, I guess the English as being the dominant is a bit of the cat that's out of the bag. Machine translation is maybe the bag, or maybe I have that reversed. This metaphor is going to get more and more mingled the more and more I talk about it. Awesome. So, Alex, we've talked about a pretty broad range of things and just a small bit of the work that's going on in and around there. Before we wrap up, are there any other things that you'd like to point us to or projects that you'd like to suggest that our audience takes a look at as, you know, perhaps as representative of some of the things that we've talked about?

47:44Yeah, definitely. You can learn more about us at dare, D-A-I-R hyphen institute.org. That's where we've got a bit of work on all our projects, all our fellows. I also mentioned Teheku Media. Check out their work. Really kind of a friend of Dare, as well as L-E-S-A-N.A-I. And check out the podcast, Mystery AI Hype Theater 3000. We've talked about a lot of kinds of things there. So, yeah, a shout out to that stuff and just the folks kind of in the orbit. Awesome. Well, thanks so much for taking the time to chat. It was great to catch up on DARE and to learn a bit about some of the work you're working on.

48:26Thanks, Sam. It was a pleasure. Thanks, Alex. All right, everyone, that's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimla.ai.com. Of course, if you like what you hear on the podcast, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening and catch you next time.

From the publisher

Today we’re joined by Alex Hanna, the Director of Research at the Distributed AI Research Institute (DAIR). In our conversation with Alex, we discuss the topic of AI hype and the importance of tackling the issues and impacts it has on society. Alex highlights how the hype cycle started, concerning use cases, incentives driving people towards the rapid commercialization of AI tools, and the need for robust evaluation tools and frameworks to assess and mitigate the risks of these technologies. We also talked about DAIR and how they’ve crafted their research agenda. We discuss current research projects like DAIR Fellow Asmelash Teka Hadgu’s research supporting machine translation and speech recognition tools for the low-resource Amharic and Tigrinya languages of Ethiopia and Eritrea, in partnership with his startup Lesan.AI. We also explore the “Do Data Sets Have Politics” paper, which focuses on coding various variables and conducting a qualitative analysis of computer vision data sets to uncover the inherent politics present in data sets and the challenges in data set creation.

The complete show notes for this episode can be found at twimlai.com/go/649.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Pushing Back on AI Hype with Alex Hanna - #649The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 49 min
Listen in VO