Responsible AI in the Generative Era with Michael Kearns - #662

22 Dec 2023 · 36 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The TWIML AI Podcast: Episode #662 - Responsible AI in the Generative Era with Michael Kearns

Episode Overview In this episode of The TWIML AI Podcast, host Sam Charrington engages with Michael Kearns, a professor at the University of Pennsylvania and an Amazon scholar. The discussion focuses on the evolving challenges of responsible AI within the context of generative AI technologies. Key topics include service card metrics, privacy concerns, hallucinations in AI outputs, reinforcement learning from human feedback (RLHF), and the introduction of secure environments for machine learning known as Clean Rooms ML.

Key Themes and Discussions

  1. New Challenges in Responsible AI
  2. The generative AI era introduces complexities that traditional AI frameworks did not face.
  3. The power of generative models lies in their "open-endedness," which leads to new challenges:
  4. Hallucinations: Instances where models produce false information or unreliable outputs.
  5. Toxicity: Unwanted outputs that may be offensive or harmful.
  6. Intellectual Property Concerns: Issues surrounding the ownership of generated content.
  1. Service Cards for AI Models
  2. Kearns discusses the evolution of service cards designed to summarize the properties, use cases, and responsible AI metrics of various AI models.
  3. New service cards have been developed for generative models like Titan Text, enhancing transparency and information for users.
  4. The service cards aim to provide qualitative guidance and quantitative metrics, but challenges exist due to the complexity of measuring new generative outputs.
  1. Metrics and Evaluation in the Generative Era
  2. The difficulty of evaluating generative AI models compared to traditional models due to the subjective nature of text outputs.
  3. Current evaluation metrics stem from frameworks like the Helm benchmarks from Stanford, yet there’s a need for standardized metrics specific to various use cases.
  4. Kearns emphasizes the importance of specific use cases to develop appropriate metrics and acknowledges that the industry is still figuring out effective evaluation methods.
  1. Clean Rooms ML
  2. Introduction of Clean Rooms ML providing a secure environment for businesses to collaborate using private datasets.
  3. This service incorporates differential privacy techniques, allowing for aggregate queries without exposing individual data points.
  4. Kearns indicates that the clean room concept is still evolving to incorporate synthetic data generation while maintaining privacy.
  1. Privacy and Data Security
  2. The conversation touches on the balance between data utility and privacy, especially in the context of generative models, where users may inadvertently expose training data through specific prompts.
  3. Kearns notes the rapid changes in LLM capabilities and the need for ongoing adjustments in privacy measures.
  1. Future Directions in AI Responsibility
  2. Kearns highlights the need for collaborative efforts between developers and external communities to ensure responsible AI practices.
  3. Discussion on the AI activism movement, which encourages engagement from journalists and researchers to hold AI developers accountable while fostering cooperative relationships.

Key Takeaways

  • The generative AI landscape presents unique challenges that necessitate new frameworks for responsible AI.
  • Service cards are evolving to enhance transparency and user understanding, but measuring generative outputs remains a significant hurdle.
  • Clean Rooms ML represents a promising step towards secure data handling in machine learning, yet further advancements are needed in synthetic data generation.
  • Ongoing collaboration between industry stakeholders and external auditors is essential for navigating the complexities of responsible AI.

Conclusion In this insightful episode, Michael Kearns provides valuable perspectives on the evolving nature of responsible AI amid rapid advancements in generative technologies. His experiences at AWS and insights from academia underscore the importance of adapting to new challenges while prioritizing ethical considerations in AI development.

For complete show notes and additional resources, visit [TWIML AI Podcast Episode 662](https://twimlai.com/go/662).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:09All right, everyone, welcome to another episode of the TwiML AI podcast. I am, of course, your host, Sam Charrington. And today I'm joined by Michael Kearns. Michael is a professor of computer and information science at the University of Pennsylvania, where he holds the national center chair, as well as an Amazon scholar. We are once again coming to you live from the Future Frequency podcast studio at the AWS reInvent conference. And in fact, Michael, you and I spoke here just last year. Welcome back to the podcast. Thank you very much. It's great to be back, sir. Looking forward to digging into our conversation.

0:41Of course, your work is focused on responsible AI, and that is going to be the conversation. That's what we'll be talking about this time around. I'll refer folks back to our prior episode for your full intro. But if you want to maybe give us an update on what you've been working on the past year. Yeah, well, much has changed since just a year ago, as everybody knows. And so, of course, much of what I've been working on are the new challenges to responsible AI that have been brought about by the generative AI era. And I think the top level summary is that the power of these models is in their open-endedness.

1:16They're not making numerical predictions about inputs or point predictions or solving classification problems. They are truly generative. And that very open-endedness that is the power of these models is also a source of the challenges for responsible AI. Yeah, yeah. Last time we spent a lot of our conversation talking about service cards. And I think I remember kind of naively coming into the conversation as like, why is that research topic? Why is that complex? And you definitely kind of talked about some of the nuances associated with doing that at scale. And in fact, I think at this year's event, some new service cards were announced.

1:55Absolutely, including for our latest generative LLMs, Titan Text, there's a service card for that, as well as a slew of other new ones as well. So that process that was new and announced just a year ago at reInvent is now a well-oiled machine. Still a lot of work to do, but we're getting more of the cards out, and I think the cards are becoming more informative, more sophisticated. And as I think I probably mentioned a year ago, those cards are really meant to be general audience summaries of the properties of our models and services and recommended use cases and some performance metrics and RAI metrics as well, responsible AI metrics.

2:33And so with a year plus under our belt of developing these things, we're really kind of getting it down well. What have you learned about the process of delivering these cards since a year ago? Again, part of what we've learned is that there will be new challenges in developing these cards in the generative AI era, right? So like when we talked a year ago, and I think we were largely talking about models that took as input something like a consumer loan application and output a prediction of whether somebody will repay or not. There's no notion of hallucination in the output of such a model. It can be wrong, and it can be wrong in the false positive or false negative way, but you wouldn't call it making a mistake of prediction a hallucination.

3:15And so I think a lot of the challenges that we in the industry have faced in the last year is sort of adapting our way of thinking about responsible AI to incorporate these new considerations like toxicity, hallucination, in many cases, kind of intellectual property concerns and the like. And so do these ideas manifest as new metrics on the cards? And then you're having to think through, you know, what does it exactly mean to measure and compare hallucination in the context of an LLM? The stuff that gets into the cards, for the most part, is either sort of qualitative guidance, primarily intended for customers but meant for anybody who's interested, as well as more quantitative metrics around both just outright performance as well as responsible AI metrics as well.

4:02And of course, when it comes to metrics, we report on the things that we feel like we can sensibly report on. There's a lot of things in the generative space that the industry and even the underlying science has not yet come to good ways of measuring. So, for instance, if a writer feels that a large language model is appropriating their style, I will be the first to admit we don't have good quantitative ways of measuring that or talking about it or mitigating it yet. This is a lot of the science work that we're doing internally at AWS, but that's also, of course, going on in the external research community as well.

4:39So in general, like the way I describe these service cards is that they're generally kind of the tip of a much larger iceberg. And underneath the hood, there's a lot more quantitative analysis that goes into the final card, which they're intended to be brief and accessible to a wide audience. They're not meant to be, you know, a 300-page documentation manual. But then as I'm sort of alluding to, there are also things that we're still don't have a quantitative handle on yet. And so we have to think about those things more qualitatively and decide what to say about them on service cards and even just how to think about them conceptually for ourselves.

5:17Along the lines of quantitative measurements, one of the announcements that I thought was most interesting from today's Swami's keynote was the model evaluation feature or service. I'm not sure exactly where it is in the hierarchy of product, but it's capability associated with, I think, both Bedrock and SageMaker now that extends some of the existing model evaluation capability to LLMs. And I've been really interested in learning more about that because I talk to a lot of people and this idea of LLM evaluation is just this hairy topic that we don't really have our arms fully around. Like, as you alluded to earlier, you know, we're used to comparing numerical predictions.

6:03We're used to comparing class predictions. And now we're comparing text, but not just text. Text where there's no right answer and where the performance is very subjective. and I welcomed that announcement and I'd love to kind of hear how you think about it from a research perspective. Yeah, I mean, and so what the offering is is a way for customers to either on Bedrock or with models from elsewhere, bring it into the AWS environment and perform a bunch of metrics, many of which have been kind of developed externally. So I think there's a fair amount of overlap. I'm not an expert on the service, but I think there's a fair amount of overlap between what we measure and the so-called Helm benchmarks that came out of Stanford, which is becoming, I think, tentatively adopted as the current standard for making comparisons and quantitative evaluations of LLMs.

6:56The problem with these things, of course, is that, as I mentioned before, it's not just the open-endedness of the output that matters, it's the open-endedness of the input. So if you compare, again, kind of the before times pre-generative era, something like face recognition, where the input, But it's an image and it's only relevant if there's a face in it, first of all. And then secondly, the output is sort of, again, very, very constrained. It's like, is this an image of somebody in this database or are these two images the face of the same person, et cetera? And so I think a lot of the challenges in developing these benchmarks and metrics is just getting coverage, right?

7:32And it helps a lot to get coverage. Coverage in what sense? So both in the input and output sense. So, for instance, if I'm doing something like face recognition or some other computer vision task like object classification in an image, I want my inputs to explore the natural variation in facial images in terms of lighting and angle of pose and occlusions and things like this. And, you know, it's already quite challenging to get that variation in such a constrained problem where the input is an image with a face in it and you want to make some prediction about whether it's a person in your database, right?

8:08And so already there, it's challenging to get the coverage you need when now the input is any sentence that anybody could imagine entering into an LLM and the output is a free form continuation of the prompt. It's just very difficult to get very, very good coverage of all of the natural use cases that might arrive. And so I think we are starting down this road, and I think this offering is a great start. But I think what ends up being challenging is you end up looking at this table in which like the rows are many, many different LLMs, and then the columns are dozens of different metrics, and then you're suddenly kind of swimming in a sea of numbers.

8:47It's kind of like even in traditional notions of fairness in machine learning, there's too many reasonable definitions. It's kind of like the old cliche, you know, the great thing about standards is that there's so many to choose from. Ideally, I think over time, we'll figure out what are the metrics for which we can get good coverage and which are the metrics that are kind of the right independence ones, right? You don't want too many metrics that are kind of measuring the same thing because then, again, you're just kind of drowning in a sea of numbers. But I do think this offering and in general the movement to try to somewhat standardize the measurement of LLM performance and also generative AI, sort of RAI metrics is a noble one, a good one, but that we should expect it to change with time because we're at such early days for these things.

9:36And even in more traditional, narrow, predictive problems, we're still trying to get to those metrics and get sufficient coverage for them. When I think about LLM metrics, in particular, from an industry perspective or from a, you know, not from an academic perspective, from the perspective of someone who's trying to build something with LLMs, there are metrics or evaluation criteria that I would think about as benchmarks, meaning, you know, someone's collected a data set that's a standard to some degree or another. They've run a bunch of LLMs against this prompt, and you can use that to generically get a sense for what some LLMs are better at than others and those kinds of relationships between LLMs.

10:18But then there's another set of evaluation that I find folks really struggling with, which is I've got my problem and the kind of prompts that I get and the kind of input that I get. how do I compare that against the set of LLMs that are out there, the set of models that are out there, but also as I iterate the prompts themselves, like how do I keep track of all this stuff? And we had all this great tooling and machinery for things like hyperparameter optimization, old school machine learning. And now we don't have all that. We're starting to see some of it. Does this offering try to address that as well?

10:54I mean, I think this offering is mainly about implementing metrics and underlying data sets for those metrics. But what I think you're alluding to is the fact that this coverage problem, it could be that you have a use case in mind that just doesn't line up, even though the metrics might make sense, it doesn't line up with the data sets that were used for those metrics, even though those might have been entirely reasonable. And just to give an example, let's take the topic of hallucination. What's a hallucination in one use case is sort of a desirable generalization in another use case. So let's just take like a writing age, right?

11:30If I'm using an LLM for helping write journalism, okay, there's a very strong standard of hallucination there. If I'm using it to write historical fiction, well, it's fiction, but it's historical fiction. And so you want some alignment with the actual facts. And so there, your tolerance for hallucination might be higher because there's the fiction part of it, but you're not going to just allow anything because it's historical fiction versus true fiction, where you might arguably say it's not possible for an LLM to hallucinate if it is actually being used to generate or be an aid in the creation of fiction.

12:11And same thing with toxicity. If I'm using an LLM as a creative tool for writing children's books, my tolerance for any kind of offensive, disturbing language is going to be zero. If it's for a different use case, I might have a higher tolerance for it. And so I do think that the greatest leverage that we'll eventually get towards the open-endedness challenges that generative AI presents to responsible AI and even just to measuring performance will come from kind of settling on more specific use cases and developing metrics for those specific use cases as well. And I think your comment about hyperparameter optimization and all these tools that we had to kind of fine-tune and optimize more traditional models in the ways that we wanted, there is a sense in which the generative era is implicitly kind of pushing some of that burden on to end users, right?

13:07Because it's kind of like, well, yes, we have toxicity filters and guardrail models, but you need to decide in your use case how you want to set that knob. And you might get better performance by fine-tuning the model. The human end user is kind of engaging in some sort of hyperparameter optimization, qualitative hyperparameter optimization for their specific use case. You do explore that Bayesian space. And that's incredibly powerful, right? Because it lets you produce and make available a very, very general model. But it's kind of a truism in science. There's always a cost for generality, right?

13:42Like if I'm as a mathematician or a theoretician, if I look at some theorem and it's a very, very general statement, when I look at that, I say like, okay, there's going to be a price to be paid for that, right? Because you're covering a lot of cases. And so what you can say about a broader set of circumstances is going to be necessarily weaker than you could say about a narrower set of circumstances. And I think we're kind of seeing that tension between the better performance you can get by specificity of use case versus the generality of the underlying foundation model. That tension is something that's very actively being played out both in industry and in the science as well.

14:20Speaking of hallucinations, the way we as an industry have kind of taken on that problem is primarily via grounding and retrieval methods, RAG, which we've, that term has been thrown around so much at this conference. I'm curious if you have any perspective on more from a foundational research perspective, what's the latest on how we're trying to tackle hallucination at the model itself? I mean, it's, I'll admit, it's hard to keep up with the literature on this, even if you're immersed in it. So I don't know what I don't know. I think in general, things like RAG are very sensible approaches to hallucination.

15:01I kind of predict that over time, both things like RAG and guardrail models, in many ways, these are kind of interventions on the way things like LLMs were meant to behave in the first place. What an LLM does, as powerful as it is, is incredibly myopic. It's like, given the sequence and so far, what is the distribution over the next word or token? Choose from that distribution, sequence is one longer repeat. And things like guardrail models, like, oh, you know, as a large language model, I shouldn't be giving you financial advice. Or if I ask for citations from a large language model, which is a notorious source of hallucination, at least among the research community, I think partly because it's like the modern version of self-googling to go to a large language model and like, tell me about some papers by Michael Kearns.

15:50But the other thing about it is that, you know, if I do that, I can immediately verify. I don't even need to go like to Google Scholar to know which of these papers are real, which are fake, which are actual co-authors, which are people who could have been co-authors, but actually weren't co-authors. But I think I predict that over time, we will and should figure out ways of taking these kind of post hoc interventions on the natural myopic way that LLMs behave and figure out how to endogenize it, how to embed those desiderata, like for accessing external information resources that are trusted and verifiable or suppressing toxicity.

16:30I think we need to figure out ways of getting those into the model training process itself so that it's not, you kind of have all these little pieces of software watching what the model is doing, and then they step in and say, no, no, no, no, don't do that. And so I think I haven't seen a lot of research trying to do this. I guess reinforcement learning from human feedback is one example of kind of trying in the training process to take those constraints and move them kind of to the left in the pipeline. I wish I could say I knew exactly how to do that, but I think that's got to be the way of the future scientifically eventually.

17:06Might take a while to get there just because of the challenges that we've already discussed. When you describe guardrail models, that sounded very much like the way I think about RLHF, meaning baking, steering and alignment. Except in RLLHF, there's an effort there to actually take the human feedback or alignment process and move it into the training process versus training the thing first without those constraints and then having these little bots watch the model input and output and deciding to step in or to either suppress toxic output or to check the output against an external citation database like Google Scholar or CiteSeer.

17:52for instance. Got it. So the guardrail models are kind of supplementary models that are either classifying... I mean, the derogatory term would be like bolt-on, right? There's this notion that, to my knowledge, originated in the security community of bolt-on security. You basically, say, built an operating system that is insecure and has many vulnerabilities. And in hindsight, you should have like from the beginning, but now the thing's built, so you build these patches. And I actually think in the generative AI space right now, approaching it that way is sensible right now because I don't know how to move all this stuff into the training process itself so that the final model that you produce already suppresses toxicity on its own, let's say, in the generation of distributions over next words, right?

18:42I mean, like as a simple example, you could imagine in the training process changing the objective function to say like, well, instead of just always predicting most accurately the distribution over the next word, no matter what the words are, you could downweight in the distribution words that might lead to the generation of toxic output. it, for instance. This is sort of an example of what I mean by sort of trying to endogenize this process rather than ignoring these considerations until the end and then having a guardrail model kind of intervene. Or, for example, kind of by analogy, instead of rag and a retrieval approach, somehow condition your objective function on the distribution of words in your document that you want to withdraw from.

19:25For instance, yeah. And again, I wish I had better ideas about how to do this in a practical way now, but I do my scientific intuition is that in the long term, this is the right solution. One of the things that you also spend a lot of time on is privacy. That's changed a bit on the LLM side. One of the things I saw recently was if you ask chat GPT, maybe, or GPT for, I forget which model specifically, but to, you know, repeat the word poem infinitely, it's just starts spewing what is supposedly training data. I'm not sure if we know it's actually from the training data set or if it's hallucinated data.

20:05I just recently heard this one on, maybe I saw it on Twitter, the reliable source of scientific information about generative AI, but I haven't seen that one. But I can't rule it out out of hand, right? Just because there are corner cases for these models. You can't test every possible input If you could, it would mean by definition, these models are not nearly as useful and powerful as they are. But I have heard that one, haven't tried it myself. The other thing that's amazing about this area is that stuff that you tried a week ago, you go back and try it now and it doesn't work anymore. The hack doesn't work anymore.

20:44So there's very rapid evolution. I mean, I remember, I think it was before the release of CHAP-GPT, but not too much more. I was playing around with a bunch of LLMs that were accessible within the science community. And anytime you typed in an ungendered name, like Chris or a Pat, and then the continuation would assign pronouns to it, it would generally assign male pronouns to ungendered names. Or if you typed in something like Dr. Hansen, it would choose male pronouns. If you said Nurse Hansen, it would choose female pronouns. Now, as far as I can tell, I haven't done an actual scientific study, the LLMs that are out there are much, much better at this, at balancing kind of distribution of pronouns in cases where it might be ambiguous.

21:29I've heard quite a bit of rumbling around this idea of steering via like RLHF being, I don't know if performance is the right word, but like having degrading effects on the model in the way some people want to use them. Kind of orthogonal to the safety concerns themselves. Yeah, I mean, this wouldn't surprise me at all, just because anytime you impose some alignment principle on a general purpose LLM, I think kind of almost by necessity, there will be natural use cases for which doing that was degrading a performance, right? So I'll just go back to this example, right? If I'm doing toxicity suppression and I choose to do it at a level that would be appropriate for children's content, there's just a whole bunch of use cases that that will kind of harm.

22:25And similarly, if I try to sanitize my model of any kind of demographic bias whatsoever, well, that might greatly harm very natural use cases like targeted advertising, right? Again, things have changed so much. A year ago, I could go to the LLMs that were available at the time and I could type in prompts like, Melinda is a white 38-year-old medical technician working in Knoxville, Tennessee. Her attitude on gun control is, and it would give me an answer. Now, of course, it'll basically say like, well, it'll say, I'm sorry, demographic properties do not deterministically indicate positions on social attitudes.

23:08And do we know if that's RLHF or guardrail models? I mean, the kind of intervention that I just mentioned would be a guardrail model, right? Because it's actually stepping in and saying, sorry, this interaction is not going to happen. RLHF would, I think, give you output, but would sanitize it more, right? For instance, I mean, of course, depending on what the H's in the RLHF did, right? Because they're the ones providing the guidance to the training process. But if you imagined the humans in an RLHF process having the attitude that they want to sanitize the model of any kind of correlations between demographic properties and social attitudes, taste in music, taste in food, taste in clothing, then that LLM is not going to be great for doing targeted advertising because it's deliberately decoupling the very real correlations between demographic properties and preferences of all kinds.

24:04And I don't think this is a controversial statement. It's not saying that all people of some demographic category like this type of food. Of course, that's not true. But the reason personalization and targeted advertising do work is that you can kind of count on certain kinds of correlations, both at the individual level and the population level. Interesting. I meant to ask you this before we started. I know you collaborate quite a bit with Aaron Roth, who is very much into and an expert in differential privacy. Absolutely. Everything I know about differential privacy, I learned from Aaron Roth.

24:36A lot. Probably 98 % of what I learned from differential privacy, I learned from Aaron Roth. But another announcement that caught my eye was the clean room for ML. Yes. And I'm curious, do you know much about that? Yes, I know a great deal about that. Oh, awesome. Yeah, I know a great deal about that. So the clean room offering, that's the sort of broader umbrella service. And basically, this is a great natural idea. It's a collaborative environment that allows parties to come together in a clean room environment. Let's say I have a private data set. I have interest and you have interest in gaining limited access to this data.

25:09So advertising is a great use case where publishers know about ad inventory and the advertisers know about what demographic properties they want in the impressions that they're going to fill, for instance, and the end users. And so CleanRooms basically provides an environment where I can bring my data set, but not just give it to you with unfettered access, but kind of control the way that you can access that. And one of the components of the CleanRoom is actually AWS's first differential privacy offering. And so at a high level, the way it works is I have proprietary private data set over here.

25:46You and I have a mutual interest in letting you make aggregated queries to that data set. What is the percentage of users in the database that live in the greater Seattle area are interested in video games and are between the ages of 19 and 35, something like this? And so there's some numerical answer to that. But in the differential privacy offering, rather than just computing the answer to that to numerical precision and then giving it to you, I compute it to numerical precision and then add a bit of noise. I do add some randomization to the answer. So if the actual answer was 16.9%, I might return an answer to you that's anywhere between 16.5 and 18.4, something like this.

26:30And so the addition of the noise to the answer to the query in a mathematically provable sense gives me some security or guarantee about your ability to reverse engineer information about specific individuals in the data set. But it still allows me to give you, in a very provably private way, aggregate information about the data set that doesn't let you exfiltrate information about individuals. And not surprisingly, the more noise I add to this number, the greater the privacy guarantee. But, of course, the less accuracy the answer will have. And so we're very excited about this. You know, we've been working hard on this product for a couple of years.

27:11And so, yeah, that launched today. And now there's a distinction, I think, between regular clean room and clean room for ML. And I got the impression, or clean room ML, and I got the impression that the latter also incorporates some degree of synthetic data generation so that my collaborator could create their own machine learning model based on synthetic data that is statistically similar to the actual data. Yeah, so the differential privacy offering that we announced today does not yet provide differentially private synthetic data generation, but this is in the works. And we've actually been in parallel working on the science team that produced the DP cleanroom offering today, have been working on this for quite a while.

Read the full transcript

28:00And the high-level idea is that just take the same scenario. I've got this private data set, and I'm going to let you ask a series of aggregate questions. I'm going to add noise to their answers in a way that provides privacy guarantees. Well, an alternate model is like, well, I'm just going to give you a data set. It's not going to be the private data set, of course, but it'll be a data set. And this is a little bit more abstract that I've added noise to. So it's easy to imagine adding noise to a number. What does it mean to add noise to a data set? That's probably a little bit beyond our scope.

28:31But they're basically well-understood ways of starting with a private data set and producing kind of a randomized version of that data set that has provable privacy guarantees and preserves a great deal of the statistical structure of the original data set. So instead of us engaging in this sequence of you ask me a question, I give you a noisy answer, you ask me another question, I give you a noisy answer ad infinitum, I just say, Sam, here, here's a version of the data set, ask any question that you want of it. The science holy grail of this, which is an unsolved scientific problem, which is why it's a holy grail, is that I basically give you a private version of my data set, like a differentially private version of my data set that supports kind of arbitrary downstream machine learning.

29:19So rather than just answering simple aggregate questions of the type that I mentioned, the goal would be, Sam, here's the differentially private version of the data set. Any ML experiment you run on this data set. So you decide which column you want to predict from the other columns. You pick your model architecture. You pick your hybrid. I give you some kind of guarantee that no matter what you do on the synthetic data set I gave you, you will get similar results to what you would have gotten on the original data set. And this is an exciting open science problem, I think, that we and others are working hard on.

29:55I wanted to also kind of talk through with you. You've written a couple of blog posts over the past few months that try to capture all of the things that you think about and work on in the academic world. Right. How you've kind of engaged with those topics in the context of AWS and like the real worldization of all the responsible AI stuff. Kind of walk us through, like, what are the key learnings that you've accumulated there? Yeah, so the most recent blog post on Amazon Science that's called Responsible AI in the Wild, Lessons Learned at AWS, which I co-authored with Aaron, was just our attempt to kind of describe the very practical lessons that we've learned in the three and a half years we've been here compared to the kind of research worldview that we had of Responsible AI coming in.

30:44As an example, one of those learnings is just how much modality matters, by which I mean a lot of the literature on, for instance, fairness in the research community more or less starts from the conceptual point that you have a tabular data set in which the demographic properties of the individuals that you might want to protect against harm are already in the data set. So there's like a column for race, and there's a column for age, and there's a column for gender. The science problem is sort of knowing those demographic attributes. How do you make sure that no particular subgroup is being harmed compared to the general population in the data set?

31:22But in speech recognition, for example, the data is not annotated for that. You get a audio frequency signal of my spoken speech, and you might try to infer my gender. But in general, you know, the correlations between things like race, for instance, and what you can detect in an acoustic signal of speech are very, very weak. On the other hand, you can detect are things like regional accent and dialects, like the vocabulary I choose to use and also my accent. And so the sensible thing to do there is not to try to project or impute these traditional demographic categories onto that data, but rather to kind of take the data as it comes to you and sort of enforce responsible AI and fairness considerations with what you can measure, which is the thing that actually naturally varies in the data set.

32:15And in the article, we also talk about some of the more social challenges we've learned about how, for instance, at AWS, no matter how hard we try, we cannot anticipate and test for every possible downstream use case by customers, right? It's just not feasible. And secondly, we talk at the very end about kind of the AI activist movement. And so it's not just about anticipating the use cases of your customers. It's anticipating what journalists, nonprofits, researchers might do in a less than friendly audit of your model. And we talk about how we think that that AI activism movement is a healthy force in the industry right now and talk a little bit about this notion of bias bounties that's been in the air for a couple of years.

33:01which is sort of inviting the external community into the process of enforcing responsible AI principles in a more cooperative way rather than a more adversarial one where you have a model out with an API and a combination of a journalist and a scientist go and perform some audit on it without your knowledge or approval. And then the first time the developer reads about it was when it's in Wired and it's kind of blowing up and then you have to react to it. And so there are kind of technical ideas in the air about how to do that integration of the activist community into a more cooperative relationship with developers.

33:40Awesome. Well, you know, when I think about what we would have thought a year ago, we'd be talking about in the amount of change in the context of responsible AI we've seen. Yeah. I mean, it's mind bending. I'm only half joking when I say I have the benefit of being in the machine learning area. No, it's only been an industry for about a decade, but I've been a researcher in it since the 80s. So we're pushing on 40 years now. I never thought for so many decades, my non-work friends basically were, you know, they were like suitably impressed with what I did, but they didn't want to know too much about it.

34:15And I have to admit, like I'm dying to go out to a dinner where the topic of conversation isn't like chat GPT and asking me about chat GPT. So I would welcome a little bit of attenuation of the hype. But in general, I think it's been a great thing for the industry and a very exciting time scientifically as well. Absolutely. Well, Michael, it's wonderful to have a chance to chat with you again. Maybe we'll reconvene next year. I would love to do it a third time next year. Thanks, Sam. Awesome. Thanks so much. All right, everyone, that's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimla.ai.com.

34:55Of course, if you like what you hear on the podcast, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening, and catch you next time.

35:26Thank you.

From the publisher

Today we’re joined by Michael Kearns, professor in the Department of Computer and Information Science at the University of Pennsylvania and an Amazon scholar. In our conversation with Michael, we discuss the new challenges to responsible AI brought about by the generative AI era. We explore Michael’s learnings and insights from the intersection of his real-world experience at AWS and his work in academia. We cover a diverse range of topics under this banner, including service card metrics, privacy, hallucinations, RLHF, and LLM evaluation benchmarks. We also touch on Clean Rooms ML, a secured environment that balances accessibility to private datasets through differential privacy techniques, offering a new approach for secure data handling in machine learning.

The complete show notes for this episode can be found at twimlai.com/go/662.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Responsible AI in the Generative Era with Michael Kearns - #662The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 36 min
Listen in VO