Explainable AI for Biology and Medicine with Su-In Lee - #642

14 Aug 2023 · 38 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Explainable AI for Biology and Medicine with Su-In Lee - #642

Episode Overview

  • Podcast Title: The TWIML AI Podcast
  • Host: Sam Charrington
  • Guest: Su-In Lee, Professor at the Paul G. Allen School of Computer Science And Engineering, University Of Washington
  • Release Date: [Date not provided in the transcript]
  • Episode Focus: Discussion on explainable AI (XAI) in the context of computational biology and clinical medicine.

Key Takeaways

Background of Su-In Lee

  • Academic Background: Trained in machine learning, specializing in high-dimensional data.
  • Current Research: Focuses on explainable AI techniques for computational biology and clinical medicine, particularly in identifying causes and treatments for complex diseases like cancer and Alzheimer's.

Explainable AI (XAI)

  • Definition: XAI refers to methods that make the outputs of machine learning models interpretable, especially critical in high-stakes fields like healthcare.
  • Importance in Medicine: Opaque AI systems can lead to decisions that impact patient health; thus, explainability is crucial for clinical applications.

Discussed Concepts

  • Collaboration Across Disciplines: Emphasizes the need for interdisciplinary collaboration between computer scientists, biologists, and medical professionals to advance research and understanding in clinical applications.
  • Bilingual Researchers: Advocates for a new generation of researchers who understand AI, biology, and medicine to drive innovative solutions.

Challenges with Current XAI Techniques

  • Feature Attribution: Traditional XAI methods assess importance based on individual features but may lack biological relevance when applied to biological data.
  • Complex Interactions: Biological data often involve complex interactions among features (e.g., genes), demanding a higher-level understanding of these relationships rather than mere feature-level insights.

New Directions in XAI

  • System-level Insights: Suggests the need for models that consider the collaboration among multiple features to provide deeper insights into biological mechanisms, rather than just attributing importance to single features.
  • Counterfactual Analysis: Discusses the use of counterfactual methods to understand model behavior by altering inputs and observing changes in outputs.

Recent Research Highlights

  • Drug Combination Therapy:
  • Focused on acute myeloid leukemia (AML) as a case study.
  • Research aims to identify synergistic effects between drug combinations using machine learning, supported by XAI techniques.
  • Statistical Analysis of Feature Attributions: Employed statistical tests to validate the significance of identified gene pathways contributing to drug synergy.

Future Directions

  • Robustness of XAI: Research into making XAI methods more robust against adversarial attacks and applicable across different data modalities.
  • Single-Cell Data: Exploring applications of XAI techniques in analyzing single-cell gene expression data for more precise medical insights.
  • Model Auditing in Clinical AI: Developing methods to audit and validate clinical AI models for dermatology and other medical fields.

Conclusion

  • Su-In Lee's insights highlight the transformative potential of explainable AI in healthcare, emphasizing the need for innovative methods that bridge the gap between complex biological systems and machine learning. The episode urges a collaborative approach to develop researchers who can operate at the intersection of these fields.

Additional Resources

  • Complete show notes available at [twimlai.com/go/642](http://twimlai.com/go/642)

Call to Action

  • Encouragement for listeners to subscribe, rate, and review the podcast for continued insights into AI and machine learning developments.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:07All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington, and today I'm joined by Suin Lee. Suin is a professor at the Paul G. Allen School of Computer Science and Engineering at the University of Washington. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Suin, welcome to the podcast. Thank you for the introduction. I'm looking forward to digging into our talk. Auk, you are an invited speaker at the 2023 ICML Workshop on Computational Biology. And we'll be talking about your talk there, which is really centered around your research into explainable AI, an important topic.

0:54But before we jump into that, I'd love to have you share a little bit about your background and how you came to work in the field. Thank you so much. So my lab is currently working on a broad spectrum of a problem, for example, developing explainable AI techniques, so that's a core machine learning. And then we also work on identifying cause and treatment of challenging diseases such as cancer and Alzheimer's disease, so that's computational biology. And then also we develop clinical diagnosis or auditing frameworks for clinical AI. And you asked about how I got into this field. So I was trained as a machine learning researcher when I was a PhD student.

1:36I was working on the problem of dealing with high dimensional data. And then at that time, when I was a PhD student at Stanford, in the field of computational biology, there was something really exciting happened, something called a microarray data. So it's a gene expression data that measures expression levels of 20 ,000 genes. And I suddenly thought that if, you know, machine learning researchers develop a powerful and effective method to identify cause of diseases such as cancer and then therapeutic targets for, you know, those diseases, then, you know, as a machine learning researcher, I can contribute hugely to the science and also medicine.

2:17And I just fell in love with this field. So that's how I got into the research at the intersection of computational machine learning and computational biology. After I got a job at the University of Washington that has very strong medical school, and then I had wonderful colleagues, amazing people who had medical data, electronic health records, and then introduced me to this field of EHR data analysis in various clinical departments, anesthesiology and dermatology, and then emergency medicine. And then I just got really interested into the possibility, the potential that AI researchers or machine learning researchers, myself and my students can contribute to medicine.

3:06That's how I got into this field of largely three fields. So one is machine learning and AI, and the second is computational biology and then clinical medicine. You probably thought that you had to deal with messy data when you were in clinical biology and computational biology, and so you saw some of that EHR data. That data can be very messy. It is. The goals of the fields are slightly different to each other, but in the future, I strongly believe that those two fields will merge, biology and medicine. So in a clinical side, researchers are already generating the biological, molecular biology data from patients.

3:46So for example, for cancer patients, you can think about measuring the gene expression levels or genetic data from those cancer patients. And then what you want is the treatment. You want the AI or machine learning models to tell you which treatment, which drug, anti-cancer drugs are going to work the best for that particular patient. For that, you definitely need biological knowledge and then actual mechanistic understanding of cancer. And what says to you that the fields will merge as opposed to kind of collaborate closely? Clearly, they need to collaborate closely. But when I think of merge, and maybe I'm taking this too far, I'm thinking of like single models that operate in both domains.

4:35Yeah, I know what you're saying. So I tell my students or, you know, other young people that to actually move the field forward, to advance this field of biology, medicine, or biomedical sciences, you really need to become a bilingual researcher or even trilingual these days, you know, computer science plus biology plus medicine. When you have one brain that really thinks like machine learning researchers and biologists and then clinical experts, usually that really helps to come up with a creative approach and that can really move the field to benefit patients. And then at the end, the ultimate goal of biology and molecular biology is to understand life better so that you can advance the health of humans, right?

5:29So I think, you know, collaborations definitely help. But at the end, we really need to think about how to produce these young researchers so that they really think like, you know, experts in this area. These things, you know, already happened earlier in computational biology than clinical medicine. And when I was doing the PhD, it was usually based on collaborations, people who were trained primarily as a machine learning researcher and people who were trained as molecular biologists who hold a pipette and they work in the wet labs and then they form a collaboration and then write papers. But then later, we see a lot of departments that's named computational biology or biomedical science departments.

6:15So it's a really healthy move for this kind of interdisciplinary fields. It makes total difference. Yeah. Your research and again, your presentation at the conference are focused on explainable AI, XAI. Tell us a little bit about some of the things that you think are most important about explainability as applied to these fields. I think we get that machine learning and models in general can be opaque and make important high stakes decisions. You need some degree of explainability. But what's unique about your take in applying applicability in your field? Right. Okay. Thank you. That's an excellent question.

6:58So the core part of explainable AI, at least, you know, this theoretical framework, it basically means the feature attributions. So imagine you have a black box model, you have a set of input, a vector X, and then you have an output Y. And then when you have a prediction, you want to find a way to attribute two features. You want to know which features contributed the most. And then there are mathematical frameworks. Our particular approach that's called the SHAP framework, it is based on game theory. So you want to find a way to understand which features are important. So that's the core of the technical side of explainable AI.

7:39And then on the other hand, if you just apply this explainable AI technique, you know, off the shelf, explainable AI algorithm to biology, mostly it's useless. It's not very useful. It's not useful in terms of biological insights. What you really want to understand is how these features, you know, collaborate with each other. Imagine that you have a set of genes as a feature. You have 20 ,000 genes, 20 ,000 expression levels are the input of the black box model. And then your prediction is which cancer drug is going to work the best for each patient. And then individual genes contributions and then gene importance scores by themselves, they are not going to be really useful.

8:23It will be only useful when some explainable AI model, explainable AI algorithm can tell you which pathway, how genes collaborate with each other, and then how genetic factors play a role into that. And then also how that leads to the good prognosis of the cancer patient and also sensitivity, the good responsiveness to that drug. So there is something missing there. And then the uniqueness of my research is that we want to develop this explainable AI method for biology and then also clinical medicine such that it can make real meaningful contribution to these fields. Another example in the medicine side is that imagine that you have a deep model, deep neural network, that's going to take you a dermatology image.

9:17So say that you find something unusual in your skin and then you take a picture, that's your dermatological image. And let's say that you want to know that has a features of a melanoma or not. So the prediction results itself is not going to be really useful. And then even, you know, the current explainable AI methods that's going to tell you which pixels, which parts of the images led to the prediction of a melanoma or not, those are not going to be very useful to understand how this black box model really works. When you try, for example, that you modify the image and then generate a counterfactual, small changes to the image such that it changes the prediction.

10:00Let's say that that changed the prediction from melanoma to normal. Only then you can understand how this model works, what the reasoning process of this black box machine learning model is like. So those examples, I'm going to show many examples like that. Basically, the message there is going to be that the current state-of-the-art explainable AI that tells you theoretically supported importance values for the features are not going to be enough to make meaningful contributions to both biological science and then also clinical medicine as well. It sounds like you're calling out a broad deficiency in the approach and kind of saying that as opposed to this feature level explainability, we need more system level or process level explainability that is more grounded in the use cases or the application than what we have available today.

11:00Exactly. Right. Yeah. The question is how to do that. So that for that, we need a new explainable AI method. So in the first part of the talk, I'm going to show many examples of what explainable AI, almost as is, can do. So those are the papers that we published a couple of years ago, and then addresses so that it addresses new scientific questions. Even explainable AI or feature attribution methods as is can be useful. So I'm going to show many examples like that in both biology and medicine. But in the second part of the talk, I'm going to show how explainably I can even open new research directions specifically for biology and health care.

11:46So those examples I showed you, the systems level insights or this counterfactual image generation that can facilitate collaboration with humans, in this case, a clinical expert. So in the second part of the talk, I'm going to show how this explainable AI can open new research directions. And then part of the second part will be I'm going to have a deep dive into our recent paper to highlight how explainable AI can help cancer medicine design, cancer therapy design. So basically how to choose two chemotherapy drugs that's going to have a synergy for a particular patient. So that's the paper that was recently published in Nature Biomedical Engineering.

12:31Before we dig into that paper, the most recent paper, can you talk us through in a little bit more detail some of the examples of the foundational machine learning research and how they contribute to the problems you're trying to solve? Okay. So some of the foundational AI methods we developed, I'm going to talk about, it can be summarized into three parts. So one is, you know, principled understanding of current explainable AI methods. So specifically feature attribution methods. So, for example, in one work, we showed that our feature attribution method, that's SHAP, it was published in Eurips in 2017.

13:16We showed that it unifies a large portion of the explainable AI literature and 25 methods following the exact same principle and all explaining by removing features. So it turned out that 25 methods, feature attribution methods that are widely used in the field and machine learning applications, they all go by the same principle. You want to assess the importance of each feature by removing them or removing subsets of them. So that helps us understand what goes on. For example, when they fail, you want to understand what goes on and also improve and then develop new explainable AI methods. So I'm going to introduce a couple of unifying frameworks.

14:04So this is about how to understand the principled understanding of feature attribution methods. So also on a computational side, we have explored many avenues to make this SHAP computation even feasible and faster. So SHAP stands for Shapely Additive. I suddenly forgot. I can't forget this. Explanations. Yes, the Shapley additive explanation. It's kind of weird because they chose the third letter of the word. Well, that's the first author, my student, Scott's choice. And then I love the name, by the way. But computing Shapley is theoretically very well supported. But then computation-wise, it's not really easy to compute.

14:52It involves exponential computation. So we need to develop approximation methods such that we can compute them in a feasible manner. So we developed many, you know, fast statistical estimation approaches. And then you want to make sure that there is a convergence and then all the theoretical, desirable theoretical properties are already there. And then also we developed approaches for, you know, specific model types. For example, ensemble tree models, and then also deep neural networks. So we have a deep SHAP and then tree SHAP. And then more recently, we also have a vision transformer SHAPLY.

15:29So that's a way to compute the SHAPLY values for transformers, vision transformers. And then there is another one that's called a fast SHAP. So the one way to make the SHAP computation more feasible is to focus on specific particular aspects of models. So for example, tree ensembles or dim neural network. They have some particular model types. There is a way to make this computation a little faster. So model-specific versions of SHAP implementation. Yes, yes. So that's another line of research. And then more recently, we also started to understand the robustness of the SHAP value. So adversarial attack.

16:17A few years ago, you know, in the field of machine learning, people, you know, researchers have tried to understand how robust the machine learning model itself, the prediction results are toward adversarial attacks. And then now we are looking into this issue in terms of the model explanations. So how feature attributions are robust. So in our most recent paper, we basically showed the removal-based approaches, including SHAP. Earlier I said many of the feature attribution methods turned out to have the same principle, which is explaining by removal. So that method is more robust to this kind of adversarial attacks.

17:01So, and then, you know, multi-modality, you know, those other kinds of issues, we are actively doing this research in terms of, you know, foundational AI algorithms also. And Shep, as you've mentioned, is broadly used both the original algorithm as well as it's the related algorithms as you described. But it's also one of the first explainability approaches to be popularized. where does it sit in terms of relevance? Are there different kind of wholly different approaches that have overtaken it in popularity or applicability based on kind of today's models and applications or is SHAP still kind of a core approach to the way explainability is looked at in practice?

17:51It's more on the later side. We believe that this removal-based approach and in this cooperative game theory, we believe in that. And then also it has the desirable properties, first of all. And then we, in our many experiments, we still see that removal-based approaches are more, you know, robust, as I said, you know, adversarial attacks. And then also in terms of various evaluation criteria, we still think that those methods are more robust than the other class, which we characterized as a propagation-based approach or gradient-based approaches. So we would prefer this removal-based approaches.

18:31But on the other hand, those approaches are very computationally, very intensive. So the way SHAP works is basically that, you know, you try all subset of features and then you add a feature of interest and then see the model, check the model output and you average across all subsets of features. So as you can imagine, it's computationally very intensive. So when we now think about foundational models or large language models, it is really large models of a lot of parameters. And then dim neural network and gradient computation is perhaps easier than trying all subsets of features, right? So practically, it's not as easy as the other class in terms of the computation.

19:15But we still want to make this computational more feasible. We want to develop various clever approaches to reduce the computation and then still maintain the desirable theoretical properties that this removal-based approach or SHAP in particular has. Got it. And so that is an example of kind of the foundational research that your lab does that contributes not only to your work on the biological science side or computational biology side, but broadly to the field. And then your more recent paper is an example of the kind of contributions you're making on the medicine side. Can you talk a little bit about that cancer paper?

20:01Yeah, sure. It is about AML. So we chose AML as an example application. So it's acute myeloid leukemia. It's aggressive blood cancer, and it's relatively common for older people. So to give you a bit of a background in general, the cutting edge in the treatment of cancers, such as AML, has increasingly become combination therapy. So the rationale here is that by choosing drugs that target complementary biological pathways, we can achieve greater anti-cancer efficacy. So basically you choose two or three chemotherapy drugs and then use them together so that when there is a synergy, usually there is a very good anti-cancer efficacy.

20:51But the issue is that choosing optimal combinations of drugs is a really hard problem. So there are about hundreds of individual FDA-approved anti-cancer drugs, which means that there will be tens of thousands of possible combinations. But when you consider pairwise combination, and there could be even more if you consider non-FDA approved experimental drugs in development or consider a combination of more than two drugs. So, and then the different patients, even patients who have the same type of cancer may respond differently to exact same drugs because of this individual, the particular genomic characteristics.

21:31So then formulate this problem as a machine learning problem. So you take this AML patient's gene expression levels. So you get the blood of the patient and then purify the cells so you have only cancer cells. And then say you measure expression levels of 20 ,000 genes. So mathematically, this is 20 ,000 dimensional vector. And then also, let's say you consider a pair of drugs, drugs A and B, and then you use various information about this drug. For example, structure of these drugs or their biological targets. There are many data sets that can tell you that information. And then you take those as a machine learning input, and then you want to predict the synergy between the drugs A and B.

22:19So in this kind of a problem, and as I said, there will be tens of thousands of pair-wise combination of those drugs. And so in this kind of situation, not only the prediction, but also explanations will be extremely important. So say you want to be able to say that drug A and B is going to work well, are going to have a synergy together because this patient, X, has gene expression levels of A, B, and C high. And then, or, you know, say expression levels of a certain biological pathway, those genes are highly expressed. So you need a set of explanation to do that. And then more importantly, if you think about, you know, all pairs of drugs, if there is an underlying principle in terms of when two drugs are likely to have a synergy, then it's going to be even more useful.

23:14So what we did in this paper was that we got the explanations. We computed the shaft values for many combinations of drugs from the machine learning model, and then we analyzed that, And then we identified the unifying principle in terms of, you know, when, in what case drugs, any pair of drugs A and B have a synergy. And then we identified a pathway. It is called stemness pathway. So it is also called, trying to find in that part of the slide, this hematopoietic stem cell-like signature. You know, cancers are sometimes more differentiated or less differentiated. If you had a family member who had cancer, you probably understand this term.

23:59So usually, less differentiated cancers have a worse prognosis than more differentiated cancers. So we identify this pathway that's really relevant to this stemness mechanism and then found the underlying principle, which basically says that it's good to have two drugs, one drug targeting less differentiated, the other one targeting more differentiated cancer, are likely to work the best. So in this project, not only our algorithm can tell oncologists or biological scientists which genes are important, which feature attributions, which features are important for drug synergy, but also by analyzing many model explanations from many patients, we can have an understanding of these underlying principles in terms of what makes a successful drug combination therapy.

24:56Cancer therapy design, I would say, this is explainable. This is an example where we can see how explainable AI can be effective in cancer therapy design. Is AML unique in having a well-understood pathway, or is that a bottleneck for the application of this technique to the broader set of cancers? Oh, so AML is just one example. I mean, this kind of a principle can be applied to many data sets. You know, computational biologists often need to work on the problem where the data are available. So, you know, as you can imagine, blood cancers, those, you know, tissues are relatively easy to, it's relatively easier to obtain, you know, blood tissues compared to other kinds of tissues.

25:45So there are many available, you know, data sets. And then also, the measurement of the drug synergy from many samples. So we happen to chose this cancer type because of the data availability. But this principle, this approach can be broadly applicable to other types of cancer. So this is one of the... I'm maybe trying to get a broader question, which is the explainability method is kind of explaining over a set of known features and pathways and processes and things like that. And my sense is that for many of the potential applications, the pathways are still a subject of research themselves.

26:33Meaning, you know, maybe there's some aspect of pathway that's known, but there are others. There are, you know, or some diseases for which there aren't pathways. And I guess I'm I'm wondering the way you think about applying techniques like this. And A, is that actually the case or am I all wrong there? But otherwise, how will you apply techniques like this in rapidly evolving fields that are very complex? That's an excellent question. Maybe you're giving an explanation and the explanation is based on the pathway as you understand it. But there's so many other things going on in the system that you really have not accounted for.

27:09Yeah, exactly. Right. So first of all, pathway is not unique to disease. So when we say, you know, pathway databases, it basically tells you the members of the genes in each pathway. That's it. I mean, it's like, you know, many, many sets of genes. We also sometimes call it gene sets. It doesn't depend on the disease. And then the way we view is that it's not like all genes need to be activated for the pathway needs to be activated. It will be only a subset of genes. We would expect only a subset of genes to be highly expressed to say that pathway is activated. And then it's extremely important for a computational biologist when we develop a method like this to get biological insights from large-scale data sets.

27:54When we develop such a method, we need to make sure that it does not fully depend on any sort of prior knowledge. And then the algorithm needs to be flexible. So that's of key importance. So in this particular example, we didn't use a pathway actually from the beginning. When the model training happens, we used genes as individual features, and then we analyzed the feature attributions and then did the statistical test to see which pathways seem to be more activated. You made a really good point. In all computational biology methods, it's really important not to make it too rigid for the existing knowledge.

28:34It needs to be flexible. And so how do you evaluate your results in this particular paper? Oh, so say that you have a feature attribution for all genes for a certain patient and then for a certain combination of drugs. And then say you will have a lot of feature attributions then, right? Combining all patients and then all pairs of drugs you considered. And then we perform the statistical test. So for example, it's a simple Fisher's exact test kind of statistical test where you see whether there is significantly large value of attribution values for certain set of genes defined by certain pathway.

29:21And then you do multiple hypothesis testing and then see whether that significance is indeed relevant. So the pathway-based analysis was done in a post-tugment. after model training and then obtaining all model explanations. So another challenge we ran into in that project that was really not addressed properly by this foundational AI field was feature correlation. So in many biomedical data sets, you will see lots of features that are correlated with each other. Many genes are correlated. It's really modular. Gene expression levels are very modular, So you easily see, you know, subset of genes that are very highly correlated with each other.

30:07So in that kind of case, SHAP values are not going to be extremely accurate because, you know, imagine that there are two genes that are perfectly correlated with each other. Then there will be infinite ways to attribute to these two genes, right? So in that paper, in that Nature Biomedical Engineering paper, we addressed it by considering ensemble model. So we ran many ensemble of model explanations. So we ran the model. In this case, it was not your team neural network. It was three ensembles. And then we averaged. We averaged the feature attributions that are from many models. And then we showed that it gives you more robust feature attributions when the features are correlated with each other.

Read the full transcript

30:54Awesome. So talk a little bit about where you see the future of your research going. That's a really important question. So in all three ways, so first of all, you know, in the foundational AI method, as I briefly mentioned, you know, this robustness issues and then also multi-model data. Let's say that you have a set of features and each feature belongs to different category. They are in different modality. and then how to attribute to these features that are in different modalities. So that's an open problem. So it was actually motivated by biomedical problem, but it's broadly applicable to other applications.

31:41And then also these emerging models of LLMs or other foundational models. And in this kind of really large models, how to actually compute the feature attributions properly. And then also we are really interested in sample-based importance. So say that you transpose, the matrix transpose of your feature matrix. So I've been talking about these feature attributions a lot, but you can also apply Shapley values to gain insights into which samples are important for your model training. So that can help us understand how, you know, foundational models in various fields or large language models rely on which training samples.

32:28So that can be really important for model auditing perspective, first of all, and then to, you know, gain insight in terms of, you know, which samples were important for these large models to behave a certain way, right? So sample-based explanation is also one of the things that we are mainly working on. In the biomedical side, there are many projects. So single-cell data science is one of the big themes in my lab now. So you obtain gene expression levels or other kinds of molecular level information at a single-cell level. So the advantage is that you will have a ton of samples. So one experiment is going to give you many samples, which is really appropriate for large scale models these days based on dim neural networks, right?

33:19So, for example, the researchers started looking into foundational model for single cell data set. So in this kind of, you know, data sets that have still, you know, high dimensional, and then researchers are now obtaining multi-omic data. So not only gene expressions, you can also obtain other kinds of, you know, genomic information. So that's going to increase the dimensionality also. and then large sample sizes, how to learn the biologically interpretable representation space. So that's one of the big questions in my lab, in the research in my lab. So all feature attribution methods at the end in the downstream prediction test, you attribute two features.

34:04And then the assumption is that each feature is an interpretable unit. In biology, as I mentioned earlier, it's not the case in biology, right? So the functional units in biology is much more interpretable than in any individual genes. So how to learn this kind of, you know, the features that have more broadly representation, feature representation space, that's biologically more interpretable. And then also how to make, you know, foundational models learned based on single cell data sets. So researchers started publishing those papers that are about applying this foundational model approach to single cell data sets.

34:47And then how to make it biologically interpretable so that you can gain scientific insights from the model results and then also audit those models to make sure that users can actually safely use them for scientific discoveries. So attribution methods for this kind of modern machine learning models so that you can gain biological insights. So that's another theme. In a clinical side, we are really interested in this model auditing. In our most recent paper that's in review, we are focusing on dermatology examples. So dermatological image is inputted into the neural network. and then you want to know whether that's, you know, the prediction result is melanoma or not.

35:35There are many algorithms out there, some published in very, you know, high profile medical journals, and then also some available through the cell phone apps. So there are many algorithms. And then we recently tested them. We just, you know, separate the held out test samples and then got the result that's a little, you know, concerning in terms of, you know, usage. So So, and then our analysis showed that explainable AI was extremely helpful. So, for example, you know, in a skin image, which part of the image led to that kind of a prediction? Or, as I said, you know, using this counterfactual image generation.

36:17So you make small changes to the input dermatology image such that it changes. It crosses the decision boundary of the classifier and then see what features were changes. So that way you can see the reasoning process of this classifier, the clinical AI model, right? So for that, there needs to be some, you know, technological development there because the feature attributions themselves are not going to be enough. It shows only very small part of the inner workings of the machine learning model. So developing methods for auditing clinical AI models, that's the research we are currently performing in the clinical area.

37:02So all three areas, we are doing exciting research. Well, Suun, it sounds like you've got a lot of work ahead of you. Yes. Yeah. Very busy. I bet. Thanks so much for joining us. Thank you. Thank you for inviting me. Thank you. All right, everyone. That's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimla.ai.com. Of course, if you like what you hear on the podcast, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening and catch you next time.

From the publisher

Today we’re joined by Su-In Lee, a professor at the Paul G. Allen School of Computer Science And Engineering at the University Of Washington. In our conversation, Su-In details her talk from the ICML 2023 Workshop on Computational Biology which focuses on developing explainable AI techniques for the computational biology and clinical medicine fields. Su-In discussed the importance of explainable AI contributing to feature collaboration, the robustness of different explainability approaches, and the need for interdisciplinary collaboration between the computer science, biology, and medical fields. We also explore her recent paper on the use of drug combination therapy, challenges with handling biomedical data, and how they aim to make meaningful contributions to the healthcare industry by aiding in cause identification and treatments for Cancer and Alzheimer's diseases.

The complete show notes for this episode can be found at twimlai.com/go/642.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Explainable AI for Biology and Medicine with Su-In Lee - #642The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 38 min
Listen in VO