In short
NVIDIA AI Podcast Episode 195 Summary
Episode Overview Title: Bojan Tunguz, Johnny Israeli on How AI and Crowdsourcing Can Advance Vaccine Distribution Description: This episode discusses the intersection of artificial intelligence, crowdsourcing, and mRNA vaccine distribution, focusing on improving the thermo-stability of these vaccines for global accessibility.
Host: Noah Kravitz Guests: Bojan Tunguz (Physicist, Senior System Software Engineer at NVIDIA) & Johnny Israeli (Senior Manager of AI and Cloud Software at NVIDIA)
---
Key Topics Discussed
The Challenge of mRNA Vaccine Distribution
- mRNA Vaccines: Highlighted the rapid deployment of mRNA vaccines during the COVID-19 pandemic and the critical issue of thermo-stability.
- Thermo-stability Issues: Traditional mRNA vaccines require ultra-cold storage, limiting access in regions without proper refrigeration.
AI and Crowdsourcing in Vaccine Research
- AI's Role: Discussed the potential of AI in drug discovery and the use of crowdsourcing to tackle scientific challenges.
- Stanford Open Vaccine Competition: A machine-learning contest aimed at solving mRNA stability problems was hosted on Kaggle.
Kaggle and Crowdsourcing
- Kaggle Overview: An online platform for machine learning competitions, datasets, and community discussions.
- Competition Structure: Competitors can earn points and rankings across various categories, including discussions and datasets.
The Stanford Open Vaccine Competition
- Purpose: To find stable mRNA vaccine candidates using machine learning models to predict RNA degradation.
- Timeline: The competition was completed in a rapid timeframe of two weeks, with the overall research conducted in less than six months.
---
Insights from Guests
Bojan Tunguz
- Background in Kaggle: A quadruple Kaggle grandmaster who discussed his experience in setting up the competition and the importance of avoiding potential pitfalls in competition design.
- Machine Learning Applications: Emphasized how crowdsourcing can leverage global expertise to address urgent scientific challenges.
Johnny Israeli
- AI for Drug Discovery: Focused on the significance of competitions in standardizing benchmarks and driving innovation in drug discovery.
- Software Development: Discussed NVIDIA's efforts in building software to streamline drug discovery processes using AI and machine learning.
---
Challenges and Future Directions
Main Challenges
- Proving Value: Difficulty in demonstrating the effective application of AI in drug discovery due to lengthy validation processes.
- Engineering Complexity: As models become more complex, there's a shift toward requiring larger engineering teams and resources.
Future Prospects
- RNA Therapeutics: Anticipation of more breakthroughs in RNA therapeutics and potential applications for seasonal flu and other diseases.
- Integration of AI: The need for ongoing collaboration between life scientists and engineers to foster a culture of innovation.
---
Conclusion The episode reinforces the transformative potential of AI and crowdsourcing in advancing vaccine distribution, especially for mRNA technologies. Both guests express optimism for future breakthroughs in drug discovery, highlighting the importance of community-driven initiatives.
Additional Resources
- Research Paper: [Deep Learning Models for Predicting RNA Degradation via Dual Crowdsourcing](https://www.nature.com/articles/nature)
- Kaggle Competition: Information available on the Kaggle platform regarding the Stanford Open Vaccine competition.
- NVIDIA Software: Explore tools like BioNemo for drug discovery on the NVIDIA website.
---
This episode provides valuable insights into the collaboration between AI and crowdsourcing in addressing global health challenges, particularly in the context of vaccine distribution.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:10Hello, and welcome to the NVIDIA AI Podcast. I'm your host, Noah Kravitz. The COVID-19 pandemic introduced many of us to the term mRNA. The rapid deployment of messenger RNA-based COVID-19 vaccines highlighted mRNA's immense potential for society. But worldwide distribution of mRNA molecules has been limited by their thermostability. In other words, mRNA vaccines traditionally have had to be kept super cold during storage and transport. But artificial intelligence and crowdsourcing might be able to change that. Our guests today are here to discuss the use of AI in drug discovery, including a crowdsourced machine learning competition called Stanford Open Vaccine.
0:54A recent research paper highlighted the immense potential of machine learning in tackling urgent problems that demand accelerated scientific discovery, like finding a vaccine to fight COVID that doesn't have to be kept super cold. Boyan Tunguz is a physicist and senior systems software engineer working on machine learning at NVIDIA. He is also a quadruple Kaggle grandmaster, and we'll explain what that means in a moment. And Johnny Israeli is senior manager of AI and cloud software at NVIDIA, where he's working on the use of AI and drug discovery. Thank you so much for taking the time to join us, and welcome to the NVIDIA AI podcast.
1:32So there's a lot to get into here, and obviously it's a very serious, important topic. drug discovery and the role of mRNA. And, you know, everyone who's lived through the past few years, I think, has probably at least heard the term. For me, it was kind of the first time I heard the term RNA relative to the work to find vaccines to fight COVID. But we're going to start, there's also a lot of fun to this story. So I think let's start on that side. Wayan, in the introduction, I mentioned that you are a quadruple Kaggle grandmaster. So maybe we can start with you describing what Kaggle is, and then we can get into the role that Kaggle played in the work leading up to the paper and everything else that we're talking about today.
2:11Sure. So Kaggle is an online machine learning competition website that is sort of a general purpose data science hub where people from all over the world compete for machine learning competitions. But it's also a place where you can find a lot of interesting data sets, solutions, or machine learning general kind of purpose material. And a quadruple Kaggle grandmaster, I'm assuming that means you're pretty good at it. But how does it work? So Kaggle is the best node for competitions. And to this day, this is where most of the interesting things happen. But over the years, they've kind of branched out to different other categories.
2:49So in addition to competitions, you can become a grandmaster in discussions, which I started participating very early on. Code itself, so you can get points as well as medals for the code. And finally, the last category that I introduced a few years back was a data set. So, you know, you have competitions, you have code, you have discussions, and you have data sets. So these are the four canonical categories. And you can, you know, compete for points, you can compete for rankings, and you can compete for all these status achievements. And I managed to get all four. I was the fourth one to do so.
3:23And there's currently only five of us. And we'll actually have a podcast in about the month. All five of us are going to be together the same. Oh, very cool. I look forward to that. And so there was a competition on, was the Stanford Open Vaccine Competition, was that on Kaggle? Yeah. So tell us about that. So I was approached, Kaggle's CEO and myself, or former CEO, I think, we had a very good relationship and he always moved it. I'm interested in looking into application of machine learning to science in general and biosciences in particular. My own background is in physics, but I have kind of drifted away from scientific work over the years.
4:01But now I kind of try to get back a little bit into that. So he knew I had interested in that. And when Riju Das from Stanford University approached him about setting up a competition in Chicago for this opium vaccine project, he reached out to me, asked me if I would be interested. And it was a very great opportunity. but you know it's a little bit of a cold sight for me because I want to compete in Kaggle but if I help set up a competition I cannot do that so you know I had to kind of forfeit uh me competing in this thing in order to actually set it up but it's a very valuable experience and I'm really glad I did.
4:37So was the the open vaccine competition specifically geared at finding a vaccine for COVID? Was it geared towards the mRNA instability problems? What was the what was the focus? Yeah I think it came out in late 2020 when vaccines were still in the semi-development stage. So we were not really sure that they're going to be out in time. And the problem with RNA molecules is that they're very unstable. So we cannot just ship them out in regular refrigerators. We have to use these very ultra-cold temperature refrigerators. So the hope was to actually, in order to make it more accessible to people around the world, we don't have access to resources and a high, very low temperature refrigeration, we could maybe make up more stable versions of the current vaccine candidates.
5:29So prior to the competition, another crowdsourcing platform that Vigil was part of designed candidate vaccine molecules. And our attempt was to find which one of these candidates would be most stable and to build a machine learning models that would be good candidates for vaccines. These were never actually used for vaccines, but that was actually a hope at the time that that's what we could do. What was that other platform? Eterly. Eterly, yeah. And so how did the competition proceed? How was the progress? So, you know, there's an interesting kind of background story involving myself in there.
6:10I was part of NVDS KG1 team, which is competitors of Kaggle. And, you know, so I was approached to be part of this competition. So I didn't want to alert them to any of this or to actually bias them in a way because I knew they were going to compete. But I'm part of this team. So I had to kind of go around the way to tell my manager to tell them that I was going to be on a project. But he cannot reveal to them what's going to be going on until the competition actually launched. So it was a very interesting kind of situation for me. And it ended up that photo by our teammates, another Kazuki won the second prize in this competition.
6:46So which was very good for NVIDIA as well. But all throughout the competition, actual competition, I could not really interact with them, you know, share any insights or any of these things during the time. But prior to the competition's launch, there's a lot more going into setting up machine learning competitions that people realize. It's not just kind of, here's the data, here's the metric, you know, go. Right. There were a lot of little details that need to be kind of hashed out. And a lot of that stuff, my old previous experience with competitions came into play. So deciding which data sets to use, whether to use some dirty data or not, which metric to use, how to split the training and test sets, and things like that.
7:27there was a lot of little details that could have tripped the competition. And there are a lot of competitions that set out well, but in the end, there's all these shakeups. So the models that seem to be doing well actually do not really generalize well. So there were a lot of potential pitfalls. And my approach and I think my biggest contribution just sort of alert the people who were setting up the competition to kind of stay away from some of these potential pitfalls. Now, in the abstract from the paper, which is out on nature.com, if you want to check it out, it's called Deep Learning Models for Predicting RNA Degradation via Dual Crowdsourcing.
8:03It mentions that the entire experiment was completed in less than six months, which to me sounds incredibly fast. Is that so? Is that actually a rapid length for a competition like this? And what were the kind of the results and takeaways? is? I think Johnny would be better to say a little bit more about how fast this was for experimental part. I'm not an experimental scientist, so he might know a little bit more. Sure. But the actual competition actually took only two weeks, which was also very rapid. Okay. The Kaggle competition part was also very rapid. Usually competitions are two to three months.
8:39That gives enough time for people to sort of understand, you know, the subject matter, understand like little tricks that are relevant to the data set and, you know, hash out things. But even from part of the competition part, it was a very rapid competition. And again, like we were hoping at the time, if we can get these models out very quickly, it can help with vaccines. Right, right. And, you know, it turned out to be a little bit too ambitious maybe, but, you know, we're still kind of hopeful that this can be very useful for future servitude. So the goal of the competition itself was to produce the best models possible.
9:14So the best ability models. So the objective of the competition in this whole project was to find the best model that predicts degradation of RNA molecules. Right. And so kind of the best model and a lot of very innovative things that kind of pushed whatever, you know, people in scientific communities were able to do quite a bit. So Johnny, my understanding is that at NVIDIA, you're working on software that can be used in drug discovery. So AI-powered machine learning and the software that's used for drug discovery. How does a competition like this figure into the work that you're doing, both, you know, specifically in this case and then kind of broadly, maybe you can describe what your work is like, what building software to aid in drug discovery is all about?
10:03Yeah, let me start with a competition question. I think competitions play a huge role in the space because to assess progress and to benchmark progress, you need a benchmark. And competitions serve that purpose. And I would add that unlike some other fields, say when you think about some of the classical fields of machine learning made, some of its initial progress with deep learning like computer vision and audio and so on, those other types of problems, they're a little bit closer to human intuition. And so you could imagine, even in the absence of a benchmark, we could, to some extent, communicate with each other our progress.
10:43It's sort of data that we intuitively understand. But then as soon as you enter the life sciences, whether it's mRNA, as in the case of this competition, or any other aspect of it, small molecules or proteins and these other things with media, it's something that very few people understand, this level of depth. And so competitions are huge. The best recent example of that is what happened with AlphaFold by DeepMind. What's catalyzed AlphaFold is the presence of the CAASPP competition sets for certain structure prediction. So I think it's a huge deal. And something that can standardize a field tends to be the precursor to breakthroughs.
11:24And then from the perspective of work that we do at NVIDIA, what we try to do is take the best work that is out there, curate, and some of the best approaches and the best algorithms, and especially ones that benefit the most from NVIDIA's platform. And we want to build them in a way that lowers the barrier to adoption, that accelerates progress, ease of use, speed, and so on. So we follow these competitions, whether it's Casp or this competition, because we want to understand where the breakthroughs are going and what can we do to democratize existing breakthroughs and hopefully enable the next one.
11:59And so in kind of a broader sense, what's the role that the software that you're working on plays in drug discovery? How does AI, how does software aid in, you know, have a little bit of an understanding of, you know, looking for the best model to try to solve for the instability of mRNA and in cases where the super low temperature refrigerators that Boyan was describing might not be readily available. But kind of in a broader sense, what's the role that software is playing? Yeah, so you can think of it and observe. In the drug discovery process, there are several data modalities that dominate the data analysis.
12:37People study small molecules, proteins, DNA, RNA, so it depends on what kind of drug discovery problem you're involved with. And, you know, there's been work with AI for quite some time using traditional machine learning methods and then more recently deep learning. And then most recently, one of the trends that's taken place is the adoption of transformer-based models and unsupervised models. They're then fine-tubed for specific products. So I would say before that, we kind of had a period where supervised models were the dominant approach. And now we're shifting to this approach that we see also in natural language processing.
13:17We have these massive unsupervised models. Some people call them foundation models. A lot of them tend to leverage the transformer. architecture. And so what is different is that in this new paradigm, there's all more data that you can take advantage of because it's unsupervised. And so data sets maybe a thousand times or even larger than that than before. And then the other trend that is happening is that those models tend to be much larger in size. So in the past, we would typically have tens of millions, maybe hundreds of millions of parameters. In the world of NLP, we're already approaching a trillion.
13:53And the And like Sciences, we already have examples where you have upwards of 10 billion parameters. And so that poses a new computational challenge in how do you enable people to do this kind of large-scale training? But then also, how do you do inference? You may need a multi-GPU, multi-note setup, and what is the right infrastructure for that? So the software that we build at NVIDIA tackles those two broad categories of problems. How do you streamline the training process? We have the Bionemo framework for that. We also have a managed cloud service where we simplify the deployment problem, the inference.
14:28So you mentioned that part of your work is keeping tabs on these competitions, curating kind of the best of what's out there to incorporate into the technology that you're building. When researchers, drug companies, medical professionals, whoever it is that's ultimately using the software that you're building, when they're using the software, are they also still crowdsourcing, to use that term, to help advance their solutions? Or what's that stage of the process like? You know, I think it comes in all flavors. And I think some users, if it's very sensitive work with substantial IP, it tends to be closed.
15:12But then you also have people doing this in academia. And Stanford is a great pioneer in the space and this competition that demonstrates that. So on the academic side, things had to be more open. And then with some startups, things can be open. So it just depends on the particular problem or part of the space there. The open vaccine competition that we're talking about, I always say there was a huge urgency. It was a worldwide pandemic and the race was on, to use that term, to find effective vaccines. Were there learnings? I mean, I assume there were, but were there specific learnings from this process that have advanced the state of the art in drug discovery, whether relative to the speed of discovery or techniques involving data set creation or, you know, training the models themselves.
16:01What were, you know, what were some of the big takeaways that might be kind of pushing the field forward now? I'm not 100 % sure about it because, you know, this is, again, just very small amount of time that I spend within the whole world of drug discovery. Right. I think the reason we do have a paper to begin with is that there were discoveries that this technique is useful and can generalize and can be readily applied. We open source all of the code for all the winning solutions, including some of the lower benchmark codes. So people, if they want to try to test the stability of the potential new RNA molecules, they can use it.
16:42So there's a lot of useful stuff that came out. I think we're still in the very early stages in in terms of both RNA therapeutics, as well as making it more stable for general purposes. So I think that, you know, paper just came out officially a couple months ago, like it presented preprint version for like a year. But I think only now people are really starting to take a look and think about how to apply this to other problems they're looking at. I think this is more my understanding than the actual subject of better expertise, is that there's potential to use RNA vaccines also for flu or some of the other seasonal infections.
17:23So I think a quicker ability to iterate through the process would actually be very helpful. Why the setting is that right now that flu vaccines take like 18 to 24 months to develop. And if we can do it in like much more period of time with more stable molecules, there can be a much bigger impact on seasonal flu as well. I think that's exactly right. Right. So I agree. So the way to look at it is RNA-based work. You know, it is less established than severe kind of bread and butter, small molecule, drug discovery. And so this challenge brings a lot of awareness to that. And it also brings awareness to what is possible with the deep learning in the space.
18:08And I think we're just getting started. Yeah, I'll continue that. For instance, all the top models for this solution use very simple RNNs, which have been around for at least a decade since, you know, deep learning revolution started. And partly the reason for that was that we didn't have enough data. So, you know, the whole competition used about 3 ,000 molecules, which when I heard about it, I was like, I don't know if you can even pull it off. And to my big surprise, it was actually very successful. but I think, you know, for RNA and artificial, we are still in the early stages compared to like DNA or proteins.
18:43So like there's a lot of unsupervised data, as Johnny mentioned, but which we can kind of leverage and come up with even better solutions than potentially this competition provided. Aside from the availability of more and more data, which obviously is a huge thing to set aside for a second in machine learning and AI broadly, let alone in RNA research. What are some of the hurdles in front of you or perhaps even out on the horizon a little bit when it comes to using AI in drug discovery that, well, both of you, but Johnny, that you're working on that are really the big things you have to clear to get to kind of those next milestones?
19:21Yeah, let me take this one and maybe Booyim could follow up on some of this in this RNA context. So challenges with machine learning, especially deep learning in the job discovery space. There are multiple challenges of different kinds. One of them is that sometimes proving value is extremely hard. So you can have a whole bunch of people jump on this trend and start doing work. And maybe from a technical standpoint, a lot of the software and algorithms, it seems to make sense. but the turnaround time to do the work with the medicinal chemists or with the people downstream that shows that it really moves the needle in terms of your costs, job development.
20:01So that can take a long time. Another aspect, more in technical terms, is this shift to unsupervised models and therefore doing them on this massive scale. That complicates things because now that means that the engineering costs, the software engineering cost in your AI operation is now completely different. Before, maybe this was a one or two engineer type problem. Maybe now this is a multi-team engineering problem. So I think there's suddenly a greater value in the engineering that goes into the tooling for training and employment, more so than before. But then it also brings this other challenge.
20:47Benchmarking can become more complicated. Let me give you an example. We have these models that what they do is they generate molecules. Just like they've generated AI sort of in the image space, and it's taken off. We have sort of a similar thing happening in drug discovery space. So the question becomes, you know, how do you test a thing that generates molecules? There's a whole bunch of, you know, metrics. I'm just kind of walking through one just to kind of demonstrate how, you know, things just become a little stranger when you have bigger models. One of the things people look at is they look at, does it generate molecules that are valid?
21:18It's like an actual chemical and not a random set of characters. And is it novel? Is it something that was not in the training data? Well, this novelty question becomes a little bit trickier when you train on massive amounts of data. Is it as simple as saying it's not present in the training set or does it need to be substantially different? Because you have seen so many molecules. you have seen a billion of them, that maybe it's not so novel what you're doing. And so we found cases where you have a benchmark. They define novelty a certain way. And then past a certain point, every model nails the benchmark.
22:00Right. And so it kind of seemed, well, did we just perfect our discovery? We didn't. It's more that as the models become more powerful, the level of rigor in how you benchmark them needs to match it. I would say one of the practical challenges is, you know, as Johnny said, that this is becoming more and more of an engineering issue. And engineers and life scientists have very different approach to problems. And I've seen a small hints of this in this experience, but I think it could become a major like culture clash between different cultures. You know, like what engineers thinks is relevant versus what, you know, life scientists think is relevant.
22:42It's like what they think is a big deal versus what other people think is a big deal. So I think we're still in the early stages of kind of merging two approaches and merging two cultures. And I'm confident that these little kinks will be ironed out as we go along. But at this stage, I think it requires a lot of more people skills than even within other situations that I've been part of. You raise actually a great point here with the cultural question. This is a very significant one. I'm glad you mentioned it. And, you know, the interesting thing is we kind of have this new breed of startups. Many people call them the tech bio startups.
23:21So it's not like a big farmer. It's not like your traditional biotech. It's sort of like a twist on the biotech in that it's very technology driven, not just in how the work is done, but how the company is organized, how priorities are set. And so this is a very new thing where I think we find that these tech bio companies, the cultural question becomes somewhat more manageable. Somehow our approach and our way of doing things and how we think about engineering is something that sort of clicks with the tech battles in a way that might be more difficult for a well-established company to streamline.
24:01So yeah, I expect it starts with... It's just like it happened with deep learning initially. It wasn't so simple for the established way of doing things to kind of ingest this thing. But all these new students and these new groups, I think it's similar. Right. When they're building with that culture in mind from the ground up. Yeah. To go back for a moment to the Kaggle competition, when I was reading about it before we got on this conversation, and correct me if I'm wrong here, but my understanding was that it was something that, and I guess there were two competitions, two crowdsource initiatives, the Kaggle and also the Eterna that was mentioned before, both involved here.
24:39but it was something that both experts in the various domains could contribute to, but also non-experts were able to make meaningful contributions to. Is that right? How would that work for a non-expert? So this is, I think, the first paper or first initiative where we actually had not one, but two crowdsourcing initiatives. I've been doing papers before where just Kaggle was part of the equation, but here you have two different platforms, Eterna and Kaggle. Eterna is like a video game that people have been playing for like i think at least 10 11 years now where the challenge is to get the right kind of folding of these rna molecules and it turns out that people's intuition is still ahead of what machines can do you know it's one of those beautiful things if you ask me that to still have advantages there right and people who have just very interesting visual intuition can actually find solutions that you know even the experts cannot really find you You always don't even have to have any kind of technical knowledge, especially for this way.
25:38It's a very interesting platform. On Kaggle, on the other hand, it's people who have in various ways come into machine learning. I always love Kaggle because there's such a wide variety of backgrounds and interests of people who participate there. People from, you know, people who have PhDs in computer science to people who have never taken computer science class in their life. they just discovered this platform that's kind of interesting to compete in. They learn some basic coding, basic machine learning skills, and then they apply it to a variety of different problems. Right, right, right. And I think what's interesting about Kaggle, and I think machine learning in general, but it's very visible in Kaggle, is that a lot of times techniques and solutions that have originally been applied to one set of problems can very easily translate to a whole different set of problems.
26:28So, you know, as Johnny has mentioned, a lot of these machine learning models were originally devised for language, various language problems, natural language processing, natural language understanding. Right. And it turns out that with very small modifications, they can be applied to DNA and RNA. So like the language of nature. Right. Applied. So it's very interesting. So people don't need to know anything about biophysics or biocatastry of RNA or DNA. They can take their, you know, natural language processing apparatus, apply it here and get pretty good results and best for students in the app.
Read the full transcript
27:05It's all coming together. As we wrap up, as we wrap up our conversation, I should say, is this project wrapped up, so to speak, or is there more work that either of you are doing specific to this mRNA stability problem? And then sort of a follow-on question to that is, what are you looking ahead to in your respective field, your respective work over the next year, a couple of years? What are kind of the big, we talked a little bit before about the hurdles, the challenges for AI power, machine learning power, drug discovery in particular, but what are some of the things that the two of you are digging your teeth into as the year unfolds?
27:46So yeah, I guess that either one of us can answer this. Right now, we are looking to developing some large language models for RNA. And it would be interesting to see if once we have them, if these embeddings that we can get from these large language models can then, you know, retrospect be applied to this problem. So I think that would be an interesting challenge to see if we can actually improve what has been done here or not, which would be very interesting to find out what Kaggle has managed to do with a very limited apparatus. It's kind of really easy that most you can get from this data set.
28:21So that would be interesting. going forward, RNA is maybe a few years, if not more than that, behind in terms of like proteins, like what kind of data we have. But it's really kind of catching up rapidly. And I think proof of its therapeutic value is there. More and more companies, more and more researchers are working on it. So I expect it will have a lot more data, a lot more interesting problems to work on in the coming years. I'll just add to that that I'm excited about the adoption of these methods broadly in drug discovery, all the way from things at the early stage, like RNA, to modalities where, you know, it's a little bit more proven, such as proteins and proteins, generative AI is taking off really quickly.
29:05But I think what I'm looking forward to is the adoption and the proven value, you know, seeing these startups and these companies not only invest in the R &D and the software, but adjust their experimental methodology, based on these methods and prove end-to-end that it's, you know, improving quality, reducing costs. I think that will catalyze, you know, that sort of next level of investment and commitment to these methods. And I think in that environment, we're going to see many more breakthroughs. Excellent. Last question for me, but are we any closer to a cure for the common cold or is that still beyond human machine understanding?
29:46You don't have to answer that. I don't think I'm aware of this kind of work applied to that problem, Boyan. I don't know if you've seen anything. No, I haven't seen anything. I think we might actually find a cure for aging before we can find a cure for COVID-19. My limited understanding of medical science would agree. Boyan and Johnny, thank you so much. I know we've only kind of scratched the surface here, but it's remarkable the advances being made, the work that the two of you are doing. And the whole role that crowdsourcing in these various competitions play, just the ability of folks to, you know, not to oversimplify it, but be able to log on to a website and actually contribute meaningfully, whether small or large, to, you know, battling these really important scientific problems.
30:33It's just mind-blowing. It's fantastic. I mentioned the paper is up on nature, but for listeners who might be interested to learn a little bit more about the specific work that either of you are doing, are there places online, do you have research pages or project home pages, social media accounts? Where are some places that people can go to learn? Yeah, this paper in particular, it also has a preprinted archive. So if you don't want to go to nature, there's an archive preprint, which I think is up to date. Maybe it says minor differences. Perfect. We also have a repo for it. And there's a Kaggle competition website where, you know, all this information also contains.
31:13Right. Great. More broadly, Drive Discovery at NVIDIA. People can go check out Bionim. We have some materials on the NVIDIA website for that. and we also have some software in early access that people can apply for and then for those who are interested we've done some work with Evozyne, a startup in the space that we shared publicly recently where they use BioNemo for some of their protein generation problems and built a really interesting model using components from BioNemo and you can check out some of the news releases around that and there's also a preprint I believe on that bio archive. Great.
31:52Well, again, Boyan and Johnny, thank you both for taking the time to come on the podcast and talk a bit about this particular project and your work in general. And it goes without saying, but best of luck to both of you and all the folks you work with and continued advances in the days and years to come. Thank you for having us.
32:24Thank you.
32:54The End
From the publisher
Artificial intelligence is teaming up with crowdsourcing to improve the thermo-stability of mRNA vaccines, making distribution more accessible worldwide.
In this episode of NVIDIA's AI podcast, host Noah Kravitz interviewed Bojan Tunguz, a physicist and senior system software engineer at NVIDIA, and Johnny Israeli, senior manager of AI and cloud software at NVIDIA.
The guests delved into AI's potential in drug discovery and the Stanford Open Vaccine competition, a machine-learning contest using crowdsourcing to tackle the thermo-stability challenges of mRNA vaccines.
Kaggle, the online machine learning competition platform, hosted the Stanford Open Vaccine competition. Tunguz, a quadruple Kaggle grandmaster, shared how Kaggle has grown to encompass not just competitions, but also datasets, code, and discussions. Competitors can earn points, rankings, and status achievements across these four areas.
The fusion of artificial intelligence, crowdsourcing, and machine learning competitions is opening new possibilities in drug discovery and vaccine distribution. By tapping into the collective wisdom and skills of participants worldwide, it becomes possible to solve pressing global problems, such as enhancing the thermo-stability of mRNA vaccines, allowing for a more efficient and widely accessible distribution process.
Don't miss this enlightening conversation on the transformative power of AI and crowdsourcing in mRNA vaccine distribution.




