In short
NVIDIA AI Podcast Episode Summary
Episode Title
Anima Anandkumar on Using Generative AI to Tackle Global Challenges - Ep. 203
Podcast Description The NVIDIA AI Podcast explores how the latest technologies, including generative AI, are shaping our world and addressing global challenges. This episode features Anima Anandkumar discussing the potential applications of generative AI in scientific research.
Guest Introduction
- Anima Anandkumar: Bren Professor at Caltech and Senior Director of AI Research at NVIDIA.
- Recognized for her contributions in developing AI algorithms for scientific applications including drug design, weather forecasting, and scientific simulations.
---
Episode Highlights
Importance of Generative AI
- Described as an “inflection point” in AI technology.
- Capable of learning not just natural languages but also the “language of nature.”
- Potential applications in predicting coronavirus variants and extreme weather events.
Key Applications Discussed
- Genomic Research:
- Generative AI can analyze DNA, RNA, and other genomic data.
- Example: Predicting dangerous coronavirus variants, facilitating quicker vaccine and drug development.
- Weather Forecasting:
- Challenges in predicting natural events due to the complexity and number of variables.
- Generative AI models can enhance predictions of hurricanes and heat waves.
Methodological Insights
- Beyond Scaling Up: Emphasis on fine-tuning models rather than just increasing compute power.
- Importance of embedding physical constraints and capturing multi-scale phenomena in models.
Responsible AI Development
- Advocacy for strengthening laws to prevent misuse of AI technologies.
- Ensuring models are used responsibly in scientific applications, especially in sensitive areas like healthcare.
---
Discussions on AI's Future and Societal Impact
AI’s Transformative Role Across Industries
- AI is reshaping job roles and workflows in various sectors.
- Emphasis on the importance of asking the right questions in research rather than solely seeking answers.
Challenges in AI and Policy Recommendations
- Need for policies that focus on the downstream impacts of AI applications.
- Importance of transparency in AI models (e.g., through Model Cards), detailing training data, intended use cases, and metrics on fairness and privacy.
Insights into the Future of Work
- Advice for young individuals entering the technology field: focus on lifelong learning and algorithmic thinking rather than specific programming languages.
- Generative AI and machine learning will augment human capabilities but not replace the need for critical thinking and problem-solving skills.
---
Conclusion Anima Anandkumar emphasizes the potential of generative AI to address significant global challenges and the importance of responsible development practices. She advocates for the integration of interdisciplinary knowledge in AI applications to ensure they are aligned with societal needs.
Listeners are encouraged to explore additional resources and discussions from Anandkumar, including her contributions to the President’s Council of Advisors on Science and Technology and NVIDIA Research.
Additional Resources
- Full recording of the PCAST talk
- Jensen Huang's Berlin Summit address
- NVIDIA Research homepage: [NVIDIA Research](http://nvidia.com/research)
---
This summary captures the essence of the podcast episode while highlighting key concepts, discussions, and recommendations that emerged during the conversation with Anima Anandkumar.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:10Hello, and welcome to the NVIDIA AI Podcast. I'm your host, Noah Kravitz. Our guest today is no stranger to NVIDIA, nor to the podcast. Anima Anankumar joined the show nearly three years ago to talk about her then-personal record of having seven of her team's research papers accepted to NeurIPS 2020. To say she's been busy since then would be a bit of an understatement. Anima is a Bren professor at Caltech and senior director of AI research at NVIDIA. Her work developing novel artificial intelligence algorithms enables and accelerates scientific applications of AI, including scientific simulations, weather forecasting, autonomous drone flights, and drug design.
0:49She's also a fellow of the ACM, IEEE, and Alfred P. Sloan Foundation, has received Best Paper awards at venues such as NeurIPS, and the ACM Gordon Bell Special Prize for HPC-based COVID-19 research. And she's part of several world-renowned foundations and councils. We'll get to all that in a moment. I could go on with a list of accolades, but I think we'd all rather hear from Anima herself. So let's get right to it. Anima Aminkumar, welcome and thank you for coming back to the NVIDIA AI podcast. Thank you, Noah. It's a pleasure to be back and great to see all the amazing things we're doing at NVIDIA and in the broader AI community.
1:28It's an exciting time to be sure. We were speaking briefly before we hit record and I think that last conversation was in the midst of the pandemic. So lots has obviously happened since then. And let's start, obviously, with AI and the world of generative AI, which, as we record this, is all over the news these days. And my non-technical friends and family members have been asking me all about, wait, wait, wait, what's this AI thing? Haven't you been paying attention for the past 20 years? So let's dive into it, starting with your recent trip to speak with the President's Council of Advisors on Science and Technology.
2:03You spoke that of AI's potential to make a huge impact on science, but obviously AI's potential impact your own purview, not to discount science by any means, but is larger than just the scientific community. So maybe you could just kind of take the reins and talk a bit about the moment we're in, the President's Council, Gen AI, and we'll kind of go from there. Yeah, it was such an honor to be part of the three speakers who spoke about AI and science to PCAST or the President's Council on Science and Technology. This was a virtual talk and panel, in fact, that's available online for everybody to go and watch.
2:44And to me, what really came out of that was this aspect of how generative AI is an inflection point in our lives and how can we harness it for all the benefits, right, to society and to humanity through scientific applications. So first of all, for the public here, you've heard of the term generative AI, but what is so exciting and different from this previous decade of AI developments is this ability to generate de-noble or from scratch, entire paragraphs of text, very realistic looking images or creative images. So we are seeing those that are about generation, whereas the previous decade was about discriminative AI, where given an image, we would reason about, oh, is this face belonging to the person who's claiming that identity or so on?
3:45So it was more about discrimination, you know, what's in the image, what's in this piece of text, rather than generating it. And generation was always considered to be very hard, because it's much more high dimensional, you're specifying some very few properties, you're saying, oh, generate an image with a cat, right? There are so many possible ways the cat can be and all the variations that we need to capture to have that ability to sample from that whole distribution. And we're able to do that. So that's an exciting time. And so what I delved into in my talk to the Science Council was to ask, you know, how can this be harnessed for scientific domains?
4:29And what is there beyond text and images? You know, because text and images are also present in science, we can harness scientific knowledge through text, right? But there's so much beyond just the English language. And one of the examples I started with was showing how instead of the thinking about generating paragraphs of English language or any other natural language, why not think about the language of the genomes? We took all the DNA data that's available, both DNA and RNA data of virus and bacteria that are known to us, about 110 million such genomes. We learned a language model over that.
5:13And then we can ask it to now generate new genomes. This was also during the pandemic where we said, oh, let's now focus on coronavirus and get to generate new variants of concern. And we were able to predict that before they emerged. And we are also able to ask, which are the dangerous ones with strong binding? And how can we reason about that and be better prepared with vaccines and drugs before they emerge? And so that's one kind of direct application of what text language models can be now brought into genome language models. And we can get the benefits immediately. You mentioned that the English language obviously is not the only thing we can input and output in these generative AI models and LLMs.
6:06And we've had a couple of podcasts. I say this, we've recorded them recently to the listener. Not sure when they've been out. If they're not out yet, stay tuned. But a couple of podcasts with folks working on AI models using LLMs for code, for writing computer software. And again, it's a language and it's maybe for a lot of people, it's not the first thing they think of, but then it's kind of an easy mental leap to think, oh, right, of course, that's a language. It's text and symbols and such. When you're talking about the language of genomes, and feel free to extrapolate into other scientific areas as well, how similar or different is the process of training and working with an AI model using the language of genomes as opposed to English or another written and spoken language?
6:53Yeah, that's a great question. And that really depends on the domain, right? In our case, this was still discrete. So it's now you have the nucleotides, you have ATTC instead of, right, like English language letters. So that way we could like kind of, you know, think of like taking in these discrete tokens or, you know, think of other abstractions like that. But the bigger challenge is that the genomes can be very long. We looked at bacteria and virus. Think about going all the way to humans. You have extremely long sequences that our current models are not really able to take in that kind of context.
7:35and some of the latest work is really looking to asking, can you take very long contexts? And so what we did was to think about these long-range dependencies being captured in the latent space of these models, and that way create more realistic generation of valid genomes, the whole idea is by looking at all existing virus and bacterial genomes, we are implicitly encoding fitness and other functions, you know, because these virus and bacteria survived, they probably thrived, right? They created havoc. So we're taking that examples and asking by forcing this into a language model learning bottleneck, can you come up with that encoding and can you come up with functionalities, right, of different genes and proteins?
8:22So that's how we can kind of get insights from the model, even though we don't have explicit aspects of fitness and other functionalities of what these genomes are doing. You mentioned working on the coronavirus, or I should say working against the coronavirus, and predicting, and forgive my lack of precise language, but predicting new variants and which ones might survive, and then obviously drug discovery as relates to coronavirus. Taking that sort of a step broader to thinking about some of the other global challenges, extreme weather is obviously a big one. I'm sitting in California in the middle of July right now.
9:06And, you know, whatever direction you look, there's wildfire threats, there's flooding, there's, it's all around us. And I think we're all more and more aware of it. Can you speak a little bit about the potential or even the progress? And we can even use that 2020 recording as a milestone, if that's helpful, but the potential and the progress of leveraging AI to tackle some of these big global scale issues. Yeah, absolutely. I mean, there are so many scientific challenges we are facing today, right? I mean, going back to even just the coronavirus example that I'm talking about, you know, yes, we can generate new variants of concern through this language model trained on genomes, but to reason, what do these new genomes do?
9:50we have to look at how do they bind with molecules in our body, for instance, right? Like, you know, if there is strong binding, they may have a bigger impact on our health compared to when there isn't. But those kind of like molecular binding dynamics is a highly multi-scale process. You know, you can think about effects all the way from quantum scale, right, to, of course, our full body, right? In all of those different levels of organization that all come into play. And that's why, you know, immunology is a very complicated subject. And for us to even, you know, make sense of, okay, you know, looking at a specific combination, right, a protein and a molecule, how well do they bind?
10:37And, you know, it's not trivial. People think about docking, but also the aspect of now the protein backbone can change in many cases as the binding is happening. So it's a highly dynamic process and that involves multiple scales. And the same is true with the extreme weather forecasting that you talked about, right? There are very small scale phenomena that happen in terms of how the clouds, you know, have turbulence or precipitation is something what people call microphysics, right? So this is very complicated and not fully understood even today, and definitely not having the scale, even in our biggest supercomputers, to be able to faithfully replicate all of this.
11:20And that was kind of my segue to saying text and image models are great, but so much of our scientific domain data cannot be captured just by that, because text limits to discrete tokens. These are continuous processes. And image models that we use it for, say, applications like mid-journey where you generate, say, new images, it's all limited to a fixed resolution, right? And mostly in the natural image generation world, you're thinking about bigger objects and you want to mostly get the shapes of the objects right and some nice looking texture. Whereas if you just go stare at a hurricane, you can't tell how it evolves, right?
12:03None of us have that ability. So it's not just about human visual perception. And we certainly cannot predict where the hurricane will go in the next week or so. So longer term prediction is even harder and chaotic. There's even inherent uncertainty there. And those are the aspects We've been working at NVIDIA, at Caltech in collaboration with many other organizations as well to say, how do we capture these multitude of scales that are present in the natural world? And even when we don't have data for all the scales, it would be just impossible to be able to simulate through current numerical methods.
12:47and also we may not know the full physics or the full equations that count them. But with the limited data we have, can we hope to extrapolate to finer scales? Can we hope to embed the right constraints and come up with physically valid predictions that make a big impact? So hold me back in if I go kind of the wrong way in asking this question. But I know that a big part of your role at NVIDIA is kind of leading the charge on developing next-gen AI algorithms. And in the five or so years that I've been lucky enough to host these podcasts and talk to all kinds of people and all kinds of stuff, one of the things that's emerged was that there was originally this drive for more compute that a lot of people were talking about.
13:36And it feels like, and I'm speaking very broadly again, so I'm eager to hear your response. It feels like more recently, there's been a little bit of a shift towards, okay, a lot of folks have enough compute. And now we're looking at developing the next gen of whether it's algorithms or going from these large, like general LLMs to fine training, maybe smaller ones, but that are more fine-tuned or trained on more specific data sets for kind of task-specific things. Your world, I imagine, is somewhat similar and probably quite different from what a lot of folks are doing out there with LLMs or other AI tools.
14:15I guess how much of your work is sort of gated or accelerated by hardware and compute availability and kind of always, are you always just waiting for the next gen of compute to be available so that you can really push towards the next big breakthrough? Or are we at a point where actually the focus, even for someone as working as in-depth as you are, let's put it that way, is also kind of more focused on, OK, I don't need more giant models that take huge resources to train. We're actually looking at sort of a finer scale tweaking of things, to put it that way. I mean, at NVIDIA, we are lucky to see the full stack.
14:55Of course. Right. Like always, you know, even look about, think about the next generation hardware before they emerge, right? Like Grace Hopper, for instance. And that's very well suited, in fact, for many of not only large language models, but also weather models and other scientific domain models that require memory. So we're always mindful of those hardware constraints and efficiency gains that we can get and design algorithms with that in mind. And scale is important, you know, with the weather forecasting work, you should, all the, you know, audience here should go and look at our CEO Jensen Wong's Berlin talk.
15:36It's the EVE Summit. It's already about a million views on YouTube within a few days. And yeah, there he puts it really nicely, right? So there is now the, you know, the next generation hardware that is going to really make these models more efficient. But you do need algorithmic design as well, you know, at least. Hopefully, we are not out of work anytime soon. I know there is the notion of like a bitter lesson that, oh, just throw in more data, throw in more compute and out comes magic. We don't need to think. Thinking, in fact, is harmful to this process, right? And I've seen these kind of comments on social media.
16:17I would say in the realm, especially of scientific domain, this will get us nowhere. I mean, it'll get you somewhere, but not a long way. To give you an example, for weather forecasting, there is a wealth of historical weather data available, and it's reanalyzed, meaning the observations are assimilated and you kind of add in through physical assumptions, right, any kind of gaps that are in observations. and we have that for several decades and tens of terabytes of such data, right, openly available. So you can consume and build models and people have done vision transformers similar to the image models, right?
16:56And for what we call medium range forecast, meaning one to two weeks, you can get some reasonably good forecast with that models. But the distinguishing features from those general purpose image models versus the models we use for scientific domains, what we call neural operators and four-year neural operators is that we are, in a principled way, able to predict at any resolution. So the idea is, you know, we have this, say, weather data that's available for training only collected at 25 kilometer, right, due to the constraints of observation and the data simulation. But we know the real world is continuous.
17:36So we need to incorporate the aspect that, oh, can we have the model also adaptively interpolate and come up with a zero-shot super-resolution prediction? So you want the flexibility to be able to predict at a different resolution, consume training data at multiple resolutions if available, because that's ultimately the underlying phenomenon we want to predict is continuous. And so that's the aspect that our models incorporate. And also in this case, because we are predicting on the globe, which is a sphere, we also incorporated the fact that the Fourier transform should be on a sphere. So this is what we call an inductive bias, right?
18:17So the more of the inductive bias of the domain we bring in, what helped this was not only to get really good forecasts in a deterministic way in the short term, but to also get a much longer term stability, meaning even if you now run this for several months, you have models produce something physically valid, meaningful, and be stable. And this is because we are embedding the right properties into the model. And so this is what I call in the realm of extrapolation is where you see the differentiation, right? So in the sense if your predictions are pretty much in the same distribution as the training data, a lot of models that fit to the data will do reasonably well.
19:06But in scientific domains, so many times you are looking for extrapolations. You want to look at longer term, seasonal to sub-seasonal forecasts. We also want to think about extreme weather events. One of the demos that Jensen showed in his Berlin Summit talk that's available on YouTube is the ability to predict extreme weather like hurricane and heat waves. And for this, you need a larger set of ensembles. Our speed-ups enable that. We have speed-up of tens of thousands of times over current weather models. So we can do a much more richer set of ensembles. but also by correctly calibrating that, meaning we are not thinking of this as just one deterministic forecast for the next few days, but probabilistically asking what is the risk assessment.
20:03This is very critical for tail events, for extreme weather events, right? What matters is the tail. And again, the models that do not capture the right physics will fail to extrapolate to tail events, because that's like the margins of the distribution, right? Whereas machine learning tries to fit to the data of like normal events. So I think that's where, especially in scientific domains, I think the better lesson doesn't hold. And we need to be much more mindful in building the domain knowledge, the domain constraints. For instance, in some of the other examples, say, you know, when we are looking at fluid dynamics and turbulence, we know the equations, We know Navier-Stokes as the equation we want to satisfy.
20:48So bringing that into the framework will also really help us really capture the fine scales well. And our neural operators are able to do that because even if our data, you know, we are forced to get only coarse scale data because, right, it's very expensive to generate it through simulations or the observations we collect is only coarse resolution. we can embed the physics in a finite resolution all directly while training the model in a seamless way. And that's the benefit that neural operators have. We're speaking today with Anima Amankumar. Anima is a senior director of AI research at NVIDIA.
21:29And as we've been talking about, her work spans the world of computer science and AI and also the scientific domain, which extends into many arenas. And she's also a Bren professor at Caltech. And I wanted to switch gears slightly and ask you a little bit about your work at Caltech and specifically your work in the Tensor Lab there. Could you share some of your current work and its alignment with the broader conversation around AI's potential? Yeah. I mean, you know, to me, Caltech is a great place where a lot of interdisciplinary work happens. It's a small community, but very tight knit. And yeah, the neural operator work started here by collaborating with Andrew Stewart, Kaushik Bhattacharya, experts in numerical methods, Kaushik in the realm of material modeling.
22:20And so, you know, again, this aspect of building the right architect model architectures and algorithms by looking at all the wealth of knowledge that has been developed, right, for almost a century with how partial differential equations are solved and incorporating that. Because, you know, again, it's very easy to go wrong in these domains, because if you don't capture the fine scales, you're completely off. So really building that into the model came about here. I also really like Richard Feynman's quote, who was at Caltech, perhaps one of the most famous professors who was at Caltech. And he says, what I cannot create, I do not understand.
23:06And that is so apt for this era of generative AI, because I really think generative AI is bringing us to this realm of both scientific understanding, right, really domain understanding. And I'm really happy Jensen also used this quote in his Berlin Summit talk as well. But to me, that code is like a showcase of the kind of long-term thinking that's been here at Caltech. That's very much the fabric of this university is to think about some of the hardest challenges and how to frame it in a way that we can make headway towards that. I get to work with scientists across multiple domains. I mentioned material modeling in chemistry, working with Francis Arnault, who is Nobel Prize winner here, thinking about how directed evolution and machine learning can really enable us to have the next breakthroughs there.
24:12thinking about seismology. You know, Celtic has had a long history of being able to predict earthquakes well in advance and understanding the physics. And there, again, these tools can be very effective because we don't know the ground truth. We don't know what happens underneath the earth, right? We only have sensors on the ground. So our ability to do very fast reconstruction and inverse modeling, what we call, what is the probability of the source of the earthquake at different locations. It's directly implied, right, is one that impacts our safety. And so again, our models are tens of thousands of times to hundreds of thousands of times faster than current simulation methods.
24:59And with that speed up, we can do now a much greater set of ensembles So we can get the right risk assessment and we can do what we call inverse modeling and also inverse design. Another example is collaboration with Kira Dario here at Caltech, where we design a better medical catheter that reduces bacterial contamination by about two orders of magnitude. Again, a great example of interdisciplinary research where we used neural operators to model the fluid flow and how the bacteria tend to swim upstream to the flow, right? Bacterial density as a function of the flow. And if we can build like triangular shapes inside the catheter pipe, then how much does it stop the bacteria from swimming upstream?
25:50You know, very simple. Right, right, yeah. Right. But the optimization doing that directly with AI in the virtual realm, you know, to the physical experiments was a huge time saver. And then, yeah, we know this collaborators went and did those 3D printed experiments in the lab, saw this benefit. And I think that's also this broader aspect of what AI can do. It can speed up simulations. It can speed up modeling. We can get better probability estimates as a result of that. And we can also invert because these AI models are differentiable. So we can explore the design space and the space of different hypotheses much more effectively.
26:35Think of it in a virtual lab before bringing that on to the physical labs. Yeah. Now, we mentioned that PCAST earlier, the President's Council. You're also a member of numerous other advisory councils and networks and such, one of which is the World Economic Forum's expert network. We're talking about Gen AI. We've been talking a lot, as you say, you've been talking a lot about the need and problems associated with longer term predictions. So this is kind of a long term and short term thing, at least from where I'm sitting. The explosion of Gen AI into kind of the mainstream consciousness and the LLMs that have been available, text and image, and people have been using them have sparked, let's just say, a lot of both positive and imaginative thoughts in the general public, and then a lot of worry and consternation over the future of job markets, economies, humanity as a whole.
27:27From your view, and sitting on these different councils, and specifically, again, the World Economic Forum's expert network, what, if any, key policies do you see the need for and to be prioritized if we're going to kind of continue advancing the field of AI in a responsible and constructive way that hopefully benefits all humanity, or at least kind of, you know, rising tide, lifting boats, as opposed to negative outcomes? What do you see as kind of the big things to be discussed and acted upon? Yeah, I mean, I think we should always be thinking about the downstream impact of these models. And I think it's very hard to think about regulating and trying to stop these models.
28:13I feel like any such attempts usually, because the system is so complex, usually puts more power in the hands of bad actors, right? So kind of the saying, when something is outlawed, it's the outlaw that becomes the law. Exactly. So we have to be right. They're not going to stop. Right. Of how to operationalize the good intent. And I do think focusing on the end applications is the key because, you know, think of LLM being used for mental health diagnosis or counseling or any kind of health advice. We should be much more mindful of what it says and how it's being exposed to patients, if at all, right?
28:59And how should the human be in the loop and all that compared to using LLM or poetry, right? So I think the use cases matter. And I understand the questions of misinformation and others are indeed very hard to tackle because we can have these models generate at scale. But again, right, I feel like bad actors with enough resources would always have access to do this if we try to limit open source and if we try to limit the broader public from being able to access and do meaningful things with the model. To me, the AI revolution so far has been a lot of open source revolution, right? And I do see the challenge with very large models, how much of it should be open source or not.
29:49That's to be determined. But I feel like first we should think about also strengthening the existing laws for all kinds of downstream application, right? And that'll really help us kind of get started on what could be the most dangerous use cases and how to limit the harmful effects based on that. I also think, you know, going back to research and trustworthy AI and best practices is so important. At NVIDIA, we launched Model Cards++, which is an add-on to the Model Cards that was proposed by Meg Mitchell and Timnit Gibru. Could you actually just, for the audience, kind of explain, I was going to do it, but you'll do a much better job, what a Model Card is?
Read the full transcript
30:33Because I think it's a fascinating and very simple as a lot of these solutions tend to be in their base, you know, idea to help with this. Absolutely. Right. Model courts is all about transparency that, you know, you want to say what was the training data that was used. Right. What is the intended use case? So if this model is now used in a scenario where it's not intended, right, then that's already a red flag. So it's all this list of things. And we added many quantifiable metrics, right? How fair is this model? How private is the model? If you care about privacy of like the, say, health records on which this is trained, you know, can we give metrics like differential privacy?
31:18So we added a number of such metrics. I mean, this is available online and some of the NGC models launched by NVIDIA have this, right? So I think this is a great effort that is started, but we need to put this on steroids for these large language models. And some of the research we are doing is also asking, how do we automate the testing of various models? Because Hugging Face, for instance, has a huge repository for source language models, right? But how biased is one model compared to the other? And that also depends on the use case. Maybe you have very specific requirements, say, for healthcare, right, where you may want some of that differentiation of different demographics, but you don't want bias.
32:03So what is undesirable and what is not? And so we are now developing tools that really gets users who are not machine learning experts, right, put in like what social groups would they like to test the bias on? What are the attributes and dimensions along which they want to do the testing? And I think we really want to see more and more of such tools that promote transparency and explainability. So we don't have a lot of time to dive into your background and it's fascinating merging of, I don't know if merging is quite the right word, but the scientific and the academic and industry and government and NGO things and all of that.
32:42But a question popped into my head earlier when you were kind of joking about, you know, we do need now and hopefully we'll continue to need people thinking about how to design the algorithms. I read something earlier this week, which kind of harkens back to earlier this year when I, as a writer and content creator by trade, had my own sort of excitement and fear at the same time of what are these LLMs going to do? And, you know, are they going to become super creator or out of work and how fast? I was reading something recently where somebody was talking about giving advice to his own child. And originally, whatever it was they said they wanted to be when they grew up, you know, the advice was, we'll go to school, you know, learn all you can, get some experience, et cetera.
33:29And then they started taking an interest in technology and computer science specifically. And, okay, well, go to school for computer science and learn the fundamentals and learn the languages that are in vogue at the time, and you can get started that way. And now with the explosion of LLMs, this person was saying they're now at a moment of not knowing what to tell their kid to do because will there even be a need for software coders by the time the 14-year-old is ready to join the workforce. So not to throw all this on your shoulders, but I'm going to anyway. Given your career path and now the point where you sit and having this deep dive hands-on view as well as the broader scope of everything we've been talking about, what advice might you give to a young person who feels right now like they're interested in a career, in technology, in science sort of broadly, but they're also aware of these advancements that feel like they might actually be accelerating the rate of change in these domains.
34:37Is the advice still to go study the scientific disciplines, the bases of computer science and how to build software and all of those kinds of things? Or has it actually shifted because of the developments with AI that you've seen and been a part of? Yeah, I think AI will have a huge impact on education, right? So I was at TED where, you know, the Khan Academy founder, Salman Khan, talked about Khan Mego and having chatbots that are very personalized and really help people kind of wade through any barriers in thinking, right, and really learn. And to me, I think learning will never stop. And in fact, the best advice I can give is be a lifelong learner.
35:24You know, that holds true for language models. That also holds true for humans. You really need to keep updating, otherwise you're out of date, right? So, and I think that's much more important in future than even now. And, you know, to me, computer science isn't about, oh, will it be Java or C or an AI-based? model, right, that's coding, but it's really about algorithmic thinking. So, you know, we still need to frame like what are the right specifications and what would we like AI to help with or even maybe do some of the tasks, but we still are the thinkers and just our notion of thinking will change.
36:05I mean, for instance, recently we released a framework called Lean Dojo that uses Lean framework for theorem proving with language models, right? So you're kind of instructing this lean theorem prover to go and try different premises, try to prove a theorem, and ultimately, right, you're guaranteed that the theorem is always going to be right, as opposed to language models that may hallucinate. So this ability, right, means, you know, all the mathematicians just completely obsolete. I don't think so. I think these really help maybe prove very laborious theorems, right, to kind of like make sure they're correct.
36:45And maybe over time may even help aid how they work together, right, the humans and the machine together. And I think we'll just go on to solving harder and harder problems this way. And I think that'll be true for every scientific domain, the neural operators and other models we're building is like really enhancing our understanding of different phenomena. And we can, you know, it's not just the intuitions of a domain scientist, it can help and aid and complete that by saying, oh, let's just look at a broader set of design possibilities. Maybe you never thought of it before, but I can, you know, AI can just very quickly rule it out or optimize through it and come up with something better, right?
37:29So, but we still humans, then what are we left to do? What we are left to do is to think about what problems to solve. This is also the research advice I give to everyone. The most important thing is the question, not the answer, right? Because the answer is always 42. Exactly. Well said. So it's really the question. And take all the time to debate, are you asking the right question? Are you solving the right problem? Don't just go by trends. I mean, the trends inform us something, but really ask, what are the unique contributions that you can make, right? And so that's, that holds true for research, that holds true for education, you know, focusing on learning and the pleasure of really understanding something.
38:12I think that'll never go away. Fantastic. Well, I'm going to carry with me what you said. You know, the LLM needs to keep learning and so do you. I love it. Anima, this is fantastic. We could talk for hours, but you clearly are working on and thinking about a million different things and have other places to get to. So we appreciate you taking the time to join us. The PCAST full recording, as you mentioned, is available on YouTube, as is Jensen's Berlin Summit address, which we encourage everybody to go check out. NVIDIA Research has a homepage at nvidia.com slash research. Where else, if anywhere, would you direct folks who want to know more about the work that you and your various teams are doing.
38:55Yeah. I mean, those are great resources. You know, sometimes I'm on social media, my website at Caltech Tensor Lab, you can find online. But yeah, it's really a great place to be here today. And thank you all. Oh, thank you. And let's do it again, maybe sooner than two and a half years if we can. Look forward to that.
39:42¶¶
39:56Thank you.
From the publisher
Generative AI-based models can not only learn and understand natural languages — they can learn the very language of nature itself, presenting new possibilities for scientific research.
Anima Anandkumar, Bren Professor at Caltech and senior director of AI research at NVIDIA, was recently invited to speak at the President’s Council of Advisors on Science and Technology.
At the talk, Anandkumar says that generative AI was described as “an inflection point in our lives,” with discussions swirling around how to “harness it to benefit society and humanity through scientific applications.”
On the latest episode of NVIDIA’s AI Podcast, host Noah Kravitz spoke with Anandkumar on generative AI’s potential to make splashes in the scientific community.
It can, for example, be fed DNA, RNA, viral and bacterial data to craft a model that understands the language of genomes. That model can help predict dangerous coronavirus variants to accelerate drug and vaccine research.
Generative AI can also predict extreme weather events like hurricanes or heat waves. Even with an AI boost, trying to predict natural events is challenging because of the sheer number of variables and unknowns.
However, Anandkumar explains that it’s not just a matter of upsizing language models or adding compute power — it’s also about fine-tuning and setting the right parameters.
“Those are the aspects we’re working on at NVIDIA and Caltech, in collaboration with many other organizations, to say, ‘How do we capture the multitude of scales present in the natural world?’” she said. “With the limited data we have, can we hope to extrapolate to finer scales? Can we hope to embed the right constraints and come up with physically valid predictions that make a big impact?”
Anandkumar adds that to ensure AI models are responsibly and safely used, existing laws must be strengthened to prevent dangerous downstream applications.
She also talks about the AI boom, which is transforming the role of humans across industries, and problems yet to be solved.
“This is the research advice I give to everyone: the most important thing is the question, not the answer,” she said.




