In short
a16z Podcast: Episode Summary - Can AI Advance Science? DeepMind's VP of Science Weighs In
Podcast Overview The a16z Podcast explores technology and culture trends, featuring discussions from industry experts and leaders. This episode focuses on the impact of AI in scientific research, with a particular emphasis on DeepMind's innovations, such as AlphaFold.
Episode Details
- Episode Title: Can AI Advance Science?
- Guests: Pushmeet Kohli (VP of Research for Science, DeepMind) and Vijay Pande (General Partner, a16z).
- Release Date: Early 2024
Key Concepts
AI's Role in Science
- AI is increasingly seen as essential for scientific exploration, moving beyond traditional methodologies.
- The impact of AI tools can lead to breakthroughs in understanding complex biological processes and structures.
AlphaFold
- AlphaFold Overview: Launched in 2021, it's a groundbreaking AI model for predicting protein structures. Utilized by over 1.7 million scientists globally.
- Significance:
- Transforms years of work (often a PhD's worth) into quick predictions.
- Facilitates advancements across fields such as drug discovery, genomics, and structural biology.
DeepMind's Vision and Initiatives
- DeepMind's focus includes various projects in biology, chemistry, and mathematics, highlighting the multidisciplinary nature of their research teams.
- Key Projects:
- Graphcast: AI for weather forecasting.
- Alpha Geometry: Addresses advanced geometry problems.
- FunSearch: Explores mathematical discoveries using large language models.
Discussions and Insights
The Cultural Shift in Science
- There is a growing understanding that AI can uncover scientific insights that were previously deemed too complex to decipher.
- The notion of what is "ridiculous" in scientific endeavor is changing; AI is making it feasible to tackle problems that human minds alone cannot.
The Future of AI in Research
- AI is not just a tool for efficiency; it has the potential to revolutionize how research is conducted, especially in life sciences and healthcare.
- The discussion emphasized the need for a paradigm shift in how scientific methods incorporate AI, allowing for faster and more comprehensive insights.
Challenges in AI and Science
- While AI presents vast potential, challenges remain in areas such as data availability, model accuracy, and ethical considerations.
- The need for robust datasets is emphasized; many biological data sources are either inaccessible or underutilized.
Economic Implications of AI in Research
- AI technologies are expected to dramatically reduce the cost and time associated with research, enabling more efficient drug discovery and clinical trials.
- The concept of "beach biotech" suggests a future where scientific research can be conducted remotely and with minimal resources.
Open Source and Collaboration
- DeepMind's decision to open source AlphaFold exemplifies a commitment to maximizing social and scientific impact.
- Open access allows for collaborative improvements in research, reflecting a culture of sharing knowledge and methodologies.
Conclusion
- The episode encapsulates a transformative moment in science, where AI is poised to significantly change the landscape of research and discovery.
- As technology evolves, the integration of AI in scientific disciplines promises to unlock new avenues for exploration and innovation.
Resources
- Follow Pushmeet Kohli on [Twitter](https://twitter.com/pushmeet)
- Follow Vijay Pande on [Twitter](https://twitter.com/vijaypande)
- Learn more about [Google DeepMind](https://deepmind.google)
- Read DeepMind's [AlphaFold whitepaper](https://deepmind.google/discover/blog/a-glimpse-of-the-next-generation-of-alphafold/)
- Access the [AlphaFold Database](https://deepmind.com/research/case-studies/alphafold)
Final Thoughts The integration of AI into scientific research marks a revolutionary era, with potential impacts on understanding life processes and developing new therapeutics. This episode highlights the excitement and challenges of harnessing AI to advance science effectively.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00AI is not sort of nice to have. It's basically almost a necessity for us to make sense and reason about any problem that we are now looking at. I think there's going to be this really fun cultural shift where 10 years ago people would say, I was ridiculous, I could try to do these things. I think 10 years from now people will be like, what's a ridiculous effort, human being to that? Like you can't like model these numbers in your head. essentially what we have entered is basically an age where a single human mind cannot comprehend the data that we are gathering about the units. One of these structures may have taken the length of a PhD right to solve a single structure and now we're talking about true scale.
0:41There have been 1 .6 or 7 million users of the Alphore database. Now if that is not a positive statement about the planet then I don't know what it is. There are 1 .7 million people interested in protein structure prediction. I'm really happy about that. The last few years have been peppered with AI announcements. Let's recap a few. April 2022, Dolly II is released. Mid -Journey and Staples diffusion fast -follow that summer. Then in November, Chattu B .T arrives. Then 2023 features the release of Claude, Lama, and Mistral 7P, just to name a few models. And we're only a quarter or so into 2024, and we're already seeing the expansion into AI music and video models faster than almost anyone could have imagined.
1:30And while much of the attention circles around creative tools, there was an AI unlock in biology that caught much attention in 2021. That was Alpha Fold 2. A breakthrough in prediction around the 3D models of protein structures was released and open source by the DeepMine team in July of that year. Since then, over 1 .7 million scientists across 190 countries have been leveraging the tool. In the meantime, the DeepMine team has been hard at work, seeing how else machine learning can expand the frontier of science across. Many areas of biology from structural biology to genomics, protein design, to selenginomics, to quantum chemistry, to meteorology, to fusion, to pio -matamatics, to computer science.
2:20They've released papers like high accuracy weather model, graphcast in November, alpha geometry in January, which approached the level of human, Olympiad, gold medalist, and other papers across materials, mathematical functions, and more, including, of course, continuing to push forward alpha -fold. And today, we have the pleasure of hearing directly from DeepMind's VP of Research focused on science. PushMeet, Coley. PushMeet sits down with myself, and A16Z General Partner Vijay Pande, who has long been part of this intersection himself, as a longtime professor at Stanford, spanning several departments from computer science, to structural biology, to biophysics, and was also the founder of the Folding at Home Project, released in the year 2000.
3:06Together, we reflect on the journey to AlphaFold. But more importantly, where are we in the trajectory of AI, meaningfully impacting the way we perform and unlock new science, from new lab economics to clinical trials to drug discovery and more. So the question becomes, can artificial intelligence help us uncover fundamentally new science? And as it already done that, let's find out. As a reminder, the content here is for informational purposes only. Should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16Z fund.
3:47Please note that A16Z and its affiliates may also maintain investments in the company's disgust in this podcast. For more details including a link to our investments, please see A16Z .com slash Disclosures.
4:04So AI has been the talk of the town. A lot of people are familiar with the consumer LLMs, they chat GBT, maybe mid -journey. But AI has been around for quite some time and it's also impacting the scientific sphere, which I think is so exciting and I think both of you do too. So push me, maybe we could just start there and talk a little bit about your background, how you kind of got into this intersection of science in AI. And also you work for DeepMind, which I feel like for one of the flagship AI companies, why have you chosen to focus more there than perhaps some of the others? Yeah, so I took a very roundabout journey into what I do today at DeepMind.
4:42I'm a computer scientist by background and was hired at Microsoft Research and worked there for a decade, mostly working on applied mathematics, solving difficult maths problems and most of them were encountered in machine learning. So I started with computer vision, computer graphics, information retrieval. And after having gone through many of these applications, was very excited about deep learning when it finally sort of emerged. I really thought that this was a game changer in terms of how machine learning is going to impact applications. Dennis Asabez, who is the CEO and founder of DeepMind, at that time DeepMind was a young starter and he reached out and said, well, we know you from some acquaintances, why don't you join us?
5:33And I said, no, I'm working on games at that time and I'm going to products and applications. And he said, well, the whole game thing is just phase one. The idea is to eventually impact science and impact applications which are the biggest challenges in the world. And the level of conviction with which he basically made his case, I was like convinced, this guy gets it. And so I moved to DeepMind in 2017 and I told him if you're very serious about real world applications, we need to make sure that machine learning systems are reliable. So in fact, when I joined DeepMind, I founded the reliability and safety sort of team at DeepMind.
6:16And around a year into it, there is sort of us be once between you're really invested into in multi -disciplinary research, where you want to apply machine learning in back full problems. And I think the most impactful area that you could work on is science. And that was a complete left field sort of suggestion the last class was in school. So I was quite skeptical to be honest and I told him like you've got the wrong guy like no background in biology or physics or chemistry but he said no I mean the way you are approaching these things it's good let's sort of give it a try and see where it goes and so we started the science program with six or seven people working on two projects and now it's almost 140 percent team and we have 10 different initiatives spanning many areas of biology from structural biology to genomics to protein design, to cell genomics, to quantum chemistry, to metrology, to fusion, to pio -mathematics, to computer science.
7:26So it's a long journey but it's started with sort of an accident. Yeah, and also a very scientific, iterative approach. I love that. Vijay, before we jump into more of those projects that push me kind of alluded to there, I'd love to hear your background and how you got into this intersection of science and AI because you also have quite the storied history there. Sure. Yeah. So from 1999 to 2015, I was a professor at Stanford and actually in a variety of departments, my home department was chemistry, but also had appointments in computer science, structural biology and was also chair of biophysics.
8:01And at that intersection, it was clear that machine learning was a very exciting tool to use. I think what really was happening early with genomics in the 90s and then just plowed all the way through was the rise of data in biology and biology becoming very quantitative. And once it starts becoming quantitative, machine learning is very natural. As Pushmeet talked about, I think, where a lot of us and myself included we've got particularly excited was maybe 2013, 2014, 2015, as deep learning was emerging. And I think machine learning before deep learning was human beings have to come up with their features and it was like a little tool with deep learning.
8:41It could be something that replaces more and more of the human part of the thinking. And actually a lot of the interesting results are emergence after that. And those emergent properties got very exciting. was clear at the time that we need a lot of compute. And so, actually early on in 2000, I founded the Fulwood and Home Distribute Computing Project. And actually, we were some of the first program GPUs. And so all of that comes together, the data, the compute. And then finally, the algorithms, once those three pieces were together, I think many of us could see that this was taking off and it was time to dive in.
9:12Absolutely. I think that brings us to this question of the why now. So you kind of already addressed it. But VJ, what gets you so excited about this intersection? and we're recording this in 2024. AI has really been around since maybe the 50s. Is it just that we have the right amount of compute? Is it that we have these unlocks when it comes to the modeling? Give us a little bit of a picture of what gets you so excited about what's to come before we dive into some of the specific examples. Yeah, if you step back, I think what we're really seeing in biology is this industrial revolution, that if you look at a biology lab, maybe even today to some extent, versus 10 years ago versus 50 years ago, will be benches and people in white coats and by petting and so on.
9:56And maybe the box little benches are a little different, but it's very, very similar. It's very bespoke and artisanal. What is shifting is that's becoming industrialized. We're seeing the rise of robotics and we're seeing with that industrialization this immense amount of data. And so AI needs data and data needs AI. And so as biology gets all that data, we can sort of lean into this. And what's mostly intriguing is that life sciences and healthcare largely has not been permeated by technology, not by IT to a great deal. And healthcare and life sciences collectively, it's almost like becoming 25 % of US GDP, these trillions and trillions of dollars going through this and none of it or very little of it being sort of revolutionized by tech.
10:39So this revolution I think is happening because of AI, AI is allowing this industrialization to happen and especially turning these spoke artisanal processes into something that is engineered and industrialized. There's one aspect of it and there's many others I talked about robotics and that's the arc that's I think exciting and it's something where I think we saw hints of it in 2015 it's probably a 25 year arc maybe 30 year arc that we're 10 years into and industrial revolutions don't happen overnight but when you look back the whole world's going to be changed And so we're living in the middle of it.
11:13And I was actually always jealous about people living in the 1920s and people going from nothing to steam trains and all the stuff. And actually now we're the ones that I think are in the center of it. It's such an exciting time. Right. You see that picture of, I think it's somewhere in New York where you have all these horses lined up, right? And back then, that just felt like the norm and then you see what, like a decade later, it's all replaced by the equivalent of cars. And so push me. Maybe we could use alpha -fold as an example here because a lot of people listening to the podcast are maybe most familiar with that paper and that breakthrough, but maybe also another great example of how that didn't happen overnight.
11:49I think most people noticed it in 2020, but it didn't start in 2020. And so maybe you could talk about that arc, what is Alpha Fold, how did it come to be, and then also where are we today in terms of its impact? Yeah, so, Ashifold, I was telling how I started my journey with the science program at DeepMine, And at that time, we had these two small scale sort of projects. One was protein structure prediction, and the other one was quadru chemistry. And other four sort of grows from that protein structure prediction project. In its simplest form, it's a very simple problem where given an amino acid sequence, which constitutes a protein, we want to understand the 3D coordinates of those amino acids.
12:29And that's pretty important because if you understand the 3D structure of the protein that informs and gives you an idea about what the function would be of that protein. And that has implications for Dr. Discovery for understanding basic cellular biology and so forth. So we started working on this problem because we thought it sort of satisfies one of our key requirements when we look into problems. That is it's real foundational root note problem. Once you solve it, it has so many different sort of implications and disease understanding in biology and so specific biology as well. And not only that, it is a classic sort of machine learning problem.
13:10You require reasoning in this problem because you are working with the very expanded solution space as well as you have access to raw material which is data. And the structural biology community had done an amazing job in sort of curating a very good dataset in the form of the PDP. So scientists all across the world had, whenever they found the structure of a protein, which sometimes took almost five years or even a decade in some cases, would diligently deposit that 3G structure in this database. And so at that time, when we started that, there were 150 ,000 odd structures, roots from extracurstrography and RIOEM, and that was like an amazing sort of dataset to start with.
13:59And not only that, the other big problem in machine learning as to how do you evaluate the machine learning model because in machine learning one of the easiest things that you can do is basically fool yourself. These models are extremely good at sort of cheat time and if you give them any sort of way to cheat they will cheat. So the protein folding community and the protein structure prediction community had this annual sort of competition called the CASP, the critical assessment for structure prediction and they would run this blind assessment, like an Olympics of protein structure prediction, where people would be given protein sequences whose structure was not known by anyone, only like one experimentalist who has deposited it, and then they would be tested.
14:45And the true generalization ability of the model would be exhibited. So we thought this problem really checked a number of key criteria which we use for taking up a problem for the very long term. So we started with a team which investigated how much progress we can make on this. We were hopeful of the mistake that machine learning can play a bottom role, but we didn't know. This was a new problem for us and we were approaching it with a lot of respect. And Pishmeet, what year was this when it started? So we started around 2017 and we took part in the critical assessment at the end of 2018. And when he entered Al -Al -Hafoal 1 in 2018, we were not really sure where would it be, maybe in the top three.
15:31But actually, performed really well. Not only was the state of the art, but performed the state of the art by a virgin. And that validated our hypothesis. The basic research philosophy here has been the multidisciplinary nature of the teams. So we had brought in some really good structural biologists and biopsychists, John Jumper, being the lead of Alpha4 was part of the team at that time. And that gave us a lot of confidence. Now, we were the best in the world, but the model was still not useful, right? It was reducing good results, but it was nowhere close to solving the problem. And then we had to sort of make a bet.
16:12Can we really go after it and solve it once and for all? What this is it? And so the first thing we had to do was start from scratch. We had to throw Alpha Fold one from the table and said this approach that we have started is not going to work. What gave you the indication that Alpha Fold one couldn't take you to the next level? Because I think even in the AI space outside of science, there are a lot of questions around, can we just depend on the scaling laws? Do we need some sort of new unlock to get to you, you know, insert problem here? Could be AGI, could be something else? What gave you the indication that this is great?
16:48We're so happy with our results, but we actually need to throw this out and start a new I forgot one had adopted a classical approach if this classical two -stage approach what the machine learning models job was it given A sequence it does not predict the 3d coordinates of the amino acid directly what it predicts is basically the distance between amino acids and then there's a second stage which was supposed to take that distance matrix and recover the 3D coordinates. So the machine learning neural networks job was restricted to find the distances between amino acid residues. And this two stage sort of model was very effective, but it was not very elegant in the sense that if you made certain errors, you will not be able to back propagate back to the neural networks.
17:41Because you found the results after the second stage and the new network would not get that supervision. So we believe that in order to be able to properly train the model, we needed end to end. We needed a model which could go directly from the sequence to the structure. Right. And that was one critical sort of element and a change that needed to be made, but it was a difficult change to make because you are starting from a much lower baseline when you are sort of building up that second end to end network. So let's fast forward. So you did throw out alpha fold one and then what happens after that?
18:18So alpha fold two, we've got this long journey where we start making progress on alpha fold two with a much lower sort of performance from alpha fold one even. We have this internal leaderboard where everyone in the team can can propose ideas and try out their ideas on the central need about to see how much of a delta each idea or each change sort of makes. And we were making studies sort of progress. And then there were times where progress would stagnate. And sometimes even for months, it would stagnate and people would ask the question, well, have you reached the limit? But over time, and I think around when the pandemic started, we got some really, really, really big dentures, where we thought we are making real progress.
19:05And if you look at the metrics as to how do you quantify protein structure prediction accuracy, it's called GTT. And we had crossed that ATGTT sort of threshold. And that was by unprecedented. And of course, that also motivated us to push it even further and later on to 90 GTT and beyond, right, which we thought is what we needed to do. And so the pandemic happened and it really sort of brought home to the whole team, the actual importance of the problem. Because we were all sort of sitting in our homes sort of shielding and they were scientists out there who said, And if you have the structure of the different SARS -CoV -2 proteins, it would be really helpful.
19:54Now, the community very quickly found the structure of the spike protein, because it was also very similar to SARS -CoV -1, but the necessarily proteins of the virus, though the structure for those was not known. And so the fact that we could compute these predictions, share it with experts who are trying to deal with the pandemic and think about and designing inhibitors and so on. It really brought to the team the real world impact and relevance that this fundamental problem has. And around September 2020, when the second cast competition ended, we got this email from the organizers who wanted to chat.
20:40And that was unprecedented. We were sort of surprised. like why did the organizers wanted to sort of chat so early on? And they were super surprised at how good the predictions were. In fact, some of them speculated maybe this team has cheated in some way. It could be so good. But apparently there was one particular sort of scientist who had submitted a protein, but did not know the structure. They had hoped that the structure would be obtained by the time the competition ended, but this structure was not known to anyone, literally anyone. And Alpha4 could give them an initial starting point which can solve the structure for all that particular protein.
21:19So they were totally amazed that such a system now existed in the gas competition and we later on sort of released Alpha4 and not only was it very accurate, it was also very efficient. So we decided to in fact find the structures for almost all the proteins that are known to scientists around one and 15 million of them and put them in a database with our partners, the European microbiology, the laboratory, the MBA and then made that as a resource that anyone can access. Yeah, that's amazing and I'd love to turn it to you, VJ. Yeah. You obviously have run a lab for a long time and you've been on the other side of this, right?
21:58All these researchers who now have access to this database, which by the way for the audience, one of these structures may have taken the length of a PhD to solve a single structure. And now we're talking about true scale. And also, again, this being deployed to all the researchers that can access it. So, Fiji, maybe you can just speak to what that really means. And also, if we can apply this to other areas of science as well. The impact of this is many fold. And I can speak to both from looking at it from the academic lens, but also from the last 10 years of investing and startups, and startups use this as well.
22:32First off, I think maybe it's worth really emphasizing the significant substructure itself. So the reason my universities like Stanford has whole departments for structural biology is that the structure is typically pretty evocative of function and other biological aspects. Perhaps the most notable example is the DNA structure. And that Watson and Crick came up with this structure. And by looking at just the structure, you can imply how DNA is replicated. and essentially how genetics works, you know, since some degree, it's a very basic of it. And so maybe that's one of the most sort of dramatic examples, but there's numerous examples where if you have the structure, you can understand a function.
23:09And so structural biology is a fun amount of part of how we understand biology from the molecular scale up. And also for drug design, often if we understand the structure and its dynamics, we can understand how the drug proteins come up with therapeutics much more sort of in an engineered fashion. So the significance of structural biology is huge. It's also a time where structural biology is an arsona because as you mentioned it used to take many years to come up with experimental structures but also new methods like cryoEM can come up with structures in much shorter time even days. And so there's an arsona going there and I think four structural biology is a field.
23:44I think we'll see this combination of new experimental methods and computational methods and I think what was most striking to me is how experimentalists were going to these databases and looking at them and using it almost like you would use the human genome database. That the human genome database takes genomics and turns it into a database lookup. That you can basically don't have to do the experiment yourself. You can just do the computational query. To some degree, I think what AlphaFull did is it took the structural biology of proteins and made it a database lookup. It's not exactly a true database lookup in the sense that this is a prediction, but as the quality of predictions gets higher and higher, it becomes kind of the same thing.
24:25So that's huge. I think the final thing that was, I think, most striking is that there's always going to be a shift from academia to industry. And maybe 30 years ago, academics would design computer chips and new types of microprocessors and so on, new architectures. We don't do that now in academia. I think that's not something that makes sense to do. That's much better done in companies, especially given the skill of what's going on. And I think what was most striking about this is that I think for multiple reasons this is something that deep mind was perfectly suited to do in a way that academic groups I think really weren't.
24:57And that shift now suggests that now I think it's a really interesting time for this to sort of leave academia and now be in the industrialized world of startups and companies. That's really interesting the relationship you're talking about of academia and industry. Something that people talk a lot about these days is whether these different AI models can really fundamentally advance science the way that you typically think of academics as the parties that are facilitating that. And so I'd love to hear from both of you, maybe starting with you, VJ. What indications, whether it's through alpha folds or other projects that you're seeing emerge, actually indicate that yes, these models, these scientific discoveries in a sense are able to help us actually push the frontier instead of actually maybe just help us be a little more efficient within the zone that we're already in.
25:43I think push me to say, well, that structure prediction is a foundational problem. But if you take, for instance, just the sort of arc of drug design, where first you have to come up with understanding the biology. The AI for biology is a very interesting area where we can maybe start to understand the nature of pathways and do this on human biology in ways that don't require experiments on human beings, which has always been one of the biggest limitations. I think we understand mouse biology really well because of all the experiments we can do, but we could never do that on human beings directly, but AI models for humans as they become more predictive, and especially just more predictive than a mouse's particular photo of human, the mouse is a model in a sense.
26:21That gets super interesting for unraveling biology. And so AI for biology is a thing. We can talk about AI for chemistry, and I think Alpha Fold is in that category where now we're trying to understand by a physical chemistry, you want to try to understand how can we quickly drug undrugable proteins? How can we come up with new antibodies and design proteins? That's a whole area. And then finally, I think AI for clinical trials is going to be really where maybe the biggest impact financially will be clinical trials could cost hundreds of millions to billions of dollars. Even a 10 % improvement on a billion dollar enterprise is huge.
26:54And that's where maybe some of the toughest problems to work on, but I think as we make impact there, I think clinical trials will be better, will be probably more easily powered, will be hopefully more successful because we'll be picking the right ones to do. And then that turns into eventually AI for personal events, which is an ascense extension of that trial. And so we're now, I don't want you to experiment on me as a mouse or a rat, but I would love to make sure I get the best drugs for me. And you and I are different and we'll respond different to drugs to be able to have that predicted.
Read the full transcript
27:26It would be huge. So I think there's the arc of that and I think we're just at the very beginning. Definitely. We talked about AlphaFold, which is very exciting and maybe the most familiar to folks, but Pishmeet, your team has also created a bunch of other papers that touch this intersection of AI and science, or you could say AI and math or AI and physics. And those are things like materials, graphcast, which has to do with weather forecasting, fun search, alpha geometry. And so I'd love to hear from you again on this probing of, are we moving the frontier forward with these different models?
27:58What are you seeing from some of these other projects that your team is working on in terms of AI helping us actually uncover new science? Essentially, what we have entered is basically an age where a single human mind cannot comprehend the data that we are gathering about the universe. And this is true in any field, you know, encounter. It is true in biology. No biologists can reason and analyze all the biological data that will be gathered. No physicists can look at and analyze all the high energy physics data that is being gathered. it, and even mathematicians cannot sort of look and analyze all the large scale mathematical simulation data that we can now compute and simulate and find out.
28:41And I think what's happened is AI is not sort of nice to have. It's basically of almost a necessity for us to make sense in reason about any problem that we are now looking at. I have examples in pure mathematics where work on topology, you describe a not in two different sort of There was a generic definition and there was a geometric definition. And mathematicians understood these characters, but never understood the connections between that. And what we showed in one of our sort of works is basically we generated a lot of data for nodes in these two characters. And somehow, Yasen even let's what can you make predictions about one character from the other?
29:22And the idea was well down there should be no. But in fact, it couldn't make predictions. And when we drill down, we found a very nice conjecture that nobody had encountered. And we work with mathematicians who then not only make that conjecture, but actually prove that there was a very elegant, nice relationship between those two characterizations. So this is like completely fundamental discoveries in mathematics that were completely unknown to mathematicians now being uncovered by a machine learning and AI model. And we are seeing this across the board in any of the scientific areas that we are looking at.
30:02We are discovering new insights, new sort of patterns that were not expected, just because the techniques to analyze the raw scale of data did not exist. I think amongst biologists, especially maybe 10 years ago and Frinter back, I think there was often a belief that biology is just so complex that it's just incomprehensible that there's no way to even understand it. The only thing you can do is run the experiment that's going to happen. And I think we're seeing the beginning of a shift where people are starting to think, well, there are complexities and there's a lot we don't know a lot to learn, but that AI actually can gather all that together and start to decipher this and to be a natural language for biology.
30:44And I think there's going to be this really fun cultural shift where 10 years ago, people would say, I was ridiculous, a computer could try to do these things. I think 10 years from now people will be like, what's a ridiculous effort, a human being to that? Like you can't, like load all these numbers in your head. That's just ridiculous to even say that. And we've seen this in other places, like chess. It seemed like impossible that a computer could be the grandmaster. And then now there's not even worth trying. It's table stakes. Yeah, yeah. And we saw it with Go. We saw it with all these other things.
31:12So I think that's just a cultural shift. But I don't think that's a bad thing. I mean, Fort Lift can lift much more than the strongest weightlifter and we view that as a positive thing. It's always going to be us and them. I think the interesting question will be is once you can do these things that we can't do, well, what do we do together with that? Yeah, and what can we do? I mean, one of the most amazing things, I think, is that deep -mind for the most part has given these models or the results of them to the community and so researchers have their hands on them. And so maybe we could talk about that, how are researchers leveraging these new breakthroughs?
31:47There's all kinds of stats around we don't have enough cancer drugs or they're insortages and those are very real things we want to fix. So Pishmeet maybe we'll start with you. What are you seeing and your team seeing in terms of this technology being deployed and how are researchers using it? Yeah, so this was another sort of fascinating journey of growth. I was not not from the natural sciences. So working on Alphabot was a learning experience, but then actually releasing Alphabot to the community was even the bigger sort of learning experience. So Alphabot database, when we were sort of building it up, we wanted it to be available everywhere in the planet to all the sort of scientists.
32:25But the scale of science was unprecedented. I was not aware of it. The Alphabot database today has been accessed in 190 countries. And there have been 1 .6 or 7 million users of the Alphold database. Now, if that is not a positive statement about the planet, then I don't know what it is. There are 1 .7 million people interested in 14 structure prediction. I'm really happy about that. I mean, all the things that are happening in the world. And in terms of the impact, it's again, like in the amazing sort of spectrum, We saw Alpha Fold being used in path -breaking, fundamental, biological discoveries.
33:07Like my personal favorite in that domain is the nuclear pore complex, the structure of basically the pore complex, like the way a nucleus controls how a material gets into the nucleus and out. I mean, that fundamental structure of that complex was not known. And the searches used Alpha Fold to structures to be able to piece together the whole complex. A recent paper from the Feng Lab showed how you could develop a molecular syringe. And again, they used Alpha42 in designing that. And there's so many other sort of areas where people have been using it for developing new vaccines and working on new antibiotics against antimicrobial resistance and synthetic biology.
33:47Like one of the key partners at the early stages was a university here in the UK, which was using sort of alpha -4 to develop and think about enzymes that could decompose plastics. So you have this whole spectrum of fundamental biology drug discovery to even synthetic biology and enzyme development that has been impacted by alpha -4. And so it was very difficult to even predict what would be the uses of the two. I think there's also just within biology, there's become a shift that I think people are sort of wrapping their heads around prediction a bit better. I think before experiment was the gold standard and that was all people wanted to hear about.
34:30I mean part is also just the zeitgeist of the time when you deal with large language models, you're basically dealing with predictions of what comes. And I think people have understood the pros and cons of predictions, but that there's massive value in having it. And I think it's It's funny that we would talk so much about the technology, but I think it's the human shifts, so the cultural shifts are the things that we're going to really need to push. And I think what gets me most excited about what Prishmits has just been talking about is the fact that I think that's the sign that we're seeing this cultural shift as well.
34:59Maybe something else you could speak to VJ that's just coming to mind as both of you are sharing more about these researchers. How does this change the economics of a lab, right? If you think about what we talked about before, is like uncovering a structure it could have taken a whole PhD. Now we have new tools and we're seeing these economics change in some of the more consumer fields and those are very obvious. How does this change the economics of research overall? One of the sort of fantasies that one of my former colleagues talked about was what we call beach biotech, where you have let's say one person at a laptop, but presumably on the beach, wherever you want to be.
35:35And you've got CROs, these contract research organizations to do the experiments, you have some AWS cloud or whatever, some GCP clouds somewhere to run your calculations, and that one person with AI, I think we're not quite there yet, but I think that's an intriguing fantasy to think about. And I think on the way to the one person's sort of aspiration is smaller teams doing way more with much less capital outweighs and building startups, I think much more efficiently, and where they get to results much more rapidly. That challenge is going to be, what I mentioned is four, is that getting to the clinical trials, speeding that up will be nice, but I think the big financial return will be on the clinical trial side.
36:17But I think the expectation is that AI for biology and understanding targets and so on, based on human data, that would also help on the trial side and in addition to anything else there. So I think put together, I think we can get to these therapeutics faster, cheaper, and hopefully better. Yeah, and maybe push me, we could tackle that directly. If you could give a sense for folks who aren't these researchers who aren't already leveraging these tools, how much does it really cost if someone does want to get a protein structure prediction or use some of the other models that we've talked about again, graphcast or materials, etc.
36:52Like what cost are we really looking at? Yeah, so for the ultimate goal database, it's literally free. We just go to the full database, sort of, to find the protein that you're interested today, now from the 250 million sort of proteins and make it and it's there. It's for free for everyone on the planet to use. So really it has democratized things in a way that scientists in Latin America or India who was working on sort of neglected tropical diseases, for instance, who had no way they could get a structural of a protein that they were interested in can now get access to these structures at the sort of click of a button.
37:33Of course, a lot of research needs to be done to take that work and towards a more focused outcome and a lot more investment is needed if you are trying to finish and accomplish the vision that Vijay has had outlined. The other four structures are start but you really need to think about how does it bind to the how do you do the liquid design, how do you solve the co -holding problem. So there is a lot of investment that is needed to make these models and make these predictions and refine them for specific applications. And we have a spin -off from DeepMind Isomorphic Labs, which is now investing in this area as well.
38:14At the same time, we are continuing and work on the foundational sort of side of things and have now released an announcement or an update on the next generation of alpha -fold, which goes beyond proteins to other biomolecules to nuclear acid, like DNA, RNA, BDM, small ligands, and so on. I think it's amazing that you've opened this up to the community. And I think something I'd love to hear both of your takes on is really the relationship of these models and them being open sourced. I mean, it's a big debate with an AI at large, but I think especially when it comes to science, there's, I think, both ends of the spectrum in a way, right?
38:56I think there's nothing more that people get excited about about this idea of curing cancer, like solving poverty and agriculture crisis, but at the same time, people also get very scared, right? I think that's where people's sci -fi nightmares come to be, right? Where they're like, oh, someone can engineer a molecule that can kill us all. And I guess starting with you, Vijay, what's your take on this relationship of AI and science and why it should be open source? I think the beauty of open source, and we see this open source for AI and biology, but AI more broadly, is that people can build on top of each other.
39:33And I think what's really remarkable about the AI field, I would say, over the last five, maybe possibly 10 years, is that it feels like an amazing result comes out once a week. And that the key part of that is that it comes out with code or GitHub repo. And that you can check out immediately. You don't even have to just believe the results you can run it yourself. people have even open sourced for his tests of things. So essentially we're building like a skyscraper where each person builds a new floor and we're going up really fast. And that's what open source can do. In the past, if it wasn't open source, I'd have to read the paper, I'd have to code it myself.
40:07And sometimes the paper may be a little vague for some detail. So I might not bother, right? And I'll just go do my thing. And so I think what open source allows us to do is to build on top of each other and build rapidly. Now, certain parts won't be open source. I think you unfortunately can't open source a drug compound because then no one's going to pay for the trial and certain things like that. Just the economics doesn't make sense given these hundreds of millions of billions of dollars and so on. So, certain parts will be closed source and there's hundreds of startups in AI in biology and AI drug design that will maybe take advantage of what's been done, develop their own methods and build them top and then that's where I think the drugs will come from.
40:46You talked about also the concern for how because this is so powerful we could maybe do a certain dangerous things with it and that's where everything there's a bit of a misconception because actually there's a huge asymmetry between the complexity of drug design for treating disease and that's a really hard problem to do but it actually turns out to be really easy to come out with chemical that actually are dangerous and toxic. In fact that's why we have phase one trials because like even the things that you thought would really hopefully not be toxic at all, turns out to be toxic. So it's actually very easy to make toxic things.
41:21And Google will teach you actually how to get rice in and how to get all this other stuff for better or worse. So I think there, the asymmetry is that if we get rid of AI for drug design, you lose all the good and you don't prevent any of the bad, which is already here. I think that's a good point that a lot of people don't think about. Fish meat, maybe you could just speak to why DeepMind has chosen to open source these models, which isn't necessarily the norm across different AI companies. There was a lot of deliberation within the team and within the company on this. I think there were a few different things that went into the final decision.
41:57One was, we wanted to, like, out of all wars, that foundational there. It was so foundational. It would, if we had kept it close source, the impact of it, like, fully leveraging them back for society, I mean that would have been difficult. It was because it's so fundamentally sort of foundational and it's very hard to even predict what are the potential sort of applications of it. Just to give you an example, when we launch Alpha Fold, couple of days later, somebody did an analysis on the uncertainty associated with the Alpha Fold predictions and figured out that in fact Alpha Fold was even though it was not trained for that was the best predictor for predicting disorder in proteins.
42:42So that was something that we would not have come up with, right? If we had kept it close so as someone barely interacting with the model in the community figured that out. So when we were thinking about it, there was of course how to maximize social impact and scientific impact of the model. The second world was responsibility. And we consulted a number of experts from structural biology, from chemistry, from doctors' recovery to figure out what is the right and responsible and safe approach here. And even considering the malicious sort of use cases, and after we had done all the due diligence that we felt that this was safe to release, and the impact of releasing it and open sourcing it in a wider sort of way would outweigh any costs that we would need to sort of model.
43:34It was decided that we should open sourcing it. And I think the decision has been validated by the impact that our whole do as had in the community. Now of course that's not true for all the different models. In fact, subsequently we have had models which we have not open sourced. But I think in the case of alcohol too, the decision was very, very clear. in favor of sharing it with the world in the most free way possible. For the ones that you haven't chosen to open source, if you're willing to share, how do you make that decision? There are a number of different factors. Both what will be the social impact, the scientific impact of releasing things versus what is the commercial cost of releasing something while leveraging it for commercial purposes or even the safety sort of argument.
44:18So just to give you the example, one of our recent models that we announced last year was alpha missense. And this is a model for predicting effect of missense variance. And what the model does, it produces state of the art accuracy in making predictions about whether missense variance are denied or path, what could be pathogenic. And in this particular case, we found that the predictions of the model for the human genome, for the human distance variant, like the 71 million of them, if we release that, that would serve most of the purposes, that a clinician or a biologist would be interested in.
45:01So we just release the predictions rather than the model, because the model had many other sort of users, you could run it on different organisms, there were other sort of commercial considerations. So it was felt that we could release the predictions, we could share the methodology but we will not sort of open source that approach. That makes sense. And I think at the very outset, you shared so many different projects or areas of scientific study that your team is working on. I'm just so curious because it sounds like there's been success across many. Are there any areas of science or mathematics that you've tried to address with this approach of using machine learning and AI that's not quite working, whether it be because we don't have the prior data set as VJ has spoken to that sets the foundation.
45:46I'm just so curious if there's limitations emerging in any of these fields that your team is running into. One specific area that I would love to have impact on, right? And I think yeah, I would eventually have impact on a system's biology. It's an incredibly important sort of problem to really understand at the system level how biological systems behave, it's just the data and the evaluation is not at a place where it is for maybe genomics, functional genomics, or for structural biology. Before we actually start an initiative in any of these areas, there is a huge due diligence process that we need to undergo because essentially you're making a very long term commitment and the careers and the impact of some of the best scientists and engineers that we have are being committed to that area.
46:40So we take that responsibility very seriously. And only when the impact, when we are confident of the impact of the problem, we are confident that we have a good evaluation metric to track progress and we have the raw material. the data or an assimilator to get good data, only then do we make that long term commitment towards a specific topic. The highlight the data issue, I think one of the biggest differences between AI for, let's say, language models or AI for video and AI for biology or for healthcare is that I think most of the interesting data in biogen healthcare is either dark, that there's all these medical records and so on that you just can't access on the internet, which would be very useful for understanding the healthcare side, for trial side and so on.
47:30It's either dark or it's never been measured. And that, oh, we need to do the experiments. I think having the data could be paramount. And that I think that's gonna be different than other places. Other places maybe the algorithms can really drive things because everyone has the same data more or less. I think here people will be differentiated by their data. And so the innovations will be innovations in AI I combine with innovations in data collection. And there are obviously things I've been in a interface for active learning and how can you use the data more efficiently and so on. But the data game I think is going to be huge.
48:02Absolutely. And Vijay, I'd love to just get your take. You've spoken to a few examples already, but what different areas do you wish that more attention was being allocated? Or do you just think there's a set of grand challenges that can and will eventually be solved with some of this technology? The fun thing about CASP, this critical assessment of structure prediction is that I think it also inspired all these other prospective trials and prospective studies. So there's a ton of that stuff to do and I think there's tests for predicting binding of small molecules. I think we'll see in time these types of methods do extremely well in those assessments.
48:39But the Holy Grail is in my mind being able to predict clinical trials is something where you know to understand how a drug works in human biology And that's a push me's point is that's a systems biology problem at the largest skill. And so that is the holy grail and I think we'll probably do it in parts. You could imagine even like models for specific organs or models for specific parts of the body and then we put them together. Make sure as of experts is pretty common these days and maybe that would be one approach. But however I guess done once that gets done to the point where these models are better than the animal models, I think that's where there's really going to be a tipping point and up point where we can just move much more rapidly, where we can sort of not get stymied with having to run these animal models which takes a long time, it's very expensive.
49:25And even there's crazy things like right now, there's a monkey shortage because monkeys are in such type of man to run these experiments. So I think there's all along roads to get there where these models of humans are more predictive than the alternatives, but I think once we get there, that will be a major intersection point. Wow, I did not know there was a monkey shortage, But I mean, it really is important to know, right? As a to your point, hopefully we get to a future where some of the things that we're doing in research today seem just so incredibly outdated because we just have better options.
49:56Push me what's next up for deep mind in terms of areas of interest. I mean, you're already working on so many things, but would love to just get a pulse on what's exciting for you too. I think what is fascinating about science and like in any of these fields is that there's so much more to work on. I mean, even the on -structural prediction, I just mentioned that the latest version of the Earth and Paul, the work there is on extending it to general biomolecules, like DNA, understanding RNA, understanding the interactions between small molecules, begins and proteins, like bigger complexes, antibodies.
50:33There's so many things that we can extend in genomics. We have worked on both gene expression, the coding part of the genome, like with the Bissens variance, and the non -coding part of the genome, right? Or like, critically gene expression, we have made progress, but we are not completely at the end of it, right? So there's a lot that we are doing in all these areas, in material science. You mentioned this model known, which was able to predict 400 ,000 normal stable compounds, which expands the number of stable compounds known by more than order of magnitude, right? But how do you now take those sort of compounds and then reason about their specific properties that would be useful in a particular application?
51:18So in any of these disciplines, we are not targeting one specific milestone. You're just saying, here is a topic and the long term sort of road map is to think about a paradigm shift in how science is done in that area and moved towards a more rational modeling based approach and tackling some of the problems that are encountered here. So there's a lot that needs to be done. And we are just trying to focus on some specific areas and then new areas come up if the raw materials are there in terms of data. And if you have to, on the valuation metric, we are constantly reviewing them as well. That's amazing.
51:56I haven't done as much research as UVJ, but I did do a summer of battery research and materials research where we were trying to discover new sodium ion transition metal materials and my summer was literally, I mean, this is when I was in college, so I wasn't very advanced, but it was literally like finding a paper that documented how to synthesize this material and the kiln, mixing it up, creating a little battery, doing it in the glove box and running it and just seeing how effective it was and obviously in cases, it was very ineffective, but every so often we found a material. It was truly just trial and error, trial and error, trial and error.
52:33And when I see papers like this that do things in a completely new way at scale, way cheaper, you don't have all of these university students just in a glove box day and night, it's so exciting. The end point for me is like, as we talked about, we're kind of in the middle of this journey and this technological journey, this cultural journey, in these cultural shifts and that it's going to feel like the big goals that I've laid out, let's say clinical trials, things and systems biology, that's so far off, right? And it's going to take a while. But we can get a lot done in 10 years collectively, 15 years, you think about where we were five years ago, 10 years ago, 15 years ago, now 15 years ago, people weren't really even talking that much about deep learning or just beginning.
53:14So the goals that we have are lofty, but I think we're right in the think of it and all that I think is very doable is just now building that tower once every time. It'll be fun to have this chat again in five years. Hopefully sooner. I think one sort of thing that having been exciting in the last few years is the rise. Of course, there's a lot of excitement about LLMs and foundational models and so forth. If you look at the impact that's going to have on science. Now, in most of the projects that I was talking to you about, you're working with structured data, data, either which was collected or in the case of some of our fusion work data that was simulated.
53:55But with the rise of foundation models and elements, that opens up the possibility of now using unstructured data to feed these models. And so that really opens the goal for a large scale ingestion of scientific knowledge into the models. And that is a very exciting direction that will I think bring a number of other problems now in the feasibility zone, which previously were not there. Of course, there are challenges with understanding uncertainty and sort of hallucination and all these sort of technical problems need to be sort of addressed, but once that is done, I think the impact that's going to have on models for scientific discovery would be amazing.
54:41So that's another reason to be excited for the future. Absolutely. And all of the problems you just mentioned are also opportunities for people to go and fix and be a part of that whole ecosystem. So this has been really wonderful push me VJ. Thank you for as you said, getting people excited about what's to come because I think these two fields intersecting what a time to be alive here in 2024 to kind of be a part of it like you said, VJ we're in our equivalent 1920s. So hopefully people in the 21 20s will look back at this fondly. Absolutely. Yeah. If you liked this episode, if you made it this far, help us grow the show.
55:19Share with a friend or if you're feeling really ambitious, you can leave us a review at ratethispodcast .com slash basic ccd. You know, candidly producing a podcast can sometimes feel like you're just talking into a void. And so if you did like this episode, if you liked any of our episodes, please let us know. We'll see you next time.
From the publisher
In recent years, the AI landscape has seen huge advancements, from the release of Dall-E 2 in April 2022 to the emergence of AI music and video models in early 2024.
While creative tools often steal the spotlight, AlphaFold 2 marked a groundbreaking AI breakthrough in biology in 2021. Since its release, this pioneering tool for predicting protein structures has been utilized by over 1.7 million scientists worldwide, influencing fields ranging from genomics to computational chemistry.
In this episode, DeepMind's VP of Research for Science, Pushmeet Kohli, and a16z General Partner Vijay Pande discuss the transformative potential of AI in scientific exploration. Can AI lead to fundamentally new discoveries in science? Let's find out.
Resources:
Find Pushmeet on Twitter: https://twitter.com/pushmeet
Find Vijay on Twitter: https://twitter.com/vijaypande
Learn more about Google DeepMind: https://deepmind.google
Read DeepMind’s AlphaFold whitepaper: https://deepmind.google/discover/blog/a-glimpse-of-the-next-generation-of-alphafold
Read DeepMind’s AlphaGeometry: https://deepmind.google/discover/blog/alphageometry-an-olympiad-level-ai-system-for-geometry
Read DeepMind’s research on new materials: https://deepmind.google/discover/blog/millions-of-new-materials-discovered-with-deep-learning/
Read DeepMind’s paper on FunSearch, focused on new discoveries in mathematics: https://deepmind.google/discover/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-models
Stay Updated:
Find a16z on Twitter: https://twitter.com/a16z
Find a16z on LinkedIn: https://www.linkedin.com/company/a16z
Subscribe on your favorite podcast app: https://a16z.simplecast.com/
Follow our host: https://twitter.com/stephsmithio
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Stay Updated:
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Podcast on Spotify
Listen to the a16z Podcast on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

