In short
Y Combinator Startup Podcast Notes
Episode
John Jumper: AlphaFold and the Future of Science
Overview In this podcast episode, John Jumper discusses his journey as a physicist-turned-computational biologist and his leadership of DeepMind’s AlphaFold team, which earned him the 2024 Nobel Prize in Chemistry for solving the protein folding problem—a long-standing challenge in biology. He elaborates on the breakthroughs in deep learning that led to the development of AlphaFold, how it works, its impact on the scientific community, and the future of AI in scientific research.
Key Themes
- Personal Journey and Background
- Jumper began as a physicist, aspiring to make significant contributions to fundamental physics.
- Transitioned to computational biology after finding his passion in applying coding and mathematical manipulation to biological problems.
- His experience working at a computational biology company fueled his interest in using AI for scientific advancements.
- The Problem of Protein Folding
- Proteins are complex structures crucial for biological functions, and understanding their folding is essential for drug development and medical research.
- Traditional methods of determining protein structures are time-consuming and require significant resources, often resulting in failures.
- AlphaFold Development
- AlphaFold 1 and 2: Key advancements in predicting protein structures with atomic accuracy.
- AlphaFold 1 showcased improvements over prior systems.
- AlphaFold 2 achieved even higher accuracy, demonstrated through the ability to predict structures effectively using significantly less data.
- The breakthrough during CASP14 (Critical Assessment of protein Structure Prediction) highlighted AlphaFold's capabilities against other methods.
- The Significance of Machine Learning Research
- Jumper emphasizes the role of research and innovative ideas in AI development:
- It’s not just about data and compute power; groundbreaking research leads to transformative outcomes.
- Ideas are a core component of machine learning advancements, amplifying the effects of data and compute.
- Impact on the Scientific Community
- AlphaFold has been cited over 35,000 times, enabling thousands of researchers to make scientific discoveries across various disciplines.
- Jumper notes that many innovations stem from the collaborative nature of science facilitated by AlphaFold's predictions.
- Accessibility of Data
- The decision to open source AlphaFold's code and to create a database with millions of protein structure predictions made it widely accessible.
- The transition from providing code to offering a comprehensive database significantly increased trust and usage among biologists.
- Future of AI in Science
- Jumper reflects on the potential for AI to continue transforming scientific processes:
- The idea of foundational models in AI can lead to broader applications in various scientific challenges.
- Emphasizes the importance of understanding and leveraging AI's potential to serve experimental biology better.
Key Takeaways
- Pivotal Role of Ideas: Innovations in AI for science hinge on mid-scale research ideas that enhance the efficacy of data and compute.
- Real-world Applications: AlphaFold serves as an amplifier for experimental biology, speeding up protein structure predictions and aiding in drug discovery.
- Community Impact: Making scientific tools open-source fosters collaboration and accelerates discovery in the scientific community.
- Broader Implications for AI: Jumper encourages a view of AI as a powerful tool for transforming scientific inquiry, with implications that extend beyond narrow applications.
Conclusion John Jumper's insights underscore the transformative potential of AI in scientific research, particularly in fields like structural biology. The development and accessibility of AlphaFold exemplify how technology can revolutionize traditional scientific methods, leading to a faster pace of discovery and innovation.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00This is something of a nice change. I've given a lot of scientific talks and no one claps and cheers when I come on. Not normally even when I come on.
0:13It's really exciting. It's really wonderful to be here. I guess I should start off assuming that not everyone in this cavernous hall knows who I am. Who am I? I'm someone who has done some work in AI for science, who really believes that we can use the AI systems, these technologies, these ideas, to change the world in a very specific way, to make science go faster, to enable new discoveries. I think it's really, really wonderful. We have the opportunity to take these tools, these ideas, and aim them toward the question of how can we build the right AI systems so that sick people can become healthy and go home from the hospital.
0:55And it's been kind of a really wonderful and winding journey for me to end up here. I was originally trained as a physicist. I thought I was going to be a laws of the universe physicist. If I was very, very lucky, I could do something that would end up one sentence in a textbook. And I did physics and I went to actually do a PhD in physics. And then kind of what I was working on didn't really grab me. I just, it didn't feel like what I wanted to do. So I dropped out. I didn't start a startup. That would have been very on point for this event. But I dropped out and I ended up working at a company that was doing computational biology.
1:36How do we get computers to say something smart about biology? And I loved it. I loved it not just because it was fun, but it was something that would let me do what I thought I was good at. write code, manipulate equations, think hard thoughts about the nature of the world, and use it toward this very applied purpose that at the end, we want to make medicines, we want to enable others to make medicines. And I really kind of became a biologist and a machine learner, actually a machine learner, because I left that job and I went back to grad school in biophysics and chemistry. And I no longer had access to this incredible computer hardware that I had when I was working at my previous job.
2:18And in fact, they had custom ASICs for simulating how proteins, this part of your body that I'll talk about, move. And since I didn't have that anymore, but I still wanted to work on the same problems, well, I didn't want to just do the same thing with less compute. And so I started to learn, and I was getting very interested in statistics, in machine learning. We didn't call it AI back then. In fact, we didn't even call it machine learning. That was a bit disreputable. I said, I'm working in statistical physics. But how are we going to develop algorithms? How are we going to learn from data and do that instead of very large compute?
2:52And I guess it turns out in terms of AI, in addition to very large compute, to answer new problems. And after this, I joined Google DeepMind and really joining a company that wanted to say, how are we going to take these powerful technologies and all kind of these ideas, and they were becoming very, very readily apparent, how powerful these technologies were with applications to especially games, but also to things like data centers and others. How are we going to take these technologies and use them to advance science and really push forward scientific frontier? And how can we do this in an industrial setting with an incredibly fast pace, working with some really smart people, working with great computer resources.
3:38And with all that, you darn well better make some progress. And it's been really, really fun. And the fact that I'm on this stage indicates that we made some progress. And I think really the guiding principle for me has that when we do this work, that ultimately we are building tools that will enable scientists to make discoveries. And what I think is really heartening about the work we've done and the part that really I think still just resonates with me at my core is there are about I think 35 ,000 citations of AlphaFold. But within that is there are tens of thousands of examples of people using our tools to do science that I couldn't do on my own, but are using it to to make discoveries, be it vaccines, be it drug development, be it how the body works.
4:29And I think that's really, really exciting. And the part I wanna talk to you about today and the story I wanna tell you is a bit about the problem, a bit about how we did it, and I think especially the role of research and machine learning research and the fact that it isn't just off the shelf machine learning. And then I wanna tell you a little bit about what happens when you make something great and how people use it and what it does for the world. So I'll start with the world's shortest biology lesson. The cell is complex. For people who have only studied biology in high school or in college, you might have this idea that the cell is a couple parts that have labels attached to them.
5:11And it's kind of simple, but really it looks much more like what you see on the screen. It's dense, it's complex. In terms of crowding, it's like the swimming pool on the 4th of July. and it's full of enormous complexity. Humans have about 20 ,000 different types of proteins. Those are some of the blobs you see on the screen. They come together to do practically every function in your cell. You can see that kind of green tail is the psyllium of an E. coli. That's how it moves around and you can see in fact how it moves around. You can see that thing that looks like it turns and in fact it turns and drives this motor.
5:51All of this is made of proteins. When people say that DNA is the instruction manual for life, well this is what it's telling you how to do. It's telling you how to build these tiny machines and biology has evolved an incredible mechanism to build the machines it needs, literal nanomachines, and build them out of atoms. And so your DNA gives you instructions that say build a protein. Now you might say your DNA is a line and so are proteins in a certain sense. It's instructions on how to attach one bead after another where each bead is a specific kind of molecular arrangement of atoms. And you should wonder if my DNA is a line and I am very much not one-dimensional, what happens in between?
6:35And the answer is after you make this protein and assemble it one piece at a time, it will fold up spontaneously continuously into a shape. Like you've opened your IKEA bookshelf and instead of having to do the hard work, it simply builds itself. And you get this quite complex structure. You can see quite typical protein, a kinase for those of you who are biologists in the audience, over there. And you can see this very complex arrangement of atoms. And that arrangement is functional. And the majority, not every one of the proteins in your body undergo this transformation and that is what functions and that is incredibly small.
7:15So light itself is a few hundred nanometers in size and that's a few nanometers in size so it's smaller than you can see in a microscope. And for a long time scientists have wanted to understand the structure because they use it to predict how changes in that protein might affect disease. How does that work? How does biology work? Often if you make a drug, it is to interrupt the function of a certain protein like this one. Now, scientists have, through an incredible amount of cleverness, figured out the structure of lots of proteins, and it remains to this day exceptionally difficult. Right? You shouldn't imagine this as, I want to determine the structure of a protein, so I shall open the lab protocol for protein structure determination.
8:04I shall follow the steps. It consists of cleverness, of ideas, of finding many ways. In this case I'm describing one type of protein structure prediction in our protein structure sorry, determination, experimental measurement, where you convince that big ugly molecule I just showed you to form a regular crystal, kind of like table salt. No one has an easy recipe for this, so they try many things, they have ideas, and it's exceptionally difficult and filled with failure, like many things in science. And you're really looking at kind of one way to get an idea of how difficult this is, just one kind of ordinary paper that we were using.
8:45I flipped to the back and it said, you know, in their protocol, after more than a year crystals began to form. Right? So not only did they do all these hard experiments, but they had to wait about a year to find out if it worked. And probably that year wasn't spent waiting. It was trying a thousand other things that didn't work as well. Once you do that, you can take this to a synchrotron, a modest thing, you can see the cars ringing the outside of this instrument, so that you can shine incredibly bright x-rays on it and get what is called a diffraction pattern. And you can solve that, and you can deposit it in what's called the PDB or the protein databank.
9:24And one of the things that enabled the work we did is that scientists 50 years ago had the foresight to say, these are important, these are hard, we should collect them all in one place. So there's a data set that represents essentially all the academic output of protein structures in the community and available to everyone. So our work was on very public data. About 200 ,000 protein structures are known. They pretty regularly increase at about 12 ,000 a year. But this is much, much smaller than the need. Getting the kind of input information, the DNA, that tells you about a protein is much, much, much, much, much easier.
10:08So billions of protein sequences are being discovered. About 3 ,000 times faster are we learning about protein sequence than protein structure. Okay, that's all scientific content, but I should talk to you about the little thing we did, which has this kind of schematic diagram. We wanted to build an AI system. In fact, we didn't even care if it was an AI system. That's one of the nice things about working in AI for science is you don't care how you solve it. If it ended up being a computer program, if it ended up being anything else, we want to find some way to get from the left where each of those letters represents a specific building block of the protein, considered in order.
10:50We want to put something in the middle, in the alpha fold, and we want to end up with something on the right. And you'll see two structures there if you look closely, where the blue is our prediction, and the green is the experimental structure that took someone a year or two of effort. If you want to put an economic value on it, on the order of$100 ,000. And you can see we were able to do this, and I want to tell you how. And there were really three components to doing this or to do any machine learning problem. And you can say you have data and you have compute and you have research. And I feel like we tell too many stories about the first two and not enough about the third.
11:35In data, we had 200 ,000 protein structures. Everyone has the same data. In terms of compute, this is an LLM scale. It's the final model itself was 128 TPU v3 cores, roughly equivalent to a GPU per core, for two weeks. This is again within the scope of, say, academic resources. But it's worth saying really, most of your compute, when you think about how much compute you need, don't get distracted by the number for the final model. The real cost of compute is the cost of ideas that didn't work. All the things you had to do to get there. And then finally, research. And I would say, this is all but about two people that worked on this.
12:20It's a small group of people that end up doing this. So really, when you look at these machine learning breakthroughs, there are probably fewer people than you imagine. And really, this is where our work was differentiated. We came up with a new set of ideas on how do we bring machine learning to this problem. And I can say earlier systems, largely based on convolutional neural networks, did okay. They certainly made progress. If you replace that with a transformer, you're honestly about the same. If you take the ideas of a transformer and much experimentation and many more ideas, then that's when you start to get real change.
13:00And in almost all the AI systems you can see today, a tremendous amount of research and ideas and what I would call mid-scale ideas are involved. It isn't just about the headlines where people will say, transformers, you know, scaling, test time inference. These are all important, but they're one of many ingredients in a really powerful system. And in fact, we can measure how much our research was worth. So someone, AlphaFold 2 is the system that is quite famous, the one that was quite a large improvement. AlphaFold 1 was the best in the world. but someone did the Al Qureshi lab did a very careful experiment where they took AlphaFold2, the architecture, and they trained it on 1 % of the available data.
13:48And they could show that AlphaFold2 trained on 1 % of the data was as accurate or more accurate as AlphaFold1, which was the state-of-the-art system previously. So there's a very clean thing that says that the The third of these ingredients, research, was worth a hundredfold of the first of these ingredients' data. And I think this is generally really, really important. That one of the big, as you're all thinking, as you're all in startups or thinking about startups, think about the amount to which ideas, research, discoveries, amplify data, amplify compute. They work together with it. We wouldn't want to use less data than we have.
14:30We wouldn't want to use less compute than we have available. But ideas are a core component when you're doing machine learning research, and they really help to transform the world. And we can even go back and we can do ablations, and we can say what parts matter, and don't focus too much on the details. We pulled this from our paper. You can see here this is the difference compared to the baseline. You take either of those. And you can see that each of the ideas that you might remove from our final system, kind of discrete identifiable ideas, some of which were incredibly popular research areas within the field.
15:06Like this work came out and a part of it was equivariant. And people said, equivariance! That is the answer! AlphaFold is an equivariant system and it's great. We must do more research on equivariance to get even more great systems. Well, I was very confused by this because the sixth row there, no IPA, invariant point attention, That removes all the equivariance in AlphaFold. And it hurts a bit, but only a bit. AlphaFold itself on this GDT scale that you can see on the left graph, AlphaFold 2 was about 30 GDT better than AlphaFold 1. And equivariance explains two or three of this. It isn't about one idea.
15:48It's about many mid-scale ideas that add up to a transformative system. And it's very, very important when you're building these systems to think about what we would call in this context biological relevance. We would have ideas that were better. We kind of got our system grinding 1 % at a time. But what really mattered was when we crossed the accuracy that it mattered to an experimental biologist who didn't care about machine learning. And you have to get there through a lot of work and a lot of effort. And when you do, it is incredibly transformative. And we can measure against this axis. were the dark blue axis, the other systems available at the time, and this was assessed.
16:29Protein structure prediction is in some ways far ahead of LLMs or the general machine learning space, and having blind assessment. Since 1994, every two years, everyone interested in predicting the structure of proteins gets together and predicts the structure of 100 proteins whose answer isn't known to anyone except the research group that just solved it, right, unpublished. And so you really do know what works. and we had about a third of the error of any other group on this assessment. But it matters because once you are working on problems in which you don't know the answer, you get to really measure how good things are.
17:05And you can really find that a lot of systems don't live up to what people believe over the course of their research. And because even if you have a benchmark, we all overfit our ideas to the benchmark, right? Unless you have held out. And in fact, the problems you have in the real world are almost always harder than the problems you train on, right? Because you have to learn from much data and you apply it to very important singular problems. So it is very, very important that you measure well, both as you're developing and when people are trying to decide whether they should use your system.
17:39External benchmarks are absolutely critical to figuring out what works, and that's what really helps drive the world forward. So just some wonderful examples of this is typical performance for us. These are blind predictions. You can see they're pretty darn good. Also important, we made it available and we thought it was, and we did a lot of assessment, but we decided that it was very important to make it available in two ways. One is that we open source the code, and we actually open sourced the code about a week before we released a database of predictions, starting originally at 300 ,000 predictions and later going to 200 million.
18:13essentially every protein from an organism whose genome has been sequenced. And this made an enormous difference. And one of the most interesting kind of sociological things is this huge difference between when we released a piece of code that specialists could use, and we got some information, and then when we made it available to the world in this database form, it was really interesting kind of, you know, you release something and every day you check Twitter to find out, or check X to find out what's going on. And what we would really see is even after that CASP assessment, I would say that the structure predictors were convinced.
18:52This obviously was this enormous advance to solve the problem. But general biologists, the people we wanted to use, the people who didn't care about structure prediction, they cared about proteins to do their experiments, they weren't as sure. They said, well, maybe CASP was easy. I don't know. And then this database came out and people got curious. and they clicked in. And the amount to which the proof was social was extraordinary. That people would look and say, how did DeepMind get access to my unpublished structure? You know, there's a moment at which they really believed it, that everyone had a protein, either had a protein that they hadn't solved, or had a friend who had a protein that was unpublished and they could compare.
19:33And that's what really made the difference. And having this database, this accessibility, this ease, led everyone to try it and figure out how it worked. Word of mouth is really how this trust is built. And you can kind of see some of these testimonials, right? I wrestled for three to four months trying to do this scientific task. You know, this morning I got an AlphaFold prediction, and now it's much better. I want my time back, You really appreciate AlphaFold when you run it on a protein that for a year refused to get expressed and purified, meaning for a year they couldn't even get the material to start experiments.
20:13These are really important. When you build the right tool, when you solve the right problem, it matters and it changes the lives of people who are doing things not that you would do, but building on top of your work. And I think it's just extraordinary to see these and the number of people I talked to. The time that I really knew this tool mattered. In fact, there was a special issue of science on the nuclear pore complex a few months after the tool came out. And the special issue was all about this particular very large kind of several hundred protein system. And three out of the four papers in science about this made extensive use of alpha-fold.
20:53I think I counted over 100 mentions of the word alpha fold in science, and we had nothing to do with it. We didn't know it was happening. We weren't collaborating. It was just people doing new science on top of the tools we had built, and that is the greatest feeling in the world. And in fact, users do the darndest things. They will use tools in ways you didn't know were possible. The tweet on the left from Yoshitaka Morawaki came out two days after our code was available. We had predicted the structure of individual proteins, but we were working on building a system that would predict how proteins came together.
21:28But this researcher said, well, I have AlphaFold, why don't I just put two proteins together and I'll put something in between? You can think of this as prompt engineering, but for proteins. And suddenly they find out this is the best protein interaction prediction in the world. that when you train on a really, really powerful system, it will have additional, in some sense, emergent skills as long as they're aligned. People started to find all sorts of problems that AlphaFold would work on that we hadn't anticipated. It was so interesting to see the field of science in real time reacting to the existence of these tools, finding their limitations, finding their possibilities.
22:10And this continues, and people do all sorts of exciting work, be it in protein design, be it in others, on top of either the ideas and often the systems we have built. One application that really I thought was really important is that people have started to learn how to use it to engineer big proteins or to use it in part of. And I want to tell this story for two reasons. One is I think it's a really cool application. but the second is how it really changes the work of science. And often people will say, science is all about experiments and validation. So it's great that you have all these alpha-fold predictions.
22:49Now all we have to do is solve all the proteins the classic way so that we can tell whether your predictions are right or wrong. And they're right about one thing. Science is about experiments. Science is about doing these experiments. but they're wrong about another thing. Science is about making hypotheses and testing them, not about the structure of a particular protein. In this case, the question was, they took this protein on the left, called the contractile injection system, but that's a mouthful. They like to call it the molecular syringe. And what it does is it attaches to a cell and injects a protein into it.
23:30and the scientists at the Jang Lab at MIT were saying, well, can we use this protein to do targeted drug delivery? Can we use it to get gene editors like Cas9 into the cell? They tried over 100 methods to figure out how to take this protein, which they didn't have a structure of, this is just kind of a rendition after the fact, and say, how can we change what it recognizes? I think it's originally involved in plant defense or something like that. and they didn't know how to do it. And they ran an alpha-fold prediction. You can see the one on the left. I wouldn't even say it's a great alpha-fold prediction.
24:06But almost immediately, they looked at that and said, wait a minute, those legs at the bottom are how it must recognize and attach to cells. Why don't we just replace those with a designed protein? And so almost immediately, as soon as they got the alpha-fold prediction, they re-engineered to add this designed protein that you see in red to target a new type of cell. and they take this system and then they show in fact that they can choose cells within a mouse and they can inject proteins in this case fluorescent proteins so there you'll see the color and they can target the cells they want within a mouse brain and so they are using this to develop a new type of system of targeted drug discovery and we see many more examples we see some in which scientists are using this tool to try thousands and thousands of interactions to figure out which ones are likely to be the case.
25:01In fact, discovered a new component of how eggs and sperm come together in fertilization. Many, many of these discoveries that are built on top of this. And I like to think that our work made the whole field of what's called structural biology, biology that deals with structures, you know, five or ten percent faster. But the amount to which that matters for the world is enormous. And we will have more of these discoveries. And I think ultimately structure prediction and larger AI for science should be thought of as an incredible capability to be an amplifier for the work of experimentalists. That we start from these scattered observations, these natural data.
25:43This is our equivalent of all the words on the internet. And then we train a general model that understands the rules underneath it and can fill in the rest of the picture. And I think that we will continue to see this pattern and it will get more general, that we will find the right foundational data sources in order to do this. And I think the other thing that has really been a property is that you start where you have data, but then you find what problems it can be applied to. And so we find enormous advance, enormous capability to understand interactions in the cell or others that are downstream of extracting the scientific content of these predictions and then the rules they use can be adapted to new purposes.
26:29And I think this is really where we see the foundational model aspect of AlphaFold or other narrow systems. And in fact, I think we will start to see this on more general systems, be them LLMs or others, that we will find more and more scientific knowledge within them and we'll use them for important purposes. And I think this is really where this is going. And I think the most exciting question in AI for science is how general will it be? Will we find a couple of narrow places where we have transformative impact? Or will we have very, very broad systems? And I expect it will ultimately be the latter as we figure it out.
Read the full transcript
27:06Thank you.
From the publisher
John Jumper on June 16, 2025 at AI Startup School in San Francisco.
John Jumper is a physicist-turned-computational biologist who led DeepMind’s AlphaFold team—and earned the 2024 Nobel Prize in Chemistry for solving protein folding, a decades-old scientific challenge.
In this talk, he shares how a deep learning breakthrough at CASP14 turned into AlphaFold 1 and then AlphaFold 2, delivering atomic accuracy predictions and revolutionizing biology. He explains the scientific puzzle behind protein folding, the key algorithmic breakthroughs, and the impact of making millions of protein structures accessible to researchers worldwide.




