In short
Podcast Summary: π¬ Beyond AlphaFold: How Boltz is Open-Sourcing the Future of Drug Discovery
Episode Overview
In this episode of the **Latent Space
The AI Engineer Podcast, co-founders Gabriele Corso and Jeremy Wohlwend from Boltz discuss their journey in the evolution of structural biology models, particularly as they transition from AlphaFold to their own open-source solutions, Boltz-1 and Boltz-2**. The conversation focuses on advancements in protein structure prediction, the importance of complex interactions, and the goal of democratizing access to these technologies.
Key Topics Discussed
- Evolution of Structural Biology Models
- AlphaFold Significance
- AlphaFold marked a significant leap in protein structure prediction, particularly for single-chain proteins.
- The models utilize evolutionary data to predict structures based on correlations in protein sequences across species.
- Challenges Remaining
- Although single-chain protein predictions have advanced, challenges remain in modeling complex interactions (protein-protein, protein-ligand) and understanding the dynamics of protein folding.
- Boltz's Approach
- Open-Source Philosophy
- Boltz aims to provide open-source tools to democratize drug discovery and structure prediction.
- They emphasize creating a community around their models to foster collaboration and innovation.
- Modeling Complex Interactions
- Discussion on the shift from regression models to generative models that can handle multiple conformations and uncertainties within molecular structures.
- Boltz-1 and Boltz-2 are designed to not only predict structural arrangements but also engage in generative protein design.
- Generative Protein Design
- Boltz-2's Capabilities
- Boltz-2 integrates structure and sequence prediction as a unified task, allowing users to design proteins based on high-level specifications (e.g., an antibody framework).
- The model also aims to predict binding affinities, which is a crucial aspect of drug development.
- Validation Strategies
- Experimental Validation
- The importance of real-world validation through collaborations with various academic and industry labs to test the efficacy of designed proteins.
- Emphasis on rigorous testing against targets with no known interactions to ensure the models are genuinely innovating rather than regurgitating known data.
- Boltz Lab Product Launch
- Infrastructure and User Interface
- Introduction of Boltz Lab, which offers a platform for scientists to utilize their models more effectively, combining user-friendly interfaces with robust infrastructure.
- The product is designed to scale efficiently, allowing numerous users to run extensive design campaigns simultaneously.
- Future Directions
- Continuous Improvement
- Acknowledgment of the need for ongoing adaptation and refinement of models as the field progresses.
- Boltz is committed to understanding and potentially modeling interactions at the cellular level, which could enhance therapeutic outcomes.
Key Insights
- Generative vs. Regression Models: Transitioning to generative approaches allows for better uncertainty handling and modeling of protein dynamics.
- Community Engagement: The growth of an open-source community is vital for advancing research and developing impactful innovations in drug discovery.
- Real-World Testing: Validation through collaboration with labs is critical for ensuring that models function in practical applications, not just theoretical constructs.
Conclusion The episode concludes with an encouragement for scientists and engineers interested in joining the Boltz mission to democratize drug discovery through open-source models. The co-founders express excitement about the potential advancements in the field and invite collaboration from the broader scientific community.
---
For more information about the podcast and access to full show notes, visit [Latent Space](https://latent.space).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe AlphaFold Breakthrough
0:38 to 2:26
Understanding the significance of AlphaFold in structural biology.
βIt's a pleasure to have with us today Gabriella Corso and Jeremy Volvin.β
Impact on Research and Careers
2:26 to 4:53
How AlphaFold shifted careers and opened new research questions.
βAnd on the second perspective, from a more personal side, seeing the structures coming out of these models, where you see this beautiful creation of life, is something that was very inspiring to me.β
The CASP Competition and AlphaFold 2
4:53 to 6:20
Exploring the importance of CASP in validating protein structure predictions.
βI was at NeurIPS when, I guess, the results of this famous competition came out.β
Evolving Understanding of Protein Structures
6:20 to 9:06
The relationship between evolutionary data and protein structure prediction.
βSo it keeps us in line about what the models can do or not.β
The Complexity of Protein Folding
9:06 to 12:05
Discussing the challenges of understanding protein folding processes.
βAnd there might be intermediate states that it's in sometimes that we're not aware of.β
Hints from Evolution and Predictive Modeling
12:05 to 14:00
How evolutionary hints inform protein folding predictions and models.
βof an insightful point i think one of the interesting things about the protein folding problem is that it used to be actually studied and part of the reason why people thought it was impossible.β
Understanding Protein Structure Predictions
14:00 to 16:44
Learn how models infer protein structures using evolutionary data and physics.
βSo this whole principle is that the structure is probably largely conserved, you know, because there's this function associated with it.β
Advancements in AlphaFold 3
16:44 to 19:40
Explore the significant advancements made in AlphaFold 3 for protein interactions.
βSo there's kind of two different things going on in the kind of coarse grain and then the fine grain optimizations.β
Generative Modeling vs. Regression in AI
19:40 to 21:53
Discover how AlphaFold 3 shifts from regression to generative modeling techniques.
βwhat were some of the key architectural and data changes that made that possible?β
Unique Characteristics of Protein Models
21:53 to 24:40
Understand the challenges and unique characteristics of protein modeling in AI.
βThis field is one of the, I argue, very few fields in applied machine learning where we still have kind of architecture that are very specialized.β
Show all 31 chapters
The Impact of AlphaFold 3's Proprietary Status
24:40 to 27:32
Learn about the implications of AlphaFold 3 not being open-sourced for research.
βPart of it is just exclusively this fact that instead of having operations that operate on the single chain, they operate on the pairwise.β
Boltz: The Open-Source Alternative
27:32 to 28:00
Find out about Boltz1 as an open-source model aiming to democratize drug discovery.
βAnd along the way, and we can talk about it more, but we realized that it was probably two ambitions to see this as an academic project.β
The Speed of Boltz1 Development
28:00 to 29:00
Learn about the rapid development and challenges faced during the Boltz1 project.
βAnd there are a lot of things that were kind of missing.β
Comparing Boltz1 to AlphaFold3
29:00 to 31:00
Explore how Boltz1 measures against AlphaFold3 and its unique strengths.
βwas something that we were exploring independently.β
Evaluating Structural Prediction Models
31:00 to 34:10
Understand the methods to evaluate and compare different structural prediction models.
βAnd then there's some progression from there?β
The Importance of Open Source Feedback
34:10 to 36:20
Discover the value of community feedback in improving open-source models.
βAnd so at the end of the day, it's critical, and this is also something across other fields of machine learning.β
Balancing Open Source and Product Development
36:20 to 40:00
Learn how Boltz balances open-source initiatives with developing practical products.
βBut I think one thing that's probably undeniable is just the pace of progress and how much better we're getting every year.β
Building the Boltz Community
40:00 to 42:03
Find out how the Boltz community evolved and its impact on the company.
βBut I just maybe open the GPT app or Cloud Code and just use it as an amazing product.β
Building a Self-Sustaining Community
42:03 to 44:50
Learn how community engagement and usability have fostered a thriving platform.
βyou know, to be able to like answer everyone's questions and help.β
Innovative Contributions from the Community
44:51 to 47:24
Discover surprising contributions from the community that enhanced the project.
βWere there any, like, if someone was doing something and you're like, why would you do that?β
Advancements in Protein Design with BoltzGen
47:25 to 52:03
Explore how BoltzGen integrates structure prediction and protein design.
βAnd that speaks to like the, you know, the power of scoring.β
Experimental Validation and Broader Testing
52:04 to 56:01
Understand the importance of broad experimental validation across various applications.
βAnd so the way that Boltzgen works is that you are basically the only thing that you're doing is predicting the structure.β
Testing Models Across Diverse Applications
56:01 to 58:29
Learn how diverse lab validations enhance the credibility of protein design models.
βtask with all sorts of different applications from therapeutic to, you know, biosensors and many others that, you know, so can we get a validation that is kind of goes across many different tasks?β
Insights into Experimental Validation Methodology
58:30 to 1:02:04
Discover the specifics of designing proteins and the success metrics used.
βAnd I was very excited about seeing like all the diverse validations that you've done.β
Introducing Bolt's Lab: Revolutionizing Drug Design
1:02:05 to 1:07:48
Understand the objectives and functionalities of Bolt's Lab platform for drug design.
βNanomolar, roughly speaking, is just a measure of how strongly the interaction is.β
Validating Innovative Agentic Systems
1:07:49 to 1:10:00
Explore the validation processes for novel protein designs outside traditional methods.
βin the LLM space, the cost of a token has gone down by a factor of 1 ,000 or so over the last three years.β
Experimental Validation Across Multiple Targets
1:10:00 to 1:11:54
Learn about the approach to experimental validation that emphasizes statistically significant results across various therapeutic targets.
βAnd so that on the one end, we can get a much more statistically significant result and really allows us to make progress from the methodological side without being steered by overfitting on any one particular system.β
Releasing Strong Binders and Open Source Philosophy
1:11:54 to 1:13:44
Discover the commitment to open-source drug discovery and the importance of clarifying the distinction between designed molecules and actual drugs.
βbasically they keep it, they do whatever they want with it.β
Navigating Challenges in Drug Development
1:13:44 to 1:16:13
Understand the complexities of drug development and the need for continuous model improvement to address unique therapeutic hypotheses.
βAnd that to some extent does mean that we need to try to go deeper and deeper in getting these models better and better.β
Collaborating with Medicinal Chemists
1:16:13 to 1:18:45
Explore the relationship between machine learning and medicinal chemists, focusing on the challenges and breakthroughs in collaboration.
βcertain proteins interfere, interact with pathways that are existing in the cell.β
The Importance of Lab Results in Convincing Skeptics
1:18:45 to 1:19:43
Learn how tangible lab results can influence skeptics in the field and the role of experience in gaining their trust.
βto understand, you know, the different things.β
Transcript
Automatic transcript. May contain errors.0:00Actually, we only trained the big model once. that's how much compute we had. We could only train it once. And so while the model was training, we were finding bugs left and right. A lot of them that I wrote. I remember us doing surgery in the middle, stopping the run, making the fix, relaunching. We never actually went back to the start. We just kept training it with the bug fixes along the way. It was impossible to reproduce now. Yeah, yeah, no, that model has gone through such a curriculum that he's learned some weird stuff. But yeah, somehow by miracle it worked out. It's a pleasure to have with us today Gabriella Corso and Jeremy Volvin.
0:42They recently founded Volts, a company trying to democratize and bring art, structure, prediction, and biology to the masses. They are both recent PhD grads from MIT and have been working on all sorts of foundational papers in generative biology. Anyway, pleasure to have you here. Thanks for coming. Thank you. I guess we're maybe, what, six years post AlphaFold 2 right now, which was kind of a big moment. Is that right? I think it was 2021. So, yeah, going on five years. Five years, yeah. Yeah, so maybe for the audience, let's go back to that moment in time and explain, what was this big moment, and why was it interesting?
1:27Why was everyone so excited? And I think you two were probably quite excited, so why were you personally excited? I would start on kind of why that was interesting, kind of from a scientific standpoint. So maybe first as a kind of introduction for the ones in the audience and not structural biologists. So the idea of structural biology is that we want to try to understand how proteins and other molecules take shape inside our cells and how they interact. And structural biology is sort of this beautiful discipline where we are somehow able to understand this minuscule structure at kind of atomic details using these incredibly complex methods like, you know, x-ray crystallography.
2:17and you know the the dream has always been of computational biology can we understand kind of the structures without having to you know resolve this crystal you know shoot x-rays and so on and so alpha fold was a real breakthrough in this problem of protein folding which is trying to understand the structure of a single protein and to me it was exciting across kind of many dimensions One, I was a computer scientist, I was working a lot on machine learning, and I saw the impact that the work similar to what I was doing could have on a longstanding scientific problem. And on the second perspective, from a more personal side, seeing the structures coming out of these models, where you see this beautiful creation of life, is something that was very inspiring to me.
3:19And so that was one of the things that led me to start working on structural biology, in particular with machine learning. Were you a structural biologist before AlphaFold came out? You did machine learning, but it was not in structural biology, so that actually shifted your career quite dramatically. Yeah, very dramatically. I was working on some pretty theoretical, methodological things. I was starting to see some of the challenges in doing some more theoretical or methodological work. and seeing the potential impact of doing excellent. AlphaFold was really a machine learning breakthrough and applied machine learning.
4:02And so that led me to want to start working in applied ML. Our group at the time was working a lot on small molecules already. And I think AlphaFold is kind of what triggered this shift to working on biologics. And at the time, I think it opened as many questions as it answered in a sense. The immediate follow-ups were, okay, can we do this on other things than proteins? Can we do interactions of small molecules with proteins, nucleic acid with proteins? Can we model more complex protein systems? And I think very rapidly, I think, after AlphaFold, people realized, I think, that machine learning could really target this problem very differently than previous methodologies.
4:48Going back to the AlphaFold2 moment, I remember this very well. I was at NeurIPS when, I guess, the results of this famous competition came out. So you wanted to talk about CASP and what it is and why it was so interesting and exciting. I think every couple of years, the goal has always been to find protein structures that are a little bit different from what's known. So CASP over the years has put in a lot of effort to gather structures from academic groups and even industry groups to try to create sort of a test set that would be difficult for different methods. And CASP 14 was when AlphaFold 2 really blew everything out of the water.
5:38The improvement was so large over the previous method and also over the previous competitions. And now CASP continues. We've had CASP-15, we have CASP-16. And what's happened now is that it's really expanding to also all these other modalities. Like I was mentioning, like protein small molecule, nucleic acid. But the goal remains to really challenge the models. How well do these models generalize? And we've seen in some of the latest CASP competitions, while we've become really, really good at proteins, basically monomeric proteins, other modalities still remain pretty difficult. So it's really essential in the field that there are these efforts to gather benchmarks that are challenging.
6:22So it keeps us in line about what the models can do or not. It's interesting you say that. In some sense, at CAS 14, a problem was solved pretty comprehensively, but at the same time it was really only the beginning. So you can explain, what was the specific problem you would argue is solved? And then what is remaining, which is probably quite open. I think we'll steer away from the term solved, because we have many friends in the community who get pretty upset at that word. And I think fairly so. But the problem that a lot of progress was made on was the ability to predict the structure of single-chain proteins.
7:06So proteins can be composed of many chains and single chain proteins are just a single sequence of amino acids. And one of the reasons that we've been able to make such progress is also because we take a lot of hints from evolution. So the way the models work is that they sort of decode a lot of hints that comes from evolutionary landscapes. So if you have some protein in an animal and you go find the similar protein across different different organisms, you might find different mutations in them. And as it turns out, if you take a lot of these sequences together and you analyze them, you see that some positions in the sequence tend to evolve at the same time as other positions in the sequence.
7:50Sort of this correlation between different positions. And it turns out that that is typically a hint that these two positions are close in three dimension. So part of the breakthrough has been our ability to also decode that very, very effectively. But what it implies also is that in absence of that co-evolutionary landscape, the models don't quite perform as well. And so I think when that information is available, maybe one could say the problem is somewhat solved from the perspective of structure prediction. When it isn't, it's much more challenging. And I think it's also worth also differentiating the, sometimes we confound a little bit, structure prediction and folding.
8:32Folding is the more complex process of actually understanding how it goes from this disordered state into a structured state. And that I don't think we've made that much progress on. But the idea of going straight to the answer, we've become pretty good at. So there's this protein that is just a long chain and it folds up. And so we're good at getting from that long chain in whatever form it was originally to the thing, but we don't know how it necessarily gets to that state. And there might be intermediate states that it's in sometimes that we're not aware of. That's right. And that relates also to our general ability to model the different...
9:16Proteins are not static. They move. They take different shapes based on their energy states. And I think we are also not that good at understanding the different states that the protein can be in and at what frequency, what probability. So I think the two problems are quite related in some ways. still a lot to solve. But I think it was very surprising at the time, you know, that even with these evolutionary hints that we were able to, you know, to make such dramatic progress. So I want to ask, why does the intermediate states matter? But first, I kind of want to understand, why do we care what proteins are shaped like?
9:55Yeah. I mean, the proteins are kind of the machines of our body. You know, the way that all the processes that we have in our cells, you know, work is typically through proteins, sometimes other molecules, sort of intermediate interactions. And through that interactions, we have all sorts of cell functions. And so when we try to understand, you know, a lot of biology, how our body works, how disease work, we often try to boil it down to, okay, what is going right in the case of our normal biological function and what is going wrong in the case of the disease state. And we boil it down to proteins and other molecules and their interaction.
10:41And so when we try predicting the structure of proteins, it's critical to have an understanding of those interactions. It's a bit like seeing the difference between having a list of parts that you would put it in a car and seeing kind of the car in its final form, you know, seeing the car really helps you understand what it does. On the other hand, kind of going to your question of, you know, why do we care about, you know, how the protein folds or, you know, how the car is made to some extent is that, you know, sometimes when something goes wrong, you know, there are, you know, cases of, you know, proteins misfolding in some diseases and so on.
11:24if we don't understand this folding process we don't really know how to intervene there's this nice line in the um i think it's in the alpha fold two manuscript where they sort of discuss also like why we even hopeful that we can target the problem in the first place and then this this notion that like well four proteins that fold the folding process is almost instantaneous which is a strong like you know signal that like yeah like we should we might be able to predict that this very like constrained uh thing that that the protein does and so quickly and of course that's not the case for you know for for all proteins and there's a lot of like really interesting mechanisms in the cells but yeah i remember reading that and thought yeah that's somewhat of an insightful point i think one of the interesting things about the protein folding problem is that it used to be actually studied and part of the reason why people thought it was impossible.
12:18It used to be studied as kind of like a classical example of like an MP problem. Like there are so many different, you know, type of, you know, shapes that, you know, this amino acid could take. And so this grows combinatorially with the size of the sequence. And so there used to be kind of a lot of actually kind of more theoretical computer science thinking about and starting protein folding as an MP problem. And so it was very surprising also from that perspective, kind of seeing machine learning so clear there is some, you know, signal in those sequences through evolution, but also through kind of other things that, you know, us as humans, we're probably not really able to understand, but that this models have learned.
13:07And so Andrew White, we were talking to him a few weeks ago, and he said that he was following the development of this and that there were actually ASICs that were developed just to solve this problem. So again, that there were many, many, many millions of computational hours spent trying to solve this problem before alpha fold. And just to be clear, one thing that you mentioned was that there's this kind of co-evolution of mutations and that you see this again and again in different species. So explain why does that give us a good hint that they're close by to each other? Yeah. Like, think of it this way that, you know, if I have some amino acid that mutates, it's going to impact everything around it, right, in three dimensions.
13:51And so it's almost like the protein through several probably random mutations in evolution, like, you know, it ends up sort of figuring out that this other amino acid needs to change as well for the structure to be conserved. So this whole principle is that the structure is probably largely conserved, you know, because there's this function associated with it. And so it's really sort of like different positions compensating for each other. I see. Those hints in aggregate give us a lot of information about what is close to each other. And then you can start to look at what kinds of folds are possible given the structure.
14:28And then what is the end state? And therefore you can make a lot of inferences about what the actual total shape is. Yeah, that's right. It's almost like you have this big three-dimensional valley where you're trying to find these low-energy states. And there's so much to search through that's almost overwhelming. But these hints, they sort of maybe put you in an area of the space that's already kind of close to the solution, maybe not quite there yet. And there's always this question of how much physics are these models learning versus just pure statistics. and like I think one of the thing at least I believe is that once you're in that sort of approximate area of the solution space then the models have like some understanding you know of how to get you to like you know the low energy low energy state and so maybe you have some light understanding of physics but maybe not quite enough you know to know how to like navigate the whole space well so we need to give it these hints to kind of get there.
15:27So you get it into the right valley and then it finds the The minimum or something. One interesting explanation about how half-fold three works that I think is quite insightful, of course, doesn't cover kind of the entirety of what half-fold does. That is, I'm going to borrow from Sergei Chinnikov at MIT. So he sees kind of half-fold. And the interesting thing about half-fold is it's got this very peculiar architecture that we have since, you know, used. And this architecture operates on this, you know, pairwise context between amino acids. And so the idea is that probably the MSA gives you this first hint about what potential amino acids are close to each other.
16:06MSA is multiple sequence alignment. Exactly, this evolutionary information. And from this evolutionary information about potential contacts, then it's almost as if the model is running some kind of diastro algorithm where it's decoding, okay, these have to be closed. Then if these are closed and this is connected to this, then this has to be somewhat closed. And so you decode this that becomes basically a pairwise kind of distance matrix. And then from this rough pairwise distance matrix, you decode kind of the actual potential structure. Interesting. So there's kind of two different things going on in the kind of coarse grain and then the fine grain optimizations.
16:51Interesting. Yeah. Very cool. Yeah. You mentioned AlphaFold3, so maybe we have a good time to move on to that. So AlphaFold 2 came out, and it was, I think, fairly groundbreaking for this field. Everyone got very excited. A few years later, AlphaFold 3 came out. And maybe for some more history, what were the advancements in AlphaFold 3? And then I think maybe after that we'll talk a bit about how it connects to Boltz. Yeah, so after AlphaFold 2 came out, Jeremy and I got into the field. And with many others, the clear problem that was obvious after that was, okay, now we can do individual chains.
17:30Can we do interactions? Interaction of different proteins, proteins with small molecules, proteins with other molecules. So, quick, why are interactions important? Interactions are important because, to some extent, that's kind of the way that these machines, these proteins, have a function. The function comes by the way that they interact with other proteins and other molecules. Actually, in the first place, the individual machines are often, as Jeremy was mentioning, not made of a single chain, but they're made of multiple chains. And then these multiple chains interact with other molecules to give the function to those.
18:09And on the other hand, when we try to intervene on these interactions, think about a disease, think about a biosensor, or many other ways, we are trying to design molecules or proteins that interact in a particular way with what we would call a target protein or target. But, you know, this problem after AlphaFold 2, you know, became clear, kind of one of the biggest problems in the field to solve. Many groups, including kind of ours and others, you know, started making some kind of contributions to this problem of trying to model these interactions. And AlphaFold 3 was, you know, was a significant advancement on the problem of modeling interactions.
18:49And one of the interesting things that they were able to do while some of the rest of the field really tried to model different interactions separately, how protein interacts with small molecules, how protein interacts with other proteins, how RNA or DNA have their structure. they put everything together and train very large models with a lot of advances, including kind of changing some of the key architectural choices, and managed to get a single model that was able to set this new state-of-the-art performance across all of these different kind of modalities, whether that was protein-small molecules that is critical to developing new drugs, a protein-protein, understanding interactions of proteins with RNA and DNAs and so on.
19:40Just to satisfy the AI engineers in the audience, what were some of the key architectural and data changes that made that possible? Yeah, so one critical one that was not necessarily just unique to AlphaFold3, but there were actually a few other teams, including ours in the field that proposed this, was moving from modeling structure prediction as a regression problem, so where there is a single answer and you're trying to shoot for that answer, to a generative modeling problem where you have a posterior distribution of possible structures and you're trying to sample this distribution. And this achieves two things.
20:20One is it starts to allow us to try to model more dynamic systems. As we said, some of these structures can actually take multiple structures. And so you can now model that through kind of modeling the entire distribution. But on the second hand, from more kind of core modeling questions, when you move from a regression problem to a generative modeling problem, you are really tackling the way that you think about uncertainty in the model in a different way. So if you think about, I'm undecided between different answers, What's going to happen in a regression model is that I'm going to try to make an average of those different kind of answers that I had in mind.
21:05When you have a generative model, what you're going to do is sample all these different answers and then maybe use separate models to analyze those different answers and pick out the best. So that was kind of one of the critical improvements. The other improvement is that they significantly simplified to some extent the architecture, especially of the final model that takes those pairwise representations and turns them into an actual structure. And that now looks a lot more like a more traditional transformer than a very specialized equivariant architecture that it was in AlphaFold3. So this is a bitter lesson, a little bit?
21:45There is some aspect of a bitter lesson, but the interesting thing is that it's very far from being a simple transformer. This field is one of the, I argue, very few fields in applied machine learning where we still have kind of architecture that are very specialized. And, you know, there are many people that have tried to replace these architectures with simple transformers. And, you know, there is a lot of debate in the field, but I think kind of the most of the consensus is that, you know, the performance that we get from the specialized architecture is vastly superior than what we get through a single transformer.
22:21Another interesting thing that I think staying on the modeling machine learning side, which I think is somewhat counterintuitive seeing some of the other fields and applications, is that scaling hasn't really worked the same in this field. Now, models like AlphaVault 2 and AlphaVault 3 are still very large models, but at the same time, in terms of parameters, they're actually not very big. They are definitely below a billion parameters. If you're here these days in LLM space, a model with less than a billion parameters, you'd think it can't do anything. But on the other hand, when you look at the computational cost of running these models, they are actually a lot more expensive than it is to run language models.
23:10Because as Jeremy was saying, we go from, instead of having quadratic operations, now a cubic operation. And so it's interesting how right now in the field, and this is maybe related to having less data or needing more inductive biases, but we have this ratio of amount of computation to parameters that is much, much higher than in other places. If I recall, AlphaFold 2 was like, what, 70 million parameters? Something like that? Yeah, it's something like that. It's quite small. It's around 100 or so. So these decisions of triangle layers and these, for AlphaFull too, this interesting equivariant architecture really were priors that baked in a lot of the physics of the system.
23:58And also co-evolution data is, I think people have argued, that is kind of almost like a database lookup of some sorts. So that provides in some sense more parameters as well. Yeah, I mean, it's more definitely the amount of pure compute flops. is very high and it's almost like more reasoning based maybe than more just like information extraction. I think one of the things that, part of the reason the LLMs are so large isn't just because of their reasoning capability, but it's also because of the sheer quantity of information that they store. And I think here there's a little bit less of that.
24:33And I think it's more about decoding this input rather than maybe memorizing as much of it. So is there a loop in the architecture that allows How does it compute more per parameter? How does that work? Part of it is just exclusively this fact that instead of having operations that operate on the single chain, they operate on the pairwise. And so you, instead of having a quadratic number of interactions, you have a cubic number of interactions. And so that, on its own, leads you to have smaller representation sizes, but more representation. That leads to more flops, but fewer parameters. On the other hand, And there is actually also this idea of, you know, somewhat similar to reasoning where you recycle kind of these operations from AlphaFold 2, but also kind of AlphaFold 3.
25:23They have this interesting framework where, you know, as we were discussing, kind of the input to the model is sort of like this initial understanding of the interactions, either from the evolution of the multiple sequence, but also potentially from what we call templates that are basically database lookup of similar structures. And so how the model works is that it decodes this and tries to understand a good potential rough structure of the pairwise interaction. And then what you can do is basically do this recycling where you feed this understanding back to the input of the model and then try to decode it again.
26:02And people do this three or four times, and in some cases I've even tried to do it tens of times. And so you can see it as a very, very early version of reasoning or trying to get. Yeah, so AlphaFold 2, really cool. AlphaFold 3, really cool. But AlphaFold 3 came with a catch. And I think this catch was important for the development of bolts and so on. Yeah, the catch was that it was an amazing paper, Nature paper. But unfortunately, they decided not to release the model. AlphaFold 2 was open source and since then was used, I think, the reported numbers is more than a million scientists. AlphaFold 3, for commercial reasons that did mine since spinoff isomorphic lab that is now trying to become a new pharmaceutical company, had decided to keep this model internal and only use it internally.
27:05And now both, you know, we were in the field and building on top of models like AlphaFold. And so now we no longer had, you know, kind of the base starting point to build on top. But even more importantly, everyone in both kind of academic research and in industry no longer had access to these incredible models that, you know, was really useful to try to understand biologists, but also try to develop new therapeutics. I decided to take the matter in our own hands and decided to try to obtain a model that was of similar accuracy. And so largely also using a lot of the information that was in the half-fold-free manuscript, we went ahead and built Pulse1, which was the first fully open-source model to approach the level of accuracy of half-fold-free.
27:59And along the way, and we can talk about it more, but we realized that it was probably two ambitions to see this as an academic project. And there are a lot of things that were kind of missing. And so we decided to also start a public benefit company to push this mission of democratizing access to these models that we started with Botswana. Quick interjection. I mean, I remember this. It was actually shocking how fast you got Boltz1 out. Like it was just like two or three months, right? I think we started in late May and it came in November. If I remember correctly. So slightly longer, but yeah, yeah, it was relatively quick.
28:47I mean, for what it's worth, like, you know, we were working on some of the some similar ideas at the time. I think like we, you know, for example, this idea of like having a diffusion model on top of this pairwise trunk was something that we were exploring independently. Now, when the paper came out, it was really clear, especially, for example, on the data pipelines. There was so much that we were not really doing, and so there was a lot to catch up on. But we were already in a place, I think, where we had some experience working with the data and working with this type of models, and I think that put us already in a good place to produce it quickly.
29:25And, you know, and I would, I would even say like, I think we could have done it quicker. The problem was like for a while, we didn't really have the compute. And so we couldn't really train the model. And actually we only trained the big model once. That's how much compute we had. We could only train it once. And so like, while the model was training, we were like finding bugs left and right. A lot of them that I wrote. And like, I would, I remember like us like sort of like, you know, doing like surgery in the middle, stopping the run, making the fix, relaunching. We never actually went back to the start, we just kept training it with the bug fixes along the way.
Read the full transcript
30:05That model has gone through such a curriculum that it's learned some weird stuff. Somehow by miracle it worked out. The other funny thing is that the way that we were training most of that model was through a cluster from the Department of Energy. but that's sort of like a shared cluster that many groups use. And so we were basically training the model for two days and then he would go back to the queue and stay a week in the queue. And so it was pretty painful. And so we actually kind of towards the end with Devon, the CEO of Genesis, and basically I was telling him a bit about the project and kind of telling him about this frustration with the compute.
30:45And so luckily he offered to kind of help. And so we got the help from Genesis to finish up the model. Otherwise, it probably would have taken a couple of extra weeks. Wow. Of weeks. Yeah. Yeah. Boltz1, how did that compare to AlphaFold3? And then there's some progression from there? Yeah. So I would say kind of that Boltz1, but also kind of these other kind of set of models that came around the same time were kind of approaching, were a big leap from, you know, kind of the previous kind of open source models and, you know, kind of really kind of approaching the level of AlphaVault 3. I would still say that, you know, even to this day, there are, you know, some specific instances where AlphaVault 3 works better.
31:36I think one common example is antibody antigen prediction, where, you know, AlphaFold3 still seems to have an edge in many situations. Obviously, these are somewhat different models. They are, you know, you run them, you obtain different results. So it's not always the case that one model is better than the other, but kind of in aggregate we still, especially at the time, so AlphaFold3 is still having a bit of an edge. We should talk about this more when we talk about Voltgen, but how do you know one model is better than the other? So I make a prediction, you make a prediction, Like, how do you know?
32:12Yeah, so easily, you know, the great thing about kind of structural prediction, and, you know, once we're going to go into the design space of designing new small molecule and new proteins, this becomes a lot more complex. But a great thing about structural prediction is that a bit like, you know, Casp was doing, basically the way that you can evaluate them is that, you know, you train the model on a structure that was released across the field up until a certain time. And, you know, one of the things that we didn't talk about that was really critical in all this development is the PDB, which is the protein databank.
32:46It's this common resource, it's a common database where every biologist publishes their structures. And so we can, you know, train on, you know, all the structures that were put in the PDB until a certain date. And then we basically look for recent structures, okay, which structures look pretty different from anything that was published before, because we really want to try to understand generalization. And on this new structure, we evaluate all these different models. And so you just know when AlphaFolus 3 was trained, you know when you're intentionally trained to the same data or something like that.
33:23Exactly, right? Yeah. And so this is kind of the way that you can somewhat easily kind of compare these models. obviously that assumes that, you know, the training. You've always been very passionate about validation. I remember like diff doc and then there was like diff doc L and doc gen. You've thought very carefully about this in the past. Like, actually, I think doc gen is like a really funny story that I think, I don't know if you want to talk about that. It's an interesting like. Yeah, I think one of the amazing things about putting things open source is that we get a ton of feedback from the field.
33:58And sometimes we get great feedback of people really liking the model. But honestly, most of the times, to be honest, that's also maybe the most useful feedback, is people sharing about where it doesn't work. And so at the end of the day, it's critical, and this is also something across other fields of machine learning. It's always critical to do progress in machine learning, set clear benchmarks. and as you start doing progress of certain benchmarks, then you need to improve the benchmarks and make them harder and harder. And this is kind of the progression of how the field operates. And so the example of DocGen was we published this initial model called DiffDoc in my first year of PhD, which was sort of like one of the early models to try to predict kind of interactions between proteins, small molecules that we bought a year after AlphaFold II was published.
35:03And now on the one hand, you know, on these benchmarks that we were using at the time, DivDoc was doing really well, kind of, you know, outperforming kind of some of the traditional physics-based methods. But on the other hand, you know, when we started, you know, kind of giving these tools to kind of many biologists, and one example that we collaborated with was the group of Nick Polizzi at Harvard, we started noticing that there was this clear pattern where four proteins that were very different from the ones that we're trained on, the models were struggling. And so, you know, that seemed clear that, you know, this is probably kind of where we should, you know, put our focus on.
35:47And so we first developed, you know, with Nick and his group a new benchmark and then, you know, went after and said, OK, what can we change and kind of about the current architecture to improve this pattern of generalization. And this is the same that, you know, we're still doing today, you know, kind of where does the model not work? You know, and then, you know, once we have that benchmark, you know, let's try to throw everything, any ideas that we have of the problem. And there's a lot of healthy skepticism in the field, which I think is great. And I think it's very clear that there's a ton of things the models don't really work well on.
36:24But I think one thing that's probably undeniable is just the pace of progress and how much better we're getting every year. And so I think if you assume any constant rate of progress moving forward, I think things are going to look pretty cool at some point in the future. ChatGPT was only three years ago. Yeah, it's wild, right? What? Yeah, it's one of those things, even being in the field, you don't see it coming. I think, yeah, hopefully we'll continue to have as much progress as we've had the past few years. So this is maybe an aside, but I'm really curious. You get this great feedback from the community by being open source.
37:06My question is partly like, okay, yeah, if you open source, then everyone can copy what you did. but it's also maybe balancing priorities, right? Where all my customers are saying, I want this. There's all these problems with the model. Yeah, yeah, but my customers don't care, right? So how do you think about that? Yeah, so I would say a couple of things. One is part of our goal with Boltz, and this is also kind of established as kind of the mission of the public benefit company that we started, is to democratize the access to these tools. But one of the reasons why we realized that Boltz needed to be a company, it couldn't just be an academic project, is that putting a model on GitHub is definitely not enough to get chemists and biologists across both academia, biotech, and pharma to use your model in their therapeutic programs.
38:03And so a lot of what we think about at Bolts Beyond, just the models, is thinking about all the layers that come on top of the models to get from those models to something that can really enable scientists in the industry. And so that goes into building the right kind of workflows that take in, for example, the data and try to answer directly those problems that the chemists and the biologists are asking. and then also kind of building the infrastructure. And so this to say that, you know, even with models fully open, you know, we see a ton of potential for, you know, products in the space. And the critical part about a product is that even, you know, for example, with an open source model, you know, running the model is not free.
38:53You know, as we were saying, these are a pretty expensive model and especially, and maybe we'll get into this, you know, these days we're seeing pretty dramatic inference time scaling of these models where the more you run them, the better the results are. But there you start getting into a point that compute and compute costs becomes a critical factor. And so putting a lot of work into building the right kind of infrastructure, building the optimizations and so on, really allows us to provide a much better service potentially to the open source models. That to say, even though we're the product, we can provide a much better service, I do still think, and we will continue to put a lot of our models open source because the critical role of open source models is helping the community progress on the research from which we all benefit.
39:47And so we'll continue to, on the one end, put some of our base models open source so that the field can be on top of it. And as we discussed earlier, we learn a ton from the way that the field uses and builds on top of our models, but then try to build a product that gives the best experience possible to scientists so that a chemist or a biologist doesn't need to spin off a GPU and set up our open source model in a particular way, but can just be like, even though I am a computer scientist, machine learning scientist, I don't necessarily take an open source LLM and try to kind of spin it off. But I just maybe open the GPT app or Cloud Code and just use it as an amazing product.
40:37We kind of want to give the same experience to scientists from the world. I heard a good analogy yesterday that a surgeon doesn't want the hospital to design a scalpel, right? So just buy the scalpel. you wouldn't believe like the number of people even like in my short time you know between a full three coming out and in the end of the phd like the number of people that would like reach out just for like us to like run up a full three for them you know or things like that just because like you know bolts in our case you know just because it's like not that easy you know to do that you know if you're not a computational person and i think like part of the goal here is also that we continue to obviously build an interface with computational folks, but the models are also accessible to a larger, broader audience.
41:24And that comes from good interfaces and stuff like that. I think one really interesting thing about Boltz is that with the release of it, you didn't just release a model, but you created a community. That community grew very quickly. Did that surprise you? What is the evolution of that community and how is that fed into Boltz? company. If you look at its growth, it's like very much like when we release a new model, it's like, there's a big, big jump. But yeah, it's, I mean, it's been great. You know, we have a Slack community that has like thousands of people on it. And it's actually like self-sustaining now, which is like the really nice part because, you know, it's, it's almost overwhelming, I think, you know, to be able to like answer everyone's questions and help.
42:06It's really difficult, you know, with the few people that we were, but it ended up that like, you know, people would answer each other's questions and help one another. And so the Slack has been kind of self-sustaining and that's been really cool to see. And that's for the Slack part, but then also obviously on GitHub as well, we've had a nice community. I think we also aspire to be even more active on it than we've been in the past six months, which has been a bit challenging for us. But yeah, the community has been really great And there's a lot of papers also that have come out with new evolutions on top of bolts.
42:45And it surprised us to some degree because there's a lot of models out there. And I think people converging on that was really cool. And it speaks also, I think, to the importance of when you put code out, to try to put a lot of emphasis in making it as easy to use as possible. And something we thought a lot about when we released the code base. It's far from perfect, but... Do you think that that was one of the factors that cause your community to grow? Is it just the focus on easy to use, make it accessible? I think so, yeah. And we've heard it from a few people over the years now. And some people still think it should be a lot nicer.
43:21And they're right. And they're right. But yeah, I think it was, at the time, maybe a little bit easier than other things. The other, I think, part that I think led to the community and to some extent, I think, somewhat the trust in the community and what we put out is the fact that it's not really been one model, and maybe we'll talk about it, after Boltz 1, there were maybe another couple of models released or open source soon after. We continued that open source journey, released Boltz 2, where we were not only improving the structure prediction, but also starting to do affinity predictions, understanding the strength of the interactions between these different models, which is this critical component, critical property that you often want to optimize in discovery programs.
44:12And then more recently, also kind of protein design model. And so we've sort of been building this suite of models that come together, interact with one another, where there is almost an expectation that we take very at heart of always having across the entire suite of different tasks the best or across the best model out there so that our open source tool can be the go-to model for everybody in the industry. I really want to talk about Bold's Gen, but before that, one last question in this direction. Was there anything about the community which surprised you? Were there any, like, if someone was doing something and you're like, why would you do that?
44:57That's crazy. or that's actually genius. I never would have thought about that. I mean, we've had many contributions. I think some of the interesting ones, like we had this one individual who wrote a complex GPU kernel for part of the architecture. The funny thing is that piece of the architecture had been there since AlphaFold 2. and I don't know why it took bolts for this person to decide to do it but that was a really great contribution. We've had a bunch of others, people figuring out ways to hack the model to do cyclic peptides. I don't know if there's any other interesting things come to mind.
45:41One cool one and this was something that initially was proposed as a message in the Slack channel by Tim O'Donnell was basically he was, you know, there are some cases, especially, for example, we discussed, you know, antibody-antigen interactions where the models don't necessarily kind of get the right answer. What he noticed is that, you know, the models were somewhat stuck into predicting kind of the antibody to interact with a part of the antigen that was incorrect. And so he basically ran the experiments. In this model, you can condition, basically you can give hints. And so he basically gave, you know, random hints to the model basically okay you should buy into this residue uh you should buy into the first residue or you should buy into the 11th residue or you should buy to the 21st residue you know basically every 10 residues scanning the entire antigen and residues are the the amino acids yeah so the first amino acids the 11 amino acids and so on so it's sort of like doing a scan and then you know conditioning the model to predict all of them and then looking at the confidence of the model in each of those cases and taking the top.
46:48And so it's sort of like a very somewhat crude way of doing kind of inference time search. But surprisingly, you know, for antibody antigen friction, it actually kind of helped quite a bit. And so there's some, you know, interesting ideas that, you know, as obviously as kind of developing the model, you say kind of, you know, wow, this is why would the model, you know, be so dumb? But, you know, it's very interesting. And that, you know, leads you to also kind of, you know, start thinking about, okay, how can I do this, you know, not with this brute force, but, you know, in a smarter way. And so we've also done a lot of work on that direction.
47:25And that speaks to like the, you know, the power of scoring. We're seeing that a lot. I'm sure we'll talk about it more when we talk about Bolzgem. But, you know, our ability to like take a structure and determine that that structure is like good, you know, like somewhat accurate, whether that's a single chain or an interaction is a really powerful way of improving the models. If you can sample a ton and you assume that if you sample enough you're likely to have the good structure then it really just becomes a ranking problem. And now part of the inference time scaling that Gabri was talking about is very much that.
48:03The more we sample, the more the ranking model ends up finding something it really likes. And so I think our ability to get better at ranking, I think, is also what's going to enable sort of the next, you know, next big, big breakthroughs. Interesting. I guess there's a, my understanding, there's a diffusion model and you generate some stuff and then you, I guess it's just what you said, right? Then you rank it using a score and then you finally, and so like, can you talk about those different parts? Yeah, so first of all, one of the critical beliefs that we had also when we started working on BOLTS1 was the structure prediction models are somewhat our field version of some foundation models.
48:50Learning about how proteins and other molecules interact, and then we can leverage that learning to do all sorts of other things. And so with Boltz II, we leverage that learning to do affinity predictions. So understanding kind of, you know, if I give you this protein, these small molecules, how tightly is the interaction? For Boltz II, what we did was taking kind of that kind of foundation models and then fine tune it to predict kind of entire new proteins. And so the way basically that that works is sort of like instead of for the protein that you're designing, instead of feeding in an actual sequence, you feed in a set of blank tokens and you train the models to predict both the structure of that protein and with the structure also what the different amino acids of that proteins are.
49:42And so basically the way that Boltz-Chain operates is that you feed a target protein that you may want to kind of bind to or another DNA, RNA. and then you feed the high-level kind of design specification of what you want your new protein to be. For example, it could be like an antibody with a particular framework, could be a peptide, could be many other things. And that's with natural language? And that's basically prompting, and we have kind of this sort of like spec that you specify. And you feed kind of this spec to the model, and then the model translates this into a set of tokens, a set of conditioning to the model, a set of blank tokens.
50:30And then basically the codes, as part of the diffusion models, decodes a new structure and a new sequence for your protein. And then we take that, and as Jeremy was saying, trying to score it, how good of a binder it is to that original target. that you can you're using basically bolts to predict the folding and the affinity to that molecule so and then that is your that kind of gives you a score exactly so you use this model to predict the structure and then you do two things one is that you predict the structure and with something like bolts two and then you basically compare that structure with what the model predicted, what Boltzsched predicted.
51:19And this is sort of like in the field called consistency. It's basically you want to make sure that the structure that you're predicting is actually what you're trying to design. And that gives you a much better confidence that that's a good design. And so that's the first filtering. And the second filtering that we did as part of the Boltzsched pipeline that was released is that we look at the confidence that the model has in the structure. Now, unfortunately, going to your question of predicting affinity, unfortunately, confidence is not a very good predictor of affinity. And so one of the things that we've actually done a ton of progress since we released Bulstian and we have some new results that we are going to announce soon is the ability to get much better heat rates when instead of you know trying to rely on confidence of the model we are actually directly trying to predict the affinity of that interaction okay just backing up a minute so your diffusion model actually predicts not only the protein sequence but also the folding of it exactly and actually kind of the way one of the big different things that we did compared to other models in the space and you know there were some papers that already kind of done this before but we really scaled it up was, you know, basically somewhat merging kind of the structure prediction and the sequence prediction into almost the same task.
52:55And so the way that Boltzgen works is that you are basically the only thing that you're doing is predicting the structure. So the only sort of supervision is we give you a supervision on the structure, but because the structure is atomic and, you know, the different amino acids have a different atomic composition, basically from the way that you place the atoms. We also understand not only kind of the structure that you wanted, but also the identity of the amino acid that, you know, the models believed was there. And so we've basically, instead of, you know, having these two supervision signals, you know, one discrete, one continuous that somewhat, you know, don't interact well together.
53:36we sort of like build kind of like an encoding of, you know, sequences in structures that allows us to basically use exactly the same supervision signal that we were using to both stew that, you know, you know, largely similar to what AlphaVault 3 proposed, which is very scalable. And we can use that to design new proteins. Interesting. Maybe a quick shout out to Hannes Stark on our team who like did all this work. Yeah. Yeah, that was a really cool idea. I mean, like looking at the paper and there's this like encoding or you just add a bunch of, I guess, kind of atoms, which can be anything. And then they get sort of rearranged and then basically plopped on top of each other.
54:21And then that encodes what the amino acid is. And there's sort of like a unique way of doing this. That was like such a really, such a cool, fun idea. I think that idea had existed before. Yeah, there were a couple of papers that proposed this and Anas really took it to the large scale. In the paper, a lot of the paper for Boltzgen is dedicated to actually the validation of the model. In my opinion, all the people we basically talk about feel that this sort of like in the wet lab or whatever the appropriate, you know, sort of like in real world validation is the whole problem or not the whole problem, but a big giant part of the problem.
55:00So can you talk a little bit about the highlights from there that really, because to me, the results are impressive, both from the perspective of the, you know, the model and also just the effort that went into the validation by a large team. First of all, I think I should start saying is that both when we were at MIT and Thomas Iacolas and Regina Barzillais' lab, as well as at Boltz, we are not a biolab and we are not a therapeutic company. And so to some extent, we were forced to look outside of our group, our team to do the experimental validation. One of the things that really honors in the team Pioneer was the idea, okay, can we go not only to maybe a specific group and trying to find a specific system and maybe overfit a bit to that system and trying to validate, but how can we test these models across a very wide variety of different settings so that anyone in the field and printing design is such a kind of wide, task with all sorts of different applications from therapeutic to, you know, biosensors and many others that, you know, so can we get a validation that is kind of goes across many different tasks?
56:27And so he basically put together, you know, I think it was something like, you know, 25 different, you know, academic and industry labs that committed to, you know, testing some of the designs from the model and some of this testing is still ongoing and, you know, giving results kind of back to us in exchange for, you know, hopefully getting some, you know, new great sequences for their task. And he was able to, you know, coordinate this, you know, very wide set of, you know, scientists. And already in the paper, I think we shared results from, I think, eight to 10 different labs kind of showing results from designing peptides, designing to target ordered proteins, peptide targeting disorder proteins, which are results of designing proteins that bind to small molecules, which are results of designing nanobodies and across a wide variety of different targets.
57:31And so that sort of gave to the paper a lot of validation and to the model, a lot of validation that was kind of wide. And so those would be therapeutics for those animals or are they relevant to humans as well? They're relevant to humans as well. Obviously, you need to do some work into, quote unquote, humanizing them, making sure that they have the right characteristics so they're not toxic to humans and so on. There are some approved medicine in the market that are antibodies. There's a general pattern, I think, in trying to design things that are smaller. It's easier to manufacture. At the same time, that comes with potentially other challenges, maybe a little bit less selectivity than if you have something that has more hands.
58:16But there's this big desire to try to design mini proteins, nanobodies, small peptides that are just great drug modalities. Okay. I think we were left off, we were talking about validation in the lab. And I was very excited about seeing like all the diverse validations that you've done. Can you go into some more detail about them? Yeah. Specific ones. Yeah. The nanobody one, I think we did, what was it? 15 targets? Is that correct? 14. 14 targets. Testing. So we, typically the way this works is like we make a lot of designs, right? on the order of tens of thousands. And then we rang them and we picked the top N.
59:03In this case, N was 15 for each target. And then we measure the success rates, both on how many targets we were able to get a binder for, and then also more generally, out of all of the binders that we designed, how many actually proved to be good binders. Some of the other ones I think involved, we had a cool one where there was a small molecule or design a protein that binds to it, that has a lot of interesting applications. For example, Gabri mentioned biosensing and things like that, which is pretty cool. We had a disordered protein, I think you mentioned also. And yeah, I think maybe those were some of the highlights.
59:43Yeah, so I would say that the way that we structure some of those validations was on the one end, we have validations across a whole set of different problems that the biologists that we were working with came to us with. So we were trying to, for example, in some of the experiments, design peptides with a target RACC, which is a target that is involved in metabolism. We had a number of other applications where we were trying to design peptides or other modalities against some other therapeutic relevant targets. We designed some proteins to bind small molecules. And then some of the other testing that we did was really trying to get a more broader sense of how does the model work, especially when tested on somewhat generalization.
1:00:37So one of the things that we found with the field was that a lot of the validation, especially outside of the validation that was done on specific problems, was done on targets that have a lot of known interactions in the training data. And so it's always a bit hard to understand how much are these models really just regurgitating what they've seen or trying to imitate what they've seen in the training data versus really being able to design new proteins. And so one of the experiments that we did was to take nine targets from the PDB, filtering to things where there is no known interaction in the PDB.
1:01:23So basically the model has never seen kind of this particular protein bound or a similar protein bound to another protein. And so there is no way that the model from its training set can sort of like say, okay, I'm just going to kind of tweak something and just imitate this particular kind of interaction. And so we took those nine proteins, we worked with adaptive CRO and basically tested 15 mini proteins and 15 nanobodies against each one of them. And the very cool thing that we saw was that on two-thirds of those targets, we were able to, from these 15 designs, get nanomolar binders. Nanomolar, roughly speaking, is just a measure of how strongly the interaction is.
1:02:11Roughly speaking, a nanomolar binder is approximately the kind of strength of binding that you need for a therapeutic. Yeah, so maybe switching directions a bit. Bolt's Lab was just announced this week, or was it last week? Yeah. This is like your first, I guess, product, if you want to call it that. Can you talk about what Bolt's Lab is and what you hope that people take away from this? Yeah. You know, as we mentioned, like I think at the very beginning is the goal with the product has been to, you know, address what the models don't on their own. And there's largely sort of two categories there.
1:02:57I'll split it in three. The first one, it's one thing to predict a single interaction, for example, like a single structure. It's another to very effectively search a space, a design space, to produce something of value. What we found, like sort of building up this product, is that there's a lot of steps involved in that. We certainly need to accompany the user through. One of those steps, for example, is the creation of the target itself. How do we make sure the model has a good enough understanding of the target so we can design something? And there's all sorts of tricks that you can do to improve a particular structure prediction.
1:03:36And so that's sort of the first stage. And then there's this stage of designing and searching the space efficiently. For something like Bolt's Gen, for example, you design many things and then you rank them. For example, for a small molecule, the process is a little bit more complicated. We actually need to also make sure that the molecules are synthesizable. And so the way we do that is that, you know, we have a generative model that learns to use like appropriate building blocks such that, you know, it can design within a space that we know is like synthesizable. And so there's like, you know, this whole pipeline really of different models involved in being able to design a molecule.
1:04:12And so that's been sort of like the first thing. We call them agents. We have a protein agent and we have a small molecule design agents. And that's really like at the core of like what powers, you know, the pulse lab platform. So these agents, are they like a language model wrapper or they're just like your models and you're just calling them agents because they sort of perform a function on behalf of those? They're more like a recipe, if you wish. And I think we use that term sort of because of the complex pipelining and automation that goes into all this plumbing. So that's the first part of the product.
1:04:46The second part is the infrastructure. You know, we need to be able to do this at very large scale for any one, you know, group that's doing a design campaign. Let's say you're designing, you know, I'd say 100 ,000 possible candidates, right, to find the good one. That is a very large amount of compute. You know, for small molecules, it's on the order of like a few seconds per design. For proteins, it can be a bit longer. And so, you know, ideally you want to do that in parallel, otherwise it's going to take you weeks. And so we've put a lot of effort into our ability to have a GPU fleet that allows any one user to be able to do this large parallel search.
1:05:24So you're amortizing the cost over your users, basically. Exactly. And to some degree, whether you use 10 ,000 GPUs for a minute is the same cost as using one GPUs for God knows how long. So you might as well try to parallelize if you can. So a lot of work has gone into that, making it very robust so that we can have a lot of people on the platform doing that at the same time. And the third one is the interface. And the interface comes in two shapes. One is in form of an API, and that's really suited for companies that want to integrate these pipelines, these agents directly in existing workflows that they have, or existing user interfaces that they have.
1:06:05and we're already partnering with a few distributors that are going to integrate our API. And then the second part is the user interface. We've put a lot of thoughts also into that. And this is when I mentioned earlier this idea of broadening the audience. That's kind of what the user interface is about. And we've built a lot of interesting features in it, for example, for collaboration. When you have multiple medicinal chemists that are going through the results and trying to pick out what are the molecules that we're going to go and test in the lab. it's powerful for them to be able to, for example, each provide their own ranking and then do consensus building.
1:06:39So there's a lot of features around launching this large job, but also around collaborating on analyzing the results that we try to solve with that part of the platform. So Bolt's Lab is a combination of these three objectives into one cohesive platform. Who is this accessible to? Everyone. You do need to request access today. We're still ramping up the usage, but anyone can request access. If you are an academic in particular, we provide a fair amount of free credit so you can play with the platform. If you are a startup or a biotech, you may also reach out and we'll typically actually hop on a call just to understand what you're trying to do and also provide a lot of free credit to get started.
1:07:20And of course, also with larger companies, we can deploy this platform in a more secure environment. And so that's more like custom deals that we make with the partners. And that's sort of the ethos of Bolt. I think this idea of servicing everyone and not necessarily going after just the really large enterprises. And that starts from the open source, but it's also a key design principle of the product itself. One thing I was thinking about with regards to infrastructure, in the LLM space, the cost of a token has gone down by a factor of 1 ,000 or so over the last three years. Is it possible that you can exploit economies of scale and infrastructure that you can make it cheaper to run these things yourself than for any person to roll their own system?
1:08:08100%. We're already there. Running bolts on our platform, especially on a large screen, is considerably cheaper than it would probably take anyone to put the open source model out there and run it. On top of the infrastructure, one of the things that we've been working on is accelerating the models. So our small molecule screening pipeline is 10x faster on Bolt's lab than it is in the open source. And that's also part of building a product, something that scales really well. And we really wanted to get to a point where we could keep prices very low in a way that it would be a no-brainer to use Bolt's through our platform.
1:08:52How do you think about validation of your agentic systems. Because as you were saying earlier, AlphaFull-style models are really good at, let's say, monomeric proteins where you have co-evolution data. But now suddenly the whole point of this is to design something which doesn't have co-evolution data, something which is really novel. So now you're basically leaving the domain that you thought was, you know, that you know you were good at. So how do you validate that. Yeah, I like every complete, but there's obviously a ton of computational metrics that we rely on, but those only take you so far.
1:09:30You really got to go to the lab and test, okay, with this method A and this method B, how much better are we? How much better is my hit rate? How stronger are my binders? Also, it's not just about hit rate, it's also about how good the binders are. And there's really like nowhere around that. I think we've really ramped up the amount of experimental validation that we do so that we really track progress as scientifically sound as possible. I don't know if there's anything. Yeah, no, I think one thing that is unique about us and maybe companies like us is that because we're not working on maybe a couple of therapeutic pipelines where our validation would be focused on those, when we do an experimental validation, we try to test it across tens of targets.
1:10:20And so that on the one end, we can get a much more statistically significant result and really allows us to make progress from the methodological side without being steered by overfitting on any one particular system. And of course, we choose, you know, we always try to choose targets and problems are sort of like at the frontier of what's possible today. So, you know, you don't want something too easy. You don't want something too hard. Otherwise, you're not going to see progress. And so, you know, this is a somewhat evolving set of targets. We talked earlier about the targets that we looked at with Boltran.
1:10:57Now we are even trying kind of, you know, even harder targets, both for small molecule and proteins. And so we try to keep ourselves on the boundary of what's possible. So do you have like infrastructure or this is like you just have a lot of different partnerships with academic labs and you're just going to keep pushing on these and driving these? We do partially this through academic labs. More and more we do this through CROs, just because of, you know, to some extent, it's also we need kind of replicability, often kind of going after the same targets multiple times and to see the progress from one month to the next.
1:11:33And speed. And speed of execution, yeah. So what happens if you start getting a bunch of really strong biters against therapeutic targets? What do you do? Release them. You mean release them in open source? Yeah, I mean, when we say we have no interest in making drugs, we're serious. When it was with the academic labs, basically they keep it, they do whatever they want with it. And with the CROs so far, we've been very releasing them. I will also say, and I think this has been a bit of the issue that I have with some of the things that have been said in the field is when we say that we design new proteins or we say that we design new molecules, go and bind these particular targets, we should be very clear, these are not drugs.
1:12:25These are not things that are ready to be put into a human. And there is still a lot of development that goes with it. And so this is kind of to us, we see ourselves as building tools for scientists. At the end of the day, it really relies on the scientists having a great therapeutic apoptosis and then pushing through all the stages of development. And we try to build tools that can accompany them in that journey. It's not like a magic box where you can just turn it and get... You get FDA-approved drugs. FDA-approved drugs. on fdm drugs um yeah but actually that brings up an interesting question that i have i've been wondering about is do you guys see yourself staying in this for lack of a better way of saying it layer or do you think that you'll start to either on the physical sense looking at different layers of the virtual cell so to speak or also you know so there's like the development process that goes, you know, sort of like design, preclinical, clinical approval, and thinking about improving the performance throughout that process based on the designs?
1:13:43Is that a direction that you guys are pushing? Yeah. So one of the things, as Jeremy said, you know, we are not a therapeutic company and we want to kind of stay not to be a therapeutic company, always be at the service of, you the different companies, including companies that we serve. And that to some extent does mean that we need to try to go deeper and deeper in getting these models better and better. One of the things that we are doing across many other in the field is now that we are really starting to be good both for small molecule and for proteins to design kind of binders, design relatively tight binders, it's starting to look at all these other properties, you know, called developabilities or atme that, you know, we care about when developing a drug and trying, can we design them from Gecko?
1:14:37The thing about those properties in some of them, you know, you need to, you know, start having an understanding of the cell. And so that's on the one end kind of why we need that understanding. But also, you know, the way that we also think about all different and complex diseases is that these models and these tools that we're building have a good understanding of kind of, you know, biomolecular interactions and kind of their interactions. Now, at the same time, every disease is often kind of unique and every therapeutic hypothesis is unique. And so you maybe want to have something that needs to hit the particular, you know, let's say target in a virus in a particular way, but you don't maybe know exactly what way you want to do.
1:15:21And so maybe in the first set of designs, you're going to try to target different epitopes in different ways. And then you're going to test them in the lab, maybe directly in vivo. And you're going to see which ones work and which ones don't. And so then you need to bring those results back into the models. And then the models can start to have a more wider understanding, you know, not just of the biophysical of the antibodies interacting with that target, but also how that is shaped within the entire cell. and so first of all you know that means on the one end that we need you know kind of these loops and this is also partially how we we design the platform to be but that also means that we also need to start understanding more and more kind of higher level things and you know i wouldn't say that we're working in any way on like a virtual cell like others are but we're definitely thinking kind of very deeply about kind of you know how does you know kind of the way that we target certain proteins interfere, interact with pathways that are existing in the cell.
1:16:25One question that has come up is you talk a lot about user interface and so on. And I think this is really important. But my experience with dealing with medicinal chemists, when you give them machine learning models, is they are the most superstitious, skeptical, pseudo-religious people I've ever talked to when it comes to doing science. Sorry for the medicinal chemists listening. They're amazing. I've worked with some spectacular medicinal chemists who just pull magic out of their hat again and again, and I have no idea how they do it. But when you bring them a machine learning model, it is sometimes quite tricky to get them to deal with it.
1:17:03How has your interaction been with this, and how have you thought about building Bolt's Lab to work with the skeptics? One of the great value unlocks for us and for our product has been when we brought to the team, Mison Chemist, his name is Jeffrey. So I think kind of like on the one end, you know, day one, you know, he obviously had a lot of opinions on kind of a lot of the ways that we should change, you know, both kind of the way that the agents work, the way that the platform worked. But it's been really amazing kind of, you know, once also we started kind of shaping kind of the platform in a better way with this feedback, how we went from, to some extent, fair skepticism to him actually using a lot more compute than any of our computational folks in the team.
1:17:55At times that he's running, he has all these sort of hypotheses. Okay, maybe I can hit this protein this particular way. I can hit it in that way, actually, Let me look at for this particular molecular space. Let me try to optimize for these particular interactions. So he ends up running several screens in parallel, using hundreds of GPUs on his own. And so this has been pretty incredible to see kind of how maybe the way that I was more thinking about a problem, which is, okay, you're just trying to design a binder, a small molecule to a particular protein. The way that he thinks about it is much more deeply and trying all these different things, these different hypotheses.
1:18:38And then, you know, once he gets the results from the model, he doesn't just, you know, take the top 15, but it really kind of looks over and, you know, kind of tries to understand, you know, the different things. And then when we select, you know, maybe some designs to bring forth, you know, he has, you know, something where, you know, both the models understand that something's good, but all himself as well. And that's why we also built kind of the platform to be an interface for, you know, this kind of, this kind of chemists and, you know, also like a collaborative experience. I think at the end of the day, like, you know, for people to be convinced, you have to show them something that they didn't think was possible.
1:19:15And until you have that aha moment, you know, I think the skepticism will remain. But then when, you know, every once in a while, I think there's like a result that like really surprises people. And then it's like, oh, wow, okay, this is actually, I can do something with this. So you just get in their hands, have them try it out and they'll be convinced. Yeah, or maybe once the lab results come back. Or maybe one of their colleagues is convinced. I think it takes going to the lab at some point. There's no avoiding that. As beautiful as the platform can be, as nice as the molecules might look that the model predicted, I think what really convinces people is hits.
1:19:54Yeah, you see the results. Exactly. Cool. Thank you for taking the time to chat with us. Is there anything that you would like your audience to know? First of all, we're just getting started, continuing to build a team. Definitely always looking for great folks, both on the software side, machine learning side, but also scientists to join the team and help us shape. On the infrastructure side too? Indeed. If you think that if you want a new challenge, because this is not just next token prediction, this is really a new engineering challenge. Exactly. No matter how much experience you have with biologists and chemistry, if you want to come help us shape what biology and chemistry hopefully will look like in five, ten years, we'd love to hear from you.
1:20:52And so go to bolso.bio and come to our team. Cool. thank you awesome thank you so much thank you
From the publisher
This podcast features Gabriele Corso and Jeremy Wohlwend, co-founders of Boltz and authors of the Boltz Manifesto, discussing the rapid evolution of structural biology models from AlphaFold to their own open-source suite, Boltz-1 and Boltz-2. The central thesis is that while single-chain protein structure prediction is largely βsolvedβ through evolutionary hints, the next frontier lies in modeling complex interactions (protein-ligand, protein-protein) and generative protein design, which Boltz aims to democratize via open-source foundations and scalable infrastructure.
Full Video Pod
On YouTube!
Timestamps
* 00:00 Introduction to Benchmarking and the βSolvedβ Protein Problem
* 06:48 Evolutionary Hints and Co-evolution in Structure Prediction
* 10:00 The Importance of Protein Function and Disease States
* 15:31 Transitioning from AlphaFold 2 to AlphaFold 3 Capabilities
* 19:48 Generative Modeling vs. Regression in Structural Biology
* 25:00 The βBitter Lessonβ and Specialized AI Architectures
* 29:14 Development Anecdotes: Training Boltz-1 on a Budget
* 32:00 Validation Strategies and the Protein Data Bank (PDB)
* 37:26 The Mission of Boltz: Democratizing Access and Open Source
* 41:43 Building a Self-Sustaining Research Community
* 44:40 Boltz-2 Advancements: Affinity Prediction and Design
* 51:03 BoltzGen: Merging Structure and Sequence Prediction
* 55:18 Large-Scale Wet Lab Validation Results
* 01:02:44 Boltz Lab Product Launch: Agents and Infrastructure
* 01:13:06 Future Directions: Developpability and the βVirtual Cellβ
* 01:17:35 Interacting with Skeptical Medicinal Chemists
Key Summary
Evolution of Structure Prediction & Evolutionary Hints
* Co-evolutionary Landscapes: The speakers explain that breakthrough progress in single-chain protein prediction relied on decoding evolutionary correlations where mutations in one position necessitate mutations in another to conserve 3D structure.
* Structure vs. Folding: They differentiate between structure prediction (getting the final answer) and folding (the kinetic process of reaching that state), noting that the field is still quite poor at modeling the latter.
* Physics vs. Statistics: RJ posits that while models use evolutionary statistics to find the right βvalleyβ in the energy landscape, they likely possess a βlight understandingβ of physics to refine the local minimum.
The Shift to Generative Architectures
* Generative Modeling: A key leap in AlphaFold 3 and Boltz-1 was moving from regression (predicting one static coordinate) to a generative diffusion approach that samples from a posterior distribution.
* Handling Uncertainty: This shift allows models to represent multiple conformational states and avoid the βaveragingβ effect seen in regression models when the ground truth is ambiguous.
* Specialized Architectures: Despite the βbitter lessonβ of general-purpose transformers, the speakers argue that equivariant architectures remain vastly superior for biological data due to the inherent 3D geometric constraints of molecules.
Boltz-2 and Generative Protein Design
* Unified Encoding: Boltz-2 (and BoltzGen) treats structure and sequence prediction as a single task by encoding amino acid identities into the atomic composition of the predicted structure.
* Design Specifics: Instead of a sequence, users feed the model blank tokens and a high-level βspecβ (e.g., an antibody framework), and the model decodes both the 3D structure and the corresponding amino acids.
* Affinity Prediction: While model confidence is a common metric, Boltz-2 focuses on affinity predictionβquantifying exactly how tightly a designed binder will stick to its target.
Real-World Validation and Productization
* Generalized Validation: To prove the model isnβt just βregurgitatingβ known data, Boltz tested its designs on 9 targets with zero known interactions in the PDB, achieving nanomolar binders for two-thirds of them.
* Boltz Lab Infrastructure: The newly launched Boltz Lab platform provides βagentsβ for protein and small molecule design, optimized to run 10x faster than open-source versions through proprietary GPU kernels.
* Human-in-the-Loop: The platform is designed to convert skeptical medicinal chemists by allowing them to run parallel screens and use their intuition to filter model outputs.
Transcript
RJ [00:05:35]: But the goal remains to, like, you know, really challenge the models, like, how well do these models generalize? And, you know, weβve seen in some of the latest CASP competitions, like, while weβve become really, really good at proteins, especially monomeric proteins, you know, other modalities still remain pretty difficult. So itβs really essential, you know, in the field that there are, like, these efforts to gather, you know, benchmarks that are challenging. So it keeps us in line, you know, about what the models can do or not.
Gabriel [00:06:26]: Yeah, itβs interesting you say that, like, in some sense, CASP, you know, at CASP 14, a problem was solved and, like, pretty comprehensively, right? But at the same time, it was really only the beginning. So you can say, like, what was the specific problem you would argue was solved? And then, like, you know, what is remaining, which is probably quite open.
RJ [00:06:48]: I think weβll steer away from the term solved, because we have many friends in the community who get pretty upset at that word. And I think, you know, fairly so. But the problem that was, you know, that a lot of progress was made on was the ability to predict the structure of single chain proteins. So proteins can, like, be composed of many chains. And single chain proteins are, you know, just a single sequence of amino acids. And one of the reasons that weβve been able to make such progress is also because we take a lot of hints from evolution. So the way the models work is that, you know, they sort of decode a lot of hints. That comes from evolutionary landscapes. So if you have, like, you know, some protein in an animal, and you go find the similar protein across, like, you know, different organisms, you might find different mutations in them. And as it turns out, if you take a lot of the sequences together, and you analyze them, you see that some positions in the sequence tend to evolve at the same time as other positions in the sequence, sort of this, like, correlation between different positions. And it turns out that that is typically a hint that these two positions are close in three dimension. So part of the, you know, part of the breakthrough has been, like, our ability to also decode that very, very effectively. But what it implies also is that in absence of that co-evolutionary landscape, the models donβt quite perform as well. And so, you know, I think when that information is available, maybe one could say, you know, the problem is, like, somewhat solved. From the perspective of structure prediction, when it isnβt, itβs much more challenging. And I think itβs also worth also differentiating the, sometimes we confound a little bit, structure prediction and folding. Folding is the more complex process of actually understanding, like, how it goes from, like, this disordered state into, like, a structured, like, state. And that I donβt think weβve made that much progress on. But the idea of, like, yeah, going straight to the answer, weβve become pretty good at.
Brandon [00:08:49]: So thereβs this protein that is, like, just a long chain and it folds up. Yeah. And so weβre good at getting from that long chain in whatever form it was originally to the thing. But we donβt know how it necessarily gets to that state. And there might be intermediate states that itβs in sometimes that weβre not aware of.
RJ [00:09:10]: Thatβs right. And that relates also to, like, you know, our general ability to model, like, the different, you know, proteins are not static. They move, they take different shapes based on their energy states. And I think we are, also not that good at understanding the different states that the protein can be in and at what frequency, what probability. So I think the two problems are quite related in some ways. Still a lot to solve. But I think it was very surprising at the time, you know, that even with these evolutionary hints that we were able to, you know, to make such dramatic progress.
Brandon [00:09:45]: So I want to ask, why does the intermediate states matter? But first, I kind of want to understand, why do we care? What proteins are shaped like?
Gabriel [00:09:54]: Yeah, I mean, the proteins are kind of the machines of our body. You know, the way that all the processes that we have in our cells, you know, work is typically through proteins, sometimes other molecules, sort of intermediate interactions. And through that interactions, we have all sorts of cell functions. And so when we try to understand, you know, a lot of biology, how our body works, how disease work. So we often try to boil it down to, okay, what is going right in case of, you know, our normal biological function and what is going wrong in case of the disease state. And we boil it down to kind of, you know, proteins and kind of other molecules and their interaction. And so when we try predicting the structure of proteins, itβs critical to, you know, have an understanding of kind of those interactions. Itβs a bit like seeing the difference between... Having kind of a list of parts that you would put it in a car and seeing kind of the car in its final form, you know, seeing the car really helps you understand what it does. On the other hand, kind of going to your question of, you know, why do we care about, you know, how the protein falls or, you know, how the car is made to some extent is that, you know, sometimes when something goes wrong, you know, there are, you know, cases of, you know, proteins misfolding. In some diseases and so on, if we donβt understand this folding process, we donβt really know how to intervene.
RJ [00:11:30]: Thereβs this nice line in the, I think itβs in the Alpha Fold 2 manuscript, where they sort of discuss also like why we even hopeful that we can target the problem in the first place. And then thereβs this notion that like, well, four proteins that fold. The folding process is almost instantaneous, which is a strong, like, you know, signal that like, yeah, like we should, we might be... able to predict that this very like constrained thing that, that the protein does so quickly. And of course thatβs not the case for, you know, for, for all proteins. And thereβs a lot of like really interesting mechanisms in the cells, but yeah, I remember reading that and thought, yeah, thatβs somewhat of an insightful point.
Gabriel [00:12:10]: I think one of the interesting things about the protein folding problem is that it used to be actually studied. And part of the reason why people thought it was impossible, it used to be studied as kind of like a classical example. Of like an MP problem. Uh, like there are so many different, you know, type of, you know, shapes that, you know, this amino acid could take. And so, this grows combinatorially with the size of the sequence. And so there used to be kind of a lot of actually kind of more theoretical computer science thinking about and studying protein folding as an MP problem. And so it was very surprising also from that perspective, kind of seeing. Machine learning so clear, there is some, you know, signal in those sequences, through evolution, but also through kind of other things that, you know, us as humans, weβre probably not really able to, uh, to understand, but that is, models Iβve, Iβve learned.
Brandon [00:13:07]: And so Andrew White, we were talking to him a few weeks ago and he said that he was following the development of this and that there were actually ASICs that were developed just to solve this problem. So, again, that there were. There were many, many, many millions of computational hours spent trying to solve this problem before AlphaFold. And just to be clear, one thing that you mentioned was that thereβs this kind of co-evolution of mutations and that you see this again and again in different species. So explain why does that give us a good hint that theyβre close by to each other? Yeah.
RJ [00:13:41]: Um, like think of it this way that, you know, if I have, you know, some amino acid that mutates, itβs going to impact everything around it. Right. In three dimensions. And so itβs almost like the protein through several, probably random mutations and evolution, like, you know, ends up sort of figuring out that this other amino acid needs to change as well for the structure to be conserved. Uh, so this whole principle is that the structure is probably largely conserved, you know, because thereβs this function associated with it. And so itβs really sort of like different positions compensating for, for each other. I see.
Brandon [00:14:17]: Those hints in aggregate give us a lot. Yeah. So you can start to look at what kinds of information about what is close to each other, and then you can start to look at what kinds of folds are possible given the structure and then what is the end state.
RJ [00:14:30]: And therefore you can make a lot of inferences about what the actual total shape is. Yeah, thatβs right. Itβs almost like, you know, you have this big, like three dimensional Valley, you know, where youβre sort of trying to find like these like low energy states and thereβs so much to search through. Thatβs almost overwhelming. But these hints, they sort of maybe put you in. An area of the space thatβs already like, kind of close to the solution, maybe not quite there yet. And, and thereβs always this question of like, how much physics are these models learning, you know, versus like, just pure like statistics. And like, I think one of the thing, at least I believe is that once youβre in that sort of approximate area of the solution space, then the models have like some understanding, you know, of how to get you to like, you know, the lower energy, uh, low energy state. And so maybe you have some, some light understanding. Of physics, but maybe not quite enough, you know, to know how to like navigate the whole space. Right. Okay.
Brandon [00:15:25]: So we need to give it these hints to kind of get into the right Valley and then it finds the, the minimum or something. Yeah.
Gabriel [00:15:31]: One interesting explanation about our awful free works that I think itβs quite insightful, of course, doesnβt cover kind of the entirety of, of what awful does that is, um, theyβre going to borrow from, uh, Sergio Chinico for MIT. So he sees kind of awful. Then the interesting thing about awful is God. This very peculiar architecture that we have seen, you know, used, and this architecture operates on this, you know, pairwise context between amino acids. And so the idea is that probably the MSA gives you this first hint about what potential amino acids are close to each other. MSA is most multiple sequence alignment. Exactly. Yeah. Exactly. This evolutionary information. Yeah. And, you know, from this evolutionary information about potential contacts, then is almost as if the model is. of running some kind of, you know, diastro algorithm where itβs sort of decoding, okay, these have to be closed. Okay. Then if these are closed and this is connected to this, then this has to be somewhat closed. And so you decode this, that becomes basically a pairwise kind of distance matrix. And then from this rough pairwise distance matrix, you decode kind of the
Brandon [00:16:42]: actual potential structure. Interesting. So thereβs kind of two different things going on in the kind of coarse grain and then the fine grain optimizations. Interesting. Yeah. Very cool.
Gabriel [00:16:53]: Yeah. You mentioned AlphaFold3. So maybe we have a good time to move on to that. So yeah, AlphaFold2 came out and it was like, I think fairly groundbreaking for this field. Everyone got very excited. A few years later, AlphaFold3 came out and maybe for some more history, like what were the advancements in AlphaFold3? And then I think maybe weβll, after that, weβll talk a bit about the sort of how it connects to Bolt. But anyway. Yeah. So after AlphaFold2 came out, you know, Jeremy and I got into the field and with many others, you know, the clear problem that, you know, was, you know, obvious after that was, okay, now we can do individual chains. Can we do interactions, interaction, different proteins, proteins with small molecules, proteins with other molecules. And so. So why are interactions important? Interactions are important because to some extent thatβs kind of the way that, you know, these machines, you know, these proteins have a function, you know, the function comes by the way that they interact with other proteins and other molecules. Actually, in the first place, you know, the individual machines are often, as Jeremy was mentioning, not made of a single chain, but theyβre made of the multiple chains. And then these multiple chains interact with other molecules to give the function to those. And on the other hand, you know, when we try to intervene of these interactions, think about like a disease, think about like a, a biosensor or many other ways we are trying to design the molecules or proteins that interact in a particular way with what we would call a target protein or target. You know, this problem after AlphaVol2, you know, became clear, kind of one of the biggest problems in the field to, to solve many groups, including kind of ours and others, you know, started making some kind of contributions to this problem of trying to model these interactions. And AlphaVol3 was, you know, was a significant advancement on the problem of modeling interactions. And one of the interesting thing that they were able to do while, you know, some of the rest of the field that really tried to try to model different interactions separately, you know, how protein interacts with small molecules, how protein interacts with other proteins, how RNA or DNA have their structure, they put everything together and, you know, train very large models with a lot of advances, including kind of changing kind of systems. Some of the key architectural choices and managed to get a single model that was able to set this new state-of-the-art performance across all of these different kind of modalities, whether that was protein, small molecules is critical to developing kind of new drugs, protein, protein, understanding, you know, interactions of, you know, proteins with RNA and DNAs and so on.
Brandon [00:19:39]: Just to satisfy the AI engineers in the audience, what were some of the key architectural and data, data changes that made that possible?
Gabriel [00:19:48]: Yeah, so one critical one that was not necessarily just unique to AlphaFold3, but there were actually a few other teams, including ours in the field that proposed this, was moving from, you know, modeling structure prediction as a regression problem. So where there is a single answer and youβre trying to shoot for that answer to a generative modeling problem where you have a posterior distribution of possible structures and youβre trying to sample this distribution. And this achieves two things. One is it starts to allow us to try to model more dynamic systems. As we said, you know, some of these structures can actually take multiple structures. And so, you know, you can now model that, you know, through kind of modeling the entire distribution. But on the second hand, from more kind of core modeling questions, when you move from a regression problem to a generative modeling problem, you are really tackling the way that you think about uncertainty in the model in a different way. So if you think about, you know, Iβm undecided between different answers, whatβs going to happen in a regression model is that, you know, Iβm going to try to make an average of those different kind of answers that I had in mind. When you have a generative model, what youβre going to do is, you know, sample all these different answers and then maybe use separate models to analyze those different answers and pick out the best. So that was kind of one of the critical improvement. The other improvement is that they significantly simplified, to some extent, the architecture, especially of the final model that takes kind of those pairwise representations and turns them into an actual structure. And that now looks a lot more like a more traditional transformer than, you know, like a very specialized equivariant architecture that it was in AlphaFold3.
Brandon [00:21:41]: So this is a bitter lesson, a little bit.
Gabriel [00:21:45]: There is some aspect of a bitter lesson, but the interesting thing is that itβs very far from, you know, being like a simple transformer. This field is one of the, I argue, very few fields in applied machine learning where we still have kind of architecture that are very specialized. And, you know, there are many people that have tried to replace these architectures with, you know, simple transformers. And, you know, there is a lot of debate in the field, but I think kind of that most of the consensus is that, you know, the performance... that we get from the specialized architecture is vastly superior than what we get through a single transformer. Another interesting thing that I think on the staying on the modeling machine learning side, which I think itβs somewhat counterintuitive seeing some of the other kind of fields and applications is that scaling hasnβt really worked kind of the same in this field. Now, you know, models like AlphaFold2 and AlphaFold3 are, you know, still very large models.
RJ [00:29:14]: in a place, I think, where we had, you know, some experience working in, you know, with the data and working with this type of models. And I think that put us already in like a good place to, you know, to produce it quickly. And, you know, and I would even say, like, I think we could have done it quicker. The problem was like, for a while, we didnβt really have the compute. And so we couldnβt really train the model. And actually, we only trained the big model once. Thatβs how much compute we had. We could only train it once. And so like, while the model was training, we were like, finding bugs left and right. A lot of them that I wrote. And like, I remember like, I was like, sort of like, you know, doing like, surgery in the middle, like stopping the run, making the fix, like relaunching. And yeah, we never actually went back to the start. We just like kept training it with like the bug fixes along the way, which was impossible to reproduce now. Yeah, yeah, no, that model is like, has gone through such a curriculum that, you know, learned some weird stuff. But yeah, somehow by miracle, it worked out.
Gabriel [00:30:13]: The other funny thing is that the way that we were training, most of that model was through a cluster from the Department of Energy. But thatβs sort of like a shared cluster that many groups use. And so we were basically training the model for two days, and then it would go back to the queue and stay a week in the queue. Oh, yeah. And so it was pretty painful. And so we actually kind of towards the end with Evan, the CEO of Genesis, and basically, you know, I was telling him a bit about the project and, you know, kind of telling him about this frustration with the compute. And so luckily, you know, he offered to kind of help. And so we, we got the help from Genesis to, you know, finish up the model. Otherwise, it probably would have taken a couple of extra weeks.
Brandon [00:30:57]: Yeah, yeah.
Brandon [00:31:02]: And then, and then thereβs some progression from there.
Gabriel [00:31:06]: Yeah, so I would say kind of that, both one, but also kind of these other kind of set of models that came around the same time, were kind of approaching were a big leap from, you know, kind of the previous kind of open source models, and, you know, kind of really kind of approaching the level of AlphaVault 3. But I would still say that, you know, even to this day, there are, you know, some... specific instances where AlphaVault 3 works better. I think one common example is antibody antigen prediction, where, you know, AlphaVault 3 still seems to have an edge in many situations. Obviously, these are somewhat different models. They are, you know, you run them, you obtain different results. So itβs, itβs not always the case that one model is better than the other, but kind of in aggregate, we still, especially at the time.
Brandon [00:32:00]: So AlphaVault 3 is, you know, still having a bit of an edge. We should talk about this more when we talk about Boltzgen, but like, how do you know one is, one model is better than the other? Like you, so you, I make a prediction, you make a prediction, like, how do you know?
Gabriel [00:32:11]: Yeah, so easily, you know, the, the great thing about kind of structural prediction and, you know, once weβre going to go into the design space of designing new small molecule, new proteins, this becomes a lot more complex. But a great thing about structural prediction is that a bit like, you know, CASP was doing, basically the way that you can evaluate them is that, you know, you train... You know, you train a model on a structure that was, you know, released across the field up until a certain time. And, you know, one of the things that we didnβt talk about that was really critical in all this development is the PDB, which is the Protein Data Bank. Itβs this common resources, basically common database where every biologist publishes their structures. And so we can, you know, train on, you know, all the structures that were put in the PDB until a certain date. And then... And then we basically look for recent structures, okay, which structures look pretty different from anything that was published before, because we really want to try to understand generalization.
Brandon [00:33:13]: And then on this new structure, we evaluate all these different models. And so you just know when AlphaFold3 was trained, you know, when youβre, you intentionally trained to the same date or something like that. Exactly. Right. Yeah.
Gabriel [00:33:24]: And so this is kind of the way that you can somewhat easily kind of compare these models, obviously, that assumes that, you know, the training. Youβve always been very passionate about validation. I remember like DiffDoc, and then there was like DiffDocL and DocGen. Youβve thought very carefully about this in the past. Like, actually, I think DocGen is like a really funny story that I think, I donβt know if you want to talk about that. Itβs an interesting like... Yeah, I think one of the amazing things about putting things open source is that we get a ton of feedback from the field. And, you know, sometimes we get kind of great feedback of people. Really like... But honestly, most of the times, you know, to be honest, thatβs also maybe the most useful feedback is, you know, people sharing about where it doesnβt work. And so, you know, at the end of the day, itβs critical. And this is also something, you know, across other fields of machine learning. Itβs always critical to set, to do progress in machine learning, set clear benchmarks. And as, you know, you start doing progress of certain benchmarks, then, you know, you need to improve the benchmarks and make them harder and harder. And this is kind of the progression of, you know, how the field operates. And so, you know, the example of DocGen was, you know, we published this initial model called DiffDoc in my first year of PhD, which was sort of like, you know, one of the early models to try to predict kind of interactions between proteins, small molecules, that we bought a year after AlphaFold2 was published. And now, on the one hand, you know, on these benchmarks that we were using at the time, DiffDoc was doing really well, kind of, you know, outperforming kind of some of the traditional physics-based methods. But on the other hand, you know, when we started, you know, kind of giving these tools to kind of many biologists, and one example was that we collaborated with was the group of Nick Polizzi at Harvard. We noticed, started noticing that there was this clear, pattern where four proteins that were very different from the ones that weβre trained on, the models was, was struggling. And so, you know, that seemed clear that, you know, this is probably kind of where we should, you know, put our focus on. And so we first developed, you know, with Nick and his group, a new benchmark, and then, you know, went after and said, okay, what can we change? And kind of about the current architecture to improve this pattern and generalization. And this is the same that, you know, weβre still doing today, you know, kind of, where does the model not work, you know, and then, you know, once we have that benchmark, you know, letβs try to, through everything we, any ideas that we have of the problem.
RJ [00:36:15]: And thereβs a lot of like healthy skepticism in the field, which I think, you know, is, is, is great. And I think, you know, itβs very clear that thereβs a ton of things, the models donβt really work well on, but I think one thing thatβs probably, you know, undeniable is just like the pace of, pace of progress, you know, and how, how much better weβre getting, you know, every year. And so I think if you, you know, if you assume, you know, any constant, you know, rate of progress moving forward, I think things are going to look pretty cool at some point in the future.
Gabriel [00:36:42]: ChatGPT was only three years ago. Yeah, I mean, itβs wild, right?
RJ [00:36:45]: Like, yeah, yeah, yeah, itβs one of those things. Like, youβve been doing this. Being in the field, you donβt see it coming, you know? And like, I think, yeah, hopefully weβll, you know, weβll, weβll continue to have as much progress weβve had the past few years.
Brandon [00:36:55]: So this is maybe an aside, but Iβm really curious, you get this great feedback from the, from the community, right? By being open source. My question is partly like, okay, yeah, if you open source and everyone can copy what you did, but itβs also maybe balancing priorities, right? Where you, like all my customers are saying. I want this, thereβs all these problems with the model. Yeah, yeah. But my customers donβt care, right? So like, how do you, how do you think about that? Yeah.
Gabriel [00:37:26]: So I would say a couple of things. One is, you know, part of our goal with Bolts and, you know, this is also kind of established as kind of the mission of the public benefit company that we started is to democratize the access to these tools. But one of the reasons why we realized that Bolts needed to be a company, it couldnβt just be an academic project is that putting a model on GitHub is definitely not enough to get, you know, chemists and biologists, you know, across, you know, both academia, biotech and pharma to use your model to, in their therapeutic programs. And so a lot of what we think about, you know, at Bolts beyond kind of the, just the models is thinking about all the layers. The layers that come on top of the models to get, you know, from, you know, those models to something that can really enable scientists in the industry. And so that goes, you know, into building kind of the right kind of workflows that take in kind of, for example, the data and try to answer kind of directly that those problems that, you know, the chemists and the biologists are asking, and then also kind of building the infrastructure. And so this to say that, you know, even with models fully open. You know, we see a ton of potential for, you know, products in the space and the critical part about a product is that even, you know, for example, with an open source model, you know, running the model is not free, you know, as we were saying, these are pretty expensive model and especially, and maybe weβll get into this, you know, these days weβre seeing kind of pretty dramatic inference time scaling of these models where, you know, the more you run them, the better the results are. But there, you know, you see. You start getting into a point that compute and compute costs becomes a critical factor. And so putting a lot of work into building the right kind of infrastructure, building the optimizations and so on really allows us to provide, you know, a much better service potentially to the open source models. That to say, you know, even though, you know, with a product, we can provide a much better service. I do still think, and we will continue to put a lot of our models open source because the critical kind of role. I think of open source. Models is, you know, helping kind of the community progress on the research and, you know, from which we, we all benefit. And so, you know, weβll continue to on the one hand, you know, put some of our kind of base models open source so that the field can, can be on top of it. And, you know, as we discussed earlier, we learn a ton from, you know, the way that the field uses and builds on top of our models, but then, you know, try to build a product that gives the best experience possible to scientists. So that, you know, like a chemist or a biologist doesnβt need to, you know, spin off a GPU and, you know, set up, you know, our open source model in a particular way, but can just, you know, a bit like, you know, I, even though I am a computer scientist, machine learning scientist, I donβt necessarily, you know, take a open source LLM and try to kind of spin it off. But, you know, I just maybe open a GPT app or a cloud code and just use it as an amazing product. We kind of want to give the same experience. So this front world.
Brandon [00:40:40]: I heard a good analogy yesterday that a surgeon doesnβt want the hospital to design a scalpel, right?
Brandon [00:40:48]: So just buy the scalpel.
RJ [00:40:50]: You wouldnβt believe like the number of people, even like in my short time, you know, between AlphaFold3 coming out and the end of the PhD, like the number of people that would like reach out just for like us to like run AlphaFold3 for them, you know, or things like that. Just because like, you know, bolts in our case, you know, just because itβs like. Itβs like not that easy, you know, to do that, you know, if youβre not a computational person. And I think like part of the goal here is also that, you know, we continue to obviously build the interface with computational folks, but that, you know, the models are also accessible to like a larger, broader audience. And then that comes from like, you know, good interfaces and stuff like that.
Gabriel [00:41:27]: I think one like really interesting thing about bolts is that with the release of it, you didnβt just release a model, but you created a community. Yeah. Did that community, it grew very quickly. Did that surprise you? And like, what is the evolution of that community and how is that fed into bolts?
RJ [00:41:43]: If you look at its growth, itβs like very much like when we release a new model, itβs like, thereβs a big, big jump, but yeah, itβs, I mean, itβs been great. You know, we have a Slack community that has like thousands of people on it. And itβs actually like self-sustaining now, which is like the really nice part because, you know, itβs, itβs almost overwhelming, I think, you know, to be able to like answer everyoneβs questions and help. Itβs really difficult, you know. The, the few people that we were, but it ended up that like, you know, people would answer each otherβs questions and like, sort of like, you know, help one another. And so the Slack, you know, has been like kind of, yeah, self, self-sustaining and thatβs been, itβs been really cool to see.
RJ [00:42:21]: And, you know, thatβs, thatβs for like the Slack part, but then also obviously on GitHub as well. Weβve had like a nice, nice community. You know, I think we also aspire to be even more active on it, you know, than weβve been in the past six months, which has been like a bit challenging, you know, for us. But. Yeah, the community has been, has been really great and, you know, thereβs a lot of papers also that have come out with like new evolutions on top of bolts and itβs surprised us to some degree because like thereβs a lot of models out there. And I think like, you know, sort of people converging on that was, was really cool. And, you know, I think it speaks also, I think, to the importance of like, you know, when, when you put code out, like to try to put a lot of emphasis and like making it like as easy to use as possible and something we thought a lot about when we released the code base. You know, itβs far from perfect, but, you know.
Brandon [00:43:07]: Do you think that that was one of the factors that caused your community to grow is just the focus on easy to use, make it accessible? I think so.
RJ [00:43:14]: Yeah. And weβve, weβve heard it from a few people over the, over the, over the years now. And, you know, and some people still think it should be a lot nicer and theyβre, and theyβre right. And theyβre right. But yeah, I think it was, you know, at the time, maybe a little bit easier than, than other things.
Gabriel [00:43:29]: The other thing part, I think led to, to the community and to some extent, I think, you know, like the somewhat the trust in the community. Kind of what we, what we put out is the fact that, you know, itβs not really been kind of, you know, one model, but, and maybe weβll talk about it, you know, after Boltz 1, you know, there were maybe another couple of models kind of released, you know, or open source kind of soon after. We kind of continued kind of that open source journey or at least Boltz 2, where we are not only improving kind of structure prediction, but also starting to do affinity predictions, understanding kind of the strength of the interactions between these different models, which is this critical component. critical property that you often want to optimize in discovery programs. And then, you know, more recently also kind of protein design model. And so weβve sort of been building this suite of, of models that come together, interact with one another, where, you know, kind of, there is almost an expectation that, you know, we, we take very at heart of, you know, always having kind of, you know, across kind of the entire suite of different tasks, the best or across the best. model out there so that itβs sort of like our open source tool can be kind of the go-to model for everybody in the, in the industry. I really want to talk about Boltz 2, but before that, one last question in this direction, was there anything about the community which surprised you? Were there any, like, someone was doing something and youβre like, why would you do that? Thatβs crazy. Or thatβs actually genius. And I never would have thought about that.
RJ [00:45:01]: I mean, weβve had many contributions. I think like some of the. Interesting ones, like, I mean, we had, you know, this one individual who like wrote like a complex GPU kernel, you know, for part of the architecture on a piece of, the funny thing is like that piece of the architecture had been there since AlphaFold 2, and I donβt know why it took Boltz for this, you know, for this person to, you know, to decide to do it, but that was like a really great contribution. Weβve had a bunch of others, like, you know, people figuring out like ways to, you know, hack the model to do something. They click peptides, like, you know, thereβs, I donβt know if thereβs any other interesting ones come to mind.
Gabriel [00:45:41]: One cool one, and this was, you know, something that initially was proposed as, you know, as a message in the Slack channel by Tim OβDonnell was basically, he was, you know, there are some cases, especially, for example, we discussed, you know, antibody-antigen interactions where the models donβt necessarily kind of get the right answer. What he noticed is that, you know, the models were somewhat stuck into predicting kind of the antibodies. And so he basically ran the experiments in this model, you can condition, basically, you can give hints. And so he basically gave, you know, random hints to the model, basically, okay, you should bind to this residue, you should bind to the first residue, or you should bind to the 11th residue, or you should bind to the 21st residue, you know, basically every 10 residues scanning the entire antigen.
Brandon [00:46:33]: Residues are the...
Gabriel [00:46:34]: The amino acids. The amino acids, yeah. So the first amino acids. The 11 amino acids, and so on. So itβs sort of like doing a scan, and then, you know, conditioning the model to predict all of them, and then looking at the confidence of the model in each of those cases and taking the top. And so itβs sort of like a very somewhat crude way of doing kind of inference time search. But surprisingly, you know, for antibody-antigen prediction, it actually kind of helped quite a bit. And so thereβs some, you know, interesting ideas that, you know, obviously, as kind of developing the model, you say kind of, you know, wow. This is why would the model, you know, be so dumb. But, you know, itβs very interesting. And that, you know, leads you to also kind of, you know, start thinking about, okay, how do I, can I do this, you know, not with this brute force, but, you know, in a smarter way.
RJ [00:47:22]: And so weβve also done a lot of work on that direction. And that speaks to, like, the, you know, the power of scoring. Weβre seeing that a lot. Iβm sure weβll talk about it more when we talk about BullsGen. But, you know, our ability to, like, take a structure and determine that that structure is, like... Good. You know, like, somewhat accurate. Whether thatβs a single chain or, like, an interaction is a really powerful way of improving, you know, the models. Like, sort of like, you know, if you can sample a ton and you assume that, like, you know, if you sample enough, youβre likely to have, like, you know, the good structure. Then it really just becomes a ranking problem. And, you know, now weβre, you know, part of the inference time scaling that Gabby was talking about is very much that. Itβs like, you know, the more we sample, the more we, like, you know, the ranking model. The ranking model ends up finding something it really likes. And so I think our ability to get better at ranking, I think, is also whatβs going to enable sort of the next, you know, next big, big breakthroughs. Interesting.
Brandon [00:48:17]: But I guess thereβs a, my understanding, thereβs a diffusion model and you generate some stuff and then you, I guess, itβs just what you said, right? Then you rank it using a score and then you finally... And so, like, can you talk about those different parts? Yeah.
Gabriel [00:48:34]: So, first of all, like, the... One of the critical kind of, you know, beliefs that we had, you know, also when we started working on Boltz 1 was sort of like the structure prediction models are somewhat, you know, our field version of some foundation models, you know, learning about kind of how proteins and other molecules interact. And then we can leverage that learning to do all sorts of other things. And so with Boltz 2, we leverage that learning to do affinity predictions. So understanding kind of, you know, if I give you this protein, this molecule. How tightly is that interaction? For Boltz 1, what we did was taking kind of that kind of foundation models and then fine tune it to predict kind of entire new proteins. And so the way basically that that works is sort of like instead of for the protein that youβre designing, instead of fitting in an actual sequence, you fit in a set of blank tokens. And you train the models to, you know, predict both the structure of kind of that protein. The structure also, what the different amino acids of that proteins are. And so basically the way that Boltz 1 operates is that you feed a target protein that you may want to kind of bind to or, you know, another DNA, RNA. And then you feed the high level kind of design specification of, you know, what you want your new protein to be. For example, it could be like an antibody with a particular framework. It could be a peptide. It could be many other things. And thatβs with natural language or? And thatβs, you know, basically, you know, prompting. And we have kind of this sort of like spec that you specify. And, you know, you feed kind of this spec to the model. And then the model translates this into, you know, a set of, you know, tokens, a set of conditioning to the model, a set of, you know, blank tokens. And then, you know, basically the codes as part of the diffusion models, the codes. Itβs a new structure and a new sequence for your protein. And, you know, basically, then we take that. And as Jeremy was saying, we are trying to score it and, you know, how good of a binder it is to that original target.
Brandon [00:50:51]: Youβre using basically Boltz to predict the folding and the affinity to that molecule. So and then that kind of gives you a score? Exactly.
Gabriel [00:51:03]: So you use this model to predict the folding. And then you do two things. One is that you predict the structure and with something like Boltz2, and then you basically compare that structure with what the model predicted, what Boltz2 predicted. And this is sort of like in the field called consistency. Itβs basically you want to make sure that, you know, the structure that youβre predicting is actually what youβre trying to design. And that gives you a much better confidence that, you know, thatβs a good design. And so thatβs the first filtering. And the second filtering that we did as part of kind of the Boltz2 pipeline that was released is that we look at the confidence that the model has in the structure. Now, unfortunately, kind of going to your question of, you know, predicting affinity, unfortunately, confidence is not a very good predictor of affinity. And so one of the things that weβve actually done a ton of progress, you know, since we released Boltz2.
Brandon [00:52:03]: And kind of we have some new results that we are going to kind of announce soon is kind of, you know, the ability to get much better hit rates when instead of, you know, trying to rely on confidence of the model, we are actually directly trying to predict the affinity of that interaction. Okay. Just backing up a minute. So your diffusion model actually predicts not only the protein sequence, but also the folding of it. Exactly.
Gabriel [00:52:32]: And actually, you can... One of the big different things that we did compared to other models in the space, and, you know, there were some papers that had already kind of done this before, but we really scaled it up was, you know, basically somewhat merging kind of the structure prediction and the sequence prediction into almost the same task. And so the way that Boltz2 works is that you are basically the only thing that youβre doing is predicting the structure. So the only sort of... Supervision is we give you a supervision on the structure, but because the structure is atomic and, you know, the different amino acids have a different atomic composition, basically from the way that you place the atoms, we also understand not only kind of the structure that you wanted, but also the identity of the amino acid that, you know, the models believed was there. And so weβve basically, instead of, you know, having these two supervision signals, you know, one discrete, one continuous. That somewhat, you know, donβt interact well together. We sort of like build kind of like an encoding of, you know, sequences in structures that allows us to basically use exactly the same supervision signal that we were using to Boltz2 that, you know, you know, largely similar to what AlphaVol3 proposed, which is very scalable. And we can use that to design new proteins. Oh, interesting.
RJ [00:53:58]: Maybe a quick shout out to Hannes Stark on our team who like did all this work. Yeah.
Gabriel [00:54:04]: Yeah, that was a really cool idea. I mean, like looking at the paper and thereβs this is like encoding or you just add a bunch of, I guess, kind of atoms, which can be anything, and then they get sort of rearranged and then basically plopped on top of each other so that and then that encodes what the amino acid is. And thereβs sort of like a unique way of doing this. It was that was like such a really such a cool, fun idea.
RJ [00:54:29]: I think that idea was had existed before. Yeah, there were a couple of papers.
Gabriel [00:54:33]: Yeah, I had proposed this and and Hannes really took it to the large scale.
Brandon [00:54:39]: In the paper, a lot of the paper for Boltz2Gen is dedicated to actually the validation of the model. In my opinion, all the people we basically talk about feel that this sort of like in the wet lab or whatever the appropriate, you know, sort of like in real world validation is the whole problem or not the whole problem, but a big giant part of the problem. So can you talk a little bit about the highlights? From there, that really because to me, the results are impressive, both from the perspective of the, you know, the model and also just the effort that went into the validation by a large team.
Gabriel [00:55:18]: First of all, I think I should start saying is that both when we were at MIT and Thomas Yacolas and Regina Barzillaiβs lab, as well as at Boltz, you know, we are not a weβre not a biolab and, you know, we are not a therapeutic company. And so to some extent, you know, we were first forced to, you know, look outside of, you know, our group, our team to do the experimental validation. One of the things that really, Hannes, in the team pioneer was the idea, OK, can we go not only to, you know, maybe a specific group and, you know, trying to find a specific system and, you know, maybe overfit a bit to that system and trying to validate. But how can we test this model? So. Across a very wide variety of different settings so that, you know, anyone in the field and, you know, printing design is, you know, such a kind of wide task with all sorts of different applications from therapeutic to, you know, biosensors and many others that, you know, so can we get a validation that is kind of goes across many different tasks? And so he basically put together, you know, I think it was something like, you know, 25 different. You know, academic and industry labs that committed to, you know, testing some of the designs from the model and some of this testing is still ongoing and, you know, giving results kind of back to us in exchange for, you know, hopefully getting some, you know, new great sequences for their task. And he was able to, you know, coordinate this, you know, very wide set of, you know, scientists and already in the paper, I think we. Shared results from, I think, eight to 10 different labs kind of showing results from, you know, designing peptides, designing to target, you know, ordered proteins, peptides targeting disordered proteins, which are results, you know, of designing proteins that bind to small molecules, which are results of, you know, designing nanobodies and across a wide variety of different targets. And so thatβs sort of like. That gave to the paper a lot of, you know, validation to the model, a lot of validation that was kind of wide.
Brandon [00:57:39]: And so those would be therapeutics for those animals or are they relevant to humans as well? Theyβre relevant to humans as well.
Gabriel [00:57:45]: Obviously, you need to do some work into, quote unquote, humanizing them, making sure that, you know, they have the right characteristics to so theyβre not toxic to humans and so on.
RJ [00:57:57]: There are some approved medicine in the market that are nanobodies. Thereβs a general. General pattern, I think, in like in trying to design things that are smaller, you know, like itβs easier to manufacture at the same time, like that comes with like potentially other challenges, like maybe a little bit less selectivity than like if you have something that has like more hands, you know, but the yeah, thereβs this big desire to, you know, try to design many proteins, nanobodies, small peptides, you know, that just are just great drug modalities.
Brandon [00:58:27]: Okay. I think we were left off. We were talking about validation. Validation in the lab. And I was very excited about seeing like all the diverse validations that youβve done. Can you go into some more detail about them? Yeah. Specific ones. Yeah.
RJ [00:58:43]: The nanobody one. I think we did. What was it? 15 targets. Is that correct? 14. 14 targets. Testing. So we typically the way this works is like we make a lot of designs. All right. On the order of like tens of thousands. And then we like rank them and we pick like the top. And in this case, and was 15 right for each target and then we like measure sort of like the success rates, both like how many targets we were able to get a binder for and then also like more generally, like out of all of the binders that we designed, how many actually proved to be good binders. Some of the other ones I think involved like, yeah, like we had a cool one where there was a small molecule or design a protein that binds to it. That has a lot of like interesting applications, you know, for example. Like Gabri mentioned, like biosensing and things like that, which is pretty cool. We had a disordered protein, I think you mentioned also. And yeah, I think some of those were some of the highlights. Yeah.
Gabriel [00:59:44]: So I would say that the way that we structure kind of some of those validations was on the one end, we have validations across a whole set of different problems that, you know, the biologists that we were working with came to us with. So we were trying to. For example, in some of the experiments, design peptides that would target the RACC, which is a target that is involved in metabolism. And we had, you know, a number of other applications where we were trying to design, you know, peptides or other modalities against some other therapeutic relevant targets. We designed some proteins to bind small molecules. And then some of the other testing that we did was really trying to get like a more broader sense. So how does the model work, especially when tested, you know, on somewhat generalization? So one of the things that, you know, we found with the field was that a lot of the validation, especially outside of the validation that was on specific problems, was done on targets that have a lot of, you know, known interactions in the training data. And so itβs always a bit hard to understand, you know, how much are these models really just regurgitating kind of what theyβve seen or trying to imitate. What theyβve seen in the training data versus, you know, really be able to design new proteins. And so one of the experiments that we did was to take nine targets from the PDB, filtering to things where there is no known interaction in the PDB. So basically the model has never seen kind of this particular protein bound or a similar protein bound to another protein. So there is no way that. The model from its training set can sort of like say, okay, Iβm just going to kind of tweak something and just imitate this particular kind of interaction. And so we took those nine proteins. We worked with adaptive CRO and basically tested, you know, 15 mini proteins and 15 nanobodies against each one of them. And the very cool thing that we saw was that on two thirds of those targets, we were able to, from this 15 design, get nanomolar binders, nanomolar, roughly speaking, just a measure of, you know, how strongly kind of the interaction is, roughly speaking, kind of like a nanomolar binder is approximately the kind of binding strength or binding that you need for a therapeutic. Yeah. So maybe switching directions a bit. Boltβs lab was just announced this week or was it last week? Yeah. This is like your. First, I guess, product, if thatβs if you want to call it that. Can you talk about what Boltβs lab is and yeah, you know, what you hope that people take away from this? Yeah.
RJ [01:02:44]: You know, as we mentioned, like I think at the very beginning is the goal with the product has been to, you know, address what the models donβt on their own. And thereβs largely sort of two categories there. Iβll split it in three. The first one. Itβs one thing to predict, you know, a single interaction, for example, like a single structure. Itβs another to like, you know, very effectively search a space, a design space to produce something of value. What we found, like sort of building on this product is that thereβs a lot of steps involved, you know, in that thereβs certainly need to like, you know, accompany the user through, you know, one of those steps, for example, is like, you know, the creation of the target itself. You know, how do we make sure that the model has like a good enough understanding of the target? So we can like design something and thereβs all sorts of tricks, you know, that you can do to improve like a particular, you know, structure prediction. And so thatβs sort of like, you know, the first stage. And then thereβs like this stage of like, you know, designing and searching the space efficiently. You know, for something like BullsGen, for example, like you, you know, you design many things and then you rank them, for example, for small molecule process, a little bit more complicated. We actually need to also make sure that the molecules are synthesizable. And so the way we do that is that, you know, we have a generative model that learns. To use like appropriate building blocks such that, you know, it can design within a space that we know is like synthesizable. And so thereβs like, you know, this whole pipeline really of different models involved in being able to design a molecule. And so thatβs been sort of like the first thing we call them agents. We have a protein agent and we have a small molecule design agents. And thatβs really like at the core of like what powers, you know, the BullsLab platform.
Brandon [01:04:22]: So these agents, are they like a language model wrapper or theyβre just like your models and youβre just calling them agents? A lot. Yeah. Because they, they, they sort of perform a function on behalf of.
RJ [01:04:33]: Theyβre more of like a, you know, a recipe, if you wish. And I think we use that term sort of because of, you know, sort of the complex pipelining and automation, you know, that goes into like all this plumbing. So thatβs the first part of the product. The second part is the infrastructure. You know, we need to be able to do this at very large scale for any one, you know, group thatβs doing a design campaign. Letβs say youβre designing, you know, Iβd say a hundred thousand possible candidates. Right. To find the good one that is, you know, a very large amount of compute, you know, for small molecules, itβs on the order of like a few seconds per designs for proteins can be a bit longer. And so, you know, ideally you want to do that in parallel, otherwise itβs going to take you weeks. And so, you know, weβve put a lot of effort into like, you know, our ability to have a GPU fleet that allows any one user, you know, to be able to do this kind of like large parallel search.
Brandon [01:05:23]: So youβre amortizing the cost over your users. Exactly. Exactly.
RJ [01:05:27]: And, you know, to some degree, like itβs whether you. Use 10,000 GPUs for like, you know, a minute is the same cost as using, you know, one GPUs for God knows how long. Right. So you might as well try to parallelize if you can. So, you know, a lot of work has gone, has gone into that, making it very robust, you know, so that we can have like a lot of people on the platform doing that at the same time. And the third one is, is the interface and the interface comes in, in two shapes. One is in form of an API and thatβs, you know, really suited for companies that want to integrate, you know, these pipelines, these agents.
RJ [01:06:01]: So weβre already partnering with, you know, a few distributors, you know, that are gonna integrate our API. And then the second part is the user interface. And, you know, we, weβve put a lot of thoughts also into that. And this is when I, I mentioned earlier, you know, this idea of like broadening the audience. Thatβs kind of what the, the user interface is about. And weβve built a lot of interesting features in it, you know, for example, for collaboration, you know, when you have like potentially multiple medicinal chemists or. Weβre going through the results and trying to pick out, okay, like what are the molecules that weβre going to go and test in the lab? Itβs powerful for them to be able to, you know, for example, each provide their own ranking and then do consensus building. And so thereβs a lot of features around launching these large jobs, but also around like collaborating on analyzing the results that we try to solve, you know, with that part of the platform. So Boltβs lab is sort of a combination of these three objectives into like one, you know, sort of cohesive platform. Who is this accessible to? Everyone. You do need to request access today. Weβre still like, you know, sort of ramping up the usage, but anyone can request access. If you are an academic in particular, we, you know, we provide a fair amount of free credit so you can play with the platform. If you are a startup or biotech, you may also, you know, reach out and weβll typically like actually hop on a call just to like understand what youβre trying to do and also provide a lot of free credit to get started. And of course, also with larger companies, we can deploy this platform in a more like secure environment. And so thatβs like more like customizing. You know, deals that we make, you know, with the partners, you know, and thatβs sort of the ethos of Bolt. I think this idea of like servicing everyone and not necessarily like going after just, you know, the really large enterprises. And that starts from the open source, but itβs also, you know, a key design principle of the product itself.
Gabriel [01:07:48]: One thing I was thinking about with regards to infrastructure, like in the LLM space, you know, the cost of a token has gone down by I think a factor of a thousand or so over the last three years, right? Yeah. And is it possible that like essentially you can exploit economies of scale and infrastructure that you can make it cheaper to run these things yourself than for any person to roll their own system? A hundred percent. Yeah.
RJ [01:08:08]: I mean, weβre already there, you know, like running Bolts on our platform, especially on a large screen is like considerably cheaper than it would probably take anyone to put the open source model out there and run it. And on top of the infrastructure, like one of the things that weβve been working on is accelerating the models. So, you know. Our small molecule screening pipeline is 10x faster on Bolts Lab than it is in the open source, you know, and thatβs also part of like, you know, building a product, you know, of something that scales really well. And we really wanted to get to a point where like, you know, we could keep prices very low in a way that it would be a no-brainer, you know, to use Bolts through our platform.
Gabriel [01:08:52]: How do you think about validation of your like agentic systems? Because, you know, as you were saying earlier. Like weβre AlphaFold style models are really good at, letβs say, monomeric, you know, proteins where you have, you know, co-evolution data. But now suddenly the whole point of this is to design something which doesnβt have, you know, co-evolution data, something which is really novel. So now youβre basically leaving the domain that you thought was, you know, that you know you are good at. So like, how do you validate that?
RJ [01:09:22]: Yeah, I like every complete, but thereβs obviously, you know, a ton of computational metrics. That we rely on, but those are only take you so far. You really got to go to the lab, you know, and test, you know, okay, with this method A and this method B, how much better are we? You know, how much better is my, my hit rate? How stronger are my binders? Also, itβs not just about hit rate. Itβs also about how good the binders are. And thereβs really like no way, nowhere around that. I think weβre, you know, weβve really ramped up the amount of experimental validation that we do so that we like really track progress, you know, as scientifically sound, you know. Yeah. As, as possible out of this, I think.
Gabriel [01:10:00]: Yeah, no, I think, you know, one thing that is unique about us and maybe companies like us is that because weβre not working on like maybe a couple of therapeutic pipelines where, you know, our validation would be focused on those. We, when we do an experimental validation, we try to test it across tens of targets. And so that on the one end, we can get a much more statistically significant result and, and really allows us to make progress. From the methodological side without being, you know, steered by, you know, overfitting on any one particular system. And of course we choose, you know, we always try to choose targets and problems are sort of like at the frontier of whatβs possible today. So, you know, you donβt want something too easy. You donβt want something too hard. Otherwise youβre not going to see progress. And so, you know, this is a somewhat evolving set of targets. We talked earlier about the targets that we looked at with, with Boltchan. And now we are even trying kind of, you know, even harder targets, both for small molecule and proteins. And so we try to keep ourselves on the, on the boundary of whatβs possible. So do you have like infrastructure or is this is like, you just have a lot of different partnerships with academic labs and youβre just kind of keep pushing on these and driving these. We do partially this through academic labs more and more. We do this through CROs just because of, you know, to some extent is also, we need kind of replicability often kind of, you know, going after the same time. So we try to, we try to keep our, our targets, you know, multiple times and, you know, to see the, the progress from, you know, one month to the next. And speed. And speed. And speed. Speed of execution. Yeah. And, So what happens if you start getting a bunch of like really strong biters against therapeutic targets? What do you do?
RJ [01:11:43]: Release them. Yeah.
Gabriel [01:11:45]: But you can release them in open source? Like,
RJ [01:11:47]: Yeah, I mean, you know, I mean, when we say we have no interest in making dress, weβre serious. Like, you know, uh, I mean, when it, when it was with the academic labs, basically the, you know, I was, they keep it, they do a lot of it.
Gabriel [01:12:02]: I will also say, and I think this has been a bit of the issue that I have with some of the things that have been said in the field, is when we say that we design new proteins or we say that we design new molecules, go and bind these particular targets. We should be very clear, these are not drugs. These are not things that are ready to be put into a human. And there is still a lot of development that goes with it. And so this is kind of to us, we see ourselves as building tools for scientists. At the end of the day, it really relies on the scientist having a great therapeutic hypothesis and then pushing through kind of all the stages of development. And, you know, we try to build tools that can accompany them in that journey. Itβs not like a magic box where, you know, you can just turn it and get FDA approved drugs.
Brandon [01:13:06]: But actually, that brings up an interesting question that Iβve been wondering about is, do you guys see yourself staying in this, for lack of a better way of saying it, layer? Or do you think that youβll start to... Yeah. Either on the physical sense, looking at different layers of the virtual cell, so to speak, or also, you know, so thereβs like the development process that goes, you know, sort of like design preclinical, clinical approval and thinking about improving the performance throughout that process based on the designs. Is that a direction that you guys are pushing? Yeah.
Gabriel [01:13:45]: So one of the things, as Jeremy said, you know, we are... We are not a therapeutic company. We want to kind of stay not to be a therapeutic company, always be at the service of, you know, all the different, you know, companies, including therapeutic companies that we serve. And, you know, that to some extent does mean, you know, that we need to try to, you know, go deeper and deeper in getting these models better and better. One of the things that we are doing across, you know, many other in the field is, you know, now that we are really... Theyβre starting to be good, both for small molecule and... For proteins to design kind of binders, design relatively tight binders, is starting to look at all these other properties, you know, theyβre called developpabilities or at me that, you know, we care about when developing a drug and try, can we design them from, from Gageco. The thing about those properties in some of them, you know, you need to, you know, start having an understanding of the cell. And so thatβs on the one hand, kind of why we need that understanding. But also, you know, the way... The way that we also think about all different and complex diseases is that these models, then these tools that weβre building have a good understanding of kind of, you know, biomolecular interactions and kind of their interactions. Now, at the same time, every disease is often kind of unique and every therapeutic hypothesis is unique. And so you maybe want to have something that needs to hit the particular, you know, letβs say target in a virus in a particular way, but you donβt maybe know exactly. So you can start to have a more open-minded understanding of whatβs, whatβs a way you want to do. And so maybe in the first set of designs, youβre going to try to target different epitopes in different ways, and then youβre going to test them in the lab, maybe directly in vivo, and youβre going to see which ones work and which ones donβt. And so then you need to bring those results back into the models. And then the models can start to have a more wider understanding, you know, not just of the biophysical of the antibodies interacting with that target, but also how that is shaped within the cell. And so first of all, you know, that means on the one end that we need, you know, kind of these loops, and this is also partially how we, we designed the platform to be. But that also means that we also need to start understanding more and more kind of higher level things. And, you know, I wouldnβt say that weβre working in any way on like a virtual cell like others are, but weβre definitely thinking kind of very deeply about kind of, you know, how does, you know, kind of the way that we target certain proteins. Interfere, interact with, you know, maybe pathways that are existing in the cell. One question that has come up is you talk a lot about user interface and so on. And I think this is really important, but like my experience with dealing with medicinal chemists, when you get the machine learning models, is they are the most superstitious, skeptical, like pseudo-religious people Iβve ever talked to when it comes to doing science. Sorry for the medicinal chemists listening. Yeah, theyβre amazing. Like, theyβre absolutely, Iβve worked with some spectacular medicinal chemists who just pull magic out of their hat again and again, and I have no idea how they do it. But when you bring them a machine learning model, it is sometimes quite tricky to get them to deal with it. How has your interaction been with this? And how have you thought about, like, building Boltβs lab to work with the skeptics? One of the great value unlocks for us and for our product has been when we brought to the team a medicinal chemist. His name is Jeffrey. So I think kind of like on the one hand, you know, day one, you know, he obviously had a lot of opinions on kind of a lot of the ways that we should change, you know, both kind of the way that the agents worked, the way that the platform worked. But itβs been really amazing kind of, you know, once also we started kind of shaping kind of the platform in a better way with this feedback, how we went from, you know, to some extent, you know, a fair skepticism to him, you know, actually using, you know, a lot of the things that we did. Yeah. So heβs doing a lot more compute than any of our computational folks in the team, you know, at times that, you know, heβs, you know, running, you know, he has all these sort of hypotheses. Okay, maybe I can hit this protein this particular way. I can hit in that way. Actually, let me look at for this particular molecular space. Let me try to optimize for this particular interactions. So he ends up, you know, running several screens in parallel, you know, using hundreds of GPUs, you know, on his own. And, you know, so this has been, you know, pretty incredible to see kind of how, you know, maybe the way that I was more thinking about a problem, which is, okay, youβre just trying to design a binder, a small molecule to a particular protein. The way that he thinks about it is, you know, much more deeply and, you know, trying all these different things, these different hypotheses. And then, you know, once he gets the results from the model, he doesnβt just, you know, take the top 15, but he really kind of looks over and, you know, kind of tries to understand, you know, the different things. And then when we select, you know, maybe some designs to bring forth, you know, he has, you know, something where, you know, both the models understand that somethingβs good, but himself as well. And thatβs why we also built kind of the platform to be an interface for, you know, this kind of chemist and, you know, also like engineers. Yeah. Collaborative experience.
RJ [01:19:09]: I think at the end of the day, like, you know, for people to be convinced, you have to show them something that they didnβt think was possible. And until you have that aha moment, you know, I think the skepticism will remain. But then when, you know, every once in a while, I think thereβs like a result that like really surprises people. And then itβs like, oh, wow, okay, this is actually, I can do something with this. So you just get in their hands, have them try it out, and theyβll be convinced. Yeah, or like maybe once the lab results come back. Or their friends. Yeah, or maybe one of their colleagues is convinced. Yeah. I think it takes going to the lab at some point. Thereβs no avoiding that, you know, as beautiful as the platform can be, as nice as the molecules might look, you know, that the model predicted. I think what really convinces people is like, you know, hits. Yeah.
Gabriel [01:19:54]: Yeah. You see the results. Exactly. Yeah. Cool. Thank you for, you know, taking the time to chat with us. Yeah. You know, is there anything that you would like your audience to know? I mean, first of all, you know, weβre just getting started, you know, continuing to build a team. And so definitely always looking for great folks, both on the kind of, you know, software side, you know, machine learning side, but also scientists to join the team and help us, you know, shape. On the infrastructure side, too. Indeed. If you think that if you want a new challenge, because this is not just next token prediction, this is really a new engineering challenge. Exactly. Yeah. If you, if no matter, you know, how much experience you have with, you know, biologists and chemistry, if you want to come, you know, help us in a shape, what, you know, biology and chemistry, hopefully weβll look like in five, 10 years. Weβd love to hear from you. And so go to boltz.bio and, you know, come join the team. Cool. Thank you. Awesome. Thank you so much. Thank you.
Get full access to Latent.Space at www.latent.space/subscribe




