Arc Institute's Patrick Hsu on Building an App Store for Biology with AI

15 Apr 2025 · 58 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Training Data - Patrick Hsu on Building an App Store for Biology with AI

Episode Overview Title: Arc Institute's Patrick Hsu on Building an App Store for Biology with AI Hosts: Josephine Chen and Pat Grady Guest: Patrick Hsu, co-founder of Arc Institute Release Date: [Insert Date] Description: Patrick Hsu discusses the potential of AI in biology beyond drug development. He presents Evo 2, a biology foundation model trained on genomic data, which can interpret mutations and assist in various applications in biological sciences.

---

Key Concepts

Introduction to Evo 2

  • Evo 2 Model: A revolutionary biological foundation model that interprets and generates genomic sequences across all domains of life.
  • Purpose: To identify evolutionary patterns, predict the effects of mutations, and design new biological systems.
  • Training Data: The model was trained on a vast dataset of genomic sequences, learning patterns that would typically take years to discover.

Applications of AI in Biology

  • Beyond Drug Development: While AI is often associated with drug design, Patrick emphasizes that its applications extend to understanding biology at all scales.
  • Identifying Mutations: Evo 2 can predict the functional consequences of genetic mutations, particularly those of unknown significance.
  • Examples of Use Cases:
  • Predicting effects of mutations in the BRCA1 gene, which is linked to breast and ovarian cancer.
  • Designing CRISPR gene editing systems.

Challenges and Considerations

  • Interpretation of Genetic Mutations: Many mutations found through genetic testing (e.g., 23andMe) are classified as variants of unknown significance (VUS). Evo 2 aims to provide insights into these.
  • Regulatory Bottlenecks: Even with advanced AI, the process of bringing drugs to market is lengthy due to regulatory requirements.

The Unifying Theory of Biology

  • Patrick introduces the idea of a unifying theory in biology, akin to the unifying forces in physics, that connects biological sequences to their functional implications through evolutionary processes.

The Future of AI in Biology

  • Building an App Store for Biology: Patrick envisions a future where various models and applications can be accessed, analogous to an app store, fostering collaborative and innovative biological research.
  • AI Agents in Science: The development of AI agents that assist in hypothesis generation, experimentation, and data analysis is anticipated to advance the scientific method.

Key Takeaways

  • Evo 2's Importance: The model aims to create a deeper understanding of biology rather than solely focusing on drug discovery.
  • Potential for Efficiency: AI can streamline various steps in the drug development process, although challenges such as regulatory hurdles remain.
  • Integration of Multi-Disciplinary Knowledge: Effective biological research will increasingly require collaboration across various scientific disciplines.

Closing Thoughts

  • Optimism in Research: Patrick emphasizes the importance of maintaining a positive outlook in scientific endeavors, which can foster creativity and persistence in achieving long-term goals.

---

Mentioned Resources

  • Evo Papers:
  • [Public pre-print of original Evo paper](#)
  • [Public pre-print of Evo 2 paper](#)
  • Databases:
  • [ClinVar](https://www.ncbi.nlm.nih.gov/clinvar/)
  • [Sequence Read Archive](https://www.ncbi.nlm.nih.gov/sra)
  • [Protein Data Bank (PDB)](https://www.rcsb.org/)
  • Articles:
  • "Machines of Loving Grace" - Daria Amodei's essay referenced by Patrick.

Conclusion This episode provides an insightful look into the intersection of AI and biology, outlining the transformative potential of Evo 2 and the need for a broader perspective on the applications of AI in life sciences. Patrick Hsu's vision of an interconnected ecosystem of biological applications highlights the future of scientific discovery driven by advanced technologies.

---

*Note: This summary is based on the provided transcript and contains synthesized information from the podcast episode.*

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00One of the things that we, you know, that the field of computational biology is often asking is, you know, if you have a genetic mutation in your genome, if I sequenced you, whether that's via, you know, 23andMe or, you know, or some other genetic tests, right? You'll find mutations in your genome. How do we actually interpret those and understand, you know, what the functional consequences are, right? Sometimes you'll get a rare genetic disease. Those are causal genetic mutations that are known to cause a devastating disorder, that might be muscular dystrophy or cystic fibrosis or breast cancer.

0:35But most of the mutations that you have, they're sort of this, we call them variants of unknown significance, which is fancy kind of scientist - You don't really know what the hell is going on, right? And it turns out the model has an opinion about those mutations and what the hell is going on with them and it turns out it's sort of state of the art in doing that. [♪ OUTRO MUSIC PLAYING [♪

1:11Today we're joined by Patrick Xu, a pioneer in genome editing, CRISPR technologies, and the emerging field of generative biology. He's the co -founder of the Arkansas Institute. where cutting edge AI and biology converge to reimagine scientific discovery. Patrick and his collaborators created EVO2, a revolutionary biological foundation model that can interpret and generate genomic sequences across all domains of life. By training on the fundamental information layer of life, DNA itself, EVO can identify patterns from genetic code at scale and predict effects of both coding and non -coding mutations that can mean the difference between health and disease.

1:52In this episode, we'll hear how Patrick's vision goes beyond creating better drugs to building comprehensive understanding of biology at all scales. Patrick, welcome to the show. Thank you for coming. Thanks for having me on. Excited to spend some time with you today. I think maybe the most obvious thing to start with is, you know, people have heard about CS and Bio for the longest time. Now it's all about AI and bio. We're the results. Like what should we actually be expecting to see? And - Where are the drugs? Why are we not seeing the drugs yet? It takes time, right? Well here's the thing. Even if we had perfect drug design molecules coming out of these pipelines and fancy models, you can design a trillion molecules, ten trillion, right?

2:42But you still have to actually test them. Right, initially in animals and then in people, right. And so that's the real bottleneck. And even if you pack top of funnel with all the things, it just takes years to actually go through the regulatory apparatus, right. And so I think there are a few intermediate checkpoints along the way in order to kind of realize this potential. But it might be worth taking a step back and just saying, you know, this is a bit of a soapbox of mine that ML for bio is not just drug design, right. This is actually ultimately, I think, a very important, but narrow part of the potential of biology and not just as a field of STEM and in the way that affects human lives, right?

3:21And do you mean basic, just like understand the human body or where else do you think the applications are? We'll end drugs that treat all of us. Yeah, one of the things that motivates me academically is the idea that we actually have a unifying theory for biology, right? So unlike the physicists who have been kind of scrimping and kind of poking for one for a century, right? We have this in biology and it's so obvious that we Find it, you know, sort of an just an obvious force, right? This is of course a solution, right? And so it acts on biology across all of its different length scales from entire planets Right biology can for example, terraform planets, right all the way down to you know ecosystems and populations to individuals, to our tissues, to individual cells, to individual molecules.

4:11That's this unifying force that is actually very deep and rich and actually you can learn a lot from. How do we activate that unifying theory? How do we put it to work? We've been thinking about this in the lab and recently have been training a series of models that we call EVO inspired by these forces of application that tries to connect biological sequences using this sort of modern sequence modeling paradigm directly to biological function, with the idea that evolution passes down its effects of natural selection throughout generations of life via DNA mutations. And so last year, of course, multiple Nobel prizes were awarded for AI, and for AI in biology, in particular for protein design, and for predicting the structure proteins to David Baker and to Demisosabis and John Jumper.

5:05But if you read those citations, they both explicitly state for proteins. And we love proteins. These are, of course, some of the most important and fundamental molecular machines, but our realization, for me, as a genome biologist, if you will, the ideas that proteins are encoded in DNA, along with RNA and with regulatory DNA, and all of the things that you need to make life. And so we asked, could we train a model on genomes with a long context models that it could reason over all the different bases and molecules are embedded inside of genomes to learn about the molecular interactions and how they lead to biological function?

5:48Now, that was very scientific or academic, right? But we can talk through specific examples of what we were able to actually do with this in a way that grandma can understand, you know, like predicting the effects of breast cancer causing mutations, right? It actually is best in class at doing this, right? Or being able to design new CRISPR gene editing systems. Or, you know, I think, you know, in addition to zero shot capabilities, I think people are building really an app store for biology on top of all of these kind of foundational layers, one of the kind of funny things and modeling today is how everyone's model has to be more foundational than someone else's model.

6:32There's a bit of a pissing contest. That's happening, right? But maybe our model is more foundational than the other models. Because you're DNA versus protein. And then below that, there are these all -atom diffusion models. So maybe those are even more fundamental. I don't really know. I think what matters are the capabilities and doing something that actually feels useful and think, you know, those are some examples of what we thought was cool and useful from the model. Can you actually walk through a couple more of those use cases like which are some of the most exciting ones? And why is it possible today with EVO but wasn't possible before?

7:06In the model's open source and so where has it been picked up? Like where are people running with it with some of those use cases that Josephine mentioned? Yeah, so the model, I can just maybe talk about what the model is. Yeah. So, you know, it's an auto -aggressive sort of multi -convolutional hybrid model, right? But you can think of it like, you know, just a really efficient long -context model that's trained at least in this version, auto -aggressively, right? And basically, it does this next token or next base prediction, and it turns out just like in natural language or in vision or in robotics and embodied intelligence, this general machine learning paradigm is able to find higher order patterns.

7:47And so just like if you're doing next word prediction, you can learn about grammar or color and world navigation. You've seen to learn some rich set of representations about biology by predicting the next base or the next amino acid residue or the next gene. The model learns something about the molecular logic that gives rise to itself. And so one of the things that we, you know, that the field of computational biology is often asking is, you know, if you have a genetic mutation in your genome, if I sequenced you, whether that's via, you know, 23 and me or, you know, or some other genetic test, right?

8:30You'll find mutations in your genome. How do we actually interpret those and understand what the functional consequences are? Sometimes you'll get a rare genetic disease. Those are causal genetic mutations that are known to cause a devastating disorder. That might be muscular dystrophy or cystic fibrosis or breast cancer. But most of the mutations that you have, there's sort of this, we call them variants of unknown significance, which is fancy kind of scientist. Yeah, we know what the hell is going on, right? And you know, it turns out the model has an opinion about those mutations and what the hell is going on with them.

9:08And it turns out it's sort of state of the art in doing that. Interesting. Wait, what's an example of one of these mutations and what, what did the model discover and how do you verify that what a discover was accurate? Yeah, yeah. So, so, you know, one example that we showcase in the paper is a gene called Brocka one, right? It's sort of a famous gene that's known to cause breast and ovarian cancer. And if you have the sort of specific causal mutations in Brocka 1, many women elect to get double mastectomies. This is obviously a serious and major life decision and medical decision for you and for your family.

9:42And the question is, if you don't have one of the known to be benign mutations, so you're fine, and you just go ahead and get an annual mammogram and just check and monitor, right? There's this entire middle distribution of these VUSs or variants of unknown significance. And there is this kind of gold standard database from the scientific literature it's known as ClinVar and it basically has a list of all the different genes that are known to cause disease and which of those mutations in those genes can cause disease state or not. And we can basically use this as a ground truth database to assess the predictions of the model for new mutations that you introduce into the gene and whether or not those would be pathogenic.

10:29And so when you develop a new type of model, you have to create a lot of the e -vows as well. And so that was actually something that we put tremendous effort into. And obviously you guys see this horizontally across AI in many different domains as folks will build things to the benchmarks. No one really, I think, likes building benchmarks. It's really gory. It requires a lot of taste. It takes a lot of time. You have to continually update them as the models get better. And we dealt with a very similar sort of challenge here, which was making eVals that would be similar to the AGI or intelligence eVals, where it actually feels meaningful when you're actually able to do it like Amy or Putnam problems or things like that.

11:16what would be the equivalent of demonstrating true biological understanding that a cell biologist would feel emotion if you were actually able to solve that, right? You know, what would it look like to make all molecular biologists feel what the NLP people felt a few years ago, right? That's sort of the core of, you know, what we're kind of noodling through. And I think you mentioned this briefly, but you know, there's, you guys are working on the DNA layer and you guys didn't actually do anything in the lab. you didn't have a lab in the loop, like there was no RL, talk us through the decision to do that, and then kind of the decision for people who are doing protein models or affinity models, a lot of them have a lot more lab in the loop, talk us through the differences between some of those models too.

12:00Yeah, so we started with DNA because we think it's the fundamental information layer of life, right? The second is it's also just pragmatically where we have the most data. Like data, yeah. Where does the data come from? It comes from the entire scientific community, right? And so there are these open source government funded and maintained databases known as the sequence read archive, where basically when you publish a paper, you have to submit your data, and all of the sequencing data that's been created over the last 25 plus years goes into these databases. And that has all the genomes that the community has ever sequenced for bacteria, for bacterial phage for viruses for humans monkeys fish flies you know the entire Noah's Arctic or menagerie whatever right we've got all those genomes and you know so so the experiment that's already been done if you will is the experiment of evolution right that you know there are different mutations that are what make us different from each other that's human genomic variation but also the mutations that make us different from chimpanzees or from worms or from bacteria.

13:14And the model can just look across this trillions of tokens large data set and learn those patterns. And that was sort of the insights of this sort of evocere of models that we've been training at ARIC is to kind of ask if you could predict that next base, right? then that might be the difference between being healthy or having a sickle -cellenemia mutation, right? Or it could predict the next amino acid residue, not could be the difference between having a catalyptically active binding pocket for a key enzyme, right, in your, you know, body physiology, or that being a null mutation where that thing doesn't work anymore.

14:01Or it could also be the next gene. So these are different levels of abstraction that completes some biosynthetic pathway, where it's a different gene that's been removed by some transposon or jumping gene mobilitarian element that's excised is some sort of viral interference, if you will. So there's just across large databases, you find new patterns, and it turns out those seem to be biologically meaningful. So if you're going from DNA and you know the function, in many ways you mentioned even the DNA model can actually predict binding affidimities. Do you even need the structure, the protein structure models at all as intermediate step, or can you just go straight from sequence to function?

14:47So structure is another way to have an abstraction of function. So you have concepts of convergent evolution, for example, where something has similar function, but they have different sequences and slightly different structures that act out the activity or function of that protein. And so this sort of sequence, the structure, the function, token, or mapping of protein language modeling think is very beautiful. It takes advantage of the central dogma of molecular biology. The full series. Right. Well, it takes advantage of just our supervised textbook understanding of how biology works. Right.

15:30Now, the interesting oxymoron with these models for me is the way that we use them is, you know, just like you use chatch -e -b -t, it's text in, text out. You know, Evo is D -N -A -N, D -N -A -Out. Yeah. And it turns out, we don't speak D -N -A very well. This sort of imagine if you were using a model like chat chatt -t -b -t in Russian, right? But 1 % of the words were in English. Right? That's kind of what it feels like. That's the vibe of using Evo is actually you don't really know what's going on. And so you have to build lots of annotators and... Interpreability. Yeah. The like techniques to try to interpret and read what's happening.

16:11And so the way that we even use and prompt the models is really primitive, right? And so the way that we do fancy prompt engineering workflows to, you know, get more utility out of these models is something that we're just in the very early innings of exploring how that works with these biological language models because we speak DNA with an extremely heavy accident. Who do you think will do some of that work? Because the model is currently open source. Do you envision a company formed around this? Will it be individuals who are at these former companies who learn how to prompt this? Like how does that ecosystem evolve?

16:47Yeah, I hope everybody, right? First of all, right? And I think the tools become useful when they meaningfully lower the energy barrier adoption. It's like a it's a catalysis type of activity where everyone uses blast for sequence alignment. Everyone uses alpha -fold to look at protein structure. Everyone uses CRISPR to do gene editing. Everyone does NGS to read DNA or RNA. So I think there will be a zoo of different models for not just modeling molecules or mapping sequence to function, but I think every step of the scientific method. Right, and so I think, you know, if 2025 is the year of AI agents, right?

17:34I think you know, there's lots of interest in agents for science and not just agents for interpreting molecules but also for doing the meta aspect of how scientists work and operate. Yeah, and we're also very excited about that at ARC and recently released some of our first AI agent work. What is the most important role for ARC to play in all of this? Um, we started the institute to be able to have this mothership that is able to attack long -term research capability breakthroughs and to be able to have the long -term thinking and the multidisciplinary expertise in order to actually execute on these goals, right?

18:13I think if you look in biology, one of the interesting things is that a lot of the biggest mechanistic or basic science breakthroughs do happen in a university context. I would say that's interestingly kind of in contrast to what happens in AI. Yeah. RCS, yeah. Yeah. What do you think that's the case? That might be a 10 -hour podcast, right? One minute. I have a lot to say about this, but I think in short, it does happen in universities. I think, you know, 30 years ago, the questions that basic science was interested in and industry was interested in were quite different, actually. And I would say today they seem to heavily overlap, right?

19:06And there are some things that, you know, folks don't typically do an university lab, like, you know, people on tennis study, PK or talks or CMC, right, or, you know, hardcore drug manufacturing type things, but folks are interested in molecular glues, induced proximity, and degraders, and new drug concepts. Folks are also interested in new machine learning models, and new delivery mechanisms, and inflammation, and stress, and all kinds of things like this. So there's much more heavy overlap, but I think the way that the type of product that that selects what you do upstream is very different.

19:45They're structurally different. You know, end of the day, you have to optimize for grant funding and first or co -first author papers and, you know, that type of stuff in the academic setting, whereas, you do have to make a drug and have, you know, dozens to hundreds of people lying behind a molecule or a program to reach the Holy Land, right? And I think that does lead to differences in strategy. I don't know personally, by the way, I think people really over emphasize what's academia and what's industry and the realities these are heavily overlapping distributions today and help people operate.

20:25I think the differences are a little overblown, but they do have fundamentally different incentives that drive different behaviors. A lot of what we try to do and blend our academic side of the house with our technical staff side of the house which is built much more like you'd see an industry. It has been part of what we hope will make ARCA model that others want to copy or replicate or propagate. Yeah. It makes sense. You've had some fun people come through on the technical side of the house, including people like Greg Brockman. Yeah, no, it's been a joy. Yeah, so Greg joined us during his sabbatical from OpenAI.

21:07We're really the first vacation that he had ever taken since starting OpenAI. Of course, it's to work for him. So actually, the funny story about this, hopefully Greg will allow me to tell it. But when he first came by Arc, and we were talking to him about biological language modeling and how the machine learning breakthroughs of all these other domains might just pour over directly lead to understanding molecules. He was really jazzed by this and also by the idea that his very specific capability and expertise would be really meaningful in actually making this happen. But initially, he was saying, this is really the first vacation I've ever taken.

21:57I promise too much. I might not be able to spend so much time on this. I need to take Anna on vacation. And then Anna gets up and goes to the bathroom. In the middle of this very long meeting, and then he's like, all right, here's my email. Get me on the repo. Of course. Of course. Yeah, just built different. It was such a joy to learn from him. What have you found works well in getting the different disciplines to work together as one unified team. Yeah, I know it's an interesting question, and we've thought about this deeply at ARC as a convening center, you know, not just between the three flagship research universities here in the Bay Area and Stanford and Berkeley and UCSF, but also between basic science and the biotech industry, but also biology and the technology sector.

22:53For example, you know, our CTO Dave Burke just started a few months ago and has really kind of been leading the computational modeling of virtual cells where we're trying to simulate human biology with these AI foundation models. And Dave used to, he actually is a PhD in biomedical engineering from many moons ago. It was most recently ran engineering at Android and Pixel. And so I think we have also built an entire operational side of the house from finance to legal to the lab ops, to facilities, to university relations and academic affairs, and we run around space, around administration, around ops.

23:35And we try to do much of that like a tech company, right? And we've also, I think, recruited in on the ops side of the house many people who maybe ordinarily wouldn't work in a basic science or discovery setting, but are kind of motivated by the mission to be able to take fundamental breakthroughs and have a product sensibility where we can get these out into the real world. And we're not optimizing it arc for nature and science papers, right? If you give Berkeley or Stanford professors millions of dollars to do more science, that's almost the default expectation and output, right? And what we really care about are things that could be tangible, right?

24:18And the real world. So does success look like just more people using the products you create or what is a success for our Institute. I think there are lots of ways that you can parse scientific productivity. Journal publications is of course a important and fundamental part of sharing work with the community and pure reviewing it and all of that good stuff. But technical blogs, code reposts, protocols, are just platforms that people can use. And I think we want to make technologies and platform capabilities that are broadly useful to be able to create new mechanistic insights, but also actually try to cure some diseases.

24:57And I think over time, if we can be an Edison shop that's inventing or finding lots of cool things, there hopefully will be real world valley in those things, and there are lots of partners who specialize in this. So you can be curing diseases earlier talking about even if you had perfect drug design or infinitely accessible, infinitely intelligent drug design. There's still a long process that has to happen after that before you can make an impact on humans. Can you talk a bit about where you see opportunities in that value chain? And if you had a magic wand and you could just accelerate progress by 10 years, which of those bottlenecks might be alleviated by things you see coming down the pipeline?

25:40Yeah, so happy to talk about this in the pharma context, which are large decentralized sprawling bureaucracies, much like universities, or the Congress or the SFCD governments. And I think some parts are incredibly functional and then I think everyone would agree in these different organizations. Some parts are less efficient. And so I think the first thing that hopefully we can see is that we can have models that can and prove efficiency in discrete individual steps. So can we have a more efficient process for target ID in particular? Can we have a more efficient process for data analysis? Can we have a more efficient process for information and literature review and summarization for molecule design, absolutely, making a better binder?

26:37And then figuring out the drug properties of those different molecules. That could be selectivity, that could be pharmacokinetics, half -life, expression, manufacturability, what have you. And there will be models or model guided approaches for each of those steps. And then there's, I mean, I think if you look at how AI is actually being used today in pharma companies, a lot of them are actually just taking massive regulatory documents, summarizing them, and then using AI to help them write more of this. So there's this compression and decompression of information in a structured fashion that has actually been leading to enterprise adoption in this setting.

27:28And I don't know, I think that says something about the process of where people find things useful. So, one example is if you talk to some pharma execs, not the drug discovery organizational leaders, but the budget people who hold the enterprise per strings, they'll say, well, I don't spend that much money on drug discovery. I actually spend most of my money on drug development. So tell me something about drug development, which is really where most of my dollars go. How can AI help me with that? right? And I think that is actually like a very deep comment right? Because it says something about where money is spent and where value can be found.

28:15You know the first thing that realizes you know our industry probably success is like 10 % right? And so you know I think a lot of the things that people talk about or complain about or comment on in drug discovery and development falls out of that fundamental statistic, right? Like, why does the FDA heavily regulate? Why is it so focused on safety? Well, if 90 % of the time it doesn't work, they're gonna care a lot about safety, right? And I think the promise of AI is, if we can go from 10 % POS to 20 or 30 or 50, as you move to those steps, can we? Do you think we can? Over time, I think here's the thing that I find kind of interesting is we have done a lot of biology with what is not that far from guess and check.

29:11If you look at what happens in the wet lab, like the actual experiments that are happening, you're just kind of in the arena trying to. Yeah, trying out random hypotheses and seeing what happens. Yeah, and this is like the missing, reasoning trace in the scientific literature is you don't know what didn't work. Everything is narrativized and written in a story of, you know, an inexorable logic and vision leading to scientific breakthrough, right? But everyone who makes the sausage knows that's not what actually happens most of the time, right? And, you know, if you actually work with the, you know, research, you know, kind of folks at the bench, right?

Read the full transcript

29:55And this is actually the case across all technical industries is that high fidelity reasoning trace of the true process is kind of not written down anywhere. And that would actually be very useful for these reasoning models and for closing the loop and doing multi agent frameworks, blah, blah, blah, blah. But that's kind of what we'll need. But if you actually look at what happens, it's guess and check. And so a model with even a modicum of predictive value would be transformative. One with even moderate predictive value, which by the way, we don't have, right? I think biology is a very pragmatic, sult of the earth, experimental discipline, right?

30:35You see this in the culture of peer review, show me the data, you can't pontificate in your discussion section because you haven't shown any of this stuff, right? If you read old papers, right? They were so clear and visionary and high -floating in a way that I think that papers today are, you know, we have this culture that is very pragmatic. I think having models that have predictive power will, you know, I think A, obviously be useful for accelerating the efficiency of science. It will also change the culture, which I think hopefully will change the culture, which I think will be really interesting.

31:14Why do you think it will change the culture? Do you have the models? Because you'll believe people's predictions or pontifications depending on how you - It's just another evidence point basically. Just like how model hallucinations could be, you know, predictions or they could be nonsense and garbage, right? And that depends on how much you trust the model, right? Got it. Why do you think it's the case? We've collected a lot of data to your point. There's obviously, we basically see what works and we oftentimes don't see what dozen, although hopefully lab notebooks are recording that in some ways you perform.

31:49Maybe. Maybe, hopefully. But somehow still very, very regularly, even when things work in cells, things work in mice, they fail in humans oftentimes. Why is that still the case? We just don't understand biology deeply enough. Why is there still that drop off and that drop off hasn't really changed over time? Well, these are imperfect models, right? And we set up this set of filters in the drug discovery process where the first show it works in cell lines, then it shows it works in primary cells or in an organoid, then it shows it works in a mouse, then it shows it works in a monkey, then tested in people.

32:28And by the time you've gone there, like five years and $100 million has gone by, and I think that's very challenging. And that's where I think predictive models will really help, right? Because the reason why we do all of these steps in linear series is because we don't have predictive power. And so we have to do things in the arena. And it just, you know, you have to do it in real life, right? And growing cells and growing animals takes months to years to actually do those experiments. And so the promise of having predictive models and just predictive power, it's that you could actually simulate things in a multi -paralleled fashion.

33:12That's the whole idea behind parts of machines of loving grace that I thought Daria really did get right. And it is the idea that if you had something that could be a trusted oracle that you could to just run 10 ,000 agents at the same time. Do we have enough data for that full closed loop to create a trusted Oracle? I think we will see more examples of this coming out of our time. Today, folks building AI agents for things are doing basically trying to close the gap between step X and step X plus one, or X plus four, or whatever. right? And the businesses are trying to find the most commercially valuable set of steps that is a set size that's as small as possible in step number in order to make a company.

34:09I think we will have agents or co -pilots at each step in the scientific method from hypothesis generation to experimentation to data analysis. And the ability to close the loop is in right the paper or or make the discovery and decide what to do next, I think is quite far away. But I think something that's very efficient at traversing the steps, I think, will really take off. And so as a concrete example, right? One of the things that we recently released at ARCA is our virtual cell Atlas, right? Which is the world's largest data set of single cells, right? That we're using for training these cellular foundation models, right?

34:51And the way that it happened was we created an agent that was essentially, it's like a crawler, kind of like, you know, kind of a search crawler, but it's able to crawl the kind of sequence read archive and then process all of the highly unstructured and messy metadata and, you know, kind of re -analyze and systematically reprocess all single cell data. And this is something that is just running on a cloud bucket instance, just cranking away, right? In a tireless fashion, right? And it's the kind of stuff that a talented computational biologist wouldn't want to do because it's so grind set, but actually the scale at which we're able to reach is community -wide.

35:37And that was that the leverage and efficiency that our team of, you know, just, you know, could achieve with one agent, I think was, it was a huge mental unlock for me. And so we wanna be at the frontier of actually deploying these and making breakthroughs. And I think the meta aspect that folks are going after right now will shake out over time. But I care about using these to actually make breakthroughs as opposed to chart the end -to -end closed loop path. You mentioned earlier the sort of the pragmatism that a lot of research papers have now versus grandiosity or something a bit more visionary back in the day.

36:20I wonder if some of that is related to the specificity of the work that people are doing now, meaning it feels like we've gotten more and more specialized over time. And I wonder if some of that is taking us away from breakthroughs, because in a lot of cases, you need the knowledge from different domains or different disciplines to achieve those breakthroughs. And I guess maybe the question is, one of the nice things about LLMs is that they can incorporate an enormous amount of information, and they're sort of inherently generalized even when you apply them to a specific domain. How much of the efficacy that we might get out of some of these models is simply related to their ability to go across all these different specialties?

37:02Yeah, if you look at at least to me the best scientist that I've had the pleasure to collaborate with or learn from They

37:16Really do two things they they they they're able to come up with really creative ideas and they're able to execute on them Yeah, the reason why they're able to come with really creative ideas is because they're able to make connections between things that other people wouldn't make. And in fact, if you got a room of 10 really smart people together to chat science, like that's a weekly lab meeting in any group, right? If you actually analyze the anthropology of what happens, there's usually a small subset of people who are hearing all the things that are being discussed and then actually saying, this is a connection, or the conceptual bridge between things, right?

37:57And so there's, you know, that's sort of like an out of distribution generalization, right? It's like this thing that I heard was really novel and let's recognize that and then try to see what can generalize out of that observation, right? And that comes from people who tend to either read a lot or reason a lot, right? And so there's some aspect of pre -training, right? You need to just read a lot of papers, right? Read a lot of chemical biology papers. Read a lot of molecular biology papers. Read a lot of AI papers. We come like physics papers, right? And do so across domains so that you can traverse those boundaries, right?

38:39I think, you know, there's this design problem of how do you build a multi -disciplinary team, right? And the reality is there, for example, you want to work at the interface of Bio -NML. They're way more ML people and way more bio people than truly bilingual ML and bio people, right? I have the extreme fortune of working with some of those at ARC, right? And, you know, they're just rare, right? But those translators can actually help you power, you know, the rest of the population. Yeah. Yeah. What do you look for in people at ARC? It depends on the role, right? I think, you know, depending on how you're trying to match specific project needs.

39:26But in a way, I could tell you all the ways that we try to intellectualize our recruiting process. But it actually comes down to very simple things, right? And it's the same thing that I look for in a research technician or an executive on, at least on the science side, right? It's really like, are you thinking about signs outside of the lab? Right. And have you done something end -to -end before? Right. And then the third is, do you have the grit to actually kind of walk the path and get it done? Yeah. Right. What do you mean by the end -to -end part? I think it's very easy to go from step one to two, or three to five, or, you know, 12 12 to 15, but going from 1 through 15, when it was down, the population significantly.

40:22And so I often say, the last 20 % of a project is actually 80 % of the work. Yeah. And it's because finishing something and then honing your killer instinct from finishing things multiple times really matters. What should we expect to be coming out of our institute in the next six months and then over the next few years? Well, I'm tremendously excited about lots of things. And I think the thing that maybe many people don't know is the degree to which we have really been trying to build biology at ARC. I think people have maybe heard about our gene editing work or our machine learning work, right?

41:03But a lot of what we're actually trying to build is this general concept of applying high throughput but scalable technologies in the context of multi -systems interactions, right? Really working at the neuro and immune interface. And so we hired two incredible scientists out of Penn last year who study the process of interoception, right? So proprioception is when you kind of close your eyes where your limbs, right? And interoception is the idea of, you know, I feel the weather in my knee or my tummy feels funny, right? A kind of midwives, tails, type stuff that actually has really deep science.

41:47And of course, it's totally unknown and how does your body talk to your brain and vice versa? And it turns out there's a deep mechanistic basis for this. And as, you know, I think when people think about programming biology, they think about in the drug paradigm of how do I get a binder that binds to this protein? Yeah, or how do I get a crisper to edit this gene? But if you think about what happens with hormones or with a ZMPIC, right, you're able to program the way that you think and feel and behave in really powerful ways that controls not just the tidy, but energy, mood, you know, you know, muscle synthesis, you know, all focus, all kinds of things.

42:30And I think how do we actually program physiology is something that I've been spending a lot of time thinking about in our lab? What's one unexpected connection that you think people don't think about? One example is the exercise, right? And so, Christoph, one of our PIs, you know, had a beautiful paper where he showed that there's a specific species of gut bacteria. that produce a certain type of molecule that connects via your interic nervous system, which is the nervous system that lines your gut that goes to your brain in order to release dopamine. And it is this functional circuit that creates runners high or exercise reward.

43:17And when you delete this bacteria, you cut off this E and S to brain circuit it, or you cut off the ability of the brain to release the dopamine at each of these steps individually, you can block the runner's height. And so it really traces in an intact animal, well, this is a mouse study, right? This full body circuit, but that also goes in reverse. So when you have deep psychological stress, right? That can lead to signaling from the brain to astrocytes that innervate your gut that releases prone flammatory cytokines that leads to gut inflammation and then can give you ulcers. So stress causes ulcers.

44:03We've actually kind of known this. Yes, but this is the mechanism of how can you treat it? Because there are folks who get recurring ulcers. There's a brain, two -body axis by which this signals, right? And I think this happens all the time, you know, some of which is conscious, most of which is unconscious. And you can actually start to figure out the dials and knobs. And that's actually in a way like a new paradigm for drugs. Yeah. Or how to think about using drugs, right? It's not just this highly reductive from suitable or, you know, kind of binder to biomarker type of thing, but you know that's giving you more of the holistic kind of almost eastern medicine flake.

44:51Exactly. Just how do I feel healthier? Yeah. Right? That I think is invoked in the longevity community today. Yeah. Like what is health span? How do I improve it? How do I improve my diet, my nutrition? So I just feel better. Right. You know, those things are also can have deep scientific grounding and needs to have that. So how do you think people will be treated in the future? You will have a full panel, you'll know exactly what's going on in your body, and then you'll decide different inputs to influence the whole thing. How do you tie in functional medicine, things that go on longevity with the drug industry as it is today?

45:30Where does that interplay? I mean, I think we'll want AI doctors, right? They are able to integrate information multimodally, right? And so just like you have your CGM that monitors your glucose with high temporal resolution, you have your oirring or your whoop that talks about your various biomarkers or you can go to quests or, you know, function health or whatever and get blood tests, right, that, you know, measure what's going on with your liver function or your cholesterol or, you know, your testosterone or estrogen or, you know, other types of hormones, right? Right now, all you really know is that that these things are going up or down and whether or not they're in standard or reference range.

46:13I doesn't tell you very much about what you're supposed to do. And I think one thing that has been really missing and personalized genetics or consumer genetics is the ability to take information content from your genome sequence. Yeah. And meaningfully integrate it with your health biomarkers in a way that gives you the genotype and the environment that can be more predictive of phenotype. That GXE equals P equation, you learn high school biology. But none of that is actually kind of accessible to mom and dad or to even us. In the setting of how do I actually live my life? And I think we need to go from measuring people with higher content approaches to connecting that to genetic signatures and make more accurate predictions.

47:08So that's sort of like the theme of our conversation today. What do you think that'll look like? Do you think that'll look like any of the existing longevity efforts? Do you think there's some new beast entirely that's going to be created to serve that purpose? Yeah. I mean, so I think if you look at 23 me was recently file chapter 11, and I think it's an amazing pioneer. And I'm visionary kind of pioneering effort in, you know, how to take genetics and, you know, put it in the hands of millions of people, right? I think the thing that I think I would really love to see in the world is something that can take all of that information with all of your different, you know, kind of body measurements.

47:52And then actually, you know, connect that to diet and sleep and give you personalized recommendations about your health in a longitudinal way. We have very fragmented data sets for being able to do this today and I think being able to collect this data at scale across populations and over time with temporal resolution will I don't want to be one of these unhinged big data will solve everything people. But it will. But it will.

48:25I do wonder if there's more stuff on the cross -functional side to your point. If you know you have a gambling addiction, I don't think anybody's thinking about maybe Mungaro's Ampick could help with that, but it does help with some of these things, right? I do think there's something about the cross -functional nature. That's fair. There needs to be some organization that actually makes accessible. Yeah, yeah, I don't think that obviously exists today. Yeah, I think folks are building it a different hands on the elephant for this, but you know, you guys should start this company. We should. It's a pretty good idea.

49:02Yeah. All right, let's ask, let's get a couple of predictions on a couple of different timescales. So we'll start with 2025 and then maybe 2030 and then maybe 2050. What is the most interesting thing we're going to see in the world of AI meets bio in 2025 by 2030 and by 2050? My hope is by the end of the year. I mean, and this is already happening, right? We can, you know, just we can design full IGG antibodies, right? Not single -chain binders like nanobodies, but just the real antibody medicines that, you know, we kind of know and love today. that we can just design their CDR regions. They're gonna bind really well.

49:44You can one -shot it, and you can kind of do point and click on that surface of your enzyme. I can just bind it, right? One shot. I think the thing that will mature over the next couple years is that we can actually design enzymes, do NOVO. I think that will be really interesting and also lots of efforts. And again, this is all in the world of proteins. And I think one of the things that most people who think about this stuff are very protein -coded And so a lot of our our our our work is to sort of zoom out from proteins and think about cells, right? And so I think building the The sort of the the PDB of virtual cells, right?

50:24It's something that we've been focusing a lot on at ARC that will take some Years from today to mature So PDB is a protein data bank, right? And it's the sort of gold standard database of Atomic resolutions solved, experimentally solved protein structures that was used by a deep -mind to train alpha -fold. And so it's the pre -training data that allows the model to reach some soda capability, like protein structure prediction, at an extreme resolution. So what is that for virtual cells, which we think would help us design better drug targets, increase therapeutic probability of success? I think that's sort of my 2030 prediction is that we have accurate and useful virtual cell models that make a cell biologist feel emotion.

51:18The 2050 idea, and hopefully this happens much, much sooner than that, there's lots of chat about scientific superintelligence, or the end to end recursion of the scientific method. I'd like to see that, right? With, you know, lab in the loop with a fully automated wet lab, that's vertically integrated. Do you think, do you think as possible of a 2050, we can simulate with 99 .9 % accuracy, you know, the impact that a particular drug is going to have on a particular target, you know, validate that in a wet lab in a fully automated way in a matter of hours, not months. What do you think the dream scenario is for going from zero to impact in the future of drug discovery if you imagine 25 plus years of technological progress?

52:13Yeah, I mean, I think we've laid out different aspects of the vision over the course of our conversation today, where things are really slow, like toxicity, long -term follow -ups. These are the, you know, it depends a lot on the disease, right? If you're doing some acute oncology thing, that's very different from some really chronic autoimmune thing, right? And so I think the only way that I can imagine you speeding this up is if you have a model with strong predictive power, right? And you know, so a lot of this hinges on our ability to make models that can actually do that. And I think you will unlock different step sizes of capability based on how good they are.

52:57Is there any reason we wouldn't have that by 2050? Yeah, basically if we make the wrong data, I would say it's like one obvious one. You know, you can model the mouse and all of its glory to great perfection. It will still not be the human. And that's, you know, one sort of, I think, you know, trivial example that is something that we still do, though, because that's just what's practical. And so I have this other soap box about how we actually just need to be doing way more experiments in humans. And what do you think it will take for us to be able to do that? Is that just a regulatory thing?

53:41Is that? I think there will be some aspect of creativity involved in addition to better regulatory innovation. So one example would be, you know, there are these kind of, you can take samples from brain dead patients, for example. And then now you can just get lungs that you can perfuse and keep alive for a week. And then just do experiments in that lung, right? And, you know, there's a paper were recently published about this. Great. Should we do Lightning Realm? Yeah, let's do it. Maybe first one, favorite new AI app that you've tried in the last three months. So this is maybe cheating, but I'm a DAU of OpenAI deep research.

54:21I find it just by far the main AI app that I find useful enough to use in my day -to -day work. And so there are lots of other fun AI apps or things that I pay attention to because I'm interested in AI. But the thing that actually changed how I work is these DPSRH models. And they have, by the way, so much more room to improve and to run. And yeah, I think it was the first time I felt real emotion thinking, oh, wow, okay, someday maybe I will be automated. Yeah, yeah. Yeah, feel that every day. Who would be on your Mount Rushmore of scientists? This is maybe a bit smarmy, but the folks I get to work with at ARC, I realize this is incredibly smarmy, but it's really genuinely, I feel so lucky to go to Lab every day and just be around, you know, just passionate, bright, kind, and incredibly ambitious people.

55:33Yeah. And it levels up my game. What do you think is going to be the killer application that scientists will use by the end of this year? Deep research. So that. All right. Nothing about you guys are going to create from work. Well, I think these virtual cell models will be incredibly useful. We, I don't think we or anyone will have working models in the sense that they, I think it will take some time for them to mature to the point where they're actually fundamentally useful, right? They're currently research problems that will ripen over some time, yeah. But we'd like to deliver those. What's the most important thing you've learned at our institute?

56:17So there's this Scottish proverb that I'm going to butcher when I pair of fries, but it's basically, you know, be happy when you're alive for your longtime dad. And, you know, I think that really hit for me when I read it. And, you know, it was one of those, you know, you're lying on the couch, the phone is six inches from your face, just beaming lucks into your eyeballs right before you're supposed to fall asleep. I don't know, that just reminded me that despite all the complexity of trying to work really hard to do useful things, you're supposed to have fun. I think we need more of that in life, not just in research labs, where I think it can be so easy to be super critical because that's the training and how you make progress is to hate on everything and everything has a problem and figure out why this can go wrong, but being optimistic and happy is not a path to mediocrity and mistakes, but it's actually how you have the emotional capacity for persistence over time to reach those long -term goals.

57:30Well, I feel like a good place to end it. No, well hopefully this podcast made you happy too. I'm always happy to see you, Buck. Thank you again. Thank you. Thank you guys.

From the publisher

Patrick Hsu, co-founder of Arc Institute, discusses the opportunities for AI in biology beyond just drug development, and how Evo 2, their new biology foundation model, is enabling a broad ecosystem of applications. Evo 2 was trained on a vast dataset of genomic data to learn evolutionary patterns that would have taken years to find; as a result, the model can be used for applications from identifying mutations that cause disease to designing new molecular and even genome scale biological systems.

Hosted by Josephine Chen and Pat Grady, Sequoia Capital

Mentioned in this episode:

Sequence modeling and design from molecular to genome scale with Evo: Public pre-print of original Evo paper

Genome modeling and design across all domains of life with Evo 2: Public pre-print of Evo 2 paper

ClinVar: NIH database of the genes that are known to cause disease, and mutations in those genes causally associated with disease state

Sequence Read Archive: Massive NIH database of gene sequencing data 

Machines of Loving Grace: Daria Amodei essay that Patrick cites on how AI could transform the world for the better

Arc Virtual Cell Atlas: Arc’s first step toward assembling, curating and generating large-scale cellular data from AI-driven biological discovery (among many other tools)

Protein Data Bank (PDB): a global archive of 3D structural information of biomolecules used by DeepMind to train AlphaFold

OpenAI Deep Research: The one AI app Patrick uses daily

More from Training Data

All 110 episodes
Arc Institute's Patrick Hsu on Building an App Store for Biology with AITraining Data · 58 min
Listen in VO