AI Got Good at Language. Now It’s Learning the Language of Life. (Eric Nguyen, Co-Founder and CEO of Radical Numerics)

4 Aug 2026 · 1 h 22 min · 31 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Radical Numerics teaches AI to read, write, and program biological function in DNA, arguing DNA is biology’s “language” and that multi-omics integration is the next frontier. The episode also covers biosecurity concerns raised by AI-designed pathogens.

Guest background

Eric Nguyen is co-founder and CEO of Radical Numerics. He helped create Evo and Evo2, generative DNA models trained on raw DNA sequences. He previously worked in tech (including Facebook), then earned a PhD in bioengineering after earlier engineering/consulting work. He describes a career pivot after major shoulder/spine/chest surgeries.

Key claims

  1. Scientists used Radical’s models to generate an AI genome from scratch, creating a bacteriophage (a working, infectious virus) not found in nature.
  2. Data isn’t the bottleneck: DNA is “internet-scale,” and current model expressivity is comparable to linear/logistic regression, leaving “vast green space.”
  3. Biology should be modeled as a unified, multimodal system (DNA plus RNA/proteins/epigenomics/metabolomics), not one modality at a time.
  4. The same capability that enables novel pathogen design must also enable defense; he says public AI tools for detecting dangerous pathogens are lacking.

Notable examples

  • Lab researchers designed an infectious bacteriophage using an Evo model.
  • “Obama neuron” style mechanistic interpretability analogy for finding cancer-relevant “neurons.”

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to DNA and AI

0:00 to 0:45

Learn about the significance of DNA and the potential of AI in biology.

“DNA is this genetic code that makes up a lot of life on Earth.”

Evo and AI-Generated Life

0:45 to 1:54

Discover how AI has created life forms and the implications for biology.

“For us, this should be pretty simple to do.”

Radical Numerics and AI in Biology

3:37 to 6:30

Understand Radical Numerics' mission to integrate AI with biological modeling.

“I think you're building such a fascinating company and a very important one.”

The Complexity of Biological Systems

6:30 to 10:45

Explore how DNA serves as a foundation for understanding complex biological systems.

“such an interesting part of the business that you're doing and how you have to sort of balance those few things.”

Philosophical Perspectives on DNA and Biology

10:45 to 14:00

Delve into the philosophical implications of using DNA as a foundation for biological AI.

“that's been lacking of how to actually incorporate all of this complexity of information.”

DNA as a Language

14:00 to 15:00

Explore the comparison of DNA to human language and its implications.

“feed into a model from a language model standpoint.”

Understanding DNA Complexity

15:00 to 17:40

Delve into the nuances of DNA's structure and its complex functionalities.

“It's complicated, but like DNA, maybe that's DNA is all you need.”

Biologist vs Engineer: Approaches to DNA

17:40 to 19:00

Contrast the methodologies of biologists and engineers in understanding biological systems.

“standpoint, um, takes, can take months, years or years, basically.”

Limitations of Traditional Genome Study

19:00 to 22:30

Discuss the challenges faced in studying genomes and the implications for biotechnology.

“So this idea is one of the reasons why we got interested in using large language models to study and to really push the field and what's possible to manipulate the genome with large language models.”

Generative AI in Biology

22:30 to 24:20

Examine how generative AI can revolutionize our approach to biological research.

“and generate novel things in that space like language or image and video.”
Show all 31 chapters

The Importance of Mechanistic Understanding

24:20 to 26:10

Highlight the significance of understanding mechanisms in both AI and biology.

“and we are especially excited about bringing some of these frontier capabilities of AI research like Mechinterp, where, you know, cool that you can use Mechinterp for understanding how, you know, language models work.”

Challenges of Applying LLMs to Biology

26:10 to 28:00

Discuss the unique challenges of using large language models in biological contexts.

“This is why we're excited about it, in large part why I love to come on these venues to be able to share why I think it's so damn exciting and how much opportunity and green space there still is.”

The Challenge of Sparse Scientific Data

28:00 to 29:16

Learn about the challenges AI faces in processing sparse biological data compared to language.

“And you kind of start seeing this like in audio and in video, right, where the density of information is not the same.”

Rethinking AI's Role in Science

29:16 to 31:05

Explore the limitations of traditional science learning and the potential of AI to directly analyze scientific data.

“to analyze raw DNA sequences or protein structure, they need to call specialized models.”

Personal Journey to Tech and AI

31:05 to 33:21

Understand Eric's upbringing and varied career path leading to his focus on technology and health.

“be used into biology and peer into it in a different way.”

Navigating Life Changes and Purpose

33:21 to 36:10

Discover how significant health challenges led Eric to reassess his career and find meaningful work.

“And I don't want to be too harsh, but it's like fundamentally a job where you have much, much, much less agency, right?”

The PhD Experience and Motivation

36:10 to 40:15

Learn about the challenges and motivations behind pursuing a PhD later in life.

“And yeah, I think part of it is having that freedom, I suppose, or being used to that kind of wandering and flexibility and not being tied to, I need to do XYZ by certain age.”

Applying AI to Biology: The Hyena Model

40:15 to 42:00

Hear about the innovative application of AI models to bioinformatics, particularly DNA analysis.

“that just because you make the decision to do a PhD doesn't mean like, it's like clear, like, oh yeah, it was the right thing to do.”

Exploring DNA as a New Frontier

42:00 to 45:58

Learn how DNA emerged as a viable data source for AI models.

“and then noticed that it was particularly good at long context.”

The Challenge of Data in Biology

46:34 to 51:09

Understand the complexities of using biological data for AI applications.

“And also what I was curious about is, is it equivalently useful?”

Overcoming Data Limitations in Bio

51:09 to 56:00

Explore strategies for utilizing data in bioinformatics more effectively.

“that we want to enable this generalization this fusion of modalities that you've seen parts of it applied to language and vision and audio sometimes you know three domains in bio there are dozens of of modalities.”

The Evolution of Bio-AI Companies

56:00 to 58:10

Learn how bio-AI companies are adapting to the challenges of the industry.

“Being able to dive into a domain, not being an expert in, say, cancer vaccines.”

Bottlenecks in Wet Lab Operations

58:10 to 1:00:20

Discover the challenges and solutions in running industrialized wet lab operations.

“So we are likely to do our own wet labs pretty soon as well.”

Applying AI to Alzheimer's Research

1:00:20 to 1:02:30

Explore how AI can recapitulate complex genetic research findings in Alzheimer's.

“Have you sequenced your own genome and tried to extract any interesting information from it yet?”

The Future of General Biological Intelligence

1:02:30 to 1:04:40

Understand the potential of general biological intelligence in various applications.

“We're going after a huge diversity from applications in pharmaceuticals, drug discovery, diagnostics, what's known as synthetic biology, so like designing whole genomes, for example, for us, and in biodefense.”

Balancing Detection and Defense in AI Models

1:04:40 to 1:10:04

Learn about the ethical considerations in AI for biodefense and pathogen detection.

“At the end of the day, we're betting on innovation to have a step change of the technology to actually make this horizontal model work.”

Dual Mission: Design and Defense in AI

1:10:04 to 1:11:42

Explore the balance between advancing AI design capabilities and ensuring defense mechanisms.

“frontier of the design capabilities as well.”

Government Awareness of Biological Risks

1:11:42 to 1:13:22

Discuss the current state of government awareness and response to biological threats.

“And so we need to bring up the defensive side to even out the playing field.”

Growing Awareness and the Future of Biosecurity

1:13:22 to 1:15:57

Delve into the evolving conversation around biosecurity and the potential for change.

“So I would challenge that, you know, there is a trend of being concerned of biological risks and biological threats, but in terms of meat behind it, not a lot.”

Innovative Applications of AI in Biodefense

1:15:57 to 1:17:48

Learn about the potential applications of AI technology in addressing biodefense challenges.

“And they're relatable to things that people in the AI community are already familiar with.”

Thought Experiments and Recommended Readings

1:17:48 to 1:19:52

Engage with thought experiments related to AI and hear book recommendations for deeper insights.

“When we can showcase there are generative problems, there are mechintert problems, there are all sorts of interesting ML problems, but people have mostly thought it was too domain specific.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Eric Nguyen:DNA is this genetic code that makes up a lot of life on Earth. Just last year, one of the things that scientists had used to generate with our models was the first AI genome from scratch. The genome is the complete set of instructions for an organism that created what's known as a bacteriophage. Something that did not exist in nature, we realized that this power, this potential to create life itself or to manipulate life itself with AI has the potential to reinvent pretty much all of biology. There's more DNA publicly online than all the text on the internet. The expressivity power of these models being applied to biology, they're basically as strong as linear regression or logistic regression.

0:35Eric Nguyen:Basically, this is the vast green space that's open for mining, basically. We don't see any tool out there publicly available on the AI side that can detect dangerous pathogens. So we initially were like, let's just make an API and just host it. For us, this should be pretty simple to do. And then when we started engaging with biosecurity experts, they basically told us like that, don't do that.

1:00Last year in a lab, researchers sat down at a computer and designed a virus using an AI model called Evo. Not a paper about a virus, but an actual working bacteriophage, built and infectious. Eric Nguyen is one of Evo's creators. He is also the CEO of Radical Numerics, a company building cutting-edge AI models for biology, starting with DNA. Since its founding last summer, Radical has raised$50 million from investors like Emergence Capital and Patrick Collison. In today's conversation, Eric and I discuss his winding path to founding Radical, the unique language of DNA, why he believes data isn't the bottleneck for biological models, and why the company that makes it possible to create novel viruses must be the same one that defends against them.

1:50I'm Mario, and this is The Generalist. This episode is brought to you by Ahrefs Brand Radar. If you've tried searching for your brand in ChatGPT or Google's AI results, you've probably noticed something. It's not always clear why certain brands get mentioned or why yours doesn't. That's the problem Ahrefs Brand Radar solves. It helps you see how your brand shows up across AI-powered answers, not just in search engines, but also on platforms like Reddit and YouTube. Instead of guessing, you get real data based on millions of actual user prompts. You can quickly check your share of voice, find out which websites or sources are influencing those answers, and see where competitors are getting ahead of you.

2:32What makes it different is that it doesn't rely on simulations or small samples. It's built on large-scale, real-world data. So you're seeing what people are actually asking AI. There's no complicated setup either. Just enter your brand and start exploring. Visit hrefs.com slash generalist to learn more about Brand Radar. The best founders aren't spending their time on expense reports. They're busy building. Brex is the agentic finance platform that makes that possible. High-limit corporate cards, banking, and AI that handles the back office automatically so your team never has to. Expenses get captured, books get closed, and spend stays in policy without anyone chasing it down.

3:15Your team gets their time back to focus on what actually moves the company forward. Vercel, OpenAI, Anthropic, Granola, and Deepgram already run on Brex. One in three startups in the U.S. does too. It's time to get Brex. Go to brex.com slash solutions slash startups. Eric, I've been really excited to have this conversation. I think you're building such a fascinating company and a very important one. I wanted to start maybe with just sort of laying the groundwork and foundation for folks to understand what it is that you're doing. I think a lot of people will maybe have thought about the confluence of AI and biology, but you're doing something quite particular.

3:58And so maybe you can just sort of give us a rough frame of what Radical Numerics is trying to build.

4:03Eric Nguyen:Awesome. Thanks for having me here. Very excited as well. So Radical Numerics is a new AI research lab. And what we do is teach AI to be able to read, write, and essentially program function into DNA. The DNA is this genetic code that makes up a lot of life on Earth. And what we think is really important is to train models that not just understand science through natural language, but that can be able to reason across the raw substrate of life itself, this genetic code, like if it was a language. And so our team, what we're known for in the past, we created these models called EVO and EVO2, which are the first generative models at scale trained using LLMs on DNA sequences, raw DNA unlabeled sequences.

4:50Eric Nguyen:One of the things that it's known for being able to open up was this idea of being able to generate novel sequences of DNA, things that did not exist in nature. And just last year, one of the things that scientists had used to generate with our models was the first AI genome from scratch. Genome is the complete set of instructions for an organism. In this case, they created what's known as a bacteriophage, something that did not exist in nature. It's also known as a virus. It's a specific type of virus. It's harmless to humans, but it did create this turning point for us as a company. We realized that this power, this potential to create life itself or to manipulate life itself with AI has both the potential to reinvent pretty much all of biology and also carries this level of inherent risk of that.

5:41Eric Nguyen:We've seen, you know, just the taste of this and the natural language and chatbot side, but imagine obviously giving the power for AI to generate and create life carries a certain level of responsibility. And so as a company, what we push on is making sure that technology is used for both improving human health and pushing the boundaries and frontier of what we call biological AI design, and also the defense. So things like biodefense and biosecurity, which basically means making sure that we're safeguarding against emerging threats that come up with the use of AI in bio. Yeah, I think that tension between, you know, riding this exponential curve of capability, what that means on the sort of amazing positive side of what it can do for people saving lives, but also, you know, what it can do in terms of creating viruses and having that detection is such an interesting part of the business that you're doing and how you have to sort of balance those few things.

6:37You know, I think a lot of folks will maybe be familiar with the lineage of like protein foundational models like AlphaFold and so on. You know, you sort of made it clear to some extent there, but, you know, maybe for to sort of parse it even more, what are sort of the key differences to think about when, you know, someone reads about AlphaFold doing something incredible and what you're really trying to do with Omni, which is the latest model.

7:01Eric Nguyen:Yeah. So Omni, as you mentioned, is our next generation of language models applied to genomics or DNA otherwise. And I think what sets apart the technology that we're building and trying to build is go to the foundation of what we think is the most informative, most plentiful source of data in biology, which is DNA. And from DNA, you hear of other things that sort of are different modalities of DNA, or the way they are basically products that come from DNA. So proteins, very important building block of life. What turns out protein is actually coming from DNA. It encodes proteins in the DNA. And I think what folks typically do in bio and bio AI is that they try to pursue and understand biology one modality or discipline or specialty, if you will, at a time.

8:00Eric Nguyen:And so proteins being a very important building block of what makes us physically ourselves, folks have pushed on making capabilities like protein structure prediction possible with AlphaFold. And I think what we're doing differently is to try to make something far more general purpose. So to move away from this idea of let's learn one specialty at a time, whether it's proteins or RNA or a specific disease at a time, these things should absolutely be pursued. And that's how we've studied biology and science for decades. But I think the point we're at with AI is that we have so much more powerful, capable tools to be able to model the complexity of biology in its entirety, an entire system.

8:44Eric Nguyen:For us, what that means is we're going to essentially not just study one modality or especially a time, we're going to model the entire space. And so that means we're going to look at proteins, DNA, RNA, epigenomics, metabolomics, all of these omics that are typically subspecialties. We think AI has gotten to the point that it can actually build this more unifying system because we're at the stage where we are bottlenecked by being able to not just, you know, understand one modality at a time or build one type of drug or understand one disease, but the complexity of how things actually behave in our body is they're inherently multimodal systems that we are now bottlenecked by not things like protein structure prediction, but understanding how these protein structures are going to behave in the cascade of effects that's going to happen in our body.

9:39Eric Nguyen:And so we, but we lack the tools to be able to model that. And so our approach is really starting with DNA as the foundation and looking at the other modalities as sort of different contexts for which this DNA is used. And so one of the examples that helps ground this is for your, let's say your body, you have DNA that essentially makes you you right that designs and codes information to make all the complexities that you are um basically DNA acts as sort of the general recipe book for your entire body but your you have different cells and different tissues in different parts of your body now the way this basically the DNA is used it grabs bits and pieces of what it needs differently.

10:30Eric Nguyen:So like in your brain cells, it's going to grab these genes, it's going to activate these transcription factors. It's all there in your DNA, but in your liver, it's going to use a different set of genes. And so basically you can think of this different contexts as something that's been lacking of how to actually incorporate all of this complexity of information. What we're doing with our models is being able to train a single system that could use DNA as a foundation, but then be able to ingest all these different biological contexts so that we can model human biology and ultimately disease and be able to treat disease better.

11:08It's sort of a version of the bet that has worked for large language models of just a general model ends up working better on these things. Unstarting with DNA, is that sort of a philosophical preference where you're starting with the foundation and so it sort of just feels like the right beginning place? Is Is that like a pragmatic decision such that when you start with DNA, you sort of get other things for free? Like, you know, that understanding almost gets baked in in a different way. Like why not start with metabolomics? Yeah, metabolomics. Yeah, metabolomics or protein structures, for example.

11:44Eric Nguyen:So I'd say the first analogy I like to use is that we look at DNA kind of like how language, natural language has been the grounding or foundation for intelligence for, you know, general AI systems. You know, technically you can learn quite a bit from just natural language, but then if you want to incorporate visual information, you could describe, you know, a thousand words and describe a scene, or you can have a shortcut and incorporate an image and then fuse that into a model. right? And so I kind of think of it's, that's a nice analogy in the sense of technically in DNA, that is all the information you need to create an organism.

12:25Eric Nguyen:And that means all the complexities of like what diseases you might have and what traits you might have, your height and your eye colors, that's all there. But because of the grand complexity of the interaction with the environment and everything, it's maybe potentially intractable to learn all just from DNA. And so think of the other modalities is like a shortcut to how does that dna evolve in in this you know biological setting well we actually have an imprint of the physical world kind of imprinting its properties via these other modalities and so like one of the modalities will be like the 3d structure of dna turns out the 3d structure of dna you know is a thing dna is also not just a string of letters it's actually you know three it's a representation of a physical thing and the structure of the dna It matters because, well, it's a physical system that has binding that's really important.

13:18Eric Nguyen:So like molecules that bind to DNA and then it changes the sort of like biological processes. Basically, it'll activate and turn on and off genes. So you may have the genes, but it may be inaccessible, like turned off physically because it's not like physically accessible. And so you could look at it as a long 2D sequence or sorry, long 1D sequence. or you can think of it as a 3D structure and things that are nearby, well, it's kind of like a shortcut in terms of information. You just have 3D contact points versus, you know, in a long sequence, you may have two points on that sequence that are very far away from each other.

13:52Eric Nguyen:But in, let's say, a natural language, you have a limitation on context, right? If you have things that are millions of letters away, it might be really hard to, you know, feed into a model from a language model standpoint. But then if you're able to feed in, for example, like an image, you can just show physically. these two things are close and therefore you have, you know, a limited reaction in terms of activating genes, for example. So these kinds of little shortcuts, I kind of think is one practical way for how the products of DNA are used in your body. And then philosophically, I, you know, I think it is an interesting question.

14:25Eric Nguyen:It's a fun question. And I think it's one of those things that is interesting to talk about with folks. Cause I remember when Greg, Greg Brockman from OpenAI, he took a four-month sabbatical to work on Evo 2 with us, which was super fun. And one of the fun questions we would ask folks joining the team would be, what are your thoughts on a universal model? Can you learn all of biology just from DNA or not? It's sort of like a fun thing with like, let's see what you think. And he was in the camp at the time, I remember that he thinks you can learn all from DNA, like all these other modalities, like maybe, maybe not.

15:01Eric Nguyen:It's complicated, but like DNA, maybe that's DNA is all you need. You mentioned this idea of like DNA as a language, obviously, and there are these sort of comparisons with human language, but there are also these sort of key differences that are useful, I think, for people to have in their minds. One is sort of the extent of the context window that you have to sort of play with. Also this notion of like frame shift mutations where this is just like a much less forgiving language. What are the things people should think of when they realize, yes, DNA is a language, but it's actually, you know, it's a very different set of problems that you're wrangling than creating an English language model here.

15:39Eric Nguyen:I think this is a crash course in a lot of other AI researchers kind of run into when they start working in bio or in some of these scientific domains. Some of these analogies and comparisons are useful for conceptual understanding. Like many things, and like, you know, the idea of models and analogies, at some point they break down in terms of, you know, how much they transfer. I think what's useful for thinking of DNA as a language is that this concept of having grammar, rules, structure, I think that's a useful one in the sense that there are patterns in DNA that, you know, as humans, we've learned to map to some kind of function or some kind of attribute in physical property that we know.

16:26Eric Nguyen:anytime we see this series of sequence of letters, it causes, you know, X or Y physical property. And so that mapping is very powerful in many ways. We can understand disease that way. You know, if you have a mutation in your body, it's kind of like having a grammar. A typo, which can lead to a disease, you know, cancer. And what's kind of wild is, you know, thinking of DNA as a language where the stakes are very high. Sometimes in certain key spots, if you have a typo in just one letter, that can mean the difference between a life-threatening disease or not. And then other parts of your DNA, actually a good chunk part of your DNA, these changes don't actually do anything or don't do anything that we know of.

17:10Eric Nguyen:And so this idea of, yes, DNAs can be like a language where the grammar really matters, but then there's so many parts of our DNA that we actually don't understand. So it's like a language that, it's like a foreign language, in our genome, there's 3 billion of these letters. So that's the equivalent of 30 ,000 books. So imagine this corpus of 30 ,000 books where most of the texts we don't understand as humans. Yes. And for us to validate like what actually causes what, like, you know, from a causal standpoint, um, takes, can take months, years or years, basically. So, um, and, and the way people study DNA in our genomes historically, it's, I found it very fascinating and it's something that I like to think of this analogy where biologists, if you want to compare historically how they thought about how to learn about bio and our genome, looking at a car.

18:05Eric Nguyen:A scenario or example of a car where if you compare an engineer versus a biologist, a biologist will learn how a car works by sort of like poking at this car and removing one part of the car at a time and seeing how the car functions. It's like, if I remove this part, does it still run? But an engineer, that's not really how they would think about how the car works. They would usually try to take the entire thing apart and then rebuild it right from the ground up. And so that's kind of analogous to how we've studied the genome. We've basically knocked out or kind of removed one letter at a time and see, did that cause something?

18:42Eric Nguyen:Did that cause a disease? Did that change a physical property downstream? And you can see how this can, you know, this combinatorial effect of one letter at a time can get very intractable and pretty much there's no way to do all these combinations. It's just too large of a search space because it's 3 billion letters and it's basically more combinations of changes you can do than like the atoms in the universe. So this idea is one of the reasons why we got interested in using large language models to study and to really push the field and what's possible to manipulate the genome with large language models.

19:15Eric Nguyen:And so that's one of the motivations that brought us to DNA in the first place, because myself, you know, trained as more of a machine learning person classically and then kind of saw this space as an opportunity to really push the field forward and rethink that, you know, do we want to make these small changes at a time or do we want to potentially generate from the ground up and really understand how life comes together from this bottoms up approach versus sort of top down. So at risk of taking us off piste here, I maybe I read or I heard you talk somewhere before about this, this idea of, you know, the car and the biologist, you know, pokes at the car and the engineer takes it apart.

19:54Do you think that's a function of actually how these fields conceive of truth seeking differently? Or is it just that biology has like a bunch of ethical constraints that mean you can't easily do that? I imagine that actually a biologist of X number hundred years ago would maybe actually take exactly that approach. We dissect animals. We have done all sorts of experiments over the years. Just from a sort of intellectual standpoint, is there something that I'm missing about the way that biology actually thinks about problems there? Yeah, it's a great question.

20:34Eric Nguyen:So, you know, in many ways, obviously oversimplifying, but I think it's a sort of representative line of thinking. I think, you know, to be fair and give credit to obviously science and the great work that scientists have done in the space and biology, there's far more that we don't know than what we know in the space. And so the complexity of being able to build something from scratch, well, there's an implied understanding of, well, if I build something from scratch, I know all the constituent components and exactly what they do. And I think for the longest time and still many ways, we don't understand all those constituent parts, right?

21:12Eric Nguyen:Like I said, in the human genome, most of it, we don't actually understand how and why it behaves the way it does. And it gives a more flavor there for the genome. roughly one and a half percent or just two percent of our of our genome codes for proteins so it's like the physical stuff that is in our body all that is basically encoded in about two percent of our genome and the rest of it is is basically regulating when to turn on and off certain patterns in their dna to make and to differentiate and to sort of like respond to the environment at the right time. So this is all the regulatory parts of the genome.

21:49Eric Nguyen:And that's historically, people have called it the dark matter of the genome. It's like everywhere. At the same time, we don't know exactly what it does. And over time, we're starting to learn more and more of it. Turns out a good chunk of our, most of our diseases are actually sort of like misspellings or typos in this sort of dark matter of the genome. That's just to say that there's a lot of the genome we don't understand. And I think at the same time, we're getting to the point, And I think it's a relatively new phenomenon that with generative AI, with the emergence of large language models, one of the cool things, one of the exciting things is that to be able to manipulate ink to control and generate novel things in that space like language or image and video.

22:35Eric Nguyen:One of the things that I think is really cool is that it showed us that to be able to create new things, whether it's text or an image or a video that did not exist before by humans, you don't necessarily need to understand exactly all the constituent parts of what makes a novel or an image or a video. You sort of the black box of magic of deep learning has allowed us to be able to control something that is clearly has structure and like a certain set of patterns that govern like image space and text space. And at the same time, it's understood to the point beyond our human capabilities. And so I think what's exciting about taking this concept and applying it to DNA, what does that mean for us?

23:25Eric Nguyen:For, you know, folks that care about biology and in human health, it means that this classical way of like, one of them's up approach that I need to understand every constituent part to be able to control and manipulate it, or affect it in some way, you know, to actuate it. We don't necessarily need to do that to have practical use cases and applications of it. Right? So that means we could jump to making new drugs, for example, it could mean that we can understand how to predict and diagnose disease. without understanding the exact mechanisms. Yeah. That being said, I think things like mechanistic interpretability, stuff that people have been pushing in the community to be able to dissect, okay, how did these language models come to their decision?

24:10Eric Nguyen:And using that as insights to understand the mechanisms for which those decisions are made, I think is a bridge for that understanding in biology. And I think that's one of the reasons why people are especially, and we are especially excited about bringing some of these frontier capabilities of AI research like Mechinterp, where, you know, cool that you can use Mechinterp for understanding how, you know, language models work. And, you know, this finding a neuron that is like the Obama neuron, if you like throttle it, you can get every sentence to generate Obama. Wait, I've never heard. Oh, you've never heard this.

24:42Eric Nguyen:Okay. So like people have identified like a specific neuron in a large language model that if you, you know, tuned it up and down in terms of its scaled factor, you can get a sentence to generate, you know, always bringing back the sentence back to, you know, incorporating Obama or not. I can't believe I haven't heard of this. It sounds like it must be a famous moment that's so funny. Yeah. It's like, you know, those geeky anecdotes that people find. The Obama neuron. Yeah. And so finding these equivalents of Obama neurons, but for, let's say, pancreatic cancer neuron, it's probably more complicated than a neuron.

25:16Eric Nguyen:But this idea that we can reverse engineer some of the usefulness and applications of applying AI to, you know, these life science applications like drug discovery and use it also to understand how disease and the science itself works. To me, this is the most impactful and probably consequential application of AI, right? When I make this pitch to other AI researchers, like, yeah, you can work on language, natural language out there at Frontier Labs, but biology has got language too. And it's, if not the most important, it's one of the most complicated languages out there. And its usefulness is unbounded.

26:01Eric Nguyen:This is the frontier in my mind. It's a mix of language meets the physical world and meets the ultimate promise of what AI can do for humans. This is why we're excited about it, in large part why I love to come on these venues to be able to share why I think it's so damn exciting and how much opportunity and green space there still is. If you were to point Fable or Sol at a bunch of relevant DNA data, let's say, where are the most obvious places it would quickly break? To better understand why you really need a dedicated approach to this. I definitely like talking about this or hearing about this topic because it's sort of the natural conclusion a lot of other AI folks jump to is like, let's just feed it all into like a chat GPT type system.

26:54Eric Nguyen:And that's how we're going to understand all these scientific domains. That's how we're going to understand everything. right? And what I would say, and like, this is personal and slash, you know, observational thing for a lot of my team is that these models and the community research community has gotten very good at figuring out the recipes on how to train LLM systems on natural language, images, and audio. I would call these the low hanging fruit. And why do I call it that? These, I would say, especially language, it's sort of by design a very efficient form of communication. Like we've designed it such that sequences pack basically the most amount of information to convey to each other, right?

Read the full transcript

27:41Eric Nguyen:And so being able to extract signal from a sequence, that's what these LLMs are in large part doing, extract signal, learn the patterns of a sequence, it's as good as it gets in terms of efficiency, in terms of how many tokens you need to feed in to be able to train a language model. And I think when you get to the physical world, that ratio of signal to noise, basically, is a completely different ballgame. And you kind of start seeing this like in audio and in video, right, where the density of information is not the same. It's starting to look like it's more sparse. And I'd say when you get to biology and science and scientific domains in general, that sparsity is just orders of magnitude there's way basically way more noise and way more essentially like ways for the model to not learn optimally basically and so these recipes that have been optimized for language you know in some ways you can say technically you can just feed it all into language models but in the same way that people have to build specialized sometimes specialized coding models or math models you can see when the specialty just becomes more efficient when you do target a domain.

28:51Eric Nguyen:I think this is even more the case in biological domains because, well, right now, the way you see frontier labs sort of doing science is that they'll orchestrate and reason in natural language space kind of the way the human does, but then still rely on calling other tools for the raw biological data. So under the hood, it still uses these classic bioinformatic pipelines or computational biomethods to analyze raw DNA sequences or protein structure, they need to call specialized models. And you can talk about this for a while, like why is it the case? A lot of it has to do with also the human annotations that available for some of this raw scientific data is way less available.

29:40Eric Nguyen:It's way sparser. And so I think the limitations on just being able to learn about science through reading textbooks and journals and lab notebooks or something. I think the labs are going to realize the same thing eventually, if not already, that there's going to be a limitation on how much you can learn about science by reading about other people doing science. You're going to have to do, at some point, they're going to figure out it's more efficient for training AI directly on the scientific data. Yes. Right. Because I guarantee you that more of the scientific knowledge or facts or observations are still left in the raw data that are not yet written about in a journal or a paper.

30:24Eric Nguyen:And that is what the frontier models are trained on. It's reading about humans doing science versus doing the science itself. And so that I think is the fundamental limitation that we are trying to overcome. right and i've spent my phd focusing on this problem right that's how i started working on the summing um initially worked on language models and computer vision actually and then got interested in this idea of let's work on something that we can open up new use cases not just do something more efficiently than what was done before or you know new capability capabilities that were already possible but i saw language models as a new potential microscope that could be used into biology and peer into it in a different way.

31:08Eric Nguyen:That's a topic on its own, too, about how we first started using language models on DNA, but just a little preview there about how I got started. Yeah, I'd love to actually spend a little time on what led you to this point. You have this article that you wrote about your experience preparing to give a TED Talk, which I really enjoyed. And in it, there was a really sort of interesting tidbit where you talk Talk about having a conversation with sort of someone at the forefront of the free range parenting movement and saying, you know, here's how I was brought up. And them basically being kind of astonished about the level of freedom that you were you were given.

31:47So I'm curious, like, yeah, what was your your upbringing like? Like, where do you think those things from the way that you did grow up maybe inform how you build or run a company?

31:58Eric Nguyen:Yeah. So when I gave this TED Talk, there was this other speaker who had this topic she gave about free-range parenting, which basically just said, like, let your kids go and grow and kind of learn on their own. But then when I told her about my childhood and how my parents kind of let me go out and, you know, not come home until 2 a.m. in middle school, so pretty young, it was pretty free-range as well. And then when I told her that, and she's like, oh, okay, that's a little extreme. Yeah, that's not even free-range. That's wild. Like, whoa, do they not care about you? So I thought that was pretty ironic that she thought that.

32:32Eric Nguyen:But yeah, for me, I think, yeah, my upbringing was different in the sense that my parents weren't really hands on. They definitely weren't like the tiger style parenting. And I think that did let me have a lot of time to explore things on my own. And then maybe that kind of carried on too for longer than I have liked. because I think before I started working on this company, I did bounce around a decent amount into different careers. I started off in civil engineering for undergrad and first masters. Then I worked as a management consultant for about four years in the energy space in this area of strategic sourcing for construction of power plants.

33:10Eric Nguyen:Super niche. How did you survive four years as a management consultant? Whenever I meet a founder who spends a long time in consulting, I'm genuinely like, how did you do it? Because there's some part of your brain that clearly is super high agency. And I don't want to be too harsh, but it's like fundamentally a job where you have much, much, much less agency, right? Totally. It was like also at the same time a humbling experience because of that. It's like you're at the mercy of clients a lot. I think what kept me going on that was I was always driven by things that I was not particularly good at.

33:41Eric Nguyen:I was always curious about like, I want to be better at that. And so at the time, I was an engineer, mostly in this idea of being client facing and understanding how, you know, quote unquote business works, I was drawn to it. I kind of wanted to understand how that world ticked. And so I thought this was a good way for me to peer into that. At the same time, you know, after that four years, actually, that was a turning point for me. I took after that role in consulting, I took a four year, what I call a medical hiatus to undergo about six major surgeries of my shoulders, spine, and chest. Oh my gosh.

34:14Eric Nguyen:And so that was, um, that was a very intense period and basically gave me a chance to reset and, you know, kind of ask myself, uh, if I got better, what would I actually want to do? What's, what would be more meaningful to me? Unfortunately, I did get better, um, after, after these surgeries. And, um, well, I think that was a, one of the, in hindsight, more thankful experiences to help me really calibrate and find something more meaningful and purposeful and really stick to it because it kind of forced me to have this downtime. And then that's when I pivoted into tech. I started coding and teaching myself to code online.

34:51And then I went to community college and sort of gradually built up and

34:56Eric Nguyen:had this idea to somehow try to use technology to improve human health in some way. I didn't know how, but that was my sort of my chip on my shoulder. I was like, I feel like I should have gone through all this for a reason. Like, you know, I'm pretty good at math and technical stuff. So maybe I can do something in this health space. Eventually, after doing a master's in computer science and working at Facebook a little bit, I went back for a PhD in bioengineering. So quite a bit around windy road and going back for a PhD in my late thirties and finishing in my early forties. Not the most common path.

35:30Eric Nguyen:In some ways, this feels like quote unquote my purpose, or it just feels like this is something I should be working on. And when I think about all the things I could be working on, I feel fortunate that I can work on something cool as tech and AI, but also something I truly believe that this can be applied to and should be applied to. And other folks also think it should be applied to. And it's a matter of how and can we get the brightest minds to work on it together? I actually look at it as a privilege to be able to work on this and try to get other folks on this journey with us. And so now I feel really thankful that we get to try this and take a big swing at it.

36:07Eric Nguyen:So yeah, that's been a little bit of the journey so far. And yeah, I think part of it is having that freedom, I suppose, or being used to that kind of wandering and flexibility and not being tied to, I need to do XYZ by certain age. I need, you know, whatever kind of rigidity to feel on track. And I think that, especially that medical hiatus I mentioned, that was definitely a forcing function for me personally. I hope it's not too personal, but definitely a forcing function for me to be taken off that track, that career track or whatever that right track is, and really reassess what is important to me, what is my identity, especially when you're not thinking about work as being that identity because I had to look at from that perspective.

36:57Eric Nguyen:And so asking myself, what do I truly value? If I don't care about titles and money or prestige, if I just had the ability to work on it, what would I do for free even? And I just kept coming back to this area. I want to make technology tools that people just did not think were possible before and apply that to things that reduce human suffering. That's as good as it gets to me. It sounds like you were living in significant pain for a long period of time. And I don't know if it was exactly like this for you, but I do find that in so many ambitious people's life, there is almost a moment like that, that perhaps acquaints you with your mortality in a different way, gives you a different relationship to suffering in a different way than you're used to.

37:48That sort of recalibrates purpose, ambition that makes you think, yeah, if I ever get back to the stage where I'm able to do something at 80 % even, I will make sure it really matters. How did you get from that point to PhD is the right thing? In some ways, there could have been probably a number of things I imagine that swum around in your head at that point?

38:14Eric Nguyen:I think there's that inner nerd in me that sort of uses school and learning as my expression or outlet. And so I think the PhD or just, I should say school in general, grad school, I always looked at it as a privilege, right? In the sense that you get to go back and focus on learning, right? And focus on like your own learning. And so I think for me, I saw grad school as a way to kind of get back on track, quote unquote, and then also use it as a way to explore and try out things. Right. And so that, that felt right. And, you know, I think for everybody, it could be different. It could just be jumping straight to the thing.

38:57Eric Nguyen:But I think there is a level of go back to the thing that felt comfortable or like what you, I thought was potentially good at before, which was, yeah, I did like learning. I did like this idea of being an expert in the space. And I think ultimately what I found that I thought was most rewarding for me was I wanted to work on things that I didn't think would exist if I didn't work on it. Which is a hard bar sometimes because a lot of people can work on a lot of things. But I think that was kind of the bar I kept on asking myself throughout the PhD. And even before the PhD, it kind of kept on going back to if that's the bar, it kind of seems like you need to be an expert or understand technology in this really deep way.

39:50Eric Nguyen:And so it checked a lot of boxes for me. I think sometimes people ask if they're thinking about grad school, why did you go back for a PhD? And sometimes it's like one answer, but I actually had many checkboxes, many reasons why I thought it made sense. And so it was multiple things that felt like PhD was the right thing to go after. pursue the time. And so, yeah, so I decided to make the leap. And also, you know, worth mentioning that just because you make the decision to do a PhD doesn't mean like, it's like clear, like, oh yeah, it was the right thing to do. I think, for example, it's always fun to tell some folks that the first half of my PhD, I was constantly wanting to drop out like every other week.

40:31Eric Nguyen:Oh yeah? It was, it was like the beginning of COVID, like the first, yeah. So it's 2020, first, first cohort that was all COVID, hard to meet people, hard to get this, you know, ball rolling in terms of like feeling like you belong there because I was, you know, older student. The work itself was challenging, like going back to school after many, many, many years. And AIs can be a difficult topic to study. Yeah. Constantly questioning like, you know, why should I be here? And does this make sense? But at some point, like halfway through, I met my co-founders at Radical Numerics Now and sometimes these types of things they just click and I started having a blast with the PhD.

41:15Eric Nguyen:In large part my teammates are brilliant co-founders are brilliant and they're just good people as well and so I think this mixture of finding the right people finding the right thing to want to apply it at this idea for us in this case was taking large language models and applying it to bio and DNA in particular. I came at it from a technology standpoint. I thought my teammates were interested in this long context problem. So long context a while ago wasn't a thing quite yet. People weren't thinking about context rot and all that stuff. And then Chris Ray's group at Stanford were really pioneers in thinking about how to extend context length and why context length matters in AI systems.

41:55Eric Nguyen:So we had an architecture called Hyena that my team developed and tried initially on language and then noticed that it was particularly good at long context. But me being in my program, which was actually bioengineering, we started thinking, what are the longest sequences out there? Let's push this to the limit. Yeah, we can keep on working on language, but cool, longer context, maybe some capabilities, use cases emerge, but what if we open up an entirely new thing? And so DNA was that thing, right? We saw DNA initially as a proof point for our long context model. So it was like we had this hammer.

42:35Eric Nguyen:Yes. Let's hit this nail. And then quickly, we basically saw it work out of the box pretty strongly. And when we showed it to folks in the domain, they basically told us like, that's cool. Can you do more of that? And then we basically leaned into it. So at first it was, can we get these models to just read DNA? So just like one way. And then, so that was a paper called Hyena DNA. my first breakout paper basically and then the evil models which um were which which had more notoriety afterwards but the question there was can we get these language models to write dna right uh mind you like we still you know this is in the regime of uh we as humans don't understand all the grammar and dna and so this idea of writing dna it just wasn't like especially from scratch like, you know, for nothing, was not done.

43:27Eric Nguyen:Like nobody had done this. And actually, I remember going around pitching this idea for like six months at Stanford to scientists in different labs. And pretty much everybody told me it was stupid, which is hilarious. Yeah, because in hindsight - Even at Stanford. Even at Stanford. Most people thought it was stupid. I heard it all. People would say things like, why would you do that? We as humans don't even understand how DNA works. How could you tell the AI models if you're right or wrong? or the data distribution of a type of data is not possible. It's too noisy. There's too many weird things going on.

43:59Eric Nguyen:There's too much dark matter in the genome. What can a language model learn by just next token prediction in DNA? Like it's not human language, not the same rules. And then the other one that I think is fun to bring up too, people would say like, even if you could, like why? Like who? We don't do that as biologists. What we do more often is we make small edits to DNA. Like, you know, that one letter change at a time and see what happens. We don't create from scratch. Like there's just limitations on that from physically in terms of like how to synthesize it. And so there's, I heard all left and right.

44:31Eric Nguyen:And then in my mind, just seeing the transformation that these language models did on in language, I truly believed in, you know, in hindsight, it's always easier, but just this idea that you could ask like why and what could you use it for? Or you can ask, yes, but if you could, what could that enable? Like, what could happen? So I just kept on being motivated by that. And then eventually found folks that were interested enough in that. And that led to EVO 1. And then that captured enough imagination. And then EVO 2 brought in Greg Brockman and NVIDIA and Jensen. And it took a life out of its own.

45:07Eric Nguyen:And so we were off to the races at that point. And now with the company, we didn't want to stop there at just DNA. The thinking was DNA is one of the modalities of biology. It's a very key one. And we started there because it's got pretty much the most plentiful amount of data. It's internet scale. And the way I like is that just for DNA, there's more DNA publicly online than all the text on the internet. Just with DNA. Yeah, I read that you mentioned that somewhere. Every revolution in AI creates one question that never changes. Can you trust the output? AI for work is incredible, but without trust, it's just leading to faster mistakes.

45:52The challenge isn't building an AI that can answer questions. It's making sure those answers are right. That's where Guru comes in. It's the AI source of truth that connects everything your company knows. So every insight, every answer, every recommendation is grounded in verified knowledge, not outdated information or hallucinations. When your teams and your AIs share one trusted foundation, everything moves faster, with fewer redos, fewer blind spots, and more confidence in every decision. Because in the age of AI, truth isn't just power, it's protection. See what Guru is doing for thousands of companies like Spotify, DHL, and Stripe at GetGuru.com.

46:34That's GetGuru.com. Where is that data? And also what I was curious about is, is it equivalently useful? Like, is it varied enough or is it, you know, I don't know, much less valuable for some reason? Like, I'm trying to understand like how much that analogy holds.

46:52Eric Nguyen:Yeah. Yeah. It holds in the sense that, yeah, in terms of raw tokens, it's orders of magnitude more. Orders of magnitude more. Wow. Yeah. Now the question of usefulness, I think this is the question, right? Right. This is... something I wish I knew the answer to. And it's a useful question, but I don't think it's the most important question. What I like to bring it back to more so is, do we think we've extracted value to the limit basically so far? Is there more we can grab from the existing data? Yep. Is there more useful things we can build from that data? And I think for that question, the answer is 100 yes um and then it becomes clear in my mind that the opportunity right there's this massive amount of data and where i think to get some balance on like the opportunity um you got this internet scale data for one of the modalities and there's more modalities that are even more than dna so it's it's just this is in my mind the next vast ocean of data um when you look at the models in in bio basically the algorithms and the expressivity power of these models being applied to biology, they're basically as strong as linear regression or logistic regression.

48:08Eric Nguyen:In my mind, when you have that vast amount of data and you have something basically not that powerful, there's a lot more on the table. Basically, this is the vast green space that's open for mining, basically. In hindsight, at some point, I think this is going to be obvious to folks. I think like Evo did for DNA, we want that same thing to emerge. We want that after the fact we want people to be like of course you train models on dna it's like language and like before the models were really weak so of course that you know that makes sense but like i said remember i would go around people talking to people and people thought it was a stupid idea so that's not obvious and i think we're at a similar moment for the multi-omics the rest of the modalities in bio where people do kind of learn training one model at a time per mode per domain yes and this idea of that we can actually integrate all of it into a single single unified system Our goal is to make this obvious to folks very soon as well.

49:02So taking a step back, very simply, like, where do you get that data? How do you make sure you're able to harness it to the extent you need to? And when you think about the data for some of these other omics, like, is that availability sort of similar scale as DNA? Is it, I don't know, half, twice, you know, just sort of the rough orders of magnitude here?

49:25Eric Nguyen:Yeah, I would say for the other modalities, other omics, there is more than DNA. So DNA is on its own, it's the biggest, but when you combine all the other modalities, it's going to be far more than the DNA. And that's because, well, you'll hear biologists say that biology is hard. a lot of things that we can't measure or it's difficult to measure. And so to get the true state of a system, let's say. And so I look at these other modalities as basically different sensors of the same thing, of the same physical world. And so because no one of them is perfect and they all have their trade-offs, people constantly come up with new modalities and they call them assays and these experiments to basically capture the physical properties of the system, but from like a slightly different angle.

50:10Eric Nguyen:And so you have a new modality that captures 3D structure or like different physical bindings to the DNA or proteins. And so every time a new modality comes out, people are like, this is going to be the key to understanding, you know, cancer or neurological disease. But I think it's less of like, you know, is this going to be the answer? The single modality is going to be the answer. It's more like all of them, right? So it's like, in my mind, when you are, let's say, trying to triangulate or trying to like, you know, use GPS, you have a bunch of sensors that like help you triangulate the system of say, like, you know, location, you don't optimize like one signal at a time, you combine them all to triangulate on the true signal position.

50:53Eric Nguyen:And so in biology, sort of analogous to that, like you don't want to use a single dialed and optimize the heck out of that. But you want to fuse them all, right and so now the next question is how do you do that and i think that's been the blocker from a technological standpoint that we as an ai lab are attacking like that's the capability that we want to enable this generalization this fusion of modalities that you've seen parts of it applied to language and vision and audio sometimes you know three domains in bio there are dozens of of modalities. And so I think this is the ultimate multimodal problem.

51:30Eric Nguyen:But are there open repositories of this data that you can start to use? Yeah. So in terms of data, there are massive amounts of public data. And I think what people like to say in biology that there's a data shortage, there's a lack of data, and there's a data problem broadly. So I agree that there's a data problem, But my hot take is there's not a lack of data. There's too much data in bio. And the reason why, before people kind of flip out and kind of make their claims otherwise, is that when you have a domain that is larger than the internet, and you have models that learn from that data, barely tapping into its potential, you know, basically as strong as linear regression, I wouldn't say that you're going to dig your way out of that by getting more data.

52:20Eric Nguyen:I'd say, I would argue that if someone or anybody in the world, you gave them all the data, all the data that you wanted in the world, if not more, I would posit that folks would not know what to do with it. Like actually how to leverage, how to build a recipe for training and scale and extracting signal in a way that scales, meaning more data, more compute, you actually see improving performance. Instead, what you constantly, constantly see in this space is tapping out and plateauing, right? And this is why it basically tops off something like linear regression. Save for AlphaFold, that's the one I think that's very successful.

52:58Eric Nguyen:But every other domain, you'll see this similar behavior. And unsurprisingly, right? Because even in general AI domains, language is pretty much the only domain you've seen that continues to scale, right? Otherwise, pretty much tap off. So unsurprisingly in bio, it's, you know, even sooner because it's, it's a very different type of data distribution. That's why we, as a lab, we focus initially on the modeling side. Yeah. Right. The technology and the bottleneck we think is actually being able to extract meaningful signal at scale from this vast amount of public data. But also we, even though I said there's like too much data, it depends on how you use the data.

53:39Eric Nguyen:And things like you get more nuanced in terms of there's obviously pre-training steps and post-training steps. So like from the pre-training step, I think there's definitely more data than that people know what to do with. the post-training, I think this is where I still think we could use more and there's a, you know, difficulty in the data. This is where data quantity, but also the labels and annotations to be able to align the language models, right? To do useful things. I think that fine distinction needs to be pointed out, right? In the same way that OpenAI and Anthropic, they don't start with a bunch of like labeled data and like try to like generate a bunch in-house.

54:20Eric Nguyen:No, they mine a bunch of public data. And then they post-train, well, in many ways, they post-train a bunch of their stuff on the proprietary stuff, the expensive stuff, right? I think what you're seeing in bio typically is people try to get around this sort of poor, you know, limited model capability problem by just using supervised training to dig their way out of it. And I think this is how you, one, you bank drop your company because generating supervised data is very expensive. For us, it's more intuitive to start with, okay, how do we actually squeeze every drop at the pre-training scale?

54:49Eric Nguyen:And then we'll take the, you know, the high valued clinically relevant labeled data as our post training. That's, that's how I think about it. So it's a much more nuanced about like, you know, is there enough data? Where do you get it? Obviously for us and for many people, more data is better. So we'll grab data everywhere. We'll partner the heck out of folks to grab data from biobanks, um, national labs, uh, pharma companies, nonprofits, um, self and biobanks. I think that people underestimate the value of creating deals around data. But data exists and showcasing how it can be leveraged better is the way we want to come about it.

55:30How much do you benefit from just underlying improvements in general models? Do you get some benefit from the fact that Mythos exists and all these sort of fundamental improvements just keep progressing on an exponential curve? How does that factor in as an input for you?

55:50Eric Nguyen:We are, in many ways, other labs benefiting from coding agents, right? So being able to parallelize many agents to do the experiments that we want, each researcher on our team is multiplied by many. And so we absolutely benefit that way. Being able to dive into a domain, not being an expert in, say, cancer vaccines. our teammates are also the barrier to entry to domain specific applications is way lower basically and so i think this just means that we're able to go after way more horizontal applications at once and so i think this is one of the reasons why perhaps companies like us didn't really find traction in the space.

56:40Eric Nguyen:So like some context there, there have been large bio-AI companies that tried to describe the scaling hypothesis in bio, but really hard to find traction in commercial viability and stuff like that. And I think now this is, in our mind, far more possible. And if not just possible, but like inevitable in many ways. And we want to showcase out this theory. And so we benefit like other folks that use these tools to basically do a lot more with less folks. And I think this is playing out very quickly. And hopefully we get to showcase to folks how we benefit from this as well. I'm curious at what point the running a sufficiently like industrialized wet lab operation becomes the bottleneck.

57:30Because if you get to the stage where you have enough data and you're able to generate potentially viable genomes, then I would imagine verification suddenly becomes the piece that's, I don't know, hardest or most bottlenecked. Am I thinking about that right? Or how do you think about that part of the process here?

57:50Eric Nguyen:Absolutely. So that's one of the key differences with the physical world in AI models in this space is that when the rubber meets the road, that's where you really get to know if the systems work or not. So yes, absolutely. This is a bottleneck and this will be a bottleneck for some time. For us, we initially partner with a lot of folks, but I think our timelines for when we do our own wet labs are accelerating because we have been able to do a lot more on the modeling side because of the pace of AI progress and our own capabilities that are just sort of accelerating in terms of emerging capabilities.

58:25Eric Nguyen:So we are likely to do our own wet labs pretty soon as well. There's luckily an ecosystem to support this. So there are these things called CROs or contract resource organizations that essentially you can outsource some of the wet lab capabilities for the more common types of experiments. And then the more bespoke things tend to require some in-house things. But yeah, we're in the same boat for a lot of these things. The idea is to do less and less of this over time as the capabilities of our models get better. But I think the first thing we want to make sure we do is sort of what I described, like extract the most amount of value we can from existing data.

59:05Eric Nguyen:And then from that regard, there's just a vast amount that I think is still untapped that we still take advantage of. And I think we're still on the relatively small scale where we need to validate with our own like feature-wide experiments. Another way to get around this is doing retrospective experiments, meaning take results that have been already published and see if you can recapitulate, meaning if you can come to the same conclusions. That's another way to kind of expedite things because if you can show that, you know, your models independently can learn those things and have these emergent capabilities, you kind of, it's like a bridge to doing some of these well-wide experiments.

59:41Eric Nguyen:And I read this one up because in our last release where we previewed Omni, our next generation models, this is indeed one of the things that we thought was really exciting. We applied it to an Alzheimer's set of experiments where this paper basically spent two years trying to understand what in your DNA, which genes are most causal to cause Alzheimer's. And it took a couple of years of really painstaking wet lab experiments for us we had a new teammate join like just a few days in and they used our model omni and just um maybe a couple days into it they're able to basically recapitulate the same results and rank which genes were the most causal to to alzheimer's and i think that was just a taste i think that was one of the because it's never seen alzheimer's specific labels and data, but it's between reading all the genomes in our training set, we're able to do some mid and post training to tease out its knowledge and get it to predict and rank which genes are most likely to cause Alzheimer's.

1:00:48That was extremely exciting for us. Have you sequenced your own genome and tried to extract any interesting information from it yet? Not yet yet for myself. Yeah, I think that'd be a fun exercise.

1:00:59Eric Nguyen:And it's something that we've been thinking about to come up with ways to make this a little more real for people. I think one of the challenges sometimes when working at AI and science level, it can get a little abstract, right? And this idea of like, oh, a lot of promises being made in the space of, you know, we're going to use AI to cure XYZ and cancer. And I think folks are, and I'm certainly one of those folks that's hungry for concrete examples to showcase, okay, what can it do now? What can it do for me personally? And I think we're getting close to showing some very, very compelling results that not yet quite ready to show folks now, but hopefully in the next X number of months that we make this more real for folks and get folks to really reimagine how powerful and beneficial this technology can be.

1:01:54Eric Nguyen:be. Well, I just ordered a very extensive genome sequencing set of tests for myself. So I offer myself this tribute when you want to do that. Maybe before we wrap up, you talked about some of the difficult challenges other companies in this space have had on the commercial side. How have you sort of sought to address that? I think what makes this especially more excited and think that we're well positioned to go after this, the breadth and the general applicability of what we're building, what we're calling a general biological intelligence, really is analogous to how I think the Frontier Labs rolled out things like ChatGPT and Claude, which is that they were able to make this tool that can be applied to a lot of things, and then they sort of let the users tell them what is the most useful and then kind of lean into that whether it was you know health questions for some of their medical questions but eventually things like coding became very clear that is where a lot of tokens were being spent and then and then once they get validations from you know people asking for more on that capability then they started building their own and you know verticalizing a little bit i think there's a little bit of that sense for us because we can build such a general purpose engine and system, we are very horizontal now.

1:03:22Eric Nguyen:We're going after a huge diversity from applications in pharmaceuticals, drug discovery, diagnostics, what's known as synthetic biology, so like designing whole genomes, for example, for us, and in biodefense. These four very, very different spaces, you'll pretty much never see a single company working on these. But this speaks to how general and capable we believe the technology that we're working on will allow us to do. And that purview of horizontalness lets us really tap into what's emerging and what is in most demand in this real-time fashion. And so I think this is going to enable us to respond very quickly and keep our, you know, mini small bets kind of open.

1:04:08Eric Nguyen:But really the focus on being able to accelerate other people's capabilities is what we're most excited about. And I think this is something that people, it's newer to folks in this space, in the life sciences. I think this idea of needing to verticalize and make a drug and take it all the way to clinical trials, like that's the only successful business model that people have been able to prove out. And I think for us, this idea of building an AI research lab, an AI research lab first grounded in innovation is our bet. At the end of the day, we're betting on innovation to have a step change of the technology to actually make this horizontal model work.

1:04:51Do you have any early hunches of where you think you're going to see sort of, I don't know, the hottest uptake on that? Like, does this look like a biodefense company for the first five years and then, you know, the capabilities catch up such that it becomes something totally different? Or yeah, how do you, what's your current picture?

1:05:09Eric Nguyen:There's two buckets. There's like the read tasks, being able to read DNA and other modalities, and there's the write. And I'd say the read, kind of like other domains in AI, being able to read and make predictions off of like embeddings is a first step and it's a little bit on the easier side. Being able to generate and write stuff is a little more challenging. So Yes. For us, I think things like being able to interpret our genome, understanding what functionality or what causes DNA in our genome is going to be useful for pharma in understanding what targets, what things to go after to make drugs for in the first place.

1:05:49Eric Nguyen:It's actually a very key area that in and of itself has typically not been a multimodal problem, but we think it's going to be inherently a multimodal problem. people that have the capabilities. So I think that ability to read and then point a lot of R &D funding for making drugs on the right targets, right diseases, and understanding where to put R &D dollars is going to be one of the first things that we're especially interested in. Things around cancer detection, cancer diagnostics, it's also a read task. This idea of characterizing like, okay, given a DNA sequence, what properties, what physical properties, what pathogenicity or things associated with disease are in it.

1:06:26Eric Nguyen:This also works really well for the biodefense side as well. So being able to see DNA and to understand is it a dangerous bioweapon or not, for example, is it naturally an API style kind of business. And so we think that is a compelling early type of use case that we're seeing a lot of folks reach out to us, for example. And then gradually, as we add additional modalities, as we have more controllability of our generation and generative models, then the idea of making acids or drugs and with multimodality in mind is a thing that we feel we have an advantage of as well. So not only just being able to design the drug, but being able to predict how it's going to behave in a system as complex as a cell, tissue, or human body are things that become more capable and possible with this thesis around multibitality.

1:07:20Eric Nguyen:And I think at that point, you're looking at the entire biological space, everything around life sciences. A couple of questions on the biodefense or biodetection piece. I think it was with EVO2 probably where you guys talked about how you purposefully kept some of the more dangerous data out of the data set. I think viruses that could infect a human, essentially, right? Yeah. And the reason for that is basically that if you added that capability and it would be a more capable model, but also more capable of creating a virus that could hurt humans. Does that fundamentally also limit its capability on the detection piece?

1:08:01Like how do you find that balance of, we want to include this because it makes it a better detector, but we don't want to include it because it also makes it more capable of creating something dangerous?

1:08:10Eric Nguyen:Yeah. I think this is such a good question and something that has evolved for us, for sure. When we first started Evo and Evo 2, we were PhD students, right? We were an academic situation scenario. The strong motivation is to open source these models, right? One of the first steps is to do what we can to safeguard them to some degree. And that we believed was appropriate to remove these eukaryotic viruses that can infect humans. But I think one other thing that is worth mentioning is that that's a type of safeguard, but there are still potentially dangerous sequences in that database. You know, certain genes and proteins from bacteria pathogens are technically still in there.

1:08:56Eric Nguyen:And so as much as one wants to safeguard everything you can, it's pretty hard to, right? And so I think for us as a company, this was a key question. And so for how we release our models now, we're much more conscious of how we provide access to the models. Like you said, the defense side and what you put in to have capabilities to be able to detect, you inherently have this conflict of being able to generate and detect. It's in large reason why we believed one company needed to work at this intersection because in our minds, it's inseparable. If you're going to build models with such capabilities to control the substrate of life, you're going to lower the barrier to making dangerous things.

1:09:48Eric Nguyen:Even if you leave the data outside the models, people can fine tune pretty simply, actually. And people have shown that for Evo too, that it's possible. And so for us, we need to push on both. And we think it's important to push the capabilities of the defense. And at the same time, folks who work on the defense, we think that you have to be on the frontier of the design capabilities as well. Otherwise, you won't know how to defend against them, right? So this inherent duality is why we have this dual mission with the company to push on both design and defense. And I think this idea of open source, it's our natural inclination.

1:10:26Eric Nguyen:Like we actually, when we first wanted to showcase our pathogen detection system on the defense side, our motivation was we don't see any tool out there publicly available on the AI side that can detect dangerous pathogens. There's no AI models really publicly available. And that seemed like such a huge gap, right? We initially were like, let's just make an API and just host it. Like this, for us, this should be pretty simple to do. and then when we started engaging with biosecurity experts, they basically told us like, don't do that, right? You're basically giving people a verifier potentially to make dangerous things.

1:11:06Eric Nguyen:So if you tell them yes or no, this is dangerous or not, that is a great reward signal. And so we've gone through a lot of soul searching and thinking like, what is the right way to roll out these two capabilities? Who are the right partners that should and should not have them? And so I don't think it's a binary moment where we figured it out. I think it's going to be evolving. I think as the capabilities evolve, we too evolve on how we roll that out. And so I think like Evo2, where at the time it felt like we're going to get more benefit from releasing it publicly, it's going to evolve how we move forward with the company right now as well.

1:11:41Eric Nguyen:And right now what it feels, what we believe to be the right balance is making sure first that are allies in the US government are on the forefront of having this defense capabilities to sort of even out this natural arms race that the offensive side design capabilities, they're just so far ahead that they can obfuscate and make things intentionally to get around detection methods. And so we need to bring up the defensive side to even out the playing field. And then this is going to constantly be this battle or this conflict, these sides going after each other. And I think depending on which is dominating over the other, then that means the strategy on how we release these things also adapt.

1:12:32Eric Nguyen:And so that's kind of our stance for now. Do you get the sense that the government, when you're talking to them about these sort of risks are fully alive and fully awake to them and sort of prioritizing it? Not as much as it should. I think opening I, Anthropic, and DeepMind, it's been on their radar more recently, and I think there is for better or worse, it's a component of fear that folks are instilling in folks, which have a little bit of a conflict around because ultimately I'm an optimist, but also I think if it's on people's radar, then there's a healthy level of concern. But right now, overall, I'd say there's a lot of talk and maybe fear-mongering, but without the action involved.

1:13:21Eric Nguyen:Yeah. So I would challenge that, you know, there is a trend of being concerned of biological risks and biological threats, but in terms of meat behind it, not a lot. I think there's folks that want to like limit how much you can talk about biology, like in some of these chatbots, but going from a natural language standpoint, understanding people's questions is one level. But at some point, you kind of need to understand the raw scientific data itself. If someone tells you, hey, this is not a virus, and then uploads a virus, you shouldn't be blind to the actual raw DNA sequence. Otherwise, you're just your hands tied behind your back, right?

1:14:05Eric Nguyen:And so that's the layer that we think is sorely behind. And there's a lot of focus on safety in other ways, but then they'll stop there. And so I think this is something that folks in DC and policy are starting to get educated around because folks in biosecurity are becoming more vocal and there's a light shined on this community. And so I think the awareness is starting to grow and there's definitely folks on Capitol Hill that we've engaged with who share more and more of this concern. Eric Horvitz, CSO of Microsoft, who's an advisor for us, has been a proponent of pushing biofense and biosecurity at this wider scale, not just in the US, but globally for decades.

1:14:48Eric Nguyen:He's been a proponent in the space for a long time. And so, you know, more voices like that, who, you know, these trusted scientific folks, you know, bringing up the potential for these AI systems to build potentially dangerous things. And I think, you know, this idea of agents going around the internet and running around and making all sorts of things, there'll be billions or if if not trillions of agents running around. These scenarios of like, oh, what happens when we give them access to these tools starts to become more real as they become more and more capable and people's lives. And so it's a concern that I don't think is going to go away.

1:15:22Eric Nguyen:And so I think people, you know, I'm an optimist in the sense that I think people will wake up to these concerns and we will rise to the challenge. And then when that happens, that's a great spot for us to showcase. Hey, look, we've been thinking about this from early on And we're interested in pushing the frontier of this technology, not just for design, but defense. And hopefully that gives folks the ability to then trust. Do you have a sense, if you and I were to have this conversation a year from now, what would you like to be different about the way this is handled? Regulation is one component in being able to identify some of the choke points where the biggest risks have been identified.

1:16:06Eric Nguyen:really i think for now i think awareness is a really good step and like taking it one step further than just the shallow awareness of like it's a concern but like show me where and um then i think once you have this awareness in a year my hope is that the focus is on the solutions um especially what we're excited about because you know we're biased as an ai research lab a focus on the research. And I think what I would love to share with folks is that the challenge, the implications, the potential impact of using frontier technology in the biodefense side, there's so many cool problems that can be applied in some of these technologies.

1:16:51Eric Nguyen:And they're relatable to things that people in the AI community are already familiar with. So things like, because I used to work on deep fake detection at Facebook, and this idea of attributing and watermarks and images, how do you know if something's been AI generated or tampered with? These are very similar techniques that we can use on the sequence of DNA side. And this is one of the things we showed in our Omni preview. We can use these models to attribute. We can use them to detect. my hope is that we build up a research community around these problems and it's it's kind of how we hoped it to be for for dna before with with our evil models for better or worse you know make for better i mean making it a a sexy space because i think it can be a sexy space right it's chat it's one of the most challenging things a huge impact if we can if we can showcase you know where it works and where it doesn't it just needs a little bit of light to be shined on it and also a bridging of the language.

1:17:49Eric Nguyen:When we can showcase there are generative problems, there are mechintert problems, there are all sorts of interesting ML problems, but people have mostly thought it was too domain specific. And so in a year from now, my hope is that these are as common at ICML and NeurIPS as all these video generators or whatnot. Well, I always like to end with a thought experiment or to, if you had unlimited resources and no operational constraints, what is an experiment you would like to run? I think scaling on DNA would be such a killer thing to run. And by that, I mean, let's collect every genome from humans and some of the annotations from it and be able to train models on things like cancer detection and modeling the evolution of cancer.

1:18:43Eric Nguyen:Because the evolution of cancer is basically DNA is mutating and then sometimes they go wrong and sometimes it's harmless. I think it's one of those spaces that if you had enough data, you can map out the entire manifold of cancer, manifold of disease. But let's start with cancer. That would be a sweet experiment. Yeah, that's an amazing one. If you could assign everyone on earth a book to read, what book would you want to make everyone read and understand? Oh, man. This one's going to be embarrassing because I don't read very much. Oh, my gosh. How could that be? You're such an intelligent person.

1:19:19Eric Nguyen:You don't read very much? I read. I read. My attention spans like other folks. I read short form. Oh, come on. I'm shocked. I mean, I read like, you know, art like M.O. articles. I'm such a word cell snob. So, yeah, that's fair. Not everyone has to read. Is there anything that you, you know, from the past that you're like, you know, that was really great? If you want to recommend a paper also, feel free. There's a book on my list that I've been meaning to read. I don't even know the title of it. So this is maybe not a great suggestion. But I know there's a book on biological warfare that talks about how during the Cold War, the you know folks in Russia and in the eastern bloc had um a bunch of bioweapon programs and during the end of cold war there's a whole uh movement to to de-weaponize and to kind of decommission these sites and then it goes into understanding like the scale in which these programs are run these are just like massive factories basically built building things that um you know are blocked from the Geneva Convention and many things uh but I think I've been recommended by many folks that this is an eye-opening book in terms of the potential threat and the scale and potential secrecy behind this type of things that in some ways can help some folks in Washington shine a light on this.

1:20:49Eric Nguyen:And so it's one of those things that it's on my wish list to read soon when I find some time. And in the meantime, recommend other folks as well. Yeah, that sounds interesting. I was just reading a book called Blitzed about the creation and usage of different pharmaceuticals during the Third Reich and how the German army was deploying so much. They were using so many different types of, yeah, of different drugs, methamphetamines to sort of like, you know, increase performance. Not biological warfare, but certainly, yeah, a confluence of chemicals and warfare. Anyway, it's a very interesting book.

1:21:30Yeah, sounds a little terrifying at the same time.

1:21:31Eric Nguyen:Oh, it's very terrifying. Yeah, yeah, yeah. Eric, thank you so much. This has been such a pleasure. My pleasure. Thank you so much for inviting me. And I had a blast. Awesome. That's it. Thank you for listening to this episode of The Generalist Podcast. Please subscribe on Apple Podcasts, Spotify, or your preferred podcast app. Ratings and reviews help others discover these discussions. So if you enjoyed the conversation, I'd be grateful if you could take a moment to leave one. For all past episodes and more, visit us at thegeneralist.substack.com. See you next time as we continue to explore the future.

From the publisher

Eric Nguyen is the co-founder and CEO of Radical Numerics, an AI research lab that has raised $50 million to train models directly on biological data. Before starting the company, Eric helped develop Evo and Evo 2, large-scale genome language models trained on unlabeled DNA sequences. Radical Numerics is now building models that can connect information across DNA, RNA, proteins, epigenetics, and other parts of biology, rather than treating each as a separate problem. Researchers have already used Evo to generate viable bacteriophage genomes, and Eric says Radical Numerics’ newer model, Omnii, matched key findings from two years of Alzheimer’s wet-lab research in a matter of days. He also believes these tools could make it easier to create dangerous pathogens, which is why the company is working on both biological design and biodefense.


In our conversation, we explore:

  • What AI models can learn by treating DNA as a language
  • Why reading scientific papers is not the same as learning directly from biological data
  • How Eric’s unusually free-range childhood shaped the way he follows his curiosity
  • Why biology may have more useful data than researchers know how to use
  • How Radical Numerics plans to connect information across DNA, RNA, proteins, and other biological systems
  • Where the company sees early opportunities in drug discovery, diagnostics, synthetic biology, and biodefense
  • Why testing AI-generated biology in the lab is still slow and difficult
  • How models that design biological systems could also help detect dangerous or manipulated pathogens
  • How to make powerful biology models safer without eliminating the capabilities that make them valuable

—

Thank you to the partners who make this possible

Ahrefs Brand Radar: Find your brand in AI results.

Brex: The intelligent finance platform.

Guru: The AI source of truth for work.

—

Timestamps

(00:00) Intro

(03:35) An overview of Radical Numerics

(06:35) From protein models to modeling all of biology

(11:08) Why they started with DNA

(15:04) The process of mapping DNA as a language

(19:47) What’s unknown, and how we learn from novelty

(26:24) The limits of language models in biology

(31:15) Eric’s free-range upbringing and path to his PhD program

(41:20) Applying long-context models to DNA and meeting his co-founders

(46:36) Biology’s untapped data opportunity

(49:02) Why biology needs multimodal AI

(55:30) How better general LLMs benefit Radical Numerics

(57:19) The challenges of biological verification

(1:02:05) Making biology more concrete

(1:04:51) Radical Numerics’ strategy and early use cases

(1:07:26) Balancing safety with capable AI models

(1:15:47) What success in biodefense looks like

(1:18:09) Final meditations

—

Follow Eric Nguyen

LinkedIn: https://www.linkedin.com/in/nguyenstanford

X: https://x.com/exnx

Website: https://erictnguyen.com

—

Resources and episode mentions: https://www.generalist.com/p/ai-got-good-at-language-now-its-learning

—

Production and marketing by penname.co. For inquiries about sponsoring the podcast, email jordan@penname.co.

More from The Generalist

All 50 episodes
AI Got Good at Language. Now It’s Learning the Language of Life. (Eric Nguyen, Co-Founder and CEO of Radical Numerics)The Generalist · 1 h 22 min
Listen in VO