AI Can Now Write DNA. What Does This Promise, And What Could Go Wrong? With Eric Nguyen, co-founder and CEO of Radical Numerics

12 Aug 2026 · 51 min · 22 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI models that “read and write” DNA as language to design gene-editing tools, generate whole viral genomes, rank disease-causing genes (e.g., Alzheimer’s), and support biodefense by detecting “deepfake viruses” whose DNA spelling is altered to evade safeguards.

Guest backgrounds

Eric Nguyen is co-founder and CEO of Radical Numerics. He helped create Evo, a DNA language model, and the company is building Omni, a multimodal genome model. He and other founders met during PhDs at Stanford doing AI research; they worked on long-context sequence architectures (e.g., Hyena).

Key claims

DNA can be treated like a language model (next-token prediction over A/C/T/G). Long-context scaling (up to millions of tokens) enables organism-level design. Multimodal context (proteins, RNA, epigenomics, metabolomics) is needed to predict side effects and environmental interactions. Compute is the biggest bottleneck.

Notable examples

AI-designed CRISPR-Cas systems synthesized and shown to cut DNA; Evo-generated a bacteriophage (“first AI genome”); Omni used in ~30 minutes to rank genes most causal to Alzheimer’s from raw DNA; “deepfake virus” concept described as AI-altered DNA that remains effective at infection.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to Gene Editing Tools

0:00 to 1:00

Explore the excitement surrounding advanced gene editing tools and their potential.

“They spent two years on very painstaking experiments.”

Understanding AI's Approach to DNA

2:12 to 3:40

Learn how AI models read and write DNA similar to language.

“Well, I guess let's start really simple.”

Advancements in DNA Context Handling

3:40 to 5:29

Discover the evolution of AI models capable of processing extensive DNA sequences.

“And so large language model, just like other types of AI.”

Applications of AI in Creating Organisms

5:29 to 7:51

Examine how AI can generate new organisms and its implications.

“It's the longest context you'll ever find, generally.”

Balancing Benefits and Risks of AI in Genetics

7:51 to 11:23

Discuss the ethical considerations of AI-generated organisms in medicine and potential threats.

“What's a good example of one type of organism that you're targeting or one type of, yeah, like what are you trying to do with that first, I guess is a good way to put it.”

Balancing Benefits and Risks of AI in Genetics

11:49 to 12:45

Discuss the ethical considerations of AI-generated organisms in medicine and potential threats.

“AI agents inside enterprise environments are growing 40 % year over year, and 7 % of organizations already had an AI agent-related security incident in the past year.”

Gene Editing Techniques and Applications

12:50 to 14:00

Understand different methods of gene editing and their medical applications.

“to being able to make these kinds of changes and edits to someone's DNA, uh, is, is to give better resistance to maybe cancers, maybe other diseases, maybe aging.”

Exploring DNA Manipulation and AI's Role

14:00 to 17:44

Learn how AI can assist in understanding and predicting the effects of DNA editing.

“You can make DNA kind of like standalone that's like separate from your body.”

DNA Context and Modeling Complexities

17:44 to 23:07

Understand the importance of contextualizing DNA within biological systems.

“And so what we've done as a company is essentially figured out how to tokenize or get into a form that's useful for large language models to be able to read those modalities.”

The Future of AI in Genome Understanding

23:07 to 28:01

Discover the challenges and potential of AI in interpreting genomic data and biology.

“But in many ways, it's sort of similar language in the sense that just, for example, DNA, looking at DNA publicly available online.”
Show all 22 chapters

Emerging Capabilities in Biology and AI

28:01 to 28:34

Explore the scaling behavior of AI in biology and the new frontier it presents.

“And so that's the opportunity that we see that other folks have largely ignored, in my opinion.”

DNA-Only Models and Their Capabilities

28:35 to 30:20

Learn how DNA-only models can design proteins and RNA, showcasing their unique capabilities.

“Like what kind of things do you see that you weren't expecting, that you didn't train?”

AI's Role in Alzheimer's Research

30:21 to 32:12

Discover how AI models can identify genetic factors related to Alzheimer's more efficiently than traditional methods.

“And then later for Omni, for what we were interested in was like clinical relevance.”

Interactions Between DNA and External Factors

32:13 to 34:10

Investigate the role of external factors in genetic diseases and how they can be modeled.

“But could a DNA model like you're building also help identify the external factors when they interact with the DNA that are harmful and how to mitigate those as well?”

Scaling Challenges in AI for Biology

34:11 to 36:26

Understand the current bottlenecks in scaling AI applications in the biological domain.

“So what would you say is the biggest bottleneck right now in terms of scaling it, as you were talking about earlier?”

DNA and Longevity: Hardware vs. Software

36:27 to 37:57

Examine the relationship between DNA and longevity, and the role of epigenomics.

“Like it's it's the main it's one of the main promises that AI says and sort of hopes to do.”

Understanding Deepfake Viruses

37:58 to 40:44

Delve into the concept of deepfake viruses and their implications for biodefense.

“I think being able to program and reprogram some of the epigenetics is something that we absolutely are excited about and very excited to showcase what we are thinking about and what we're doing over this year.”

Proactive Approaches to Biodefense

40:45 to 42:01

Learn how AI can be used both to create and to detect biological threats proactively.

“And what we're trying to showcase that's possible and really change the dialogue, which is you don't have to just wait for this to happen and then respond.”

Using AI for Biosecurity and Countermeasures

42:01 to 45:49

Learn how AI can help in detecting and responding to biological threats.

“Like, are they, like, and if so, like, how could they or anyone else, be they researchers, be they government scientists?”

The Promise of Precision Medicine

45:50 to 47:57

Explore the potential of AI in creating bespoke therapies and the future of precision medicine.

“It's gone through different hype cycles.”

The Future of Gene Therapy

47:58 to 49:58

Understand the current state and future possibilities of gene therapy powered by AI.

“And I think it will make it more possible over the next five years.”

Concluding Thoughts and Resources

49:59 to 50:38

Wrap-up discussion with resources for further exploration of biological AI.

“In the limit, we would have to figure it out, right?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Eric Nguyen:What's really exciting or was really exciting last year that folks really took us to the next level and it was part of our initial goals which was creating these gene editing tools which are very complex and very useful and interesting. The most ultimate machinery is in tests is can you create a whole organism from scratch and indeed scientists actually did use Evo to create the first AI genome so a complete set of DNA for what's known as a bacteriophage. They spent two years on very painstaking experiments. Our teammate spent 30 minutes using our model to be able to rank the DNA from these subjects and identify this and rank the same genes that were most causal to Alzheimer's.

0:41Eric Nguyen:It's all based on the spelling of the DNA. So they'll take a sample, they'll sequence it, and basically take the letters and match it to a database. Now, a deepfake virus, there's a line of work that showcases that it's possible to use AI to essentially change the spelling. Welcome, humans, to the Neuron AI Explained. I'm Corey Knowles here with Grant Harvey. How are you today, man? Doing good. Doing good. And I'm especially well today because we're talking about a really interesting topic. So most generative AI writes text, image, or maybe code. Well, today's guest, Eric Nguyen, helped create Evo, which is a model trained to understand and generate DNA.

1:23His new company, Radical Numerics, is building even more capable systems for genomic research, disease detection, and biodefense. I want to learn more about that. Today, we're going to unpack what these models can actually create, how close they are to producing useful new medicines, and whether our ability to detect AI-engineered biological threats can keep pace. Before we get started, please take just a quick second to like today's video and subscribe to the channel so you never miss an interview or one of our live streams. Real quick, I want to take a moment to thank the sponsor of today's video, BeyondTrust's AI agent security.

1:56You can learn about them later on in the episode and in the link in the description below. And on that note, Eric, welcome to the Neuron. It's great to have you.

2:05Eric Nguyen:Great to meet you guys. Thanks for having me on the show. And very excited to talk about what we're building here at Radical Numerics. Excellent. Well, I guess let's start really simple. When you say that an AI can read and write DNA, what is it actually doing and how is that similar and different from how ChatGPT reads and writes language? Yeah, great question. So actually, in many ways, it's very similar to how ChatJPT works. Conceptually, what we're doing is treating DNA also like a language. And in many ways, it is like a language. It's got sort of grammar, structure, in terms of vocabulary, but it happens to have a very different set of letters, right?

2:53Eric Nguyen:It's got these letters A, C, T, and G. But those are just really stand-ins for what they represent, chemical compounds. But they do have a certain way that where you have them basically ordered in a specific sequence, it changes its function. It changes its properties. And so it makes those changes basically following a set of rules. And for humans, it's just really been hard to understand what those rules are because, well, it isn't natural language. And so the hope is that using AI, we can essentially train it to understand these underlying physical rules. And ultimately, if we can understand, then we can control.

3:33Eric Nguyen:And that's going to help us understand how to treat disease and ultimately improve human health. For people who are a little more technical, is this a language-based model or is it an entirely new type of model? Yeah. So it is a language model. And so large language model, just like other types of AI. Now, how you get it to work on DNA, that's sort of the trick of the trade, because it is a different type of data distribution. That being said, what we, our team is known for doing, we, the founders of Radical Numerics, there's four of us, we met during our PhDs at Stanford doing AI research. And we were working on essentially novel architectures, new models to handle sequences, not just language, but novel types of sequences.

4:25Eric Nguyen:And what our models were particularly good at, we called this model at the time hyena. It's a new type of architecture. It was very good at long context, which now is a very common vernacular that people are familiar with in language models. How long are we talking? How many tokens here? Yeah. So back in 2023, the longest context models at the time for DNA and language were on the order of a few thousand. Like in DNA, it was like 2 ,000 tokens. Wow. We had created the first language model that could handle a million contexts. Wow. Yeah. And so now our newest models are now at two million. But back then we were pushing the envelope already to a million.

5:10Eric Nguyen:And that was about 500 times longer than previous language models on DNA. And so what we started to get people to feel that was possible was this idea that we could train AI to understand DNA, which is basically the longest sequence of them all. It's the longest context you'll ever find, generally. What's like, just for context, what's the average DNA strand? How long is that? In an order of magnitude, let's say. Yeah, so the best example is us, right? So we are made of DNA ourselves. All our cells have DNA, the instructions to encode an organism. That is 3 billion base pairs. They're called base pairs.

5:55Eric Nguyen:Base pairs long. And so the way we tokenize and feed our language models DNA is one token equals one letter. And so that means you have a 3 billion token context that you ultimately want to be able to read. Yeah. Wow. So I'm curious, what does the model understand about, like, biology as opposed to just thinking of, like, predicting the next likely sequence? What does it understand about the world around it? Yeah. As you said, so it first basically learns like language models. It's like this, you know, think about next token prediction or in some cases mass language modeling. But through this pre-trained self-supervised task, what it's able to basically learn is this implicit understanding of the rules and grammar of organisms in life.

6:53Eric Nguyen:And so we've trained it on DNA across the tree of life, like pretty much every type of organism on Earth. And there's a lot of shared information and information about what an organism is, what it needs to survive, what is conserved, meaning what is basically most useful for survival and what's passed down. and specifically for us, we can take advantage of those implicit things it's learned and essentially you can think of it like prompt engineering the model to be able to generate new stuff, like understanding the rules and grammar of DNA. Can you use that to then essentially autocomplete new functionality or new organisms with the hope that you can essentially create medicines or treatments for people that did not exist before.

7:48Eric Nguyen:Wow. That is so cool. That is. What's a good example of one type of organism that you're targeting or one type of, yeah, like what are you trying to do with that first, I guess is a good way to put it. Yeah, absolutely. So starting with the maybe chronology of some of the interesting things that we were able to create with Evo, the first example was being able to generate the first CRISPR-Cas system using AI. And so CRISPR-Cas is this gene editing tool that exists in nature. but also there are many more in nature that we haven't discovered yet or we might want to come up with synthetic ones that have different functionality and being able to they're like molecular sources they can edit dna and what we're able to do is essentially train evo on a bunch of existing crispr cast systems and then essentially ask the model all right given what you've seen before, can you generate something more efficient or efficacious?

8:54Eric Nguyen:And so essentially by prompt engineering the model, it's in many ways like hallucinating or like coming up with new variants. And then the scientists that we're working with actually took those designs and synthesized them and physically created them in the real world and tested their functionality. And indeed, they were able to cut DNA. They're able to work and function like editing DNA, just like other types of CRISPR-Cas systems. So that was the first. And I think what's really exciting or was really exciting last year that folks really took us to the next level. And it was part of our initial goals, which was not just creating these gene editing tools, which are very complex and very useful and interesting, but sort of like the most ultimate machinery is and test is can you create a whole organism from scratch?

9:41Eric Nguyen:And indeed, scientists actually did use Evo to create the first AI genome. So a complete set of DNA for what's known as a bacteriophage. It's basically, in other words, a virus, which when folks hear about that, a couple things come to mind. You would be like, no, we should not be doing this. This is what Daria is always talking about, like trying to stop. Yeah. So naturally people have this reaction, right? Yeah. Yeah. And so, yeah, peeling back a couple of layers, bacteriophages are harmless to humans. They cannot attack humans and they actually have beneficial use, meaning because they attack bacteria, they can essentially be used like antibiotics.

10:25Eric Nguyen:And so they do have beneficial use cases. So in this particular case, what's exciting is that there's a potential to create these precision custom antimicrobials. microbials. And I think that was exciting for the scientists to showcase that, hey, we can create something that did not exist in nature. It has human benefit. But also, right, there is this inherent threat where, yeah, if one can make such an organism, have such control over the manipulation of life itself, right, then the idea of creating something harmful becomes very top of mind for folks, right? And I agree. In many senses, being able to lower the barrier to make biology controllable, programmable, inherently has this inseparable ability to make things that are dangerous as well as beneficial use.

11:22Eric Nguyen:And so that's why as a company, we actually have a dual mandate for us to push on the capabilities of biological AI design, for the beneficial use, but also make sure that it's used on the defense side, on the biodefense side. And this was a key turning point for us, this being able to generate a whole genome, in this case a virus, that we saw this ability and also opportunity to apply these models and get ahead of it and make sure that these are safe in biodefense applications as well. Okay, quick security reality check. AI agents inside enterprise environments are growing 40 % year over year, and 7 % of organizations already had an AI agent-related security incident in the past year.

12:01The danger here isn't that agents can invent some new permissions. It's that they can weaponize the permissions that are already sitting there. So an agent may launch with the same cloud keys, SHH keys, and login tokens as the person running it, often with no scoping, no exploration, and nobody watching. That's why BeyondTrust built AI agent security. It discovers every agent running in your environment, including shadow AI that nobody approved, and traces each action back to the human or agent behind it. Then it steps in before damage happens by blocking destructive commands, limiting credential access, and requiring human approval for sensitive moves.

12:36One policy works across Cloud Code, GitHub Copilot, Cursor, and more. So be among the first to secure your AI coworkers before they act. Visit beyondtrust.com and learn more about Beyond Trust's AI agent security. I'm curious. I assume the long-term approach. to being able to make these kinds of changes and edits to someone's DNA, uh, is, is to give better resistance to maybe cancers, maybe other diseases, maybe aging. Uh, but what I'm wondering is, is this an approach where this is editing a piece out, replacing it? Is this adding a new piece in that strengthens resistance for it? Is this, uh, How does something like that even work?

13:25I realize that's maybe a little grammar school, but I want to make sure I understand.

13:30Eric Nguyen:It's a great question. I think this is such a foreign concept for folks that may not be in the sciences, and so totally reasonable. I think what's first kind of setting the stage to how these models work, they focus on the design side, right? So these models are generating sequences and helping to control or program functionality into DNA. And so how you use that, that knowledge, that design capability can take many forms, right? You can make DNA kind of like standalone that's like separate from your body. And you can think of that more like a drug or a separate biological machinery on its own.

14:11Eric Nguyen:And then also you can make something all the way to a whole organism like that. There's a huge range. And so that's one category which is kind of like separate from our bodies, right? So like not on our raw bodies. And then there's a whole category of, well, if you want to manipulate the DNA of something like an animal or a person, that's a whole other category of like gene editing, right? So then that's a category where we as humans now have the ability to edit our DNA in vivo or in a cell or body. I think where AI interfaces with that ability, right? Because AI doesn't have the physical components.

14:52Eric Nguyen:But what it can do is tell you if you make these edits, what are all the cascading side effects that can happen or that you can anticipate? Right? Because that's the biggest fear. One of the biggest fears for gene editing is that if you make an edit where your intention is to cure a disease of some sort, what are the unintended consequences that cascade and that you can't account for. And so being able to know what edits to make and how it plays out is what AI can help with. That's awesome. Yeah, that is so cool. And is that something that we can, I guess it's something we can't really fully know at the DNA level, but we can get a lot closer to knowing it, right?

15:35Because I assume at some point you're going to have to simulate the entire body and how it works together as a system before you can really predict all of those consequences. And each unique body is different, right? So we probably have to do it for everybody at some level. This might be a naive way of thinking about it. But how are you thinking about that problem of like, can you handle it all at the DNA level? Do you have to model things at the systems level? How do you think about it?

16:03Eric Nguyen:Yeah, such a great question. and actually a really big motivator for what we're trying to create at Radical Numerics. So, and it's also a little bit of a philosophical question, this idea of like, is DNA all you need to know to understand everything about your body? Actually, some people say yes, some people say no, right? So this is actually an open question in some ways. For us, I think DNA is the foundation, is the starting point. But DNA is not in a vacuum. So it's interacting with the world. It's interacting in your body. And how the DNA is used in your body, in different parts of your body, in different time periods of your body, this is sort of measured in different – what I call different modalities or different ways of manifesting.

16:55Eric Nguyen:Manifesting like the right kinds of proteins that are made from certain genes from your DNA. Gene expression is changed by a lot of different other physical factors around your DNA. And these are captured very differently than, you know, somewhat similar to DNA, but in different contexts. And so what I see as helpful in understanding is how do we provide that extra biological context with DNA so that we can understand how the complexity in the system of our bodies, how it behaves when a disease happens or when a drug is taken. And being able to model all of these other functional properties has been what's missing for some of these DNA-only models.

17:44Eric Nguyen:And so what we've done as a company is essentially figured out how to tokenize or get into a form that's useful for large language models to be able to read those modalities. And so specifically those modalities are like proteins, protein structure even, like RNA, epigenomics, which is like the really environmental kind of measurable things around your cells. Metabolomics, all of these things are basically, in my mind, somewhat like different languages that have been sort of learned individually on their own throughout science and biological history. But now for the first time, we have this ability with potential with AI to not model one specialty or one siloed modality at a time, but to integrate it all into a single system that can actually measure the complexity of something like our bodies.

18:37And that's what we're building toward. And I think approaching it from a long context perspective is the key because to be able to do that, you need to be able to hold so much context in your head. me anthropomorphizing the data centers here or whatever, like the models. You need a data center in your head, Joe. Yeah, you need to hold so much information together at one point. I feel like$3 billion is just the starting point for how much context you need to do that.

19:06Eric Nguyen:Yeah, yeah, absolutely. Yeah, so we like that analogy of like, oh, it's different contexts, especially for biology. Like people think of context usually in these language models in a similar way. Like the more context you feed to a language model or chatbot, you know, the more specific it's able to answer your questions, the more it can adapt, right? Like, oh, for example, if you ask a chatbot to like write you a story, let's say, if you ask it to just broadly write you a story, it'll write something, but it would just kind of go for, you know, a bunch of directions. But if you ask, if you give it more context, like, oh, I want a mystery story set in London in the 1950s, you're providing more context and it's able to generate things that are basically following that regime, that pattern.

19:49Eric Nguyen:And in similar ways, what we're doing on the language models for biology is also feeding additional context, both for DNA, like sequence lengthwise, but also in these other modalities. So like give it context for a specific cell type. Give it context for a specific disease state. And that also shows up in different similar sequence information from biology. It's like it's not language. It's like this physical world property kind of stuff. In many ways, it can be kind of thought as like time series kind of cell data for models. But anyways, essentially for a lab, an AI research lab at Radical, we've figured out how to tokenize and discretize the data so that it's fed in a native way to these language models, that it can actually learn all of these contexts so that it's like in a certain disease state.

20:38Eric Nguyen:It understands what the methylation patterns are for someone who has Alzheimer's. Alzheimer's. And another person who has cancer, this is what the shape of their DNA through chromatin accessibility looks like. And it's shown up in these modalities and different phenotypes, molecular phenotypes through sequence information. And so being able to feed those different contexts, you now have a big context adapting machine to these different states of biology. That's what I was wondering is I think the idea of how do you pre-train a model on the volume of DNA it would have to take, I presume, in order for it to be able to do what you're doing.

21:24I mean, like, that has to be just a monster undertaking. And simultaneously, I can't help but wonder, is, like, any of the Human Genome Project stuff, is that a thing that enables this in some way that was done over the years?

21:38Eric Nguyen:Absolutely. So those are a great question. So the Human Genome Project was really one of the first steps for this endeavor, right? The Human Genome, what it was trying to do was basically trying to come up with the actual sequence of letters that makes up the genome. And we thought, you know, as humans that once we could do that, this is going to unlock. We've unlocked the world. We've unlocked the world. We're going to solve biology. We're going to understand everything that there is about. We've won biology. Yeah. We've solved. It's solved. Yeah. And so I think what was, you know, it was extremely successful in many ways.

22:16Eric Nguyen:It has opened up a lot of directions and made, you know, a lot of treatments possible and understanding of biology far possible. But it did have its limits, right? Yes. It's able to show what letters are. You can physically, like, read what letters there are in our genome and the exact orders of the letters and combinations. But being able to interpret what those letters mean was really the next step and still is a giant step for what we need and have been doing to really understand human biology. And so that's been the challenge left open. Now, being able to train and hopefully use AI to live up to that challenge, as you said, it can be in many ways a very big undertaking.

23:06Eric Nguyen:And one of the reasons is that the scale is a different scale in terms of data and context length. But in many ways, it's sort of similar language in the sense that just, for example, DNA, looking at DNA publicly available online. There's more DNA online than all the text on the internet. So it's like – it's internet-scale data. Whose DNA is this? So it's not necessarily human data. There's human data, but mostly it's other species. So it's the tree of life. You know, organism, microorganism, mostly microorganisms actually. Mammals, plants, fungi, archaea. So it's the diversity of life you can capture online, essentially.

23:56Eric Nguyen:It's not annotated. So people also in many ways don't know what the grammar is, right? And so one of the hopes is that being able to use a language model, you can distill a lot of that underlying grammar into the models. And then the challenge is then how do you figure out how to do something useful with it? And that's a big challenge for us and an opportunity for us as an AI lab to be able to harness what these models can learn and ingest, but then be able to tease out what it's learned and do something incredibly useful and beneficial with it. That is the name of the game. Yeah, because if they could annotate it, even just the existing data, if you can use AI to annotate the existing DNA that's known.

24:45yeah it's unlocking a whole other level if it could just go back and add definitions to everything based on everything it's learned this is what this is, this is what that is that would be pretty cool give me an index then you can see all the patterns across everything and then maybe you do have a key to unlock the tree of life yeah

25:04Eric Nguyen:this is an active area people do indeed do this there are companies a non-profit in particular comes to mind that works on annotating DNA from nature using AI models, ingest and be able to feed in annotations of known things and leverage that to annotate the world, annotate the biological world. Very, very ambitious endeavor. What's your philosophy as far as the future of how far this goes at the language model level? Are you imagining like one big God model to kind of like, like, you know, kind of like the anthropic model where you train just like a massive, massive size model and then slowly you distill that down?

25:49Or do you think that there would be a future where there's lots of smaller models, like you make like a model specific to the organism that you're trying to create? Maybe we make a human model. Maybe we make a butterfly model, whatever. And then even down to the point where could there ever be a model of me so that my doctor could essentially have a model that helps knows everything inside my body? Corey knows that too. Maybe it's a model that if you go the God model route, maybe it's a fine tune on top of the existing main model. But what's your philosophy on how that's going to play out? Yeah.

26:28I think what's incredibly exciting for biological sequences is the scaling laws that we're seeing is exhibiting similar emergent

26:44Eric Nguyen:capabilities as language. And I don't think people know it for language necessarily either. I think people have been scaling obviously constantly right now and still are. And people still wonder, is that going to be tapped out at some point? So, like, ultimately, I don't think people do know what the limitations are. It's just that current trajectory. It seems to be. Keep going, right? I think there's a similar sense in bio that – well, actually, I'd say in many ways people think that things don't scale in bio. But I think what we're going to showcase as a lab very soon, hopefully, is that there is a lot more scale that people have left on the table.

Read the full transcript

27:21Eric Nguyen:Like we showed this for DNA initially with EVO 1. And I think for EVO 2, things started to somewhat taper a little bit, or at least people didn't know how to scale it. And so what we're doing as a lab now is creating this multimodal version of these genome language models. And I think when we have multimodal versions, it's really able to provide that additional context to continue that scaling curve. And I don't know what the ultimate limit is, but we absolutely do believe right now they can be far, far larger, orders of magnitude. And so we've only seen small scale so far. And the things that we're able to see and have emerging capabilities right now across modalities in biology makes me feel very optimistic that we can see similar scaling behavior to language, which means a long way to go.

28:16Eric Nguyen:And so that's the opportunity that we see that other folks have largely ignored, in my opinion. They've really focused on language and code, for example. And I think what's exciting is that this new frontier language of biology, in my mind, is clearly going to be the next frontier in AI. What does – and I have to ask this – what does emergent behavior look like from this type of model? Like what kind of things do you see that you weren't expecting, that you didn't train? Yeah. Yeah, a great question. So I think the first one for Evo, which really just was a teaser, for the longest time, people have thought that to understand biology, you have to really focus on one specialty, one modality at a time, as I was kind of mentioning before.

29:09Eric Nguyen:And so there were models that were like focused on proteins. So you have things like AlphaFold and ESM, they're like protein language models. And then you had separate models for like RNA. And then you also had separate models for DNA. And the thing that we first showed and were aspiring to show was that, hey, technically all these things come from DNA. Like if you're an organism, all you have really is your DNA to start with and everything else is made from it. And so can we – what's the limit of what we can learn just from DNA? And so that's why we fed – that's why we trained a DNA-only model.

29:41Eric Nguyen:And what we were able to show was that, indeed, these language models on DNA can be competitive and understand proteins, can understand RNA entirely in one system. So that's the first capabilities. And we showcased that, like, you can design biomolecular shineries that have the ability to design proteins and RNA simultaneously. So that was the first time that was done. All the way to, you know, Evo was able to do the phage, right? So that's the first organism. So you can't do that with a protein model. You can't do that with an RNA model. You need a DNA model that can understand across all those modalities.

30:20Eric Nguyen:And so this ability to learn across things beyond DNA was sort of the first emerging capability. And then later for Omni, for what we were interested in was like clinical relevance. Like, okay, cool. You can do these things for science. What can emerge that's potentially useful for human health? There was this one project, an application that we showcased last month in our preview where we had a new teammate join like two days before. and they took Omni and they applied it to this, the model to this data from a paper, this paper that was trying to understand Alzheimer's. And this Alzheimer's paper was trying to identify what parts of your DNA, which genes, are most likely to cause Alzheimer's, right?

31:05Eric Nguyen:And so what they did was they spent two years doing these wet lab validations where they would knock out genes, knock out and mutate base pairs, and then see what the phenotype, what diseases emerge and what's most causal or like correlated. They spent two years on very painstaking experiments. Our teammate spent 30 minutes using our model to be able to rank the DNA from the data, from these human subjects, and identify this and rank the same genes that were most causal to Alzheimer's. And it did this without being trained to identify or be told what's Alzheimer's versus healthy. It was just trained on a bunch of raw human DNA And we're able to tease out the model's outputs to be able to basically look at the statistics of its predictions and use that to identify these anomalous mutations that are most likely to cause Alzheimer's, which was like such a wild moment to us that a model like that can learn that kind of thing without being told that explicitly.

32:08What would be interesting for me from the human health perspective is there's obviously lots of things that lead to diseases, chronic diseases, health issues that you might have that are genetic-based. But could a DNA model like you're building also help identify the external factors when they interact with the DNA that are harmful and how to mitigate those as well? Like what's causing the cancer or something? Yeah, like sometimes it's from the DNA and, you know, there's mutations and stuff. But sometimes there's external factors that also play a part.

32:44Eric Nguyen:Yeah. Yeah, such a key question. And I think this gets to the debate about is DNA all unique, so to speak. And I think technically it's possible, theoretically, and that's my opinion. But I think it becomes computationally intractable to actually come up with how DNA, you know, predict how it's going to respond to every possible environmental condition. It's too complex. And so what I think make what makes sense more is, luckily, we don't have just DNA, we have all these other modality, we have all this other sequencing information that biologists, scientists, pharma companies, you name it, that they do collect to find potential, basically physical evidence from the environment that's imprinted into these other modalities.

33:33Eric Nguyen:So it's imprinted to things like so like, you know, let's say for cancer, if someone smokes, obviously it has some interaction with the body. It does show up. Those interactions shows up in these different modalities. It may show up in changing gene expression. It may show up in changing the shape of your DNA, which are, you know, in terms of sequencing information, it's like ATAC-seq or chromatin accessibility. And these things can – these are other modalities that can be fed into a model like Omni, Which basically, that's why we use this analogy. It's like you're feeding additional context, but via these other types of biological modalities.

34:10Right. So what would you say is the biggest bottleneck right now in terms of scaling it, as you were talking about earlier? Is it the intelligence? Is it compute? Is it high-quality biological data, maybe simulation or just access to the labs? Or is it our ability to validate what the AI produces that's holding back the scaling?

34:35Eric Nguyen:Yeah, great question. So it's a multifactorial thing, in my opinion. I think that for us, the biggest bottleneck is the compute. We've spent the last nine months researching that technology to be able to adapt large language models to read in these different types of biological sequences. And so now we've overcome that constraint because in my mind, the modeling side was the biggest constraint. Like, for example, being able to – how do you feed in DNA, RNA, proteins, epigenomics into a single model? Like, it's not clear. And I think to most folks, it still is unknown, right? And so we've sort of cracked that in terms of recipe V1.

35:20Eric Nguyen:And so scale, because what we're seeing now is the scaling laws, it's still continuing to improve at this medium scale. And so compute has been the bottleneck. So we're very much on the compute hunt for us. And then, as you said - Got some heavy hitting competitors there for it. Exactly, exactly. Which has obviously its own set of challenges, but - I have a funny aside on this. I wonder if people would be more okay with data centers in their neighborhood if they know it was going towards a project like this, as opposed to just like generating images or minting Sam Altman billions of dollars. Yeah.

36:01Eric Nguyen:100%. Right. I mean, that's my belief. That's my hope. And I think maybe not the one in your backyard, because I think that always is going to get people feeling close to home. But overall, I think it does. I think you'll have a hard time. You'll have a lot more folks being supportive of channeling that compute for, obviously, human health applications. Right. Like it's it's the main it's one of the main promises that AI says and sort of hopes to do. And at the same time, it's what's kind of somewhat disappointing is that for Frontier Labs, it's kind of like a side bet. Right. It's like, oh, we're going to we're going to focus on chat.

36:44Eric Nguyen:We're going to focus on on on coding and then get to that stuff later. Yeah, because that'll kind of fall off just naturally. And so we think that's we think this is too important of a domain to be a side bet. This has to be the main bet for us, at least. Agreed. This is the promise. Like, you know, we're all here because, you know, we'd like to live longer and feel like living longer. I don't want to just extend the bad years. I want additional years of like 35 to 40 is what I want, tacked on to the end, you know. And I say that a little tongue-in-cheek, but not entirely. Do you think solving aging is a DNA problem?

37:26Eric Nguyen:I think ultimately it can be a big breakthrough can come from DNA. But I see DNA in this case.

37:39Eric Nguyen:Kind of borrowing from other folks in the longevity space, DNA is kind of like the hardware component. And then the epigenomics is kind of like the software. In that sense, the DNA doesn't change throughout your life so much. Yeah. But the epigenomics and the other modalities do, and so it's kind of like the software. And so I do see these models, these multimodal models, which is like DNA plus these other modalities, as tools to push on longevity as well. I think being able to program and reprogram some of the epigenetics is something that we absolutely are excited about and very excited to showcase what we are thinking about and what we're doing over this year.

38:25Very cool. Well, I notice a little bit of a topic shift. I noticed Radical Numerics has used the phrase deepfake virus. And I'm wondering what qualifies as a deepfake virus and like how does it differ from something that is naturally evolved? Yeah. So that term especially can sound scary, right?

38:51Eric Nguyen:Yeah. I think what's useful for that term is that folks do have – a lot of folks have that sense of what a deepfake is for like images and videos. And it's this idea that AI can generate something that looks similar to a real thing, right? Like in this case, like deep fake images, like a person perhaps or a voice. I think what's similar to on the virus side, it's not that it actually like looks like visually, but it's that it functions like a virus. But actually, the looks or the spelling of the letters, like the DNA of a virus, is actually changed. And why this is important for something like biodefense in particular is that security and being able to detect if harmful DNA is found in the environment or found in a hospital or found in a research lab.

39:44Eric Nguyen:It's all based on the spelling of the DNA. So they'll take a sample, they'll sequence it, meaning they'll read the letters and basically take the letters and match it to a database. Have I seen this virus before? Does it look like COVID in terms of the letters? Now, a deep fake virus, what we described in our preview model last month is that there's a line of work that showcases that it's possible to use AI to essentially change the spelling, but still make the virus just as effective in terms of infection or virality. And so that makes it very alarming to folks in the national security level or safety in terms of population level because now you have the potential to intentionally design things that can obfuscate and get around any kind of safeguards that folks have in place, which is concerning and only growing concern because AI gets even more capable, right?

40:42Eric Nguyen:And so that's why you see folks at the Frontier Labs like Anthropik and DeepMind and OpenAI mention that biological risks and misuse is one of the top, if not the top concern for more emergent, more capable AI. Ultimately, we're optimists. And what we're trying to showcase that's possible and really change the dialogue, which is you don't have to just wait for this to happen and then respond. You can be proactive. You can actually use the same technology that enables these capabilities to actually defend. And so what I mean by that is that the same technology that can generate potentially something dangerous is actually the same technology that's going to be best suited to detect that danger because it understands the grammar.

41:31Eric Nguyen:It's underlying the same technology. And so what it takes is a group, an entity, a research lab, whatever, that's going to treat this as a first-class citizen. And, you know, being able to not just work on the, you know, pushing out the capabilities, but also, okay, now that it's possible, what do we do about it? Who do we talk to? Who do we work with? Who are the allies? And how do we get this kind of technology in their hands to make sure that it's safe? I am wondering that a bit about, like, who you're working with on this. And, like, are you going, you know, directly? Are you working directly with pharmaceutical companies?

42:03Like, are they, like, and if so, like, how could they or anyone else, be they researchers, be they government scientists? How could they use the model that you're creating to respond quickly to a viral threat, for example? Is there a way that this can help us speed up the time to market for whatever type of cure it is, whether it's a gene editing thing or a vaccine or I don't even know what the possibilities are here?

42:30Eric Nguyen:Yeah, it's a great question. It's something that we are absolutely interested in and working on. So broadly, for biosecurity, biodefense is another term for it, there's a few different steps or pillars, I should say. The first is like detect the threat, be able to identify, you know, from the environment, a sample and see is something dangerous in it. And then the second step is like once you've noticed it's a threat, figuring out where it came from, right? Is it coming from nature? Is it coming from a specific country and a specific lab that's intentionally doing it? And the third, which is where rubber meets the road, what do you do about it?

43:13Eric Nguyen:It's this area of countermeasures. And so the hope is that with AI, you can detect, attribute where that sequence came from, and come up with a design countermeasure and response. in a similar way that people had used this for COVID to design, in this case, not AI, but they were able to design a vaccine for a flu virus, right? And what you want to do here, what we're doing and interested in is being able to come up with systems that could basically take in a sequence of interest in, say, let's say a bacteria or a virus and be able to create something on demand that counteracts that. And I think what's really concrete that we can kind of share about what's possible here and what we're doing is I mentioned about the bacteriophages before.

44:06Eric Nguyen:Like Evo can treat bacteriophages. And so we are, too, working on bacteriophages as a company because they can be used as an antibiotic or antimicrobial in this case. And so a scenario where you, let's say, come up with or find a new bacteria that is infecting someone, and sometimes people get this antimicrobial resistance. So they have some kind of superbug, and they don't respond to existing antibiotics. This is a growing concern, actually, and very prevalent. What you can do is actually design a custom phage or a bacteria phage that can counteract that bacteria on demand. And so see that bacteria, then be able to design a phage that counteracts it.

44:53Eric Nguyen:And there you have a precision therapy that's enabled by AI. And really only possible by AI because of the speed. The bottlenecks are going to be like the manufacturing and the wetland validations. But there are circumstances where this can be sped up too. So we are very optimistic about this. And we're excited to share more about this later. That is really cool. That is really cool. Do you think this could enable personal medicine at the limit? Like where you go into your doctor or you say, hey, I'm really sick right now. And they take whatever measurements, samples they need to take from you and then could design something maybe hundreds of years from now, like the next day or a couple days from now.

45:37A bespoke medication for you? Yeah. Yeah. I mean it's probably not necessary for most things, but for certain things. Yeah.

45:44Eric Nguyen:Yeah, at the limit, this is one of the promises. Absolutely. Precision medicine, which has been sort of talked about for a very, very long time. It's gone through different hype cycles. But I think there is going to be a new emergence and promise of this. I think it's going to be powered by AI. That I do believe. And it won't just be DNA based. So DNA is sort of like the starting point. But you need the real-time ground conditions in these other modalities that was mentioned before. And I think with that combination, if you can actually read a snapshot of someone's multi-omics, is that the way to describe it, that's your best shot at being able to design precision medicines.

46:33Eric Nguyen:And I think that is the ultimate application of some of this technology. Yes. Eric, before we go, I'm curious. Looking, say, five years ahead, what in your eyes is the most beneficial use of biological AI that you think is actually scientifically plausible, not just theoretically plausible? Yeah. So I mentioned that, you know, this ultimate use of this precision therapy, precision medicine is an ultimate, but it's not the only ultimate. And I actually think in five years, you'll see more of this to some degree. And that's gene therapy. So not just a precision medicine, but a precision treatment that isn't temporary, but that is permanent.

47:23Eric Nguyen:And that's being able to manipulate the DNA itself. And this is how AI can really help. And I mentioned before, what's one of the biggest roadblockers to this type of technology that can modify and alter the DNA permanently. And this has been done for things like sickle cell, for example. It's the first gene therapy to be successful. This, in my mind, is only going to be more prevalent, right? The throttle, the bottleneck is, again, knowing what DNA to change and all the cascading effects. And with models like Omni and multimodal genome language models, this is the best chance to make that possible.

48:05Eric Nguyen:And I think it will make it more possible over the next five years. You think there's a watershed moment in that window somewhere where things are unlocked at a faster rate and enabled by? Yeah. It's a tricky one. There may be. And I realize I'm asking for a crystal ball. Yeah. I think it's one of those things where progress is going to feel slow and then all of a sudden it may feel like it bursts the wall. Well, how prevalent is gene therapy today? Like how many people are actually getting it? I mean, a ballpark, if you know. It's not that prevalent. Yeah. I think it's in the U.S. It's probably dozens.

48:45Eric Nguyen:Wow. However, there are hundreds in clinical trials and waiting to be approved. And so that's what it feels like in terms of it can feel slow and then it can also feel all of a sudden. But it's anyone's guess about when that moment is. All right. Last question for you. Any theories on the origin of life?

49:10As a DNA scientist, you're the closest I've got to someone who's an expert in this area.

49:15Eric Nguyen:At this point, it's like, what is life, man? Origins of life. I can't even begin to think about what that even best guess is. But, you know, I hope that we get closer with it by using AI and using some of these models to get a deeper understanding of life, you know, on Earth. But then moving beyond what even other possible life forms there are and yet alone understanding the origin of all life. That's something I wish I had a better pulse on. Love to be around for that one. Yeah. I love it. Just solve aging and then we can all work on it together. That's right. That's right. In the limit, we would have to figure it out, right?

50:01Eric Nguyen:That's right. Well, Eric, thank you so much for joining us today, man. It's been fantastic. Thanks so much, Corey. Thanks, Grant. Where can viewers go to learn more about Radical Numerics and the work you all are doing today? Yes. We try to make a very concerted effort in building a research community around this biological AI that we're building, especially around the biodefense. And so we do have a series of blogs that describes our technology and our applications. And so RadicalNumerics.ai is where you can go. Excellent. Last thing we want to thank again, the sponsor of today's video, Beyond Trust's AI Agent Security.

50:34You can learn more about them by clicking the link in the description below. If you like today's video, please take just a moment to like and subscribe. Don't forget to check out the Neuron's other projects, including our daily newsletter read by more than 700 ,000 people just like you. The Neuron Academy as well, where you can go and learn about AI and how to use it in your work and your life, as well as our newest sister newsletter, Robotics Insider. Thank you for joining us. We hope you'll be back next time, and that's all we have for today. So farewell for now, humans.

From the publisher


AI can now read DNA like language, generate working biological sequences, and help scientists identify disease-related mutations. Radical Numerics CEO Eric Nguyen explains how genome language models work, why his earlier model Evo helped scientists design CRISPR systems and bacteriophage genomes, and how the company’s newer model, Omnii, combines DNA with RNA, proteins, and other biological signals. The conversation explores personalized medicine, gene therapy, AI-designed biological threats, “deepfake viruses,” and whether the same technology that creates new biology can also defend against it.

Learn more about Radical Numerics: https://www.radicalnumerics.ai/

More from The Neuron: AI Explained

All 106 episodes
AI Can Now Write DNA. What Does This Promise, And What Could Go Wrong? With Eric Nguyen, co-founder and CEO of Radical NumericsThe Neuron: AI Explained · 51 min
Listen in VO