🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)

23 Sep 2026 · 1 h 32 min · 36 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eric Nguyen (Radical Numerics) explains “genetic language models” (GLMs) and their evolution from hyenaDNA (long-context DNA reading up to ~1M bases) to Evo/Evo2 (generative genomics) and Omni (a more aligned, multi-task model). He argues biosecurity is an AI arms race: design models are advancing faster than defenses, so defenses must be trained using similar underlying models to detect pathogenic sequences and enable surveillance.

Guest backgrounds

Brandon (Atomic AI) builds RNA therapeutics; RJ Honecki is CTO/co-founder of Mirroromics. Eric Nguyen is CEO/co-founder of Radical Numerics; PhD in Chris Ray’s group; early generative genomics work including Evo/Evo2 and long-context genomic modeling.

Key claims

Omni improves across many genomics tasks versus specialist models, largely due to mid-training/post-training “alignment” (Q&A-style task formatting, structured prompts, and likely RL). Variant-effect prediction is especially strong in non-coding regulatory regions (where many diseases reside). Data leakage is mitigated via bioinformatic QC/deduplication against benchmarks.

Notable examples

Evo generating new CRISPR-Cas systems (DNA→RNA/protein co-design) and generating a functional bacteriophage genome; Omni benchmarking on ClinVar/Trachem-style variant tasks; “chain-of-thought” analog using RNA aptamer fitness trajectories; potential applications in VUS diagnosis, antimicrobial phage design, and rare-earth mineral–binding protein discovery.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Arms Race in Bio-security

0:00 to 1:00

Explore the dynamic between design and defense capabilities in bio-security.

“The design side is going to get more capable.”

Understanding Genetic Language Models

1:50 to 2:56

Learn about genetic language models and their significance in genomics.

“So, Eric, let's talk about Omni and the blog posts that you guys did about the benchmarking.”

Hyena DNA Model and Long Contexts

2:56 to 5:06

Discover the Hyena DNA model and its ability to process long DNA sequences.

“And so we felt that it was a big opportunity to train AI on the genome.”

Evo Model and Generative Genomics

5:06 to 6:12

Explore the Evo model's ability to generate DNA sequences and their implications.

“like if you've heard of the phrase context rot, you know, the longer the input you put into a language model, it starts to deteriorate.”

Implications of Generative DNA Models

6:12 to 10:10

Understand the potential risks and ethical concerns of generative DNA technology.

“Models that we saw were really small, short contexts, so they could only pick up small patterns and limited context.”

Advancements in Omni Model

10:10 to 14:00

Learn how Omni surpasses previous models and addresses genomic tasks effectively.

“And I think indeed, a lot of companies, a lot of frontier labs are also being concerned about this emerging risk of AI models being capable of designing biological sequences.”

Aligning Genomics Models with Tasks

14:00 to 15:10

Learn how genomics models align to various tasks to optimize output.

“But then taking those learned embeddings or features and pointing at specific tasks or, you know, a bunch of tasks really in aligning it, meaning have it show you the output in a way that is meaningful to you.”

Technical Aspects of Model Structuring

15:10 to 17:50

Explore the technical structure and fine-tuning of AI models in genomics.

“And it was just a preview in that sense because we're still actively training and incorporating additional techniques into the model, like additional modalities.”

Understanding DNA Variants and Effects on Disease

17:50 to 22:50

Discuss how specific DNA variants can influence health and disease causation.

“i expect this type of output whether it's score prediction score or design is largely the mid and post training.”

Challenges in Predicting Non-Coding Variants

22:50 to 24:25

Learn about the difficulties in predicting the impact of non-coding DNA variants.

“And actually, most of these cases, it's state of the art.”
Show all 36 chapters

Benchmarking Genomic Models

24:25 to 26:50

Understand the comparisons and benchmarks used for evaluating genomic models.

“First, just while we're here, for the listeners, there's this column on this benchmark chart called Bordzoi, reference number four in the blog post.”

Model Training and Probability Scoring

26:50 to 28:00

Examine the training process and scoring metrics for genomic model outputs.

“This time, we're not describing the exact makeup, but Evo was autoregressive.”

Model Ranking Metrics and Training Techniques

28:00 to 29:50

Learn about how models are trained to predict disease variants and the importance of ranking metrics.

“We use every tool in the toolbox, essentially.”

Data Quality Control and Leakage Prevention

29:50 to 31:53

Discover how data leakage is managed and the measures taken for quality control in model training.

“which people are basically fine-tuning, you know, using them scores from EVO2.”

Supervised Methods and Performance Improvements

31:53 to 33:48

Explore advancements in supervised methods that enhance model performance beyond traditional benchmarks.

“And so, yeah, we actually have steps to QC the data quite extensively.”

Chain of Thought in AI and its Applications

33:48 to 36:05

Understand the concept of chain of thought and its implications for AI reasoning and biology.

“I believe Jason Wei at OpenAI showcased the first examples.”

Experimental Validation with RNA Aptamers

36:05 to 37:58

Learn about RNA aptamers and the model's ability to optimize sequences based on fitness scores.

“And so we took a data set, a large data set of aptamers.”

Scope of Model Training and Genomic Applications

37:58 to 41:35

Discover the model's training scope across various genomic domains and its applications in human health.

“So I, I mean, I'm wondering like Genomes carry lots of different information.”

Biodefense and Antimicrobial Resistance Solutions

41:35 to 42:00

Explore innovative approaches to tackle antimicrobial resistance using engineered bacteriophages.

Innovative Uses of Bacteriophages

42:00 to 45:00

Explore the potential of bacteriophages in combating antimicrobial resistance and superbugs.

“In particular, they mentioned that folks had used Evo to generate the first AI genome, a bacteriophage.”

Data Collection for RNA Structure

45:00 to 48:00

Understand how RNA data is collected and its implications for model training.

“I can maybe provide a better context on this if you want.”

Iterative Experimentation in Biotechnology

48:00 to 51:00

Learn about the iterative processes of biotechnology experiments and their efficiency.

“Following up on the rare earth mineral extraction, I find that is an interesting use case for this.”

Designing Proteins for Rare Earth Extraction

51:00 to 54:00

Discover how proteins can be engineered for selective binding to rare earth minerals.

“Folks have used EVO models to do this with toxin and anti-toxins in a very similar technique.”

Advancements in Long Context AI Models

54:00 to 56:04

Delve into the complexities of long context models and their applications in AI.

“Yeah, so I could speak to the evobacteriophage just a little bit because it's actually a separate group that worked on that.”

Journey into Long Context and DNA

56:04 to 1:03:12

Learn about the exploration of AI models handling long context and their application in DNA analysis.

“If you want to distract our researchers at Radical Numerics, this is how you nerd snipe them.”

MechInterp and Biological Insights

1:03:12 to 1:09:58

Discover the emerging field of MechInterp and its potential to extract biological insights from AI models.

“in addition to some of these more scientific questions.”

Fusing Language with Biological Signals

1:09:58 to 1:11:05

Explore the challenges and opportunities in integrating language with biological signals for AI models.

“Sidebar on that, is language, natural language, one of the most?”

The Importance of Biosecurity in AI Design

1:11:05 to 1:13:08

Discussing the dual mandate of ensuring responsible AI design and biosecurity.

“For us, we as a company thought it was very important to have a dual mandate, what we call a dual mandate.”

Pillars of Biodefense Strategy

1:13:08 to 1:15:09

Outlining the key components of a biodefense strategy against pathogens.

“The first one is around detection and surveillance, right?”

Characterizing Pathogenic Sequences

1:15:09 to 1:17:31

Understanding how to identify and characterize potentially dangerous sequences.

“sequence space, meaning the letters matching up exactly, but also function space, right?”

Challenges in DNA Synthesis Safety

1:17:31 to 1:18:47

Discussing the risks associated with DNA synthesis technologies and the need for safety measures.

“And this runs the scientific community, right?”

Balancing Innovation and Safety in Biodefense

1:18:47 to 1:21:31

Examining the tension between scientific innovation and safety in synthetic biology.

“I'm just curious about, you know, we've all, everyone who is working in bio has tried to use Fable and instantly got picked out on literally everything you type in.”

The Arms Race of Biological Technology

1:21:31 to 1:24:00

Understanding the dynamic between offensive and defensive strategies in biotechnology.

“Does it have to be perfect to be useful there?”

The Dynamics of the AI Arms Race

1:24:00 to 1:28:08

Explore the competitive landscape of AI in bio-security and its implications.

“It's similar to the cybersecurity community.”

Addressing Bottlenecks in AI Research

1:28:08 to 1:30:31

Discuss the key challenges and philosophical shifts needed in scientific work.

“Because ultimately, we do think it's going to be an engine for discovery and human health improvement.”

Embracing Innovation in AI for Health

1:30:31 to 1:31:49

Learn how to balance cutting-edge AI technology with health advancements.

“felt like there was a choice that they had to make sometimes to either work on the frontier of AI technology, and that was like consumer-related apps or enterprise-related apps.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Eric Nguyen:The design side is going to get more capable. The defensive side needs to try its best to get ahead. So I think inherently there is this arms race style dynamic that the defensive side has been far, far lagging. And so what we want to do is bring the defensive side to par, essentially. We felt it was important as a lab that a team that was both building the design capabilities is actually also best suited for building the defense capabilities because they're basically the same models. A model that is good at generating turns out is also very good at discriminating or predicting if a sequence is pathogenic or not.

0:34Eric Nguyen:For us, we as a company thought it was very important to have a dual mandate. It's this idea of essentially being cognizant and feeling responsible for the capabilities that we're enabling on the design side. So if we're going to create models that can design function into sequences, we believe and see a gap in companies being able to safeguard that technology. Welcome to Lane Space. I'm Brandon. I build RNA therapeutics at Atomic AI. I'm joined by my co-host RJ Honecki, CTO and co-founder of Mirroromics. Today, it's a pleasure to have with us Eric Gwynn, CEO and co-founder of Radical Numerics.

1:16Eric got his PhD in Chris Ray's group. He spent a lot of time thinking about how to do long context genomic models before long context or genomic models were cool. He was the first author and I think basically visionary behind the Evo generative model, one of the first generative genomics platforms, developed Evo2, which naturally led into Radical Numerics. Thank you for being here. Did I miss anything?

1:46Eric Nguyen:That sounds great. Cool. Welcome. Thank you. So, Eric, let's talk about Omni and the blog posts that you guys did about the benchmarking. But I want to hear first, okay, what is a genetic language model? Why do I care? What does it do? And then let's talk about the top line results from the blog post. So, a genome language model, or GLM, is a large language model trained on DNA sequences. So very much like natural language and chatbots you see, but not trained on words or natural language, but on the raw fabric of life, which is these sequence of letters that make up DNA. And we ourselves, our company, our team is known for creating the first generative genomics models, which are models trained on DNA, not just to read, but also write, meaning able to generate new sequences of DNA.

2:42Eric Nguyen:And we felt this was an area that was overlooked and that if AI could read and write DNA, it could change a lot. Scientific discovery and understanding of human health and how to treat it. And so we felt that it was a big opportunity to train AI on the genome. What kind of things can you potentially do with a model like this? Great to start for us when we first started working on DNA models. we worked on this model called hyena dna which is a large language model but it used a convolution instead of a tension so a little more technical details dna has this property that well it's very long right at the time these large language models had limited constraints on context right being able to fit long sequences and so we were looking for a more efficient algorithm to be able to handle something like DNA.

3:40Eric Nguyen:And so we came up with this, what we call the hyena operator, uses convolutions. Long story short, it let us process longer sequences, in this case up to a million, and at the time was the largest context for a language model. And what we did with it was essentially used it to read DNA, predict function. So given a sequence of DNA, string of characters, we predict its regulatory function, its effect on a genome. And this is interesting to scientists because a lot of the DNA in our bodies, you know, perhaps people are less aware, but actually we don't know a lot that much about our genome. It's, we know it, obviously, it encodes the information for making us, us and how, you know, all the different complexities and potential diseases.

4:30Eric Nguyen:At the same time, the grammar, sort of grammar rules about how the combination of those letters are sort of formed, what they encode and how they encode function and traits is not fully understood. And so the hope was using these DNA models, language models, to be able to map some function from the raw DNA sequence. And so our first generation models was able to show that, yes, we can train AI to be able to read and understand to some degree DNA sequences, and especially what we called the longer range interactions, meaning over sequences, you know, if you use chatbots, for example, like if you've heard of the phrase context rot, you know, the longer the input you put into a language model, it starts to deteriorate.

5:14Eric Nguyen:And so being able to pick up long range information and sort of patterns, motifs, grammar over long sequences was what we were trying to accomplish. And so we showcase in that first generation of hyena DNA that that was possible over a million contexts. And then really what started the field now known as generative genomics was the model called Evo. And Evo, we should try to showcase there was this idea of not just reading DNA, but being able to generate it. And so we wanted to accelerate essentially how biologists and scientists have learned from biology and particular genomics. And we felt like this whole field of generative AI being applied to language, great, accelerated obviously our understanding and ability to manipulate the natural language.

6:06Eric Nguyen:But here's this other language, DNA, the genome, that we don't understand. And it's barely being applied with AI at the time, a few years ago. Models that we saw were really small, short contexts, so they could only pick up small patterns and limited context. And none of them generated DNA. So they all just would read. And we felt that the idea of generation was so powerful and transformative in natural language. What if we could bring that to biology and DNA in particular? What can you accomplish that you can't do in a lab? So what does that unlock for you if you could do that very well? I think one of the first things that we showcased that got folks sort of intrigued by the potential of this was a CRISPR-Cas system.

6:55Eric Nguyen:So it's an enzyme that's able to cut DNA itself. And I think what was particularly enabled by the Evo models was the ability to generate over not just one modality or one type of sequence, but spanning multiple modalities and spanning multiple scales. So CRISPR-Cas, it's a molecule made up of both RNA and proteins. And so at the time, you hadn't really seen models that can generate multiple modalities. They had protein language models that can generate proteins. Sometimes you had RNA models that could generate RNA, but you didn't have a single system to sort of co-design. We showcased that a single DNA model, sort of the foundation of both of those, right from DNA, you can get RNA and proteins that we can design a single system to generate and also function in the real world.

7:52Eric Nguyen:So we asked Evo, we showcased it a bunch of natural CRISPR-Cas systems and essentially asked it, can you make a new one? And we're able to sample from that model that we trained. And indeed, we showcased that Evo was able to discover a new CRISPR-Cas system. And, you know, folks were intrigued by it and cover Science Magazine and later get to give a TED talk about the work, which is interesting because obviously the audience is very general. And so trying to make, you know, what is CRISPR-Cas system? What is DNA? How do you generate it? Why would you generate it? All sorts of fun topics. And I think the more even intriguing, exciting thing that scientists eventually, just last year, showcased what you can do with a generative DNA model was to generate the first genome from scratch using AI.

8:38Eric Nguyen:So this is something not possible by humans, right? Humans usually, you can think of it like copy and paste parts of other genomes or other DNA, put it into something else. but they would just take out small motifs that you know they know the function and they understand the rules but to build something from scratch in the ground up had not been done before at the genome level and so evo turns out was able to generate a functional genome and in this case it was it's known as a bacteriophage also known as a virus and this this was a key turning point for i think the scientific community and for us as a company at radical numerics because we felt And this was such an indicative manifestation of the potential, right, to create a whole organism, not existing in nature, but also the potential harm that that means as well, right?

9:28Eric Nguyen:If you can control, if you can manipulate the fabric of life itself, control its function, what kind of implications does that mean? What are you enabling into the world? And so we actually got a lot of feedback, a lot of comments, a lot of outreach from folks both excited and concerned about this capability, this kind of capability, and just the trajectory, right? This is the early stages, the first thing one can sort of project and imagine what this could lead to. And so we felt as a company, it was important to not only push on the biological design capabilities of these models, but also the ability to use them as defensive tools for the potential of misuse and biological risk that emerges.

10:11Eric Nguyen:And I think indeed, a lot of companies, a lot of frontier labs are also being concerned about this emerging risk of AI models being capable of designing biological sequences. And, you know, at the same time, it's being mostly attacked from like a natural language standpoint, like safeguards and things like CLOD. You know, if you talk about viruses, it'll just like shut you down, which is great. I think it's to some degree you need safeguards at the natural language level. But I think what you also need clearly is safeguards at the biological sequence level too. So you need models that not just can understand language and, you know, the trajectory of your chat, but to understand the substrate itself is the next step in ultimate limit, right?

10:55Eric Nguyen:If you can have models can understand, you know, if the sequence is pathogenic or a virus, that's the level of defensive capabilities that you want. And then, you know, being able to push that out into surveillance systems, national security of, you know, being able to monitor, you know, emerging sequences in the environment. this is what that kind of capability makes possible and then we think bringing this frontier technology to that community as well is also important just as important as using this for human health which is what we primarily focus on. There was this evolution that you mentioned there's the hyena DNA and then there's EVO, EVO2 and now Omni.

11:37Eric Nguyen:Can you just talk about a little bit about Evo and Evo 2 struggled to meet sort of specialized models across many tasks, whereas in the blog post, you talk about how across a wide range of tasks that Omni is actually able to outperform them now. So can we just talk a little bit about that? So yeah, Evo was intriguing to folks in many ways, showcase the potential for applying to multiple types of modalities. But it's still, it's still in many ways underperformed sort of the specialist DNA models, especially on human genetics or genomics. And so although EVO was competitive, it still wasn't state of the art or kind of pushing the needle.

12:21Eric Nguyen:And so, you know, some parts of the community thought like, why use a giant LLM when I can use these smaller, more specialized models? And so what we wanted to do with Omni was amongst many things. But one of the first things was to showcase this idea of mid-training and post-training or broadly alignment. So the way we think about language models and bio and Jonas in particular is that mostly we've only seen base models trained. So they're pre-trained, but they're basically unaligned. So in the analogous space for natural language, it's like you're doing all the pre-training, but to make it actually useful in the real world and answer questions that users actually want and that is in the form factor they find actually informative, there's a bunch of alignment and post-training and mid-training done to get the models to be production ready and actually useful.

13:15Eric Nguyen:And so we felt EVO was just showcasing the potential of that pre-training, but Omni is a step of actually making it useful for folks like scientists, right? And so we spend a lot of time on alignment and mid and post-training, which is essentially showcasing tasks in the form that people generally would want to understand about genomics. You know, given a wild type and a mutation sequence, you know, help me find the causal variant, right? These types of questions and form factors for how you might want to analyze genomics doesn't just emerge necessarily easily on its own from pre-training. Pre-training is, you know, this next token predictions task or infilling.

13:58And basically, I think of it as like the raw pattern making ability that you're teaching it is in the pre-training.

14:06Eric Nguyen:But then taking those learned embeddings or features and pointing at specific tasks or, you know, a bunch of tasks really in aligning it, meaning have it show you the output in a way that is meaningful to you. takes a little bit of teasing and manipulating that so far, like the furniture labs are the ones that drive that research in the natural language community. And so we wanted to bring a lot of that research and more to genomics. And so I think Omni is really just a preview to showcase that potential, right? And I think once you do that, even just a little bit, you know, we were surprised that it did start being state of the art and pushing the boundaries, not just being affected of multiple tasks just broadly like EVO was, but actually pushing the frontier of each of those areas of variant effects prediction, causal mutations for disease.

15:02Eric Nguyen:It could start actually being useful for human genomics. And so we're really excited to share that with folks. And it was just a preview in that sense because we're still actively training and incorporating additional techniques into the model, like additional modalities. but I think we were basically too excited and we wanted to get this in the hands of folks faster and some of the feedback we got in early interest lots of hospital systems, non-profits that have tons of genetic data and for example, they know there's some kind of condition or symptom for the patient but they can't figure out which parts of the DNA are causing it so they've got these VUSs or variants of unknown significance that we're extremely excited to apply these models to and actually help diagnose a lot of these patients is one example of a real-use application.

15:54I'd like to talk more about the applications in a bit, but I am curious just from a technical standpoint, what does it look like? What does it mean to align a genomics model? I can imagine with large language model, there's a sort of natural chain of thought. You know, RL, there's like a clear paradigm here. I'm curious, what does it look like for a genomic model, which is not maybe better described as something like fine-tuning, in a different language.

16:19Eric Nguyen:Yeah, I mean, I think in many ways one could describe it as fine-tuning, but then introducing, I'd say, sort of the key components are the right structure of the inputs. So feeding them in a certain sequence so that the model is aware that a certain task is being asked of it. So there's a mix of special tokens to basically you can think of as like, if you're going to do disease prediction for disease A, expect this special token, right? just kind of like a way to prompt it. If you expect it to do design, have another special token and then showcase the examples kind of like in a chain of thought manner, meaning showcase a sequence of desired outputs and the directory of it.

17:03Eric Nguyen:This is a little vague sort of intentionally because it's part of our secret sauce that we're still developing. And I think over time, we want to showcase more and more of it. But in many ways, it does mimic a lot of the natural language community. A lot of it is fine tuning, but really it's also carving out specific data sets that you want it to focus on and then structuring the questions or tasks in specific ways as opposed to pre-training. is really pre-training is really just feeding it in and everything in really and just doing next token prediction or mass skin filling if you're doing mass language modeling and it has no sense of like this q and a type structure where you have a question you know a prompt and then an output and you can think of mid and post training as starting to showcase given this type of input i expect this type of output whether it's score prediction score or design is largely the mid and post training.

18:01Eric Nguyen:The post training also includes things like reinforcement learning too. But I think the bigger steps are, you know, introducing structure of like questioning and answering. Do the models have multiple heads that are task specific or are you training one set of heads or whatever that can answer multiple questions at the same time? Just change the input tokens or whatever. Yeah. Broadly, it's a little, you can, I think we're flexible on this. The idea, yeah, sometimes you can use different heads, but the idea for us is to unify more so. And so I think early experiments, we did have different heads, but in some cases, the different heads do better.

18:45Eric Nguyen:In some cases, a single head, a single model does better. And so I think we're flexible on that. But I think broadly, the direction that we are moving toward is a single. And the reason for a special motivation for that is that we're trying to unlock a lot of modality, a lot of generalization. And I think the more unifying we're able to make these models, I think that's when you see more emergent capabilities happen. And in large part, that's what motivated the DNA work. We felt like a lot of models were specialized into other modalities, a little more downstream from DNA. So like RNA or proteins or molecules, we largely felt DNA is the foundation and that from DNA, you can learn a good deal of other modalities, potentially all of them.

19:38Eric Nguyen:And I think other modalities is sort of like additional context that you're showing the model. That's how I kind of view it philosophically in my head. But yeah, the idea that single models unifying across modalities, scales is what the lab builds toward. Can you walk through a few of the tasks that you talk about in the blog and just explain and remembering to narrate for the listener-only audience, but talk about some of these top-line results and maybe dig into them a little bit? Sure. One of the areas that we thought was really interesting for showcasing in this particular release for Omni was on understanding variants and their effects, which are essentially in DNA, a change in a position, you know, changing the letter of one of the other three letters in your genome in DNA, sometimes can cause a disease and sometimes it doesn't do anything.

20:42Eric Nguyen:Actually, many times it doesn't do anything. But there are specific areas in your genome that if you have a different variant or a different letter there, it can cause disease. And in many cases, because the combinations of these changes in the genome, over 3 billion letters, right, is so vast that for clinicians and scientists, we actually know only a very small portion of which variants are causal to disease or not. And so there's these benchmarks from folks who collect variants, different hospital systems and clinics, and some are known and some are still unknown. Folks have created some benchmarks from ClinVar or Trachem to basically the idea is given a mutation in a DNA, can you tell if it's going to cause a disease or not?

21:35Eric Nguyen:And so this is a very good setup for DNA models because they're probabilistic. And so when you do make a change, it basically can modify its confidence or probability of predicting an X letter. And in this case, we've sort of leveraged that predictability of these models. They've essentially seen and been trained on so much DNA, in particular human DNA, they kind of understand what's common in this and usually common or conserved across different other folks usually typically means healthier. And so if it's less common, you can think of it this way, it's less common. The model can pick it up and sort of predict that it's potentially pathogenic or disease causing.

22:18Eric Nguyen:And so we've taken some of these benchmarks. And when you introduce a variant or a mutation, there's different types of mutations and variants. So sometimes you can delete a letter altogether. You can just flip it. You can remove big portions of the DNA. But these are generally single variants. And in this case, this is where previous DNA models really struggled, especially on humans. And so we showcase that not only is it competitive or capable for humans, but in many cases, it's state of the art. And actually, most of these cases, it's state of the art. And I think the exciting part is that areas where other models that were, you know, currently on the frontier, they still were lagging behind quite a bit in terms of where in the genome.

23:06Eric Nguyen:So in the genome, there's coding and non-coding regions, like protein areas, protein regions. Meaning areas that are actually coding the structure of a protein versus other areas that do other things like regulate what genes are expressed. Yeah. And so these non-coding regions are largely regulatory. They kind of control how much or when to use a certain gene or turn them on. And in many cases, these non-coding regions, these regulatory regions, variants there, mutations there are much harder to predict if they cause disease or not. And I think what's exciting about the new generation of models we're building with Omni is that that's where we shine, especially.

23:47Eric Nguyen:The models are able to pick up mutations and be able to distinguish if it's disease-causing in these non-coding and especially long-range areas. So I think this is particularly exciting to a lot of geneticists that have struggled to use traditional bioinformatic tools or statistical methods. Because they've largely focused on coding regions, which is only about 1.5-2 % of the genome. And it turns out many, if not most of the diseases, are in these non-coding regions. And so there's been a real desire to build models that can actually pick up these variants of disease-causing variants in these non-coding regions.

24:24I have several questions. First, just while we're here, for the listeners, there's this column on this benchmark chart called Bordzoi, reference number four in the blog post. I think for maybe some historical context, could you talk about what this column represents? And, you know, maybe this also helped give context for, you know, the omni column on the right.

24:49Eric Nguyen:Yeah. Yeah. Great point. So what we show in this benchmark here is really taking some of the representative models or the strongest models in the deep learning side and also in the traditional methods. So we have Evo2 is the latest previous genomic model that our team had worked on. And then Borzo is, it's also a DNA model, but a very different kind. Essentially, it's a supervised model that predicts from DNA functional genomic tracks. So it too is inherently multimodal, but it's not a language model. So it doesn't predict like a next token prediction. It goes from a DNA sequence directly to a functional genomic track like chromatic flexibility or gene expression.

25:29Eric Nguyen:And these tracks have been annotated extensively by people writing their dissertations and whatever. Yeah, yeah. Right. So the big difference there is that it's a supervised task, right? So it requires labeled outputs. And in our case, these language models, they do not require labeled outputs, right? So you're doing raw pre-training on unannotated sequences. So, you know, why is that desirable? Well, there's a lot more data, a lot more genomic data that's not annotated. Actually, most of it, pretty much in many ways, almost all of it is not annotated. And so being able to learn from an unsupervised manner, hugely desirable, right?

26:06Eric Nguyen:For us, we wanted to showcase the benefit of pre-training on raw genomic sequences and, you know, comparing it to state-of-the-art models in other spaces in DNA. And so your prediction is when you say it's unsupervised, how does the unsupervised property work? Like, how do you convert the output of whatever your model is to an actionable, like, ranker, score, whatever? What I talked about before is, you know, to train the Omni model, it's pre-trained, so it's unsupervised. But when you're pointing it at specific tasks, there is a supervised step. So it's, you know, taking a smaller data set that is labeled.

26:46Eric Nguyen:but essentially what we're doing is using the likelihood scores so the raw outputs of the language model which basically you can think of it like a probability for predicting what the next letter is we can essentially showcase the probability score the likelihood score for the mutation versus the wild type so that's seen in the reference genome versus the mutation in this particular case and then we'll have two scores and then you can think of it as like using a ratio of the two to showcase basically how different are you from how different is this mutation from um from a normal or baseline basically and then that's you could think of it as like a surprise factor um that the model is able to use and leverage and then we can use that to essentially um score an actual prediction for disease or not does that make sense yeah so uh Omni autoregressive, or is it diffusion or something else that you can't tell me?

27:45Eric Nguyen:Yeah. This time, we're not describing the exact makeup, but Evo was autoregressive. It was the first large-scale autoregressive. And so I think for us, we don't tie ourselves down to a specific training objective. We use every tool in the toolbox, essentially. Okay. But for this specific benchmark, you're going along and you're just using the likelihood distribution of the tokens. And some tokens are, the model thinks these are unlikely. And that is probably because some evolutionary constraint. Like this doesn't show up often. And because it doesn't show often across genomes, it is probably going to cause problems and people will not survive.

28:33So on. And so you think that is basically your ranking metric or something.

28:39Eric Nguyen:Yeah, it's one interpretation of how the model is thinking about it. And very similar to in natural language, you can describe the same kind of paradigm. And there is additional case in mid and post training to leverage more than that, I suppose, because we can teach it specific structure and benchmarks so that it can build on top of what you just said. described, which is like what's common in nature, but also because it's a specific task for disease variant prediction, then the model has additional training introduced during mid-training to showcase and to add additional learning power, essentially.

Read the full transcript

29:20Eric Nguyen:What are some examples of that? Again, probably secret sauce to some extent, but can you give just a gist of what that looks like? What are the kinds of things you would throw in there? We would actually throw in the score to itself. So, you know, like I mentioned a ratio, it's a ratio of wild type versus mutation. I would say that's more of a zero shot method where you don't even have to do any mid-training. And that's what EVO2 is doing in particular in this column. So EVO2 is not fine-tuned, essentially. There's an EVE column which people are basically fine-tuning, you know, using them scores from EVO2.

29:55Eric Nguyen:That's from Goodfire. And that's also, you know, folks that we greatly respect and they kind of showcase that these models are able to be state-of-the-art when you can fine-tune them as well, not just zero-shot. And then our model is introducing sort of a step about that, not just fine-tuning, but also introducing structure into, by structure, I mean the format of these benchmarks into the model itself, that Q &A style formatting during mid-training, which is what gives us an extra boost even. I see. And extra boost, but I think the other benefit too, that we didn't emphasize too much in the blog, but I think it's really convenient for practical use for scientists, is that you're doing this without taking the embeddings and then tapping on ahead and then doing some regression, which is the extra step.

30:43Eric Nguyen:It's an extra hurdle. Can you imagine if Chachapiti, every time you ask a question, you had to fine-tune it for a certain domain. We've done it so that the model is flexible during mid-training to be trained on many tasks at once. And so the, that fine tuning, that, that last step of training, the embeddings doesn't have to be done. It's out of the box at that point. You just prompt in a certain format and it will, um, you know, that certain format tells it which task you're going to do. And then we'll output the answers in that, in the desired format, basically. How careful were you in designing this post-training scheme to avoid kind of data leakage with the ClinVar, CrateGem, RNAGem, and so on?

31:23Like these data sets, like how confident are you that there is no data leakage, either like accidental or something upstream and that, you know, because I would not be surprised if a lot of these sequences showed up also in your training data, even in a, you know, unsupervised sort of way.

31:41Eric Nguyen:Yeah. The short answer is we're extremely cognizant of the risk of data leakage and extremely hard to not mislead or, you know, be careful. And so we have bioinformaticians that are able to basically comb through the data and curate, dedupe, and align sequences to make sure that things that are similar potentially to what's in the benchmark are not there. And if they are there, we remove it. And so, yeah, we actually have steps to QC the data quite extensively. Yeah, cool. Yeah, I guess maybe before we move on, I think it's really cool seeing that there are these supervised methods, which previously several of these numbers were, let's say, within the error bars, if not just straight up beating what came before them.

32:24And now you have significantly improved upon that. Yeah, we're super excited.

32:29Eric Nguyen:And I think for the longest time, there's this area of this other method called CAD, which has been state-of-the-art. And state-of-the-art for a reason, which is what they, sort of by design, they'll take the best methods and kind of do it on ensemble right so they'll take they'll take up another even if the best method is another previous model um they'll mix it with like an svm and just like throw the kitchen sink at it and so you can see why it would be the best right and so that was the bar for us we're like if they're going to throw the kitchen sink on it like we're not going to you know cherry pick one model and say we're better than that we need to beat what's possible humanly possible now, like across everything.

33:05Eric Nguyen:And so, yeah, our researchers were setting their sights on that to see if they can actually improve performance across every method. I'd be interested to see, there's some discussion of chain of thought. And that broke my brain a little bit when I was first hearing about that. I'd be really interested to hear about what that even means. I had to pour through the blog post to really understand that. Yeah. Yeah. I think this is really just a taste of where we think the design capabilities can move toward and be more usable for folks. So Chain of Thought stems from natural language community. I believe Jason Wei at OpenAI showcased the first examples.

33:55Eric Nguyen:and really what the breakthrough there was showcasing that these language models perform better when you just show your work essentially you show the steps of how you came to a conclusion or an argument and it turns out even if they were like simple steps but it just gave the model a chance you know maybe it's sort of like philosophic who knows exactly why it works but essentially feeding more tokens in and giving it more scratch space to think and so people just think this is the sort of the beginning of reasoning for these language models, this ability to kind of get to an answer by thinking to itself by itself.

34:37Eric Nguyen:And so it seemed quite successful in language, very successful. And that's why you have a lot of agents that just spent tons of tokens, right? Just showing its work, right? In many ways, it stems from this chain of thought paradigm. and in biology we saw very little of that we saw we started to see some of that in um protein design a little a little um and so we wanted to push that and and showcase that you know well at first explore is that possible in dna like what does that even mean in genomics because you don't really have words that describe you know it's thinking um so how do how do you take that pair that same paradigm and introduce it to a DNA language model.

35:20Eric Nguyen:And so what we did was a simpler version in 70 main ways. We had this data set of RNA aptamers. So just think of it as these desired sequences with some kind of fitness score associated with them. So we took this large data set that had RNA input and a fitness score. The fitness score go high. It's good. The simplest version. It's a big data set. And so what we wanted to showcase was that if we show the model progressively better RNAs in a series of steps with its score, right? So you have like low scores first, and then you gradually move up the chain. Can the model continue that trajectory on its own?

35:58Eric Nguyen:And then, you know, in the final step, does it self-optimize to a point where it's like the best score it can get? That was the experiment. Can we do that? And so we took a data set, a large data set of aptamers. We held out a portion of the best performing ones. and we showed it only the lower ones, but then we ranked it, right? So we showcased lower scores with the RNA abtomers and then progressively got higher and then asked the model to just like continue with that pattern. And it turns out it was able to recapitulate some of those higher scores that we had not shown it yet. We were actually in the process of validating the wet lab right now.

36:33Eric Nguyen:So we didn't get to show it here, but we wanted to know, right? Actually, can it not just do this in silico, which it can, it showcased that it was able to continue this trajectory and create designed plausible aftermers with higher fitness scores and now we think this is uh you know obviously if this works in the lab we think this is a hugely hugely valuable paradigm that can be that can be pretty much applied to every other type of biological sequence that's usually what you have you have sequence you have some kind of fitness score desired output and if we can get models to eventually learn that structure and you know basically just show a series of progressively stronger sequences the model can then predict the rest that's a very powerful paradigm honestly a bit surprised about this um that specifically this task saw strong improvement maybe my personal bias is coming in here but rna is uh somewhat notorious for not having good co-evolution data in terms of like constraining structures, right?

37:35For, you know, viral genomes, oftentimes there is strong evolutionary pressure, but for mammalian, usually RNA does not have evolutionary pressure. And I think the community has seen that very, very clearly the predominant pressure is like RNA will code, you know, will carry information, coding information. So I, I mean, I'm wondering like Genomes carry lots of different information. They code for proteins, they have regulatory elements, and, you know, different types of genomes have different types of structure. So I'm wondering, where do you think this capability might have emerged in this language model?

38:20Eric Nguyen:That's a good question. And honestly, I'm not sure. Like, we're surprised, too. one because the model is pre-trained on on dna and it's really just like mid-trained um on rna very very you know in a small way yeah i think what your intuition about the dna having a lot evolutionary effects or information is probably where and so i think this is hinting at the idea of why we think it's so important to pre-train on on dna genomes and genomes first um and then sort of add additional modalities on top because you get a lot of transfer and you want in the modality generalization is the thing that we're working toward you know what i didn't talk about as a company for the company as well is this idea of a building toward general biological intelligence where we are unifying a lot of the different so-called like languages or modalities of biology at the end of the day they stem from dna and i think people have not exploited that fact as much It's usually really specialized, domain-specific modality-specific models and not leveraging a lot of inherent shared structure from other modalities.

39:31Eric Nguyen:And so one example of that is like, you know, when people talk about virtual cells, for example, there's a little bit tangent, but, you know, they tend to focus on just RNA and, you know, transcripts and they want to generalize to describing an entire cell. But obviously a cell is a lot more things than that. in my mind, if you want to learn a system, you want to learn from all the signals or sensors of that world or that system. You know, if it's a cell and you want to fuse, you want to understand the DNA, you want to understand the metabolomics, the epigenomics, the proteomics. And that's when you get closer to like, quote unquote, a virtual cell.

40:09Eric Nguyen:And in our minds, we don't even want to stop at just the cell, but we want to fuse all of these sensors across all of biology. Will it get us to a super intelligence that understands every component of minutiae biology? Who knows? But I am confident that this type of paradigm will get us a hell of a lot further than we are now. Like that's my bar. Can you make something far more useful than now? So I'm curious, are you focusing on eukaryotic cells? Are you focusing on like human genomes? Have you gone so far as to do viral genomes? I mean, there's a lot of DNA viruses, but it seems like plausible that RNA, that there's a lot of RNA sequence, virus sequences out there.

40:52And I'm not sure fundamentally they would be much different in terms of training. I'm curious, like, what's the scope of that, if you can talk about it?

41:00Eric Nguyen:Yeah, absolutely. We are interested in all domains of life. so here we focused on humans in particular because we thought this was an area that of previous models evo and evo2 were not as strong and sort of got a lot of feedback from folks asking you know what are these models useful useful for they can't understand human genomics because it's too complex of grammar and rules and dna it's too noisy it's too there's too many repeat characters and all that stuff so we wanted to showcase we think this is actually useful and it can be applied to humans and it's sort of the most complex of the complex in some ways but i think there's opportunity to apply these models generally to to every form of of life um so we absolutely are interested in pro carryouts and in viral um viral in particular we care about especially for biodefense and biosecurity especially and um i think there are also lots of therapeutic applications that we can learn from microbial life of you know maybe obviously for some folks.

42:01Eric Nguyen:In particular, they mentioned that folks had used Evo to generate the first AI genome, a bacteriophage. It turns out you can use bacteriophages potentially for AMR or antimicrobial resistance. You know, if you have a superbug bacteria infection, which in the world is about 2 million deaths from bacteria infections still, the idea of using viruses, designed viruses to target specific bacteria has been done for a long time, particularly in Eastern Europe. And there's a potential to make a new class of antimicrobials that are not like antibiotics, but very similar that can be used just like it. And so I think we're gravitating toward things that are high impact and the potential to save lives.

42:48Eric Nguyen:And so we don't stop at just one type of genome. I think we're interested in anything that's beneficial to humans. Maybe going back to my question about Solex and RNA predicting kind of a chain of thoughts of RNA evolution. I'm curious, was this model trained on RNA sequences or sequences which might have evolutionary pressure on RNA structure? Only during mid-training. So pre-training is all just genomes and DNA. And so the only time we introduced RNA was for this specific task and only RNA from this data set. So not even outside. I see. So this really was something along the, there is something encoding RNA structure in this model to some degree, maybe.

43:38Or either that or implicitly.

43:40Eric Nguyen:Yeah, implicitly. I would say implicitly. That's cool. Yeah, because I'd say sequences, as we know from proteins, implicitly should learn structure from just sequence. yeah and so we don't we don't add in 2d or 3d information at this point but we absolutely plan to yeah so um just so i understand first of all the chain of thought idea is this is a demonstration of it but the idea is that anything that you can get sort of a training set that has a sequentially better um measurement of some sort is maybe a candidate for this technique yeah and so can you just describe for this particular experiment just so we can understand how we're mapping chain of thought to this data set and what the data sets have to look like can you just describe how is the for this i know this this wasn't your data set but um how was the data collected in such a way that you could map accurately from sort of fitness or whatever to a particular phase or part of the data set.

44:47Eric Nguyen:Yeah, so I'm less familiar with how the data was actually sent us or generated from the experimental point viewpoint, but they are validated from a wet lab in the real world when it was collected. So it has some kind of fitness score, I believe through... I can maybe provide a better context on this if you want. I mean, so the idea here is you just generate a bunch of random sequences and then you take those sequences. And so you have like an aptamer structure, which is essentially like a switch with RNA, which sort of when something binds to it, it will do something like cleave off a sequence.

45:24And you can use an NGS, like Big Generation Sequencing Readout. I'm very high throughput. So you create lots of these different sequences. I guess in this case, they were targeting a specific HIV protein or genome or something. And if it binds, you basically get the signal of like you get more reads of that. And so the more time, the more sequences which are floating around, kind of like the more fitness, the more likely it is to bind. And then you take those and then you mutate them again and you kind of iterate on this. Right.

45:57Eric Nguyen:So you have this iterative experiment where you're progressively using the Petri dish to basically identify the most fit thing. So, and then what you're doing here is you're basically doing this same experiment in silico. And they're actually doing very high throughput. Like, I think there's like 10 to the 11 or something sequences, some like really high number of kind of sequences explored in parallel for this experiment. But the key is the biology of the experiment is actually doing the filtering, right? Yes, yes. Yeah. And so you can imagine other types of experiments where you could apply the same kind of idea where you're doing this like progressive refinement of something which is very common in biological lab work.

46:39Eric Nguyen:And so if you're capturing those intermediate states and you can maybe feed them into the model and do that kind of thing. Absolutely. Yeah. So another example is for the antimicrobial resistance. I think that's something we're very interested in as well. The ability to selectively target specific bacteria strains or kill bacteria strains. Yeah, that's measured basically a score zero to one of how effectively that is done. And so showing progressively more effective phage genomes and their associated scores for effectiveness fits that paradigm very well. There's another design task that we're working with a national lab to do this with as well.

47:20Eric Nguyen:And an area that is quite different for us, but it's on designing proteins to extract rare earth minerals. And so it turns out that you don't just care about proteins that could bind to something, but you want it to be selective. You want it to bind to one type of rare earth mineral. and so you have scores associated with how much affinity or you know binding affinity for each type of mineral we want to essentially do the similar thought this is a similar exercise with rare earth minerals and show it a series of progressively you know desirable scores not just for binding to this but like you know lower binding scores for other ones you can selectively do it so i think there's the creativity in which you can showcase sequence and design desirable desirable sequence with, you know, some kind of fitness or functional output is a relatively intuitive way for people to design, you know, just prompt by just by prompt engineering, essentially, which I think is very exciting to explore more.

48:21Following up on the rare earth mineral extraction, I find that is an interesting use case for this. In fact, maybe one where you probably wouldn't have a comparative advantage compared to some other techniques because it seems like your strengths are probably in larger scale design across like entire organisms but focusing on individual proteins that may be one much more of a like a structural task and i'm curious like going to talk about clinvar the argument here is that the model now is a really good statistical representation of what type of mutations are common or not common yeah and i think in order to do design as a if you want to design structures I think you want to understand structure.

49:04If you want to understand disease, I think that is, I think, more natural. For many diseases, it's much more natural in terms of like a population genomic sort of way. So I'm curious, where do you think your strongest competitive advantage is, you know, in using this strategy? And do you think that your model understands structure in addition to function?

49:27Eric Nguyen:Yeah, I'd say what drew us to this particular application of rare earths and our strengths in general, why we thought it might be suited for it is two things. One is, I think, in areas where context matters. So yeah, you're right. So a protein design task where you're just designing structure, maybe not naturally, where we see ourselves, you know, competitively advantaged. but in the rare earth's case our hypothesis is that context matters and what matters here possibly is certain microorganisms with proteins that you know have the function that we desire we can potentially prompt and provide us context for this is the neighborhood in which you know tell omni this is the neighborhood of genomes or microorganisms that you should search for new proteins.

50:21Eric Nguyen:So it's really sort of like a mining exercise and what we can showcase to our models that other folks can't, like protein infrastructure models, is that we can feed in non-coding regions before that gene of interest or that protein of interest and then ask for the model to provide variants essentially. So like mutate this protein but know that you're in this microorganism but show me different variants that you've seen in nature or combine different things that you've seen in nature given this context. And I think that is one reason why we can make very evolutionary diverse sequences and potentially phages.

51:00Eric Nguyen:Folks have used EVO models to do this with toxin and anti-toxins in a very similar technique. They've prompted on things upstream from the proteins of interest and then asked the model to kind of generate a bunch of plausible other ones. And I think this case for rare earths, that's super exciting because then we can come up with plausible variants and then test them relatively simply for what they bind to and selectively bind to which i think is really interesting for this this um this partner who cares for example about securing the supply chain of rare earths for the u.s strategically and so we thought that's absolutely worth something that um is worth supporting so your point is not just you're designing a protein but you're designing an organism which generates a protein and this protein has an action but or is it that that the the way to design this is you need to understand how hit through the tree of life interactions of proteins with rare earths have or with certain minerals have occurred yeah i'd say in this particular case for like the rare earths we're not really interested in the organism like designing the whole organism but i think the organism does tell us about what proteins are plausible and their selectivity is not or the way they evolve not just around the protein itself but also the regulatory elements around it and it helps us narrow and provide additional context i guess i think of the genome or dna broadly as sort of the imprint of the physical world into dna and so there's parts of the dna that we want but the things around those parts that we want also tell us a little bit about the context of where it came from and how it came to be in its function.

52:39Eric Nguyen:And so I just think of it, you know, like going back to that word, but context, I think context matters, um, in many of these applications, or at least in some of them. Yeah. I think context matters a lot in biology. Um, I think maybe one of our big bottlenecks is the lack of context and how we as humans mostly approach biology in terms of a very engineering, like let's isolate individual systems and a systems biology approach is very hard to get any sort of meaningful quantitative predictive power so maybe going off on a tangent but i i mean i'm curious you know going to context and thinking about how context scales to an organism you're talking about right now two million length context right yeah you know i think the human genome is roughly you know a thousand times longer um so uh or but even like a lot of you know say bacterial genomes if you're trying to engineer them are are quite a bit longer than that so how do you leverage something which has a long but still finite context compared in a to to do synthetic biology across like large organisms maybe actually maybe for context how did the evo2 bacteriophage design work which was probably much more than 2 million for that as well.

54:00Eric Nguyen:Yeah, so I could speak to the evobacteriophage just a little bit because it's actually a separate group that worked on that. But that context was actually pretty short. Actually, the reason why they started with phages is because it's amongst the shortest genomes. And so I believe it was something around 6 ,000 base pairs. Oh, wow. That's really short. Yeah, yeah, yeah. So extremely short. Viruses are insane. They're incredibly efficient. They packed a lot in there. Yeah. Yeah. Yeah. But yeah, no, I think your question about how do you get longer context with something smaller is a very key question that we, I think, broadly, the AI community is constantly trying to fix and be creative about.

54:42Eric Nguyen:And so, I mean, I think that's a large part why we as a company are an AI research lab first, because we think the innovation needs to be constantly pushed. It's not a space where we can just grab open source models and expect that, you know, many of the tasks that we care about are just going to be solved. We want to continually push the envelope. And so context is one of the key researchers that we drive. And I would say that's probably how we got our name because we worked on long context before it was a thing, I guess, like 2023, 2022. who long context. You and your team, I mean, your collaborators have a long history of these state-space models doing, you know, pushing context.

55:25Like what seemed insane at the time. Yeah. I mean, now I think routine and books the big slabs. Yeah, exactly. But at the time was like orders of magnitude longer. I don't know if you want to talk about that a bit. And also I think one, maybe one thing which I think is really fascinating is how in for both this model and other ones, you really go on a first principle way, like diving into the architectures of how GPUs work and designing models, which are, you know, exploiting the architecture of GPU in addition to, I mean, it's not just, oh, we're building longer. It's like, what can we do special given the compute constraints we have to push the boundary?

56:03Eric Nguyen:Yeah, absolutely. If you want to distract our researchers at Radical Numerics, this is how you nerd snipe them. You talk about, you bring up long context and kernels, GPUs, then they're like, wait, did someone say kernels? And they start trying to figure out, you know, how to make things fast. Yeah, long context, a special place in my heart because that's what I focused on in my PhD at Stanford. And that's how we started thinking about DNA and, you know, going back a little bit for fun. We were looking at working on language models in general. And then we noticed our models were good at long context.

56:39Eric Nguyen:And so that's the progression of like how we started working in the space was like, oh, these models seem to be really efficient on long context. And then Michael Pauly, my lab mate, he worked on the first design of Hyena, this convolutional architecture. And then we started thinking like, let's push this further. Let's see what new applications open up if we really lean into long context. And we asked, what's the longest sequence out there? and and eventually unsurprisingly we landed on dna we're like dna's got to be the longest three billion base pairs and you know um and um we started thinking like okay what's being done there like what kind of context lanes are people doing there they're doing super short they're doing like one or two thousand base pairs or tokens at time it's like a you know way smaller than what one would want for dna i mean most human transcripts are like 3k or so so that's not even And like, you know, most, that's not even what you need to represent like a protein or most of the time.

57:39Eric Nguyen:And so it's clearly a need there and overlooked. And we saw it as a way to initially, like, let's see if we can do something that we had no idea if it was going to work and no idea who would want it. And so we just started tinkering around. And it turns out out of the box, relatively out of the box, it was doing pretty well at reading DNA. And then it led to, okay, let's people keep asking about like DNA. They didn't really care about our language stuff as much. and they would say like, can you, can you do longer? Can you, you know, what can you do with it? They always ask, what can you do with it?

58:09Eric Nguyen:We didn't know really. Um, and then for some reason, latched on to this idea of like, well, no one's writing DNA. Can we, can we get it to write DNA? And that was a really simple question, but And in hindsight, it almost seems obvious that, yeah, you would want to write DNA and design it. But when I was first pitching the idea of Evo to folks, I spent six months, which I guess in bio words, not that long, but I spent six months going around saying like, hey, if I generate DNA, like, would you find that useful? Like, what would you do with it? Would you back us? Like, would you want to be a part of this?

58:41Eric Nguyen:And crazy enough, most, almost every scientist at Stanford I talked to thought it was a stupid idea. it was i was like this would be so cool like generating dna um how how much would we accelerate the field and then people would say like what would you do with it i'm like i don't know and then i would get that comment or comments like that's not possible like we as humans don't understand the rules how could you expect an ai to learn it you can't even tell if it's right or wrong like you can't tell the ai yes that's right or wrong how can you expect it to learn it or or there's too many repeat characters, or DNA is too noisy of a distribution.

59:19Eric Nguyen:There's no real rules, and there's just a bunch of junk in there. I heard all the reasons, and I was just so stubborn about it. There's got to be a use case from being able to generate DNA. It just feels right, and I didn't know what it was. So it really was an experiment of what happens. And so when we first trained Evo, I remember we had no idea if it was going to work. Like we had no idea. We didn't know what thing was going to emerge. We just thought, let's just train a big one, which is kind of ridiculous. But somehow, you know, folks that are, they're like, sure. Yeah, why not? Why not? Let's see what happens.

59:54Eric Nguyen:Let's pay some money for the GPUs and let these crazy kids train a model. And I remember when the first result came back and it kind of gave us chills. We're like, oh, maybe there's something going on. which was somebody took the model checkpoint, the first one, and threw it at Protein Gym, one of the protein benchmarks, and turned out to be competitive with protein-specific models. And we were like, okay, that is pretty surprising because we never told the model what's proteins versus not. And there's actually, it was, you know, there's also protein and DNA, right? So there was one piece, I think it was just surprising that it was actually competitive with protein-specific models.

1:00:39Eric Nguyen:And this was like the first experience. And then they just kind of kept on coming one after another, like, oh, this competitive here, oh, it's state-of-the-art on RNA and DNA. And we started seeing like, oh, it can learn across different modalities, like not just DNA. And then we felt like, okay, there's something there. And so it kind of went from there. And then, you know, folks at NVIDIA were like, let's back to Evo 2, let's make this even bigger. And then Greg Brockman from OpenAI was like, I'll take a break from OpenAI and take a four-month sabbatical. and like help these crazy kids out. And, you know, then we're Slack messaging Greg Brockman at 3 a.m.

1:01:11Eric Nguyen:trying to debug our code, which is wild. Yeah. So it just, it went, the trajectory was very surprising in many ways. But at the same time, you still got a lot of feedback like, you know, what are these models good for? What are they, you know, what are they useful for in the real world? And so that's really motivated us to start a company. We thought what we showed was just really just the taste from like an academic flavor. In a similar way, the language models, when they first too, you know, if natural language came out, people asked similar questions like, what are these things good for? Oh, cool.

1:01:43Eric Nguyen:It can write some jokes for me. Is this going to lead to like an all, you know, all comes to seeing AGI that can automate everything? They did not think that, right? They thought it's like a toy. It's got emerging capabilities and they would extrapolate the potential. And in many ways, I think what we're seeing here is even more exciting or reminiscent of that trajectory for DNA. But is EVO 2 even, was showing the bacteriophage. And that wasn't a persuasive. I mean, that's almost a scary example, right? As you mentioned. So the people didn't see that. Like that isn't a light bulb moment for people.

1:02:20Eric Nguyen:Yeah. So, I mean, and some people, right? I think either fair or, you know, rough critiques, you can even, you know, play devil's advocate and say what it's generating is kind of, you know, pretty close to nature and you're kind of recapitulating just kind of a small variance. And so there's, I think there's a lot of ways to critique and kind of, you know, minimize the potential, which can be fair argument. Like, so I think at this point, we think what is more important is not just pushing on the science and like, cool, this can be done and like, kind of leave it there as a sort of thought experiment.

1:02:58Eric Nguyen:But like, how do we actually use this to improve human understanding of disease, improve treatments, make better rare earth mineral extractors? How do we actually make this useful is what we care about now as a company in addition to some of these more scientific questions. So maybe that's a good segue. That brings up two things for me. One is the McIntur stuff that's in the end of the blog post And then also the biosafety stuff. Those are both, I think, important applications. So let's do McInterp first. Can you talk a little bit about, and this is actually kind of building on the work that Evo did as well, I think.

1:03:41Eric Nguyen:I remember in the Evo paper, there was some McInterp work in which they were discussing using the, reversing the question and using the model to extract insights about biology. And this has actually become a common theme with, I think, maybe biology more than any other domain, is that people, because it's a scientific question and it's learning patterns about the world, you can actually say, okay, well, what patterns did you learn? Yeah, so I think Mechaterp is an emerging field in bio that we're extremely excited about. And admittedly, we are on the early side, I'd say. So we're building up that team.

1:04:21Eric Nguyen:But so far, what we've showcased and been excited about is looking, really analyzing the embeddings and some of the activations in the model. And so this idea of Mechinterp for bio is borrowing a lot from the natural language community currently. You could think of what the models are doing is compressing a bunch of information it's seen, right? And so in this compression, it's really basically distilling it down into the key components, key patterns that helps it understand the data or learns the distribution of the data. And so what we're trying to do is probe the models to see what did the model distill into its weights and its activations or sort of like the outputs.

1:05:00Eric Nguyen:And so for us, we started off with a lot of the outputs of the model, so the activations, and we wanted to see what kind of structure, what kind of visualizations can we see that help us understand some of the complexities of DNA, which is a ton, right? And so some of these complexities can range around GC content, certain motifs of transcription factors. There's a bunch of regulatory types of patterns and motifs that the model, we believe, has to pick up to be able to do its tasks, right? To understand whether a disease is caused by a variant. And generally, it's going to compress all these different motifs and distill them into the model weights.

1:05:40Eric Nguyen:And so our job is to then find these and see if we can distill a pattern or structure that can be generalized to other cases where we don't understand the patterns. So that's sort of largely the goal. So in this case, GC means the two nucleotides, the fraction of those in a sequence. Exactly. Yeah, so the fraction of the G and C letters in the genome, which is, I guess, amongst the more simpler things, but also, you know, simple things as repeats, number of retiefs. I think transcription factor, TF motifs, is another one that is especially interesting for folks because eventually we're going to get to a case where we can design transcription factors, meaning we'll have certain transcription factor patterns via like chip seek modalities.

1:06:24Eric Nguyen:Transcription factors are, well, you can go ahead. Transcription factors are molecules that combine to DNA and they would alter essentially the gene expression pattern. And so it has this regulatory effect that doesn't modify the DNA itself, but can modify sort of the effects of DNA and the products that DNA makes. It can have a lot of implications on, well, pretty much everything in your body. So it can modify disease states, it can modify, I guess there's a lot of aging related research around transcription factor design. And so we think being able to understand some of the motifs via DNA, but also additional modalities will eventually let us be able to design transcription factor patterns as well.

1:07:09Eric Nguyen:And so I think this is very exciting. In many ways, the combinatorial space of learning these transcription factors, like what binds and where they bind and what effects it causes, is just far too vast to be able to do this in a manual way. And so we want to take a data-driven approach to learn some of these motifs. The complexity here is partly because the transcription factors are themselves coding genes. And so that you can write those can regulate each other. And so that you have this. That's where that combinatorial effect. Yeah. So like just narrating for the listeners only, we're kind of marching through these increasingly complex and higher level factors all the way from GC content.

1:07:51Eric Nguyen:We started now, we're looking at disease, which is maybe the most complex thing you could or you were pointing towards in this analysis. Yeah. So I think what we our first idea for previewing this was and we're working toward this broadly is this idea of mapping the manifold of disease. Right. So manifold, there's like the sort of representation space of what disease looks like to a model in terms of the output scores or embeddings. We believe that there's a lot more structure that can be gleaned from understanding some of these outputs. And so, you know, mapping this manifold or the landscape of what the structure of disease and obviously many diseases.

1:08:35Eric Nguyen:And so I think would be a big, exciting area of research for us to actually drive motivation for mechinterp in bio. I think we're just really scratching the surface because if you can get this much of, clean this much of insight potentially from just DNA, which is in my mind just one of the sensors that you want to ultimately fuse into, you know, modeling biology, then being able to do something similar across all modalities is something like, by all modalities I mean protein, RNA, epigenomics, you know, the attack, chromatin accessibility and methylation patterns. all these other different types of sort of molecular phenotypes around DNA just presents such a huge opportunity that is kind of laid in front of us that is all green space, green green field.

1:09:28Eric Nguyen:No, I haven't seen anybody do this level of sophisticated techniques from, you know, machine learning, deep learning into what I think is going to be the most important, impactful area of understanding and applications for AI. That's what we're excited about. And I think we're just showing a preview, a very simple preview from the DNA-only models. But in our next generation of models, which will be increasingly multimodal, we're talking dozens, it's a very exciting moment for us to showcase. Sidebar on that, is language, natural language, one of the most? Not yet. Not yet. Yeah, but it will be.

1:10:05Eric Nguyen:So I'd say there's a lot of questions about how to fuse that with bio, but I actually think it's pretty will be relatively straightforward I think the more tricky the tricky part for us is actually how to fuse the biological signals more so fusing language there's a lot of examples with that with like with image and video space so we feel pretty good about that and we've done some early experiments with language and I think that will make it extra accessible for folks when you can connect it to language but I think the part that recipe that folks still are trying to figure out is how to do this across biological modalities.

1:10:43Eric Nguyen:It just seems like a natural way to be able to do chain of thought, right? Exactly. Yeah, absolutely. And I think, you know, especially when you start having chain of thoughts from orchestrating tools, things like cloud science, I think the idea of incorporating language is already in a lot of people's minds. If you don't have any questions, let's talk about the... Bios security. For us, we as a company thought it was very important to have a dual mandate, what we call a dual mandate. It's this idea of essentially being cognizant and feeling responsible or wanting to feel responsible for the capabilities that we're enabling on the design side.

1:11:26Eric Nguyen:So if we're going to create models that can design function into sequences, we believe and see a gap in companies being able to safeguard that technology and make sure that it's used responsibly increasingly more. I was just at a panel last night, a panel on AI scientists, agents that can do scientific discovery. And one of the last questions was, what are some of the biggest risks or doomsday scenarios with AI learning about science? and every one of them talked about biological weapons. And at the same time, I was curious because I was like, okay, so then what are any of these folks doing about that?

1:12:08Eric Nguyen:And basically, I didn't hear anything about that. And these companies, I won't say their names, and I look at the companies, they don't have big efforts in those spaces. So anyways, we felt it was important as a lab that a team that was both building the design capabilities is actually also best suited for building the defense capabilities because they're basically the same models. A model that is good at generating turns out is also very good at discriminating or predicting if a sequence is pathogenic or not. So we felt it was not just from a principle standpoint necessary to work on both biosecurity and design, but that it was strategically, it just made sense as well.

1:12:48Eric Nguyen:And so we felt this resonated with our team, but also the broader community, folks in the U.S. government and abroad even, that there was a clear gap and need for a type of entity to exist to do this. And so, yeah, we felt it was necessary to build it into our mission. And so for us, what we have done and plan to do, we look at it, or I should say biodefense, biosecurity sort of has three or four different pillars of, you know, strategy for biodefense. The first one is around detection and surveillance, right? You know, broadly, can you detect from the environment if a sequence has a pathogen in it, something that can cause disease?

1:13:29Eric Nguyen:Let's say from, you know, swabs at an airport, a nasal swab or a sewage system, right? Collecting samples. The next is attribution, which is once you detect danger, can you figure out where it came from? Is it natural? Is it from, you know, a random country abroad? Is it engineered by a human from a specific lab in a country? And that helps you figure out, you know, what to do about it, right? So this next idea is around countermeasures. So once you've detected, figure out where it's from, what do you do about it? Can you make a counteragent? Can you make an antiviral or antimicrobial? And then the fourth one's generally around deterrence, but that's more of like a government kind of level thing.

1:14:09Eric Nguyen:But yeah, so we focus primarily on the first three. so we create tools that can given a sequence detect if it's pathogenic but also what we felt was missing from the community was not just detect you know broadly if it's pathogenic but characterize the heck out of it meaning what parts of the sequence are dangerous what genes for example attributing where it came from being able to not just look at its you know sequence and match it to a database but be able to attribute its signatures and then also a big component to make this effective in the first place. The community, the biodefense community broadly focuses on sequence matching.

1:14:50Eric Nguyen:So they'll take a sequence and they'll basically align it to a known database and say, have I seen this before? Does it match this list of known pathogens? But I think what's emerging and a concern for a lot of labs is, well, one, new stuff, right? If it's not on your list. And two, things that were intentionally obfuscated to not be detected in sequence space, meaning the letters matching up exactly, but also function space, right? Because basically models that we're enabling now, they will be able to be function aware or structure aware. And so that means for concretely, you can have a sequence that has the same function, like a pathogen, but actually look different in terms of the letters.

1:15:36Eric Nguyen:And you can imagine, basically, what we have on the screen here still is this Mechinterp thing, and there's this sort of manifold that the model constructs internally that is kind of coding for these, among other things, functions. And so you can imagine how it would be able to say, oh, well, that's maybe genetically quite different or at least somewhat different, but it still has a similar function. Exactly, exactly. So what we talk about in our defense blog is this idea of things that can function similarly. They can start having separation in terms of what the sequence looks like while maintaining the same functionality, right?

1:16:18Eric Nguyen:So this can happen in nature sort of naturally, but also what these AI models allow you to do is also intentionally do that as well. So have the same function, functional capability, but have diverse letters, basically, diverse spelling, but describe the same thing, basically. And so in this case, this work from Microsoft called Paraphrases. So it's like, you know, kind of rewording things where they showcase that you can, for example, use protein language models to essentially keep the same structure, which structure implies similar function, but then change the spelling. Right. And not just that, they wanted to test that if you have this capability and you send this through existing detection systems, would it break the system?

1:17:05Eric Nguyen:Like, would it actually detect it or not? And I think one of the interesting facts that maybe the general public, but most biologists know, is that there's these DNA synthesis companies, right? Where you can basically send a design of sequences and get back a DNA molecule like it's an Amazon package. Like you just send it off and they'll send you the physical DNA of design. They'll manufacture it for you. And this runs the scientific community, right? It's the pipeline that allows people to do research and understand biology and make drugs and everything. So it's prevalent and it's public. And so I think one of the concerns for folks is, well, and one of the concerns for us when we first started working on this was in a world, for example, where agents are prolific online.

1:17:54Eric Nguyen:presumably billions and trillions of patients building and taking all sorts of actions on the internet it's wild that they don't have any tools to basically tell it if it's making anything dangerous or not and so that was literally our first motivation we should probably make something that can detect if something is dangerous or not so you guys have a filter for these manufacturers that they can say what am I making here is this dangerous or whatever and maybe you could have an exception if you were like some licensed lab or something and I'm doing something dangerous, I know I'm doing it, please let me do it anyway or something.

1:18:29Eric Nguyen:All sorts of cases. So that's one scenario. To be fair, many of the DNA synthesis companies have detection tools, but I would strongly hypothesize that they're not AI-based and fairly, they're probably not as robust. Yeah, they're all pattern matching mostly. Sorry. Yeah, please, please. Maybe I should ask this question later. I'm just curious about, you know, we've all, everyone who is working in bio has tried to use Fable and instantly got picked out on literally everything you type in. My website, for example. Yeah, yeah. You know, you can't do anything in Fable without. But in terms of people like scientists exploring synthetic biology and creating new sequences and designing new sequences, it seems like it'd be very hard to get, you know, to have an ROC curve, which you can live on, that doesn't like impede novel scientific research, which for legitimate purposes, how do you avoid, you know, even with an F1 of 0.99, which you're not even close to right now, I think, you know, that still could easily, if there are billions of sequences, you know, generated a day or at least like a year.

1:19:43I mean, And I think you could really have a lot of, it seems like a very hard balance to.

1:19:48Eric Nguyen:Yeah, it's a tough question, right? So I think broadly, the way we look at it is, you know, if we thought about like, how do we 100 % stop the dangerous design? I think it's a harder question to ask. I think the question we asked is, on the capabilities design side, there's plenty of folks pushing the frontier of that. when we look at the defense side, do we see frontier technology being applied there? And the answer to that was no, right? So we saw this huge gap and we wanted to sort of aid, come to its defense, I said, come to its aid to give it a boost, right? So I think that perspective, it's an easy choice for us to say, let's push, let's bring that to a better, you know, head-to-head match against the design side.

1:20:37Eric Nguyen:Is it going to solve everything? Well, I think that's what we're going to aspire to be. But realistically, there's always going to be cases where it can get around, right? And I think that's a big motivation for why we think safety for language models and chatbots is one layer, right? But also, you're right, there may be things that kind of just get out there, get past that anyways. And so what do you do about those cases? It's already past the chatbots, right? It's already aided someone into making dangerous sequences. So it's out there. I think the cool thing about our tools is that what we're building is tools for the folks that care about things that's already out there, out in the environment, that's made it somewhere.

1:21:20Eric Nguyen:And now there's this whole ecosystem that we want to build into that does the surveillance, that does the attribution, that does the countermeasures. We want to boost that community, right, and build stronger tools for that space. Does it have to be perfect to be useful there? I don't think so. I think we can be helpful and move the biodefense community forward in terms of bringing AI technology and AI frontier technology to their aid without being perfect. So that's kind of how we would do it. Maybe another way to say, so threat model, it's not even clear what threat model you're actually trying to defend against at this point, but the capabilities don't currently exist at all.

1:21:57So, you know, you provide something and now this can be worked on with in terms of a larger regulatory framework or government or, you know, nonprofit, whatever larger framework. And it now provides you a tool that the community can build upon. Yeah. Even if it's not perfect, it's it provides a starting point. And if you don't have a tool, then you can't do anything.

1:22:22Eric Nguyen:Yeah. The bar is the improvement, right? So the bar is like, where are you at now? Can we move the needle? Yeah. Can we move beyond sequence-based matching, alignment matching? Absolutely. I think there's tons. And I think the interest is only growing. I think we're hitting a lot of chatter at the regulatory side, you know, from different politicians and different folks in think tanks. It does seem like folks are mobilizing. So we're optimistic that this gets out and is more top of mind for folks to actually take action as opposed to just talking about it. Because right now it does feel like there's a lot of talk.

1:22:54Eric Nguyen:and one of the reasons why we felt like let's build tools and like put it out and give access to people because there's been a lot of talk about AI companies saying like we should do biofdefense but then like okay what does that mean like what are you going to do about it and so that was our approach so I can just steel man this for a minute though can you go to the other diagram where you have the yeah and if you click on them then they show other similar compounds or similar genomes so arguably just steel manning the the opposing viewpoint if you are on the frontier which it seems like you are and you certainly strive to be you're moving the frontier of the attack and the defense at the same time so like right now that what you're doing clearly helps but that that maybe if you're on the frontier then that doesn't really matter for in the long term so how do you think about that?

1:23:47Eric Nguyen:I mean, I guess you just have to nerf what people are doing or something. Yeah. The mindset of what was hoping to get folks to start thinking about it as less of a, oh, let's make a biodefense tool and like call it day. It's basically an arms race, right? It's similar to the cybersecurity community. You're going to make better technology, especially with something like Fable that can potentially attack. And that means you're going to have this back and forth. The design side is going to get more capable. The defensive side needs to try to get ahead. Right. And then that, then that just motivates other folks to do past that too.

1:24:25Eric Nguyen:Right. Uh, on the design side. So I think inherently there is this arms race style dynamic that, um, the way it appears to us is that the defensive side has been far, far lagging. And so what we want to do is bring the defensive side closer to, to par essentially. So that, That's the mindset I picture or the framing I think about it. The cybersecurity analogy is interesting, but I think it differs in some key points. So first of all, I think that with cybersecurity, with a sufficiently strong model, you might actually be able to close all loopholes, which are not sociological. There are certain ones which will always be hard, always be ways of getting around things.

1:25:06But you could, in principle, catch every single exploit. And I think that might be possible in the future. and then you can patch them. And as long as people say updated, you know, you're secure. We have fixed genomes, right? So you can't patch a human. I mean, so once something, so I think that in some sense, the, you know, attack surfaces or the way that you defend against that is much higher or much harder. But then maybe the converse is that the, it seems much less likely that someone would have the incentive to go on offense to the same degree. and also the barrier to entry to success, even if you can print out arbitrary DNA and the process of going from that to making a successful virus, especially one which doesn't kill the person designing it, is actually quite large.

1:25:58So I guess maybe I'm curious, what is the single, from your opinion, what is the single biggest threat that we actually have? What would the thing which keeps you asleep or keeps you awake at night, is there something in particular or is this like you just think this is something we need to build and let's build it?

1:26:18Eric Nguyen:One, it's hard to get in the mindset of a person wanting to design a battle weapon. Yeah. So we're not trying to necessarily think of all the potential people or scenarios that bad actors might work on. What I think broadly, what I worry about is we're lowering the bar for how much expertise is needed and the speed at which folks can engineer these kind of things. And that means the volume is going to just exponentially increase at some point. And so there's a mix of, yes, there's intentional worries because there's certainly state actors that have had biological war programs, very, very, very large ones.

1:26:59Eric Nguyen:And one of our advisors on a company has physically seen these facilities and decommissioned them. And so we've heard a lot of stories about states actually being motivated to create such weapons. That is one concern. And I think during peacetime, it's less scary. During wartime, it's particularly scary. I think the unintentional ones are also things that, in my mind, potentially more likely in the near term, where folks do try to generate things and control function and inadvertently things that maybe they tried to make a certain... thing to understand and to study, but it got out, right? Because these things can be hard to contaminate, to contain, for example, and it's just a leakage.

1:27:45Eric Nguyen:So I think those scenarios potentially seem the most likely. Does it keep me up at night? Not necessarily. I'm much more of an overall an optimist. I think building this type of technology is ultimately a game that you weigh out the pros and cons and the costs and benefits. I think the benefits far outweigh the potential harm. And so that's why we work on it. Because ultimately, we do think it's going to be an engine for discovery and human health improvement. But at the same time, we just felt the defensive side was sort of losing this arms race. And so we work on it as well and try to push the frontier.

1:28:23Eric Nguyen:But I'd say overall, I think as a community of researchers, but also folks in the policy side, I think the community is resilient enough and has the ability to mobilize to get ahead of it. So I'm very optimistic about it. We have two questions that we like to ask every guest. So the first one is if you could, by fiat, remove a bottleneck that is important to you, what would that be? Interesting. Could I give two answers on this one? Sure. the first one is um kind of a cop-out because every al lab says this like gpus gpus we've done that a few times we can use gpus um the other one which i think is a little more philosophical i think in this space and what we're trying to do aspire to do is reinvent how scientists do their work in the space and i think one hurdle we run into is folks who and you see this in many domains folks who are the more expertise you have in something the more pessimistic you become about that space and i think you especially see this in bio where you know a disease area or modality so well and then someone introduced something else new and you're like oh but what about this and this and that and they're very pessimistic and probably rightfully so but also i think what i've noticed at the company, what we strive to do and the folks that we try to bring in are domain experts that do know that field, but also are still dreamers, meaning they still do have that imagination and desire to change how things are done.

1:30:06Eric Nguyen:I think that barrier, we see that a lot in the field. I think if we embrace that more, we can see a lot more progress and step change that I would love to see. Brings us to the last question, which is, yeah, is there something that you want the audience to take away a single message? Yeah, I think one message is, I think folks have had, especially into the AI research community, felt like there was a choice that they had to make sometimes to either work on the frontier of AI technology, and that was like consumer-related apps or enterprise-related apps. And like, they just had to work on chatbots.

1:30:46Eric Nguyen:And that's the cutting-edge technology. but also do you know i think when people think about the true potential what ai can do and i think a lot of it is about improving human health understanding our biology but people have felt like they have had to choose and like if i work in that space i can't work on the frontier of ai and what i would like people to take away is that you don't have to choose i think we can work on things you truly care about that you think will push humanity and work on cutting edge technology. And that's what we're trying to build at Radical Americs. Yeah. And I mean, clearly, and I encourage people to read the blog posts, the architecture, the Mech and Turb.

1:31:27Eric Nguyen:There's a lot of innovation that is going into building these models. And I firmly believe that biology is really on the forefront of AI. Awesome. Yeah. Glad you feel that way. Amazing. Thanks for chatting. Yeah. Thank you for making the long trip. Anytime. Appreciate the invite. I had a blast. Great. Thank you.

From the publisher

The OpenAI → Hugging Face attack has people asking “what else do we need to worry about?” and Anthropic’s filters flag two things: cyber-security and biology. The natural question is: what about bio-security, then?

Clem Delangue argues that cyber-warfare defensive capabilities need to be open and to keep pace with frontier models’ attack capabilities

Radical Numerics co-founder Eric Nguyen sat down with us and explained why the same models that increase biological capability can also keep defense from falling behind.

Building a virus from scratch

While he was at Stanford, Eric couldn’t get traction on Genomic Language Models (GLMs) for a long time. Biologists didn’t believe it would work, didn’t think they could verify the output, and didn’t see important applications beyond what they could already do. He kept pushing, eventually helping lead the development of Evo and contributing to Evo 2 at Arc Institute. Those models were later used by a separate Arc/Stanford team to generate entire bacteriophage genomes that were synthesized into functional viruses!

Long context unlocks biological intelligence

Early ChatGPT spit out poems and email, and early DNA language models like Evo and Evo-2 could build a genome from scratch. DNA is different, however, from natural language in that it has a very small alphabet (4 characters ACTG) and that its sequences are very long:

* 60K for an average human gene

* long being up to 2.3M

* the whole human genome around 3B.

Innovation in long-context models made this possible about 3 years ago (footnote: striped hyena), long before the frontier labs were building 1M+ context models.

Now Eric and other AI x Bio luminaries have founded Radical Numerics to build and scale GLMs to tack a wide range of biological problems, extending well beyond generating DNA.

Thinking in DNA

Their GLMs already do pretty well with RNA and protein because there are clear markers in the DNA sequence for genes (RNA sequences the perform many functions) and specific genes that encode proteins. This means that the models already generalize to multiple “languages,” before even attempting to train in other modalities, such as 3d protein structure, epigenetics and natural language.

If a model thinks in the DNA language, maybe it understands the imprint that environment left on different genomes as well? Perhaps the model has learned the functional relationship between different sequences, and could extrapolate to new sequences based on that?

And so what we wanted to showcase was that if we show the model progressively better RNAs in a series of steps with its score, right? So you have like low scores first and then you gradually move up the chain. Can the model continue that trajectory on its own? And then in the final step, does it self optimize to a point where it's like the best score it can get? That was the experiment. Can we do that? And so we took a data set, a large data set of aptamers. We held out a portion of the best performing ones and we showed it only the lower ones, but then we ranked it, right? So we showcase lower scores with the RNA aptamers and then progressively got higher, and then ask the model to just like continue with that pattern. And it turns out it was able to recapitulate some of those higher scores that we had not shown it yet.

So, voila: chain-of-thought, thinking in DNA!

The arms race

But much as long-context inference, chain-of-though and multi-modal perception unlocked sophisticated reasoning in natural language LLMs, these capabilities in GLMs are enabling increasingly sophisticated “biological intelligence,” and along with it, greater danger.

According to Eric, defense is currently losing this battle, but Radical Numerics argues to push the frontier harder!

I won’t spoil the details for you. In the episode we talk in detail about:

* Biosecurity as an arms race — and how defense can keep up

* The genome as the imprint of the environment on DNA

* Going truly multi-modal

* How chain-of-though works when you “think” in the language of DNA



This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

More from Latent Space: The AI Engineer Podcast

All 247 episodes
🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)Latent Space: The AI Engineer Podcast · 1 h 32 min
Listen in VO