In short
Episode Notes: Long Context Language Models and their Biological Applications with Eric Nguyen - #690
Episode Overview In this episode of *The TWIML AI Podcast*, host Sam Charrington interviews Eric Nguyen, a PhD student at Stanford University, focusing on his research involving long context foundation models, particularly the Hyena and Hyena DNA models, and their applications in biology. The conversation covers key concepts in language modeling, architecture motivations, model training, and the future of these technologies.
Key Concepts and Discussions
Background of Eric Nguyen
- Field of Study: Bioengineering with a focus on machine learning.
- Research Interests: Applications of machine learning in biology, particularly regarding long context sequence modeling.
Hyena Model
- Architecture: A convolutional-based language model designed to handle long context lengths efficiently.
- Challenges Addressed:
- Standard transformers struggle with long sequences due to their quadratic time complexity (O(n²)).
- Hyena aims to reduce computational costs while maintaining expressiveness.
- Training and Structure:
- Utilizes a parameter-efficient implicit kernel for convolutions.
- Explores Fast Fourier Transform (FFT) for computational optimization, reducing the complexity from O(n²) to O(n log n).
Limitations of Transformers
- Time Complexity: As sequence length increases, resource demands become impractical for standard transformers.
- Exploratory Gaps: The need for models that can effectively process longer context sequences beyond the capabilities of existing architectures.
Hyena DNA
- Genomic Application: A model pre-trained on 1 million DNA tokens, designed to capture long-range dependencies within DNA sequences.
- Architecture Evolution: Transitioned to a 7 billion parameter hybrid model (Evo) that combines convolutional layers with attention mechanisms to leverage strengths from both architectures.
Evo Model
- Goals: A larger foundation model incorporating learnings from Hyena DNA, focusing on genomic sequences from a variety of species.
- Key Features: Designed for DNA sequence generation and understanding long-range dependencies critical for biological functions.
Applications in Biology
- Potential Uses:
- Drug Development: Designing novel drugs and molecules.
- Gene Editing: Applications in CRISPR-Cas systems for precise genetic alteration.
- Challenges:
- Ensuring quality and reliability of generated DNA sequences to prevent harmful mutations.
- Evaluating model performance against biological benchmarks.
Evaluation and Performance
- Benchmarks:
- Focus on predictive and generative tasks, assessing model sensitivity to mutations and fitness scores.
- Comparison with state-of-the-art models, demonstrating competitive performance in various biological tasks.
- Zero-shot vs. Few-shot Performance: Initial focus on zero-shot capabilities to demonstrate generalizability across modalities.
Future Directions
- Scaling: Research on scaling laws for DNA models to improve performance.
- Complex Organisms: Future work aims to model more complex organisms like mammals to enhance therapeutic applications.
- Multi-modality Potential: Exploring integration of different biological modalities for broader applications in health and therapeutics.
Key Takeaways
- The Hyena architecture provides a promising alternative to transformers for long sequence modeling, particularly in biological applications.
- There is significant potential in using AI for genetic research, drug design, and understanding complex biological systems.
- The development of advanced models like Evo represents a critical step toward leveraging AI in healthcare and genomics.
Additional Resources For complete show notes and further details, visit: [TWIML AI Podcast Episode #690](https://twimlai.com/go/690).
---
This markdown note captures the essence and detailed discussions from the podcast episode, outlining the key concepts and future directions in the field of long context language modeling and its applications in biology.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:28All right, everyone. in biology, in particular hyena and hyena DNA, and most recently Evo. To get us started, I'd love to have you share a little bit about your background and research interests. Sure. So I am in the bioengineering department at Stanford, but my background is mostly in the machine learning side. I started off in computer vision and started working with state-space models and seeing if they can be used for some of these vision and language domains, and then started moving on to see if we can extrapolate them and use them for biology applications, in particular for their long context purposes.
1:04And so I've been exploring these applications of hyena in particular, which is a convolutional-based language model, to the domain of biology, in particular DNA. You know, let's start by talking a little bit about hyena and some of the motivations for that architecture in addressing longer context lengths for language modeling and other sequence models. Tell us about the big challenge there and the motivation. So our lab has been interested in this long sequence task for quite some number of years. And it started with Albert Gu, who was interested in state-space models at the time. And so initially their work in our lab was applying these models to continuous signals typically.
1:48So, like time series and audio. And so, when language models started to become really popular, we too were interested in applying it to language, but we noticed there was a gap in applying some of these early state space models. And so, the Hyena architecture was really focused as an offshoot of these state space models to see if we can get it to work on language. And so, we focused on getting it to work on these discrete signals that we think of as language. So, Michael Polley, the lead author in the Hyena architecture, and Stefano Mazzaroli, the other co-author on the Hyena paper, we started focusing on what's the gap and getting this gap to close between transformers.
2:24And so that was the initial work that they focused on. And it's when we started getting excited, we were able to sort of match similar quality, but then explore this opponent, which was being able to fit longer sequences. And so that's when we got started getting excited about new applications. So the obvious question is transformers have proven to work very well. Why the need for an alternative? What is the challenge that is presented with the transformer architecture in dealing with longer sequences? Yeah. So, you know, at the time when we were working on it, this was still in the age where not many folks were thinking about long context in particular.
3:01And so there's this key constraint with transformers, which as a sequence length grows, the time complexity, right, the algorithmic efficiency of the operation has this squaring law that grows. And so when you double the sequence length, the computational requirement quadruples. And so when the sequence gets very long, if you care about things like DNA or, you know, entire code bases, then it becomes intractable for these models. It can become intractable for, you know, these classic transformer models. And so the motivation is, is there a different type of operation or primitive that can be hopefully as expressive, but will reduce that complexity in computation for these longer sequences?
3:44And you mentioned one of the early attempts to try to address that in particular with state space models. I'll refer folks to my interview with Dan Fu from actually just about exactly a year ago. Language modeling with state space models was the title of that episode. And so this work builds on some of that state space work. Does it retain some of the kind of hierarchical approach that made that that made the state space work? Yeah, it uses a good portion of it in terms of its skeleton. So the differences that we took from that paper with Dan Fu on state-space models for languages was how we parameterize the convolutional kernel.
4:29So a key way to make convolutions work for long sequences in particular is to move on from this sort of explicit parameterization of kernels to an implicit version. And so what that allows you to do is essentially have a parameter efficient way to use convolutions. And so the state space model would use state space models for this kernel. And then the original hyena work we would use, we sort of generalized it to just use an MLP. So an MLP that would parameterize the filters for a convolutional kernel. And so the thinking was to have an unconstrained version that can learn perhaps a more expressive kernel for the convolution.
5:05And your research group published a really interesting blog post last year talking about this scope of research from deep to long learning is the post that I'm referring to. And in that post, there's a really interesting graphic, and we'll link to all of this in the show notes. But one of the elements discussed in that post and the centric to this graphic is the idea of the FFT as kind of a new primitive that's driving this alternative family of models. You know, talk about the role of the FFT in Hyena. So for Hyena, we rely on long convolutions. And so typical convolutions, they'll have short filters, right?
5:48And so the time complexity of having a kernel stride across an input is still basically a linear operation. But when you have a convolutional kernel like Hyena, the key difference is that we're going to have a filter the same length as the input. So a global convolution. Now, the challenge is the time complexity will increase. So you essentially get back a time complexity that's similar to a tension, like a n squared time complexity. The kernel is the same length as the input. And so you lose that benefit of a lower time complexity if you just apply the convolution in this standard way that people think of in the space domain.
6:27Instead, we'll take advantage of a convolutional theorem for folks familiar with signal processing, which states that a convolution done through the Fourier domain is the equivalent as an element-wise multiplication in that space to its convolution in the space domain. And so what we'll do is we'll use the FFT fast forward transform to essentially convert the input kernel to the frequency domain, do an element wise multiplication, and then bring it back to the space domain. And that whole operation allows us to reduce time complexity from an n squared to an n log n operation, which is for long sequences near linear.
7:06here. And do you retain the benefit of being able to accelerate using GPUs? Yeah, that's a good question. So in general, GPUs are a lot more optimized for matrix multiplication. And so there's probably some room for improvement in terms of fully utilizing the hardware. But in general, we can get, you know, fairly efficient convolutions using GPUs. And that's what we do for language model training for Haina. Now, there's follow-on work from Dan Fu, which I also helped out on, called Flash FFTConf, which focuses on trying to bring the... Flash attention plus this FFT work? Exactly. Bringing the sophistication from hardware-aware algorithm design from Flash attention to the FFT.
7:51And so that was really exciting work as well. And going from state-space modeling to the long-sequence convolution, Do you lose some explainability or mechanistic understanding in doing that? Having a states-based model that describes explicitly how the state evolves from one state to the next seems like it would have some of those benefits. Is that true at all? And do you kind of abstract away from some of that with a convolutional approach? Yeah, that's a good question. I haven't focused too much on the explainability of what these state-space models are learning in terms of their state-space representations.
8:34So it could be, but I think what we've focused on more so is using sort of attention matrix-like view of the convolutions. so even though we use convolutions we can also still produce an attention map you know quote unquote which allows us to essentially see what inputs are being attended to but but more so in convolutional like what's lit up right in terms of attributing weights essentially so that's not been a constraint so far for interpretability talk about how far you've pushed the hyena model in terms of input size or context length? Yeah. So in the initial work, we started off with 64 ,000 tokens for the hyena hierarchy paper on synthetic tasks.
9:21And then in a follow-on work, we applied it to DNA sequences in that it's in constant need of having long-range dependency modeling. When we created a model called Hyena DNA, which we scaled up to 1 million token context, which at the time was the longest language model context since then. Lots of folks have caught on and found this kind of work also interesting and also get Transformers to work as well with Long Context. So interesting to see the work evolve or the space evolve. Can you talk a little bit about the landscape of Long Context's enablement on the Transformer side? What are the approaches that folks are taking there and what's working?
10:04Do you have a sense for what Gemini is doing to, you know, get at 1 million, 2 million context windows and how those approaches will fundamentally differ from swapping out transformers for convolutions? Yeah, so I haven't. So, I think I've heard some rumors about how Gemini might do it. But, yeah, I actually don't have a strong idea myself. there's you know talk of using things like ring intention or some kind of sliding window attention where it's attending to a local uh window and then having a hierarchy where these local windows can get aggregated into more meaningful uh groups so i yeah i'm not sure actually but um i think I suppose, yeah, I think there's a lot more that could be explored.
11:00I think that sort of my takeaway is that folks have focused a lot on the memory reduction of using attention for long context. But there's sort of no way to get around, or at least I'm aware of, of reducing the time complexity. It's still going to be an N by N type of operation or N squared operation. and so you can reduce the hardware requirement but then the amount of time is still going to be pretty key at least with the current setups I've been seeing and so I think that's sort of motivated folks to look at these alternative architectures if they do care about long context including Mamba, obviously very popular and we're sort of placing our bets on a slightly different approach with hyena variants but in general I think quality of these long context model with a transformer or alternative architecture seems to be something that people are still trying to improve like you can fit it in context but does it you know generate high quality sequences throughout the long context for example still an open question of something we're observing and is hyena comparable to transformers from a memory utilization perspective?
12:19Or, well, I guess the implication is that it's much more efficient.
12:25Yeah, so in terms of, you know, if we compare to vanilla transformers, vanilla attention algorithm, yeah, we're definitely in reduction in terms of memory and time. The, I guess, bar is raised when you look at flash attention, you know, a very clever, sophisticated algorithm that's essentially learned how to linearize the memory by doing it in blocks, right? And so, I think that is a hard bar. Like, attention is super fast. And so, I think there's still room for improvement to get convolutional-based models to that degree of sophistication. But in certain... Yeah. So, the sort of answer is depends on which settings you care about.
13:04There's certain very long sequences that will still take very long on flash attention or transformers, which could be faster on the convolutions. But the memory side, the full benefit is not there yet because they probably have some more room in terms of the systems optimization. And you alluded to this a moment ago. In training Hyena, the training corpus was synthetic language as opposed to a traditional training data set. Is that am I remembering that correctly? And what does that mean? So, in the original hierarchy paper, we trained on language for, I believe, 2K context on the pile. So, like a pre-training perplexity comparison with transformers.
13:50And that essentially was stacking up favorably for Hyena. And then to push specifically the long context, at the time, there wasn't a lot of great benchmarks for analyzing long context or, you know, very useful ones that we thought. And so, we came up with or used synthetic tasks that focused on a specific task, which is recall. So, being able to look up portions, sort of like a key value dictionary lookup, look up values earlier in the sequence. And then we use a simple synthetic task of key value scores, like letters and numbers, essentially, just to see if we can do this simple task, but over long range.
14:31And so that sort of was a prototype for even fitting in lung context, but then that really motivated to look for useful applications and that kind of segue to the hyena DNA work. Got it, got it. So let's jump into the hyena DNA work. How did you identify biology as an application for this problem? So, part of it was I was always looking for an impactful application in healthcare or biology when I started my PhD, sort of like the right opportunity or right technology to dig in deep. And so, I sat in a number of different labs on campus at Stanford who work at the intersection of biology and computer science.
15:14And long sequences and long-range dependencies kept on coming up in the world of DNA where people can model it just like a language, like natural language in the sense that you have letters, you have a vocabulary, you can feed them to language models. And then you can ideally learn all the interactions and grammar associated with that language to do useful tasks like prediction or perhaps generation as well. Can you talk about how the long-range dependencies kind of manifest from a biological perspective? And how does that differ from natural language? Sure. So, I think the really neat thing that makes DNA an interesting problem area or application to work on for machine learning is that it does physically have these long-range dependencies.
16:03So, in DNA, you have this sequence of letters, essentially. They're actually chemical compounds, but we can represent them as letters, the four bases, A, C, T, and G. And in general, when you're modeling DNA, there's processes that where what you care about are essentially things binding to DNA and kickstarting the process of creating proteins. And so, the process of having something bind to DNA can be influenced by what we call motifs that are present in the DNA that will essentially alter the probability of things binding or not. Now, the interesting thing for long-range dependencies is that some of these interactions can occur over millions of base pairs away from each other.
16:51And so, having a model that could essentially identify the presence of certain motifs and how they affect downstream binding of certain molecules can mean the difference between someone expressing some kind of disease or altering functions or their traits, essentially. And so it's really desirable to capture these long-range dependencies within a single model. So can you talk a little bit about how one might use a long-range genomic foundation model and how that differs from the way someone might use something like an AlphaFold and kind of compare and contrast those approaches? Sure. So AlphaFold, AlphaFold 3 in particular, very powerful, very successful model and has probably the most impact with machine learning and biology so far.
17:39Now, as powerful as it is, it is a supervised task, which is specifically designed to predict structure, right? So, given a sequence of amino acids, and now more recently including DNA, they want to predict the 3D shape of that sequence because it has a 3D shape analog in physical space. For AlphaFold, it's specifically a supervised task. And so, it can be difficult to modify it for use of different purposes. So the nice thing about language models and foundation models, they're meant to be this general learning representation machine that could be modified to be applied to a number of different downstream tasks.
18:17And that could be more difficult or perhaps more challenging for an outfold model that's restricted. And so some of those tasks for a biological foundation model that's language-based can include the prediction tasks, right? So, let's say you have a sequence, what's its function, affect gene expression or how much of a certain protein does this particular sequence expect it to make, all the way to the generative side as well. And so, these language models, right, for natural language can generate lots of text and answer prompts and questions, right? But for DNA or proteins even, what you can imagine is, and what they're currently being used for, is to design novel drugs or novel molecules, proteins, or systems that could be designed just from the sequence space with these models.
19:09got it can you talk a little bit about um how the model is trained the human genome has what 3.2 billion nucleotides which is quite a bit less than the uh the corpus size for a state-of-the-art language model nowadays um but the language is a lot simpler in terms of the number of tokens like what are what's the training process and what are the um elements that come into play Yeah, so you're right. So our initial work started off with training on the human genome, which is made up of 3 billion base pairs or nucleotides. And the vocabulary, you know, depends on how you want to tokenize. But for our case, we use the simple vocabulary of just four bases, or you can consider like the bite level.
20:00And that can make it challenging for a number of ways. But yeah, for focusing on our initial use case, we'll train by just taking in an arbitrarily long sequence from this human genome and then do next token prediction or next nucleotide prediction, sort of a causal style. But we'll do this essentially billions of times, just like natural language. And so, yeah, that is a challenge for, I guess, scale, right? So, you know, initially, this could be seen as like a first attempt or prototype to focus on the human genome. But more desirably, we're following work for Evo in particular, we grabbed millions of genomes and millions of species to try to train a much larger DNA foundation model across different species.
20:52In that case, the good thing about DNA, it's one of the most common or most widely available sources of biological data out there. And so there's definitely trillions of tokens also in this domain that we can reach to to train a foundation model. Let's maybe take a step back and have you introduce EVO. Is it simply this idea of taking the hyena DNA approach and architecture and incorporating additional species? Or did the architecture itself evolve to enable you to do that? So, Evo is sort of our next stage in training large DNA foundation models, taking some of the learnings we had from hyena DNA and scaling it much larger.
21:37So, hyena DNA was about a 7 million parameter model, which in language models are very small. And then Evo is a 7 billion parameter model. So, it's about a thousand times bigger. And to get that to scale effectively, we did have some architecture improvements in this work. So, a couple of key ones is that it's a hybrid model. So, we actually mix in, reintroduce some attention layers with these high-end convolutional layers. And then thinking there is sort of to get the best of both worlds where convolutions can get, you know, the efficiency gain and then attention has somewhat stronger abilities in terms of recall ability.
22:18And so we'll take advantage of those as well. And so we also have some modifications in terms of how we parameterize the convolutional kernels themselves and then change some of the MLP blocks as well. But the overall training style and core of this convolutional style language model from Hyena is in large part used in Evo as well. Okay. And you talked about EVO, the EVO dataset spanning many species. Do you retain the human species from hyena DNA and then incorporate the phage and prokaryotic and those types of sequences? Or is it only the latter? I'm trying to wrap my head around, does it make sense to have all of that in one model?
23:13Yeah, that's a really good question. So, we actually made the sort of strategic choice to focus on just prokaryotic and phage genomes. And so, which, you know, folks who are not familiar with that biology, essentially microbial life. And the thinking there was we're going to focus on these genomes that are somewhat shorter. They're shorter and they're simpler life forms. And their DNA, the language of their, the grammar of their DNA follows more clear rules than humans or mammals, for example. Humans and mammals are quite complex organisms as we can imagine. And so, having a large model trained on humans and mammals or what we call eukaryotes is a level of complexity that we're still working our way up to and actually in follow-on working for EVO2 is what we plan to do next.
24:05But we sort of wanted to focus on simple life forms first because sort of the thinking so far before we started working on EVO was we were unsure that a large DNA model could work in general. We weren't sure if some of the tasks that we were going to do or go after could actually, specifically, to be able to generate and design DNA was a key task. So we wanted to start with simpler organisms first. Dig into that task in a little bit more detail. In terms of generating and designing DNA, what are some examples of reasons why you might want to do that and ways that it might be applied? Sure. Sure.
24:43So, one of the initial questions before we started the EVO project was, is it possible to have a language model or foundation model design an entire genome from scratch? And the motivation for that was a lot of work previously had, you know, or a lot of research in general allows us to write DNA, read DNA, edit any sequence of DNA. But to be able to design DNA, we still didn't know enough about how DNA worked to be able to design. And so, we wanted to use language models to see if we can help design through a foundation model. And so, you know, why would you want to do that if you could? the thinking was if we could design an entire genome then we can perhaps control the function of these genomes as well and so some of the ideas were can we target cancer cells if we can design organisms to to have that function can we make better biofuels can we make antibiotics these are some of the the sort of the moonshots we thought that could be used for um you know manufacturing or climate or therapeutic purposes, for example.
25:51And now that transformer-based approaches are accommodating longer context lengths, could those be as easily applied to these problems? We've seen attempts to apply them to these types of problems. Is it simply, quote-unquote, a choice of kind of training corpus and going through the motions, or are there unique adaptations of this approach that lends itself to the biological use case? So that's a good question. We get this question a lot. And it depends. You know, I think people have different answers. But I certainly, you know, having trained these DNA models for the past year and a half, certainly have my biases and my opinions, which I do believe that convolutional-based models have this inherent inductive bias advantage over attention models.
26:46And I think some of these reasons is that DNA is not natural language, for one. Natural language is a very dense, rich type of sequence or type of data distribution. In DNA, the information is far more sparse and over long range. It's noisier because there's, you know, like the biologists, there's a lot of talk about junk DNA in your genome where people don't know what it does or it might have no function in your genome. And so what we observed is that these convolutional models, we hypothesize that they're better at filtering out this noise and being able to parse the signals over the long ranges better than attention.
27:27perhaps because convolutions from signal processing and they're built to filter out information really well and let's use that analogy but yeah there's other things about modeling dna that make it challenging like being able to tokenize at the single character level or the single nucleotide level in general we've seen transformers just do a number of different works that are trying to model language at the character level and it usually underperforms and they've not been able to use long sequences, for example. And so in particular, it's important to have this single character tokenization for DNA because a single character change could mean the difference between having a disease or not.
28:07And so you want that sensitivity and resolution at that level that transformers, I haven't seen, I don't think anyone's seen a successful large scale language model that's tokenized at the single character level yet. Okay. One of the properties that is characteristic of transformer-based models is hallucination, of course. Is there anything kind of inherent to the hyena models that yields different characteristics in terms of propensity to hallucinate? Ah, that's interesting. I don't think I've played around with natural language enough to be able to identify human interpretable hallucinations, I suppose.
28:55But yeah, I mean, in general, it should be susceptible, I imagine, to these hallucinations, just like transformers. I guess maybe another way to ask the question is, is hallucination a function of the architecture more so than the probabilistic nature of the approach in general or the training approach? If you have a take on whether approaches like hyena or the hyena approach is for some reason less susceptible to it. And I think the context is that hallucination in language can be bad. hallucination in DNA generation sounds scary. Yeah, fair enough. Yeah, I'd say that these systems, they're probabilistic machines.
29:40You're teaching them the probability of the next token. So because it's probabilistic, these probabilistic errors could accumulate and then hallucinations can abound. And I don't think that's going to change by the architecture or the layer that we swap in. It's still fundamentally a next token prediction with a probabilistic mechanism. But yeah, you're right. So in terms of biology, this could be scary as one way to think of it, but it also could be a source of variability, right? Or like a controlled evolution, which that's usually the analogy that folks in machine learning try to describe these generative models in biology.
Read the full transcript
30:16You're controlling... Meaning it's an opportunity for creativity and to use it for the generative tasks that you envision? Is that where you're going? Yeah, exactly. At the same time, it's still not a deterministic thing, right? Because if it was, then you wouldn't have much variety. So, it could be also viewed as a feature, not necessarily a scary thing. And so, to overcome that fear, perhaps, we definitely filter out, for example, in the same way in language you would want to filter out good quality generations. Luckily, in biology, there's a lot more heuristic software tools to be able to filter out high-quality sequences to perhaps appease some of these concerns.
30:56In the original hyena work, we talked a little bit about the synthetic tasks. Do you take advantage of, and you just mentioned some of the existing tools and taking advantage of existing biological tools on the filtering side of things. Do you also use existing simulation tools and synthetic data generation tools? Is that a part of what you've done with HyenaDNA or EVO?
31:31Yeah, that's a good question. We haven't done so much on the synthetic side, actually. But that is something that we're interested in, in particular to see if we can create some synthetic tasks for long-range abilities. I think that's an ongoing search for ideal tasks on that front. But I think in biology, there seems to be so many different applications that we could try that we're trying to be more thoughtful and choosing the most relevant or impactful tasks. So, there isn't a shortage of tasks is what I mean. So, we haven't had to look for synthetic ones. It's more so prioritizing which biological meaningful tasks to go after first.
32:17And in the note of tasks, folks listening may be familiar with CRISPR, which is a gene editing tool that's become popular over the past several years. Is there a direct application of this work to CRISPR? Exactly. And so, the thinking there is that normally, let's say like a drug or pharma company might try to look for existing molecules in nature by just mining a bunch of genomes or DNA from different microbial life. What they'll do is just take the exact copy. They'll just look for it and take the exact copy. What we want to do instead is to learn what's out there and then essentially, conceptually how I think about it, is interpolate between what's seen before to hopefully design something not existing in nature, right?
33:05Something new with the desired function, you know, hopefully that works better, that's more efficient, that enables, in our case, for example, gene editing tools that can work on different modalities. So maybe not just DNA, but RNA, different types of biological sequences, different levels of on and off targets. Like, does it cut DNA where you want it versus cut DNA in areas that you don't want it? Being able to engineer and design these new tools could open up new therapies, for example. That's why we're interested in finding new ones. Mm-hmm. And to make sure I'm understanding this correctly, you talk about CRISPR as a tool and a family of tools and kind of cutting DNA in places.
33:52It's not like it's using a laser to cut the DNA. It's a chemical sequence. And it's a particular chemical sequence in a family of chemical sequences. And you could use generation here to identify chemical sequences in the distribution of that family that may have different properties. Exactly. Yeah. So stepping back to give folks a sense of what these CRISPR-Cas systems are. So in general, they're defense systems. So they're sequences of RNA and proteins. And bacteria will have these sequences in their own genome and they'll use it as a defense mechanism for viruses. And so, what happens when a virus tries to come in and inject its own viral DNA into a bacteria, this CRISPR system essentially will recognize a pattern of this DNA and then essentially cut up that DNA and render it innocuous.
34:53And so, that's sort of this defense system that will both identify DNA and then cut DNA. And so, as humans, we've sort of hijacked that system to use it for gene editing, essentially, as a therapy. And so, specifically what it's made up of is RNA and proteins. So, the RNA is going to be used to recognize where to cut and the protein is going to be the thing that cuts. And so it's this complex of RNA and proteins that we're trying to generate because it's also just DNA as well. Talk about in the case of both hyena DNA and EVO evaluation and evaluation criteria, benchmarks, performance. What have you seen and what are you comparing against?
35:44Yeah, so I'll focus on the EVO one. that's most top of mind and relevant in this case, for the benchmarks that we look at, we break them down into two areas, or I guess our tasks are broken into two areas, the predictive tasks and the generative tasks. And the predictive tasks are the more benchmark-oriented ones that machine learning folks are used to seeing, which is given a sequence, can it predict some kind of function or fitness score associated with that sequence? So, for example, if they have protein sequence, in the lab, we'll have a data set that measures some kind of fitness score associated with the sequence, meaning what's its solubility or stability or binding affinity, metrics related to sort of its drug properties.
36:32And so what we'll do for this EVO work is to see if the model is sensitive to these mutations or perturbations that we introduce and see if it changes the fitness score in some meaningful way. And so we'll take, for example, a sequence of proteins and add a mutation into a portion and see if that mutation can be picked up in terms of how the outputs of the model correlate with the fitness score for that molecule. And so, specifically, there's a dataset called Protein Gym that we apply this to. And the nice thing about this sort of task setup is that we can probe the model to see if it's learned useful features, zero shot or without fine tuning.
37:16And so, what we saw for our work or our performance is that a model like Evo was comparable with state-of-the-art protein models in terms of this perturbation and correlation to fitness score. we thought that was interesting even though it's just competitive or in some cases a little bit better the more compelling part is that evo is trained on raw unannotated dna sequences whereas protein models are trained specifically on just protein sequences so a clear advantage and the main difference there is that dna has both protein sequences and non-protein sequences so if you train your model on the entire dna or genome your model has to basically tease apart and figure out what's a protein and what's not a protein.
38:02And the thinking before this was that most models could not do this. They had to focus. They had to be one model for proteins, one model for DNA, one model for RNA to be competitive. And what we were trying to show is that you can have a single model that trains across all of DNA that can learn all three modalities. Okay. Is there something you can say about, or did I miss in there you saying in terms of relative performance to existing things, state of the art on some number of things? So for proteins, we were competitive with state of the art, so about the same. And then for RNA and DNA modalities, we were able to show that EVOS has a significant performance boost or state of the art on these zero-shot tasks compared to other domain-specific models for DNA or RNA.
38:50And so we focus on those three. Those are the central dogma of biology, what they're known for as. So we try to have a little bit of a benchmark on each of these domains for the predictive tasks. And so just as strong as proteins and then state of the art on DNA and RNA. Okay. And in terms of compute budget, how does this approach compare to the prior state of the art? yeah so when we talk about state of the art it's more so a comparison of absolute performance versus controlling for parameter size or architecture so they do have different parameter sizes and data sets and so it's hard to make that exact comparison and so yeah it's most yeah it's less so focusing on the architecture comparisons well maybe how would you characterize the efficiency of these models relative to other approaches that someone trying to address these tasks might typically use?
39:53Yeah. So, so I take that back. So we did compare, we do a scaling laws analysis of pre-training on DNA across the prevailing DNA architectures. So we do a comparison with transformers, Mamba, and then Hyena and its hybrid variant as well. And so on that front, we do compare performance across architectures, mainly on perplexity in terms of how they scale with perplexity. And so it was interesting to see that these convolutional-based models being particularly strong on DNA over transformers. And again, this kind of alludes back to noting the difference between DNA and natural language being sparser in terms of information and also noisier at the same time.
40:36what did you observe in terms of zero shot performance versus few shot performance and is fine tuning a thing that makes sense in this context and if so how so we focus initially on zero shot tasks in this work because we wanted to showcase um i guess most work previously in biology usually focuses on fine tuning but we sort of thought that was um uh i don't know primarily because you're like a crux taking it off the shelf natural language model and trying to fine tune it to uh be useful in the biological context um i mean that's not necessarily because i think these models still also train from scratch but i think i think it was more so that previous models we observed didn't focus on generalizability and so they typically would always fine tune just because our hunches and what we observe in many models is that the zero-shot performance wasn't that compelling.
41:35They weren't practical for many applications. In this particular work, because we had scale and because we were focusing on generalizability, we wanted to see how well it could generalize zero-shot across different modalities in biology. That being said, we did have some fine-tuning results, which in many of the tasks were also very strong, quite competitive, if not more competitive than the zero-shot results we observed before. Okay. Awesome. Awesome. Talk a little bit about where you see it all going. What's next on your research agenda in this particular line of work and kind of how do you see it evolving more broadly?
42:17So this is something I've been thinking a lot about. So we're certainly interested in both, you know, biological perspective and a machine learning perspective, how to improve. And so from a machine learning side, we absolutely do think scale is going to be really important. That's why we did a set of DNA scaling laws to see, do you get some kind of gained observed that you observe in natural language, also in DNA. And indeed, we do see this similar trajectory for DNA, which is something we're going to focus on. From the biological perspective, trying to model more complex organisms is another key area.
42:53So mammals and humans, we definitely want to build foundation models there because the applications for human health and therapeutics is just huge. If we can be able to have a language model that understands the grammar of our own DNA, as well as we've learned for prokaryotes or microbial life, I think that the application space is far more than what we're aware of right now, I'd say. And so, in general, my sort of philosophy or thinking or hope is that in similar ways that we've seen folks like OpenAI merge modalities in language and vision and gave us things like Sora and Dali, I think there's a potentially bigger opportunity in biology to merge modalities into designing sequences or therapeutics or different drugs because there's so many different things in assays or types of sequences that try to measure similar things in our body.
43:52And so being able to merge those signals to design or generate or predict useful things in DNA or other biological sequences, I think the potential is far greater and potentially far more impactful. And do you see the hyena architecture as getting you there in the same way that transformers are applied to multimodality or will it need significant evolution to accommodate that? I mean, there's always room for improvement, I'd say. I think it's a place where we put our bets because in a resource-constrained world, we've gotten good results in terms of longer sequences, but is it enough to, for example, have a single ginormous biological foundational model that will solve biology.
44:43No, it's not enough to get us there. And so I do think there's room for innovation on many, many fronts, including architectures down the road. Awesome. Awesome. Well, Eric, thanks so much for taking the time to share with us a bit about what you're working on and Hyena and Hyena DNA and Evo in particular. It was a great conversation. Really appreciate it. Awesome. Super fun. Really fun to be here. Thanks, Sam.
From the publisher
Today, we're joined by Eric Nguyen, PhD student at Stanford University. In our conversation, we explore his research on long context foundation models and their application to biology particularly Hyena, and its evolution into Hyena DNA and Evo models. We discuss Hyena, a convolutional-based language model developed to tackle the challenges posed by long context lengths in language modeling. We dig into the limitations of transformers in dealing with longer sequences, the motivation for using convolutional models over transformers, its model training and architecture, the role of FFT in computational optimizations, and model explainability in long-sequence convolutions. We also talked about Hyena DNA, a genomic foundation model pre-trained on 1 million tokens, designed to capture long-range dependencies in DNA sequences. Finally, Eric introduces Evo, a 7 billion parameter hybrid model integrating attention layers with Hyena DNA's convolutional framework. We cover generating and designing DNA with language models, hallucinations in DNA models, evaluation benchmarks, the trade-offs between state-of-the-art models, zero-shot versus a few-shot performance, and the exciting potential in areas like CRISPR-Cas gene editing.
The complete show notes for this episode can be found at https://twimlai.com/go/690.




