929: Dragon Hatchling: The Missing Link Between Transformers and the Brain, with Adrian Kosowski

7 Oct 2025 · 1 h 14 min · 23 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Adrian Kosowski (Pathway) introduces “Dragon Hatchling” (BDH), a post-transformer architecture meant to bridge transformer attention with biological neuroscience, emphasizing Hebbian learning, attention as synapse-level dynamics, and “positive sparse activation” for more brain-like, efficient reasoning and potentially longer-context lifelong learning.

Guest background

Adrian Kosowski is a neuroscience-and-ML researcher at Pathway, blending biological neuroscience concepts (e.g., Hebbian learning, synapses) with machine learning architectures.

Key claims

  • BDH reconciles transformer attention with Hebbian-style learning by modeling attention as local, synapse-driven selection rather than GPU-friendly context lookup.
  • BDH is “missing link” because it’s more biologically plausible and may explain reasoning mechanisms.
  • Sparse activation: in their ~1B-parameter “hatchling,” about 95% of artificial neurons are silent per input, yet performance matches or exceeds dense 1B baselines (e.g., GPT-2 scale).
  • BDH supports scalable long context without “peanut-memory” compression, via flexible state/context handling.
  • Interpretability: concepts emerge as compact “grandmother synapses”/sets (monosemantic-like), enabling easier feature attribution than dense transformers.
  • Multilingual composability: English and French models can be concatenated (“side by side”) to form a working multilingual model out of the box.

Notable examples

Hebbian “neurons that fire together wire together”; “grandmother cell” analogy; “cocktail party effect” for attention; L2 vs L1 norm analogies using latkes vs hamantashen; tower of macarons illustrating sparse concept mixing.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introducing the Dragon Hatchling

0:45 to 2:58

Discussion about the breakthrough architecture BDH and its significance.

“This episode of Super Data Science is made possible by AWS, Anthropic, Dell, Intel, and Gurobi.”

Understanding Hebbian Learning

2:58 to 4:59

Exploration of Hebbian learning and its implications in neuroscience.

“It's rich in both machine learning and artificial neural network concepts, transformer-related concepts or ways of learning in artificial neural networks.”

The Evolution of Attention in AI

4:59 to 8:04

Discussion on the historical evolution of attention mechanisms in AI.

“And when thinking about the brain or about natural models, at the back of our heads, it's always good to have an idea of how many neurons there are.”

Contrasting Biological and Machine Attention

8:04 to 12:20

Comparison between attention mechanisms in biological systems and AI.

“Let me try to recap back to you the history that you've just given so far.”

Deep Dive into Attention Mechanisms

12:20 to 14:00

In-depth discussion on how attention works in both brains and machines.

“But one thing to look at if we look at the natural kind of systems at the brain is that you can look at attention at two levels.”

Understanding Attention in Machine Learning

14:00 to 16:02

Learn about the historical interpretations of attention mechanisms in transformer models.

“So these interpretations of attention have changed historically.”

The Dragon Hatchling Architecture

16:08 to 17:56

Explore the connection between transformer models and brain-like architecture in AI.

“And so hopefully this now gives a bit of context around the history, neuroscience, machine learning of those two different branches, is kind of where they overlap, where they don't.”

Addressing Limitations of Transformers

17:56 to 21:23

Discuss the potential of the dragon hatchling architecture in overcoming transformers' limitations.

“And the connection to the transformer is such that we rely on an attention mechanism, which is at a very general level of the same.”

Efficient Context Management in AI

21:23 to 23:26

Understand how the baby dragon hatchling can manage context effectively.

“So the idea is that the baby dragon family, starting with the baby dragon hatchling, BDH, we could have a model that has no limits on its context window, it sounds like.”

Sparse Activation vs. Dense Activation

23:26 to 28:00

Learn about the differences between sparse and dense activations in machine learning models.

“It's orders and orders of magnitude less.”
Show all 23 chapters

Introduction to Sparse Positive Activation

28:00 to 29:02

Learn about the new approach of using sparse positive activation in neural networks.

“So these topics have been explored from the point of view of the presence of the brain.”

Understanding Transformer Architectures

29:02 to 31:15

Explore the architecture of transformers and how they scale.

“Yeah, and so just kind of recap those two worlds.”

Vector Spaces and Concept Representation

31:15 to 33:49

Discuss the difference between dense and sparse vector spaces in concept mapping.

“So here the advantage is we can kind of look back at what the Transformers are doing from this perspective and somehow see the way it organizes concepts and maps them into a vector space.”

Differences in Reasoning Patterns

33:49 to 35:58

Examine how reasoning differs between transformer architectures and the BDH model.

“where you compose multiple words to create a new noun, for example.”

L1 vs. L2 Norms Explained

35:58 to 38:00

Delve into the implications of using L1 and L2 norms in machine learning.

“Sometimes one is seen as a view of the other, because it's possible to go mathematically between the two worlds.”

The Cake Analogy for Norms

38:00 to 42:00

Understand the concept of L1 and L2 norms through a culinary metaphor.

“Well, they are at least they set the scene.”

Understanding High Dimensional Concepts Through Macarons

42:00 to 48:00

Learn how high dimensional concepts can be represented through a tower of macarons and the significance of color mixing in this context.

“So this is, and at this point, we would need like high dimensional pastry to be able to.”

Insights on Neural Representation in BDH

48:00 to 53:10

Explore how the BDH architecture provides insights into neural representation and the concept of grandmother synapses.

“And so hopefully we'd have a lot of theory now under our belts.”

Multilingual Models and the Power of Concatenation

53:10 to 56:00

Discover how the BDH architecture allows for seamless concatenation of multilingual models and its advantages over traditional transformers.

“And those synapses that are responsible or that are related to specific contexts activate in those settings.”

Exploring the BDH Architecture

56:00 to 1:02:08

Learn about the Baby Dragon Hatchling architecture and its advantages over traditional Transformers.

“and it performs comparably to GPT-2 despite requiring far less compute.”

Comparison with Mamba Architecture

1:02:08 to 1:06:30

Discover how BDH differs from the Mamba architecture and its potential as a transformer replacement.

“We are providing the architecture publicly, a simplified version of the architecture, which nonetheless performs reasonably well, comparable to the transformer as provided in our paper.”

The Future of BDH and Innovations

1:06:30 to 1:09:46

Hear about future advancements in BDH and its implications for machine learning.

“You and the Pathway team must be delighted to have made this discovery, this invention.”

Book Recommendation

1:09:46 to 1:10:01

Get a book recommendation related to the episode's themes.

“Wow, what a mind-expanding episode with Adrian Kosofsky on the brain, on machine learning, and how Pathway's exciting new BDH architecture could provide the missing link between the two fields.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:What on earth do transformer killer architectures have to do with potato latkes, with hamantuschen, a dessert I had never even heard of before this episode, and a tower of delicious macarons? In today's episode, you'll find out. Welcome to the Super Data Science Podcast. I'm your host, Jon Krohn. We've got Adrian Kosofsky on the show today to tell us about an exciting innovation, potentially a transformer replacing architecture called BDH that him and his team at Pathway have come up with. Adrian is absolutely brilliant at blending biological neuroscience with machine learning, and you will get to benefit from all of that intelligence in today's episode with some helpful takeaways as well.

0:49Jon Krohn:This episode of Super Data Science is made possible by AWS, Anthropic, Dell, Intel, and Gurobi. Adrian, welcome to the Super Data Science Podcast. Delighted to have you here. It's your third time on the show. John, it's a great to be here, and thank you for having me again. Yes, and for the second time, we're recording in person together. That's amazing. Happy to be in New York. Yeah, thank you for flying out, made the trip from San Francisco to New York very quickly. And the reason why, the reason why, so we were recording on a Monday and we decided on Friday that we were going to do this episode.

1:26Jon Krohn:You flew from San Francisco to New York on Sunday. And the reason for doing all this is that you at Pathway have had a huge breakthrough. And yeah, so, so you have this paper that is just released out of embargo to the public. it's called the dragon hatchling the missing link between the transformer and models of the brain and uh so we'll of course we'll have a link to the archive i assume it will be an archive by the time this episode is out definitely and uh yeah there's so one of the first things to talk about is that um you know your bit your paper title is the dragon hatchling but your abbreviation for this breakthrough this new architecture you're calling it bdh even though there's no b in the name so you want to explain this uh explain the name of the breakthrough first before we get into the details that's that's definitely um one thing to explain because we're working on a new architecture in the new family of reasoning models uh this is a family called baby dragon and here we are announcing the hatchling which is a baby dragon which has just appeared so it's no longer an egg It's just hatched, but it has a lot to offer for the audience for a conceptual point of view.

2:51It's already something new and worth looking at. So we decided to go public with it.

2:57Jon Krohn:Nice. I love it. And so I've read your paper. It is fascinating. It's rich in both machine learning and artificial neural network concepts, transformer-related concepts or ways of learning in artificial neural networks. But it also leans heavily on biological neuroscience, which is interesting. John, it's a pleasure to discuss with you, given you have a background both in machine learning and in neuroscience. So happy to have this conversation. Yeah, exactly. So let's let's start with something called Hebbian learning, which is a big it's a it's kind of a big anchor in your concept. And so I actually to make sure that I didn't mess up explaining Hebbian learning at all.

3:40Jon Krohn:I mean, you could do it as well. So it's named after a Canadian psychologist named Donald Hebb. And when I was doing a PhD in neuroscience, like you just mentioned, this idea of Hebbian learning is a big part of the way that we think about the way a brain in a living organism learns. and you might have some kind of other summary statistic on this or summary explanation of this, but the key idea that I take away from Hebbian learning is it's this idea that neurons, so brain cells in your brain that fire together. If something that you're thinking or something that you're seeing or something that you're hearing causes two brain cells to fire at the same time, that increases the probability that they will have a connection formed between them or that that connection will strengthen.

4:33Jon Krohn:And if a brain cell is firing well, another isn't, but they have an existing connection, that existing connection is likely to wane, is likely to kind of fade away. What do you think of that explanation? That's where happy learning works. It is a very quick introduction to what is actually a super fundamental process. One thing to take note of is the different timescales. And when thinking about the brain or about natural models, at the back of our heads, it's always good to have an idea of how many neurons there are. We're talking about like 80, 100 billion neurons and the timescales at which they operate.

5:14So if you think about all the processes, whether you like to take a perspective, which is more on the physical side of dynamical systems, so whether you want to go into the chemistry of what's actually going on, there are different processes which can take place in the timescale of seconds, minutes, hours, days, and those that go on throughout our lifetime. So all of the processes that you mentioned, John, which is about creating links, which is about strengthening links, which is about also links waning, take place at different timescales and may be governed by different dynamical processes.

5:46Jon Krohn:Very interesting. And so in this dragon hatchling or in this baby dragon family of models that you're developing, including in the dragon hatchling that you've now just published on, it seems like Hebbian learning is a big part of the theory that you've used to develop it. Do you want to tell us more about this? Yes, I think we should look back a little bit about the history of the study of intelligence as such before we actually get to Baby Dagon Hatchling. Because in a way, the beginnings of computational science and studies of the brain, like looking back into the 1940s, the days of Alan Turing, it all started together.

6:33And since then, advances in machine learning and in neural science have sometimes gone hand in hand, sometimes drifted a little bit apart. One topic, one keyword which was seen as a hope to reconcile them was around recurrent neural networks. Recurrent neural networks were seen as a hope for, at the same time, the next machine learning architecture and a model which could attempt to explain what's happening in the brain. Eventually what happened, as we all know it today, we are in an era where the transformer is the keyword. We have large language models which are based on the transformer. And this architecture, this approach, the transformer, is harder to reconcile with what goes on in natural biological processes.

7:23I see. So here the approach that we took is a foundational one. We started looking at the concept of attention. Now, attention is a concept which actually comes from neuroscience and came into the world of machine learning from the 1990s until the pattern that we see today in the context of natural language processing. It has undergone a long evolution. And the approach that we took was really to reconcile the understanding of attention in natural systems as captured, for example, by Hebbian learning and the understanding of attention as the keyword behind the transformer. I see.

8:04Jon Krohn:I see. Let me try to recap back to you the history that you've just given so far. So you were saying that some artificial neural network architectures like the recurrent neural network were designed with a biological inspiration in mind. And the idea was that that could actually end up being a good model of how our brains work, of how biological brains work, of how human brains work. and then um we ended up having this architecture with the attention is all you need paper uh you know i'm getting close to a decade ago now actually which is wild that attention is all you need paper uh described an architecture called the transformer and at the time was all people who worked at google that had come up with this idea of this transformer and and it's what you're saying is that the attention mechanism that the transformer has, and that has now become widespread in all of the large language models of the frontier labs that are kind of at the cutting edge, whether open source or proprietary, that attention mechanism is kind of, it's designed with GPUs in mind and efficient compute.

9:18it's not designed to be a model of the way that attention works in biological brains. That's it. And I think what we can discover is actually it's worth studying attention as a concept because as you mentioned, there's like a GPU-friendly implementation of attention, which has been since the attention is all you need paper, maybe even a little bit earlier than that. But the actual question of what attention is, it's a somewhat more fundamental concept because you can look at attention as not just an implementation, it's a mechanism which allows you to efficiently manage your context to think in a contextualized manner.

10:03So based on our introduction, our listeners will be able to also approach the next things that we'll be discussing differently. for example even the word attention by now it's clear that we're discussing it in the context of machine learning and natural systems and not for example some other meaning as in attention, attention, whatever else context is important

10:24Jon Krohn:and so it seems to me like it could be a good opportunity now to give people a bit of an idea of what attention is in a biological organism and it's something that we experience every moment right now hopefully people are listening to this podcast. Hopefully their attention is mostly focused on the sound of my voice and then your voice when it comes on. But we also have, my understanding is that in terms of at least attending to conversations, the human brain on average has the ability to listen to 1.6 conversations at the same time. And so there are things like the cocktail party effect where you're deeply engaged in conversation with, you know, some person or a group of people.

11:10Jon Krohn:And then in another corner of the room, someone says your name. And so there was some, you know, you weren't consciously attending to that other conversation happening, but it was being processed to kind of a lower level of attention than the main conversation you're focused on. But when your name came up, boom, your attention pops over there. And so that kind of, so attention is, I guess, where your kind of where your conscious experience is focused on any given time. And we only have so much bandwidth of attention to provide. And so I could, for example, be using that, say 0.6. If I have, if I'm 100 % focused on the conversation that you and I are having right now, Adrian, I could still have that extra kind of like 60%, maybe kind of running on, what am I going to say next?

12:02Jon Krohn:Kind of attending to internal conversations. And so, yeah, so this attention can be external, it can be internal, and it doesn't necessarily, in fact, it's hard to kind of imagine how it could have anything to do with the transformer architecture. Indeed, it's a path which needs bridging. We've made some steps in this direction. But one thing to look at if we look at the natural kind of systems at the brain is that you can look at attention at two levels. You can look at the level of the entire system. Here we speak of what we consciously feel or what our brain as the whole is actually attending to or listening to or focused on.

12:40And we can also look at it at the micro level for each neuron specifically. specifically what is this neuron focusing on and the the spoiler here the hint is that actually what neurons care the most about is their connections to their neighbors which means that this is basically what they what they are who they connect to is in in some sense what they are and if they want to pay attention to some of their neighbors to what some of their neighbors are saying more than two others, this is the attention mechanism, which is encoded through the neuron connections. These are called synapses, which change in time, which can open up communication.

13:23They potentiate and allow communication between neurons more directly at super short timescales. And as we speak, the neurons in the parts of the brain which are responsible for processing conversation, for understanding, for reasoning around it, are all the time acting on these synapses. So here we have this picture. Imagine it as a massively dynamical system, a bit like a social network of neurons which just decide actively which of their friends, which of their neighbors they are listening to and which they are not listening to. So this is natural attention. And indeed, at the other extreme, we have attention as it is understood in machine learning.

14:01So these interpretations of attention have changed historically. The one that perhaps listeners will be most familiar with, is attention as it is explained in the context of the transformer implementation on GPU. So you have a certain number of vectors of attention in each layer of the model which is running. And then you search, given the current, we say token, the current element of conversation, we search for elements in the past which somehow relate to it and then at a very intuitive level at at the basic level of of an architecture like the transformer we are listening we are connecting to words sounds that have been heard in the past and in higher levels of the architecture it's a bit like we are connecting to something that's not as well defined some higher higher order concepts which appear.

15:03So this is basically the understanding of the tension, which is very much about context lookups and searching. So it's like a search data structure, which is very different from at least in description, from the local mechanism that we know appears in the brain.

15:30Jon Krohn:instances deliver 20.8 petaflops of compute, while the new Tranium 2 Ultra servers combine 64 chips to achieve over 83 petaflops in a single node, purpose-built for today's largest AI models. These instances offer 30 to 40 % better price performance relative to GPU alternatives. That's why companies across the spectrum, from giants like Anthropic and Databricks to cutting-edge startups like Poolside, are choosing Tranium 2 to power their next generation of AI workloads. Learn how AWS Trinium 2 can transform your AI workloads through the links in our show notes. All right, now back to the show.

16:11Jon Krohn:Really well explained there, Adrian. That was fantastic. Very easy for me to follow. And so hopefully this now gives a bit of context around the history, neuroscience, machine learning of those two different branches, is kind of where they overlap, where they don't. And so your paper now, the dragon hatchling paper that's out, it is described as the missing link between the transformer models of the brain. So how does this paper blend together what we've been talking about so far and form a missing link? We introduce a post-transformer architecture. It's an architecture which relies on attention and which fundamentally has properties of a massively parallel system of neurons.

17:01So you can look at it as a system, a network of a large number of neurons which also communicate with each other. They are artificial neurons, but they do have some of the properties of natural neurons, especially in what pertains to how they address the problem of attention. And in some sense, this architecture is the missing link in the sense that it is more biologically plausible. It is closer to the brain. And it also hopefully this is something to be verified in experiments, but hopefully explains some of the mechanisms of the functioning of reasoning in the brain, or at the very least, provides a plausible explanation of how the brain could use certain mechanisms to achieve performance known from machine learning models like the transformer.

17:53So this is the connection to the brain. And the connection to the transformer is such that we rely on an attention mechanism, which is at a very general level of the same. It is implemented in a different way. There are a couple of elements which we may dive into, which are technical. It is what is known as a state-space model. State-space models are a form of reconciliation of some of the concepts of recurrent neural networks and the transformer. the fact that that it is a state space model means that attention does not have to be viewed like a lookup structure you can also view it in this local perspective of actually paying attention to certain certain concepts you don't have to look you don't have to look at it as looking back in

18:47Jon Krohn:time so you have the state space interpretation excellent all right so with this kind of state space interpretation and this dragon hatchling architecture that you have that could potentially replace a transformer it sounds like it's exciting not only because it could be a more efficient replacement for the transformer and could be something that some years from now is kind of the go-to thing when we build a large language model instead of going to the transformer as your choice, you go to some, some animal from the baby dragon family. And so that seems like one part of it that's really exciting.

19:29Jon Krohn:And then the other part you mentioned there is that it could end up being that this could help us understand the way that learning happens in biological systems better. And so it could potentially not only be a machine learning breakthrough, but a neuroscience one as well. That's, that's exactly the hope. And that's why we're approaching it. There's one more aspect, which is perhaps worth mentioning, which is that it's kind of pointless replacing the transformer in places where the transformer performs well. But there are places where the transformer does have its limitations and the human brain is able to overcome them.

20:04These pertain to lifelong learning, to reasoning over long periods of time, learning with experience. So these are areas where a human mind can spend years potentially diving into a subject to perfect the state of the art and to push the state of the art forward. This does not have to be scientific state of the art. It can concern any kind of work activity, any specialized task in which the human mind is able to do it, whereas the transformer has its limitations. There's a lot of ongoing discussion about specifically reasoning models, synthetic reasoning models today, whether they are able to extend reasoning beyond patterns that they have seen in their training data, whether they are able to generalize reasoning to more complex reasoning patterns and longer reasoning patterns.

21:01The evidence is largely inconclusive with a general no as the answer currently machines don't generalize reasoning as humans do. And this is the big challenge where we believe architectures that we're proposing may make a real difference.

21:23Jon Krohn:Fantastic. So the idea is that the baby dragon family, starting with the baby dragon hatchling, BDH, we could have a model that has no limits on its context window, it sounds like. So you could theoretically have as much learning as you want and be able to attend efficiently over all of that information. That's true. That's true. And the answer is yes. One thing to always bear in mind is that there's no such thing as a free lunch in the sense that, especially for those of the audience who are familiar with post-transformer architectures which have appeared in the past, there's often this attempt to somehow make context infinite but at the same time compress it.

22:18Just have a kind of a memory the size of a peanut and hope for it to be able to somehow compress long, long reasoning. chains of thought or long sequences of text. In our case, the way the architecture is designed, and obviously we invite everyone to have a look at a paper and take a look into the details, but the idea is really that there's so much space, so much flexibility in manipulating storing context that this is not a bottleneck and at the same time it's efficient um so here for uh for the sake of analogy if you take again the human brain one parameter that i mentioned at the beginning is the scale of it it's like it's 80 100 billion new neurons but also if you look at the way the brain represents its state here we are talking about synapses they're like 100 trillion of them not clear exactly how many but it's 100 trillion and i can assure you that none of us process 100 trillion tokens in a lifetime.

23:29It's orders and orders of magnitude less. So we're talking about a system which does have the storage space. It does have the state in sufficient quantity to be able to process long context, but needs to do so efficiently so as not to waste time on doing some operations which don't move it forward.

23:49Jon Krohn:Yeah, and so we'll talk about that in a moment later in the episode. what you're starting to allude to there is this idea that something that's a characteristic of transformer-based architectures is that they tend to be densely activated. So you're activating either all the neurons, or more recently it's become in vogue to kind of have modules of neurons, but then you still have dense activation within those modules. And when I say dense, I mean that you're basically flowing information through all of the neurons in the network or all the neurons in that module. And this is computationally expensive, energy expensive.

24:29Jon Krohn:If the human brain were to do that with the trillion connections that we have in it, we wouldn't have enough energy to support that. And so we have a much sparser activation system in our brain, in the human brain, which is that, yeah, so only, you know, a relatively small portion of neurons are being activated at any one time. And there's all kinds of things. So then you can do things like functional magnetic resonance imaging, fMRI studies, or EEG, electroencephalogram studies, to be able to see what parts of the brain are active during particular types of thought. And that you can do that is demonstrative of the fact that we have a sparsely activated brain.

25:10Jon Krohn:And so a big thing that you've done with BDH is that this now allows for sparse activation. So you can correct me if I'm wrong on the stat, but it looks like 95 % of the artificial neurons in BDH are silent at any given time. They're not firing. And yet the model is able to rival. So right now you've only been experimenting with it being a little baby dragon hatchling. So it's a relatively small network. It's about a billion parameters, which is about comparable to GPT-2. GPT-2 is densely activated. So all of the neurons are being activated at any given time to any given prompt, whereas in yours, only about 5 % are active at that time, much more like a brain.

26:02Jon Krohn:How, yeah, maybe explain a bit more about that and how central that efficiency principle is to BDH, as well as, you know, the implications for biological brains. Absolutely. So, John, indeed, the idea of working with models at a billion parameter scale here is largely for the sake of demonstration of certain tasks that we can work with. And here we were looking at the core tasks that can be considered advanced from the point of view of cognition related to language, related to translation of tasks that need attention. We are demonstrating it at this scale. Indeed, the aspect of sparse activation is a fascinating one.

26:50It's an aspect which is quite deep because there's a bit of a chicken and egg story to unravel here. In fact, there are two worlds which work well, which can be compared to each other. However, there's a gap between them. So there's no middle ground. You can be in one world or the other world. One world is the world of dense activations, which is the world the Transformer is in. The other world is the world of sparse positive activations. This is the world that Baby Dagon Hatchling is in. Incidentally, the use of sparse positive activations in biologically inspired models and their discovery in brain function is a topic which has been ongoing at least since the 90s with brilliant work regarding visual function, olfactory function, basically sense of smell even for the food fly, which relies on sparse positive activations.

28:02activation. So these topics have been explored from the point of view of the presence of the brain. But the use of sparse positive activation for reasoning is something that's new. We believe we are the first work that has actually scaled it to transform a like or beyond transform a performance. And the fact that we are at one billion means it scales onwards, meaning all the key points have been achieved and we have achieved the scale. The implementation is GPU efficient, by the way, so just to reassure everyone there's no, especially for inference, it has a lot of points which allow us to outperform the transform on certain hardware configurations quite significantly.

28:48So this is actually super encouraging. But maybe one more thing to say while introducing the topic, and maybe we can dive in just a little bit on the technical side what these two worlds mean. But first, one main difference between the two worlds, which I'd like to kind of highlight.

29:05Jon Krohn:Yeah, and so just kind of recap those two worlds. So on the one hand, you have the densely activated world of the transformer, where, as I was talking about earlier, when any prompt goes through in a fully dense network like GPT-2, all of the neurons, we run compute through every single neuron, every single connection, on every input that goes into the model. Whereas the other world is the world that the baby dragon hatchling, BDH, lives in, where some small percentage are activated. I would say to analyze the transformer, you can actually dive deeper. Because for the transformer, essentially, if you look at the architecture, which is behind the transformer, and for most open source models, The basic part is the GPT2 architecture.

Read the full transcript

29:56Nothing much has changed. If you look at all the family of LAMA models, LAMA 3, LAMA 4, all of them, this is like GPT2 almost unchanged. Yeah, just scaled up. Just scaled up. So here, the scaling of the transformer is actually a bit of an interesting story because there's no one single recipe for scaling the transformer. It ends up being done differently. You have a different number of attention heads. You have a different number of layers, et cetera, et cetera, et cetera. So the transformer scales in different ways and it's a little bit of know-how how to scale it. Or actually it's from the point of view of a theoretician or somebody who would want to apply something like computational complexity and ask what's the limit of a transformer.

30:38It's actually rather hard to decide how we should scale it up. But there's one dimension which appears to be fixed, to have converged, and that's the size of the attention head vector dimension in a transformer. So this does not scale even as the models get larger. It has stopped scaling. And here the intuition is that basically all concepts that the transformer works with have to be mapped into a vector space of about 1 ,000 dimensions. I see.

31:20Jon Krohn:I see. So we have this constraint that's appeared in the transformer that will limit its ability to have nuanced reasoning because it doesn't seem like we can scale the dimensionality of the vector space of the tension heads beyond a thousand. So this is actually a point that we address a bit more mathematically in our paper because we take the opportunity to analyze both, you know, baby dag and hatchling and also kind of take the reverse approach to history and look at transformers and approximation of baby dag and hatchling because it's actually easier to take this direction to say that we're coming up with a simpler and somewhat cleaner mathematically architecture cleaner in terms of the number of moving blocks.

32:06So here the advantage is we can kind of look back at what the Transformers are doing from this perspective and somehow see the way it organizes concepts and maps them into a vector space. The kind of question which comes up, and this is, I think, for a large part of the audience who are familiar with vectors, with vector spaces, with notions even quite far from the transformers, such as, let's say, preference vectors, you can manipulate them as if it was a linear space. So you can add vectors together, you can subtract vectors, you can have the opposite elements, you can have a negative vector.

32:52And this is the essence of a vector space. By contrast, if you go into the sparse positive spaces, the way you actually compose concepts is somewhat different. You don't work so much with linear combinations of vectors. You work more with bags of concepts, so bags of words, combinations. It works a bit like a tag cloud. It looks a bit like an association set, so a number of elements put together which form a whole. It's a bit how you form sentences by putting together words to get the meaning correctly or in some languages which have a bit more of a tendency to play, especially Germanic languages, German in particular, where you compose multiple words to create a new noun, for example.

33:53This is a place where you have this kind of compositional effect. And there are a number of differences. One difference to this type of representation, apart from the points that you discussed, John, that efficiency and so on, is actually that you don't look so much towards negatives or opposites. Starting with a simple example, if you show somebody a piece of work that's been done badly, and you tell them, look, this is what you're supposed to do, but do the opposite to get a good effect, there's no such thing as take the opposite and find you know the the opposite to it likewise we don't like in reasoning patterns there's no symmetry between being attracted towards a certain reasoning pattern and being repelled from it so if i tell you now don't don't think about the color blue don't think about the color blue you'll be consciously thinking about the color blue just to just to try to compensate for it and avoid it.

34:59But it's not the mechanism of switching off, of damping down. So we are in a different vector space, in a different space.

35:09Jon Krohn:Right. And so what you're saying there is that, so transformer architecture, it doesn't behave in the same way as a brain with these kinds of negative activations, but your BDH does, your new architecture. is that so there's a there's a term that you mentioned earlier uh and we kind of glossed over it i wouldn't mind trying to dig into this a bit more you described the bdh architecture as positive sparse and so now and now we've been talking about negatives can you explain a bit more about this positivity thing the question is actually a profound one um because the uh the correspondence between the two worlds, the world of dense vector spaces versus the world of sparse positivity.

35:54It's like two worlds, which are sometimes complementary. Sometimes one is seen as a view of the other, because it's possible to go mathematically between the two worlds. For the audience who is familiar with, again, with like hands-on machine learning, the vector world is the world where the L2 norm rules. That's the kind of king or queen of norms, the L2. If you go into the sparse positive world, you suddenly go into the world of probabilities. The world of probabilities, the world of chance, because the concepts that you're working with start to have interpretations of likelihoods, or at least some value between 0 and 1, which reflects how much you are drawn to a given concept.

36:50So in L2, you don't have it. In L1, in the probabilities, you have like an L1 norm interpretation.

36:57Jon Krohn:Okay, and then just to quickly say this, so in either of these cases, whether we're talking about L2 norm or L1 norm, these are ideas that have been around for decades in machine learning. And in either case, they're an additional factor that we add into our models that allows our models to generalize better to data that they haven't seen before. This is the idea that you introduce a concept that hasn't been seen before. And the question is, where do you place it compared to other concepts? So we can walk through this. I will be using as props a number of pastries. The first two... Apologies to our audio-only listeners.

37:39Jon Krohn:This segment, Adrian brought out a set of delicious treats, which smell fantastic, and it's very hard not to eat his props. And we have three different plates or serving dishes of desserts or of food. And they're critical to this explanation, I suppose. Well, they are at least they set the scene. And I will explain as best I can to our listeners. So the debate, what's better, like L1 or L2 from the point of view of working with basically any kind of dynamical system representation. This is a debate which is prevalent in many fields. And so what would better look like? Like if one is better than the other, what is the outcome?

38:36So, there are differences, and I'll explain the differences. It's hard to say which one is better, and that's why I'm bringing in cakes, because everybody has their own preference for cakes. And here, it's kind of, you know, there will be applications, there will be situations in which every cake is suitable. The first two cakes in the selection are actually not due to me. They're due to one of the more renowned figures in quantum information, quantum computing, Scott Aronson, who had this idea that the L2 norm world compares to a round potato cake like Latke. And the L1 norm compares to this triangular cake, the Ham and Daschen.

39:26and the explanation a little bit is to understand you know how you can of course it's it's it's trivial taking a high dimensional space and reducing it into into what we can see what we can feel but you can try to do it so in a in a vector space you'll have you'll have this kind of round thing it's like like a ball in which you have vectors and you you move around inside this kind of circle, you have positive, you have negative. In fact, if you move to the L1 spaces, then at this point you have corners. You start to have sharp corners. I see. And these are the kind of concepts that you are connecting.

40:07So the kind of well-defined entities are in the corners of a triangle, for example, if you take the lowest example possible, the smallest example possible. and the combinations of these are in the center. So you have this effect of combining corners to mix them in. I see.

40:30Jon Krohn:So the potato latkes are representative of the L2 norm. Yes. And so it's not a coincidence then that potato latke is relatively homogeneous, that it's all kind of the same kind of substance throughout. whereas with this other kind of treat which I must say is new to me latke is very familiar with I think they're delicious these are ham and tushen? yes nice and then that's interesting so the tushen part at least that's pocket in German yeah so I think both of them obviously come from the diaspora cuisine but you can get both in central Europe and then of course in New York and New York obviously I mean it's like we can get everything in New York but anyway I'm bringing up these two because as long as you stay in discussion between these two, it's a bit of a debate, as some of you may have seen, like debates between the two kinds of cakes, academic debates at most.

41:26Academic debates. Academic debates. So you have to reduce your problem to the debate between the two. But the kind of thinking here is that, you know, if you are in this L1 norm world. A commentation world with the pointy corners. Yeah, that's kind of like, okay, it's fine, but it only starts to get interesting when you go to higher dimension because sparsity needs higher dimension because you want to be choosing a few concepts from a very, very large group of concepts. So this is, and at this point, we would need like high dimensional pastry to be able to. to have to send this.

42:11Jon Krohn:They need a baker that can bake in more than three dimensions. That's it. That's it. So technically, yeah, technically here we're talking about three concepts, right? And we have a triangle. The three corners on our triangle. The best I could do was go one dimension higher and here we are escaping Central European cuisine to the best of French patisserie. Still available in New York. Of course. Of course. So here we have a structure which is a triangle but one dimension higher at least okay so yeah so in this case here for people who aren't watching the video version we have a stack it's kind of like you know people order those seafood platters with like lots of layers of seafood where the the bottom layer is the biggest and maybe there's like some lobster or some crab on there and then you have a medium tier that's a little bit smaller and you got the shrimps on there and then a top tier with scallops that's the smallest or something.

43:08It is a kind of representation of a pyramid. At least that was the objective to have like a pyramid. And the interesting thing is that when you look at this kind of structure, first of all, the first thing you notice is that you don't eat the center. In the center, you just have a plate. It's the things that are kind of outside on the walls that are interesting. So you have combinations of smaller numbers of corners, which give you the interesting elements. In fact, although we went from what is essentially a structure in two dimensions to a structure in three dimensions, we are not combining four concepts, but we're still combining usually three concepts to get the desired outcomes.

43:54Also, one thing about the specific type of construction...

44:00Jon Krohn:I don't think we even mentioned here is that this is... So Adrian didn't bring a seafood tower into the studio, which might not have been the most considerate thing to do. That would have been pretty interesting. But it's a tower of macarons. It is a tower of macarons. And these macarons are beautiful round objects. So you can think of each macaroon as having a certain radius, because indeed it does have a certain radius. And you can have concepts which somehow have a certain radius of its action. Right. Again, for those of you who are more into into machine learning, this has a vibe of a nearest neighbor's kind of detection field.

44:43Right.

44:44Jon Krohn:And that sounds similar to the idea earlier of neighboring brain cells being the thing that are of most interest to a given local brain cell. Yeah, so it is the kind of representation that we would be looking at. So again, this is kind of a little bit trivialized, but the other thing that we are trying to represent is how you can have a combination of different concepts and how it gives rise to kind of a new concept and how you can represent it in this past representation. And perhaps the most visual way we can think of it is through colors, because it's not like these concepts are positive or negative.

45:23I would want to leave the vector space world in which we talk about location and length and so on, but more about mixing colors. If you mix yellow with red with blue, you'll probably get something very brownish. But in some cases, if you mix too many colors, it's a bad thing. And if you mix just the right colors and just the right amounts, you get exactly what you want to get. So the whole kind of effect, the visual impact comes from the fact that you are just not mixing the whole palette together, but you are actually picking a small number of colors to mix in a palette, to place in one place in the painting.

46:01And that way, these sparse mixtures of colors allow you to represent the concepts that you want to represent. And only then do you kind of place them in the space.

46:15Jon Krohn:And so is that why there's some specific color ordering to your tower of Macaron, to your four-tier? Is that providing us with a fourth dimension of information? So I would say that the two, ideally, in an ideal world with, let's say, perfectly arranged towers, there would be a certain duality between the color and the location, meaning every color would have its location on the pyramid. So you could use one information or the other. but somehow the thing I want to hint at is that the location is less important than the actual, you know, the colour that this carries. I expect you could make the same analogy with taste and mixing taste.

47:02Again, if you're a good chef, you don't want to mix all the different possible tastes but you either want to have a sparse combination of tastes and you won't have a macaron with everything in it from chocolate to blueberries but you will be trying to isolate individual one or two tastes and put them together. Nice.

47:18Jon Krohn:Would have messed with your tower if I took a bite of a macaron. I could take one from the back so it doesn't change anything for the audience. Nice. That's one of the things, and this is, you know, how we go towards sparsity. You don't have to represent all of them, but I can still tell you the one that you've taken should have had that color. So that's... Those are very good macaron. Thank you, Adrian, for bringing those into the studio. So I'll dig into the latkes and the hamantushin later. We'll be demonstrating sparsity. Nice. Fantastic. Okay, so I think we've hit on a lot of the key ideas around your new architecture, the BDH.

48:00Jon Krohn:And so hopefully we'd have a lot of theory now under our belts. So then maybe some of the questions I have next can maybe, we'll see if they can be a bit more rapid fire. um in the biological brain there's been an idea there was an idea that was in vogue a few decades ago it's fallen out of vogue a little bit now but it was the idea that you would have like a grandmother brain cell that would fire when you saw your particular grandmother or heard her voice that this you know particular neuron would fire and in subsequent years it's become clear that But for something that complex, you don't just have one brain cell amongst the 90 billion in your brain that fires, but you have a set of brain cells.

48:46Jon Krohn:And those set of brain cells firing together gives you this representation of your grandmother. And so this seems similar, this kind of having a grandmother cell seems similar to something that you talk a lot about in your BDH paper, which is your discovery that there are specific neurons that fire for the idea of a currency or for the idea of a country. And the implication there is that it might be much easier with BDH to interpret what the neural network is processing than with a dense activation like a transformer-based architecture. Definitely. And definitely this is two. The way to look at it, first of all, again comes back to positive activations.

49:37One of the things to note about positive activations is that combinations are easy to express. You don't have to work so much with doing things like projections or finding negative and positive coefficients to say that I want this part of my network to activate in a positive way and that one in a negative way and balance it all together. You have full interpretability because you just say I want this set, I want this combination to fire and then this represents my grandmother. Interestingly enough, if you dive into optimization of this type of architectures and find the natural patterns that they optimize for, there's a pattern that the more important concepts are represented by generally smaller sets.

50:34So for those in the audience who are familiar with concepts of network science, of systems, this general search for power laws in systems, power law distributions in which the more important concepts are somehow more compactly represented. And you can actually, for the more important concepts, even in a relatively small network, even at like 100 million scale or even below that, find in our network an individual synapse which is sufficient evidence for a concept being mentioned. So this touches on a notion known as monosemanticity or being responsible for one concept and figuring on one concept.

51:18And this we see. Perhaps one more thing, John, to say is when you mentioned the notion of grandmother cells and how they appear, they appear spontaneously. And in architecture, it's not architected. There's no genetic code that says, hey, I want to have cells responsible for different operations. They evolve, they emerge in the course of training. We have no control over where they will be, but we find that they emerge and that you have this very clear location of signals passing through the artificial brain as a function of what we are talking about. Right.

52:02Jon Krohn:So kind of like your multicolored tower of Macaron, you don't have control over where kind of the orange and yellow Macaron end up in the tower. But you can bet that they will kind of end up aggregating together somewhere in the structure. You just don't know where. Yeah, it's pretty fascinating. And as we continue with this tower and it becomes sparser, you also end up with patterns in which the most important kind of locations are filled. So those more important concepts are filled and the less important ones are not that much filled. Maybe one thing to say for the audience who is kind of interested in technical details, but one fascinating technical detail here is that in our model, which we've been able to find and would actually shed some light on how these things work, is the concept of a grandmother synapse rather than the grandmother neuron, which means that if you think of how the state of the system works, the state of the system, the context that we are listening to, is represented by synapse activation potentiation.

53:12And those synapses that are responsible or that are related to specific contexts activate in those settings. So we have specific synapses which react to specific notions.

53:26Jon Krohn:Yeah, and just for our listeners who don't have a neuroscience background, a synapse is where two neurons meet and chemicals pass between them. And that is where learning must happen. Chemicals cross over this tiny little gap, the synapse, between two brain cells. And it's kind of analogous to the idea of the parameter in an artificial neural network model. Yeah, this is how we see it. And like being able to align the two is actually. Nice. All right. So next question for you related to this, something that I found fascinating about your paper, about your BDH paper, is you were able to concatenate, literally just like a concatenate operation.

54:16Jon Krohn:You could have one neural network trained on one language, let's say English, and you could have another language trained on, let's say French in honor of the macarons here. and with your architecture, and this seems like a rare thing to be able to do with an architecture that could be the building block of a large language model. You can just concatenate those English and French language models together. And because of the sparse activation, it just works and it's a multilingual model. That's the spirit. And I think this touches on so many different aspects, which I think are good to highlight because it's something new.

54:57It's new in many senses. As I mentioned before, the transformer, while obviously being an amazing breakthrough in the focus of machine learning and AI in general, does have its limitations in the way we understand its scaling. So if you have like two transformers and you put them side by side, there's no really clear way how to connect them. In BDH, this is much easier in the sense that the model scales in one dimension. We call it the number of neurons, N, and it's like the size of a bane. And then if you want to put two such banes together, you can do it. Depending on what you do, it will be a little bit like a mix of the skills that you had.

55:43or you can also do some post-training for the combined Bain and make sure it coordinates properly. But definitely, if you just put the Bains side by side, you have a model which out of the box has understanding for the different languages or is able to map them into concepts in English, for example, and to work with them.

56:02Jon Krohn:That is very cool. All right, so with all of these incredible novel capabilities of BDH relative to Transformers, So the positive sparse activation that we've talked about, this ability to concatenate that comes out of that, the energy efficiency that comes out of it and compute efficiency that comes out of it. Where are you today? It kind of sounds like you've, you know, with this paper, with BDH, with the Baby Dragon Hatchling paper, we're talking about a billion parameter model, which is about the size of GPT-2 from OpenAI, which is now some years old. and it performs comparably to GPT-2 despite requiring far less compute.

56:46Just to reassure the listeners to this point, we are looking at models which at a given scale are on par with models of a given scale. So really it's given all the progress that has happened in the state of the art, we use that progress obviously. So the 1B models that we produce are comparable or outperform the 1B models out there. The reason why we focus on this 1B scale for demonstrations is that this is a scale at which we are able to achieve instruction following and to start testing other capabilities of the model which is able to actually follow instructions and to have the basic capabilities that we would expect of a language model.

57:43And this is really for the ease and speed of experimentation. There's nothing particularly stopping us from releasing a super large model like in the 70-80 billion scale larger. The kind of question which is super pertinent is why do it? because if you're in the world of language models, just language models, there's a certain market, which we could call a bit of a commodity market for the kind of chatbot-like applications, discussions, and so on.

58:18Jon Krohn:Right, so your Claude, your Gemini, your ChatGPT, they're all kind of competing in the same space. I think the switch that most of us are kind of most aware of is if you are working with a reasoning model or not. Usually you are kind of explicitly aware of the switch, especially with models like GPT versus 0103 with Claude, etc. You have this option to go into reasoning mode. And this is the place where we don't want to just yet launch a non-reasoning model, which is super large, because there's actually not our objective here. What we are doing is we are entering reasoning models. We are entering it from the moderate scale, obviously.

59:09But this is a scale where we can display the advantage of this architecture. I see. I see. Notably, yeah.

59:17Jon Krohn:So yeah, so the most promising avenue for you for moving forward with this baby dragon family is into reasoning models. So models where you don't just have tokens output being spit out to your screen immediately, but there's multiple phases of reasoning happening in the background, refining your answer, ensuring accuracy. Yeah, that's where you see the most potential. That's it. Lots of consideration, lots of introspection, and also something that we see as extremely pertinent is the ability of reasoning models to work with contextualized inputs and to process them. so if you think of breaking the barriers the limits of like 1 million token context but you have a reasoning model which goes through billions of tokens of context here you're in a space in which you can for example ingest a contextualized data set private to enterprise like a documentation of an entire technology which is like 1 million pages of paper, 1 million sheets of paper that's 1 billion tokens you ingest it in a matter of minutes given enough hardware in this architecture.

1:00:31And with that in hand, you can start actually making sense of large data sets in the way you would expect of reasoning models. Again, maybe for the developer audience out there, I'm sure you're familiar with the use case of AI-assisted coding in general. and this is perhaps for currently the frontier use case. We are looking at the next generation of use cases like this, but to focus on this use case for a moment, the complexity of having an AI code assistant increases with the amount of pre-existing code with the size of the code base. And usually it's much easier to have a model which contributes a piece of new code to just invent things without actually having internalized everything that was created before its action.

1:01:30So it's basically doing a project on the side of its own, then to have basically a model which is able to control and contextually operate in an environment which requires understanding of a large code base. And again, code bases are perhaps the frontier example, but they're still the easiest kind of example that we're looking towards.

1:01:57Jon Krohn:Exciting. So for people who want to get their hands on this architecture or on models based on the BDH architecture today, can they do that? Are you providing these publicly? We are providing the architecture publicly, a simplified version of the architecture, which nonetheless performs reasonably well, comparable to the transformer as provided in our paper. We have our internal enhancements, of course, and especially those that allow this architecture to work faster, the transformer, especially in the inference generation and so on. These we keep internal, so you can lay your hands on the architecture.

1:02:33You can play with it. I believe it's an excellent playground for anybody interested in understanding models, so all topics of introspection, of observability, of getting a feel of how the model works. and what we will be releasing in the near future. So kind of follow on announcements, just hinting at them, will be related to inference efficiency and reasoning on especially enterprise datasets. Excellent.

1:03:08Jon Krohn:And so fantastic for you to provide this for our audience so that they can implement the baby dragon hatchling themselves. so there's there's been hubbub before about a kind of architecture that could potentially replace the transformer as the go-to building block within a large language model so for example mamba was something that was pretty big a couple years ago and you know there was hype around you know this being superior to the transformer i can't even now remember why do why was the mamba considered superior? Was it compute efficiency? It was compute efficiency in the case of long context and bypassing limitations on long context.

1:03:50Jon Krohn:Right. But yet today, I haven't heard somebody talk about Mamba in over a year. As far as I know, there's not a single frontier model that uses Mamba instead of a transformer. So what makes BDH different in your view? How could BDH actually be a transformer killer? So there are two aspects at least which I'd like to say in which we are more fundamentally different than the transformer. The first is something that I hinted at before, that in order to get a non-transformer architecture to work, you have to make a number of changes to make it work. You can't just say, I'm changing one thing. So the change that's most often talked about is taking attention, as we discussed attention before.

1:04:38Attention is an operation which viewed in vector spaces allows finding approximate closest neighbors due to the due to the soft max operation used in the transformer at least as it is implemented in the transformer and one attempt was basically to take the soft max out of it that's that's an operation which is called linear attention and linear attention brings us into the world of state space models so this is like one change but there are in fact five changes that you have to put together to make it work not one but five so being in sparse positive large dimension and all this has to be put together to make these models work and to find ourselves out of the ridge like in crossing the valley so to speak between two to from one place of local optimum which is the transformer to a place further out and there things start to happen so this is kind of one thing a second point is also that you can be completely radical with use of architecture there's a lot of talk of you know of supplementing the transformer with one or two layers which are different which which which help a little bit so like the hybrid transformer and mamba for example a hybrid transformer and something here there's no need for this you can call all baby go all the way baby dragon you can have like this entire architecture all consistent and what this means is that you don't complicate your picture compared to the transformer, you simplify it.

1:06:10When you simplify it, you simplify the way it works with hardware. Whether it is GPU, whether it is a GPU alternative, so one of the AI accelerators out there that do tensor-based processing, it's DPU or Lithuanian or others, the layout of memory is radically simplified with BDH. And you find that some of the bottlenecks in memory transfers of the transformer disappear because of this architecture so in some sense even if um if you take an improvement of the transformer which doesn't have a bottleneck but you still keep stick to the transformer you will have the transformer based bottlenecks with bdh you can start refresh and start sharding start arranging things differently for more optimal performance

1:07:00Jon Krohn:Well, really exciting times, Adrian. You and the Pathway team must be delighted to have made this discovery, this invention. It's a first. We're super excited about it. More news coming up soon. Definitely, definitely it's a good place to be. Congratulations. It's a big deal. It's incredible for me to think about all of the disparate knowledge that you and your team at Pathway concentrated together to be able to come up with this breakthrough. It's inspiring. And yeah, I look forward to seeing BDH in Frontier LLMs all over the world a few years from now. Exciting times indeed. Before I let you go, you'll surely remember this because you've been on the show before.

1:07:46Jon Krohn:I do always ask my guests for a book recommendation. I have no choice this time, John. I mean, I could think of different books, but since we're talking about Dagens, I have to tell you the backstory, like why dragons? I see. So these dragons are due to Terry Pratchett. Terry Pratchett, gotcha. Yeah, so it's the Discworld series. And here it'd be kind of hard best to pick one specific book. I think you can start from the beginning. So it's Color of Magic Light Fantastic. The Color of Magic Light Fantastic. The first two books, The Color of Magic and the second part. Very nice. All right, thank you for that recommendation.

1:08:25Jon Krohn:I'm sure we have some Terry Pratchett lovers out there listening already. And yeah, and then finally, where should people be following you or Pathway to get the latest on your breakthroughs or Pathway's breakthroughs, your latest thinking after today's episode? We'll be sharing on our website a series of updates. Please don't hesitate to subscribe to that. That's pathway.com. And of course, follow us on social media. Fantastic. Yeah, we'll have the links in the show notes for you. Yeah, what an exciting time. Crazy how quickly these innovations come, and it must be a real trip to be experiencing those innovations come up within your consciousness, within your biological neural network, and your attention mechanisms.

1:09:14It is exhilarating, and the interesting part is that you actually get to understand yourself better as you work on this type of topics. So that's kind of at a meta level, it's kind of all in-folding as well.

1:09:31Jon Krohn:Love it. Well, I hope to get the opportunity to check in with you and Pathway again on the podcast in the near future and see how these brilliant innovations are coming along. Thanks for joining me here in New York today, Adrian. Pleasure anytime, John. Thank you. Wow, what a mind-expanding episode with Adrian Kosofsky on the brain, on machine learning, and how Pathway's exciting new BDH architecture could provide the missing link between the two fields. As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Adrian's social media profiles, as well as my own at superdatascience.com slash 929.

1:10:11Jon Krohn:Thanks, of course, to everyone on the Super Data Science Podcast team, podcast manager Sonja Breivich, media editor Mario Pombo, partnerships manager Natalie Zyski, researcher Serge Massis, writer Dr. Zara Karchet, and our founder Kirill Aramanko. Thanks to all of them for producing another outstanding episode for us today for enabling that outstanding team to create this free podcast for you. We are deeply grateful to our sponsors. You can support the show by checking out our sponsors' links, which are in the show notes. And if you yourself are interested in sponsoring an episode, you can get the details on how to do that by making your way to johncrone.com slash podcast otherwise help us out by sharing this episode with folks that would love to hear about it review it on your favorite podcasting app or on youtube subscribe if you're not a subscriber but most importantly just keep on tuning in i'm so grateful to have you listening and hope i can continue to make episodes you love for years and years to come until next time keep on rocking it out there and i'm looking forward to enjoying another round of the super data science podcast with you very soon Thank you.

From the publisher

Breaking news: Jon Krohn welcomes Adrian Kosowski to the show to talk about the groundbreaking research happening at Pathway. Adrian and his team demonstrate how they have brought attention in AI closer to the way the brain functions, creating, in essence, a “massively parallel system of [artificial] neurons” that communicate with one another and exhibit properties similar to natural neurons. The goal is to move beyond the current limitations of transformers, where reasoning can be generalized across more complex and extended reasoning patterns, approximating a more human-like approach to problem-solving.

This episode is brought to you by the Trainium2, the latest AI chip from AWS, by ⁠Dell⁠, by ⁠Intel⁠, by and ⁠Gurobi⁠.

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/929⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

In this episode you will learn:

(01:27) Pathway’s ground-breaking new biologically inspired architecture

(20:40) Limitless context windows

(34:39) BDH architecture as positive space

(53:11) Building multilingual models

(1:01:07) How to access the BDH architecture

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
929: Dragon Hatchling: The Missing Link Between Transformers and the Brain, with Adrian KosowskiSuper Data Science: ML & AI Podcast with Jon Krohn · 1 h 14 min
Listen in VO