In short
Eye On A.I. Episode #326 Summary: Zuzanna Stamirowska on Pathway's AI Systems
Podcast Overview Host: Craig S. Smith Guest: Zuzanna Stamirowska Focus: The innovative approach of Pathway in enabling AI systems to work with live, real-time data.
Key Highlights
Introduction
- Traditional AI applications often depend on static datasets, which can quickly become outdated.
- Pathway aims to allow developers to build AI systems capable of processing real-time data streams, ensuring models remain current and relevant.
The Core Problem
AI Memory Limitations
- Current AI models, particularly language models, lack memory, leading to limitations in their performance and reasoning capabilities.
- Example given: Describing an AI system as an intern who, despite being intelligent, cannot learn or improve over time due to the lack of memory.
Pathway's Mission
- To integrate memory into AI systems, enhancing their capacity for continuous learning and improved reasoning over time.
- Zuzanna emphasizes the importance of memory not just for learning but also for reasoning.
Zuzanna Stamirowska's Background
- Complexity scientist with experience in studying emergent phenomena and complex systems, focusing on how local interactions lead to global order.
- Co-founder of Pathway and involved in the development of the BDH (Baby Dragon Hatchling) architecture.
Key Concepts Discussed
- Transformers and Groundhog Day Effect
- Current transformer models reset their memory daily, limiting their capabilities.
- Zuzanna compares this reset mechanism to a Groundhog Day scenario.
- BDH Architecture
- Inspired by brain structures, it allows for memory incorporation within neural networks.
- Operates on a graph structure rather than traditional layers, facilitating organic growth and learning.
- Learning Mechanism
- Utilizes Hebbian learning principles, focusing on the connections (edges) between nodes (neurons) rather than the nodes themselves.
- Connections strengthen through relevance and interaction, resembling how memories form in the human brain.
- Neural Networks vs. Graph Structures
- Unlike conventional neural networks that require structured layers, Pathway's architecture emphasizes local dynamics and connections.
- The model is designed to be efficient, allowing for rapid learning and minimal computational costs.
Performance & Productization
- Pathway is currently partnering with NVIDIA and AWS to productize their technology.
- The first applications are expected to target areas with limited but critical datasets, such as healthcare claims resolution.
Addressing AI Hallucinations
- Pathway's architecture aims to reduce hallucinations in AI outputs by maintaining context and consistency over time.
- The hope is that through improved memory and reasoning, AI can avoid falling into nonsensical outputs.
Future Directions
- Pathway focuses on reasoning capabilities and potential applications in contexts where real-time adaptability is essential and where datasets are limited.
- The goal is to foster a model capable of generating innovative ideas and generalizations rather than merely recomposing existing information.
Conclusion
- The conversation highlights the potential of new AI architectures like BDH to address longstanding issues in AI, such as memory limitations and hallucinations, while emphasizing continuous learning and reasoning.
- Zuzanna invites listeners to explore the research paper available on GitHub, showcasing the innovative approaches being developed at Pathway.
Additional Resources
- [Pathway GitHub](https://github.com/pathway) - Access the research paper and further information about their AI systems.
- Contact Zuzanna Stamirowska via LinkedIn for research inquiries or further discussion.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOExploring Self-Improving AI
2:23 to 3:00
Discussion on the challenges and potential of self-improving AI systems.
“I'm not sure I entirely understand everything yet, but I've been interested in self-improving AI along with everybody else.”
Memory and Reasoning in AI
3:00 to 4:52
Insights into the importance of memory for enhancing AI reasoning capabilities.
“So we actually do follow the work of the guys at Manifest AI.”
Complexity Science and AI Applications
4:52 to 7:48
Zuzanna shares her background in complexity science and its relevance to AI.
“yeah why we're looking at those uh those problems uh before jumping into you know what we've done because i think the most maybe the most interesting thing would be to establish some kind of common ground.”
Transformers and Memory in AI
7:48 to 11:34
Discussion on transformer architecture and the need for memory in AI systems.
“what led me to it was game theory, and game theory played on graphs.”
The Future of AI with Memory
11:34 to 14:00
Exploration of how memory can shape the future capabilities of AI.
“and this is something that we kind of saw early on, they're missing memory.”
Understanding Neural Connections
14:00 to 15:00
Learn how connections between nodes create memory in neural networks.
“And as we started this discussion with talking about locality, the concept of locality and how the system works is very key.”
Post-Transformer Architecture Explained
15:00 to 16:00
Discover the architecture that mimics brain function on GPUs.
“So what we, so maybe what we're doing right now, we're actually discussing the architecture, the post transformer architecture that we published in the paper in September, which works a bit like the brain on GPU.”
Graph Structures in Learning
16:00 to 17:00
Explore the significance of graph structures in neural learning.
“we have neurons a bit like in the brain and then the neurons are linked with synapses.”
Optimization of Neural Communication
17:00 to 18:00
Understand how optimization occurs in neural networks for effective communication.
“sends over a message to its neighbors over those wires.”
Conceptual Connections in Neural Networks
18:00 to 19:00
Learn about how concepts are connected in a neural network context.
“more optimized kind of space of connecting things.”
Show all 37 chapters
Adding New Knowledge in Learning Systems
19:00 to 20:00
Examine how new knowledge is integrated into evolving neural systems.
“That was my case kind of actually when I started this research.”
Duality of Operators and States
20:00 to 21:00
Delve into the distinction between operators and states in learning systems.
“And they're talking about, you know, evolutionary systems that can grow new neurons, in effect, to add knowledge and, you know, avoid the catastrophic forgetting of fixed LLMs.”
Fast and Slow Weights in Neural Systems
21:00 to 22:00
Understand the role of fast and slow weights in neural computation.
“but is there a function in the edge then is, you know, rather than just.”
Interconnections and Initialization Rules
22:00 to 23:00
Explore how connections are initiated and evolve in neural networks.
“This in fact comes from quantum physics.”
Fixed Network Size and Synapse Density
23:00 to 24:00
Discover how a fixed network size can function with extensive synapses.
“Yeah, this is actually a very good question.”
Local Dynamics in Neural Interactions
24:00 to 25:00
Learn how local dynamics influence interactions within neural networks.
“So I mean, we've been looking at it from both perspectives.”
Addressing Catastrophic Forgetting
25:00 to 26:00
Understand the challenges of catastrophic forgetting in neural learning.
“So there is a, I mean, by now I think we know, we tested a great number of them.”
Organic Structure of Neural Networks
26:00 to 27:00
Explore the organic nature of network structures in contrast to traditional layers.
“But since it's a graph, right, you can pack a lot of synapses between them.”
Understanding Neural Networks and Graph Structures
28:00 to 29:06
Explore the differences between traditional neural networks and graph structures in AI.
“um but but let's say that yeah we see this they mentioned as being very very helpful Yeah, and the structure of the traditional neural network is in a very, it's very structured.”
Thresholding and Activation in Neural Networks
29:06 to 30:19
Learn how neuron activation relates to thresholding and information relevance.
“So for example, like why certain neurons fire up?”
GPU Implementations and Matrix Structures
30:19 to 31:29
Discover how GPU-friendly implementations affect matrix structures in AI models.
“And there will be squares and they're dense.”
Bridging AI and Neuroscience
31:29 to 33:01
Examine how AI models can mimic brain activity and reasoning processes.
“in the implementation that is GPU-friendly?”
Network Dynamics and Learning
33:01 to 34:19
Understand how connections in AI networks evolve during learning processes.
“That sounds a little bit like mixture of experts.”
Scale-Free Networks and Resilience
34:19 to 35:46
Find out how scale-free networks contribute to resilience and efficiency in AI.
“This is also maybe coming really deeply from the complex systems intuitions.”
Performance of Language Models in AI
35:46 to 37:08
Analyze the performance of AI language models compared to traditional models.
“So yeah, some could be if you don't use them.”
Efficient Use of State in AI Systems
37:08 to 38:14
Learn how state is efficiently utilized in AI networks for better performance.
“right so so uh yeah so but how how how does it perform uh for example in uh is it's not auto-aggressive right or is it uh how how does it perform on language i mean you're you've built a language model with this, right?”
Concept Representation in Neural Networks
38:14 to 39:46
Discuss how concepts are represented and emerge from neural connections.
“So these are the benchmarks that we are the tasks that we're showing in the paper.”
Plasticity and Resilience in Neural Networks
39:46 to 41:18
Delve into the idea of synaptic plasticity and its role in network resilience.
“connections between finally not so many neurons.”
Function Shapes and Neural Connections
41:18 to 42:00
Explore how functionality can influence neural connections and their roles.
“talk about so this is this is this is this is a beautiful question i would like to see your questions are actually very deep.”
Exploring Network Resilience and Functionality
42:00 to 44:20
Learn about how network functions shape their structure and resilience.
“As the system lifts, I mean, we have this effect, and then there's this kind of update to the slower weights that just, I mean, change at slower frequency.”
Creativity and Consciousness in Neural Networks
44:20 to 46:30
Discuss the relationship between complexity in neural networks and creativity or consciousness.
“And then as for concepts appearing, I mean, yeah, there is a number of things that happen with the fact that we're just working with the dimension n, but this is maybe jumping across topics.”
Generalization in AI Systems
46:30 to 49:20
Understand the importance of generalization in AI and potential eureka moments.
“of how consciousness appears in the brain or how it can, you know, appear ultimately in a lens.”
Productization and Use Cases of Emerging AI
49:20 to 53:00
Explore the productization process of AI and potential use cases in various industries.
“And then there is, of course, a lot of work to be done for this.”
Addressing Hallucination in AI Models
53:00 to 56:06
Delve into strategies for reducing hallucination in AI and maintaining consistency.
“And as you go on, you don't fall over very easily.”
Pathway's Unique Approach to AI Scaling and Memory
56:06 to 59:15
Explore how Pathway's AI prioritizes reasoning over scale for efficient learning.
“In hardware sense, it all sits on the chip.”
Integrating Different AI Models for Enhanced Performance
59:16 to 1:02:20
Discover the potential of combining BDH and transformer architectures in AI systems.
“Do you think this would work with traditional LLMs?”
The Future of AI Research and Collaboration
1:02:21 to 1:07:22
Learn about the importance of research collaboration and the role of AI in understanding AI.
“Then there is, however, another just anchoring maybe on the merging type of or gluing angle.”
Transcript
Automatic transcript. May contain errors.0:00Craig Smith:So it's a bit like having an intern that you hire on the first day of her job. And she may be brilliant, but she stays an intern forever on the first day of her job. Doesn't get more context, doesn't get better with experience, doesn't get better over time.
0:13Zuzanna Stamirowska:Do you think there will be creativity in this kind of a network? We'll be able to think beyond the fixed parameters of an LLM and come up with new ideas. This episode is brought to you by Tasty Trade. On Eye on AI, we talk a lot about how artificial intelligence is changing how people analyze information, spot patterns, and make more informed decisions. Markets are no different. The edge increasingly comes from having the right tools, the right data, and the ability to understand risk clearly. That's one of the reasons I like what Tasty Trade is building. With TastyTrade, you can trade stocks, options, futures, and crypto all in one platform with low commissions, including zero commissions on stocks and crypto, so you keep more of what you earn.
1:11Zuzanna Stamirowska:The platform is packed with advanced charting tools, backtesting, strategy selection, and risk analysis tools that help you think in probabilities rather than guesses. They've also introduced an AI-powered search feature that can help you discover symbols aligned with your interests, which is a smart way to explore markets more intentionally. For active traders, there are tools like Active Trader Mode, One-Click Trading, and Smart Order Tracking. And if you're still learning, Tasty Trade offers dozens of free educational courses, plus live support from their trade desk reps during trading hours.
1:58Zuzanna Stamirowska:If you're serious about trading in a world increasingly shaped by technology, check out Tasty Trade. Visit tastytrade.com to start your trading journey today. I'm going to myself. Tasty Trade Inc. is a registered broker-dealer and member of FINLA, NFA, and SIPC. I mean, I've done some reading. I'm not sure I entirely understand everything yet, but I've been interested in self-improving AI along with everybody else. and continual learning and those things. And post-Transformer Architectures, I've spoken to a few people. Carl Friston, I don't know if you followed his work, but he's working with a startup called Versus.
2:54Zuzanna Stamirowska:And then there's a company called Manifest AI in New York City. I don't know if they're at all related to what you're doing, uh but in any any case i'm i'm interested and you guys are looking at this specifically for uh robotics applications or no not necessarily so first of all craig thank you so much for having
3:20Craig Smith:me uh it's a pleasure to be to be here um and i mean you also had a chance to talk to wonderful guests. So we actually do follow the work of the guys at Manifest AI. So it was actually great to hear them on your podcast as well. We don't necessarily focus on robotics. In fact, what we focus on and the part of maybe some sort of obstacles that we see for AI that are pretty big that we are lifting and hoping to lift are linked to memory. so not just continual learning not just self-adaptive systems but memory and then memory leading uh to enhanced reasoning so it's slightly more than just you know allowing for a larger state or or the capacity of learning over time it's part of it but memory can also can also power reasoning and having let's say memory organized in a different way we can get to representation uh let's say something that folks call call that internal representation that may potentially be more interesting and easier um okay uh can can you uh give us a little your
4:41Zuzanna Stamirowska:background before we start talking about uh baby dragon hatchling pretty pretty creating name of
4:51Craig Smith:course so i would love to um actually i would love to also maybe set the stage a little bit yeah why we're looking at those uh those problems uh before jumping into you know what we've done because i think the most maybe the most interesting thing would be to establish some kind of common ground. For what we see is happening, I think many people across a number of labs by now also are seeing and some of your guests as well. So as for my background, I'm actually a complexity scientist. So I worked at the Institute of Complex Systems of Paris. I am specialized in studying emerging phenomena. So seeing how some sort of global order emerges from small local interactions and how locality matters in appearance of phenomena.
5:50Craig Smith:So, and this may happen on different sorts of topologies. I mean, in complexity science, we usually look at the graphs. So graphs is just a graph is a network, right? You have like dots, which represent the nodes, and then you have edges that link to notes. This can be your road network in a city, right? This can be your internet, this can be your social network, this can be your network connecting neurons. And something happens on those links, right? There's some sort of, I don't know, cars travel, mail is being sent, right? Water travels through pipes, whatever. But something happens. I mean, sometimes there is some operation being done on the link.
6:33Craig Smith:There's some sort of function. that then also drives the evolution of this network. So it's a system that may be evolving following the local interactions that happen on the network. So I know for roads, we would look at, okay, somebody built a road in one place, but maybe they built a road between two places where two types of objects were produced to exchange, right? But then there was a settlement somewhere. So, okay, then we built a road again. And since there was a road, people started producing things because it was easier to trade with that location so you sort of start to have this loops you know uh in a way that well that's that's the network structure uh what's happening or is it uh is it the other way around kind of chicken and the problem thing is it's an evolving system uh where things happen locally not are not necessarily usually always planned by a super designer at the global level that says what should be happening, which function should be done where.
7:37Craig Smith:You'd rather have this organic emergence. And this is what I actually focused on early on in my career. Before that, I worked. And actually, what led me to it was game theory, and game theory played on graphs. So you can have players, but players play on not like one-to-one or many-to-many, but they play on potentially complex topology and not against everybody at every step. But I also was trained at the French school, at the school for French politicians. In my past, I also got a chance to teach for example, at the same place. Yes, Siong Spoh. So yeah, this is a very useful type of education as I kind of see it right now being a CEO of a startup.
8:33Craig Smith:So, and then, you know, when you look at AI and many implications that it has, and probably trying to project where we can get with AI as a society and also as countries and how this may impact maybe even the concepts of how states collaborate, et cetera. I mean, perhaps at least I have some anchoring to start to think about this. So this is my background. And I have a chance right now at Pathway. I mean, I'm a co-founder and also a co-author of the BDH architecture. I have a privilege of leading a team of fantastic researchers in AI with sometimes various backgrounds. So our CTO, Jan Horowski, was the first person to apply attention to speech recognition.
9:34Craig Smith:So that was in the pre-transformer era, and he was at Google grade. Our CSO, Adrian Kosowski, is a quantum physicist and a theoretical computer scientist at the same time. and he has his PhD at 20 and you know there's there's entire team built of algorithms missions people have background from physics as well and I'm I'm just I'm just you know extremely actually proud to be to be able to do
10:05Zuzanna Stamirowska:them and and what were you doing I mean you know you wrote this paper on uh dragon hatchlings uh i'm actually looking i don't remember when that was but uh and and it was uh that was in september was that before uh or after you founded
10:28Craig Smith:uh pathway no we found this pathway way before so we've been working on this um on the problem of memory and intelligence for a long time. And this is just, let's say, the first glimpse into what we're doing internally. So if you look at LLMs, right, and this is something that I think is becoming more and more public, I'd say, or maybe more people are becoming aware of this, and in general public, like all the LLens are pretty much based on the transformer architecture. That was a huge unlock for AI when that enabled all of us pretty much to have the imagination of what could be possible with AI.
11:21Craig Smith:The transformers are probably not, and I mean, as Pathway would argue, are not the last architecture in this entire AI market shift. And specifically from what we're seeing, and this is something that we kind of saw early on, they're missing memory. So they work a bit like a Groundhog Day. I mean, they wake up every day with their memory of their interactions, of whatever problems they were solving being completely wiped out. So it's a bit like having an intern that you hire on the first day of her job. And she may be brilliant, but she stays an intern forever on the first day of her job. So doesn't get more context, doesn't get better with experience, doesn't get better over time.
12:22Craig Smith:and then also isn't capable of staying focused on a task for a very long time without falling into hallucinations. And then, of course, I mean, with the capacity of staying, let's say, focused and coherent on solving a task, the complexity and the value of tasks that AI can accomplish will be growing. And there's a lab that is called meter that is measuring, like the length of human tasks that AI can do right now, successfully. And we are at a couple of hours right now. And of course there are so many tasks that would in reasoning, for example, would take it way long, will take us way longer, potentially also with changing inputs, data coming from the environment, right?
13:14Craig Smith:Not all the problems are closed form kind of math, math problems. some are open highly dependent on context some may depend also on small data um so yeah so this memory but i mean we see it also from the market perspective as let's say one of the vectors of the of the a market shift and there are people who of course work on the topic of memory from the hardware side um and we are working on this from the algorithm just to back up a little
13:44Zuzanna Stamirowska:bit and and pull out um uh to a wider view so you're you're working i mean these are your language models are uh neural networks uh but the the difference is instead of the uh information being coded and i'm guessing now or from my understanding being coded in the weights of the nodes uh the it's the edges the the connections between nodes that that strengthen over time and and in effect uh create uh memory is is that am i getting that at all right it's a bit more uh in a sense it's a bit more uh it's a bit more because of the learning mechanism
14:41Craig Smith:that we have. And as we started this discussion with talking about locality, the concept of locality and how the system works is very key. So we have a graph structure, a graph kind of structure of connections between neurons. This is true. So what we, so maybe what we're doing right now, we're actually discussing the architecture, the post transformer architecture that we published in the paper in September, which works a bit like the brain on GPU. So it follows the rule of Hebbian learning. This is a simple brain-like model, and we managed to get it to GPU, and it actually trains even better than transformer when we compare it to the GPT-2.
15:35Craig Smith:So this is the transformer architecture. We compare architectures, not yet models, to models, right? At one billion scale and has a number of great properties. We designed it, in fact, to have memory and then to align it with the way we see reasoning, let's say, unfolding. So we can discuss it a bit later on. So to give you maybe an image of how in this system learning works, we have neurons a bit like in the brain and then the neurons are linked with synapses. You can imagine wires or you can imagine roads. In a traditional transformer, everybody is connected to everybody. Imagine a conference call where everybody is connected to everybody and if you want to double the number of participants, you have to quadruple the complexity of everything that's happening.
16:41Craig Smith:Well, here not everybody is connected to everybody. We have a graph structure of relevant connections that are trained actually with signal that comes into the system. So whenever a signal comes to the system, the neuron, which was, let's say, triggered, sends over a message to its neighbors over those wires. And they're not connected to everybody. They're connected, let's say, just to their friends, like you on Facebook, for example, right? They send over their message, the message just to their friends. According to some rule of a threshold, the friend is interested or not, so fires up, potentially, receiving the message.
17:27Craig Smith:And then in the following step, And this connection between them becomes stronger because it's been proven to be relevant. And it helps to create certain shortcuts or kind of wider roads, okay, like high-speed roads between certain places in the network. And in that way, during learning, you get optimization of the space. I'm not saying this very formally right now, so please don't hold me to it, but it's somehow more optimized kind of space of connecting things. And we find that this structure optimizes itself naturally during training for communicability. And this is, in fact, very often a feature of complex systems that are grown organically.
18:26Craig Smith:You need to The systems naturally like to balance communicability and, let's say, the cost of maintaining the links. But it should be usually easy for you and quick to reach everybody who's important on the network quickly. And then in this structure, the links can appear according to... We don't control the rules according to which they appear. So it may be the information, you know, you were at the cafe with a friend and that friend told you something very interesting about neuroscience. That was my case kind of actually when I started this research. And you will have a connection between the taste of that coffee and, you know, something that happened in the brain of a mouse.
19:16Craig Smith:And you may have this connection. It's not a very formal one. It's the one that you kind of have. and then it may help you to do some sort of informal thinking and just kind of exploring this space of, I don't know, related concepts according to whatever types of links that you could have found while learning.
19:38Zuzanna Stamirowska:Let me just ask some dumb questions to try and triangulate where this sits in the graph of my knowledge. Because I've been, I had a conversation recently with David Ha at Sakana AI, And they're talking about, you know, evolutionary systems that can grow new neurons, in effect, to add knowledge and, you know, avoid the catastrophic forgetting of fixed LLMs. In your baby dragon hatchling system, two questions. Is that what happens as you learn new things? Does the system add new nodes or is it all in the strength of the connections along the edges? And how does that strengthening or weakening, mathematically, is there?
20:58Zuzanna Stamirowska:Because in a standard LLM, there's a function in the node that encodes knowledge and could be wiped out with retraining. but is there a function in the edge then is, you know, rather than just.
21:26Craig Smith:This is, this is a very good question. This is, this is one that our CSO likes very much. And I'd say it's mostly coming from him. So you have something, and I mean, this will be, this will be perhaps like pretty technical, but an important distinction for intuitions when, when one is thinking about such systems. and I think it's not very often the case when we talk about LLMs, it's a duality between operators and the state. The intuitions, the way it was actually apparently, I mean, I'm just repeating after our CSO because this is his part. This in fact comes from quantum physics. Just, I mean, historically in science, how it's being developed.
22:14And in this case, the state sits on the edge and on the fast weight, as we call it.
22:23Craig Smith:And then we have the equivalent of it, which is a slow weight, which is a parameter. And this is kind of key. And what happens is that neurons are computations, but really the state sits on the edges, on the fast edges mostly. So this is how we look at this. And what happens is that while with every step you actually also update, you use the network topology and waits for message passing. but then you also after you've done it you update your uh your fast weights and you update the the slow weights let's say at the slower frequency um and this is yeah this is exactly and if you're
23:19Zuzanna Stamirowska:if there is no um connection between uh two nodes uh but as you said you know maybe there's a node for a coffee shop maybe there's a node for for I mean this is obviously not realistic analogy but maybe there's a node for you and for the person that you're talking to but you the the person you're talking to is connected to the coffee shop and you're connected to that person how does the system know to build a connection, an edge, or whatever?
23:59Craig Smith:Yeah, this is actually a very good question. So I mean, we've been looking at it from both perspectives. Important to know that concepts, let's say, if you think about concepts, concepts sit more on the edges than on the nodes. So we're looking more at the synapses where the concepts are sitting. And then we would say that, OK, actually, if a concept is important they will have one node responsible for that concept one note one synapse actually responsible for that concept and this is what we find what we demonstrate with one like with an example in the paper um you know we have we have like the concept of currency and we find this this synapse uh firing uh firing up um so yeah we sorry can you repeat the question yeah i mean how
24:47Zuzanna Stamirowska:How do you build a connection between two nodes if there's no connection?
24:53Craig Smith:Yeah, exactly. So the nodes, in practice, we usually keep the N as fixed. We keep the N as fixed and there is an initialization rule at the beginning. So this is usually when you start growing at work, you need to kind of start with, have an initialization rule for the very kind of first basic structure to appear that then starts to iteratively kind of grow and enhance. So there is a, I mean, by now I think we know, we tested a great number of them. There's like a simple initialization rule, which is linked to, okay, how many attempts you have creating your your first links um at the very beginning and then as as as the network
25:40Zuzanna Stamirowska:uh how does the network grow i mean this is what i was talking to david ha about how do you or or it don't doesn't grow and the network so our network you may imagine it doesn't grow so much i
25:53Craig Smith:mean the n the the the number of the size of network defined as the number of neurons in the brain, in practice for us stays fixed. But since it's a graph, right, you can pack a lot of synapses between them. And then you can also, and so in the brain, we're looking at hundreds of trillions of synapses for billions of neurons. So this gives us like a very, very, very large state that we can use it access and use efficiently at any time, because we don't always need to use the entire state. We actually use it very locally. Thanks to the local dynamics, we may have huge state and use very little of it at every step because of the local interactions, like the local dynamics.
26:45Craig Smith:This is key. This is very different to everything else that happens in the ASP for us. I know, because we are really focusing on the local dynamics and we found a way to kind of maquillage this and apply some makeup and make it work. This is kind of sparse things to make them work on GPU and make them seem like dense matrices, but actually being kind of preserving mathematically the sparsity of interactions that we have and I'm happy to jump into this. So yes, so point is local interactions. And if you do it, you effectively have the context size, which is the size of your model. But you don't use always all of that.
27:29Craig Smith:You use just the tiny bits that are necessary for you. And then in terms of catastrophic forgetting, well, one of the arguments is that once you learn new skills, you can have so much space to pack things that you don't impact you don't necessarily impact the ones that were linked to to your previous task uh i mean it's too early for me you know to to to talk more about uh like the solving the curse of catastrophic forgetting and continued learning um but but let's say that yeah we see this they mentioned as being very very helpful
28:08Zuzanna Stamirowska:Yeah, and the structure of the traditional neural network is in a very, it's very structured. There are layers and there's the width of each layer. This is much more organic. Is that right? It's not in layers. It's in a graph. So there's, yeah. Exactly.
28:39Craig Smith:So it's in the ideal scenario, right? And the pure thing to really build your intuitions and to really explain how and why it works. It's exactly a graph with local dynamics that work like a rumor spreading on networks. For technical listeners, I would invite you to maybe have a look at John Kleinberg's and Eva Tardosh's papers about rumor spreading. So for example, like why certain neurons fire up? I mean, the rule is linked to thresholding. So you need to reach a certain threshold, in fact, of relevance of information to activate the neighbor, et cetera. But it's really like spreading a rumor on the internet, right?
Read the full transcript
29:26Craig Smith:You cared enough to pass over to your friends and and etc some people didn't care or like a bit like epidemics um so this is this is the logic then when you bring it to gpu and i get get questions about this because of course you know was being seen uh like with the paper and online and github is just a small portion of uh of our work um what do you what do you see uh what do you see in the code and in the very GPU-friendly implementation that we have there are ultimately, we have to talk about matrices. And it looks very much like a transformer. But there are important differences that happen.
30:11Craig Smith:Attention is linear. So when you look at sizes of, let's say, our matrices, the traditional transformer has this large square. right? And there will be squares and they're dense. Whereas for us, it looks, the only dimension that we care about is the N. That's the number of neurons effectively. And the objects we work with look like very kind of narrow, but very long snakes. Okay. So we have those huge squares in transformers that grow quadratically. And we have this very long snail, which effectively we only care about the length, which is n. And this is the number of our neurons. And then to make it sparse, because of course, we had to, as I said, we had to apply some makeup to make it work on GPU that really likes dense matrices.
31:08Craig Smith:We actually multiply it by a positive sparse vector of activations. And this, once you actually multiply one for the other, you get back your graph.
31:25Zuzanna Stamirowska:Yes. Just.
31:28Craig Smith:So this is for the listeners who might be asking, okay, where is the sparsity hidden in the implementation that is GPU-friendly? Actually, mathematically, it is there. So you get it back, and hence you can actually do other things like spot the synapses responsible for specific concepts, et cetera. I mean, sparsity is very much there, even though if you just look at the code like this, you may not spot it because it's hidden in one specific multiplication. And the fact that we have sparse positive vectors of activations, this is a very big difference that our CSO loves to explain as well. I mean, I think once he used a tower of Macaron to explain the kind of the space of those factors.
32:17Zuzanna Stamirowska:And you talked about you need a certain amount of signal to activate a neuron. That sounds like spiking neurons is...
32:30Craig Smith:It's an edit. This is not... I mean, intuitively, I mean, not in all the implementation details, but intuitively, it's not too far. So we're getting to, so what actually response to our paper was especially good from the neuroscience community, because indeed we kind of started to show the way of bridging very non-organic AI with the models and the thinking of how neuroscientists see brain activity of course we're not getting to the levels of chemical interactions right like reactions that happen in the brain um but we we are offering like a plausible brain-like explanation for for
33:21Zuzanna Stamirowska:for how reasoning may appear and then the sparse uh uh the the activation of when the network is working. I understand that it's very localized. You're not using the entire network. It's whatever the edges connect. That sounds a little bit like mixture of experts. Am I wrong there?
33:54Craig Smith:This isn't just at a very, at infinitely small, I mean, not infinitely, but really tiny, granular scale. So if you were, because mixture of experts, you know, is more kind of defined from top down. Here, imagine that your expert can be even just, you know, one, potentially could be just one edge, right? And you don't decide expert groups that are being decided on their own. kind of organically while training. This is important. This is also maybe coming really deeply from the complex systems intuitions. We really don't impose any structure on the system. The system does what the system wants to do.
34:42Craig Smith:And that's the point. That's the magic of it. This is also why we believe this kind of, ultimately the space that's created it can be very interesting for reasoning. Because at the same time, the concepts appear, the more important ones are strengthened. And it gives you a topology, potentially with shortcuts, to explore. And if you believe that reasoning is a way, some sort of search in the space, then this gives you an interesting topology to explore. And it's aligned with the method of learning itself.
35:17Zuzanna Stamirowska:And then as the system learns, edges become stronger or connections become stronger or new connections are formed. So the graph becomes denser, not in the number of nodes, but in the number of connections.
35:45Craig Smith:And yes, plus some may fade over time.
35:51Zuzanna Stamirowska:Yeah.
35:52Craig Smith:So yeah, some could be if you don't use them. I mean, you have only positive activations, but you may have fading as well. So in a way empirically what we find is that we get like almost a scale-free type of distribution of degrees of nodes, which gives it. This is a type of structure that we find in the real-world organic networks quite a lot. Because this is one that's known to be pretty resilient and with good communicability, and also it behaves similarly wherever you zoom in the network. So it has similar properties, irrespective of its size.
36:39Zuzanna Stamirowska:and how how you said you you built uh like a billion parameter model using this architecture is that right how how is its uh not only how is its performance but i understand memory and in transformers the intention mechanism but how the memory here is the strength of the connections right so so uh yeah so but how how how does it perform uh for example in uh is it's not auto-aggressive right or is it uh how how does it perform on language i mean you're you've built a language model with this, right?
37:39Craig Smith:Yes. So this is, we train it on predominantly on language. I mean, we're not talking about the vision models, for example, right? Like there is an entire line of research and right now attracting a lot of attention, which is actually linked to vision and world models per se. We actually believe that there is maybe a way to bridge the two. But yes, we're looking at the language model and we compare it on traditional language tasks. So these are the benchmarks that we are the tasks that we're showing in the paper. We're looking at translation, we're looking at how it behaves at traditional language tasks.
38:28a bit the way in which transformers were developed.
38:32Craig Smith:So the question of scaling it further, in fact, and the features that we see more likely coming out of it, out of this architecture scaled, well, scaled or developed into a model, into a product, is, let's say, the power of your output with way less data. So with a way smaller model and way less data, we should be able to get to the same results. I mean, compared to GPT-2, so I mean, comparing apples to apples, architectures to architectures on the same data sets, exactly same data, I mean, we're actually learning sometimes even faster. Event transformers and scaling loss are preserved. Knowing, however, that we don't need to grow this size of neurons so much, and we can be adding, I mean, we have state, right?
39:33Craig Smith:Which is huge. And which is huge, well, we can actually use it efficiently. So it's both packed efficiently, because it's packed in the graph between the connections between finally not so many neurons. And then because of the local dynamics, we only access a bit of it at every step or. And again, some some more dumb questions so I can
40:04Zuzanna Stamirowska:understand this. I mean, in in the brain, these local networks that are activated represent
40:19concepts, perceptions, all of that, memories,
40:25Zuzanna Stamirowska:and in your system, but it's not,
40:35Zuzanna Stamirowska:let me think, how am I going to say this? The knowledge that's stored there depends on the connections between these neurons. And the knowledge kind of emerges from those connections, right? It's obviously not stored explicitly. uh and does does that is there some plasticity in that like if you cut a bunch of connections you still have the knowledge maybe it's got to grow stronger connections again but yeah what can you
41:22Craig Smith:talk about so this is this is this is this is a beautiful question i would like to see your questions are actually very deep. So definitely none would qualify as dumb question. This is dramatically important. I'll take it in two ways. So the very concept that we have here is synaptic plasticity. So this means that not only a message is passed through the network or through a connection, but then this connection becomes stronger because it was triggered. And this goes on perpetually somehow, right? As the system lifts, I mean, we have this effect, and then there's this kind of update to the slower weights that just, I mean, change at slower frequency.
42:12Craig Smith:But then for the plasticity of the network, I mean, this is a brilliant topic, actually, for complexity science and network resilience in general. And one intuition that's actually this one came from me. I mean, most of this research actually came from my colleagues and our team. But this one came from me is that function shapes the network. And you actually may have, and this is something that the neuroscientist at that coffee shop told me, you may sometimes swap the nerves and the function of them though, like let's say the example he was giving me was the auditory with the visual nerve. But then the mouse will turn out just great.
43:04Craig Smith:Because the stimulus that was kind of given shaped the network because the network's job is to transport it in a way. So it's getting these local functions of what the network, what the edges, what the nodes are doing that then help you to potentially even compensate. So I don't think we've run deep enough studies of what's happening even in our model right now in terms of its resilience to deletion of nodes. of nodes, but I've done a number of studies of how such systems work and how scale-free networks behave when you delete. So two intuitions is that the function shapes network, there's a good chance that some functions will be overtaken by the locality, and this will also happen according to the rules of local connections, like let's say common neighbors.
44:08Craig Smith:between the nodes. And this is like some redistribution maybe of a function is likely to happen. And yeah, so this is a very deep topic, very, very interesting. And then as for concepts appearing, I mean, yeah, there is a number of things that happen with the fact that we're just working with the dimension n, but this is maybe jumping across topics.
44:32Zuzanna Stamirowska:Well, actually, I was going to jump, but, well, I will. So you're in complexity studies and you're studying emergent properties. And the obvious thought is that in consciousness studies, you know, a lot of people think consciousness is an emergent property of the complexity of neural interactions. Do you have any thoughts on that? I mean, I know that's in left field, but are there, or rather in a less sort of unanswerable way, Do you think there will be creativity in this kind of a network in that the network will be able to think beyond the fixed parameters of an LLM and come up with new ideas?
45:46Zuzanna Stamirowska:and this is something I was talking to David Ha about with evolutionary systems in general. They're looking across a landscape, not simply following gradient descent to a local optimum or something. uh so so they can jump over across to find um you know ideas that that didn't wouldn't necessarily emerge from a fixed llm so yeah
46:27Craig Smith:yeah absolutely so i i don't think i feel nearly qualified enough to have a position of how consciousness appears in the brain or how it can, you know, appear ultimately in a lens. Just somehow completely by accident I started reading about Heidegger over Christmas. And I feel, I mean, I don't even know the advances in philosophy lately, but the definition of being and thinking may be somehow impacted by what's happening with the reasoning models. So, I didn't say just my kind of cultural thoughts. When it comes to generalization, of course, this is the main goal. This is why we're doing all of this.
47:20Craig Smith:because the real innovation wouldn't be just recomposing things that exist but pretty much seeing the loophole and this is then when you know there is something interesting that should be added how do you know that there is something interesting that should be added that you need to actually somehow modify what's your most modify your topology right and yes so we are there are two ways to maybe generalize the general the most imminent generalization that we are after is generalization over time so this is that this is this is the kind of capacity to to maintain coherent reasoning over over time without falling into hallucination and just doing something completely silly and that's the intuition and then generalization in terms of having really i mean internally we call it the model which is a real innovator right one that can come up with eureka moments and this eureka moments they don't they don't result from very formal thinking i mean i believe that very few mathematicians really kind of start their proof and work it's formally step by step reaching reaching the the conclusion i mean usually they have in different different bits of information of conviction filling in dilemmas on the way sometimes or even like strategic proofs uh like so i know mathematicians who actually design ways to strategic ways to to get to their proofs So, yes, and I think that this will go through having some sort of like an internal representation that allows for it.
49:09Craig Smith:Our hope is that having this plasticity and the structure which is not 2D, is not 3D, is actually somewhat efficient and emergent will make it easier. And then there is, of course, a lot of work to be done for this.
49:26Zuzanna Stamirowska:So what are next steps? I mean, this is not productized yet. It's still in the research phase. Is that right? And where are you guys headed?
49:37Craig Smith:Correct, knowing that you're actually right now working on productization. Because, I mean, naturally it locks a lot of interesting properties. I mean, some of the first ones are infinite context windows. and hence reasoning in some use cases can be, you know, jump from zero to one. So we actually partner with NVIDIA and AWS, something that we announced at RainVent in December. And the moment our, you know, the first model will be ready, it will be immediately available to AWS customers. And yeah, so this is, I'd say, this is being productized as we speak, knowing that we're productizing along yeah every search and what kind of use cases do you think the first iteration would would be applied to um so the very good use cases for this are the ones that are links to small so we're looking highly valuable small data where uh we're actually want to we were doing for example research about um about or are you making new designs for uh anyway we actually saw a use case like this for for uh in the nuclear space right you don't have a lot of uh a lot of documents but those that exist are actually very very highly valuable and you you would like to get some ideas for um for new for the current engineers right and ideally you won't learn it very effectively because there's just not enough data.
51:15Craig Smith:And then we're looking at cases where you have continued learning and data that kind of keeps on changing where you want to deliver value. And then also reasoning, which is slightly more kind of complex, which is highly personalized. So the use case that we look at is, for example, healthcare claims resolution. This is because of the fact that, okay, it has to be very personalized. It is complex reasoning because you need to, within the context of, you know, what's linked to the process of the claim resolution, all the information that was given, you kind of need to resolve it. and also you need to be able to explain why and how.
52:06Craig Smith:And because we kind of see, we have some advantages in terms of interpretability because we kind of see the synopsis that light up. And then also in reasoning, I mean, you can have layers of, let's say, explaining of what happened. We have some big benefits for the regulated industries. so but journey speaking think about small data uh with time changing elements yeah and and potentially next best action uh suggestion uh yeah which is you know or like a sort of reasoning and in terms of um
52:52Zuzanna Stamirowska:the traditional traditional i mean they've only been around for a few years but llm uh weaknesses hallucination does this uh eliminate the the problem of hallucination uh because hallucination comes happens during inference when i mean from my simplistic understanding when the probability distribution of the next token doesn't include the correct token and uh and and gives something that's in the distribution but then that error is propagated uh as as you go along um is is first of all is that right but second of all uh in that you're not uh predicting uh the next token from a probability distribution you're you're you're i don't know if reading's the right word but you're you're looking at uh an existing uh the activation of an existing sub network within the network or within the graph uh yeah how does that affect uh hallucination
54:22Craig Smith:yeah so it's just actually one of the big motivations for for this work um was to limit hallucinations right and and especially hallucinations over over time so important to note because i think we didn't stress it at the beginning all the dynamics that we that we talked about happen at inference it all happens at inference uh so indeed this is this is this is the hope right that you because you keep the memory, you can keep consistency. And as you go on, you don't fall over very easily. I mean, that's like a structure to hold on to. I mean, it was too early for me to communicate results.
55:13Zuzanna Stamirowska:uh yeah and uh and the uh uh you you said that it follows the same scaling laws or or similar scaling laws to transformers which was was a huge advantage for transformers and and you know what they did is just build bigger and bigger and bigger networks and it got better and better or better at least to a certain point do you think the same will happen with uh bdh with baby dragon
55:48Craig Smith:what is yeah actually the paper is dragon hatchlings it's dragon hatchling and the the acronym is bdh and i think people are very very just tempted to to to put a b explain to me um so it's dragon hedge link uh we are not playing the scaling game in the sense that we don't need to grow the end so much um and the ultimate goal but this is you know this is this is the goal of almost almost everybody for for us that the path goes through through memory uh but the ultimate goal is to to get to generalization and generalization in reasoning so so yeah we are focused on reasoning models and and and seeing actually not the scale per se because the power won't be coming from the scale um so yeah point is we don't need to we're not playing the game of having our n you know being humongous because i mean that that's precisely not not the reason why we're doing it I mean, this is a way more efficient way of learning and then evolving and storing memory.
57:04Craig Smith:In hardware sense, it all sits on the chip. And this is a big advantage. And if you look at those, if it sits on the chip, you do less lookups. And you may divide your compute for reasoning 10 times for the same output token. right? Plus you can probably do reasoning on cases which are more complex because you have consistency or like consistency, you have capacity to work for longer without falling off into hallucinations. So it may be a more sustainable, I mean, we believe it's a more sustainable way forward also in terms of compute and how it distributes. It's a distributed system. of the there is a part of theoretical computer science like a very big community actually that that deals with distributed systems um and it is it so for those listeners if they're here it is a
58:03Zuzanna Stamirowska:it's effectively a man like a distributed system uh the uh is is there the reason you're not looking at scaling is because the number of uh connections between the nodes or the neurons in this graph uh can can you can pack enough in that you don't need to scale the number of parameters or or i mean it could be that if you scale uh you'll you'd be able to contain that much more information It's possible.
58:45Craig Smith:It's possible. So I'm not saying that we don't need to reach the scale. Our success is not blocked by the size of N. I would say it this way. So for us focusing on larger N, it's probably fun. But I mean, our goal is in fact to get to this reasoning in places where it wasn't possible before. And so I'd say we have a slightly different objective. And for this, as of now, we didn't see the scale as being the main bottleneck. And yes, this structure allows us to pack a lot in the connections, not in the end, but in the synapses, and then keep it apart enough because we activate only parts of it and such that when we touch one part, we don't necessarily touch the other part of the of the model so i mean hence hence some of the nice properties for against catastrophic forgetting um and like the plasticity is very very important in in this entire work
59:49Zuzanna Stamirowska:you know with transformer architectures uh once they were validated and they exist then there's this whole second layer of activity in combining models or combining different kinds of models into larger systems. Do you think this would work with traditional LLMs? You know, maybe, I don't know, maybe the transformer architecture handles one thing and the post-transformer architecture handles something else.
1:00:31Craig Smith:Oh, in that sense. So, I mean, in terms of maybe putting models together, I mean, at the engineering level, I believe there will be different use cases for which different models will be used. And, you know, for your traditional knowledge type of use cases where we have language models, I mean, they are doing phenomenally well. and I believe that a lot of the right now chatbot types of use cases, I mean, they will stay with the ELMs and they will be doing great. So, you know, I believe that the models based on BDH and other architectures, I mean, they will be used in use cases that, for example, require consistent reasoning over a longer time, right?
1:01:20Craig Smith:So with fin data. So we're definitely looking at enterprise use cases naturally, or deep innovative research, where you were potentially with also changing inputs, right, over time, because you may want to kind of maybe iterate along with the environment as things progress. So I know in the engineering I imagine you can use different models for different bits and it's kind of a question of just building the system in which you can do it. On the more fundamental way, can you glue BDH with a transformer and have kind of one model that somehow works together? Because I mean, this is what some people who modify attention do.
1:02:08Craig Smith:This is not the case. This is not the case. So this is actually either transformer or BDH, because it's so different, so fundamentally different. It's not, let's say, a plug into a transformer or a plug-in. Then there is, however, another just anchoring maybe on the merging type of or gluing angle. Because we have just this dimension n, like just the neurons right it's a graph and it's a distributed system you you can totally see that you're taking those brains and you kind of train them separately and you build them together and then you maybe have a couple of runs of training together such that they form connections between them and then becomes even stronger and in this sense you can get you know not not a separate mathematician and a separate separately trained uh computer scientist but all of the sudden you actually have a mathematician and computer scientist who combines the intuitions of the two disciplines in one um or you know this can work for finance and law you're talking about
1:03:19Zuzanna Stamirowska:training uh separate uh graphs and then merging them is that right yeah then gluing them and we
1:03:27Craig Smith:glue them along this one dimension n uh we actually show a small experiment in the paper for this where we actually didn't do the kind of the round of training when they were glued together but we actually on language we do we're going to glue the two the two models train into two different languages and then all of a sudden they start to speak uh you know some sort of like i mean like the language that make makes sense but mixes the words of of of of the two languages point is it's very it's very simple with uh with this architecture because because you effectively you look at the graphs and you just put them together.
1:04:05Craig Smith:So it feels a bit like composable programs.
1:04:09Zuzanna Stamirowska:And it sounds, you know, again, Sakana AI came up with this, I think they call it model merge, where they can put two models together. And that's part of their evolutionary strategy that they'll have a bunch of models, models generate output, they pick the best models and merge them into, presumably you could do that as well with BDA?
1:04:48Craig Smith:Presumably you could. I mean, I haven't explored it in full honesty, but yes, you could technically because you can merge them And if you want to train evolution, do some sort of kind of evolutionary pruning of what's your best model, then probably you could do it. I don't know if it has a purpose in our case, but technically, yes, the merging is it.
1:05:15Zuzanna Stamirowska:That's all fascinating, Susanna, really. And yeah, because Transformers, it seems, have kind of hit a plateau or are hitting a plateau. So it's fascinating to see new architectures emerging
1:05:37Craig Smith:or new strategies emerging.
1:05:40Zuzanna Stamirowska:Yeah, okay, well, let's leave it there. Is there anything I didn't ask that you'd like listeners to know maybe where they can go and find BDH? Yeah,
1:05:55Craig Smith:thank you. Thank you so much for this conversation. It was great. And I mean, these were pretty deep questions. I really appreciate it. So, I mean, you can reach out to us. I mean, for me, LinkedIn is probably the easiest. Also, research at pathway.com. This is for all the kind of research types of questions. I mean, we're very, very happy to answer there. Please note that there is, the paper is available on GitHub. To digest it, it's pretty long. I mean, Craig, I think you've got it. It's very deep. It's very dense and very long. Because it has to, it covers a lot of intuitions, actually a lot of proofs as well.
1:06:38Craig Smith:because we're showing the link between transformer and the brain that goes through expressivity. So all of this goes through computer science. If you're to go computer science, that is not always the thing that's being the most, let's say, studied in the AI community. So I can strongly encourage something that we've seen working a lot very well with many people. Actually, by now, LLens are doing a phenomenal job of organizing this paper, especially the reason the most advanced reasoning ones. Once we ask them to properly read the paper with the proofs and then break it down, they really do an amazing job.
1:07:16Zuzanna Stamirowska:So I would strongly encourage you to use AI to kind of read about AI. Yeah. I read it on archive. It's the same paper, right?
1:07:28Craig Smith:It's the same paper, and the paper is publicly available, of course. And then, I mean, there are also other podcasts you know that are that way more technical with with other members of of our team so i would also uh i would strongly invite you to to listen to this if you if you have questions uh this one was great i mean thank you thank you so much
From the publisher
This episode is sponsored by tastytrade.
Trade stocks, options, futures, and crypto in one platform with low commissions and zero commission on stocks and crypto. Built for traders who think in probabilities, tastytrade offers advanced analytics, risk tools, and an AI-powered Search feature.
Learn more at https://tastytrade.com/
In this episode of the Eye on AI, Craig Smith speaks with Zuzanna Stamirowska about how Pathway is enabling AI systems to work with live, continuously updating data.
Most AI applications rely on static datasets that quickly become outdated. Pathway takes a different approach, allowing developers to build AI systems that process real-time data streams, keeping models, knowledge bases, and AI agents constantly up to date.
Craig and Zuzanna explore why real-time data may be critical for the next generation of LLM applications, RAG systems, and enterprise AI infrastructure, and what it takes to build AI that can operate in a constantly changing world.
Subscribe for more conversations with the researchers and builders shaping the future of AI.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) The Core Problem: Why Today's AI Lacks Memory
(03:16) Pathway's Mission to Bring Memory Into AI
(04:53) Zuzanna's Background in Complexity Science
(10:30) Why Transformers Reset Like "Groundhog Day"
(14:34) The Brain-Inspired Dragon Hatchling Architecture
(23:59) How the Network Learns and Builds Connections
(37:38) Performance vs Transformers on Language Tasks
(49:37) Productizing the Technology With NVIDIA and AWS
(54:23) Can Memory Solve AI Hallucinations?




