This AI Grows a Brain During Training (Pathway’s AI w/ Zuzanna Stamirowska)

6 Jan 2026 · 49 min · 20 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: The Neuron - This AI Grows a Brain During Training (Pathway’s AI w/ Zuzanna Stamirowska)

Podcast Overview

  • Title: The Neuron: AI Explained
  • Host: Grant Harvey and Corey Noles
  • Description: The Neuron covers the latest AI developments, trends, and research, providing digestible and informative insights into AI.

Episode Details

  • Episode Title: This AI Grows a Brain During Training (Pathway’s AI w/ Zuzanna Stamirowska)
  • Guest: Zuzanna Stamirowska, CEO & Cofounder of Pathway
  • Main Focus: Introduction of Pathway's Dragon Hatchling architecture (BDH) - a novel AI model that incorporates memory and reasoning akin to a human brain.

Key Topics Discussed The Limitations of Current AI Models

  • Groundhog Day Loop: Current AI models (like Transformers) lack memory and time awareness, essentially "waking up" without retaining prior knowledge.
  • Lack of Memory: Traditional language models do not remember past interactions, reducing their ability to learn over time.

Introduction of BDH Architecture

  • Brain-Like Structures: BDH employs brain-like neurons and synapses to create a model that can remember, adapt, and reason over time.
  • Temporal Reasoning and Continual Learning: The architecture allows for true temporal reasoning and the ability to learn continuously, improving adaptability.

Key Features of BDH

  • Emergent Structures: The model can develop complex structures spontaneously during training, resembling brain functionality.
  • Role of Memory: Memory is treated as an essential component linked to time, where connections between neurons strengthen based on shared interests or activities.
  • Adaptive Learning: BDH can “get bored” with repetitive tasks, prompting it to strengthen connections with relevant new information.

Advantages Over Transformers

  • Infinite Context: BDH can maintain context over longer periods, enhancing its reasoning capabilities.
  • Efficient Scaling: BDH can "glue" together models trained on different datasets, merging their knowledge effectively.
  • Improved Interpretability: Unlike Transformers, BDH allows for observable neural activity, akin to a CCTV inside the model, providing insights into its functioning.

Real-World Applications

  • Early Adopters: Organizations such as Formula 1, NATO, and the French Postal Service are exploring BDH's capabilities in real-world scenarios.
  • Model Deployment: Pathway aims to make BDH available to AWS customers, emphasizing the infrastructure's readiness for integration.

Philosophical Considerations and Future Directions

  • Path to AGI: The model is seen as a step towards Artificial General Intelligence (AGI), focusing on reasoning as the core function of intelligence.
  • Safety and Control: Discussions on the safety of AI models in terms of controllability and preventing unwanted learning are ongoing. Pathway is exploring methods to rollback systems to checkpoints if necessary.

Conclusion The episode presents a forward-looking perspective on AI, highlighting the potential of the BDH architecture to revolutionize how AI models learn and operate. As the technology continues to evolve, it raises questions about the implications of creating systems that think and reason like humans, fostering a deeper understanding of intelligence itself.

Additional Notes

  • Guest’s Background: Zuzanna shares her journey from studying game theory and complexity science to co-founding Pathway, underscoring the interdisciplinary nature of the work.
  • Cultural References: The name Dragon Hatchling is inspired by Terry Pratchett’s "The Color of Magic", adding a whimsical touch to the technical discussion.
  • Community Engagement: Listeners are encouraged to subscribe to the Neuron newsletter and follow Pathway’s research updates for more insights into AI advancements.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to AGI and Podcast

0:00 to 0:29

Learn about the podcast's focus on artificial general intelligence and the guest's expertise.

“We believe we're on a faster way to AGI.”

Zuzana Stamirowska's Journey in AI

1:02 to 5:40

Explore Zuzana's unique path from politics to AI and her passion for complexity science and game theory.

“And today we dig into what live AI really means, why Pathway is banking on it, and whether this could be the next major architectural leap in AI.”

The Role of Time in AI Development

5:40 to 7:40

Understand the significance of time in AI systems and how it impacts memory and intelligence.

“So you see, when you have a system, like the one that's growing and changing, right, you kind of need to have a notion of time.”

Limitations of Current AI Models

7:40 to 10:10

Discover the limitations of transformer models and the challenges faced in achieving true memory in AI.

“that actually kind of measures the benchmarks, the equivalence of, okay, like how the level of human tasks, let's say that the LLMs can do with, let's say, 50 % success rate.”

The Future of AI Beyond Transformers

10:10 to 13:30

Discuss the potential next steps for AI technology beyond the current transformer models.

“And having a brain, let's say, which is set once.”

Understanding 'Baby Dragon Hatchling'

13:30 to 14:01

Learn about the concept of 'Baby Dragon Hatchling' and its inspiration from literature.

“And I have a theory, but I want to hear the actual explanation and I'll tell you my theory.”

Dragons and AI: A Mythical Analogy

14:01 to 19:08

Learn how the concept of dragons is used to explain AI architectures.

“And I mean, this book is worth at least 10 business books, I think, that you could find out there for so many different reasons.”

Understanding Brain-Like Architectures in AI

19:09 to 28:00

Explore how AI models mimic brain functions through simple neural structures.

“And you always have to work with the hardware that you have, with the materials that are possible whenever we see big technological shifts.”

Rethinking AI Model Scaling

28:00 to 28:50

Explore the shift from scaling parameters to enhancing problem-solving capabilities in AI.

“This is not the game of scaling of more parameters and more data, because this is fundamentally not where the value is to come from.”

The Brain vs. AI: Memory and Efficiency

28:50 to 30:20

Discuss how AI and the human brain manage memory and efficiency differently.

“We're looking at this getting better at puzzle solving and reasoning and hopefully in as general way as possible to get it closer to the way that humans reason, work and ultimately innovate.”
Show all 20 chapters

Innovative AI Models: Gluing Brains Together

30:20 to 33:10

Learn about the novel approach of combining separately trained AI models.

“So effectively, you get to something that operation works like infinite context.”

Real-World Applications and Partnerships

33:10 to 34:50

Discover how leading organizations are exploring AI in real-world scenarios.

“And then, of course, the value will be allowing them to train together a bit more such that they create really like a one unity because they also create links between each other.”

Navigating Towards General Intelligence

34:50 to 37:40

Examine the journey towards general AI and the importance of reasoning.

“we have built for the dragon nest to make sure that we actually can feed data at low latency very efficiently and actually do cool things that are necessary once you put such models and live intelligence in production.”

Addressing Safety in AI Development

37:40 to 40:20

Discuss potential risks and safety measures as AI approaches human-like reasoning.

“On a philosophical level, let's go a different turn.”

Controlling AI Learning and Behavior

40:20 to 42:00

Learn about methods to manage what AI systems learn and how to revert unwanted changes.

“This is something we should get with AI, right?”

Understanding Information Spread in AI Models

42:00 to 43:26

Explore how information cascades can impact model performance and the ability to reverse them.

“And you have observability that traditional LLMs don't have, essentially, with your CCTV.”

The Future of AI in Space Exploration

43:26 to 44:39

Discuss the potential of AI and TPUs in advancing space technology and exploration.

“Some would argue that it can kind of generalize, but...”

AI and the Evolution of Civilization

44:39 to 45:55

Learn about the transformative potential of AI in shaping future societies and civilizations.

“And I see AI, in fact, not as an end in itself, but as a crucial tool that will help us lift a number of obstacles that will then allow civilization 2.0.”

Rapid Changes in AI Technology

45:55 to 46:45

Understand the unprecedented speed of advancements in AI technology and its implications.

“example about, oh, so what will happen in AI next year?”

Reflection on AI Progress and Predictions

46:45 to 47:25

Reflect on past predictions regarding AI developments and current understandings of its pace.

“I mean, you know it better than anyone because, you know, your job is to really get an understanding and translate it, right?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Corey Noles:We believe we're on a faster way to AGI.

0:02Grant Harvey:Whenever two neurons were interested by something, the connection between them becomes stronger. And this is memory. We actually saw the emergence of just this kind of brain appearing. We can actually glue two separately trained models together and they become one. I remember we all rushed into the office and I see the brain and it was like, whoa.

0:28Zuzanna Stamirowska:Welcome, humans, to the Neuron AI Explained podcast. I'm Corey Knowles, and joined, as always, by my partner in crime here, Grant Harvey. How are you, Grant?

0:37Corey Noles:Doing good. Doing good. Thanks for having me.

0:39Zuzanna Stamirowska:Of course. Of course. Quite the guest. You come every time. You just never stop showing up, right?

0:44Corey Noles:I just got to do something different, you know, and I'll switch it up.

0:47Zuzanna Stamirowska:Well, we have an incredibly fascinating guest today that we're both really excited about. Grant, you want to tell them about her?

0:54Corey Noles:So we have invited Zuzana Stamorovska, CEO of Pathway, one of the boldest challengers to the reigning transformer-based AI paradigm. And today we dig into what live AI really means, why Pathway is banking on it, and whether this could be the next major architectural leap in AI.

1:12Zuzanna Stamirowska:Zuzana, welcome to the Neuron. We're so excited. Hello.

1:15Grant Harvey:Hi, Corey. Hi, Grant. I mean, thank you so much for having me. Great to see you guys.

1:19Zuzanna Stamirowska:Well, I guess to get started, one of the first things that kind of stood out to us was how did you go from studying at a French school for politicians to complexity science and AI?

1:31Grant Harvey:So have you guys seen that movie, The Beautiful Mind? Yes, I love it. There's this scene where actually he gets a Nobel Prize and they're like, oh, those people who bring him pens, right? And I remember my dad always cried at that scene. He was like, for him, it was like the most beautiful romantic thing that, you know, it was so funny. I mean, a big guy always crying. And then, I mean, actually, when I was studying at Sion Spa, so that is that school for, you know, French kind of presidents or whatever. I actually went to Stockholm School of Economics and I got the chance to take a course in game theory.

2:02Grant Harvey:And I think, oh, game theory. Well, this sounds cool. You know, having seen John Nash in the movie.

2:09Corey Noles:Right.

2:10Grant Harvey:And all of that. and I actually took a course in game theory, and I remember I was sitting there, coming from a very different background than the other students, in a way, and I just saw all the results of the games without doing the math, and actually, the guy who was teaching it was sitting on the Nobel Prize Committee, so it was just an amazing course. It was just so beautiful. I became obsessed with it. I understood that, okay, that was like, I felt like fish in the water, finally, as if somebody, you know, finally showed me the real thing I should be doing that just felt so natural, and I said, okay, this is it, I mean, there is nothing else I can or I should be doing in my life.

2:45Grant Harvey:At the same time, I was training in the kind of management consulting because this is what folks do at Stockholm School of Economics. So, I mean, I got a lot of exposure to all of this. But then I knew, OK, OK, how do I make it happen? And I guess I was like lucky enough, you know, to actually I have met John Nash once. So that was. That's awesome. That is kind of cool.

3:05Corey Noles:Well, what was the context or how did that happen?

3:07Grant Harvey:And there was a conference in Lisbon and he was actually a speaker there. Oh, wow. Yeah. And then I actually went to, I had an option actually to go to Ecole Polytechnique for like my master's, et cetera. And yeah. And then, so I did my master's specializing in game theory on graphs. And game theory on graphs actually very quickly evolves into complexity science. Once you do it, I mean, we have, you know, small particles, big structures. It's more interesting, more fun if the structure keeps on changing. And then you try to play a game on like an infinitely changing structure that keeps on growing.

3:43Grant Harvey:I mean, this sounds tricky. It is.

3:47Corey Noles:What? Like it's hard?

3:50Grant Harvey:Yeah. But of course, thoughts, you know, for like a pretty long time, we're trying to crack it and kind of just bring it to some more universal levels of math.

4:00Corey Noles:Yeah.

4:00Grant Harvey:Especially in particle physics and this sort of stuff. I mean, at the end of the day, you have small particles bumping, like doing something between them, right? Sometimes in space, bumping into each other, sometimes having connections, like in a graph, or between neurons, and you kind of send things over. And then this gives rise to small folks doing something, you know, give rise to society, or like, I don't know, a big phenomenon, or intelligence, I mean, you name it. But somehow, once you get to the math, it starts to look somewhat similar. I mean, I think physicists and mathematicians would kill me for it to say that, you know, it's just all the same, same.

4:34Grant Harvey:But somehow you get those intuitions and some sort of like toolkits that help you to kind of put things at levels of abstraction that make it way less complex. I mean, you know, things that all of a sudden start to look like just a sphere. And you just say that it's like, you know, in infant dimensions, but who cares? It's like a sphere. so this is kind of yeah doing a lot of graphs and then playing games on it inventing games like this is just that you know i was getting a kick from it so i did my studies at the call politique this is like a very funny it's also a military school like engineering military this is where like for physicists point carré for example like was so you know pretty big physics names like that noble prizes go there pretty often um yeah and then then i actually did research there as well and then And then, yeah, then I worked at the Institute of Complex Systems of Paris, where, yeah, we're kind of, you know, trying to figure out how we get to global phenomena from very local small interactions between things.

5:37Corey Noles:And then how did you go there to founding Pathway?

5:40Grant Harvey:So you see, when you have a system, like the one that's growing and changing, right, you kind of need to have a notion of time. If things are to evolve and emerge, they kind of need time. If you look at, well, IT systems in general, then intelligence and kind of AI right now, especially, they're kind of deprived of the notion of time. Right. And this bugged us. So it's not just me, you know. I mean, I was fortunate enough to over the years, you know, to meet wonderful people like our CSO, Adrian Kosovsky. You know, the guy had the PhD at 20, right? It's like quantum physicist, mathematician, and theoretical computer scientists, you know, top level.

6:20Grant Harvey:Amazing. The guy is just crazy. Ian, who, you know, was at Google Brain. And then all of a sudden, all those guys, you know, are kind of jumping off the cliff, dropping 10 years to do this thing with me.

6:32Corey Noles:This was pretty cool. So was time the core insight or problem you were trying to solve, the three of you?

6:40Grant Harvey:Right now, all the models that we see that are out there are built on one type of architecture, one type of technology. And that was an absolute kind of algorithmic breakthrough. And this is a transformer. So the transformer was like fundamentally built for language. Funnily enough, one of the cultures of transformer actually was like the first check in pathway. But this technology is by definition deprived of the notion of time and memory. So Pathway right now is building the first post-transformer frontier model, which is tackling this fundamental problem of lack of memory in AI. Memory is linked to time, of course, because you need to remember things over time.

7:20Grant Harvey:You need to remember how you were thinking, how you were solving something, for example. You need to remember to see consequences, right? You need to remember to stay coherent while problem solving. The more you know, the longer you can stay focused on a task. I mean, this means memory that kind of requires time, right? And we kind of know right now there is this lab called meter that actually kind of measures the benchmarks, the equivalence of, okay, like how the level of human tasks, let's say that the LLMs can do with, let's say, 50 % success rate. And right now, the length of those tasks is at like 2 hours, 17 minutes for GPT-5.

8:00Grant Harvey:So, I mean, if you were, like, we could say that current LLMs are kind of reliving their Grand Hawk Day every day. So they don't have memory as such. The way it works is that they're trained ones with a lot of, a lot of, a lot of, a lot of data, right? To the point that by now we know we've exhausted all the data readily available on the internet for training. This is where they get their power from because these are like fundamental language models. So they actually managed to produce something new that they didn't necessarily see in the training data very explicitly, right, from having so many kind of samples of data.

8:37Corey Noles:And everything is like a relationship to everything else, right?

8:40Grant Harvey:Yeah, exactly. So they've seen so much. And these are language models, and they don't have memory.

8:48Zuzanna Stamirowska:And I guess what they're calling memory now is essentially something that's getting pushed into the system prompt each time you run a query. Is that correct?

8:57Grant Harvey:Yes, correct. So right now when you have an LLM and you kind of start it, it's trained once. We have these like huge models and they'll tell you how many parameters, right? And the bigger, the better. That would be usual, the kind of the vibe that they would be giving. And then every time you use it, it's as if the model was waking up and always using the same brain that was set during training. Training is very expensive because of data and compute because you need to produce a huge model. Then it works better, of course. you ask your question you add some context like i know your private document like some whatever you're asking about right uh then you get your answer but then you relive the ground hook day every time and you have a sort of memory cory exactly kind of as you said but it's more like as if you were leaving leaving sticky notes for yourself from the day that you know you

9:43Corey Noles:wouldn't remember yeah right or the movie memento where he's like tattooing it on his arm yeah i Exactly.

9:49Grant Harvey:So it's like, you know, there's a big difference between having a library of knowledge versus having internalized it. Yeah. Such as you actually create a framework based on this that you can kind of adapt to new situations. Okay. This is a very big difference of having like evolving memory, contextualized evolving memory. That's your own, right? Yeah. And having a brain, let's say, which is set once.

10:16Corey Noles:So before we get into what you're doing differently than transformers, based on that thesis or that problem set, let's say, do you think that there's a plateau on the meter chart where only transformer-based language models can only get to like X hours of consistent tasks? Do you think that there's like, and what hour mark would that be or day mark would that be?

10:40Grant Harvey:So I guess right now we're kind of, you know, pushing. In fact, it's not through LLMs per se that we are pushing and we would even have folks from OpenAI saying some things that the way forward is through reasoning, right? Yeah. So how far can we get with reasoning? So reasoning is reasoning. It's even less related to Transformer per se. so there I wouldn't put a bar necessarily but just I mean given the math like the memory is not there so it's difficult and kind of somehow tiresome to actually try to trick transformer into having memory so like what I like to talk about is like epicycles so you guys like before we had you know Copernicus and the proper theory of solar system people were observing the moon and to make sense of the observations or trying to kind of design some sort of orbit that would be maybe like this because that was the only way that they could explain the observations, right?

11:42Grant Harvey:It was like cumbersome, pretty ugly if you think about this, but then every time they got a bit better, they were getting, you know, like champagne or, you know, they would party. And the thing is, well, sometimes you just need to swap things kind of around and then the orbit is actually just in orbit, right? It kind of looks good. It starts to make sense

Read the full transcript

12:06Corey Noles:when you switch the perspective.

12:08Grant Harvey:Yeah, things kind of start to fall into place. So, I mean, yes, we believe that there are just, you know, some things that we need to roll back to. I mean, Transformer opens, it's an amazing, absolutely amazing innovation which opened the entire market and actually it's done two things. One is a technological innovation, right? Yeah. Scientific and technological innovation. Second, with the go-to market that happened, it managed to tickle the imaginations of everybody. And this is huge for scientific innovation. Just think about this. Oh, yeah. But we are still early in this kind of AI market shift.

12:48Grant Harvey:So, so far, I mean, 0.7 % of GDP was spent on this AI technological shift. If you compare it to other such shifts that, you know, like in the past century, I mean, just the telecom in the 90s took over 2 % of GDP to be accomplished. Wow. And I'd say that probably AI is more fundamental, right? Yeah, me too. So we're super early and well, Transformers most likely not. I mean, as Bradford would say, it's not the ultimate technology to get us all the way through it. And yeah, we need something else. There's a lot to be done.

13:29Corey Noles:All right. So we have to ask about the name Baby Dragon Hatchling. How did you come up with it? And I have a theory, but I want to hear the actual explanation and I'll tell you my theory.

13:39Grant Harvey:This is like Dragon Hatchling. So in the paper, like specifically, it's Dragon Hatchling and the abbreviation is BDH. And of course, the question is where does the B come from? Dragon Hatchling per se is coming from Terry Pratchett's The Color of Magic.

13:55Corey Noles:I love it. Okay, cool. I have not read that one, actually. I need to.

14:01Grant Harvey:I strongly recommend it. And I mean, this book is worth at least 10 business books, I think, that you could find out there for so many different reasons. Dragons being probably even a smaller one. But I mean, specifically in The Color of Magic, there are dragons that appear more the more you think about them. And it was actually very funny for, I mean, we found it funny for reasoning models, you know, because we literally had to reason about reasoning to build a reasoning architecture. Actually. So the dragon started to appear. Yeah, yeah, exactly. And well, it's like what you see publicly, right, is an architecture in the paper.

14:46Grant Harvey:So this is a hatchling and this is it.

14:48Zuzanna Stamirowska:I love it.

14:49Grant Harvey:And then, yeah, we do get some questions about why BDH. And the truth is, and everybody tries to put something that would naturally fit the B. The very simple truth is that, I mean, I just thought that AI dudes really like three-letter acronyms. I agree. They're also kind of...

15:08Corey Noles:You're not wrong.

15:09Grant Harvey:Easy to pronounce and it worked. But I had one person, a physicist, who came to our office and he said, listen, I read your entire paper. I read everything. And because I really... I think I still need to read the appendix because I still don't know where the B is coming from. And you had to explain it to yourself. I'm like, well played, well played. Oh, I love it.

15:32Corey Noles:Is it, so is the B because it's a small version and you're going to grow it?

15:38Grant Harvey:So to be perfectly honest, it's like the most truthful explanation is really just the three-letter acronyms. Yeah, oh, gotcha. The B per se comes from the fact that like the model working on, like the working name is Baby Dragon. Yeah. So architecture is Dragon Hatchling and then you have Baby Dragon because it's already, you know, somehow grown. And it was just inherited, like, it inherited the bee because our internal name was Baby Dragon. And then, you know, we do have some dragons flying around the lab, so... I love it. We even have a random name generation for dragons. Oh, that's amazing.

16:12Corey Noles:Like, in Dracoric? Like, how nerdy are we getting here? Like...

16:17Grant Harvey:Oh, no, no, no. I, like, we literally have an LLM dude because we have versions whenever you have you have versions you're getting you two no test model like one thing against the other and stuff so we have we literally have a random dragon names generator that's cool I love it yeah my my theory uh was

16:33Corey Noles:that if this is truly continual learning it's kind of a dragon in a sense because it could be very powerful and dangerous if we're not careful but I imagine it's more like a dragon in Game of thrones where they're controllable um so i guess the question is you know one we'd love to know how how it works and and two you know if it is continual learning how do you control something

16:57Grant Harvey:that's potentially powerful you actually hit something here as well because part of why we venture towards dragons at all is that well this is a mythical creature right yeah uh that nobody believed it could exist this is where okay we're in the business of building dragons uh so i think It was a realization we had very early on. And indeed, we're talking about continual learning, in time, long horizon reasoning, and adaptation over time, right? To new data, new learnings, et cetera. So the way it works, it's actually a bit like a brain that works on silicon. So a bit like a brain, I mean, there's a concept, this is very nerdy, the concept called Hebbian Learning.

17:44Grant Harvey:It's like a very simple principles of how the brain works. So I guess you all can imagine that the brain has neurons, which are like little cells, right? And then connections between neurons, which are synapses. And this is kind of it. So we have kind of dots or like neurons and then links between the neurons. And this is like a big network that we have in our heads. And this is a very simple model because, you know, somebody who's like a neuroscientist or a biologist would say, well, yeah, but you have all those chemical reactions, all this and that. I mean, here we're really just looking at the very stripped down, very basic, very simple, you know, structure.

18:18Grant Harvey:Like, we know that birds fly and they have wings, right? Yeah. We don't need to know how they've really, like, moved, how many bones they have in them and stuff. No, it's just literally something to fly in to have, like, the surface of wings. And there we go. Right. So we're looking at those kind of little neurons, almost like particles, that entities, and they have links between each other. and they're actually passing signals between them. Oh, wow. So this is why you have connections, you have this structure. This structure, we know, has to be dramatically efficient. Right. Why? Why? Because, well, our heads are somewhat limited in space.

19:00Grant Harvey:We walk on two feet and we kind of fall over, so our brains kind of get larger.

19:05Corey Noles:Yeah. Right.

19:07Grant Harvey:So it has to be very efficient. we know it is very efficient in terms of power but it does offer this kind of capabilities of lifelong learning keeping kind of very like infinite context pretty much so we know that there exists a physical system that is capable of doing those kind of dragon like things it's not fully impossible so this we know question is how to make it work and especially how to make it work on the hardware that we have right now. Right. And you always have to work with the hardware that you have, with the materials that are possible whenever we see big technological shifts. I mean, it's usually some sort of inflection points where many things come together.

19:54Grant Harvey:I mean, so much compute with this algorithm, all of a sudden this gives us a boom. So what we did is we looked a little bit at Transformer and thought like, okay what is it really missing to like from from the brain right like to to to get closer to the brain and then and then yeah that was actually adrian you know our um our chief scientific officer who who went who went on this journey like literally like with very strong conviction that it has to be local interactions looking at the brain i mean we have to have those like small particles and our model like our architecture bdh the way it works is that you really have small neurons neurons are connected whenever you have a new bit of information like as folks call it tokens but you have some new information coming in only the neurons that are interested and connected light up okay so one neuron gets information passes it on to its neighbors those with whom he's connected not everybody not everybody lights up just the neighbors if they care enough about this thing, they light up as well.

21:04Grant Harvey:So this is the principle of neurons, you know, who are connected to them. They fire together.

21:11Zuzanna Stamirowska:Does that vary based on how connected it is? For example, the idea being that some information would be more tightly tied to point X?

21:22Grant Harvey:Yeah. Okay. Yeah. And actually, but this structure emerges naturally. We don't set it. Wow. It just comes from data. It emerges. We actually saw in our lab, this was an amazing moment. We actually saw the emergence of just this kind of brain appearing.

21:38Zuzanna Stamirowska:It was like, whoa.

21:39Grant Harvey:I remember we all, it was like late in the evening, and we all just rushed into the office because we were doing something else. And then Adrian just calls us, hey, look at this. And then I see the brain. And he's like, whoa. And I remember that's my brother.

21:53Corey Noles:I almost have chills thinking about that. It's just like learning that on its own.

21:57Grant Harvey:I remember I immediately texted my brothers and investors and I, look at this.

22:03Zuzanna Stamirowska:Oh, my God. Yeah. That's awesome.

22:06Grant Harvey:It was huge. I mean, emergence also for complexity scientists, right? Emergence is what we love to see. This is a spontaneous order. You think that something's just so random doing, the gods know what, but you can strip it down to such very simple fundamental rules that, yeah, you will see this larger order appearing. And this is what we got, like the structure of the brain somehow appearing naturally from those very local, honestly, message passing between neurons. As we do it on social networks, for example, we say something to our friends, right? Imagine this rumor spreading dynamics. This is how kind of learning works here.

22:45Grant Harvey:So whenever two neurons were interested by something, the connection between them becomes stronger. and this is memory yeah

22:58Corey Noles:that's right because that's sort of like how the hippocampus works right where it's like I'm going to do a terrible job

23:06Grant Harvey:you use it more it becomes stronger I mean this is just a principle and it's only positive activation so there's no like positive and negative it's only positive it gets stronger something is not used you know over time it will start fading um but but journey speaking you the connections that were useful become stronger and and this is kind of it and then i mean this structure is you know it's actually very efficient because it's like a brain so it's computationally efficient it distributes nicely it gives so many nice properties that unlock a number of things you know that then for for us even from the engineering standpoint in terms of how it scales how it distributes how you can run it on many machines, etc.

23:51Grant Harvey:But it's like a scale-free graph structure. So point is, even if we go beyond the scales that we've seen in data and tests, we scientifically know how it will behave. It's very different from the transformer, at least as we see it now, because before transformer it hasn't been studied, it would be difficult to study. For this, because we know how the emergence works, we know that, I mean, Yeah, it's scale-free, same laws we'll be holding, you know, above what we've seen in tests and kind of data until now.

24:28Corey Noles:Does that mean that it's also more interpretable at some level? Like you can kind of understand what it's going to do or no?

24:35Grant Harvey:Yes, in a way. So specifically, we do see very precisely neural activity. And because we see the neurons when they care about something, right? We just see them.

24:45Zuzanna Stamirowska:Yeah, right, right.

24:45Grant Harvey:So like for LLMs right now, for transformers, I've also tried to build MRI machines to scan the brain, whereas we sort of have a CCTV inside of the brain.

24:56Zuzanna Stamirowska:I was just going to ask you about that. I remembered you making the CCTV analogy when we talked recently. What does it look like?

25:04Corey Noles:Literally, we see the neurons that fire up when they fire up on something.

25:08Grant Harvey:So in the paper, we're showing things that fire, like synapses and neurons that fire up for the notion of currency. And you have like a dollar one, right? That you fire up, firing up. Of course, there is two scientists who may be listening. As an end of the system, there is compression of information while learning. So, you know, it's not always super clear. Like some concepts may be fairly large and you will have like a lot of things firing up, for example. But dream speaking, we see this neural activity and we see even more. We see neurons getting bored. so when you keep on repeating something to them uh we just see their activity yeah whatever uh just going down in learning like think about this especially as you go get older i mean kids get learn very quickly but also because everything is new to them everything is a surprise it's worth learning um with as we get older like things you know we stop thinking about many things We stop noticing them because they're just so obvious.

26:11Grant Harvey:I mean, our connections are strong enough, let's say, for that one thing, like eating soap or whatever. We shouldn't be eating soap. Bad idea. So there's this element of surprise that actually somehow shows that something is valuable and worth remembering. So it was actually for us very funny to see this surprise effect literally on your activity.

26:32Zuzanna Stamirowska:So will it, in the same way that the brain over time, if there are areas that are not being used that they can weaken, will the same thing happen in a model?

26:46Corey Noles:And if that's a dumb question, please say so.

26:49Grant Harvey:No, no. So, I mean, yeah, actually, so you'd be getting some sort of like fading connections that they're not used for very long. But this is more a topic of, okay, how to transfer also to long-term memory, right? Yeah. So, yes, because there are some things that, again, it's not to work like a database, right? Okay. If there's a database, then in deployment, this is something you plug in. Right. If you want to store absolutely everything forever, right? This is less of a problem. But for reasoning and having, let's say, your space to explore when you reason, you want to build it in such a way that you have the most relevant and compact structures.

27:36Corey Noles:Well, then I guess what I want to know now is, So, you know, you've proven this BDH works at GBT2 scale, last I read, with 1 billion parameters. Is that correct? What's the path to scaling it to, say, 100 billion parameters? What needs to happen to get there or grow larger?

27:54Grant Harvey:Of course, first of all, we do it. There are no reasons actually for it not to scale. And like scaling laws are inherited from the transformer. But there's also no big need to scale. This is not the game of scaling of more parameters and more data, because this is fundamentally not where the value is to come from. The value is to come from faster learning how to solve problems that haven't been seen in the training data. Like, this is what we, this is where we want to get to. And actually, if we can show better learning out of smaller data, well, this is the kind of value that we want to prove.

28:37Grant Harvey:So actually, I hope that very quickly, you know, we'll be more looking at models that are very small but capable of producing results comparable to the big ones.

28:48Corey Noles:Love that.

28:49Grant Harvey:That's awesome. We're not looking at scale and root for scaling. We're looking at this getting better at puzzle solving and reasoning and hopefully in as general way as possible to get it closer to the way that humans reason, work and ultimately innovate. Because if you look at the real innovator, like the best ones that I know because I kind of have them on the team, right? It's not about seeing what's there, but it's seeing what's not there and what could be there.

29:20Corey Noles:Let's say, for example, it's learning something really complicated, like where it's, I mean, you're a complexity scientist. You know more about this than I do, but something like really complicated. does it at some point run out of brain power or how like because i'm used to thinking of you know parameters as this thing that's like oh this is like it can retain a lot more information like and if this thing is just continuously running at what point does it reach its limit of what it can think of or do we just not know that we don't know but i don't think those limits

29:50Grant Harvey:work in this way i don't think i mean right now we do cap the number of neurons that's like the problem number that you may get there you may imagine models where you would be adding them this is all right you could have a model that's that's growing technically um but but it's like right now in transformers it's not the reasoning power doesn't come so much from this size per se okay when you think about this we actually do have a lot of compute power because of how much place we have in the brain because of the structure so these models but if you if you look like the brain now i would need to check again but i mean for the synaptic synaptic connections the brain are in the trillions right and i think it could be yeah like one thousand one some folks would say could be having one thousand trillions or something like this of synaptic connections in the brain, this gives you a lot of memory and a very efficient structure.

30:53Grant Harvey:So effectively, you get to something that operation works like infinite context. I mean, to be super scientifically precise, yes, context and BDH is limited by the size of your brain. So the number of neurons and then connections between them. But this network structure allows you to encode so much. Well, even the human brain fits in a cereal bowl.

31:17Zuzanna Stamirowska:There you go. It's really not that big considering the amount of information it can store and its speed and its capacity, I would think.

31:26Grant Harvey:Exactly. And then when you think about this, you keep your memory close to the core. Actually, exactly at the core. So you don't need to do look-ups for technical people. You don't spend your energy on all of that. You don't need additional compute for this. So it becomes very efficient from this point of view. You have memory directly there. and yeah, it's like in memory on a chip. Okay. Wow, yeah. And then second thing is that you don't fire up the entire model every time, but you fire up only those guys who are connected and who care. Right, that makes sense. And for most cases, yeah, you don't fire up the entire brain.

32:04Grant Harvey:I guess when you fire up the entire, like I don't know how much brain people are using, but I mean, small percentage, right? Yeah. At any time, but that's exactly the point. Why would you be using the full thing? Not for every task, right?

32:17Zuzanna Stamirowska:I don't need that to pick up a soda or to, you know, answer the door and say hello. You know, there's definitely lightweight tasks for sure.

32:26Grant Harvey:Exactly. I mean, you may have some that are, you know, more specialized in something. Cool thing, however, like with our heads, what we're not capable of doing is that we cannot exactly glue two brains together. It would be very difficult. And with those models right now, since like they scale only with a number of neurons, So there's just this one dimension to something we're showing in the paper. We can actually glue two separately trained models together fairly easily. And they become one. So we showed in the paper, we have like one model trained in one language, the other one trained in another language.

33:02Grant Harvey:We just put them together, even without leaving them, you know, to train a bit. They actually start producing sentences that mix up the two languages pretty well. And then, of course, the value will be allowing them to train together a bit more such that they create really like a one unity because they also create links between each other. But yeah, this gluing is a bit like Lego blocks that you can put together. So you could imagine a model trained in finance, another one in legal, you know, in an enterprise. And then you put them together, you get this like super, well, I don't know what the person with, you know, good legal and finance background would be doing.

33:40Grant Harvey:Maybe we shouldn't venture there. Right. Too much power. I guess some superpowers that, I mean, we would definitely want. It's like, you know, us and having Adrian. It's a different story when you have, if I were to have a physicist, a mathematician, and a computer scientist on a team. Sure, cool, but still communication, different intuitions, and all of that. Versus having one person who kind of has all of this in one efficient structure, which is the brain.

34:08Zuzanna Stamirowska:Something I'd like to ask, and this will shift gears a little bit, but I know you already have some early adopters and people who are working with this. Like, I understand NATO, the French Postal Service, and maybe Formula One. Is that correct?

34:21Grant Harvey:So, these are like, we have history of actually working with pretty amazing accounts.

34:26Zuzanna Stamirowska:Wow.

34:28Grant Harvey:But as of now, they don't have the models deployed. They have the Dragon's Nest, if you wish. so I mean you have to understand that there is if you are to bring life intelligence you'd have life data as well they need the way to connect and you need your dragon to feel cozy you know in its environment so those guys I mean as as of now they're using like layers of technology that we have built for the dragon nest to make sure that we actually can feed data at low latency very efficiently and actually do cool things that are necessary once you put such models and live intelligence in production.

35:09Grant Harvey:Yeah. And yes, this is like, of course, cases are pretty cool, right? I mean, sometimes I joke that, I mean, how are we exactly getting only the coolest customers? Like, do we have a cool, cool factor and kind of regeneration? There's a lot of cool there.

35:22Corey Noles:It's the dragon factor.

35:23Zuzanna Stamirowska:I can't think of anyone with more data than what Formula One collects. it's unreal and definitely the cool factor

35:36Grant Harvey:it's really cool for them it's like from the strategy of the race you have so many things that can go wrong every car is like a prototype it can blow up of course NATO I can't talk about this freely but I mean this is very fundamental I imagine that there is if you played any sort of war game like I know harpoon or whatever and just imagine getting real-time information from the field and then kind of informing the strategy and stuff. So this is all cool. And then, you know, boys and girls like buses, ships.

36:13Zuzanna Stamirowska:Oh.

36:13Grant Harvey:Yeah. All of that. And, I mean, you know, we have these complex systems of moving parts and, like, this is how the world works, really, right? We have a lot of moving, living things that are interconnected it and ideally would like to predict what's happening out of it and somehow see the patterns in this chaos, hopefully be able to control it. And yeah, for this, you kind of need a number of technologies that come together. And we do have a nest. That's the point. Our baby dragons have nests.

36:47Zuzanna Stamirowska:That's awesome.

36:48Corey Noles:What's the roadmap then? And, you know, are we thinking we're going to be Lego blocking a bunch of different models together? I mean, as far as the company goes, like, where do you see this going in production?

37:01Grant Harvey:Yeah, so actually, just like a week ago, we announced a partnership with NVIDIA and AWS at ReInvent. Yes. So point is, the moment we're ready, it's going to be available to AWS customers kind of right away, all built on the infrastructure and kind of made in such a way that it will be easy to plug in and test it and adopt it. So this is one. I mean, this should happen sometime next year. So this is it. And then, I mean, we have our own roadmap. I mean, we believe we're on a faster way to AGI and working towards getting there as fast as possible and keeping all the focus on this.

37:42Zuzanna Stamirowska:On a philosophical level, let's go a different turn. Do you see this approach as, and you just mentioned AGI, as a step toward general intelligence as a whole? Are there any safeguards or things along those lines that you're thinking you might need, as it begins to think more like a brain, operate like a brain?

38:05Grant Harvey:First of all, I think it's important to say that we're looking at reasoning and we see reasoning as a primary function of intelligence. LLM says we've seen them for like literally language models, right? As we've seen them for child code use cases, summarization, some sort of like search. I mean, these are great use cases and brought us here, but they are not the primary function of intelligence. And we would say that reasoning is the primary function of intelligence. This is not just us, actually. I think by now, everybody from all the labs would tell you that reasoning is this. Yes. So we're looking at reasoning, then with reasoning capability of solving and inventing very tough problems.

38:50Grant Harvey:So the North Star is to get to an innovator who sees what's not there, as opposed to just seeing what was there and recomposing. Right. Yeah. So that would be the true generalization. And then you have generalization over time. So going towards solving problems in environments and over like periods of time that were not tested, that were not seen in data, etc. In terms of safety, I think I'm not yet there to give very definite answers. But I'd say that part of what we're doing is actually gaining a way better scientific understanding of how the models are working and why. For us, mapping, like very precisely mapping, understanding this passage from micro interactions and having something that we know we call the equations of reasoning.

39:39Grant Harvey:to having a scape-free structure of known laws that govern it, this is kind of important and it's important for safety. Then some discussions that we have, let's say more internally, are around getting to provable risk levels of how such system will behave and if it will be kind of do-what-I-mean-machine or not. because so the risk the risks that are we're looking at that are you know more controllable maybe for us is just making sure that yeah these models will not just out or just by themselves won't venture into doing something just completely silly because they're hallucinating right something so off policy and it will be just ridiculous because they would stop working like a predictable human

40:31Zuzanna Stamirowska:yeah because if you think

40:33Grant Harvey:about this if you're hiring someone you observe them for what couple of hours max yeah you see you kind of see their credentials you know how they were kind of taught at Princeton probably same curriculum

40:48Corey Noles:you know

40:49Grant Harvey:you get them and you give them tasks and you safely assume that they will perform in a certain way right without getting completely crazy and blowing up the planet earth yeah And this is a fair assumption. This is something we should get with AI, right? Agreed. So at least mathematically, we should have some sort of, you know, comfort that, okay, we know how those Princeton graduates function. And then, however, we still don't control the objectives, right? So if somebody puts AI to a bad objective, I mean, this is something that, I mean, as a startup for now, I mean, I have no influence over.

41:33Grant Harvey:this is definitely something that as we get closer to HDI we'll have to be resolved at many levels and I'm sure we won't be the only voice to take part in this discussion, I hope we'll actually have something to say

41:48Corey Noles:How do you prevent it from learning something you don't want it to learn? And maybe you're not at the point where you can't do that yet

41:55Grant Harvey:No, this is actually not too bad I think the easiest is that you can roll back to a checkpoint Nice. So this is fairly simple.

42:04Zuzanna Stamirowska:And you have observability that traditional LLMs don't have, essentially, with your CCTV.

42:10Grant Harvey:Knowing that physicists, I mean, it looks like a cascade. Imagine like an epidemic spreading in a graph, right? So you have a system, you have an epidemic spreading. This is how information is out of spreading, right? So if you have a very small cascade, very small epidemics, you probably don't care because it wasn't relevant enough for the model. potentially. But if it's small, you can still reverse it.

42:36Corey Noles:You can literally reverse it. You could quarantine it, essentially. But if it's big,

42:42Grant Harvey:this is when also the information is very relevant and you get to a point when you wouldn't be even able to say where this came from if you were somewhere in this graph looking around, you wouldn't be able to say where it came from. Then it's non-reversible per se. Just from a physics point of view. So, yeah, but you roll back to checkpoint and you could just checkpoint your kind of a little over time. So, you know, if you infuse something you didn't want to have in the data and the model, you take it out.

43:13Zuzanna Stamirowska:What is the one AI capability that you're most excited to unlock with BDH that isn't possible now?

43:22Corey Noles:With the current architecture.

43:23Zuzanna Stamirowska:Yeah. Yeah, with Transformers specifically.

43:26Grant Harvey:This is very easy, like generalization. Yeah. But I mean, for me... True generalization. True generalization.

43:33Corey Noles:Some would argue that it can kind of generalize, but...

43:36Grant Harvey:No, true generalization. So getting to innovator level. I mean, why? I mean, I think there is a path towards space exploration for us.

43:47Corey Noles:As in like your models in space?

43:49Grant Harvey:No, I mean, although I do know a guy who's kind of working TPUs in space. This is so cool.

43:56Zuzanna Stamirowska:That is cool.

43:57Corey Noles:That is cool.

43:58Grant Harvey:Yeah.

44:01Corey Noles:You mean the space more abstractly?

44:04Grant Harvey:Yeah, I mean TPUs in space, like where we'll be building data centers.

44:08Zuzanna Stamirowska:Hey, in space.

44:09Grant Harvey:That's a big discussion lately. Yeah, it's a very serious thing, right? And I think timelines are pretty short. And within two years, they assume they will have actually TPUs in space. Wow. So, yeah, but I think to just make space travel actually work, we will need so many scientific unlocks and especially on the energy front. So there's a couple of very fundamental problems in technology and science that we need to crack. And I see AI, in fact, not as an end in itself, but as a crucial tool that will help us lift a number of obstacles that will then allow civilization 2.0. So when I talk to friends, you can compare this transition to maybe the moment when people started agriculture.

45:05Grant Harvey:you know it's like before we're kind of like hurting and moving from places to place and then kind of humanity settled down and then i was able to build culture and civilization uh so like ai the moment we becomes you know everywhere and this powerful will allow us to build civilization

45:25Zuzanna Stamirowska:2.0 you know and we spend a lot of time talking about the next five years the next 10 years the next 20 and we forget that like time's going to go a long way what's going to happen a hundred years a thousand years five thousand years where will where will humanity be as a result of what we're experiencing right now at that point and that's uh i can ramble on that for days but look at the

45:52Grant Harvey:pace yeah sure look at the pace right yeah just i remember this time last year folks asking me for example about, oh, so what will happen in AI next year? And I was saying, well, reasoning models are going to be everywhere and they're going to be the main thing. And people by then didn't exactly know what reasoning models were. Yeah.

46:11Zuzanna Stamirowska:Correct. They look at you like you're crazy when you say it, don't they?

46:14Grant Harvey:Yeah. I mean, I'm happy my predictions work, but thing is just that this space right now, the speed of transition, let's say of shift is like unprecedented. In hardware, I know Railway, that was a huge one that changed the world. I mean, you have huge infrastructure projects that take just so much time to be built, right? So, I mean, it's like so, so slow. Whereas here, it's like from month to month, the landscape is different. I mean, you know it better than anyone because, you know, your job is to really get an understanding and translate it, right? Yeah.

46:50Corey Noles:So hard.

46:51Grant Harvey:Yeah.

46:53Zuzanna Stamirowska:So many times there's a thing that we think of as being way in the past, and then we go and we look at something and we're like, well, that was eight weeks. That was eight weeks to go. You didn't have those two months ago?

47:08Grant Harvey:At my PhD defense, I had a guy from United Nations who was there, and he actually commented and gave me a book entitled that every time I found the meaning of life, they change it.

47:25Zuzanna Stamirowska:So true. Well, Susanna, thank you so much for joining us today. We were both so excited to meet you and look forward to watching your tech grow for many, many years to come. If our viewers want to learn more, how do they go find Pathway?

47:41Grant Harvey:Yeah, so, I mean, please go to our website, pathway.com, and then you can connect with me on Twitter and LinkedIn. And then, yeah, I mean, please follow our research papers. This is what we do. We'll probably also kind of start some sort of, like, blog to get some of these messages kind of closer and maybe take them, you know, out of the paper. Yeah. And, like, paper is pretty long and deep. But, yeah.

48:07Zuzanna Stamirowska:This is it. And guys, thank you so much.

48:08Grant Harvey:This was so much fun.

48:09Zuzanna Stamirowska:Excellent. Thank you for having me. Well, everyone, that wraps up another episode. If you haven't yet, please take just a moment to like and subscribe today's video. That way we can continue bringing you more cutting edge interviews from some of the most fascinating people at the frontier of AI today. Also, don't forget to check out the Neuron newsletter that started all of this and join 600 ,000 or so other people who read it every morning. Thank you again to Zuzanna and Pathway. And that's it for today. Farewell for now, humans.

From the publisher

Imagine an AI that doesn’t just output answers — it remembers, adapts, and reasons over time like a living system. In this episode of The Neuron, Corey Noles and Grant Harvey sit down with Zuzanna Stamirowska, CEO & Cofounder of Pathway, to break down the world’s first post-Transformer frontier model: BDH — the Dragon Hatchling architecture.


Zuzanna explains why current language models are stuck in a “Groundhog Day” loop — waking up with no memory — and how Pathway’s architecture introduces true temporal reasoning and continual learning.


We explore:

• Why Transformers lack real memory and time awareness

• How BDH uses brain-like neurons, synapses, and emergent structure

• How models can “get bored,” adapt, and strengthen connections

• Why Pathway sees reasoning — not language — as the core of intelligence

• How BDH enables infinite context, live learning, and interpretability

• Why gluing two trained models together actually works in BDH

• The path to AGI through generalization, not scaling

• Real-world early adopters (Formula 1, NATO, French Postal Service)

• Safety, reversibility, checkpointing, and building predictable behavior

• Why this architecture could power the next era of scientific innovation


From brain-inspired message passing to emergent neural structures that literally appear during training, this is one of the most ambitious rethinks of AI architecture since Transformers themselves.


If you want a window into what comes after LLMs, this interview is essential.


Subscribe to The Neuron newsletter for more interviews with the leaders shaping the future of work and AI: https://theneuron.ai

More from The Neuron: AI Explained

All 106 episodes
This AI Grows a Brain During Training (Pathway’s AI w/ Zuzanna Stamirowska)The Neuron: AI Explained · 49 min
Listen in VO