#302 Karl Friston: How the Free Energy Principle Could Rewrite AI

19 Nov 2025 · 1 h 3 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye on A.I. Podcast Episode Notes

Episode Title

#302 Karl Friston: How the Free Energy Principle Could Rewrite AI

Host

Craig S. Smith

Episode Description In this episode, Craig Smith discusses the Free Energy Principle (FEP) with neuroscientist Karl Friston, exploring its implications for artificial intelligence (AI) and how it may serve as a blueprint for future AI architectures. The conversation touches on the limitations of transformer-based systems and the potential of new architectures like Axiom, which are inspired by brain-like learning processes.

---

Key Concepts and Themes

  1. Free Energy Principle (FEP)
  2. Definition: A theoretical framework that explains how living systems, including the brain, maintain order in the face of chaos.
  3. Core Idea: The brain predicts sensory inputs, and when there is a discrepancy between prediction and reality (prediction error), it updates its models. This process is seen as a continuous loop of inference rather than static learning.
  1. Neural Architecture and Active Inference
  2. Active Inference: A method of learning where agents update beliefs based on the minimization of prediction errors.
  3. Axiom Architecture: A new AI architecture by Verses AI, which employs FEP to create agents that learn from the environment much like the human brain, focusing on minimizing surprise rather than merely predicting outcomes.
  1. Limitations of Transformer-based Systems
  2. Scaling Issues: Transformers face reliability and scaling limits with only incremental improvements in performance.
  3. Inefficiency: Current generative AI models struggle with issues like hallucinations and overconfidence in predictions.
  1. Continuous Learning and Catastrophic Forgetting
  2. Continual Learning: Axiom's architecture allows for continuous learning and adaptation without suffering from catastrophic forgetting, a common issue in traditional neural networks.
  3. Model Growth: Models can dynamically grow and adapt based on new information and experiences without losing prior learning.
  1. Role of Uncertainty in AI Systems
  2. Explicit Uncertainty: Unlike traditional models, Axiom integrates uncertainty into its learning framework, enhancing reliability and reducing overconfident predictions.
  3. Bayesian Belief Updating: The framework utilizes Bayesian principles to refine beliefs about the world based on new sensory information.

---

Discussion Highlights

Neuroscientific Insights

  • Cognitive Processes: The brain continuously makes inferences and updates its understanding of the world, which is crucial in understanding mental disorders as failures of inference.
  • Biomimetic Principles: Axiom aims to replicate these biological processes within AI models to enhance their efficiency and effectiveness.

Practical Applications

  • Real-world Implementations: Axiom's principles are already being tested in logistics, robotics, and autonomous agents. For example, optimizing cab driver placements demonstrates improved efficiency through active inference.
  • Robotics and Autonomy: Friston discusses the potential for robots to operate effectively in dynamic environments, showcasing the practical implications of these AI principles.

Future Directions

  • Integration of Language Models: The episode discusses whether active inference models could eventually learn language and understand meaning through embodied experiences, similar to human learning.
  • Emerging Research Community: The conversation indicates that while FEP and active inference are established in academia, their translation into commercial applications is still in early stages.

---

Conclusion This episode provides a deep dive into the intersection of neuroscience and artificial intelligence through the lens of Karl Friston's Free Energy Principle. It highlights the challenges of current AI systems, outlines promising alternatives, and emphasizes the need for AI architectures that closely resemble biological processes for true innovation in the field.

---

Additional Resources

  • Verses AI: [Verses AI Website](https://verses.ai/)
  • AGNTCY: [AGNTCY Website](https://agntcy.org/)
  • Follow Craig Smith: [Craig Smith on X](https://x.com/craigss)
  • Follow Eye on A.I.: [Eye on A.I. on X](https://x.com/EyeOn_AI)

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We change our mind 100 milliseconds by 100 milliseconds as we engage with the world and we continue making sense making sense making sense making sense making sense. And that's a process of inference. It's not necessarily learning, which will be a much slower process. And if the way that our brains work can be cast as inference, that means if our brains don't work properly, as in mental disorders, then we can understand that as false inference. The brain is in the game of basically trying to predict what it's sensing. So if there's a prediction error, that's newsworthy. So it can now use the prediction error to update, to revise its explanation, its representations of the latent causes.

0:39So in the moment, I literally change my mind. I literally change my mind on the basis of the prediction errors that have been elicited by comparing what I predict and what I actually saw. Hi. The transformer algorithm created by a team at Google in 2017 has been the core of generative AI. But scaling transformers has hit practical theoretical limits. New releases are providing only incremental improvement, and as a result, many researchers are looking into new architectures, aiming for better memory use, better generalization, and longer context reasoning. Among those researchers is the team at Versus, a startup implementing the ideas of Carl Friston, who is the most highly cited neuroscientist globally and one of the most cited living scientists overall.

1:40Tristan is the creator of the Free Energy Principle, a sweeping attempt to explain how all living systems, including brains, maintain order in a chaotic world. His work has inspired entire branches of computational neuroscience and new directions in artificial intelligence. Versus, where he serves as chief scientist, was started by some of his students. Its core project is called Axiom, a groundbreaking AI architecture built on the principles of Friston's work, taking inspiration from how the human brain predicts and adapts to its environment, not by memorizing data like today's models, but by constantly minimizing what Friston called surprise, which is a measure of the difference between prediction and reality.

2:37It's an ambitious attempt to create machines that learn and reason the way living systems do, grounding intelligence in physics rather than statistics. In fact, this is my second conversation with Friston because I was completely lost the first time around, and he says himself that he has trouble speaking in the vernacular. So while some listeners will be able to understand what he's saying, most are better off letting the jargon wash over them, focusing instead on the higher concepts and how they're implemented in an AI model. You may not understand everything, but you'll come away with a picture of how and why post-transformer architectures are being developed.

3:26I hope you find the conversation as fascinating as I did. But first, I want to give a shout out to our sponsor. Build the future of multi-agent software with Agency. That's A-G-N-T-C-Y. Now an open source Linux Foundation project, Agency is building the Internet of Agents, a collaborative layer where AI agents can discover, connect, and work across any framework. All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting. Agency also provides open, standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows.

4:31collaborate with developers from cisco dell technologies google cloud oracle red hat and more than 75 other supporting companies to build next generation ai infrastructure together agency is dropping code specs and services no strings attached visit agency.org to contribute that's A-G-N-T-C-Y dot O-R-G. My name's Carl Friston. I'm a professor of neuroscience at University College at London and chief scientific officer of Verses, a cognitive computing company in Canada and America. Verses is using your free energy principle in its models. And to talk about free energy, I thought we would start with physics, move on to neuroscience, and then discuss how it's implemented in, or the free energy principle is implemented in machine learning.

5:40My understanding is that free energy is a measure of the usable energy in a system. Can you talk about that? It's a very confusing concept. it gets more confusing when it moves when you take that principle to neuroscience but can you talk first of all in layman's terms what free energy is in physics yes i can or perhaps i should qualify that i'm not sure i can do that in layman's terms i have a reputation for nothing so you're gonna have to hold my hand in in trying to um make this accessible um so the free energy we're talking about here is closely related to the thermodynamic free energy that you're talking about, which is the amount of available work in a system that is not, if you like, entailed or trapped or locked in because of the entropic or the disorder of the system.

6:44However, the particular mathematical free energy, the variation free energy we're talking about, is a purely information theoretic construct. It actually inherits from Richard Feynman's work on quantum electrodynamics, and he was trying to solve a problem, which in fact Versys is also trying to solve in one sense, which is finding the most efficient paths, the most efficient courses in his context. These were the sort of paths of small particles, for example. So trying to find a probabilistic description of the paths of least action. And he solved that almost intractable problem by turning an impossible marginalization or integration problem into an optimization problem.

7:34And what he did to do that was introduce this notion of variational free energy. And then that was taken up decades later by people like Geoffrey Hinton and David McKay and the like. And they noticed that this information theoretic quantities mathematical information theory quantity that has exactly the same components as a thermodynamic free energy as an expected energy that's supplied by some world model or some probabilistic specification of the thing that you are trying to model minus the entropy was exactly the same that you would need to optimize things from the point of view of um crypto analysis and um with a nod to the russian notion of algorithmic complexity and that led through to universal computation in in terms of crypto analysis um and also was exactly what you needed if you wanted to make something like a Helmholtz machine work.

8:40So if we go right back to sort of, you know, the 1990s, you'll find sort of prototypical variational free energies at the heart of the things like the Helmholtz machine, like the sleep-wake algorithm. And indeed, you can track that right through to things like variational autoencoders. So this is like a universal objective function, and it has two bits to it um one bit is the entropy term and the other bit is this constraint or expected energy which is this basically um something that enforces um an accurate account if you're trying to model something say you're taking a sort of a neural network that was trying to predict the next word or trying to compress information in an image to predict the image and thereby generate new kinds of images, then you're going to want to do that as accurately as possible.

9:42But by including the entropy term into the free energy, then you're effectively applying a kind of Occam's razor. you're sort of not committing to a particular explanation but you're including or you're trying to um um not overfit the data because you're trying to maximize the dispersion or the uncertainty about your particular explanation for this content or for the uh for these data i slipped that in uh again this is not vernacular and i apologize but there's a beautiful link between uh minimizing free energy which is just a sort of expected energy minus entropy which means that you're implicitly under constraints maximizing the entropy which of course is james's maximatory principle which is a cornerstone of measurement physics and you know it is that sort of if you like built into your cost function it is that sort of uncertainty aspect that means you preclude overfitting and elude many of the problems that we see in current uh say deep neural networks um optimize using reinforcement learning so this free energy principle is a first principle account of how to make neural networks or neural networks used as agents or as classifiers or inference machines as efficient as possible, going right back to the path integral formulation of Richard Feynman in terms of finding the paths of least action, the paths of maximum efficiency, the paths of least effort that include, if you like, this cost of overfitting, technically a sort of complexity cost.

11:37So I don't have to change my mind too much in order to account for these data, which means that I don't want to commit to a particular explanation for these data. Does that make sense? It does, only because I've done a lot of reading. I'm thinking for listeners, one of the confusing things for me when we first spoke was the word energy when you apply it to machine learning. What you're in machine learning, energy is, is the, tell me if I'm wrong, but it's the, it's the difference between an input and the prediction is it's, it's an, it's in effect an error function. Is that right? Yeah, no, absolutely.

12:30Indeed, you can sort of look at this first principle approach to optimization in terms of error correction or error minimization. Indeed, in some parts of the life sciences, you get things like predictive coding, which is basically sort of scores the accuracy of your machine in terms of the prediction error. And that's an important part of this free energy. And indeed, you could actually read free energy in the context, in the setting of predictive coding as a precision-weighted prediction error. It's just the mismatch. So that's absolutely right. So just to address that, sort of unpack the notion of energy in this information theory context.

13:21And energy is just a potential. and a potential is just the sort of negative log probability so it's just a measure of the implausibility of something so if something is very implausible or very surprising or very unlikely or very uncharacteristic and then that has a very high potential it has a high energy and you really want to minimize that you know that energy um in order to produce an unsurprising account or render this content unsurprising given what you might have predicted if you knew the causes of this particular content yeah and and in neuroscience i mean you're a neuroscientist uh you saw this principle as uh as you know sensory input coming into the brain and and the brain trying to fit that in its model of the world.

14:17And if it's surprising, there's a lot of energy there, free energy, and the brain then has to update its model of the world to match the sensory input that's coming in. Is that right? No, that's absolutely right. That's a perfect description of predictive coding formulations of this kind of free energy minimization. that the brain is in the game of basically trying to predict what it's sensing, the sensory content. And if it's got the perfect prediction, clearly the free energy will be zero and there will be no prediction errors. So if this is a prediction error, that's newsworthy. So it can now use the prediction error exactly to say to update, to revise its explanation, its representations of the latent causes of these data.

15:12So, one can describe that in terms, if you're a statistician, that would be called Bayesian belief updating. If you're a philosopher, it would be sort of resolving prediction errors. It does really emphasize a sort of key point about this biomimetic principle of how we make sense of things and indeed act upon the world. We do so in a very sort of inside-out way. We generate predictions based upon hypotheses about what could have caused this content, what could have caused this sensory input. Just use the mismatch, the ensuing free energy flight, the implausibility in relation to your predictions of your actual sensory content to update online and over time.

16:00So now sort of introducing a distinction between inference. So in the moment, I literally change my mind. I literally changed my mind on the basis of the prediction errors that have been elicited by comparing what I predict and what I actually saw. But also, over time, the much longer times I can learn in the spirit of machine learning to be a better predictor given this kind of world. A couple of things about the neuroscientific view that I've found fascinating is your view is that this is an instantaneous, continual loop, that the brain is constantly doing this, and that some mental illnesses may be caused by an error or an inability of the brain to update correctly or to minimize the free energy principle.

16:56Absolutely. And in part, it was exactly that observation that motivated much of this work in the context of neuroscience and in particular computational psychiatry. And you could even go further and say that all mental illness could be construed as some kind of false inference. So you made the point that this process is ongoing, it's continuous. So it is something that we change our mind millisecond by millisecond, or at least 100 milliseconds by 100 milliseconds as we engage with the world and we continue making sense, doing our sense-making of the world. And that's a process of inference. It's not necessarily learning, which will be a much slower process.

17:44and that if if the way that our brains work can be cast as inference then that means if our brains don't work properly as in mental disorders then we can understand that as false inference and that makes a lot of sense in this you know and by false inference i mean exactly what you would have done at school when i don't know if you remember doing type one and type two errors when doing t-tests you know you can make a type 1 error and infer something is there when it's not of course that's a very apt description of an hallucination or a delusion but you can also make type 2 errors type 2 false inferences inferring something is not there when it is and of course that's a precise explanation for many neglect syndromes in psychiatry dissociative syndromes, what used to be called hysterical syndromes.

18:35So nearly everything in mental health and mental disorders can be construed as inferring properly or some failure of inference. And then getting into the weeds of the actual mechanisms of this predictive coding is really quite informative and sort of pointing to the particular failures of message passing entailed by this view of the brain. So I'm generating predictions in an inside sort of way, and then using the outside-in prediction errors to correct and update and perform this kind of inference. And another thing that fascinated me from what I read, that this cycle or this process matches the anatomy of the brain.

19:19I mean, that there are layers like in the visual cortex, and information is passed both back and forth, up and down, which is one of the reasons why backpropagation has never been validated as a process that happens in the brain because there isn't that kind of pathway for updating weights going down. Am I wrong on that? No, and I think you've hit a really central argument. and focus of both the machine learning community and people invested in predictive coding formulations. So just to sort of highlight the importance of that question, it is certainly the case that the functional anatomy, the computational anatomy of the brain has a hierarchical structure.

20:16And indeed, you could argue this is where deep learning comes from. It just means that the world model, the generative model, that is entailed by the structure of our brains has a hierarchical depth and of course then that you know naturally translates into deep neural networks and deep rl for example however there's a fundamental distinction between back propagation on a deep neural network and the kind of optimization and message passing between the hierarchical levels in the brain exactly as you say it is all local so as you move from one hierarchical or say layer in the neuronal brain architecture you are in exactly the way you describe sending messages to the layer below which are the predictions and then receiving messages from the layer below which are the prediction errors that revise the representations equipped with uncertainty that then generate the predictions so there's lots of local message passing crucially the objective function that describes the dynamics implied by that message passing is local so now you get into a message passing scheme which is much more sympathetic you know um well the message passing scheme um under predicted coding has this locality and it now acquires a biological plausibility but also a massive increase in efficiency which you don't get with back propagation in terms of sending all the way to the top and then all the way back again which is a very non-local sort of a violation not just of the principles of biological brain function but also the principles of least action that we started with because if you can do it locally you're doing it very very efficiently and you know interestingly in the past five years i would imagine people in machine learning have been looking towards predictive coding as an alternative to backpropagation of errors and certainly establishing that in terms of the efficiency it is at least as good as if not more efficient than backpropagation there are issues about scaling it up to the enormous dimensionality of deep rl but from a mathematical perspective this is the right way to do it because it is just much more efficient it is just so much and it just relies upon uh local um you know local optimization well the beliefs in this context um are used in a sort of technical sense of bayesian beliefs um so just to try and um you know not make that vernacular but certainly um qualify the notion of belief updating in the sense of basically updating so beliefs are read in my world or in the world of active inference and the free energy principle as um conditional probability distributions so are there sufficient statistics so they're not sort of propositional beliefs not something you can talk about these are just sort of mathematical descriptions of a representation which is equipped with uncertainty so there's content and the confidence that you have about this so that's another bright line between the kind of networks that we have in our brain and indeed that are employed in active inference for example versus axiom where the nodes actually represent beliefs in terms of both the content and the uncertainty or the inverse uncertainty which would be which would be the precision as opposed to a neural network which just has a value it just has a cascade of functions that you can interpret as representing content, but it is not equipped with uncertainty.

24:05So once you've got one of these, I repeat, biomemetic kind of neuronal networks in play, of the kind that you might find in predictive coding, then you can talk about the changes in the values of the neural network in terms of belief updating and of course beliefs about what but beliefs about the causes of your content beliefs about the causes of what is observable um so now there's a there's a distinction between the latent causes the unobservable causes of observable data sensations images um you know whatever you know whatever is can actually be explicitly measured or observed so when we talk about belief updating, what we're talking about now is a base-optimal way of updating your beliefs about the causes, the unobservable causes of your sensations.

25:04Of course, we can only sense a very small part of the entire universe. So we have to, on the basis of some very sparse sampling, build beliefs and update those beliefs continually about the states of affairs out there beyond the brain-bound skull or indeed beyond my deep neural network. So inference just is a description of the process of Bayesian belief updating. And Bayesian belief updating is just mathematically something that can be described as minimizing your uncertainty, minimizing your prediction error, minimizing your free energy in the right kind of way that includes this measure of uncertainty.

25:49Another, when you move into machine learning, this Bayesian active inference, it's happening continually. You're working with probability distributions, not weights in a, at least in, I mean, you're working with versus. And at least in Versus's implementation, you're working with the parameters in the model are probability distributions, not weights of individual nodes or neurons that influence the outcome more or less. Am I off there? No, no, no, I don't know. You're spot on. Yeah. So just to sort of contrast. So a deep RL or a standard neural network, including things like transform architectures that underwrite large language models or most of generative AI, these are basically just universal function approximators.

26:58They just map through a series, a cascade of functions, values afforded by the input to generate some transformed values, which are the output. Those could be predictions of the next token. There could be some action. That is not how the brain works. It is not how active inference works. And it's not, as instantiated, say, in Axiom, which was used to demonstrate the fundamental increases in efficiency if you do it using belief updating. So the equivalent architecture is certainly in place. There is a deep structure to it. I mentioned sort of world models and degenerative models before. The structure of the network determines the structure of your model and the implicit conditional dependencies and all the contingencies in there.

27:57But as you say, each node now stands in for a probability distribution, a belief, and therefore what you're passing around, the influence of one node or another node, is determined by and driven by the optimization of your probabilistic beliefs or representation. And that's really important because, for example, if I don't have a representation of uncertainty, then I don't know what I don't know. And more importantly, I don't have any way of judging or evaluating the goodness of an action in terms of its ability to reduce my uncertainty, to reduce my prediction errors, to reduce my expected free energy.

28:56but if i've now got a representation at each and every level of a deep generative model a deep world model i can work out if i went and solicited or looked over there so for example you and i are extremely skilled at this you know every 250 milliseconds every quarter of a second we deploy our visual apparatus by looking over here foveating foveating over here. That's an incredibly skillful act, simply because we are choosing to look at the parts of the visual scene that will have the greatest information gain or resolve the greatest amount of uncertainty, given what we believe at the moment, in this personal Bayesian sense.

29:43So what we're talking about now is a mechanics of optimal smart data mining, knowing where to go and get the right kind of information that enables you to resolve your uncertainty so that you can now have a better model of the world in this probabilistic sense. And Axiom showcased that in the sense that it spoke to what could be construed as two of the main problems with the current tech on offer with generative AI, which is basically the inefficiency and the lack of reliability. so the inefficiency can be resolved by appeal to these first principal accounts as exemplified by things like predictive coding or other implementations of active inference because now you've got this ability to resolve uncertainty in the most efficient way in accord with these principles of least action you can now not only outperform deep RL on the benchmarks as you know i think there's a 60 improvement in performance uh demonstrated by axiom but more importantly you do it much more efficiently so um i think there's a you know using three percent of the compute and that translates into efficiency in many many different ways so not only is it informationally more efficient or statistically more efficient it also is It's thermodynamically more efficient, so you're using less power.

31:23And it's also sample efficient, so you need a fraction of the data. So this is, if you like, by appealing to these biomimetic principles, or from a physicist's point of view, principles of least action, that sort of underwrite the way that we deal with our world and make sense of our world, You can resolve the efficiency problems that are currently plaguing or preventing the right kind of enterprise update of the opportunity to be AI, for example. Where you'd expect it to be deployed, it's not being deployed. And of course, the reliability is, you know, you can't get hallucinations if you've got an explicit uncertainty quantification.

32:11if you're not sure, you're uncertain about a particular recommendation or a particular classification, then because you quantified it, then you inure yourself or you protect yourself against making overconfident predictions or classifications and the like. Build the future of multi-agent software with Agency. That's A-G-N-T-C-Y. Now an open source Linux Foundation project, Agency is building the Internet of Agents, a collaborative layer where AI agents can discover, connect, and work across any framework. All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting.

33:19Agency also provides open standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows. collaborate with developers from cisco dell technologies google cloud oracle red hat and more than 75 other supporting companies to build next generation ai infrastructure together agency is dropping code specs and services no strings attached visit agency.org to contribute that's A-G-N-T-C-Y dot O-R-G. I should have said at the beginning, Axiom, that you've mentioned a couple of times, is versus is the startup that you're working with, and Axiom is the model implementing the free energy principle and Bayesian active inference and all of that.

34:28Is that right? Yeah, yes. it's a sort of a minimal implementation just to you know just to demonstrate what you could achieve and the problems you can solve if you commit to you know to these biomimetic principles you know and there are many aspects that we could talk about you know there's one of them interesting which people sort of latch on to and indeed speaks to one of your other questions about sort of the architecture of the brain i mean the brain is not just you know a deep neural network it has a lot of factorial structure it has a lot of modularity some parts do sort of perception other parts do planning other parts do um memory um so all of these different aspects of uh inference and sense making and decision making all um imply and require a certain kind of structure so your axiom was built to show that if you put this kind of structure, agentic kind of structure, into a world model, into a generative model, and then you apply these variational principles of least action to the local updating, the local optimization, then you can get this kind of performance enhancement.

35:46And to my mind, probably the more important aspect is this efficiency. you could have guessed that from the mean that you and I can drive a car on 20 watts but if you have to train a large language you may need a nuclear power station this is to my mind this is the important part of the solutions to the current problems that are inherent with the current direction of travel with Jared and Alec. Yeah. And a couple of things that fascinated me is this is a dynamic model. It expands and contracts depending on the data that it's looking at. Is that right? So, and it learns continuously because the probability distributions are being refined, but they're not being overwritten.

36:49And that's a problem with neural networks that new data coming in can, that's why they can't achieve continuous learning because of catastrophic forgetting. And can you talk about that, the promise of continual learning in this kind of a system? Yes, I can. But you've already summarized the bullet points. So I'll just speak to those. No, absolutely. So we do not suffer from catastrophic learning. We will grow our models, our world models, in a way that is sufficient to account and optimally explain what's going on. And in growing them, we're now talking about something you implicitly implied, that the gerontine model in and of itself has an optimal size.

37:44For example, it could have, if you were just using a simple layered deep neural network in this instance, how many layers do I have? And in each layer, how many hidden units? And in my world, how would you factorize or modularize within each layer? So all these architectural aspects determine the structure of the model, which just is a specification of the cause-effect structure statistically in the world you're trying to model. so if you remember where we started with the free energy and the importance of that putting entropy into the objective function to implement sort of occam's uh razor what that means is there is an optimum complexity of the structure for any given set of content or world in which you are navigating predicting classifying actor acting in so you need to find the right sort of complexity of the structure, the right depth of the model, the right sort of factorial structure that is sufficiently complex to provide an accurate account, but not too complex.

38:52So this is exactly Einstein's, you know, keep everything as simple as possible, but no simpler. So there's a sweet spot. And in principle, and indeed in practice, one, and indeed Axiom actually demonstrated this under the hood though it was never it wasn't sold as one of the key issues but to my mind it was a really interesting aspect that it grew itself and just to the right level of complexity before if you're like opening up new slots or new um parts of the model for objects that it had never you know it had never seen so you've got this notion now that you can um increase or decrease the structural complexity of the model in exactly the right way so you don't overfit and this contrasts with deep rl because in deep rl you just start with something that's too complicated you start with billions of parameters and then fondly hope you can reduce the number of parameters say connection strengths or weights by some kind of say pruning or dropout or putting noise in and doing mini batching in a way that eludes those sharp minima that you get from the overfitting, simply because you've got too many parameters, too many connection weights in your model.

40:11So conventional approaches use this sort of top-down approach. They start off with an overly expressive model with too many parameters, 99.9 % of which will be redundant, and then try and find some engineering heuristics to eliminate the redundant, to reduce that complexity in a way that you get for free if you'd use the free energy. But to use the free energy, you have to encode the uncertainty. And of course, deep RL doesn't have that. It doesn't have the capacity to represent the uncertainty unless you're dealing with a variation autoencoder. So this is another motivation for using the free energy as your ultimate objective function because it now gives you a way of scoring not just the quality of your current beliefs in the moment or the parameters of your generative model the connection strengths as you assume a simulation accumulate data and build better and better models but also it scores the quality of the structure of your model and that speaks to the possibility of growing models uh so growing models from scratch just so you can get to the right level of complexity apt for the kind of content that you're that you're you're dealing with and this is you're not it is a fascinating and i think really important possibly vexed um problem uh at the moment which in the life sciences would come under the rubric of structure learning learning the right kind of structure automatically in the face of data by optimizing your model with respect to this free energy score that can also be well provides a bound in machine learning i should say this is the elbow this is the evidence lower bound in fact it's another bound because it's a negative in physics.

42:10And therefore, you can understand this sort of model growing and shrinking and adapting to new worlds and new contingencies as maximizing model evidence. Or if you treat your model as an agent, then the agent looks as if it is gathering evidence for its own model. And philosophers sometimes call this self-evidencing, a little twist on the philosophical, simply because you're trying to optimize the model evidence that is bound by the elbow or the variational free energy so this notion of you know gathering evidence can be played right through not just to beliefs in the moment beliefs about states of affairs in the world or hidden states but also the parameters of your model that would be optimized through learning but also the very structure.

43:07In statistics, we call that Bayesian model selection. So it's selecting the model, the structure, that has the highest evidence in the face of the content or the data which it's trying to explain. And if you pursue that notion, what you've now got is a mathematical image of natural selection. So you can now look at natural selection as nature's way of doing Bayesian model selection where you and I are the hypotheses and the structures that provide evidence that this is the kind of thing that can live in this eco-niche. So again, we come back to biological principles. Nature's already done that.

Read the full transcript

43:52Nature's already done its Bayesian model selection and free energy minimization to create you and me as existence proofs of generalized intelligence right now in its implementation in axiom it's not a language model i mean it doesn't uh the the data that's coming in is primarily uh frame by framed video data or or it's something like that sequential data the inferences are uh decisions and i understand it can be applied for example, could this also work on language? I guess two questions. If you have this model, I mean, this is, you know, Jan LeCun's theory that if you have a model that has direct experience of the world, it'll eventually, and if it can learn continuously, It'll eventually catch up with human knowledge.

45:00But so much of human knowledge is embedded in text. That's what makes language models so intelligent, is that as inefficient as it is, they have that knowledge base to work from. So can this be applied to language? And if not, or even if it can, if you ran a model like this theoretically long enough, would it eventually learn all of the knowledge that humans have accumulated today? That's a challenging question, which I could take in many directions because it's such an inviting question. First of all, it would only learn language if it was exposed to people speaking that language. So there's a certain aspect of active inference that rests upon a coupling between the agent and the world that it is trying to model and gather evidence for.

46:18So if that world includes other agents like itself, then there will be emergent language. And certainly exposing this agent to communication with other agents like itself, it would eventually learn a language, just like you and I do and just like our children do. Would it do that just by being exposed to large cohorts of textual language of the kind that are used to train large language models? I suspect not. I agree entirely with Jan. I think he's on exactly the right track here, with one small exception.

47:00I think he would argue and I would certainly argue as would most of my colleagues in the life sciences and philosophy that to build the right kind of generative model where the words have meaning you have to be embodied and situated so that you have to experience the world as part of that world in order to endow meanings to the actual language. So we're moving beyond just the statistical structure of language and the predictability of the next word to now read the words as labels or some representation of some latent causes or states that are part of your gerative model. And in that sense, if you exposed an active inference agent to the world as richly encountered by people like you and me, or things like you and me, then yes, it might indeed learn language.

48:07but in no way different from the way that your children learn language or my children learn language it would have to be exposed to the lived world so that the meaning is all in the interactions not just with the world but also with other things in that world that are like me and i keep emphasizing the like me because um for language or communication in general to have any utility there has to be a common ground a shared narrative a shared reference frame so from the point of view of the free energy principle active inference or indeed artificial intelligence research that basically means we're talking about artifacts or agents or artificial intelligences that are natural ones and that have a shared world model and if you have a shared world model then you can share your beliefs through some kind of communication you can only do that with a shared reference frame.

49:04So I think this interactive aspect becomes incredibly important when understanding utility of language. And you could turn your question on its head and ask, why are large language models so potent? They are beautiful things. The probably most beautiful invention of this century. Why are they so alluring? And I think it is just because they generate content that resonates with our shared narrative and that we can understand and we can anthropomorphize them um so i you know but is that enough um to um is that enough to build a world model i i don't think it is i'd be much more with with yan and and and a lot of I repeat, I'll tell you to other colleagues, with machine learning, that you actually have to have something that's embodied and situated, embedded, and dynamically coupled to and engaging in its environment.

50:09Yeah, although the two could be combined at some point. The other direction I wanted to take your question. Yeah, absolutely. So does that mean that there is no role for all this beautiful engineering that we have access to, especially over the past few years. No, if you can, in the context of active inference and building the right kind of generative models, you can certainly install transformer architectures and deep neural networks. And the way that you would do that, and in my world that's called deep active inference it's literally putting a um a standard machine learning neural network uh into one of these technically what they're called factor graphs but um say neuronal networks with message passing and belief updating and the way that you can um massively accelerate the efficiency of this kind of predictive coding or belief updating or active inference is if you can map directly from content to the beliefs about the causes of the content then you then you can very much accelerate the efficiency of these continual learning and continual inference machines but you can only do that if it's learnable so on the you know as it says in machine learning on the tin it has to be learnable which means that it has to be context is sensitive.

51:47So if you can find some mappings between content and probabilistic beliefs, very much like a variational auto-accur that goes from data at the bottom to some posterior, some variational density with expectations and variances on it at the top. If you can learn that because it's learnable, context-sensitive, then you would actually have a hybrid, or you'd be using conventional technology machine learning to augment active inference schemes but notice you can only do that when it when it's learnable which means it has to be context insensitive it has to be you know has to have a certain symmetry and be exactly the same mapping over time so you can look at this as basically using machine learning to learn how to infer

52:44and leverage that by, I repeat, putting that inside these message passing schemes. The other way that would be really interesting is to map from the actual posterior basing beliefs of your machine, beliefs about not only states of affairs in the world you know what did you cause this particular image but also about its intentions and actions so if it if it has agency which simply means that the world model of the geratin model entails or part of the geratin model includes the consequences of my actions in the future which in turn implies that i can plan by choosing a particular course of action, which in turn means that I've now got true intentionality.

53:36I have intended counterfactual states in the future that are part of my gerative model. You get a large language model to tell you not only what this agent thinks about the current states of affairs, but what it intends to do, and why it's doing this as opposed to doing that. I mean, that would be wonderful if one could sort put a large language model to interrogate and to render what is private to the active inference neural network now publicly available because it can now broadcast to the user or to other artificial intelligences, not only its beliefs about what's going on, but also what it intends to do and why it intends to do that.

54:21Yeah. In which direction do you think the research will go both directions or, I mean, which direction are you working on? Well, Versus at the moment is really sort of focused on establishing, socializing this approach and also serving sort of customer needs. There are a few sort of lighthouse customers and demonstrating the improvements or the solutions to the problems of inefficiency and unreliability in alternative current options, both in the context of complex system modeling. So one example would be the work with analog and trying to optimize the placement of cab drivers. So imagine that you're waiting for a cab and there's some scheme that is deploying cab drivers to your location.

55:32and yet this is a you know an incredibly difficult problem to optimize because it depends upon the traffic flow it depends upon whether tennis swift is in town depends upon where your drivers are so this is a complex system that has a cause-effect structure that can now be learned on the basis of historical data and ongoing learning if you put one of these active inference machines trying to explain what the cause-effect structure of traffic flows and the deployment of your drivers. And we applied this sort of first principle approach, and in simulation we're able to get at least 30 % more rides in simulations for the cab drivers.

56:19So that would be one example. the other direction of travel i think speaks to your you know what we were talking about in terms of um um imbuing ai with the authentic kind of agency that rests upon having intentions and the ability to plan um i think this is probably best demonstrated in a versus work in robotics for example, looking at benchmarks in the habitat setup, where currently robots can now understand and infer the scene around them in terms of, for example, very simple notions like, there are objects and have in mind as part of the gerative model the consequences of reaching out or moving objects to the extent that if you can build a sufficiently deep gerative model with a separation temple scales and planning you can simulate and indeed realize practically and indeed the robotics team versus R &D have done this, you can get robots to sort of clean up a room or go and get something from the fridge and do it with no...

57:46to do it instantaneously because this is inference. You don't learn to take something from the fridge. You infer, what do I need... what am I going to do if I'm the kind of thing that needs to retrieve a bottle of milk? Well, where is a bottle of milk? Well, it's likely to be in the fridge. How do I result by uncertainty that the bottle of milk is in the fridge well i have to open the door well how do i open the door well if i if i'm that's kind of thing that's opening the door i should expect to feel my my actuators you know sending these imu signals and then you just get this predictive coding uh sort of technology just to make all that happen um so that you know that would be another direction of travel probably more apt less for work with people um doing complex system um you dealing with complex systems such as traffic flows or our work with investment, people responsible for investing, for pensions who want to sort of maximize their return but minimize their volatility.

58:49That kind of application, I think, speaks to the efficiency of being able to have good world models of the system that your customer is dealing with. as distinct from building agents that have a situation awareness and can use active influence to actually do stuff in the moment, which would obviously speak to things like robotics and autonomous vehicles. So those are two, to my mind anyway, two ways of leveraging and exploiting these first-principal biometic approaches. I mean, it seems that there's so much that you can do with this. Sorry, how widespread is research on this active inference or first energy principle in modeling?

59:46I mean, is it, for example, with your academic hat on, do you have a team of PhD students working on this? Are there other institution or labs working on it? Right. That's a nice question. Yeah, so in the life sciences, and particularly in the neurosciences, active inference as basically the technical version of something called predictive processing is now the main paradigm and has been since the turn of the century. So last century, we had behaviourism and behavioural psychology that gave rise to reinforcement learning in behavioural psychology that then was the inspiration for much of, say, Sutton and Barton, and then subsequently Q learning and reinforcement learning.

1:00:39there was a move in the latter half and 20th century to now away from behaviorism sort of you know thinking about cognitive processes belief updating from from you know from the point of view of this conversation and then at the turn of the century there was this inactive movement um where people were emphasizing the embeddedness the inactiveness the extended aspects of cognition in exchange with the environment at that point in academia certainly the cognitive neurosciences and in philosophy and and the neurosciences predictive processing you know became the major theme active inferences are so if you like the technical aspect of that so in that world um it's now the standard model in industry company uh and and um applications of a commercial sort i think we're just at the beginning um there are um there are now institutes there's the active infrastructure institute based in california that's now a sort of not-for-profit registered charity, I think, in America.

1:01:56We're just about to have our sixth international workshop on active inference. But these are little baby conferences. They're dwarfed by Neurips, for example. So this is a very embryonic community with lots of bright young things, lots of promise. But from a translational perspective into industry, I think we're at very early days and for my body it'd be a very exciting time just to answer your final question all my PhD students grew up and left me seriously I actually started my flexible retirement from university last year so I have 80 % full time employment now but I have no students left they all grew up and what happened to them they ended up in verses which is one reason why I'm so committed to verses.

1:02:53Keep an eye on all my bright favorite young things. That's it for this week's podcast. I hope you learned something, and I'll be doing more interviews with researchers working on post-transformer architectures because I think that's the future. See you next time.

From the publisher

This episode is sponsored by AGNTCY. Unlock agents at scale with an open Internet of Agents. 

Visit https://agntcy.org/ and add your support.


How could Karl Friston's Free Energy Principle become a blueprint for the future of AI?

In this episode of Eye on AI, host Craig Smith sits down with Karl Friston, the neuroscientist behind the Free Energy Principle and advisor to Verses AI, to explore how active inference and brain inspired generative models might move us beyond transformer based systems. They unpack how Axiom, Verses' new architecture, uses probabilistic beliefs and message passing to build agents that learn like brains instead of just predicting the next token.

We look at why transformers face scaling and reliability limits, how Free Energy unifies prediction, perception, and action, and what it means for an AI system to carry explicit uncertainty instead of overconfident guesses. Learn how active inference supports continual learning without catastrophic forgetting, how structure learning lets models grow and prune themselves, and why embodiment and interaction with the real world are essential for grounding language and meaning.

You will also hear how Axiom can sit beside or beneath large language models, how explicit uncertainty can reduce hallucinations in high stakes workflows, and where these ideas are already being tested in areas like logistics, robotics, and autonomous agents. By the end of the episode, you will have a clearer picture of how Karl Friston's Free Energy blueprint could reshape AI architectures, from enterprise planning systems to embodied agents that understand and act in the world.


Stay Updated:
Craig Smith on X: https://x.com/craigss 
Eye on A.I. on X: https://x.com/EyeOn_AI  

More from Eye On A.I.

All 266 episodes
#302 Karl Friston: How the Free Energy Principle Could Rewrite AIEye On A.I. · 1 h 3 min
Listen in VO