Why the Next AI Breakthrough May Come from Physics with Max Welling - #774

25 Aug 2026 · 56 min · 20 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Max Welling (CUSP AI) argues the next AI breakthroughs for science may come from physics—especially equivariant neural networks, ML force fields, and links between generative AI and thermodynamics. He explains how CUSP uses physics-informed AI to design materials for climate and energy, and how their platform supports an end-to-end loop from molecule generation to simulation to lab/self-driving labs. He also discusses his book connecting generative AI math (diffusion, probabilistic models) to non-equilibrium statistical thermodynamics, plus a separate line on wave-based neural networks inspired by brain waves and phonons.

Guest backgrounds

Max Welling is a theoretical physics PhD; previously worked at Microsoft Research (AI for Science lab in Amsterdam) and founded/led startups (including one acquired by Qualcomm). He co-founded CUSP AI in 2024 with Chad Edwards.

Key claims

Equivariance improves molecular/material predictions; ML surrogates can accelerate quantum-force calculations by 3–4 orders of magnitude; CUSP’s agentic pipeline plus ML force fields enables faster candidate filtering; self-driving labs can raise experimental throughput (e.g., ~100 experiments/day); diffusion/thermodynamics share deep information-theoretic mathematics.

Notable examples

metal-organic frameworks (MOFs) for CO2 capture (Nobel-winning class, not discovered by him); semiconductor perovskites; battery/fuel-cell materials; PFAS removal; CUSP open-source “COPS” (universal particle simulator) compiling ML force fields into JAX for GPU-efficient MD; diffusion models used to accelerate free-energy calculations (protein–drug binding) and thermodynamics concepts used to reduce diffusion noise.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Max Welling's Journey in AI and Physics

0:45 to 2:47

Max discusses his background in physics and how it led to his interest in AI.

“And typically you need to use quantum mechanics to compute these forces because a large contribution comes from the electrons and electrons are very light.”

Founding Cusp AI and Company Progress

2:47 to 5:24

Max shares insights about founding Cusp AI and the company's growth.

“Actually, first I spent two years at Microsoft Research as a VP because they were building their AI for Science lab in Amsterdam.”

Addressing Climate Concerns with AI

5:24 to 7:04

Discussion on the role of AI in addressing climate change challenges.

“So let's dig into the technology and the approach that you're taking to kind of apply your, this original set of work that you developed to materials.”

Innovative Materials for Carbon Capture

7:04 to 10:39

Exploration of metal organic frameworks and their applications in carbon capture.

“And so we've been working, the first project we've been doing was improving the materials that take out this carbon dioxide from the atmosphere.”

Collaborative Approaches in Material Discovery

10:39 to 12:29

Max explains the importance of partnerships in material synthesis and discovery.

“We're doing a lot of work on perovskites, which is materials for solar panels, improved solar panels.”

The Process of Identifying New Molecules

12:29 to 14:01

Insight into the end-to-end process of identifying and generating new molecules.

“and then we can find customers for that IP.”

Navigating the Molecule Generation Process

14:01 to 17:47

Learn how AI is used to generate and filter molecules with desired properties.

“So we have a very large database of materials, which we have all ingested into this database.”

Experiments in Self-Driving Labs

17:47 to 19:41

Discover the advancements in experimental setups using self-driving labs for materials research.

“And so the experimental loop is much, much faster for those.”

Engagements with Academic Labs

19:41 to 21:06

Understand the collaborative projects with academic institutions and their outcomes.

“So basically, the setup is that we bring some customers and we can do experiments in that lab, but they're lab scientists and they can also bring their partners and they can use our platform.”

Publishing Research and Open Source Initiatives

21:06 to 23:15

Learn about the publishing efforts and new open-source frameworks developed in materials science.

“So myself, I have one day at the university, and of course, I still work in that day with students, and there we publish everything.”
Show all 20 chapters

Foundation Models in Chemistry

23:15 to 26:19

Explore the concept of foundation models in chemistry and their applications.

“Are the representations of the molecules Like, are there standard formats for representing these things that someone working in the space would already have?”

Future Directions for CUSP and Material Science

26:19 to 28:00

Discuss the future goals and innovations at CUSP in material science research.

“In this case, the use of fine-tuning, how analogous is it to the process of fine-tuning in LLM, you know, reinforcement fine-tuning with traces, you know, language-based traces?”

Exploring Material Science and AI Integration

28:00 to 30:20

Learn about the future of AI in material science and the importance of self-driving labs.

“You turn this into a bunch of tokens using your graph neural network.”

Generative AI and Thermodynamics Connection

30:20 to 32:36

Discover the relationship between generative AI and thermodynamics through Max's book.

“It's talking about generative AI and stochastic thermodynamics.”

Entropy and Information Theory in Physics and AI

32:36 to 36:29

Understand how entropy relates to information theory in both physics and machine learning.

“Or is that statement, you know, more kind of specific and concrete?”

Applications of Machine Learning in Chemistry

36:29 to 39:28

Learn how machine learning can improve chemical calculations and drug interactions.

“So, of course, what happens is the information in the universe doesn't go away, but it's transferred from something you know, the system, to the heat bath where you've completely lost this information, unrecoverable.”

Cross-Fertilization Between Physics and AI

39:28 to 42:00

Explore how concepts from physics can enhance machine learning and vice versa.

“But in the other direction, it's also true, right?”

Exploring Waves in Neural Networks

42:00 to 46:32

Learn how waves can enhance the function of neural networks and memory tasks.

“because you know concepts like heat um and work and entropy production these things we don't use when we talk about machine learning models.”

The Edge of Chaos in Neural Networks

46:32 to 53:56

Discover the significance of operating at the edge of chaos for optimal neural network performance.

“It's moved this information around in what we call channels or memory channels, or you You can call them capsules where this information gets sort of moved around and stays stable.”

Cross-Fertilization of Physics and AI

53:56 to 54:25

Understand how physics and AI can benefit from each other's methodologies in research.

“as a design principle for neural networks.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Max, it's so great to be on the line with you again. It's been a while. It's great to be back, Sam. I'm looking forward to our discussion. I am as well. Our audience can look up our conversations from, I think, 2019 and 2020, where we covered what you were working on at the time and I think still echoes into your work today. Geometric neural networks and gauge equivariance neural networks and the like. But I'd love to have you kind of catch us up on what you've been up to since. It's been quite a while. Yeah, actually, the aquivariance theme has definitely continued. So, in fact, I found out that aquivariance was used very fruitfully in chemistry and material science.

0:49So in chemistry and material science, people train neural network models to predict the forces on atoms because people want to evolve atoms forward in time in order to compute their properties, which is called molecular dynamics. And typically you need to use quantum mechanics to compute these forces because a large contribution comes from the electrons and electrons are very light. And so you need to treat them with quantum mechanics. But, you know, if you get 10 electrons or more, in the case it becomes completely unfeasible to solve the so-called Schrodinger equation. So people have come up with approximations like density functional theory known as DFT.

1:34And the inventors of that got the Nobel Prize for that. But what now people do is they train surrogates. So they provide data using this expensive approximation to quantum mechanics. And then they use neural networks to shortcut the computation. So to predict the outcome of that computation, but at a much more accelerated pace. So in other words, three orders or four orders of magnitude acceleration, more efficiency relative to these quantum mechanical approximations. And in those models, because the world is three-dimensional symmetric, so if I rotate a molecule, all the forces will rotate with it.

2:17And so we could now put the same ideas that we use for images, we could put them in these molecules, these models that predict the forces, and we could use equivariance. And so that's why I kind of, also because my background is in science, I did my PhD in theoretical physics, I thought, okay, this is a perfect unification of my old sort of passion and my new passion. I can put it together, and that's when I started to be interested in AI for science. So that led pretty directly to the founding of Cusp AI. Actually, first I spent two years at Microsoft Research as a VP because they were building their AI for Science lab in Amsterdam.

2:59And so I helped that along. But after two years, I wanted to start a startup. I already did a startup a while ago, but this was actually the startup that got acquired by Qualcomm. And then I spent some time at Qualcomm. and I really like startups, the dynamical environment and the impact you can make. And I wanted to do it the Silicon Valley way together with my co-founder, Chad Edwards. And so we started Cospi in 2024. Talk a little bit about the progress that you've made since then. What is the kind of shape of the company today? It's been a huge ride, actually. It's a roller coaster. So we started, I think, about two years ago.

3:42So May, spring 24. Yeah, we started with a good initial investment of about 30 million from which we could hire an excellent team. So the team has grown to about 50 people right now across different geographies. geographies. So there's a headquarter both in Amsterdam and in Cambridge. Actually, the headquarter officially is in Cambridge, but the two initial labs were Amsterdam and Cambridge because Chad is from Cambridge, you know, I'm from Amsterdam. We now also have labs in London and Berlin. And we're also expanding into Asia and North America. Got it. And no surprise, your list of advisors is a bit of a who's who with Jeff Hinton and Jan LeCun at the top of the list?

4:35Yes. Yeah, the advisors are actually fantastic. So we have Jeff Hinton and Jan LeCun. We added to that also Martin Van Dambrink and Lord Brown. So Lord Brown is the former CEO of BP. And Martin Van Dambrink is the former president and CTO of ASML. They're both retired. and they like to spend their time with new startups and help them along. And then there's Verity Harding. She's working for the UK government and also DeepMind, or maybe formerly at DeepMind. And then Kristen Person, who has sort of initiated the materials project, and she's also advising on that side. Got it. So let's dig into the technology and the approach that you're taking to kind of apply your, this original set of work that you developed to materials.

5:38How did you get started with that effort? So I think the opportunity is, so to me, there's a deep fascination with the fact that there isn't sheer infinite amount of possibilities in which you can put together atoms. And the universe has only figured out so many of them because, you know, they form naturally, I guess, in the universe. But there's many more that you can design yourself with all sorts of exotic properties. When we started this company, both Chad and I were kind of concerned about the climate. We still are. And so we felt there's a strong need to accelerate the energy transition to more sustainable energy sources, as well as trying to take the carbon dioxide that's in the atmosphere out.

6:27So not many people know that by the time, but 2015, of course, we really like to be completely carbon neutral. But after that, there is still 50 to 100 years where we have to take out every year about half of what we currently put in. So that's 20 gigaton a year. That's an enormous amount. I think the size of the Lake of Geneva filled with sort of liquid carbon dioxide, a gigantic amount. And we don't have the technology for that because it's actually very hard to take it out because it's so dilute in the atmosphere. It's very expensive, too. And the two factors which are more expensive are energy and the material, the sorbit material that you use in order to take the carbon dioxide out of the atmosphere.

7:15And so we've been working, the first project we've been doing was improving the materials that take out this carbon dioxide from the atmosphere. And these are called metal organic frameworks. Actually, this is the material that this year won the Nobel Prize in Chemistry. And yeah, so we built a platform, and I can go into much more detail, but we built a platform that designs this material with the use of AI. So you can sort of think of it as a search engine, but it's not searching over existing documents. It searches over known and unknown materials. So it actually completely designs them from scratch if there's no suitable ones already available.

7:54And there's many components. It's a Gentix. So there is an agent sitting at the core of it who orchestrates a long computation. And it searches through existing databases. It generates entirely new molecules. and it also evaluates all these molecules with all sorts of tools. And in this generation and evaluation, you know, equivariance and all sorts of methods that we have developed over the years in my academic lab play a very important role. To be clear, you mentioned a molecule that won the Nobel Prize. Were you involved in the discovery of that particular molecule? No. No, I wish. Those were chemists who did that, who discovered it.

8:43I think Professor Kitagawa, Professor Yagi, and Professor Propson. I think those are the three. I'm not quite sure of the last one. But this is a class of molecules where there is a metal complex, a metal node with all sorts of atoms along it at the vertices of a graph. And then there is so-called linkers, which are also organic complexes, which are connecting these vertices on the graph. And they're extremely porous. So they have a very large holes in the middle with an enormous surface area. And so if you blow sort of air atmosphere through it, the molecules in the air, which is water, nitrogen and carbon dioxide.

9:30The carbon dioxide is only a small fraction of that. they tend to stick to the sides. And so then what you need to make is a molecule. You need to design a molecule where actually only the carbon dioxide sticks inside these pores and the rest goes through. So that's a design question. And also when it's full, you want to shake it or heat it to get it out so that you can actually reuse that particular material. What we added to this is a way to fine tune or design a molecule for a particular purpose. So people have actually made maybe around 100 ,000 of these molecules in labs and verified their structure.

10:18And so what we can do is we can now basically come up with an entirely new molecule for a very specific task and then make it in the lab and then use that for that. for that particular task. I should say, CUSP is not only working on MOPS. In fact, this was just the first set of molecules that we worked with. We have, hence, expanded to semiconductors. We're doing a lot of work on perovskites, which is materials for solar panels, improved solar panels. We look at semiconductors for new chip materials. We look at battery materials, fuel cells. We also look at removing PFAS from water. And the current set of molecules we use for that is, again, metal organic frameworks.

11:12Do you partner with other companies that have an interest in these particular molecules? or are you out exploring? And then if you find something, you will find partners and maybe license to them. What's the thinking around the business model? So we like to work with partners because there's a very broad class of materials and every class has its own super experts that focus on those particular areas. And they're either in academia or they're in companies. And of course, we also like to partner on the actual synthesis of these materials. So we partner with academic labs, but also with the labs inside of these companies.

11:50And so we build an ecosystem or a network where our engine can actually help in all of these different material classes, design the materials, and then we work with those companies to actually make it. So that's a partnership model. But there's also internal projects that we run. So for instance, the project on metal organic frameworks for carbon capture we ran self-funded internally. And then we also have another project now in the semiconductor side where we run it. And so if we discover something fantastic with self-funded, so we got the IP and then we can find customers for that IP. I think the most important thing is that we discover something that gives people certain confidence that we can do this.

12:41We actually really own this process and we know how to do it. And so then the cost that we work with customers to design materials for their specific needs. You talked a little bit about kind of the generative or agentic nature of this scanning scientific literature and, you know, using that to identify potential molecules. But, you know, we've also talked about, you know, some of the geometric implications of your work. You alluded to, you know, potentially the use of simulation. Can you talk a little bit more about the kind of end-to-end process of identifying these molecules? Yeah, happy to.

13:20So I guess there's a sequence of things that happens, right? So the first thing is, like in a search engine, you actually type a request. So you basically say, I want a material, and these are all the properties that it should have, and these are the things it should not have, these are the properties it should not have. And so you give this as a query, and then you could also tell if you have prior knowledge it's about how you want this particular search to happen, you could sort of tell the system, maybe use these tools and maybe sequence it in this way. So you can also give it some instructions.

13:55And then it goes through a process of steps. So the first step is it will look through its database. So we have a very large database of materials, which we have all ingested into this database. Lots of it is sort of exclusive licenses from the big publishing houses. and then they will start to look through all of this literature whether something exists out there that has these properties or which is close to having these properties. And so if it doesn't, then it will have to go into a new phase. So you can hold a conversation with this agent and talk about it. So it's already quite useful. But typically then the next step is that you go to a generative model.

14:37So in this case, it's the same generative model that generates images or video. And so, but in this case, it will generate molecules for you. And so you tell it, I want molecules with these very specific properties, the conditioning statement, it's called. And then it will start to generate these molecules, often hundreds of thousands of them, because it's quite cheap in the computer. And then comes the next phase, which is out of these generated molecules, we now have to sieve out the ones which look very promising. And this can be a very expensive step. In some sense, you build a multi-scale digital twin of the process that you really want these molecules to operate in.

15:25And so the first step is basically you relax the molecule to its ground state to make sure that it's the best energy state. Then you do a bunch of checks, like is it charged? If I shake it, will it fall apart? How big are the pores inside, if that's important? You know, all sorts of things that are easy to compute fast to throw away, you know, a whole bunch of things that do not look promising. And so then. So filtering it down. It's definitely a filtering step. Yes. And then go to the next step. So we have a pipeline that fine tunes or distills machine learning force fields for that particular material.

16:07So this is a process by which we take all the data there is about that material. We have a foundation model that's trained on a much wider range of data. And then we distill this force field into this. It's a very efficient force field for this particular class of problem. And we use that in an MD loop, typically the molecular dynamics loop, to simulate the molecule as it wiggles around and moves around, from which you can often compute very key properties. of that particular molecule those properties then often go into a partial differential equation at a higher scale and or in a process that actually models the device in which you want this up this material to operate and so that's again more expensive and so again you want to do this with fewer and fewer candidate materials and then at the very end you go to you know to an experiment right you go so now now actually do the experiment now that's more expensive and even slower um and that's you should do this only with order 10 materials at most um in the old way and so then uh so and then you get a candidate what we are currently the current i would say revolution that's happening in this space is self-driving labs where the amount of experiments you can do is much much faster so you could do maybe 100 experiments a day um and then the game is is more like the agent figures out what the settings of the experiment should be.

17:37The experiments are done, the data comes back in. And then you have the data from the experiment and the data from your computations. You combine them to set the next stage for the experiments. And so the experimental loop is much, much faster for those. And that's a very interesting development that we are now integrating our platform with. Got it. And so how many molecule classes and individual molecules have you kind of gone, you know, all the way through this cycle with? Yeah, we are engaged in a few of those, but, you know, the question is a little bit, what do you mean by all the way? The different materials are at different stages of maturity.

18:22So one of them, we went all the way to actually doing the lab experiments. Another way on semiconductors is on its way. And probably in a few months, we'll start doing the experiments. I meant all the way to lab experiments. And I was curious if you have enough data to say that your hit rate with this process is higher versus lower, or if there's anything you can say qualitatively or quantitatively about the candidates that you produce relative to, you know, the traditional approach? Yeah, so I definitely think that, you know, there's definite evidence that these things are much more efficient.

19:06So some of our scientists, they have said things like, we can do now in a few days what took a PhD before, but that's more in the digital domain. These are actually scientists that are more simulation-based work. We have two projects going which have experimental pieces to it, but we have a contract with a big national lab in Asia, which I cannot quite say the details of yet, where we have many more of these experiments planned out. So basically, the setup is that we bring some customers and we can do experiments in that lab, but they're lab scientists and they can also bring their partners and they can use our platform.

19:54And then we can collect data that way. And I think, oh, and then there is one other one that we're currently doing in Amsterdam on perovskites, which is running right now. So that's also a lab engagement. There is one. Perovskites is what? Sorry, yeah, proskise is a material class that's a semiconductor crystal structure that you would put on silicon typically on top of the normal solar cells. And that can help you filter out or basically convert a much larger amount of the energy in visible light to energy. Right. So we also have something with catalysis with DTU, which is the Danish Technical University.

20:39so that's on currently running on catalysis and we have a project that's about to start them off so there's quite a few engagements with labs that are either running or are starting to run but they haven't finished completely and does your lab or does CUSP publish are you active are you still active in kind of academic publishing around you know materials science now? So myself, I have one day at the university, and of course, I still work in that day with students, and there we publish everything. I'm very interested in all sorts of things which have to do with AI for science, so definitely yes.

21:21But also, COSP actually publishes. So we have interns that work with our scientists where we publish the results. So we, for instance, recently had results on machine learning force fields with uncertainty prediction. property predictors. And the most important thing, I think, that we've recently released open source and published a blog post about is our new molecular dynamics framework called COPS, confusingly. So that's universal particle simulator. And so the reason why we did this is actually together with NVIDIA is that, you know, Now, with this new development of these machine learning force fields that I talked about, these neural networks that replace quantum mechanical calculations for the forces, you cannot run them very efficiently in the current sort of MD simulators, the simulators that evolve a material or a molecule forward in time.

22:27because you need to run these neural networks on GPUs. And typically, these simulators, they don't run on GPUs, they're more on CPUs. And you want to run things in parallel. And so what was built by our team is a framework where you can compile these force fields into JAX, which is Python-based. and then it will actually very efficiently run these MD simulators and you can run them in parallel as well so that you use your GPUs. The utilization of your GPUs is high and so we released this open source a few days ago actually or a week ago at iClear and yeah we got a lot of excited responses to that.

23:15Are the representations of the molecules Like, are there standard formats for representing these things that someone working in the space would already have? And then your machinery just works on those existing ones? Representations, you can think of that as maybe a foundation model for chemistry. So what you do is, and this is also work that we did together with Meta. So what you do in that case is you train yourself one of these force fields. So you take a molecular structure. So that is basically the position of the atoms and their class, like if it's oxygen or hydrogen. And of course, you want this to be equivariant because rotations and translations don't matter.

24:04And then you map this into a latent space where they get represented by some code that is meaningful. And if you do a machine learning force field, you will then actually predict the forces and the energies from that. But you can stop at that intermediate level, and then you have a representation for that particular material. And you can train this on a very wide range of materials and chemistry. And from that representation, you can build property predictors. You can use it to condition your generative models. There's all sorts of uses for that foundation model. but it's also the starting point for distilling, let's say, models for specific material classes.

24:48And to be clear, are these foundation models trained per molecule or per class or are they very broad? Yeah, so that's a great question. So a foundation model, almost by definition, is trained on a very broad class of materials. The best data set for that is the materials project data set and Omole from Meta. So they have generated a very large number of DFT calculations, these quantum mechanical calculations, to create that data set. And those calculations are being used to train these force fields that I talked about, or those representations, those foundation models. And then if you say, but I'm actually interested in this particular material, What you can then do is start from that very broad representation, this foundation model, and then fine-tuned model for that particular material class.

25:44That way it is specialized for that material class, but that way it's also very fast. Because you need these things to do very fast. Because in a molecular dynamic simulation, you have to call them many, many times. One little step in an MD simulator is a femtosecond. It's a tiny step. And so you have to call them many times to make any progress. And you mentioned distillation earlier. Is that where distillation comes in? You're trying to get to a smaller model that's more focused on the molecules that you care about? Absolutely. That's what distillation is called. You distill the big model into a small model.

26:19In this case, the use of fine-tuning, how analogous is it to the process of fine-tuning in LLM, you know, reinforcement fine-tuning with traces, you know, language-based traces? Like, what is your data set that you're fine-tuning and your process that you're fine-tuning with here? So, actually, it's quite related to an LLM. In fact, the models we use are very related to an LLM. So you can basically think of the sequence, the information that you have. You can sequentialize it. And then from there, you can actually map it into this latent space. So there's either, typically you can do like a graph neural network, which actually looks at the three-dimensional structure, or you can use it as an LLM, like a sequential version.

27:13Like often molecules are represented as a, you know, as a string called a smile string. So then it becomes a one-dimensional string and you can use that to represent the molecule as well. So you can use both LLM style models as well as graph neural network style models for this. So your base model might be similar to an LLM, but you've got a very unique tokenizer in the case of representing molecules. It goes even further because you can basically also look at combinations of language and graphs and molecules. So you can sort of have, you can learn from, you know, let's say the literature where there's a lot of text.

27:54And then you can, every time a molecule is mentioned, you can then actually use the molecular representation, which is then the little graph neural network that represents. So that becomes a token. You turn this into a bunch of tokens using your graph neural network. And so you can really start to combine these things, actually. What do you see the work that you're doing at CUSP going in the future? Like what are the kind of near and midterm and longer term things that you're most excited about with that work? Okay, so we're constantly expanding the tools that we need for the different material classes.

28:32So we know how to train these tools, but every new material class has to be retrained. We're also going up the stack. So we're going into more coarse-grained representations, like larger scale digital twins, all the way up to modeling the actual reactor device in which this material is sitting. The effort that's most important for us right now is connecting the platform to self-driving labs. So to really generate large amounts of data from the self-driving lab, that's important. And I think ultimately it's very important to go through the entire process of predicting the molecule, making it in the lab, scaling the material, and then actually putting it in an actual device and getting a customer excited about that particular material and willing to pay money for it.

29:30So this basically means that this last part of scaling is something that still needs to be done. But getting there as quickly as possible is, I think, absolutely key. And so I'm really looking forward to discovering a material that is unique. It could not have been done by AI. and that makes it into an actual semiconductor device or something like this or a new solar cell. And we can say, you know, AI actually helped, you know, was a very important piece of the design of this particular material. I think that's a very exciting proof point that will help the entire industry forward. You also have a book that you're working on that has a really interesting title.

30:20It's talking about generative AI and stochastic thermodynamics. what is that what's the connection between generative AI and thermodynamics and tell us a little bit about the direction you're taking with the book right so maybe first a bit of a history on this so I started this almost two years ago now even before Cusp so I had a bit of a lull between you know when I stopped working for Microsoft and the startup Cusp I started so it was about six months between that and I started me and my wife were on vacation in Italy and sort of I cannot do absolutely nothing and so in the mornings I'm just going to write a book I had my cappuccino in the sun and I would start a book so wonderful, the best vacation you can have and in the afternoon we would hike through the mountains and enjoy life and so and then I was teaching a course in a town called Muizenberg, which is a Dutch word, but probably people say Muizenberg, but it's close to Cape Town in South Africa.

31:34There is the African Institute for Mathematical Sciences. And so I had the pleasure of spending a couple of weeks there teaching from these kinds of ideas that I was writing a book about. And then I recruited two students to help me out actually finish it because starting is easy, finishing is hard. And you need some help. And so it took two more years to actually finish it. It's a huge amount of work. But what is the book about? So the book is about actually the mathematics that describes modern generative AI, including probabilistic models, including diffusion models and many other things. That mathematics turns out to be equivalent to the mathematics that describes modern non-equilibrium statistical mechanics or thermodynamics.

Read the full transcript

32:28And how specific is that statement? Is that statement, you know, everything is a PDE at the end of the day? Or is that statement, you know, more kind of specific and concrete? It's surprisingly specific. and I say surprisingly because I don't think, maybe many people think of it as an analogy. So that's what I think is a really good question because I think this runs much deeper. So I think it's more than an analogy in my mind. I think there is a very, very deep connection and this has to do with the fact that we are talking about information theory in the end, at the core of physics is information theory and in the core of machine learning is information theory.

33:15And both of these are described. So if you talk about information theory with loss of information, in other words, processes where you lose information as you are evolving over time. So there's an observer that tries to describe a process, but there's so many degrees of freedom to keep track of, you can't. And so what happens is that you're losing information and you need to capture that by probabilities. And that mathematics is the core of both thermodynamics as well as machine learning. And I think that's the core statement. And this thing, I think, you know, there's this concept of entropy, which is maybe interesting in both of these theories.

34:04And so maybe it's good to talk a little bit about that. So entropy is basically the surprise that you find for a particular state. Do you know the precise state or you only have a very vague probability distribution for that particular state? And that's true. But in physics, you think of entropy as maybe a thing you can measure of the system because you have to subtract entropy from energy, actually, to get free energy, which is the energy you're free to use to do useful work in the world. And so it feels like a real thing, not like something that I don't know about the world. In a Bayesian statistics, this is exactly right.

34:52It's like you talk about subjective statistics because it describes what I do not know about the world and I need to ascribe probabilities to the things that I do not know about the world. But in my view, and a number of physicists as well, in particular, E.T. Jaynes is a famous one. Entropy in physics also precisely describes all the information you're missing about the world. And so there's this very deep connection between these two fields. Maybe the last thing I want to say about this, there is this famous theorem by Landauer, who basically said that if you erase a bit from a device, you must radiate you know, K, L, and T which is just a number amount of energy as heat to the environment so here you can see the direct relationship I'm deleting a bit which is a piece of Shannon information from a chip and he determined when you do that you must radiate heat to the environment and so you can see the direct relationship It's kind of a conservation of information along the lines of conservation of mass, which, you know, is very foundational in physics.

36:09Yeah. So, in fact, the statement is the second law of thermodynamics says that the entropy must always increase, which means, you know, you either keep the level of information the same, which is when the entropy stays the same, or you lose information in the process, and that's when the entropy goes up. So, of course, what happens is the information in the universe doesn't go away, but it's transferred from something you know, the system, to the heat bath where you've completely lost this information, unrecoverable. And so, yeah. Anyway, it's a bit technical, but the idea is that information theory is behind both of these theories, which basically makes the math very, very similar.

36:52And many of the tools which have been developed in one field have their exact analog in the other one. And this book is a lot about finding these different, you know, this kind of dictionary between the two fields, right? There's something called stochastic normalizing flows in one in the machine learning, and then there is escorted free energy estimation in the other field, and they turn out to be exactly the same methods. And so ultimately, who is the book for, and how do you see the book changing the way they view the world or the way they're able to do the things they do? For actually the, you know, for both sides.

37:34It's written for the machine learner who is interested to learn a little bit about, you know, thermodynamics and non-equilibrium thermodynamics. And I think, and in reverse, right, it's for also for the physicist who wants to get into machine learning and build on the things they already know. And I just want to point out that the first paper that was written about diffusion models actually had the word non-equilibrium thermodynamics in the title. So the authors of that paper actually already knew about this connection. So diffusion models really are a process by which you take structure and you destroy it, which is typically what happens in the world.

38:13The entropy goes up. And then we try to reverse that backward in time, which is to start with noise and create structure, which is our generative models. And the way I think it will evolve into the future is, of course, there's, you know, or the way it can be used fruitfully is, first of all, it can help people in physics to use these tools for machine learning. And let me give you one example. So it's a very important problem to compute the free energy difference between, let's say, an unbound system, which is like a protein and a drug that you want to neutralize the protein. And so you want this drug to attach or bind to the particular protein.

39:00And so you want to figure out how much does it want to bind to this particular pocket on the protein. So you need to compute the free energy difference between the unbound state where these two things are far apart and the bound state where they're together. That's a very important problem. And it's something that a chemist would typically want to do. But now you can use modern machine learning methods like diffusion models in order to accelerate those particular calculations. So that's a clear way in which you can start from a machine learning method and help the chemist do their calculations better.

39:34But in the other direction, it's also true, right? Because there is things which have been developed in stochastic thermodynamics, like a concept called counter-diabetic driving. It's something that they have figured out on how to do very efficient control of particular physical systems. And those concepts can now be used in diffusion models in order to do a better job at generating images. because it turns out that that's very helpful in bringing down sort of statistical error or noise in the diffusion process. And so it can actually help you to build better diffusion models. So this is a cross-fertilization between these two fields.

40:20And do you see on the machine learning side the primary target of the analogy or relationship as being diffusion models? or is it more fundamental than that? It's more fundamental than that. In fact, the book goes into variational autoencoders. It goes into MCMC methods, Markov chain Monte Carlo methods, which are used to sample, sort of sample from particular distributions. It goes into free energy estimation. It goes, there's many different applications and cross-fertilizations that can happen. So I think as a field, we struggle a little bit with, you know, we're creating these ever more complex, you know, foundation models, our understanding with them, our understanding of them, you know, still somewhat surface level, like in terms of their mechanics.

41:17you know we've um you know thermodynamics as a field is you know much more mature do you see us being able to use this relationship to better understand the the models kind of at their core absolutely at least it gives a different perspective in trying to understand things so um i would say the field of stochastic thermodynamics actually itself quite new interestingly so it is still very actively researched so thermodynamics is systems in equilibrium stochastic thermodynamics is systems out of equilibrium so that's actually quite new field um but it it gives you a different way to try and understand um how diffusion models work because you know concepts like heat um and work and entropy production these things we don't use when we talk about machine learning models.

42:18But I think they give you a very new and interesting perspective on how to think about what's going on and how to also improve them. And so, yeah, I think, and in reverse, it's also true, right? Machine learners have their own way of thinking about how to improve models. And that could also help the physicists to think about problems. And I found both parties actually to be very interested in the other party's work. So that's good. Maybe to take the next step in getting even more abstract, you recently did a keynote at iClear. And in addition to talking about your work at Cusp and the book, you put forth another, you know, analogy, let's call it, or opportunity to kind of learn from the physical world.

43:08than that is in looking at waves and applying that to machine learning and stochastic systems. Talk a little bit about that work. Yeah. So this is, I think, very exciting in the sense that, so the thing that, so there's two reasons why we thought we needed to think more about waves in science. And that is because, first of all, waves are actually seen in the brain now because we've gone from single electrode measurements to maybe hundreds or thousands of electrode measurements. And of course, in a single electrode, you cannot see a moving wave, a traveling wave. But if you put hundreds or thousands in them, you can actually see these traveling waves.

43:49And people have now observed them everywhere. And then the question becomes, is this functional or is just a side effect of something? That's the first reason. The other reason is that in neural networks, you have this phenomenon called over-smoothing, which means that you start with a piece of information and you have thousands of layers. And this information basically is exponentially suppressed as you go through it. So the input and the output of the neural network become independent of each other. Of course, it's something that you really do not want because you want the output to say something about the input, like maybe what's the class of this particular input.

44:30But it's always been a little hard to make sure that the information travels all the way from the input to the output layers through thousands of layers. People have used many tricks to do that, but I feel many of the tricks we have developed in deep learning is exactly to try to do that. And the other application is in reasoning. So if you have to reason about the problem, you have to sometimes keep things in memory. Of course, we kind of write it maybe to shorter memory and then we read and write from this kind of memory. It's a kind of more stable piece where things don't get mixed up so fast.

45:07So you need some kind of memory.

45:13And basically, so the question becomes, how do you communicate very deep into time or reason very deep into time? Or how do you communicate with things that are very far away in your brain or take information from a material that has very long-range interactions. And also in materials, materials use waves for long-range interactions, which are called phonons. So phonons are lattice vibrations, and they carry information over long distances in a material. In fact, if you think about the universe, the only reason we can see very far and deep into the universe is because light waves hit our instruments in our eyes.

46:02And that is, again, waves that are doing the communication over these long distances. And so what's interesting is that in neural networks, we don't use this tool at all. So we just create maps and the maps, you know, map numbers to new numbers. And so the question becomes, can we actually start using this phenomenon, this wave phenomenon that we see in the brain? Can we also start to use it in neural networks? And so this has been the work that we've been developing. And especially for, so we've basically designed neural networks that naturally create these waves. It's moved this information around in what we call channels or memory channels, or you You can call them capsules where this information gets sort of moved around and stays stable.

46:51And we've been very successful in, for instance, tasks that require memory. So a task where a neural network or an RNN, basically, the task is you take a number, you hold it in memory, and then at some random other point, another number arrives. You hold it in memory, and then you add them. And then another random point, you're asked to produce the sum. and that point could be very far into the future. So you have to do the computation and you have to hold things in memory. And these models that naturally work with these waves, they can actually, they could do these tests much, much better than RNNs.

47:26When I hear RNNs and long running memory, I think about challenges like exploding gradients and then I think about kind of the oscillatory nature of waves and that that has some inherent dampening properties that allow you to kind of access, you know, memory further into the future without, you know, these exploding gradient challenges. Is that part of the mechanism at work here? That's a very good point. So if you think of this, a neural network as a dynamical system, so it is basically you think of the layers as time and you're trying to get a signal from the beginning and you're trying to propagate it through the layers over time.

48:07So a couple of things can happen. There's different types of dynamical system. The first one is a stable system where basically if you take two different inputs, they map to this. After a while, they collapse onto a same point. And then from there on, that one point moves forward. So basically, every converges onto what's called a point, stable point. And then from there on, basically, that's what propagates. that's not what you want because information gets lost because everything gets mapped to a point now that's an imploding gradient if you wish, that's something where the information gets lost there's another extreme where you take two input points, I say two images that are similar and then they start to sort of wildly move around and that's called chaos it's a non-linear system so this starts to wildly move around And you also lose the information, but for a different reason.

49:08The information is still there, but you cannot track it numerically. Your numerical precision is too small. And that's actually what we mean when we say entropy is increasing. It basically becomes random. It's a random process. And we all know that a Markov chain is a random process that depends on the previous step. it will lose information about the input because we can't track it. That's also not what you want. So what you really want is sitting somewhere in the middle, somewhere that's not too stable and not too unstable, and that's called edge of chaos. And people have found that, in fact, neural networks that perform best operate at this edge of chaos, and people have also found that, unsurprisingly, the brain operates at the edge of chaos.

50:00And we found that if you include these waves in your neural network, you very naturally operate in this regime, Edge of Chaos, and you don't have to fine tune the system to be there. So it's very easy to get to this Edge of Chaos regime. And then maybe if we have time, we can talk about how we do this. And this is done by spontaneous symmetry breaking. But that's a rather technical discussion, I think. And so what is spontaneous symmetry breaking? Okay, this is a very interesting phenomenon in physics. So here's another example of something quite deep from physics that we can start to use to understand neural networks, which is one of your previous questions.

50:48and so spontaneous symmetry breaking is where a system that has a particular symmetry let's say perfect translation symmetry so think of a gas as perfect translation symmetry or a liquid better and then the liquid at lower temperature the liquid will then actually condensate into let's a solid. You can go from water to ice, right? And if you look at ice, you know, it has a crystal structure. So it has less symmetry now because it's now the continuous translation group has now been turned into a discrete translation group where you can only translate over the lattice spacing. And so you've actually broken the symmetry into something smaller.

51:36And there is a very deep theory from physics that says when that happens, when a continuous symmetry breaks into a discrete symmetry or breaks into a smaller symmetry, then you'll have new wave-like modes which can propagate without using any information or with very, very little information. And these are these phonons or light, or these are the waves that can travel very long distances. They don't decay. They don't disperse. They hold a shape. You start with a certain shape, then the shape can just translate and just move. And so we thought that's a really good candidate for these traveling waves.

52:23And so what we did is we created the neural network with the large symmetry. We just bake it in. It's not like rotation or translation. that we typically use. We just give every neuron some extra dimensions and we say there is some symmetry. And then we do something, we generate random weights, which generates random activations. And if you make the weights large enough, the distribution from which you sample the weights large enough, the original symmetry gets broken. And that's complex to explain, but the symmetry gets broken. And now you're in this broken symmetry phase where you have these goldstone modes or these waves which can travel without using any energy.

53:10And they do. So you find them when you create that. You get these waves where you can just see things oscillate around very naturally. And then the idea is you can now train the neural network to make use of these oscillations which are basically for free. They're running there without costing any energy. They're very stable, so they oscillate all the way from the beginning of the neural network to the end of the neural network. And you can now design the system so that it has these waves and it has this edge of chaos behavior. So information from the beginning, the input of the neural network, propagates all the way to the output of the neural network through these wave-like patterns that are created by this phenomenon called spontaneous symmetry breaking.

53:56And so here we use something, a deep result from physics, as a design principle for neural networks. Well, I love that you're getting to play with all of these tools from physics and bringing them into machine learning and trying to find the correlations and ways that we can take advantage of them. There's many more to come, I think. There's a very rich field of mathematics and physics that can all be leveraged by machine learning. But it's beautiful because the cross-fertilization goes in two directions, right? because with our platform, we are using AI tools to help the scientists discover new materials and accelerate their simulation tools by using AI.

54:40And so it's really a healthy cross-fertilization between these two communities, which I find very exciting. Well, Max, it's been great catching up with you. You have a lot going on. You're working on a lot of different angles. it's been enjoyable to learn about them looking forward to keeping in touch well thank you sam it was a pleasure again to talk to you thanks so much see you next time see you next time

From the publisher

The conventional wisdom in AI is that the next breakthrough will come from more compute, more data, and larger models. But what if the next leap comes from somewhere else?

In this episode, Max Welling—co-founder and CTO of CuspAI and professor at the University of Amsterdam—argues that physics may provide some of the ideas behind the next generation of AI systems.

We begin with CuspAI’s work using generative AI to design entirely new materials for semiconductors, batteries, carbon capture, and clean energy. Max explains how foundation models for chemistry, agentic workflows, simulation, and automated experimentation are dramatically accelerating the search for new materials and reshaping scientific discovery.

The conversation then broadens into a deeper question. Beyond giving AI new scientific problems to solve, can physics also teach us how to build better AI? Max explores surprising connections between machine learning and thermodynamics, why waves may become a new computational primitive for neural networks, and how concepts like symmetry breaking and statistical physics could inspire AI architectures beyond today’s scaling paradigm.

🗒️ Full show notes: https://twimlai.com/go/774.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Why the Next AI Breakthrough May Come from Physics with Max Welling - #774The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 56 min
Listen in VO