πŸ”¬Why There Is No "AlphaFold for Materials" β€” AI for Materials Discovery with Heather Kulik

24 Mar 2026 Β· 35 min Β· 23 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT Β· Add to Claude

In short

Podcast Summary: Latent Space - Episode with Heather Kulik

Episode Overview

  • Title: πŸ”¬Why There Is No "AlphaFold for Materials" β€” AI for Materials Discovery
  • Host: [Latent Space](https://latent.space)
  • Guest: Prof. [Heather Kulik](https://cheme.mit.edu/profile/heather-j-kulik/)
  • Focus: The intersection of AI and materials science, exploring the challenges and opportunities in materials discovery and design using AI techniques.

Key Themes and Concepts

  1. Importance of Materials Science
  2. Materials science is foundational to product development and innovation.
  3. Significant advances over decades have made modern materials possible, influencing everyday products.
  1. AI's Role in Materials Discovery
  2. Heather Kulik has pioneered the integration of AI and computational tools in materials science before the trend gained popularity.
  3. Successful implementation requires a deep understanding of both AI techniques and domain expertise in materials science.
  1. Notable Achievements in AI-Driven Materials Design
  2. New Polymer Design: Kulik's group used AI to develop a polymer that is four times tougher than previous versions, showcasing the potential of AI to reveal unexpected chemical phenomena.
  3. 22-Atom Ligand Challenge: A recurring test to evaluate AI's performance in ligand design, highlighting AI's current limitations in complex chemical tasks.
  1. Challenges in the Field
  2. Data Scarcity: The materials science dataset is not as rich as that in biology, making predictions and generalizations more difficult.
  3. Human Intuition in Science: Despite advances in AI, human scientists are still crucial for complex understanding and intuition, indicating the need for collaboration between AI and human expertise.
  1. Limitations of Current AI Models
  2. AI models can struggle to replicate human-level intuition in chemistry, as illustrated by the 22-atom ligand challenge where AI often fails to meet specific criteria.
  3. Trustworthiness of literature: Kulik's team noticed discrepancies between reported experimental data and interpretations in scientific papers.
  1. The Future of AI in Materials Science
  2. Kulik emphasizes the need for more quality datasets to improve AI predictions in materials science, as many existing datasets are based on noisy approximations.
  3. The potential for a new breakthrough akin to AlphaFold for biology is hindered by the complexity and variability of materials chemistry.
  1. The Changing Landscape of Academia
  2. As private companies ramp up AI efforts in science, the role of academic researchers is evolving, with a greater need for creativity and problem-solving beyond sheer computational power.
  3. Collaborative efforts and access to high-throughput experimental labs are essential for academic researchers to contribute effectively.

Conclusion and Call to Action

  • Engagement in Research: Kulik encourages aspiring researchers to dive into chemistry and materials science while effectively utilizing AI as a supportive tool.
  • Resources: Kulik developed a software tool called MOL Simplifier for transition metal complex structure generation, available on [GitHub](https://github.com/) and other platforms.

Key Takeaways

  • Materials science is crucial for technological advancement; understanding and integrating AI into this field is a growing necessity.
  • AI has immense potential but also significant limitations, particularly when it comes to complex chemical and materials challenges.
  • The future of materials discovery will depend on high-quality datasets, interdisciplinary collaboration, and the curiosity of researchers willing to explore beyond conventional boundaries.

For more insights and discussions, check out the [full episode on YouTube](https://youtu.be/KSCCKCz2x04).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Questioning the Need for Traditional Chemistry

0:00 to 0:14

Discusses the relevance of traditional chemistry in the age of AI.

β€œThere's a school of thought that why should I bother to learn chemistry or physics or whatever when ChatGPT, you know, as PhD level understanding of that anyway.”

Understanding Ligand Binding

0:25 to 0:41

Explains the binding of ligands to metal complexes.

β€œAnd what that means is that some combination of atoms and it's going to bind to the metal and it's going to change its properties.”

Accelerated Discovery of New Materials

1:28 to 2:20

Discussion on how AI is used to accelerate material discovery.

β€œAnd yeah, maybe to get started, can you just tell us about one of the coolest things you've done in Europe and in for an AI engineering audience?”

Unexpected Discoveries in Polymer Networks

2:20 to 3:20

Sharing a surprising discovery in polymer material toughness.

β€œSo we were able to screen with artificial intelligence, a set of thousands, tens of thousands of materials where each individual experiment, if it were done in the lab, would have taken months to years.”

The Role of Quantum Mechanics in Material Design

3:20 to 4:43

Describing how quantum mechanics influences material properties.

β€œof some of the promise of AI and materials discovery.”

Transition to Machine Learning in Chemistry

4:43 to 5:50

Transitioning from traditional methods to machine learning in research.

β€œSo we weren't the first ones to discover that phenomenon on its own, the general phenomenon that putting little places that could break to make the network stronger.”

Active Learning in Material Discovery

5:50 to 8:19

Explaining the concept of active learning in material research.

β€œSomewhere around 2015, 2016, I realized it was a bad idea to call things cheminformatics, and it was a good idea to start calling things machine learning.”

Applications of Metal Organic Frameworks

8:19 to 9:38

Discussing the uses and potential of metal organic frameworks.

β€œAnd usually just even for a not so accurate machine learning model, you get, you know, at least a hundred to a thousand fold speed up for every dimension you're optimizing over.”

Understanding Metal Organic Frameworks

9:38 to 10:36

Explaining what metal organic frameworks are in layman's terms.

β€œBut they're used for all sorts of things, even drug delivery.”

The Evolution of Catalysis and Quantum Modeling

10:36 to 12:21

Exploring the historical context of catalysis and quantum modeling.

β€œbe combined in basically infinite ways to create very precise chemistry.”
Show all 23 chapters

Predicting Method Accuracy with Machine Learning

12:21 to 13:15

Discussing how ML can predict quantum mechanical methods.

β€œAnd that's what I would have normally been doing before I got started in AI.”

The Limitations of AI in Chemistry Revisited

13:15 to 14:01

Addressing the ongoing debate about AI's role in chemistry education.

β€œThat's probably going to be in the best part.”

Molecular Design with AI

14:01 to 15:49

Explore how AI models like LLMs can assist in molecular design, specifically in ligand identification.

β€œLike how do you find a new ligand that can go into a transition male complex?”

Challenges in Machine Learning for Chemistry

15:50 to 17:52

Discuss the data challenges in machine learning applications in chemistry and areas needing attention.

β€œBut one of my favorite things, if someone can get in one shot an LLM to generate me a 22 atom ligand, I would love to see it.”

Benchmarking in Materials Science

17:53 to 21:45

Investigate the similarities and differences between CASP in protein science and data challenges in materials science.

β€œSo in the protein world, there's CASP, right?”

The Role of Processing in Material Development

21:46 to 24:47

Understand the importance of processing in material science and how it influences the development of new materials.

β€œBut I think that there's also an extent to which that just pure process and automation, good operational practice, those are important things.”

Integrating Textual Information with AI

24:48 to 28:00

Learn about the integration of textual information from literature into AI models for materials properties prediction.

β€œAnd right now, no potentials are really robustly encoding all of that bonding, especially with respect to metal-organic bonding.”

Challenges in Literature Extraction with LLMs

28:00 to 28:32

Learn about the challenges of using LLMs for accurate literature extraction in materials science.

β€œBut the other would be just, you know, people interpret their results in different ways.”

Bias in Computational Methods for Discovery

28:32 to 29:31

Explore how computational methods can bias the discovery process in chemistry.

β€œAnd what about the way that it might bias the discovery process, right?”

Building Generative Models for New Discoveries

29:31 to 30:26

Understand the process of training generative models based on existing literature.

β€œAnd maybe it won't get all of them, but maybe some of those discoveries that we think are new in the most recent 20 years, maybe some of them are trivial for a model to generalize to, whereas others are not.”

Collaborative Initiatives in Materials Science

30:26 to 31:24

Discuss potential initiatives for collaborative and high-throughput data collection.

β€œWhat would your dream be if you could organize something which really will drive the field forward in your mind?”

Philanthropic Efforts in the Biotech Space

31:24 to 32:25

Examine the role of philanthropy in advancing material science research.

β€œSome research subfields are trying to do that, but it's not really developed across material science.”

Academic Resource Limitations vs. Corporate Power

32:25 to 33:38

Delve into the resource challenges faced by academics compared to corporate entities.

β€œYeah, that kind of brings up the question.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00There's a school of thought that why should I bother to learn chemistry or physics or whatever when ChatGPT, you know, as PhD level understanding of that anyway.

0:13Heather Kulik:ChatGPT is super good at Wikipedia level chemistry knowledge. I'm really interested in molecular design. Like how do you find a new ligand that can go into a transition male complex? And what that means is that some combination of atoms and it's going to bind to the metal and it's going to change its properties. The thing I constantly do every time an LLM is updated is I just ask it, please design me a ligand that has 22 atoms. I can never get an answer that has 22 atoms. Hi, we're really excited to have Heather Kulig here. She's a professor of chemical engineering at MIT. Heather has done some amazing work in material science and computational chemistry.

0:57But we're particularly excited to have her today because she has, for almost her entire career, been working on the intersection of using data-driven methods, AI, and applying them to improve materials and understanding materials. And she has a lot of really interesting opinions about what works and how do you approach these problems to get the most out of them. So, yeah, we're really excited to have you here. And yeah, maybe to get started, can you just tell us about one of the coolest things you've done in Europe and in for an AI engineering audience?

1:36Heather Kulik:Yeah, so my group, we work a lot in accelerated discovery of new materials. When I first started out, we were just really using AI to make predictions we'd normally make with computational models, just make them faster. But the question I would often get when we were doing that was, okay, but what's surprising? What's sort of something from AI that like, I wouldn't have already known if I were a really smart chemist or a really smart material scientist. And, you know, you make all these computational predictions. Has anyone actually made in the lab something that you predicted? Recently, I was able to do a really nice demonstration where the answer to both of those questions, you know, was very clear from the work.

2:20Heather Kulik:So we were able to screen with artificial intelligence, a set of thousands, tens of thousands of materials where each individual experiment, if it were done in the lab, would have taken months to years. And through AI, we uncovered this sort of unexpected chemical phenomenon that led to a emergent property in what's known as a polymer network, so plastics, that would make the polymer about four times tougher. And when we showed the design that AI had come up with to the experimentalists, they were really surprised. They would have never come on this on their own. And then we were able to convince them to make it in the lab.

3:02Heather Kulik:And in fact, it was this tougher material. And where this has applications is if we can make plastics tougher than we, you know, can get more use out of them. And it'll ultimately address some of the problems we have with overall durability and use of plastics. So I think that's an example of some of the promise of AI and materials discovery. Cool. So can you dig into that a little bit? What was the surprising chemical discovery there? So it's sort of hard for me to think about how to explain it without getting too deep into the chemistry. But basically, these are molecules that have to break apart.

3:42Heather Kulik:And when they break apart, they make the overall structure that they're in tougher. So a little part of the material breaks and that helps to dissipate the force. Normally, the way you would think about making it easier to break apart these small molecular components might be to create a hinge so they can kind of peel open instead of sliding apart. But what we discovered was that there was a fully quantum mechanical phenomenon. There was really no way for us to predict this, you know, based on anything else, where the electrons just move around in a different way so that at this moment where the molecule is going to break apart, it's a lot more stabilized.

4:19Heather Kulik:These types of concepts, they're sort of similar to what's kind of known about how catalysts and enzymes work, but it had never before been shown in these polymer materials. So this is sort of like the fuse in the Bay Bridge that sort of like allows the bridge to keep its structural integrity during an earthquake by having a controlled break. Is that kind of... Yeah, yeah. So we weren't the first ones to discover that phenomenon on its own, the general phenomenon that putting little places that could break to make the network stronger. That was published in Science Magazine a couple years ago, But the specific way we came up with to design the material to do this, that was our new contribution.

5:02How did you, you mentioned that, you know, you started off in accelerating kind of existing methods using, you know, sort of enhanced computation. What caused you to take that leap to more machine learning based methods?

5:17Heather Kulik:So, you know, I was drawn to data driven discovery pretty early on, sort of before I even knew the phrase machine learning. and I guess I was just really excited by what you could learn from patterns and data. Back then we were trying to call it cheminformatics and just sort of trying to think about you know in what ways could you unearth trends in data because I started my career actually working kind of one molecule at a time or one material at a time and I was just impatient. I I wanted to be able to sort of understand not just one molecule at a time and write one paper about it, which is something people would have been happy to do back when I was starting my career in the mid-2000s, but to actually kind of unearth broader trends in how you understand how material is going to behave.

6:09Heather Kulik:Somewhere around 2015, 2016, I realized it was a bad idea to call things cheminformatics, and it was a good idea to start calling things machine learning. And I had a brilliant student, Jean-Paul Janais, who's now, I think, an assistant director at AstraZeneca in Sweden, running their inverse design program. He and I originally talked about all sorts of ways of thinking about materials design, and he very quickly adapted that into training neural networks. works. And that's sort of when, you know, I thought we were in the first sort of hype cycle, the first wave, but I think compared to what's going on right now, it was a tiny baby wave.

6:54I read in your paper that that was actually a class project or something.

6:59Heather Kulik:Yeah, yeah, that's right. You know, he just said, I have to do something for my homework. And that's how we got into it. I've also read in your paper that you've done a lot of work, like slightly more recently on active learning and using. Can you talk a little bit about that? Yeah, yeah. So even that polymer example I was giving, that would have been active learning in principle, but we sort of stopped after one generation because we had exhausted the space. But I think one of the areas where machine learning kind of just with what's out there right now has the most promising chemical sciences is in solving multidimensional challenges.

7:37Heather Kulik:So right now we're working on a project in metal organic frameworks where we're trying to solve trade-offs relevant for direct capture of CO2 from the air. And so in order to find a material that's good for that, we would worry about its cost, its stability in, say, aqueous humid environments, its ability to take in CO2 over other molecules, its mechanical stability. Is it going to hold up under force? Is it thermally stable? Can you heat it up and will it be okay? I'm just naming a few, but in total, right now in an active learning campaign, we're working on seven different objectives. And usually just even for a not so accurate machine learning model, you get, you know, at least a hundred to a thousand fold speed up for every dimension you're optimizing over.

8:31Heather Kulik:So the real promise is going to be in searching for that needle in a haystack with, say, seven objectives and doing something where you're not waiting for the models to be accurate before you start doing that optimization, that's really the promise of active learning. Yeah, that has an interesting parallel in my mind to the pharma world where you have a lot of computation work in the discovery process, but that actually getting it the drug out to in people's hands is often the bottleneck for a drug. And also, you know, what happens to the drug when it sits on the shelf for three months, that kind of thing.

9:11Heather Kulik:Yeah. Are these medical organic frameworks, what are the kind of things that they're used for? They're used most in gas storage sensing and separations. They're used in combination with polymer composites. They have really strong promise for CO2 capture especially, but people have looked at them for catalysis. The limitation on catalysis has been, you know, how stable are they? So one of the things we've spent a lot of time on is trying to be able to predict their stability. But they're used for all sorts of things, even drug delivery. You know, what they have the opportunity to do is really place precise chemical groups in specific orientations that can ultimately allow for what's known as host-guest interaction.

10:02Heather Kulik:So basically kind of create a glove to have a targeted interaction with a guest molecule in the metal-organic framework. I see. And just for the non-chemist metal-organic framework, Legos for chemistry? Yeah, yeah. Metal organic frameworks, I think, are going to be a little bit more of a household name among some engineers because the discoverers of those materials just won the Nobel Prize in chemistry this year. So as much as that can make something in chemistry a household name, but they're basically like Tinker Toys or Legos, and they have different building blocks that can be combined in basically infinite ways to create very precise chemistry.

10:47I see. Maybe for context, could we, could we step back in like, what are the techniques you were using before you started, or maybe in parallel with machine learning? And how does machine learning help you advance those? Like, what are the roles of the two?

11:01Heather Kulik:So I started my career studying what's known as transition metal catalysis. If you look at the periodic table, the middle of it contains a bunch of metals. A good example would be iron. And all of those things sitting in the middle of the periodic table, they have what's referred to as an open shell. So the electrons in those materials are not paired and they're not, they're as a result more reactive. Normally, like the way that you understand how they're going to behave, so that for instance, they give rise to, you know, Different combinations of these metals give rise to the catalysts that are used in a large number of transformations, including the things that, say, feed and sustain most of the world's population, such as the Haber-Bosch process for ammonia synthesis.

11:52Heather Kulik:And the way, going back 20, 30, 50 years, that people understood these materials and could enable their rational design is through quantum mechanical modeling. Quantum mechanical modeling, by using approximations to the Schroinger equation, it can be very accurate, but it's very computationally costly. And so a single quantum mechanical prediction, depending on the level of fidelity used, could take hours to days to weeks. And that's what I would have normally been doing before I got started in AI. Some of what we do these days is accelerating those quantum mechanical predictions, as well as looking at, you know, an area that I'm particularly excited about is that not all quantum mechanical approximations are equal.

12:40Heather Kulik:And you can actually use ML models to kind of predict what the best approximation to use is depending on the material studied. Is that like closer is better or is it not really distance related? In terms of which method is the right method to use? So it actually turns out to be quite complex. You can't just determine it from heuristics. So we actually, in one area, use the quantum mechanical wave function as inputs to neural networks to actually predict what is the right method to use and learn that mapping. I see. That's probably going to be in the best part. That sounds like a challenge. The cool 22 atom-ligand challenge.

13:23Go. I have a spicy question I want to ask. So there's a school of thought that why should I bother to learn chemistry or physics or whatever when to LGBT as PhD level understanding of that anyway? And shouldn't I just focus on being really good at using AI for stuff? So I want to hear your thoughts.

13:47Heather Kulik:My personal experience is that, and this will date itself immediately, is that chat GPT is super good at Wikipedia level chemistry knowledge. But one of my favorite things to actually throw at GPT as an anecdote is I'm really interested in molecular design. Like how do you find a new ligand that can go into a transition male complex? And what that means is that some combination of atoms and it's going to bind to the metal and it's going to change its properties. And so the thing I constantly do every time an LLM is updated is I just ask it, please design me a ligand that has 22 atoms. So the first time I've done that, there are many ligands out there that have 22 atoms.

14:34Heather Kulik:And then I say I want it to bind to the metal with two nitrogen atoms. I can never get an answer that has 22 atoms. So then you can try a range and see how many times you can get a range. And so that's maybe a trivial thing, but that's something that an expert chemist could do in a second. So there are really good introductions to chemistry that I think you can get through conversations with an LLM. You can get a lot of insight into an area you're unfamiliar with. And for sure, things have improved a lot. Like when I first tried typing in, you know, which exchange correlation functional should I use for this type of chemistry?

15:11Heather Kulik:The answers were completely wrong. They looked right, but they were completely wrong. I think things have gotten better because that knowledge is out there on the Internet. It's in the training data. But I think there's a lot of things that probably backing up a moment. You should learn chemistry well enough to know when when these models are right or wrong. And if you don't know any chemistry at all, it's hard to know if you're assessing correctly. But I think that there are a lot of things that you don't have time to do a deep dive into that you can now get from, say, an LLM that can augment knowledge.

15:48Heather Kulik:But I think you have to start from somewhere and then use it as a tool rather than starting from zero and relying blindly on what an LLM will say. But one of my favorite things, if someone can get in one shot an LLM to generate me a 22 atom ligand, I would love to see it. What do you think the biggest gaps that machine learning has from your experience? That like if you are an aspiring ML engineer with looking to take on a new problem from the machine learning side, what do you think someone could work on which would really help the chemistry side? There are a lot of challenges out there where the data sets aren't large enough or diverse enough, and so I think they've attracted less interest.

16:32Heather Kulik:So the ones closest to my heart are reactivity predictions, so predicting which reactions will occur and why, especially in complex phenomena, like in multiple elements and sort of predicting those transformations. another thing that i think um there's not enough data on is just more diverse chemical bonding and more diverse uh chemistry um for me that's transition metals but there's also questions of warm dense materials sort of exotic phenomenon we have really good data sets out there for really boring chemistry um so we have you know probably even if you're not a chemist you're familiar with organic molecule data sets and organic molecules binding to proteins.

17:22Heather Kulik:Those are the common data sets out there. There's lots of challenges out there where the physics is much more complex and the things like how does matter behave when you shine light on it and you excite it into excited states, all sorts of things like that receive relatively little attention because, you know, there may not be a benchmark or a leaderboard yet for that. And so maybe it's on osteochemists to generate more data sets so those leaderboards are out there. But there's definitely, you know, a lot of interest in chemistry for which there has been less attention. So in the protein world, there's CASP, right?

17:59And people have been working on that for a while, and this led to AlphaFold, like kind of without CASP, AlphaFold probably wouldn't exist. Is there like an equivalent to CASP in the material science world?

18:10Heather Kulik:So there are all sorts of repositories of fairly low fidelity DFT data on crystalline materials. So materials project, open catalyst project, these do provide good leaderboards. But some of the limitations that are the data comes from not very high fidelity density functional theory. So I'd say that's a second challenge is that we're all the smartest ML engineers right now are learning on data that is not going to be reflective of experiment. There aren't big experimental data sets, for example. One of the advantages of things like CAASPP is that it comes from an experimental ground truth, whereas that aspect just isn't available in materials as much.

18:57We talked about Casp and, you know, the role of Casp and AlphaFold. Do you think that there is like a problem, a way of phrasing this, that we could start collecting data at scale, that we could, you know, really have a community challenge, which breaks open some open problem in your mind? And maybe like, maybe actually even stepping back beyond that, what would you want to have if there was like an alpha fold for materials? What would you want it to do?

19:29Heather Kulik:One kind of murky area, so maybe I'm not going to directly answer this question. One murky area for us is electronic structure calculations are expensive, and they should, in principle, give you the right answer. They should, from first principles, give you the right answer of how a material is going to behave. And a lot of people are scaling these up right now with machine-learned interatomic potentials on training data. and every time someone comes out with kind of a new data set trained on a and they call it a foundation potential foundation model it looks really good and then you get it into your lab and you say okay i want to use it for this problem i'm really excited about and it starts doing kind of wacky things like molecules fall apart i won't name names but um there was one that made a huge splash this summer and people started declaring oh this method is dead this method is dead we're all going to just use these neural network models now.

20:32Heather Kulik:It's only in my hands, the one I'm still not naming is only about five times faster than my fastest DFT calculation on a GPU. And it also doesn't work all the time. So I would say we need a more transparent way of trying to figure out if these models can really replace conventional physics-based modeling. if they could if I could just give up ever doing a DFT calculation again and just rely on machine learning potentials and if they were you know two orders of magnitude faster than the traditional approach that would change that would change how we're doing science but there needs to be a little more rigor on what we consider you know just fitting data when that data maybe lacks quality or there needs to be a little bit tougher requirement for how we say this model can really replace the physics-based modeling.

21:30Yeah. So one of our pieces is that the interface between bits and atoms is really the bottleneck, right? Where you have to, the actually activity of trying things in the lab is the bottleneck. And you've addressed that to some extent in active learning. But I think that there's also an extent to which that just pure process and automation, good operational practice, those are important things. so that if you can push to automation on the one side, but on the other side, that creates brittleness. So how do you think about kind of bridging that gap to experimental chemistry and using that sort of as a nature's computer to figure out things for your design process?

22:23Heather Kulik:Yeah, so there are a lot of really smart people working in high-throughput synthesis and experimentation and autonomous labs. I think the thing that's interesting to me in that space, at least, is that there are some types of experiments that, at least as of the last conference I went to on this, are really hard for autonomous high-therapy experimentation, but are really easy for human and vice versa. And then there's, you know, the serendipity that a human might experience in the lab that a couple of people have tried to think about, like, well, how do you introduce that noise into high throughput experimentation?

23:05Heather Kulik:So I think that's a challenge. Your question also brought to mind another point that I am by no means an expert on. but most people who actually work on getting materials to the device scale say something that would be in your television or something like that is they will tell you that it's not just the material it's the process and I think we're at we're at ground zero we're nowhere when it comes to like well how do we machine learn not just the structure and the properties but also the role that processing plays. I don't think we know anything about how to do that. Maybe for non-experts, like with protein structure, it's really easy to imagine like, oh, you can see these proteins, and we can run some simulations and see them wiggling around.

23:51And the structures look really pretty. What does the data look like for material science? There's the computations, like DFT, I think, gives you something which looks like a crystal structure you can imagine. But then there's also like, is there experimental data where you can observe that crystal structure? or is this mostly sort of like kind of probes where you're measuring individual properties which are kind of collective and not fine-grained?

24:14Heather Kulik:So experimental structures are available, and the example I was giving is something we know is stable, and we've seen a structure of it before, and it will fall apart with some of these models. The challenge here is that what AlphaFold has done really well is predict structures of globular proteins, primarily with 20 natural amino acids. I could actually point to lots of cases where alpha-fold fails too for more interest in chemistry. The challenge is that you have a lot more than 20 building blocks when it comes to materials. And so there's lots of different ways to think about chemical bonding.

24:50Heather Kulik:And right now, no potentials are really robustly encoding all of that bonding, especially with respect to metal-organic bonding. Yeah, maybe a different way of saying it is like with alpha-fold, I mean, alpha-fold is solving ground state structures. It's not looking at dynamics, which is, I think, consistent with some of your statements about needing quantum mechanics for catalytic enzymes. So, but even, you're saying, even at just kind of ground state properties, you're saying that just there are too many parameters and there's not like a clear set of interactions, which is limited to a small number of building blocks.

25:27Heather Kulik:The bonding is highly variable across all of material space. Now there's simple regions of material space. You can pick aluminum. Aluminum is very boring and you can write down, people in the 60s could write down on paper, you know, how you need to model aluminum. That's something that is pretty easy to fit a neural network potential to. But then if you want to get over to iron oxide and then if you want to get over to high entropy alloys, there are definitely cases where people are using these methods but I'd say a big challenge is is that there's no real way to know if when you go to bigger land scales and time scales there's no real way to know if you're right or wrong the experimental data is not there experiment even interpreting say looking at an image of an experiment surface which you would want to do it requires some degree of an interpretation of that image so it's just it's just hard to know from experiment or from other computations if these types of models are correct.

Read the full transcript

26:31Heather Kulik:And they're certainly not correct across all of chemical space. And I'd say they could fail more catastrophically than AlphaFold obviously fails, though there are definitely failures of AlphaFold too. Switching gears a little bit, I read in your paper also that you had done some work with integrating textual information from papers and into your, so it's kind of the AI that we all know and love right now. Can you talk about what kind of lift that gives the models and how did you actually do that integration? Yeah, so we started, I guess, about five years ago. So when we first started doing it, we were just doing sort of standard natural language processing and graph digitization.

27:13Heather Kulik:These days we use LLMs, but just to try to extract from the literature data sets of properties, wherever people are widely reporting properties. And what we noticed is that there's a lot you can learn from these models. So you can, even on the scale of a few thousand data points, you can then do things like predict the temperature at which a moth will break apart based on experimental reports. But one of the funniest things I think we noticed is that you can get the temperature at which a material will break down two ways. One, you can get it from the graph. And two, you can get it from what the authors say about how they interpret the graph.

27:51Heather Kulik:And those two things do not line up. So people, you know, one of the challenges I think with literature extraction from papers is one would be the obvious mistakes people make, you know, no one's perfect. But the other would be just, you know, people interpret their results in different ways. And so if we're building models based on those interpretations, that's a challenge. In terms of LLMs, they've come a long way in terms of literature extraction, but they're still definitely sensitive to false positives. And I think the amount of time we spend checking on LLMs to make sure that the data we're ingesting is accurate definitely is an overhead on those types of workflows.

28:31I see. And what about the way that it might bias the discovery process, right? Because you have this known literature. Your job as a chemist kind of sorta is to find new stuff. But if, so if you're emphasized, if your computational method is pulling in literature, then maybe it's biasing you towards the previously reported results instead of something, you know.

28:54Heather Kulik:You know, one of the ways we try to address that is we try to train a model on that literature, but then apply it to new structures that have never been seen before and try to really look at how far we can extend the model. But we are trying to answer this in general. There are repositories out there of experimental data where you can have a sense of when it was published, what the structure is, what it was used for. And we're really trying to build generative models on top of that now to try to be able to say, well, if I know about the first 30 years of a field, can a model trained on that predict the next 20?

29:31Heather Kulik:I think that's an open question. and what model is best. And maybe it won't get all of them, but maybe some of those discoveries that we think are new in the most recent 20 years, maybe some of them are trivial for a model to generalize to, whereas others are not. I think in an ideal case where we have the available literature data and we don't know, we could use uncertainty quantification to then identify, okay, these would be the most interesting materials to get into our data set. I see. In those data sets, just for people who are interested in getting involved, What are some of the best ones to get started with?

30:06Heather Kulik:I don't know about the best. We've curated a few thousand data points of middle organic framework, thermal stability, as well as middle organic framework, activation stability, water stability. Other groups have curated other measures of stability. They're all out there. They're on our website. That kind of thing. Awesome. Do you imagine there being useful to create an initiative or a multi-institutional funding source or something which really is trying to get data in a high-throughput automated way? What would your dream be if you could organize something which really will drive the field forward in your mind?

30:49Heather Kulik:I think the National Science Foundation has one initiative. I've also heard about things with foundations before, sort of being interested in putting together cloud labs. So things that users can on demand make use of high throughput automation. I definitely think having user facilities where a computational researcher like me could design an experiment and have it executed would be awesome. having all that data collected in sort of a public way would be great you know the way that research right now gets published into papers it's very hard to then extract back out we spend a lot of energy trying to get it back out and so some of this is a need also for you know maybe systematization of how results get reported so that they can be machine learning ready from from day one when they're published.

31:45Heather Kulik:Some research subfields are trying to do that, but it's not really developed across material science. But for sure, you know, I think there will be more sort of shared facilities where people can make use of data from high-throbot experimentation, and that would be really, really awesome. I don't know if it'll come from companies donating equipment from National Science Foundation or from, you know, private foundations. Yeah, there is a large, like, philanthropic push in the biotech space, it seems like people haven't quite picked up on this as such an important field, like especially with things like materials for climate change.

32:24You can imagine in particular a very important problem that we could use a lot of cushion. Yeah, that kind of brings up the question. There's been a ton of very recent materials investment for private companies, startups. Where does that leave in your mind, the role of the academic income I see?

32:45Heather Kulik:I ask myself that all the time, or more recently in the past year. So in particular, there's a lot of compute that companies have access to that academics don't. So I ask myself, you know, what can we do that's more creative, that doesn't require just brute force compute. And I think there is, there is like a lot of stuff that we can still do, but we have to ask those questions. For sure, Microsoft, Meta, those, those ones are kind of like the companies that have basically infinite resources. And as an academic, I don't have infinite resources, you know, but we have an interest in problems that, you know, haven't crossed the radar of those companies yet.

33:34Heather Kulik:And I think as long as we, you know, whenever someone poses a problem to me now versus a few years ago, I try to make sure that we're not just in the process of trying to do something that throwing a lot of compute at it would solve it. Yeah. I think we're kind of running out of time, but would like to give you an opportunity, call to action. What would you like our listeners to know about, do? What should they do to get involved or something that you're really passionate about? I think I will stick to something kind of niche. So I think there is still a place for chemistry. I will say that. But my group develops a code for transition mal complex structure generation metal organic framework screening.

34:21Heather Kulik:It's called MOL Simplifier. When we're working on MOFs, we call it MOF Simplify. There's website versions of it that you can look up and not install anything, but it's also on Conda and GitHub. And if you do have an interest in transition mal complexes, you know, just try it out. It includes machine learning predictions, but it also makes novel structures. And I'm just really interested to hear ever if people are using it. I know a lot of companies are using it, but we sort of find out sort of after the fact. So if you're interested more in this material space, I'm definitely interested and open to feedback.

34:56Grateful. Awesome. Getting involved. Thank you very much. Take care, doctor. Thank you.

From the publisher

Materials science is the unsung hero of the science world. Behind every physical product you interact was decades of research into getting the properties of materials just right. Your gym clothes contain synthetic fibers developed over decades. The glass screen, diodes, and chip substrate technology needed to read this blog post were only viable due to many teams of material scientists.

Our guest Prof. Heather Kulik was one of the first material scientists to realize that there was alpha in combining computational tools with data driven modeling β€” she did AI for science before it was cool. She has a hard-fought perspective for how to succeed in this field. Yes, she believes the wins are real. To get there you must work hard to deeply integrate domain expertise with AI techniques, and also maintain a discriminating mind. Ultimately what matters is you succeed in the lab, and nature doesn’t care about how hyped a model is. These lessons personally resonated with the Latent.Space Science team and our own experience.

This episode is a must watch for all aspiring AI for science practitioners. A few highlights:

Designing new polymers with AI: Heather’s group recently used AI to design new polymers that are significantly stronger. These materials were created and tested in the lab, and the scientists who built them were surprised by the designs. The AI had figured out certain building blocks could break in a novel way. The AI discovered a purely quantum mechanical effect, and after convincing their lab collaborators to actually synthesize it, the material turned out to be four times tougher!

The twenty-two-atom ligand challenge: When asked about the role and need of human scientists, Heather points out that AI has a strong understanding of academic chemistry, but is still lacking intuition. Every time an LLM is updated, Heather asks it to design a ligand that contains exactly twenty-two heavy atoms. She has yet to find one that can succeed at this seemingly simple task that any expert could do in a second! Is this the chemistry counterpart to counting β€˜r’s in strawberry?

Side note: Heather joked that this comment would date itself immediately, so we decided to see if this was still true three months after recording. We found some interesting results! We asked both Claude and ChatGPT to design a 22 atom ligand for both a metal-organic framework (MOF) and a Kinase protein.

* For the Kinase, both models got it right: Claude pulled out RDKit in a python script and iterated on several designs, whereas ChatGPT just one-shotted it.

* For MOFs, both models got it wrong, generating ligands with 21, 23, or 24 atoms, yet stubbornly not getting 22 atoms.

Is there something different about how LLMs reason in the materials and bio domains?

Materials vs biology: The two biggest domains of AI in science have been biology and materials. We asked Heather if there could be an AlphaFold moment for materials. Her answer reframes how we should think about the field:

* First, the datasets in material science are woefully lacking in comparison to the bio world. The closest to ground truth in most cases are noisy DFT datasets. These are just approximations to the real world! The datasets that are accurate are all boring, as Heather quipped β€œWe have really good datasets for really boring chemistry.” Furthermore, good experimental structures are hard to come by and require interpretation. So generating generating high-quality, novel datasets at scale would really drive the field forward.

* More philosophically, AlphaFold is making predictions in a fairly limited space: there are just twenty amino acids. Sure, even here AlphaFold doesn’t get everything right, but it seems plausible that one could learn the entire design space. For materials, each element is a new set of interactions and chemistry, with little to no transferability. This is a massive open problem in material science that we hope some of the smartest AI scientists will want to work on!

The difficulties of trusting the literature: Heather’s team has spent the last few years using NLP and later LLMs to extract data from literature. Even a few thousand data points from these papers can be valuable for guiding her group’s work. One surprising result: sometimes the reported values for a property (say temperature) do not match up with the graphs in the papers! So there’s lots of potential in using LLMs to mine data from the literature, just do it with care.

The role of academia in an ever-changing world: One theme that has been running through many of our conversations has been the changing role of the academic β€” and the scientist β€” in science. When startups are raising $100s of millions and hyperscalers and Big Pharma are all ramping up AI-for-science efforts, the academic researcher needs both resources and judgement about problems to chase more than ever.

Resources include data that is organized for machine learning, access to high throughput experimentation labs, and compute resources. These are all things that academics can build together. More importantly, Heather emphasizes curiosity about problems that haven’t hit the radar of the heavily capitalized AI companies. After so many years on the forefront of AI for Science, Heather’s judgement that Chemical Engineering and Material Science still need curious people asking questions with no clear path to money is a welcome beacon in the AI fog.

Full Video podcast

Is on Youtube!



This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

More from Latent Space: The AI Engineer Podcast

All 247 episodes
πŸ”¬Why There Is No "AlphaFold for Materials" β€” AI for Materials Discovery with Heather KulikLatent Space: The AI Engineer Podcast Β· 35 min
Listen in VO