๐Ÿ”ฌ Automating Science: World Models, Scientific Taste, Agent Loops โ€” Andrew White

28 Jan 2026 ยท 1 h 14 min ยท 39 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT ยท Add to Claude

In short

Episode Summary: Automating Science: World Models, Scientific Taste, Agent Loops โ€” Andrew White

Podcast Overview The Latent Space podcast is aimed at AI engineers, providing insights into the latest developments in the field, including foundational models, AI agents, and more. In this episode, co-hosts RJ Honicky and Brandon Anderson welcome Andrew White, a prominent figure in AI for science, to discuss his journey and the implications of AI in scientific discovery.

Episode Highlights

Introduction to Andrew White

  • Former academic turned entrepreneur, Andrew White co-founded Future House and Edison Scientific.
  • His work focuses on integrating AI into scientific research, particularly in automating the scientific method.

Key Topics Discussed

  1. The ChemCrow Story
  2. ChemCrow, an AI agent using GPT-4 and cloud lab automation, raised alarms about the potential misuse of AI in creating bioweapons and led to significant governmental attention, including presentations to the White House.
  1. Scientific Taste as a Frontier
  2. Discusses the challenges of Reinforcement Learning from Human Feedback (RLHF) in evaluating hypotheses.
  3. Emphasizes the need for more nuanced feedback mechanisms that account for human preferences in scientific discovery.
  1. Kosmos: The Autonomous Research System
  2. Kosmos serves as a full scientific agent that integrates literature search, data analysis, and experimental design, evolving its world model based on findings.
  3. Highlights the importance of including data analysis to refine hypotheses rather than relying solely on literature.
  1. Critiques of Molecular Dynamics and DFT
  2. Argues that these methodologies are overrated, stating they don't accurately model complex real-world behaviors, such as reactions in catalysts.
  3. Points to AlphaFold's success as a paradigm shift, indicating machine learning's ability to outperform traditional simulation methods.
  1. The E3 Zero Reward Hacking Story
  2. Discusses the complexities of creating molecules with specific atom counts, leading to unexpected outcomes, including the generation of unfeasible nitrogen compounds.
  3. Highlights the absurdity of the AIโ€™s creative ways to exploit reward systems in generating chemical hypotheses.

Future Directions for AI in Science

  • White anticipates that scientists will transition from traditional roles to a form of "agent wranglers," utilizing sophisticated AI tools to enhance the speed and efficiency of scientific discovery.
  • Discusses Jevons Paradox, suggesting that automating science could lead to an increase in scientific inquiry rather than job displacement.

The Role of Human Intuition in Science

  • Despite advances in AI, White acknowledges that human scientific taste and context are irreplaceable, emphasizing the need for human oversight in interpreting AI-generated hypotheses.
  • Expresses the challenge of balancing AI's capabilities with the nuanced understanding that human scientists provide.

Key Takeaways

  • AI's Impact: AI has the potential to significantly enhance the efficiency and scope of scientific research.
  • Need for Human Insight: The integration of human intuition and scientific taste is crucial in interpreting AI outputs.
  • Continued Exploration: The future of science will likely involve increased collaboration between human scientists and AI tools, leading to a more rapid pace of discovery.

Conclusion Andrew White's insights provide a thought-provoking perspective on the intersection of AI and scientific research. The discussion highlights both the opportunities and challenges posed by AI in automating the scientific method, emphasizing the enduring importance of human insight in navigating the complexities of scientific discovery.

For more information and detailed notes, you can visit the Latent Space website at [latent.space](https://latent.space).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Limits of Molecular Dynamics

0:00 to 0:57

Discover the challenges faced by MD in protein folding and a surprising breakthrough.

โ€œMD was supposed to be the protein folding solution.โ€

Andrew White: From Academia to Startups

2:17 to 3:00

Learn about Andrew White's transition from a professor to co-founding startups.

โ€œWe're going to get into all those points.โ€

Research on Non-Fouling Materials

3:00 to 4:33

Explore Andrew's PhD research on non-fouling materials and its implications.

โ€œAnd so the goal of my PhD was trying to find what are called non-fouling materials.โ€

Merging Simulations with Experiments

4:33 to 7:20

Understand the challenges of combining simulations with laboratory experiments.

โ€œYeah, it's kind of like some rejection is like immune-based.โ€

The Influence of Machine Learning

7:20 to 9:24

Find out how machine learning began to impact research in chemistry and materials.

โ€œSo I'm writing a book about like how you can apply these methods in chemistry.โ€

AI's Role in Chemistry Research

9:24 to 10:29

Discover how Andrew's work with GPT-4 influenced chemistry and AI's impact on science.

Automating Science: Founding Edison

10:29 to 12:37

Learn about the founding of Edison and the vision for automating scientific research.

โ€œthe president on their schedule for a 30-minute block.โ€

The Future of AI and Science

12:37 to 14:06

Explore the potential challenges and progress in the future of AI-driven science.

โ€œBut then I did resign my tenure position when we co-founded Edison.โ€

The Challenges of Automating Science

14:06 to 16:46

Explore the difficulties and progress in automating scientific discovery.

โ€œLike I'm, I feel like I'm always miscalibrated in this domain, but it's always hard to predict progress.โ€

Defining Automation in Scientific Discovery

16:46 to 19:52

Learn about the cognitive processes involved in scientific discovery and automation.

โ€œIt means like putting all the papers in one spot, like getting APIs wrapped around everything.โ€
Show all 39 chapters

The Role of Human Preferences in Science

19:52 to 24:12

Understand the impact of human taste and preferences on scientific research.

โ€œIt's like a broad category of all these things.โ€

Learning from Experiments and Results

24:12 to 28:00

Discover how testing hypotheses can lead to unexpected outcomes in research.

โ€œribosutal was a very good medicine and had a mechanism that I think is novel, although there was lots of debate on X because I think in 2012, there was a master's thesis which proposed this mechanism on like page 38.โ€

Tiling Trees and Hypothesis Evaluation

28:00 to 28:40

Exploring the brute force method of tiling trees and its impact on hypothesis testing.

โ€œI need to have like, I don't know, some kind of substrate.โ€

LLMs in Hypothesis Filtering

28:40 to 29:30

Discussion on the role of large language models in filtering scientific hypotheses.

โ€œYeah, but there's a lot of gotchas and I think people can miss those, but I think they're actually pretty good.โ€

Challenges of Data Interpretation

29:30 to 30:20

Understanding the complexities and biases in data interpretation and context.

โ€œI can also think from my own life, multiple cases where the data in some sense was there and you had two people who were both experts and very smart people who looked at it and drew very different interpretations.โ€

Bixbench and Human Disagreement in Analysis

30:20 to 31:20

Analyzing the bioinformatics benchmark Bixbench and human consensus in data analysis.

โ€œIt's in some frontier alums when they release their system card, they'll mention Bixbench.โ€

The Dark Arts of Medicinal Chemistry

31:20 to 33:00

Exploration of biases and superstitions affecting medicinal chemistry practices.

โ€œThat is like a spot where like, you know, there's so much superstition.โ€

World Models and Scientific Agents

33:00 to 35:00

Introduction to the concept of world models in scientific agents and their implementation.

โ€œSo I glanced at the paper and one of the things that jumps out is that there were certain class of problems for which it was only 50 some percent accurate.โ€

Development of Cosmos

35:00 to 36:20

The iterative process of developing Cosmos for automating scientific discovery.

โ€œI'm probably some fancy word for it, but I'm like a Lego guy or something.โ€

Agent Interaction and Scientific Discovery

36:20 to 38:20

Understanding how different agents work together in the scientific discovery process.

โ€œand then go like uh come up with experiments that you could do in a wet lab yeah and this is our inventory list and then go analyze all the data then go back and repeat the process right so that's like what Robin was.โ€

Calibrating the World Model

38:20 to 39:40

How the world model serves as a memory system for scientific analysis.

โ€œAnd so in Cosmos, we basically, we had all the pieces sitting around.โ€

Critique of Molecular Dynamics and DFT

39:40 to 40:10

Challenging the assumptions behind molecular dynamics and DFT methods in research.

โ€œbetween a Git repo and what a world model is.โ€

Limitations of Simulation in Science

40:10 to 42:01

A discussion on the limitations and misconceptions surrounding simulations in scientific research.

โ€œhelp you guys pump up your views here so i i think molecular dynamics is overrated in fact coming from someone.โ€

Challenges of DFT and MD Simulations

42:01 to 43:18

Discusses the limitations of Density Functional Theory (DFT) and Molecular Dynamics (MD) in simulating complex materials.

โ€œLike in DFT, you simulate water at 330 Kelvin when you want room temperature water.โ€

The Rise of AlphaFold and Protein Folding

43:19 to 44:42

Explores the impact of AlphaFold on protein folding and the shift in computational biology.

โ€œThere's a, I don't know what there's a word, but the counterfactual is basically a group called DESRES, D.E.โ€

Surprising Efficiency in Protein Folding

44:43 to 46:20

Highlights the unexpected efficiency of protein folding models using experimental data.

โ€œThat's what DeepMind did is they took X-ray crystallography data.โ€

Risks of AI in Chemical and Biological Domains

46:21 to 48:22

Evaluates the risks associated with using AI in potentially harmful chemical or biological applications.

โ€œI think I guess the short answer is that there is very.โ€

The Role of AI in Material Procurement and Safety

48:23 to 50:48

Discusses how AI might influence the procurement of materials and its implications for safety protocols.

โ€œBut then I think now is the next frontier is like, can it somehow help you with real-time protocols, troubleshooting more in the loop and more, especially in the computational side of things?โ€

The Future of Science and AI Automation

50:49 to 53:02

Speculates on how automation will change the landscape of scientific research and the role of scientists.

โ€œAnd so I think there's a lot of like weird second order things that we don't pay attention to in AI safety.โ€

Perceptions of Scientific Roles Amidst Automation

53:03 to 55:49

Considers the evolving role of scientists in a world where AI increasingly contributes to research.

โ€œMaybe there'll be somewhat an increase, but like there is a finite amount of like time people will be spending in cars.โ€

The Role of Humans in Science

56:00 to 57:20

Discussion on the necessity of human involvement in scientific processes.

โ€œAnd so I think if they're going to be consumers of science, they're also going to be some of the producers who are involved in the process by itself.โ€

Exploration and Bias in AI Science

57:20 to 59:30

Exploration of the biases in AI-driven science and the importance of human exploration.

โ€œDon't want to be contrarian, but yeah, be contrary.โ€

Natural Language as the Future of Chemistry

59:30 to 1:01:20

The discussion revolves around the belief that natural language will bridge chemistry concepts.

โ€œSo we think Cosmos is, we think it's great, but there's a very large amount of area for improvement.โ€

Limitations of Language in Science

1:01:20 to 1:04:30

Exploration of the challenges of describing complex scientific concepts using language.

โ€œAnd the only way to bridge all that information is natural language.โ€

The EtherZero Project: Lessons Learned

1:04:30 to 1:10:02

A humorous recount of the EtherZero project and its unexpected discoveries in chemistry.

โ€œthey have like, you know, relativistic effects because it's like there's a whole bunch of stuff around it.โ€

Discussion on Nitrogen Compounds

1:10:02 to 1:10:52

Learn about the challenges and breakthroughs in synthesizing complex nitrogen compounds.

โ€œLike, look, Andrew, like it's not actually impossible.โ€

Creative Model Approaches in Chemistry

1:10:52 to 1:11:56

Discover how AI models can creatively approach chemical reactions and the challenges faced.

Challenges in Training AI Models

1:11:56 to 1:12:41

Understand the difficulties in training AI models with verifiable rewards and data handling.

Insights into Engineering and Chemistry

1:12:41 to 1:13:33

Gain insights into the engineering aspects and nomenclature of medicinal chemistry.

โ€œI actually, I used to know all the names of these modifications, but I think it's like DAPO is one modification and like the clipping we did was special.โ€
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00MD was supposed to be the protein folding solution. There is a great counterexample. The counterfactual is basically a group called DESRES, D.E. Shaw Research. They had similar funding to DeepMind, probably more actually. They tested the hypothesis to death that MD could fold proteins. They built their own silicon. They built their own clusters. They had them taped out all themselves. They burned into the silicon the algorithms to run MD. They ran MD at huge speeds, huge scales. I remember David Shaw came to a conference once on MD and he flew in by helicopter and this pretty famous guy, kind of rich.

0:38And he gave an amazing presentation about the special computers and special room and outside of Times Square and like what they can do with it. It was beautiful, amazing. And I always thought that protein folding would be solved by them, but it would require a special machine. Maybe the government would buy like five of these things and we could fold, you know, maybe one protein a day or two proteins a day. And when AlphaFold came out and it's like, you can do it in Google CoLab, you know, or on a GPU or desktop, it was so mind blowing. I forget like that protein folding was solved. I always thought that was inevitable.

1:07But the fact that it was solved and on like your desktop, you can do it was just completely floored, changed everything. This is the first episode of the new AI for Science podcast on the Lease in Space Network. I'm Brandon. I work on RNA therapeutics using machine learning at Atomic AI. My name is RJ Haneke. I'm the co-founder of Mira Omics, where we build spatial transcriptomics AI models. The point of this podcast is to bring together AI engineers and scientists or bring together the two communities. These are two communities which have been developed independently for quite some time, but there's been some attempt to combine them.

1:43And only now, after many years, are we starting to see some of the big developments start to play out in the real world and start to solve key scientific problems. There's no like one size fits all solution. You need domain expertise. You need people on both sides of the aisle who can really talk to each other and really work together and understand both the modeling and all of the real subtleties of the system you're actually trying to work on. We hope that we can connect these communities and that we can provide a starting point for this new era of AI and science to move forward. So without further ado, let's get started on the first podcast.

2:21we're really happy to have in the studio today andrew white co-founder of future house and newly formed startup edison scientific um rather than introduce him i'll let him introduce himself uh hey i'm andrew from san francisco former professor now running two startups uh one that's a non-profit research lab and one that's a for-profit venture-backed company and we're trying to automate science. We're going to get into all those points. I'm really happy to be here. Thanks for having me on. I want to know personally about jump from academia to industry and quasi industry. So I would love to hear that story.

3:02Yes. I guess that's the whole story, right? So I did my PhD at University of Washington and I worked in a group with I think 19 people doing experiments and like two people doing simulations and i was working on a topic uh called molecular dynamics um which i think is actually suddenly becoming interesting again as everyone's looking for ways to generate data from first principle simulation and molecular dynamics you know covers uh basically everything that's molecules moving around and dynamic systems so like biology of course the complement in material sciences things like density functional theory where you You can model chemical reactions in these like solids.

3:43So I was working on that. We were going to biomaterials. And so the goal of my PhD was trying to find what are called non-fouling materials. So in biological systems, whenever you put like a foreign object into the body, it will trigger a response. And that response called the foreign body response basically encapsulates it in like this layer of collagen. This actually is exploited for some implants. like if you get a heart uh sorry pacemaker installed like it coats it with this collagen so that if you go to change the battery you can almost change the battery out like without even bleeding because like the body has like completely encased and this is great for pacemakers but for like a glucose sensor or like a you know brain cognitive interface bci is what they call it now yeah they're it's not so great and so that's why some of those things have like a limited lifetime because eventually your body treats it like a wound and heals.

4:33Rejects it. Yeah, it's kind of like some rejection is like immune-based. Okay. And so that's where like if the body can see anything on it, like if it can see like some ligand that it combined to with antibodies, then you get this like inflammation, which is like a rejection response you see in organ transplants. But with materials, the body's just like, oh, there's just like a wound or there's just something here and it just covers it up. I think the research in that field has gone on a long time since I left my PhD. And there was a lot of theories about it's related to the mechanical properties of the material.

5:05Like if it's spongy, there's things like if it's trabecular, like it has a bunch of little pores in it. We worked on the theory that it had to do with how hydrophilic the material was. But anyway, so I was the only one working on computers in this group. I couldn't figure out how to connect what's on the computer with what's done in the lab. because you can make like a simulation of whatever 10 ,000 particles, 10 ,000 atoms. It's like, well, this is not going to model the human body. It's a lot more atoms involved. So I had a good time. We did some cool stuff, some bioinformatic stuff. I learned a lot.

5:39But then when I did my postdoc, I was like, okay, we're going to try to merge experiments and simulations. So I worked on this theory called maximum entropy. And it's about like, how do you take complex simulations and match them to limited observations? and it's like the inverse of machine learning machine learning is like you have simple models you're going to a lot of data where i had like complicated models i'm trying to fit to very little data yeah it was fine it's great we wrote some papers it was useful and then i wrote like i started my research group at university of rochester on applying these methods to model peptides yeah i'm always like too early for things we studied peptides for i don't know four or five years and it was a cool niche field not that popular now peptides are like the hottest thing ever i think there's even like a peptide rave i heard about a couple yeah weeks ago but but when i was an assistant professor nobody cared about peptides um so we worked a lot on different ways to combine them we looked at like uh different experimental methods that we could do these molecular dynamic simulations peptides and then in 2019 i was like out on a sabbatical at ucla um they have a place called the institute of pure and applied mathematics there which is like this institute where people can go and do a sabbatical and learn new methods and they were happening to be doing they happen to be doing um machine learning for physics i think the name of it was like some like symmetric thing it's like machine learning for physics and physics of machine learning okay it's kind of a cool yeah yeah concept right but like jan lecun was there and like um frank noe was there who's a big guy in europe in in this field i don't know it's just terence tau came by it's really great group yeah and everyone's kind of jamming it was like 2019 so like they're not really been the big hit um especially in in uh in non-computer science fields.

7:19And then I came back from that and I was like, well, I got to teach class on this. So I'm writing a book about like how you can apply these methods in chemistry. It was very kind of niche field because every machine learning class that my PhD students could take at the time, this is when I was a professor at University of Rochester, it was always would end in like, okay, this is an RNN and this is like what you need to know, or like, this is how you do image classification. But in chemistry, it's all about graphs, right? It's all about how do you represent these graph structures it's all about symmetry and geometry yeah and that was like not a thing it was very popular but you had max welling on before and yeah the godfather of geometric deep learning yeah so i wrote this textbook about like this these methods and there was a bunch of interesting mathematics to it i had a good time and stuff and then um uh i think i was following you know the news in the space and codex the original codex came out and um i had been looking at transformers for a while i just tinkered with them and we started trying them on doing some chemistry tasks we were really impressed actually and we wrote a benchmark and this is like 2019 we wrote a benchmark of verifiable rewards in 2019 maybe it's 2020 by then but like yeah here's the head of the curve like here's a here's a function and um uh sorry there's like a task which is like i have a body of a function for like a markov chain monte carlo simulation it's missing some pieces complete it and then we had like a verifier that would see is it a valid mcmc simulation yeah yeah we wrote this paper ended up coming out i think 2021 2022 um because it took a long time to bank enough questions but i wrote an opinion piece about how transformers could change um how we think about chemistry and things like this and how we teach it and then opening eye some people there um lama uh was there she saw this paper and they reached out like hey we're building this new model and we think it'd be great to red team it to see like what could happen with these models if they're applied to chemistry or biology and so i was a red teamer for gpt-4 and i was using it like nine months before release was august so gpt-4 came out in march and i was using in august yeah and then like the react miracle paper came out um i think shenyu he wrote that paper and i plugged it in with gpt-4 like in the fall you know and i was like wow there's so much stuff coming out with react yeah and it was really exciting and then so when gbd4 came out i released this paper called chem crow i work with philippe schwaller and in um switzerland on this and ibm so that was like react applied to chemistry yeah and what we had is we had like there's a cloud lab that ibm built in switzerland yeah so we had like gbd4 operating the cloud lab and then it was like i had written a literature research agent that did like agentic rag again nobody knew what agent to grag was um i think actually uh harrison chase had like written a blog post about some ideas there and so i i stole some of the ideas really smart guy um and basically we applied that and we saw some really cool stuff it was really exciting and then i you know we wrote the paper it set off this crazy storm of like everyone was had a lot of anxiety about ai progress yeah and I ended up visiting the White House.

10:23I guess my paper was like the only time a preprint or peer-reviewed paper was presented to the president on their schedule for a 30-minute block. And the National Security Advisor at the time, Jake, God, I was confused. No, sorry. One of them is a talk show host, and one of them is like the National Security Advisor. I forget which is which. That guy. That guy, yeah. He like had a presentation about our paper, and they presented it to all the, because there was a big tech CEO summit at this time where they sent out Sam Altman and some other CEOs out there. This is the future of chemistry is language or a different one?

10:57This is the ChemCrow papers. Oh, ChemCrow. That's right. Yeah, yeah. Sorry. I probably should name these things. Yeah, yeah. And so it was crazy. And they had me go out there and then I met a lot of three-letter agencies I didn't really want to meet. And I'm like, you know, there's like somebody from one three-letter agency is like, how does this change explosives? You know, another three-letter agency is like, how does it change breakout time for nuclear weapons research? Yeah, yeah. I was like, guys, I'm not really sure. But it turns out that there was not that many people, world experts on AI and science.

11:28Right. So what's the answer? Yeah. Great. Good question. We'll come back to that. Okay. Yeah, let's come back to that. In the end, I had a lot of energy, a lot of excitement about this area. So I took a sabbatical from the University of Rochester. And it was Sam Rodriguez. And Sam had been talking to Eric Schmidt and Tom Khalil, who was also a national security counsel at the Obama administration, about how to scale up these ideas. And so Sam had this concept of focused research organizations, which is how do you do science not in academia, not in one of these kind of near monopoly tech companies, these big labs.

12:08and want to try this idea out. And I was like, hey, we should do this around agents for science or AI for science. I love Sam. He pushes me to come up with really lofty ambitions. So we decided to automate science as the goal instead of like, see what fun stuff we could do with agents and science. But I think that was maybe the real mission. But of course, automating science is the long-term mission. Yes. And so that was what led to Future House. And it was a very long-winded... Yeah. No, no, that's great. So you chose to leave a tenure track position to do this. I was on sabbatical, which is a beautiful concept.

12:40But then I did resign my tenure position when we co-founded Edison. Yeah. Okay. And I think it was, you know, I had been on sabbatical for a very long period of time. And so at a certain point, I just had to resign my tenure. So I resigned tenure in June. Okay. Oh, so that's only recently. Yeah, only recently. And you just felt like this is the direction of my career. Yeah. Yeah. I mean, I got tenure and I had these early career awards, like the NSF career award. It was great. And I think academia is really exciting. But I just thought that right now, this kind of area, like A, ever science is just, I think, A, difficult to do in academia and B, so exciting.

13:20But I think you can take bigger bets. And I think having a tenured position and writing research grants is maybe not the biggest bet you can take on a field. Yeah. So now we have a venture-backed startup called Edison, which we spun out of Future House. And we took a lot of the ideas and we're trying to do this at an even bigger scale right now. And so Edison was always kind of the plan. Going back to Sam's idea of a FRO or a fundamental research organization, he always had this goal of let's do fundamental research in this tightly scoped nonprofit, which can kind of explore. And then you have that as a natural arm for spinning off.

13:59yeah um you know venture backed yeah i think that's right um i think some things that like make that not as clean these days is how expensive ai research is and how expensive gpus are so i don't think we can repeat it many times from future house it might be like an n of one thing right now it just maybe not i don't know if venture capital keeps growing then maybe i mean maybe we can but yeah i think we took a lot of the ideas of future house another thing like i think we expected it to be harder to automate science. And actually it's really hard. Like I'm, I feel like I'm always miscalibrated in this domain, but it's always hard to predict progress.

14:36Um, and I think that I overestimate the speed of things on month scale and I underestimate things on year scale. So the two years from like 2023 to 2025 was an enormous amount of progress. Yeah. Um, and always felt like things were not going as fast as I thought, but when you look back on it, wow, like there's a lot of progress. And so I think the idea that I think in future house and our sam actually regrets us writing this but in the original marketing or like the announcement is like it's our 10-year mission automate science and like now it's like okay yeah yeah so two years later we had cosmos and um things are going so much faster and it also this kind of this this thing you notice it in um in san francisco is like it's actually kind of hard to find problems which are like so hard that like they are a challenge for language models but not so hard that they're impossible.

15:27And we're in this gray zone. And actually, I feel like that's where we are right now is that we can actually automate so much of the scientific method because it turns out, especially in a field like biology, which is very empirical limited, the top 1 % guesser of what they think will happen in an experiment, the top quintile or quartile, they're about equal. And so even if we wait 10 years and get even smarter models, I don't think it's going to really change the fact that we're ready to automate a lot of science with existing LLMs. What do you mean by automate science? That's like a pretty loaded state of brains.

15:57There's lots of things. There's many ways of thinking about that. So we try to draw a line between what I would call groups that are trying to model something like the cell or how proteins fold or like how antibodies can be designed or like maybe the virtual cells. Like an example, if they're trying to use machine learning or AI to model some very specific system, we're trying to automate the like cognitive process of scientific discovery. making hypotheses, choosing experiments to do, analyzing the results from experiments and using it to update your hypotheses or your confidence in those hypotheses.

16:29And then leading to like a world model, which of like, okay, this is how I understand this process to be. And then that, you know, begets new hypotheses or new experiments. We want to automate that sort of loop. We thought that we would have to build up like a whole new organization from the ground up for agents. So it means like automated labs. It means like putting all the papers in one spot, like getting APIs wrapped around everything. But over time, like the models have gotten better and better that we had to like, you know, stop and rethink. We don't actually have to hold our hands so much anymore or like they don't actually need to necessarily have an automated lab.

17:01They can like write an email to a CRO or something or they can like tell you what experiment to do and you can take a video of you doing it and then show it to the model and they can be like, okay, well, this is, you know, what happened. So it's been a really interesting experience of like sometimes, you know, we over-engineer things and sometimes actually basically just mostly over-introoted. So I always think about systems and scientific processes as a system. I always think of it in terms of constraints, right? And like what is a bottleneck in the system? So what is your hypothesis about this, right?

17:34Like in my mind, not knowing a ton, but in my mind the constraint of the scientific process is the work you do in the lab. And that's sort of notably missing from, Well, not entirely. You mentioned automating lab and whatever. So how are you thinking about this? Yeah, I think you're right. Basically, the best model, whatever, Opus 7 or GPT-10, it really can only propose the first experiment, maybe slightly more clever. But at a certain point, you just need information, right? Like some little calculations you can do that there's more atoms in the brain than you could ever simulate even if you had all the energy from the sun, right?

18:14Like I think it seems maybe a thousand brains in real time with all the energy in the sun because just too much information. Yeah. So science really hits these like bottlenecks where you just actually have to go measure things. Yeah. We definitely think about maybe like lab in the loop sort of situations. Like one of our papers, which was called Robin, is that we like had one of our agents propose an experiment. We did the experiment and then we had our agent analyze the experiment and propose the next experiment. Yeah. And that kind of loop I think is where you want to get to. Yeah. So what is the bottleneck in that?

18:41I don't think it's like the intelligence of the first experiment. I think the bottleneck might be something like right now. I think the bottleneck is something silly, like knowing what's the lead time on all the reagents that you need and what is available in the lab. Like, you know, I think whether GPT 5.2 Codex Max or Opus 4.5 is going to do better is probably doesn't matter. It's just a matter of like, which one's going to have all the information about what's in the lab and how much will it cost, how long will it take. And also, I guess the kind of frontier that I think about for these models is taste, which is like a lot of science.

19:15I mean, of course, we want to accelerate technology. We want to improve the economy. We want to improve people's life expectancies. We want everyone to be happier. But a lot of what is done in science is based around like human preferences. Like why do people study, I don't know, a particular worm? Well, like there is a theory that by studying the worm, it has led to good medicines or it's led to discovering new genes, but also people studied in the past, people's careers depend on that worm and people want to write papers about that worm. And so there's a human element to some of this. And I think that models don't capture that so well about knowing what is an exciting result and what is a boring result.

19:50I see. So I think that's like a scientific taste. It's like a broad category of all these things. How do you, like, do you try to quantify taste in any way? I mean, I know that I have some, like, fun anecdotes about this, but maybe, yeah, just like hear what you... Yeah, actually, we sat on this idea. really sad enough but we like argued about it for a long time sam and i usually every monday morning at eight o 'clock in the morning sam and i meet and we're both you know caffeinated and ready and we argue about stuff like this and we had a lot of mondays where we talked about scientific taste and in the end we're like okay let's just let's just do the dumbest thing which is to like have our agents make hypotheses and put them in front of humans and have them be like i like this one or like that one right so we just did like whatever rlhf on on on hypotheses and we learned a lot about how bad our LHF is with people.

20:39Just like people pay really attention to the tone, to the details, to like how many specific facts or figures on the hypothesis, right? Like actionability about like if the experiment is feasible. But what people didn't really pay attention to is like, I don't know how to describe it, but like if this hypothesis is true, how does it change the world? If the hypothesis is false, how does it change the world? This like, how much information do you gain? It's not really information, but like impact or something. and that really didn't come through from those things so they're like okay well this is maybe one strategy and so we had to go back and think about it more and um we then took a pause from that research and then we made cosmos and then cosmos like has baked into it taste right like at the end of the day there will be some report and we're working on generalizing this basically at the end of the day like okay i made these discoveries and a person would like great i'm gonna download that one or i like that one right i don't like this one and that rolls up to some hypothesis that came earlier in the process.

21:34And so we think we can get to end-to-end on this as opposed to human preferences. So you mean the feedback loop is the click? It could be the click. It could be also like the you know, we do an experiment. Sometimes in Cosmos you could ask to end an experiment and you can go see what the experiment is success or failure or something like that. But I guess we brought it out of this kind of like hard-to-quantify is this a good hypothesis or a bad hypothesis and into this like you can see some downstream consequences of the hypothesis So yeah, humans have, I think, a very strongly well calibrated nose for science.

22:08Like, I mean, maybe you could argue there are sociological effects like across the community. But ultimately, oftentimes people really like good scientists know right off the bat, like, is this going to be likely to be useful or not? How long, how many attempts did it take before you started to see results that to yourself seemed useful? Like even working on this for, I guess, two years now. You know, I think when the co-scientist paper came out from Google, I think it was a really interesting idea to do this like tournament style or just pairwise ranking of hypotheses, right? So I think co-science is a very interesting counter example to what we built.

22:48What we built is something with either lab in the loop or data analysis in the loop or literature research in the loop where you're like iterating on an idea. I think co-scientists took this like very different approach of like, Like, let's list all the ideas and then try to come up with a filtration process to come up with the best hypotheses. So co-scientists will produce these very long reports of like, oh, we like really tested this idea with lots of dialogue and it's very interesting stuff. And I was really impressed with the paper that came out. And then we had this Robin paper. And one of the things that came out of the Robin paper is that the hypothesis that people thought was best was not the one that led to success in that paper.

23:25interesting it was in um age-related macular generation or ocular uh age-related macular basically it's like part of the eyes you're going blind because you have this accumulation of debris in the eye and can't clear it out that's one of the major cause of blindness in people over yeah uh 60 ollie who works on the hill yeah yeah he'll cringe when he hears me say that but something like that something like that sorry ollie um and that one like we went to optometrists or ophthalmologists, I'm actually confused on that as well. Sorry, Ali. But essentially, you know, what hypotheses do you think are good hypotheses?

23:59What do you think would lead to like a good mechanism for treating dry MD? Yeah. And yeah, there was, they agreed to be on the top 10, but beyond that, it was kind of noise. Yeah. And then, you know, what we found was ribosutal was a very good medicine and had a mechanism that I think is novel, although there was lots of debate on X because I think in 2012, there was a master's thesis which proposed this mechanism on like page 38. I actually think it was a typo. I think they meant wet AMD. But anyway, I won't belabor the point. I will concede that maybe there was one reported example of it in the past.

24:38That was a really eye-opening experience for me because that was the first really serious test where we really went to the lab and we spent like four weeks on a battery of experiments to see what hypothesis led to a good mechanism and a good repurposed drug. And it was not as correlated with human opinions as I expected. And so since then, I think that I have a lot more faith in these like verifier in the loop kind of scenarios where you have either data analysis, literature search, or you're running a unit test or whatever, you're going and running the experiment. Anything like that, I think, is going to give you a higher signal than the sort of vagaries of like, oh, this is a higher opinion or we like this one better.

25:19Yeah. Maxwellian called it nature's computer. Yeah. It's like you have this computer computational cycle you're running and the nature is part of that. Yeah. I'm curious. So you said that there is a paper which maybe like could propose, maybe propose where this molecule came from. But like, do you have some way of interpreting or like understanding where that hypothesis originated in the absence of that? is there a traceable thought train? Yeah, yeah, yeah. Actually, this is something we pay really close attention to. At Future House and at Edison was provenance of information. So our first sort of agent was PaperQA.

25:57Sorry about the name. PaperQA sounds like an email set, but it wasn't an agent. It really does. PaperQA has every sentence that it outputs has a citation to a page, right? So it's like a lot of provenance. And then we basically built along a philosophy for everything. So Robin, which is the name of this, I don't know, workflow or something, you can call it, that led to this result on rippocetal being a good therapeutic for dry MD. It has like data analysis that goes, shows you like which line Python code led to the result here. And then that is like, okay, then it goes to this other model, which says, well, based on this literature finding and this result from the data analysis, I believe this is the right thing.

26:36but you know where does the original idea come from like going after these rock inhibitors which is the mechanism for the target was basically enumeration and so this is like if you can't be smarter you can be i don't know you can try more times of course yeah yeah and i think that was like the theory of of the robin paper was that we can put out a whole bunch of hypotheses and then we can filter them just like i think about how co-scientists did is you go for a filtration process but the difference is that in co-scientists their filtration process was other lm sort of ranking it with rubrics or like personas.

27:07And our filtration process was like literature search and data analysis. Like here's some data. Is it consistent with the data? Go see if anyone's discovered in the paper and literature or if they've disproven it. And I think that's the easy way to succeed in AI over humans is you can try more ideas faster. Something I've heard people say, and maybe I've experienced this in my own life, um like sometimes hypotheses are kind of cheap especially in you know biology yeah it's many ways actually easy to come up with what you think could be happening yeah and um it seems like to me verifying is oftentimes a big bottleneck in maybe the biggest bottleneck like if you have lots of hypotheses you know and it costs you know one one hundredth of your runway to test each one of them or something you don't have any shots on goal yeah yeah so how do you make sure that like you you are actually enriching for good hypotheses literature and data analysis right you know like yeah yeah there was a time when we used um something called tiling trees and tiling tree is like a literal brute force method invented by ed boyden sam's phd advisor and basically the idea is okay i want to like accomplish x okay i could try these methods and then like once you pick i'm going to try this method then you split into like two different paths i'm going to use this method or not use this method.

28:25I'm using this method. I need to have like, I don't know, some kind of substrate. I'm going to try this substrate or this substrate or this substrate, right? And you can basically try to really like tile space of all the data. We tried some early experiments there and you're right, you run into this thing where some of the hypotheses that come out just don't make any sense. And like you are going to waste a ton of effort if you actually test them all. Nowadays, I actually would argue that if you go to an LLM and you ask it to evaluate, you know, hypotheses, including some garbage ones, it will probably do as good of a job as an expert in the field and filtering them out.

Read the full transcript

28:56That's not always the case. Yeah, I've actually seen that myself. Yeah, but there's a lot of gotchas and I think people can miss those, but I think they're actually pretty good. And so I'm not as worried about hypotheses that can fail fast by an expert looking at them. I think now the filtration process really happens in literature. And I think the filtration process happens in looking at like, you know, like a biobank data or like, you know, what do we know from GWAS or something? You know, other sources of existing data as much as you can draw upon. Yeah. So, yeah, with regards to existing data, another contrarian take is that oftentimes the hardest part is just understanding the context of data and where it comes from and how do you interpret it.

29:37I can also think from my own life, multiple cases where the data in some sense was there and you had two people who were both experts and very smart people who looked at it and drew very different interpretations. In fact, when we were interviewing Heather Kulik, she had some fun stories about using LLMs and she would find that there would be raw data in a paper which wouldn't agree with the conclusions of the actual paper. And it's straight from the paper. It's not even like cross paper talk or something. Man, I'm going to be a really boring interviewer and be like, yes, you're right. You know, like this is a hard question.

30:16I think, you know, to give you something concrete, we have a bioinformatics benchmark we call Bixbench. Bixbench, like we put it out. We've updated a few times. It's in some frontier alums when they release their system card, they'll mention Bixbench. It's like one of the things they test on. And, you know, we're getting to 60%, 70 % correctness on Bixbench. And we found that actually we're at the point where humans disagree at this level. Like humans only agree 70 % of the analysis. And so it's true that like when it comes to analyzing data, like humans do not agree 100 % of the time. There is a certain amount of like choice that goes into it.

30:57And, you know, we try to, so Edison is a for-profit company. We like trying to sell some of this stuff to the companies and we'll go to some companies like, oh, we never impute data. Imputing data is bad, like, or, you know, whatever. And like, well, we'll have to change our ages so we don't impute data for them. But then some of the companies like, oh, yeah, we impute data. It makes everything easier. Right. Or in and, you know, you want to know what the real modern dark arts are that like AI resistant area of the world is like medicinal chemistry. That is like a spot where like, you know, there's so much superstition.

31:26Oh, yeah. Yeah. Everyone. Yeah. Everyone is like pseudo religious. yeah exactly you have to be the survivor i feel those burnt out but the religions never agree two medicinal cannabis will have completely different viewpoints about like a functional group yes exactly and i remember this is a talking to somebody who works at cro and they're like oh whenever like company x orders anything we never put boron on any of the compounds because they hate boron because there was one program that was killed because there was a boron and you know somewhere in the core and it led to some toxic side effects so no boron for this company this company they like love things to be fluorinated or something because they love think it's great for the AdMet properties, right?

32:02And so there's like all this stuff where you reach the point where you're at, I don't know, human bias level or human disagreement level. And I think we're getting to that point in data analysis. And so of course you will see then that if I take the raw data from a paper and I analyze it myself, I will get a different conclusion. One of the cool tricks you can do is this back to this brute force thing is that I can go to our agent and I can run it a hundred times and I can take the consensus like analysis. Or I can say, even if you make these three different choices in your data analysis, you get the same conclusion, right?

32:31Or this conclusion is somehow sensitive to those choices. And then you can, there's even like words, it's like epistemic versus aleatory uncertainty, right? It's like, this is aleatoric, which means like, I think it's noise from the data or this is epistemic uncertainty, which means like, I think there's some choices that are being made. There's some model differences that lead to the disagreement. Anyway, there's like a Donald Rumsfeld formulation of this as well. Like, no, no, no. It's an aleatoric, epistemic. debate there. Interesting. This kind of digging into your cosmos. Yeah. So I glanced at the paper and one of the things that jumps out is that there were certain class of problems for which it was only 50 some percent accurate.

33:16Oh, yeah. Yeah. And can you talk a little bit about that and how that like, OK, so if I'm just raw getting 50 percent accurate answers and then I'm going into the wet lab and being like okay try this and then it's like ah like the the stupid thing did told me to do it well how do you i would say first of all that 50 it's actually pretty good because it's rare that experiments in the lab are actually coin tosses right they're usually a lot more outcomes than you know than than binary yeah sure okay yeah but but that particular number was uh human agreement in the interpretation of the results okay and so we asked people to evaluate different aspects of cosmos we had them evaluate like the data analysis decisions we have people ask it to evaluate the literature like is these do you agree with its finding the literature that number that was 50 that came from cosmos's interpretation of uh some of the analysis yeah so like it might go in literature and find this result and then would say wow this is super exciting this is amazing or my do data analysis like this is a novel discovery really excited about it yeah and then people would disagree that's actually not interesting or like i don't agree with the interpretation of it so it's like picking bad problems maybe yeah in the the negative of class and so i think it's like that that 52 or 55 whatever it is that's um interpretation and so i agree i think that's where like i was saying i think the frontier right now is scientific taste yeah and so that's what we're working on right now is how do you get that interpretation to match could you step back and just introduce cosmos on a high level yeah yeah um i would actually be in even curious to hear starting from like chem crow and uh you know you have uh paper qa avery ether zero yeah yeah i'd like to hear a little bit of the the lineage and how those different decisions were made what were the key learnings and how did you get to where you are now yeah so i could retcon and tell a really great story about how we arrived at cosmos but i will say that like to a large extent we just try a lot of stuff and sometimes it works and sometimes it doesn't okay you know i'll say that we're very i'm i'm a builder like i like to like build things piece by piece.

35:19I'm probably some fancy word for it, but I'm like a Lego guy or something. My vision was that we would make an agent that does this part of the scientific process, an agent that does this part of the scientific process, whatever. And so we had like ChemCrow, which is going to help us with setting up our medicinal chemistry work. We had ProteinCrow, which we haven't released. I don't know if we will ever release, but ProteinCrow is like designing proteins we might need for some part of our workflows. Or we had a data analysis agent. Is that LLM or that's a... It's an agent, so an LLM plus tools.

35:47Okay. Or we had Ether Zero. It was like, okay, we noticed that the frontier models can't work with molecules very well. So let's make a model with intuition for medicinal chemistry. And that was what led to Ether Zero. But then Sam actually really pushed on us to like, let's just even do the whole thing. Let's just try to build an AI scientist. Let's just try the whole thing. And that was what led to Robin. and um robin was like let's just take these agents we already have and we'll just put them in like a a work basically it's like you could express it in a concise python file of like you know try a whole bunch of ideas then go see if they all filter through literature or if they've been disproven and then go like uh come up with experiments that you could do in a wet lab yeah and this is our inventory list and then go analyze all the data then go back and repeat the process right so that's like what Robin was.

36:30And, um, we like, we came across cosmos. We're trying to like, understand what is the process that Robin is, is automating. And it came from this idea of like a world model, which is that when we first started Edison, we were thinking like, what do we, what do we want to change about this? Like what is new here? And so we spent some time thinking about, well, the scientific process, like what is actually going on in like my brain, which is that I have some understanding of, of the world or the phenomena I've studying. And that's my world model. And then a lot of the actions I take are about trying to update that world model.

36:59And it's something that changes over time. And so this is like this ability to change over time, but it's also something that is practical. Like I can use it to make predictions about, I know from this experiment, this will happen. That's why it's like a model and not just like, you know, memory or like a bunch of like papers, just like that. It's like, it's supposed to operate. In Cosmos, we tried this idea out and actually Ludo, who's the first author on paper, We tried a whole bunch of ideas around world models and we kind of thought they weren't really appropriate. Like, well, we tried a lot of different ways to do this.

37:33We tried, you know, method A, method B, method C, and they're okay. And so we all just had to take a break. Ludo, like his project didn't work on trying to do this world model stuff. He's like, I'm going to keep trying it. Ludo's a very stubborn person. So he tried it for like, I don't know, a week or two weeks. Then he was kind of like quietly. He's like, hey, can you guys come take a look at this? And we're like, wow, this is actually really cool. And then we like started building on it and jamming really. And I think what Ludo figured out is that you have to get this like experiment loop thing.

37:59You have to be able to let it in the data analysis agent is what got us in the loop. So if you put that in the loop of like, it can really update this world model because the, we were trying to build it around literature before. And when you build it around literature, there's just like not really experiments you can do and then see the results for. That was like our surrogate was, was literature. It just wasn't working. But data analysis actually really lets you explore ideas. And so that was what led to Cosmos. And so in Cosmos, we basically, we had all the pieces sitting around. We've been working on world models, we've been working on a data analysis agent, we've been working on a literature agent.

38:27And then we've been working on, you know, we built a platform for scientific agents. So we had things that can write a LaTeX report. We had things that can make nice plots. Then we put that all together and like a world model was like sort of the glue that allowed it to fit together. An analogy is like encoding agents. Like GitHub is sort of the glue. Like there's some shared repo and everyone works on the repo and like software engineers have spent whatever, lots of brain cycles thinking about what's the way to coordinate, you know, and organize working on code together for a long time. So the world model is actually like a memory system?

38:59Yeah, you can think of it as a memory system. We think about it as a model. So like it actually, you can put in input and it will output predictions. We think about calibration. But like really it is a set of like a big bundle of information that we accumulate over time that's distilled in some way. And that is like what allows us to do this. And I think you can think about like a GitHub repo is like, it's a distillation, right? Like really there's a long graph of commits that lead up to it. And like the current file system in that GitHub repo, or keep saying GitHub, I'm such a corporate shill here.

39:31Your Git repo is like a distillation of all of the work that people put into the PRs, into the commits. And so I think there's a nice analogy between a Git repo and what a world model is. I see. And I think that's just sort of what allows us to automate scientific discovery so well. Can you talk about like kind of how you implement a world model or is that sort of like secret that's our like secret sauce right now yeah that's fine yeah no it's fine people have asked so one thing that's notably missing yeah is the like simulation right yeah or dynamics or or or like uh bolts or yeah i want to help you guys pump up your views here so i i think molecular dynamics is overrated in fact coming from someone.

40:17Yes. That goes in the intro. In the thumbnail, you know. And DFT is overrated. In fact, DFT may be even more overrated than electronics. I think these methods... You mean for materials or for biology or for both? For materials. Okay. And I can explain more about that. Basically, MD and DFT have consumed an enormous number of PhDs and scientific careers at the altar of the beauty of the simulation. Also, random interjection. Once I did an estimate, I think pre-ChatGPT, something like 20 % of the world's computing power just went to simulating water. Oh, my fucking God, water. Yeah, yeah. I had to deal with so many water simulations.

40:57I did DFT simulations of water, and they are so annoying. I used these big computers from the Department of Defense, and I spent, like, I don't know, five months. And by the way, this is pre-LLM training days. Five months of compute is actually a really long time. I simulated water with quantum effects with a grotus mechanism for how a proton hops through water. And it's on YouTube. It's my number one YouTube video. And it represents like... Until now. And it represents like, I don't know, a million CPU hours of compute. It was one of the biggest computes that I... Probably the biggest one I've done in my life so far.

41:33Maybe Ether Zero is bigger, but it took a lot more work. Anyway, and what's the point? What did you learn? All I learned was like what set of hyperparameters reproduce some physical effects of water, but none of it was de novo. Right. And this is the this is the issue with with molecular dynamics and DFT is that they don't model the world correctly. And so we have to invent little stories we tell ourselves about we're like making good inductive biases and then it models the world more correctly. Like in DFT, you simulate water at 330 Kelvin when you want room temperature water. Is room temperature 330 Kelvin?

42:07No, it's not. That's a little too hot. Right. And so the issue is that people just make up these things or like, I don't know, GGA or B-LIP or B-3-LIP, all these different methods people. They're clearly empirical. And then they bolt it on to DFT and they say, look, it's a first principles method, right? But actually you made a whole bunch of choices and you overfit to the validation data to get this to work. And that's, I think MD and DFT are like that. Because if you go look at the catalysts, you know, what catalysts change the world, none of them are single crystal materials that are really well suited for DFT.

42:45They're always like they have grain boundaries, they have dopants, they're complicated, right? And you never capture DFT. So I think this is one of the fundamental, I don't know, dichotomies of the world is that simulations stimulate really boring things really well. They don't simulate interesting things very well. And so that's why I don't do DFT and MD anymore. What about somewhere like the machine learning stuff like AlphaFold and... AlphaFold was trained on x-ray crystallography data. And I think, you know, this is the story of MD is that MD was supposed to be the protein folding solution.

43:22There is a great counterexample. There's a, I don't know what there's a word, but the counterfactual is basically a group called DESRES, D.E. Shaw Research. they had you know similar funding to deep mind um probably more actually they tested the hypothesis to death that md could fold proteins they built their own silicon they built their own clusters they had them taped out all themselves they burned into the silicon the algorithms to run md they ran md at huge speeds huge scales yeah i remember david shaw came to a conference once on MD and he flew in by helicopter and was like this pretty famous guy kind of rich and um he gave an amazing presentation about the special computers and special room and outside of Times Square and like what they can do with it is beautiful amazing and I always thought that protein folding will be solved by them but it would require a special machine maybe the government would buy like five of these things and we could fold you know maybe one protein a day or two proteins a day.

44:20And when AlphaFold came out and it's like, you can do it in Google Colab, you know, or on a GPU or desktop, it was so mind blowing. I forget like that protein folding was solved. I always thought that was inevitable. But the fact that it was solved and on like your desktop, you can do it was just completely floored, changed everything. Like the bitter lesson on steroids. Yeah. I don't even know what it is, but it's like, imagine chat GPT came out, but instead it was like, oh, you can just run it on your phone or locally on your own desktop. Like that's the level of like shock that came out and it gets down to this thing that humans are really bad at estimating problems that aren't human-made problems protein folding we all thought was like would require a huge amount of compute very challenging problem was hardest problem in the world right and it turns out that you can actually do it on i don't know i think the numbers are now like 10 000 gpu hours you can train a good protein folding model it's actually turned out to be barely an inconvenience therefore why not oh oh therefore protein folding was highly efficient based on experimental data They took X-ray crystallography.

45:14That's what DeepMind did is they took X-ray crystallography data. Desiree has tried the first principles method. And it's like a nice head-to-head comparison. Two very well-resourced groups. They both tried different ideas. And the machine learning on experimental data beat out first principle simulation by a very large margin. And so why isn't like Bolts or whatever inside of Cosmos? Like why isn't there a tool that can run Bolts? Oh, we have Bolts inside of... We have Bolts Gen. yeah yeah we have that inside of cosmos okay it is i mean i think in the version that we have uh for people to just sign up and use it's not in there yeah but like uh you know you can imagine that you can just modal or lambda or tamarind or 310 there's all these companies that basically wrap a lot of these these like um uh deep learning protein design tools or chemistry design tools they wrap them in an api you just give that to give it to cloud code if you want you can give it to cosmos and you can be like hey you know if you want to design a protein for x use these tools Your mechanism, it sounds like, or one of the primary mechanisms that has been successful is like it like enumerate a whole bunch of possibilities and filter.

46:19Yeah. And so how do you think about serendipity and out of out of distribution thinking and getting there and how far have you gotten and what's left? Yeah, that's a great question. I think I guess the short answer is that there is very. So this is the domain of seaborne. So chemical, biological, radiological, nuclear weapons or I don't know, safety. Yeah. This domain has been explored a lot in history by a lot of organizations. And I would say that there was a big question mark for us a few years ago was like, how much of this stuff is intellectually bottlenecked? Like how often are people like, oh, wow, I want to cause harm, but I need to know like some facts.

46:59And could LLMs make that easier or go faster or anything like that? I think, you know, the first set of answers in 2023, I think was basically no, is that like, you know, you can go find the synthesis route for many dangerous compounds on Wikipedia. People know what are the targets in the human body that like are targeted by most biological weapons. It's not really that much of a mystery. So I don't think there was a lot of like, there's a lot of new ground when LLMs first came about. And then there's a lot of concern about like laboratory protocols is that could agents or LLMs reveal some tacit knowledge that like maybe people couldn't find on Wikipedia or like maybe for making something, there's some technique that is acquired when you scale it up in size or something.

47:45Or maybe there's like some way to get around like tracking lists by ordering different compounds. So in that, I think, was really well tested by a few different labs. not me, but there were some groups that spun up that started making tests for this and labs pay attention to. I think it's really been put into process where LLMs will shut down or be filtered in those scenarios. But I think that is actually an area where there is some risk. And so I think that's something that people pay attention to for open source models. And there's still, I think, some discussion there. But I think to a large extent, it's not really been greatly accelerating in practice, or at least I haven't seen much evidence of it.

48:22And again, I think it comes down to the fact that it's not really available, but if you look hard enough, you can find most of the information you would need to get up to no good in the public domain already. But then I think now is the next frontier is like, can it somehow help you with real-time protocols, troubleshooting more in the loop and more, especially in the computational side of things? there are some scenarios that are now coming into focus that could be more dangerous or more intellectually bottlenecked and so i think people are trying to pay attention to that to some extent there was like a first wave that we thought this could unlock a lot of stuff and i don't think it came to pass yeah i think there's now an emerging sort of second wave of like there are some actually new scenarios that were just too far-fetched to consider two years ago that i think are now realistic um some smart people are paying attention to it but i don't think it's solved yet.

49:15I don't know. It's very vague. So I guess like one kind of differentiator, there's a lot of talk about AI safety in like the broader LLM, you know, ASI space. And, you know, there it's jokes about paper or paper click maxing robots or something. But like the core threat here is more like a malicious actor using this as a tool to accelerate something dangerous. And like kind of the first order hypothesis is that you basically already have to be an expert to effectively create a bioweapon or a chemical weapon and a non-expert or an expert already know how to do this yeah i think you know so so each of the categories in the cbrn they're all a little different but i think to a large extent it's a lot of like pushing material around you know the classical example nuclear is like it's a lot of a lot of centrifugation yeah a lot of ultra a centrifugation, a lot of high pressure or high RPMs.

50:10And so it's just, you can maybe get smarter about how to set up, you know, the economy of scale to do that with an LLM. But to a large extent, I think you can call your friend and country X and they can tell you what are the steps. It's not, I don't think it's that much of a secret. It's just a lot of like moving material around. And I don't think it's accelerant, meaningfully accelerated. Now, with that said, there are all kinds of like, you know, dumb dual use things of like, maybe you want to call a company that makes centrifuges and you want to make sure that they sell you them and they go through some KYC steps and maybe an LLM can get you through the KYC faster.

50:46And that's like a dumb thing that like, okay, like, yes, like, you know, email makes it so that you can order centrifuges off the internet more easily. Is email like a dual use technology? Like, yeah, to some extent it is. And so I think there's a lot of like weird second order things that we don't pay attention to in AI safety. of like, does it make KYC easier? Does it make it easier for people to know where to order this from? Or like, what is the expected price? Or like, what should you order first, right? All those like sort of simple logistical things, I think are accelerated by AI just as like a consequence of AI being an accelerating technology.

51:19But certainly, I mean, shit guys, there's some scary stuff. And I try not to think about it too much. Yeah. I don't know, I guess, I don't want to get too political, but I do think that right now the United States government is maybe taking a slower, less intensive look at safety. But there's definitely people, I think, in other spaces than the U.S. government thinking about it hard. Do you think it's a thing people need to spend more time on? I do get waves of angst about AI. I'm sure many people living in San Francisco do get a little bit of waves of it. And sometimes I think that there isn't enough work being done on it.

52:00And then sometimes I think, wow, I need to mellow out And like, you know, we have lots of time to think about it. What is my opinion on it then? I don't know. I think my opinion is not formed fully. Yeah. You and Sam have done a lot of thinking about funding science and future of science. You have you've been vocal about the reproducibility crisis and other things. First question, why this focus research organization or for it? Yeah. FRO. Yeah. Yeah. What does that get you that you don't get from academia or, you know, big lab or whatever? A nice network of people. And I think Edison is like a real, of course, I think Edison's going to be great, but I think it's a mystery of what's going to happen.

52:46So I don't think we've had as much friction there as you might expect. But yeah, this is all stuff that Sam and I think about all the time. It's like, how do you balance stuff like this? How do you balance the economics? um you know there are some there are some venture-backed companies that are having cash salaries over a million dollars and it's like insane to me yeah that you would use all of your cash from your equity financing you know in these insane salaries but they can in terms of like total spend on gpus it can still be a total a small fraction of your burn so sometimes it kind of make sense yeah yeah that's that's one way to think about it so so like you this is a good uh lead-in to you are automating science yeah in some capacity yeah so where does that leave scientists so i think um this is uh jevin's paradox we can try here is uh um so uh let me start with the contrast here is that, you know, if we automate, you know, taxi cab drivers, there's not going to be an increase in people needing to go places.

53:52Maybe there'll be somewhat an increase, but like there is a finite amount of like time people will be spending in cars. And so there's an upper limit. So when you automate that, that's like a scarcity thing. It's basically you're displacing jobs when you automate driving. In science, I don't think there is a finite appetite or a finite capacity for science. I don't think science is like a scarcity thing. Like there's, you know, 100 more discoveries left to be made and then we'll be done. And so like we're displacing jobs. I think instead, actually, if we can make science go much, much faster, there will be no, there will be no decrease in demand.

54:25There'll be actually, I think, an increase in demand that will match whatever automation amount we have. And so my vision for what a scientist would be in the future is that they will be, I don't know, like agent wranglers or cosmos wranglers of like, okay, they're exploring 100 ideas simultaneously, or they're like working with systems like ours to make 10x discoveries 100x discoveries, because I think there's an unlimited amount of scientific discoveries to be made. So there's no like scarcity set where basically we will displace them all. Now that's kind of like, this is what I would tell when I talk to a first year PhD student yeah everything's gonna be just fine you know but then when it gets into the nuts and bolts i i do agree that this is going to be like a really hard thing where like if i am ceo of a company that makes science like a pharma company or material science company or something like that or a r &d arm at ibm i think well i could spend you know a million more dollars on on compute for the ai scientist or could hire 10 more people i might just choose to go with the ai scientist is because, you know, to a large extent, like hiring people is hard, right?

55:30And hiring an AI scientist is probably a little bit easier. And so I think that there could be some friction. But another thing is like, science is in some ways closer to art in the sense that like, there is a large number of people who just appreciate good science. Like if you get published in Nature, it's not because it's really going to be world changing. Of course, that's part of it. But it's also because like people are like, wow, this is really interesting science. So I think the enjoyers of science are also scientists. And so I think that it's kind of hard to imagine a scenario when there's not scientists as the consumers of science.

56:05And so I think if they're going to be consumers of science, they're also going to be some of the producers who are involved in the process by itself. I don't know if that makes any sense. Yeah, you've touched on this. The question in my mind is just what does a scientist do then? There's a great short story by Ted Chiang, I think in like 2003 or something. and it's about like well at first scientists were displaced and they became like the uh interpreters of like what the ai scientists are doing like the scientists read the ai scientists like papers and then you know translate them for whatever popular science or something and then after that like they couldn't read the papers anymore and so they were left behind and so they had nothing to do and they just sat around and but the problem is that science is like you know you you have to translate science to make any impact like science cannot exist by itself i do agree there's like engineering can exist by itself like if you give some kind of system a goal of like making me a material that i can make a space elevator out of you could be not participating in the beginning the process in the middle of the process and you just come by the end and be like okay all this recipe like science of like what's the origin of life or like is there water on other planets or, you know, why is some catalyst better than another catalyst?

57:14That has to be hitting human eyes and human brains at some point. So I think a human has to be involved in the process. Don't want to be contrarian, but yeah, be contrary. Why does a human have to be involved? Why does a human have to be involved? Well, a human has to be involved at least some point to be like, yes, this is good science or this is bad science. Okay. So it goes back to taste. Yeah. But I don't know. Maybe you're right. Maybe there is no point for humans. Maybe we'll be like no what is it sora uh you know like the ai slop app but i think in sora there's still humans at the end clicking the videos or something yeah so so the the sort of analogy kind of brings up an interesting point like is it possible that like due to the biases of ai science if we really go full in science that you know there still is a market for kind of boutique human science like you know there's still people who want to you know paint things the old-fashioned way but But more to the point, does it become even more important to have a human who is actively doing their own exploration?

58:10Because there will be like large blind spots and biases due to the models that just you'll never be able to overcome because this is sort of baked in now due to your training data. And without a human, you'll always get stuck and there will be a blind spot that will never. bio which is a company in oakland um or in emeryville and they do really cool stuff with automation i think they're going to be testing this theory of like okay maybe if that's the bottleneck we can see evidence of it because they're going to start doing really well um it could be true i still though want to say all of those i in my mind are still sort of scoped in terms of like r &d for pharma or bio but they're not like none of them are attempting to answer big fundamental questions and maybe there's like different levels when i think about that and you seem to be, um, it seems like the future, how the, the focus of future house in Edison is much more towards like, you know, sort of R and D and sort of end run science.

59:08But, um, you know, I, I have some background in, you know, fundamental physics. Um, you know, it's like, is there any thought about like, how do you like take on, you know, dark matter candidates? And like, I just, you know, think the data to really give us a complete story is just not there yet. You know what? I'm sure everybody at every company is the biggest critic of their own product. So we think Cosmos is, we think it's great, but there's a very large amount of area for improvement. So with Cosmos, there's like an open access to everybody version. Do you provide access to other labs that is less open?

59:51um we have a version of cosmos that has like um bigger resources like it can run for longer it uses gpus um so like basically when it does data analysis it'll have a gpu so we use that for things like um like machine learning experiments you know if you want to know like this question about whether it's better to pre-train first on noisy data or not yeah um we have like pre-release models that are coming out and we try those. But yeah, so I guess like, yes, we do. And we do have like research partnerships with companies where we like build something specific for them. And that is something we think about.

1:00:28But broadly, I would say Cosmos that's on the website is pretty close to what is the best we have internally. Yeah. I have a question. So you previously have stated that you think that language is the natural, um, language, what is it? Language of chemistry. The future of chemistry is language. Yeah. Yeah. Um, okay. So I wonder, do you still believe that? Good question. I think, uh, I w I would say yes. I still believe that, um, that so, so in that article, that opinion article, my, my point was that, uh, you know, at the time when I wrote that article, I think maybe three years ago now or something, maybe 2023.

1:01:10It was that we have models for predicting solubility of compounds. We have like data about our large populations and we have like papers and we have code. And the only way to bridge all that information is natural language. And the argument was that like humans, like, you know, whenever we can't bridge information, like if I can't talk about my code or I can't talk about some idea to you, I will invent words until I can get the point across. Right. And that humans are always innovating on language to make it represent all known observations and people innovate in language to represent whatever code pattern they have.

1:01:44Right. Like this is like the only shared activity we've been doing for this long is like coming up with words to represent everything we know. And so I think that for that reason, natural language is the only possible way to connect all the different pieces of data we need in biology, medicine or any domain for that matter. I think there's some caveats to this of like, you know, you can make an argument. Like if Jan LeCun were here and he would make an argument about like, you know, world models or like vision or embodiedness, right? Like there's arguments against natural language that like, you know, that maybe there's something more that it's not the complete story.

1:02:18Or maybe natural language imposes limitations that cannot exceed because you're stuck in this abstract space that was invented by humans and you can't escape it until you can like touch something. Yeah, I mean, it is an abstraction, right? And like scientists basically work exclusively in abstractions to some degree. I just, I find, I found that interesting because it seems like most scientists, you're right. Like when they explain things, they explain things through language. But many conversations, maybe most at some point result in people drawing diagrams or something. Like, you know, chemistry, like biochemistry largely or medicinal chemistry is oftentimes a, it's a language of graphs, right?

1:02:56Or, you know, I mean, bonds are abstractions, yes, but, like, they're pretty good abstractions for most, for many cases. Or, like, you know, geometry, you know, thinking about, you know, protein is like the geometry of a protein. It's like, I think that that's how people will, a lot of scientists like to think about things. And so I find it interesting that, like, yeah, that you are focusing primarily in this language. Like, have you thought about essentially a multimodal version of this? Like where, you know, when it comes along a smile string, it doesn't just say, oh, this is a smile string, but like this is a graph.

1:03:31This is a representation of some higher like abstract object. You're absolutely right. And the problem with these is like, I don't know, Jacob's ladder or something, whatever you want to call it. It's like, yes, you can say that you can call a molecule by its name. You can show the graph. And if you go to a molecule like ferrocene, well, it doesn't really have bonds, but like part of it. And so then you're like, well, we need to draw it visually. and then you go to molecule like i don't know glycine betaine on this dihedral angle right and so like it's not actually this thing i drew it's actually an ensemble between this thing and this thing right then you go to benzene you're like well not only is it like a ensemble of these different conformers it actually has electron density and you can't really ignore the electron density in benzene you like need to treat it correctly and it's like well you can't actually represent the electron density that way you actually have to look at the correlation of the electrons individually right because you can't really model benzene with like dft right or functional, you have to actually look at the electron correlation.

1:04:25These are the electron correlation. Like, well, you know, you can model electron correlation, but, you know, actually these things, when they're in a solution, they have like, you know, relativistic effects because it's like there's a whole bunch of stuff around it. So you really got to have the relativity in there. And you're like, well, you got the relativity and you have the electron correlation. You can have the bonds and you have the conformers, but you really need to think about the cosmic radiation background because like, you know, it does actually impact everything and there is some energy there, right?

1:04:47And before you know it, you've ran out of, you know, You've ran out of compute or whatever resource you're using to model this. And so I think you have to draw the line somewhere. Natural language, like I said, is that humans have worked for a long time to make it be the, you know, what's the word? Like the least abstract or the, you know, it's somewhere on the border of like it's still abstract enough that you don't need to know all these details. But it's still granular enough or concretized enough that you actually can make use of it. there may be some other representation like multimodal might turn out the video or maybe i don't know there's some other like fusion that you can make i like natural language because we all work really hard to make it right at that boundary and i do agree sometimes sometimes ideas slip and they can't be in language you have to get out the whiteboard or ideas slip and you have to wave your hands around you know or maybe then then you need that that uh degree of freedom to communicate just digging in on this a little bit more like uh famously quantum mechanics is like undescribable, right?

1:05:50Like there's, there's an argument that you cannot understand quantum mechanics with words or in, in with our preconceived understanding of the physical world, because it doesn't behave like the macroscopic world. And so the only way to understand it is through mathematics. Right. Um, and I largely see language as the joint key of science as well, but I wonder if that's not true for many domains in quantum mechanics is just the one that hits you in the face i mean i don't know i actually i think the there's like seven principles of quantum mechanics or five or something like this that you can actually express pretty concisely in language i agree that like you need to actually look at the consequences of them you need some mathematics um i don't know i actually i don't know this is like a challenge i think you could actually describe a lot of quantum mechanics and language Sure, sure.

1:06:40But but I see your point. And yeah, I guess I'm a realist. Like I when I talk to my kids, you know, maybe I will be like, OK, let me draw for you. I don't I don't make sure in our house everything is described with natural language. So I agree with you there. I think maybe we can be a little a little flexible with with natural language and include equations and smiles strings in it. And I think we can get a little bit farther. So maybe that's OK. But some people, I think, like optionality. You know, like, oh, it could be this or it could be that. I'm somebody I like to take strong opinions and see how much farther they can get me.

1:07:17And I think in my career, it's actually been better for me to take strong opinions, which in my deepest of hearts, I know that are maybe not correct or not fully correct. But once you take these strong opinions, it just, you can sort of move many steps down the road once you take these strong opinions. And like, for example, at Future House, we took the opinion that scientific agents are the future. And that allows us to skip a lot of steps because a lot of other people were like, we need to build a foundation model for X. Yeah. And we just skipped all that. Right. And I think if you also were unopinionated and you had optionality, like I can think of a famous example of a different company that liked the optionality and they wasted a lot of time on foundation models or something.

1:07:53Then I think you get stuck. So that's one of my strong opinions is that natural language is a way to join all these different domains. It may not be a correct opinion. It may be more subtle or more complicated, but it's allowed me to get very far. I'll drop it someday and maybe find a new one. Yeah. Not yet, though. That's my meta opinion on the matter. The Ether Zero story on your blog, I find hilarious and kind of awesome. Yeah. You know, when I was a kid, I loved the like genie slash monkey paw concept of be careful what you wish for, but you just might get it. Yes. Maybe just like quick story.

1:08:32Can you just talk about that? That was just a really fun... EtherZero was a hell of a project because conceptually it was a very short project of like, hey, people have made a lot of progress and verifiable rewards in math and code. Let's say we can do it in chemistry. So chemistry is like not a verifiable field, right? Like, of course, you can go test something in the lab, but then we had to think about all these ways that we can make chemistry verifiable. And one of the ones we settled on was like, make a molecule that has like three nitrogens, two oxygens, ten hydrogens or something. And we thought that was like a pretty verifiable question.

1:09:09But every time we would train a model, it would find some new, insanely weird trick to generate these molecules. And I'll just tell you, one of the examples was that it would make these molecules and we would do some checks to make sure it had the right bonds, the right number of electrons, the right number of atoms and stuff like that. But it would just solve the problem in any way possible. So it would just put all the nitrogens over here, put all the auctions over here just like things that don't look good yeah so we started coming up with these rules of like oh let's check to make sure it followed these good practices or these good practices and we found ourselves into this like you know it's like the opposite of the bitter lesson like i don't know the boutique lesson where you like try to make everything custom but one of the things it kept doing is it kept putting these nitrogens in a row and it put like one nitrogen two nitrogen three nitrogen all in a chain and this is like you know if you have three nitrogens it's like explosive you know two nitrogens it's like bad and like four nitrogens you can't make and i kept telling everyone like it would make these like six nitrogen compounds and they're just they're just literally impossible and they're not possible and many of the people on the team were like computer scientists like on this team and one of them like one day sent me that like this is on the cover of nature today on nature's website somebody made a six nitrogen compound and this is like somebody's like career to deliver this compound because this is the most unstable like insane compound you can make it's some ridiculous setup and like the spectroscopy to get that proven was like very difficult.

1:10:26And this, I don't know how they did it. It was amazing accomplishment. Like, look, Andrew, like it's not actually impossible. And it was so funny to me that like our model was sitting here, spitting out these six nitrogen compounds in like, you know, 2024 or 2025. And like the paper just happened to come out that year. Like mankind had finally made a six nitrogen compound. Do you think that those were actually synthesizable even under these extreme circumstances? No, our model was just, it was just reward hacking. Okay. it was just the the model was so creative and ways to reward hack like one of the one another one we did was um you know we wanted it to make sure that the when it would propose a reaction like make this compound tell me how to make this compound we would try to make it sure that all the reagents were purchasable like you could purchase them they were not like made up yeah um and the reason we came with that is that originally we just like take the end compound and then like remove one atom and be like here's you buy this and then put the atom on it's like okay i wish it was like that um so they have to be purchasable and then we'll be like we thought it might be hard if they're all purchasable because sometimes you actually order things custom or something so we'll just make sure one purchasable so the first thing it starts doing is putting nitrogen in there because nitrogen is purchasable and it like has no participation in the reaction right like oh my god okay so they're like okay it has to be purchasable it has to participate in the reaction they started putting like acid base chemistry we'll just put an acid here acids are purchasable and it'll move one atom and they're like okay fine can't be that everything has to be purchasable then we find ourselves and i'm like sitting there one day building this like ridiculous catalog of purchasable compounds and bloom filter so i can go fast enough on our training loop and i'm like why am i doing this how did i get here and i don't know it was really funny because um pre-training or training transformers you know on just data like just supervised training where you just have the inputs and the outputs directly very nice relaxing you know like things are always robust you know things go pretty smoothly when we do these verifiable rewards where you have to like write a bulletproof verifier it is really difficult and we had so many models trained only to find out they were hacking some other like random thing in our setup it's really hard and i and i i don't envy the frontier labs that have to do this at a very massive scale because we had a lot of adventures in ether zero and you guys I should read the blog post.

1:12:41Definitely read the blog post. It's very fun. It's a great read. GRPO? We did make some modifications to GRPO. Yeah. I actually, I used to know all the names of these modifications, but I think it's like DAPO is one modification and like the clipping we did was special. And we explored a lot of that stuff. Yeah. And it was also one of these things where like you think the hypers are wrong, the algorithm is wrong, and then you find out it's just because like you had some. somehow sorted the reagents when you made your training data but when you made your test data you didn't sort them alphabetically and the model was just like barfing because its whole strategy was to exploit something in the way you sort of things so yeah we explored a lot of different methods and it was um i learned a lot about chemistry a lot about nomenclature um and actually there's a i learned a lot about medicinal chemistry as well more than i ever wanted to awesome if you want to do some like engineering just check out edison scientific and they have I think they're hiring with lots of interesting things, everything from scientists to infrastructure engineer.

1:13:44Thanks, Andrew, again. Yeah, thank you very much for joining us.

From the publisher

Editorโ€™s note: Welcome to our new AI for Science pod, with your new hosts RJ and Brandon! See the writeup on Latent.Space for more details on why weโ€™re launching 2 new pods this year. RJ Honicky is a co-founder and CTO at MiraOmics (https://miraomics.bio/), building AI models and services for single cell, spatial transcriptomics and pathology slide analysis. Brandon Anderson builds AI systems for RNA drug discovery at Atomic AI (https://atomic.ai). Anything said on this podcast is his personal take โ€” not Atomicโ€™s.

โ€”-

From building molecular dynamics simulations at the University of Washington to red-teaming GPT-4 for chemistry applications and co-founding Future House (a focused research organization) and Edison Scientific (a venture-backed startup automating science at scale)โ€”Andrew White has spent the last five years living through the full arc of AI's transformation of scientific discovery, from ChemCrow (the first Chemistry LLM agent) triggering White House briefings and three-letter agency meetings, to shipping Kosmos, an end-to-end autonomous research system that generates hypotheses, runs experiments, analyzes data, and updates its world model to accelerate the scientific method itself.

The ChemCrow story: GPT-4 + React + cloud lab automation, released March 2023, set off a storm of anxiety about AI-accelerated bioweapons/chemical weapons, led to a White House briefing (Jake Sullivan presented the paper to the president in a 30-minute block), and meetings with three-letter agencies asking "how does this change breakout time for nuclear weapons research?"

Why scientific taste is the frontier: RLHF on hypotheses didn't work (humans pay attention to tone, actionability, and specific facts, not "if this hypothesis is true/false, how does it change the world?"), so they shifted to end-to-end feedback loops where humans click/download discoveries and that signal rolls up to hypothesis quality

Kosmos: the full scientific agent with a world model (distilled memory system, like a Git repo for scientific knowledge) that iterates on hypotheses via literature search, data analysis, and experiment designโ€”built by Ludo after weeks of failed attempts, the breakthrough was putting data analysis in the loop (literature alone didn't work)

Why molecular dynamics and DFT are overrated: "MD and DFT have consumed an enormous number of PhDs at the altar of beautiful simulation, but they don't model the world correctlyโ€”you simulate water at 330 Kelvin to get room temperature, you overfit to validation data with GGA/B3LYP functionals, and real catalysts (grain boundaries, dopants) are too complicated for DFT"

The AlphaFold vs. DE Shaw Research counterfactual: DE Shaw built custom silicon, taped out chips with MD algorithms burned in, ran MD at massive scale in a special room in Times Square, and David Shaw flew in by helicopter to presentโ€”Andrew thought protein folding would require special machines to fold one protein per day, then AlphaFold solved it in Google Colab on a desktop GPU

The E3 Zero reward hacking saga: trained a model to generate molecules with specific atom counts (verifiable reward), but it kept exploiting loopholes, then a Nature paper came out that year proving six-nitrogen compounds are possible under extreme conditions, then it started adding nitrogen gas (purchasable, doesn't participate in reactions), then acid-base chemistry to move one atom, and Andrew ended up "building a ridiculous catalog of purchasable compounds in a Bloom filter" to close the loop

Andrew White

Future House: https://futurediscovery.org

Edison Scientific: https://edison.science

X: https://x.com/andrewwhite01

Kosmos: https://edisonscientific.com/articles/announcing-kosmos

Chapters

00:00:00 Introduction: Andrew White on Automating Science with Future House and Edison Scientific
00:02:22 The Academic to Startup Journey: Red Teaming GPT-4 and the ChemCrow Paper
00:11:35 Future House Origins: The FRO Model and Mission to Automate Science
00:12:32 Resigning Tenure: Why Leave Academia for AI Science
00:15:54 What Does 'Automating Science' Actually Mean?
00:17:30 The Lab-in-the-Loop Bottleneck: Why Intelligence Isn't Enough
00:18:39 Scientific Taste and Human Preferences: The 52% Agreement Problem
00:20:05 Paper QA, Robin, and the Road to Cosmos
00:21:57 World Models as Scientific Memory: The GitHub Analogy
00:40:20 The Bitter Lesson for Biology: Why Molecular Dynamics and DFT Are Overrated
00:43:22 AlphaFold's Shock: When First Principles Lost to Machine Learning
00:46:25 Enumeration and Filtration: How AI Scientists Generate Hypotheses
00:48:15 CBRN Safety and Dual-Use AI: Lessons from Red Teaming
01:00:40 The Future of Chemistry is Language: Multimodal Debate
01:08:15 Ether Zero: The Hilarious Reward Hacking Adventures
01:10:12 Will Scientists Be Displaced? Jevons Paradox and Infinite Discovery
01:13:46 Cosmos in Practice: Open Access and Enterprise Partnerships

More from Latent Space: The AI Engineer Podcast

All 247 episodes
๐Ÿ”ฌ Automating Science: World Models, Scientific Taste, Agent Loops โ€” Andrew WhiteLatent Space: The AI Engineer Podcast ยท 1 h 14 min
Listen in VO