Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem

4 Aug 2026 · 47 min · 25 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Chai Discovery co-founders Josh and Matt argue drug design is a scaling problem and should be made more “engineering-like” via de novo molecule generation, rigorous evals, and simplification of model complexity.

Guest backgrounds

Josh (OpenAI early team; GPT-1/GPT-2/scaling laws; grew up programming; later focused on AI for DNA/protein). Matt (switched from pure math/theoretical CS into deep learning for protein structure prediction; earlier work on protein folding; antibody-focused results).

Key claims

Biology progress can be hill-climbed like ML if models are verifiably evaluated in the lab. Chai builds models from scratch (not fine-tuning GLMs). “Simplicity” and scaling compute/data/models guide iteration; biology’s wet-lab error bars require larger step improvements.

Notable examples

Protein folding competitions (2018 step change; AlphaFold; later 2020 with “AlphaFold2/LFOLD2”); diffusion models enabling simultaneous structure+sequence generation; antibody success rate rising from ~0.1% (1 in 1000) to ~15% with CHI-2; glycosylation motifs as hidden target features.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Revolutionizing Drug Discovery with Engineering Principles

0:49 to 1:41

Discover how Chai Discovery aims to change drug discovery into a more engineering-like process.

“A lot of the medicines that we have today were kind of discovered quite randomly.”

Understanding the Drug Development Process

1:41 to 3:23

Explore the evolving boundary between engineered processes and real-world testing in drug development.

“The same way we have that in code and in modern engineering allows us to iterate really quickly and try to bring that into biology and into drug discovery.”

The Evolution of Protein Folding and Drug Design

3:23 to 4:28

Learn about the advances in protein folding and their impact on drug design over the years.

“So at first what you do, these are kind of like bring me back to my original days in my advisor's lab.”

Antibody Design and the Challenges Ahead

4:28 to 7:09

Understand the challenges of antibody design and the breakthroughs needed to overcome them.

“And what's pretty remarkable is that with relatively minimal information, like you could build a machine learning system that could actually produce something that really looks like a protein.”

The Role of Diffusion Models in Protein Generation

7:09 to 10:47

Explore how diffusion models are enhancing protein generation and their implications for drug design.

“Like my background was was never biology.”

Building an Interdisciplinary Team for Drug Innovation

10:47 to 14:01

Learn how Chai Discovery assembles a diverse team to tackle complex drug design challenges.

“But at the time, like I think like it wasn't really as developed enough to work for our problems.”

Navigating Drug Design Challenges

14:01 to 15:10

Learn about the complexities of designing next-gen antibodies and the importance of engineering in drug discovery.

“We were even, I think, worried when we started that trend because for some of the next generation formats, the complicated antibodies like multispecifics, like they didn't even work with Chi2.”

Aha Moments in Protein Design

15:11 to 20:38

Discover the breakthroughs in protein folding and generation, and how these advancements are affecting therapeutic properties.

“So we've kept that bar really high while also trying to level that with great research talent and people that can actually push the frontiers of what's possible.”

Improving Molecule Generation

20:39 to 22:46

Understand how success rates in molecule generation have improved from 0.1% to 15% and the implications for drug discovery.

“So it actually just means the bar is really high in terms of the step changes that you want to see with the models.”

Scaling Challenges in Biology

22:47 to 24:48

Explore the scaling laws in drug design and how different teams approach challenges in the drug development process.

“a 24th module, like we can actually like unlock that new target.”
Show all 25 chapters

Future of Drug Discovery

24:49 to 26:50

Envision the future of the pharmaceutical industry with computer-aided design and its potential to revolutionize drug development.

“And there's usually not one bottleneck at Chai.”

Business Model Decisions in Drug Development

26:51 to 28:00

Discuss why the company chose to enable the existing drug development industry instead of developing drugs themselves.

“Let's imagine a lot of innovation has flowed downstream into some of the wet lab parts of the process.”

Future of Drug Design

28:00 to 28:40

Understand the evolving strategies in drug design and the importance of specificity.

“And I think that means that the future is really bright.”

Business Model Decisions in Pharma

28:40 to 30:00

Explore the rationale behind adopting a supportive infrastructure model over direct drug development.

“Because I think a lot of times folks think about chai and isomorphic and the same neighborhood.”

Learning from Pharma Partnerships

30:00 to 31:30

Discover key insights gained from working with pharmaceutical partners and their expectations.

“We've we've always wanted to just partner broadly with the ecosystem.”

The Competitive Landscape of Pharma

31:30 to 32:50

Investigate the competitive pressures in pharma and how AI impacts drug development efficiency.

“Have you any surprises or any learnings from working with these partners?”

Data Sources for Drug Modeling

32:50 to 34:40

Learn about the primary data sources utilized for training drug design models.

“As our models get better, we get better results to our customers and so on.”

Integrating Data and Model Improvement

34:40 to 36:00

Understand the cyclical relationship between data generation and model accuracy in drug design.

“Josh, interestingly, was taking like the exact opposite approach.”

Navigating the Competitive Drug Discovery Space

36:00 to 37:40

Discuss strategies for staying ahead in a competitive drug discovery market.

“Sometimes people ask us at CHI about, these days at CHI, like which paradigm are you actually going after?”

Working at Chai: Insights and Challenges

37:40 to 39:40

Learn about the culture and challenges faced by employees at Chai in the drug development space.

“partnerships that are making your models better.”

Reflections on Success and Impact

39:40 to 42:00

Explore the emotional highs of achieving impactful results in drug development.

“But now we just need to continue to hone in on making these things even better.”

The Excitement of Progress in AI for Drug Design

42:00 to 42:56

Learn how recent breakthroughs in AI are impacting drug design and patient outcomes.

“You'd much rather have the pain of growth.”

Challenges and Responsibilities in Drug Development

42:56 to 44:15

Discover the balance between innovation speed and quality assurance in drug development.

“to change the world in a pretty profound way.”

The Name Behind Chai Discovery

44:15 to 46:00

Understand the rationale behind the name Chai Discovery and its significance.

“We also wanted a name that like biotech companies have such complicated names.”

Future Aspirations and Innovations at Chai Discovery

46:00 to 46:42

Explore upcoming projects and innovations that Chai Discovery aims to achieve.

“It's actually really easy and it's very motivating when you're making progress.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Josh Meier:One of our big guiding principles is just simplicity. So like when you look at a model like let's say CHI 1, I think there are 23 distinct sub-modules in CHI 1. And like when you're trying to iterate on something like that, it gets really hard. Because you're like, I kind of need to understand each of these sub-modules independently. I need to understand all their behaviors, their dynamics. And like that doesn't actually scale that well. So like you can start to think, okay, how do I simplify this? How do I identify like what's really important? And once you have that, like kind of the whole research process and like identifying these types of scaling directions becomes a lot simpler.

0:48Josh Meier:We're here with Josh and Matt, two of the co-founders of Chai Discovery. Chai is engineering molecules with AI. so it's something of a foundation model lab for biology sounds like a big idea but let me just

1:01Matt McPartlon:start there what is the big idea thanks for having us on on the show um one of the exciting things we're trying to do at chai is to make the drug discovery process look a little bit more like engineering and we've seen all the work happening with alums for code generation for instance right and that works really well because code's a very simple abstraction right i'd like to kind of get across what you want to do uh biology doesn't look like that today it's a lot of like trial and error I think if we go back to the early days of just like modern biotech, right? A lot of the medicines that we have today were kind of discovered quite randomly.

1:34Matt McPartlon:It was a bit serendipitous. And what we're trying to do is allow us to industrialize that process a little bit more and try to come up with the tools that we'd need in order to have abstraction layers. The same way we have that in code and in modern engineering allows us to iterate really quickly and try to bring that into biology and into drug discovery.

1:52Josh Meier:So let me ask you a question on that then. So there's, and tell me if this is a reasonable way to frame it. There's almost this boundary between that which can be engineered and that which needs to be tested in the real world. And it feels like that boundary has been moving over time. Like the proportion of the drug development process that can be engineered as opposed to trial and errored seems to be increasing. Is that a reasonable way to think about it? If so, can you talk about like what specific developments have pushed that boundary over time?

2:22Matt McPartlon:I think that's one way to think about it. The way that we often look at this is like the lab is an important part to verify that what you're doing is correct. And actually verification is a very big theme in AI as well. If you can evaluate that your model works, you can verify it, then you can start to hill climb that and you can make progress on it. So I think the lab is a very important part of this. And the question is, how do we take drug discovery and make it look a lot more like drug design? So one of the reasons why we call this field drug discovery is we're often looking for a needle in a haystack.

2:50Matt McPartlon:We'll screen millions, billions of molecules, try to find one that works. And if we can instead put in what is like the dream state of the molecule you want, and then the model can materialize that, that's going to be really powerful. So it's not even a question of like reducing lab testing. I mean, maybe that happens as a result. I actually might even take the opposite side of the coin and we can talk about that where maybe we'd actually do even more lab testing because the ROI will increase. The same way there's more software engineers or there's more demand for software engineers now that they become more productive.

3:16Matt McPartlon:There might be more demand for the lab. um but i think the key part is how do we just change the paradigm here and how do we we make

3:22Josh Meier:it more design oriented can we talk about what is state of the art today and then maybe let's take a little trip down memory lane five years ago three years ago one year ago like what have been some of the major breakthroughs and like how has state of the art changed over the recent years yeah i think the the field has like it's really evolved especially like over the last decade um so it It really wasn't until, like, you can actually look back, there's like this biannual protein folding competition. So every two years, usually a bunch of academic groups who compete in this protein folding competition, what would happen is like, you kind of hold out these protein targets that no one's ever seen before, they never get deposited publicly, and then all these groups compete so you can predict the proteins the best.

4:06Josh Meier:And like, really, it wasn't until 2018 where you started to see this like big step change in performance, and then finally again in like 2020 with LFOLD2, but really like it was kind of like the advent of deep learning that really like kicked off this whole field. So at first what you do, these are kind of like bring me back to my original days in my advisor's lab. So we worked on like one of the first systems to do this with deep learning. What we would try to do is predict like the distances between amino acids and a protein and then we'd have some render which then took in like these kind of noisy and complete looking distances and then emitted some protein from that.

4:41Josh Meier:And what's pretty remarkable is that with relatively minimal information, like you could build a machine learning system that could actually produce something that really looks like a protein. And for the most part was, was correct. Like there was still pretty big gaps. It didn't quite have the resolution of like what we have today, almost like with the early image generation models, where it was like a little bit grainy, a little bit blurry. And now it's like, oh my gosh, like this is, this is crazy high definition images like in seconds. So we kind of like saw that same evolution happen over time so it started with protein folding and then i guess once that became more realizable uh there were kind of these other sub problems that people wanted to solve so now given a protein structure can i design say like a sequence which might fold to that and that's getting closer and closer to like the drug design problem so you're you're kind of like now thinking okay how do i start engineering proteins that like match a certain shape or might perform a specific function.

5:36Josh Meier:And so people kind of studied that problem independently. And that was kind of like 2021, 2022 time. And then kind of these ideas began to merge together. It was really like the advent of diffusion models where we started to be able to like, okay, I can now generate a protein structure and a sequence kind of simultaneously. And I can start making these like what we'd call prompts more and more realistic. So I can now prompt these models on a target that might have a certain shape. And I can say, hey, I want to bind over here, kind of like add more real world constraints on the problem what did you guys see in 2024 that made you think that was the right

6:07Matt McPartlon:time to start the company yeah there were a lot of discussions that went into it i remember one of these uh early discussions actually back in matt's house when we were uh you know matt was was showing me some of these results on like antibody antigen like structure prediction matt was just talking about protein folding um but for for a long time people thought that protein folding for antibodies was just like too hard of a problem people were like there was not enough data for antibodies, like in the protein databank, for instance, to solve this problem. And people as a result thought that antibody design was going to be out of reach.

6:37Matt McPartlon:A lot of the work that people even did on protein design in the early days with deep learning, it wasn't antibodies. It was these different class of proteins called mini proteins, which are actually really interesting in their own right, but they're not what most of the drug industry is looking at. So the holy grail was whether we could design antibody proteins on the computer, especially ones that had all that therapeutic function. And our thinking was that if you couldn't predict what an antibody looks like, how are you ever going to design one?

7:03Josh Meier:Traditionally, people have thought like, like biology is scary. Like there's so much to know. I'm terrified of biology. I am as well, honestly. Like my background was was never biology. I studied like pure math and started my Ph.D. in theoretical computer science. And it was only like after my third year that I ended up switching into like deep learning, protein, structure, prediction, all of this stuff. So it was like totally new field to me. seemed insane but like at the end of the day it's much simpler and like the the problems are much more like interconnected uh than one might think so like i think people like when they start going into the field they're like oh man what's an antibody what's a mini protein where these are all just like sequences of amino acids at the end of the day like these are just like different types of prompts for the model um but like in the same way we might like have a math problem that goes into chat gbt chat gbt can both answer your math math problem and like help you with your english homework.

7:55Josh Meier:So like really we have the same type of thing going on with our models. Like we just have some way of representing these sequences of amino acids. Then we have a way of like designing, predicting those as well. And I think like in that lens, like things, things become a lot more clear. And Matt, you mentioned your background a little bit. Josh, talk a bit about your background and then more broadly, in order to pull this off, you have a bunch of different disciplines that kind of come together. So can you just talk a bit about like your background and that of some of the other core members of the team and how these things all fit together?

8:25Matt McPartlon:Yeah, I've been excited about AI and biology since I was a kid. So I guess I'm like, Matt, I didn't start with theoretical CS and get into that way. But I went to a high school with a stem cell lab. So I was just always excited about biology as a kid. And I grew up as a programmer. I really started my career at OpenAI. So it was on the early team there. It was a nonprofit back then. So it was a pretty good time to be there. We did GPT-1, GPT-2, scaling laws. And the question was, like, if the models can learn to speak English, German, French, why can't they learn to speak DNA and protein? And that was kind of my research agenda since then.

9:02Matt McPartlon:I think that sort of intersects with around the time, like, Matt got into the field as well. And I think it was a pretty important time as well, right? Because if you look at the kind of methods that we were bringing in, like, there's been a lot of these, you know, changes on the edges, if you will, right? And, you know, Matt talked a bit about the history of what's happened in the fields here. But, you know, it's all about like, how do we find like the right deep learning architectures with the right compute configuration and the right model architectures to make this happen? What are the right tasks to apply it to?

9:29Matt McPartlon:We're talking about how we even knew that like, you know, 2024 is the right time to start the company. As Matt was saying, you know, these are all different like kinds of amino acid sequences. And people thought that, you know, antibody class of problem was going to be too hard. And we started to see the first signs of life that actually this was starting to work. I think a lot of it fueled by some of the new architectures we're bringing to the problems, where I think it was diffusion models back then.

9:51Josh Meier:The first time anyone was able to generate reasonable looking proteins was the advent of diffusion models. So it was pretty crazy. There were a bunch of generative modeling approaches that would kind of work if you had a bunch of data. So people got these working for images. There were a bunch of tricks to make this better and better along the way. We've definitely borrowed a lot of those ideas in our domain as well. but it was really like once diffusion models came around. Is there an intuition for why diffusion models work? Yeah, so diffusion is not a one-step process. So I think kind of up until this point, the main generative design paradigm was called variational autoencoders.

10:31Josh Meier:And in that case, you're saying, I want to just compress my input distribution. So you have some proteins, you want to make these look like fuzzy Gaussian vectors. that task is just like really hard. And maybe today if we tried like super hard, I think we'd crack it. But at the time, like I think like it wasn't really as developed enough to work for our problems. What turned out working really well was just kind of giving the model more time to think and showing it like more examples. Like here's like a slightly broken looking protein. How do you make it better? And you can kind of break that protein more and more and more and you can make it look more and more noisy, more and more broken and teach the model just quick little shortcuts.

11:09Josh Meier:it's all right here's how i make it slightly better you can just keep asking over and over again make it slightly better make it slightly better and kind of like breaking the problem down to that scale worked really really well for for biology yeah the make it slightly better reminds me of a game that i like to play with my daughters where we have chat gpt give us a unicorn and then we make it stronger and we just keep telling it to make it stronger and by the time we're done we have the strongest unicorn in the world so about the same right yeah that's uh it's that easy. Okay, maybe not the same.

11:40Matt McPartlon:Yeah. So I feel like in this domain, you need to assemble kind of like a quadrilingual group of people, like an Avengers squad of chemistry people, biologists, AI folks. And so that's a challenge. How have you guys gone about finding people, convincing people to join the team? And who are your superstars? So we've been really pragmatic about this at Shai. If you look at the founding team, it was mostly AI researchers. So people who had worked on either scaling models or getting them to work in this domain. But really with each generation of model, the kind of people we've needed for the next milestone has changed, or I'd say probably has expanded, right?

12:17Matt McPartlon:So if you look at Chai 2, right, that's the point when we started to bring in some of the most incredible, like, you know, antibody engineers and scientists in the world. One of the scientists, Andy Young, actually when we hired him, people asked us if we had pivoted into building a full stack drug pipeline because they're like, you'd be crazy not to do that if Andy joined your team. But I think Andy has enjoyed running more antibody campaigns in the past couple of months than he's probably run in his whole career, which is very cool to see. You have folks like Nathan Rollins on the team. Nathan was actually homeschooled and then started college very early on.

12:51Matt McPartlon:So you joined David Baker's lab, who won the Nobel Prize for protein design when he was 14, started his PhD when he was 18, and has so many creative ideas. As the model started to get better, we needed to build up a product team, because while the researchers might get the models to point that they're very powerful, you need to build the right product interfaces so that the models are actually useful. And that's where we started to bring on people who have built some of the most exciting products we know about today. Like our co-founder Jack worked at Stripe. Munaz, who was one of the top 10 code contributors at Stripe.

13:22Matt McPartlon:Neil, who ran his own cybersecurity company before security started to become very important as we deploy this to our big partners. As we started to scale up, We brought in people who've really done a lot of the GPU hacking, if you will, in order to scale up our systems that they don't break when we're running them at scale. We had an email or a Slack message from one of our hyperscalers the other day where we had a cluster. I think that's an issue, too. And they're like, oh, the GPUs got too hot. I think you guys are running too many. And we're like, isn't that the point? Right? That's probably, I was like, good, we're doing our job at least.

13:53Matt McPartlon:Yeah. It's like, how do you convince those people to join? I think, again, a lot of it comes back to the results and like a clear need. We didn't hire antibody engineers before in an antibody design model. Like, what are those folks going to do? We were even, I think, worried when we started that trend because for some of the next generation formats, the complicated antibodies like multispecifics, like they didn't even work with Chi2. So it actually took a couple of weeks when some of those people showed up before the models could work at a point that they could work on some of these interesting case studies.

14:19Matt McPartlon:But fortunately, the progress was fast enough to kind of bring that online. So I think we're always like evolving that team and going for, you know, that next milestone. We've got the team very small as a result, too. so this way you know everyone is a little bit like slightly over capacity i think which means we have to prioritize it forces us to work on the things that really matter most i think on the

14:37Josh Meier:research side as well one of the founding engineers kevin woo he had the first uh i think it was the first protein diffusion model like ever and that speaks to kevin's speed of execution like he is a heck of an engineer and like i think engineering has just always been like important since day one So really, even our researchers, they're all excellent engineers. And we really care about building a high-quality code base. At the end of the day, we are technically a software company. We're AI researchers. We're protein designers. We're a lot of things. But our deliverable is some piece of software.

15:12Josh Meier:So we've kept that bar really high while also trying to level that with great research talent and people that can actually push the frontiers of what's possible.

15:21Matt McPartlon:And you've had a number of, it seems like, aha moments in the field. Like alpha fold was an aha. We can figure out how a protein folds. And then the diffusion models, aha, like we can generate proteins. And it seems like your latest models are a kind of another aha moment. We're not only generating molecules that look like proteins, but they also have therapeutic properties. Like they can bind really tightly. They have really high hit rates. Can you talk sort of about the quality of the molecules that your models are producing now and everything that sort of went in to those models to make them able to do that.

15:54Matt McPartlon:Yeah. If we look at the quality of the molecules that's coming out, if it goes back to one of the theses when we started the company that we really wanted to focus in on this de novo generation of the molecules. If you look at what a lot of the drug discovery AI work was at the time, it was about how do I take a molecule and just make it a little bit better, which we just talked about in a sense, but it was doing it with a lab-in-the-loop style, where I take some data, try to make it better that way. And the question we started with is, well, can we actually just do all that on the computer? Is there a way that maybe there would be enough data or we could collect enough data so that we could just zero shot a molecule that has a lot of these properties?

16:32Matt McPartlon:So the first thing we needed to do to get there was to design molecules with really high success rates. When we started the company, the state of the art for antibody design was about like a 0.1 % binding rate. So one in a thousand of the molecules you design would actually bind in the lab. So first of all, that means you have to screen a lot of molecules to find some good ones. It also means that like the gradient you get on your process is actually quite weak as well. So for many targets, you won't get any hits. For the ones that, you know, you do get some hits, you won't have enough to actually see whether you're having like the drug-like properties.

17:01Matt McPartlon:So we really focused in on how do we just make this process more accurate? We got to, with our CHI-2 model, about like a 15 % success rate. So now if you screen a thousand molecules, you're getting 150 back. Now you can start to get some like interesting statistics on the properties of the molecules, right? And a lot allowed us to iterate on that, build the right evaluations around that in the lab, and really try to hill climb that as well. And we're getting to a point now where we can actually bake in a lot of these different properties from the start into this engine. And then maybe if you think about how does this happen or how are we approaching this as well and why do we think it's going to continue improving?

17:34Josh Meier:Yeah. So honestly, I think Josh and I were both surprised at how quickly this worked. When we were originally budging this, we're like, ah, maybe 20 % hit rate in three or four years. we were like really shooting for like a one percent we thought one in a hundred would be amazing we're like this is going to be a groundbreaking thing we had a philosophy on like an approach that we wanted to take and it just like ended up working really well and that approach is like very similar to what's worked in the rest of machine learning so people kind of treat biology as this like bespoke problem or like bespoke field but really it's like the same principles as like self-driving lms well same principles but one of the things we've talked about before is how you manage to find scaling laws.

18:13Josh Meier:And it's one thing to tokenize a string of text. It's another thing to tokenize biology. Can you say a couple words about, without giving away any of the magic, you know, can you just say a couple words about that challenge and how you guys solve that? Yeah. So one thing that I really liked about Chai is like, we're a very bitter, less impaled company. So like, we really believe in like scaling data, scaling models, scaling compute. In order to do that, obviously you need to identify scaling laws. otherwise you're just kind of like wasting time and resources. And I think like without giving away too much, like one of our big guiding principles is just simplicity.

18:46Josh Meier:So like when you look at a model, like let's say Chai 1, so like they're like, I think there are 23 distinct submodules in Chai 1. And like when you're trying to iterate on something like that, it gets really hard because you're like, I kind of need to understand each of these submodules independently. I need to understand all of their behaviors, their dynamics. And like that doesn't actually scale that well. So you can start to think, okay, how do I simplify this? How do I identify what's really important? And once you have that, the whole research process and identifying these types of scaling directions becomes a lot simpler.

19:18Matt McPartlon:One of the other things too, I think it's interesting about this, is we take a lot of these lessons from what's worked in the rest of the deep learning space. But as you're pointing out, the data itself is different. The models at Chai are completely built from scratch. We're not fine-tuning GLM or something like that on some protein data. We build everything from the ground up. I think a lot of the company building process, though, is taking a philosophy and actually sticking with it and iterating on that and just having some guiding principles. When you're building like a blue sky research company, if you will, you know, Chai is almost like one of these neolabs, right?

19:47Matt McPartlon:Like we have this big AI problem we're going after where as we make progress on it, you know, that opens up opportunity for our customers. But if you're going to work on something so open ended that way, you need some principles to guide you. And I think we've done a very good job on like tracking those principles in the company, working on it. So a lot of the things that Matt's saying - Can you share them?

20:04Josh Meier:What are the guiding principles?

Read the full transcript

20:05Matt McPartlon:So I think simplicity was one of the ones that Matt mentioned. It's this like bitter lesson pilness of like, you know, scaling compute and data and models. It's being really rigorous. That's something that's so important in this space. You can fool yourself so easily in biology. Like the error bars in the wet lab are actually quite large as well. So it's actually a little bit different than if you look at like cogeneration, for instance, if you look at Sweebench, people were like, oh, this is maybe like a year ago. People were like, oh, I got 16%, then 17%, then 18%. I mean, in biology, if you're like plus or minus 5 % in your lab, that might all be the same.

20:39Matt McPartlon:So it actually just means the bar is really high in terms of the step changes that you want to see with the models. But you also need to be really honest with yourself about whether you're making progress or not. So you could come up with some fancy model that looks like it works well in one or two new tasks, but it's very important to show that that works more generally if you're actually trying to build a product that can bring the field forward. Yeah. And that's, I think, pretty interesting because biology is one of those inherently not so verifiable domains. And you guys have been really good at sort of showing your progress to customers and to people like us who know very little about biology.

21:13Matt McPartlon:And so can we talk a little bit about the evals and the verifiable part of the model progress? How do you guys know that your models are getting better? Well, I would say, first of all, that I actually think this is one of the more verifiable domains. It's actually a very objective readout if you look at something even like CodeGen, right? Like maybe it's a verifiable task, like did my code compile? Did it solve these unit tests? But how do you think about the taste, right? Like did I write some really sloppy code that can't be maintained? Like what does that look like? When we think about designing a molecule in the lab, we can actually be quite specific about many of these properties, right?

21:46Matt McPartlon:So maybe we get a molecule that binds the target, but can I manufacture it, right? That might be your version of like some tech debt, but you can measure that. And I think those evals, again, actually make this domain more verifiable. Maybe it takes a little bit longer to validate it, right? It's not like five seconds to get a readout and run a unit test. You might have to spend a couple days, a couple weeks in the lab to get that readout. But at least you can be honest with yourself. Yeah. When we talk about progress and how good the models are, there's a domain of targets in biology that you can, as you mentioned, sort of address with traditional screening methods.

22:16Matt McPartlon:They take a long time. They're very slow and rudimentary. And then there are targets that just aren't addressable with existing methods. They're not druggable for whatever reason. And so where are we in terms of model progress, in terms of, you know, working on existing targets and generating molecules faster? That's one end of the spectrum. And then the other end of the spectrum is unlocking novel biology, new targets, and things that we couldn't drug before.

22:40Josh Meier:This kind of goes back to, like, kind of, like, our core modeling philosophy. So, like, there have actually been, like, plenty of times where, like, man, if we had a 24th module, like we can actually like unlock that new target. And we're like, is that really something that we want to maintain long term? Is this incremental or is this like actually a compounding improvement? Will this actually like help us generalize to the broader class of like these, this whole class of targets that we really can't hit? And so like our philosophy has been like, all right, let's just like continue to focus, like identify your scaling laws.

23:09Josh Meier:Like there comes a point where like if the model is able to push loss down even further, it has to understand something very intrinsic about the target that it's operating on. So one example we were talking about the other day, Paul and I, maybe the way that we're looking at certain glycosylation sites on proteins, we're like, oh, we might want to represent them differently or something like this. And we're like, well, even if we didn't represent them, there are certain sequence motifs that will tell the model there should be a glycosylation site here. And in order to drive Lost Down further, the model should just have to learn that.

23:40Josh Meier:So there are all these hidden features of targets where, if you really believe in scaling laws, you believe the models will get there, these types of targets should just unlock with better models. And of course, you still have to take this very seriously and you still need all the proper validation. You need to really challenge yourself and make sure that this is truly working. But I think our approach has always been with better models, we should be able to unlock a lot of these targets.

24:06Matt McPartlon:One of the other interesting things is if we look at, we talked about the interdisciplinary nature of this. If we look at the different teams at Chai, what people will call a hard target is actually different in literally every team. So on the science team, it might be, you know, like a undruggable GPCR target. No one's gotten something that has like, you know, modulated that in a functional way. On the ML research team, it'll be something like, oh, there's like the model just can't fold this thing up. It doesn't know what it looks like. And then on the product team, it might look like, oh, I've got all these like, you know, modifications, like my glycosylations and it's a membrane protein.

24:36Matt McPartlon:How do I represent that to the user? And actually the fact that it's different for each of these groups, I think is a feature rather than a bug. And it means that if we want to make broad progress over here, everyone is kind of pushing in parallel on these different ways. And that means that there's very smooth progress that we can make all the time. And there's usually not one bottleneck at Chai. It's not like, oh, if we only had that one extra module, things would work. Or if we only tried to push this into the product in some way, we could unlock it. We're trying to build this unified solution.

25:03Matt McPartlon:Because at the end of the day, the goal of the company is to build a computer-aided design suite for molecules. It's not to make one or two molecules. It's not to get a pipeline of like five interesting therapies that we bring to market, it's the change of the way that medicines are discovered. And if we're going to do that, we need to work on all the hard targets, regardless of how you define hard.

25:20Josh Meier:Say more about this idea of computer-aided design suite for molecules. What does that mean?

25:24Matt McPartlon:So at its core, it goes back to this point about making biology look more like an engineering discipline. So we're not going and fishing something out of a large library or doing a ton of trial and error. You want to be able to specify upfront the principles that go into designing your molecule. And then I have an engine that can actually realize that into some molecule that we're going to go and create in the lab. And look, you still might do some iteration on the lab and on your model because maybe your hypothesis was wrong. But what we want to speed up is actually, again, have that computer-aided design suite so that you can go from idea to testable hypothesis very quickly.

25:57Matt McPartlon:And if that loop right now takes something like nine months, I don't know, to go and discover your molecule versus if it takes nine weeks or it takes nine days, each order of magnitude just scales in a very big way the number of ideas you can really sort through. I think that's ultimately how the field is going to converge on better medicines. It really comes back to people sometimes talk about, do we care? And you kind of noted on it, it's not like, do we care about speed or do we care about the difficulty of the targets? At some point, they converge in this way as well. Because a hard target, if we can iterate through hypotheses a lot faster, then maybe it'll be easier to crack it.

26:31Josh Meier:So if we can maybe detach ourselves from the present reality and go far enough into the future that we're not thinking present forward, we're actually thinking future back, 2035, 2040, 2100, whatever you want. What does the industry look like? Like, let's imagine that computer-aided design suite for molecules has become a standard. Let's imagine a lot of innovation has flowed downstream into some of the wet lab parts of the process. Can you paint the picture of what the industry might look like 10 years from now?

27:03Matt McPartlon:Yeah, I think it's going to be a really exciting time. And we can look at this on a couple of angles. So first of all, the quality of the medicines that we develop will hopefully go up. There's a lot of molecules we put into the clinic today that it's really hard to discover a molecule. You find something that's like 80 % of the way there. Maybe I'm just going to advance it anyways to hit my timelines. It's probably going to benefit some patients. But then you get beat a year later by someone else. and it's really just not the most efficient spend of resources in the industry. You've got the kind of diseases that are just too hard to go after today.

27:33Matt McPartlon:People have been trying to drug Alzheimer's forever and unfortunately haven't made as much progress as we'd like. You have things that just aren't economical to go after today. Think about personalized medicines, rare diseases, things where maybe the patient population is going to be smaller. But again, if we can iterate through those ideas faster, if we can launch something faster, if we can do it in a less expensive way, then those probably come into reach as well. So I think there are just so many different axes that that we're able to push on. And I think that means that the future is really bright.

28:04Josh Meier:It used to be like you either want to be first in class or best in class. Yeah. Now it's like you want to be last in class because like you actually just want to like be the final answer. So I think like a lot of what you'll see is just like way more intentionality in the types of drugs that you're designing. Like this drug will be super specific to the disease of interest. It won't have like certain interactions, like certain negative interactions that a lot of drugs today do. A lot of this stuff is actually like you're able to model most of this compensation, like maybe not today, but there's definitely a path towards getting there.

28:34Josh Meier:And I think that's like one of the most exciting things for me. Very cool. Can we talk about a business model decision that you guys made? Because I think a lot of times folks think about chai and isomorphic and the same neighborhood. Isomorphic is developing drugs. You guys are enabling the existing industry to develop drugs more efficiently, better, faster, cheaper than they had before. Why did you decide to go down that path versus the isomorphic path?

29:00Matt McPartlon:First of all, I think both of these paths are great and they can create tremendous value. We've always been really excited about building infrastructure for the industry. Our bet when we started to go back to those results in Matt's house a few years ago was that this is how most future drugs are going to be discovered. And if that's the case, somebody needs to go and build that infrastructure to make it happen. I think part of this, too, is a lot of our founding team, including Matt and I, had worked in companies before where we had built these full stack drug pipelines. And we were building AI models.

29:31Matt McPartlon:We were putting those drugs into the clinic. Again, we're pretty bitter lesson-pilled as well. And our thinking was, as the models get better, we want to be spending more of the money making better models as opposed to diverting those resources into clinical trials and things like that. And one of the interesting things about our business models, as the models get better, it actually wins us the right to continue investing more in them. Right. And you have partnerships with with these pharma companies that are paying off today and allow us to invest in it. So it's a much more scalable business in that way.

29:58Matt McPartlon:And I just think about the ultimate impact that we can create for the world is a lot larger. We've we've always wanted to just partner broadly with the ecosystem. It goes back to this point about fooling yourself in biology. If you work on a small number of drug targets, you might come up with the most exquisite molecules, creating a ton of value for the world by doing that. But you might miss the forest for the trees because maybe your model doesn't generalize to the other 500 molecules that people are going to make that year. And if you go and partner with people, you just can't fool yourself.

30:26Matt McPartlon:Like you look at our partners, Eli Lilly, Novartis, Argenix, Pfizer. These are not companies that are, you know, they take this stuff for granted. Like you have to really deliver on these partnerships for them to take you seriously. And that means it can't just work like some of the time. Like when we ship models at Chai, they really have to work. They have to deliver value to our partners. and it's not like, oh, we made some molecule. It doesn't fully work. We're going to have our chemists like clean it up a little bit. So it's almost a harder business to pull off. I think that's one of the reasons why you haven't seen it pursued many times.

30:55Matt McPartlon:If you work on a drug pipeline, again, like there might be ways to fix things up later. There will be the proof of like what happens in the clinic, of course. But when you have this partnering based model, you have to be really rigorous about your models. Your models have to work really well because otherwise those partners are not going to come easily. So it's made our life harder, I think, in many ways. But I think it's also the more rewarding path if we can get it to work.

31:14Josh Meier:Well, and what have you learned? I mean, you know, you're out of the lab, so to speak, and in the real world, you know, delivering real value for actual pharma companies. What have you learned as you started to work with these partners in terms of any surprises in terms of how their needs might have differed from what you expected or how their level of sophistication around this might be different than what you expected? Have you any surprises or any learnings from working with these partners?

31:35Matt McPartlon:Yeah, so when we went into these partnerships, I think a lot of people told us that like pharma doesn't know how to use AI. These are not like a tech forward industry and things like that. And to be honest, that hasn't really been our experience. I think that these are, again, they're very rigorous partners, very rigorous customers, right? And they're going to test every one of our claims right before they start to deploy these things. But when they see the data, they go all in, right? Because pharma is an innovation industry. And it's interesting just to think about even the whole economics of the pharma industry.

32:04Matt McPartlon:If you build a product in pharma, right, like a drug, right, you only have exclusivity on that drug for a certain period of time, right? And you have to continue to reinvest. Eli Lilly is a trillion dollar pharma company right now. If they don't get more blockbuster drugs, they will not be a trillion dollar pharma company forever. And I think that forces these companies to really be on their game of adopting new technologies and deploying them and trying to stay ahead. Pharma is a very competitive arena. There's a lot of people trying to bring these drugs to patients. By the way, it's great for patients, but it means that you have to be on top of your game here.

32:36Matt McPartlon:And I think that means that once you're through that door and your models are working, we've seen an upscale this adoption very quickly and people thinking about how to use the models in incredibly creative ways.

32:47Josh Meier:I really like the incentive structure that it creates as well, like taking more of the partnership model. As our models get better, we get better results to our customers and so on. So I think that's just a nice side effect for us as well. Chai has been like incredibly focused. And like really as the models get better, you can kind of like iterate on a better model with better data and so on. So like a lot of what we see and like a lot of what our partners are using the model for, like generating binders, antibodies, so on. Like we're also like doing a lot of dog fooding in house. Like we have a whole science scene that's using the models and just trying to understand what they can do.

33:20Josh Meier:And a lot of that comes back to like, okay, now that like the model, now that we've unlocked use case X, like what type of new data can we generate? How can we make the model better that way? And really like that's been the long-term vision of Chai is like, we never thought of like, we're going to build this one model that's going to solve all of this. Like we know that they're just like in other fields, like they're going to be multiple iterations of this model. And as you get better and better models, like you get this flywheel effect, I guess. And you mentioned what data can we generate? Can we touch on data for a minute?

33:48Josh Meier:You guys can't exactly just go scrape the internet and have all the data you need to build your models. Where does the data come from? How does it kind of compound over time? Can you just say a word about data as an input to your models? Yeah. So I think the primary source of data or the gold standard source of data is this protein database. So this is like, it's actually just like legit lab scientists who since 1970 have just been depositing crystal structures of proteins and other molecules over time. And really without that, structure prediction and design wouldn't have been a thing. So there are these, so that's like one source of structural data.

34:30Josh Meier:When I started in the field, I really took like a structure-pilled approach. I was like super interested in predicting structure. Like, how do you think about a machine learning model that can output 3D coordinates? Like that's, that's not, doesn't look like an LLM. It doesn't look like an image model. This is really in its own class. Josh, interestingly, was taking like the exact opposite approach. So he, he was like one of the original authors on ESM and that was like a really like a seminal work and understanding how to apply language models to protein sequences. What's really cool about that work is if you can train a language model to understand protein sequences, what ends up happening is it ends up kind of like representing the 3D structure internally.

35:04Josh Meier:And there's like a really interesting reason why this happens. So like in order to predict like missing amino acids, like same way that you'd predict like next word in a sentence, I want to predict next amino acid in a protein. In order to do that effectively, like you really need to understand like, okay, what does amino acids like immediate microenvironment look like? because that kind of tells you, okay, what are like the compatible amino acids with everything surrounding it? And in order to do that well, you need to understand the protein's 3D shape. So I was taking like this really structure-based approach and Josh was even more bitter lesson to me.

35:34Josh Meier:He's like, we're just going to like look at the sequences and this is just going to emerge. So yeah, the two main sources of this data, again, like protein database for structures and then like these massive, massive, maybe even like order of like trillions of tokens sequence databases. And what you can do once you have really good models is just run them on the sequence database to get new structures out. So again, you have this compounding effect. As your models get better, they get more and more accurate at predicting these structures. And then you have more and more training data for the next series of models.

36:03Matt McPartlon:Sometimes people ask us at CHI about, these days at CHI, like which paradigm are you actually going after? I think one of the things I love about our team is that it's actually just like neither. We're very pragmatic. We want to solve the problem. We don't really care is it a sequence approach, is it a structure approach. In practice, it's going to end up being both, of course. I think you'd be surprised, but there might even be more like biological sequence tokens on the internet than like English language tokens. Now, a lot of that data is not that useful. It might not be redone. It might be very noisy, but there's a lot of data out there.

36:31Matt McPartlon:There's a lot of art in bringing it together. I think also the exciting thing is as the models have gone into a point now where we can design things in the lab, write to new targets, for instance, we can actually use the models to generate data as well. So there's a lot of exhaust from all the experiments that we're doing at CHI, which also helped to make the models even better. So it's a similar kind of takeoff that we saw with LLMs. Like when I was at OpenAI, we worked on reinforcement learning of like a GPT-1 architecture. Like did not work because the models weren't like good enough. But once the base model got good enough, then you could start to do those kinds of experiments.

37:01Matt McPartlon:And I think there's a similar analogy that's starting to happen in our world now as well, where the models have reached a point where there's actually a renewed interest in data. And like, how do we actually bring the models into the loop on like making that happen? I think that creates another really like interesting cycle on compounding improvement of the models. Yeah. We're talking a bit about compounding improvement in the models. And you said something about how pharma as an industry is an incredibly competitive landscape. And it's interesting because I think there's been a renowned interest in using AI and using ML to generate proteins and molecules.

37:30Matt McPartlon:And so your arena has actually become quite competitive. And you guys have obviously done an amazing job at staying at the frontier. It's been the year of deployment for you. You've locked up a number of pharma partnerships that are making your models better. But how do you guys think about the competitive landscape and staying at the front of the frontier? Well, first of all, I think it goes back to looking at the results and not fooling yourself and being rigorous. So one of the reasons why we do a lot of evaluation of the models, we mainly do it just to hill climb in the models itself. I think if you look at many of the capabilities that we've brought in have kind of been like first in the field, you look at our CHI 2 model, right?

38:04Matt McPartlon:Like getting to success rate of the models that you didn't have to do large library screening anymore to see results. A couple of months later, showing how we could bring in a lot of these developability properties we've talked about before, like the manufacturability of the molecules. So in many cases, we are pushing the model forward and trying to see these capabilities emerge. And then we try to quickly lock those in. How do we make those capabilities even more pronounced so that they become production ready and we can ship them to our customers? I think one of the things I like to tell the team is that it's not like we are head to head with other model providers or something like that, where all of us, I think, are working against nature.

38:39Matt McPartlon:Nature has actually been a pretty good baseline. People have done drug discovery a certain way for a long time. And Matt talked about how maybe you don't want to add module 24 to the CHI model, but people have added module 240 to the existing wet lab protocols, and they have been tuned quite considerably. One of the scientists on our team, Andy Young, he was one of the first people working on yeast display at MIT actually two decades ago. And he's got 20 years experience in Pfizer and Genentech, really honing in these methods, has a drug approval to his name and antibodies. And I think you look at someone like that and like, you know, Andy knows how to make a good antibody with existing tools.

39:15Matt McPartlon:And that is actually the bar that we need to clear. Now, of course, I think the ceiling on AI is going to be a lot higher than what we've managed to do before. Otherwise, what would be the point of doing this? We didn't start the company just to make, you know, a 10 times faster mouse, right? We started this to make breakthrough medicines that weren't possible before. But ultimately, that is the bar that we needed to clear in order to get adoption. I think we hit that inflection point a couple months ago. That's why you've seen a lot of these big pharma announcements. But now we just need to continue to hone in on making these things even better.

39:44Josh Meier:For somebody who's listening who thinks, wow, this sounds pretty cool. I wonder what it would be like to work at CHI. What is the best thing about working at CHI? And what is the worst thing about working at CHI? Yeah, I can speak to some of this. I'll think of this on the fly, the worst thing. So I think like the best thing is just like how actually like mission driven everybody is. Like everyone is so dedicated to what they're doing. I've worked at other companies, like the closest I've ever seen to this is like maybe some of the guys in my PhD lab. But like everyone is just like incredibly, incredibly motivated.

40:20Josh Meier:We all work really hard. There's like an obvious shared goal. And I think that's really rare to see. And I think this goes back to just like the focus that we've had since the beginning and like we've always had like kind of a clear philosophy a clear plan on how things are going to get there how things are going to get better and really like every night chai is very bought into this it's pretty amazing just to see like the amount of dedication that that everyone's putting in uh least favorite thing about chai uh not directly on top of dandelion chocolate maybe no maybe the next office yeah

40:51Matt McPartlon:Yeah.

40:53Josh Meier:Actually, so probably the least favorite part now is just like, I guess it's getting things to work at scale. And really, I didn't even know what that meant. Actually, when we started HCI, we had 128 GPUs. And I was like, this is like the most scale possible for a lab. Like, this is crazy. I came from like, you know, my PhD group where there were four of us sharing eight GPUs. And I was like, I just felt GPU rich. It was crazy. So now, at CHI, we have a lot more infrastructure to maintain. We have a lot more compute resources. Luckily, GPUs are parallelizable, but that also brings up its own set of problems.

41:29Josh Meier:So just how do you keep a cluster healthy over time? How do you get that large training run? So how do you keep that running for months on end? And even when it does die, how do you automatically resume these things? How do you keep all of your communication down? How do you optimize the models and make the best use of the resources that you have? So I think these are a lot of problems that they continuously pop up. They're good problems to have, but I think they're also really difficult to solve. Yeah, and I'm just excited to work on this. Yeah, I remember one of the CEOs that we've worked with a couple times, Frank Slootman, had this line about, you either have the pain of failure or the pain of growth.

42:03Josh Meier:You'd much rather have the pain of growth. Yeah, that's exactly right.

42:06Matt McPartlon:I think my favorite part is probably the results. And that sounds a bit cliche, but there's nothing. That's working. It's working, right? And just knowing that you're – I think many of us in the company, we've been working in AI for a long time. There's all this experience you've built up. And to know that you're applying it to something that really matters, like just even – I think just take Matt and – like we've been working on this problem for like 10 years, right? And a couple of years ago, I'm looking at something like, yeah, we're writing some cool papers. Like everyone is like celebrating this.

42:33Matt McPartlon:Are we actually making the world better? Like is this actually going to impact some patients? And I think now the answer is like actually yes. Like we have reached the point where this is going to make a big difference in the world. And every time you get one of these breakthrough results, anytime there's a new feature on the product that makes our lives of our customers easier, whenever there's like new lab results coming back from the science team, it's just always so honestly exhilarating to realize like this is actually going to change the world in a pretty profound way. I think that's also then maybe comes to the least favorite side, right?

43:03Matt McPartlon:Like, you know, we're running the company and it's like we have real partners that are relying on this and like things have to work, right? And you ship a new model generation. How do you make sure that there's no bugs in that? How do you make sure you don't have regressions, right? This is no longer just like a, again, the blue sky research problem of like, oh, we got some cool results and we move on. We've had to have really high priorities on like, you know, having production level code bases as the team grows. How do we make sure that the code base is in a state that more people can contribute to this?

43:28Matt McPartlon:So something that one of our other co-founders, Jack, likes to say is that if you want to move fast in the long term, you sometimes have to just move a little bit slower in the short term, right? and make sure that you are building something, again, that goes back to that compounding idea. It goes back to, we don't add module 25 to make the next thing work. So sometimes, you know, you're like so excited to get to the next result and you just want to jump into it. But we have real partners, some of the biggest companies in the world that are now relying on us. And it's important that we realize that, we take that responsibility to heart and we make sure we're building systems that, you know, continue to work.

44:03Matt McPartlon:I have a burning question. Why are you guys called Chai Discovery? Chemistry and AI. But we love Chai-Chi as well. It's a lot of Chai-themed stuff in the offices.

44:12Josh Meier:That's a good question.

44:13Matt McPartlon:I didn't know that either.

44:14Josh Meier:Josh is the visionary. That was all him. It is a very user-friendly name.

44:18Matt McPartlon:Yeah. We also wanted a name that like biotech companies have such complicated names. Yeah. We wanted something that's going to be much simpler. Yes. We're trying to make this whole thing simpler, right? Yeah. So we need a simple name to go along with it.

44:28Josh Meier:Awesome. What are you guys most excited about in the next six to 12 months?

44:33Matt McPartlon:I think for me, it's just the deployments that are happening. So we've announced a couple of these partnerships, and I'm really excited just to hear about the results that our partners are bringing online. It goes back to this point of making a real difference and also why I'm so happy to see how these partnerships are going, even posting the agreement. And as we're working with these folks, the models are not just sitting on a shelf somewhere. They're actually being used on real programs, people trying to approach devastating diseases where, you know, if CHI could give them a molecule that works, it could really change the lives of patients.

45:06Matt McPartlon:So I'm really excited to see how that goes. Just the pace of progress here is incredible, but also just like the pace of the models. Like a year ago, you could not zero shot a molecule and like, you know, have a good sense that your program was going to work. Like now that's changed. Someone might zero shot a molecule and be like, I think we're going to bring this program to the clinic now. And then a year later, they might even have some of those first molecules going into patients. And just the speed of that is just incredible. And sometimes you get some shivers thinking about this, that like, okay, my model is going to...

45:36Matt McPartlon:Matt has some patents from our last company about just generating the molecule on the computer, and these things are now in patients. And just to think about the scale, I don't know, a few years from now, do we have dozens? Do we have hundreds of chai molecules going into people? It's a bit mind-blowing to think about what that might look like for patients.

45:53Josh Meier:I think one of the things that, again, what makes Shai unique, the dedication, people are like, man, you work a lot. And aren't you burnout or whatever? It's actually really easy and it's very motivating when you're making progress. Seeing the progress they were making and just thinking, man, the next model is going to be even better than the last. We identified this new thing, so on. That's so incredibly motivating. For me, it's more of just like what what can we unlock next and like how do we make these things more controllable and like when someone comes to us with a certain target instead of just like hoping we get good affinity or something like that you know can we actually control this can we say like we want exactly a 10 nanomolar binder things like that like there are a lot of technical things that i think are we're like kind of right on the brink of solving um and for me it's like it's really motivating just to like pin those things down and just like get all of this over the line and see kind of where that leads to next.

46:43Josh Meier:Awesome. Matt, Josh, thank you for engineering biology and thank you for sharing your story with us today. Thank you guys. Thank you. Thanks for having us.

47:12Thank you.

From the publisher

Most people treat biology as a bespoke, messy science. Josh Meier and Matt McPartlon, co-founders of Chai Discovery, treat it as an engineering problem. They make the case that drug design obeys the bitter lesson: scale data, models, and compute, and the model can learn what a hand-built pipeline simply couldn't capture. The results are concrete: Chai-2 pushed de novo antibody design from a sub 0.1% hit rate to 16%, turning a needle-in-a-haystack search into something more like designing a key to fit a lock. Josh argues, counterintuitively, that biology is more verifiable than code, and explains why the goal should be more lab experiments, not fewer. Their bet: a design suite that collapses drug discovery from nine months to nine days, and arms the pharma industry rather than competing with it.

Hosted by Pat Grady and Sonali Singh, Sequoia Capital

00:00 Introduction
01:52 From Discovery to Design
03:25 Protein AI Breakthroughs Timeline
06:04 Why Start in 2024
10:13 Diffusion Models Intuition
11:41 Building the Avengers Team
15:22 Hit Rates and Scaling Laws
25:01 Molecular CAD Vision
25:24 Faster Design Loops
26:32 Future Drug Discovery
28:37 Platform Business Model
31:14 Partnering Reality Check
33:44 Data Flywheel Explained
37:16 Staying Ahead at Scale
39:44 Culture and What's Next

More from Training Data

All 110 episodes
Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling ProblemTraining Data · 47 min
Listen in VO