🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)

21 Jul 2026 · 1 h 30 min · 37 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Zara Therapeutics’ X-Cell “virtual cell” model for causal drug discovery, emphasizing that causal prediction needs causal perturbation data (not just observational single-cell data). They describe how genome-wide CRISPR perturbation screens plus single-cell RNA-seq train a diffusion-based foundation model that predicts gene-expression responses more accurately than linear baselines, and can generalize across unseen cell contexts.

Guests (backgrounds)

Bo Wang, SVP and Head of Biomedical AI at Zara; previously Associate Professor at University of Toronto. Ci (Chu) Wang? (spelled “Tsutru/Chu” in transcript), SVP of AI-enabled discovery at Zara; leads high-throughput biology; previously ~10 years at AI+big data+biology (InSitril; Verily/Google X).

Key claims

  1. Descriptive foundation models (e.g., trained on Cell by Gene) don’t outperform linear models on causal counterfactual perturbation tasks.
  2. X-Cell is trained on causal Perturb-seq data (pooled CRISPR knockdowns + scRNA-seq) to enable causal prediction.
  3. Diffusion language modeling + diverse biological priors improves generalization to unseen contexts.

Notable examples

  • “Wow moment” heatmaps aligning linear baseline vs ground truth vs XL predictions for JNC changes; XL matches ground truth better.
  • “Seven genome-wide perturb-seq campaigns” and “16 cell types / ~25M cells” (Pisces/Orion dataset mentioned).
  • Context-universal perturbations observed across screens.
  • Perturb-seq details: CRISPR guide RNA barcodes identify which gene is silenced; Cas9 targets promoters to shut off transcription.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Initial Impressions of XL Prediction

0:00 to 0:22

The discussion begins with an impressive prediction model that outperformed traditional baselines.

“It's visually very clear to see that XL prediction is much more similar to ground truth than the linear baseline.”

Introducing Guests from Zara Therapeutics

0:52 to 2:04

The guests Bo Wang and Ci Chu from Zara Therapeutics introduce themselves and their roles.

“We're really happy to have in the studio with us today, Bo Wang and Si Chu from Zara Therapeutics.”

Zara's Mission in AI Drug Discovery

2:04 to 3:18

Discussion about Zara's mission to enhance drug discovery using AI and high-throughput experimentation.

“My first name is incredibly difficult to pronounce unless you speak Mandarin.”

AI Platforms at Zara: An Overview

3:18 to 5:42

Overview of the three AI platforms at Zara aimed at improving drug design and patient outcomes.

“And at the core of our mission, we're using AI platforms to generate better therapeutics to advance patient care.”

Integration of AI Models in Drug Discovery

5:42 to 7:16

The conversation centers on how Zara integrates various AI models to improve drug discovery processes.

“I know there's a lot of interest right now in that third thing, maybe called translation from the lab to the clinic.”

Challenges in Data Collection for Drug Discovery

7:16 to 11:02

Exploration of the challenges faced in collecting sufficient and quality biological data for predictive models.

“where we're mostly working on computers.”

Introducing X-Cell: Zara's Virtual Cell Model

11:02 to 13:16

Discussion on X-Cell, Zara's new AI model designed to predict responses to genetic perturbations.

“Excel is Zara's first virtual cell models.”

Concept of Virtual Cells and Their Applications

13:16 to 14:00

The guests explain the concept of virtual cells and their unique approach to predicting cellular responses.

“then that has an impact on the larger phenotype of the cell, what the cell looks like, does, et cetera.”

Understanding Virtual Cells and AI Models in Biology

14:00 to 16:49

Learn how AI models can mimic cell responses and the concept of virtual cells.

“or predict the cell expressions or cell functions after certain interventions.”

The Challenge of Defining Virtual Cells

16:49 to 19:48

Explore the complexities and definitions surrounding virtual cells in biology.

“And before these foundation models, what happens in single-cell domain is that for every task, biologists have to choose the so-called specialist state-of-arts approaches.”
Show all 37 chapters

The Importance of Causal Data

19:48 to 22:59

Understand why causal data is crucial for predictive modeling in biology.

“Or can we even describe the spatial changes at different cellular resolutions?”

High-Throughput Biology Techniques

22:59 to 26:48

Discover high-throughput methods for generating causal datasets in biology.

“I think the field has come of age to do these at scale technique that we call high-throughput biology.”

Scaling Experimentation in Gene Expression

26:48 to 28:00

Learn about the engineering challenges in scaling CRISPR and RNA-seq experiments.

“So these are also recent technologies in the last decade that have been scaled that can let you read out the expression level of all 20 ,000 genes simultaneously from each cell.”

Scaling CRISPR Techniques for High-Quality Data

28:00 to 29:00

Learn about the challenges of scaling CRISPR and RNA sequencing to massive cell quantities.

“but you've used this to scale a simple perturbation response, which is individually maybe not all that interesting, to this massive scale of basically an arbitrary number of cells.”

Stem Cells and Their Experimental Implications

29:00 to 31:00

Explore the role of stem cells in drug discovery and the implications for research.

“Techniques that are published in academia used to be all about handling fresh cells.”

Advancements in Cell Type Differentiation

31:00 to 33:00

Discover innovative techniques for differentiating iPSCs into multiple cell types.

“as the first two datasets, which is the world's largest perturbsic data release at the time.”

Virtual Cell Models and Future Directions

33:00 to 36:20

Understand the significance of virtual cell models and their potential in research.

“It's not just the total number of cells or total number of sequencing reads.”

Spatial Transcriptomics: Enhancing Biological Insights

36:20 to 39:40

Learn about spatial transcriptomics and its role in understanding cellular interactions.

“We had Ron Alpha and Dan Baer from Noetic as guests recently and viewers who want to hear a little bit more about that.”

Foundational Models for Single Cells

39:40 to 42:00

Delve into the architecture of single-cell foundation models and their training methods.

“Getting back to Excel, this, you know, presumably can inform a spatial model as well, right?”

Diffusion Language Models in Gene Expression

42:00 to 43:48

Learn how diffusion language models generate gene expression data without assuming order.

“For DNA sequences, the order of ATTG make total sense to us, right?”

Understanding Diffusion vs. Autoregressive Models

43:48 to 46:01

Explore the differences between diffusion and autoregressive models and their applications in gene expression.

“So that's why we switched it from a CGBT-like model to the current Excel model, which is using diffusion language models.”

Incorporating Biological Priors in Modeling

46:01 to 49:09

Discuss the importance of integrating diverse biological priors in enhancing model accuracy.

“Similar to how like an image diffusion model kind of refines the image over and over again.”

Evaluating Model Contributions and Dataset Quality

49:09 to 53:08

Understand how data quality and model architecture contribute to performance in gene expression modeling.

“Do you now need to provide all of that context in order for the model to work?”

The Vision for Causal Prediction in Biology

53:08 to 56:00

Learn about the aspirations of virtual cell modeling for causal predictions in complex biological systems.

“I'm sure there's different choices of architecture and have different ranks of contributions.”

Predicting T-Cell Responses with X-Cell Model

56:00 to 58:07

Learn how the X-Cell model predicts outcomes in T-cells using unique data.

“We actually generated the data expressly for this purpose.”

Generalization Across Cell Types

58:07 to 1:02:19

Discover how the X-Cell model generalizes predictions across various cell types.

“And so we're very excited to follow up on those hits and validate them in the lab.”

Challenges of Virtual Cell Models

1:02:19 to 1:04:58

Explore the challenges and potential of virtual cell models in drug discovery.

“I mean, I think some of maybe your own models might also have had trouble beating linear baselines in the past.”

The Future of AI in Biology

1:04:58 to 1:10:01

Understand how AI can enhance biological research and the evolving role of scientists.

“Yeah, I think the field suffers from a lack of consistent and uniformly accepted benchmarks.”

Predicting Combinatorial Gene Perturbations

1:10:01 to 1:12:06

Learn about the capabilities of AI models in predicting gene interactions and their implications.

“are trained on single gene perturbations, but once the model is trained, you can actually predict combinatorial perturbations just on the model in silico, right?”

The Changing Role of Scientists in the AI Era

1:12:07 to 1:14:27

Explore how the integration of AI in research is transforming the role of scientists and the academic landscape.

“And certainly you can imagine students probably face 10x anxiety.”

Innovation and Funding in Academia vs. Industry

1:14:28 to 1:17:16

Understand the challenges and advantages that academia faces compared to industry in terms of research funding and innovation.

“Why do you think that, why shouldn't money just go to industry?”

The Importance of Open Science and Collaboration

1:17:17 to 1:20:56

Discuss the significance of open science and collaborative efforts for advancing research in virtual cells and biotechnology.

“So, you know, just thinking about the lab workflow that we do, a lot of these are building upon innovations that were first pioneered in academia as well.”

Future Directions in Academic Research and AI

1:20:57 to 1:24:03

Examine the future of academic research in the context of AI and the importance of interdisciplinary collaboration.

“Let's put all the resources together to generate next generation of virtual sales models.”

Symbiotic Innovation in Biology and AI

1:24:03 to 1:25:19

Exploration of how academia and industry can collaborate in biological data generation.

“I think we'll enter a field of an era of symbiotic innovation and cross-pollination of ideas.”

Waving the Magic Wand: Overcoming Bottlenecks

1:25:24 to 1:26:36

Discussing potential breakthroughs in protein measurement technology for improved data.

“which you could say maybe is AI and, you know, sort of high throughput experimentation or however you want to define that.”

The Importance of Temporal Dynamics in Cells

1:26:39 to 1:27:34

Highlighting the need for technology to measure cell states over time for better modeling.

“My hope is, I hope to see a breakthrough in sequencing technology not just the reduced cost, but sequencing technology that can sequence the same cells at different time points.”

Innovations in Measuring Cell Transcriptomes

1:27:36 to 1:28:48

Discussion on current limitations and future possibilities in cell transcriptome measurement technologies.

“Would you be okay with even just partial, like, small snippets of genes or maybe three prime regions of a small number of transcripts?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Ci Chu:And what really blew my mind away is when I saw the model make prediction, just print out the heat map of the JNC changes, look at the actual raw data, and line up the linear baseline prediction, the ground truth, and the XL prediction all together. It's visually very clear to see that XL prediction is much more similar to ground truth than the linear baseline. This is a wow moment I was talking about in the beginning. This is the first time that someone can put together not just one perturbsy, but seven genome-wide perturbsy campaigns together. Something that jumped out to us biologists right away is that some of the perturbations are context universal.

0:41Ci Chu:Hi, I'm RJ Haneke, CTO of Mirroromics. This is Brandon Anderson, who builds RNA therapeutics at Atomic AI. And this is the Latent Space AI for Science podcast. One of the themes that has run through the podcast is how the lab and experimentation and the real world have probably the biggest impact and have the most relevance to whether something is AI for science or something like B2B SaaS. We're really happy to have in the studio with us today, Bo Wang and Si Chu from Zara Therapeutics. At Zara, they're building with a bunch of other people, a AI drug discovery platform. They're using high throughput experimentation system to collect very large data sets and then training AI models that can predict the way that your cells in your body will respond to drugs and therapeutics.

1:42Ci Chu:Really happy to have you. Big fan of your work. Why don't you two introduce yourselves to the listeners?

1:50Bo Wang:Hello, everyone. My name is Bowen. I'm SVP and head of biomedical AI at Zara Therapeutic. I joined Zara about eight months ago. And before that, I was associate professor at the University of Toronto in Canada.

2:03Ci Chu:And I'm Tsutru. My first name is incredibly difficult to pronounce unless you speak Mandarin. So I go by Chu, as in Chewbacca or Pikachu. That's your favorite fictional character. I'm the SVP of AI-enabled discovery at Zara. I joined about more than two years ago when I was still in stealth mode. And here I lead the high-throughput biology group, generating the kind of data that will feed our AI models and also think about their applications. Before this, I spent about a decade at the intersection of AI and big data and biology. Previously, I worked at InSitril, leading the in vitro discovery platform there.

2:45Ci Chu:And before that, I was at Verily, which spawned out of Google X. Okay, so you are at Zara, the company which is on the Pareto frontier of confusing names and mega rounds. So Zara is, I think, kind of came out of stealth like a few years ago and just really big org kind of out of nothing. So I'm curious if you can explain a little bit about what is Zara's mission? What is their thesis statement? Like what is special about Zara? and kind of where you're going in the future? Yeah, Zara is an AI-enabled drug discovery company. And at the core of our mission, we're using AI platforms to generate better therapeutics to advance patient care.

3:31Ci Chu:And so we will be making drugs using different AI capabilities. There are three main AI platforms that we're building here. The first one is Protein Design, work that spun out of our co-founder, Dr. David Baker's group from UW. A lot of the current generation of protein designers are here in the company. So there, the thinking is to use advanced AI technology to develop molecules against previously undruggable targets. The second AI platform, I guess we'll spend a lot of time talking about today, is the one that Boleyn and I have been working on for quite some time and just released a preprint on.

4:10Ci Chu:That's the virtual cell or foundation model of biology work. There, the hope is to build an AI model to predict biology, exactly like you said, and predict what genes and drug molecules will affect cell biology. And the third piece, which we're beginning to build now, is patient representation models. And the goal there is to have AI models that can understand which patients will respond to which therapeutics. So hopefully together, these platform technology will help us make better drugs faster and with a higher success rate than previous technologies to transform what is used to be artisanal trial and error in the past into more and more into an engineering discipline.

4:51Bo Wang:I think what sets Zera different is not just the one billion. But also, I think Zera is one of the very few AI native companies for drug discovery that works from end to end of all sections of drug discovery. From as early as, you know, Target ID, and the M-Protein Design, Small Molecules, and two phase one, two, three clinical trials, We aim to use AI to accelerate every part of the drug discovery so that not only we increase the success rate of developing drugs, but also greatly reduce the cycle time so that we can have new drugs instead of every 10, 20 years. So hopefully we can have the cycle time so we have more useful drugs for patients.

5:41Ci Chu:That's really interesting. I know there's a lot of interest right now in that third thing, maybe called translation from the lab to the clinic. Where are the bottlenecks? You have these three models. What are the bottlenecks that you're addressing? And sort of like, how are you doing that? Why are you doing it that way?

6:00Bo Wang:There is an AI native company. Almost every part of the sections of drug discovery, we're trying to use AI to revolutionize how we develop drugs. So the early part, we built causal foundation models, or sometimes we call it virtual cell. Proteins, we have state-of-the-art protein engineering models, and we have also patient representation learning models. And I think what Zara is trying to do is not only we develop AI models, but also we create the right data sets to empower these models. And I think what's really made me excited to work at Azera is we always aim to connect three AI models together instead of letting them work individually by their own.

6:50Bo Wang:So when we design virtual cell models, we look for connections to that. Can we find targets that is easier to apply the protein engineering models? And then even when we design the cellular causal models, can we connect to patient representations? What are the right patient data to connect to the cellular models so that we have something to show clinical utilities? So I think what really makes me excited is, before I joined that, I'm kind of a professor in computational biology department or computer science department, where we're mostly working on computers. We look at the data, look at arrays, etc.

7:32Bo Wang:But once coming to Xera, what really excites me is that I get to talk to people like Chu, lots of drug hunters, extremely experienced drug hunters, to really understand their pinpoint. So when we design AI models, we think about questions that really excites biologists. So later, maybe we can talk about how one of the rewarding signals I receive after we develop X-Cell is that like it's a wow moment from biologists that this is the first time biologists actually find the model can predict exactly how these unseen cell lines kind of respond to different perturbations. So that's kind of the part really excites me is the integration of kind of dry lab or AI models to wet lab or the biology or even eventually to the clinical side.

8:23Ci Chu:With this clinical model, I know you guys are aiming to take a drug all the way to FDA approval and beyond. Where do we stand now? I don't know if you're able to talk about this, but are you able to collect data from clinical trials and tie that back yet? As Bo said, I think if you think about drug discovery process, it's easy, right? You just need to find the right target, make the right molecule, and find the right patients to give them too. Of course, each of those steps are incredibly difficult to get right. And so far, like I said just now, it realized a lot of untrodden error and guesswork.

9:01Ci Chu:And the main issue, I think, is that we don't have the right biological data, really the power, the training of a predictive model. And in protein design space, I think that's where we have seen the most rapid progress so far. That's partially because we have a lot of data, high quality data over 70 years curated by the entire community. People deposit protein structures into a database called PDB. We also have a lot of sequence data collected over the years from different genomes that can help inform the model as well. And it's these high-quality data that are collected and accumulated that ushered in this revolution in protein design and alpha fold and other folding models.

9:45Ci Chu:In the other domains, such as clinical model prediction, such as virtual cell, we are nowhere near the same kind of massive data that are high quality. And I think it's mainly a data limitation issue. So to your question, that's why we're very invested in generating this data, particularly causal data in cell biology in a lab. And that's, I think, what made it possible to innovate on the algorithm side as well to usher in virtual cell models. On the patient side, it's a very interesting question. perhaps that's one of the hardest data to get because getting access to high-quality patient samples is difficult in itself.

10:24Ci Chu:Getting it matched to the right clinical annotation so that you can actually learn the difference, the bridge between molecular data and clinical response, that's even harder. And you might be able to do that across different disease severities, but it will be harder to collect the right data to predict which drug treatment will or will not respond in a particular patient or not. And so that takes a lot of thought and a lot of careful curation to generate data out of. So we're beginning to go into that area, but hopefully we'll be able to share more soon. Awesome. Maybe we should switch gears now.

11:04Ci Chu:You just released Xcel. Why don't you guys describe? I'll butcher it.

11:10Bo Wang:Excel is Zara's first virtual cell models. It is an AI model that can predict the response to genetic perturbations. Certainly we can extend it to other type of interventions such as drug perturbations, chemical perturbations, etc.

11:27Ci Chu:So can you just describe for the non-biologists that are listening, what is a perturbation? What do you mean by that? In our cells, when Bo talk about genetic perturbations, our cell, human cell typically have 20 ,000 genes. Not all cells express every gene equally. That's why your eye cell, your skin cell, your heart cell, even though they share the same genome, they function very differently. A lot of that is determined by selective gene expression that determines the type and the state of the cell. So what we do is to build a model that you can in silico ablate certain genes from the cell. That is in silico perturbation.

12:13Ci Chu:That's to say, if I reduce the expression of this gene in a cell, what is the implication for the rest of the cells? What's the biological consequence? You basically turned the knob down on one gene. That's right. And then what happens to all the other genes in that cell? Correct. and the hope is of course to predict the effect on all the other genes but maybe even more things than gene expression such as the function of the cell. And that's important that that's therapeutically relevant because a lot of drugs are inhibitors and they function through exactly that turning down the activity of a protein or a gene.

12:47Ci Chu:And so we can start with gene perturbation prediction. The hope is that we can also go to pathway inhibition prediction so on and so forth. So a pathway is just a set of genes that all kind of talk to each other by this gene expresses a protein, that protein has some impact on another gene, and so forth and so on. There's this long-chain reaction of genes and proteins, and then so that's called a pathway. And so if you interrupt that or somehow change it, then that has an impact on the larger phenotype of the cell, what the cell looks like, does, et cetera. That's exactly right. Yeah. So you have what you call the virtual cell, or you're creating a virtual cell.

13:30Ci Chu:And virtual cells are very popular. A lot of people are interested in this concept. But I think your approach is somewhat unique or separate from what other people are doing. Can you explain what do broadly people mean when they say virtual cells? What are some of the distinct other strategies? And then what is your specific strategy that you're going for?

13:50Bo Wang:Certainly virtual cell is a very high-level term to describe an AI model that is able to predict or describe what cell looks like or predict the cell expressions or cell functions after certain interventions. It's a very high level concept. First of all, it was not a novel idea. We had virtual cell project almost 20 years ago, but back then, sometimes we call it virtual cell 1.0, is that people are trying to derive differential equations to trying to use mathematics to describe what's the response for certain pathway interventions, as you just mentioned, and by fitting these equations to different observations.

14:40Bo Wang:And largely speaking, that was a failed attempt in the sense that the biology is just way too complicated to write in a few predefined set of differential equations. Moving forward, with the rise of language models, I think that the idea of using AI models to mimic how cell responds to different interventions by data-driven approach started to get popular. And I think three years ago, almost just four months after Chagibit was released, our lab at University of Toronto published one of the early foundation models of single-cell genomics called SCGBT. It can kind of interpret it as a GPT-like model for single cells.

15:26Bo Wang:And it quickly become very popular in the sense that for the first time we have a foundation model that is able to tackle different downstream tasks using the same model, such as we can use the same model to integrate different batches of single cell RNA-seq. We can use the same model to predict multi-omic integrations.

15:45Ci Chu:Let's define those things. So batches integrate different batches of RNA-seq. So you have different equipment, you're all collecting, maybe you're collecting the same. Or data from different labs. Yeah, different labs, different time of day. Correct. Different phase of the moon, whatever. And those actually have a big impact on the data that you collect. And so there's a big problem of how do I even compare this data set to that data set when there's all this other differences that have nothing to do with the gene expression and just how I measured it.

16:17Bo Wang:We call that batch effect. We certainly want to remove the batch effect while preserving the cell types, which are more important biology we want to reserve.

16:26Ci Chu:So this is sort of like analogous to the tank problem in image classifiers, for example, as sort of the models pick up on these crazy spurious features, which have nothing to do with what you actually care about, the underlying biology.

16:38Bo Wang:The core idea of integrating different batches is to keep the biological signals while removing the batch effect. And before these foundation models, what happens in single-cell domain is that for every task, biologists have to choose the so-called specialist state-of-arts approaches. And with foundation models such as SGVD or geneformers, what we hope to bring is that one model that solves all the tasks in single cells. And with the popularity of foundation model, lots of researchers come together under CZI, Chen Zagreb Institute, and we published a perspective paper at Journal of Cell to coin, for the first time, coin the term virtual cell, almost virtual cell 2.0, in the sense, let's use data-driven approaches.

17:32Bo Wang:If we cannot describe, let's learn it. So that's the idea of virtual cell, so that general spin can we build a language model or language type of model to predict what the cell types look like, how the cell responds to different interventions, and eventually we can replace all the cellular experiments by simply running simulations on computer without even running the actual experiments?

17:59Ci Chu:Maybe for a bit more context, you can think about this as, so a virtual cell is just a general concept, but you think cells have 20 ,000 genes in them. And in most human cells, I think what, roughly, you know, 4 ,000 to 5 ,000 are usually active at any given time or expressed at reasonable levels. So you look at a normal cell, you might have 4 ,000 to 5 ,000 genes doing things. And so your question is, in many cases, the way medicine works is you target a protein or you target something which makes proteins more common or less common, or they stop the protein from doing something. And your goal is, given some number of genes which are in a cell, every cell has a different composition of genes, what is going to change?

18:48Ci Chu:You know, will some pathway die off? Will some pathway grow? And how does this, you know, from this you could predict how medicine is going to work by just understanding how changing one specific gene or some cluster of genes could change everything. Is that a correct understanding?

19:08Bo Wang:Yeah, that's a correct high level understanding about virtual cell. What's happening for this field is that we are lacking a concrete definition of virtual cells. And people almost equate foundation model with virtual cell. But in my view, virtual cell is probably a much broader concept than just foundation models. Foundation models mostly provide a reliable, semantic, meaningful representation of cells. But I think virtual cell is more dynamic in the sense that can we build AI models, even predict the development of different cell states across different times? Or can we even describe the spatial changes at different cellular resolutions?

19:55Bo Wang:In my understanding is that we are really at the early stage to develop such comprehensive virtual cell models. And the foundation model is really just the starting point.

20:06Ci Chu:AI models always begin with the data. You are building a high throughput experiment or have built and are continuing to develop a high throughput experimentation system. Can you that sounds really cool and really complicated? Can you tell us what that entails? Like, what are you doing? What are the experiments that you're running? How does that inform the building of an AI model? And why do this rather than pick up the Celex gene database, which is a collection of gene expression data that has been aggregated over the public data sets? Yeah, great question. I want to pick up where Bo left off.

20:44Ci Chu:I think Bo said something pretty profound, going from a representation model, the foundation model of biology, to a virtual cell. And the key difference there is perturbation prediction or dynamic processes in biology. That's a causal concept. For that, I think we need causal data. And if you look at Cell by Gene, that's a fantastic data set that curated in the beginning more than 33 million cells, now a lot more than that. And at the time when SGBT was trained on that data set coming out of Bose lab in Toronto, that was mostly an observational profiling data set. It's a descriptive data, not causal, and mostly profiling healthy human donors.

21:32Ci Chu:And so the model that was trained on this data set is very, very good at doing descriptive tasks, such as harmonizing across batch effects, removing effects from different labs, different technologies. But I think both us and many others in the field have found that these models that are trained on descriptive data do not yet outperform linear models on causal tasks, perturbational tasks, what we call counterfactual tasks. If I did this to the cell, then what would happen? That makes intuitive sense to a biologist because the correlation data in the descriptive data set can be fit with many, many possible causal structures.

22:14Ci Chu:In a very simplistic case, let's say you observe gene A, B, C all go up and down together in your descriptive data set. You can infer that A regularly B and C. That's why when A goes up, B and C also go up. You might also say that B regulates A and C, and that would be perfectly reasonable as well. You could also say that A regulates B and C is completely regulated by something different. You see the problem there. And there's n number of ways to fit a causal regulatory network into this group of data. Fundamentally, we believe our original data are underpowered to learn causality truly. And this is why we realized pretty early on that we need to really start building causal data set to train a causal model.

22:59Ci Chu:So what are the ways to do that? I think the field has come of age to do these at scale technique that we call high-throughput biology. And there are many ways to generate these causal data at scale. The technique that we have focused on is something called PerturbSeq. So for the listeners who are not familiar with that technology, it combines high-throughput pulled CRISPR perturbation together with single-cell RNA-seq technology to build 2D datasets. Let me break that down. Please. So we just talked about in a cell, there are at least 20 ,000 analytes to measure. These are the genes. These are both the features to measure.

Read the full transcript

23:40Ci Chu:These are also the lever to perturb the cells with. So for clarity, let's call them perturbations and gene expressions. For perturbation on one axis and the features that you measure that describe the cell on the other axis. Perturbseek is a technique that leverages the latest breakthrough in lab biology, CRISPR-Cas9. These are bacterially derived enzymes that allows you to disrupt gene expression in mammalian cells, in human cells, for example. And we can do so in one at a time fashion. So I can take out one gene at a time. Of course, that would be incredibly difficult to scale if I want to do all 20 ,000 gene expression knockout in one single experiment.

24:29Ci Chu:I probably need a huge factory, a lot of robots to do that. Or you can do them in a pooled fashion. And I love pooled experiments. These are hyperscalable. So we have lab tricks that allow us to disrupt one gene per cell, but do all 20 ,000 genes across many, many cells in one single pooled experiment. perfectly scrambled so there's no batch effects, there's no play-to-play differences. So that's all of the perturbation throughput. Basically use some sort of communitarian trick to first perturb all the different genes and different combinations, and then you can read them out and do some math on it, and you basically pull out a whole bunch of different experiments in one experiment.

25:12Ci Chu:Correct. It requires barcoding technology, and that barcode is actually achieved by directly reading out what kind of CRISPR guide RNA is present in which cell. So for CRISPR-Cas9, this bacterially derived machinery to work in mammalian cells, you just have to deliver two things to each cell. You have to deliver the protein, the Cas9 protein that does the job, and you have to deliver an address barcode encoded by a short piece of RNA called a guide RNA. and the guide RNA tells the protein where to go in the cell purely via Watson-Crick-based pairing, ATCG. So it matches a part of the gene. It's sufficiently long to say this will match the correct gene and then that guides it to connect to the right and reduce the expression of that particular gene in the cell.

26:05Ci Chu:Correct. We designed this guide to go to the promoter part of the gene. That's the beginning stretch of every gene. before the transcription starts. And if we bring the Cas9 protein to there, armed with the right effector, the silencer, that promoter will get shut off and that gene will never be transcribed out of again. So effectively we tune down the expression level of that gene. And so all you have to know is figure out which guide RNA is in which cell and that can be done using genomic readouts. That's the barcode. And you can then infer which gene is being silenced in which cell. So that's the way you scale throughput on the perturbation side.

26:43Ci Chu:On the readout side, it's a 2D data set, right? So we'll just talk about one of the dimensions. On the readout side, we leverage single-cell RNA-seq technologies. So these are also recent technologies in the last decade that have been scaled that can let you read out the expression level of all 20 ,000 genes simultaneously from each cell. So armed with both high-throughput CRISPR perturbation and high-throughput single-cell RNA-seq technologies, all of a sudden we can generate these 2D datasets where we systematically perturb or knock down every single gene in the human genome in the cell type.

27:18Ci Chu:And we read out this impact on every other genes in the same cells. So we generate these 2D-rich datasets, not that different than the type of PDB data that trained alpha-fold models. If you think about that, that's hundreds of thousands of protein entries. If those are the rows, columns are the XYZ coordinate of every single amino acid. That's also a 2D dataset. And I think it's these type of rich 2D datasets that power the training of foundation models of biology. I find it really fun how you have turned a fairly straightforward assay in using, this is NGS sequencing, next generation sequencing, very high throughput.

28:00Ci Chu:but you've used this to scale a simple perturbation response, which is individually maybe not all that interesting, to this massive scale of basically an arbitrary number of cells. I think you did 25 million or something. So it's actually a lot more than that. So 25 million is what came out of the most stringent quality filtering. It's actually as much of a scientific challenge to figure out how to do CRISPR and single-cell RNA-seq as it is an engineering challenge. In the first part of the experiment, oftentimes we have to harvest tens, if not hundreds of millions of cells. And they go through various quality funnels to give Bo and team the highest quality data at the end.

28:42Ci Chu:That's incredibly difficult to do because, as you can imagine, all of these techniques have been published by Academia before. And they work very well in small-scale experiments. But when you think about scaling them to a genome-wide perturbation, we're talking about handling hundreds of millions of cells. Techniques that are published in academia used to be all about handling fresh cells. Cells are still alive. And that may be okay if your entire experiment takes only an hour or two. It's not quite easy to handle cells across a 14-hour day that's hundreds of millions of cells. And so by the end of the day, I used to joke with him, you can easily detect stress signals from the cells and from your scientists in the lab.

29:25Ci Chu:And quickly we realized that's not the way to do this data generation. Machine learning is very quality dependent and we want to give the highest quality data to our AI teams. So we're putting a lot of engineering thought and industrialize the whole workflow step-by-step, introduce chemical fixations so that we lock the state of the cells in at the beginning of this experiment, but figure out ways that it doesn't disrupt all of the molecular biology steps afterwards. It doesn't impact data quality. So that we can do all of these data generation in a time-shifted operational manner that's very not prone to batch effects.

30:04Ci Chu:One thing that you didn't mention is that you're using some sort of stem cells. And so obviously, like, you don't have, you know, like brain cells or blood cells. Or if you did, then you would have a big combinatorial effect on that. So how are you, why are you convinced that working on stem cells, which are, you know, my understanding is that there are actually blood cells that have been sort of the stem cell behavior has been unlocked on them. And that causes some sort of stress on the cell as well. So you have these like sort of not quite blood cells that are stressed. stressed and then why are we convinced that that is a good proxy for a brain cell or whatever you're studying?

30:52Ci Chu:Yeah, not quite. So we didn't actually start with stem cells. That was more of a later development. When we started data generation, so we put out the method that I talk about as well as the first two datasets, which is the world's largest perturbsic data release at the time. last June in the preprint, we call a dataset XLS Orion. That was actually generated from two cell lines. A lot of this field's early work started with cell lines. These are cells. One of them is a cancer cell line. The other is just a cell line. These are immortalized cells. Some of them are derived from cancers, hence cancer cell lines.

31:30Ci Chu:Others are just derived from primary cells but have been immortalized, many times grown for many, many years in various labs. People start with these cell lines in the beginning, as you can imagine, because those are easy to do. It's easy to scale, easy to grow a lot of cells out of. Turns out the ability to grow millions of cells is actually critical for doing these large experiments. So we started there first. And they actually still capture the characteristics of the cell types that are derived from colorectal cancer, as well as hematopoietic cells. But later on, in the most recent preprint, we actually expanded to many more cell types.

32:10Ci Chu:Now, some of these are still cell lines. Our T cells, we chose to use cell lines. But some of these have now gone into primary cells. So we did one experiment in iPSC. These are induced pluripotent stem cells. And another experiment in, and we think this is the most ambitious and coolest screen that we've done to date. This is a pen differentiation multi-cell type stem cell project. So effectively, we differentiate iPSC into 10 different cell types in one single experiment without restriction. And we did a genome scale perturbation across them. So you can imagine, instead of just generating 10 ,000 different biological experiments, we did 10 ,000 by 10 cell types.

32:51Ci Chu:So it's almost a library-on-library experiment. Why are we doing this? We think that in the beginning phase of data collection, as Bo said, I think we're just in the early days of virtual cell building, context and diversity and richness of the data matters. It's not just the total number of cells or total number of sequencing reads. It's about bits per dollar in the information content. So we want to scale not only in the genetic perturbation landscape, but we also want to scale across biological context so that we can give our AR teams the best rich data set to build a generalizable model on.

33:29Ci Chu:Is there any thinking about, so you say context, but obviously these cells in these experiments have been sort of de, I forget the term, but they've been separated from their cohorts, right? Is there thinking about using spatial transcriptomics or other sort of technologies, imaging-based technologies, to build models with perturbations, but in the context of the cells that it lives near? Great question. So we're thinking about that in a couple of ways. Number one, that's actually exactly why we want to build a virtual cell model in the first place. You might think that, well, you can already do exhaustive screening in these cell lines.

34:09Ci Chu:Why do you still need the model? You can just do the experiment and generate the data. Certainly, if your query is just about cell biology and cell lines, you're right. We don't need a model, right? At least for genetic screening, we can just do the experiment. But you're also correct that oftentimes good targets, biological insights are not about cell lines. These are about primary cells, about cells in their native physiological context in organs, or even multi-organ coming together and have some emergent properties. A lot of immunological diseases are that way. You cannot do exhaustive high-throughput experimentation in animal systems or in organs or in all of these complex translational models.

34:50Ci Chu:You can do some experiments, and these are expensive and high-stake. The ability to build a model that can be trained on massive data where it is possible to scale and be trained in a way that can be fine-tuned and transferred to make high-quality causal predictions in these complex models so that we can go into the lab and have the highest quality hypothesis possible to validate. I think that's the whole point about building a virtual cell model.

35:15Bo Wang:But from AI side, I think you're absolutely right that I believe the future virtual cell model should be able to incorporate multiple modalities, not just RNA expressions. Spatial single cell RNA stick is already a popular technology. Even for SGBT, we actually have an extended version. We call it SGPT spatial, that is specifically designed for spatial single-cell omics. And we also have papers on early attempts to try to take the H &E images, trying to predict the gene expressions. There's already some signals you can find. So eventually what I predict is that a virtual cell model will be able to integrate not only RNC, can integrate more functionally related, for example, proteinomics or other regulatory side ofomics such as ataxic to overall combine all your descriptive omics datasets to predict the future states of the cellular functions.

36:19Bo Wang:I think that's probably the future for virtual cell model.

36:22Ci Chu:We had Ron Alpha and Dan Baer from Noetic as guests recently and viewers who want to hear a little bit more about that. I think we go in quite in depth there. So if you want to, so background, you can go to that. But can you explain a little bit about what spatial transcriptomics and spatial proteomics are? So maybe a bit of a history lesson here. Before we had single-RNA-seq, we had RNA-seq, and before that we have microarray technologies. What RNA-seq and microarray used to do is take a chunk of my tissue, grind it all up, put it in a blender, imagine, make a smoothie out of it, and take all of the RNA from different cells in that piece of tissue and measure all of their expression levels.

37:03Ci Chu:It is great. For the first time, you can measure gene expression all 20 ,000 at a time. We used to have to do them one at a time. But it is not great in that we don't know which RNA came from which cell. And this is particularly a problem if you're dealing with a multicellular piece of tissue. You want to attribute RNA to the immune cell, to the skin cell, to the fibroblast, to the keratinocytes, but you can't because you want everything up in a smoothie. What single-cell technology allows you to do is analyze them cell by cell. So now I can attribute RNA gene expression to the cell that they originate from.

37:39Ci Chu:But there's still a problem. I don't know spatially where these signals come from. And for many diseases, it matters, right? In immuno-oncology, for example, you want to know when T cells are close to tumor cells or when a T cell is not able to penetrate the solid tumor, what is the difference between them? Or when a T cell is attacking the tumor cell, when a T cell is not, what is the difference about that? And for that, you need spatial information. You need to observe cell in situ in their context. And so now there are different technologies that solve that problem. Essentially, take that chunk of tissue.

38:18Ci Chu:I don't have to grind it up anymore. I just make a cross-section, lay it down in a piece of slide, and I can measure its morphology using standard techniques like H &E staining. I can then also measure many protein expression using multiplex IF assays, immunofoescence assays. Ultimately, I can also look at the gene expression up to genome-wide in all of these cells in their native spatial coordinates by using some of the latest spatial omics assays. So you have the XY coordinates of every cell, but also all of the molecular analyses that we talked about earlier. And that's an exciting new direction for genomics field.

38:58Bo Wang:Certainly, can you imagine that the spatialomics adds more difficulty to AI modeling? Because you have, instead of looking at the individual cells, you have to look at the neighboring niche cells to better kind of learn the representation that is spatially cohesive. That is a challenge the current spatial foundation model are facing.

39:20Ci Chu:But that context is going to be crucial for understanding, let's say, cancer, where the interaction of immune cells and cancer cells and non-immune or cancer cells is absolutely crucial.

39:33Bo Wang:You can extract spatial aware biomarkers to predict some of the clinical response. I think that would be extremely important to build such models.

39:43Ci Chu:Getting back to Excel, this, you know, presumably can inform a spatial model as well, right? Because you have like one cell in one place. If you can imagine, okay, I can just throw away the coordinates and just do inferences on one cell at a time. And now I can create a more complicated model that does that, but it also knows who its neighbors are or something.

40:07Bo Wang:You're absolutely right. But the current version we were releasing, we're not dealing with spatialomics. However, definitely our ongoing work and the next version of Excel will be able to infer the spatially aware representations for different cells. I see.

40:23Ci Chu:We've talked about the data collection a bit. Let's talk about the architecture, get some red meat for the for the AI engineers listening in.

40:32Bo Wang:Sure. Let's get to the history of virtual cell modeling, particularly virtual cell 2.0. I think our SGP kind of sets the foundation for most of foundation models of single cells is that we adopted kind of autoregressive training, extremely similar to how Chaggbt is trained on languages, right? We use its next token predictions. So we mimics the way how Chaggbt is trained on languages to train the single cell foundation model on cells. by doing that, you have to assume an inherent order of genes. The way we assume the order of genes is by attention mechanism. There's many other methods that are using different orders of genes, some as simple as just rank the genes based on the expression values.

41:21Bo Wang:There's also more complicated kind of methods to rank different genes. But inherently, you have to assume an order of genes.

41:29Ci Chu:And just to be clear, so when you talk about genes, those are intrinsically ordered right they're a sentence spelled out in atgc right so genes themselves have this the nucleotides and there's this long chain and that makes a lot of sense to have an order to them but what we're talking about is something different that's the expression data the expression levels so the expression level means how it's just a count for each gene of how many of these genes did I see when I was measuring.

42:03Bo Wang:For DNA sequences, the order of ATTG make total sense to us, right? But for expression data, they're literally just matrices. So it's really hard to assume an inherent order of genes. Even if we shuffle the order of genes, I think the biology doesn't change much. However, because of the way kind of language model is trained, And everybody has kind of preset tricks to train such models. So it's easy to adopt. That's how all the foundation models are started for single cells. And then I quickly realized that with diffusion language models, we actually don't need to assume the order of genes. Instead, we can have a bidirectional diffusion process to generate such high-dimensional gene expression data sets.

42:51Bo Wang:So just to think about it, what's the difference between autoregressive training versus diffusion language models? You can think of autoregressive training as typing. For example, I like coffee. You have to type I and then like and coffee. There's inherent orders. But diffusion language model, you can treat it as editing. You iteratively generate a sentence from a very vague, very rough sentence. And then you can iteratively refine it. So same thing with gene expressions. You can generate a very rough representation of the gene expressions and then iteratively from noisy representation to more refined representations.

43:31Bo Wang:So you kind of iteratively edit the gene expression predictions until it minimizes the losses. So this is a very different philosophy to generatively predict the response after perturbation. It turns out it actually fits more to a single cell RNSEC. So that's why we switched it from a CGBT-like model to the current Excel model, which is using diffusion language models.

43:57Ci Chu:When I think of transformers, they're fundamentally objects which operate on sets. The community spends a lot of time trying to make them things which have some sort of causal ordering to them. But if you just naively take a transformer, it's a set operation, right? So given that, why think about this in terms of diffusion or autoggressive LLMs? Why not have your initial prediction strategy be something like take just a set of genes, each of which has its own kind of one hot encoded identity, and then use that as sort of a prediction? That seems like a much more natural architecture to me. And it's not just your work.

44:37Ci Chu:A lot of people work on things like this. And I have been somewhat confused why there's this bias in the community about this.

44:44Bo Wang:So I think what you were referring to is more related to representation learning, where you can take sets of genes and trying to project to low dimensional latent space. But what we care about for building generative modeling for virtual sales, because you want to predict the dynamics of sales. So we want to have a generative model. So that's why we're mostly using decoder-only architectures in order to generate the full transgretomics instead of just a predefined small set of genes. Because you want to model the whole gene-gene regulatory networks, which are extremely kind of high dimensional, right?

45:22Ci Chu:So just to be clear, input is genes plus a perturbation. Output is new gene expression levels. Is that gene expression levels plus perturbation is input. output is?

45:34Bo Wang:Plus times cells.

45:35Ci Chu:Like for each cell?

45:36Bo Wang:Correct. That is correct. Right. Okay.

45:38Ci Chu:And so the way that I think about, the way I think about diffusion language models, and you can correct me here because I don't know a lot about them, but the way I think about them is they're like BERT, but you do it over and over again. Is that kind of a good thing to think about?

45:55Bo Wang:Yeah, that is a rough understanding of how diffusion language model works. Yeah.

45:59Ci Chu:So you just apply the diffusion processes, like basically unmasking or editing once over and over again. Similar to how like an image diffusion model kind of refines the image over and over again. In this case, I'm using BERT. So it is a transformer, basically.

46:19Bo Wang:It is a transformer.

46:20Ci Chu:But it's like repeatedly updating the sort of sentence in this case, which is a bunch of expression levels over and over again.

46:31Bo Wang:That is correct. Actually, in our paper, we show that as the number of diffusion steps goes on, the loss function keeps decreasing. The fitness of the prediction to the ground truth keeps increasing. So this which means the model start to understand how iteratively refine the predictions.

46:52Ci Chu:I see. We're talking about diffusion versus autoregressive. What I noticed in the paper, there's a bunch of discussion of preconditioning using a whole bunch of stuff. Can you want to talk a little bit about that?

47:04Bo Wang:Another major innovation we made in Excel is the way we incorporate prior knowledge into the model. So incorporating biological priors has always been a good idea in biology in general, because biologists spend decades to understand some of the biologists already. How do we tell the model some of the prior knowledge, some metadata about the cells? Before Excel, what people do is they try to incorporate a single type of priors. For example, GEAR, using gene record network as a prior to predict the perturbations. SGBT sometimes trying to incorporate PBI as prior as well. Excel, to my knowledge, is one of the first models and trying to incorporate extremely diverse sets of biological priors.

47:56Bo Wang:So in our preprint, we incorporate five types of priors, including literatures, as simple as just ask SGBT, tell me everything about this gene. And then we embed as an output as the embedding. So that's a gene PT. Exactly, that's a gene PT. And we also incorporate PPI, protein-protein interaction networks. We also incorporate DEVMAP, which is cancer-related essential gene information, morphology information. We even try to incorporate CGBT embeddings, which is basically cell types. So with a set of prior knowledge as conditions to the model, the model starts to have more accuracy in terms of context-specific predictions.

48:43Bo Wang:And what's more interesting to us is that by looking at the weights of different priors, we can actually understand which prior knowledge are more important to these particular cell types. So it adds more interpretability to the models. So we find that combining diffusion language model plus a very diverse set of prior knowledges, Excel does much better in generalizing two unseen contexts. So this is some of the AI innovations we made for Excel.

49:13Ci Chu:Do you now need to provide all of that context in order for the model to work? Or those are like preconditioning that it can also do without if you want?

49:23Bo Wang:We don't need to incorporate this prior knowledge anymore because these are already learnable parameters inside the models. However, what you suggest is more promptable or in-context learning for virtual sales. We can do that as well, basically by adding more conditions into the prior knowledges so that to prompt the model to predict towards certain directions.

49:45Ci Chu:In other words, the model now takes advantage of the learning using the priors that you provided during training and doesn't need them, but has some advantage because you provide them during training. But you can even get more advantage if you are able to provide those priors during inference. Yeah. Okay. Wow. Nice. How much does that matter? I mean, whenever I see big machine learning papers with tons of things thrown in, I'm always wondering, where's like the big alpha and where's the little alpha? How much are, you know, is this just some, are these adding this little bit of incremental performance boost?

50:20Ci Chu:Or, I mean, are all these actually crucial to general generalization?

50:24Bo Wang:So there's multiple factors we have to consider. How much contribution the data contributed? How much of the contribution the AI architectures contributed? even for the architecture. What's delta from switching to autoregressive training to diffusion algorithm? What's delta from the prior knowledges? Certainly all of these needs very specific ablation studies. From empirical experiments, we find that the qualities, the amount of the datasets matter the most. This is why we were extremely excited to publish the Pisces datasets, which has 16 different cell types and across 25 million cells. And it's genome-wide.

51:10Bo Wang:It's kind of a huge tensor if you really think about it from a computer perspective. Genome-wide perturbation, genome-wide transcritomics plus number of cells plus at times number of conditions. So it's massive tensors. And because of the poor screening technology, we don't have batch effects. So you don't need the model to climb the hill of batch effect. That's already a advantage. So we find that trend on perturbation datasets, high-quality perturbation datasets, already gives a big boost to the models. We also did an ablation that if we train all the virtual cell models out there, including state cell to send us original STGVD on the same datasets, what's the data we are observing?

51:57Bo Wang:We reported the results there as well. We find that switching from autoregressive training to division language models give a significant improvement over some of the harder tasks, particularly generalized to unseen tasks. And the prior knowledge, more or less condition specific. For certain cell types, some of the prior knowledge make a huge difference, but for certain cell types, the data seems to be marginal. We are thinking about how to better incorporate the prior knowledge. We still believe that let the model know a big chunk of existing biology should be helpful. But maybe it's the way we incorporate the prior knowledge through cross-attention, limited the scope of the metadata.

52:50Bo Wang:But I think it's certainly a research topic. But overall, if we have to give an order, my order would be the quality among scale of the data sets and then the architecture and then the prior knowledge. Personally, this only applies to our Excel. I'm sure there's different choices of architecture and have different ranks of contributions.

53:13Ci Chu:First of all, this is really fascinating, very cool model. I hope everyone has a chance to look at the paper. um there's a lot of obviously a lot of resources that were put into doing this i don't know if you guys can disclose how much it's a lot of money whatever it was operating wet lab probably very complicated training runs um i think there's a couple four billion parameter model is that right or 4.9 yeah for five billion parameter model so much larger model um probably took a lot of GPUs to train. What's the lift that you get from this effort versus let's just put the money into wet lab work and the sort of traditional pipeline that basically has been the status quo up until now?

54:00Ci Chu:Biology is a multi-scale discipline. There are cells, there are DNA sequences on the most fundamental level. There are cells, there are multicellular pieces of tissues, co-cultures, you have tissues, you have animal systems, and finally you have human. I think we would like to be able to do causal protection towards the right of this spectrum, ultimately do causal protection in human, know what drugs will work in which patients, but that's very difficult to collect high-throughput data on. And so the whole vision of virtual cell is to generate data where it is possible so that we can transfer the causality protection towards the right towards the more translational, the more complex systems.

54:45Ci Chu:Certainly, you can mine the data already. We generated a lot of data, as Bo said, seven screens, 16 different biological context, genome-scale perturbation. There's a lot of good ideas in that already. There is a figure that we put out in the preprint that just looking to inactivation of T-cells. We already saw some, We saw known biology, TCR complex. We also saw some palliative new biology, which we're very excited to validate in the lab. Some of that were actually also caught out in a very recent screen last December published from Alex Morrison Lab, also in the Bay Area. So very excited to see that.

55:24Ci Chu:But the hope is to not just mine the existing data. The hope is that the model can generalize and we will be able to do in silico experiment into the future. Nobody knows before how much data and what kind of data are needed to do that. The whole field is waiting for the demonstration that the model can beat linear baseline in perturbation prediction. And it can generalize out of context, not just within a cell line you have training data on, but out of that context. That's why you need a model. So what's very exciting for us is that in this preprint, we saw that generalization capability. A few demonstrations.

56:05Ci Chu:We first did in T-cells. We actually generated the data expressly for this purpose. We generated a resting T-cell perturbation screen. So these are T-cells in their baseline condition, not activated. And now we have an activated T-cell perturbs-seq. So just T cell activation means I'm trying to kill something? No, these are regulatory T cells. But yes, we activate their receptors so that they're starting to proliferate. They become more active. They can do their physiological job. And critically, we only train the model on the resting T cell. And we told the model, hey, this is how the active T cell looks like.

56:51Ci Chu:now go and predict what all of the perturbation are going to do in this active T cell. And the model have not seen how perturbation works in active T cells. And we set up a couple of rigorous tests. One, linear baseline. Took the perturbational delta in the resting case, just transposed that linearly onto the active T cell, and that's our linear baseline. Essentially, think about this as a combinatorial perturbation prediction problem. One of the perturbation is activation of the cell. The other is all of the genomic perturbations. Can I just linearly add the two effects together? That would be a linear baseline.

57:28Ci Chu:And second, we apply other models from the field. And last, we applied X-cell. Critically, X-cell has not seen active T-cells. And it's able to make accurate prediction, not only on the known biology, the TCR complex, predicting their effect accurately, that these are going to inactivate the T-cells, which is exactly what we would expect to see. but also it predicted the putative T-cell inactivators that we found in the screen correctly as well. So that's very exciting to us, and that suggests the possibility that we might be able to use these virtual cell models completely out of context, in an unseen context, and predict new biology.

58:07Ci Chu:And so we're very excited to follow up on those hits and validate them in the lab. Just very briefly, a couple other cases that we saw exciting generalizing capability of this model. Remember, we did a multi-cell type differentiated IPSC experiment. There we specifically held out one cell type from training. So the model has not seen that cell type. Train on the other cell types as well as the rest of the data sets. The model made very good prediction across thousands of genes, thousands of perturbations in that unseen cell type. So again, suggesting the model's ability to generalize out of cell type.

58:43Ci Chu:And the last experiment I think we're very excited is that we trained this on T-cell cell line. But there was just very recently a primary T-cell petropecic published from Alex Marston's lab. That's an impressive amount of work. It's not easy to do this scale screening in primary cells. Very few labs have that kind of capabilities. Much easier to do that in T-cell lines. Again, the model is able to generalize out of cell lines into primary cells and make accurate predictions there. So they actually perturbed primary cells, not cell lines? Primary T cells harvested from donors. And we were able, for multiple donors, and Xcel trained on just one T cell cell line is able to make predictions across multiple donors from primary T cell experiments.

59:33Ci Chu:This is a validation of the whole theory, right? That you can train on these slightly weird cells and that it will be good because, you know, you're covering the domain well enough or whatever it is that you're able to actually predict in real cells that come directly from real people. That's right.

59:53Bo Wang:Yeah. I think building virtual cell is not to replace biological experiments, as you mentioned. And what we're trying to do, really the holy grail of virtual cell, is to have a model to generalize to unseen contexts that is harder or even impossible to conduct biological experiments on, right? So far, Excel focusing on cell lines. And eventually, we want to extend to more complicated biological systems such as animals, organoid, and eventually, as we mentioned before, to patients, to human biology, right? And bearing a few numbers, 90 % of disease has no cure. And most of the drug failed at the phase three clinical trials on patients.

1:00:40Bo Wang:And the success rate of these three trials is as low as five to 10%. And phase three means? The final stage on the patient trials.

1:00:48Ci Chu:So that's when you generalize from toxicity in phase two to efficacy in phase three. From a small cohort. Yeah, from small cohort. Into a much larger cohort. Oh, sorry. Toxicity is one, right? Two is small cohort. And so a lot of - To a large cohort. So the generalization problem, okay, this drug, I've very carefully selected my patients and it works pretty well. And now I get a bunch more patients and suddenly it doesn't work very well. And that's the big problem that you're -

1:01:16Bo Wang:The promise of virtual cell is that can we build such a model that learn all the causality biology so that can be grounded to predict the response eventually on patients so that we can, for certain drugs, we can select the right patients to conduct the clinical trial zone. So this is a long-term vision, but we already see some early hopes that Excel trend on diverse set of causal data sets can already generalize to some unseen cell types. So certainly there's a lot of experiments to be done to validate this model. So even continuously fine tune this model. But I think that we certainly see some early hopes.

1:01:59Ci Chu:So you were talking about linear models and this brings up this famous or infamous arc challenge about, you know, perturbation. And there's been this theme about complicated foundation models oftentimes not beating linear baselines. And I'd like to get your take about that is in terms of what is this, you know, first of all, is this different? I mean, I think some of maybe your own models might also have had trouble beating linear baselines in the past. Is there something different about your current data strategy or where you're going and where's the field going? And what is the role of foundation models versus, or in virtual cell models versus, these simple baselines.

1:02:42Bo Wang:A few things. First of all, those benchmarks, as you mentioned, are conducted on ReproLocal datasets, which are very small datasets. And the metrics people report are mostly MAE. Certainly you can imagine because single cell datasets are so sparse, the average profiles of the cells, certainly you can imagine is a great kind of local optimum to minimize the MAEs. This is why sometimes the average profile of cells has lower MAEs even than technical replicates, which are considered of ground truth for kind of perturbation experiments. So that itself shows that that metric is not reliable. However, most of these benchmarks are still comparing foundation models that trend on static expression data sets, such as cell-by-genes, SGB is often benchmarked against.

1:03:42Bo Wang:Internally, we also find that when it comes to MAE, sometimes SGB kind of fail to outperform linear models just because of the reasons I just stated. And what sets Excel different from these static expression models, such as SGB or gene formers, is that we actually, instead of trend on gene expression data sets, we trend on causal data sets. We train on massive amount of genome-wide perturbation data sets so that it learns better about the dynamics of the interventions. And in our preprint, we extensively compare with linear model as well. And as Chu mentioned, linear model totally failed to extend to unseen cell types.

1:04:23Bo Wang:You can quickly imagine why. And I believe that foundation model or other more complicated AI models that train on the right data will outperform these linear models in harder tasks, particularly in generalization tasks. And that's why I keep mentioning the right data set with the right AI model will lead to huge improvements. But I think the field still needs to see more biological validations to be more convincing that the virtual cell direction is the right one.

1:04:58Ci Chu:Yeah, I think the field suffers from a lack of consistent and uniformly accepted benchmarks. What gets measured will get improved. And in our paper, we measured, I think, one of the metrics that we put a lot of thought into and saw the model really shine is metrics around gene expression changes. So, you know, Pearson Delta, the similarity between predicted changes and ground truth changes upon the perturbation. That's very hard to cheat on. You have to really get the changes right. And what really blew my mind away is when I saw the model make prediction, just print out the heat map of the gene description changes, look at the actual raw data, and line up the linear baseline prediction, the ground truth, and the XL prediction all together.

1:05:47Ci Chu:It's visually very clear to see that XL prediction is very much more similar to ground truth than the linear baseline. This is a wow moment I was talking about in the beginning.

1:05:58Bo Wang:That's right.

1:05:58Ci Chu:And it's not hard to understand why. When we put, so this is the first time that someone can put together not just one perturbsy, but seven gene-wide perturbsy campaigns together. Something that jumped out to us biologists right away is that some of the perturbations are context universal, meaning that the genes do the same thing in regards of the cell types you experiment in. Might not be surprising to you that these are your housekeeping genes, right? Of course, they do the same thing every cell. And then there are all of these other clusters of genes that have very context-specific functions.

1:06:33Ci Chu:They do different things in different cells. Again, not hard to imagine why. In iPSCs in our stem cells, we saw developmentally relevant genes, genes that are important for neuronal differentiation. They only light up in iPSC experiments, of course. Right? That makes sense. So think about biology. It's so complicated. You have to capture these context-universal perturbation effects. You also have to somehow learn the context-dependent perturbational effects. It's not hard to then see why a very sophisticated, nonlinear model is able to capture and learn all of those biology much better. Are your perturbations always single gene perturbations?

1:07:11Ci Chu:Or do you have more? Because my understanding of regulatory networks is oftentimes, sometimes it can be a single gene does a ton of things. For example, I think males are just differentiated due to one gene being enabled at like day seven of embryo development or something. And that differentiates everything. It's this one gene. But then sometimes you have large networks of genes which all are very redundant, which allows for more subtle feedback mechanisms and so on. So I could imagine a lot of single gene perturbations as being kind of irrelevant and that you might want to start having a more commentatorial strategy here.

1:07:52Ci Chu:Yeah, that's a great question. And Bo and I have thought about this a lot. Actually, it's interesting that you brought up reproductive biology. I study a lot of female cells, the dosage compensation mechanisms. There are female cells have two X chromosomes. Male cells have one X chromosome. To match the dosage output from the X chromosomes, the strategy that the mammalian cells employ is one gene that produces a RNA that does not encode for any protein, just a non-coding RNA. That RNA wraps around one of the female X chromosomes and turns down most of the gene expression from that chromosome, shove it away in the corner of the nucleus, and it's called a bar body.

1:08:35Ci Chu:It is never heard from again. And so I absolutely agree with you. One gene can do a lot. But in biology, you also have redundancy. You have compensation. You have all kinds of mechanisms where knocking down one gene is not sufficient to always see a phenotype. What if four genes redundantly do the same thing, right? Taking out one is not going to be sufficient. So where we started with one cell type at a time, loss of function, single gene perturbation, and look at only RNA expression as the output, we're expanding the platform along all of those axes. So that's what we do today to build a scaffold of the data for training models like Excel.

1:09:16Ci Chu:We are now beginning to grow in all three axes of the platform, going beyond transcopotomics alone to look at multimodal data, going beyond just one gene perturbation alone to look at also pathway activation and inactivations, turning on and off entire cascade of gene, chain reactions. And also going beyond just cell lines, monocultures into more and more complex translationally relevant systems into primary cells, into organoids, and doing even direct in vivo perturbation swings. So we believe with all of that expansion, the data will be all the more exciting to train models on.

1:09:54Bo Wang:This is also why we incorporate PPI networks as the prior knowledge into our model. And although the model right now are trained on single gene perturbations, but once the model is trained, you can actually predict combinatorial perturbations just on the model in silico, right? So in the sense that you can just perturb the tokens of two genes at the same time and see what's the response. Certainly, without training the actual combinatorial perturbation data sets, the accuracy may not be there. But at least with the existence of such models, we can start to generate hypotheses using incitical perturbations.

1:10:33Ci Chu:How do you see the role of the scientists changing in the age of AI? And you may, because you're not using, your focus is not language models themselves and agentic science and things like that, then you may have a different, slightly different take. You're building, you know, these very specific models. but still one how are you and your students able to maintain such a high pace and i suspect it may have something to do with generative ai partly but also and how do you see the role of the

1:11:06Bo Wang:scientist the academic changing yeah that's a great question so my official split of time is 80 on zara 20 on my university affiliations but turns out the reality is 100 on zara 100

1:11:24Bo Wang:You invented a time machine. That's the answer. So the way I try to keep up is my lab use lots of agentic AI trying to monitor all the AI papers every day. And every week we have lab meetings. We're trying to discuss different topics in AI for biology, AI for healthcare, et cetera. And it's, to be very frank, even as a professor, I find it's extremely hard to catch up. The pace of AI is just so incredibly fast. And to the point that sometimes I feel anxiety, waking up says, oh my God, these people, so many people published and what happened to our existing unpublished work? And certainly you can imagine students probably face 10x anxiety.

1:12:13Bo Wang:So sometimes I try to encourage students to really kind of using different tools, trying to stay focused, you know, find is a niche area that we become experts on. Right. But with the era of generative AI, now agentic AI, I find the way people, at least academics, does science are extremely different now. Overall, most of professors or students in academics start to be very struggling in terms of fundings, in terms of the pace of publications. And that's why you can see lots of major breakthroughs come from industry, right? Like AlphaFo for example. So how academics survive or even thrive at such era of agentic AI is certainly something everybody's thinking about.

1:13:06Bo Wang:We see lots of faculties left university and joined industry simply for the reasons of resources, right? If you are doing research on AI, do you have enough GPUs at your school? It's the first question you should ask. When a student joins a professor's lab, the first question they often ask is, how many GPUs do you have? So certainly in that sense, industry has a major advantage over academic labs. But I think what academics labs have advantages on is really the kind of the pace of innovations and also the niche areas of these specific academic labs can be extremely expert on. So also having the freedom of thinking sometimes also make you easier to innovate on ideas that maybe industry people didn't even think about.

1:14:03Bo Wang:Overall, I think the whole field needs to be a lot more innovative to catch up. And hopefully with the help of different tools and hopefully the government start to invest more into academic because I still deeply believe that academic is the main source of innovation for the whole field and particularly when it comes to biotech. So hopefully we see more investment into academia so that we stay afloat.

1:14:33Ci Chu:I agree with you. But why? Why do you think that, why shouldn't money just go to industry? Why should the government put any money into academia?

1:14:43Bo Wang:I still believe the power of academic freedom. This is actually the original motivation we have. such the existence of academic professors who not only we teach, but also we do research. And there's also benefits of teaching and research at the same time, in the sense that when you teach a subject, you actually have to become the experts on, and that forces you to keep updating your knowledge base and then find easy ways to convey your knowledge to students. And by doing that, you actually start to innovate on different ideas. And myself, for example, I teach a big class in University of Toronto about deep learning and neural networks.

1:15:29Bo Wang:It's a gigantic class with 600 students every year. And by teaching that, it forces me to update the slides, lectures every year. And myself reading lots of materials, trying to update myself so that I can find ways to convey some of the knowledge to students, right? So that's how I keep updating myself whenever there's transfer. I remember vividly we have GP2, we updated the lecture. than different language models, how the multi-thread GPU communication is used in training, larger scale neural networks, et cetera. So I forced myself to just update the models. And by talking to different students, we really generate lots of novel ideas to apply the cutting edge AI models to very specific niche areas in biology or in healthcare.

1:16:22Bo Wang:Maybe it's unique to Canadian academic system. By being a professor in academic, we also have access to lots of healthcare data sets, which are very hard for industry to access due to many illegal reasons or regulatory reasons. And that's why you see some of the papers we publish through academic hospitals in Canada, where we develop some of the state of arts foundation model for ultrasound images. So that's what I mean, that there's a certain level of freedom of academic thinking that really drives lots of innovations. And I still believe, maybe it's biased, but myself still believe that having a certain level of freedom of academic thinking will lead to lots of kind of innovations that is unsinkable in

1:17:16Ci Chu:industries. I 100 % agree with Bo there. So, you know, just thinking about the lab workflow that we do, a lot of these are building upon innovations that were first pioneered in academia as well. CRISPR, of course, was discovered in academia. CRISPR applied to mammalian high-throughput screening, also demonstrating academia first. Singles RNA-seq, this kind of droplet encapsulated singles RNA-seq first demonstrating academia. Then different companies tried to build it up into commercial offerings. Putting all of these together to do perturbs you get first, also pioneered in academia, right? Chris Bach's lab, Jonathan Weissman's lab, Abby Ricker's lab, many of the pioneers in academia.

1:17:56Ci Chu:And then I think these, especially on the lab side, this innovation takes so long and the discovery process can be so accidental, right? That it's perhaps not ideal for pure industry to take on. But once they show early promise, scaling them and robustifying them and generating data that's not only massive, but high quality, especially for AI scale. I think that's something that can be very well done in industry, both the mindset as well as the resources that we can support. That sometimes can be hard for academia labs to match. Zara has been very generous with releasing your data sets and your models.

1:18:43Ci Chu:You seem to be very committed to open science. Given some of the things you were just talking about and the discussion about what is the best, most important data strategies for virtual cells or understanding human biology, where do you think academia should go next, now and next? since Zara has probably a budget comparable to probably dozens or hundreds of bio labs right now. What do you think that if you are an academic, a professor, especially in a wet lab,

1:19:17Bo Wang:what would you want to be focusing on? First of all, myself is a deep believer of open science. That's why all the models we talk about here that are open source. You can find all the data weights from my lab GitHub. It's very thorough. Yeah, thank you. And the reason I believe open science is that, as Chu mentioned, most of the time academic lab, we started an idea and we prototype it. It's not scalable. It's not even a good product. And industry can take it to scale things up. This is also why SCGBT quickly become one of the most widely used single cell foundation model in pharma companies. So that's very encouraging to us.

1:19:58Bo Wang:And this is why Zara is also start to open source some of the data sets, some of the models. Part of the reason is that we believe virtual sale is such an early field. And it doesn't help to withhold certain data sets or certain models because it's so early. A better win-win situation is everybody gets to in this field, start to contribute data together, start to contribute models together to exchange ideas so that this field can move forward in a much faster pace. We see successful examples in protein space, right? Because of the availability of open source data sets in PDBs, therefore we have models such as AlphaFold, RosettaFold.

1:20:40Bo Wang:Again, AlphaFold to also open source, therefore we can quickly iterate different models. That's why you see kind of a booming situation in protein space. we want to do the same thing for virtual sales. Let's put all the data sets together. Let's have the same standard protocols to generate high-quality data sets. Let's put all the resources together to generate next generation of virtual sales models. When it comes to academic labs, Stanley Chun can comment on what lab, how what lab academics can survive. But from dry lab perspective, I do encourage all the dry lab AI researchers in universities start to collaborate with industries so that they can get more resources to develop their own ideas.

1:21:25Bo Wang:And with the era of agent AI, now everybody can code. So it's more important to have a right taste about your project so that you don't just let just burn tokens without purpose, right? So we want more academic students, academic professors to have higher taste of research so that we make the right utility of agenting AIs.

1:21:52Ci Chu:As a professor, your job is to have taste, obviously. But as a student, how do you develop taste as in a world where so much of the thinking scientific process could be essentially outsourced to an LOM? or, you know, some there's not a world where you're forced to bang your head against something and learn taste by the hard way.

1:22:15Bo Wang:This is why we need academic training where you get into a field you know nothing about. And hopefully after you graduate, become the expert about this particular topic in the world. Right. This is why you have to go through different programs to talk to your peers, talk to your professors to get an idea about what's a good research taste to begin with. And but more importantly, I always teach students that the best way to learn something is to just program it. By programing, you kind of know what's the details hidden in all the mathematical equations in the paper, which often you admit. But in the era of agentic AI, things are slightly different in the sense that we used to spend lots of time coding, a little bit of time just debugging.

1:23:05Bo Wang:But now we let the agent do most of the coding, but we spend most of time debugging, which seems to be definitely interesting to me. And we had lots of discussion with the students in the lab, what's the best way to spot bugs by AIs. So how do you find places where AI is particularly good at? Also, how do you find places AI are still limited at is kind of you need lots of China era as well. And in the end, you still need to kind of validate your model using real world evidence, right? So that's why collaboration with wet labs, collaboration with clinical teams to validate your model, provide a feedback signals to your taste, as we discussed, is a very important training program.

1:23:58Ci Chu:On the web lab side, academic has extremely important roles to play. I think we'll enter a field of an era of symbiotic innovation and cross-pollination of ideas. Just like in AI field, I think we have great ideas coming out of academia still all the time. But industry now increasingly are contributing new ideas on architecture, on all of that as well. Well, in the WellApp side, certainly industry seems to be able to scale this type of data generation quite well. But biology is so much more than just cell-based PerturbSeq. Beyond RNA-Seq, we would like to measure many other analytes, right? Proteins, metabolites, lipids, protein-protein interactions.

1:24:45Ci Chu:How do we do that at scale? beyond individual cells, we would like to be able to measure cell-cell interaction, spatial, cell in their native context, or even whole animal level in vivo perturbations. Again, how do we do that at scale with lots of great innovation coming out of academia? Actually, just one great paper last week. And so how do we connect all of these together? I think we have many years of work ahead of us to fully crack data generation for all biology. And that, I think we need to scale, the industrialization, the innovation from industry, we also need that from academia. I think we'll together move this field to the next level.

1:25:23Ci Chu:One question that we've been trying to ask everyone is in your field, which you could say maybe is AI and, you know, sort of high throughput experimentation or however you want to define that. If you could wave your magic wand and have a bottleneck removed for you or a key problem solved, what would that be? protein okay if there is a way to do protein sequencing or high super protein measurement the same scale that we can do genomics that will be amazing you know i'm training genomics field but if i can do that i would incorporate that technology in a hard speed rna is amazing it foreshadows which proteins are going to get made but protein by and large are the functional units in a cell not only does their abundance matter their post-translational modification matter, their localization in the cell matter.

1:26:17Ci Chu:If you can measure all of those things, their conformational states, their modifications, their abundances, their localizations at scale, single cell or even spatially, I think such data sets would be incredibly useful to train the next generation of initial models. I know there's a lot of innovations in that direction. Can't wait to see that come of age.

1:26:39Bo Wang:My hope is, I hope to see a breakthrough in sequencing technology not just the reduced cost, but sequencing technology that can sequence the same cells at different time points. I think this is much lacking right now because in order to sequence the cell, you have to kill the cell. So can we have a technology that can measure the cell states at different time points for the same set of cells? I think that will bring a very different dimension to the data set so that we can start to measure the temporal dynamics of cells. So far, everything we measure, everything we model is extremely static. So can we have a technology that measures different cell response at different time points for the same set of cells will unlock massive opportunity to kind of model the dynamics of cells.

1:27:31Bo Wang:To me, that is a real virtual cell.

1:27:33Ci Chu:That's a really interesting idea. I never would have thought about that. Wow. Would you be okay with even just partial, like, small snippets of genes or maybe three prime regions of a small number of transcripts?

1:27:45Bo Wang:Yeah, we can start with a small set of gene panels to begin with, right? But eventually, since we're talking about the magic one, eventually, if we can have a system that observe how sales evolve at different time points and we have enough data to actually model such development, I think that would be real virtual sale modeling.

1:28:11Ci Chu:There's a small attempts in, oh sorry, earlier attempts in just, for example, sucking out portions of the cells, taking almost biopsies from the cell to do, you know, a fraction of the cytoplasm measurements. So that might be similar to the idea you talk about. There's also work from Paul Pellini's lab to have the cells secrete little vesicles. And they harvest that in the cell culture media to measure what the cells are producing longitudinally. But there hasn't been technology that can actually measure the entire cell transcriptome while still keeping the cell. You can't have the cake and eat it.

1:28:49Cool.

1:28:50Ci Chu:Yeah, thank you for taking the time to chat with us. It's been great. We learned a lot. It was a lot of really interesting discussions. And I especially really appreciate your commitment to open science and all of the cool models and data you've released. Is there any last thoughts you have or anything you'd like the audience to know to follow up with?

1:29:08Bo Wang:Overall, I think virtual sale is such a new and fast-moving field. We hope to have more and more people join us. And our Excel paper is out. And we look forward to receiving your comments and the feedbacks. And also, we're hiring.

1:29:23Ci Chu:Yeah, we are always looking for talented engineers, technologists, biologists, drug hunters, AI scientists, computational biologists. So look on Xero.com, look for the open roles. We'll be happy to chat with you. Thank you. Thank you very much. Great. Thank you.

From the publisher

Bet on information

If test loss flatlines after 1.5B parameters while training loss continues to drop as you scale, that tells you that your model is limited by the amount of information in your data.

Training on a single, smallish data set exposed an information gap: the 3.1B model falls off the scaling trend. Neither parameters nor compute will improve performance past this wall. For predicting changes to gene expression, you need more information rich data.

This is what Chu and Bo’s teams have done, and here is what ~30x the information buys you:

Now we can scale with parameters and training compute! We don’t know how much this effort costed, but we can guess that data collection experiments and infrastructure was a few tens of millions, and compute + headcount + research was a few million. The budget looks like a RL rollout budget, rather than a data rich pre-training one.

We were lucky enough to have the two central figures in this story on our podcast. Taking the lead from Ci Chu and Bo Wang, Xaira Therapeutics is betting that information rich data is the key to AI-driven drug development. Chu was recently promoted to Chief Discovery Officer and Bo to Chief AI Scientist, underscoring just how strategic Xaira considers this bet.

Reverse engineering the human cell

If you had to figure out how a human cell works, what would you do? A good place to start might be by documenting what genes are expressed (e.g. what RNA is floating around) in different kinds of cells, in different circumstances.

That is CELLxGENE, a database of 168M cells built by Chan Zuckerberg Institute that maps each cell to a count of how many times 20K-30K genes were detected in that cell, plus detailed metadata about every cell. A ~4 trillion-entry matrix.

If the Protein Data Bank (PDB) unlocked structural biology models [link Boltz, BioHub], CellXGene has done the same thing for Virtual Cell models. Like PDB, CELLxGENE has inspired a zoo of AI models of RNA expression; so much so that RNA expression models have become synonymous with Virtual Cell models. Bo Wang built one of the most influential, scGPT, that became the starting point for Xaira’s new model.

RNA expression ≠ Virtual Cell

Models trained on CELLxGENE describe the relationship between cell types and cell states, but they are not good at predicting what will happen if we make changes to RNA expression. Changes in gene expression are highly correlated, and its is difficult (impossible) to figure out what causes what in most cases.

If you could “turn the dial down” on one gene at a time, however, then you would be able to observe what is upstream and downstream of a given gene. You could tell if A → B & C or B → A & C or B → A, C → B → … If you did this for all of the genes, then maybe you could train a model that could predict what would happen to a cell if you change a gene (e.g. with a drug or a gene edit). Or maybe you could figure out the least invasive way to change a particular gene’s expression.

X-Atlas → X-Cell

This is exactly what Chu and Bo’s teams have done. The data set is called X-Atlas and the model is called X-Cell.

In this episode, we discuss:

* Why the team abandoned autoregression for diffusion

* The CRISPR-based experiments that run millions of tests in parallel, and generate the raw data for X-Atlas and X-cell

* Generalization to real lab experiments in real human cells

* Beating the linear baseline that has outperformed previous models

* Justifying a kitchen-sink of priors, and how that stacks up vs. data and architecture

Bo also shared with us some of the (major) advantages he has as an academic vs. industry leader, and how his labs keep up with the breakneck pace of AI innovation.

Check out the full episode on YouTube, or your favorite podcasting platform!



This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

More from Latent Space: The AI Engineer Podcast

All 247 episodes
🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)Latent Space: The AI Engineer Podcast · 1 h 30 min
Listen in VO