In short
Causal AI—how to answer causal questions (not just correlations) using causal graphs, probabilistic programming, and modern tools; includes a “three-rung ladder of causation,” links between Bayesian networks/graphical models and causal inference, and how LLMs can act as “causal knowledge bases.”
Guest backgrounds
Dr. Robert Usazuwa Ness is a Senior Researcher at Microsoft Research AI. His work focuses on statistical and causal inference for controllable, human-aligned multimodal models. He founded AltDeep.ai (teaches advanced ML). He has a PhD in statistics from Purdue University and authored Causal AI (Manning, March).
Key claims
AI has been dominated by correlation-based learning; causal reasoning requires explicit causal assumptions and the right data/variables. Probabilistic programming tools (PyTorch+Pyro, Stan, NumPyro) help separate statistical complexity from causal assumptions and handle latent confounders. LLMs (e.g., GPT-4.0) can outperform traditional methods on causal discovery/causal reasoning benchmarks by synthesizing causal knowledge from text, though they can hallucinate missing assumptions.
Notable examples
A gaming scenario where “side quest engagement” correlates with spending, but the guild confounds the relationship; causal modeling would avoid the wrong conclusion. A “vacuuming robot” example illustrating the need to explain user upset to improve behavior.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOExploring Causal AI with Dr. Ness
0:45 to 2:36
Discussion about Dr. Ness's background and his research focus on causal AI.
“Welcome back to the Super Data Science Podcast.”
Causal Narratives in AI
2:36 to 4:25
Dr. Ness explains the significance of generating and understanding causal narratives in AI systems.
“Robert, welcome to the Super Data Science Podcast.”
Humans vs. AI in Understanding Causality
4:25 to 6:32
Discussion on how humans and animals inherently understand causality compared to AI systems.
“the book was after reading his book, The Book of Wire.”
Traditional vs. Causal Inference in Statistics
6:32 to 8:05
Comparison of traditional causal inference methods in statistics and their implications for AI.
“And so there are these models are very much rooted in classical statistics.”
Cognitive Science and Causal Reasoning
8:05 to 10:41
Insights into how cognitive science informs our understanding of causality and AI reasoning.
“of having this kind of statistical association that's due to causality and due to non-causal relationships.”
Integrating Causality into AI Systems
10:41 to 14:00
Discussion on the predominant methods for incorporating causal elements into AI systems.
“Like if I you mentioned your dog, right?”
Connecting Causal Inference and Probabilistic Models
14:00 to 18:25
Learn how causal inference relates to probabilistic graphical models through software tools and techniques.
“or more generally probabilistic models, including deep probabilistic models.”
The Role of the Do Operator in Causal AI
18:25 to 22:37
Discover how the do operator in PyMC simulates causal interventions in data analysis.
“Let me try to summarize back some of the things that you said for our audience.”
Causal AI vs. Traditional Bayesian Approaches
22:37 to 28:00
Explore the distinctions between causal AI and Bayesian methods, including practical applications and assumptions.
“And so it's interesting because so far in this episode, we've been talking about Bayesian libraries that I'm familiar with.”
Introduction to Causal Inference in AI
28:00 to 28:58
Learn about the role of causal inference in AI decision-making and experiments.
“Because one thing that you can do with artificial intelligence is automate decision processes.”
Show all 23 chapters
Inductive Bias and Causal Questions
28:58 to 30:43
Explore how inductive bias influences causal reasoning in machine learning.
“And so that could come in the form of, as you mentioned, Bayesian priors.”
Challenges of Data Collection for Causal AI
30:43 to 33:19
Understand the limitations and considerations when collecting data for causal analysis.
“workflows that we're using in, say, for example, deep learning, where we're saying, like, all right, well, look, my algorithm seems to be making some, be able to, seems to be able to answer some causal questions.”
Challenges of Data Collection for Causal AI
35:31 to 35:44
Understand the limitations and considerations when collecting data for causal analysis.
“Next, develop and prepare to scale by rapidly building AI and data workflows with container-based microservices, and then deploy and optimize in the enterprise with a scalable infrastructure framework.”
Real-World Case Studies in Causal AI
35:52 to 42:01
Discover practical examples of causal AI applications in gaming and other industries.
“So are you able to give a couple case studies maybe of real-world situations?”
Understanding Causal AI with PyTorch
42:01 to 48:09
Learn about the integration of causal AI and Python tools like PyTorch for modeling.
“We'd have to basically make the assumption that being in a guild doesn't matter or not.”
Causal Reasoning and Generative Models
50:29 to 56:00
Explore the relationship between causal reasoning and large language models in AI.
“And so despite the shortcomings of large language models of LLMs, in your paper, causal reasoning and LLMs opening a new frontier for causality, you found that it's causal reasoning outperforms existing methods.”
Understanding Causal Inference with Generative Models
56:00 to 58:20
Learn how generative models apply causal inference in various contexts.
“the world, which, you know, typically, you know, if we're using Stan, we're using bugs, those models don't have, there's no built in kind of understanding of the meaning of words or a feature name.”
Exploring Multimodal Models and Causal Reasoning
58:20 to 1:01:30
Discover the integration of multimodal models and their implications for causal reasoning.
“And then kind of connect this all together.”
Audience Questions: Causal AI Concepts Explained
1:01:30 to 1:07:50
Engage with questions on causal AI models and Judea Pearl's ladder of causation.
“Let's jump now into some of the audience questions.”
Practical Workflow in Causal AI
1:07:50 to 1:10:04
Understand the typical workflow for tackling causal AI problems in practice.
“And hopefully Doug McLean finds that answer helpful as well as many of our listeners.”
Understanding Causal AI Methodology
1:10:04 to 1:14:30
Learn how to specify causal assumptions and utilize statistical methods in causal AI.
“That does side question engagement have an effect on in-game purchases?”
Book Recommendations from Dr. Ness
1:14:30 to 1:16:36
Discover book recommendations that provide insight into causality and thinking.
“something that I forgot to prepare you for is that at the end of every episode, I ask our guests for a book recommendation.”
Recap of Key Points and Insights
1:16:36 to 1:18:42
Review the main concepts discussed regarding causality and AI systems.
“by Richard Feynman recently, and it's just great.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:This is episode number 909 with Dr. Robert Usazuwa Ness, Senior Researcher at Microsoft Research AI.
0:14Jon Krohn:Welcome to the Super Data Science Podcast, the most listened to podcast in the data science industry. Each week, we bring you fun and inspiring people and ideas exploring the cutting edge of machine learning, AI, and related technologies that are transforming our world for the better. I'm your host, John Krohn. Thanks for joining me today. And now, let's make the complex simple.
0:48Jon Krohn:Welcome back to the Super Data Science Podcast. In today's episode, we're covering the fascinating field of causal AI, and we have exactly the right guest, Dr. Robert Osazua Ness, to be leading us on that journey. Robert is a senior researcher at Microsoft Research AI. His research focuses on statistical and causal inference techniques for controllable, human-aligned, multimodal models. He's also founder of AltDeep.ai, where he teaches professionals advanced topics in machine learning. He holds a PhD in statistics from Purdue University in Indiana. In addition to the above, Robert is the author of the book, Causal AI, which was published by Manning in March.
1:27Jon Krohn:I will personally ship five physical copies of Causal AI to people who comment or reshare the LinkedIn post that I publish about Robert's episode from my personal LinkedIn account today. Simply mention in your comment or reshare that you'd like the book. I'll hold a draw to select the five book winners next week. So you have until Sunday, August 3rd to get involved with this book contest. Today's episode will resonate most with hands-on practitioners like data scientists, statisticians, and AI engineers. In it, Robert details the three-rung ladder of causation that determines what types of causal questions you can actually answer with the data that you have, the surprising connections between Bayesian networks, graphical models, and modern causal AI, why AI systems have been dominated by correlation-based learning, and what's stopping them from adopting causal reasoning like humans and animals naturally do, how tools like PyTorch, Pyro, and Do-I are revolutionizing causal inference by separating statistical complexity from causal assumptions, and how LLMs like GPT-4.0 can act as causal knowledge bases outperforming traditional causal methods in some scenarios.
2:33Jon Krohn:All right, you ready for this deep episode? Let's go.
2:42Jon Krohn:Robert, welcome to the Super Data Science Podcast. It's a delight to have you on the show. So we have a fantastic episode scheduled today. I can't wait to get into the questions. We also had a lot of audience questions. We're going to try to get through as many of those as we can at the end of the episode. So those who asked their questions on social media beforehand, yeah, we'll get to you. But in the meantime, I've got tons of questions. So our researcher Serge Massis prepared amazing research on this episode. I was so excited when I worked through it. You're a researcher at Microsoft Research focused on causal AI and probabilistic machine learning.
3:18Jon Krohn:You're also founder of an educational platform called Alt-Deep. I'll talk about that more a bit later. And you're the author of the book, Causal AI, which was released just a couple months ago by Manning. Fantastic book. And in a glowing recommendation of your book, the Turing Award winner and creator of causal calculus, none other than Judea Pearl, who many of our listeners will know, said that your book, Causal AI, is a timely resource for building AI systems that generate and understand causal narratives. Could you unpack for us, Robert, what it means for AI systems to generate and understand causal narratives?
3:56Jon Krohn:How does the word narrative fit into all of this? What does it mean for it to be causal? Yeah, this seems like kind of an interesting starting point. Yeah, I mean, I think when Uta Pearl was saying that he was looking at the chapter in the book, is the last chapter that is looking at the intersection between causality and foundation models, namely, in this case, large language models. And I talk a little bit about multimodal models as well. And it's interesting because I think the first time I got the inkling to write the book was after reading his book, The Book of Wire. And he has this, so most of that book is really kind of bread and butter causal inference, but written as a popular science book.
4:42But he has one snippet in there, very short, but it talks about what it would be like if there were a robot that could understand causality if there was an artificial intelligence that had the ability to do causal reasoning. And it was a cool example. It had to do with a robot waking you up in the morning because it decided to vacuum and then you were upset about it and it had to understand why you were upset in order to improve its behavior. And I was thinking to myself, well, like, yeah, and what else? And it was, you know, so I said, okay, well, we need a book that's really taking an AI approach to looking at causality and using tools like PyTorch to actually write algorithms.
5:25And so that's kind of, maybe that was just the point that the book was initially conceived.
5:31Jon Krohn:Very cool. And so the concepts of causality, it seems to me like they should be more intuitive and analogous to human reasoning than something like correlation. And so like in that example that you just gave there, where you're talking about a vacuuming robot, you know, when a human or even when my dog makes a mistake, it's pretty easy for us to have an intuition around what's causing that. My dog knows that I'm annoyed at him because he's barking. He knows that I'm not just arbitrarily annoyed, and he knows that if he stops barking, I will stop saying stop barking. And so animals seem to kind of intuitively have these built-in systems for understanding causality.
6:20Jon Krohn:Yet the AI systems that we've built up until now, they're dominated by correlation-based learning. So why do you think AI systems develop like that? And what's stopping it from adopting more causal building blocks? If we think about, well, if we look at traditional causal inference, as you might see in a econometrics textbook or a causal inference textbook that's targeting people who work in epidemiology or taking more of a statistics approach. And so there are these models are very much rooted in classical statistics. and they're thinking about, okay, well, we have, you know, there's some kind of causal relationship between these variables, but there might be some confounding, which means that they might share some common cause.
7:12A statistical association is coming both through this causal relationship that they share as well as through this common cause. And we want to figure out how to distill the statistical association coming from a direct causal relationship from the overall kind of background noise of statistical dependence that includes that causal dependency as well as a non-causal dependency through that confounder. And, you know, so that's everything I just said there is very kind of statistical language, right? And packed in with some causal assumptions about the structure of causality between these variables.
7:55And, you know, as, say, insofar as we do have some use of, say, for example, deep learning and causality, it was still very much based on this kind of basic problem of having this kind of statistical association that's due to causality and due to non-causal relationships. And then what deep learning was doing when it was introduced into this space was to say, well, let's make sure we can scale it up with more data and work with, say, nonlinear relationships, maybe work in higher dimensions, et cetera. But still basic underlying framework. Now, kind of going back to your example about how humans or animals reason causality, reason about causality, there is a space in research that thinks about this thing a lot, these types of the ways that animals and people reason causally or otherwise quite a bit, which is in cognitive science, right?
8:54So in cognitive science, researchers try to understand, hey, how is it that people are reasoning about cause and effect? How are they making causal decisions? How do they understand why some outcome happened? And in other words, what are the causes or what is the main causes that led to this outcome, etc.? And it's a fascinating area of study because it's less about understanding actual ground truth, right? So, for example, if you think about causality in practical terms, let's say that there's a pandemic and people are sick. And I want to propose a vaccine that's going to help people.
9:41Jon Krohn:Well, you know, you really want me to be right. Right. If I say that there's a causal relationship between administering this vaccine and people not dying of this illness, the burden of proof is quite high. And so we're looking at that same type of rigor that we take in statistics that we may be thinking about statistical hypotheses, thinking about falsifiability, et cetera. But in this cognitive science field, the question is not what's true in the world, but the question is how do we write algorithms that do what we think is happening in those people's heads or in those animals' heads? And it's an interesting space because it turns out that humans, despite what kind of conversations you might have at your Thanksgiving table, humans reason really well about causality, relatively speaking, especially about what some call intuitive physics.
10:40Right. Like if I you mentioned your dog, right? Like if you walk into a room and there is something has been knocked down off of the table and you can see some evidence about the spread of the debris on the floor, you might make some pretty good inferences about whether it was your dog or your kid that knocked it down and how it was knocked down, etc. etc. Similarly, humans tend to be really good at what some call folk psychology, which is to say, you know, you and I could be sitting together in a cafe and watching people across the cafe having an argument in hushed voices and without actually hearing what they're saying, probably get a good guess about what it is they're arguing about, at least a theme about what they're arguing about.
11:27And so, you know, so there are certain domains where humans tend to reason about cause and effect fairly well. And so one of the interesting things that we can talk about from an AI perspective is to say, OK, how can we write algorithms that emulate those reasoning processes to kind of. yeah but this is in contrast to say classical statistics which is much more concerned about type 1 error false positive
12:03Jon Krohn:right right yeah so I guess so what you're saying is that there is a branch of study in cognitive sciences mostly where people are trying to figure out how the way that humans animals have these intuitions around causality we try to figure out some ways of packaging that into the way that models work so that they can do some causal inference, whether that's a statistical model or a machine learning model. And so it sounds like that's a relative niche, you know, relatively niche academic pursuit from the way that you're kind of explaining it. So if the way, yeah. I'm not sure I think so. I mean, I think particularly in AI, People are stealing ideas from other sciences all the time, right?
12:53Like in some sense, reinforcement learning is like the Pavlovian learning that you mentioned your dog and a dog quite learned in terms of responding to stimuli without actually really understanding a lot about how cause and effect is happening under the hood. or we borrow a lot from physics and other domains when designing loss functions and machine learning architectures. So I think this is just another type of approach is to say, hey, these guys are doing something here. What happens if we combine it with the stuff that we're doing?
13:32Jon Krohn:But I guess where I'm going with my question is when, would you say that the predominant approaches to causal AI today, maybe they don't really have to do with these cognitive science approaches. Like maybe some of them do, but are there kind of, are there predominant, what are the predominant ways? And this is maybe too big of a question and maybe we'll kind of get to this answer through other questions that I asked today. But what are the predominant ways that we add causal elements into an AI system? So one of the things that I observed in practice that made me really want to write the book was that people weren't drawing the connection that seemed pretty obvious to me between kind of graphical causal inference, like causal inference with a DAG, Directed Acyclic Graph, and probabilistic graphical models in statistical machine learning.
14:33or more generally probabilistic models, including deep probabilistic models. And, you know, insofar as they had very common origins, right? Like, you know, we already talked about Perl, you know, and Perl was a huge contributor to the world of Bayesian networks. And Bayesian networks in terms of, you know, if you take a Bayesian network and you interpret the edges in the Bayesian network as being causal, you get a causal model. In fact, a lot of the earlier kind of causal graphical models were based on the theory. It was very much part of probabilistic graphical modeling theory. And the do calculus, for example, that you mentioned is based on ideas of what happens when you fiddle with a graph in ways that simulates an intervention, for example.
15:29and actually doing an action on a data generating process. And so at some point, that area of probabilistic graphical models, we started developing this software that would implement it, right? So some of your listeners might remember tools like JAGs or bugs that were where we take something that looked like a graphical model and write it out as a program, and it was using inference like just sampling-based inference. And then the interesting thing about some of these probabilistic programming approaches is that they would use for loops, they would use control flow that you wouldn't really be able to do in a kind of conventional Bayesian network.
16:18But this technology continued to develop into tools like Stan, for example. And then also tools like, say, Web People or Gen and Julia, other probabilistic programming languages, languages like Pyro, which use PyTorch, or in the case of NumPyro, use NumPy and Jax as the inference engines, allowing you to kind of import learning algorithms from deep learning, for example, using stochastic variational inference. And so, but these all had a common root, right? So, and some of those probabilistic modeling approaches, they were very much adopted, in fact, developed by people in the cog-sci space. In fact, even going back to like Bugs and Jags, there was some famous Bugs or Jags book back in the day that was written by cognitive scientists, even though most statistics departments were using it.
17:19And, you know, so now this technology develops quite well. You can go into PyTorch and using a language like Pyro or an extension of PyTorch like Pyro, you can write fairly sophisticated, deep, latent variable models that are using modern deep learning architectures, doing things like variational autoencoders. using deep learning to do inference with things like stochastic variational inference, but are just as capable as of modeling any Bayesian network you saw in the 90s and thus capable of building a causal model that's built on a graph. And the fact that they can deal with latent variables makes them pretty powerful because I think one of the reasons that causal inference is a field kind of moved away from generative models is because they didn't do a very good job at dealing with latent variables, which is what a confounder is, right?
18:16It's something that's confounding your causal inference because it's unobserved and you need to deal with it. But now that's no longer an issue, right? Like these models can handle that kind of problem fairly well.
18:26Jon Krohn:Okay, nice. Let me try to summarize back some of the things that you said for our audience. So a lot of it made sense to me because I come from a time when I was doing my PhD many years ago now. JAGs and bugs were the kinds of tools that people were using for Bayesian inference. Now we have come into a time where tools like Stan are more common. You also mentioned Pyro, NumPyro, for those of us who work in Python, which is a lot of our listeners. PyMC too. PyMC also has some causal abstractions. For sure, PyMC. And so if you want to learn more about those kinds of tools, we had an amazing two-hour-long episode on Stan with Rob Trangucci back in episode 507 some years ago now, four years ago now.
19:09Jon Krohn:and we've also we've more recently had episodes on PyMC so I can really quickly look that up so for example episode 585 with Thomas Vicky he's so Thomas Vicky is CEO of PyMC labs and he yeah so we've got a great episode on PyMC there as well and you know to plug his stuff you know they added you know and a the main causal attraction being something that can model an intervention called a do operator. They added a do operator to PyMC. And so if you're a PyMC fan and you search that, you'll find some pretty good PyMC tutorials for causal reasoning when, I believe, for uplift modeling or like multimedia mixture models.
19:55Jon Krohn:Nice. And so things like this do functionality, what that is allowing us to do with a Bayesian model is to be able to try to simulate whether like try to simulate that a variable. So if you have a whole bunch of data where, uh, you know, all the data have already been collected. And so you can't really run a real experiment in the real world. This do operator allows you to kind of simulate one of your variables as potentially being the instrument, being the cause, like it, Like if you were running, going back to that vaccine example, you, you know, if you're running an experiment, you give people a placebo, you give some people a placebo, you give some people the real, the real treatment.
20:41Jon Krohn:And that is kind of the ideal. That's what we'd ideally like to be able to do is to run an experiment to determine whether that treatment does actually cause, you know, in your case, there are reduction in disease rates that it's an effective vaccine. but a lot of the time we've already collected the data. We just have a bunch of data. And so it sounds like what you're saying is that something like the do operator that we could implement in a tool like PyMC would allow us to, in some circumstances, simulate afterward, post hoc, whether that variable is in fact an instrument. Is that right? Yeah.
21:18So, you know, so a randomized experiment, you know, obviously it's a code standard. And the reason it's the gold standard is because, well, barring certain things that can come up in the actual implementation of the experiment, you don't need to make many causal assumptions between the treatment and the outcome. in a causal model any model that allows you to model intervention we can call that a causal model and what the causal model is doing is in exchange for explicitly providing some assumptions about the causal structure of the data generating process we're getting the ability to simulate the effects of an intervention into that process like we would like we would get by randomizing in a clinical trial or in a randomized experiment, a randomized control trial.
22:16And so it's not a free lunch here, right? It's not like these causal models are doing something magic. It means you don't have to run experiments. It just means that in exchange for adding some assumptions about the causal structure of the system, you can simulate what would happen if those assumptions were true.
22:36Jon Krohn:Nice. Okay, I got you. And so it's interesting because so far in this episode, we've been talking about Bayesian libraries that I'm familiar with. It seems like a lot of the examples are in a Bayesian realm. So I'm hoping that there's a funny answer to this or that you find this a funny question. But based on what you said so far, why wasn't your book called Causal Bayesian Statistics instead of Causal AI? So what's the distinction there? There's a little bit of Bayesianism in the book. And, you know, the way I think about it, I didn't go too far into Bayesian domain just because, well, number one, It's important to disentangle things when you can.
23:12One of the things that I like about or what I was trying to accomplish with my book was to say, here are the kinds of assumptions that you're doing with statistics. And here's the kinds of problems that you're dealing with when you're trying to scale up to larger data or you're working with higher dimensions. And here's the causal problems that you're trying to solve. and rather than kind of squishing these all together and trying to figure out how many, you know, a lot of these books, it feels like you're having to get a whole new master's degree just to kind of solve some of the problems of the book.
23:42We can say like, all right, well, you know, this is, you know, these things we're going to separate out. You can either use a library to do this for you, or you can, you know, rely on your existing knowledge or go and deep dive into that if you need to. And then the causal stuff is over here.
23:59So I think of Bayesianism as being in that kind of statistics box, which is to say that. But another way of thinking about it is that one of the things you're doing with Bayesian model is that you are injecting assumptions about the data generating process into your model. In this case, typically what's happening is that you're injecting your assumptions in the form of priors on unknown elements of the model. And so you're thinking about it. And it's kind of reflecting your certainty about the values or the structure of those elements as a modeler as opposed to what's actually true in the data generating process.
24:43And then from the causal perspective, we're injecting assumptions, but usually in the form of causal assumptions, say, for example, with a causal graph or alternatively in the form of, say, mechanistic assumptions between how the variables are connected. But so both the Bayesian and the causal approach are thinking about we need to inject some assumptions into our model to get better inferences.
25:16Jon Krohn:Okay, nice. So when we're talking about causal AI, what kinds of libraries would we use to do causal AI? What kinds of problems can we solve with causal AI that we might not be able to with other approaches? Like, you know, maybe there's examples from your book that you can provide us with that are kind of illustrative, illustrative. I feel like I'm pronouncing that word wrong, but that illustrate the value of causal AI. Like why should our listeners be, what kinds of circumstances could our listeners find themselves in where they should be thinking about using a causal AI approach? What do they get if they do that?
25:49Jon Krohn:And how would they do it? What kinds of libraries or approaches would they use? So there's the traditional answer to the question, which is to say that anytime you need to do an inference or that's not just about, say, predicting some variable given another variable, but actually trying to understand what would happen if you intervene in that system, right? So if, for example, again, if I needed to figure out if this vaccine, if some medicine or some, say, you know, some supplement that people were taking were actually having an outcome, having an effect on some outcome that we care about, say for example you know does does some new exercise supplement have an effect on on muscle gain you want to isolate out all the factors that have you know that's people who are already taking the supplement are taking maybe they're in the same maybe going to the same gym and they listen to the same workout podcasts and they have the same diet or something like that and so all these things you want to control for like you would an experiment and then you say like well i can't run an experiment here but i can make some pretty pretty good assumptions about the way that the system set up, given those assumptions, what can I do?
26:58All right, now you're asking causal questions. Once we get into the realm of AI and the kinds of questions that we typically think about in machine learning, it's interesting. It gets a little tricky because, in my experience, what tools like deep learning are really good for is kind of brute forcing their way through a lot of problems with a lot of data. And so really what you're trying to do is to say, like, what are the set of questions that can't be answered in this particular domain with just more data, right? Or with just some kind of clever architecture. When would you want to be thinking about using causal inference or causal AI?
Read the full transcript
27:47I think the easiest place to start is to think about when would you use causal inference in the first place? And then thinking about how you could use AI to either scale it up to larger data sets or higher dimensional data sets or to automate it. Because one thing that you can do with artificial intelligence is automate decision processes. And so again, when would I use causal inferences? When I'm thinking about an experiment that I'd like to run and maybe the experiment itself is expensive to run or it's infeasible. or I can run all the experiments I want, but there's some kind of opportunity cost, right?
28:27And so anytime I can simulate the outcome of an experiment in exchange for providing some causal assumptions for how the data generating process is set up, then I'm deep into causal inference territory. And then, you know, I think there's the question of, you know, where do we start to apply causality, causal reasoning in AI problems? You know, and the way I think about that is in machine learning, we're often thinking about, we often use this term inductive bias, right? So we're thinking about like, you know, what inductive bias means is that given some data and some induction, and let's say, for example, a prediction that I want to make, I need to have some kind of assumptions that are guiding the direction of the inference, right?
29:26And so that could come in the form of, as you mentioned, Bayesian priors. It could also come in the form of, say, the architecture of the model. Say, for example, using convolutions with max pooling to look at invariances of the identity of shapes as they have different positions within an image, for example. And so what causal inference theory does is it tells us what kinds of causal questions can be answered with a given set of data and a given set of assumptions. There might be some causal questions that you can't answer given your assumptions, even if you get more data. or you might say given my data and my assumptions I can't answer this question so I need more assumptions or I need more variables in my data rather than having a greater quantity of data, a greater spread of the data across observables and so that type of intuition you can bring into typical workflows that we're using in, say, for example, deep learning, where we're saying, like, all right, well, look, my algorithm seems to be making some, be able to, seems to be able to answer some causal questions.
30:58According to the causal inference theory, these conditions must be satisfied in order for that to be true. So if I know that, you know, if I'm doing my ablation study, or I'm trying to understand if my results are going to transfer to a new domain, that intuition from the causal theory is definitely going to help me as opposed to just kind of crossing the river by feeling for stones, to quote Deng Xiaoping.
31:26Jon Krohn:Nice, great quote there. So the idea here is, so we had already covered kind of earlier that any kind of circumstance where you want to be really careful about how a variable is impacting another. Not just that one is correlated with the other, but you want to be able to determine that X actually causes Y. And so it sounds like for the most part, it's kind of, is it often the case that there's more heavy lifting required or that you need to be more careful about the way that you collect your data or maybe the particular data that you collect in order to be able to do causal AI? Or can you just use data that you've already collected, make some assumptions?
32:13Jon Krohn:I guess what I'm trying to get at here, is there anything special? If somebody wants to do causal AI, are they going to have to do something special with their data or the way that they collect their data? Or is it just a matter of making assumptions? I realize I'm asking a vague question. I don't know the answer, but maybe there's some kind of examples in your head that you can provide that could help answer my question and give us like a clear picture of when we can use causal AI or when we can't. You know, so one thing that I think happens a lot in practice, particularly in industry, say you're working as a data scientist, you often don't get a lot of control over what data is collected, right?
32:53Somebody sends you an Excel spreadsheet or points you at some data warehouse and they say, you know, make some sense of this, right? And we all know that's not an idea, that ideally we would like to be there at the point of data collection and perhaps provide some input into what variables get collected. if we're going to be thinking about causality yeah we need to be we need to make sure that we are making causal assumptions that reflect what we honestly believe about the data generating process as opposed to you know what's convenient to what variables we happened to collect in our data set right so you know one thing i often tell students when i teach is to say like the causal inference doesn't care about what data you're stuck with.
33:48If your data is insufficient to make certain types of causal conclusions, you can't coerce the data into doing so. We know this from statistics. When I was doing my PhD, one professor said something that stuck with me, which is to say like you don't have any good methods. There is no method in statistics that's just going to remove bias from your data, right? Like you can get more data, you can use a fancier model, but it's not going to extract the bias. The only way you can deal with that is by, say, modeling it explicitly with some assumptions that aren't coming from the data but are coming from you.
34:30And that's the same issue that we have in causality, right, Which is to say, you know, we talk about, you know, confounding. Does this supplement affect muscle gain or are there some confounding factors? What we're saying is it's a question of bias, right? There's a confounding bias. There's something that's, you know, we look at this association between the treatment and the outcome and we want to interpret it causally. But there's some other signal leaking in that's biasing that conclusion. And so in that case, you know, we either have to rely on assumptions or we have to collect more data that's covering things that we're not yet covering.
35:31Jon Krohn:Next, develop and prepare to scale by rapidly building AI and data workflows with container-based microservices, and then deploy and optimize in the enterprise with a scalable infrastructure framework. Visit www.dell.com slash superdatascience to learn more. That's dell.com slash superdatascience. Okay, okay, cool. So are you able to give a couple case studies maybe of real-world situations? It could be examples from your book or maybe examples from your research or examples from Microsoft or, I don't know, some other job you've had in the past or some client you've worked with where you're able to kind of explain to us, these were the data that we had, these were the assumptions we made, and this is the conclusion that we were able to draw.
36:19Jon Krohn:We were able to make this causal conclusion that if we hadn't done that, we wouldn't have been able to achieve this commercial outcome or this research outcome or something. That's an interesting question because of all the practical examples that I could think about from Microsoft are definitely confidential because they all involve products. There must be something from your book. Everything in your book is hands-on. It's got lots of PyTorch, lots of real-world business examples. so yeah there must be lots of examples that you've already published in there even yeah so in the book one example that i lean on a lot is this and it's a simple example because it's it's meant to be there to be understood um but it's an example of um you're trying to imagine that you're a data scientist at a gaming company say like a you know an online game where you're you know an online role-playing game, for example.
37:23And in this game, you can go on quests, but you can also do kinds of side quests, right? With your clan or your, what do they call that? Like your crew or whatever. I'm sorry, I'm bad at video game lingo, but with your guild, I don't know. So you remember, you remember and so like let's suppose that uh your question is does engaging in more side quest um increase the amount of money that a player uh pays for on virtual items within the game right and and so the question and so in the example i show just kind of in the statistics are all laid out pretty clearly that if you just kind of look at the correlation between engagement in side quests and amount of money spent on in-game items, it sounds like that there's a strong, there's a kind of belief in the book, it's a positive relationship such that more side quest engagement leads to more purchases.
38:33But when we account for a third variable, in this case, it's membership in the guild, right? So we can assume that people who are in the guild are probably going to go engage in side quests together or, you know, or discourage each other from engaging in side quests. So guild membership being a cause of side quest engagement. And of course, maybe people in a guild pool their resources or say, you know, hey, John, you should buy the sword and I'm going to buy the healing potions, right? And so that is going to be a cause of in-game purchases. And so this becomes like a confounder if you're not measuring it directly.
39:10And so kind of in the book, I walk you through kind of how you would, you know, how you would kind of get the answer wrong and report the wrong answer to your boss. If you were to just say, look at the correlation between side quest engagement and in-game purchases, how, if you ran, if you ran like a, a, an AB test, a randomized test that you would get the right answer, but this would involve essentially forcing people to engage in side question gauge. If you have to randomly, somebody logs onto the game, you're like, hey, here's some side quests, do it. And now, because you've randomly assigned them to that group, and they don't want to do it.
39:47Or maybe some people are getting a better experience than other ones. And so sometimes in the real world, running the randomized experiment is infeasible.
39:57Jon Krohn:Yeah, it sounds like an economics kind of situation where the human knows that they're in an experiment if they're being forced to do one thing or another. What you're trying to get at is if the human in the wild in this kind of economic free world that we live in, which the video game is simulating there on a slightly smaller scale. If you're trying to answer some question, running a real controlled, randomized controlled trial is not going to be feasible, but you might be able to use the data that you have already collected that you've observed and so then yes in your case there that you're describing there's if there's feasibility one thing to jump in there there's feasibility but there's also you know ethics right right like you know if i wanted to understand the effects of caffeine on miscarriages clearly i can't ethically run that experiment um not on humans and then uh and then there's a question of just the old-fashioned question of cost, right?
40:59But just kind of say like, and this happens again, not just in kind of theoretical questions or questions in economics and medicine, but also in an industry, things like optimizing a video game. But please, sorry, I interrupted. Please continue.
41:17Jon Krohn:No, no, no, not at all. It's your episode. I'm just trying to divine some information. Okay, so now we have a clear understanding of a kind of causal problem that we might want to solve. So I'm a data scientist at a gaming company and I want to figure out whether the users, whether the users of this game, when they tend to engage in more side quests, does that cause them to spend more money on in-game assets that they can be buying? And so there are potentially confounding variables out there that like things like being a member of a guild that you mentioned there. And so if we weren't collecting that guild data, we'd have to have more assumptions.
42:01Jon Krohn:We'd have to basically make the assumption that being in a guild doesn't matter or not. And so it seems like, so this has made clear that there's a lot of assumptions, more thinking potentially about your problem that you need to do if you're engaged in causal AI. So that's a great thing to understand about this. But to kind of get into the nuts and bolts, your book does a great job of using PyTorch code, using examples to make causal AI or causality in general, which is often a very theoretical, difficult to understand topic, because of all of your examples and use of code in the book, it makes understanding causal AI more intuitive.
42:45Jon Krohn:And so let's say that you were the data scientist, Robert, at this gaming company. what python tools would you use um to then to then to then do causal ai how would you model this in order to uh in order to come up with a causal conclusion so one of the things that i was i had mentioned that i was trying to do with my book was to separate out the abstractions that have to do with statistics and uh computing right like scale it up um you know algorithmic complexity from the causality. And what's cool about, you know, like the libraries that we have today is that they can actually help us. If we're able to separate those abstractions, then we get to focus on one thing while leaving the nuts and bolts to be handled essentially by the library.
43:38You know, to some extent you saw this, you mentioned you interviewed somebody who talked about Stan. And what's cool about STAN is that that inference algorithm Hamiltonian Monte Carlo is, I mean, you can go in there and understand it's not, you know, well, it is physics, but it's not rocket science. I'm like, well, it kind of is rocket science. But you can still kind of just specify your model, specify what are the parameters, what's the model. And as long as you satisfy a certain set of requirements, I think mainly that all these things have to be continuous. that the inference will kind of just work for you without you having to like go implement your own inference algorithm.
44:24It's the same thing here, right? Where like if you can, as long as you can kind of specify your causal assumptions in some cases in the form of a graph, for example, then you can rely on say graphical causal inference
44:47algorithms from, say, probabilistic graphical models to kind of handle the inference there for you. If you're implementing it in PyTorch, I have plenty of PyTorch examples in the book, as long as you can incorporate your causal assumptions in the structure of the model in your algorithm and writes a kind of basic inference algorithm that has a differentiable loss function, then PyTorch is going to handle all of the nuts and bolts of the inference for you, right? That's kind of why we invented PyTorch, to say, well, if I can differentiate it, then I can get a gradient, then I can just turn it into an inference problem there.
45:42And so that's why, you know, so I do have examples there in PyTorch that are saying, okay, well, let's just, let's not worry about whether or not, you know, we need to use linear regression here or propensity scores or double machine learning or instrumental. These are all different types of kind of statistical methods for doing the inference you want. And you can learn all these things. Great. And there's great books for that, right? I think I can name a couple off the top of my head. But you can also say, let's work with some libraries that are just going to handle that stuff for us under the hood and kind of treat it as either an objective function to be optimized or as a configuration parameter in some model specification and then focus on our ability to think causally and write that thinking down in the form of a model and focus on the actual domain network modeling as opposed to all the inference stuff that we need to do to get that to work.
46:52And so to answer your question, I talk a lot in the book about using deep probabilistic models like modeling with libraries like Pyro as well as some more conventional tools like the DoY, DoY from the broader PyY suite, which is a big collection of causal inference libraries. And so even in DoY, right, there are different types of statistical techniques that you can use to estimate a causal effect. But at the end of the day, you're thinking more about what are your modeling assumptions and can you answer the question given your assumptions and your data and then if you can, you want to get to an answer and then all of the various statistical approaches you can take to arrive at that answer given your assumptions and your data.
47:45You can kind of just toggle between them and see which is giving you more stable results, for example.
47:51Jon Krohn:Okay, so it sounds like a lot of these things we could do at a lot of causal AI modeling we could do in Python with PyTorch but that might be unnecessarily low level. And so there are other libraries, a high level. And sorry, exactly. And so there are other libraries out there like Pyro, which we talked about earlier in the episode. Pyro was just an extension of PyTorch. So that's still in PyTorch land. Oh, I see, I see, I see. And then PyY, is that kind of a similar? No, so PyY, you know, you want to work with kind of conventional numeric data or categorical data that you have in like a data frame.
48:34Do-I is definitely a way to go. You would go with Pi Tours if you really need to work with neural nets inside your causal model for whatever reason or that you really need to work with. Say your variable instead of being like treatment, aspirin or placebo, outcome, healed or not healed, maybe your treatment and the outcome variable are, you know, it's a vector or a matrix, right? Because you're working with, say, media, for example, you're working with some kind of rich, high-dimensional data. There's still just variables, but now because you're working in higher dimensions and the relationships are all nonlinear and you need a lot of data, you want to kind of be working with tools that were designed for that.
49:18Well, if you're working with the kind of things that come into a Pandas data frame or an R data frame, you know, things like do I are our way to go.
49:27Jon Krohn:On this podcast, I'm always going on about how Claude Code is mind-blowing, but now Claude Cowork is making my jaw drop as well. For example, I recently wanted to quantify how healthy my sales pipeline is for my AI consulting business. I simply asked Claude to estimate my sales for the coming quarter, and it brought info from relevant Google Sheets and my Gmail to create a professional spreadsheet of clients with estimated revenue for each one. Whoa, this might've taken me a day. Instead, it was done flawlessly with Claude Cowork in minutes. Claude is the AI for minds that don't stop at good enough.
49:58Jon Krohn:It's the collaborator that actually understands your entire workflow and thinks with you. Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. Ah, and you'll appreciate that I can ask Cowork to show me data such as my sales spreadsheet, and it provides an interactive chart right in the conversation. For problems worth solving, get started with Claude at Claude.ai slash superdata. That's Claude.ai slash superdata, and check out Claude Pro, which includes access to all of the features mentioned in today's episode.
50:28Jon Krohn:Claude.ai slash superdata. Okay, cool. Thank you for that. And so my last kind of big technical question that I have for you before we get to some of the audience questions is that in the last few years, AI agents, generative AI, and LLMs have all been widely featured in causal conferences and papers, and you've been actively contributing to research at this intersection. And so despite the shortcomings of large language models of LLMs, in your paper, causal reasoning and LLMs opening a new frontier for causality, you found that it's causal reasoning outperforms existing methods. So tell us a bit more about this, about the paper, about the relationship between generative models and causal reasoning, because that isn't something, you know, generative models, agentic AI, they're the hottest kind of topics we can be talking about right now.
51:20Jon Krohn:And that hasn't come into the episode yet. So I'd be interested in hearing your thoughts on that intersection and kind of what's possible there. You know, so in the book, well, in the paper that you mentioned, we're essentially trying to interrogate the causal reasoning abilities of a large language model. And that paper had focused mostly on GPT-40. And, you know, so we were looking at, we had some benchmarks for answering causal questions. we created some benchmarks for what we call causal discovery, which is to say you have two variables. Is there a causality between A and B? And if so, does A cause B or does B cause A?
52:06And essentially what was happening there is that it works pretty well. It was doing so without data, right? So as opposed to trying to infer a causal relationship based on statistical association data, it's inferring a causal relationship based on relationships it's already learned semantically from or semantically is wrong. Kind of correlations is learned between tokens in natural language text from the training data. Right. And so and so in that in that in that sense, what the what the large language model was doing was kind of acting as a kind of causal knowledge base. Right. So, you know, if you wanted to if you wanted to infer whether or not there was a relationship between, let's say, for example, in biology, some some genetic product like a protein and some condition.
53:04condition, right? And some papers were written on this and maybe, you know, the LLM is able to kind of synthesize to some extent the information across all these papers into some kind of coherent statement about the causal relationship between these two variables. And, you know, so there was a lot of that, you know, so in the last chapter of the book, I talk about the first half of it.
53:33Jon Krohn:I say, how can we use these foundation models to kind of be oracles for causality, right? So if, for example, I want to know what the causal structure of this system might be, I can prompt it to propose a DAG. And then given that DAG, I can prompt it to propose some DOI code that allows me to implement that DAG in Python and run the analysis. If I get a bug on that code, I can plug it back into the AI and tell it to fix it for me. And then I can, similarly, if I, you know, a lot of what we're doing in theory is to take some things that we would say in natural language, like, hey, I think this vaccine might have a positive effect on this, on preventing this illness, right?
54:17And actually formalizing that. What we have to do is formalize those into variables and to, and relationships between those variables, right? So like there's this, there's this step of taking your assumptions and formalizing them into you know math that you can write down or symbols that you can write down and and then and apply operations to so you can apply the theory and everything and that's the language model is pretty good at that right like so if i say you know in the i talk about this this uh this kind of factual idea called the probability of necessity and saying like you know or you know if i were if i were working at netflix and i want to understand if uh you know i did some promotion and people uh watched this show and i want to understand if they would still watch the show had i not done the promotion right can i how would i formalize this in causal terms and um and then the language vinyl does a pretty good job in doing that for you um of course it still hallucinates too so like i also asked it's like okay well how would i estimate now that i formalize this in the math you know how do i estimate this quantity from data and it gave me an answer.
55:23But the answer, it wasn't, it wasn't, it wasn't wrong. It was wrong, but in a very subtle way, it was wrong. You know, such that it would have been right had it specified certain very strong assumptions, but it didn't mention those assumptions. And so like, it was like, wow, this would have gotten you in trouble as you actually try to apply this. And so, you know, we still have the same types of cautions that we have with language models when it comes to using the reason causally.
55:47Jon Krohn:I see. So the overall, the interesting twist here is that when When we're talking about generative AI and causal AI together, what you're describing, at least there, is where you're using the generative model to use its understanding of the world, which, you know, typically, you know, if we're using Stan, we're using bugs, those models don't have, there's no built in kind of understanding of the meaning of words or a feature name. And so when we're using these other kinds of approaches, you know, in PyTorch with do with Pyro, with Stan, with Bugs, with all of those kinds of libraries, we have some table of numbers, typically, that we are specifying some assumptions or we might be able to use a graph to describe relationships between those variables.
56:38Jon Krohn:But ultimately, we're working with just numbers. Whereas when in the kind of evaluations that you were doing, you were seeing, okay, how does a generative model like GPT-4.0 how does it use its understanding of the world to draw some causal inference? So that's kind of more like the kind of intuition that you were describing where, okay, you walk into a room, your dog is on the floor with, you know, peanut butter on his face, and the peanut butter jar has been knocked off the counter. What happened? It's like an example of the dog thing. If we were going to kind of throw a technical word. You might think of that as, say, root cause analysis.
57:22And there are algorithms for doing root cause analysis. Let's say you have a whole bunch of data about events that happened on a network and there was a data breach. You're trying to do some root cause analysis. Then there's you, the guy who's staring at all these logs, trying to figure out what's going on. Or you can take those logs and paste them into an LLM and say, tell me what happened. Or you can take those logs and paste them into an NLM and say, take this data, isolate the key variables, and apply this root cause analysis inference algorithm to it, or at least give me the code to implement it.
57:57And so in the same way, you can say, here's a problem that I'm thinking about. What's the right DAG to use here? Okay, given the right DAG, how would I write this up in do I, for example? In the same way, you can say, given this problem, how would I write this up in stand? Because it's seen a lot of stand code. Right. So that's one aspect. But I think, you know, to me, from a research standpoint, the more interesting aspect is to say, like, well, how can I use causality to make even cooler generative AI? Right. So, you know, so there's an example of that in a book, kind of a basic example where I'm saying, let's imagine that we can kind of take, right, have a separate, say, generative model for each node in the DAG that's conditional on its parents in the DAG.
58:44right? And then kind of connect this all together. And so you can still get, you know, by implementing it as a graph that reflects causality, you still get all the benefits of the theory, but you can also generate like you would from a generative model. But there are other really interesting ideas about where you could take this. Say, for example, with multimodal models, right? So if they say if we have a model that's incorporating natural language and maybe say video, for example. Well, there's some causal reasoning behind how you and I talk, right? There is some causal signal that's being extracted from natural language.
59:35And insofar as a video is a time series and causes precede effects in time, there's some causal signal kind of hiding in that video data as well. And when we combine these into a multimodal model, what kinds of new interesting causal things is that model learning? And can we extract it? Can we constrain it so that we can get certain guarantees? And one of the things that I'm working on is kind of looking at the space of generative AI for video games. and saying like, you know, to what extent can we get this generative AI to understand the underlying game mechanics or the underlying game physics, right?
1:00:18So rather than just generate really lifelike or really generate images of Minecraft that look a lot like real Minecraft, you know, can we train this model in a way that's going to learn something about the underlying mechanics of how Minecraft works. So maybe that it can generate things that honor those Minecraft mechanics but still create environments that we've never seen in the training data. Or in a way that makes this game kind of composable with other models. So let's say if I train a model on game A and I train a model on game B and I make it such that their mechanics are kind of are explicit in certain ways that they learn the right mechanics, kind of then combine them to kind of create a new game.
1:01:12Right. So those are that's more to me. Those are kind of more interesting applications, but they're they're they're still kind of fresh.
1:01:22Jon Krohn:Tons of fascinating possibilities with generative AI in this causalized in this causal AI space is so many other AI spaces. It's really fascinating to hear some of those examples. Let's jump now into some of the audience questions. We've got one here from Dr. Doug McLean. He's got a PhD in applied math and works as the lead data scientist at a food company in the UK called Gregg's, which is famous for its delicious pastries, though they have lots of other foods as well. But anyway, so Doug, technical guy, he's got a good technical question here. He says, could you make sure Robert comments on Judea Pearl's ladder of causation?
1:01:58Jon Krohn:So rung one is association, rung two is intervention, doing, and rung three is counterfactual. And so there's some terms there that you might, you mentioned counterfactual in some of your other responses, but you might need to dig into the definition of that a bit for our audience. And so the reason why he's asking you to comment on this ladder of causation, again, rung one is association, which I think is kind of like correlation. Rung two is intervention. So the being able to make some assumptions about causal effect. And then rung three is counterfactual. He says, to be blunt, I'm really stumped when it comes to causal modeling.
1:02:36Jon Krohn:It seems you need to know what to expect first before running any analysis. I guess he's basically saying here that this ladder of causation or the way that you set up causal models, there seems to be often to set up the model effectively, you already have to have some understanding of what's going on. And he finds that confusing. Yeah. And so to answer that last point directly is that, you know, I'd say that you want to model, I think oftentimes the way that people are trained to think about data, particularly again, in an industry where we unfortunately often have less control over how the data is collected, causality is kind of asking you to think more about the data generating process than the data.
1:03:27And so when you don't actually have control over that process, it can get a little bit frustrating because it's kind of out of your hands, right? But, you know, so typically we kind of get some data set and we're thinking about like, okay, well, what are the transforms I should apply to this? What should I, maybe I should discretize this. Maybe I should apply a long transform to that.
1:03:44Jon Krohn:And so we're thinking directly in manipulating the data to make it more amenable to the models that we want to use. And causality is saying, like, you know, data aside, you need to tell me what your assumptions are about the underlying data genning process. So now going to the causal hierarchy. I go into it very extensively in the book, I'd say, probably more so than any other book.
1:04:11So, again, at level one of association, we're just kind of in kind of plain old vanilla statistics and correlation land. And then level two intervention, this is what we're talking about, say, when we're trying to emulate an experiment when we're asking what if questions like, what if I were to take this vaccine? Would it prevent me from getting sick? And then level three is counterfactuals. And here we're asking questions where we're imagining what might have been different. So say, for example, I didn't get vaccinated and I got sick. Would I have gotten sick had I been vaccinated? So this is actually the same as kind of level two, which is say, what if I had taken the vaccine?
1:05:02What if I take the vaccine? Would I get sick? But now we're going to condition on two pieces of additional information. The fact that I didn't take the vaccine and the fact that I did get sick having not taken the vaccine. Right. So in level two, I'm saying, what if I take the vaccine? Will I get sick? And level three, I'm saying, what if I take the vaccine? Putting aside tense, putting aside present tense, past tense. Say, A, take the vaccine. B, get sick, yes or no. But now we're going to condition. And so that's the same as the first case. But now we're going to condition on two extra things.
1:05:43The facts that had you not, that's given not A, that not B happens. Right. And so the counterfactual and reason we call it counterfactual is because the the. Either the hypothetical condition taking the vaccine or the hypothetical outcome getting sick or not is in conflict with some data that we've actually observed, which just makes it a bit more challenging model. But at the same time, we're still just conditioning on some some some some extra evidence. And so in terms of what is required to answer those kinds of questions requires, very generally speaking, some additional causal assumptions.
1:06:37Some of those assumptions can be specified entirely in the form of a DAG. And some of them can't often.
1:06:44Jon Krohn:Frankly, the more interesting ones can't rely on DAG-based assumptions alone. We need to make additional assumptions about mechanism. But if, you know, him being a mathematician who understands things, like for example, you might make, you know, so not only does A cause B, which you represent with a graph, but maybe you make it the additional assumption that B is monotonic in A, right? That's for a change in A, there's a corresponding change in B, and that change is constant, or that change is going one direction no matter where you're at with A.
1:07:23And so it's just asking for more assumptions. It's not asking you to know the answer. It's just asking you to say, well, for certain questions, you can get away with lesser assumptions. And for certain questions, you need to make a few more. And so we can think about our assumptions as falling out on different levels of the hierarchy. And so you might sometimes for some interesting questions, you need to make level three counterfactual assumptions. Nice.
1:07:51Jon Krohn:Okay, that is super helpful. And hopefully Doug McLean finds that answer helpful as well as many of our listeners. We have time for one more question here. So this is from Adriana Salcedo. She is a flight attendant in Bavaria in Germany. But she is training to become a data scientist or an AI engineer. So she's been taking lots of courses online for over a year now. She's a regular listener and regular commenter. and yeah, the last time I talked about her on air I said maybe we'll get her on air at some point because it's such an interesting journey and so yeah, we haven't scheduled that yet but maybe someday, Adriana so Adriana had three questions I think you've actually answered two of them the first was around what types of problems does causal AI give you a clear advantage of over non-causal machine learning approaches and we've talked about that a lot in this episode already with examples of when you need to understand if one variable is just causing another, not correlated.
1:08:50Jon Krohn:And then she had questions around whether you need domain knowledge in a particular area to apply causal AI. I'm guessing that for the most part, the answer is yes. But she has a cool follow-up, which is, if it is essential, then could LLMs like GBT4.0 help us automate that part of the process or understand some of the domain knowledge that is required. And so that ties into, you answered that earlier. You said it can help you suggest what the directed acyclic graph might look like in a particular domain or help you put the code together in Stan or in PyTorch. So I think that that is kind of already answered.
1:09:33Jon Krohn:The one question that we haven't really talked about and I am really fascinated by is what does a typical causal AI workflow look like in practice? So when you set out to tackle some causal AI problem, how do you go about it? What's your workflow like? Just theoretically different workflows, but I'll talk about the main workflow that most people kind of get exposed to when they first get exposed to the space, which is to say you kind of start with a question that you want to answer, right? That does side question engagement have an effect on in-game purchases? Then you're going to specify your causal assumptions in some shape or form.
1:10:18In my book, I focus on doing so in the form of a graph, right? Or potentially a structural causal model, which is a graph plus extra assumptions about how variables relate to one another. But let's just stick to a graph for now. So, you know, you say like, okay, well, I think that guild membership causes side quest engagement. I think that side question engagement is a cause of in-game purchases. I also think guild membership is a cause of in-game purchases as well. The next thing I want to do is make sure that I'm going to pick my library, let's say do Y, and I'm going to specify that DAG, encode, say as a network X object and then pass it to some constructor in a model class in doY and it's going to give me a model object in Python.
1:11:08The next thing that I'm going to do is say, okay, well, I have this question and I'm going to now kind of look at my data. So I have some data on certain variables in the system and then I'm going to do something that's called identification. And so what's going to happen is say, okay, given this, identification means that given these assumptions here in the form of the DAG, given this data that covers certain variables that are in the DAG and has some, and the DAG has maybe some unobserved variables and some of them might be confounders, common causes between the cause and the effect of interest.
1:11:44Can I answer this question? And let's suppose the answer is no. Then we say like, okay, well, you know, can I get better data? Can I, for example, do something like, you know, can I, can I, yes, not even better data. Can I observe additional variables that would help me answer this question? Let's just be able to say like that. And so that could be, for example, getting an instrument, something that is a cause of side quest engagements, but is an indirect cause of in-game purchases and only media and side quest engagement mediates that relationship completely. Maybe then I could say, then in that case, I can kind of do something called an intramental variable analysis, or maybe there's some kind of intermediate variable between side quest engagement and outcomes like what we call a mediator, and we could do something called front door analysis.
1:12:40But you might say, okay, is there something else I can add to this data? I can collect this data that can help me get the answer. Great. Now you have that thing. Now you can answer the question. The next step is you're going to pick some kind of, given that you can answer the question, you want to pick some kind of statistical approach to do the actual estimation of the causal relationship. And those are going to have the same old statistical trade-offs you've seen elsewhere. So maybe you're using a linear regression type of model and you might get something, maybe you're kind of leaning really heavily on say the linearity assumptions or the standard or like constant error assumption.
1:13:19Or maybe you do something like propensity scores, maybe use something called double machine learning. And, you know, one of them gives you, one of them has really fat confidence intervals with the other one is a bit unstable, all kinds of stuff that you're thinking about. But these are all good old fashioned stats questions, right? Like you've already done all the causal heavy lifting. you know so so you've done that and then maybe you do some sensitivity analysis um at the end and do why we call it refutation just to see kind of how sensitive your results your conclusions are are to the assumptions and the uh that you're making at every step so like maybe you miss some variables in your dag maybe your um maybe your data is a bit small or maybe you're leaning too hard on linearity this or that and so you can kind of test all those things to see how robust your results are to violations of those assumptions.
1:14:07Jon Krohn:Excellent. That was super helpful. I wish I had thought to ask that question earlier in the episode because that provides so much context around what causal AI is and how you do it. Super helpful. And yeah, great question, Adriana. So thank you so much for taking all this time with us, Robert, today. You've been very generous with your time. We've gone over our scheduled recording slots. So thank you for that. But very quickly before I let you go, something that I forgot to prepare you for is that at the end of every episode, I ask our guests for a book recommendation. And so yeah, so this is something other than your own book.
1:14:44Jon Krohn:Do you have anything for us? Do you have a specific kind of domain in mind? Not necessarily. You know, if it happens to be, I mean, you have actually already mentioned some books. There's a Judea Pearl book that you mentioned earlier in the episode. on causality that could be a great one. But we also just sometimes people, you know, give us a novel that they've enjoyed recently or whatever. Yeah, so off the top of my head.
1:15:12Okay, so recently I've been reading. So yeah, if you're kind of new to causality and you just kind of want a light read, I think the Book of Y is a good call by Yudapurl.
1:15:25Jon Krohn:And so it's pronounced Udaya? I think it's called pronounced Yuda. Yuda. Yuda. Okay. Oh man, I've been butchering that for years, for a decade. Okay. But yeah, the book of wine. I mean, when I talk to him, I just go back to Burroughs. I don't, he's never correct. That's a good way to go. That's what I'll do from now on too. But I think, so recently I've been reading, I think it's called, you know, Surely You Just, Dr. Feynman or Mr. Feynman. But the book by Richard Feynman. Isn't it, it's Surely You're Joking, Mr. Feynman. Surely You're Joking, Mr. Feynman. is it Mr. It's it. Yeah. Even though he had a PhD, the book is Mr.
1:16:00Jon Krohn:Fineman. Right. I mean, Mr. Fineman. Yeah. I mean, to me, uh, you know, his ability to kind of, uh, boil complex concepts down into really clear explanations is something that I aspire to. Um, and so, um, I've been trying to reread that book and kind of, um, I mean, that book is not so much about explanations. which was kind of more about his way of thinking. But I say that would be my recommendation for folks. It's a very, it's a light read. It's not a technical book. Yeah, he's a lot of fun. I've been watching some lectures by Richard Feynman recently, and it's just great. Fantastic. Thank you so much, Robert, for taking all this time with us today.
1:16:45Jon Krohn:For people who want to follow you after today's episode or connect with you, what's the best place to do that? Good question. I occasionally lurk on LinkedIn, but I should be dialing back to social media recently. It's one of those times where social media can get a bit much, but LinkedIn is good. Yeah. And it's hard to be doing things like writing books and writing papers if you're too absorbed by social media. It's a tricky balance. So yeah, thanks again, Robert, and hopefully catch you again in the future. Thank you for an illuminating episode. Thanks for having me.
1:17:28Jon Krohn:Great thanks to Dr. Robert Osazu-Aness for coming on the Super Data Science Podcast and teaching us so much. In today's episode, he covered how well humans and animals intuitively understand causality, AI systems developed using correlation-based learning because traditional causal methods were rooted in classical statistics and didn't scale well to high-dimensional data. He talked about how libraries like Pyro, DoY, and PyMC now enable causal AI by handling statistical inference automatically, letting practitioners focus on specifying causal assumptions rather than implementation details. We talked about Perl's causal hierarchy.
1:18:03Jon Krohn:Level 1 is correlation or association. Level 2 is intervention, asking what-if questions. And Level 3 involves counterfactuals, asking what would have happened given the observed evidence. He also talked about how LLMs can propose causal graphs, generate analysis code, and synthesize causal knowledge from training data, though they still hallucinate and require careful validation. and finally he provided us with a typical causal problem a typical causal problem workflow where he starts with a causal question specifies assumptions often as a directed acyclic graph he checks if your data can answer the question chooses a statistical method for estimation and then performs sensitivity analysis to test robustness yeah so as always you can get all the show notes including the transcript for this episode the video recording any materials mentioned on the show, the URLs for Robert's social media profiles, as well as my own at superdatascience.com slash 909.
1:19:00Jon Krohn:Thanks, of course, to everyone on the Super Data Science podcast team, our podcast manager, Sonja Breivich, media editor, Mario Pombo, our partnerships team, which is Nathan Daly and Natalie Zajski, our researcher, Serge Massis, writer, Dr. Zara Karche, and yes, our great founder, Mr. Kirill Arimenko. Thanks to all of them for producing another deep episode for us today for enabling that super team to create this free podcast for you. We are deeply grateful to our sponsors. You can support the show by checking out our sponsors links, which are in the show notes. And if you'd ever like to sponsor the podcast yourself, you can see how to do that at john crone.com slash podcast.
1:19:37Jon Krohn:We've got that link in the show notes, of course. But yeah, lots of other ways to support us. We really appreciate it. You can share this episode with someone who might enjoy it review the episode on wherever you listen to it or watched it subscribe if you're not a subscriber but most importantly just keep on tuning in i'm so grateful to have you listening and i hope i can continue to make episodes you love for years and years to come until next time keep on rocking it out there and i'm looking forward to enjoying another round of the super data science podcast with you very soon
From the publisher
Researcher at Microsoft Robert Usazuwa Ness talks to Jon Krohn about how to achieve causality in AI with correlation-based learning, the right libraries, and handling statistical inference. When dealing with causal AI, Robert notes how important it is to keep aware of variables in the data that may mislead us and force inaccurate assumptions. Not all variables will be useful. It is essential, then, that any assumptions are grounded in a deeper understanding of how the data were gathered, and not what appears in the dataset. Listen to the episode to hear how you can apply causal AI to your projects.
Additional materials: www.superdatascience.com/907
This episode is brought to you by Trainium2, the latest AI chip from AWS and by the Dell AI Factory with NVIDIA.
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.




