In short
Practical AI Podcast Episode Notes: Causal Inference
Episode Overview
- Title: Causal Inference
- Hosts: Daniel Whitenack and Chris Benson
- Guest: Paul Hünermund, Assistant Professor at Copenhagen Business School
- Date: [Date not specified in transcript]
- Description: The episode discusses the significance of causal inference in AI, particularly amidst the increased focus on large language models (LLMs). Paul Hünermund introduces the concept, shares relevant trends, and provides practical methods for implementing causal inference in data science.
Key Concepts Causal Inference
- Causal inference seeks to determine cause-and-effect relationships rather than mere correlations.
- Commonly contrasted with standard machine learning, which primarily identifies correlations and patterns without establishing causality.
- Causal inference is essential for answering "why" questions within enterprise contexts, such as the impact of certain actions on business outcomes.
Differences Between Causal AI and Standard Machine Learning
- Causal AI:
- Focuses on cause-and-effect relationships.
- Requires background knowledge and domain expertise to distinguish between alternatives.
- Emphasizes experimental and observational methods.
- Standard Machine Learning:
- Predominantly correlation-based, using algorithms for prediction.
- Lacks interpretability regarding causal relationships.
Importance of Causal Inference in Business
- Causal inference is relevant for decision-making and forecasting outcomes of interventions within businesses.
- Examples include evaluating the effectiveness of marketing strategies or HR policies.
- Addresses fundamental AI challenges like fairness, robustness, and explainability.
Practical Implementation Approaches to Causal Inference
- Experimental Methods:
- A/B testing as a form of experimentation.
- Directly manipulates variables to observe effects (e.g., testing different marketing strategies).
- Observational Methods:
- Analyzes existing data without intervention.
- Techniques include:
- Regression Discontinuity Design
- Difference-in-Differences
- Nearest Neighbor Matching
- Directed Acyclic Graphs (DAGs)
Tools and Resources
- DoWhy: A Python library for causal inference that facilitates modeling, causal effect estimation, and testing assumptions.
- Causality-focused MOOCs and Literature:
- "The Book of Why" by Judea Pearl
- "Causal Inference: The Mixtape" by Scott Cunningham
Challenges Faced by Data Scientists
- Data scientists often feel a mismatch between the questions asked by stakeholders (often causal) and the tools they have (often correlational).
- There is a growing interest in understanding causal methods among data scientists, emphasizing the need for education, tools, and collaboration.
Key Takeaways
- Causal inference allows organizations to make informed decisions based on the anticipated effects of actions.
- Understanding the difference between causation and correlation is crucial for effective data analysis.
- Continuous learning and adaptation of new tools and methodologies are essential for data scientists to leverage causal inference effectively.
Future Trends
- Increasing integration of causal inference in AI research and practical applications.
- Growing interest in heterogeneous treatment effects and causal discovery from observational data.
- Emphasis on understanding the interactions between treatments and their implications for policy and business strategies.
Conclusion
- Causal inference plays a vital role in enhancing the interpretability and effectiveness of AI applications in various domains.
- As the field evolves, data practitioners are encouraged to adopt causal methods to improve decision-making and strategy formulation.
---
For further discussions, check out the conversation linked in the episode description or join community discussions on [Practical AI's changelog](https://changelog.zulipchat.com/#narrow/stream/456003-practicalai).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:06Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related technologies are changing the world, this is the show for you. Thank you to our partners at Fastly for shipping all of our pods super fast to wherever you listen. Check them out at Fastly.com. And to our friends at Fly, deploy your app servers and database close to your users. No ops required. Learn more at fly.io.
0:43Welcome to another episode of Practical AI. This is Daniel Whitenack. I'm a data scientist building a tool called PredictionGuard. And I'm joined as always by my co-host, Chris Benson, who is a tech strategist at Lockheed Martin. How are you doing, Chris? I am doing very well. I've been watching you building PredictionGuard from afar and looking forward to hearing more about it in the days ahead. It's a fun one. We'll talk about it in more detail soon and the causal reasons why I ended up doing those things that I'm doing. Could that be a transition? Yes. Speaking of cause and effect and causal things, we're really privileged today to have with us Paul Hunermund, who is an assistant professor at Copenhagen Business School.
1:34Welcome, Paul. Hi, Daniel. Hi, Chris. Thanks for having me.
2:07or what, like for a business person, like they're wanting to know what is the attribution? What is the behavior behind this thing that I'm seeing? So you're an expert in causal AI, causal machine learning and have been doing research in this area and are very well versed in it. And I'm wondering if you can, as you start out here, just like give a brief understanding to everyone about like, what do you mean when you say causal AI or causal machine learning? And how maybe is that differentiated from what people might commonly think of when they think of AI or machine learning? So, I mean, there are many names, causal AI, causal machine learning, causal inference, I think, is the more traditional term.
2:53But I think the basic idea is pretty intuitive for everyone who works with data, is that if we look at correlations and patterns in the data, sometimes they can produce quite surprising. and probably nonsense results. I mean, we all know the story about ice cream sales and shark attacks that are highly correlated over the course of the year, chocolate consumption and Nobel Prize winners in the country, probably driven by Switzerland predominantly. Storks in the babies, right? Like the stork population and fertility rates are correlated. With these examples, we usually use them in the classroom as sort of like a caveat, right?
3:37Wait a minute, a correlation is not causation. People have heard this term. But then causal inference and causal machine learning is really this idea of taking causality seriously and trying to build tools, algorithms that allow you to draw causal inference from data to distinguish cause and effect and weed out the kind of nonsense correlations. And yeah, that comes with a different tool set. So, I mean, you can approach this from a purely algorithmic point of view and you would probably apply different tools than standard machine learning. It goes even deeper. There's a whole epistemological point about it.
4:18Well, if you want to do causal inference, you cannot do this in a purely model three way. you actually need background knowledge, expert domain knowledge in order to do this, in order to, for example, distinguish between possible alternative explanations. And that is a whole paradigm shift in terms of how we approach data and how we approach machine learning. Well, and then my last word on this is what's the difference to standard machine learning is that, well, standard machine learning, the bulk of it is really correlation based. I mean, all the tools that we're having, deep learning, support vector machines, and so on, those are predictive tools.
4:59Prediction means correlation, finding, detecting patterns and data. And so they're not suitable for it. I mean, there's a branch of AI that is reinforcement learning. Maybe we could talk about this later that goes more in the direction of actually intervening yourself. The learner intervenes itself in the environment. So that goes in the right direction. but is not getting there full way. This is the main difference to standard machine learning. That kind of gets to the what, I guess, of what is causal AI, causal machine learning, causal inference. I guess the next question that's probably a good foundational one is the why.
5:40And maybe this is even exacerbated in recent times. I don't know if you've seen this with all of this sort of hype around large language models that are incredibly non-interpretable or produce things that are very factually incorrect or unexplainable. But how should a data scientist working in an enterprise, let's say, why should they care about causal inference rather than just making good predictions, let's say? Yeah, so maybe it helps if I've defined first what I mean exactly with causal inference because there's actually a neat definition by James Woodward, which is a philosopher of science, who said that causal inference or causal machine learning is a special kind of prediction problem.
6:26So in a sense, it is a prediction problem. But here we are predicting the likely impact of an action, intervention, or manipulation. Really this idea of I do something like I increase the chocolate consumption in a country. Will that produce more Nobel Prize winners in this context? I think if we approach from that perspective, you immediately see the value for business. Because in business, we are always asking these kind of questions. What if we do X? What if we implement this new HR policy? What if we enter a new market? Should we invest in this product or another product? So business always involves actions, interventions, and we want to forecast, predict the likely outcomes of this.
7:13That would be what causal inference people call the interventional level. Then one level above that is the counterfactual level. So counterfactual meaning we're reasoning about two states of the world. Had I not taken the aspirin this morning, would my headache be worse today still? These kind of questions. And they are also very relevant in sort of hindsight retrospective. retrospective. Was it, again, HR policy that we implemented that improved employee satisfaction and so forth? Immediately relevant in all sorts of domains in the business world. Specifically, in AI, we're talking about fundamental problems in AI, which is fairness, robustness, explainability.
8:02I believe that causal AI has to say something in all of these domains. There's also an immediate practical value in this. Let me ask you a question on, I want to throw another term that we haven't used yet that sometimes gets thrown in in casual causal conversations, say that 10 times fast, determinism versus non-determinism, because we have a habit of applying that at a high level with AI models and say, oh, they're non-deterministic. And there's a certain expectation over the years in training AI models that are non-deterministic in terms of you do have that disconnect in terms of understanding that causality from beginning to end.
8:43Can you kind of distinguish a little bit between the two terms in the sense of if someone is just kind of getting into this, they're early in data science, and they're trying to go, wait a minute, I thought AI was non-deterministic, but yet causality is explainable. How do those fit together? How do those as slices of perspective on an AI model? How do they work together where you have determinism, non-determinism, and causality or not? And what are the implications of those? When I put out the definition, I use the term likely impact. And that already hints at that also causal inference, the approaches that we have being our probabilistic framework.
9:19So there's no determinism in this. What we're interested in is still a probability or contrast rest of two probabilities, the probability of a certain outcome that I care about if I had taken the specific action or if I hadn't. So these are factual questions, but it's still, there's no determinism in the sense that if I implement something, it will always work or we will always have success with this product and so forth. There's an interesting, I think, sort of intellectual history because the frameworks that we're having causal inference directed acyclic graphs developed by Judea Pearl and these kind of people they built on actually earlier work in AI like Bayesian nets for example that were at that point still a purely predictive tool but a way or tool to deal with complexity in terms of probabilities because yeah expert systems were too rigid we figured that out in the 70s and 80s and building on that, then people immediately, once you reason probabilistically, people made the shortcut, the mental shortcut in reasoning in terms of cause and effect, probably because it's so intuitive to us, but the tools were actually not ready for it yet.
10:32So that was the intellectual history, how we move from probabilistic AI frameworks to causal inference. And again, I think people immediately started to think that way because causality is such a fundamental concept for human thinking right we learn it very early in our development babies can think causally there's some psychology work that we picked that up at the age of two or so pets sometimes can think causally probably so it's a very fundamental concept have you found like in practically interacting because i know you're involved sort of in the data science community as well and have helped run events and other things related to this topic.
11:13Have you found data scientists are sort of, because I could see some data scientists were so in the mindset of like, we're making a prediction, we probably understand that we're thinking about correlations in many cases, I think, but then it's sort of scary for us to think about like, well, I don't know if I want to like put out if I tell my executive, like, this is the reason that something happened, like, I can see the value of it, but how confident can I be in that? And that also gets to maybe people's, I think, people during COVID and other times realized maybe how rusty they were on, like, basic statistical and probabilistic type of concepts where, you know, everyone was all of a sudden thinking about medical trials and such.
12:04Have you found this sort of hesitation amongst data scientists as you've interacted with them? And what maybe is some steps that data scientists can take to gain confidence in initial thinking and education around this topic? Yeah, so we talked with a lot of data scientists from industry practitioners, and I don't think there's hesitation. It's actually the opposite. There's lots of interest. Of course, this is sort of a new topic. You need to tool up in a different area. So, I mean, that's a step that you need to take. But many people are very curious. And we just simply wanted to understand sort of where are we right now?
12:44And we had a hunch that this is the toolbox that we know, predictive analytics, correlational AI. And then based on that, what kind of questions practically do you address? Also, maybe in the interplay with the broader organization, right? What is it that the executives want to know? What do they approach you with? And is there sort of a mismatch between the methods that you're working with and the questions that are asked? And in our interviews, and we did some qualitative analysis on this too, we could clearly see this kind of mismatch between many questions that are asked. Do we actually have this causal component to it because actions, forecasting interventions is so ubiquitous.
13:28The standard tools, right, are not up to the task for this. And that creates actually this interest in approaching causal inference and looking beyond what we currently do. So one interview is stuck to my mind was an IT consultant that was working a lot in the data science field. And he said, like, yeah, most of the questions that our clients asked are causal questions in the end. But what we do in the end with them is always some form of predictive analytics, deep learning and so forth. That always created this kind of tension in the projects that he was working in. So that was very eye opening for us.
14:10It is now time for a changelog news break. The team at Suno AI is helping change the game in text-to-speech realism by releasing Bark, a transformer-based text-to-audio model that can generate highly realistic multilingual speech as well as other audio, including music, background noise, and simple sound effects. It can also laugh, sigh, cry, and make other non-word sounds that people make. Crazy, right? here's an example that includes sad and size meta tags. My friend's bakery burned down last night. Now his business is toast. And here's one more with laughter. I don't like PyTorch, Kubernetes, or Schnitzel.
15:00And xylophones flummox me. You can still hear some digital artifacts and blips here and there, but we're getting closer to synthesized audio that's indistinguishable from the real thing. And that's cool slash scary. You just heard one of our five top stories from Monday's Changelog News. Subscribe to the podcast to get all of the week's top stories and pop your email address in at changelog.com slash news to also receive our free companion email with even more developer news worth your attention. Once again, that's changelog.com slash news.
15:52Well, Paul, you've described really well how to think about generally this sort of causal inference, causal AI, causal machine learning, the importance of it. And you mentioned that in doing causal inference, you have sort of a different tool set or maybe different algorithms that are applied. I know that one thing that, of course, I've done before and know about from various data science positions is like experimentation or hypothesis testing, like A-B testing. I know that only scratches the surface you were talking about you know direct acyclic graphs and other things so could you give us like a broad sketch of like currently what are the like main categories of approaches within causal inference and how can we think about those like from really broad categorization traditionally people that divided the field into experimental and observational methods and experimental would be the A-B testing that you're talking about.
17:02One of our interviews even called it the big hammer that tech companies swing around A-B testing and it's applied a lot sometimes together with some form of multi-armed banded reinforcement learning type of approaches but often just in a plain vanilla way and that's great because well experiments are easy in many domains to set up, easy to understand, and you don't need a lot of background knowledge for it. You simply try out different things, such shades of a button on a website, classic example. But in other domains, it's really not that simple because, well, experiments can be very costly. They can be unethical in many questions.
17:44I think you mentioned the COVID pandemic earlier. That was an interesting example to observe because when we tested the vaccines, of course, we did the standard clinical trials, which is an experimental method and maybe testing if you want, right? That costs a lot of money, but we have these procedures for it and we need to approve drugs in that way. But then after we rolled out the vaccines, immediately there were follow-up questions like, for example, where is the vaccine more effective? Is it for older population or a younger population? Or in which way do we need to roll out scarce vaccines and so forth?
18:22These kind of questions they were not included in the control trial we didn't have experimental evidence for it we needed to answer this based on ex post data so people picking up vaccines and then seeing where they're most effective that was interesting to see because many of the questions that we ask in practice do involve this observational causal inference and with observational causal inference I mean we don't actively intervene ourselves but we passively observe the data and still want to get cause and effect out of it, although we haven't designed the experiment also. So in a sense, we're then trying to mimic a thought experiment, if you want, with observational data.
19:03And that creates all sorts of problems because, well, those people that picked up the vaccine earlier probably are those that thought they have the most to gain from it, for example. So there's this sort of self-selection bias or confounding bias in this, and we need to address all of these things. These are the two main categories and then within those categories we have all sorts of different techniques, algorithms like experimental design is an entire course catalog at our university in the observational fields. So for example, I originally come from an econometrics background and in econometrics or in economics we ask a lot of calls of questions and then we have tools like regression discontinuity design difference in differences, nearest neighbor matching, and so forth.
19:51The new kid on the block are the computer scientists, and they catch up fast in causal inference, and they develop these techniques like directed acyclic graphs, causal reinforcement learning, so all sorts of exciting streams of literature coming up these days. I'm really trying to absorb what you're saying, and it's very interesting. I'm kind of wondering if I have a problem today, like before our conversation and I want to go through your kind of the typical data prep and model training and model testing to deploy. But now I've listened to you and I want to start implementing causal approaches into my workflow.
20:30How does my workflow change? What does it look like with typical tools now? And where might gaps be in the typical tool chain that we currently have? How do we make it practical and go do it after the show? Starting from the epistemological challenge is that we cannot do causal inference in a purely data-driven way. We cannot just optimize a target function and or look at our confusion matrix loss function in that sense. We need to complement this with background knowledge that, well, in the simple examples, right, it's not just ice cream sales and shark attacks. There's a third variable lurking, right, which we need to consider, which is weather probably or sunshine.
21:12So this is in the simple case, but now imagine a problem that you approach for the first time, you do exploratory research, so you don't have this good theory, so we need to do something about this. Then a lot of the standard challenge that we are having, collecting good data, maybe designing an A-B test are the same, but this additional step of bringing in background knowledge, and there it depends, I guess, right like how also the data science team is structured in an organization do we need to bring in outside stakeholders do we maybe need to talk with the marketing people or the logistics people depending on the project often at the moment data science teams are almost this kind of in-house consulting type right and there are for example not that many mixed teams that could bring in this background expert domain knowledge practically speaking i mean there are all sorts of tools out there in the standard software languages.
22:10It's a little bit scattered, probably, the landscape, so you really need to know what kind of libraries are out there. For example, in Python, the do-y package by Microsoft really became a standard industry-wide because they also have this kind of causal inference pipeline implemented in the package which starts from modeling a specific domain or phenomenon to applying the causal inference algorithms, getting causal effects out, and then also refuting the model or challenging the model. So you have this kind of step-by-step procedure that can really help you in getting started and getting results quickly.
22:49One of the things that you said that I wanted to ask kind of a clarifying question was, you kind of talked about going to that kind of external source, the extra authority, if you will. A lot of practitioners, you know, these days are kind of starting on their own if they're not on a big data science team. Like, you know, I work in a big company and we have tons of data scientists. So this probably doesn't apply to me in that capacity. But a lot of people in startups are out there trying to kind of delve into new businesses and stuff. And they may not have access to kind of outside the data expertise to apply.
23:23Do you have any tips or guidance on like if you're that practitioner and you're trying to solve a problem for which you don't have that external expertise? How would you go about tackling that? How would you go about saying it's me, myself and I, you know, joking around and I'm going to this is a way I can apply causal approaches when I don't have a lot of resources available to me? First of all, I would say you're never really alone. So I think outside of the box a little bit. I mean, often it doesn't take that much. You can just approach people and maybe talk with them for an hour and get the insights out that you need.
24:00Consulting probably the scientific literature on a certain topic can help too in sort of figuring out alternative explanations that you can then bring to the data and test to the data. We're also not completely helpless in the sense that everything has to come from the theory. There are data-driven approaches, so that will be the area of causal discovery. that we can apply to get closer to kind of a causal model based on the relationships that we find in the data. We know that that never gets us 100 % all the way. So we will always need to complement it with some form of background knowledge, but it can already help.
24:37And then I would say, I mean, talking also to practitioners, I think sort of the 80-20 rule, I think it's called, right, applies. I mean, already getting closer to something causal is often good enough, and we should get away from this idea that it's a zero one, right? That either it's causal or it's not. Often we get closer to the truth. And if not, we can do, for example, there are whole tools on sensitivity analysis that we can challenge our assumptions and see how robust they are. And I think in practice, this already helps tremendously. You mentioned kind of reaching out to practitioners and the community around this.
25:14Could you describe a little bit? I know, like I mentioned earlier, there are some resources that you kind of helped co-found and run over time related to this. Could you mention those so that people could find those as they're looking into the topic? Based on what we identified in where the field is, we actually saw the need for more exchange between different academic fields because causal inference is such a general purpose technology almost. It's applied in various different fields. I mentioned economics, computer science, epidemiology health sciences but then also practitioners so it's really like a mixed group so we set up the annual causal data science meeting we started in 2020 so had to do it online because of the covid pandemic and then realized that is really an easy way to get people into one well virtual room in this case and there was lots of interest from practitioners and we're going to have the third iteration of this this year in November, so still some time.
26:15But hopefully listeners will make a mental note. Well, there are also good teaching tutorials out there, many blog posts, online courses that you can sign up to books like the Book of Why by Judea Pearl. Maybe not really a textbook, but really drives home the idea of why causal inference is so important and has really nice historical anecdotes because Judah is really a giant in this field. Causal Inference, the mixtape by Scott Cunningham. If you have more of an econ background, perhaps The Effect by Nick Huntington-Klein. So these are all beginner-friendly textbooks that you can pick up. And then trying out the different packages like DoWy in Python, for example.
26:58There's the startup called Geminus that is developing causal inference software. and they're having free trial versions where you can start out drawing your directed acyclic graphs and see how answers change if you change assumptions, for example. So I think that is usually the best way to learn and pick this up.
27:34Well, Paul, I'm selfishly going to present you maybe with a scenario and do some sort of on the fly problem solving. I figure you're you're probably good at that being a professor and always solving problems with students and others and colleagues. So I mentioned my wife runs a business. It's a candle manufacturing business. And there's actually this sort of like why question that we've been talking about a little bit. So last year, to give context, I just logged into Shopify. So last year, they had 87 ,837 orders. Each of those orders, or at least most of them, when they ship them included a free sample two ounce candle.
28:20It's like a freebie add on. and over time the assumption has always been oh people really like that and it's sort of like part of the package that they get it increases reorder value right like they see the package and they're like oh cool I got like a free candle and now I'm gonna I love these people forever and I'm gonna reorder right well the question has come up obviously at this scale of orders that's a lot of free two ounces. And even just the savings of those would be huge. So how might you as a practitioner or someone thinking about this problem, both in terms of the experimental or the observational approaches, what might be some ways to dig into this?
29:06Obviously, it's very expensive to send out like get different packaging and do like a experiment at that scale. So it'd be nice to know without doing like a large scale shift in packaging and that sort of thing. Any tips for me? In this case, the big advantage is that you or your wife, you are actually controlling this process, right? So you decide in which packaging to put the free add-on. And in that sense, you immediately understand the selection process or the treatment assignment, how we would call that here. And in that situation, I think an experimental approach would be the way to go. Then it becomes more of a statistical question like how much your sample or how large your sample needs to be in order to draw robust conclusions.
Read the full transcript
29:58And if it's just the yes-no questions about a free sample or not, the experiment can probably be quite small. That relates a little bit to the COVID example that I discussed earlier. probably you want to broaden that up and think about for example heterogeneous treatment effects in the sense that well is it high volume customers that like this free add-on most or is it more like the casual shoppers right and suddenly you have four groups that you're catering to because well high volume low volume and treatment and control these are problems that always come up in causal inference so the tools like causal random forest developed exactly for that problem how do you efficiently partition the population in order to reduce costs associated with an experiment right you want to be as cost efficient as possible but also still get robust conclusions out a similar problem arises then of i mentioned earlier robustness of findings right so transfer learning is a big topic in AI.
31:05So maybe let's assume you've done this experiment at this point in time and you found that robust treatment effect, right? Like so people react to it, it increases reorder value. The question is then in six months from now, will the world have changed or will these results still be valid, right? And maybe you're thinking about that in six months, not so much has changed of the business, right? But for example, platforms like booking.com, right? Hotel bookings for them is very relevant because you have people that book hotels in the summer for leisure travel are very different than business travel, for example.
31:45So you have this kind of problem. Can you transfer causal knowledge that you've obtained to a different domain? Because that would save you on experimentation costs, for example. So yeah, all sorts of interesting questions in that domain. but the big advantage is you can actually run experiments here. Like in other domains, we are reliant on data where people self-select into something like, for example, the standard question in economics about, for example, the returns to a college education. There we could not randomly assign people to colleges or whether they can go to college or not. There we have to rely on them self-selecting into a sort of treatment and control group And then it's always the question whether we have really an apples to apples comparison or is it perhaps an apples to oranges.
32:35I just wanted to say I think you're lucky that you got Daniel's example question because coming from my industry, I would have had to ask, I don't know, hypersonic missile design or something and I don't think we want to go there. This is a great thing about the podcast, right? We get to have the expert on and I get to selfishly ask the question that helps me in my day to day. So excellent way of getting some free consulting in there. Yeah. So wanted to actually take you back to something that you mentioned a little while ago. We were kind of talking about the benefits of causal inference and you brought up reinforcement learning, but we were generally talking about kind of fairness, bias, robustness, the impact of causal on those.
33:18Could you kind of go back to that point and kind of talk a little bit about what that means? These are huge topics that are in all of the different branches of AI right now, and it's on everyone's mind, especially with all the advances this year. How does causal affect that worldview of doing these amazing things in these different branches of AI, but doing it without bias, doing it fairly, such as that? I'll start with fairness, because that's actually the very first example that I use in my own course, Causality, Causal Inference course here at Copenhagen Business School. It's a case taken from Google, actually.
33:54So a while ago, I think in 2019, well, already earlier, the story goes longer, but they have been accused of underpaying women in their organization, right? So there we have a classic example of like a protected attribute like gender, race, and so forth. And we want to prevent bias in some form of automated or semi-automated decision making, right? And that comes up all the time. I mean, in loan acceptance models, for example, we want to remove bias and so forth. So to make the story quick is they have been accused of underpaying women in their organization. Then they did a fairly sophisticated analysis, published a white paper.
34:33And the result was of that analysis that they found that they actually underpaying men. At least they thought so. And not only men, but actually high level software engineers. So high seniority software engineers at Google. and then because they're committed to fairness in their organization they actually raised salary levels for these high level software engineers based on analysis so it also had a practical component to it or like a policy implication we cannot analyze this case here in detail but if you do that analysis is very likely that they actually did a sort of fairly common causal inference mistakes so they conditioned on some variables that are downstream of so that are affected by gender like occupation for example right and then if you have discrimination already at that stage that for example women don't have it so easy to get into high level positions for various reasons that we know of then that will be a classic mistake and you can produce these kind of again nonsensical correlations in the end like the sharks and the ice cream that's one example that can actually easily transport to other kinds of questions like I mentioned algorithmic bias.
35:44And that's a causal question because if you don't understand how variables in your model causally interact and relate to each other, you cannot answer this question. You cannot decide how to correctly analyze the data. Robustness I mentioned, so the transportability transfer learning kind of aspect of experimental knowledge and their causal inference techniques have been developed. also dealing with selection bias in data. So a data set that might not be a representative sample of the population that you care about, but is measured with some form of selection bias because only happy customers answer your consumer survey or unhappy customers, but no one in between, right?
36:26These questions. And then lastly, explainability. I think explainability almost comes for free with causal inference. I mean, don't get me wrong. Causal inference is a hard task, But once you solve it, explainability almost comes for free because, well, I mentioned the book of wine, right? So causal questions are always related to why questions, counterfactual as well, right? Like, why did my headache go away? Was it because I took the aspirin this morning? I mentioned this example. This is the way we reason. This is the way we explain, for example, things to other humans. And so there's an immediate connection to explainability.
37:04too. That's a really great way to think about this. And it gets me thinking like, what will be the impacts of these two fields as they interact more over the coming years? And I'm wondering, from your perspective, because you're so plugged into the research that's going on in this area, but also the practical side of this and how data scientists are beginning to use these techniques, what, as you look forward to the next, let's say, year or whatever time period you want to have there like what gets you excited or what trends would you like to highlight that maybe people should be thinking about in this field or maybe it's just things that you're excited about in terms of new new opportunities or new methods or whatever it might be yeah just uh attended a conference last week in tubing in germany the causal learning and reasoning conference and it was just exciting to see how many young minds there were attending the conference.
38:04So it was a very young audience, a lot of grad students and computer science specifically there and a bunch of seniors as well. But that really showed me that this seems to be the next big thing in AI and people have confirmed that to us, that there's more and more interest in the academic side, But we also see that in practice in the industry. Yeah, so I'm excited about, well, experimental design, what I mentioned earlier, right? Heterogeneous treatment effects, for example, not only being satisfied with having one average treatment effect or average causal effect, right? One number, and this is what we're expecting.
38:45but actually making this more fine-grained and opening up, answering questions like, is it all people, young people that benefit more from vaccines? In which way do we need to roll that out in the most effective way? Then on the observational side, I think causal discovery is really promising. And this is really the idea of how far can we go with just simply trying to get out causality from observational data? And we will never get 100%. I mentioned that, but how far can we go? So one big challenge in that area is, for example, to have good benchmarking data sets. In machine learning, that's usually easy.
39:22You divide a sample up into a training and benchmarking data set. With causality, that's not so easy. You often need an experimental benchmark, for example. A lot of work has been done in genomic research where you can knock off genes, for example, in an experimental way. So that is really exciting. There's new work on causal root cause analysis by Amazon, for example. So figuring out what are the causes of actually outliers, even in an engineering system. So you mentioned, Chris, that you're working in that area. So I've seen, for example, companies in the defense industry thinking about this problem of root cause analysis.
40:00Lastly, perhaps, because originally, before I came to causal inference, I was actually trained in economics, like I mentioned, but specifically innovation economics. So this idea of how do we produce knowledge as a society? How does knowledge spread across society? And so in causal inference, there's new lines of work thinking about actually interactions between treatments. So not just the idea that I take a pill and I get an outcome from that, but it's like you take a pill and that reduces the viral load in our community. And that's why I actually have a lower likelihood to get sick, for example.
40:36So these kind of interactions between people are really important, I think, in many domains and specifically also in the way knowledge spreads across networks. So that is something I'm really exciting about. Awesome. Well, I am really, really happy that we got to have this conversation on the podcast because I think it highlights something that's a real complement to many of the things that people are exploring around deep learning and large language models and other things. This is a really important piece of the sort of practical side of what data scientists are doing in the enterprise. So, yeah, thank you so much.
41:14And thank you for your research on the topic and also engaging the community around this. It's really great and really happy to have had you on the podcast. Thank you. I really enjoyed the conversation.
41:33Thank you for listening to Practical AI. Your next step is to subscribe now, if you haven't already. And if you're a longtime listener of the show, help us reach more people by sharing Practical AI with your friends and colleagues. Thanks once again to Fastly and Fly for partnering with us to bring you all Change Talk podcasts. check out what they're up to at fastly.com and fly.io and to our beat freaking residents breakmaster cylinder for continuously cranking out the best beats in the biz that's all for now we'll talk to you again next time
From the publisher
With all the LLM hype, it’s worth remembering that enterprise stakeholders want answers to “why” questions. Enter causal inference. Paul Hünermund has been doing research and writing on this topic for some time and joins us to introduce the topic. He also shares some relevant trends and some tips for getting started with methods including double machine learning, experimentation, difference-in-difference, and more.
Changelog++ members save 3 minutes on this episode because they made the ads disappear. Join today!
Sponsors:
- Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
- Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs.
- Changelog News – A podcast+newsletter combo that’s brief, entertaining & always on-point. Subscribe today.
Featuring:
- Paul Hünermund – Website, LinkedIn, X
- Chris Benson – Website, GitHub, LinkedIn, X
- Daniel Whitenack – Website, GitHub, X
Show Notes:
- How Can Causal Machine Learning Improve Business Decisions?
- Causal Inference is More than Fitting the Data Well
- Causal Data Science in Practice
- Causal Discovery
- DoWhy Github
- The Book of Why
- Causal Data Science Meeting
- Paul’s study on causal ML adoption in industry (incl. an overview of useful software packages in Table 3)
- Causal Data Science MOOC on Udemy
Something missing or broken? PRs welcome!




