Position: Probabilistic Modelling is Sufficient for Causal Inference

3 Jan 2026 · 12 min · 4 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Argues that causal inference (interventions and counterfactuals) can be done with standard probabilistic modeling by writing a full joint probability over “observed,” “intervened,” and “counterfactual” worlds, making Pearl-style causal operators unnecessary except as shorthand.

Guests/backgrounds

No explicit podcast guests; the episode discusses researchers Judea Pearl (causality requires leaving statistical vocabulary), Andrew Gelman (causal framework overcomplicates), and David McKay (mandate: write down probability of everything).

Key claims

Confounding makes naive conditional expectations misleading (Simpson’s paradox). A “twin model” approach breaks links (e.g., make T independent of Z) to model interventions. Counterfactuals require deciding what to hold fixed (e.g., Z and possibly stable traits/noise).

Notable examples

Aspirin dosage vs headache severity/duration; sugar-pill scenario where observed data suggests harm despite zero causal effect; casino/dynamic dice example with variable-dependent structure.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Causality Debate

0:46 to 4:30

Exploring the conflict between traditional probability theory and causal frameworks in understanding causality.

“He thinks you have to enrich your language with a dull operator.”

Simpson's Paradox and Confounding

4:31 to 6:42

Discussing how confounding can lead to misleading conclusions in causal inference using the aspirin example.

“This is key because it means the data from the messy world can teach us about the clean one.”

Probabilistic Framework and Twin Models

6:43 to 10:48

Introducing the twin model approach to solve causal questions by modeling two worlds simultaneously.

“In fact, this whole area really highlights why the joint model is clearer.”

Conclusion on Causality and Probability

10:49 to 12:26

Summarizing the benefits of using a probabilistic approach to understand complex causal relationships.

“They're just specialized dialects of this much bigger, more powerful language.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. We're jumping straight into what feels like one of the most fundamental fights happening right now in statistics and machine learning. We're talking about the war over causality. That's right. I mean, if you're building models, if you're trying to figure out what's going to happen next, you hit this wall. And the question is always the same. Do you really need these special dedicated tools, you know, like Judea Pearl's famous Dota operator, to figure out cause and effect? Or is good old probability theory, the stuff we all know and use, actually powerful enough to handle it all?

0:34Because this isn't just some small academic squabble. You have on one side Judea Pearl himself saying, and this is a quote, there is no way to answer causal questions without snapping out of statistical vocabulary. He's very clear. He thinks you have to enrich your language with a dull operator. But then you have someone like Andrew Gelman, a giant in the field, who basically says the whole causal framework is just complicating things unnecessarily. It's a total standoff. And so our mission in this deep dive is to try and cut through that confusion. We think we can show how a single probabilistic approach can handle it all.

1:11Interventions, even counterfactuals. And it all comes down to this one really elegant rule. It's from David McKay, who said the one overarching mandate is just. Always write down the probability of everything. And when we say everything, we mean everything. You have to write down the full master joint probability distribution over all the worlds you care about. The one you see and the one you're imagining. That joint probability, that's the whole machine. That's the whole machine. Okay, so let's get specific. Why does almost everyone think standard probability just isn't enough? Well, it's because the simple, the most intuitive statistical approach fails.

1:48And not just fails little, it feels spectacular. Right, especially when you ask an interventional question, something like, what happens if I go in and actively change something? Exactly. Let's make this concrete with an example we can use throughout. Let's say we're studying aspirin and headaches. We've got three variables. Z is the initial headache severity. T is the aspirin dosage someone takes. And Y is the final headache duration. So if I were being totally naive, I'd just look at the data and see what? The average headache duration for people who took a certain dose. That's it. You calculate the conditional expectation, EYT.

2:22Yeah. You're asking for the group of people who took dose T, what was their average outcome? And that is a trap. A huge trap. It's totally misleading. And the reason is one word, confounding. Let's spell that out. Okay. So in the real world, who takes a high dose of aspirin? The people with the worst headaches. Exactly. People with a high initial severity, Z, are the ones who choose a high dose, T. And of course, people with a high initial severity Z are also likely to have a longer headache duration Y, no matter what they do. So the severity is driving both the dosage choice and the outcome. It's the hidden culprit.

2:57And this leads to something called Simpson's paradox. We can even imagine a scenario where the aspirin is just a sugar pill. It has zero effect. The parameter for efficacy, let's call it Degelo, is zero. So it does literally nothing. Nothing at all. But even in that world, if you just look at the data at EYT, you will see that people who took a higher dose of aspirin actually had longer headaches. It's just, it's wild. So the data tells you aspirin makes headaches worse. When in reality, it does nothing. The initial severity is just completely skewing the correlation. And that's it. That's the whole argument for why you need something more than just looking at associations in your data.

3:36Precisely. That failure is the cornerstone of the argument that statistics is not enough for causality. Okay, so here is where it gets really, really interesting. Because the probabilistic framework doesn't invent new math to solve this. It just gets clever. It solves it by carefully modeling the system. Not as one world, but as two worlds at the same time. We call it this twin model approach. Like a parallel universe. Exactly. In universe one, you have the observed world. People are acting naturally, choosing their aspirin dose based on how bad their headache is. Isn't it that's the messy, confounded data we just talked about?

4:11Right. Then in universe two, you have the hypothetical intervened world. We'll call the variables there T star, Y star, Z star. This is our clean, controlled experiment. And the magic is in how you connect them. It is. First, you assume the basic physics are the same in both worlds. The true efficacy of the aspirin, the parameter theta, is shared. This is key because it means the data from the messy world can teach us about the clean one. Okay, that makes sense. Second, the rule for how a headache works is the same. The way Y star depends on the dose T star and severity Z star is the exact same physical mechanism as in the first world.

4:49It's invariant. So what's the one thing that changes? The only thing that changes is how the dose gets assigned. In the hypothetical world, we break the link. We say that dose T star is assigned randomly, maybe by a doctor. It's now completely independent of the severity Z star. You've broken the confounding. You've snapped the connection. So if I'm understanding this, the whole famous dodal operator, is that just a shorthand for saying let's define a new world where T is independent of Z? That is exactly what it is. And the probabilistic approach does that explicitly. We write down the master joint probability over both worlds, and when you do that, the confounding just vanishes.

5:30The math correctly isolates the true efficacy parameter,$2. So you solve Simpson's paradox. You solve it with nothing but careful, explicit probability theory. Fair. Okay, so we have to address the elephant in the room. if the joint distribution is all you need. Why did brilliant people invent a whole new language with Doodle annotation and the Doodle calculus? Right. That's the practical question because writing down the joint probability for every single variable sounds, I mean, it sounds like a nightmare. Aren't these tools a massive shortcut? They absolutely are. And we're not saying they are useful.

6:04What we're saying is that we should think of them as incredibly powerful, convenient, well, syntactic sugar. Like a macro in a programming language. Well, perfect analogy. The Deudol operation is just a very clean, simple way to say you're modifying the joint distribution in that specific way by breaking the links to a variable. And the Deudol calculus, that seems like heavy math. It is, and it's brilliant. It gives you these algebraic shortcuts to figure out things like which variables you need to control for to remove confounding, your adjustment sets. But you're saying anything it can do, you could in principle derive from the base probabilities.

6:39In principle, yes. It's just a set of proven theorems built on top of probability theory plus that definition of an intervention. Okay. In fact, this whole area really highlights why the joint model is clearer. There's this issue called non-identifiability. Meaning the data alone can't tell you the truth. Exactly. You can have two different causal theories, two different graphs, that produce the exact same correlations in the data you can see. The data can't tell you which theory is right. So what do you do? Well, the probabilistic model forces your hand. To write down the full joint distribution, you have to be explicit.

7:15You have to commit to one of those theories about how Z, T, and Y are connected. Yeah. You put your assumptions right there on the table. There's no ambiguity. Okay, we've handled interventions which are about populations, like a clinical trial. But what about the really hard one? The deepest level of causality? The counterfactual? The what-if question for one specific person. Right. My friend calls me. She says she took dose T and her headache lasted all hours. I know her actual outcome. And I wonder, what if she had taken dose T star? Would her headache A star have been shorter? This is harder because you're not averaging over a population.

7:53You are modeling her. So, another twin model. Another twin model, but with a critical difference. We still model her observed world, T-Y, and her hypothetical world, T star, Y star. But now some of her personal unobserved traits have to be shared between those two worlds. Her initial headache severity, Z. Exactly. Z has to be the same in both scenarios because it's her headache. So what we do is we take the facts, she took dose T and got outcome Y, and we use them to update our belief about her specific Z. Then we use that inferred Z to predict her hypothetical outcome, Y star. And this is where the probabilistic model is really flexible, right?

8:30Because you have to decide what else to share between the two worlds. This is the crucial part. What about the random noise, that little epsilon variable, epsilon? The fudge factor. The fudge factor, yeah. Should that be shared too? It depends on what you think it represents. If that epsilon represents her stable personal traits, her age, her pain tolerance, then yes, of course you should share it. It's part of her. But what if it's not? What if it represents pure randomness? Like a car alarm went off outside her window at a critical moment, or there was a tiny variation in that specific pill she took.

9:06Then it makes no sense to assume that same random event would have happened if she'd taken a different pill at a different time. It makes no sense at all. And the issue is that many classical approaches, like some structural causal models, kind of force you to share all the randomness by default. They can be very rigid. The probabilistic framework gives you that fine-grained control. You, the modeler, decide precisely which parts of reality stay fixed when you jump to the counterfactual world. And you do it all within the language you already know, the language of joint probability. Which brings us back to the original question.

9:40If this probabilistic way of thinking is so powerful and flexible, why is this debate still raging? I think it's got to be about semantics. It's what the words mean. Critics often say statistics is insufficient, but it seems like they're using a really narrow definition of statistics. That's it, exactly. They're often talking about methods that just look at simple correlations in a static data set, you know, statistics from decades ago. But that's not what the field is anymore. Not at all. I mean, that definition is just historically inaccurate. Modern statistics and ML are all about non-static situations.

10:15Think about domain shift or off-policy reinforcement learning. Right. These are fields entirely about what to do when the distribution of your data changes. Which is the exact same problem causality is trying to solve. So it's like saying a screwdriver can't hammer a nail so the entire toolbox is useless. Perfect. When really the toolbox has this amazing power tool that can handle it just fine. And for you, for the learner, this is great news. The probabilistic approach is more general and it lowers the barrier to entry. You don't have to learn a whole new separate language. You just extend the language of probability, which you probably already know, to cover these new hypothetical settings.

10:52The causal frameworks aren't wrong. They're just specialized dialects of this much bigger, more powerful language. So let's wrap this up. We've gone from interventions all the way to counterfactuals, and we've managed to do it all using the core tools of probability. And the big takeaway really is that expanded McKay rule. Right. It's not just write down the probability of everything. It's always write down the joint probability over all the settings you're interested in. The observed world, the intervened world, the counterfactual world, all of them. And when you see it that way, these powerful causal tools like the doula operator become what they are.

11:30Not new mathematics, but brilliant, efficient shorthands. They're shortcuts. And this whole viewpoint just opens the door to modeling systems that are way too complex for traditional static graphs. Give me an example. Imagine trying to model a casino game where, I don't know, the number of dice you roll depends on a previous card you drew. The number of variables itself is changing. Trying to draw that as a fixed graph would be a mess. A complete mess. But if you can write down the joint distribution, the probability of drawing that card, and the probability of all the outcomes from the resulting dice rolls, you can still ask causal questions about it.

12:06So the final thought for you to take away is this. What complex dynamic system are you looking at in your own work? Maybe it's user behavior that changes every quarter, or a supply chain that reacts to real-time events. Could you start to unlock it not by looking for a special tool, but just by trying to write down that one single overarching joint probability?

From the publisher

This paper argues that probabilistic modelling is sufficient for causal inference, challenging the belief that specialized causal notations like the "do-operator" are strictly necessary. By advocating for a "write down the probability of everything" approach, the authors demonstrate that interventional and counterfactual questions can be solved using standard **Bayesian Networks** and joint distributions. They reinterpret traditional causal tools, such as **Structural Causal Models**, as useful syntactic shorthands rather than distinct mathematical requirements. The text suggests that the perceived gap between statistics and causality stems from a **semantic confusion** that unnecessarily narrows the definition of statistical inference. Ultimately, the authors promote a **unified framework** where causal reasoning is treated as a flexible application of existing probabilistic principles.

More from Best AI papers explained

All 475 episodes
Position: Probabilistic Modelling is Sufficient for Causal InferenceBest AI papers explained · 12 min
Listen in VO