DoubleGen - Debiased Generative Modeling of Counterfactuals

27 Sep 2025 · 13 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

DoubleGen, a framework for generating unbiased counterfactuals (“what if” outcomes) from biased observational data by debiasing confounding and avoiding misspecification.

Guest backgrounds

No guest names or bios are provided in the transcript.

Key claims

Naive generative models learn correlations, not causal effects. Existing counterfactual methods like IPW and plug-in are singly robust (fail if their one auxiliary model is wrong). DoubleGen is “doubly robust,” using two auxiliary models so validity holds if either the propensity model or outcome model is correctly specified.

Notable examples

CelebA faces—smiling correlates with lipstick (56% vs 38%), makeup (47% vs 30%), and female label (65% vs 52%). DoubleGen maintains low kernel arcface distance (KD ~0.68) under misspecified models, unlike plug-in (KD 2.17). Amazon reviews—rare synthetic intervention (~3.5%); DoubleGen beats IPW (Wasserstein error 0.12 vs 0.17) and scores higher on MAUVE (~0.82–0.83 vs 0.72).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Counterfactuals

0:45 to 2:10

Exploring the necessity of generating unbiased counterfactuals from observational data.

“It's designed specifically to try and generate these, well, unbiased counterfactuals using the often messy observational data we actually have.”

Confounding Bias Explained

2:10 to 4:20

Discussion on confounding bias and its implications in AI training models.

“The Celebi data set the celebrity faces seems like a good way to see this bias visually.”

DoubleGen Overview

4:20 to 6:20

Introducing DoubleGen, a framework designed to generate unbiased counterfactuals using dual auxiliary models.

“So how does DoubleGen avoid that fragility?”

DoubleGen's Robustness

6:20 to 8:10

How DoubleGen's dual model approach mitigates risks associated with model misspecification.

“Did they actually show this works, especially when things go wrong?”

Application in Images and Text

8:10 to 10:20

Examining how DoubleGen performs in experiments and its advantages over traditional methods.

“The visual results in the paper really drive this home.”

Causal Insight in AI

10:20 to 12:20

The importance of causal insights in AI and implications of DoubleGen for future applications.

“If you connect this to the bigger picture, DoubleGen is pushing generative AI towards genuine causal reasoning.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If you've spent any time marveling at the current generation of AI, the massive language models, the image generators, they are just incredible mimics, aren't they? Absolutely masters of imitation. They reflect the world they've seen in their training data with, frankly, stunning accuracy. Right. They mirror the patterns perfectly. But here's the really tricky part, the next step. Can they answer what if questions? You know, not just predict based on what is, but model what would be if we intervened. We're talking about counterfactuals. Exactly. That's the key term. Data that would have arisen if things had been different, if an intervention had happened.

0:38Which, of course, it didn't in the real data we have. Precisely. So we need to generate that missing piece. And that's what this deep dive is all about. A new framework called DoubleGen. It's designed specifically to try and generate these, well, unbiased counterfactuals using the often messy observational data we actually have. And messy is the right word because real world data isn't like a clean lab experiment, is it? Not at all. It's inherently biased. Think about a medical trial scenario you mentioned earlier. Yeah, the example where maybe a new drug only goes to the sickest patients. Exactly.

1:12So if you just train an AI on that data, naively, it might look at the outcomes and think, wow, this drug is linked to people being really sick. Maybe even conclude it's harmful. Because it's mixing up the drug's effect with the fact that only very ill people got it in the first place. That's it. That's confounding bias. It pops up because the groups receiving or not receiving an intervention are different in systematic ways. It's almost never random in the real world. Okay, so we know we have confounding. And data scientists have ways to try and correct for that, right? Yeah. But that introduces its own potential headache.

1:47It certainly can. It's what we call misspecification bias. You try to fix the confounding using, let's say, an auxiliary statistical model. But what if that model itself is wrong? Ah, so your correction tool is flawed. Right. If your assumptions about how the confounding works are incorrect if the model is misspecified, then your correction might actually make things worse, or at least not truly fix the bias. Okay, let's make this concrete. The Celebi data set the celebrity faces seems like a good way to see this bias visually. It really is. Let's say we ask the model, show me this person smiling.

2:19A simple request, seemingly. Seems simple. But smiling in that data set is heavily correlated or confounded with other things. The source data shows this clearly. Smiling faces were way more likely to also have lipstick. What were the numbers again? It was something like 56 % of smiling faces had lipstick versus only 38 % of non-smilers. Wow, that's a big gap. And it continues. Makeup, 47 % on smiling faces, just 30 % on non-smiling ones. And the biggest one, female label, 65 % for smilers versus 52 % for non-smilers. So smiling is tangled up with being female, wearing lipstick, wearing makeup in the data.

2:57Exactly. It's correlation, not causation. So a standard naive model told to make someone smile doesn't just add a smile. It also adds a bit more lipstick and maybe makes the face subtly more feminine. Right. Because that's the pattern it learned. Precisely. It reflects the bias baked into the observations. And that's where DoubleGen comes in. The goal is to isolate the causal effect of the smile itself. What would this specific person look like smiling, stripping away those confounding factors? Getting the true counterfactual. Okay, so let's unpack DoubleGen itself. We have older methods. You mentioned inverse probability, weighting, IPW, and plug-in estimation.

3:33Why aren't they enough? What makes DoubleGen doubly robust? Yeah, that's the crucial difference. Those existing methods, IPW and plug-in, they're generally singly robust. Meaning? Meaning they rely entirely on getting one auxiliary model correct. IPW, for example, hinges on correctly estimating the propensity model, the probability of receiving the treatment. Plugin relies on getting the outcome model right, how the outcome changes with the treatment. So basically you have one shot to get that helper model right if you mess up that one model. The whole correction fails. The bias comes flooding back in.

4:07It's a single point of failure, which, let's be honest, is pretty risky when you're dealing with complex, real-world data where your models are almost certainly imperfect approximations. Okay, that makes sense. It feels fragile. So how does DoubleGen avoid that fragility? It changes the game by building the correction differently, right into the generative models training. And critically, it uses two auxiliary models working together. Two models. Okay, what are they? What do they do? So first, there's the propensity model. Just like an IPW, this one tries to figure out the probability that someone got the intervention.

4:41Like, why was this celebrity smiling? Based on their other features, the confounders. It tackles that selection bias. Right, the who gets what and why. Okay, model one. And the second. The second is the outcome model. This one's a bit more complex. Technically, it's often a conditional transport map. But think of it as modeling how the outcome should change given the intervention after accounting for the confounders. It predicts what the face should look like if it smiled, controlling for those other factors. Hmm. Okay. Two models sounds more complicated. Why is that better? Is it just redundancy?

5:15It's a very specific kind of redundancy. It's designed mathematically so that DoubleGen gives you a valid, unbiased, counterfactual, even if only one of those two models' propensity or outcome is correctly specified. Ah, so if your propensity model is a bit off. The outcome model can still carry the load, potentially. And if your outcome model is misspecified. The propensity model, if correct, can still lead you to the right answer. It's like having a backup system built in. It acknowledges that we probably won't get both helper models perfectly right in practice. That sounds much more, well, robust.

5:49It anticipates failure. Exactly. And this isn't just for pictures. The paper shows it works across different types of generative models. Diffusion models for images, yes, but also flow matching models, and importantly, autoregressive language models. So the kind of models powering chat GPT and other large language models. That's right. This double robustness principle could be applied to make LLMs generate more causally sound text, too. It's broadly applicable. Okay, this is where it gets really interesting for me. The proof is in the pudding, right? Let's talk experiments. Did they actually show this works, especially when things go wrong?

6:23They did. Back to the Celebe faces. First, the baseline. When everything was set up correctly, both propensity and outcome models well specified, DoubleGen performed great. It was right up there with plug-in and IPW and actually showed better recall. Recall meaning? Meaning the faces it generated covered the true diversity of potential smiling faces better. It was 0.74 recall versus about 0.65 for the naive method that just learned from smiling faces only. Okay, so it's good when things are good. Now, the stress test. What happened when they deliberately messed up one of the helper models? Let's take the plug-in method first.

7:01It relies only on the outcome model, right? What if that model was wrong? That's where you see the single robustness fail spectacularly when they fed the plug-in method a misspecified outcome model. Disaster. Pretty much. They measured the quality using something called the kernel arc face distance KD. Lower KD is better means the generated faces are closer to the true unbiased distribution. Okay, so what were the KD numbers? With the correct outcome model, plug-in got a KD of 0.68. Respectable. But with the incorrect, misspecified outcome model, the KED shot up to 2.17. Wow, more than tripled.

7:40So the generated faces became much less realistic or much more biased again. Exactly. That huge jump signifies a major failure. The confounding bias, the extra lipstick, the feminization likely crept back in strongly because its only correction mechanism was broken. And double gen in that exact same scenario with the bad outcome model. Rock solid. Because its propensity model was still working correctly, DoubleGen maintained its performance. Its KD stayed right down at.68. Okay, that's the double robustness in action. One model fails, the other sees it. Precisely. The visual results in the paper really drive this home.

8:12You see the biased outputs from the standard model versus the clean, de-biased counterfactuals from DoubleGen. It's a clear difference. That's compelling for images. But text is a different beast. Did they manage to show this works for language models too, like with those Amazon reviews? They did, using a clever semi-synthetic setup. Real reviews, real product features, but they synthetically assigned an intervention. Imagine it like flagging certain reviews but did it rarely, only about 3.5 % participation. Why synthetic? Because then they knew the actual ground truth for the counterfactual. They knew precisely what the distribution should look like if the intervention had been applied differently, allowing for rigorous evaluation.

8:51Smart. And because the intervention was rare, the basic naive model must have struggled. Big time. For instance, the way they set up the synthetic intervention meant it was less likely for items in the books category. So the naive model, trained only on the few intervened examples, learned to underrepresent book reviews when ask-generate counterfactuals. It just didn't see enough of them. It inherited the synthetic bias. So how did DoubleGen fare against, say, IPW when they started introducing errors into the helper models for text generation? Again, the W robustness paid off. When they misspecified the propensity model, for example, DoubleGen did better than IPW.

9:31They used Wasserstein error to measure this lower as better. And the scores? DoubleGen got a Wasserstein error of 0.12, whereas IPW was higher at 0.17. So DoubleGen's generated text distribution was closer to the ground truth. Okay. And what about when the other helper model, the outcome model, was wrong? Still strong performance from DoubleGen. They use another metric here called MAUVE, which looks at both the quality and diversity of the text. DoubleGen consistently scored around 0.820.83 on MAUVE, even with misspecified models, the naive model. It was stuck down at 0.72. So across images and text, when you introduce realistic imperfections in the modeling process, DoubleGen holds up much better.

10:14That's the core finding. It provides a more reliable way to generate these what-if scenarios from observational data. This feels like a really significant step beyond just imitation for AI. Absolutely. If you connect this to the bigger picture, DoubleGen is pushing generative AI towards genuine causal reasoning. Instead of just reflecting the world of biases and all, it's about rigorously modeling how the world would look under different conditions. It's moving towards prediction based on intervention. And it's not just a neat trick. There's solid theory behind it. You mentioned theoretical guarantees.

10:46That's right. The research provides proofs that, under certain assumptions, DoubleGen can achieve what's called Oracle optimality, meaning it performs as well as if you magically had access to the true, perfect, counterfactual data the oracles view. It also achieves minimax rate optimality, which basically means it's as statistically efficient as possible for this type of problem. It's theoretically sound. Okay, so let's bring this home. For you listening now, maybe you're using LLMs for analysis or you're in research, maybe even drug development. What's the big takeaway here? The big takeaway is that if you're using generative AI trained on real world observational data for anything involving what if questions or predicting the impact of interventions, policy changes, medical treatments, business strategies, you need to be thinking about causal inference and robustness.

11:35Because otherwise. Otherwise, these incredibly powerful models might just be confidently feeding you biased predictions, reflecting the historical confounding in the data rather than the true causal effect you're interested in. Frameworks like DoubleGen, this doubly robust approach, are becoming essential for trustworthy AI in these high-stace applications. We need to move beyond naive imitation. Demand more than just pattern repetition. Demand actual causal insight. That makes perfect sense. And maybe one final thought to leave you with. The paper draws a fascinating link. These causal inference problems.

12:08Mathematically, they look a lot like problems where you simply have data missing at random. Which suggests that this idea, this double robustness framework, might actually have much broader applications. Think about all the machine learning tasks where you have massive data sets, but large chunks are incomplete or missing. DoubleGen's principles might offer a powerful new way to handle that missingness and improve predictions across a whole range of ML problems, well beyond just counterfactuals. It could be a key to unlocking more reliable insights from imperfect data everywhere.

From the publisher

The academic paper introduces **DoubleGen**, a novel, doubly robust framework designed to adapt standard generative models—such as diffusion models, flow matching, and autoregressive language models—to generate **counterfactual data**. Unlike existing methods that are only singly robust and susceptible to bias if auxiliary models are misspecified, DoubleGen remains valid if either the propensity score or the outcome model is correctly specified. The research addresses the challenge of **confounding** in observational data, where models trained naively might internalize skewed relationships, leading to inaccurate counterfactual predictions (e.g., predicting outcomes if everyone received a new treatment). The authors provide **theoretical guarantees**, including minimax rate optimality for DoubleGen diffusion models, and demonstrate the framework's effectiveness and **robustness to misspecification** through experiments generating counterfactual celebrity faces and product reviews.

More from Best AI papers explained

All 475 episodes
DoubleGen - Debiased Generative Modeling of CounterfactualsBest AI papers explained · 13 min
Listen in VO