In short
GenAI-Powered Inference (GPI), a statistical framework (Yamai & Nakamura, Harvard; paper dated July 8, 2025) that uses generative AI to do causal/predictive inference on unstructured data (text, images) by extracting low-dimensional “internal representations” (R) and then learning a deconfounder to reduce hidden confounding.
Guest backgrounds
No guests are named in the transcript; it’s a host-led deep dive.
Key claims
GPI processes existing data (not generating new data), uses open-source models with deterministic outputs for replicability, avoids fine-tuning, and improves precision/robustness while handling confounders via a learned deconfounder.
Notable examples
Reanalyzes Weibo censorship (75k posts/4k users) using Llama 3 and Gemma; finds stronger self-censorship and stronger repeat-censorship effects, better covariate balance, and more efficient use of the full sample. Reanalyzes Danish election face traits (7k+ candidates) using Stable Diffusion 1.5/2.1; dominance is not significant under GPI, and results are more stable. Revisits a conjoint experiment on political arguments (3,300 participants; 336 arguments) using LLM representations; authority appeals are most persuasive, ad hominem least, and effects are more precise than the original parametric model.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding GPI and Its Importance
0:42 to 2:50
Discover how GPI helps analyze complex unstructured data more efficiently.
“And GPI isn't about making like pretty pictures or fancy text.”
The Process of GPI: From Data to Insights
2:50 to 5:01
Learn about the steps in GPI, including using internal representations and deconfounders.
“So using these R representations gives you three main advantages.”
Real-World Application: Censorship on Weibo
5:01 to 8:10
Examine how GPI analyzes the effects of censorship on social media platforms.
“Much more broadly accessible for researchers.”
Investigating Facial Traits and Election Outcomes
8:10 to 11:19
Understand the impact of facial traits on election results through GPI analysis.
“So GPI gets you similar answers, but with more statistical power, using all the data.”
Persuasiveness of Political Arguments through GPI
11:19 to 14:01
Explore how GPI enhances the analysis of political arguments and their effectiveness.
“Still large, but way, way smaller than the raw pixel data of the photo.”
Understanding GPI's Methodology
14:01 to 14:36
Learn how GPI improved upon traditional models of persuasiveness using LLMs.
“How did the original study analyze that?”
Key Findings from GPI Analysis
14:37 to 16:33
Discover the main takeaways about persuasive arguments from GPI's analysis.
“Again, they used LLMs to process the arguments.”
Precision and Robustness of GPI
16:34 to 17:55
Explore how GPI's approach achieves greater precision in persuasive analysis.
“especially given that GPI doesn't impose those restrictive assumptions the original model did.”
Future Directions of GPI
17:56 to 18:38
Examine potential future applications and discoveries enabled by GPI.
“So for you, the listener, it offers a path toward more reliable insights, especially in fields where you've got information overload, hidden variables.”
Broadening GPI's Reach
18:39 to 20:08
Consider the implications of GPI extending beyond text and images.
“We know it's adjusting for something important in the text or image, but figuring out precisely what that something represents.”
Transcript
Automatic transcript. May contain errors.0:00The world of generative AI, well, it's just exploding, isn't it? Transforming everything, art, code, you name it. It really is everywhere now. But what if it could do more than just, you know, create? What if it could help us find deeper truths and really complex data, like for actual statistical analysis? That's exactly where things are headed, moving beyond just generation. So today, we're doing a deep dive into something I think is genuinely groundbreaking. It's called Gen AI Powered Inference, or GPI for short. And this comes from a fascinating new statistical framework outlined in a paper from Kosuke Yamai and Kentaro Nakamura at Harvard.
0:37Just published actually July 8th, 2025. Very recent. Yeah. Cutting edge stuff. Exactly. And GPI isn't about making like pretty pictures or fancy text. It's about using these powerful AI models to pull out real actionable insights from data that's traditionally, well, messy, hard to analyze. Unstructured data, basically. Text images. Right. Think of it as a kind of shortcut for you, the listener, to get really well informed about this cutting edge research. It promises to shake up how we look at causal and predictive analysis with this kind of data. It's got huge potential. Okay, so our mission for you today is to unpack this ingenious framework, figure out why it's such a big deal, and look at how it's already changing our understanding of some real world stuff.
1:23All right, let's get into it. Let's do it. So the big problem this research tackles, it's working with that unstructured data we mentioned. You know, huge amounts of text, maybe thousands of images. Right. Data that doesn't fit neatly into spreadsheets. Exactly. And when you're trying to figure out, say, a causal effect, like does A cause B, or even just a predictive effect, does A predict B? These big, messy data sets often hide things, confounders, right? These unobserved factors that just mess up your analysis. It's like trying to see through fog, like you said in the intro. You know something's there affecting things, but you can't quite pin it down.
1:58Yeah, that's a great analogy. So how does GPI cut through that fog? Well, what's really fascinating here is how it uses these readily available, often open source, generative AI models. Things like large language models, diffusion models for images. OK. But here's the key. It's not using them to create brand new data from nothing. It uses them to process the existing unstructured data. Ah, okay. So processing, not generating from scratch. Precisely. Their strength lies in their ability to sort of replicate or understand that unstructured data. And in doing that, they extract what the researchers call low-dimensional representations.
2:37Low-dimensional representations, like a summary. Exactly. A much smaller, distilled version of the original text or image. They call these internal representations, or just R for short. So it's about distillation, you said. Why is that R value so crucial? What does distilling it down achieve? Right. So using these R representations gives you three main advantages. First, R is way smaller, dimensionally speaking, than the original data. Okay. So less data to crunch. Makes it much more computationally efficient. I mean, imagine, we'll see an example later, taking a photo with hundreds of thousands of pixels down to a vector of like 16 ,000 numbers.
3:15Wow. Okay. That's a big reduction. Huge. Second, R already contains really rich, relevant information. Why? Because it's built from all the data the original Gen EI model was trained on. It sort of knows things about the world already. Ah, leveraging the model's preexisting knowledge. Exactly. And third, and this is big for research, GPI uses open source models and makes sure they run deterministically, meaning you get the same R every time you run the same input. That helps with replicating studies, Rob. Hugely important for scientific replicability. Okay, that makes sense. So they get these distilled R representations.
3:50Yeah. What's the next step in the process? So then they apply pretty standard machine learning techniques to these R representations. And the goal is to identify something they call a deconfounder. A deconfounder. So that's the thing that deals with the hidden factors. You got it. Think of the deconfounder as GPI's way of automatically finding and neutralizing those hidden variables, those confounders we talked about, that could bias your results. It's like a smart filter for those hidden biases. And this is really critical because trying to directly control for super high dimensional data, like raw images or the full text of documents, that can actually introduce bias.
4:29How so? It can violate something called the positivity assumption in statistics. It's a technical point. But basically, you need enough overlap between your different groups. If you try to control for too many things directly, especially complex things, you might not have that overlap, leading to biased estimates. Right, I see. GPI cleverly gets around this. It doesn't need to explicitly model how the original data was generated. And crucially, it doesn't require you to fine-tune the big Gen AI models themselves. Which makes it much more accessible, I guess. You don't need massive resources to fine-tune.
5:00Exactly. Much more broadly accessible for researchers. So it's like it automates a lot of the really hard statistical work and avoids some common pitfalls. That sounds, well, potentially like a game changer. It really could be. OK, let's make this concrete. Let's dive into the first real world example they looked at. Trying to figure out the effects of government censorship on social media. Right. A really complex issue. Yeah, because censored posts almost by definition are different from uncensored ones. Right. That creates this inherent confounding bias. it's hard to compare apples to apples.
5:34Absolutely. So this application, it actually reanalyzes an earlier study, one by Roberts and colleagues from 2020, looking at censorship on Weibo, the Chinese social media platform. Okay, Weibo. What did the original study do? The original collected a lot of data, over 75 ,000 posts from more than 4 ,000 users via the Weiboscope project. They were trying to see if users who got censored were then, you know, more likely to get censored again, or if they perhaps started self-censoring, posting less potentially sensitive stuff. Right. So how did GPI tackle this specific data set? What did Imae and Nakamura do?
6:08They took two well-known open source large language models, Llama 3, the 8 billion parameter version, and Gemma 3, a slightly smaller one. And they used these models to essentially regenerate or process each Weibo post. The goal was to extract its internal representation, that R value. Specifically, they used the representation associated with the last token of the post. Why the last token? In many LLMs, the final token's representation acts as a kind of summary of the entire sequence, capturing the overall meaning. So it's a compact way to get the essence of the post. Gotcha. So they got these representations.
6:43Which were 4096 dimensions for Lama 3, 1152 for Gemma 3, still big, but much smaller than the raw text. Then they fed these representations into neural networks to estimate the deconfounder and model the outcomes like future censorship or posting activity. Okay, so what did GPI find? Did it differ from the original analysis, which I think used something like text matching? Yeah, the original used text matching. And here's where it gets really interesting. GPI found something the original study didn't really pick up on. Prior censorship significantly reduces how much users post afterwards. Ah, some clear evidence of self-censorship, people getting quieter after being censored.
7:23Exactly, a chilling effect. The original analysis didn't show this clearly. This suggests, you know, maybe traditional methods were underestimating this impact. GPI seems more sensitive here. Wow. Any other differences? Yes. GPI also found a much stronger effect of prior censorship on the likelihood of being censored again in the future. So a stronger link between being censored once and facing it again. And were these findings solid? Like, did both AI models give similar results? Yes, remarkably consistent across both LMA3 and GEMA3. That's important, right? Shows it's not just an artifact of one specific model.
7:58Definitely adds confidence. And what's more, GPI's estimates using the full sample of data closely match the results from the original study's matched sample. But GPI's results were statistically more efficient, more precise. So GPI gets you similar answers, but with more statistical power, using all the data. That's the takeaway. Advantages in both robustness and precision. But how did they, you know, prove GPI was doing a better job handling those confounders? That seems key. Good question. They check something called covariate balance. It's a standard technique. Essentially, you check if your adjustment method successfully balanced out known potential confounders between the groups you're comparing, like censored versus uncensored users.
8:41OK, how did they check that balance? They calculated the correlation between a score derived from GPI, the estimated efficient score, and a known potential confounder. In this case, they used the proportion of certain keywords related to censorship or sensitive topics within each post. Right. And what did they find? GPI consistently produced very small correlations with large p-values, especially using the full data set. That's what you want to see. It indicates that GPI effectively adjusted for that keyword confounder. And the original method. In contrast, the original text-matching approach, when they checked its balance, still showed a statistically significant correlation for one of the outcomes.
9:17A small p-value, suggesting there might still be some residual confounding left over, GPI seemed to clean it up better. Fascinating. So GPI seemed to handle the hidden text features more effectively. Okay, but GPI isn't just for text, right? You mentioned images. The next application looks at faces, specifically politicians' faces. That's right. Moving from text to the visual domain. We all kind of do this subconsciously, right? Look at a face and make snap judgments, trustworthy, competent. Psychology research backs this up. It does. So this application revisited a study by Lindholm and colleagues from just this year, 2024.
9:52They looked at whether facial appearance predicts election results in Denmark, local and national elections. OK. Danish elections. What did the original study involve? They gathered photos of over 7 ,000 candidates. Then they used a specialized AI model, a fine-tuned convolutional neural network, to rate these faces on traits like attractiveness, trustworthiness, competence, dominance, etc. So, quantifying facial traits. Then what? Then they ran a pretty standard statistical analysis, ordinary least squares or OLS regression, to see if these quantified traits predicted the candidates' vote counts after controlling for basics like age, gender, education.
10:32But what was the limitation there? The potential limitation was that the original analysis might not have accounted for other visual cues in the photos beyond those specific traits they measured. Latent features, things the researchers didn't explicitly code for, that might also predict votes. Ah, more potential confounders, but visual ones this time. So how did GPI step in? GPI processed each candidate's photo using two versions of Stable Diffusion. That's a popular open source AI model for images, version 1.5 and 2.1. Okay. Stable diffusion. And it extracted the R representation again. Exactly.
11:05But here, instead of just the last token, like with text, they took the full internal representation the model generated for the image, which is usually a complex tensor, like a multidimensional grid of numbers. To duplicate it. It is, but they flattened it. So a 16 by 16 by 64 tensor became a single vector, a list of numbers, with 16 ,384 dimensions. Still large, but way, way smaller than the raw pixel data of the photo. Okay, a compressed visual summary. Then what? Then they re-estimated the predictive effects of those facial features. Attractiveness, competence, etc. But this time, adjusting not only for the observed things like age and gender, but also for these latent features captured in the GPI representation, the R value derived from the whole image.
11:48Right, using the deconfounder idea. Yeah. So what did GPI reveal about faces and votes? Any surprises? A couple of key things. First, the GPI estimates were more robust. When they added or removed the control variables, like age and gender, the GPI results for the facial traits didn't bounce around as much as the original OLS results did. More stable findings. That's good. Definitely suggests it's better handling all the underlying factors. Second, and quite interestingly, GPI consistently found that facial dominance was not statistically significantly linked to election outcomes. Really? Because the original study found a link.
12:22Sometimes. The original OLS results were inconsistent. Sometimes dominance seemed predictive, sometimes not, depending on whether covariates were included. GPI consistently said no significant effect. So GPI challenges that idea that a dominant-looking face necessarily helps you win elections. It certainly casts doubt on it, suggesting the previous findings might have been less stable. This robustness likely comes from GPI's better ability to account for all those other subtle visual cues in the photos. Beyond just the measured traits, it gives a more nuanced picture. And again, consistent across both versions of stable diffusion.
12:59Yes, consistent across 1.5 and 2.1. More evidence for the reliability of the GPI approach itself. All right, one more application to really show the versatility here. This one involves integrating GPI into a more complex kind of statistical model, looking at political arguments. Yeah, evaluating the persuasiveness of different rhetorical strategies. This is a big topic in political science, communication. Trying to figure out what kind of argument actually convinces people. Exactly. So this work revisits another study, this one by Blumenau and Lauderdale from 2022. They did what's called a conjoint experiment.
13:32Conjoint experiment. Yeah. They showed over 3 ,300 participants pairs of political arguments. There were 336 unique arguments in total, varying things like the policy issue, the rhetorical style used like appealing to authority, or maybe using statistics and whether the argument was for or against a position. Okay, lots of variations. And participants chose? They just had to say which of the two arguments presented they found more persuasive. Simple enough. How did the original study analyze that? They used a specific type of statistical model, a parametric structural model. That basically means they made some strong assumptions about the mathematical form of how persuasiveness works, and they only adjusted for confounders they could manually identify, like how long the argument was or maybe its emotional tone.
14:19Okay, so assumptions involved and maybe not accounting for all hidden factors in the text itself. Precisely. That left the door open for potential residual confounding and relied on those potentially restrictive functional form assumptions. So how did GPI improve on this? Did it get rid of those assumptions? It allowed them to relax those strong assumptions. That's the key here. They didn't have to assume a specific mathematical formula for persuasiveness. How did they do that? Again, they used LLMs to process the arguments. They regenerated all 336 arguments using three different models this time.
14:51LMA38B again, a much larger LMA 3.370B model, and GEMA 31B again. Okay, getting representations from each. Yep, extracting the unique internal R representations for each argument from each model. Then they estimated a semi-parametric version of the original structural model. Semi-parametric, meaning? Meaning partly parametric, partly not. They used a flexible neural network to model persuasiveness based on the rhetorical element being tested, and the GPI-derived deconfounder, which captures all those subtle text features. No rigid assumptions about the functional form. Very clever. So what were the main takeaways about persuasive arguments from this enhanced GPI analysis?
15:30Well, GPI confirmed some things, but with more certainty. It found that appeals to authority were indeed the most persuasive type of rhetoric used in these arguments. Citing experts works. What about the least persuasive? Ad hominem attacks. Arguments that attack the person making an opposing point rather than the point itself. GPI found these significantly reduced persuasiveness. No surprise there, maybe. Yeah, seems logical. But how did GPI's findings compare overall to the original study? The general patterns were largely consistent, which is reassuring. But the GPI estimates were substantially more precise.
16:08The statistical uncertainty around the estimates was much smaller. More precise, meaning they were more confident about the results. Exactly. For instance, the original study, despite having similar point estimates, didn't find statistically significant effects for appeals to authority or arguments based on cost-benefit analysis. The results were too uncertain. But GPI did. GPI did find significant effects for both of those thanks to its increased precision. This is pretty remarkable, especially given that GPI doesn't impose those restrictive assumptions the original model did. Usually, fewer assumptions mean less precision, but here, GPI delivered more.
16:43That is impressive. And consistency across the different LLMs, even the big 7 dB parameter one. Yes, absolutely. They checked the estimated persuasiveness scores for each rhetorical element across all three LLMs, LMA 8B, LMA 7DB, GEMMA 1B, and found they were highly correlated. So it doesn't seem overly sensitive to the specific AI model used to generate the representations. Which again speaks to the robustness of the underlying GPI framework. Okay, let's pull this all together. What does this mean for you listening? We've seen GPI applied to Chinese social media censorship, faces in Danish elections, political rhetoric.
17:20Yeah, quite a range. It's this new statistical framework, Gen.A.I. powered inference, using open source AI models not just to create, but to analyze. To analyze that messy, unstructured data for real, causal, or predictive insights. I think it really boils down to making these incredibly powerful AI tools accessible and genuinely useful for rigorous scientific research. That's the bridge it's building. Right. By pulling out these internal representations, these R values, and using them to find and adjust for those tricky unknown confounders, GPI gives us a more robust, more precise way to get clearer answers from complex data, whether it's text or images or potentially other things too.
18:03So for you, the listener, it offers a path toward more reliable insights, especially in fields where you've got information overload, hidden variables. Yeah. Things that usually obscure the truth. Yeah. It's almost like GPI gives you special glasses to see the unseen forces shaping the data. Does that make sense? It does, yeah. It's revealing hidden structures. And the researchers themselves, Ima and Nakamura, they point towards some really exciting future directions, things that really get you thinking. Oh, yeah. The potential is huge. Like, first, interpreting the deacon founders. GPI finds these latent confounding features, right?
18:36But what are they exactly? Right now, it's a bit of a black box. We know it's adjusting for something important in the text or image, but figuring out precisely what that something represents. That could unlock even deeper understanding, couldn't it, about human behavior or society? Absolutely. Imagine knowing the specific subtle visual cue in a face, or turn a phrase in an argument, that GPI identified as a key confounder. Second, they mentioned discovering new treatments. What if GPI could go beyond just adjusting for features we already suspect are important? What if it could sift through the unstructured data and actually discover entirely new features, new aspects of the text or image that are powerful predictors or even causes of outcomes, features we haven't even thought of measuring?
19:23That sounds like AI-driven discovery. Potentially, yes. Yeah. And third, we focused on text and images because that's what this paper did. But the framework seems general enough. Exactly. It could naturally extend to other kinds of unstructured data. Audio, video. Think about analyzing tone of voice from audio recordings or complex events in video footage. Wow. The potential really does seem immense. Bridging that gap between these super complex AI models and science that's transparent, repeatable, and gives real insight. That's the goal. Imagine really pinpointing which specific elements of, say, a speech or an image are truly driving an outcome.
20:01Not just correlation, but maybe getting closer to actual causal or predictive power. Cleaned of confounding. That's the promise. Definitely a thought for you to mull over until our next Deep Dive.
From the publisher
This paper introduces GenAI-Powered Inference (GPI), a novel statistical framework for both causal and predictive analysis of unstructured data, such as images and text. GPI utilizes open-source Generative AI models to extract low-dimensional representations from high-dimensional unstructured data, which are then used in conjunction with machine learning techniques to quantify causal and predictive effects while also providing estimation uncertainty. This approach distinguishes itself by not requiring fine-tuning of generative models, thereby offering computational efficiency and broad accessibility. The paper demonstrates GPI's versatility through applications including analyzing social media censorship, predicting electoral outcomes based on facial appearance, and assessing the persuasiveness of political rhetoric, consistently showing enhanced robustness and precision compared to existing methods.




