In short
The episode explains why the common “two-step” workflow—use AI/ML to generate variables from text/audio, then run a standard OLS regression—can produce biased causal estimates and misleading confidence intervals. It centers on the April 2025 paper “Inference for Regression with Variables Generated by AI or Machine Learning” and introduces two fixes: analytical bias correction (recenter using classifier error rates) and joint maximum likelihood estimation (estimate AI outputs and regression jointly via Hamiltonian Monte Carlo).
Guests
No specific guest names or backgrounds are provided in the transcript.
Key claims
Standard errors may look correct, but first-order bias can shift the entire confidence interval (undercoverage). With large datasets, small AI misclassification errors can dominate.
Notable examples
Remote work wage premium using Distilbert on Lightcast job postings (99% accuracy; ~1% false positives; effect jumps from 36 to 64 log points after correction). CEO time-use leadership vs management using LDA on 916 CEOs (two-step works when per-CEO behavioral data is dense; fails when 90% of time units are removed; joint estimation recovers). Fed communication tone using BERT on 200 FOMC meetings (R-squared rises from 0.0425 to 0.1429 after joint estimation; attenuation bias corrected).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Dangers of the Two-Step Strategy Trap
1:30 to 4:10
Explore the risks of plugging AI outputs directly into regression models and the implications for data analysis.
“Because this goes far beyond just a theoretical problem for academics.”
Understanding First-Order Bias in AI Analysis
4:10 to 6:20
Discuss the concept of first-order bias and its effects on confidence intervals in regression analysis with AI data.
“Usually, when statisticians see a constructed variable being used in a regression, they worry about what is called a generated regressor problem.”
Case Study: Remote Work and Wage Inequality
6:20 to 10:00
Examine a case study on the wage gap between remote and in-person work using AI-generated data.
“We know the researchers looked at the Lightcast dataset, which is a massive repository containing hundreds of millions of online job postings.”
Proposed Solutions for Bias Correction
10:00 to 14:00
Learn about two mathematical solutions for correcting bias in AI data analysis and their practical applications.
“The sheer precision of the AI lulls you into a false sense of security.”
Introduction to CEO Time Use
14:00 to 14:19
Learn how CEO time use illustrates the effectiveness of AI in regression.
“Solution two is your expensive, heavy-duty software upgrade that keeps your entire system from quietly drifting off a cliff.”
Analyzing CEO Behavior with AI
14:19 to 16:28
Discover how AI classifies CEO behaviors and its impact on revenue.
“This application is incredibly instructive because it actually shows us exactly when the naive two-step strategy works and doesn't fail.”
Simulation Outcomes and Data Limitations
16:28 to 17:04
Understand the implications of reduced data on CEO behavior analysis.
“The naive two-step strategy completely failed.”
Central Bank Communication Analysis
17:04 to 18:34
Explore how AI analyzes Federal Reserve communications and its market impact.
“Okay, let's look at the final and perhaps most globally impactful application of the material, central bank communication.”
Understanding Attenuation Bias
18:34 to 19:33
Learn about attenuation bias and its effects on market analysis.
“They ran the joint estimation model on the exact same underlying data.”
The Golden Age of Unstructured Data
19:33 to 20:30
Reflect on the implications of unstructured data in decision-making.
“We've gone from the wages of food service workers in San Diego to the daily schedules of manufacturing CEOs, all the way to the global impact of the Federal Reserve.”
Show all 11 chapters
Challenges of Dynamic AI Models
20:30 to 21:40
Consider the challenges posed by continuously learning AI models.
“The standard two-step strategy is easy to code, it's easy to explain, and it's fast.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to this deep dive. Imagine for a second that you are building this massive data-driven strategy for your business. Right. You're sitting on a goldmine of unstructured information, like millions of customer emails, years of job postings, decades of internal company reports. A lot of data. Exactly. A ton of data. And you decide to use cutting-edge AI to read all of these documents and just categorize them for you automatically. For what everyone is doing right now. Right. Then you take all of that beautifully organized AI generated data and you plug it into a standard statistical model to find cause and effect relationships.
0:40You hit run. Your software tells you the results are highly significant. I feel great about it. You feel incredibly confident. Well, what if the very way you plug that data in, what if it structurally guaranteed your confidence intervals were entirely wrong? It's a terrifying thought for literally any data scientist or business leader. But it is the reality we are unpacking today. Because in a world of just absolute information overload, our knowledge is really only as valuable as our understanding of the structural limitations behind it. That's so true. Today we are examining a really groundbreaking April 2025 research paper.
1:16It's titled Inference for Regression with Variables Generated by AI or Machine Learning. And it's by Laura Battaglia, Timothy Christensen, Stephen Hansen, and Simon Sacker. And the mission for our deep dive today is simple. We are going to shortcut you to being the smartest person in the room regarding AI-generated data. I love that. We're going to explore exactly why plugging AI outputs directly into standard regressions is statistically dangerous, how this hidden bias radically distorts real-world economic findings, and most importantly, we will walk through the two brilliant methods the authors of this paper propose to fix it.
1:51Because this goes far beyond just a theoretical problem for academics. I mean, this is an issue of first-order importance for anyone relying on algorithmic tools for data analysis. Which is basically everyone. Which is almost every major industry at this point, yeah. Okay, let's unpack this. Let's start with part one, the two-step strategy trap. The trap, right. This strategy is the absolute gold standard practice right now in modern economics and data science. Step one, you use an AI or machine learning algorithm to quantify your unstructured data. So turning text into numbers. Exactly. You might train a model to turn the raw text of a job posting into a simple binary label like remote work or not remote work.
2:32Or maybe you feed an audio recording into a neural network to spit out a continuous tone of voice metric. Right. Then step Step two, you take that shiny new variable and you plug it into a standard ordinary least squares regression, an OLS regression as a covariate. You want to find a trend or make a forecast. On paper, it sounds perfectly logical. It does. You have a tool that extracts data and a tool that analyzes data. You just connect them. But that logical assumption is exactly where the trap lies. The workflow is modular, which engineers absolutely love. But researchers are naively treating these AI-generated variables as perfect, regular, numerical data.
3:10They aren't perfect at all. They inherently contain measurement error. Even a highly tuned, state-of-the-art AI makes mistakes. The research introduces a really brilliant asymptotic framework to explain what they call a sequence of DGPs, or data-generating processes. Sequence of DGPs. Okay, for those of us who haven't taken advanced econometrics recently, what does that actually look like in practice? Yeah, think of it as a mathematical model for a very realistic modern scenario. As your sample size increases, as you feed the algorithm more and more data, the sampling uncertainty shrinks. Makes sense.
3:45But the measurement error from the AI does not vanish. Ah. The sequence of DGPs framework models a situation where the AI's measurement error and the statistical sampling error remain comparable as the data set grows to infinity. Meaning, as you feed the beast more data, the tiny areas the AI makes don't just wash out in the noise. They scale right alongside your data collection. They do. And this scaling error leads to a massive misconception in the field. Usually, when statisticians see a constructed variable being used in a regression, they worry about what is called a generated regressor problem.
4:21Which is what? They assume the variance is going to be inflated, meaning they think the standard errors of their results are going to be too wide. Okay. But the paper proves mathematically that this assumption is actually completely false in this specific AI context. Wait, really? The standard errors are actually fine? The standard errors are perfectly consistent. the asymptotic variance, which is just a fancy way of describing how the data spreads out as your sample size approaches infinity, is the exact same as if you ran the regression on the true, perfect underlying data. Wow. The real problem is something much more insidious.
4:55It's first-order bias. The OLS estimator has the wrong centering. So the width of your confidence interval is totally correct, but the entire interval is shifted to the wrong place on the number line. Exactly. It reminds me of using a sniper rifle. Oh, I like this. Go on. Imagine you have a beautifully machined rifle and you fire a grouping of shots at a paper target. The grouping is perfectly tight. All the bullets are hitting exactly in the same tiny cluster. Right. That tight cluster is your standard error. Exactly. And because it's tight, you feel highly confident in your shot. But because the scope on the rifle is calibrated two inches to the left, your tight grouping is completely missing the actual bullseye.
5:37Yes. You consistently miss the truth while feeling incredibly confident that you hit it. That leads to massive undercoverage of the true effect. That is a brilliant analogy for first-order bias. You are highly confident in a biased result. Right. And because the measurement error in AI is often what statisticians call non-classical, meaning the errors the AI makes can actually be correlated with the true values, it is trying to measure you might not even know the direction of the bias. Your misaligned scope could be making you underestimate an effect or drastically overestimate it. Let's make this abstract math incredibly concrete.
6:13We'll move into part two, the first of three fascinating real-world applications explored in the material. This one involves remote work and wage inequality. A very hot topic. Definitely. We know the researchers looked at the Lightcast dataset, which is a massive repository containing hundreds of millions of online job postings. They wanted to analyze the wage gap between remote and in-person work. How did they set up the AI for this? They started by taking a small sample of those job postings and had human beings manually label whether each job offered remote work or not. Then they used that human-labeled data to fine-tune a large language model called Distilbert.
6:53Essentially, they taught the AI to read the text of any job description and output a label, remote or not remote. And from what I understand, the model performed phenomenally well. Yeah. achieved a 99 % test set accuracy. Which is incredible. Right. For most data teams, a 99 % accurate AI reading job postings is time to pop the champagne. That data is treated as basically as good as ground truth. It is very tempting to declare victory there. But the authors wanted to prove how dangerous that 99 % can be. Yeah. They zoomed in on a highly specific subset of the data to test the two-step strategy.
7:26they looked at 16 ,315 job postings specifically for accommodation and food services. That's NAICS Code 72 in San Diego, California. So a very specific, sizable local market. What happens when they run the naive two-step strategy on this San Diego data? Well, they replaced the unknown true remote status with the AI's 99 % accurate prediction, and they controlled for occupation fixed effects to make sure they were comparing apples to apples. Right. The standard regression estimated a 36 log point positive effect of remote work on wages. OK, let's translate 36 log points for a second. In plain English, a 36 log point increase translates to roughly a 43 percent wage premium for remote workers in that sector.
8:10That's a huge premium. It sounds like a massive, highly significant finding. The finding appears solid until you look under the hood of that 99 percent accuracy. What's fascinating here is even with just a 1 % error rate, when your data set is massive, like those 16 ,000 postings in San Diego, the sampling error of the regression gets incredibly small. Right, the tight grouping of shots. Exactly. That means the tiny 1 % misclassification error from the AI completely takes over the math. It distorts the inference entirely. To figure out exactly how much distortion was happening, the authors had to audit the AI, right?
8:46They manually checked a random sample to find the false positive rate. They did. They took a thousand postings. The AI classified 26 of them as remote. When human readers checked those 26 specific postings, they found that nine of them were actually mistakes by the AI. Wow. They weren't remote jobs at all. That gives an expected false positive rate of just.009, less than 1%. Here's where it gets really interesting. When the authors took that tiny.009 error rate and applied a mathematical bias correction to their original regression, The estimated effect of remote work on wages didn't just shift by a fraction.
9:21It exploded. It jumped by 76%. The effect went from an estimated 36 log points all the way to 64 log points. And a 64 log point jump is close to a 90 % actual wage premium. It is a staggering difference in the real world economic takeaway. Missing a massive jump like that simply because of a hidden 1 % AI error is wild. If you are a business leader or a policymaker and you are relying on highly accurate AI classifiers to make big decisions on large data sets, you are very likely underestimating the true effects of whatever you are measuring. You're trusting the scope on the sniper rifle without realizing it is severely misaligned.
10:03The sheer precision of the AI lulls you into a false sense of security. But thankfully, the material doesn't just point out the trap, it provides a map to get out of it. Which brings us to part three, the two proposed solutions. The paper gives us two entirely different mathematical ways to realign that scope. Let's start with the first one, which they call the analytical bias correction. How does this work? This is an incredibly elegant, lightweight fix. If you know the false positive rate of your AI classifier, you can mathematically recenter the confidence interval. Just slide it over. Right.
10:35We just saw the result of this in the remote work example. But the absolute brilliance of this approach is the methodology used to find that false positive rate in the first place. You do not need a massive, wildly expensive validation sample where humans have to read and label thousands of random data points to find the true values. Which is huge because that is usually the biggest bottleneck for any data team. Nobody wants to pay analysts to hand label 10 ,000 documents just to double check their expensive AI. Exactly. The authors show that you only need to audit a tiny fraction of the AI's predictions, specifically the positive predictions.
11:10In the San Diego example, out of 1 ,000 postings, the AI only flagged 26 as being remote. The researchers only had to read those 26 specific postings to establish the false positive rate. That's nothing. It scales beautifully for modern, massive data sets where finding the needle in the haystack is the whole point of using AI. You let the AI find the needles, and then you just manually double-check the handful of needles it hands you. So for the business leaders listening, solution one is your quick, cheap, manual audit. It's highly practical. But what if you were in a situation where you can't easily audit the false positive rate?
11:47It happens. Like maybe the unstructured data is too complex or the output isn't a simple yes or no binary label, but a sliding scale of sentiment. What do you do then? That is where the second solution comes in. It is the heavy duty fix, joint maximum likelihood estimation. Okay. Joint maximum likelihood estimation. Instead of the traditional two-step strategy where the AI does its job, hands off a number, and goes to sleep while the regression takes over, joint estimation treats the AI model's probability outputs and the downstream regression as one massive connected statistical system. Wait, so treating the AI and the regression as one giant breathing system sounds amazing, but also incredibly complicated.
12:27Can standard software even handle that kind of math? You're asking it to look at thousands of AI probabilities and regression lines simultaneously. It is technically demanding. The model linking the unstructured text to the hidden true variable acts like an observation equation in a state space model. You are jointly estimating the regression parameters and the latent variables at the exact same time. The authors solved this computational hurdle using something called Hamiltonian Monte Carlo. Hamiltonian Monte Carlo. That sounds like a casino game designed by theoretical physicists. What is it doing under the hood?
13:02It is a type of Markov chain Monte Carlo algorithm. Think of it like trying to find the absolute lowest, deepest point in a complex, pitch-black rolling landscape. I'm picturing it. Instead of just randomly guessing where the bottom is by throwing darts in the dark, Hamiltonian Monte Carlo uses the gradient, the actual slope of the mathematical distribution, to guide it. Imagine dropping a marble into a large, curved bowl. Gravity naturally pulls it down the slope to the lowest resting point. Exactly. This algorithm uses the shape of the probability distribution to sample from it highly efficiently.
13:36And thanks to modern probabilistic programming languages, researchers can now essentially specify this complex likelihood in code, and the software automatically compiles it to perform the heavy lifting. To put that in perspective, solution one is your quick manual check. But if you are building an automated million-dollar algorithmic trading system where the AI has to run constantly in the background without human intervention, then solution two is what you need. Solution two is your expensive, heavy-duty software upgrade that keeps your entire system from quietly drifting off a cliff. A perfect summary.
14:10Let's see how these methods play out in the wild. Right. Moving into part four, the material looks at two more fascinating applications that really push these concepts to the limit. First up, CEO time use. This application is incredibly instructive because it actually shows us exactly when the naive two-step strategy works and doesn't fail. The setup here is fascinating. The paper looks at a massive 2020 study involving 916 actual CEOs in the manufacturing sector. Right. The original researchers had survey responses detailing the CEO schedules broken down into 15-minute intervals for a given week.
14:44Which is very granular. Extremely. They use a specific kind of AI, an LDA topic model, to take 654 different features of how a CEO spends their time, internal meetings, client calls, factory site visits, and compress all that noise into two main behavioral topics, management versus leadership. Okay. Then they regress the firm's overall sales on this new leadership variable to see if leader CEOs drove more revenue than manager CEOs. If we connect this to the bigger picture, the sequence of DGPs framework we discussed earlier, we can understand why the two-step strategy survived here. It comes down to the ratio of measurement error to sampling error.
15:20How so? Well, in this specific data set, there were only 916 CEOs. In statistical terms, that is a relatively small sample size for a firm level regression. But for each individual CEO, there were many, many 15-minute survey responses. Ah, I see. So the AI had a massive amount of behavioral data per person to figure out if they were a manager or a leader. Because it had so much data per person, the measurement error for classifying each CEO was extremely low. The measurement error was low relative to the sampling error of having only 916 companies. The noise from the small number of companies drowned out the tiny AI errors.
15:58Oh, that makes perfect sense. Because of that ratio, the two-step strategy worked fine. The standard OLS confidence intervals were valid. However, the authors of the current research didn't stop there. They ran a simulation to stress test their theory. What did they do? They artificially dropped 90 % of those time-use survey units. They basically asked the math, what if we only observed half a day of the CEO's behavior instead of a full, detailed week? You're starving the AI of the data it needs to make a confident classification. And the moment they did that, the measurement error spiked. The naive two-step strategy completely failed.
16:34It couldn't find any significant effect of CEO behavior on sales. The confidence interval shifted into irrelevance. Just like the misaligned sniper scope. Exactly. But when they applied the new join estimation method to that same starved low data data set, it successfully corrected the bias. It dug through the noise and found the true link between leadership and sales. That is such a powerful validation of the math. It proves the heavy-duty fix works even when your data is sparse. It really does. Okay, let's look at the final and perhaps most globally impactful application of the material, central bank communication.
17:12Specifically, how do financial markets react to the exact tone of Federal Reserve announcements? This is one of the classic problems in modern economics. We all know that the Federal Open Market Committee, or FOMC, releases highly scrutinized written statements about interest rates. And we know those statements instantly move billions of dollars in long-term bond yields. Right. But how do you mathematically quantify the tone of a dense, dry economic statement? The researchers replicated a sophisticated method using a BERT classifier, another large language model, on 200 FOMC meetings stretching all the way from 1995 to 2023.
17:48That's a lot of meetings. Nearly three decades of them. They trained the AI to read the dense paragraphs of these statements and labeled them as either hawkish or dovish. Then they took that AI-generated sentiment score and regressed a metric for the long-term yield curve, known as the path factor, against it. The raw statistical findings here are a masterclass in why this methodology matters to the bottom line. When they ran the standard, naive, two-step approach on this Federal Reserve data, they found a relatively weak positive effect. How weak? The R-squared value, which tells us what percentage of the market's movement is actually explained by the sentiment of the statement, was a tiny.0425.
18:28So using the standard method, it looks like the Fed's tone barely nudges the long-term market. But then they applied the heavy-duty fix. They ran the joint estimation model on the exact same underlying data. Yeah. The results completely transformed. The measured effect size of the sentiment nearly tripled. The R-squared jumped from.0425 all the way to.1429. What was happening in the first test was a phenomenon called attenuation bias. Attenuation bias. Let's break that down for the listener. Think of it like trying to listen to a beautiful piece of music on a radio station with a lot of static.
19:01The static represents the measurement error from the AI classifier. Okay. That static doesn't just add annoying noise. It actually washes out the true signal, making the music sound much quieter and weaker than it really is being broadcast. The naive AI approach was dramatically muddying the waters of the financial analysis. Wow. By treating the AI classification and the market regression as a joint system, the math was able to filter out the static, allowing the true strong signal of the market reaction to punch through. So what does this all mean? We've gone from the wages of food service workers in San Diego to the daily schedules of manufacturing CEOs, all the way to the global impact of the Federal Reserve.
19:42What is the grand takeaway for you, the listener? It's a big question. To me, it's that we are living in a golden age of unstructured data. We can finally measure complex human concepts like tone, sentiment, remote work prevalence and leadership styles at a massive, unprecedented scale. But as AI tools become perfectly invisibly integrated into our data pipelines, we absolutely cannot treat their numerical outputs as ground truth reality. We really can't. If you do, your confidence intervals are a mirage. You might be making million-dollar decisions based on a statistical ghost. The overarching lesson here is about how we must interact with technology as analysts and leaders.
20:22Critical thinking and understanding the why behind the math is our strongest defense against information overload and automated bias. Definitely. The standard two-step strategy is easy to code, it's easy to explain, and it's fast. But as we've seen across multiple distinct industries today, easy can be wildly expensively inaccurate. Taking the extra step to audit your false positives with an analytical correction or implementing a joint estimation model isn't just a statistical nicety for academics. It is the literal difference between seeing reality clearly and acting on a distorted reflection of it.
20:56Well said. I want to leave you with a final, slightly provocative thought to chew on. The incredible fixes we discussed today from this research, the analytical bias corrections, the complex joint likelihood models, they all rely on one crucial assumption. static error rates. They require knowing or mathematically estimating exactly how often the AI hallucinates or misclassifies in a given frozen data set. But what happens tomorrow? That's the million-dollar question. We're increasingly plugging in self-updating, continuous learning AI models into our pipelines. Their measurement errors aren't static.
21:32They are dynamic, shifting every single day based on real-world feedback and continuously updating training weights. The target is always moving. Exactly. How do we calculate a reliable confidence interval when the ruler we are using to measure reality is changing its own shape in real time? Something to think about next time you look at a dashboard and feel highly confident. Thanks for joining us on this Deep Tide.
From the publisher
This research investigates how using artificial intelligence (AI) or machine learning (ML) to generate variables for economic regressions can lead to biased estimates and invalid statistical inference. While researchers often treat AI-generated outputs as standard data, the authors demonstrate that measurement error in these variables—even from high-performance algorithms—shifts the centering of confidence intervals, making them unreliable. To address these distortions, the paper introduces two practical solutions: a mathematical bias correction that does not require ground-truth validation data and a joint estimation framework that models the latent variables and regression parameters simultaneously. The effectiveness of these methods is illustrated through diverse applications, including job posting classifications, CEO time-use analysis, and central bank sentiment indexing. Ultimately, the study provides a robust toolkit for economists to maintain statistical integrity when integrating modern computational tools into empirical research.




