In short
The episode explains a January 2026 economics paper that introduces “equivalent sample size” (ESS) to quantify how many real human survey observations an LLM effectively replaces when predicting human behavior, including causal effects.
Guest backgrounds
No guests are named; the episode is presented as a two-host discussion.
Key claims
ESS is a defensible metric (via blockout cross-validation and sequential hypothesis testing) that avoids data leakage by using ChatGPT-4 (cutoff April 2023) against PSID 2021 waves released publicly only in Oct 2024. LLM predictive value varies sharply by behavior.
Notable examples
Home ownership ESS 590–676 (LLM ≈ ~600 human profiles). Drinking habits ESS 41–59. Hourly wage ESS 17–20. Smoking habits ESS = 1 (ML beats LLM after one data point). For causal inference (Medicaid and primary care visits), they use a “transformed outcome” under unconfoundedness to apply ESS to counterfactuals.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroducing the Equivalent Sample Size Metric
1:30 to 5:00
The hosts explain the concept of the Equivalent Sample Size (ESS) and its significance in AI research.
“We are unpacking this really groundbreaking recent paper.”
Understanding the Research Methodology
5:00 to 10:00
Discussion on how researchers used human data from the PSID to test AI predictions and the potential issues of data leakage.
“So how did the researchers put this to the test?”
Evaluating AI Predictions vs. Traditional Algorithms
10:00 to 13:00
The hosts discuss the results of AI predictions compared to traditional machine learning algorithms on various human behaviors.
“A machine learning model like a random forest only needed to look at 20 real people to become better at predicting hourly wages than a billion parameter language model.”
Statistical Validation of Results
13:00 to 14:00
The segment covers how researchers ensured the reliability of their ESS results through robust statistical methods.
“cross-validation, you isolate one block, those 50 people, and you hide them.”
Understanding Incremental Improvements in LLMs
14:00 to 15:28
Learn how algorithms determine the necessary sample size for confidence in LLM predictions.
“the traditional algorithm just keeps getting incrementally higher, and they refuse to stop the game until they are mathematically certain the algorithm has permanently taken the lead.”
Exploring Conditional Average Treatment Effects
15:29 to 16:44
Discover the complexities of predicting causal relationships with LLMs.
“We've established how ESS works for predicting simple static facts.”
The Challenge of Counterfactual Predictions
16:45 to 19:19
Understand how LLMs tackle the difficulties of counterfactual scenarios in predictions.
“Do you just tack a hypothetical onto the end of the character sheet?”
Evaluating AI's Predictive Capabilities
19:20 to 20:26
Learn how the transformed outcome trick allows LLMs to simulate alternative realities.
“And then they compare the errors against the synthetic score.”
The Ethical Implications of Relying on AI Predictions
20:27 to 21:39
Consider the potential risks of depending solely on AI for understanding human behavior.
“Keep in mind, this research evaluated an older static model, Chad GPT-4, from 2023, and it found it was worth about 600 observations for certain demographics.”
Transcript
Automatic transcript. May contain errors.0:00I want you to imagine just for a second that you're a researcher. You've got this fascinating hypothesis about human behavior, and you have the analytical skills to test it. But there is one massive catch. You have absolutely zero budget. Yeah, which is pretty common, honestly. Zero dollars to work with. Right, exactly. Normally, to do this kind of work, you'd go out and pay, I don't know, 500 real people to take a survey. Or you'd buy access to a massive proprietary data set. But you just can't afford either of those things. So you're looking at the blinking cursor on your computer and you wonder, what if I just asked an AI to role play as those 500 people?
0:38Just bypass the humans entirely. Yeah. You open up a chat window and you type, pretend to be a 37-year-old architect from Tennessee. Do you own your home? And you just record what the AI says. The big question is, I mean, is that actually scientifically valid? It sounds a bit like a desperate shortcut. Yeah. Or maybe a thought experiment from like five years ago. But this is the actual reality of what's happening right now across economics, sociology, even corporate market research. People are actively using artificial intelligence as a synthetic stand-in for human beings. Which is wild to think about.
1:12It is. But the problem we've faced up until very recently is that nobody had a rigorous mathematical way to measure this. Like, is treating an AI like a focus group a brilliant cost-saving innovation? Or is it a catastrophic methodological mistake? And that brings us to the mission of our deep dive today. We are unpacking this really groundbreaking recent paper. It was published in January 2026 by economists Wayne Gao, Supjin Han, and Annie Lang. And what they've done here is just genuinely wild. Yeah, they didn't just sit around philosophizing about whether AI can simulate humans. They actually created a brand new metric.
1:49It finally puts a hard, undeniable number on exactly how much human data an AI is actually worth when it comes to predicting human behavior. So they've essentially engineered a universal ruler for AI's predictive power. Exactly. And before we can even look at how the AI actually performed in their experiments, and I'll tell you the results are heavily polarized, we really have to understand the mechanics of this new ruler. They call this metric the equivalent sample size, or ESS. Okay, so let's unpack this ESS concept, because understanding this is basically the engine for the rest of our deep dive, The way I visualize equivalent sample size is it's like a trivia contest between two very different types of competitors.
2:31I like this analogy. Go on. So on one side of the stage, you have a really well-read generalist who isn't allowed to study for this specific game. That is your pre-trained large language model, like a standard chat GPT. It just brings whatever general knowledge, assumptions, and, you know, internet reading it already has in its head. Right. And the crucial thing to remember about that generalist is that while it has this vast amount of background information, I mean, it's read Wikipedia, it's read Reddit, financial blogs, all of it. It hasn't seen the specific localized answers for this particular quiz.
3:01Right. And then on the other side of the stage, competing against the generalist, you have a student who walks in knowing absolutely nothing, just a total blank slate. But this student is actively being handed flashcards containing the exact answers for this specific test. This student represents a traditional machine learning algorithm. And when we talk about traditional machine learning, we're talking about algorithms like a random forest or a lasso regression. Right, the classics. Exactly. And let's just pause on those algorithms for a second to explain why they are the blank slate student.
3:34A random forest, for instance, operates by building these complex decision trees based strictly on the raw data you feed into it. If you feed a zero data, it cannot make a decision. It literally has no intuition. None. And a LASA regression is basically trying to draw the most efficient mathematical line through a scatterplot of data points. Again, no data points, no line. These algorithms possess zero common sense. They rely entirely on the real-world data points, which are your flashcards in this analogy, to figure out the patterns. Okay, so the game is afoot. The machine learning model starts with zero knowledge, but it learns rapidly as it flips through the flashcards of real-world data.
4:13The equivalent sample size, the ESS, is the exact minimum number of real-world flashcards that the blank slate student needs to look at before they finally outscore the AI generalist. That's a great way to put it. So if the traditional algorithm needs to study 100 human profiles to beat the AI's out-of-the-box guess, then the AI's ESS is 100. Which makes this metric incredibly practical for anyone trying to build a predictive model today. If you're that broke researcher we talked about or a startup trying to forecast consumer behavior, knowing the ESS tells you the literal physical data value of an LLM.
4:49Yeah, you have to know if you're getting the equivalent of a massive, expensive, hard to acquire data set for the cost of a few API calls or if you're getting something that is just statistically useless. But a ruler is useless if you don't actually measure anything with it. So how did the researchers put this to the test? They needed a massive baseline of human truth, which they found in the panel study of income dynamics or the PSID. The PSID is, well, it's somewhat legendary in the field of economics. It's one of the longest running longitudinal household surveys in the United States. It basically tracks families across generations.
5:26Wow, that's a lot of data. Oh, it is. It's packed with incredibly granular demographic and economic data on real human beings. We were talking age, education level, occupation, state of residence, religion, their parents' education, their income, just everything. And the researchers took all those rigid spreadsheet data points and translated them into natural language personas. They literally wrote prompts that read like a tabletop role-playing character sheet. Yeah, they did. A prompt would look like, you are a 37-year-old male living in Tennessee in 2020. You work in architecture and engineering.
5:59You've completed 14 years of education. You are married and have one child. and then they would throw down the gauntlet and ask, given your persona, do you own a house or do you rent? Forcing the AI to step into the shoes of a specific demographic profile like that, it tests its pre-trained understanding of how the world actually functions. But doing this with an AI immediately raises a massive methodological red flag. Wait, what kind of red flag? The risk of data leakage. Oh, because the PSID isn't exactly a secret, right? Right, exactly. If the AI read the 2021 PSID survey results while it was being trained by its developers, it's not actually predicting anything.
6:38It's just remembering the answer key. That makes total sense. If it's just regurgitating what it memorized, the whole experiment is pointless. Data leakage completely compromises an experiment like this. If the model memorized the answers, you aren't measuring predictive power. You're just measuring its hard drive. So to get around this, the researchers deployed a very strict temporal safeguard. What did they do? They evaluated an older, static model. Specifically, they used ChatGPT4, which had a hard knowledge cutoff of April 2023. They then tested it against the 2021 waves of the PSID data. The really clever trick here is that the 2021 wave wasn't actually released to the public until October 2024.
7:17Oh, wow. So it's physically impossible for the AI to have read it during training. Exactly. Those specific survey answers were locked in a statistical vault. Any prediction the AI makes has to come from synthesizing its broader knowledge, not from doing a quick search engine lookup. OK, but here's where it gets really interesting to me and frankly, a bit weird. Is it strange to ask an AI to roleplay a human to guess their hourly wage? Like, does the AI actually understand human economics or is it basically just playing a massive statistical game of Mad Libs? That's the million dollar question, isn't it?
7:49Yeah, I mean, are we trusting a glorified autocorrect to do economic forecasting? Well, your skepticism hits on the core philosophical debate about AI right now. Does it actually understand? In a human cognitive sense, absolutely not. But in a statistical sense, a large language model functions as a massive compression engine for human culture. A compression engine. Okay, I like that. Yeah, it has digested billions of articles, census reports, economic analyses. When you prompt it to be that 37-year-old architect, it isn't just picking words at random. It is activating a multidimensional web of correlations it learned from the internet regarding age, location, profession, and income.
8:32Okay, so let's see how good that compression engine actually is. Let's look at the flashcards. They tested the AI against the traditional machine learning algorithms on four specific human behaviors, and the findings are honestly all over the map. They really are. Let's start with the big winner, which was home ownership. Right. So for predicting whether someone owns a home or rents based on that character sheet, the LLM was astonishingly effective. The data shows its equivalent sample size was between 590 and 676 observations. Let that sink in for a second. The traditional machine learning algorithm had to study roughly 600 real human profiles, 600 flashcards before it could guess homeownership as accurately as the AI did right out of the box.
9:14It's wild. For this specific task, querying the AI is mathematically equivalent to deploying a survey team to interview over 600 people. If you have no budget, that is an incredible victory. So the LLM has basically internalized the socioeconomic markers of homeownership almost perfectly. It has. But the victory lap ends abruptly when we look at the other behaviors. Let's take the second one, which is drinking habits, basically predicting whether a person drinks alcohol. And we see a steep drop off here. A huge drop off. The ESS fell to between 41 and 59 observations. The AI still brings some predictive power to the table, but the traditional algorithm learns the patterns of drinking behavior and overtakes the AI much, much faster.
9:56Okay, and then we get to hourly wage. The AI's ESS was a measly 17 to 20 observations. That is practically nothing. No, it's really not. A machine learning model like a random forest only needed to look at 20 real people to become better at predicting hourly wages than a billion parameter language model. And then we hit the absolute floor. The fourth behavior, which was smoking habits. The AI completely failed here. The ESS was exactly one. Let's clarify what an ESS of one actually means because it's pretty brutal. It means the traditional machine learning algorithm outperformed the LLM after being fed a single data point.
10:35The LLM brought zero predictive advantage. Wow, so it was effectively no better than just flipping a coin. Exactly. A total coin toss. You know, it actually mirrors our own human blind spots in a way. If you handed me that character shade, I could probably guess if a 45-year-old married architect from Tennessee owns a home. The demographic clues are loud. Very loud, yeah. But guessing if they smoke cigarettes, that is a total coin toss for me too. The AI has the exact same intuition gaps we do. And let's dig into why that gap exists because it goes right back to how these models are built. An LLM is a text prediction engine trained on the internet.
11:09It only knows what humans write about. Right. It's learning from our collective output. Exactly. Think about the Internet. There are endless forums, census reports, mortgage calculators, articles discussing the economic milestones required to buy a house. Home ownership is deeply textualized and tied to structural variables like age and income. People blog about buying a house. They don't write thousands of essays about their daily smoking habits cross-referenced with their demographic metadata. of data. Precisely. Smoking is a highly individualized micro habit. It's noisy data. Because people don't write about it in relation to their exact demographics in the same structured way, the AI lacks the textual representation to learn the correlation.
11:53That makes so much sense. Yeah, the AI's pre-trained knowledge is fundamentally bound by what we choose to publish online. You cannot just make a blanket statement like, oh, AI is good at predicting human behavior. You always have to ask which behavior. Okay, I have to play devil's advocate here on behalf of anyone listening. Why should we trust these specific numbers? Like, 580 for homes, one for smoking. If I shuffle a deck of cards, sometimes I get three aces in a row just by pure chance. Right, the fluke argument. Yeah. How do we know these ESS numbers aren't just statistical flukes based on how the researchers fed the data into the models?
12:27Well, the researchers anticipated that exact critique. They didn't just want to provide a rough estimate. They wanted mathematically defensible certainty. So to prove it wasn't a fluke, they developed a really robust statistical procedure utilizing what they call blockout cross-validation. Okay, blockout cross-validation. Let's strip away the jargon for a second. How does that actually work in practice? Right. I mean, imagine taking all the thousands of people in the PSID survey and dividing them into distinct isolated blocks. Let's say blocks of 50 people. Okay, so we've got our groups. Right.
12:59In blockout cross-validation, you isolate one block, those 50 people, and you hide them. You train your traditional ML algorithm on everyone else, and then you test it on the 50 people you hid. Got it. And then what? Then you rotate, you hide a different block of 50, train on the rest, and test again. You average the errors together to ensure the algorithm is truly being tested on fresh data every single time. Okay, so that prevents the algorithm from just memorizing the whole data set, but how do you calculate the ESS from that without guessing? They use a framework called sequential hypothesis testing.
13:33They start the race with the ML algorithm trained on a tiny block of data, say 10 people. They test it against the LLM and ask, statistically, can we definitively rule out the possibility that the ML algorithm is as good as the LLM? And at 10 people, I'm guessing the ML's error rate is still huge? It is. So they reject the null hypothesis and move the hurdle higher. They train the ML on 20 people. They test again. So the researchers basically created a game where the hurdle for the traditional algorithm just keeps getting incrementally higher, and they refuse to stop the game until they are mathematically certain the algorithm has permanently taken the lead.
14:10That is the perfect way to describe it. They sequentially increase the training size of the algorithm. They only stop at the exact sample size where they can no longer mathematically prove that the LLM is better. Wow. Yeah. This creates a one-sided confidence interval. When the paper states the ESS for homeownership is 590, they aren't guessing. They are saying with 95 % mathematical certainty that the AI is worth at least 590 observations. But wait, I remember from my college stats class that statisticians hate it when you run tests over and over again. It's called the multiple testing problem.
14:44Oh, yes. It's a huge issue. Right. If you run 100 different statistical tests just by random chance. One of them will look significant even if it's garbage. So how did they avoid triggering massive statistical penalties for testing the algorithm at 10 people, then 20, then 30? The multiple testing conservativeness problem is a real trap, but they bypassed it because their hypotheses are nested. Think about it logically. If an algorithm trained on 500 people still can't beat the LLM, an algorithm trained on 10 people definitely couldn't either because the structure is perfectly nested like Russian dolls.
15:19They can stop the sequential testing at the very first sign of a statistical tie without triggering those false positive penalties. Oh, that's incredibly smart. It is. It's a very elegant theoretical breakthrough. Okay, so the math is rock solid. We've established how ESS works for predicting simple static facts. Do you own a home? Do you drink? But the research goes further. They push this metric into the much more complex world of causal inference, what economists call the conditional average treatment effect, or CAIT. Right, because predicting a static outcome based on a persona is one thing.
15:53Predicting a causal relationship, like how a specific intervention changes a behavior, is entirely different. Let's use an analogy to break down the conditional average treatment effect. It's the difference between predicting if it will rain today, which is standard ESS, You're just looking at the clouds and guessing the state of the world versus predicting exactly how much wetter the ground will get if we manually turn on a sprinkler instead of waiting for the rain. I love that. That what if scenario, the manual intervention of the sprinkler is the treatment effect. Exactly. In medical or economic research, evaluating that sprinkler is almost always done through randomized controlled trials.
16:32You give a treatment to one group, withhold it from a control group and measure the difference. Right. They wanted to know if we can use an LLM to estimate this effect without spending millions of dollars to run the actual trial. So how do you prompt the AI for a what if? Do you just tack a hypothetical onto the end of the character sheet? They did exactly that. They took that same persona, the 37-year-old architect, and added a crucial counterfactual scenario. The prompt literally says you are a subject in a randomized controlled trial in which the treatment is the provision of health insurance, specifically Medicaid.
17:08What would be the effect on your primary care visits? Hold on. This is where the whole scoring system seems to break down for me. How so? Well, to calculate the equivalent sample size, you need the A.I.'s guess, the M.L.'s guess and the true answer from the human data to see who is closer. But this is a counterfactual. We can't physically observe the same 37-year-old architect both with Medicaid and without Medicaid at the exact same time. We don't have a parallel universe to check the true answer. You can't calculate ESS without an answer key. You have hit on the hardest mathematical headache in the entire paper.
17:43How do you measure the accuracy of a counterfactual when the true individual effect is fundamentally unobservable? Right. So how did they do it? They solve this using an econometric technique called the transformed outcome trick. Okay, I'm going to need you to explain the transformed outcome trick because it sounds like literal statistical magic. It's not magic, but it requires an assumption called unconfoundedness. Unconfoundedness. Say that three times fast. Unconfoundedness basically means we have to assume there are no invisible hidden variables secretly driving the results. Like someone having a secret family history of illness that makes them desperately want to go to the doctor, but which isn't written down anywhere in their demographic data.
18:25Exactly. We have to assume the data we have is the whole picture. Once we assume that, we look at the actual outcomes we do see in the data, the people who really got Medicaid and the people who really didn't. Okay. Tracking with you so far. The trick is mathematically combining the outcome we observe for a person and dividing it by the probability that a person with their specific demographics would receive this treatment in the first place. This creates a new synthetic variable, which is the transformed outcome. Ah, I see. So you are basically taking the messy real world data and mathematically hammering it into a proxy score that represents the treatment effect.
19:03Yes. Even though the true what if for that specific individual remains unknown, this synthetic transformed outcome acts as a mathematically valid stand-in. That's brilliant. It is. They train the ML algorithm to predict this synthetic score. They have the LLM predict its counterfactual prompt. And then they compare the errors against the synthetic score. It allows them to apply the exact same ESS ruler to deep causal questions. So they built a mathematical bridge allowing us to evaluate AI not just as a static guesser of facts, but as a simulator of alternative realities. If you are a policymaker simulating a tax change or a business simulating a price hike, you now have a mathematical framework to test if your AI simulator is actually grounded in reality or if it is just hallucinating.
19:49It fundamentally expands the utility of the ESS metric. But the core lesson remains the same. Businesses and researchers are rushing to use AI to simulate customers and forecast trends. This research proves that you absolutely cannot trust it blindly. Right. For some variables, like structural demographics, the AI is a powerhouse worth 600 human surveys. But if you pivot slightly and ask it to predict a localized habit, it might be worth exactly zero. You have to figure out the equivalent sample size for your specific unique problem before you make any decisions based on an AI's output. It is not a magic eight ball.
20:22It is a tool with massive, highly specific blind spots. And that leads to a rather profound question to leave you with. Keep in mind, this research evaluated an older static model, Chad GPT-4, from 2023, and it found it was worth about 600 observations for certain demographics. But AI models are scaling exponentially. What happens when next generation models hit an ESS of 6 ,000 or 60 ,000? Are social scientists and businesses just going to stop polling humans entirely? If the ESS becomes high enough across all human behaviors, the financial incentive to run real expensive human surveys completely disappears.
20:59But if we stop asking real people about their lives and only rely on the AI's synthetic predictions whose version of reality are these models actually reflecting back at us? Wow. If the models are trained on past internet data and we stop collecting new human data because the models are getting good enough, we risk locking our understanding of human behavior into a simulated past. That is a genuinely chilling thought. We started with the dream of a researcher with no budget trying to simulate 500 people. It turns out that dream is statistically valid right now for certain structural things. But as that ESS ruler stretches longer and longer, we might just simulate ourselves right out of the equation entirely.
21:35Something to think about the next time you ask an AI for a prediction. Thank you.
From the publisher
This research paper introduces the equivalent sample size (ESS) as a novel metric to quantify the predictive value of Large Language Models (LLMs) compared to traditional human-provided data. The authors define ESS as the specific amount of domain-specific training data a machine learning algorithm requires to match the accuracy of a pretrained, fixed LLM. To estimate this value, they developed a statistical inference procedure utilizing block-out cross-validation to compare LLM performance against error curves of models like Random Forests and Lasso. Applying this method to the Panel Study of Income Dynamics, the study reveals that LLMs effectively substitute for hundreds of observations in tasks like predicting homeownership but provide negligible value for others, such as forecasting smoking behavior. Ultimately, the framework offers a standardized way for researchers to determine when an LLM can serve as a reliable surrogate for human data versus when traditional data collection remains essential.




