From Unstructured Data to Demand Counterfactuals: Theory and Practice

14 Jan 2026 · 14 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Demand estimation for differentiated products using unstructured-data proxies (images/text) in random-coefficients logit/BLP-style models; proposes a post-estimation bias correction plus proxy diagnostics to recover valid counterfactual substitution patterns.

Guest backgrounds

Not specified in the transcript (no names, roles, or prior work mentioned).

Key claims

ML embeddings/proxies can systematically miss latent product attributes, causing model misspecification and biased parameters/counterfactuals. A computationally light correction reparameterizes with a composite parameter (gamma) to absorb proxy mismeasurement, enabling efficient inference (closed-form standard errors) even with data-dependent proxies. Two Lagrange-multiplier diagnostics (LM1, LM2) assess proxy quality and embedding dimension adequacy.

Notable examples

E-book case study with observed “second choice” as ground-truth counterfactual; review-data proxies improved closest-substitute prediction from ~40% to 60–70% after correction; LM1/LM2 guided better proxy selection.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Complexity of Consumer Choices

0:46 to 1:54

Discussing how consumer preferences are influenced by unstructured data.

“If option A, your favorite car, is suddenly gone, which option are you, the consumer, going to pivot to?”

Understanding Proxies in Economic Models

1:55 to 3:19

Explores the importance of proxies and their potential pitfalls in models.

“When those proxies, no matter how sophisticated the ML model is, fail to capture the true, subtle things driving our choices.”

The Dangers of Flawed Proxies

3:20 to 5:46

Examining how poorly measured proxies can disrupt predictions in models.

“You can have the most elegant model structure in the world, but if your inputs, your proxies are core, the whole thing is corrupted.”

Introducing a Correction Toolkit

5:47 to 8:43

Introducing a toolkit for correcting biases in economic models.

“You've just wasted the entire effort of being sophisticated.”

Case Study: E-books and Proxy Accuracy

8:44 to 11:19

Real-world example showcasing the effectiveness of the correction toolkit.

“This correction handles that dependency without any issue.”

Conclusion: Ensuring Accurate Predictions

11:20 to 13:30

Summarizing the importance of accurate proxies for reliable economic predictions.

“the accuracy for predicting the closest substitute just jumped.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You know, if you've ever had to analyze a market, I mean whether you're comparing say electric vehicles or sizing up subscription services. Or even something like choosing a school for your kid. Exactly. You're wrestling with what is really the central problem of modern economic and social modeling. And that problem is how to accurately estimate consumer demand for what we call differentiated products. And this isn't just some academic exercise. That core estimation is the absolute backbone for studying these huge high stakes issues. Things like the consumer impact of a corporate merger or forecasting a new product launch.

0:38Right, or even trade policy. So our mission at the end of the day is to map out these substitution patterns. We want to know the counterfactual. The what if. The what if. If option A, your favorite car, is suddenly gone, which option are you, the consumer, going to pivot to? So we're kind of like detectives, and our evidence is the product attributes. For a long time, that evidence was pretty straightforward. It was neat numerical variables, price, horsepower, miles per gallon. Sugar content in your cereal. Exactly. Quantifiable, objective stuff. But the reality of what drives our choices today is, well, it's a lot more subtle.

1:13It's subjective. It's really hard to quantify. We're talking about visual design, the feel of an app's interface. The vibe of a brand. The vibe. How do you put vibe into a spreadsheet? You can't. You can't. So researchers have necessarily pivoted. They're now using unstructured data, huge amounts of product images, text from reviews, forum discussions, and using machine learning to synthesize all that messy information. And these ML models, they then compress all this complexity down into simple, low-dimensional numerical variables. These are the famous embeddings. Embeddings, or, and this is the term we'll probably use more today, proxies.

1:51They're a proxy for the real thing. And here's the risk. This is what we're diving into. When those proxies, no matter how sophisticated the ML model is, fail to capture the true, subtle things driving our choices. Then the whole model falls apart. The predictions become invalid, systematically biased. So today we are exploring a new, computationally light toolkit designed to fix this bias after the fact, ensuring all the work you put into a sophisticated model doesn't just get quietly erased by bad inputs. Let's really unpack this problem of the proxy. Okay, so what's essential to understand here is the kind of model we're trying to save.

2:28Think of things like the famous BLP framework, which uses what are called random coefficients logit models. They're designed for one specific, very important job. And that's capturing preference heterogeneity, right? That's the big leap from a simple model to a sophisticated one. Precisely. A simple model just assumes that if, say, a car's price goes down, everyone's preference shifts in the same direction. But a random coefficients model gets that some of us really care about horsepower, while others are all about the design or fuel efficiency. So it tries to capture all those different flavors of consumer taste.

3:02Exactly. It allows for much more complex, realistic substitution patterns. Okay, so if I use a proxy for style that's poorly measured, the model can no longer tell if I chose a car because I have a high preference for style, or for some other random reason, the whole mechanism just breaks. That's the danger. You can have the most elegant model structure in the world, but if your inputs, your proxies are core, the whole thing is corrupted. It's a really insidious kind of flaw. How is this different from, say, a standard measurement error, like a typo in the price data for one car? Why is this product-level flaw so much harder to fix?

3:39A standard measurement error, you know, it often averages out or maybe it only affects a single observation. But here, the flaw is baked into the product's very description in the model. Think of it this way. Unstructured data, like a high-res product photo, is immensely complex. Turning that into a low-dimensional embedding is like, well, it's like taking that beautiful photo and compressing it into a tiny, low-quality thumbnail. Right. The thumbnail gives you the basic shape, but you lose all the texture, the subtle colors, the design nuances that are, scientifically speaking, the true latent attributes driving your preference.

4:12And if every single product's proxy is a flawed thumbnail, that loss of detail is systematic across your entire dataset. It creates this persistent mismeasurement of differentiation. So the paper calls this model misspecification. It is. The parameters that describe how features affect consumer choice become biased. And at the end of the day, that means you get biased parameters and completely useless counterfactual predictions. And this is all made worse by the fact that we're using machine learning, right? These proxies come from these black box models, and we can't really justify assumptions about how the ML model messed up.

4:49We just know it did. That's why the traditional fixes don't work. You need a solution that is totally agnostic about where the mismeasurement came from. It just needs to know it's there and how to correct for its effect on the final prediction. This brings us to the real cost of getting this wrong. The simulation findings you mentioned, they sound like a huge warning sign for anyone using these advanced models. It's the ultimate failure of sophistication. It really is. When those proxies get too noisy, when your thumbnails are just too blurry to tell products apart, the model's internal machinery just starts to collapse.

5:19So instead of capturing those rich, complex substitution patterns... It reverts. It starts to fall back to the restrictive, simplistic predictions of a basic logic model. Wow. So you spent all the computational power, the time, the brain power to build this incredible engine to capture human choice. And the low quality of your fuel, the input data, just quietly nullifies the entire effort. You end up with the simple, inaccurate answer you were trying so hard to avoid. You might as well have used a much simpler model from the start. You've just wasted the entire effort of being sophisticated. That's why a fix is so critical.

5:53Okay, let's pivot to that solution. This toolkit, it promises a practical low effort repair. It almost sounds too good to be true. Well, the elegance is in its approach. It's a post estimation bias correction. So you're not rebuilding the whole model or retraining the proxies. I'm thinking of it like a calibration. You know, if you have a high-tech thermometer that you know is always off by some weird complex factor, you don't throw it out. Right. You take its reading, the naive estimator, and you just apply a known, carefully calculated correction term to get the real temperature. That is a perfect analogy for what it does to the predicted counterfactuals.

6:30The core insight involves reparameterizing the model with something they call a composite parameter, which they denote as gamma. So what is this gamma actually doing in terms of, you know, consumer utility? Okay, so the model is normally trying to estimate how an attribute, let's say style, affects a consumer's utility. Since the style proxy is flawed, the model gets that relationship wrong. This composite parameter, gamma, it steps in and captures the combined effect of both the true structural parameter and the mismeasurement error from that proxy. Okay, so instead of trying to perfectly fix the flawed input, we just estimate the total effect of that flawed input on utility, and then we adjust the prediction based on that relationship.

7:12Precisely. By absorbing the error into this gamma parameter, the correction method can stay agnostic about the black box mechanics of the ML proxy. It just needs to see how the flawed inputs and the model structure interact to produce the bias, and then it can correct the predictions. That flexibility seems absolutely critical. But the research also emphasizes that this isn't just about removing bias. The corrected estimator is designed to be efficient. What does that actually mean for a practitioner? Efficiency means precision. It means confidence. If you think about your estimates as darts thrown at a target, efficiency or having the lowest possible asymptotic variance means your darts are clustered as tightly as possible around the bullseye.

7:55You get more statistical power. I could be just as confident with maybe a smaller data set. In a sense, yes. It means your estimates are as precise as they can possibly be. The corrected estimator is designed to maintain that high precision, even if your initial model parameters were estimated inefficiently because of the whole complicated setup. And the practical benefits here seem huge. No complex, time-consuming bootstrapping to get standard errors. None. It's computationally light. You get simple, closed-form expressions for your standard errors. This means valid inference is just there. It saves an enormous amount of time and resources.

8:32And it also accommodates proxies that are data dependent, which is a big deal. So if I fine tune my ML model on the same choice data I'm using for my demand model. Which is a very common practice to get a better proxy. This correction handles that dependency without any issue. Exactly. It makes sure that your process of improving your inputs doesn't accidentally invalidate your statistics down the line. Yeah. You can use all your data to make the best possible proxy and then trust the correction to stabilize the results. The fix is elegant, I'll give you that. But if the true attributes are latent, are unobservable, how do we choose the best inputs to begin with?

9:09If I have 10 ML models giving me 10 different proxies, how do I vet them? That's where the second part of the toolkit comes in, the diagnostics. Since you can't see the truth, the toolkit gives you two simple LM, that's Lagrange multiplier statistics, that act like litmus tests for your proxy quality. Okay, let's take Diagnostic 1, LM1. What's the specific question that's answering for a researcher? LM1 asks simply, is this proxy set good enough to even use? It basically assesses how close the parameters you get from the proxy are to the true latent relationship. It's a measure of how much pollution the proxy is adding to your model.

9:49So if I'm comparing a proxy from text reviews versus one from product images, I can run both through this diagnostic. And you just choose the one with the lower LM1 score. It gives you concrete, measurable feedback that guides you toward the proxy that's the strongest approximation of what consumers actually care about. It's your garbage in, garbage out detector. And diagnostic 2 LM2, this one sounds like it's about whether we've compressed the data too much. It is. It's about dimension. It asks, is the number of attributes I'm using correct? Often you use a technique like PCA to shrink a massive embedding down to just a few variables.

10:21But how many do you keep? Keep too few, and you throw away the nuance that actually explains the substitution patterns. Right. So LM2 tests the hypothesis that the dimension you chose is adequate. If the LM2 statistic comes back with a rejection, it's telling you your proxy is too restrictive. The real substitution patterns are happening in a higher dimension than your two or three proxies can capture. And to really bring this home, the researchers showed how powerful this is with a really compelling real-world example using e-books. This case study is the payoff, really. They had experimental data where consumers made a first choice, and then after that choice was removed, they were immediately asked to make a second choice.

10:59So that second choice is the perfect ground truth for the counterfactual. It tells you exactly what the closest substitute is. It's the real answer. So they could test how well the model predicted that second choice. both with the naive biased estimates and then with their corrected estimates. And the results were... Dramatic. Dramatic. For the specifications using review data, which is incredibly complex, messy data, the accuracy for predicting the closest substitute just jumped. The uncorrected model using the naive proxies was accurate only 40 % of the time. Barely better than a coin flip. Pretty much.

11:34But once they applied the simple bias correction, the prediction accuracy shot up to between 60 % and 70%. That's a 30-point leap. That changes everything. That's the difference between giving a business useless advice and giving them trustworthy forecasts that could drive multi-million dollar decisions. And the diagnostics worked in this case. Perfectly. The LM1 diagnostic correctly pointed them to the best performing specifications. It successfully guided them to the proxy sets that minimize contamination and maximized accuracy. It's proof that the vetting tools actually work in practice. So let's synthesize this.

12:10We've established that relying on these very sophisticated demand models is, well, it's useless if your inputs, these proxies from unstructured data, are flawed. And we've introduced a simple post-estimation toolkit that fixes this with minimal computational cost. And we've shown that it ensures the investment you make in that complexity actually pays off with reliable, precise predictions. And this works across the board market level data, individual choice data, even standard numeric data if you suspect mismeasurement. Which brings us back to that really terrifying simulation warning. If you use poor proxies and you don't apply these corrections, you risk spending enormous resources on a complex model.

12:48Only for the measurement flaw to quietly override all your hard work. And it pushes your results right back to the simple, inaccurate predictions the model was supposed to save you from in the first place. It is the subtle trap of modern econometrics. Complexity without correction just gives you simplicity that's probably wrong. This toolkit makes sure the complexity works for you, not against you. So think about the products or services in your field. Are your current proxies capturing the real secret sauce of differentiation, the nuances consumers truly value? Or are you just relying on a low-quality blurry thumbnail?

13:24Because the success of your next big decision might just depend on answering that question and then applying the right calibration.

From the publisher

This paper introduces computational toolkit designed to correct errors in demand estimation when using unstructured data, such as images or text, to represent products. Because researchers often use machine learning embeddings as proxies for true product attributes, these approximations can introduce statistical bias that leads to inaccurate predictions of consumer behavior. The authors propose a bias-correction method and diagnostic tests to ensure that these proxies adequately capture the dimensions of differentiation driving market choices. Their approach is efficient and lightweight, integrating easily into standard workflows without requiring extensive re-computation or optimization. Simulations and empirical applications demonstrate that this method significantly improves the accuracy of counterfactual predictions, such as how consumers react when a product is removed from the market. Ultimately, the research provides a rigorous framework for economists to use high-dimensional, complex data while maintaining the integrity of their structural models.

More from Best AI papers explained

All 475 episodes
From Unstructured Data to Demand Counterfactuals: Theory and PracticeBest AI papers explained · 14 min
Listen in VO