Learning-to-measure: in-context active feature acquisition

19 Oct 2025 · 16 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Learning-to-measure (L2M) for active feature acquisition, aiming to learn a general policy that decides which next data feature to measure to reduce prediction error while accounting for real-world data costs and missingness.

Guest backgrounds

No guests are mentioned in the transcript; it’s a solo “Deep Dive” discussion.

Key claims

Real data acquisition has tradeoffs (cost, time, risk). Standard AFA fails with systematic retrospective missingness and doesn’t scale across tasks. L2M frames feature acquisition as an in-context, sequence-based decision problem and uses uncertainty to choose features; it uses a surrogate objective that matches conditional mutual information choices (Theorem 4.2).

Notable examples

Medicine—CT biopsy vs blood test; clinical tasks on MIMIC-IV (LOS, mortality, readmission). Benchmarks include Gaussian processes, MedPabric, Miniboone, and MNIST pixel-by-pixel acquisition.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Active Feature Acquisition

0:45 to 1:30

Exploring the concept of active feature acquisition in machine learning.

“And this whole process, deciding which bit of information to go after next, which feature to acquire, that's what we call active feature acquisition, or AFA.”

The Meta AFA Problem and L2M

1:30 to 2:38

Introducing the Learning to Measure framework and its scalability.

“Because, you know, traditional AFA methods, they sounded good on paper, but they ran into some pretty serious walls when you tried to apply them to, say, real clinical data.”

Challenges in Traditional AFA

2:38 to 4:14

Discussing the limitations of traditional AFA methods in clinical settings.

“You end up just learning the old bias strategy instead of finding the truly most informative next step.”

Innovations of L2M

4:14 to 6:12

How L2M addresses key issues in AFA and improves feature acquisition.

“not just for one task, but across a whole distribution of tasks.”

Practical Strategies of L2M

6:12 to 8:01

Examining the surrogate objective and its implications for feature selection.

“Okay, let's dig into that decision-making process a bit more.”

Empirical Validation of L2M

8:01 to 9:21

Reviewing the testing phase and outcomes of L2M in various applications.

“And positivity, basically, you need enough historical examples of every possible feature acquisition choice being made in various contexts.”

L2M's Performance Against Traditional Methods

9:21 to 13:01

Comparing L2M's effectiveness with traditional task-specific agents.

“Choosing a feature, acquire test A or acquire test B, that sounds like a discrete yes-no decision.”

Limitations and Future Directions

13:01 to 14:00

Discussing the limitations of L2M and potential areas for future research.

“So bringing it back to the clinic, on those MIX tasks like predicting mortality or length of stay, did the adaptivity of L2M actually help?”

Limitations and Future Challenges in AFA

14:00 to 15:25

Explores the limitations of the current AFA framework and outlines future challenges.

“It tackles that huge scalability bottleneck by using pre-trained sequence models, pushing AFA much closer to that foundation model paradigm we talked about.”

Potential Impact on Medical Diagnostics

15:25 to 16:03

Discusses the transformative potential of AI in medical diagnostics and personalized inquiry.

“Still, even with those future steps, the potential here feels significant.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. Today we're diving into, well, a bit of a painful reality check for anyone actually using machine learning out there in the messy real world. And that's the fact that data, it's not free. It's often expensive, time-consuming, and sometimes even risky to get your hands on. Absolutely. Think about medicine, for instance. Yeah, exactly. Yeah. Imagine you've built this powerful model, maybe predicting patient outcomes. Your model says, hey, feature X gives a slightly better prediction. But feature X means doing an invasive biopsy. You know, expensive, takes time, not without risk.

0:35Meanwhile, feature Y, maybe just a basic blood test, gives you a pretty good result almost immediately, and it's cheap. Right. And standard machine learning models. They often just assume you have all the data all the time for free. The real world just isn't like that. It's all about tradeoffs. And this whole process, deciding which bit of information to go after next, which feature to acquire, that's what we call active feature acquisition, or AFA. That's the core idea. So our mission today is to explore this really fascinating new approach called learning to measure or L2M for short. And what's really exciting here is that L2M doesn't just tackle AFA for one specific problem.

1:13No, it goes bigger. Much bigger. It aims to solve what the researchers call the meta AFA problem. Essentially training a single general agent that learns the smartest way to acquire features across potentially dozens of different tasks all at the same time. And that scalability piece, that's really why L2M feels like such a step forward. Because, you know, traditional AFA methods, they sounded good on paper, but they ran into some pretty serious walls when you tried to apply them to, say, real clinical data. Messy, retrospective data, like electronic health records. Exactly. Two main bottlenecks, really.

1:48The first one, and it's a biggie, is what they term systematic retrospective missingness. Okay, systematic missingness. What does that mean in practice? Well, you look at historical patient records, right? Features are missing. But crucially, they aren't missing randomly. There was usually a reason a certain test wasn't run. Like a clinical guideline. Precisely. Maybe a doctor only orders an expensive CT scan if the basic X-ray results look worrying. So the very fact that the CT data is missing for many patients tells you something it's linked to the results of the observed X-ray. Ah, so the missingness itself carries information, but it's biased by past decisions.

2:25Exactly right. And if you just try to use standard techniques like imputation to fill in those missing values, well, you're essentially training your AFA model on data that already reflects those potentially suboptimal historical choices. You end up just learning the old bias strategy instead of finding the truly most informative next step. You got it. Standard methods don't really calibrate on how informative a feature might be in context. They mostly just try to guess its likely value based on what's already there. That's a critical distinction. If the historical data shows feature Z is often missing when patients seem less sick, a naive model might think Z is unimportant, when actually Z was just never measured because the doctor didn't think it was necessary based on other signs.

3:09That's the exact kind of bias L2M is designed to handle. And then there was the second major headache, the single task limitation. Right. So you'd build an AFA agent just for predicting, say, how long a patient might stay in the hospital. Yep, length of stay. And then if you wanted another agent to predict, maybe the risk of mortality or readmission. You had to start all over again, new training, new data. Completely from scratch. Yeah. Which, you know, takes time, costs money, needs specific labeled data for that new task. It just doesn't scale. And it really flies in the face of this whole foundation model idea that's becoming so popular, where you want one big model adaptable to many situations.

3:47Absolutely. L2M needed a way to generalize the whole process of asking questions, of acquiring information. Okay, so those are the challenges. This baked-in systematic bias and the lack of scalability. How does L2M actually structure itself to get around these and tackle this meta-AFA idea? Well, the core goal of meta-AFA is pretty straightforward conceptually. Learn a policy, a strategy that's good at acquiring features to reduce prediction error, not just for one task, but across a whole distribution of tasks. So learning the general skill of information gathering. Kind of, yeah. And L2M does this by taking a page out of the book of large language models, actually.

4:26It frames feature acquisition as an in-context decision problem. In-context, like how you give a prompt to GPT. Very similar idea. It happens in two main stages. First, there's pre-training. Here, they use a sequence model specifically, a modified transformer architecture. Why a sequence model? Because it's naturally suited to this problem. Think about it. You acquire features one by one. That's a sequence. A sequence model can handle inputs of varying lengths. Sometimes you've seen two features, sometimes ten. It learns patterns from these sequences. Okay, that makes sense. So this sequence model is pre-trained on lots of different tasks with lots of different patterns of missing data.

5:03And it gets really good at one crucial thing, figuring out how uncertain the prediction is, given only the features it's seen so far. Ah, the uncertainty quantification part. That sounds key. It's the absolute linchpin. If the model is already pretty confident about the outcome with the current data, why spend resources getting more? But if it's uncertain? Then it needs to know which specific feature will give it the biggest reduction in that uncertainty, the best bang for its buck. Exactly. And that robust uncertainty estimate feeds directly into the second stage, meta-training. Here, a separate policy network is trained.

5:39Its job is to look at the current state, the current uncertainty, and greedily pick the next feature that promises the biggest drop in that uncertainty. And the scalability magic happens because this is all in context. That's the key innovation. Because they're using these powerful pre-trained sequence models, you don't need to retrain everything for a brand new task. You just feed the initial observations for your new task into the existing pre-trained model as the context. And it adapts its acquisition strategy on the fly. Precisely. It uses its general knowledge about uncertainty and information gain, learned across all those pre-training tasks, and applies it in the context of this new problem.

6:17No massive retraining needed. It generalizes the how to ask. Very clever. Okay, let's dig into that decision-making process a bit more. How does the policy network actually choose the best feature? What's the ideal mathematical strategy? Right. So the theoretical ideal comes from a field called Bayesian experimental design. The perfect greedy strategy involves maximizing something called the conditional mutual information or CMI. Conditional mutual information. Sounds complicated. It sort of is, but the intuition is simpler. Imagine you have several features you could acquire next. The CMI basically asks, which one of these, if I learned its value, would give me the biggest expected reduction in my uncertainty about the final outcome I'm trying to predict?

7:02So it's quantifying the expected information gain from each possible feature. Exactly. You want the feature that offers the maximum potential surprise, the biggest decrease in what you don't know. Okay, that makes perfect sense. Maximize information, minimize uncertainty. But if CMI is the gold standard, why don't they just calculate that directly? Why build this complex L2M system? Ah, well, that's where the missy reality bites again. Calculating CMI accurately is really hard, especially with the kind of retrospective data we talked about, the data with all that systematic missingness. Because the assumptions break down.

7:37Precisely. To even get a reasonable estimate of CMI from that kind of historical data, you have to make some very strong, often untestable assumptions. Things like missing at random, MAR. Which means you assume that any systematic reasons for missingness can be fully explained just by looking at the features you did observe. Yeah, which is a huge leap of faith in many real world cases. You also need assumptions like exclusion restriction, meaning the act of measuring a feature doesn't somehow directly change the outcome itself. And positivity, basically, you need enough historical examples of every possible feature acquisition choice being made in various contexts.

8:14Okay, so calculating true CMI is often intractable or at least relies on shaky ground. L2M needed a practical alternative, a kind of shortcut. Exactly. And they found a really smart one. They used what they call a surrogate objective. Instead of trying to directly maximize the complex, hard-to-estimate CMI, they trained the policy network to do something much more concrete. Minimize the expected prediction error after hypothetically acquiring a potential next feature. Oh, I see. So instead of asking how much information will feature A give me, it asks if I acquire feature A now, how much lower will my prediction error be on the very next step?

8:50Precisely. It's a one-step-ahead prediction quality check. And the really neat part is the theory specifically, they prove in theorem 4.2, that optimizing this much simpler practical surrogate objective actually leads the policy to choose the same action as if it had maximized the ideal CMI. Wow. Okay. That's an elegant way to bridge the gap between theory and practice. It really is. It turns a potentially impossible optimization problem into something computationally feasible. And one quick technical point. Choosing a feature, acquire test A or acquire test B, that sounds like a discrete yes-no decision.

9:29How do they optimize that using standard machine learning tools, which usually like smooth continuous gradients? Good question. They use a pretty standard trick for this called the Gumbel Softmax relaxation. You can think of it like smoothing out the sharp edges of that discrete choice. Making it temporarily soft so gradients can flow through. Exactly. It turns the hard decision into a probabilistic one that's differentiable, allowing them to use powerful gradient-based optimizers like Atom to train the policy network effectively. Got it. Okay, theory covered. Clever hacks explained. Now for the payoff, did it actually work?

10:01Let's talk about the empirical results. Did L2M deliver when tested? Yeah, this is where it gets really compelling. They didn't just test it on one thing. They threw a diverse range of problems at it. Everything from synthetic tasks like Gaussian processes where they knew the ground truth to standard real-world tabular data sets, things like MedPabric for cancer data, Miniboon from particle physics. Okay, diverse benchmark. And even image tasks. They used MNIST, the handwritten digit data set, but framed it as acquiring features sequentially, like revealing blocks of pixels one by one to identify the digit.

10:37Interesting application. But what about the clinical side? You mentioned health records earlier. Right, and that's probably the most impactful validation. They use the MIMC4 dataset. For listeners who don't know it, MI is a huge, incredibly rich database of de-identified patient data from intensive care units, really complex, messy, real-world data. And what tasks did they try on MIM? Core clinical prediction tasks. predicting length of stay, LOS in the ICU, predicting patient mortality, and predicting unplanned read admission to the hospital. Really meaningful outcomes. And how did L2M stack up against other methods?

11:13They compared it against some strong task-specific AFA baselines methods like GDFS, DI, even deep reinforcement learning approaches like DQN, all trained specifically for each task. So comparing the generalist L2M against specialists. Exactly. And the findings were pretty striking. L2M consistently matched and often surpassed these specialized task-specific agents in terms of prediction accuracy. Even though the baselines were custom trained for just that one job. Yes. But maybe even more importantly, L2M consistently provided more reliable uncertainty estimates. The model wasn't just accurate.

11:48It had a better sense of when it was likely to be wrong. And for a doctor, using an AI system, knowing the model's confidence level is arguably as important as the prediction itself, right? Absolutely. Trust comes from calibrated confidence. And they measured this reliability using things like log loss and mean squared error. L2M consistently did better. Did the advantage change depending on how much data was acquired? It did. And in an interesting way. The performance gap often widened as more features were acquired sequentially. So, in more complex scenarios where you need multiple steps of acquisition, L2M's ability to make consistently good choices really shone through.

12:25It seemed more robust over longer decision horizons. That robustness sounds like a major win, especially in challenging data situations. It really was. They specifically tested scenarios with scarce labeled data, meaning the model had only seen a small number of fully observed examples as its initial context and scenarios with very high rates of retrospective missingness. mimicking really patchy historical records. And L2M held up better. Significantly better. This suggests that the meta-learning across diverse tasks really pays off when data is limited or messy. It learns a more general, robust acquisition strategy that doesn't overfit to the quirks of one specific data set's missingness patterns.

13:04So bringing it back to the clinic, on those MIX tasks like predicting mortality or length of stay, did the adaptivity of L2M actually help? Did it learn to skip unnecessary tests? Yes, definitely for the more complex tasks like LOS and mortality prediction. The ability to choose features instance by instance, based on the specific patient's data seen so far, was clearly beneficial. It learned, for example, not to acquire certain costly features if cheaper ones already gave a confident prediction. Which translates directly to potential cost savings and faster decisions. Exactly. Interestingly, though, for the somewhat simpler task of readmission prediction, the adaptive nature of L2M offered less of an advantage over simpler policies, which is also informative.

13:49Yeah, it shows the model isn't just being complex for complexity's sake. It learns when that fine-grained adaptivity is actually needed. Precisely. Okay, so wrapping up this deep dive, the main takeaway seems to be that L2M offers a really promising path towards generalizable, reliable, and scalable active feature acquisition. It tackles that huge scalability bottleneck by using pre-trained sequence models, pushing AFA much closer to that foundation model paradigm we talked about. I think that's a fair summary. It's a strong step forward. Of course, the paper is also good about pointing out the remaining limitations and where future work needs to go.

14:23Right. What are the key ones? Well, the big one, the elephant in the room, is still the reliance on that missing-at-random MAR assumption. Because you can't truly test MAR in practice on real data. You need ways to diagnose when it might be badly violated. Exactly. Developing robust diagnostics for MA violations is a critical next step for real-world deployment. Another major area is scaling the pre-training. They showed proof of concept, but getting this to work across truly massive, diverse data sets from different domains is the next engineering challenge. Makes sense. And what about applying this in real time?

15:00Ah, yes. That's a key limitation currently. The framework presented is for time and variant settings. It assumes the features, once measured, don't change. But in many real-world scenarios, like an ICU. Things are constantly changing. Vital signs, lab results. Exactly. So extending L2M to handle these time-varying dynamics is crucial future work before it could be used for, say, continuous patient monitoring and decision support. Still, even with those future steps, the potential here feels significant. Stepping back. If this kind of technology matures and scales, what does it really mean? What does it mean for, say, the future of medical diagnostics if we have an AI that can tell us, moment by moment, the single most informative and cost-effective test to run for this specific patient, for this specific question?

15:47It could fundamentally change how we approach information gathering, couldn't it? Yeah. Moving away from fixed protocols towards truly personalized, efficient inquiry. Making sure we never waste time, resources, or patient discomfort collecting data that just doesn't add value. A fascinating thought to end on.

From the publisher

This paper introduces Learning-to-Measure (L2M) to address the challenges of meta-Active Feature Acquisition (meta-AFA), a sequential decision-making problem. Traditional AFA methods often struggle with scalability because they are designed for a single task and fail when trained on retrospective data containing systematic missingness in features. L2M overcomes these limitations by formalizing the meta-AFA problem to allow learning acquisition policies across diverse tasks and leveraging a pre-trained sequence-modeling or autoregressive approach to provide reliable uncertainty quantification. By coupling this uncertainty quantification with a greedy policy that maximizes conditional mutual information, L2M can select the next feature to acquire in-context without requiring retraining for every new task, demonstrating superior performance, especially when labeled data is scarce or missingness is high.

More from Best AI papers explained

All 475 episodes
Learning-to-measure: in-context active feature acquisitionBest AI papers explained · 16 min
Listen in VO