In short
Explains “WIMHF” (What’s In My Human Feedback), a method to make human preference data interpretable by discovering sparse, human-readable features behind which response wins in pairwise comparisons. It contrasts measurable vs expressed preferences, shows conflicts across datasets, and applies WIMHF to safety fixes and personalization.
Guests
No guest names or backgrounds are provided in the transcript; only the host(s) are speaking.
Key claims
Human preference signals are driven by a few sparse concepts (about four features per pair). WIMHF features recover 84% of decision accuracy vs dense embeddings and 67% of signal vs a full black-box reward model. Dataset generation strategy determines measurable preferences; mixing datasets can create contradictory signals (“digital schizophrenia”).
Notable examples
Emoji usage as a measurable feature; narrative prose vs lists (48% negative preference; most subjective, high variance); sustainability discussed when irrelevant (34% win-rate drop due to irrelevance, not ideology); “anti-refusal” in Elmarina/Chatbot Arena where refusals were dispreferred, leading to a safety crisis. Surgical label flipping for ~1,000 examples improved Reward Bench 2 safety from ~9% to 46% (+37%) without hurting helpfulness, and shifted arena leaderboard fairness. Personalization via style tuning used ~16 training conversations per annotator.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Human Preference Data
0:45 to 2:59
Exploration of how human preference data guides AI behavior and the issues of interpretability.
“So we've basically created this enormously powerful GPS for AI, but the map inside is totally blank.”
WIMHF: A New Approach
2:59 to 6:30
Introduction of WIMHF and its role in interpreting AI responses through sparse autoencoders.
“WMHF relies on something called sparse autoencoders, or SAEs.”
Measurable vs. Expressed Preferences
6:30 to 11:18
Distinction between measurable and expressed preferences in AI training and their implications.
“Measurable versus expressed preferences.”
Conflict Zones in AI Training
11:18 to 13:32
Examination of how mixed data sets create conflicting preferences in AI models.
“And that takes us straight to the most compelling practical application.”
Personalization and Safety with WIMHF
13:32 to 14:03
Discussion on how WIMHF enhances safety and personalization in AI by identifying subjective preferences.
“Okay, so moving beyond safety, WIMHF is also proving essential for personalization, specifically by identifying the most subjective preferences, areas where people really truly disagree.”
Understanding Variance in Personalization
14:03 to 14:59
Learn how personalization can be achieved without ideological biases.
“One person wants quick bullet points for efficiency.”
X-ray Vision into Reward Models
14:59 to 15:41
Discover how WIMHF reveals biases and enhances data curation.
“I mean, we started with the black box of the reward model, and now, using WIMHF and Sparse Autoencoders on that difference vector, we basically have X-ray vision.”
The Challenge of Global Alignment in AI
15:41 to 16:12
Explore the complexities of achieving consensus in AI responses.
“and maybe more importantly, how much control should we, the user, ultimately have over these fundamental stylistic choices in the AI we interact with every day.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. If you're interacting with any AI today, whether you know it or not, you are relying on a concept called alignment. Right. And this is really the foundational challenge of making these powerful large language models, LLMs, not just smart, but helpful, harmless, and, you know, most importantly, aligned with what we actually want them to do. And that alignment is built on human preference data. Yeah. It all comes down to that. We collect millions of pairs of responses. We show a person two options, say A and B, and just ask them, which one do you prefer? And that label gets fed into the system, creating what we call the reward model.
0:37But it's a black box. The model learns to chase this reward. But here's the problem. We don't really know why it starts exhibiting these strange, subtle behaviors like becoming overly confident or sycophantic. The preferences are just opaque. So we've basically created this enormously powerful GPS for AI, but the map inside is totally blank. We know the model got to the right destination, but we have absolutely no visibility into the thousands of tiny turns and decisions it made along the way. Well, that ends today. We are unlocking that black box using a brand new technique called WIMHF. That stands for What's in My Human Feedback.
1:16WIMHF fundamentally changes how we approach this whole problem of interpretability. Instead of, you know, manually guessing what features might matter, is it politeness? Is it detail? Is it humor? WIMHF automatically discovers the hidden features that distinguish responses and drive human choice. It just gives us the feature list after the fact. Okay, let's unpack this. I think a quick analogy will help anchor the whole deep dive here. Let's think about that classic internet debate. Is a hot dog a sandwich? Right. So suppose we give two AI responses to that question. Response A includes, let's say, a hot dog emoji and a thinking emoji.
1:51Response B is just plain text. When WIMHF looks at those two, it knows that they consistently differ along one clear axis, emoji usage. We call that the measurable preference. It's a feature that actually varies between the two options. But the critical question is, does the human judge actually care about the emoji? That's the key. If WIMHF finds that the preferred win rate is, say, 15 % higher for the response that uses emojis, that becomes the expressed preference. Exactly. WMHF separates what can be measured from what actually drives human choice. And that distinction is the key to principle alignment.
2:26It is. Absolutely. So for this deep dive, we're first going into the engine room, how WMHF works, using this fascinating concept called sparse autoencoders. Then we'll explore the deep, surprising conflicts WMHF reveals across seven major training data sets. And these conflicts are essentially giving our AIs digital schizophrenia. And finally, we'll look at two really practical high-impact applications, surgically fixing severely unsafe data and enabling highly data-efficient personalized AI experiences for you, the user. So let's get into the mechanics. WMHF relies on something called sparse autoencoders, or SAEs.
3:04Right, and we know autoencoders are generally used to compress data, but how does WMHF use them to discover these hidden human interpretable features? That sounds like a heavy lift. It's actually a very clever bit of engineering. WIMHF trains the SAE not on the text itself, but on the difference between the two candidate responses. The difference. Yeah, think of it like this. The AI represents every piece of text as this massive, complex numerical vector. It's called an embedding. For any given pair of responses, we just take the numerical representation of response A and literally subtract the representation of response B.
3:38So you're left with a single numerical value. A single difference vector, exactly. And that vector captures only what is unique and contrasting between the two options. It strips out the prompt, the general context, anything the two responses have in common. It isolates the variation. Precisely. And by feeding the SAE only this difference vector, we force the autoencoder to focus exclusively on those contrasting elements. The SAE is then designed to be sparse, which means when it reconstructs this difference, it has to do it using only a tiny handful of conceptual components. And those components are the features we can actually read, right?
4:12the human interpretable ones. Exactly. The result is a set of these fine-grained, sparse concepts. Things like answers directly without clarifying questions, or avoids legal disclaimers, or uses excessive markdown-style formatting. And on average, for any given input pair, only about four features actually light up. Only four. Just four. It's remarkable, really, because it confirms this intuition that even though text is complex, the reasons we prefer one response over another usually boil down to a few key differences. That sounds, I mean, almost deceptively simple. We're replacing these massive, dense reward models with a handful of sparse, readable concepts.
4:53But how accurate can these simple features actually be? If we're trying to decode a black box that controls trillions of parameters, I'm a little skeptical that just a few concepts can capture the whole picture. That's a crucial question, and the validation here is key. The research shows that these interpretable features capture a huge amount of the overall decision signal. They achieve 84 % of the prediction accuracy you'd get just from using the dense, complex text embeddings directly. 84 % is impressive. But here's the benchmark against the full black box. Relative to random chance, WIMHF's interpretable features capture 67 % of the signal achieved by a fully trained, state-of-the-art black box reward model.
5:3567%. So if we can explain two-thirds of human preference using maybe four or five features per decision. That is a monumental win for interpretability. It is. But let me push back a bit on that. If 33 % of the signal is still locked away in the black box, does that mean WIMHF is missing the really complex stuff, the ethical or highly abstract decisions, and only explaining surface-level things like tone and formatting? Well, that's what's so fascinating. WIMHF is highly effective at identifying the stylistic and topical drivers of preference, which it turns out often overwhelmed those truly abstract philosophical differences.
6:11That remaining 33 % likely includes the most nuanced context dependence and rare complex value judgments. But WMHF successfully proves that the majority of human choice, that's 67%, is driven by surprisingly simple identifiable factors. Now that we understand the mechanism, let's go back to those core concepts, Measurable versus expressed preferences. Let's focus on what the measurable side tells practitioners. Right. The measurable preference is key because it defines the limits of what a model can even learn from a specific data set. If you're running our hot dog experiment and every single response the AI generates uses emojis.
6:48Then there's no variation to judge. Exactly. The data set has zero measurable preference for or against emojis. So the features WIMHF discovers here are entirely dependent on how the responses were generated in the first place, the sampling strategy. So this gives crucial feedback on data set construction. It does. We looked at two data sets, PRISM and Community Alignment, which were generated using very different strategies. PRISM was built by taking inputs from 21 different LLMs using high temperature sampling. So high randomness, lots of variety. Tons of variety. And because of that, WIMHF found PRISM's measurable features centered around fundamental things like style, tone, and refusal behavior.
7:29You know, did the model answer or did it decline to answer a controversial topic like abortion? The features reflect that structural diversity from using many different models. And the contrast with the other data set. The Community Alignment Data Set, or CA, used a single LLM but explicitly prompted it for diverse values. So because the underlying engine was the same, WIMHF found that CA's measurable features reflected more specific topic diversity. So things like? Things like the difference between discussing sustainability issues versus recommending, say, luxury vacation spots. This is incredibly useful for you, the developer, because you don't want to waste time and money collecting labels if your response generation strategy is flawed.
8:07You can run WIMHF on your raw unlabeled data to check if your model is producing the right kind of variation before you pay for thousands of human labels. Right. And then moving to express preferences, WIMHF confirms some things we sort of already knew. Humans overwhelmingly prefer responses that are direct, on topic and structured. For instance, on that community alignment data set, the feature uses narrative prose, not lists, resulted in a massive decrease in win rate, a negative 48 % preference. Wow. So people really hate long paragraphs when they just want information. They really do. Nobody wants to read a novel when they need a recipe.
8:47But this brings us to what feels like the most dangerous part, the conflict zone. Developers frequently mix data sets to increase scale and robustness. But WMHF is showing that we are inadvertently creating contradictory signals for our models. We're basically giving the AI digital schizophrenia. We tested features across major data sets and saw preferences just flip direction dramatically. In crowdsourced, informal environments like Reddit and Elmarina, there is a strong express preference for informality, jokes, casual language. But in more curated or safety-oriented data sets. The preferences flip.
9:23Jokes and informal tone are actively dispreferred in HHRLHF, PRISM, and CA. So imagine mixing those. You train a model to be funny and informal half the time and professional and formal the other half. The model just learns to be confused. It's guessing the right tone based on noisy correlations. Exactly. And that's where you start to see dangerous outcomes like reward hacking. Where'd you see that? On the safety data set, HHRLHF, WIMMHF, pinpointed a consistent dispreference for responses that express uncertainty or, and this is crucial, ask clarifying questions. So the model is getting penalized for saying, I don't know, or I need more detail to answer that.
10:04It's being punished for responsible communication. Absolutely. And this puts immense pressure on the model to feign certainty, regardless of the prompt. This supports prior independent findings that training on HRLHF can increase overconfidence. It's like training a student to believe that saying, I don't know, is always a failing grade. That makes perfect sense. Now tell me about that strange finding in the community alignment data, the one about sustainability. Why would an emphasis on environmental issues be so strongly dispreferred? A 34 % drop in win rate sounds almost like an ideological rejection.
10:38Well, what's fascinating is that WIMHF helped confirm it wasn't an ideological rejection, it was a correlation issue. The feature emphasizing environmental sustainability was only active in response pairs, where the user's prompt had nothing to do with the environment. Oh, I see. So the annotators weren't rejecting sustainability, they were rejecting irrelevance. Exactly. But the model doesn't see that nuance, it just sees when this topic comes up, the response loses. And that's the risk generalization. The model internalizes this negative association, that sustainability is a losing feature, and might then actively suppress or avoid that topic even in future prompts where it's highly relevant and important to you.
11:17It shows how crucial WIMHF is for understanding why a preference exists, not just that it exists. And that takes us straight to the most compelling practical application. WIMHF didn't just find conflicts. It uncovered a genuine safety crisis hiding in one of the most widely used public data sets. Chatbot Arena or Elmarina? This was the biggest red flag in the entire study. The strongest express preference in Elmarina was a massive dispreference for all refusals of user requests. That led to a 31 % reduction in win rate. Wait, say that again? A dispreference for refusals? Yes. Annotators were consistently choosing the response that generated toxic, sexual, or otherwise harmful content over the response that correctly implemented safety guardrails.
12:04That's more than just a preference. That's actively training the model to be dangerous. The crowdsourced evaluation was systematically rewarding unsafe behavior. And this is where WIMHF moves from being a diagnostic tool to a surgical one. Because WIMHF identified the precise feature, this anti-refusal bias, it can pinpoint the specific data points that are misaligned. So you don't have to retrain everything or relabel thousands of points blindly. No, you just flip the labels for the top 1 ,000 examples identified by that single anti-refusal feature. And what happens when you feed that surgically clean data back into a model?
12:38The results are dramatic. Models trained on this safety-adjusted arena data see their safety accuracy on a standard benchmark. Reward Bench 2 just skyrocket by 37%. They go from essentially random, around 9%, up to 46%. Wow. And critically, this intervention preserves overall performance on helpful tasks. A 300 % gain in safety performance just from adjusting 1 ,000 data points based on one interpretable feature, that is transformative. And this adjustment fundamentally changes the widely reported LLM rankings on the arena leaderboard, doesn't it? It absolutely does. That safety bias penalizes models that refuse harmful requests, which gives an unfair advantage to less safe models.
13:18When we remove that bias using WIMHF, top performers who adhere to safety, like LOD 3.5 Sonnet, see significant jumps. they gain over 112 ELO points. So WIMHF is crucial for fair model evaluation too. It is. Okay, so moving beyond safety, WIMHF is also proving essential for personalization, specifically by identifying the most subjective preferences, areas where people really truly disagree. Right, we measure subjectivity by looking at the variance on a given feature, how much the preference differs among individual annotators. Wim HF found that the single most subjective preference in the entire community alignment data set was, again, that choice between narrative prose versus itemized lists.
14:02It had a really high variance score. So that confirms our earlier point. Style is the most divisive factor. One person wants quick bullet points for efficiency. Another person hates them because they feel robotic. Exactly. And because Wim HF separates style from, say, political or value judgments, we can engage in selective personalization. So we can confidently personalize a model's writing style. Without running the risk of tuning it toward ideological echo chambers. And that's the big concern when we talk about personal AI. Right. And how data efficient is this approach? Does personalization require thousands of conversations for every user?
14:36Not at all. By isolating and tuning only these formatting features, the stylistic differences, WMHF yields statistically significant gains in preference prediction, with only about 16 training conversations per annotator. 16. That's it. That's it. It makes fine-grained, controllable, and fast personalization entirely practical for mass deployment. This has been incredible. I mean, we started with the black box of the reward model, and now, using WIMHF and Sparse Autoencoders on that difference vector, we basically have X-ray vision. We've uncovered deep conflicts between popular data sets, exposed dangerous safety biases like the anti-refusal preference in Chatbot Arena, and demonstrated real utility in surgically curating bad data and enabling fine-grained, controlled personalization.
15:23Yeah, WIMHF gives practitioners the framework to step back from this chaotic process of just mixing data and hoping for the best. It moves us toward principal dataset construction and alignment, where we ensure our models are learning exactly what we intend them to learn, not just the noisy correlations or deeply entrenched biases hiding in that preference signal. And if the most subjective human preference is simply deciding between paragraphs and bullet points, a matter of presentation, what does that say about the inherent challenge of achieving a single global alignment where everyone agrees on what the best AI response looks like?
15:56and maybe more importantly, how much control should we, the user, ultimately have over these fundamental stylistic choices in the AI we interact with every day. That's something to chew on as these models become more embedded in our lives. Indeed. It puts the power back in the hands of the user. Thanks for joining us on the Deep Dive.
From the publisher
This paper introduces a method for automatically decoding hidden preferences from language model training data. By utilizing sparse autoencoders, the method translates complex text embeddings into a small set of interpretable features that explain why human annotators prefer one response over another. The research reveals that feedback datasets often contain conflicting signals, such as Reddit users favoring informal jokes while other groups disfavor them. Notably, the authors demonstrate that What’s In My Human Feedback? (WIMHF) can identify misaligned or unsafe preferences, such as a bias against model refusals in certain benchmarks. These discovered features allow developers to curate safer datasets by flipping harmful labels and to personalize model behavior based on specific user stylistic choices. Ultimately, the work provides a human-centered diagnostic tool to make the black-box process of model alignment more transparent and controllable.




