Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

3 Oct 2025 · 16 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How RLHF preference learning hides “hidden context” from annotators and, via its aggregation math, implicitly implements a social choice rule (Borda count), leading to strategic manipulability and jailbreak vulnerabilities; proposes distributional preference learning (DPL) to model disagreement and reduce risk.

Guests

No guest names or backgrounds are provided in the transcript.

Key claims

Standard RLHF preference learning doesn’t neutrally average preferences; it effectively runs a Borda-count-like rank aggregation. This favors compromise/least-offensive responses and can be gamed by strategic annotators. The resulting structure can create exploitable weaknesses that later enable jailbreaks.

Notable examples

Annotator choices influenced by mood, time pressure, or UI details (e.g., font color). Polarizing options where a “meh” consensus response wins under Borda count. DPL models bimodal ratings (e.g., peaks at 10 and 0) and reduces jailbreak vulnerability.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Hidden Context in RLHF

0:45 to 3:00

Exploration of hidden context in reinforcement learning from human feedback and its implications.

“It turns AI training from what we thought was pure optimization into something, well, something closer to political science, like running a flawed election almost.”

Varied Preferences and Hidden Context

3:00 to 5:00

Analysis of how varied human preferences lead to hidden context issues in AI training.

“Well, first, and maybe the most obvious one, is just varied preferences.”

The Political Science Connection

5:00 to 8:00

Comparison of AI preference learning to social choice functions and voting systems.

“Why the big leap to political science and voting theory?”

The Consequences of Board Account

8:00 to 11:00

Discussion on how board account voting systems impact AI training results and decision-making.

“Preference learning, using feedback from diverse humans, tacitly implements a social choice function.”

Jailbreak Vulnerabilities in AI

11:00 to 13:00

Examining how vulnerabilities in AI arise from the aggregation methods used in RLHF.

“It's saying your method creates a specific exploitable structure because of how it handles diverse opinions.”

Introducing Distributional Preference Learning

13:00 to 14:00

Introduction to distributional preference learning as a solution to the issues discussed.

“It stopped trying to pretend everyone agrees or that there's one simple utility function.”

Understanding RLHF Vulnerabilities

14:00 to 14:38

Learn how the standard RLHF method exposes vulnerabilities in AI models.

“Then we learned that the standard method, RLHF, doesn't just average this out.”

The Importance of AI Alignment

14:39 to 16:06

Explore the complexities of AI alignment as a social choice problem.

“So the really big takeaway for you listening to this is probably that getting AI alignment right, it's not just a coding challenge.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If you've used one of the big AI chatbots lately, you know the really sophisticated ones, you've probably noticed how well how aligned they feel helpful, mostly safe. Yeah, they've come a long way. They have. And they didn't get there just by, you know, scraping the whole Internet. A lot of that refinement comes from a technique called reinforcement learning from human feedback. RLHF. It's become the gold standard pretty much for alignment anyway. Exactly. But here's the twist. The very thing that makes it work so well, that human feedback loop, it also seems to introduce a, well, a fundamental instability.

0:36That kind of hidden problem. Yeah, researchers have put a name on it. They call it hidden context. Okay, let's unpack this. Let's do it. So our deep dive today is really about this hidden context, how it doesn't just add noise, but actually fundamentally changes the math of how these AIs learn. Right. It turns AI training from what we thought was pure optimization into something, well, something closer to political science, like running a flawed election almost. That's a great way to put it. And we're drawing directly from some really interesting research here, stuff that not only flags this mathematical issue, but actually offers a potential fix, a way to make these systems hopefully safer.

1:15It's a fascinating finding, honestly, because it does sort of undercut some common assumptions. You know, when you simplify preference learning, which is really the core of our LHF. Right. The part where humans rate the AI's answers. Exactly. On the surface, it seems solid. You get thousands of opinions, rank responses, tune the model. You assume it averages things out, finds the best policy overall. Makes sense. But the core finding in this paper is that unavoidable things in collecting that data, stuff that just happens when you ask humans for judgments. Like having a bad day or just clicking quickly.

1:48Precisely. That stuff doesn't just add a bit of random error. It introduces real fundamental mathematical problems. Okay. Meaning the AI model you end up with isn't just slightly off. It's actually optimized based on a mechanism, a rule that nobody really intended to put there. Okay. Let's nail down this hidden context idea first. So imagine you're an annotator, right? You're looking at two chatbot responses, A and B. You pick A, say it's better, the system notes, A, B. Simple enough. But what if you picked A because response B used a font color you just find annoying? Or maybe you saw A first, the phone rang, you came back and just clicked A to get it done fast.

2:31Right, those behind-the-scenes factors. Exactly. Your mood, maybe your screen brightness, how much time you had, none of that gets recorded. But it absolutely affected your choice. And that's the hidden context. It's the data influencing the feedback that isn't explicitly captured in the training data itself. Got it. And when you scale that up, thousands of annotators, all with their own hidden context, that failure to capture the full picture, it creates some serious issues. The research actually breaks this down into three main types of hidden context that are particularly critical. What are they?

3:04Well, first, and maybe the most obvious one, is just varied preferences. People are different. Sure. They just are. They have diverse values, sometimes conflicting ones. What I think is a safe response, you might find overly cautious. What you find helpful, I might find too wordy. Right. So the AI is trying to please everyone, but everyone doesn't agree on what pleasing means. Exactly. And when it tries to find some compromise between, say, witty and factual or concise and thorough, those underlying preference differences become a hidden context the model has to somehow resolve. Okay. So varied tastes.

3:40That makes sense. What's the second type? The second is more about that internal human state you mentioned, the cognitive processes, processes, things that lead to what the paper calls seemingly irrational behavior. Ah, like the font color example or being tired. Precisely. Did you skip breakfast? Are you feeling rushed, just plain fatigued? This stuff is real. It affects the judgment. But the AI system, it's completely blind to it. Invisible data points. Totally invisible. And the third type is also really problematic. Issues with data combination. What does that mean? Combining data sets. Yeah.

4:11Imagine you have one set of labels collected with a focus on, say, making the AI really harmless and another set collected trying to make it super engaging and entertaining. OK. Now you mix that data together to train one single model. The difference in the original instructions, the goals of those labeling tasks, that becomes another layer of hidden context. The AI is forced to somehow average out, be harmless and be entertaining. I see. So it's not just that people are messy. It's that the system is trying to find a single best answer based on input that came from different questions, different moods, different values.

4:49Exactly. It treats all this incredibly diverse, messy input as if it all came from one consistent source trying to maximize one clear goal. But it still feels a bit like, OK, it's messy data. Isn't that always a problem in machine learning? Data quality. Why the big leap to political science and voting theory? That seems dramatic. dramatic. It does seem like a leap at first, but that's the core insight here. And here's where it gets really interesting, because when you take all that diverse human feedback, all those preferences, and you try to combine them to pick one best AI behavior, you are fundamentally performing what's called a social choice function.

5:23You're running an election. You are absolutely running an election, an election to decide the winning AI policy or response style. Hang on, though. If we're using human preferences, aren't we supposed to be aggregating them? Isn't that the whole point of RLHF? Why is it such a shocker that it acts like social choice? Ah, good question. The shock isn't that it aggregates. It's how it aggregates. What kind of election is it running? Right. See, the central really surprising proof in this research is that the standard way we do preference learning in RLHF, it doesn't just average things out neutrally.

5:57It doesn't find the response that maximizes overall happiness or utility, like you might expect. Okay. So what does it do? It implicitly forces all that hidden context, all those varied preferences through the rules of one specific, very well-known and actually quite controversial voting system. And that system is, don't tell me. The board account. Board account. OK, I vaguely remember that from a Poli Sci class ages ago. Refresh my memory. Why is finding that inside our AI training loop a big deal? Right. So board account is a rank system. Let's say you have options A, B, and C. Your top choice gets, say, three points.

6:31Your second gets two. you last gets one point everyone ranks them you add up the points highest score wins okay sounds democratic ish ish is the word yeah it's different from say just picking your top choice it tends to favor consensus candidates the big issue is that board account can produce some really counter intuitive results how so well imagine an option let's call it response a that nobody loves everyone ranks at second or third just hmm okay the lukewarm option exactly now imagine response b is ranked first by, say, 40 % of people they love it. But the other 60 % absolutely hate it, rank it dead last.

7:07Right. Polarizing. Very. In a board account, guess who often wins? Let me guess. Response A, the meh one. Often, yes. Because it avoids those last place votes. It picks up points consistently by being everyone's second or third choice. Even if zero people actually want it as their first choice, it's the ultimate compromise candidate sometimes. Ah. So developers think they're training the AI to be the best, like, maximizing genuine usefulness or appeal for many users. Right. Aiming for high utility. But because the underlying math is accidentally running a board account, they're actually training it to be the least offensive option.

7:45The one that generates the fewest strong objections, even if it's kind of bland or mediocre. You've got it. That's the core problem identified. Yeah. The system is aggregating values using this specific political rule, not maximizing expected utility. And if we connect this to the bigger picture, this analysis formally proves something crucial. Preference learning, using feedback from diverse humans, tacitly implements a social choice function. It's not a bug, it's a feature of the math, but it's a feature with serious political science implications baked right in. Wow, okay. So, every time we train one of these big models using standard RLHF, we're basically running an election, an unintentional one, using a voting system known for weird outcomes.

8:28That's a takeaway. And that sounds like it leads directly to the really scary part, the security implications. It does. Because if it's a social choice function, especially one like board account, well, political scientists have known for ages that these systems can be manipulated, right? Exactly. It introduces what they call strategic manipulability. Meaning? Meaning annotators, the people giving the feedback, now have an incentive to misreport their preferences. not tell the truth about what they like, but vote tactically to try and steer the final outcome closer to what they want. They learn to game the election.

9:02Precisely. Instead of saying, I honestly prefer A over B, they might think, hmm, I know the system uses board account. If I rank C last, even though I actually think it's okay, it might hurt the chances of B winning, which I really don't want. Wow. Okay. And this isn't just some theoretical problem about slightly skewed results. Not at all. This leads directly, the research argues, to concrete vulnerabilities in how RLHF is deployed. Specifically, they link this mathematical structure, the board account aggregation, directly to the very practical, very worrying problem of jailbreak vulnerability in LLMs.

9:36Okay, connect those dots for me. Yeah. How does gaming the board account during training let someone jailbreak the AI later? Think about safety alignment. That's just one set of preferences the model is trying to learn, right? Right. Be helpful, but also be harmless. Right. Now, if a small group of annotators, maybe even one really dedicated attacker who understands this board account thing is happening, realizes the system is manipulable. They can vote strategically during training. Yes. They can systematically misreport their preferences. Maybe they consistently rank safe responses lower than they actually feel they should be.

10:09Or they exaggerate their preference for certain kinds of risky outputs in specific contexts. To push the compromise outcome away from safety. Exactly. They push the border winner, the final learned policy, subtly away from strong safety constraints, maybe creating little weaknesses or loopholes, often tied to unusual or specific kinds of prompts. Ah, the edge cases. Right. Later, an attacker who knows where those strategically created weaknesses are, what kinds of prompts exploit them, can use those specific inputs to bypass the safety filters. That's the jailbreak. The vulnerability isn't just bad data.

10:45It stems from the mathematical properties of the aggregation rule being used. That is sobering. It's like finding out the voter machines have a known flaw that allows strategic voters to influence the outcome way more than they should. It's a very good analogy. So this research isn't just saying data is noisy. It's saying your method creates a specific exploitable structure because of how it handles diverse opinions. That's the key insight. It identifies the mechanism of failure. Okay, so So, pathway forward. If the standard approach has this baked-in flaw, how do we aggregate human preferences better without falling for these manipulations or weird compromises?

11:23Well, this is where the paper introduces its proposed solution. A new class of methods they call distributional preference learning, or DPL. DPL, okay. And the aim here is really to tackle these problems. The hidden context. The board account effect. The manipulation incentive head-on. What's fascinating here is how DPL changes what information we even try to capture from the humans. Oh, so. It moves away from just getting that single preference point, A is better than B, full stop. That forces everything into one number, losing all the nuance, right? Yeah. All the why. And that loss of nuance is what fuels the hidden context problem and lets board account take over.

12:01Right. It hides the disagreement. Exactly. So instead of forcing that single point, DPL tries to capture the underlying uncertainty, the variation, the distribution of opinions. Okay, practically what does that look like? How do you capture a distribution instead of a single vote? Well, the DPL methods discussed, they work by estimating not just a single score for each option, but a distribution of possible score values. So if response A gets rated, DPL doesn't just record the average score, say, 5. It might record something showing, okay, the average is 5. But look, there's a big peak at 10 from users who loved it and another big peak at zero from users who hated it.

12:37A bimodal distribution, perhaps. Exactly. It shows the disagreement, the spread. It tells the whole story, not just the simplified average. And by capturing that distribution, the model can actually account for that hidden context, the varied preferences that the board account just kind of smooths over, often incorrectly. So it acknowledges the diversity instead of ignoring it. Precisely. It stopped trying to pretend everyone agrees or that there's one simple utility function. It models the full spectrum of opinion, warts and all. Does it work? Do we see better results? That's the encouraging part.

13:11The empirical results they show are pretty strong. Applying DPL methods to RLHF for training chatbots, it successfully identifies that hidden context in the data. Okay, good first step. But more importantly, it delivers on the practical side. Using DPL significantly reduces the chatbot's vulnerability to jailbreaking. It seems to close off those seams, those weaknesses that the board account method inadvertently creates. That's huge. So it's not just theoretically better. It actually makes the AI safer in practice. That's what the results indicate. Yes. It directly addresses the vulnerability stemming from that implicit social choice mechanism.

13:50The paper got attention at ICLR 2024, which suggests the community is taking this seriously. OK, let's try and synthesize this. So what does this all mean? We started with this idea that hidden context, different human values, moods, biases is basically unavoidable when you train AI with human feedback. Right. It's inherent. Then we learned that the standard method, RLHF, doesn't just average this out. It accidentally implements a specific voting system, the board account. Which has no mathematical properties. Including being vulnerable to strategic manipulation, this creates real security risks, like making AI models easier to jailbreak.

14:26And finally, there's a potential path forward. Distributional preference learning, BPL, which tries to capture the spread of opinions, not just a single average, to account for that hidden context and reduce vulnerability. That sums it up nicely. So the really big takeaway for you listening to this is probably that getting AI alignment right, it's not just a coding challenge. It's not just about bigger models or more data. No, it's deeper than that. It's fundamentally inherently a social choice problem. How do we aggregate diverse, messy, sometimes conflicting human values into a single coherent policy?

15:02The rules we choose for that aggregation matter immensely. They have direct consequences for safety, for reliability, just like the rules of voting matter in a democracy. Absolutely. And that parallel opens up some really profound questions, doesn't it, about ethics? About governance. Well, this research basically says, look, your current training method is a voting system. It happens to be board account, one known to be gameable. Now, if the people building these powerful AI systems are even unintentionally implementing a specific social choice function. A function with known flaws. Then maybe the focus needs to shift.

15:37Maybe AI governance shouldn't just be about technical fixes after the fact, trying to patch the vulnerabilities created by the aggregation method. maybe we need to think more deliberately, more ethically, about the design of those aggregation rules themselves. Like consciously choosing the voting system we use to train our AI, understanding its properties and biases. Exactly. Should we be explicitly designing the voting rules, the social choice functions that shape AI behavior, rather than just discovering them by accident? That's a heavy question, something to definitely mull over next time you get a strangely agreeable answer from a chatbot.

From the publisher

The paper "Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF,"** was submitted to **arXiv.org** and presented at ICLR 2024. The paper, authored by Siththaranjan, Laidlaw, and Hadfield-Menell, addresses the challenge of **"hidden context"** in preference learning, particularly in **Reinforcement Learning from Human Feedback (RLHF)**, where unrepresented data can skew model training. The authors **prove that standard RLHF methods** implicitly aggregate preferences using the **Borda count voting rule**, which can lead to counter-intuitive results and vulnerabilities like incentives for annotators to misreport their preferences. To mitigate these issues, they introduce **Distributional Preference Learning (DPL)**, a new class of methods shown to reduce jailbreak vulnerability in large language models. The source also contains a brief **system message** confirming the completion of a scheduled database maintenance.

More from Best AI papers explained

All 475 episodes
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHFBest AI papers explained · 16 min
Listen in VO