In short
Pluralistic AI alignment using Direct Preference Optimization when user preferences vary in unobserved ways; argues standard binary preference data (A vs B) is mathematically insufficient and proposes ternary preference collection plus a two-step training/aggregation pipeline.
Guest backgrounds
No guests mentioned; it’s a solo “Deep Dive” episode with no identifiable guest speakers.
Key claims
One-size-fits-all reward/policy assumptions in DPO-like methods fail under preference heterogeneity. Binary comparisons hide distinct latent preference types (non-identifiability). Identifiability requires ternary choices among three options.
Notable examples
Synthetic MPI personality dataset with two opposite groups P1 and P2 (P2 = −P1) that look identical under binary data but separable with ternary data using EMDPO. Global Opinion QA dataset (Britain, Indonesia, Mexico, Pakistan) where EMDPO clusters country-level viewpoints without explicit labels; MMRA aggregation yields zero positive regret for Britain/Indonesia/Pakistan and small regret for Mexico.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Challenge of Uniformity in AI
0:45 to 1:40
Discussing the limitations of assuming a single definition of 'good' in AI models.
“And if you try to force this uniform standard onto that diversity, what happens?”
The Problem of Binary Comparisons
1:40 to 2:48
Exploring how binary comparisons fail to capture diverse user preferences.
“Well, it mostly comes down to our reliance on binary comparisons.”
Understanding Non-Identifiability
2:48 to 3:46
Revealing how binary data structures obscure user preferences.
“They looked at this data set called MPI, which models personality traits, and they created two hypothetical groups.”
Introducing Ternary Preferences
3:46 to 4:24
The solution of using ternary preferences to better capture user diversity.
“Our standard data collection is actively obscuring the very diversity we need to capture.”
Framework for Diversity in AI Models
4:24 to 6:32
Detailing the proposed two-step pipeline for accommodating diverse preferences.
“It's a change in how we ask the question, not necessarily some massive algorithmic change initially.”
Step One: Expectation Maximization in Action
6:32 to 8:00
Explaining how EMDPO utilizes ternary data for preference discovery.
“Why not just, I don't know, average the outputs?”
Step Two: Fair Aggregation of Outputs
8:00 to 10:34
Describing how the MMRA ensures fairness in combining diverse models.
“They showed pretty strong validation, starting with that data requirement.”
Key Takeaways on AI Alignment
10:34 to 11:08
Summarizing the importance of shifting to ternary preferences for better alignment.
“So, to recap for you listening, the big takeaway seems to be, if you want alignment that respects diversity, you need to shift gears.”
Provocative Questions for the Future
11:08 to 12:12
Contemplating the implications of grouping human preferences in AI alignment.
“And this framework, EMDPO plus MMRA, gives you a way to do that pluralistic alignment without needing to estimate those often complex and sensitive explicit reward models.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. Today we're tackling a really fundamental challenge in AI alignment. It's this assumption that a large language model, you know, an LLM, should have just one definition of good. We're going to unpack some research that argues pretty convincingly, I think, why that whole one-size-fits-all idea just doesn't work in our diverse world. And maybe more importantly, what the surprisingly simple fix might be. Yeah, it's really about moving past that single model thinking. You know, standard alignment methods, even the newer ones like DPO, direct preference optimization, they kind of implicitly assume there's one reward function everyone agrees on.
0:38Yeah. And that all users have, well, basically the same preferences, like human feedback is just one voice. Right. But we know that's not true. People want different things, value different things. And if you try to force this uniform standard onto that diversity, what happens? You risk the model only listening to the biggest group, right? The majority, which, well, that sounds like a recipe for bias and just a worse experience for lots of people. Exactly. That's the core problem. So this deep dive is really focused on pluralistic alignment. How do we actually find and cater to these, let's call them unobserved preference heterogeneities?
1:11These hidden factors could be culture, politics, just personality that shape what someone thinks is a good AI response. We'll walk through a framework that tries to discover these hidden user types, train models for them, and then combine them fairly. But the starting point is maybe the most surprising bit, the data we usually collect. It might be fundamentally flawed for this task. And the fix involves just one more choice. Okay, well, saying the foundational data is broken, that's a big statement. Let's unpack that. What is the basic limitation of the standard preference data, the kind we use for LLMs right now?
1:44Well, it mostly comes down to our reliance on binary comparisons. You know, the setup. You're shown response A, response B, and you just pick the one you like better. This research connects that setup to working in conometrics, specifically something called the random coefficient logic model. And the conclusion is pretty stark. Binary comparisons, just A versus B, are provably insufficient to reliably figure out these underlying hidden diverse preferences. Hold on, provably insufficient? Binary choices everywhere, market research polls. Surely if you ask enough people A vs B, you map out the landscape.
2:22You'd think so, intuitively. But mathematically, when preferences aren't uniform, when they're heterogeneous? The answer is no. Even with like infinite data points, infinite annotators giving you AVSB choices, you still can't reliably recover the true distribution of those different underlying preferences. The math just doesn't allow it with only two choices. Okay, provably insufficient. That really shakes things up. The math itself is hiding the diversity. How do they actually demonstrate this? Is there an example? Yeah, they use a really clever sort of adversarial setup with synthetic data. They looked at this data set called MPI, which models personality traits, and they created two hypothetical groups.
2:58Let's call them P1 and P2. And the key thing is P2 is defined as the exact opposite of P1. Mathematically, P2 equals minus P1. So like if P1 loves cats and hates dogs, P2 hates cats and loves dogs. Total opposite. Exactly like that. Now, when you ask these two completely opposite groups to make binary choices, A versus B, something strange happens in the math. Their opposite preferences actually lead to the exact same overall probabilities of choosing A over B at the population level. They effectively cancel each other out. So if you only look at the average choice or how often A is picked over B in total, these two radically different groups look identical.
3:36The model can't tell them apart. Ah, I see. Non-identifiable, you called it. Precisely. It's non-identifiable. The binary comparison data structure literally hides their distinct preferences from view. It's like taking a picture where opposite colors average out to gray. Okay, that's the aha moment. Our standard data collection is actively obscuring the very diversity we need to capture. So if binary comparisons fail us here, what's the solution? It's almost anticlimactic in its simplicity. Yeah. Theoretically profound. Identifiability actually being able to reliably see that underlying preference distribution requires comparisons over three or more options.
4:13They call it ternary preferences. Instead of just A, V, S, B, you need to ask the user, which do you prefer among A, B, and C? Wow. Just adding one more option to the choice that unlocks the ability to see this complexity. That's surprisingly practical. It's a change in how we ask the question, not necessarily some massive algorithmic change initially. Okay, so now we have this richer ternary data, but we need a system to actually use it, right, to build these diverse aware models. How does that work? So the framework proposes a two-step pipeline. It moves away from the single model idea towards, well, a more pluralistic system.
4:45Step one is about discovery and specialization. Step two is about fair aggregation. Okay, let's tackle step one first. It's called expectation maximization direct preference optimization, or EMDPO. Bit of a mouthful. How does EMDPO take this ternary preference data and figure out who's who and what they want? Think of expectation maximization, the EM part. That's kind of smart sorting algorithm for dealing with mixed data, like our pool of annotators with different hidden preferences. It tries to soft cluster annotators based on these unobserved factors, these latent types, which the paper calls ziversadors.
5:21It's an iterative process. So as I'm providing my AVSBVSC choices, the system is learning, hmm, this person seems to fit type 1 or maybe type 2. Exactly. The E-step expectation calculates the probability that you, the annotator, belong to each of the, say, typical possible hidden groups based on your choices. Then the M-step maximization uses that information. It trains a specific LLM policy tailored for each group taller, weighting your feedback based on how likely you are to belong to that group. Ah, so it's learning the groups and training specialized models for those groups at the same time.
5:53Precisely. EMDPO outputs an ensemble of LLMs, Pi-1-Moyers or Pi-2-2, up to table each bitter of those is optimized for a discovered preference type. And a really nice feature is that it stays reward-free. It builds on DPO's stability, avoiding the need to explicitly model a reward function, which can be tricky and unstable in traditional RLHF. That's clever. Okay, so we've got our specialized team of LLMs, each one great for its particular group. But often, you need one model for deployment, right? A user isn't going to announce, I'm preference type Z4 when they ask a question. We need a single policy that's robust and crucially fair.
6:30This brings us to step two, aggregation. Why not just, I don't know, average the outputs? Yeah, averaging or just picking randomly from the ensemble might maximize the average outcome across everyone. But it could leave some groups, especially smaller ones, feeling really poorly served. Their specific needs get averaged out. That's where the fairness part comes in, with min-max-regret aggregation or MMRA. Manx regret. The name suggests it's focused on minimizing the worst case scenario for any group. How do they define regret here? So regret for a specific group dollars is defined like this. Imagine the group's ideal policy is their specialized model.
7:05If we instead deploy a single combined policy dollar, the group dollar will likely get a slightly lower expected reward than they would with their perfect model. That difference, that loss in potential reward compared to their ideal is their regret. It measures how much they lose out by using the general policy dollar. Got it. So if the general policy is great for group A but terrible for group B, group B has high regret. Exactly. And MMRA doesn't try to maximize the average reward. Its goal is to minimize the maximum regret experienced by any subgroup. Find the group that's potentially worst off under policy dollars and try to make that outcome as good as possible.
7:42It focuses on equity, ensuring no single preference group is severely underserved by the final aggregated model. That feels like a much more, well, ethical approach to creating a general policy from diverse specialists, making sure no one gets left completely behind. So does this whole pipeline ternary data, EMDPO, MMRA actually work in practice? What did the experiments show? They showed pretty strong validation, starting with that data requirement. Let's go back to that tricky MPI data set with the P1 and P2 opposite personalities. Right, the ones that looked identical with binary data. What happened when they used ternary data with EMDPO?
8:16The difference was clear. EMDPO trained with ternary preferences was significantly better at distinguishing P1 and P2. It achieved better reward margins for those groups, better accuracy in identifying them, compared to when it only had binary data. It really backs up the theory. You needed those three choices to pull apart the hidden types. The simple data fix really did unlock the potential. Okay. Theory confirmed. And what about the fairness part, the MMRA aggregation? Did it protect those opposite groups? It did. When they applied MMRA, specifically a version called MMRA Lightweight, it performed much better than just averaging or using a standard DPO model.
8:55It managed to almost have the maximum regret for those P1, P2 groups. That's a big improvement in fairness for groups that were previously invisible. Having the regret, yeah, that's tangible evidence. Okay, what about a more realistic scenario? Something with, say, diverse cultural views? For that, they used the Global Opinion QA dataset. This uses real-world polling data from different countries, Britain, Indonesia, Mexico, Pakistan, to simulate annotators with very different viewpoints. Here, the latent types are basically national perspectives. That sounds challenging. You've got potentially very different worldviews, maybe unequal amounts of data from each country.
9:35For sure. But EMDPO handled it surprisingly well. It managed to cluster the annotators effectively based on their likely country-level preferences. It achieved good reward margins, comparable accuracy to benchmarks, and interestingly, sometimes even better than models explicitly told which country the data came from, especially when the true labels were imbalanced. It suggests EMDPOs successfully found the underlying structure without needing explicit labels. That's impressive. And the fairness. How did MMRA do when combining these diverse national perspectives? The MMRA results were quite striking.
10:07The final aggregated policy you produced achieved zero positive regret for Britain, Indonesia, and Pakistan. Mexico had only a small amount of regret. Essentially, the MIRA policy managed to create a single output strategy that was almost as good as the specialized policy for three out of the four groups and only slightly less good for the fourth. Compared to averaging or standard DPO, it demonstrated a much more equitable balance across these diverse viewpoints. Okay, that really ties it all together. So, to recap for you listening, the big takeaway seems to be, if you want alignment that respects diversity, you need to shift gears.
10:41Stop chasing one single perfect model. Instead, think about an ensemble of models, trained on latent types discovered from the data that's the EMDPO part. Then combine them fairly using something like MMRA. But the absolute key, the starting point, is realizing binary preferences, A versus B, are insufficient. You need ternary preferences, AVIUS BVSC, to even make identifying that diversity possible in the first place. Your data collection design is crucial. Yeah, exactly. And this framework, EMDPO plus MMRA, gives you a way to do that pluralistic alignment without needing to estimate those often complex and sensitive explicit reward models.
11:21It's a more direct, stable route to building LLMs that are hopefully not just capable, but also more equitable. Agreed. But, you know, this work, as impressive as it is, naturally leads to another thought, maybe a provocative one to leave you with. This whole framework assumes we can sort people into a, you know, a limited number of discrete groups, dalatypes. We're essentially putting people into buckets. It's a necessary approximation mathematically for the model to work right now. Right. But what if human preferences aren't really bucketable? What if they exist on a true continuum where everyone's slightly different?
11:55Can our alignment methods ever truly capture that infinite granularity? or are we always going to be dealing with these kinds of approximations? And if we are always approximating, how do we decide what level of regret is ultimately acceptable for those whose unique views don't perfectly fit a discovered type? Something to chew on. It's definitely a deep question for the future. For now, thanks for diving into this with us. And thank you for listening. We'll see you next time on The Deep Dive.
From the publisher
The academic paper claims that pairwise-comparison-based RLHF is incapable of learning heterogeneous preferences, whereas tenary comparisons can. They propose **Expectation-Maximization Direct Preference Optimization (EM-DPO)**, a clustering algorithm that discovers latent user preference groups and trains an ensemble of specialized LLMs for each group. Crucially, the authors establish a theoretical link to econometrics, arguing that **binary comparisons are insufficient** for identifying heterogeneous preferences, demonstrating the necessity of collecting **ternary preferences** (preferences among three options). Finally, the paper introduces **MinMax Regret Aggregation (MMRA)** to combine the ensemble models into a single "fair" policy that minimizes the worst-case performance loss across all identified user subgroups, ensuring equitable deployment.




