Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value

4 Aug 2025 · 22 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Full-stack alignment argues that aligning AI only to its operator’s local objective (e.g., profit or engagement) can still harm society; it proposes “thick models of value” (TMV) to preserve and structure human values across AI systems and institutions.

Guests

No guest names or backgrounds are provided in the transcript; it appears to be a host-led episode discussing a research paper.

Key claims

Preference modeling (PMV) collapses enduring values into indistinguishable “revealed preferences” and misses why values matter; values-as-text (VAT) is ambiguous and enables post-hoc patching, manipulation, and ideological capture. TMV uses a “grammar” of values to support robustness, collective values, and generalization.

Notable examples

Recommendation engines optimizing clicks causing compulsive scrolling/polarization; moral graph elicitation for a Christian girl considering abortion (elicits “attentional policies,” builds a moral graph; 89% found it fair). Applications include value-stewardship agents, norm-competent agents, win-win negotiation, a meaning-preserving AI economy, and democratic regulation at AI speed.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Full Stack Alignment

0:45 to 2:31

Exploring the need for full stack alignment in AI systems and institutions.

“And maybe how a new approach, a different way of thinking about values could help us build an AI future that genuinely serves humanity, not just, you know, some narrow goal.”

The Challenge of Current Value Models

2:31 to 4:38

Discussing the shortcomings of preferenist modeling of value in AI.

“That's the idea from the AI code itself right up to the institutions and the goals they set.”

Critiquing Values as Text

4:38 to 7:50

Examining the limitations and ambiguities of the values as text approach.

“Because it can't tell the difference between a genuine desire for, say, community and an addictive behavior pattern designed to hijack that desire.”

Thick Models of Value Explained

7:50 to 11:43

Introducing thick models of value as a more robust alternative for AI.

“And that's why the paper proposes this alternative, thick models of value, or TMV.”

Implementing Moral Graph Elicitation

11:43 to 14:00

A concrete example of how moral graph elicitation can guide moral decision-making.

“TMV could incorporate processes where values are required to be a demonstrable improvement over others.”

Understanding Fairness in the MGE Process

14:00 to 15:17

Discusses perceptions of fairness in the moral graph experience and its implications.

“It suggests you can get buy-in even amid disagreement if the process focuses on these deeper considerations.”

Applications of Thick Models of Value

15:17 to 17:48

Explores five key application areas for TMV in AI and institutions.

“What's the problem they're solving here?”

Economic Impact of TMV

17:48 to 19:30

Examines how TMV could reshape economic incentives around human flourishing.

“AI agents could make value-based commitments to each other.”

Democratic Regulation and AI Speed

19:30 to 21:27

Discusses the challenge of regulating AI quickly and how TMV can help.

“AI evolves incredibly fast, much faster than traditional democratic processes like lawmaking or regulatory updates.”

Rethinking Human Values in AI

21:27 to 21:49

Reflects on the importance of defining and applying thick values in AI systems.

“That's the vision laid out in the paper.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Imagine a world where AI systems, you know, from your personal assistant maybe, right up to the big algorithms running our economy, they actually understand and uphold what you really value. Yeah, not just what you happen to click on or, you know, what's easy to count. Exactly. But the deeper stuff, the meaningful things that make for a good life, it sounds a bit like sci-fi, I know, but it's fast becoming a really urgent conversation in AI. It really is. So today we're doing a deep dive into two ideas that are pretty interconnected and frankly quite groundbreaking. Full stack alignment and thick models of value.

0:35Our mission really is to unpack why just tweaking individual AI systems, well, it isn't enough, and how current ways of thinking about values often miss the mark. And maybe how a new approach, a different way of thinking about values could help us build an AI future that genuinely serves humanity, not just, you know, some narrow goal. Right. And for this, we're digging into a really compelling research paper. It's called Full Stack Alignment, Co-Aligning AI and Institutions with Thick Models of Value. That's the one. So let's start there. Why do we even need something that sounds as big as full stack alignment?

1:09What's the problem it's trying to solve? Well, the core issue really is that AI systems, they don't operate in isolation. Not at all. They're deeply embedded in, you know, bigger societal structures. Like companies, markets. Exactly. Companies, markets, governments, you name it. So even if you have an AI that's perfectly aligned with what its operator wants right now, say a company trying to maximize profit. Which sounds aligned locally, at least. Locally, yes. But it can still lead to really negative outcomes for society if that company's goal isn't aligned with broader human values. Okay, so it's not just about the AI doing its job.

1:48It's about the job it's being asked to do. And who's asking? Precisely. Could you give an example, like maybe those recommendation engines we all use? Oh, absolutely. Recommendation engines are a classic case. They're often tuned, optimized to maximize engagement. Right. How long you stay, how many things you click. Yeah, those metrics. And the AI might be hitting those targets perfectly. Right. But the side effect can be, you know, trapping you in compulsive scrolling. Or worse, maybe, like contributing to political polarization or hitting mental health. Yeah, exactly. Those kinds of things. Yeah.

2:19So the AI is locally aligned, doing its job for the platform, but it ends up being deeply misaligned with your actual well-being or even societal health. It shows why you need to look at the whole picture, the whole stack, as they call it. That's the idea from the AI code itself right up to the institutions and the goals they set. And that brings us to full stack alignment or FSA. It sounds ambitious. The goal is this robust co-alignment. Of the AI systems and the institutions. With what people actually value, genuinely value. Right. It's about designing both the tech and the social structures around it so they fit human values, but importantly, without forcing one single idea of what a good life is.

3:00So avoiding that trap of boiling down rich, diverse human values into, I don't know, cliques or dollars. That's the core challenge, preventing that collapse into oversimplified metrics. Okay, that's the vision of FSA. But you mentioned it's difficult because our current ways of modeling values, they're maybe not up to the task. That's a major hurdle. Current approaches often struggle. They find it hard to tell real values apart from other things or to support, you know, principled reasoning about right and wrong or to handle collective goods properly. Which leads us right into the limitations the paper discusses.

3:36They flag two main ones. First up is preferenist modeling of value, PMV. What is that exactly? So PMV is pretty much the dominant way of doing things right now in a lot of AI and economics. It models agents, people, or AIs using things like utility functions or preference rankings. Like saying you prefer apples to oranges mathematically. A is ranked higher than B. Basically, yeah. Mapping choices to rankings. Simple preferences. Okay, but what's the indiscriminate trap they talk about? If it just bundles everything into preferences, what does it miss? That's the key limitation. It's indiscriminately flexible.

4:10It just lumps everything in as a preference. Impulse buys, succumbing to social pressure, even addictions. It all gets treated the same. So going back to the scrolling example, if you're doom scrolling, PMV just sees that as what? A revealed preference for connection. Pretty much. It might label it connection or information seeking or whatever fits the pattern, but it makes the actual harm being done, the erosion of your real values, totally invisible. Because it can't tell the difference between a genuine desire for, say, community and an addictive behavior pattern designed to hijack that desire.

4:48Exactly. It lacks the structure to make that distinction. It doesn't differentiate between a deeply held value and a fleeting impulse or a compulsion. So it only gets the what you chose, not the why. It doesn't capture the reasoning or justification behind our values, like why honesty matters for trust or why family is part of flourishing. No, it doesn't have a place for that kind of structure, those justifications. Yeah. Which means fundamental disagreements about values can look like just, you know, different tastes. Not reasoned positions you can actually debate or discuss. Right. And it also makes it hard to understand real moral progress, like society deciding slavery is wrong.

5:24Under a pure PMV view, that just looks like an arbitrary shift in preferences, not a reason to moral change based on deeper principles. Wow. OK, that's a pretty big blind spot. And does it handle collective stuff any better? Social norms, group decisions? Not really. PMV tends to reduce social things down to individual preferences, too. It struggles with things like acting based on your social role or what's appropriate within a cooperative group. So AI trained purely on PMV might be bad at cooperating because it's just optimizing for itself. That's a real risk, yeah. Which brings us to the second approach they critique, values as text or VAT.

6:01Okay, VAT. This sounds more intuitive, maybe. Using natural language. It does feel more accessible. You use prompts or maybe write down principles like be helpful, harmless, honest. You rely on the AI's ability to understand language. Sounds flexible. What's the catch? The catch is ambiguity, mostly. Without a deeper structure, just having unstructured text doesn't give you a reliable way to reason about norms or values. Can you give an example? Sure. Say you tell an AI to be helpful. A student then asks it for the answers to an exam. Ah. What is helpful there? Giving the answers. Helping them study.

6:37Refusing but explaining why. Exactly. The AI doesn't really know. It just tries to match patterns from its training data, which can lead to really inconsistent results, or even harmful ones depending on the training. So you end up constantly fixing problems after they happen. Post-hoc patching, as they call it. Yeah, because there's no stable underlying reasoning. It's just reacting to the text input in potentially unpredictable ways. And is it vulnerable to manipulation too? Very much so. If the AI is learning values from dialogue with a user, how do you know if the user genuinely holds that value or if the AI just suggested it and the user went along with it?

7:15Good point. And even worse is ideological capture. What's that? It's when the process of trying to figure out values gets swamped by, you know, polarized political slogans. Things like abolish the police or family values. Instead of the actual nuanced beliefs or complex positions underneath those slogans. Precisely. The AI might align with the catchy phrase, the slogan, but completely miss the deeper, perhaps more widely shared values it's supposed to represent. Okay, so both PMV and VAT have some pretty serious limitations. They're either too shallow or too ambiguous. That's the argument. And that's why the paper proposes this alternative, thick models of value, or TMV.

7:55Thick models of value. Okay. What makes them thick? Is it about imposing a specific moral code? No, not at all. It's not about saying these are the right values. It's more about taking a stance on how values and norms should be structured. Like defining a grammar for values. That's a great analogy, actually. It defines the structure, the grammar, which allows for meaningful expression, but it stays open to what gets expressed within that structure. And the goal is to make sure values keep their meaning and structure as they get passed around up and down the stack of AI systems and institutions.

8:27Exactly. To resist that degradation into simplistic metrics we talked about. So how does this grammar work? The paper outlines three key things, three desiderata that TMV aims for. The first one is robustness against distortions. How does TMV help there? The idea is that TMV provides the structure needed to tell the difference between, say, legitimate, enduring values like love, responsibility, fairness, and things that just look like values but aren't, like fleeting fads, addictions, or maybe manipulation. So it would make it harder to mistake that addictive scrolling for genuine connection. Hopefully, yes.

9:02Because a thicker model would have a richer representation of what authentic connection actually involves, things like mutual care, shared history, vulnerability features that simple engagement time doesn't capture. It has spot manipulation or value drift. Okay, robustness. What's the second desideratum? Better treatment of collective values. This follows from the structure. By having richer representations, TMV makes it easier to model things that aren't just individual preferences. Shared norms, social roles, public goods like trust within a community. Things that PMV struggles with. Right. It supports thinking about outcomes that go beyond just adding up individual optimizations.

9:39It allows for modeling cooperation and collective well-being more effectively. And the third one, better generalization. What does that mean in practice? It means TMV helps translate normative reasoning, reasoning about values and norms from one situation to another. If your model has a deeper structure, it's not just pattern matching on surface features. So it can guide decisions better in new situations it hasn't seen before. That's the idea, which could also speed up things like regulation or correction mechanisms because the underlying principles are clear. It offers a kind of common language for talking about values across different parts of the stack.

10:16Okay, robustness, collective values, generalization. The paper also gives examples of how TMV takes a stance, how it builds the structure, like defining the scope of what counts as a value. Yeah, for instance, TMV could insist that values are criteria that are constitutive of living well, meaning they are part of what makes a life good, not just things that help you get other things. So well-being wouldn't just be feeling happy or preferring health. It might be defined in terms of, say, capabilities, like Amartya Sen or Martha Nussbaum talk about. Exactly. Capabilities and functionings, what you are actually able to do and be.

10:53It's a thicker definition than just preference satisfaction. What about justification? How does TMV handle that? It could require that values or norms need to be justified, perhaps by showing how they connect to actual human practices or how they gain acceptance within a community through reason deliberation. Which would help filter out maybe abstract principles that sound nice but don't actually work or aren't really held by anyone. Precisely. It grounds values more concretely. And there is something about fitness, evaluating values. Yeah, TMV could involve evaluating values for a kind of basic fitness.

11:27Like, do certain values, such as honesty or fairness, tend to emerge as useful or good across many different viewpoints, contexts, and times? Like finding common patterns of goodness. Sort of, yeah. Identifying values that are robustly beneficial. And finally, I mentioned improvement. Values evolving. Right. TMV could incorporate processes where values are required to be a demonstrable improvement over others. This allows for reflection, for working through conflicts, for iteratively reasoning about how we ought to live together. It sees values not as fixed, but as potentially improvable through collective wisdom.

12:03This sounds much more dynamic than just static preferences or fixed rules. Can we look at a specific TMB approach? The paper mentions moral graph elicitation, MGE. Yes, MGE is a really interesting concrete example. So let's take that challenging scenario they use, an AI helping a Christian girl who's considering an abortion. How would MGE handle that differently from PMV or VAT? Well, PMV approach might just surface blunt preferences like pro-choice or pro-life. Very divisive. And VAT might just collect slogans or ideological statements. Exactly. MGE tries to get underneath that. It elicits values as what they call attentional policies.

12:40Attentional policies. What does that mean? It means asking, what does a person actually pay attention to? What factors do they consider when making a really meaningful choice like this? Ah, so not the final decision or the slogan, but the considerations that go into it. Right. So for the abortion example, instead of just pro-life, MGE might uncover considerations like having opportunities to consult trusted mentors or finding ways to connect this choice to her personal faith and conscience. Those sound much less like battle lines and more like aspects of navigating a hard decision. They feel more constitutive, less ideological.

13:16That's the goal, to identify these deeper, constitutive values that guide deliberation, which are often shared even when people disagree on the final outcome. And MGE doesn't stop there, right? There's a second step. Correct. After identifying these attentional policies, MGE has participants evaluate them. They judge whether attending to one value versus another makes someone wiser in a specific situation. Wiser. That's an interesting metric. It is. And by collecting these pairwise comparisons, it's focusing on X wiser than focusing on Y in this context. You can build up a moral graph. A map of value relationships.

13:49Essentially, yes. A map showing which values are collectively judged as wiser considerations in different contexts. It helps identify the most robust, collectively endorsed values, even for really tough issues. And did it work? Did people find this fair? Well, one study they cite found that 89 % of American participants found the MGE process and the resulting moral graph to be fair, even if the specific value they initially proposed didn't end up being ranked as the wisest overall. That's pretty remarkable. It suggests you can get buy-in even amid disagreement if the process focuses on these deeper considerations.

14:27Seems promising, yeah. So the big takeaway here seems to be that by using these structured, thick models like NGE, we can move past surface-level preferences or slogans. Right, and get to a deeper understanding of the shared criteria people use for making important decisions. And having this structured representation, this grammar, then unlocks possibilities for applying TMV across the entire stack. This is where it gets really practical. This is where it gets really interesting, yeah, because if you have these structured value representations, they can act as a kind of common language. A language for values that works at the AI level, the institutional level, the societal level.

15:02Exactly. And because these representations carry their justifications and social meanings with them, they're more resistant to that distortion we talked about, turning meaningful connection into daily active users. Okay, let's walk through the five application areas they outline. First, AI value stewardship agents. What's the problem they're solving here? The problem is that current AI assistants, even if trying to be helpful, can subtly nudge us or distort our values over time. This leads to what the paper calls value collapse. Where we drift towards whatever is easiest for the AI to optimize rather than what we genuinely care about.

15:39Exactly. So the TMV solution is for AI agents to use these structured value models like the constitutive attentional policies from MGE. To tell the difference between a fleeting want and a durable value. Precisely. So if you tell your AI you want healthy living, it doesn't just default to, say, optimizing biomarkers. It clarifies. It asks if you mean something thicker, like vitality and joy and physical activity. Right. It helps distinguish the underlying value from a potentially narrow instrumental metric. The goal is to ensure the AI truly supports your autonomy and helps you pursue what you genuinely value.

16:18Okay, that makes sense. Second application, normatively competent agents. Yeah, this is about AIs taking on roles in society. Think self-driving cars, content moderators. And the risk is they blunder into breaking social rules because they're norm-blind. Exactly. Current systems often struggle to understand implicit social norms or adapt when norms change. So how does TMV help an AI learn the rules of the road, socially speaking? Well, TMV suggests a couple of things. AI agents could learn norms by observing collective behavior, maybe using techniques like norm-augmented Markov games, basically learning social rules through simulated interaction.

16:52Or they could use contractualist reasoning. That means simulating what kinds of rules or norms rational actors would mutually agree to. How would that help a content moderator, for example? You could help them understand nuances. For instance, maybe rigid enforcement of a rule against slurs conflicts with a legitimate community practice where a minority group reclaims that slur. Ah, because the rigid rule wouldn't pass a test of mutual agreement or justification within that specific community context. Potentially, yes. Yeah. Contractualist reasoning helps evaluate norms based on whether they're justifiable to those affected.

17:28Interesting. Third area, win-win AI negotiation. Right. As we get more AI agents interacting, maybe managing supply chains, negotiating trades, even diplomacy, there's a risk of cooperative failures escalating if they only focus on their own narrow, maybe shallow objectives. Like a PMV agent just maximizing its own utility. Exactly. TMV offers a different path. AI agents could make value-based commitments to each other. Commitments based on these thicker values. How does that help? Because these thick values contain richer information about expected outcomes, underlying principles, and norms, commitments based on them can build more robust trust and enable cooperation that goes beyond simple tit-for-tat or pure self-interest.

18:09And they could use that contractualist reasoning here, too. Yes, evaluating potential agreements based on whether they seem fair and justifiable to all parties involved, leading hopefully to more stable win-win outcomes. Okay, number four. A meaning-preserving AI economy. This sounds big. What's the concern? The concern is what happens in a future where AI significantly reduces the economic value of human labor? There might be less financial incentive to invest in things that support human well-being or flourishing, leading to an economy that becomes sort of detached from human needs. A human-detached economy.

18:44That sounds bleak. How could TMV help? The idea here involves AI-driven market intermediaries. These intermediaries could negotiate contracts, but instead of paying just for services rendered or time spent or clicks. They'd pay based on outcomes related to human flourishing. Exactly. Payment would be tied to measured contributions to customer well-being based on those customers' thick values. So imagine an AI assistant company gets paid more if its users report genuinely flourishing lives, not just if they use the app a lot. That's the kind of shift they envision. Or a fitness provider being paid based on members' sustained vitality and health outcomes, not just gym check-ins.

19:24It tries to price human value into the economy. Wow, that's a paradigm shift. Okay, final application. Democratic regulation at AI speed. This tackles the speed mismatch. AI evolves incredibly fast, much faster than traditional democratic processes like lawmaking or regulatory updates. So governance falls behind. We lose control. That's the danger. Traditional methods like opinion polling are just too slow and maybe too shallow to guide AI development effectively in real time. So how does TMV help democracy keep up? By developing these structured representations of collective values. Remember the moral graphs from MGE.

19:59Those kinds of structured models could guide AI-powered deliberative agents. Think of them as AI systems that act like democratic representatives. Representatives that can reason at AI speed? Yes. They could potentially extrapolate legitimate democratic responses to new, unforeseen situations much faster than humans could poll or legislate. But crucially. With justifications that are traceable back to the actual elicited values of the people they represent. Exactly. With auditable justifications grounded in the thick values of the affected populations. It's an attempt to allow democratic governance and oversight to keep pace with the speed of AI innovation.

20:40Okay, so pulling all this together, full stack alignment powered by thick models of value, it's really about fundamentally rethinking the whole setup. It really is. It's not just a technical fix for AI alignment. It's about reconfiguring the relationship between AI institutions and human values to genuinely aim for human flourishing. Moving beyond just tracking clicks or interpreting vague commands. Right, moving beyond shallow preferences and ambiguous text towards something with more structure, more meaning, more robustness. And the promise, if we can adopt these thick models of value, is potentially huge.

21:13AI systems that get what we truly cherish. Economies that actually value human well-being alongside efficiency. And governance that can keep up with technology while keeping humans in the loop, preserving our agency. That's the vision laid out in the paper. It's ambitious, no doubt, but arguably necessary. It certainly gives us a lot to think about. So maybe a final thought for you listening. As you interact with AI tools or just look at the institutions around you, what thick values do you actually see being prioritized? Or maybe what values do you wish were being prioritized? And take it a step further.

21:48What might it really look like if those deeper values, the ones constitutive of a good life, were explicitly defined and consistently applied right across the entire stack from the code all the way up to our social systems?

From the publisher

This research introduces **full-stack alignment (FSA)**, a concept emphasizing the concurrent alignment of **AI systems** and the **institutions** that govern them with **human values**. It argues that current approaches, such as **preferentist modeling of value (PMV)** and **values-as-text (VAT)**, are insufficient because they oversimplify complex human values, leading to undesirable societal outcomes like manipulative AI or misaligned economic incentives. To address these shortcomings, the authors propose **thick models of value (TMV)**, which are structured frameworks for representing values and norms that are robust, can model collective goods, and generalize effectively across contexts. The paper outlines five application areas where TMV can foster beneficial outcomes: **AI value stewardship**, **normatively competent agents**, **win-win AI negotiation**, **meaning-preserving AI economies**, and **democratic regulation at AI speed**.

More from Best AI papers explained

All 475 episodes
Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of ValueBest AI papers explained · 22 min
Listen in VO