In short
The episode argues that large language models often fail in strategic settings because they assume opponents are perfectly rational “Nash-type” optimizers, creating “Nash traps” where the mathematically correct Nash equilibrium action lowers real payoffs against bounded-rational humans. It uses 12 normal-form strategic games (e.g., inventory stocking, corporate pricing, advertising bids) with 186 human participants and six frontier models (GPT 5.2, Claude Sonnet 4.5, Gemini 3 Flash; plus DeepSeq v3.2, Grok 4.3, Quen3).
Key claims
only 13.4% of humans play Nash; 45.2% are level-1 and 33.9% level-2. Example: in inventory stocking, Nash “Action C” yields payoff 1.49 vs human “Action A,” while adapting to play “Action B” yields 8.95. Fixes: behaviorally informed prompts reduce Nash play but don’t reliably improve best-response; supervised fine-tuning with a router (“TrapAware SFT”) achieves ~97% payoff recovery in traps and preserves Nash behavior (~100%) when not a trap.
Guests
No guest individuals are named in the transcript.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Nash-type Reasoning
0:50 to 2:11
Discover how AI's assumptions about human rationality can lead to errors.
“We are exploring the hidden assumptions AI models make about human rationality in competitive environments.”
Levels of Human Strategic Reasoning
2:11 to 4:50
Learn about different levels of human reasoning in strategic games.
“Well, the data set reveals a Maslow strategic mismatch right out of the gate.”
The Nash Trap and Its Implications
4:50 to 7:35
Explore how Nash traps emerge in AI-human interactions and their consequences.
“Which means we need to look at how the human on the other side is actually operating.”
Prompting AI: A Double-Edged Sword
7:35 to 12:00
Understand the challenges and limitations of prompting AI to adjust strategies.
“deployed into a human population where almost half the people are playing at level one.”
Supervised Fine-Tuning for AI Improvement
12:00 to 14:00
Learn about supervised fine-tuning as a solution for AI's strategic failures.
“It comes down to how large language models actually function.”
Understanding Supervised Fine-Tuning
14:00 to 15:30
Learn how supervised fine-tuning enhances model performance in pricing games.
“For the ultimate solution, we turn to supervised fine-tuning, or SFT.”
Direct SFT Methodology and Its Tradeoffs
15:30 to 17:10
Explore how Direct SFT works and the tradeoffs it introduces in non-trap scenarios.
“The unintended consequence was a degradation of performance in the non-trap games.”
Introducing TrapAware SFT
17:10 to 19:00
Discover TrapAware SFT and how it selectively applies AI strategies based on context.
“And when that specialized policy activated during the tests, it recovered 97.0 % of the available payoff.”
Router Mechanism and Performance
19:00 to 21:00
Understand the router mechanism in TrapAware SFT and its impact on performance.
“We spend a tremendous amount of cultural energy worrying that artificial intelligence is going to outsmart us.”
The Future of AI Negotiation
21:00 to 21:44
Consider the implications of AI agents negotiating with each other using advanced strategies.
“Will we see supercomputers intentionally mimicking messy, unpredictable human flaws just to trick the opposing AI into abandoning its mathematical equilibrium in order to secure a better deal on a used car?”
Transcript
Automatic transcript. May contain errors.0:00Imagine you are buying a car. You are walking onto the lot, and instead of dealing with the high-pressure stress of haggling yourself, you pull out your phone and deploy an AI agent to negotiate the price for you. Yeah, that sounds like a dream, honestly. Right. And you would want that AI to be perfectly logical. Just a cold, calculating machine that knows exactly how to map out the absolute mathematical best deal. Of course. You want it to be ruthless with the math. Exactly. But what if its flawless textbook logic actually cost you thousands of dollars? And simply because the car dealer across the desk is, you know, a messy, unpredictable, emotional human being.
0:42It is a scenario that completely flips our expectations of what makes an intelligence artificial or otherwise actually effective in the real world. It really does. Welcome to today's Deep Dive. We are exploring the hidden assumptions AI models make about human rationality in competitive environments. Right, because we tend to assume that mathematical perfection automatically wins, but the recent experimental data shows that isn't always true. Exactly. So we will be digging into this massive data set covering 12 normal form strategic games. And for those who might need a quick refresher, normal form games are scenarios where everyone makes their choice at the exact same time without knowing what the other person is doing.
1:20Kind of like a really high stakes game of rock, paper, scissors. Yeah, exactly. But applied to things like inventory stocking, corporate pricing competitions, and adipizing auctions. And this data set tested real human participants alongside six frontier AI models, including GPT 5.2, Claude Sonnet 4.5, and Gemini 3 Flash. And the mission today is to understand why the AI's pursuit of theoretical perfection fails so spectacularly against human unpredictability, and most importantly, how these findings show we can actually fix it. Okay, let's unpack this, because this completely shifts how we should be thinking about delegating our daily tasks to algorithms.
1:59We trust these frontier models to be infallible decision makers, you know. We do, yeah. But to set the stage, we first have to look at the baseline of how an AI views the person sitting across the negotiating table. Right. What is its underlying assumption? Well, the data set reveals a Maslow strategic mismatch right out of the gate. When you drop these frontier models into a strategic game, they overwhelmingly act as what game theorists call a Nash-type reasoner. Ah, named after John Nash, the mathematician behind Nash equilibrium. Exactly. And a Nash-type reasoner operates on a very specific premise.
2:34They assume they are playing against a fully rational strategic optimizer. So the AI basically assumes you're going to play a perfect game. Right. And so it computes and plays the mathematically perfect equilibrium response. Across the board, looking at the tested models like DeepSeq v3.2, Grok 4.3, and Quen3, the data shows they almost entirely default to this hyper-rational behavior. Okay, I'm going to push back on why this is framed as a negative. If I want an AI to negotiate my salary or bid on a commercial property for me, shouldn't it be playing the most mathematically perfect game possible?
3:08I get why you'd think that. It seems like having a brilliant poker bot that plays statistically perfect hands based on pure probability. That sounds like a good thing. What's fascinating here is that playing a theoretically perfect strategy is only optimal if your opponent is also playing a perfectly rational strategy. Oh, I see. Think about the poker bot scenario you just brought up. If it is playing a mathematically flawless game, it expects its opponents to only bet when the statistical probabilities dictate they have a strong hand, right? Right. It expects everyone else to be doing the same math.
3:39But what happens when that bot goes up against an amateur who just drank three beers and decides to go all in on a terrible hand just for the thrill of the bluff? Oh, wow. The bot folds, doesn't it? Because it assumes the human must have this statistical upper hand. You have no framework for understanding an irrational bluff from someone who is just acting on emotion. Exactly. And the underlying data reveals this assumption gets even more entrenched when the models are explicitly prompted to reason out their answers. Wait, really? You'd think asking it to think it through would help? You would assume that asking an AI to think step by step, giving it time to consider context, might make it account for human quirks.
4:19But the exact opposite happened. When prompted to reason, their tendency to lock into Nash equilibrium strategies actually increased. Oh, because giving them more time to think just means they pull harder from their training data, which is full of formal logic and game theory textbooks. It just reinforces their rigid logic. Precisely. And giving the games a real-world business context, like framing the math as a corporate negotiation rather than just abstract numbers, well, it didn't meaningfully change their behavior. They still acted as if the human on the other side was a flawless supercomputer.
4:54Which means we need to look at how the human on the other side is actually operating. Because to see why this is a collision course, we have to contrast the AI's choices with the data from the 186 actual human participants who played these exact same 12 games. Right. And to contextualize the human data, we use a concept from behavioral game theory called level thinking. It's essentially a framework for measuring strategic depth in decision making. Let's break down how those levels work for the listener. Sure. So it starts at level zero. A level zero player is completely non-strategic. They act randomly without any consideration for the rules of optimization.
5:30It's your chaos, okay. Yep. Then you have level one. A level one player assumes their opponent is a level zero random player, and they just play the easiest, best response to that randomness. And then level two assumes the opponent is level one and plays the best response to that. Exactly. It builds on itself. So in a game of rock, paper, scissors, a level zero person just throws a hand blindly. But a level one person thinks most people throw a rock on the first try, so I'll throw paper. And the level two player thinks, well, they know most people throw rocks, so they are going to try paper, which means I'm going to throw scissors.
6:04It is an endless recursive loop of psychological outguessing. It really is. Now, a Nash type player, which is what the AI is doing, is engaging in infinite, perfect recursive reasoning. It assumes the loop goes on forever and settles at the absolute mathematical equilibrium. But the human data paints a very different picture of our cognitive bandwidth. Yeah. We are definitely not infinite recursive reasoners. No, we are definitely not. We are a very messy distribution of bounded rationality. Bounded rationality meaning like we just get tired and guess. Essentially, yes. It means that human beings have limited cognitive energy, limited time, and we often settle for a good enough strategy rather than computing the perfect one.
6:47That makes total sense. And in the aggregate data for the human participants, a mere 13.4 % of humans played the Nash strategy. Only 13.4 % played mathematically perfect. Wow. It is incredibly humbling to realize that the vast majority of us are not the strategic optimizers we think we are. So where does the rest of the population fall? The bulk of the human participants were lower-level reasoners. So 45.2 % were level one, another 33.9 % were level two, and 7.5 % were just purely random level zero. That's crazy. Nearly half the population is sitting at level one, essentially making straightforward, localized decisions based on the assumption that the other party is acting somewhat randomly.
7:30So we have frontier AI models stubbornly locked into infinite-level Nash reasoning, deployed into a human population where almost half the people are playing at level one. And when you force a mathematically perfect algorithm to interact with someone relying on bounded instinct, you get a structural breakdown. Game theorists call this operational disaster the Nash trap. Here's where it gets really interesting, because this is where the implications get very real. What exactly happens to the payoffs when a Nash-obsessed AI falls into one of these traps against a level one human? Well, a Nash trap occurs when playing the theoretically perfect equilibrium action is actively detrimental to your actual payoff.
8:10And that's because your opponent is behaving imperfectly. So you are making the right move, but because the environment is messy, your right move actually hurts you. Exactly. Let's use a specific example from the data. Game one, which is inventory stocking. We have two players trying to stock warehouses. The AI's textbook, perfectly calculated Nash action, is a choice we will call Action C. Okay, action C is the genius math move. Right, but the most common human action, the level 1 behavior, is action A. In this scenario, action C represents an optimized supply chain choice that requires the partner to also make an optimized choice to ensure maximum throughput.
8:47Which the human is not doing. No, the human is playing A, which is a simpler, less optimized stalking method. So if the AI stubbornly plays its perfect Nash action C against a human who plays A, the data set shows the AI earns a dismal payoff score of just 1.49. Only 1.49? It essentially gridlocks the supply chain because the AI is waiting for a perfect synergy that the human just never provides. Exactly. But if the AI had anticipated the human's bounded rationality like, if it had looked at the board, realized the human would likely play the suboptimal action A and adjusted its strategy to the empirical best response, it would have played action B.
9:24And what would have happened if it played B? It would have earned a massive payoff of 8.95. Wait, really? A leaf from 1.49 to 8.95 simply by abandoning perfection? By trying to outsmart the human with advanced textbook theory, the AI actually punishes itself. It really does. It is like a chess grandmaster trying to apply advanced chess theory to a game of Monopoly with a 7-year-old. The genius strategy actually makes you lose because it's completely miscalibrated for rolling dice and buying Boardwalk. If we connect this to the bigger picture, the grandmaster strategy requires a highly predictable, rational opponent to work.
10:01And the experimental data shows this isn't an isolated anomaly. Eight out of the 12 game designs tested were classified as potential Nash traps. Wow. Eight out of 12. Yeah. In two-thirds of these competitive scenarios, spanning pricing algorithms to advertising bids-perfect logic, failed against human reality. So this means a Nash trap isn't just some rare glitch. It is a fundamental feature of human-AI interaction. It is defined jointly by the game's mathematical mechanics, the specific behavioral quirks of the human, and the AI's programmed objective. Right. The AI has to understand the exact population it is dealing with to actually be effective.
10:39Okay, so if the AI is failing because it doesn't understand our human flaws, the logical next step would be to wonder, can we just tell the AI that humans are messy? That is exactly what the researchers wondered. The data tested inference time prompts to see if AI could be talked out of the Nash trap on the fly. Like just explaining it in the prompts before it makes a move. Exactly. They used a behaviorally informed prompt where the models were explicitly fed the exact human statistics we just outlined. They provided the prompt with the breakdown. 7.5 % random, 45.2 % level one, 33.9 % level two.
11:14So they fed it the raw statistical parameters of our messiness. Yes. And they explicitly instructed the model to factor in this bounded rationality. Well, it seems like having that hard data would snap the AI out of its textbook obsession, right? It did snap it out of the textbook obsession, but it created a completely new problem. The prompt successfully reduced Nash play by a staggering 76 to 84 percentage points. Oh, wow. So it completely abandoned the perfectly rational strategy. It did. However, it only increased the correct empirical best response by about 11 to 12 points. Wait. So we can easily convince the AI to abandon its perfect logic, but we can't easily convince it to pick the right alternative.
11:55The AI stopped playing Nash, but it basically just panicked and guessed wrong. That's essentially what happened. It comes down to how large language models actually function. These models are fundamentally next-token predictors trained on vast amounts of text. When you give them a text prompt saying humans are irrational, they easily abandon the token pathway that leads to the Nash equilibrium. Right, because they understand the concept of irrationality. Exactly. But to find the empirical best response, the model has to take those abstract behavioral probabilities, map them onto a complex multidimensional payoff matrix, and calculate expected values for every possible move.
12:34Oh, I see. And LLMs struggle heavily with multi-step arithmetic over unseen matrices unless they have been explicitly trained on that specific kind of calculation. Yes, moving away from Nash isn't enough, you know. The AI must mathematically navigate toward the calibrated human response. Open-ended prompting just leaves the model scrambling to translate behavioral data into a concrete mathematical action. So did they try any other prompts that tried to bridge that calculation gap? They did. They tested a welfare-aware prompt. Here, the models were instructed to explicitly consider three distinct outcomes.
13:08Their own expected payoff, the human's expected payoff, and the joint payoff of the interaction. forcing the AI to map out the consequences for the whole ecosystem rather than just solving for its own isolated logic. And this did perform a little better. It increased the best response rate by up to 18.7 points, and it reduced the massive joint payoff losses that occurred when the behaviorally informed AI just panicked and guessed wrong. But it's still just a patchwork fix. Exactly. It didn't solve the underlying limitation. Prompting is really just a band-aid that creates unpredictable behavior.
13:41We tell it the human is a level one player, it abandons the math, and then chips over its own computational limits trying to adjust. Which means we need to fundamentally rewire how the AI evaluates these decisions at a structural level. Since prompting is just a band-aid, how do we actually fix the AI so it perfectly navigates these traps? For the ultimate solution, we turn to supervised fine-tuning, or SFT. Okay, let's dive into supervised fine-tuning because the data set outlines a very specific experiment here. They were fine-tuning a Quen 3.59b model on a held-out benchmark of pricing games, right?
14:17That's correct. And for those unfamiliar, supervised fine-tuning is the process of taking a base model and explicitly training it on a highly curated set of examples. You are updating the model's core neural weights using input-output pairs so that it instinctively associates certain complex game matrices with the correct human-exploiting move. So it doesn't need to perform the raw arithmetic from scratch every single time. Exactly. And the experimental data explored two different methodologies for this. The first being direct SFT. Direct SFT. How does that work? Direct SFT functions as an always-on policy.
14:50The model was trained so that whenever it encountered a game, it would assess it and hit the human-calibrated best response if it was a trap and play the standard Nash strategy if it wasn't. It attempts to blend both capabilities into one unified decision-making process. It sounds like teaching a single algorithm to be both a street smart negotiator and a textbook mathematician simultaneously. How did that always on policy perform? Well, it was highly effective in the traps, recovering 95.2 % of the available payoff gain. Wow, 95.2%. It learned to exploit the human flaws beautifully based on its fine tuning.
15:27It did. But anytime you try to make one algorithm do two contradictory things, there is usually a tradeoff. Ah, of course. What was the catch? The unintended consequence was a degradation of performance in the non-trap games. Because it was an always-on policy, tinkering with the model's core weights to make it sensitive to human error caused it to kind of overthink straightforward math problems. Yeah, in scenarios where it should have just played the standard perfect Nash strategy because no trap existed, its accuracy dropped to 83%. So it started looking for human unpredictability even when the environment demanded pure logic.
15:59It basically got too paranoid. You could say that, yeah. Which brings us to the need for a system that knows exactly when to deploy which mindset. This is where method two comes in, and it is the holy grail of the data. TrapAware SFT or TASFT. TrapAware SFT. I like the sound of that. It treats alignment as a selective deployment problem, right? Exactly. Instead of forcing the AI to use one giant blended set of weights for everything, it implements a specialized router. Explain the mechanism of this router. How does it actually diagnose whether a game is a trap or just a standard logic problem?
16:34The router acts as an upfront classifier. Before deciding on a move, the system takes the game's mathematical payoff matrix and runs a rapid internal calculation. It computes the theoretical Nash equilibrium payoff, and then it computes the expected payoff against a simulated baseline of typical human behavior. Okay, so it compares the textbook math against the messy human reality. Right. And if the gap between those two numbers is large-like, if playing the Nash strategy leads to a severe drop in the expected payoff against a human, the router mathematically flags the game as a trap. And once it flags it as a trap, it actively turns on the specialized human-adapted policy weights that were fine-tuned to exploit that specific bounded rationality.
17:19Exactly. And when that specialized policy activated during the tests, it recovered 97.0 % of the available payoff. That is incredible. Yeah. But the brilliance of this router is what happens when the math doesn't show a large gap, right? If the router determines the game is not a trap, it doesn't use a fine-tuned, human-adapted model at all. No, it completely bypasses the fine-tuned weights and defaults to a strict symbolic Nash fallback. It relies on pure algorithmic math, achieving 100 % target accuracy for preserving Nash behavior when necessary. Wow. 100 % accuracy. Yeah, because it separates the diagnosis of the environment from the deployment of the strategy.
17:57It is brilliant. This structure operates exactly like modern self-driving cars. Like 90 % of the time when you are cruising down an empty highway, the car is driving using rigid, perfectly algorithmic rules. It is essentially using a Nash fallback, just calculating geometry and speed. That's a great analogy. But when you pull into a crowded supermarket parking lot, the car's sensors detect an environment dense with human unpredictability. The system dynamically routes control to a different set of protocols. It stops looking for perfect geometric driving lines and starts anticipating that a pedestrian might, you know, erratically step out from between two parked cars.
18:36The Trapperware SFT is basically the AI equivalent of shifting into parking lot mode. I love that. And it highlights the core belief here, which is that knowledge must be applied correctly. Effective human AI strategic interaction requires more than just having a powerful reasoning engine that can process billions of parameters. Right. It's about context. Exactly. It requires calibrated expectations of the population you're interacting with. If an AI system does not possess a structural mechanism to recognize when to be a mathematical genius and when to anticipate human foolishness, its intelligence remains incredibly brittle.
19:10We spend a tremendous amount of cultural energy worrying that artificial intelligence is going to outsmart us. But the immediate operational danger highlighted in these findings is that AI is going to outsmart itself by giving us far too much credit. Yeah, it operates on the assumption that we are perfectly rational actors. But anyone who has spent time navigating rush hour traffic or observing corporate bidding wars understands that pure rationality is rarely the driving force behind human decision making. Which means the era of AI acting as a purely rational textbook optimizer has to end for these tools to be genuinely useful to you in real world scenarios.
19:47Whether an AI agent is managing supply chain logistics or negotiating a freelance contract on your behalf or optimizing an advertising budget, that AI must fundamentally understand that it is operating in a messy, level one human world. It definitely needs specialized training. Yeah, it requires specialized routed architecture, like the Trapperware SFT, to dynamically adapt to our bounded rationality. It has to be able to mathematically read the room. That selective adaptation, driven by a structural understanding of human limitations, is the necessary bridge between artificial intelligence in a controlled laboratory setting and effective autonomous agency out in the real world.
20:26Which leaves us with one final, slightly mind-bending thought to mull over as we wrap up today's deep dive. As these AI agents become more deeply integrated into our lives and we begin handing over more of our daily high-stakes negotiations to them, We are going to see a shift in the ecosystem. We certainly are. What happens when two highly trained, trap-aware AI agents start negotiating with each other? Oh, man. The strategic layers of that interaction would be absolutely fascinating to observe. Right. Because if both agents are utilizing routers designed to detect and exploit the Nash trap, will they both attempt to project an aura of bounded rationality?
21:03Will we see supercomputers intentionally mimicking messy, unpredictable human flaws just to trick the opposing AI into abandoning its mathematical equilibrium in order to secure a better deal on a used car? It's possible. The ultimate sign of advanced artificial intelligence might actually be its ability to perfectly simulate human irrationality. So what does this all mean? It means the future of AI isn't just about artificial intelligence. It's about artificial empathy for our own logical flaws. and perhaps the weaponization of those flaws. Thank you for joining us for this deep dive into the hidden dynamics of human AI strategy.
21:39Until next time, keep an eye on your AI and consider double-checking its logic before you let it haggle for you.
From the publisher
This paper investigates a critical strategic mismatch between Large Language Models (LLMs) and human decision-makers in competitive environments. Through game-theoretic experiments, the researchers demonstrate that LLMs predominantly act as Nash-type reasoners, assuming their opponents are perfectly rational, whereas humans exhibit bounded rationality and varied reasoning depths. This overestimation of human sophistication often leads LLMs into a Nash trap, where equilibrium play fails to maximize payoffs against actual human behavior. To rectify this, the authors propose supervised fine-tuning methods, including Trap-Aware SFT, which calibrates model responses to empirical human benchmarks. Their findings suggest that effective human–AI alignment requires models to possess not just high reasoning capabilities, but also calibrated expectations of human behavior. Ultimately, the study advocates for a selective deployment architecture that preserves equilibrium play while adapting strategies when human interaction makes it more profitable.




