Reasonably reasoning AI agents can avoid game-theoretic failures in zero-shot, provably

24 Mar 2026 · 20 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Whether off-the-shelf, zero-shot AI agents can avoid game-theoretic failures (e.g., price wars, persistent overcharging) and converge to stable Nash equilibria without manual fine-tuning, even with stochastic, noisy, partially unknown payoffs.

Guest backgrounds

No guests are named in the transcript; only hosts/speakers discuss the research.

Key claims

Standard LLMs are posterior belief samplers, not perfect utility maximizers, yet can still converge if they perform Bayesian learning plus asymptotic best-response learning. A prompting scaffold called posterior sampling best response (PSBR) outperforms one-step “social chain of thought” (SCOT). PSBR remains robust under Gaussian-noise feedback where ~25% of observations mislead about true rewards.

Notable examples

Prisoner’s Dilemma, Battle of the Sexes, Promo (alternating marketing), Samaritan (altruism/moral hazard), Lemons (adverse selection/trust). Reported success: SCOT collapses (0% on complex cooperation); PSBR reaches ~92.5–98% success in clear games and ~71–98% under noisy unknown-payoff settings.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Unstable Nature of AI Agents

0:46 to 2:20

Discussion about how AI agents can lead to unstable pricing strategies.

“And the stakes here, I mean, they really cannot be overstated.”

Understanding Super Competitive Pricing

2:21 to 4:11

Exploration of how AI can maintain artificially high prices without collusion.

“So they just kind of silently agree to overcharge us.”

Game Theory and AI Behavior

4:12 to 5:32

Analyzing the failures of AI models in achieving Nash equilibrium.

“They know an incredible amount of vocabulary.”

The Limitations of Current AI Solutions

5:33 to 7:05

Challenges faced in training AI models for strategic decision-making.

“Which feels like asking a calculator to invent calculus on the fly.”

The Concept of Reasonably Reasoning Agents

7:06 to 8:26

Introduction to RR agents and their capabilities in adapting strategies.

“Standard economic game theory assumes that players are perfect, calculating expected utility maximizers.”

Testing AI Models in Game Scenarios

8:27 to 10:23

Overview of various complex games used to test AI models' stability.

“At first, the AI has no idea where the walls are.”

Comparing SCOT and PSBR Methods

10:24 to 12:28

Examination of two reasoning methods and their effectiveness in AI.

“And finally, they tested lemons, which is based on the classic economic problem of adverse selection.”

Success of PSBR in Complex Cooperation

12:29 to 14:09

Highlighting PSBR's superior performance in achieving cooperation.

“Which brings us to the centerpiece of this study, PSBR.”

Understanding Continuation Values in AI

14:09 to 18:08

Learn how AI agents can navigate complex games using continuation values and the implications of future simulations.

“It almost perfectly executed highly complex cooperative strategy Zero Shot.”

Implications for AI in Real-World Economies

18:12 to 19:30

Explore the broader implications of AI reasoning and cooperation in today’s economy.

“The noise is basically filtered out by the sheer force of long-term reasoning.”
Show all 11 chapters

AI vs. Human Conflict Avoidance

19:30 to 20:04

Ponder the potential for AI to surpass human conflict resolution abilities.

“It fundamentally changes how we view the deployment of autonomous systems.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Right now, autonomous artificial intelligence agents are actively negotiating the price of that digital ad you just scrolled past. Yeah, they really are everywhere. Exactly. And they are, you know, dynamically calculating the cost of the airline ticket you're about to buy. But behind the scenes, these hyperintelligent systems are, well, they're fundamentally unstable. Which is pretty concerning when you think about it. Right. Because without human intervention, they are highly prone to triggering chaotic price wars or just quietly keeping consumer costs artificially high. So welcome to this deep dive.

0:36Our mission today is to explore a really profound question regarding the very fabric of our digital economy. And it's a big one. It is. We want to know. Can independent AI models learn to play nicely and rationally in complex, high-stakes environments entirely on their own without us having to manually program them to do so? And the stakes here, I mean, they really cannot be overstated. We are deploying these models into the wild at just an astonishing rate. Yeah, it's incredibly fast. Right. But by default, standard AI does not know how to sustain stable, rational strategies when it has to interact with other AI systems.

1:12The fundamental tension we're unpacking today is that we are essentially handing the keys of our digital markets. You know, automated negotiations, supply chain logistics, high frequency trading over to systems that, out of the box, behave erratically when they're forced to compete or cooperate. It's wild, but I promise you, by the end of this deep dive, you will understand exactly how AIs can autonomously evolve to avoid these destructive game-theoretic failures and actually find stable, cooperative outcomes entirely on the fly. Okay, let's unpack this. Why is this such an urgent problem for your everyday digital life?

1:47Well, just consider the digital landscape you interact with daily. The research we are examining looks at autonomous algorithms that operate in environments where their strategic choices have real, tangible consequences. Like actual money on the line. Exactly. Real money. Studies show that when these algorithms are left to their own devices, they can accidentally sustain what we call super competitive prices. Which means what exactly? For the average person. It basically means they figure out how to keep prices artificially high for consumers, but without ever formally colluding or breaking antitrust laws in a way we can easily regulate.

2:21Oh, wow. So they just kind of silently agree to overcharge us. Yeah, exactly. And they also generate these wild, unpredictable profit margins that can literally destabilize entire market sectors. That's terrifying. And the alarming detail in the research is that off-the-shelf large language models, right, the incredibly smart systems you are already familiar with, like GPT or CLOD, they frequently fail to exhibit stable equilibrium behavior. Right, which is the baseline goal in these scenarios. Yeah, in game theory, you always look for a Nash equilibrium, which is a stable state where no player has an incentive to change their strategy, assuming the other players keep theirs the same.

3:00But standard models just fail to find the equilibrium on their own. They really do. Instead of finding that balance, they resort to incredibly brittle heuristics. Like basic shortcuts. Exactly. They rely on simplistic, short-term rules of thumb that completely fall apart the moment they're placed in a complex, repeated interaction with another agent. And to understand why this happens, we have to look at how these models are built, right? I mean, they are exceptional at pattern recognition and text generation, but they are not inherently strategic planners. No, not at all. And the proposed solution in the tech industry up to this point has been this massive, massive undertaking called post-training or fine-tuning.

3:40Which is just a huge amount of work. It is. The idea is to basically take every single AI model and manually train it on how to behave in specific strategic games. Essentially sit the model down and force feed it data on how to achieve a Nash equilibrium in a pricing game or an auction game, hoping it internalizes the lesson. Which brings up a glaring logistical nightmare. I mean, think about it like this. Imagine taking a group of highly intelligent but completely socially inexperienced toddlers and just placing them on the floor of the stock exchange to trade. That is a brilliant way to picture it.

4:14Total chaos. Right. They know an incredible amount of vocabulary. They can process massive amounts of information instantly, but they have absolutely zero understanding of the long-term consequences of their actions. They just scream and throw a tantrum if they don't get exactly what they want right this second. Which leads to total market collapse. Exactly. And the tech industry's current solution is essentially trying to send every single one of those toddlers to a universal, standardized boarding school. But there isn't just one AI model out there. There are thousands of independently developed AI models created by competing companies all across the globe.

4:51You cannot possibly coordinate a universal fine-turning procedure to ensure every single developer's AI plays nice. You really can't. And what's fascinating here is that the holy grail of this research completely bypasses that universal boarding school idea. Oh, thank goodness. Right. The researchers set out to discover if these models possess intrinsic reasoning capabilities to adapt and find stability, zero shot. And just to clarify, zero shot means the AI receives absolutely no explicit prior training or fine-tuning for the specific game it's playing, right? Exactly. No cheat sheet. The research explores whether these off-the-shelf models can fix themselves and find a Nash equilibrium purely by reasoning through the problem in real time.

5:33Which feels like asking a calculator to invent calculus on the fly. I mean, since we can't practically retrain all AIs, how does the research propose they actually achieve this? Because the paper introduces this framework, defining what they call reasonably reasoning agents or RR agents. Yeah, the RR agent framework. To qualify mathematically as an RR agent, an AI must demonstrate two fundamental capabilities. The first is Bayesian learning. Okay, break that down for us. You can think of this as the capacity to develop a theory of mind. The AI must be able to observe what its opponent has done in past interactions, update its internal beliefs based on that historical evidence, and form a working hypothesis about what the opponent's overarching strategy actually is.

6:16So it learns from hustry to map out the other player's mindset. Precisely. It's reading the room, so to speak. Okay. And the second capability? The second is called asymptotic best response learning. Once the AI has that theory of mind, it needs the ability to eventually learn and deploy an optimal counter strategy over time. The research proposes that if an agent possesses these two distinct capabilities, Bayesian learning and asymptotic best response, they will naturally evolve toward a Nash equilibrium along the realized paths of play. Meaning no manual fine-tuning required? None at all. That is a theoretical foundation.

6:52The agent continually updates its beliefs about the opponent and continually refines its response until both players just sort of lock into a stable, mutually beneficial pattern of behavior. Okay, I have to challenge this premise entirely just based on how these systems actually function. Standard economic game theory assumes that players are perfect, calculating expected utility maximizers. Right, the perfectly rational actor. Yes. But large language models are absolutely not that. As we established, they are stochastic. stochastic. They sample probabilities to generate their next output. Sometimes they hallucinate entirely new concepts, or they make noisy random choices based on their internal temperature settings.

7:29They absolutely do. So how can a machine that essentially rolls dice to pick its next word be considered a perfectly rational economic actor? That is a great question, and that pushback cuts to the very heart of why this paper is so significant. The research explicitly acknowledges your point. Standard language models are not expected utility maximizers. By their very architecture, they are what the researchers term posterior belief samplers. They evaluate a spread of probabilities and sample an action from that spread. The mathematical breakthrough detailed in the research is extending classical Bayesian learning to accommodate this highly stochastic behavior.

8:07Wait, so they mathematically prove that an AI does not need to be a perfect calculator to achieve stability. Exactly. But executing that kind of stability with a randomized text generator just sounds totally contradictory. It does. But it comes down to how the sampling behaves over time. Picture the AI's Bayesian updating like navigating a complex maze in the pitch dark. Okay, I'm picturing it. At first, the AI has no idea where the walls are. Its beliefs about the opponent are wide and uncertain, so its random sampling leads to scattered, highly erratic behavior. It's bumping into walls constantly.

8:42Because it's basically just guessing. Right. But as it observes more interactions, as it takes more steps and gathers more historical data, its posterior belief begins to concentrate. It starts building an accurate mental map of the maze. The mathematics in the paper prove that as long as the AI's belief concentrates onto a single correct hypothesis over time, its random sampling naturally stabilizes into an optimal response. Oh, wow. Yeah. The erratic noise simply fades away as certainty increases and rational behavior emerges from a stochastic system. That is, I mean, the theoretical math proving that a randomized text generator can become a rational actor is an incredible leap.

9:19But we need to look at how a language model actually executes this theory of mind in real time. Because the theory is only half the battle. Right, because executing that level of strategic foresight requires a massive amount of compute. You are essentially asking the model to process an exponentially expanding tree of possibilities. So to figure out if models can actually do this, the research puts specific methodologies to the test using five complex repeated games. And the testing ground is where this gets highly practical. They ran simulations using off-the-shelf models to see if they could naturally stabilize.

9:55They tested the prisoner's dilemma, which is your classic test of mutual defection versus cooperation. The classic game theory staple. Exactly. Then they tested the Battle of the Sexes, which is this complex coordination game where two players want to align but prefer different outcomes. They included a marketing game called Promo, which requires players to take turns, offering promotions to maximize joint profit. Okay. Then there was Samaritan, a game of altruism that tests moral hazard, essentially. If I always help you, well, you just become lazy. Right. The freeloader problem. Yeah, exactly.

10:28And finally, they tested lemons, which is based on the classic economic problem of adverse selection. Think of it like buying a used car where the seller holds all the secret information about the car's true quality. Okay, so to navigate these games, they tested a baseline model, which is basically just asking the AI what it wants to do against two very specific reasoning methods. One is called SCOT, which stands for social chain of thought. Right. And the other is PSBR, or posterior sampling best response. And just to be clear on the mechanics here, we are not changing the AI's underlying code for these methods.

11:03We are prompting the AI to generate a scratch pad in its context window. It literally uses text generation to write out its reasoning before finalizing its actual move. Exactly. And let's look at Scott first. Scott, the social chain of thought, is an intuitive approach, but it is fundamentally myopic. Scott prompts the AI to predict the opponent's very next move and then act based solely on that immediate prediction. Just one step ahead. Just one. The data reveals this works adequately for simple stage game equilibriums. Like, in The Prisoner's Dilemma, Scott easily deduces that the safest immediate play is mutual defection.

11:40It looks one step ahead, sees the risk of betrayal, and chooses to protect itself. But Scott completely collapsed when the games required nuance. Like, when the AI was prompted to follow complex, non-trivial cooperative paths, like coordinating the alternating marketing promotions in the promo game, or sustaining trust over a long period despite information asymmetry in the Lemons game, Scott scored a flat 0%. Zero? It failed utterly to maintain cooperation. And that failure highlights exactly why predicting only one step ahead is completely insufficient in digital markets. If you're playing the lemons game and you only look one step ahead, you will always assume the seller is trying to rip you off with a defective product right now.

12:20So you never buy anything. Exactly. Trust is never established and the market completely collapses. Scuts cannot build trust or fear long-term retaliation because it has absolutely no concept of the future beyond the current turn. Which brings us to the centerpiece of this study, PSBR. Okay, here's where it gets really interesting. Let me give you a structural metaphor to help picture the difference in how these models process reality. Using SCUT is like navigating that dark maze we talked about by only looking at the ground directly an inch in front of your toes. Right. You won't trip over the immediate rock, but you will absolutely get lost in the macro structure of the labyrinth.

12:59Using PSBR, though, is like running a full decision tree simulation in your mind. The AI isn't just looking at the next step. It is explicitly simulating the next 20 intersections in its head. It recognizes that if it betrays its partner right now for a quick advantage, it's going to lead to a ruinous, mutually assured destruction price war 10 moves down the line. That structural simulation is the core of posterior sampling best response. PSBR works by having the AI sample a hypothesis about the opponent's entire overarching strategy, not just their next localized action. So it's looking at the whole playbook.

13:35Yes. And once it has that hypothesis, the AI explicitly rolls out candidate self-strategies into the simulated future within its text generation scratchpad. It literally plays out the entire continuation of the game in its head, weighs the long-term mathematical outcomes, and selects the overarching strategy that yields the highest return. And the performance data on this is staggering. Where the baseline models and Scott failed completely scoring that flat 0 % on complex cooperation, PSBR achieved 92.5 % to 98 % success across the board. It's a massive difference. It almost perfectly executed highly complex cooperative strategy Zero Shot.

14:13It figured out how to take turns in the promo game and how to build trust in the lemons game simply by simulating the future in its context window. And the mechanism behind that massive jump in success comes down to a concept called continuation values. Implementing a complex, repeated game equilibrium requires an agent to weigh the value of the ongoing relationship against the temptation of immediate betrayal. Right. Like, is it worth ruining a good thing for a quick buck? Exactly. Scott fails because it literally cannot grasp the concept of future punishment for a short-term deviation. PSBR explicitly models future play, which forces the AI to internalize long-horizon incentives.

14:54It mathematically sees that cooperating today is the most selfishly beneficial move it can make because it prevents a retaliatory price war tomorrow. Which is an incredibly elegant solution for games where the rules are clear. But let's ground this in a real-world scenario. Digital markets, like an algorithmic pricing engine on a massive e-commerce platform, are messy. Very messy. An AI doesn't have a neat rulebook or a perfect payoff matrix. It might lower a price and maybe sales drop anyway just because it happens to be a slow Tuesday or a competitor ran an unannounced ad campaign. How does this future simulation hold up when the market feedback is actively lying to it?

15:32What happens when the AI doesn't actually know the exact rules or the exact rewards of the game it's playing? This represents the ultimate stress test in the research. They extended their mathematical theory to environments characterized by unknown, stochastic, and private payoffs. Meaning the AI is entirely flying blind. It isn't handed a clean spreadsheet explaining exactly how many points it gets for every single action. No, not at all. Instead, the agent only observes its own privately realized, highly noisy rewards. The research uses Gaussian noise to simulate this reality. Gaussian noise, okay.

16:05Right. So imagine the AI makes a strategic pricing decision. The market feedback it receives is corrupted. Roughly one in four observations might actually be directionally misleading about what the true optimal choice was. One in four. Yeah. It is navigating that dark maze, but now the physical feedback is unreliable. You bump into a wall, but your sensors tell you it's an open hallway. Wow. It's like trying to figure out if your brand new multi-million dollar marketing strategy is working, but a quarter of your customer feedback surveys are intentionally filled out by internet trolls trying to skew your data.

16:37That's a great way to put it. I have to imagine that in a fog of Gaussian noise where the scoreboard lies to you 25 % of the time, the AI's ability to maintain that 98 % cooperation rate just shatters. Yeah. Is the AI essentially building the airplane while flying it? Like it has to figure out what game it's playing, what the score actually is, and D, what the opponent's secret strategy is all simultaneously. If we connect this to the bigger picture. Yeah. Yes. That is exactly the burden placed on the system. Under this extreme uncertainty, the researchers modified PSBR. Now, the AI doesn't just sample a hypothesis for the opponent's strategy.

17:14It also samples a hypothesis for its own payoff matrix from a vast menu of possibilities. That is a lot of variables. It is. It simultaneously guesses the fundamental rules of the universe it inhabits, while guessing the opponent's mindset and projecting both into the future. Okay, so what was the success rate under those conditions? because slots completely failed even when the rules were clear. Right. Even in this thick fog of noisy, misleading data, PSBR remains remarkably robust. While Scott and the baseline models completely collapsed again, failing entirely at sustaining complex cooperative paths with a 0 % success rate, PSBR still succeeded 71 % to 98 % of the time across the complex games.

17:59Wait, really? 71 % to 98 %? Yes. That is profound. Even when the world is lying to it, the model can simulate a path to cooperation. This data point is arguably the most significant revelation in the entire paper. It proves that reasonably reasoning agents do not need a pre-programmed, perfect understanding of their environment to achieve strategic stability. They just figure it out. Yeah. By simply observing historical interactions, continuously updating their Bayesian beliefs about both the rules and the opponent, and simulating the future, their public behavior eventually converges to a rational, stable, Nash equilibrium.

18:34The noise is basically filtered out by the sheer force of long-term reasoning. So what does this all mean for you as someone navigating an economy increasingly run by these algorithms? Well, it means we don't necessarily need to panic about the logistical impossibility of coordinating a universal alignment or fine-tuning procedure for every single AI model deployed in the global market. Which is a huge relief. Exactly. We don't have to force every developer into a standardized regulatory bottleneck. The research strongly suggests that by simply outfitting off-the-shelf AI with the right reasoning scaffold, specifically PSBR, which forces them to look ahead and mathematically weigh the future, they naturally figure out how to cooperate.

19:15They find strategic stability on their own, even in environments that are noisy, confusing, and full of misleading data. We don't have to send the toddlers to a universal boarding school. We just have to teach them how to run a decision tree simulation before they act. It fundamentally changes how we view the deployment of autonomous systems. We are moving from a paradigm of trying to hard-code good behavior to simply providing a computational space for rational behavior to emerge naturally. And it leaves us with a rather profound question to consider long after we wrap up today. Oh, I like the sound of this.

19:48If these AI agents operating in a complete fog of noisy data, corrupted feedback, and unknown rewards can naturally evolve to recognize that long-term cooperation is mathematically superior to short-term betrayal, might these artificial systems eventually become better at avoiding destructive conflicts than human beings themselves? That is a staggering thought to leave on. Thank you so much for joining us for this deep dive. We've taken those toddlers off the stock exchange floor and hopefully given you a clear mechanistic view of how the algorithms running our future markets might just learn to play nice entirely on their own.

20:22We'll see you next time.

From the publisher

This research explores whether AI agents can autonomously reach strategic equilibria in repeated interactions without specialized training. The author proves that "reasonably reasoning" agents—those capable of basic capabilities such as Bayesian learning and asymptotic best-response—naturally converge toward Nash equilibrium play, where posterior-sampling behaviors of off-the-shelf models guarantee asymptotic best response. The study further demonstrates that these agents successfully navigate environments, even when payoffs are unknown or stochastic, by inferring the game structure from private observations. Empirical simulations across various scenarios, such as the Prisoner’s Dilemma, confirm that advanced reasoning capabilities enable stable, predictable cooperation. Ultimately, the paper suggests that sophisticated AI naturally possesses the intrinsic mechanisms necessary for reliable decision-making in complex economic markets.


More from Best AI papers explained

All 475 episodes
Reasonably reasoning AI agents can avoid game-theoretic failures in zero-shot, provablyBest AI papers explained · 20 min
Listen in VO