#53 - John Schultz - Why Google Made ChatGPT, Gemini & Claude Play 900,000 hands of Poker...

6 Feb 2026 · 36 min · 19 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Win-Win with Liv Boeree - Episode #53

Episode Overview In this episode of the Win-Win Podcast, host Liv Boeree interviews John Schultz, a research engineer from Google DeepMind. They discuss the AI Game Arena project, which tests various large language models (LLMs) on games like poker and Werewolf. The conversation delves into the project’s aims, the capabilities of LLMs, and the implications of using games as benchmarks for AI development.

---

Key Themes and Discussions

  1. Introduction to the AI Game Arena
  2. Purpose: The AI Game Arena is a significant research initiative that evaluates how different LLMs perform in various games, aiming to establish benchmarks for AI capabilities and safety.
  3. Games Tested: The project includes over 900,000 hands of poker and 37,000 games of Werewolf.
  1. John Schultz’s Background and Role
  2. Expertise: John has a rich background in games and game theory, having played poker professionally and involved in significant AI projects.
  3. Role in the Project: He spearheads efforts to benchmark LLMs against traditional games to assess their performance.
  1. Comparison of LLMs and Narrow AI
  2. Narrow AI vs. LLMs: The discussion contrasts specialized AI (like those designed for chess or poker) with general-purpose LLMs capable of handling diverse tasks, emphasizing the goals of achieving AGI (Artificial General Intelligence).
  1. Game Performance and Learning Mechanisms
  2. Skills Transfer: Exploration of whether poker skills in LLMs can translate to success in other games, such as Werewolf or chess.
  3. Internal Understanding: Analysis of whether LLMs truly understand concepts like expected value or if they merely replicate learned patterns from vast datasets.
  1. Emergent Behaviors in Gameplay
  2. Diverse Playing Styles: LLMs exhibited varying styles, from aggressive to passive, raising questions about the underlying mechanisms driving these differences.
  3. Misreads and Hallucinations: Notable examples of LLMs misreading poker hands, particularly with flushes, reflecting potential flaws in reasoning or perception.
  1. The Role of Social Games
  2. Werewolf as a Case Study: The choice of Werewolf highlights the complexity of social deception and communication, assessing LLMs' abilities in persuasion and collaboration as well as their potential risks.
  1. Ethical Considerations
  2. Deceptive Behavior in AI: Addressing concerns about training LLMs on games that involve deception and manipulation, and the importance of ensuring that AI behaves ethically in real-world applications.
  1. Future Directions in AI Cooperation
  2. Cooperative Game Theory: Discussion on the need for LLMs to learn cooperation and find win-win scenarios, as well as potential future research to measure and enhance these capabilities.

---

Key Takeaways

  • The AI Game Arena serves as a crucial testing ground for evaluating LLMs in various gaming contexts, providing insights into their reasoning and adaptability.
  • Understanding the nuances of LLM gameplay helps inform the development of safer and more effective AI models.
  • Continued exploration in social games and cooperative strategies can contribute significantly to the advancement of AI towards beneficial outcomes for all.

---

Additional Resources

  • Kaggle Game Arena: [Main Hub](https://www.kaggle.com/game-arena)
  • Poker Benchmark: [900,000 Hands](https://www.kaggle.com/blog/game-arena-poker)
  • Werewolf Benchmark: [Social Deception Game](https://www.kaggle.com/benchmarks/kaggle/werewolf)
  • DeepMind Polaris Library: [GitHub Repository](https://github.com/google-deepmind/polaris)

---

This episode presents a fascinating intersection of artificial intelligence and game theory, encouraging listeners to consider the broader implications of AI's role in society and how cooperative strategies can shape future technologies.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding the AI Game Arena

0:48 to 2:52

Discussion on the AI game arena and its significance in testing LLMs like poker and werewolf.

“And I guess to get us started, can you explain what is the game arena exactly and what your role in this project has been?”

John's Background and Game Preferences

2:52 to 5:56

John shares his background in games, his experiences with poker, and insights on game theory.

“And obviously, so those have been around for a few years now.”

The Evolution of AI in Gaming

5:56 to 7:20

Exploration of how AI has progressed in gaming and the implications of LLMs in complex games.

“It's always tough knowing how much generalization you're going to get out of something ahead of time.”

Generalization of AI Skills Across Games

7:20 to 9:15

Discussion on whether proficiency in one game translates to others for AI models.

“I mean, that's certainly like an open research question.”

Understanding AI Decision-Making

9:15 to 11:30

Examining how AI displays its thought processes in decision-making during games.

“Yeah, you know, I think people are most focused on things that probably have direct scientific economic impact.”

Evaluating AI's Game Abilities

11:30 to 14:00

Discussion on how AI's performance in games compares to math and coding evaluations.

“here's the you know situation and then just like you know make your move and you know one of the that we just found there was that the models, the thoughts weren't very informative.”

Understanding Context Windows in LLMs

14:00 to 15:00

Learn about the significance of context windows in language models and poker.

“Sometimes like, you know, you'll see when going through a lot of the thoughts, they screw things up.”

Adapting Strategies in Poker

15:00 to 16:00

Explore how poker models adapt strategies by utilizing hand histories.

“we got some room to work with, but you know, in retrospect, even that was, uh, that was very little.”

Emergent Behaviors of AI Models

16:00 to 17:00

Discover unexpected behaviors exhibited by AI models during poker games.

“And we reveal the opponent cards at the end of each hand, regardless of whether it went to showdown, in order to sort of like accelerate their ability to adapt to their opponents over the course of those 100 hands.”

Playing Styles of Different AI Models

17:00 to 18:00

Analyze the varying playing styles of different language models in poker.

Show all 19 chapters

Sources of Poker Knowledge in AI

18:00 to 20:00

Understand how LLMs acquire poker knowledge from literature and reasoning.

“in other ways of the llms yeah i'm not necessarily sure that there's and i i would be i would be cautious about reading too much into a connection there.”

Misreads and Errors in Poker Models

20:00 to 22:20

Examine specific misreads and errors made by AI during poker matches.

“is like not applying that very effectively to, you know, new game situations.”

Exploring the Game of Werewolf

22:20 to 24:20

Delve into the reasons why Werewolf was chosen for testing LLMs.

“Everybody's done it at some point or another.”

Communication and Language in Werewolf

24:20 to 27:20

Learn how language and communication are crucial in the game of Werewolf.

“Well, I mean, that was really the social aspect and the language aspect, I think, were the two main drivers behind Werewolf, right?”

Concerns About Deception in AI

27:20 to 28:09

Discuss the implications of training LLMs on deceptive games like Werewolf.

Exploring Deceptive Behavior in AI

28:09 to 29:53

Discuss the implications of LLMs exhibiting deceptive behavior in games.

“Or is it like the wrong way to think about the problem?”

The Importance of Cooperative Games

29:53 to 31:14

Examine the potential benefits of training AI models in cooperative game scenarios.

“So I think, you know, this is important science, essentially, that needs to be done.”

Challenges in Game Complexity for AI

31:14 to 33:00

Identify types of games that present challenges for AI performance.

“Yeah, I mean, those are things that we're definitely thinking about.”

The Future of AI in Real-Time Strategy Games

33:00 to 34:56

Discuss the difficulties AI might face in mastering real-time strategy games like Starcraft.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So if you haven't seen yet, there is this huge battle called the AI game arena where all of the major LLMs, the AIs that you and I use every day, are being tested against each other as some of the most interesting games in the world. This is an experiment never run before at this scale. There are 10 different LLMs playing over 900 ,000 hands of heads-up poker, and they're also playing over 37 ,000 games of werewolf, which any of you fellow games nerds will know is a very statistically significant sample size. And I'm thrilled to be partnering with Google this week to bring you guys an exclusive interview with one of the main brains behind it, DeepMind engineer John Schultz, where we'll be learning not only the how they're doing this experiment, but also, more importantly, the why.

0:48On that note, let's delve in.

0:54John, thank you so much for joining. You work for Google DeepMind. And I guess to get us started, can you explain what is the game arena exactly and what your role in this project has been? Yeah, hi. Yeah, so it's a real pleasure to be here. Yeah, I'm a research engineer at Google DeepMind. My background's in games and game theory. This is a partnership with Kaggle to, yeah, just have a like premier benchmark for, you know, modern LLM's frontier models on how they do it, you know, various classic games. What are your favorite games personally to play? Because you clearly, you're clearly a games player.

1:31I imagine you've had to study all of these inside out. I mean, you mentioned you played a lot of poker anyway in life, right? Yeah, so I did live in Vegas for five years. Played poker out there after my master's. Yeah, it was a good time. And poker's a great game. And I love Vegas. And it was actually while I was out there that AlphaGo happened. And that was sort of like the moment for me. That was the wake-up call that, like, you know, AGI was on the horizon. In part because I had been looking into the state-of-the-art chess engines while, like, almost just by happenstance, right when that came out.

2:04And I was sort of dismayed to learn that it was basically still the same algorithms that were used to defeat Kasparov in 97. Obviously, they had been fine-tuned and everything and were much more performant, but the same general approaches. And then I also knew how far off a lot of people thought Go was at the time. Many people said it was 20 years off. There were many people who said that it would never happen. And then all of a sudden it happened. And it happened very convincingly. and as soon as AlphaGo came out it was pretty clear to me that that technique would work on a broad class of games and then not long after AlphaZero came out and that's exactly what they did and so yeah that that was that was you know big moment for me and and really what what motivated me to uh to get into this my my favorite game I have to plug is uh is Gin Rummy the thing that got me into Gin Rummy so I was just always into you know classic card games I grew up playing hearts like like on Microsoft and that was yeah I just just I still love hearts I think it's a great game but when I got into poker you know obviously one of the greatest poker players ever Stu Unger you know three times world series of poker champion and many people said he was the greatest uh gin rummy player uh ever and so yeah it just felt like kind of another classic game that I had to you know had to understand so you know I was around playing poker professionally when Labrata spot first came on the scene which was like a big deal because it's the first superhuman level poker AI.

3:30And obviously, so those have been around for a few years now. And frankly, they're already starting to wreak havoc on online poker. But this is quite a different thing, right, which people might not realize, because these are general models that are playing, you know, these sort of everyday chatbots that so many people are using. So why are you testing general purpose LLMs. And what does it mean for general models to start being okay at these kinds of games? You're right. So it's a good point. And I see a lot of confusion around this in poker and also like chess and other games. Obviously in chess, we've had Deep Blue beat Garry Kasparov in 1997.

4:10And so we've had superhuman chess engines for a long time. We've had superhuman and poker for a while now. But those are very specialized systems built for that sole purpose. They couldn't really do anything else outside of that. And as you noted, these are general purpose models. You can interface with them through text, ask them anything. Yeah, so we're not expecting these to be as good as the specialized models. But the goal is to build AGI, to have generally intelligent systems. And if you wanna have a generally intelligent system, it should be able to perform at least reasonably well in a wide variety of settings, including these game settings.

4:51And I think it's a really interesting time when the models are finally starting to perform reasonably well. It depends on the domain, but in various games. I think it was an interesting situation because I think maybe the community in some sense, the AI community had felt like, well, a lot of these games, they were solved already. And so people were more interested in various other things like math and coding. And, you know, obviously there's a lot of importance to those. And, you know, some people are focused on them for good reason. But a lot of games that really like lag behind. So, you know, when we first launched the game arena at the end of the summer, a lot of the models are like frequently making illegal moves in chess.

5:33We're seeing a lot less of that now. And in fact, starting to see some quite good moves. So I think it's, yeah, we're at this like really interesting phase transition where the models are starting to pick up on these games. Yeah, because as we said, these are sort of playing, I would say from what I've seen at poker thus far, they're playing at weak to medium level players, but far, far from the domain of experts, top poker players. Do you expect that like once an LLM, for example, gets good enough to be beating top pro poker players in poker, that that will also translate to expert abilities at other complex games, like werewolf, chess, diplomacy, etc.?

6:16Yeah, it's a good question. It's always tough knowing how much generalization you're going to get out of something ahead of time. We've seen the models improve in some domains, and you get some generalization around that, but just because you get good at one thing doesn't necessarily mean you're going to automatically get something else for free. So, and I mean, you know, the fact that like the models were like very good at things like math and coding before they were, you know, good at games, I think is an indication of that. But, you know, games are a really broad domain and you can sort of gamify everything in some sense.

6:54Right. So, yeah, I think, I think it's, you know, the models get better at broader classes of games. It's probably going to track with, you know, general abilities in some sense. um so yeah it's certainly something i mean that was one of the goals of the Kaggle game arena is to just have good clean benchmarks on this so we we can measure that progress um and i think it's very useful to have uh experts uh like yourself you know really be able to like like like track this progress right because that's that's one of the things that we want to see is you know how well not just how well they're doing against each other but like you know inspecting the reasoning and seeing you know how well does this stand up to uh expert scrutiny yeah what i love about the way you're sort of presenting this information is not only do we get to see what the plays they make, you know, the way you would see normal poker players play online, but we also get to see the chain of thought, as it's called, of each model written on the side.

7:48What I'm curious to ask you, because I've sort of got my own opinion on this, but I wonder what you think, is how precisely do these thoughts that they're displaying actually map onto what's truly driving the model sort of underneath the surface to make these decisions? I mean, that's certainly like an open research question. And yeah, there's a lot of people from the interpretability community working on that. I don't really have like a strong, you know, opinion on that myself as to, yeah, exactly how that all maps on. But, you know, if you look at things like, yeah, in chess and poker, I mean, like chess, I think is sort of a good one because you kind of get a good, like you see what it's saying about, well, if I move here, then this will happen and this will happen and that.

8:31And yeah, I mean, I think, look, to some extent, you're certainly seeing the thought, like, you're definitely getting some sense of what it's thinking. I don't know that, like, the external thoughts obviously fully capture, you know, what's going on in, like, internal model weights and everything. But yeah, I mean, that's going to be another, you know, interesting question that people are working on all sorts of techniques to analyze what's sort of going on inside the models as they think through their decisions. So benchmarks and evals are generally now front and center of any new model release.

9:06How do these models' game-playing abilities sort of rank in relevance to, for example, math code and language evals? Yeah, you know, I think people are most focused on things that probably have direct scientific economic impact. I think that's probably going to remain the case. So like, you know, coding is certainly front and center, but, you know, and other things like visual understanding, those sort of things. But, you know, a lot of benchmarks, like everybody now, and when a new benchmark comes out, everybody's always saying, well, you have to take this with a grain of salt because of this and this and this.

9:45And people kind of, you know, everybody's caveating everything. And yeah, you know, games are one of those things. I mean, they have their own caveats around them, but in many ways they are resistant to saturation in ways that other ones aren't. So for instance, there's always questions about whether a particular eval has leaked into the training data. And obviously people work hard to make sure that that doesn't happen, but then there's always questions like these things are out there on the internet, it's very hard to know. So, but games are the kind of thing where you can't train on every single chess position just because there's, you know, like combinatorially far too many, similarly with poker and other games.

10:28And then even if, you know, you could always just create new games that don't really exist out there. And so it's clearly not in the training data, or you could take games that are in the training data, but then like change the rules on them a little bit so that it doesn't necessarily, it could even hurt you, right? If you were like overfitting to what was out there. So there's lots of things about games that make them a useful eval in a time when, you know, evals are constantly being saturated. So the prompt you used for poker is very long. I'm actually going to put it up on the screen for the viewers.

10:57You guys can sort of, I'd recommend you pause and take a look through it if you want to look at it closely. We can also link to it beneath the video. But John, I'm curious, how did you decide specifically on what to put in the prompt and whatnot? like did you have specific limitations in terms of how long it could be um or sort of the level of poker specific complexity you could include yeah so good question a lot of thought goes into the prompts and um so we wanted to accomplish a number of objectives uh with the prop with the poker prompt so first we had just tried like some very simple prompt they're just you know here's a here's the you know situation and then just like you know make your move and you know one of the that we just found there was that the models, the thoughts weren't very informative.

11:45And so one of the things that we're trying to accomplish with the eval is to be able to inspect the thoughts to see how the models are coming to their decisions. So that's why ultimately we decided on something where we were saying explicitly, break down your decision making in terms of these concepts that that are just well-known poker concepts. And we're also very clear in the prompt to say that you should play to maximize your expected value. We like to start off with a game theory optimal strategy, but exploit your opponent if you can. And the reason for that is if you just say, here's a hand of poker, play it, right?

12:25You're not really specifying the objective there. It's perfectly reasonable for the model to go, So like, well, I should just, you know, have fun here and do whatever or and maximize for that or like. Or just maximize for like, yeah, entertaining the reader, because some models are just meant to be entertaining or they're meant to be useful. But it's like often optimized for just like how they think the user will keep coming back and asking more questions. Right. Exactly. So and if you haven't specified the objective, it's kind of fair game for the model to respond within whatever you've said.

12:57Right. So we want to be very explicit, like, yeah, this is what you should do. You should play to maximize EV and justify your answer in terms of standard poker concepts. And now in terms of the exact wording of a prompt or things like that, I mean, first off, it's impossible to sort of ablate everything, right? You'd have to just rerun everything all the time. So, you know, we really just tried to like frame everything as sort of straightforward as neutrally as possible while framing the, you know, the objective. So that was like kind of the main overarching goal there. Do you think they have any internal understanding when you say maximize expected value?

13:39I mean, to the extent, right, there's always, you know, caveating like, you know, what internal understanding means and that. But yeah, certainly they, you know, like they know what that means. Like, and if you ask them about - Because they can do an EV calculation, right? If you put it on, like what, or tell me what the expected value formula is. They know what it is. Right, they'll be able to do all those things. Sometimes like, you know, you'll see when going through a lot of the thoughts, they screw things up. But they also do a lot of things right. And they certainly will, like they are familiar with lots of stuff about poker history, poker theory, like all this stuff in the abstract, like they can tell you about.

14:17Obviously, applying it is, you know, is a different question. But no, no, in terms of all that, they are certainly familiar with it. Yeah. And in terms of your, you know, context windows are kind of everything. And a lot of these LLMs now have huge context windows, right? Like what is like a million tokens now on Gemini? Right. Yeah. how long were the context windows you were using because it you know again looking through the hands certainly it seems like they're building up a history when they're playing in one of these matches they they remember previous things that it's their opponent has done so was there a limit on context window you used right so so first i mean i remember like uh you know and when language models wasn't that long ago and when you know they first started coming out were like you know talking about 512 tokens then it was like you had a couple thousand tokens and it felt like oh now we got some room to work with, but you know, in retrospect, even that was, uh, that was very little.

15:07Um, and I'm sure at some point we'll even feel the same way about, uh, you know, even, even a million tokens now. Um, although you can do it, you can do a lot with a million tokens. Um, yeah. So, uh, well that was, so, so one of the key things that we did here was we wanted to, you know, obviously poker is a lot about adapting, uh, to your opponents. So we wanted to make that one of the features of the competition. So we wanted to include hand histories. And the hand histories themselves, I don't know the exact token counts off hand, but there's not too many tokens per hand history. And so we, but there's still obviously, they do add up.

15:48And so for that and other reasons, just in terms of like infrastructure wise and running episodes and everything, we separated all the hands into like 100 episode chunks. So, you know, the models play 100 hands and over the course of that 100 hands, they see every previous hand in context. And we reveal the opponent cards at the end of each hand, regardless of whether it went to showdown, in order to sort of like accelerate their ability to adapt to their opponents over the course of those 100 hands. So with that, it was definitely not bumping up against the, you know, the absolute limits of the, you know, the context windows there.

16:27Were there any very surprising emergent behaviors that you did not expect? Yeah, I don't know that it was so surprising. I guess some of the models are particularly aggressive. Like Grok's VPIP is way up there. What was it? Do you know? I'd have to double check off the top of my head, but it was over 90 something. and it was um yeah it was it was playing every hand i mean the funny thing was that uh i don't know some of the times that like models would take stands sometimes in good spots sometimes not in good spots uh yeah i mean um but but they they yeah they definitely like would recognize when when someone was playing you know tons of hands and things like that so um yeah i think that's like in terms of getting principled adaptation i'm not quite sure that that they're there yet but uh they definitely made note of it why why do you think it is then that some of the models have like because it does seem like they have truly different playing styles like as you say grok was um especially aggressive bordering on reckless and um i think i can't remember which one it was it was it claude opus was just incredibly uh passive i believe um but why why is it do you think that like what what would be the mechanism that is making them become more like having almost like a personality and does it match up with the general personality in other ways of the llms yeah i'm not necessarily sure that there's and i i would be i would be cautious about reading too much into a connection there.

18:12Although it is the case that you know, training models can kind of, all sorts of behavior can be affected by all sorts of things, right? So it's very hard to know exactly what causes a model to play a certain way. I think it's certainly interesting. I think it would be, you know, I'd love to see people kind of look into it even more and see if they can find it. But I would, yeah, I would just be careful about reading too much into anything specific because i just think there's like sort of too many factors to be able to pinpoint that yeah and in terms of like where they're getting their their poker skills from or lack of it in some cases uh is it do you think it's mostly coming from just the fact that they have read basically every poker book out there they've read every poker forum they've probably just found all these hand histories online you know in in the data they were originally fed with um and and therefore they're sort of actually developing an internal abstraction of the game or is it more that they actually just have a very small amount of poker specific knowledge but they just have these generally broad reasoning capabilities and that is sufficient for them to play decently well yeah it's a little bit of both right they've definitely seen lots of you know various things and like again they're familiar with a lot of uh you know, poker concepts.

19:27They do have a good amount of general reasoning capability, even in like sort of unseen, you know, domain. So if you kind of put that together, they're going to do, they're going to do reasonable, you know, somewhat reasonable things. Yeah, I don't know. They definitely have a bit of both. They definitely like, you know, familiar with general principles, good at general reasoning, you know, Combine that and you're going to get something. But I think one thing that we've seen, and I've seen working on games with models, is that you can't just... So if you take a model and you just, let's say, fine-tune it on some strategy for a particular game or something like that, like a strategy guide of some sort, it will sort of regurgitate information that sounds similar to the strategy guide, but doesn't necessarily map on...

20:21is like not applying that very effectively to, you know, new game situations. I think it's sort of similar to a person where if you, I don't know, you know, you, I would study like some new other game or subject in math or physics that I'm not particularly familiar with, and you just sort of read the book, and then someone gives you a problem. You're not like right away going to be able to necessarily like perform well on that particular task, right? you kind of have to work through problems in that domain whatever that given domain is so i think that's where uh rl has been really powerful and we've just seen like rl mature a lot uh with llms in the past year um and yeah i think that's that's where we're seeing a lot of gains just across the board from the models come from were you surprised at the results of the poker tournament i don't know that it was uh i'm not sort of too surprised by anything with the language models i mean you know you kind of never know what you're going to get i was a little surprised to see them misread some hands especially around flushes they love thinking they flop to flush when they have like king high and they're like and i you know and they're reasoning you see it they're like i i have flop to flush so i'm gonna get it in now and it's like no you have a king high flush draw you're the king of clubs and that's it and and yeah there's been a number of those it seems like flushes in particular they really really struggle with um i wonder why yeah do you have any idea why that would be any suspicions i actually am not quite sure and one of the things that that that is interesting about that is that models tend to uh you know sometimes tokenization can always potentially be an issue that can trip up models.

22:07But I don't know that that's the reason here. And yeah, I don't necessarily have a good explanation for that. But I can say as someone who's played a lot of poker, I've certainly misread my hand before. We've all been there. Everybody's done it at some point or another. Not making excuses for the models, but it's something that does happen. Yes, it's a surprisingly human behavior, actually, because they really should have the perfect information. They don't have eyes that go bleary. You know, they don't get tired, right? So it is odd that they seem to consistently do that much more than any other sort of odd behavior that I've noticed.

22:52But that's one of the things that like something like tokenization where sometimes people see models make stupid mistakes and they sign all the models just doing something stupid or whatever. But, you know, it's hard to know exactly how the information is being perceived. And sometimes there's artifacts of training and architectures and things that cause these little blind spots and they manifest in these sorts of ways. so yeah i mean sometimes it can be a failure of reasoning but sometimes it's also like these weird failures of perception and again hard to know exactly what but were there any other strange hallucinations you noticed no i actually it was really the flush one was was actually the one that stuck out to me uh stuck out to me as well those are the ones in terms of just like pure misreads uh that that jumped out yeah well i mean i guess they're not numbers right um okay Okay, switching to Werewolf specifically.

23:49So can you talk a bit about why you guys chose to test the LLMs on Werewolf itself? Because it's a game I love. I play regularly here at home with friends in Austin. You know, poker is obviously a much more general game than a game like chess because it tests you on uncertainty. It's a game of incomplete information and so on. and it feels like werewolf is an even more general game than that because there's sort of there is game theory there is some you know uncertainty that you can try and quantify but it's a very social game and it's a game of you know trying to decipher people's words there you know and it's sometimes you need to collaborate and sometimes you need to try and deceive so why werewolf and like were there any interesting sort of like how did you even think about trying to do this Right.

24:40Well, I mean, that was really the social aspect and the language aspect, I think, were the two main drivers behind Werewolf, right? So, you know, chess is obviously the sort of like abstract structure, board game, you know, card games are sort of something similar. There's like, in those cases, like rules and cards and things like adult. Here, it's really like mainly just communication and language. And so that's, yeah, just like a sort of fundamentally different type of game and relies on, yeah, persuasion, these sorts of things. And I think that's what makes it really, really fascinating. What's also interesting about Werewolf is there is presumably way less literature available online about how to play the game.

25:28Not that many people know about it. You know, whereas like chess, there's presumably reams of training data. Same with poker you can use to um that would have gone into these llms so at the same time these models some of them played well decently well like it seems like they know like they've studied games so where do you think um where do you think their reasoning abilities are coming from is this like a sign of like actual general reasoning like of an understanding of what it takes to persuade somebody or do you think it's just purely like they're they're kind of just spitting back the little literature that they have found yeah it's a good question so well first off there's a lot of literature out there on persuasion generally right so and just like social dynamics these sorts of things so i i mean that all factors into uh you know like like sort of their their training and and they've seen that so uh even though they may not have seen as much on werewolf in particular they can bring all that to bear.

26:33One of the things though, also that language models are, so language models, because they're primarily, you know, trained on language, they are quite good at like language specific things. And it's, I think it's one of the gaps between something like chess, where there is this, you know, abstract board game object. Knight D2, yeah. Right. And like models, for instance, up until recently would have a really tough time even understanding some of the notation, whether you gave it to them in some sort of like an ASCII board format or the fen string is another representation or, you know, these things themselves because the models, you know, relative to maybe just like language, you know, data, like they just weren't as familiar with that and had a tougher time with that.

27:15But yeah, a game like Werewolf where it's purely in natural language is sort of like that's right in their wheelhouse. And that's why models were like originally, you know, people would put them in these sort of just games or settings where it's like you're in some room and there's something and these sort of like those things the models could actually work through pretty well because again it was like all in natural language so those are games where i felt like there was actually maybe the models in some ways are a little bit ahead of where they are in some games that are a little bit more just like yeah these these sort of pure abstract uh structures just for some pushback you know i can imagine that some people might find it concerning or at least um worth considering that by training llms or at least testing llms on games of you know largely of deception like werewolf and poker that we could in some way be sort of encouraging deceptive or manipulative behavior from LLMs.

28:17And I think that is a fair concern. What would you say to that? Do you agree with it? Or is it like the wrong way to think about the problem? So I'll say a couple things. First, I mean, when they're playing these games, like they know they're playing a game. So I think, you know, them exhibiting these various behaviors in the context of a game is not necessarily something that's problematic. I mean, we really just don't want models doing things that would violate our expectations in other settings. And so we should have, we do have all sorts of evals and testing to make sure to minimize the risk of that as much as possible.

29:03I think it's important to use these as evals to see how good they are at these, like seeing what capabilities they have with regards to that is important in its own right. So, yeah, I mean, obviously we want to train models to be helpful, useful, trustworthy, all of these things. But I think it's certainly fair game to have them in the context of playing games and other things. And in many ways, I think these are sort of like the safest settings in which to explore this behavior. Yeah, I would agree. And I mean, also, again, for viewers, it's not like you guys, these are all your models. You're actually testing all of the models of different companies, competitors even.

29:46And you're trying to create a sort of healthy sandbox environment whereby we can actually learn about these. So I think, you know, this is important science, essentially, that needs to be done. Yeah, I agree. And I also think it's important that, you know, okay, so a lot of these games are sort of, you know, competitive games in that. I hope we get, like, you know, more of the way of, like, cooperative games, you know, in the future where we want models to, you know, look for win-win opportunities. And especially as we have, like, more agentic behavior of models, you know, out in the world. And I think that's something that's really important to benchmark and, you know, to encourage in models is to, yeah, look for wins, find, you know, maybe the Nash Equilibrio that wouldn't be accessible otherwise, except if you, you know, knew how to cooperate and how to find, you know, the best outcomes.

30:39Yeah, I'd actually love to dig into that a little bit more because that was, obviously, that's the theme of this whole channel, right? Like, how do we find better, as you say, Nash Equilibrio? where more people benefit, you know, essentially just like shift the Pareto frontier in different areas so that more people benefit. So are there any particular lines of research that you are looking into, or are there any particular promising games out there that you think could better, essentially better inform LLMs how to optimize more for cooperation and coordination over competition? Yeah, I mean, those are things that we're definitely thinking about.

31:17and I mean I'm sure many of your listeners will be aware but you know it's it's it's a really interesting problem in game theory right where so if you have a two-player you know zero-sum game any equilibrium solution is a best response you know to any other and there's no incentive to deviate the problem is when you have general sum and player games everybody can be in an equilibrium and then there might be some other equilibrium with a better payoff for everybody but nobody unilaterally has an incentive to deviate to get there. But you could imagine ways where if, you know, the models could identify that equilibrium and then say to everybody, hey, look, you know, that exists and we should all try and get there.

Read the full transcript

31:57And, you know, you find a path there where there's sort of like cooperation along the way. And if nobody, you know, kind of deviates, then, you know, you can get to a place that you otherwise wouldn't be able to get to. Yeah, that would be a huge win. On Werewolf, we actually collaborated with the game theory team in DeepMind with some of their more recent developments on metrics and rankings. And it uses a library which was recently open source, which anybody can use, called Polaris. And you can find them in the GitHub DeepMind Polaris library to, in this case, measure how well each model does within each role.

32:38like, you know, relative to the performance of all the other models. So to just kind of give, you know, more insight into where the strengths and weaknesses of models are. Are there any games that you expect general models to not do well at for quite a long time? So I don't know, maybe one thing I would say, you know, in terms of which games would be like hardest, which may be the kind of the longest holdouts. I do think games that are, that just have like lots and lots of pieces, maybe like just tons of structure to them just in terms of kind of keeping track of everything like probably going to so one for instance you know a game like Stratego is going to be one that's probably going to be you know it'll probably be just tougher and take a little bit longer for the models to get to like you know strong performance on I think they'll eventually you know I trust they'll get there but I wonder because someone I got into Stratego a couple of years ago and I mentioned it to a friend who plays a lot of chess and he was like oh yes it's like chess for children it's baby chess it's super simple and like kind of he kind of shat on it a little bit um but I agree because I was like actually there's tons of pieces it was I I found it quite complex but I mean I can't tell if that's just like a chess purist thing they they tend to be very puritanical about their game I I get it it's a great game um but it's funny you mentioned it uh I want another one that like comes to mind is something like actually Starcraft which of course you guys built like basically a superhuman starcraft ai um but i wonder how well because it's such a because a it's very fast in real time right and again presumably there's many different variables massive state space complexity so for an llm to be able to play that sufficiently well i'm curious whether you think it translates right yeah i was thinking initially more in terms of the like sort of tabletop strategy games but yeah when you get into real time games that's sort of a whole nother dimension um we basically have agi if or even more right if we can do that and i i i would say i mean yeah you know just in terms of the latency alone that's yeah that that's one of the uh going to be one of the tricky parts there um but i yeah i mean you know yeah definitely definitely has its own challenges but um you know i i expect there'll be solutions uh before too too long.

34:56Amazing. Thanks so much, John. It's a real pleasure. Thank you. So there we go. Huge thank you to John for joining today. Keep an eye out tomorrow for an interview that we'll be doing with Doug Polk, the greatest heads up player of all time, who's also been involved in this project. And he has a lot to say on how these different LLMs actually play. And of course, if you want to check out the games themselves, head on over to Kaggle, where they're all being uploaded, You can watch some of the replays. There's just tons of stuff to dig into. I'll link that below. And I will see you tomorrow.

From the publisher

Which LLM is best at POKER? What about the social deception game Werewolf? This week I'm collaborating with ⁠ @googledeepmind ⁠ and ⁠ @kaggle ⁠ to explain their new “AI Game Arena” project, designed to test all the top LLMs at various games.

The Game Arena is a massive research project to create new AI benchmarks, and understand what really makes these general-purpose AIs tick. It also helps us evaluate how far along the path to true AGI we are, NBD.So who better to speak to than Deepmind research engineer John Schultz, one of the main brains behind the project.

We discuss why some LLMs are so much better than others, their internal understanding of the games (or lack of it!) and why games are a useful way to evaluate both capabilities and safety of frontier models.And friends - do please note that for the first time ever on this channel, this post is part of a paid partnership (woohoo I made it!).

So while everything I said in this interview, is absolutely my own views, there are also direct financial incentives behind this particular interview.


Chapters

00:00 - Intro

00:55 - What Is Game Arena?

01:25 - John’s Favorite Games

03:18 - Narrow AI Vs LLMs: What’s The Difference?

05:58 - Will LLM Poker Skills Transfer To Other Games?

09:00 - Benchmarks And Evals

10:50 - Prompt Design For Game-Playing LLMs

13:32 - Do LLMs Understand Concepts Like Expected Value?

14:24 - Context Window Length17:27 - Why Do LLMs Have Different Playing Styles?

18:41 - Where Does Their Skill Come From?

21:09 - Surprising Results

23:50 - Why Werewolf?

25:48 - Where Does LLM Reasoning Come From?

27:49 - Could LLMs Learning Social Games Be Dangerous?30:58 - Can Games Teach LLMs Cooperation?

32:48 - Which Games Will Be Hardest To Learn?


Links:

♾️ Kaggle Game Arena (Main Hub): https://www.kaggle.com/game-arena

♾️ Introducing the Kaggle Game Arena (Official Post): https://www.kaggle.com/blog/introducing-game-arena

♾️ Poker Benchmark – 900,000 Hands: https://www.kaggle.com/blog/game-arena-poker

♾️ Werewolf Benchmark (Social Deception Game): https://www.kaggle.com/benchmarks/kaggle/werewolf

♾️ Google DeepMind (Research Organization): https://deepmind.google

♾️ Polaris Library (DeepMind, Open Source): https://github.com/google-deepmind/polaris


Credits:

♾️ Hosted by Liv Boeree

♾️ Produced by Luca de Vico


The Win-Win Podcast:

Poker champion Liv Boeree takes to the interview chair to tease apart the complexities of one of the most fundamental parts of human nature: competition. Liv is joined by top philosophers, gamers, artists, technologists, CEOs, scientists, politicians and more to understand how competition manifests in their world, and how to change seemingly win-lose systems into Win-Wins.


Podcast links:

♾️ Website: https://www.winwinpodcast.com/

♾️ Youtube: https://www.youtube.com/playlist?list=PLWgq0OZMtwtOIyMsVM_vksqdfWcM-b68S

♾️ Spotify: https://open.spotify.com/show/03bGVUaFZmJUmEvSHNDPdI?si=64379cc23696454f

♾️ Apple Podcasts: https://podcasts.apple.com/us/podcast/win-win-with-liv-boeree/id1724791350

♾️ Pocketcast: https://play.pocketcasts.com/podcasts/7f708340-d17c-013b-f46e-0acc26574db2

#winwinpodcast #AI #poker #kaggle

More from Win-Win with Liv Boeree

All 58 episodes
#53 - John Schultz - Why Google Made ChatGPT, Gemini & Claude Play 900,000 hands of Poker...Win-Win with Liv Boeree · 36 min
Listen in VO